Skip to content

solvi.core.knowledge.risk

What to do with a prediction: protection by default (Protect), justified risk as an option (RiskBudget), and a learned gate with an expiry and a floor (LearnedGate).

Knowledge as protection (the default) and justified risk (an option): what to do with a Prediction.

from solvi.core.knowledge import Protect, RiskBudget
policy = Protect()                                   # the verdict is final: refuse → avoid, unknown → System 2
policy = RiskBudget(max_risk_per_episode=1.0, min_gain_ratio=1.0, min_support=3)
d = policy.decide("descend", prediction, gain=0.3)  # → RiskDecision(choice="take" | "avoid" | "ask_s2", ...)
policy.new_episode()                                 # the per-episode budget starts again

The RiskPolicy protocol: decide(action, prediction, gain=0.0, *, key=None) → RiskDecision; new_episode(). Every decision records its rationale: the risk estimate, its support, the expected gain, the budget left.

Protect: "accept" → take, "refuse" → avoid, "unknown" → ask System 2. Knowledge used only as hard gates. RiskBudget: a refused action is taken when its expected gain is at least min_gain_ratio × its estimated risk and the risk fits what is left of the episode's budget; a refusal resting on fewer than min_support cases (and not on a written spec's rate, support -1) is ambiguous → System 2; a hard prediction (an instant-harm or policy rule) is never taken in any mode. A key is charged once per episode (the same fight, step after step).

Why both exist: knowledge used only as gates removes the failures it targets and can lower a metric that rewards risk (in a dungeon game: fewer deaths of the targeted kinds, less depth reached, heroes starving instead). Justified risk won back part of it there, not all. So protection stays the default; RiskBudget is an option to measure on your own metric. benchmarks/knowledge/risk_dungeon.py shows the mechanism on a toy dungeon.

LearnedGate: a gate learned from failures is bounded. A depth gate learned as "the median death depth at this level − 1" tightens itself: deaths happen where the gate lets the hero go, so each death lowers it further (risk_dungeon.py, arm protect_unbounded: 6.70 → 2.99 levels from the first to the last third of the streams). Here a learned limit reads only the last expiry episodes, never falls below floor, and is reopened by evidence: a level whose recorded arrivals show a low failure rate is predicted with that rate and support, which a RiskBudget can take; Protect still refuses it. A bounded gate does not make the decline go away by itself: with it, Protect still fell 6.74 → 5.02 there, as learned refusals accumulated.

RiskDecision dataclass

RiskDecision(choice: str, reason: str, risk: float = 0.0, support: int = 0, gain: float = 0.0, budget_left: float | None = None, action: str | None = None)

What a risk policy decided about one action, and why: the choice, the reason, the prediction's risk and support, the expected gain, and the episode's budget left after it.

RiskPolicy

Bases: Protocol

What to do with a prediction.

You implement: decide(action, prediction, gain=0.0, *, key=None) → RiskDecision (take / avoid / ask_s2 with its reason, the risk, support, gain and budget left) and new_episode().

You get for free: every decision on a refused or unknown action recorded with its rationale; a hard prediction is yours to never take (Protect and RiskBudget never do).

Stability: stable.

Protect

Protect()

The verdict is final (the default): accept → take, refuse → avoid, unknown → ask System 2.

RiskBudget

RiskBudget(max_risk_per_episode=1.0, min_gain_ratio=1.0, min_support=3)

Justified risk within a per-episode budget (see the module docstring). max_risk_per_episode: the sum of the risks taken in one episode (in the predictions' unit, e.g. death-equivalents); min_gain_ratio: take only when gain ≥ ratio × risk; min_support: fewer cases behind a refusal → ask System 2.

LearnedGate

LearnedGate(name, *, expiry=10, floor=0, near=1, min_support=3, margin=1, prior=0.5)

A limit on a level (a depth, an amount, a distance) per context (a number: an experience level, a stage), learned from failures and bounded.

gate = LearnedGate("depth", expiry=10, floor=lambda xl: xl + 1, near=1, min_support=3, margin=1)
gate.failed(context=xl, level=depth)          # a failure (a death) at this level
gate.arrived(context=xl, level=depth, failed=False)   # evidence: reached this level; failed there or not
gate.end_episode()
gate.limit(xl)                                 # max(floor(xl), median failure level nearby − margin) or None
gate.predict(xl, depth)                        # Prediction: accept below the limit, else refuse with the rate

expiry: only failures of the last expiry episodes count (a gate cannot be tightened by old deaths forever); floor: the limit never falls below floor(context) (a number or a function of the context); near: failures at contexts within ±near count; min_support: failures needed to set a limit at all, and arrivals needed before their rate replaces the prior; prior: the risk predicted past the limit without enough arrivals.

limit

limit(context)

The highest allowed level at this context, or None (no limit learned: not enough recent failures).

evidence

evidence(context, level)

(failure rate (Laplace), number of recorded arrivals) at this level and a context nearby.

predict

predict(context, level)

Prediction for going to level at context: accept within the limit; past it refuse, with the measured arrival failure rate and its support when there is enough evidence (else the prior, support = arrivals).