solvi.core.knowledge.risk¶
What to do with a prediction: protection by default (Protect), justified risk as an option (RiskBudget), and a
learned gate with an expiry and a floor (LearnedGate).
Knowledge as protection (the default) and justified risk (an option): what to do with a Prediction.
from solvi.core.knowledge import Protect, RiskBudget
policy = Protect() # the verdict is final: refuse → avoid, unknown → System 2
policy = RiskBudget(max_risk_per_episode=1.0, min_gain_ratio=1.0, min_support=3)
d = policy.decide("descend", prediction, gain=0.3) # → RiskDecision(choice="take" | "avoid" | "ask_s2", ...)
policy.new_episode() # the per-episode budget starts again
The RiskPolicy protocol: decide(action, prediction, gain=0.0, *, key=None) → RiskDecision; new_episode(). Every decision records its rationale: the risk estimate, its support, the expected gain, the budget left.
Protect: "accept" → take, "refuse" → avoid, "unknown" → ask System 2. Knowledge used only as hard gates.
RiskBudget: a refused action is taken when its expected gain is at least min_gain_ratio × its estimated risk and the
risk fits what is left of the episode's budget; a refusal resting on fewer than min_support cases (and not on a written
spec's rate, support -1) is ambiguous → System 2; a hard prediction (an instant-harm or policy rule) is never taken in
any mode. A key is charged once per episode (the same fight, step after step).
Why both exist: knowledge used only as gates removes the failures it targets and can lower a metric that rewards risk (in a dungeon game: fewer deaths of the targeted kinds, less depth reached, heroes starving instead). Justified risk won back part of it there, not all. So protection stays the default; RiskBudget is an option to measure on your own metric. benchmarks/knowledge/risk_dungeon.py shows the mechanism on a toy dungeon.
LearnedGate: a gate learned from failures is bounded. A depth gate learned as "the median death depth at this level
− 1" tightens itself: deaths happen where the gate lets the hero go, so each death lowers it further
(risk_dungeon.py, arm protect_unbounded: 6.70 → 2.99 levels from the first to the last third of the streams). Here a
learned limit reads only the last expiry episodes,
never falls below floor, and is reopened by evidence: a level whose recorded arrivals show a low failure rate is
predicted with that rate and support, which a RiskBudget can take; Protect still refuses it. A bounded gate does not
make the decline go away by itself: with it, Protect still fell 6.74 → 5.02 there, as learned refusals accumulated.
RiskDecision
dataclass
¶
RiskDecision(choice: str, reason: str, risk: float = 0.0, support: int = 0, gain: float = 0.0, budget_left: float | None = None, action: str | None = None)
What a risk policy decided about one action, and why: the choice, the reason, the prediction's risk and support, the expected gain, and the episode's budget left after it.
RiskPolicy ¶
Bases: Protocol
What to do with a prediction.
You implement: decide(action, prediction, gain=0.0, *, key=None) → RiskDecision (take / avoid / ask_s2 with its
reason, the risk, support, gain and budget left) and new_episode().
You get for free: every decision on a refused or unknown action recorded with its rationale; a hard prediction is yours to never take (Protect and RiskBudget never do).
Stability: stable.
Protect ¶
The verdict is final (the default): accept → take, refuse → avoid, unknown → ask System 2.
RiskBudget ¶
Justified risk within a per-episode budget (see the module docstring). max_risk_per_episode: the sum of the risks taken in one episode (in the predictions' unit, e.g. death-equivalents); min_gain_ratio: take only when gain ≥ ratio × risk; min_support: fewer cases behind a refusal → ask System 2.
LearnedGate ¶
A limit on a level (a depth, an amount, a distance) per context (a number: an experience level, a stage), learned from failures and bounded.
gate = LearnedGate("depth", expiry=10, floor=lambda xl: xl + 1, near=1, min_support=3, margin=1)
gate.failed(context=xl, level=depth) # a failure (a death) at this level
gate.arrived(context=xl, level=depth, failed=False) # evidence: reached this level; failed there or not
gate.end_episode()
gate.limit(xl) # max(floor(xl), median failure level nearby − margin) or None
gate.predict(xl, depth) # Prediction: accept below the limit, else refuse with the rate
expiry: only failures of the last expiry episodes count (a gate cannot be tightened by old deaths forever); floor:
the limit never falls below floor(context) (a number or a function of the context); near: failures at contexts
within ±near count; min_support: failures needed to set a limit at all, and arrivals needed before their rate
replaces the prior; prior: the risk predicted past the limit without enough arrivals.
limit ¶
The highest allowed level at this context, or None (no limit learned: not enough recent failures).
evidence ¶
(failure rate (Laplace), number of recorded arrivals) at this level and a context nearby.
predict ¶
Prediction for going to level at context: accept within the limit; past it refuse, with the measured
arrival failure rate and its support when there is enough evidence (else the prior, support = arrivals).