solvi.experimental.oncalib¶
Calibration on the fly from the outcomes an agent sees (experimental: importing it warns ExperimentalWarning). A
guarantee recalibrated continuously does not keep its promise — read the risk note below before using it. The stable
way is explicit: System.outcome(...) stores the labels, System.guarantee(..., corrections=True) recalibrates on them
when you decide to.
EXPERIMENTAL — calibration on the fly from the outcomes the agent sees. Importing it warns (ExperimentalWarning); its API may change or it may be removed, and nothing stable in solvi imports it.
from solvi.experimental.oncalib import OnTheFly # ExperimentalWarning
live = OnTheFly(system, "move", max_risk=0.05, every=50, window=500)
res = system.ask(state)
... # later, the environment shows what the decision led to
live.outcome(res, "west", note="Link lost a heart going east") # stored as an "outcome" label (System.outcome);
# every 50 labels: recalibrated on the last 500
live.drifted() # a drift flag: recalibrate now, on the labels since the flag
What it does: every outcome is stored as a label of its decision through System.outcome (source "outcome", the
same source check as every label). After every every new labels of the question — and when drifted() is called —
the question's guarantee is set again (System.guarantee) on the last window outcome labels (after a drift flag, only
those recorded since it), as long as there are at least min_labels. Each recalibration is kept in history with its
report and the ids of the labels it used.
RISK — why this is not in the stable path. A guarantee's promise ("P(answered alone and wrong) ≤ 5%") holds for
decisions drawn like its calibration examples, calibrated once. solvi's own measurements found the promise broken when
a guarantee was recalibrated continuously, which is why the stable path does not do it. Why:
- the outcomes an agent sees are not a random sample of its decisions (it sees the consequences of the moves it made,
more often of the bad ones; a decision that abstained has no outcome at all), so the calibration set is biased;
- the threshold is moved again and again on overlapping windows, each time to the edge of the promise on the latest
labels: the promise of one calibration does not hold for a threshold chosen by many — the error drifts above the
stated level without any single step looking wrong;
- a window short enough to follow a change is too short for a strict method (the guarantee abstains on everything,
or weak="raise" refuses to calibrate).
Use it to explore, never to make a promise you report. For a promise: calibrate explicitly, on labels drawn at random
(or all the outcomes of a period), on a schedule or after a drift flag — System.guarantee(question, examples,
corrections=True) — and say on what it was calibrated.
OnTheFly ¶
Recalibrate a question's guarantee from the outcomes the agent sees (EXPERIMENTAL; see the module docs and its risk note).
system: a System with storage; question: the question whose guarantee is recalibrated. every: recalibrate after this many new outcome labels; window: on the last this many (None: all); min_labels: fewer → not recalibrated (the labels are still stored). The other keyword arguments go to System.guarantee (max_risk=, max_error=, method=, delta=, signal=, ...).
outcome ¶
Store the outcome as a label (System.outcome) → its stored id; recalibrate when every new labels came.
drifted ¶
A drift flag (a DriftMonitor, an open-set gate): from now on only the labels stored after it count, and the guarantee is recalibrated at once when there are enough of them.
labels ¶
The outcome labels of the question the next recalibration reads → [correction dicts], oldest first.
recalibrate ¶
Set the question's guarantee again on labels() → its calibration report, or None when there are fewer than min_labels (the guarantee in force stays).