Skip to content

solvi.experimental.oncalib

Calibration on the fly from the outcomes an agent sees (experimental: importing it warns ExperimentalWarning). A guarantee recalibrated continuously does not keep its promise — read the risk note below before using it. The stable way is explicit: System.outcome(...) stores the labels, System.guarantee(..., corrections=True) recalibrates on them when you decide to.

EXPERIMENTAL — calibration on the fly from the outcomes the agent sees. Importing it warns (ExperimentalWarning); its API may change or it may be removed, and nothing stable in solvi imports it.

from solvi.experimental.oncalib import OnTheFly                  # ExperimentalWarning
live = OnTheFly(system, "move", max_risk=0.05, every=50, window=500)
res = system.ask(state)
...                                                 # later, the environment shows what the decision led to
live.outcome(res, "west", note="Link lost a heart going east")   # stored as an "outcome" label (System.outcome);
                                                                 # every 50 labels: recalibrated on the last 500
live.drifted()                                      # a drift flag: recalibrate now, on the labels since the flag

What it does: every outcome is stored as a label of its decision through System.outcome (source "outcome", the same source check as every label). After every every new labels of the question — and when drifted() is called — the question's guarantee is set again (System.guarantee) on the last window outcome labels (after a drift flag, only those recorded since it), as long as there are at least min_labels. Each recalibration is kept in history with its report and the ids of the labels it used.

RISK — why this is not in the stable path. A guarantee's promise ("P(answered alone and wrong) ≤ 5%") holds for decisions drawn like its calibration examples, calibrated once. solvi's own measurements found the promise broken when a guarantee was recalibrated continuously, which is why the stable path does not do it. Why: - the outcomes an agent sees are not a random sample of its decisions (it sees the consequences of the moves it made, more often of the bad ones; a decision that abstained has no outcome at all), so the calibration set is biased; - the threshold is moved again and again on overlapping windows, each time to the edge of the promise on the latest labels: the promise of one calibration does not hold for a threshold chosen by many — the error drifts above the stated level without any single step looking wrong; - a window short enough to follow a change is too short for a strict method (the guarantee abstains on everything, or weak="raise" refuses to calibrate). Use it to explore, never to make a promise you report. For a promise: calibrate explicitly, on labels drawn at random (or all the outcomes of a period), on a schedule or after a drift flag — System.guarantee(question, examples, corrections=True) — and say on what it was calibrated.

OnTheFly

OnTheFly(system, question, *, every=50, window=500, min_labels=30, **guarantee)

Recalibrate a question's guarantee from the outcomes the agent sees (EXPERIMENTAL; see the module docs and its risk note).

system: a System with storage; question: the question whose guarantee is recalibrated. every: recalibrate after this many new outcome labels; window: on the last this many (None: all); min_labels: fewer → not recalibrated (the labels are still stored). The other keyword arguments go to System.guarantee (max_risk=, max_error=, method=, delta=, signal=, ...).

outcome

outcome(decision, value, *, note=None, by=None)

Store the outcome as a label (System.outcome) → its stored id; recalibrate when every new labels came.

drifted

drifted(why='drift flag')

A drift flag (a DriftMonitor, an open-set gate): from now on only the labels stored after it count, and the guarantee is recalibrated at once when there are enough of them.

labels

labels()

The outcome labels of the question the next recalibration reads → [correction dicts], oldest first.

recalibrate

recalibrate(why='asked')

Set the question's guarantee again on labels() → its calibration report, or None when there are fewer than min_labels (the guarantee in force stays).