solvi.drift¶
Drift: has the stream of decisions moved away from the one the thresholds were calibrated on?
Drift: has the stream of decisions moved away from the one the thresholds were calibrated on?
A threshold from act_guard holds "for inputs like the calibration examples". When the inputs change, the promise is kept
by escalating more — or the model goes on answering alone and is wrong more often — and nothing says so. A DriftMonitor
compares the last window decisions of one question with a reference window and names what moved.
from solvi.drift import DriftMonitor
mon = DriftMonitor(window=100) # the first 100 decisions observed are the reference
for text in stream:
d = part(text)
rep = mon.observe(d) # or mon.observe(d, label=truth) when the truth is known
if rep["drift"]:
... # rep["why"]: what moved; recalibrate (act_guard), ask for labels
mon = DriftMonitor(window=100).set_reference(decisions, labels) # or: an explicit reference
Without labels, four signals of any decision: the share answered alone (two-proportion z test, Fisher's exact test once
it is small), the distribution of the answers (chi-square test of homogeneity, with the total-variation distance as its
effect size; answers expected fewer than 5 times in a window are pooled into one bin, and when fewer than two bins are
left the signal is not tested — the report's "not_tested" says so: a question with K answers needs a window of about
5 × K), the mean confidence and the mean act probability (z tests). With labels (at least min_labelled in both
windows): accuracy with escalations counted as errors, and — among the labelled decisions answered alone — the
calibration error (ECE) and coverage_at(accuracy), by a permutation test run once per window decisions (both are
noisy on a hundred cases). A signal is flagged when its test is significant AND its effect is at least the minimum — on
a large window a tiny shift is significant and is not drift. drift is true when at least min_signals signals are
flagged.
Two kinds of test. The window tests above compare the last window decisions with the reference at every decision;
they are repeated and there are several of them, so each is held to ¾·alpha / (signals tested × decisions in
horizon) — a union bound, which makes them slow on real streams. A sequential test runs beside them (sequential=True,
the default): CUSUMs (Cusum) on the share answered alone, the mean confidence and the mean act probability, each in
both directions — the increment of a decision is its shift from the mean of the reference and of every decision since,
in the reference's standard deviations, minus allowance (0.25) — and a flag when one reaches h, set by simulation on
streams drawn from the reference (with the reference's own error of its mean and spread): at most alpha / 4 of them
flagged within horizon decisions. A CUSUM flag also needs the shift since its change point to be at least the signal's
minimum. So on a stream that has not changed, the chance of a false flag within horizon decisions (default 1,000) is
at most alpha (default 1%), as far as each test's p-value is exact and the reference stands for the stream; a longer
stream gets alpha per horizon.
It is a signal, not a verdict: what to do — recalibrate, escalate, ask for labels — is the caller's. Simulated on independent decisions (benchmarks/drift_simulation.py): 6 of 1,152 unchanged streams were flagged within 1,000 decisions; a fall of the share answered alone from 73% to 13% was flagged 23–27 decisions later whatever the window; a change of the mix of three answers from 1:1:1 to 1:8:1, which only the window tests see, 67 / 80 / 110 decisions later with window 50 / 100 / 200. More in docs/guide.md, "Drift".
Observation
dataclass
¶
Observation(value: object, confidence: float, alone: bool, act: float | None = None, correct: bool | None = None)
One decision as the monitor reads it: the answer, its confidence, whether it was given alone, the act probability (when the model has an act head) and, when the truth is known, whether the answer was right.
of
classmethod
¶
From a solvi Decision, a Response with question (the question's Result, and the act probability from its
decision's record in the trace), a Result (answer / confidence / status — it carries no act probability), a dict
of the same fields, or an Observation.
Cusum ¶
One-sided CUSUMs over several channels at once: S_j = max(0, S_j + x_j) for each step's increments x (one per
channel; NaN leaves a channel as it is), a flag when any S_j reaches h. Each channel's increments are a
log-likelihood ratio or a standardized shift minus an allowance, so S_j stays near 0 while nothing changes and climbs
after a change. h is set by simulation (Cusum.calibrate): the (1 − alpha) quantile of the largest S over horizon
steps of sims streams drawn from the null — no union bound over the looks, which is what makes it faster than a
window test repeated at every decision. Used by DriftMonitor (shifts of the share answered alone, the mean
confidence and act probability) and by solvi.openset.OpenSetGate (the share of inputs from outside).
calibrate
classmethod
¶
make_null(rng, sims) → a function giving one step's increments of sims null streams, an array (sims,
channels). sims: at least 2,000 and 20 / alpha. → a Cusum with h such that at most alpha of the simulated
streams reach h within horizon steps; .rate is that share.
DriftMonitor ¶
DriftMonitor(window=100, *, alpha=0.01, horizon=1000, min_n=None, min_signals=1, min_share=0.1, min_tv=0.15, min_shift=0.05, min_accuracy_drop=0.1, min_ece_rise=0.05, min_coverage_drop=0.15, min_labelled=30, accuracy=0.9, keep=1000, seed=0, sequential=True, allowance=0.25)
See the module docstring. window: the decisions compared with the reference (and the size of a reference taken
from the stream). alpha, horizon: the chance of a false flag within horizon decisions of an unchanged stream is at
most alpha. min_n: the window must hold this many decisions before anything is
flagged (default: half the window, at least 20). The minimum effects: min_share (answered alone), min_tv (answers),
min_shift (mean confidence / act probability), min_accuracy_drop, min_ece_rise, min_coverage_drop.
keep: the reports of the last keep observations are kept in history (0: none). sequential: also follow the
share answered alone, the mean confidence and the mean act probability with CUSUMs (a Cusum; it gets a quarter
of alpha, the window tests the rest); allowance: the shift (in the reference's standard deviations) below
which a CUSUM does not climb.
set_reference ¶
Set the reference explicitly (calibrate names the thresholds with a promise elsewhere in solvi): decisions (Decision / Result / dict) and, when known, their correct answers.
observe ¶
Take one decision → its report. Until the reference is full (window decisions, unless calibrate() set it)
the decision goes into the reference and the report's phase is "reference". A Response: with question= (its
act probability is read from the trace); a bare Result of a model's decision has none, so the act signal is
not tested for it — said once in a warning, and in every report's "not_tested" when the reference has it.
report ¶
The last window against the reference → {"phase": "monitor", "drift", "flags": [signal], "why": [text],
"tests": {signal: {"reference", "window", "p", "level", ...}}, "not_tested": {signal: why}, "window": stats,
"reference": stats}. "level": what p must be below — alpha shared among the signals tested and the looks of
horizon decisions (see the module docstring).
window_stats ¶
What a window of observations looks like → {"n", "answered", "answers", "confidence" (mean, std), "act" (mean, std, n) when there is an act probability, and with labels "labelled", "accuracy" (escalations count as errors), "alone" (n, accuracy, ece, coverage_at)}.