Skip to content

solvi.drift

Drift: has the stream of decisions moved away from the one the thresholds were calibrated on?

Drift: has the stream of decisions moved away from the one the thresholds were calibrated on?

A threshold from act_guard holds "for inputs like the calibration examples". When the inputs change, the promise is kept by escalating more — or the model goes on answering alone and is wrong more often — and nothing says so. A DriftMonitor compares the last window decisions of one question with a reference window and names what moved.

from solvi.drift import DriftMonitor
mon = DriftMonitor(window=100)                 # the first 100 decisions observed are the reference
for text in stream:
    d = part(text)
    rep = mon.observe(d)                       # or mon.observe(d, label=truth) when the truth is known
    if rep["drift"]:
        ...                                    # rep["why"]: what moved; recalibrate (act_guard), ask for labels

mon = DriftMonitor(window=100).set_reference(decisions, labels)      # or: an explicit reference

Without labels, four signals of any decision: the share answered alone (two-proportion z test, Fisher's exact test once it is small), the distribution of the answers (chi-square test of homogeneity, with the total-variation distance as its effect size; answers expected fewer than 5 times in a window are pooled into one bin, and when fewer than two bins are left the signal is not tested — the report's "not_tested" says so: a question with K answers needs a window of about 5 × K), the mean confidence and the mean act probability (z tests). With labels (at least min_labelled in both windows): accuracy with escalations counted as errors, and — among the labelled decisions answered alone — the calibration error (ECE) and coverage_at(accuracy), by a permutation test run once per window decisions (both are noisy on a hundred cases). A signal is flagged when its test is significant AND its effect is at least the minimum — on a large window a tiny shift is significant and is not drift. drift is true when at least min_signals signals are flagged.

Two kinds of test. The window tests above compare the last window decisions with the reference at every decision; they are repeated and there are several of them, so each is held to ¾·alpha / (signals tested × decisions in horizon) — a union bound, which makes them slow on real streams. A sequential test runs beside them (sequential=True, the default): CUSUMs (Cusum) on the share answered alone, the mean confidence and the mean act probability, each in both directions — the increment of a decision is its shift from the mean of the reference and of every decision since, in the reference's standard deviations, minus allowance (0.25) — and a flag when one reaches h, set by simulation on streams drawn from the reference (with the reference's own error of its mean and spread): at most alpha / 4 of them flagged within horizon decisions. A CUSUM flag also needs the shift since its change point to be at least the signal's minimum. So on a stream that has not changed, the chance of a false flag within horizon decisions (default 1,000) is at most alpha (default 1%), as far as each test's p-value is exact and the reference stands for the stream; a longer stream gets alpha per horizon.

It is a signal, not a verdict: what to do — recalibrate, escalate, ask for labels — is the caller's. Simulated on independent decisions (benchmarks/drift_simulation.py): 6 of 1,152 unchanged streams were flagged within 1,000 decisions; a fall of the share answered alone from 73% to 13% was flagged 23–27 decisions later whatever the window; a change of the mix of three answers from 1:1:1 to 1:8:1, which only the window tests see, 67 / 80 / 110 decisions later with window 50 / 100 / 200. More in docs/guide.md, "Drift".

Observation dataclass

Observation(value: object, confidence: float, alone: bool, act: float | None = None, correct: bool | None = None)

One decision as the monitor reads it: the answer, its confidence, whether it was given alone, the act probability (when the model has an act head) and, when the truth is known, whether the answer was right.

of classmethod

of(d, label=None, question=None)

From a solvi Decision, a Response with question (the question's Result, and the act probability from its decision's record in the trace), a Result (answer / confidence / status — it carries no act probability), a dict of the same fields, or an Observation.

Cusum

Cusum(names, h)

One-sided CUSUMs over several channels at once: S_j = max(0, S_j + x_j) for each step's increments x (one per channel; NaN leaves a channel as it is), a flag when any S_j reaches h. Each channel's increments are a log-likelihood ratio or a standardized shift minus an allowance, so S_j stays near 0 while nothing changes and climbs after a change. h is set by simulation (Cusum.calibrate): the (1 − alpha) quantile of the largest S over horizon steps of sims streams drawn from the null — no union bound over the looks, which is what makes it faster than a window test repeated at every decision. Used by DriftMonitor (shifts of the share answered alone, the mean confidence and act probability) and by solvi.openset.OpenSetGate (the share of inputs from outside).

step

step(x)

One step's increments → the channel with the largest S (its index).

top

top()

→ (channel name, S, steps since it last stood at zero) of the largest S.

calibrate classmethod

calibrate(names, make_null, alpha=0.01, horizon=1000, sims=None, seed=0)

make_null(rng, sims) → a function giving one step's increments of sims null streams, an array (sims, channels). sims: at least 2,000 and 20 / alpha. → a Cusum with h such that at most alpha of the simulated streams reach h within horizon steps; .rate is that share.

DriftMonitor

DriftMonitor(window=100, *, alpha=0.01, horizon=1000, min_n=None, min_signals=1, min_share=0.1, min_tv=0.15, min_shift=0.05, min_accuracy_drop=0.1, min_ece_rise=0.05, min_coverage_drop=0.15, min_labelled=30, accuracy=0.9, keep=1000, seed=0, sequential=True, allowance=0.25)

See the module docstring. window: the decisions compared with the reference (and the size of a reference taken from the stream). alpha, horizon: the chance of a false flag within horizon decisions of an unchanged stream is at most alpha. min_n: the window must hold this many decisions before anything is flagged (default: half the window, at least 20). The minimum effects: min_share (answered alone), min_tv (answers), min_shift (mean confidence / act probability), min_accuracy_drop, min_ece_rise, min_coverage_drop. keep: the reports of the last keep observations are kept in history (0: none). sequential: also follow the share answered alone, the mean confidence and the mean act probability with CUSUMs (a Cusum; it gets a quarter of alpha, the window tests the rest); allowance: the shift (in the reference's standard deviations) below which a CUSUM does not climb.

set_reference

set_reference(decisions, labels=None)

Set the reference explicitly (calibrate names the thresholds with a promise elsewhere in solvi): decisions (Decision / Result / dict) and, when known, their correct answers.

observe

observe(decision, label=None, question=None)

Take one decision → its report. Until the reference is full (window decisions, unless calibrate() set it) the decision goes into the reference and the report's phase is "reference". A Response: with question= (its act probability is read from the trace); a bare Result of a model's decision has none, so the act signal is not tested for it — said once in a warning, and in every report's "not_tested" when the reference has it.

report

report()

The last window against the reference → {"phase": "monitor", "drift", "flags": [signal], "why": [text], "tests": {signal: {"reference", "window", "p", "level", ...}}, "not_tested": {signal: why}, "window": stats, "reference": stats}. "level": what p must be below — alpha shared among the signals tested and the looks of horizon decisions (see the module docstring).

window_stats

window_stats(obs, accuracy=0.9)

What a window of observations looks like → {"n", "answered", "answers", "confidence" (mean, std), "act" (mean, std, n) when there is an act probability, and with labels "labelled", "accuracy" (escalations count as errors), "alone" (n, accuracy, ece, coverage_at)}.