solvi.openset¶
Inputs from outside the calibration set: a threshold sized for the share of them, estimated as the stream goes, a
change flag (the Cusum of solvi.drift), and leave_out to simulate outside inputs from labelled examples.
Inputs from outside the calibration set: an open-set gate that keeps a promised error rate when some of the inputs are of a kind the decider has no answer for (a new topic, a product it never saw), and a detector that notices when the stream changes.
Every promise of act_guard / calibrate_for / System.guarantee holds "for inputs like the calibration examples". An input whose right answer is not among the options is outside that: whatever the decider answers is wrong, and a threshold calibrated without such inputs lets a share of them through, and the error among the answers given alone can end up well above the promised rate. The gate sizes the threshold for a share of such inputs and follows the share as the stream goes:
from solvi.openset import OpenSetGate, leave_out
sim = leave_out(calib, make) # make(kept options) → a decider without the others: the left-out
# inputs play "new kinds", any answer to them is wrong
gate = OpenSetGate.calibrate(known_signals, known_right, sim["novel"], max_error=0.05)
system.guarantee("intent", promise=gate, signal="act") # or, outside a System: gate.gate(decision)
How the threshold is sized. For a share π of inputs from outside, the error among the answers given alone at threshold t is (π·U + (1 − π)·W) / (π·U + (1 − π)·A), with U the share of outside inputs whose signal reaches t, A the share of known inputs that do and W those of them that are wrong. It grows with π. For each π of a grid the gate takes the lowest t at which that error passes learn-then-test on a mixture of the calibration examples in exactly that share — known and left-out ones, stratified, as many as there are — a binomial test at delta / 16 over 16 thresholds (quantiles of the known signals); a larger share never gets a lower threshold. So, for each share on the grid, with probability ≥ 1 − delta over the calibration sets, its threshold keeps the promise on every stream in which at most that share of the inputs are outside (the error grows with the share), the outside inputs being like the left-out examples and the rest like the calibration ones.
What share to size for. The gate reads one indicator per decision: is the signal below the cut c — the signal at which
known and left-out inputs differ most? Over each of the windows in track (the last 25 and the last 200 decisions) it
bounds the share of such signals from above (Clopper–Pearson at delta) and turns the bound into a share of outside
inputs through the known and left-out shares below c; the threshold is the one for the largest of these shares, never
below the one for min_share (0.1). The short window catches a sudden change within a few dozen decisions, the long one
a small share. A share above the largest one any threshold can serve escalates everything.
The flag. One-sided Bernoulli CUSUMs on the same indicator, one for each share in design (0.1, 0.3, 0.6), against q0,
an upper bound of the share below c among known inputs; a flag when any of them reaches h. h is set by simulation at
calibration: on streams of horizon decisions in which the indicator is Bernoulli(q0), the chance of a flag is at most
alpha (solvi.drift.Cusum, the detector DriftMonitor uses too: at least 2,000 and 20 / alpha simulated streams, a
fixed seed — the report gives h and the simulated rate; alpha below 1e-4 is refused). After a flag the share is
also estimated from the decisions since that CUSUM last stood at zero (its estimate of the change point; at most the
last window), and the largest estimate is used. A DriftMonitor (solvi.drift) can be passed as a second detector: its
flag starts the same estimate (from its window); the flag and its reason are in every record's state. The flag tells;
the estimate keeps the promise — it moves the threshold before any flag.
What it does not do: it does not know what the new inputs are and does not learn them (labels, a new option and a recalibration are the caller's). The left-out inputs stand in for the real outside ones; when those look more familiar to the decider than the stand-ins did, the bound is optimistic — a group of new intents can be harder to tell from the known ones than every left-out fold. Between a sudden change and the moment the short window sees it the threshold is the one for the share seen before: a stream that jumps to a large share of outside inputs gets a few wrong answers in that time, which matter when little is answered after it. How to use it: docs/guide.md, "Inputs from outside the calibration set".
OpenSetGate ¶
OpenSetGate(error, delta, shares, thresholds, cut, f_known, f_novel, q0, design, cusum, min_share, window, report, monitor=None, track=(25, 200), min_track=50)
See the module docstring. Made by OpenSetGate.calibrate(...); stateful: observe(signal) after every decision it gated (System.guarantee does it), threshold_of() → the threshold for the next one.
calibrate
classmethod
¶
calibrate(known_scores, known_correct, novel_scores, *, max_error=0.05, delta=0.1, min_share=0.1, shares=SHARES, track=(25, 200), min_track=50, alpha=0.01, horizon=1000, design=(0.1, 0.3, 0.6), window=500, min_support=10, monitor=None, grid=16, seed=0)
known_scores / known_correct: the decider's signal and right / wrong on calibration examples like the stream's known inputs (the deployed decider on held-out labelled examples); novel_scores: its signal on inputs whose answer is not among its options (leave_out(...)["novel"], or real outside examples). max_error, delta: the promise. min_share: the share of outside inputs the threshold is always sized for (0: the plain learn-then-test threshold while the stream looks like the calibration examples). shares: the grid of shares. track: the windows (decisions) the share is estimated over, each once it has min(window, min_track) decisions; () or None: only after a flag. alpha, horizon: the detector's false flags (≤ alpha within horizon decisions of a stream like the calibration examples). design: the shares of outside inputs the detector's CUSUMs are tuned to. window: at most this many decisions since the change are used for the estimate after a flag. min_support: a threshold must let through this many known examples. monitor: a solvi.drift.DriftMonitor whose flag also starts the estimate. grid: thresholds tried per share (learn-then-test, Bonferroni). → OpenSetGate
observe ¶
One decision's signal (after it was gated) → the state after it.
gate ¶
Gate a solvi Decision (or a number) as System.guarantee would: below the current threshold the decision escalates with the reason; its extra["open_set"] records the threshold and the state; the gate observes it. → the decision (a number: whether it may be answered alone).
run ¶
The gate over a stream of signals, from a fresh state, without touching this gate's → [(threshold, alone, level)] and the flag (decision number or None). For backtests on labelled data.
leave_out ¶
The leave-options-out simulation: the options are split into folds groups; for each, a decider is made without
that group (make(kept options) → a callable decider: a DecisionPart, or anything with decide(list) / call) and
asked about every example — the examples of the kept options give known signals (right or wrong), those of the
left-out group give signals of inputs whose answer is not among the options. examples: [(input, correct option)];
label: the option as the decider names it, from an example's correct answer (default: as given). → {"known":
(signals, right), "novel": signals, "groups": [left-out options per fold]}. Signals: the act probability when the
decider gives one, else the confidence.