Skip to content

solvi.guarantee

A guarantee on any signal: a calibrated threshold with a stated promise on a question's answer (System.guarantee), a computed fact or any scalar (calibrate).

A guarantee on any signal: a calibrated threshold with a stated promise on a question's answer — given by a rule, a fitted head (fit), a model decision — on a fact the catalog computes (a trust score, an agreement share), or on any scalar outside a System.

report = system.guarantee("pair", examples, max_risk=0.01)          # conformal risk control on the answer's confidence
report = system.guarantee("pair", examples, max_error=0.02)         # learn-then-test: error among the answered ≤ 2%
report = system.guarantee("act", examples, max_risk=0.03, signal="trust", correct=judge)     # a computed fact
report = system.guarantee("correct", examples, max_error=0.3, answer="yes")   # one-sided: "yes" alone when P(yes) ≥ t

from solvi.guarantee import calibrate
p = calibrate(scores, correct, max_error=0.05)                      # any scalar: p.threshold, p.allows(s), p.report

examples: [(init_state, correct answer)] held out from whatever fitted the answer (or folds=k for a fast head fitted on them). The question is asked on each (nothing stored, nothing counted), the signal and right / wrong are collected, and the threshold is chosen by one of three methods — each makes a different promise, for inputs like the calibration examples (exchangeable with them: the same stream, not a new domain):

method       parameter  promise
"crc"        max_risk=r     P(answered alone and wrong) ≤ r — a share of ALL inputs, answered or escalated; on average
                        over calibration sets (conformal risk control: (wrong answered + 1) / (n + 1) ≤ r)
"ltt"        max_error=e    the error AMONG the answers given alone ≤ e, with probability ≥ 1 − delta over the
                        calibration set (learn-then-test: a binomial test per threshold, Bonferroni over ≤ 64)
"empirical"  max_error=e    none: the error among the answered was ≤ e on the calibration examples only

max_risk= selects "crc", max_error= selects "ltt" (method="empirical" to ask for the plain one). groups= gives a threshold per group (a fact name, a hierarchy of fact names, a function of facts, or "answer": the answer the question would give — "among the inputs answered 'match', ..."), with the promise inside every group (after HG-CRC; see solvi.calibration).

What is checked before a threshold is set, so that a promise is never made on a signal that cannot carry it: - separation: the signal must rank the right answers above the wrong ones on the calibration examples (a one-sided Mann–Whitney test at 5%); one that does not — a model approving its own answers at chance — raises (weak="warn": a warning, and the warning is recorded in the guarantee); - support: a threshold must have at least min_support calibration examples at or above it (default 10); none that does → everything escalates, and the report says why; - feasibility: with n examples, conformal risk control cannot certify a risk below 1 / (n + 1) — the report says so.

The report gives, on the calibration examples, both "risk" (answered alone and wrong, a share of all) and "error" (among the answered), the share answered, the threshold's support and the separation (AUROC, p). Every answer of a guarded question records the promise, the signal's value and the threshold in its result (extra["guarantee"]) and in the trace (a hashed record "guard:"); below the threshold the question abstains with the reason. Replay with the System re-derives the verdict; a guarantee changed since the decision is a mismatch.

GuaranteeWarning

Bases: UserWarning

A guarantee was set on a signal that did not show it can carry it (weak="warn"), or on too few examples.

Promise

Promise(method, level, delta, threshold, nodes, by, report, why, min_group, signal='signal')

A calibrated threshold on one signal with the promise it keeps — the result of calibrate(). allows(signal, group=None) → answer alone? threshold_of(group) → (threshold, node, info). report: what it does on the calibration examples; record(): what every gated decision records.

threshold_of

threshold_of(group=None)

The threshold for an input of this group (a value or a path; the deepest calibrated group on its path, else the rest of the stream) → (threshold, node, info); without groups: (threshold, None, None).

allows

allows(signal, group=None)

Answer alone at this signal? (False for a missing or non-finite signal.)

QuestionGuard

QuestionGuard(question, promise, signal=None, answer=None, groups=None)

A Promise (or another object with its interface) bound to a question of a System: which signal it reads, which answer it lets through (answer=: one-sided), how groups are read. Applied to every answer of the question by System.ask.

value

value(result, vals, trace)

The signal of one answer → float, or a reason it cannot be read (str).

promise_text

promise_text(method, level, delta=None, groups=False, n_groups=1)

What a method promises, in words (recorded with every decision it gates).

separation

separation(s, ok)

Does the signal rank right answers above wrong ones? → {"auroc", "p" (one-sided Mann–Whitney), "n_right", "n_wrong", "tested"}; untested (p None) with fewer than MIN_CLASS of either.

calibrate

calibrate(scores, correct, *, max_error=None, max_risk=None, method=None, delta=0.1, groups=None, min_group=100, min_support=10, weak='raise', signal='signal')

A threshold with a promise on any scalar signal: answer alone when signal ≥ threshold. scores: the signal per calibration example (None / NaN / −inf: never answered alone; +inf: always — a forced answer); correct: was the answer right (booleans). max_risk= → conformal risk control ("crc"), max_error= → learn-then-test ("ltt"), or method="empirical" with max_error=; the promises are in the module docstring. delta: learn-then-test's confidence and, with groups, the binomial bound of conformal risk control per group (delta=None there: on average per group). groups: one group (a value or a path) per example; min_group: a group with fewer examples is pooled with its parent. min_support: a threshold must have this many examples at or above it. weak: "raise" (default) or "warn" when the signal does not separate right from wrong. → Promise (its threshold is inf — everything escalates — when none keeps the promise; report["why"] says why).

guard_question

guard_question(system, question, examples=None, *, max_error=None, max_risk=None, method=None, delta=0.1, signal=None, answer=None, groups=None, min_group=100, min_support=10, correct=None, folds=None, seed=0, weak='raise', promise=None)

Put a calibrated threshold with a stated promise on a question of system (System.guarantee): every later answer is let through alone only when its signal ≥ the threshold, else the question abstains with the reason; the promise, the signal and the threshold are recorded with every answer. examples: [(init_state, correct answer)], not used to fit the answer (or folds=k: the question's fitted head is refitted k times, each example scored by a head that did not see it — the head that answers is fitted on all of them, so the promise is then approximate). False removes the question's guarantee. signal: "confidence" (default — the answer's confidence), "act" (the act probability of the model decision that answers it), a fact name (any number the catalog computes or the input gives), or a function (result, facts) → number. answer: one-sided — the question answers this alone when its signal (default: the answer's probability, P(answer)) ≥ the threshold, whatever the most probable answer is, and abstains otherwise ("return the query when P(right) ≥ 0.43"). correct: a function (result, label) → bool for answers judged otherwise than by equality (a quote that overlaps the gold one). groups: a fact name, a hierarchy of fact names, a function of facts, or "answer". promise: an already calibrated Promise to attach instead of calibrating here. The other arguments, the methods and their promises: calibrate() and the module docstring. → the calibration report ({"threshold", "n", "answered", "error", "risk", "support", "separation", "promise", ...}; for an attached promise, its report).

apply_guards

apply_guards(system, results, trace, vals, flow, live=True)

Gate the answers of the guarded questions (called by System.ask and System.answers_of). live: record each verdict in the trace (and let a stateful gate observe the signal); False — re-derive the verdicts of a recorded trace (a stateful gate's recorded threshold is taken as it was).

replay_guard

replay_guard(r, system, vals)

A recorded guard verdict (kind "guard"): its inputs must be the recorded facts, the verdict must follow from the recorded signal and threshold, and with the System its guarantee must be the one the question has now.