Skip to content

Confidence, calibration and abstention

How confidence is computed (what it means for each answer kind: the table in Answer primitives):

  • Rule answers: the minimum confidence among all extracted values and model decisions the rule's inputs depend on (and the rule's own, for a model-backed rule). A rule over dict inputs and computations only has confidence 1.0.
  • Learned heads: the head's probability of the chosen answer, multiplied by the same extraction confidence.
  • Forced answers (failed hard checks): 1.0.

Calibrate a question on held-out examples (Platt scaling on the confidence logit):

system.calibrate("total_band", heldout)       # [(init_state, correct answer)], as fit takes them

After calibration, ok answers of that question carry the calibrated confidence. With fewer than 10 usable examples, when all held-out answers are right (or all wrong), or when the confidences do not vary (a rule over plain inputs: always 1.0), only a constant shift is fitted: the calibrated confidence is then the share of right answers. The correct answers are written as for fit (True / False for a yes/no question). The held-out examples are run without the question's previous calibration, are not saved to the storage and do not count in system.stats, so calling calibrate again replaces the parameters with a fit of the same kind.

A threshold you choose yourself ("answer automatically at confidence >= 0.95") promises nothing on new inputs. For a threshold with a promise, use system.guarantee (below).

A guarantee on any question: System.guarantee

part.act_guard and calibrate_for put a promise on a model's decision. system.guarantee puts one on a question, whatever answers it — a fitted head (fit), a rule over computed facts, a model decision — and on any number the question's answer can be judged by: its confidence, a fact the catalog computes (a trust score, the share of candidates that agree), the act probability, or a function of your own:

import random
from solvi import Answer, Catalog, Question, System

cat = Catalog()

@cat.fn
def margin(score: float) -> float:          # any number the catalog computes can carry a promise
    return abs(score - 0.5)

system = System(cat, [Question("refund", "Refund the order?", Answer.yes_no())])
rng = random.Random(0)

def draw(n):                                # (input, correct answer); near 0.5 the answer is a coin flip
    return [({"score": (x := rng.random())}, x + rng.gauss(0, 0.1) > 0.5) for _ in range(n)]

system.fit("refund", draw(400), features=["score"])
report = system.guarantee("refund", draw(600), max_error=0.05)     # held out: not the 400 the head was fitted on
report["threshold"], report["answered"], report["error"], report["risk"]     # 0.84, 0.78, 0.019, 0.015
r = system.ask({"score": 0.52})["refund"]
r.status, r.why        # "abstain", "confidence 0.6229 < 0.8406 (threshold of the guarantee: error among the answers
                       #  given alone ≤ 0.05 with probability ≥ 0.9, ...); would have answered 'yes' (...)"
r.extra["guarantee"]   # {"method": "ltt", "promise", "n": 600, "signal": "confidence", "value", "threshold", ...}

The other forms:

system.guarantee("refund", examples, max_risk=0.02)                    # P(answered alone and wrong) ≤ 2% of ALL inputs
system.guarantee("refund", examples, max_risk=0.02, signal="margin")   # on a fact the catalog computes
system.guarantee("refund", examples, max_risk=0.02, groups="answer", delta=None)   # inside each answer ("yes", "no")
system.guarantee("act", examples, max_risk=0.03, signal="trust",       # right / wrong by a judge of your own
                 correct=lambda result, label: overlaps(result, label))
system.guarantee("correct", examples, max_error=0.3, answer="yes", folds=5)
# one-sided: "yes" alone when P(yes) ≥ the threshold (it may be below 0.5), else abstain; folds=5: the fitted head
# was fitted on these very examples, so each is scored by a head refitted without it (the promise is then approximate)
system.guarantee("refund", False)                                  # remove it

Which promise each method makes, for inputs like the calibration examples (the same stream, not a new domain):

method given as promise
"crc" (conformal risk control) risk=r P(answered alone and wrong) ≤ r — a share of all inputs, answered or escalated, on average over calibration sets
"ltt" (learn-then-test) error=e the error among the answers given alone ≤ e, with probability ≥ 1 − delta over the calibration set
"empirical" error=e, method="empirical" none: the error among the answered was ≤ e on the calibration examples only

Which one to use: "≤ 5% of the answers we give are wrong" is error= (learn-then-test). risk= is the cheaper promise and the weaker one when most inputs are easy: when most pairs are non-matches, "1% of all pairs" can be met while a predicted match given alone is wrong far more often — groups="answer" puts the promise inside each answer. groups also takes a fact name, a hierarchy or a function of facts, as act_guard does.

What is checked before a threshold is set, so that a promise is never made on a signal that cannot carry it:

  • separation: the signal must rank the right answers above the wrong ones on the calibration examples (a one-sided Mann–Whitney test at 5%). A judge at chance — a model approving its own queries, say — is refused with a ValueError (weak="warn" sets it anyway, warns, and records the warning with every answer);
  • support: a threshold must have at least min_support=10 calibration examples at or above it; when none does, everything escalates and report["why"] says so (a "75% right" threshold resting on one example is no threshold);
  • feasibility: conformal risk control cannot certify a risk below 1 / (n + 1) with n examples, and learn-then-test often lets nothing through on a hundred: the threshold is then inf and the report says why.

The report gives both shares on the calibration examples — "risk" (answered alone and wrong, of all) and "error" (among the answered) — with "answered", the threshold's "support", "separation" (auroc, p), "base_error" and, for conformal risk control, "must_escalate_at_least". The facts the signal and the groups read become required parts of the question (requires), so every flow computes them. Every answer of the question records its verdict as a hashed trace record (guard:<question>: the signal's value, the threshold, the promise); below the threshold the question abstains (safeguard low confidence) and says what it would have answered; a forced answer (a failed hard check) and an abstention are left as they are. Replay with the System re-derives the verdict; a guarantee recalibrated since the decision is a mismatch ("the question's guarantee changed"). solvi.guarantee.calibrate(scores, correct, max_error=...) does the same for any scalar outside a System: p.threshold, p.allows(score), p.report.