Skip to content

solvi.calibration

Reliability, expected calibration error and coverage at a target accuracy.

Is a confidence honest? Reliability, expected calibration error and coverage at a target accuracy — for any model or question (solvi answers, a decider, a head).

from solvi.calibration import coverage_at, ece, reliability, threshold_for
coverage_at(conf, correct, 0.9)     # share of cases you can answer automatically at ≥ 90% accuracy
threshold_for(conf, correct, 0.9)   # the confidence threshold that gives it (e.g. for Question(min_confidence=...));
                                    # None when fewer than min_n=10 cases stand behind it

conf — confidences in [0, 1]; correct — whether each answer was right (booleans or 0/1). evaluate(system, question, examples) collects both from a System on labelled examples (abstentions count as not covered).

Thresholds with a guarantee (used by DecisionPart.act_guard, calibrate_for(method="ltt") and conformal):

crc_threshold(score, wrong, risk=0.1)       # P(answered alone and wrong) ≤ 10% of all questions
ltt_threshold(score, wrong, error=0.1)      # error among the answered ≤ 10%, with probability ≥ 90%

They hold for inputs like the calibration examples, not under a shift of domain: calibrate on your own labelled data.

reliability

reliability(conf, correct, bins=10)

Equal-width confidence bins → [{"lo", "hi", "n", "confidence", "accuracy"}] (bins without cases are left out).

ece

ece(conf, correct, bins=15)

Expected calibration error: the case-weighted mean |confidence − accuracy| over equal-width bins (0 = honest).

coverage_at

coverage_at(conf, correct, accuracy=0.9)

The largest share of cases a confidence threshold can let through with accuracy at least accuracy: the cases with confidence ≥ t for the best t (equal confidences are taken or left together, so the result does not depend on the order of the rows); 0 if no threshold reaches it.

threshold_for

threshold_for(conf, correct, accuracy=0.9, min_n=10)

The lowest confidence threshold t (an observed confidence) such that the cases with confidence ≥ t — all of them, ties included — are right at least accuracy of the time: answer when confidence ≥ it. min_n: at least so many cases must be at or above the threshold (default 10); a threshold that rests on fewer says nothing about new inputs. None if no threshold reaches the accuracy on at least min_n cases (so it can be None while coverage_at is above 0: pass min_n=1 for the threshold of coverage_at whatever its support). Empirical — no guarantee on new inputs; see ltt_threshold for one.

accuracy_at

accuracy_at(conf, correct, threshold)

Answer only when confidence ≥ threshold → (accuracy of the answered part, coverage).

summary

summary(conf, correct, accuracy=0.9, bins=15, min_n=10)

{"n", "accuracy", "ece", "coverage_at", "threshold"} in one call ("threshold": threshold_for with min_n — None when fewer than min_n cases support it, whatever "coverage_at" says).

evaluate

evaluate(system, question, examples, accuracy=0.9)

Ask a System on labelled examples [(init_state, correct answer)] → summary() over the answered ones plus "answered" (share not abstained) and "accuracy_all" (abstentions counted as wrong); also "conf" and "correct" lists.

crc_threshold

crc_threshold(score, wrong, risk=0.1)

Conformal risk control: the lowest threshold t such that (Σ 1[score ≥ t and wrong] + 1) / (n + 1) ≤ risk; answer alone when score ≥ t. Guarantee (inputs exchangeable with the calibration examples): P(answered alone and wrong) ≤ risk — a share of ALL questions, not of the answered ones. inf when the examples are too few or too hard for the risk.

separation

separation(score, correct)

Does a signal tell right answers from wrong ones? → (AUROC, z) over the examples: the probability that a right answer's score is above a wrong one's (ties count half), and the one-sided Mann-Whitney z of that against chance (0.5); (None, None) when the examples are all right or all wrong. z < 1.645: not better than chance at the 5% level — a threshold on the signal then escalates right and wrong answers alike, and the answers it lets through are wrong about as often as all of them (act_guard warns).

check_rate

check_rate(name, v, zero=False)

A target error rate / risk / delta must be strictly between 0 and 1 (0 cannot be certified from finitely many examples; 1 promises nothing) → ValueError with the name otherwise. zero=True also allows 0 (an empirical target: no error on the calibration examples).

ltt_grid

ltt_grid(score, size=LTT_GRID)

The default learn-then-test grid: the distinct finite scores, or size of them at evenly spaced quantiles when there are more. It reads the scores only, never the labels, so fixing it on the calibration examples keeps the guarantee; it follows the scores' own scale (an LLM's confidence packed near 1 as well as an act probability).

ltt_threshold

ltt_threshold(score, wrong, error=0.1, delta=0.1, grid=None)

Learn-then-test: the lowest threshold t on a fixed grid such that the error AMONG the cases answered alone (score ≥ t) is ≤ error with probability ≥ 1 − delta over the calibration examples (binomial test, Bonferroni over the grid). A stronger promise than crc_threshold, so it often allows no automatic answers at all (inf). grid=None: at most 64 quantiles of the distinct calibration scores (ltt_grid; before 0.7 a fixed linspace(0.2, 0.995, 32), which let nothing through for a decider whose scores sit above 0.995).

conformal_quantile

conformal_quantile(scores, alpha)

The ⌈(n + 1)(1 − alpha)⌉-th smallest score; inf when there are too few (n < 1/alpha − 1).

set_scores

set_scores(p, ordinal=False, has_unknown=False)

Non-conformity of every answer of one question, from its probabilities p (the last one "not stated" when has_unknown): 1 − p (LAC); ordinal=True — for scores and numbers — the mass added before an answer while an interval grows from the mode towards the more probable neighbour ("not stated" competes as one more candidate), so the answer set is always one contiguous interval.

group_path

group_path(g)

A group as a path from the top of the hierarchy: a tuple of strings. A scalar is a one-level group, a tuple or list a path (domain, task); the path stops at the first None; None alone is the whole stream ().

group_nodes

group_nodes(paths, min_group=100)

Which groups get a threshold of their own. Deepest level first: a group whose examples not yet taken by a group under it number at least min_group becomes a node and takes them; the whole stream () takes the rest. → ({node: [example indices]}, the node of each example). Depends on the group sizes only, not on the labels.

node_of

node_of(path, nodes)

The node whose threshold applies to an input of this group: the deepest node on its path (a new or small group falls back to its parent's, then to the whole stream's).

loss_budget

loss_budget(n, risk, delta=None)

The most examples of n that may be answered alone and wrong for a threshold to certify risk: conformal risk control (delta=None: (k + 1) / (n + 1) ≤ risk, a promise on average) or a binomial test at level delta (k with P(Binomial(n, risk) ≤ k) ≤ delta: the risk ≤ risk with probability ≥ 1 − delta). −1: no threshold can.

certify_groups

certify_groups(losses_of, paths, risk=0.1, min_group=100, delta=0.1)

Group-wise risk control over a hierarchy: losses_of(indices) → (candidate thresholds ascending, answered-alone- and-wrong count at each — non-increasing) for those examples. Each node (group_nodes) takes the lowest candidate whose count is within loss_budget(n, risk, delta / nodes) — Bonferroni over the nodes, so with probability ≥ 1 − delta the promise holds in every group at once; delta=None: conformal risk control per group (each group on average, no correction needed). → ({node: {"threshold", "n", "index"}}, node of each example); inf where no threshold certifies.

group_thresholds

group_thresholds(score, wrong, groups, risk=0.1, min_group=100, delta=0.1)

certify_groups for one signal (answer alone when score ≥ the threshold of the input's node) → {node: threshold}; apply with node_of(group, nodes). groups: one per example — a value, a path (domain, task) or None.