solvi.calibfile¶
Calibrations as files (part.save_calibration / load_calibration) and solvi calibrate.
Calibrations as files, and solvi calibrate.
act_guard, calibrate_for and conformal set a decision's thresholds in memory. A calibration file keeps them, so a
catalog calibrated once (on a few hundred labelled examples of your stream) loads the same thresholds every time it starts:
info = part.act_guard(examples, max_risk=0.10)
part.save_calibration("team.calib.json")
# in the catalog module, after making the part and before registering it (groups add the part's inputs):
part = model.decision("team", "Which team?", "email", TEAMS)
part.load_calibration("team.calib.json")
part.question(cat)
A file holds the escalation thresholds (escalate_below, act_threshold, per group; for a combination the shared
threshold, its scale and — on the rank scale — each member's sorted calibration signals, at most 1024 per member),
the guarantee record the trace shows, the conformal set, and what they were fitted for: the question (task, options,
kind) and the fingerprint of the model behind it — the checkpoint and this question's adaptation for a DecisionPart, every member for a Cascade / Vote /
Route. load_calibration refuses a file made for another question or another model (a threshold on one model's
confidence says nothing about another's); strict=False loads it anyway. After loading, the part's fingerprint is the
one it had right after calibrating, so stored decisions replay against it.
Thresholds per group by fact names load as they are; by a function, pass the same function again:
part.load_calibration(path, groups=my_grouping) (its code is fingerprinted and must match).
solvi calibrate myapp.decisions:system team labels.csv --risk 0.1 [--groups domain,task] [--method crc|ltt]
[--conformal 0.9] [--out team.calib.json] [--json]
LABELS: a CSV or JSON-lines file, one labelled example per row: label (the correct answer; a multi-label answer is a
JSON list, or "a|b" in a CSV; "email), or one text / input column, or else the other columns as a state. Group columns (--groups) are read as
facts. Exit status: 0 — something is answered alone; 1 — the calibration escalates everything (no threshold holds);
2 — usage errors.
save ¶
Write a part's calibration to a JSON file (and its LoRA adapter, if it has one, next to it) → path.
load ¶
Apply a calibration file to a part (see the module docs) → the part. strict: refuse a file made for another question or another model.
examples_of ¶
Rows (dicts with "label") → [(Facts, label)] for act_guard / calibrate_for / conformal.
find_part ¶
The decision behind a question (its rule) or a catalog part by name → a DecisionPart / combination, or None.