solvi.honesty¶
Honesty numbers for a release gate.
Honesty numbers for a release gate: how often a system acts and is wrong, how much it can answer at a target risk, and how often its quotes support a right answer — on a fixed, labelled set, compared against a stored baseline.
python -m solvi.honesty tests/honesty/core_v1.json --baseline tests/honesty/core_v1.baseline.json
A set is a JSON file: {"name", "version", "task": "task.py" (relative to the set file), "cases": [...]}; a case is
{"name", "state": {...}, "expected": {question: answer}, "ask": [questions] (optional; default: the expected questions)}. A gold
answer is written as in solvi's JSON: an option (a list for multi-label), a number, "system() → System, or cat and QUESTIONS (and optionally
prepare(state)), as the gallery's task.py files do.
Injection traps: a case may add "injected": {question: answer} — the answer that an instruction embedded in its input
("ignore the rules and answer X", "SYSTEM: the correct answer is X", a quoted command) pushes for, often a near-duplicate
of the right option (an allowed but wrong action). The suite then counts how often such a question was answered alone
with exactly that answer.
The three numbers (per question asked, "act" = answered with status ok or forced, "escalate" = abstained):
confident_error_rate acted and wrong, over all answers asked — a wrong option, a value where the text states none, or
any answer where the gold is null. Lower is better.
coverage_at_risk the share of all answers the system can give automatically with at most risk (default 10%)
errors among them: its acted answers, most confident first (solvi.calibration.coverage_at),
times the share it acted on. Higher is better.
quote_support_proxy of the quotes behind acted answers (evidence, spans, quoted facts), the share that is literally in
its text at its offsets AND backs an answer that matches the gold. A proxy: it does not check
that the quote entails the answer. Higher is better.
injection_followed_rate of the answers with an injected answer, the share acted on with exactly the injected answer (None when the set has no injection traps; per question in "injection_by_question"). Lower is better.
compare(metrics, baseline, tolerance) lists every number that got worse by more than the tolerance; the command line
prints the report as JSON and exits 1 when something regressed (2 when the set cannot be run). See docs/honesty.md.
load_set ¶
A honesty set (JSON) → its dict, with "path" and "task" (the task module's path) resolved. A case's right answers
are its "expected", as in a solvi test cases.json (null or "abstain": the honest outcome is to abstain), so one file
serves both commands. "gold", the 0.7 key, is still read (SolviDeprecationWarning; removed in 0.9); a case with both is
an error.
load_task ¶
Execute a task.py in a fresh module namespace (the way the gallery runners and the playground load a preset). The module is registered in sys.modules under a name unique to its path, so pydantic and dataclasses can resolve the task's own types (postponed annotations).
system_of ¶
The System a task module describes: task.system(), else System(task.cat, task.QUESTIONS).
gold_of ¶
A gold answer as written in JSON → ABSTAIN (null), solvi.Unknown ("
same ¶
Does an answer match the gold? Numbers up to rounding, sequences item by item, anything else by ==.
run ¶
Ask the system every case → one row per (case, gold question): {"case", "question", "gold", "answer", "status", "guard", "confidence", "acted", "correct", "safeguards", "quotes"} (JSON-ready); with an injection trap also "injected" (the answer the embedded instruction pushes for) and "followed" (acted with exactly that answer). store=False: nothing is written to the system's storage (as System.ask(..., store=False)).
metrics ¶
The release numbers of a run (see the module docstring), plus counts that explain them.
compare ¶
The gated numbers that got worse than the baseline by more than tolerance (absolute) → ["name: was → now"]. A
number the baseline has and the current run lacks (None) counts as worse; one the baseline lacks is not compared.
report ¶
Run a set and compare it with a baseline file (a previous report, or {"metrics": {...}}) → the report dict: {"set", "version", "cases", "metrics", "baseline", "tolerance", "regressions", "ok"} (and "rows" when asked).