solvi.testing¶
Decision regression tests from cases.json.
Decision regression tests from cases.json files — the format the gallery entries use — for solvi test and pytest.
A cases.json sits next to the task.py it tests (a module defining cat and QUESTIONS, or system(), and optionally
prepare(state)), or names it: {"task": "path/to/task.py", "cases": [...]}. When the System needs more than the catalog
(a head fitted or a rule list learned from examples), {"system": "run.py:system", "cases": [...]} names the function that
builds it (paths relative to the cases file). A case:
{"name": "legal threat",
"state": {...}, # init_state (prepare(state) runs first, if the task has one)
"expected": {"priority": "high", "tags": ["refund_request"], "route": null},
"status": {"priority": "forced"}, # optional: ok / forced / abstain per question
"safeguards": {"priority": ["hard_check"]}, # optional: the safeguard kinds that fire for a question (exactly)
"ask": ["priority", "route"]} # optional: ask only these questions (default: all)
("gold" — the key a honesty set uses — is read as "expected", so one file serves solvi test and solvi honesty.)
An expected answer is written as in solvi's JSON: an option, a list for multi-label, a number, "
from solvi.testing import run_path
for file in run_path("gallery"):
print(file.path, file.passed, [c.problems for c in file.cases if not c.ok])
fuzz(system, state) asks with random, missing and wrong-typed fields and reports exceptions that escape solvi and answers
given with a confidence that is not a number in [0, 1]. See docs/testing.md.
Suite
dataclass
¶
A cases file, loaded: its task module, the cases, and where the task came from.
CaseResult
dataclass
¶
One case of a cases.json file after solvi test: its name, the problems found (none: ok), the answers given
{question: [answer, status]} and the time it took in ms.
FileResult
dataclass
¶
One cases.json file after solvi test: its CaseResults, passed (how many are ok), ok (all of them, and the
file loaded), error (why the file or its task could not be loaded).
load ¶
A cases file → Suite. The task is {"task": ...} in the file (relative to it), else task.py next to it.
check ¶
Deprecated (removed in 0.9): run_case(system, case, state).
run_case ¶
Ask one case and compare → CaseResult. Nothing is written to the system's own storage (a test input is not a decision) unless store=True. Exceptions from ask are reported as a crash, not raised. A case that cannot fail is a problem too: a key that is not one of CASE_KEYS (a misspelled "expcted"), a status or safeguards entry for a question that was not asked, a case that expects nothing.
run_file ¶
Run every case of a cases file → FileResult (with fuzz_n > 0, each case is also fuzzed; crashes are problems). store=True: the cases (never the fuzz mutations) are saved to the system's storage, when it has one.
run_path ¶
Every cases file under the paths → [FileResult].
mutations ¶
n mutated copies of a state → [(description, state)]: a field removed, set to None, to a value of another type (number ↔ text, a list, a dict, NaN, a huge number), text cut short or garbled, an unexpected extra field; one level into dicts and lists too. Deterministic for a seed.
fuzz ¶
Ask with n mutated states (see mutations) → one line per problem: "crash — mutation: exception" for an exception that escaped System.ask, "invalid — mutation: ..." for an answer given with a confidence that is not a number in [0, 1] (NaN, inf). Abstentions are fine. Numeric warnings the mutations provoke are silenced.