solvi.testing.conformance¶
Conformance checks for the extension points of solvi.core (see Building blocks): each runs
your implementation through the parts of solvi that rely on it and fails with the guarantee that does not hold.
Conformance checks for the extension points of solvi.core: what the docs say you get for free, tested on your class.
from solvi.testing.conformance import check_storage, check_slow_path
def test_my_store(tmp_path):
check_storage(lambda: MyStorage(tmp_path / "decisions"), reopen=lambda: MyStorage(tmp_path / "decisions"))
Each check runs the implementation through the parts of solvi that rely on it and raises ConformanceError (an AssertionError, so pytest shows it as a failed assertion) naming the guarantee that does not hold; it returns a dict of what it checked. solvi's own test suite runs every check on every built-in (tests/test_conformance.py).
check_storage(make, reopen=None) TraceStorage: chain verifies, records round-trip, replay, query, redact,
tampering is caught, a reopened store is the same store
check_slow_path(path, states, ...) SlowPath: Thought of its mode, replay matches, records round-trip through
JSON, cost recomputes, budget respected, decisions replay from a store
check_strategist(strategist, catalog, ...) Strategist: a Flow in an executable order, deterministic, hard checks in
the flow, decisions replay, the plan recorded when it says so
check_head(make, options, rows, answers) Head: probabilities over its options, deterministic, teach changes the
fingerprint, System.fit(head=) answers from it and replays
check_decider(decider, inputs, ...) Decider: Decisions within its options, probabilities, deterministic,
identity recorded, replay
check_extractor(extractor, system, texts) Extractor: literal quotes at their offsets, deterministic, read fields
replay
check_monitor(make, stream, changed=None) Monitor: reports, flagged, a change flagged, reset forgets it
check_environment(make, seeds=(0, 1)) Environment: the same seed and actions give the same outcomes
check_action_model(make, transitions, ...) ActionModel: Predictions, determinism, the fingerprint moves with what it
learned, never contradicts an outcome it observed; held-out comparison
ConformanceError ¶
Bases: AssertionError
An implementation does not keep a guarantee its protocol promises (the message says which).
check_storage ¶
make: a function () → a fresh, empty store (a TraceStorage subclass); reopen: () → the same store opened again (its records read back from where make() wrote them), when the backend persists. → what was checked.
check_slow_path ¶
path: a SlowPath; states: inputs to think about; question: the question (default: the path's, or its System's only one); price: as Dispatcher's; system1: a System 1 for the same question — then every state is also dispatched (supervise=1: the path checks every answer) through a store, and the stored decisions must replay. → what was checked.
check_strategist ¶
strategist: the planner; catalog, questions: a System's; states: inputs (dicts of given facts). The System is
built with System(catalog, questions, strategist=strategist) (which refuses a planner that leaves a hard check's
then out of its question's flow). → what was checked.
check_head ¶
make: options → a fresh, unfitted head; rows: [{fact: value}]; answers: the right option per row; features: the facts it may read (default: the first row's keys). → what was checked.
check_decider ¶
decider: a Decider (called with the facts it reads); inputs: [{fact: value}]; system, states: a System using it and inputs to ask it — its decisions must replay and record the decider's fingerprint. → what was checked.
check_extractor ¶
extractor: an Extractor; system: a System whose entry point (question=, or its only one) reads the fields; texts: messages to read; textin: TextIn's other options (today=, patterns=, synonyms=, ...). → what was checked.
check_monitor ¶
make: () → a fresh monitor; stream: decisions of an unchanged stream (Responses, Results, Decisions or dicts the monitor reads); changed: decisions after a change — when given, the monitor must flag it. → what was checked.
check_environment ¶
make: () → a fresh environment; policy: (state, actions) → the action to take (default: the first one). → what was checked.
check_action_model ¶
make: () → a fresh action model (solvi.core.knowledge.ActionModel); transitions: what an environment did — [(state, action, args, accepted, effect)], observed in order; held_out: more of them, only predicted (the result reports how the predictions compare). Checks: the protocol's methods; every prediction is a Prediction with a verdict in accept / refuse / unknown, a risk ≥ 0 and an integer support, plain JSON in to_dict(); learning is deterministic (two fresh models fed the same transitions give the same predictions and the same fingerprint); what it learned changes its fingerprint; and it never contradicts an outcome it observed (on a (state, action, args) the environment answered one way every time, it predicts that answer or "unknown" — never the other). → what was checked, with the held-out comparison (precision and recall of refusals over answered ones, abstention).