Skip to content

solvi.testing.conformance

Conformance checks for the extension points of solvi.core (see Building blocks): each runs your implementation through the parts of solvi that rely on it and fails with the guarantee that does not hold.

Conformance checks for the extension points of solvi.core: what the docs say you get for free, tested on your class.

from solvi.testing.conformance import check_storage, check_slow_path

def test_my_store(tmp_path):
    check_storage(lambda: MyStorage(tmp_path / "decisions"), reopen=lambda: MyStorage(tmp_path / "decisions"))

Each check runs the implementation through the parts of solvi that rely on it and raises ConformanceError (an AssertionError, so pytest shows it as a failed assertion) naming the guarantee that does not hold; it returns a dict of what it checked. solvi's own test suite runs every check on every built-in (tests/test_conformance.py).

check_storage(make, reopen=None)            TraceStorage: chain verifies, records round-trip, replay, query, redact,
                                            tampering is caught, a reopened store is the same store
check_slow_path(path, states, ...)          SlowPath: Thought of its mode, replay matches, records round-trip through
                                            JSON, cost recomputes, budget respected, decisions replay from a store
check_strategist(strategist, catalog, ...)  Strategist: a Flow in an executable order, deterministic, hard checks in
                                            the flow, decisions replay, the plan recorded when it says so
check_head(make, options, rows, answers)    Head: probabilities over its options, deterministic, teach changes the
                                            fingerprint, System.fit(head=) answers from it and replays
check_decider(decider, inputs, ...)         Decider: Decisions within its options, probabilities, deterministic,
                                            identity recorded, replay
check_extractor(extractor, system, texts)   Extractor: literal quotes at their offsets, deterministic, read fields
                                            replay
check_monitor(make, stream, changed=None)   Monitor: reports, flagged, a change flagged, reset forgets it
check_environment(make, seeds=(0, 1))       Environment: the same seed and actions give the same outcomes
check_action_model(make, transitions, ...)  ActionModel: Predictions, determinism, the fingerprint moves with what it
                                            learned, never contradicts an outcome it observed; held-out comparison

ConformanceError

Bases: AssertionError

An implementation does not keep a guarantee its protocol promises (the message says which).

check_storage

check_storage(make, *, reopen=None)

make: a function () → a fresh, empty store (a TraceStorage subclass); reopen: () → the same store opened again (its records read back from where make() wrote them), when the backend persists. → what was checked.

check_slow_path

check_slow_path(path, states, *, question=None, price=None, system1=None)

path: a SlowPath; states: inputs to think about; question: the question (default: the path's, or its System's only one); price: as Dispatcher's; system1: a System 1 for the same question — then every state is also dispatched (supervise=1: the path checks every answer) through a store, and the stored decisions must replay. → what was checked.

check_strategist

check_strategist(strategist, catalog, questions, states)

strategist: the planner; catalog, questions: a System's; states: inputs (dicts of given facts). The System is built with System(catalog, questions, strategist=strategist) (which refuses a planner that leaves a hard check's then out of its question's flow). → what was checked.

check_head

check_head(make, options, rows, answers, features=None)

make: options → a fresh, unfitted head; rows: [{fact: value}]; answers: the right option per row; features: the facts it may read (default: the first row's keys). → what was checked.

check_decider

check_decider(decider, inputs, *, system=None, states=None)

decider: a Decider (called with the facts it reads); inputs: [{fact: value}]; system, states: a System using it and inputs to ask it — its decisions must replay and record the decider's fingerprint. → what was checked.

check_extractor

check_extractor(extractor, system, texts, *, question=None, **textin)

extractor: an Extractor; system: a System whose entry point (question=, or its only one) reads the fields; texts: messages to read; textin: TextIn's other options (today=, patterns=, synonyms=, ...). → what was checked.

check_monitor

check_monitor(make, stream, *, changed=None)

make: () → a fresh monitor; stream: decisions of an unchanged stream (Responses, Results, Decisions or dicts the monitor reads); changed: decisions after a change — when given, the monitor must flag it. → what was checked.

check_environment

check_environment(make, *, seeds=(0, 1), steps=30, policy=None)

make: () → a fresh environment; policy: (state, actions) → the action to take (default: the first one). → what was checked.

check_action_model

check_action_model(make, transitions, *, held_out=None)

make: () → a fresh action model (solvi.core.knowledge.ActionModel); transitions: what an environment did — [(state, action, args, accepted, effect)], observed in order; held_out: more of them, only predicted (the result reports how the predictions compare). Checks: the protocol's methods; every prediction is a Prediction with a verdict in accept / refuse / unknown, a risk ≥ 0 and an integer support, plain JSON in to_dict(); learning is deterministic (two fresh models fed the same transitions give the same predictions and the same fingerprint); what it learned changes its fingerprint; and it never contradicts an outcome it observed (on a (state, action, args) the environment answered one way every time, it predicts that answer or "unknown" — never the other). → what was checked, with the held-out comparison (precision and recall of refusals over answered ones, abstention).