solvi.diff¶
solvi diff and shadow mode.
solvi diff and shadow mode: which decisions change with a new catalog or model, and why.
diff(storage, system) re-runs stored decisions (a TraceStorage) with another System — a changed rule, a new model — and reports, per decision, the questions whose answer, status, safeguard or confidence changes, with the steps that changed the answer — those on a path of differing outputs from the answer back, where a difference starts — and why each differs (the part's code or declarations changed, its model changed, it is new, or none of these: a non-deterministic or external source). A step that was added or changed and gives nothing different downstream is not named.
Shadow(current, candidate) answers with the current system and runs the candidate on the same inputs, storing the candidate's response with its differences; the answer returned is always the current system's.
DiffReport
dataclass
¶
Shadow ¶
Answer with current, run candidate on the same input, keep what differs.
shadow = Shadow(current, candidate, storage=SQLiteStorage("shadow.db"))
res = shadow.ask(state) # current's response, exactly as current.ask(state) (saved to current's storage)
The candidate's response is saved to storage (when given) with meta {"shadow_of": the current response's stored id,
"current_catalog", "diff": compare(current, candidate)}; the candidate never writes to its own storage. A candidate that
fails is counted and recorded, never raised. stats: {"asks", "agree", "differ", "errors"}; changed: the last keep
differing inputs [{"stored_id", "shadow_id", "questions"}]; report() as text.
causes ¶
The steps that changed question's answer between two responses to the same input, in flow order → [{"step",
"name", "old", "new", "why"}]. From the answer step (and a hard check that decided it) back through the recorded
inputs, only through steps whose output differs: a step that gives the same output as before stops the walk — what
changed above it did not reach the answer — so a step that was added (or changed) and changes nothing downstream
is not a cause. Of the steps on such a path, the causes are those where a difference starts: its own code,
declarations or model changed, it is new or no longer runs, or none of what it reads differs. [] when the answer
step itself gives the same output (see first_difference for what is reported then).
compare ¶
The differences between two responses to the same input → {question: {"old", "new", "changed", "first_step",
"causes"}} for the questions whose answer, status, safeguard (the guard that settled it, or the safeguard events
concerning it) differ — or whose confidence moved by more than confidence (None: ignore confidence). "causes": the
steps that changed the answer (see causes), "first_step" the first of them in flow order (see first_difference).
Questions asked in one response only are listed with the other side None.
diff ¶
Re-run stored decisions with system (e.g. a new catalog or model) and report what changes → DiffReport.
Each stored decision is loaded (typed values restored with system), its recorded input is asked again for the same
questions (store=False: nothing is written to the system's own storage), and the two responses are compared (see
compare). A stored decision whose input did not come back as it was (see solvi.schema: an untyped enum or object, an
untyped date stored by solvi ≤ 0.7.1) is listed under errors — "could not be re-run" — not as a changed one. filters: TraceStorage.query filters (question=, since=, ...) to pick the decisions; limit: at most this many.
The system's learned parts may learn from these asks as from any other (System(learn=...)).