Skip to content

The trace and verification

Every step of the flow is a Record in res.trace.records:

Field Content
step, kind, name step number, part kind (fn, check, extract, rule), fact name (answer:<question> for rules)
inputs input name -> hash of the input value
value the computed value
quote (start, end, source) for extracted values
confidence, error extraction confidence; error text if the part failed or had missing inputs
prev, hash hash of the previous record (or of init_state for the first) and of this record
producer, tried for a fact with several producers: the one used, and every producer that ran with its outcome
tried_models for a fact with several producers: the model-backed producers that ran and were not used (rejected, or a shadow run) → {"type", "id", "fp"} and the probs they proposed; absent when no model was turned down
provenance, model, probs the value's provenance kind (record.origin gives the default when not stored), the model that produced it ({"type", "id", "fp"}), a decision's probabilities

Answers from a learned head are records too (kind="head", after the flow's steps): the answer, its probabilities and the head's fingerprint.

What hashing costs: every given value and every computed value is put in canonical form and hashed once per ask, however many steps read it; the input's hash is taken when the ask starts. The time grows with the size of the values — about 0.3 ms per thousand floats of a list (a list of plain floats, strings or ints takes a fast path; the input is written as JSON once, its values' hashes taken from the same text) — so a decision over a large input is slower than the "about 0.3 ms" of a small one (README, Speed; both from benchmarks/ask_speed.py). A part should not change a given value in place: the hashes describe the input as it was given.

res.trace.fingerprint records what decided: the catalog's fingerprint, the questions' and the fingerprint of every part in the flow (see Catalog fingerprint, solvi diff and shadow mode); it is not part of the hash chain (a stored response is covered by the store's chain).

res.trace.init keeps init_state and res.trace.init_hash its hash; res.values[name] is a computed value. res.trace.to_json() / Trace.from_json(text, catalog=cat) store and load a trace (typed values are restored, see Types); the loaded trace replays like the original.

res.trace.replay(system) independently re-executes every step from the recorded init_state, and checks (pass the System: a bare catalog replays the steps but cannot check the answers, and the result says "answers": "unchecked"):

  • that each record still hashes to its stored hash and links to the previous one;
  • that each step's inputs match the recorded input hashes;
  • that recomputing the part gives the recorded value (and the same quote offsets);
  • that quotes lie within the source text (and, for model-backed parts, are literally the quoted text);
  • for model-backed steps, that the recorded model fingerprint matches the catalog's current model (see below);
  • with the System: that each stored answer and its status are the ones the trace gives — the hard checks, rules, heads and constraints applied again to the recorded facts. An answer edited after the run is a mismatch of kind answer.

It returns {"ok": bool, "steps": int, "mismatches": [(step, name, reason), ...], "models": [(step, name, verdict), ...], "answers": "same" | "differ" | "unchecked", "catalog": "same" | "changed" | "unrecorded"} (with "changed_parts" when the catalog changed since the trace was recorded — for information: a changed part that still re-computes the recorded value is not a mismatch).

Each mismatch is the triple with a .kind, so a report can tell damaged data from a catalog that moved on: integrity (the hash chain, a record's hash, the input's hash or a recorded input hash does not verify), recompute (a step no longer gives the recorded value), model_changed, missing_part (the part or a producer was renamed or removed since: a mismatch, not an exception, and the steps after it are still checked), missing_input (a part now reads an input the trace does not hold), flow (a planned step is not recorded), not_restored (a hash or a step does not verify because a value it rests on did not come back from storage as it was — its type is neither one the dump restores nor declared: no verdict on the data, see serialization) and error (replay_all: a stored record could not be loaded, or the replay raised). When there are mismatches the result also has "kinds", the count per kind, and "summary", one line: "data damaged: the hash chain or a record does not verify", "data intact, catalog changed (parts missing)", "data intact, catalog changed", "data intact, model changed", "data intact, steps do not recompute", "not verified: values stored without their type did not come back as they were (no verdict on the data)" or "replay failed (no verdict on the data)". replay_all gives the same two keys and the catalog verdict for every stored decision that does not replay, and solvi replay prints them.

Continuing the README quickstart:

rep = res.trace.replay(cat)
assert rep["ok"]

res.trace.records[0].value = 6               # tamper with a recorded value
print(res.trace.replay(cat)["mismatches"][0][:2])   # (1, 'days_requested'): the altered step is named

Replay also catches consistent tampering, where the value is changed and all hashes of the chain are recomputed: the recomputation from init_state no longer matches, and the altered step is named.

Replay needs the same catalog code. Parts that call external systems (databases, APIs) must return the same values on replay, or their steps will be reported as mismatches.

Storing decisions: TraceStorage

from solvi import SQLiteStorage, System

store = SQLiteStorage("decisions.db")          # or JSONLStorage("decisions.jsonl"); storage="decisions.db" also works
system = System(cat, questions, storage=store)
res = system.ask(init_state)                   # saved; res.stored_id is its id

store.get(res.stored_id)                        # the Response, loaded back (typed values restored)
store.query(question="refund", answer="no", since="2026-09-01")
store.query(safeguard="grounding")              # every decision where a model's quote was rejected
store.query(model="solvi-base")                # ... this model ran in a step (id, type or fingerprint), used or rejected
store.replay_all(system)                        # [] when every stored trace replays against the current catalog
store.verify()                                  # the chain across stored records

A stored response loaded without a catalog (store.get(id) on a store opened with no catalog=, or Stored.response()) still audits and reports: its flow keeps each part's name, kind and inputs and, for a check, whether it is a hard one — a record stored before that was kept shows such a check as "hard or soft: not recorded". Replay, diff and counterfactuals need the System.

Two backends ship without dependencies: JSONLStorage (an append-only file, one record per line; several processes may append on POSIX systems, where each append holds a file lock — on Windows one writing process) and SQLiteStorage (stdlib sqlite3; index tables by question, answer, status, safeguard kind, model and time; several processes may write to one file). Two more take an optional dependency and keep the same tables: PostgresStorage("postgresql://user@host/db") (pip install 'solvi[postgres]', psycopg 3; tables named solvi_* — prefix= to change; several services may write: each append locks the head table for its transaction, so the chain cannot fork) and DuckDBStorage("decisions.duckdb") (pip install 'solvi[duckdb]'; one writing process; query the tables with DuckDB next to JSONL or Parquet files). storage="decisions.duckdb" or a postgresql:// URL work too. A stored record holds the answers, the safeguards, the models, the whole response (res.to_dict()), the time and your own meta (store.save(res, meta={"ticket": 42})). teach stores its corrections in the same chain (store.corrections()). query(answer=...) matches the stored form of an answer: answer=True finds a yes/no question's "yes". len(store) counts every chained record (decisions, corrections, redaction marks); len(list(store.iter())) the decisions. Every store has close() and is a context manager (with SQLiteStorage("decisions.db") as store:); the file extension is read in any case (decisions.DB is SQLite).

What a stored decision costs and how durable it is depends on the backend. JSONLStorage appends one line and flushes it to the operating system: the cheapest, but by default not synced to disk, so a power failure can lose the last records (the chain stays verifiable up to them) — JSONLStorage(path, fsync=True) syncs every record before save returns, at the cost of one disk sync per decision. SQLiteStorage commits a transaction per record (the record, its index rows and the new head): durable once save returns, and slower than an unsynced JSON line (a disk sync per commit). DuckDBStorage and PostgresStorage commit per record too; with PostgreSQL the cost is mostly the round trip to the server. Measure on your machine before storing every decision of a high-volume stream.

Method Returns
save(res, meta=None) the id of the stored record (System(storage=...) calls it on every ask)
get(id), record(id) the stored Response; the stored record as a dict
iter(), query(question=, answer=, status=, safeguard=, model=, since=, until=) Stored records (.id, .time, .answers, .response()) in stored order; since <= time < until
head() {"count", "hash"} of the chain
verify(anchor=None, signature=None, candidates=None) {"ok", "count", "head", "legacy", "problems": [(seq, id, reason)]} (+ "signature")
signature(alg="syndrome") 64 bytes that later name the one changed record (see below)
replay_all(system) the stored decisions whose trace no longer replays, with the mismatches
quarantine(fact, value=...) the stored decisions whose answers rest on this fact (with this value), and the path from the fact to each answer
where_is(fact, value=...) (forget in 0.7) a report: decisions resting on a given fact, and records that only hold it; nothing is deleted
redact(id, by=, note=) erase a stored record's content (the response, its input and trace, the meta) and keep the chain: the record keeps its place, hash and id, is marked redacted, and a record of kind redaction naming it is appended; verify() — also with an earlier anchor or signature — still passes, and iter / query / replay pass the record over. What is left (the answers, the time, the safeguards) is still covered by the record's hash, which is taken over the lasting fields and the digests of the answers and of the content: an answer edited in an erased record does not verify. keep_answers=False erases the answers too

The chain across records. Each record stores the hash of the record before it, and its own hash covers its content and that link. Editing a stored decision, deleting one, inserting one or changing their order breaks the chain at that point, and verify() names the record. A line of a JSONL store that is not a record of the chain — unreadable JSON, or a JSON object without a hash that another tool appended — is reported by verify() as one problem and passed over by iter, query, report and replay_all; the records after it still verify. Cutting records off the end leaves a shorter chain that is still consistent, so the store keeps its head (count and last hash) next to the log (decisions.jsonl.head, or a table in SQLite) and verify() checks it. Someone who can rewrite the whole store and its head can rebuild a consistent chain: publish store.head() somewhere else from time to time (a ticket, a log you do not control, a signed message) and check with store.verify(anchor=head). verify() reads the stored head first and checks the records against it, so it can run while the store is written to: a record appended meanwhile is not reported as damage. verify() needs no catalog; replay_all(system) re-computes every stored step, which also catches a value changed inside a stored trace with every hash recomputed.

Which record changed: a signature. The chain and the anchor tell that a store was rewritten, not where: an edited record with every hash after it and the stored head recomputed shows up only as "the record at the anchor differs". Keep a signature next to the head (preview), and verify names the edited record and restores its content hash:

sig = store.signature()                  # {"alg": "syndrome", "count", "root": 2 numbers}: 64 bytes of plain JSON
...
v = store.verify(signature=sig, candidates=backup_records)
v["problems"]                            # [(1, "3f9a…", "the record differs from the signed one (its original …")]
v["signature"]                           # {"index": 1, "digest": original content hash (hex), "match": the backup record}
solvi verify decisions.db --sign sig.json          # write the signature (only when the store verifies)
solvi verify decisions.db --signature sig.json     # later: names the changed record and its original content hash

solvi.signature works on anything: sign(res) / res.signature() for one response's trace (position 0 is the input, position i the record i−1), sign(items) for a list, locate(obj, sig) → the position or None, repair(obj, sig, candidates=...) → {"index", "digest", "match"} (for a trace record a candidate may be a plain value), extend(sig, new_items) after appending. Each record is reduced to its content hash h_i (without prev, hash, id, so a recomputed chain does not move the other records). The default code, alg="syndrome", keeps S0 = Σ h_i and S1 = Σ (i+1)·h_i mod a 256-bit prime: one change at k by d moves them by d and (k+1)·d, which gives k and the whole original hash.

A signature carries its "alg". Up to 0.7 a second code, alg="octonion" (each hash written into 4 octonions times an element of its position, multiplied in order — 32 floats), was part of the package; it located exactly as the syndrome code on a store, larger and slower, and is now an experiment in benchmarks/octonion_signature.py, kept for future signatures of tree-shaped objects (derivations), where its non-associativity sees a change of brackets that sums cannot.

Change Result (stores of 2–500 records, benchmarks/trace_signature.py; the syndrome code and the octonion experiment alike)
one record edited, the chain and the head recomputed located and its content hash restored: 2000 of 2000, 0 wrong
two or three records edited detected 1500 of 1500, located 0 (NotLocatable), never a wrong record
two records swapped, one deleted or inserted in the middle detected, not located
records cut off the end / appended after signing "signed items missing" / not covered: sign again or extend
Records syndrome: sign / locate octonion experiment (benchmarks/): sign / locate
1 000 1.3 / 1.4 ms 13 / 16 ms
10 000 13 / 14 ms 175 / 149 ms
50 000 66 / 67 ms 0.80 / 0.76 s
size 64 bytes 256 bytes

The signature restores the record's content hash, not the record: to get the record back, pass candidates (a backup, a replica). It is an error-locating code, not a MAC — anyone who can rewrite the signature can forge it, so keep it where you keep the head.

Provenance over the store. store.quarantine("fx_rate", 1.37) lists the stored decisions whose answer depends on that value of that fact — through the recorded inputs of each step, from the answer back to the fact (a hard check that decided an answer counts), with the path — so you can re-decide or review them. store.where_is("email", "a@b.c") answers "what would removing this input touch": the decisions resting on it and the records that merely hold it. Neither changes the store: deleting a record would break the chain by design.

Existing journals. A 0.5 journal file keeps working: its old lines stay at the start of the file, are skipped by get / query / iter and counted by verify() as legacy; new lines are chained after them.

Catalog fingerprint, solvi diff and shadow mode

system.fingerprint() → {"catalog", "questions", "parts", "models"}. A part's fingerprint covers its declarations (kind, inputs, hard / then, options, validate, min_confidence, declared types — a pydantic model by its fields, an Enum by its members — and the type of its model) and its code: the function's syntax tree without decorators, docstring, comments or formatting, the simple values it closes over or reads as module constants, and the module's own functions it calls. A model's weights are not in the catalog's fingerprint: they are the model's own fingerprint, recorded with every step it produced. Every trace records the catalog's and the questions' fingerprints and those of its flow's parts; store.query(catalog_fp=fp) finds the decisions made by one catalog. Limits: a module constant rebound after the first ask, or a part edited in place, is not noticed within the process; code the fingerprint does not reach (another module's functions, a database) changes nothing in it.

solvi diff. "We changed a rule — which decisions change?"

from solvi.diff import diff

rep = diff(store, new_system)              # re-runs every stored decision (or diff(store, s, question="refund", since=...))
print(rep)                                 # per question: how many changed and how (yes → no: 12); per decision the
rep.changed                                # steps that changed the answer and why: the code changed, the model changed,
rep.ok                                     # a new step, or none of these (a non-deterministic or external source)

rep.changed is a list of {"id", "seq", "time", "questions": {question: {"old", "new", "changed", "first_step", "causes"}}}. causes are the steps that changed that answer, in flow order, each {"step", "name", "old", "new", "why"}; first_step is the first of them. They are found from the answer step (and a hard check that decided it) back through the recorded inputs, only through steps whose output differs: a step that gives the same output as before stops the walk, so a part that was added or edited and changes nothing downstream — a new rule that scores 0 — is not named, and a decision that changes because of a threshold is attributed to the rule that holds the threshold. Of the steps on such a path the causes are those where a difference starts (its own code, declarations or model changed, it is new or no longer runs, or nothing it reads differs); with several changes at once each is listed, and which of them alone would have changed the answer takes a diff against a system with only that change. The header lists the parts whose code changed since the decisions were stored and the parts that ran now and not then.

Each stored decision's recorded input is asked again for the same questions (store=False: the new system's own storage is not written) and compared with the stored response: answer, status, the guard that settled it, the safeguard events concerning it, and a confidence change above confidence=0.01 (None ignores confidence). The same from the shell:

solvi diff decisions.db --system myapp.decisions_v2:system        # exit status 1 when something changes
solvi replay decisions.db --system myapp.decisions:system         # every stored trace against the current system
solvi verify decisions.db --anchor 1204:3f9a...                  # the chain, against a head kept elsewhere

--system is module:attribute or file.py:attribute (a System, or a function returning one); python -m solvi ... works too.

Shadow mode. Run a new version next to the current one before switching:

from solvi import Shadow, SQLiteStorage

shadow = Shadow(current, candidate, storage=SQLiteStorage("shadow.db"))
res = shadow.ask(state)                    # the current system's response, exactly as current.ask(state)
print(shadow.summary())                     # candidate agrees on 981, differs on 19; refund: 'no' → 'yes' ×12, ...

The candidate runs on the same input after the current system; its response is stored in the shadow store with meta {"shadow_of": the current response's stored id, "current_catalog", "diff"}, never in its own storage, and a failing candidate is counted (shadow.stats["errors"]), never raised. The candidate runs in the same thread: it adds its own time to each ask.