The trace and verification¶
Every step of the flow is a Record in res.trace.records:
| Field | Content |
|---|---|
step, kind, name |
step number, part kind (fn, check, extract, rule), fact name (answer:<question> for rules) |
inputs |
input name -> hash of the input value |
value |
the computed value |
quote |
(start, end, source) for extracted values |
confidence, error |
extraction confidence; error text if the part failed or had missing inputs |
prev, hash |
hash of the previous record (or of init_state for the first) and of this record |
producer, tried |
for a fact with several producers: the one used, and every producer that ran with its outcome |
tried_models |
for a fact with several producers: the model-backed producers that ran and were not used (rejected, or a shadow run) → {"type", "id", "fp"} and the probs they proposed; absent when no model was turned down |
provenance, model, probs |
the value's provenance kind (record.origin gives the default when not stored), the model that produced it ({"type", "id", "fp"}), a decision's probabilities |
Answers from a learned head are records too (kind="head", after the flow's steps): the answer, its probabilities and the
head's fingerprint.
What hashing costs: every given value and every computed value is put in canonical form and hashed once per ask,
however many steps read it; the input's hash is taken when the ask starts. The time grows with the size of the values —
about 0.3 ms per thousand floats of a list (a list of plain floats, strings or ints takes a fast path; the input is
written as JSON once, its values' hashes taken from the same text) — so a decision over a large input is slower than
the "about 0.3 ms" of a small one (README, Speed; both from benchmarks/ask_speed.py). A part should not change a given value in place: the hashes describe the input as it was given.
res.trace.fingerprint records what decided: the catalog's fingerprint, the questions' and the fingerprint of every part
in the flow (see Catalog fingerprint, solvi diff and shadow mode); it is
not part of the hash chain (a stored response is covered by the store's chain).
res.trace.init keeps init_state and res.trace.init_hash its hash; res.values[name] is a computed value.
res.trace.to_json() / Trace.from_json(text, catalog=cat) store and load a trace (typed values are restored, see
Types); the loaded trace replays like the original.
res.trace.replay(system) independently re-executes every step from the recorded init_state, and checks (pass the
System: a bare catalog replays the steps but cannot check the answers, and the result says "answers": "unchecked"):
- that each record still hashes to its stored hash and links to the previous one;
- that each step's inputs match the recorded input hashes;
- that recomputing the part gives the recorded value (and the same quote offsets);
- that quotes lie within the source text (and, for model-backed parts, are literally the quoted text);
- for model-backed steps, that the recorded model fingerprint matches the catalog's current model (see below);
- with the System: that each stored answer and its status are the ones the trace gives — the hard checks, rules, heads
and constraints applied again to the recorded facts. An answer edited after the run is a mismatch of kind
answer.
It returns {"ok": bool, "steps": int, "mismatches": [(step, name, reason), ...], "models": [(step, name, verdict), ...],
"answers": "same" | "differ" | "unchecked", "catalog": "same" | "changed" | "unrecorded"} (with "changed_parts" when the catalog changed since the trace was
recorded — for information: a changed part that still re-computes the recorded value is not a mismatch).
Each mismatch is the triple with a .kind, so a report can tell damaged data from a catalog that moved on:
integrity (the hash chain, a record's hash, the input's hash or a recorded input hash does not verify), recompute
(a step no longer gives the recorded value), model_changed, missing_part (the part or a producer was renamed or
removed since: a mismatch, not an exception, and the steps after it are still checked), missing_input (a part now
reads an input the trace does not hold), flow (a planned step is not recorded), not_restored (a hash or a step does
not verify because a value it rests on did not come back from storage as it was — its type is neither one the dump
restores nor declared: no verdict on the data, see serialization) and error
(replay_all: a stored record could not be loaded, or the replay raised). When there are mismatches the result also has "kinds", the count
per kind, and "summary", one line: "data damaged: the hash chain or a record does not verify", "data intact,
catalog changed (parts missing)", "data intact, catalog changed", "data intact, model changed", "data intact,
steps do not recompute", "not verified: values stored without their type did not come back as they were (no verdict
on the data)" or "replay failed (no verdict on the data)". replay_all gives the same two keys and the
catalog verdict for every stored decision that does not replay, and solvi replay prints them.
Continuing the README quickstart:
rep = res.trace.replay(cat)
assert rep["ok"]
res.trace.records[0].value = 6 # tamper with a recorded value
print(res.trace.replay(cat)["mismatches"][0][:2]) # (1, 'days_requested'): the altered step is named
Replay also catches consistent tampering, where the value is changed and all hashes of the chain are recomputed: the
recomputation from init_state no longer matches, and the altered step is named.
Replay needs the same catalog code. Parts that call external systems (databases, APIs) must return the same values on replay, or their steps will be reported as mismatches.
Storing decisions: TraceStorage¶
from solvi import SQLiteStorage, System
store = SQLiteStorage("decisions.db") # or JSONLStorage("decisions.jsonl"); storage="decisions.db" also works
system = System(cat, questions, storage=store)
res = system.ask(init_state) # saved; res.stored_id is its id
store.get(res.stored_id) # the Response, loaded back (typed values restored)
store.query(question="refund", answer="no", since="2026-09-01")
store.query(safeguard="grounding") # every decision where a model's quote was rejected
store.query(model="solvi-base") # ... this model ran in a step (id, type or fingerprint), used or rejected
store.replay_all(system) # [] when every stored trace replays against the current catalog
store.verify() # the chain across stored records
A stored response loaded without a catalog (store.get(id) on a store opened with no catalog=, or
Stored.response()) still audits and reports: its flow keeps each part's name, kind and inputs and, for a check,
whether it is a hard one — a record stored before that was kept shows such a check as "hard or soft: not recorded".
Replay, diff and counterfactuals need the System.
Two backends ship without dependencies: JSONLStorage (an append-only file, one record per line; several
processes may append on POSIX systems, where each append holds a file lock — on Windows one writing process) and SQLiteStorage (stdlib sqlite3; index tables by question, answer, status, safeguard kind, model and time;
several processes may write to one file). Two more take an optional dependency and keep the same tables:
PostgresStorage("postgresql://user@host/db") (pip install 'solvi[postgres]', psycopg 3; tables named solvi_* —
prefix= to change; several services may write: each append locks the head table for its transaction, so the chain
cannot fork) and DuckDBStorage("decisions.duckdb") (pip install 'solvi[duckdb]'; one writing process; query the tables
with DuckDB next to JSONL or Parquet files). storage="decisions.duckdb" or a postgresql:// URL work too. A stored record holds the answers, the safeguards, the models, the whole
response (res.to_dict()), the time and your own meta (store.save(res, meta={"ticket": 42})). teach stores its
corrections in the same chain (store.corrections()). query(answer=...) matches the stored form of an answer:
answer=True finds a yes/no question's "yes". len(store) counts every chained record (decisions, corrections,
redaction marks); len(list(store.iter())) the decisions. Every store has close() and is a context manager (with
SQLiteStorage("decisions.db") as store:); the file extension is read in any case (decisions.DB is SQLite).
What a stored decision costs and how durable it is depends on the backend. JSONLStorage appends one line and flushes
it to the operating system: the cheapest, but by default not synced to disk, so a power failure can lose the last
records (the chain stays verifiable up to them) — JSONLStorage(path, fsync=True) syncs every record before save
returns, at the cost of one disk sync per decision. SQLiteStorage commits a transaction per record (the record, its
index rows and the new head): durable once save returns, and slower than an unsynced JSON line (a disk sync per commit).
DuckDBStorage and PostgresStorage commit per record too; with PostgreSQL the cost is mostly the round trip to the
server. Measure on your machine before storing every decision of a high-volume stream.
| Method | Returns |
|---|---|
save(res, meta=None) |
the id of the stored record (System(storage=...) calls it on every ask) |
get(id), record(id) |
the stored Response; the stored record as a dict |
iter(), query(question=, answer=, status=, safeguard=, model=, since=, until=) |
Stored records (.id, .time, .answers, .response()) in stored order; since <= time < until |
head() |
{"count", "hash"} of the chain |
verify(anchor=None, signature=None, candidates=None) |
{"ok", "count", "head", "legacy", "problems": [(seq, id, reason)]} (+ "signature") |
signature(alg="syndrome") |
64 bytes that later name the one changed record (see below) |
replay_all(system) |
the stored decisions whose trace no longer replays, with the mismatches |
quarantine(fact, value=...) |
the stored decisions whose answers rest on this fact (with this value), and the path from the fact to each answer |
where_is(fact, value=...) (forget in 0.7) |
a report: decisions resting on a given fact, and records that only hold it; nothing is deleted |
redact(id, by=, note=) |
erase a stored record's content (the response, its input and trace, the meta) and keep the chain: the record keeps its place, hash and id, is marked redacted, and a record of kind redaction naming it is appended; verify() — also with an earlier anchor or signature — still passes, and iter / query / replay pass the record over. What is left (the answers, the time, the safeguards) is still covered by the record's hash, which is taken over the lasting fields and the digests of the answers and of the content: an answer edited in an erased record does not verify. keep_answers=False erases the answers too |
The chain across records. Each record stores the hash of the record before it, and its own hash covers its content and
that link. Editing a stored decision, deleting one, inserting one or changing their order breaks the chain at that point,
and verify() names the record. A line of a JSONL store that is not a record of the chain — unreadable JSON, or a JSON
object without a hash that another tool appended — is reported by verify() as one problem and passed over by iter,
query, report and replay_all; the records after it still verify. Cutting records off the end leaves a shorter chain that is still consistent, so the store
keeps its head (count and last hash) next to the log (decisions.jsonl.head, or a table in SQLite) and verify() checks
it. Someone who can rewrite the whole store and its head can rebuild a consistent chain: publish store.head() somewhere
else from time to time (a ticket, a log you do not control, a signed message) and check with store.verify(anchor=head).
verify() reads the stored head first and checks the records against it, so it can run while the store is written to:
a record appended meanwhile is not reported as damage. verify() needs no catalog; replay_all(system) re-computes every stored step, which also catches a value changed inside
a stored trace with every hash recomputed.
Which record changed: a signature. The chain and the anchor tell that a store was rewritten, not where: an edited
record with every hash after it and the stored head recomputed shows up only as "the record at the anchor differs".
Keep a signature next to the head (preview), and verify names the edited record and restores its content hash:
sig = store.signature() # {"alg": "syndrome", "count", "root": 2 numbers}: 64 bytes of plain JSON
...
v = store.verify(signature=sig, candidates=backup_records)
v["problems"] # [(1, "3f9a…", "the record differs from the signed one (its original …")]
v["signature"] # {"index": 1, "digest": original content hash (hex), "match": the backup record}
solvi verify decisions.db --sign sig.json # write the signature (only when the store verifies)
solvi verify decisions.db --signature sig.json # later: names the changed record and its original content hash
solvi.signature works on anything: sign(res) / res.signature() for one response's trace (position 0 is the input,
position i the record i−1), sign(items) for a list, locate(obj, sig) → the position or None, repair(obj, sig,
candidates=...) → {"index", "digest", "match"} (for a trace record a candidate may be a plain value), extend(sig,
new_items) after appending. Each record is reduced to its content hash h_i (without prev, hash, id, so a recomputed
chain does not move the other records). The default code, alg="syndrome", keeps S0 = Σ h_i and S1 = Σ (i+1)·h_i mod
a 256-bit prime: one change at k by d moves them by d and (k+1)·d, which gives k and the whole original hash.
A signature carries its "alg". Up to 0.7 a second code, alg="octonion" (each hash written into 4 octonions times an
element of its position, multiplied in order — 32 floats), was part of the package; it located exactly as the syndrome
code on a store, larger and slower, and is now an experiment in benchmarks/octonion_signature.py, kept for future
signatures of tree-shaped objects (derivations), where its non-associativity sees a change of brackets that sums cannot.
| Change | Result (stores of 2–500 records, benchmarks/trace_signature.py; the syndrome code and the octonion experiment alike) |
|---|---|
| one record edited, the chain and the head recomputed | located and its content hash restored: 2000 of 2000, 0 wrong |
| two or three records edited | detected 1500 of 1500, located 0 (NotLocatable), never a wrong record |
| two records swapped, one deleted or inserted in the middle | detected, not located |
| records cut off the end / appended after signing | "signed items missing" / not covered: sign again or extend |
| Records | syndrome: sign / locate | octonion experiment (benchmarks/): sign / locate |
|---|---|---|
| 1 000 | 1.3 / 1.4 ms | 13 / 16 ms |
| 10 000 | 13 / 14 ms | 175 / 149 ms |
| 50 000 | 66 / 67 ms | 0.80 / 0.76 s |
| size | 64 bytes | 256 bytes |
The signature restores the record's content hash, not the record: to get the record back, pass candidates (a backup,
a replica). It is an error-locating code, not a MAC — anyone who can rewrite the signature can forge it, so keep it where
you keep the head.
Provenance over the store. store.quarantine("fx_rate", 1.37) lists the stored decisions whose answer depends on that
value of that fact — through the recorded inputs of each step, from the answer back to the fact (a hard check that decided
an answer counts), with the path — so you can re-decide or review them. store.where_is("email", "a@b.c") answers "what
would removing this input touch": the decisions resting on it and the records that merely hold it. Neither changes the
store: deleting a record would break the chain by design.
Existing journals. A 0.5 journal file keeps working: its old lines stay at the start of the file, are skipped by
get / query / iter and counted by verify() as legacy; new lines are chained after them.
Catalog fingerprint, solvi diff and shadow mode¶
system.fingerprint() → {"catalog", "questions", "parts", "models"}. A part's fingerprint covers its declarations (kind,
inputs, hard / then, options, validate, min_confidence, declared types — a pydantic model by its fields, an Enum by
its members — and the type of its model) and its code: the function's syntax tree without decorators, docstring, comments
or formatting, the simple values it closes over or reads as module constants, and the module's own functions it calls. A
model's weights are not in the catalog's fingerprint: they are the model's own fingerprint, recorded with every step it
produced. Every trace records the catalog's and the questions' fingerprints and those of its flow's parts;
store.query(catalog_fp=fp) finds the decisions made by one catalog. Limits: a module constant rebound after the first ask,
or a part edited in place, is not noticed within the process; code the fingerprint does not reach (another module's
functions, a database) changes nothing in it.
solvi diff. "We changed a rule — which decisions change?"
from solvi.diff import diff
rep = diff(store, new_system) # re-runs every stored decision (or diff(store, s, question="refund", since=...))
print(rep) # per question: how many changed and how (yes → no: 12); per decision the
rep.changed # steps that changed the answer and why: the code changed, the model changed,
rep.ok # a new step, or none of these (a non-deterministic or external source)
rep.changed is a list of {"id", "seq", "time", "questions": {question: {"old", "new", "changed", "first_step",
"causes"}}}. causes are the steps that changed that answer, in flow order, each {"step", "name", "old", "new",
"why"}; first_step is the first of them. They are found from the answer step (and a hard check that decided it)
back through the recorded inputs, only through steps whose output differs: a step that gives the same output as
before stops the walk, so a part that was added or edited and changes nothing downstream — a new rule that scores 0 —
is not named, and a decision that changes because of a threshold is attributed to the rule that holds the threshold.
Of the steps on such a path the causes are those where a difference starts (its own code, declarations or model
changed, it is new or no longer runs, or nothing it reads differs); with several changes at once each is listed, and
which of them alone would have changed the answer takes a diff against a system with only that change. The header
lists the parts whose code changed since the decisions were stored and the parts that ran now and not then.
Each stored decision's recorded input is asked again for the same questions (store=False: the new system's own storage
is not written) and compared with the stored response: answer, status, the guard that settled it, the safeguard events
concerning it, and a confidence change above confidence=0.01 (None ignores confidence). The same from the shell:
solvi diff decisions.db --system myapp.decisions_v2:system # exit status 1 when something changes
solvi replay decisions.db --system myapp.decisions:system # every stored trace against the current system
solvi verify decisions.db --anchor 1204:3f9a... # the chain, against a head kept elsewhere
--system is module:attribute or file.py:attribute (a System, or a function returning one); python -m solvi ...
works too.
Shadow mode. Run a new version next to the current one before switching:
from solvi import Shadow, SQLiteStorage
shadow = Shadow(current, candidate, storage=SQLiteStorage("shadow.db"))
res = shadow.ask(state) # the current system's response, exactly as current.ask(state)
print(shadow.summary()) # candidate agrees on 981, differs on 19; refund: 'no' → 'yes' ×12, ...
The candidate runs on the same input after the current system; its response is stored in the shadow store with meta
{"shadow_of": the current response's stored id, "current_catalog", "diff"}, never in its own storage, and a failing
candidate is counted (shadow.stats["errors"]), never raised. The candidate runs in the same thread: it adds its own
time to each ask.