solvi.storage¶
TraceStorage: stored responses and traces, queries and replay; JSONL, SQLite, PostgreSQL and DuckDB backends.
TraceStorage: stored responses with their traces, a hash chain across them, queries, a replay of everything stored, and provenance questions over the store (which stored decisions rest on a fact found to be wrong; what a forgotten given fact touches).
Every stored decision is one record (a dict of plain JSON):
kind "ask" the 0.5 journal line's keys — init_hash, answers {question: [answer, confidence, status]}, flow, records [[step, name, hash]] (and producers) — plus index fields (guards, safeguards, models) and the whole response (Response.to_dict(): answers, flow, trace), which get() loads back and replay_all() re-checks; kind "teach" a correction (System.teach): {"teach": question, "init": ..., "answer": ...} and, when given, its "source" ("outcome", "rule"; none: a human), "by" and "of" (the stored id of the decision it corrects); kind "update" a learning update (System.learning): what changed, the gates' results, the state to roll back to; every record seq (0, 1, 2, ...), time (seconds since the epoch), meta (optional, yours), prev and hash.
A record is written with every dict's keys in their own order (a decider reads a dict's keys in that order, so a
stored trace gives its dicts back as they were). hash is the SHA-256 of the record's canonical JSON (keys sorted,
without id and hash) — it does not depend on the order the record was written in — which includes prev, the
hash of the record before it: editing, deleting, inserting or reordering a stored record breaks the chain at that point
(verify()). Cutting records off the end leaves a shorter chain that is still valid, so the store keeps its head — the
count and the last hash (head()) — next to the log and verify() checks it; a head you published elsewhere
(verify(anchor=head)) also catches a rewrite of the whole chain together with the stored head.
Backends: JSONLStorage (append-only file, one record per line; what System(storage="file.jsonl") writes), SQLiteStorage (stdlib sqlite3; indexed by question, answer, status, safeguard, model fingerprint and time; several processes may write), PostgresStorage (psycopg 3; the same tables, several services writing) and DuckDBStorage (duckdb; the same tables in a DuckDB file, for analytics).
UntrustedLabel ¶
Bases: ValueError
A label from a source outside TRUSTED_SOURCES (the model, the system itself, an unknown process).
Stored
dataclass
¶
One stored record: its id, position, time (seconds since the epoch), kind ("ask": a decision; "teach": a correction; "redaction": the mark of an erasure, see TraceStorage.redact) and the record itself.
TraceStorage ¶
A store of responses and their traces with a hash chain across the stored records (see the module docstring).
save(response, meta=None) → id; get(id) → Response; record(id) → the stored dict; iter() / query(...) → [Stored]; corrections() → the teach records; head() → {"count", "hash"}; signature(); verify(anchor=None, signature=None); replay_all(system); quarantine(fact, value=...); where_is(fact, value=...).
system: the System (or a Catalog) used to restore typed values (dates, enums, models) when loading responses;
System(storage=...) sets it to that system when it is not set. clock: a function → seconds since the epoch.
head ¶
{"count", "hash"}: the number of chained records and the last record's hash (GENESIS when empty). Publish it somewhere else (a ticket, a log, a signed message) to later catch a rewrite of the whole store: verify(anchor=...).
save ¶
Store a response (its answers, flow and whole trace) → its id. meta: your own JSON data kept with it (a
ticket id, a user) — hashed into the chain like the rest.
save_correction ¶
Store a correction (what System.teach records) → its id. label_source (source= in 0.7; stored as the
record's "source"): where the label comes from — "human" (a person
corrected or confirmed the answer), "outcome" (what really happened: the parcel was lost, the loan defaulted) or
"rule" (code rejected a model's proposal and decided instead); anything else is refused (UntrustedLabel): the
system's own answers are never labels. by: who (a user, a reviewer, a process); of: the stored id of the decision
it corrects. Records of 0.6 have no source: they are human corrections.
get ¶
The stored response with this id, loaded back with system (default: the store's; see Stored.response).
iter ¶
Stored records in order (kind "ask": responses; "teach": corrections; None: all) → iterator of Stored. A redacted record (see redact) has no content left and is passed over; redacted=True: it is yielded too.
__len__ ¶
The number of chained records — decisions, corrections and redaction marks (head()["count"]); the decisions alone: len(list(store.iter())).
close ¶
Release what the store holds open (a database connection; nothing for a JSON-lines file). A store is also a
context manager: with SQLiteStorage("decisions.db") as store: ... closes it at the end.
corrections ¶
The stored corrections → [{"id", "time", "question", "init", "answer", "source", "by", "of"}] (feed them to fit / learn_rule, a CorrectionMemory or System.learning). source: "human" (also every record without one), "outcome", "rule" — or, for a record written around save_correction, whatever it says (solvi.memory and System.learning refuse anything outside TRUSTED_SOURCES).
query ¶
query(question=None, answer=ANY, status=None, safeguard=None, model=None, since=None, until=None, catalog_fp=None)
Stored responses that match every filter given → [Stored] in stored order. question: asked this question; answer: answered this (with question: that question's answer; None matches an abstention); status: "ok" / "forced" / "abstain" (with question: of that question); safeguard: a safeguard of this kind fired ("grounding", "hard_check", "low_confidence", ...; with question: one that concerns it); model: a model with this fingerprint, id or type produced a step; since / until: stored in [since, until) — seconds since the epoch, a datetime, a date or an ISO string; catalog_fp: decided by the catalog with this fingerprint (System.fingerprint()["catalog"]). An answer is matched in its stored form: answer=True finds a yes/no "yes" (see _answer_key).
report ¶
A human-readable report of the stored decisions in [since, until) (optionally of one question): counts by
answer, status and safeguard, the escalation rate, the guarantee coverage of the answers a model took part in,
changes of the catalog's and the models' fingerprints, and examples stored ids per answer, escalation and
safeguard. format: "md", "html" (one self-contained page) or "data" (a dict); other filters as query. See
solvi.report.
redact ¶
Erase a stored record's content — a person's data that must go — and keep the chain whole.
The record keeps its place, its time, its hash and its id, so every link after it still verifies; its content
(the response with its trace and input, the meta; a correction's input and answer) is removed and it is marked
redacted with who and why and the digest of what was removed. A record of kind "redaction" is appended that
names the erased record and its hash: the erasure is itself in the chain, with its time. keep_answers=False
removes the answers too (their digest stays in the mark). The record no longer replays and is passed over by
iter / query / replay_all / reports; record(id) returns what is left. → the redaction record's id.
What is left stays verified: a record's hash is taken over its lasting fields, the digest of its answers and the digest of its content (record_body), so verify() recomputes the hash of a redacted record like any other — an answer edited in it afterwards, or a record passed off as redacted with other answers, does not verify — and a signature taken before the erasure still holds. A record written by solvi ≤ 0.7.1 (format 1) has one flat hash: redacting it leaves its kept fields unverifiable, which verify() lists under "unverified". What it cannot do: copies made before (a backup, an exported report, a published anchor's holder) are not touched, and derived state (a correction memory, a fitted head) keeps what it learned — rebuild those.
signature ¶
The signature of the chained records (solvi.signature.sign): {"alg", "count", "root"} — with the default "syndrome" code two numbers (64 bytes). Keep it where you keep the head: verify(signature=...) then names the one record that changed — even when every hash after it and the stored head were recomputed — and restores its content hash.
verify ¶
Check the chain across stored records: each record's hash, its link to the record before it, the sequence
numbers, the stored response against its own summary, and the stored head (a cut-off tail). anchor: a head()
taken earlier and kept elsewhere — the chain must still contain it (catches a rewrite of the whole store).
signature: a signature() taken earlier and kept elsewhere — the records it covers must be the ones signed; when one
changed, its seq is named and signature in the result holds solvi.signature.check's answer (the original content
hash; candidates: records, e.g. from a backup, one of which may be the original → its "match").
→ {"ok", "count", "head", "legacy", "problems": [(seq, id, reason)]} (+ "signature" when given; + "unverified":
[(seq, id, reason)] — redacted records of format 1, whose kept fields no hash covers). Records written before the
chain (0.5 journal lines) are counted in legacy and not checked.
A store that is being written to verifies as it stands at one moment: the stored head is read first and the
records are checked against it, so a record appended while verify runs is not reported as damage.
replay_all ¶
Replay every stored trace (or those matching query filters) against system (a System or a Catalog): each
step is re-computed from its recorded inputs (see Trace.replay). → the ones that fail: [{"id", "seq", "time",
"mismatches": [(step, name, reason)], "models": [(step, name, verdict)], "catalog", "kinds", "summary"}] — empty
when all replay. Each mismatch has a .kind, and "summary" tells damaged data from a catalog or a model that
changed since (see solvi.runtime.Mismatch). A stored record that cannot be loaded is one mismatch (0, "load", ...),
a replay that raises is (0, "replay", ...): both of kind "error", no verdict on the data. "note" (a record of
format 1 whose model step does not recompute): it was stored with sorted dict keys, see LEGACY_ORDER.
quarantine ¶
A fact (or a given input) found to be wrong: the stored decisions whose answers rest on it — through the trace's
provenance graph (each step's recorded inputs, from the answer back to the fact; a hard check that decided an
answer counts). value: only where the fact had this value (compared by its hash as the consumers recorded it).
→ [{"id", "seq", "time", "questions": {question: {"answer", "status", "path": [fact, ..., "answer:question"]}}}].
Nothing is changed: re-decide or review these.
forget ¶
Deprecated (removed in 0.9): where_is(fact, value) — it never deleted anything.
where_is ¶
Where a given fact (e.g. a person's data) is held, and what removing it would touch (forget in 0.7) a given fact (e.g. a person's data) would touch — a report only: nothing is deleted (deleting a
stored record breaks the chain by design; keep the report as the record of the request, and erase the records it
lists with redact).
→ {"fact", "value", "dependent": the stored decisions whose answers rest on it (as quarantine), "stored": ids of the
stored records that hold it without an answer resting on it (responses and corrections), "deleted": 0}.
JSONLStorage ¶
Bases: TraceStorage
Append-only JSON lines, one record per line (the file System(storage="file.jsonl") writes). The head (count and last hash)
is kept in <path>.head. Threads of a process may write; several processes may too where the system has advisory
file locks (POSIX: an append takes an exclusive flock on the file, reads what other processes appended since it
last looked, then writes its record and the head) — on Windows keep to one writing process, or use SQLiteStorage.
Lines of a 0.5 journal at the start of the file are kept and skipped (verify reports them as legacy).
fsync=True: flush every record to disk before save returns (slower; without it a power failure can lose the last
records, which the OS had not yet written). index=False: open from the stored head (checked
against the file's last line) without reading every record — a long file opens at once, for a process that only
appends; get(id) then scans the file. A head that does not match the last line is not trusted: the file is read.
SQLiteStorage ¶
Bases: _SQLStorage
SQLite (stdlib sqlite3): one row per record with the record's JSON, plus index tables — answers (question, answer,
status), safeguards (kind, question), models (fingerprint, id, type) — and the time. The head is kept in the meta
table. Appends run in a write transaction, so several processes may write to one file. Each record is one committed
transaction: durable once save returns (unlike JSONLStorage without fsync=True), at the cost of a disk sync.
PostgresStorage ¶
Bases: _SQLStorage
PostgreSQL (psycopg 3, pip install solvi[postgres]): the tables of SQLiteStorage, named with prefix
("solvi_records", ...), created when missing. Several processes and services may write: an append locks the head
table (LOCK TABLE ... IN SHARE ROW EXCLUSIVE MODE — readers are not blocked) for its transaction, so the chain has
no forks. conninfo: a connection string ("postgresql://user@host/db") or an open psycopg connection in autocommit
mode (the store runs its own BEGIN / COMMIT).
DuckDBStorage ¶
Bases: _SQLStorage
DuckDB (pip install solvi[duckdb]): the tables of SQLiteStorage in a DuckDB file (":memory:" for none) — for
analytics over the stored decisions next to JSONL or Parquet files (store.db.sql(...)). One writing process at a
time (DuckDB's own rule); threads of that process are fine.
check_source ¶
A label's source → itself, when it is trusted: "human", "outcome" or "rule"; else UntrustedLabel.
record_body ¶
What a record's hash is taken over. Format 1 (solvi ≤ 0.7.1): the record without id and hash. Format 2: the
fields that stay for ever (KEPT), the digest of its answers and the digest of everything else (its content: the
response, the input, the meta) — three parts, so that redact can remove the content, or the answers too, leave their
digests in the record's mark, and the hash still recomputes: what is left of a redacted record is verified like any
other record, and a record cannot be passed off as redacted with other answers.
record_hash ¶
The hash of a stored record: SHA-256 of the canonical JSON of its body (record_body; it covers prev).
plain ¶
A value as stored: JSON data (sets sorted as vhash sorts them, dates as ISO strings, "not stated" as "
chained ¶
Is this a record of the chain (a dict with its hash)? A JSON object without one — a line another tool appended to the file, a hand edit — is not: it is passed over by iter / query / replay_all, and verify reports it.
open_storage ¶
A TraceStorage from a path: .db / .sqlite / .sqlite3 → SQLiteStorage, .duckdb → DuckDBStorage (in any case), a postgresql:// (or postgres://) URL → PostgresStorage, anything else → JSONLStorage; a TraceStorage is returned as it is.