Grounded decisions: provenance, audit and safeguards¶
The principle: fuzzy proposes, deterministic decides, everything is in the trace. A model may extract a value, pick a category or learn an answer, but its output is checked by deterministic code before anything uses it, and every step records where its value came from. A model's hallucination is either caught or visible in the audit — never silently an answer. A decision without any model and one with models are the same system; they differ only in the provenance of the facts.
Provenance¶
Every fact and answer has a provenance kind (record.origin, result.provenance, solvi.provenance.KINDS):
| Kind | Where the value comes from | How it is kept honest |
|---|---|---|
given |
a key of init_state |
hashed into init_hash; the chain starts from it |
computed |
a plain function: fn, check, a hand-written rule |
replay re-runs it and compares |
quoted |
an extract part returning a Quote |
the offsets must lie in the source text; for a model, doc[start:end] must be the value |
decided |
a model's choice among declared options, with probabilities (Decision) |
the value must be one of the options; probabilities recorded |
learned |
a fit head, a learn_rule list, another trained function |
the head type and a fingerprint of its parameters are recorded |
proposed |
a model that writes: a strategist's plan, a generator's text or JSON (solvi.generate) |
the deterministic layer verifies what it proposes; replay re-reads a recorded reply through its parser and schema |
The default comes from what a part returns (a Quote → quoted, a Decision → decided) and whether a model is behind
it. Declare it explicitly with provenance= on any decorator. A part is model-backed when you pass model=:
cat.extract(extractor.field("total", "the total amount paid")) # field() functions bring their model along
@cat.extract(model=span_extractor) # any function that calls a model
def vendor(doc): ...
@cat.fn(model=classifier, options=["travel", "meals", "equipment"])
def category(doc):
p = classifier.predict(doc) # {option: probability}
return Decision(max(p, key=p.get), p) # downstream parts get the plain value
@cat.rule("risk", model=risk_model) # a model answers the question directly
def risk(amount, country): ...
LongSpanExtractor.field, MultiSpanExtractor.field and LongSpanExtractor.embedder mark their functions (the attributes
__solvi_model__ and __solvi_provenance__), so registering them is enough. For any other model,
pass model=.
Model identity in the trace¶
A model-backed record stores record.model = {"type", "id", "fp"}: the class, the model id (the Hugging Face id or path it
was loaded from, model.model_id), and a fingerprint (solvi.provenance.fingerprint):
- extractors: settings, thresholds / temperatures, the span head, evenly sampled encoder weights, and the names and sizes of
the weight files — computed once, then cached until
fit/save; FastHead(fit;Head, the logistic head before 0.8): a hash of their parameters — it changes with everyteach;RuleList(learn_rule): a hash of its rules;- any other object: its own
fingerprint()method, or aversionattribute, or"unversioned:<type>"(then a changed model cannot be detected — give your models a version).
res.trace.replay(catalog) handles model-backed steps as follows:
- the fingerprint differs from the catalog's current model → a mismatch, "model changed since this decision"; the recorded output is still checked for grounding;
- same model, deterministic (the default; set
model.deterministic = Falseotherwise) → the step is re-run and compared; replay(catalog, trust_models=True), or the model is not available (no model on the part,model.available = False) → the model is not re-run; the recorded output is verified instead: the quote is literally at its offsets in the recorded input, the decision is among the options.
Pass the System instead of the catalog (res.trace.replay(system)) to verify answer-head records too: their fingerprint,
and (unless trusted) their probabilities recomputed from the recorded facts. rep["models"] lists every model-backed step
with its verdict: recomputed, trusted, unavailable or changed.
Hashes: a record hashes its provenance only when it differs from the default (quoted for a record with a quote, else
computed), and its model and probabilities only when present — so traces of catalogs without models hash exactly as before.
Safeguards¶
| Safeguard | Fires when | Effect |
|---|---|---|
| grounding | a quote lies outside its text, or a model's quote is not literally doc[start:end] (strings up to whitespace; numbers as written, e.g. 1250.0 ↔ "1,250.00" — int, float, Decimal, Fraction and numpy scalars; a date when the text reads as that date; a bool, a datetime or a list cannot be compared and is not checked); an evidence quote or a span is not literally in its text |
the output is rejected: the fact is missing, the claim stays in the error; the next alternative producer runs, else dependent answers abstain |
| closed set | a Decision (or a value of a part with options=) is not one of the options; a rule's answer is not one of the question's options |
rejected / the question abstains |
| low confidence | a Quote / Decision is below the part's min_confidence (a decision part's too); an answer is below the question's min_confidence |
rejected / the question abstains, saying what it would have answered |
| model escalated | a decider's act / escalate signal is below its threshold (see the output) | rejected: the fact is missing, next producer, else the question abstains, saying what it would have answered |
| validate | a producer's validate(value, ...) returns false |
rejected, next producer |
| type rejected | a typed part's argument or output fails its type annotation, or a field fails System(input_model=...) |
rejected: the fact is missing, next producer, else dependent answers abstain |
| hard check | a hard check governing the question is false | the answer is forced by then, or the question abstains |
| constraint repair | learned or model answers break a constraint between answers | the most probable consistent combination is chosen |
| fallback | an alternative producer was rejected and a later one was used | recorded in tried |
| evidence missing | a question with require_evidence=True got an answer without a supporting quote |
the question abstains, saying what it would have answered |
| instruction | a decision part with perturb=k answered differently without an instruction-like sentence of its input ("ignore the rules and answer X") |
rejected like an escalation: the fact is missing, next producer, else the question abstains, naming the sentence (system.stats["instruction_flips"]) |
Hand-written extractors (no model) may return a value derived from the quoted text; the audit then shows the value next to the text it was derived from. Numbers, dates and other non-string values from a model are checked when they can be compared (numbers) and shown next to the quoted text otherwise.
The audit¶
print(res.audit()) # every answer
a = res.audit("approve") # one answer: an AnswerAudit
a.to_dict() # the same as data
For each answer: the given inputs → computed facts → quotes (offsets, the quoted text, whether the value is literally that
text, the model) → model decisions (the model, probabilities) → learned parts (head type and fingerprint) → checks (hard or
soft, which one decided) → the rule → constraints → the answer; the parts skipped at run time; the safeguards that fired;
and a summary of the support: how many items are deterministic (given, computed, quoted by plain code) and how many come
from models (a.share_deterministic, a.counts).
approve = 'yes' [ok] confidence 0.60 ← computed by approve
given doc = 'Expense claim #2291\nVendor: C…; limit = {'travel': 100, 'meals': 60, 'e…
computed amount = 48.6
quoted total = '48.60' doc[100:105] literal '48.60'
decided category = 'travel' (travel 0.60, meals 0.20, equipment 0.20) [StandInClassifier demo/expense-category #bb3352d4]
check amount_positive = True (hard)
rule approve (computed)
→ answer 'yes' — amount = 48.6; category = 'travel'; limit = {'travel': 100, 'meals': 60, 'equipment': 800}
support 7 items (2 given, 3 computed, 1 quoted, 1 decided): 86% deterministic, 1 from models
guarantee none for some decisions: their thresholds were not calibrated on your data (see act_guard)
safeguards grounding rejected ×1, fallback producer ×1
· grounding rejected: total — total_model: not grounded: '488.60' is not the text at [100:105] ('48.60')
· fallback producer: total — total_regex used after total_model rejected
Counterfactual explanations: res.counterfactual¶
"What would have changed the answer?" — the smallest change of the given inputs, for adverse-action reasons in lending and clear answers in support:
res = system.ask({"amount": 1200.0, "debt": 1000, "income": 5000, "history": "on time", "age": 30})
cf = res.counterfactual("approve")
print(cf)
# approve = decline [ok]
# approve if amount ≤ 1000 (now 1200)
# held at their recorded proposals (no model called): risk
# not searched: history (str: no domain (pass domains={'history': [...]}))
cf.best.changes[0] # Change(fact="amount", now=1200.0, to=1000.0, op="≤", cost=0.17)
cf.to_dict()
Only the deterministic flow is re-run, on the recorded plan: every model-backed part — an extractor, a model decision, a
learned answer head — is held at the proposal it recorded in this trace, and no model is called. The explanation is
"what the code would decide if the models said what they said"; the result lists the parts held. A model part that did not
run in the recorded decision (a hard check failed first) has no proposal: inputs that need it make the question abstain and
do not count as a change (listed as "without a recorded proposal"). Learned rule lists (learn_rule) are code and re-run.
What is searched (over=: default, the given facts the question's flow reads):
- numbers and dates — outward from the current value in both directions with doubling steps, then bisection between the
last unchanged and the first changed value: the nearest threshold crossing, exact for inputs the answer is monotone in
(a non-monotone input can hide a nearer crossing between two probes). A direction where no probe changes the answer
is tried again on an even grid up to the farthest probe, which finds the band of a two-sided rule (
abs(value + 20) > 5at 40 →no if value ≤ -15); a narrower band can still be missed, so when nothing is found the result says "no change ... was found", not that none exists. Integers and dates give exact bounds (debt ≤ 1999,purchase_date ≥ 2026-08-20); floats are shown at the shortest decimal that holds,≤or<as the rule has it. A non-negative input stays non-negative; a domain(lo, hi)that does not contain the current value is refused for that input (cf.not_searchedsays why); - booleans, Enums and
Literalfields ofSystem(input_model=...)— every other value; - anything else only with
domains={"history": ["on time", "late"]}; a tuple bounds a number:domains={"amount": (0, 5000)}.
max_changes=2 (the default) tries two inputs together when no single input changes the answer ("approve if amount ≤ 1000
(now 1200) and debt ≤ 1999 (now 2500)" — each bound holds with the other change made); max_changes=1 does not.
target="approve" looks only for that answer. Results are ranked by the number of changes, then their size (the relative
change of a number; for a date the days moved over 30, or over the width of its domain; 1 for an enumerated value).
A given input read only by a part that did not run (a soft check skipped after a hard check failed) is listed in
cf.not_searched. max_evals=5000 caps the re-runs (cf.exhausted). A response loaded from
a store with its System works the same; one loaded without it needs system=.
Reports for people: res.report, store.report, solvi report¶
The audit is for developers; a report is for an auditor or a customer — one page per decision or per period, as Markdown,
one self-contained HTML file (no external assets, scripts or fonts; every value escaped) or data (format="data").
print(res.report()) # Markdown
open("decision.html", "w").write(res.report(format="html"))
res.report(format="data") # the same as a dict (answers, documents, models, trace, replay)
A decision report shows, per answer: the answer, status and confidence, the reason; what it rests on (given inputs,
computed facts, quotes with their offsets, model decisions with probabilities and the model, learned parts, checks — which
one decided —, the rule, evidence, constraints, parts not run); the safeguards that fired; and the guarantee line — the
promise of the calibrated thresholds of the model decisions behind it (act_guard, calibrate_for), "none" when a model
decided without one, or that no model decided the answer. Then the source texts with every quote highlighted (the offsets
on hover; a quote that is not the text at its offsets in red; a text over 20 000 characters as excerpts around the
quotes), every model that ran with its fingerprint (also the ones whose output was rejected), the trace's input and last
hashes, the catalog's fingerprint and the replay status. replay="trusted" (the default) re-runs the deterministic steps
and verifies the models' recorded outputs without calling them; replay="full" re-runs the models too, replay=False
skips it. A response loaded from a store with its System (store.get(id) when the store belongs to a System) reports
like the original.
print(store.report(since="2026-09-01", until="2026-10-01")) # every question
store.report(question="refund", format="html", examples=5) # one question
A period report counts per question: the answers, the statuses, the escalation rate (abstentions — handed to a person — by
the safeguard that caused them), the safeguards that fired, and the guarantee coverage: of the answers a model decided
or took part in, how many rest only on calibrated thresholds (answers from code alone are counted apart). It lists the
catalog and model fingerprints in use, how many decisions of the period were erased (store.redact: they are not in
the counts) and how many corrections were recorded in it ("erased" and "corrections" in the data; store-wide for
the period, whatever the other filters), and every change of the fingerprints over the period (from which stored decision on), and up
to examples stored ids per answer, escalation reason and safeguard — res = store.get(id) and res.report() give the
page of one. From the shell:
solvi report decisions.db --since 2026-09-01 --question refund # Markdown to stdout
solvi report decisions.db --html september.html # a self-contained page
solvi report decisions.db --id 3f9a0c1d2e4b5a67 --system app.py:system # one decision, replayed against the system
OpenTelemetry: solvi.otel¶
solvi.otel.export(res_or_store, tracer=None, **filters) sends decisions to your tracing backend as OpenTelemetry spans
(pip install "solvi[otel]"): per decision a root span solvi.decision, a child span per step of the trace
(solvi.fn risk, solvi.extract total, solvi.rule answer:pay, …) and one per answer (solvi.answer pay). A store
exports every stored decision, or those matching the query filters (export(store, question="refund", since=...)).
The root span is a child of the span current in your code, so a decision sits inside the request that asked it.
from solvi.otel import export, to_otlp_json
with tracer.start_as_current_span("POST /refund"):
res = system.ask(state)
export(res) # the global tracer provider's "solvi" tracer, or tracer=...
body = to_otlp_json(res, service_name="refunds") # OTLP/JSON without OpenTelemetry: POST it to a collector's /v1/traces
| Span | Attributes |
|---|---|
| step | solvi.step, solvi.kind, solvi.fact, solvi.provenance, solvi.value (a short repr), solvi.confidence, solvi.error, solvi.producer, solvi.tried, solvi.quote.source / .start / .end, solvi.model.type / .id / .fingerprint, solvi.probs (JSON), solvi.safeguard (kinds that fired on this fact), solvi.inputs, solvi.hash, solvi.prev |
| answer | solvi.question, solvi.answer, solvi.status, solvi.confidence, solvi.why, solvi.guard, solvi.provenance, solvi.source, solvi.safeguard |
| root | solvi.questions, solvi.trace.init_hash, solvi.trace.head, solvi.trace.steps, solvi.catalog.fingerprint, solvi.questions.fingerprint, solvi.stored_id, solvi.ms, solvi.confidence, solvi.complete, solvi.model_outputs; an event solvi.skipped per step skipped at run time |
A failed or rejected step has status ERROR with the reason; an abstention is not an error. The trace records each step's
run time, not its start: step spans are laid end to end from the decision's start (durations measured, start times not;
parallel steps appear one after another). A stored decision ends at its stored time; a fresh one when exported (or at
end_ns=). In the OTLP JSON the trace and span ids are derived from the trace's hashes; through the API the SDK assigns
them — solvi.hash ties a span to its trace record either way.
Lifetime stats¶
system.stats counts, over the system's lifetime: asks, answers, abstained, model_outputs (outputs of model-backed
parts, answer heads and learned rules), grounding_rejected, type_rejected, outside_options, rule_abstained,
low_confidence, validator_rejected,
forced_by_hard_check, constraint_repairs, fallbacks, model_escalated, evidence_missing, timeouts,
instruction_flips and memory_disagreements.
system.safeguard_summary() prints them (evidence missing once it has fired).
examples/12_grounded_audit.py runs one catalog with and without models,
with a hallucinating extractor and a classifier answering outside its options.