Skip to content

Grounded decisions: provenance, audit and safeguards

The principle: fuzzy proposes, deterministic decides, everything is in the trace. A model may extract a value, pick a category or learn an answer, but its output is checked by deterministic code before anything uses it, and every step records where its value came from. A model's hallucination is either caught or visible in the audit — never silently an answer. A decision without any model and one with models are the same system; they differ only in the provenance of the facts.

Provenance

Every fact and answer has a provenance kind (record.origin, result.provenance, solvi.provenance.KINDS):

Kind Where the value comes from How it is kept honest
given a key of init_state hashed into init_hash; the chain starts from it
computed a plain function: fn, check, a hand-written rule replay re-runs it and compares
quoted an extract part returning a Quote the offsets must lie in the source text; for a model, doc[start:end] must be the value
decided a model's choice among declared options, with probabilities (Decision) the value must be one of the options; probabilities recorded
learned a fit head, a learn_rule list, another trained function the head type and a fingerprint of its parameters are recorded
proposed a model that writes: a strategist's plan, a generator's text or JSON (solvi.generate) the deterministic layer verifies what it proposes; replay re-reads a recorded reply through its parser and schema

The default comes from what a part returns (a Quote → quoted, a Decision → decided) and whether a model is behind it. Declare it explicitly with provenance= on any decorator. A part is model-backed when you pass model=:

cat.extract(extractor.field("total", "the total amount paid"))    # field() functions bring their model along

@cat.extract(model=span_extractor)                                  # any function that calls a model
def vendor(doc): ...

@cat.fn(model=classifier, options=["travel", "meals", "equipment"])
def category(doc):
    p = classifier.predict(doc)                                    # {option: probability}
    return Decision(max(p, key=p.get), p)                           # downstream parts get the plain value

@cat.rule("risk", model=risk_model)                                 # a model answers the question directly
def risk(amount, country): ...

LongSpanExtractor.field, MultiSpanExtractor.field and LongSpanExtractor.embedder mark their functions (the attributes __solvi_model__ and __solvi_provenance__), so registering them is enough. For any other model, pass model=.

Model identity in the trace

A model-backed record stores record.model = {"type", "id", "fp"}: the class, the model id (the Hugging Face id or path it was loaded from, model.model_id), and a fingerprint (solvi.provenance.fingerprint):

  • extractors: settings, thresholds / temperatures, the span head, evenly sampled encoder weights, and the names and sizes of the weight files — computed once, then cached until fit / save;
  • FastHead (fit; Head, the logistic head before 0.8): a hash of their parameters — it changes with every teach;
  • RuleList (learn_rule): a hash of its rules;
  • any other object: its own fingerprint() method, or a version attribute, or "unversioned:<type>" (then a changed model cannot be detected — give your models a version).

res.trace.replay(catalog) handles model-backed steps as follows:

  • the fingerprint differs from the catalog's current model → a mismatch, "model changed since this decision"; the recorded output is still checked for grounding;
  • same model, deterministic (the default; set model.deterministic = False otherwise) → the step is re-run and compared;
  • replay(catalog, trust_models=True), or the model is not available (no model on the part, model.available = False) → the model is not re-run; the recorded output is verified instead: the quote is literally at its offsets in the recorded input, the decision is among the options.

Pass the System instead of the catalog (res.trace.replay(system)) to verify answer-head records too: their fingerprint, and (unless trusted) their probabilities recomputed from the recorded facts. rep["models"] lists every model-backed step with its verdict: recomputed, trusted, unavailable or changed.

Hashes: a record hashes its provenance only when it differs from the default (quoted for a record with a quote, else computed), and its model and probabilities only when present — so traces of catalogs without models hash exactly as before.

Safeguards

Safeguard Fires when Effect
grounding a quote lies outside its text, or a model's quote is not literally doc[start:end] (strings up to whitespace; numbers as written, e.g. 1250.0 ↔ "1,250.00" — int, float, Decimal, Fraction and numpy scalars; a date when the text reads as that date; a bool, a datetime or a list cannot be compared and is not checked); an evidence quote or a span is not literally in its text the output is rejected: the fact is missing, the claim stays in the error; the next alternative producer runs, else dependent answers abstain
closed set a Decision (or a value of a part with options=) is not one of the options; a rule's answer is not one of the question's options rejected / the question abstains
low confidence a Quote / Decision is below the part's min_confidence (a decision part's too); an answer is below the question's min_confidence rejected / the question abstains, saying what it would have answered
model escalated a decider's act / escalate signal is below its threshold (see the output) rejected: the fact is missing, next producer, else the question abstains, saying what it would have answered
validate a producer's validate(value, ...) returns false rejected, next producer
type rejected a typed part's argument or output fails its type annotation, or a field fails System(input_model=...) rejected: the fact is missing, next producer, else dependent answers abstain
hard check a hard check governing the question is false the answer is forced by then, or the question abstains
constraint repair learned or model answers break a constraint between answers the most probable consistent combination is chosen
fallback an alternative producer was rejected and a later one was used recorded in tried
evidence missing a question with require_evidence=True got an answer without a supporting quote the question abstains, saying what it would have answered
instruction a decision part with perturb=k answered differently without an instruction-like sentence of its input ("ignore the rules and answer X") rejected like an escalation: the fact is missing, next producer, else the question abstains, naming the sentence (system.stats["instruction_flips"])

Hand-written extractors (no model) may return a value derived from the quoted text; the audit then shows the value next to the text it was derived from. Numbers, dates and other non-string values from a model are checked when they can be compared (numbers) and shown next to the quoted text otherwise.

The audit

print(res.audit())            # every answer
a = res.audit("approve")      # one answer: an AnswerAudit
a.to_dict()                   # the same as data

For each answer: the given inputs → computed facts → quotes (offsets, the quoted text, whether the value is literally that text, the model) → model decisions (the model, probabilities) → learned parts (head type and fingerprint) → checks (hard or soft, which one decided) → the rule → constraints → the answer; the parts skipped at run time; the safeguards that fired; and a summary of the support: how many items are deterministic (given, computed, quoted by plain code) and how many come from models (a.share_deterministic, a.counts).

approve = 'yes'  [ok]  confidence 0.60  ← computed by approve
  given       doc = 'Expense claim #2291\nVendor: C…; limit = {'travel': 100, 'meals': 60, 'e…
  computed    amount = 48.6
  quoted      total = '48.60'  doc[100:105] literal '48.60'
  decided     category = 'travel'  (travel 0.60, meals 0.20, equipment 0.20)  [StandInClassifier demo/expense-category #bb3352d4]
  check       amount_positive = True (hard)
  rule        approve (computed)
  → answer    'yes' — amount = 48.6; category = 'travel'; limit = {'travel': 100, 'meals': 60, 'equipment': 800}
  support     7 items (2 given, 3 computed, 1 quoted, 1 decided): 86% deterministic, 1 from models
  guarantee   none for some decisions: their thresholds were not calibrated on your data (see act_guard)
  safeguards  grounding rejected ×1, fallback producer ×1
              · grounding rejected: total — total_model: not grounded: '488.60' is not the text at [100:105] ('48.60')
              · fallback producer: total — total_regex used after total_model rejected

Counterfactual explanations: res.counterfactual

"What would have changed the answer?" — the smallest change of the given inputs, for adverse-action reasons in lending and clear answers in support:

res = system.ask({"amount": 1200.0, "debt": 1000, "income": 5000, "history": "on time", "age": 30})
cf = res.counterfactual("approve")
print(cf)
# approve = decline [ok]
#   approve if amount ≤ 1000 (now 1200)
#   held at their recorded proposals (no model called): risk
#   not searched: history (str: no domain (pass domains={'history': [...]}))
cf.best.changes[0]            # Change(fact="amount", now=1200.0, to=1000.0, op="≤", cost=0.17)
cf.to_dict()

Only the deterministic flow is re-run, on the recorded plan: every model-backed part — an extractor, a model decision, a learned answer head — is held at the proposal it recorded in this trace, and no model is called. The explanation is "what the code would decide if the models said what they said"; the result lists the parts held. A model part that did not run in the recorded decision (a hard check failed first) has no proposal: inputs that need it make the question abstain and do not count as a change (listed as "without a recorded proposal"). Learned rule lists (learn_rule) are code and re-run.

What is searched (over=: default, the given facts the question's flow reads):

  • numbers and dates — outward from the current value in both directions with doubling steps, then bisection between the last unchanged and the first changed value: the nearest threshold crossing, exact for inputs the answer is monotone in (a non-monotone input can hide a nearer crossing between two probes). A direction where no probe changes the answer is tried again on an even grid up to the farthest probe, which finds the band of a two-sided rule (abs(value + 20) > 5 at 40 → no if value ≤ -15); a narrower band can still be missed, so when nothing is found the result says "no change ... was found", not that none exists. Integers and dates give exact bounds (debt ≤ 1999, purchase_date ≥ 2026-08-20); floats are shown at the shortest decimal that holds, ≤ or < as the rule has it. A non-negative input stays non-negative; a domain (lo, hi) that does not contain the current value is refused for that input (cf.not_searched says why);
  • booleans, Enums and Literal fields of System(input_model=...) — every other value;
  • anything else only with domains={"history": ["on time", "late"]}; a tuple bounds a number: domains={"amount": (0, 5000)}.

max_changes=2 (the default) tries two inputs together when no single input changes the answer ("approve if amount ≤ 1000 (now 1200) and debt ≤ 1999 (now 2500)" — each bound holds with the other change made); max_changes=1 does not. target="approve" looks only for that answer. Results are ranked by the number of changes, then their size (the relative change of a number; for a date the days moved over 30, or over the width of its domain; 1 for an enumerated value). A given input read only by a part that did not run (a soft check skipped after a hard check failed) is listed in cf.not_searched. max_evals=5000 caps the re-runs (cf.exhausted). A response loaded from a store with its System works the same; one loaded without it needs system=.

Reports for people: res.report, store.report, solvi report

The audit is for developers; a report is for an auditor or a customer — one page per decision or per period, as Markdown, one self-contained HTML file (no external assets, scripts or fonts; every value escaped) or data (format="data").

print(res.report())                          # Markdown
open("decision.html", "w").write(res.report(format="html"))
res.report(format="data")                    # the same as a dict (answers, documents, models, trace, replay)

A decision report shows, per answer: the answer, status and confidence, the reason; what it rests on (given inputs, computed facts, quotes with their offsets, model decisions with probabilities and the model, learned parts, checks — which one decided —, the rule, evidence, constraints, parts not run); the safeguards that fired; and the guarantee line — the promise of the calibrated thresholds of the model decisions behind it (act_guard, calibrate_for), "none" when a model decided without one, or that no model decided the answer. Then the source texts with every quote highlighted (the offsets on hover; a quote that is not the text at its offsets in red; a text over 20 000 characters as excerpts around the quotes), every model that ran with its fingerprint (also the ones whose output was rejected), the trace's input and last hashes, the catalog's fingerprint and the replay status. replay="trusted" (the default) re-runs the deterministic steps and verifies the models' recorded outputs without calling them; replay="full" re-runs the models too, replay=False skips it. A response loaded from a store with its System (store.get(id) when the store belongs to a System) reports like the original.

print(store.report(since="2026-09-01", until="2026-10-01"))            # every question
store.report(question="refund", format="html", examples=5)               # one question

A period report counts per question: the answers, the statuses, the escalation rate (abstentions — handed to a person — by the safeguard that caused them), the safeguards that fired, and the guarantee coverage: of the answers a model decided or took part in, how many rest only on calibrated thresholds (answers from code alone are counted apart). It lists the catalog and model fingerprints in use, how many decisions of the period were erased (store.redact: they are not in the counts) and how many corrections were recorded in it ("erased" and "corrections" in the data; store-wide for the period, whatever the other filters), and every change of the fingerprints over the period (from which stored decision on), and up to examples stored ids per answer, escalation reason and safeguard — res = store.get(id) and res.report() give the page of one. From the shell:

solvi report decisions.db --since 2026-09-01 --question refund          # Markdown to stdout
solvi report decisions.db --html september.html                          # a self-contained page
solvi report decisions.db --id 3f9a0c1d2e4b5a67 --system app.py:system   # one decision, replayed against the system

OpenTelemetry: solvi.otel

solvi.otel.export(res_or_store, tracer=None, **filters) sends decisions to your tracing backend as OpenTelemetry spans (pip install "solvi[otel]"): per decision a root span solvi.decision, a child span per step of the trace (solvi.fn risk, solvi.extract total, solvi.rule answer:pay, …) and one per answer (solvi.answer pay). A store exports every stored decision, or those matching the query filters (export(store, question="refund", since=...)). The root span is a child of the span current in your code, so a decision sits inside the request that asked it.

from solvi.otel import export, to_otlp_json

with tracer.start_as_current_span("POST /refund"):
    res = system.ask(state)
    export(res)                               # the global tracer provider's "solvi" tracer, or tracer=...

body = to_otlp_json(res, service_name="refunds")   # OTLP/JSON without OpenTelemetry: POST it to a collector's /v1/traces
Span Attributes
step solvi.step, solvi.kind, solvi.fact, solvi.provenance, solvi.value (a short repr), solvi.confidence, solvi.error, solvi.producer, solvi.tried, solvi.quote.source / .start / .end, solvi.model.type / .id / .fingerprint, solvi.probs (JSON), solvi.safeguard (kinds that fired on this fact), solvi.inputs, solvi.hash, solvi.prev
answer solvi.question, solvi.answer, solvi.status, solvi.confidence, solvi.why, solvi.guard, solvi.provenance, solvi.source, solvi.safeguard
root solvi.questions, solvi.trace.init_hash, solvi.trace.head, solvi.trace.steps, solvi.catalog.fingerprint, solvi.questions.fingerprint, solvi.stored_id, solvi.ms, solvi.confidence, solvi.complete, solvi.model_outputs; an event solvi.skipped per step skipped at run time

A failed or rejected step has status ERROR with the reason; an abstention is not an error. The trace records each step's run time, not its start: step spans are laid end to end from the decision's start (durations measured, start times not; parallel steps appear one after another). A stored decision ends at its stored time; a fresh one when exported (or at end_ns=). In the OTLP JSON the trace and span ids are derived from the trace's hashes; through the API the SDK assigns them — solvi.hash ties a span to its trace record either way.

Lifetime stats

system.stats counts, over the system's lifetime: asks, answers, abstained, model_outputs (outputs of model-backed parts, answer heads and learned rules), grounding_rejected, type_rejected, outside_options, rule_abstained, low_confidence, validator_rejected, forced_by_hard_check, constraint_repairs, fallbacks, model_escalated, evidence_missing, timeouts, instruction_flips and memory_disagreements. system.safeguard_summary() prints them (evidence missing once it has fired). examples/12_grounded_audit.py runs one catalog with and without models, with a hallucinating extractor and a classifier answering outside its options.