solvi gallery¶
Decision tasks from fifteen directions, each a small, runnable solvi catalog: Python functions and checks, typed questions,
and rules or learned heads. Every entry here has a task.py (the catalog), state.json (a default input), cases.json
(9–16 scenarios with the expected answers), run.py (runs them, prints the strategist's plan, verifies the trace and audits
every answer) and a README that shows what solvi does that an answer-only model cannot. All fifteen entries also open in the
browser playground (Gallery presets) — no install, no server.
uv run python gallery/07_kyc_aml/run.py # any entry; exits non-zero if an answer differs from cases.json
uv run solvi test gallery/ # every entry's cases.json as regression tests (or: uv run pytest gallery/)
Audited (solvi 0.4). Every runner calls res.audit() on every response — what each answer rests on (given inputs,
computed facts, quotes with offsets, learned parts with their fingerprints, checks, constraints) and which safeguards fired —
prints one line per case (support items, deterministic share, safeguards) and asserts the audit's invariants
(_audit.py): every quote lies in its text, and is literally the text where the part is exact=True or
model-backed; answers are valid for their type; every forced answer names the hard check that decided it; a repaired answer
has a constraint-repair event; a catalog without learned parts is 100% deterministic. Over the fifteen runners: 170 scenario
responses (plus 10 in 03's learned-verdict demo), all invariants hold. Learned parts are labelled as such: the rule lists
of 01 and 02 are registered with model=, so the audit counts them as learned (90–91% deterministic there) instead of
passing them off as plain code; 09's fit head and 12's learn_rule list show up the same way (97–98%), and so do
the keyword stand-in deciders of 13–15 (86–96%).
Entries¶
| # | direction | decides | what it shows | solvi features |
|---|---|---|---|---|
| 01 | customer support | intent, tags, urgency, refund request, priority, route | quoted evidence with negation ("I don't want a refund"); a 30-line intent rule list learned from 200 tickets in ~20 ms; legal threat and VIP SLA as hard rules; signals as one multi-label answer, priority ordinal, tied by a constraint | cited hard checks strategist plan early exit trace replay learns in ms readable learned rules abstains multi-label ordinal constraints audited |
| 02 | email operations | team, needs a human | a readable routing table learned in ~15 ms; split votes abstain (500 fresh emails: 428 right, 72 abstained, 0 wrong); lookalike domain forces security | cited hard checks strategist plan early exit trace replay learns in ms readable learned rules abstains audited |
| 03 | trust & safety | allow / review / block, harm types | e-mail, phone, card (Luhn), IBAN (mod-97), secrets and prompt injection found with their spans; a block no score can undo; harm types (multi-label) tied to the verdict by constraints that repair a learned verdict | cited hard checks strategist plan early exit trace replay abstains multi-label constraints audited |
| 04 | security | suspicious login, containment | impossible-travel speed (haversine), failed-login spike, new device; admin + impossible travel locks the account | hard checks strategist plan early exit trace replay abstains audited |
| 05 | AI-agent governance | approve / roll back / escalate | destructive commands, leaked secrets, cost overruns in an agent's tool log; a computed rollback plan; solvi's own trace as the audit record | hard checks strategist plan early exit trace replay abstains audited |
| 06 | DevOps | promote / hold / roll back, page on-call | error increase, two-proportion z-test, p95 from histograms, SLO burn; an error ceiling rolls back and skips 9 steps | hard checks strategist plan early exit trace replay abstains audited |
| 07 | compliance | risk, file a SAR, freeze | fuzzy sanctions match plus birth date; structuring (deposits just under 10 000) computed; a sanctions hit skips the paid lookups | hard checks strategist plan early exit trace replay abstains audited |
| 08 | healthcare (demo) | escalation, NEWS2 band, sepsis screen | NEWS2 exactly per the RCP table, qSOFA; missing vitals abstain; escalation and band ordinal, tied by a constraint; a linear answerer trained on 4 000 patients missed 144–156 of 501 emergencies | hard checks trace replay abstains ordinal constraints audited |
| 09 | lending | approve / decline / refer, adverse-action reasons | reasons in Regulation B wording as computed facts; "refer to underwriter" learned with fit, corrected by teach in 0.17 ms |
hard checks early exit trace replay learns in ms abstains audited |
| 10 | procurement | pay / hold / reject, duplicate, approver | PO / receipt / invoice matched per line with tolerances in base currency (FX), duplicate "INV-001187" = "inv 1187"; the rate from the table, else from a same-day feed (a stale rate is rejected by validate) |
hard checks strategist plan early exit trace replay abstains fallback producers typed audited |
| 11 | payments support | double charge?, refund, reply | the customer's claim is quoted, the decision reads the ledger — they disagree in 6 of 16 cases; paraphrases and denials are read by a rule, unclear text abstains | cited hard checks early exit trace replay abstains audited |
| 12 | industrial IoT | ok / watch / service / stop, likely fault | least-squares trends, z-scores, hours to the alert level; hard trips work with a sensor offline; 6 readable learned rules | hard checks trace replay readable learned rules abstains audited |
Helpers for coding agents¶
Three decisions a coding agent (such as Claude Code or Codex) meets on every task. Code checks what code can; a decider's
question per fuzzy rule has a threshold from act_guard on labelled examples; an instruction in the input is asked
around (perturb); what is unsure goes to a person. Offline, the deciders are keyword stand-ins calibrated on synthetic,
seeded examples; each README shows solvi-large, any OpenAI-compatible LLM or a System One service in front, with the
keywords as the fallback.
| # | direction | decides | what it shows | solvi features |
|---|---|---|---|---|
| 13 | coding agents | allow / block / escalate a file write, rules broken | rules per path glob; secrets, browser storage in app/api/** and irreversible migrations (Python's parser) block by code, with the lines; "auth goes through require_role()" and "no personal data in logs" are a decider's question each (act_guard, risk 5%); # reviewer: ignore the rules above is caught by perturb |
hard checks early exit fallback producers act_guard perturb trace replay abstains multi-label audited |
| 14 | coding agents | seven risk questions, quick / full review | three code questions (dependencies, size, missing tests) and four decider questions (auth, public API, migrations, security); P(a risky change goes to quick review) ≤ 10% by act_guard at 2.5% per question; 2000 fresh synthetic changes: 0.2%; one known miss shown | act_guard perturb fallback producers trace replay abstains audited |
| 15 | coding agents | one skill, none, or a person |
"none", "not sure" and a wrong pick kept apart among look-alike skills; a skill the user names is cited, one named in pasted text is not; ties escalate with both candidates (min_margin, conformal); a production deploy needs the word | cited hard checks fallback producers act_guard conformal perturb trace replay abstains audited |
More directions in examples/: HR leave approval (01), e-commerce fraud with a learned head (02), accounts payable (03), refunds with a hard 30-day rule (04), a game agent that never loses (05), logistics routing with learned rules (06), receipts with a fine-tuned model (07), contract review by field description (08), an insurance claim desk with generated strategies, early exit and parallel services (09), learning in milliseconds (10). Games: the arcade.
Against an answer-only model¶
Entries 01–06 were also run through Laya 0.3.20, a model that answers typed questions directly
(zero-shot, its own presets for 01–03 and the rules written into its questions; RTX 3060; scored on the answers both were
asked — 01's tags and 03's harm, added with solvi 0.4, are not scored):
| entry | solvi correct | Laya correct | time per decision, solvi (CPU) / Laya (GPU) |
|---|---|---|---|
| 01 support triage | 50 / 50 | 33 / 50 | ~1 ms / 63 ms |
| 02 email routing | 18 / 18 | 11 / 18 | 0.65 ms / 42 ms |
| 03 content guard | 30 / 30 | 18 / 30 | 0.76 ms / 56 ms |
| 04 security alert | 18 / 18 | 8 / 18 | 1.05 ms / 78 ms |
| 05 agent trace audit | 20 / 20 | 12 / 20 | 0.95 ms / 62 ms |
| 06 release rollout | 18 / 18 | 7 / 18 | 0.93 ms / 64 ms |
Read this as a demonstration, not a benchmark: the cases were written together with the catalogs, and the answer-only model saw
the rules only as text. The structural differences hold regardless of the numbers — solvi's hard rules cannot be overridden,
numbers are computed instead of guessed, every extracted value is quoted with its offsets, the trace re-verifies, and the
system abstains instead of guessing. For entries 01–06, compare_laya.py and compare_laya.out.txt reproduce and record the run;
measured document benchmarks are in docs/benchmarks.md.