Examples¶
Each file is a self-contained script. Run from the repository root with uv run python examples/<file>.
| file | domain | what it shows | needs |
|---|---|---|---|
| 01_leave_request.py | HR | rules and hard checks over a plain dict; the strategist skips catalog parts the questions do not need | core |
| 02_shop_order.py | e-commerce | two answers by rules, one ("suspicious?") learned from labeled history with fit |
core |
| 03_invoices.py | accounts payable | fields extracted from invoice text with quotes, checks on computed facts, a learned risk level | core |
| 04_refunds.py | customer support | a learned yes/no that a hard "within 30 days" check always overrides | core |
| 05_tic_tac_toe.py | games | an agent from small functions over the board; never loses (checked exhaustively) | core |
| 06_learned_rules.py | logistics | learn_rule: a readable if-then list learned from labeled addresses |
core |
| 09_strategy_at_scale.py | insurance | a different generated plan per question set on a big catalog, early exit on hard checks, parallel slow services; timed | core |
| 10_learn_in_milliseconds.py | any | fit: learn a question in milliseconds and absorb each correction instantly with teach |
core |
| 11_answer_types_and_constraints.py | trust & safety | multi-label and ordinal answers, constraints between answers, joint decoding | core |
| 12_grounded_audit.py | expenses | one catalog with and without models: provenance, res.audit(), a hallucinated quote caught, a decision outside its options, a changed model, safeguard stats |
core |
| 13_decide_model.py | customer support | a decider model as a catalog part: bias correction without labels, few-shot shift with teach, "other" as a threshold, abstention, constraints, audit, escalation for a target error rate, a JSON ticket |
core (stand-in); the real model with solvi[onnx] or solvi[model] |
| 14_typed_catalog.py | customs | typed facts: type hints checked between producers and consumers at registration, a pydantic request (System(input_model=...)), answer types from the rules' return types, type_rejected → fallback / abstention, the response as JSON that loads back and replays |
core |
| 15_typed_decisions.py | customer support | typed decisions: the questions as a pydantic model's fields (choice, ordinal score, yes/no, multi-label) about a pydantic ticket, four answers (one forward pass when the model shares passes), act / escalate, a hard check, a constraint and a rule over the model, audit and stats | core (stand-in); the real model with solvi[onnx] or solvi[model] |
| 16_primitives.py | insurance claims | answer primitives, each a value and a confidence: "not stated" (Maybe[bool]) vs abstain, a span parsed into a float, evidence quotes checked in the text (require_evidence), a ranking (Rank), an estimate with an interval (Estimate) — from plain rules and from a decider with the typed v2 contract; the confidence table, JSON round trip and replay |
core (stand-in); a typed v2 checkpoint with SOLVI_DECIDE_MODEL |
| 17_model_strategist.py | payments | the code strategist and the model strategist (experimental): a dead end the deterministic strategist cannot plan around, the cheapest verified plan with declared costs, a model's proposal checked (a bad one falls back), the plan in the trace; aliases for another team's names accepted by examples and targeted questions | core (stand-ins); the trained model with SOLVI_STRATEGIST |
| 18_several_models.py | customer support | several models, one decision (solvi.multi): a cascade small → large, a vote of two model families, a route by code — each under one guarantee from act_guard (P(answered alone and wrong) ≤ 10%), with cost per question; the audit lists every stage and the trace replays |
core (stand-ins); checkpoints with SOLVI_DECIDE_SMALL / SOLVI_DECIDE_MODEL |
| 19_agent_guard.py | accounts payable | an agent's tool calls through solvi.agents.Guard: the agent proposes, solvi checks (catalog, types, arguments quoted from the conversation, instructions hidden in tool outputs, policies, an authorizer with act_guard and perturb) and allows, denies or escalates; a person's approval recorded; every decision stored and replayed |
core (a scripted agent and a stand-in authorizer) |
| 20_vote_across_families.py | customer support | a vote of two model families behind the System One API (stand-in servers started in-process): each alone and the vote under one guarantee (act_guard, P(answered alone and wrong) ≤ 10%) — the vote answers more alone than either model; a sure mistake of one family escalates; a hard check first; the audit and the replay |
core (stand-in servers); real servers with systemone(URL, MODEL) |
| 21_verified_chart.py | reports, press releases | a verified chart (solvi.charts, preview): a proposer writes a chart spec with a quote per value; code checks every number (the quote, the unit, the scale), the chart type (a pie only for shares of a whole) and a stated total; a careless model's swapped digit, invented share, percentage points drawn as percent and unquoted value are dropped with reasons; a deterministic accessible SVG; the trace replays to identical bytes |
core |
| 22_coding_agent_hooks.py | software development | a coding agent behind solvi hook (Claude Code's hooks): solvi hook install in a temporary project, then three edits (allowed; denied with the rule and the line; a comment that tries to talk past the rules, denied and flagged) and two prompts for the skill picker, run through the installed commands as Claude Code runs them; the store verified, one decision audited and replayed. The rules it installs: coding_agent_rules.toml |
core |
| 07_receipts_model.py | expenses | a receipts-tuned ModernBERT extractor cites each field; rules decide | solvi[model] |
| 08_contracts_by_description.py | legal | fields defined only by description, read from a whole contract, cited or "absent" | solvi[model] |
Models for 07 and 08 load from Hugging Face (solvi-ai/extract-receipts, solvi-ai/extract-base) or from a local directory
given in SOLVI_MODEL. Examples 13, 15, 16 and 18 use the decider from SOLVI_DECIDE_MODEL (a checkpoint folder or a Hugging Face id) when
it is set, and a keyword stand-in otherwise.