| Costs from measurements in the strategist |
System.costs already measures every part's run time; feed it to the cost-optimal planner so that, with no hand-written cost=, solvi picks the fastest of equivalent producers (e.g. a local table over a 300 ms feed) and adapts when a source slows down. Opt-in like producers="equivalent", with a switch to freeze the choice; the trace records why a path was chosen. Offline, 0.4-style plans were up to 48% costlier than optimal without declared costs. |
done (0.6) |
| Several questions per pass by default |
The decider answers many questions about one state in one forward pass (≈2–3× faster). Trained with a consistency loss, the answers still differ from one-question-per-pass in ~3–10% of cases on real states, so it stays off; next: an architecture where questions are independent queries over a text encoded once. |
research |
| Honest act thresholds in the shipped models |
The act thresholds shipped in 0.5.0 are over-confident (measured 32–39% real error at a nominal 10%). 0.5.1 ships thresholds calibrated with conformal risk control and a table "threshold → measured error" in each model card. |
done (0.5.1) |
| Honesty suite |
A fixed test suite for abstaining, "not stated", act/escalate and traps that fails a release (or a fine-tune, merge, adapter) if it gets worse; three numbers per release: rate of confident errors, coverage at 10% risk, share of evidence quotes that support the answer. |
done (0.5.1) |
| Honest labels in the trace |
A quote not yet checked for support is marked as such; the runtime records the level of guarantee of each verdict and prints its conditions next to it. |
done (0.5.1) |
| Guaranteed escalation thresholds |
Calibrate on your own labelled stream and get a statistical guarantee on answering alone: act_guard(examples, risk=0.10) (conformal risk control) or max_error= (learn-then-test), plus conformal answer sets (part.conformal(...)) as a short list for the person who handles an escalation. Measured: the guarantee holds on the calibrated stream but breaks under domain shift, so calibration must be on your own data. |
done (0.5.1) |
| Guarantees per group |
act_guard(examples, risk, groups=["domain", "task"]): a threshold per group of a hierarchy, small groups pooled with their parent, a Bonferroni-corrected bound so that the risk holds inside every group at once — not only on average over a stream where a hard group can be far over it (28% at a 10% promise in our simulation). For one part and for Cascade / Vote / Route; the trace records which group's threshold applied. |
done (0.7) |
| Robustness to instructions in the input |
A message that says "ignore the rules and answer X" can push a decider to an allowed but wrong option. perturb=k asks again without instruction-like sentences (deterministic rules) and escalates when the answer changes; the honesty suite gets injection traps (embedded instructions, near-duplicate distractor options) and gates the share answered alone with the injected answer. Measured on solvi-decide base: 5–15% of injected messages followed without it, 0–1% with it, no extra pass on inputs without such sentences. |
done (0.7) |
| TraceStorage |
An abstract store for responses and traces: save, get(id), query(question=, answer=, safeguard=, model=, since=, until=), iter, replay_all. Backends: JSONL (append-only, no dependencies; replaces today's journal=) and SQLite (stdlib, indexed by question, answer, safeguard, model fingerprint, time). A hash chain across stored traces makes deleting or editing a stored decision detectable. Uses: audit reports for a period, re-checking every stored decision after a model or rule change, drift, feeding human corrections to teach. Later backends: Postgres (several services writing), DuckDB (analytics; it reads JSONL/Parquet directly) — done for 0.7; RocksDB only if a high-rate key-value log is really needed. |
done (0.5.1); Postgres, DuckDB done (0.7) |
| Which record changed: a trace signature |
solvi.signature: 64 bytes kept next to the head; when a stored record was edited and every hash after it and the head recomputed, verify(signature=...) names the record and restores its content hash (candidates= finds the original in a backup). A syndrome code (two sums mod a 256-bit prime); the positional octonion code, kept for future tree-shaped (derivation) signatures, left the package in 0.8 for benchmarks/octonion_signature.py. One changed record: 2000/2000 located, 0 wrong; several: detected, not located. |
preview (0.7) |
| Async execution |
await system.aask(state) next to ask: async def catalog parts (database lookups, HTTP APIs, model servers) are awaited, independent steps of the flow run concurrently, early exit stops pending calls (a failed hard check stops paid lookups; with speculate=True lookups start at once and are cancelled), per-part timeouts turn into abstentions; sync parts still work (inline or in a thread). The trace keeps flow order, so replay and hashes do not depend on timing. Worth it because solvi is called from async web servers and agents, where a blocking call stalls the event loop, and because Pyodide in the browser needs async for network calls. Plain CPU parts gain nothing: the sync ask stays the default. |
done (0.6) |
solvi check: catalog lint |
Catches catalog mistakes before they bite: a hard check with then= that is not in its question's flow (it only governs a question when the question reads it or lists it in checkpoints), parts nothing uses, questions no input can answer, cycles, type conflicts, constraints that cannot all hold, silent defaults (x or 0, .get(k, 0)) in functions that read the input. |
done (0.6) |
| Command line for a whole project |
solvi init scaffolds a project (a typed catalog, regression cases that pass, a README, a CI workflow); solvi ask runs one decision from a JSON state or a text and prints the answers, audit or report; solvi calibrate runs act_guard on a labelled file and saves the thresholds to a file the catalog loads, refusing it for another model (part.save_calibration / load_calibration); solvi models lists, pulls and checks deciders (capabilities, fingerprint, latency, accuracy). |
done (0.7) |
solvi test and a pytest plugin |
Decision regression tests from cases.json (the format every gallery entry already uses): expected answers, statuses and safeguards per case, run in CI; plus input fuzzing to find crashes and unhandled values. |
done (0.5.1) |
Shadow mode and solvi diff |
"We changed a rule — which decisions change?": re-run stored decisions (TraceStorage) with a new catalog or model and list the answers that change and why; run a new version in the shadow of the current one before switching. The catalog's fingerprint (hash of its code) goes into every trace. |
done (0.5.1) |
solvi serve: HTTP and MCP server |
solvi serve catalog.py:system exposes the questions as an HTTP API (FastAPI, OpenAPI schema from the same pydantic types) and as an MCP server (each question a tool), and a decider behind the System One API (POST /v1/systemone); traces go to TraceStorage — done for 0.6. Guarding an agent's tool calls (solvi.agents.Guard: the agent proposes a call, solvi checks it — catalog, types, grounding in the conversation, instruction-like tool outputs, policies, an authorizer under act_guard and perturb — and makes it, denies it or escalates it, every decision a stored trace), adapters for PydanticAI, LangGraph and the OpenAI Agents SDK, and an MCP proxy (solvi serve --guard --upstream) — done for 0.7; hardened before release: approvals bound to the call and the reasons shown, scan_user for pasted injections, locale= for ambiguous numbers, guard_run_config() so the OpenAI Agents guard sees a run's own tool outputs. Measured on AgentDojo (97 tasks, two open models): default settings cut successful attacks by 93–97% but blocked honest tasks that take payees or recipients from tool outputs (67% → 51%, 84% → 56% solved). Added for 0.7, opt-in: tool_values="escalate" (such a value goes to a person instead of a refusal), a "url" matcher, require_request policies for actions with no user-given value, and a wider detector of booking / event / visit commands. All together with a reviewer: attacks 0.6% / 1.6%, honest tasks 72% / 75%, a person asked in 28–31% of honest tasks. Next: measure the new detector rules and intents on attacks they were not written for; a reviewer UI that shows the injected instruction next to the escalated value. |
done (0.6); agent guard and adapters: preview (0.7) |
| Hardening for 0.7 |
solvi serve with a bearer token, request size / JSON depth / time limits, errors that never carry a traceback or a path, CORS off by default, --decider never downloading without --pull; strict JSON everywhere (non-finite floats tagged); ruff and pyright in CI; an ask overhead benchmark against 0.5.0-style settings. |
done (0.7) |
| Verifiable specialists |
Small models for one job each, under one contract (solvi.specialist): the model proposes a typed spec, code checks it against the source, code renders it, the trace replays to identical bytes; what does not verify is marked, never invented. First: verified charts (solvi.charts) — every number on the chart quoted from the text, with its unit and scale; a deterministic, accessible SVG. Next: tables and slides (the same checks, HTML / PPTX output), then checking layers over open speech models (speech to text: silence has no words, forced alignment, a term dictionary; text to speech: numbers and dates normalised by rules, the synthesised audio transcribed back and compared). |
charts: preview (0.7); tables, slides, speech: planned |
| Counterfactual explanations |
The smallest change of the inputs that would change the answer ("approved if the amount were ≤ 1000", "refused: 40 days since purchase, the limit is 30"), computed by search over the deterministic parts; adverse-action reasons for lending, clear answers for support. |
done (0.7) |
| Human-readable reports and OpenTelemetry |
An HTML / Markdown report of a decision or a period for an auditor or a customer: the answer, what it rests on, quotes highlighted in the document, safeguards that fired; export of traces as OpenTelemetry spans. |
done (0.7) |
| Browser check of every Space on each release |
Automated smoke test of playground / arcade / documents / realms after a release: tools/smoke_spaces.py (headless browser, one preset per Space), run by .github/workflows/smoke-spaces.yml by hand or after a release is published. |
done (0.7) |
| Gallery: helpers for coding agents |
Three runnable entries for decisions a coding agent meets on every task: a pre-edit rule check (rules per path glob; code checks where code can, a decider's question per fuzzy rule under act_guard and perturb; allow / block / escalate), review triage (seven risk questions, quick review only for a confident "no" to all, P(a risky change goes to quick review) ≤ 10%), a skill picker with an honest "none" (ties escalate with their candidates). Offline with keyword stand-ins and synthetic labels. Next: calibrate solvi-large and an LLM decider on labelled real edits, changes and prompts, and publish what the promise costs in escalations. |
done (0.7.1) |
| System One client for hosted services |
solvi.systemone over a hosted decision model (e.g. through OpenRouter): extra_body for provider routing and user, "not stated" as an option of its own, multi-label questions as one yes/no per option, cost and latency per decision, and no replay re-call by default (deterministic=False). |
done (0.7.1) |
| Local thinking decision models (Jeeves) |
solvi.systemone over Jeeves: its options (max_think, nothink_threshold, think, return_reasoning) through extra_body, reasoning tokens and the service's latency recorded, the reasoning text kept for the audit (never read for the answer); tested against a stand-in server with Jeeves's request validation. Measured inside solvi on the vs-LLM benchmark, with and without thinking (docs/vs_llm.md). |
done (0.8); measured on the vs-LLM benchmark (0.8) |