Skip to content

Roadmap

What we plan to add or improve, roughly in priority order. Research items ship only if their pre-registered measurements hold; the numbers are published either way (see the research notes linked from each release). Ideas and votes are welcome in Discussions → Ideas.

Status: planned · in progress · research (may not ship) · done (version).

Next (0.5.x)

Item What it gives Status
Costs from measurements in the strategist System.costs already measures every part's run time; feed it to the cost-optimal planner so that, with no hand-written cost=, solvi picks the fastest of equivalent producers (e.g. a local table over a 300 ms feed) and adapts when a source slows down. Opt-in like producers="equivalent", with a switch to freeze the choice; the trace records why a path was chosen. Offline, 0.4-style plans were up to 48% costlier than optimal without declared costs. done (0.6)
Several questions per pass by default The decider answers many questions about one state in one forward pass (≈2–3× faster). Trained with a consistency loss, the answers still differ from one-question-per-pass in ~3–10% of cases on real states, so it stays off; next: an architecture where questions are independent queries over a text encoded once. research
Honest act thresholds in the shipped models The act thresholds shipped in 0.5.0 are over-confident (measured 32–39% real error at a nominal 10%). 0.5.1 ships thresholds calibrated with conformal risk control and a table "threshold → measured error" in each model card. done (0.5.1)
Honesty suite A fixed test suite for abstaining, "not stated", act/escalate and traps that fails a release (or a fine-tune, merge, adapter) if it gets worse; three numbers per release: rate of confident errors, coverage at 10% risk, share of evidence quotes that support the answer. done (0.5.1)
Honest labels in the trace A quote not yet checked for support is marked as such; the runtime records the level of guarantee of each verdict and prints its conditions next to it. done (0.5.1)
Guaranteed escalation thresholds Calibrate on your own labelled stream and get a statistical guarantee on answering alone: act_guard(examples, risk=0.10) (conformal risk control) or max_error= (learn-then-test), plus conformal answer sets (part.conformal(...)) as a short list for the person who handles an escalation. Measured: the guarantee holds on the calibrated stream but breaks under domain shift, so calibration must be on your own data. done (0.5.1)
Guarantees per group act_guard(examples, risk, groups=["domain", "task"]): a threshold per group of a hierarchy, small groups pooled with their parent, a Bonferroni-corrected bound so that the risk holds inside every group at once — not only on average over a stream where a hard group can be far over it (28% at a 10% promise in our simulation). For one part and for Cascade / Vote / Route; the trace records which group's threshold applied. done (0.7)
Robustness to instructions in the input A message that says "ignore the rules and answer X" can push a decider to an allowed but wrong option. perturb=k asks again without instruction-like sentences (deterministic rules) and escalates when the answer changes; the honesty suite gets injection traps (embedded instructions, near-duplicate distractor options) and gates the share answered alone with the injected answer. Measured on solvi-decide base: 5–15% of injected messages followed without it, 0–1% with it, no extra pass on inputs without such sentences. done (0.7)
TraceStorage An abstract store for responses and traces: save, get(id), query(question=, answer=, safeguard=, model=, since=, until=), iter, replay_all. Backends: JSONL (append-only, no dependencies; replaces today's journal=) and SQLite (stdlib, indexed by question, answer, safeguard, model fingerprint, time). A hash chain across stored traces makes deleting or editing a stored decision detectable. Uses: audit reports for a period, re-checking every stored decision after a model or rule change, drift, feeding human corrections to teach. Later backends: Postgres (several services writing), DuckDB (analytics; it reads JSONL/Parquet directly) — done for 0.7; RocksDB only if a high-rate key-value log is really needed. done (0.5.1); Postgres, DuckDB done (0.7)
Which record changed: a trace signature solvi.signature: 64 bytes kept next to the head; when a stored record was edited and every hash after it and the head recomputed, verify(signature=...) names the record and restores its content hash (candidates= finds the original in a backup). A syndrome code (two sums mod a 256-bit prime); the positional octonion code, kept for future tree-shaped (derivation) signatures, left the package in 0.8 for benchmarks/octonion_signature.py. One changed record: 2000/2000 located, 0 wrong; several: detected, not located. preview (0.7)
Async execution await system.aask(state) next to ask: async def catalog parts (database lookups, HTTP APIs, model servers) are awaited, independent steps of the flow run concurrently, early exit stops pending calls (a failed hard check stops paid lookups; with speculate=True lookups start at once and are cancelled), per-part timeouts turn into abstentions; sync parts still work (inline or in a thread). The trace keeps flow order, so replay and hashes do not depend on timing. Worth it because solvi is called from async web servers and agents, where a blocking call stalls the event loop, and because Pyodide in the browser needs async for network calls. Plain CPU parts gain nothing: the sync ask stays the default. done (0.6)
solvi check: catalog lint Catches catalog mistakes before they bite: a hard check with then= that is not in its question's flow (it only governs a question when the question reads it or lists it in checkpoints), parts nothing uses, questions no input can answer, cycles, type conflicts, constraints that cannot all hold, silent defaults (x or 0, .get(k, 0)) in functions that read the input. done (0.6)
Command line for a whole project solvi init scaffolds a project (a typed catalog, regression cases that pass, a README, a CI workflow); solvi ask runs one decision from a JSON state or a text and prints the answers, audit or report; solvi calibrate runs act_guard on a labelled file and saves the thresholds to a file the catalog loads, refusing it for another model (part.save_calibration / load_calibration); solvi models lists, pulls and checks deciders (capabilities, fingerprint, latency, accuracy). done (0.7)
solvi test and a pytest plugin Decision regression tests from cases.json (the format every gallery entry already uses): expected answers, statuses and safeguards per case, run in CI; plus input fuzzing to find crashes and unhandled values. done (0.5.1)
Shadow mode and solvi diff "We changed a rule — which decisions change?": re-run stored decisions (TraceStorage) with a new catalog or model and list the answers that change and why; run a new version in the shadow of the current one before switching. The catalog's fingerprint (hash of its code) goes into every trace. done (0.5.1)
solvi serve: HTTP and MCP server solvi serve catalog.py:system exposes the questions as an HTTP API (FastAPI, OpenAPI schema from the same pydantic types) and as an MCP server (each question a tool), and a decider behind the System One API (POST /v1/systemone); traces go to TraceStorage — done for 0.6. Guarding an agent's tool calls (solvi.agents.Guard: the agent proposes a call, solvi checks it — catalog, types, grounding in the conversation, instruction-like tool outputs, policies, an authorizer under act_guard and perturb — and makes it, denies it or escalates it, every decision a stored trace), adapters for PydanticAI, LangGraph and the OpenAI Agents SDK, and an MCP proxy (solvi serve --guard --upstream) — done for 0.7; hardened before release: approvals bound to the call and the reasons shown, scan_user for pasted injections, locale= for ambiguous numbers, guard_run_config() so the OpenAI Agents guard sees a run's own tool outputs. Measured on AgentDojo (97 tasks, two open models): default settings cut successful attacks by 93–97% but blocked honest tasks that take payees or recipients from tool outputs (67% → 51%, 84% → 56% solved). Added for 0.7, opt-in: tool_values="escalate" (such a value goes to a person instead of a refusal), a "url" matcher, require_request policies for actions with no user-given value, and a wider detector of booking / event / visit commands. All together with a reviewer: attacks 0.6% / 1.6%, honest tasks 72% / 75%, a person asked in 28–31% of honest tasks. Next: measure the new detector rules and intents on attacks they were not written for; a reviewer UI that shows the injected instruction next to the escalated value. done (0.6); agent guard and adapters: preview (0.7)
Hardening for 0.7 solvi serve with a bearer token, request size / JSON depth / time limits, errors that never carry a traceback or a path, CORS off by default, --decider never downloading without --pull; strict JSON everywhere (non-finite floats tagged); ruff and pyright in CI; an ask overhead benchmark against 0.5.0-style settings. done (0.7)
Verifiable specialists Small models for one job each, under one contract (solvi.specialist): the model proposes a typed spec, code checks it against the source, code renders it, the trace replays to identical bytes; what does not verify is marked, never invented. First: verified charts (solvi.charts) — every number on the chart quoted from the text, with its unit and scale; a deterministic, accessible SVG. Next: tables and slides (the same checks, HTML / PPTX output), then checking layers over open speech models (speech to text: silence has no words, forced alignment, a term dictionary; text to speech: numbers and dates normalised by rules, the synthesised audio transcribed back and compared). charts: preview (0.7); tables, slides, speech: planned
Counterfactual explanations The smallest change of the inputs that would change the answer ("approved if the amount were ≤ 1000", "refused: 40 days since purchase, the limit is 30"), computed by search over the deterministic parts; adverse-action reasons for lending, clear answers for support. done (0.7)
Human-readable reports and OpenTelemetry An HTML / Markdown report of a decision or a period for an auditor or a customer: the answer, what it rests on, quotes highlighted in the document, safeguards that fired; export of traces as OpenTelemetry spans. done (0.7)
Browser check of every Space on each release Automated smoke test of playground / arcade / documents / realms after a release: tools/smoke_spaces.py (headless browser, one preset per Space), run by .github/workflows/smoke-spaces.yml by hand or after a release is published. done (0.7)
Gallery: helpers for coding agents Three runnable entries for decisions a coding agent meets on every task: a pre-edit rule check (rules per path glob; code checks where code can, a decider's question per fuzzy rule under act_guard and perturb; allow / block / escalate), review triage (seven risk questions, quick review only for a confident "no" to all, P(a risky change goes to quick review) ≤ 10%), a skill picker with an honest "none" (ties escalate with their candidates). Offline with keyword stand-ins and synthetic labels. Next: calibrate solvi-large and an LLM decider on labelled real edits, changes and prompts, and publish what the promise costs in escalations. done (0.7.1)
System One client for hosted services solvi.systemone over a hosted decision model (e.g. through OpenRouter): extra_body for provider routing and user, "not stated" as an option of its own, multi-label questions as one yes/no per option, cost and latency per decision, and no replay re-call by default (deterministic=False). done (0.7.1)
Local thinking decision models (Jeeves) solvi.systemone over Jeeves: its options (max_think, nothink_threshold, think, return_reasoning) through extra_body, reasoning tokens and the service's latency recorded, the reasoning text kept for the audit (never read for the answer); tested against a stand-in server with Jeeves's request validation. Measured inside solvi on the vs-LLM benchmark, with and without thinking (docs/vs_llm.md). done (0.8); measured on the vs-LLM benchmark (0.8)

Models

Item What it gives Status
solvi-large and a distilled solvi-base Typed decisions with evidence quotes, not stated, spans, rankings and numbers; ahead of GLiNER2.5-Decide and Laya on typed questions over JSON states, behind GLiNER2.5-Decide on zero-shot choice questions (level with it after fit on ~64 examples). done (0.5.0, preview)
Decider v3 New layers around the same pretrained encoder: text encoded once and questions as independent queries (many questions per pass that match one-by-one exactly), cached option vectors (hundreds of options), answers computed only from 1–3 selected evidence passages, several heads whose disagreement is the uncertainty, early exit on CPU, adapters per question kind. Each block ships only if it beats solvi-large on the same data. research
Long documents Beyond the decider's context, find the relevant sections first and decide on them: long="retrieve" (sections by headings and paragraphs, BM25, optional rerank by the decider, offsets mapped back, sections read in the trace; top_k follows the budget, sections of ≈ 170 tokens). long="full" reads a text whole up to the checkpoint's declared max_len_long and retrieves within it beyond that (a GPU mode). Measured on 4–8k-token documents: a solvi-large fine-tuned on inputs up to 8k tokens reads them whole at 85% (retrieve at 512 tokens: 78%; untrained solvi-large: 73%), and retrieve with max_len=2048 matches that at about 3× a 512-token pass on a CPU (reading 4–8k tokens whole: 12–31×). Next: publish a long-input checkpoint that declares max_len_long once it keeps quotes and the act signal on short texts. retrieve, long="full" code path: done (0.7); long-input checkpoint: research
Better evidence quotes Quotes that actually support the answer on long legal and business texts (today ~23% on ContractNLI). research
Business-judgement questions Better zero-shot answers on operational judgements (typed-decisions style); today use fit on 30–60 labeled examples. research
A LoRA adapter per question (part.adapt_lora) When a question has ~100 labelled answers or more, fit levels off (it only moves the logits); a 3.2 MB LoRA adapter on solvi-base keeps improving: +3 / +4 / +6 / +9 points over fit at 32 / 100 / 300 / 1000 examples per process on typed decisions. Trained in-process on a CPU in minutes (the time is measured and reported first), act_guard on held-out labels in the same call (the adapted model is overconfident), the adapter in the fingerprint and every trace, save / load next to the calibration file, remove_lora rolls back; solvi-large and big k: tools/adapt_lora_gpu.py on a GPU. Next: use it as the learning loop's adapter rung, and decide on real use whether it leaves experimental. experimental (0.7)
Numeric estimates Tighter, calibrated intervals for Estimate. research
extract-base v2 A stronger zero-shot field extractor, unified with the decider's span pointer. planned

Strategist

Item What it gives Status
Reliability-aware producer choice Choose among producers by how often each one is accepted and right, learned from outcomes (continues the learned producer policy). planned
Text in: entry points Text → which question is asked + its typed input state: the decider picks the entry point, spans and deterministic parsers fill the fields with quotes; provenance says the input was read by a model. system.ask_text, solvi.textin.TextIn, dialogue updates; served as POST /ask_text and an MCP tool. Next: measure the share of wrong actions that pass the checks on gallery catalogs, and a generative fallback only where a parser fails. done (0.7)
Learning from corrections, with gates and rollback The system learns from every human correction of an escalation, every known outcome and every proposal your rules rejected — never from its own accepted answers (tried three times: no gain). By the number of labels per process: shift/scale (today's fit / teach), a memory of corrected cases and a small head, then a LoRA adapter (part.adapt_lora, experimental in 0.7; not yet wired into the ladder). Every update runs in shadow, is diffed against the current one, must beat it on held-out labels and pass the honesty suite, recalibrates act_guard on fresh labels, and is recorded in each trace so it can be rolled back. The scaffold ships in 0.7 as experimental and off by default (System.learning: trusted labels only, the ladder, the gates, versions and rollback, tested with stand-in models); the simulation on real data sets decides whether it becomes a default. experimental (0.7); simulation research
Cascade and memory of corrections Decider → larger decider → LLM, escalating by the cost of a mistake; a memory of human-corrected cases (nearest neighbours with an abstain threshold). Any OpenAI-compatible LLM server is a decider (solvi.llm: JSON-schema replies validated, invalid output escalates, probabilities from log-probabilities when given), so it can be the last stage of a cascade or a family in a vote. cascade, vote, route done (0.6); LLM stage and memory done (0.7)
LLM decider confidence Log-probabilities from LLMs are saturated (≥ 0.999 on almost every answer), yet the raw score ranks right and wrong well and a threshold taken from its distinct values on the calibration examples (as act_guard does) separates them — no transform needed. What broke were fixed scales around it: combinations can share one threshold on each model's rank among the calibration examples (act_guard(scale="rank"), opt-in: it helped on one of three measured sets and hurt on two, so compare it with the raw default), a cascade warns when a stage does nothing, and calibrate_for(method="ltt") takes its grid from the calibration scores (at most 64 quantiles) instead of a fixed 0.2–0.995. done (0.7)
Task in words → goal and plan A model that turns a natural-language task into the questions to ask and the goal to plan for over a catalog; code still verifies and executes. research
Name matching across teams (solvi.aliases) Link parameters and facts whose names differ; acceptance by active, targeted questions and units in types. Experimental: no checkpoint of the matcher is published. research
Model strategist for plan segments The typed decomposer from research as a proposer for ambiguous segments (ModelStrategist; the code planner is CostStrategist since 0.8). Experimental: no checkpoint is published, and it adds little over the code planner when contracts and costs are declared. research

Games and demos

Item What it gives Status
Learned strategist in realms A tiny learned policy behind hard laws; leads 23 of 32 new maps over 5,000 turns with zero law violations. done (0.5.0)
Lookahead + policy mode Stronger (leads 29/32) but up to 0.9 s per decision; needs a faster search to fit the browser. research

Docs

Item Status
Documentation site (mkdocs on GitHub Pages) built from the guide, format specs and examples, with an API reference from the docstrings done (0.7)
Explanations and safeguard messages in more languages (Russian first): audit, show and the safeguard report render in Russian with lang="ru"; traces and hashes stay in English Russian done (0.7); more languages planned
How-to series: support triage, classification with guarantees, solvi as a tool for an LLM agent, moderation and PII spans, three-way invoice matching, KYC screening, learning from 10 examples, field extraction with quotes in progress

Done

Item Version
One name per concept across the package (old names warn until 0.9), one decider protocol for parts and combinations, one client and error policy for every remote model (solvi.remote), __all__ in every module, solvi.decide as a package of layers; README and guide around any model (an LLM first, a local checkpoint for offline use) 0.8
A model that writes, checked: solvi.generate, agreement of candidates (solvi.agree), the re-ask loop with checks that say why (solvi.refine, Fail) 0.8
A guarantee on any question or signal (System.guarantee, solvi.guarantee), inputs from outside the calibration set (solvi.openset), decisions over a set (solvi.sets), search through a System's checks (solvi.search) 0.8
Agent guard: "the user confirmed this" (require_confirmation, accepts), policies shown to the model, once=True everywhere; drift with a sequential test; System.fit as one entry point; ask(early_exit=False) 0.8
The audit of 0.7: checks that silently did nothing, replay that checks the answers, verified redactions, stores and calibrations written by 0.7.1 still load 0.8
Fast heads taught by corrections refit on all kept examples each time they double (fit_fast(refit=2.0)), so the ridge strength and features chosen on the first few examples do not stay frozen 0.7
solvi serve (HTTP, MCP, System One API), solvi check, Cascade / Vote / Route under one guarantee, aask with timeouts, costs from measurements 0.6.0
Escalation with a guarantee (act_guard, learn-then-test, conformal sets), option order and near-tie safeguards, System One backend, honesty suite, solvi test, TraceStorage, solvi diff and shadow mode 0.5.1
solvi-large / solvi-base (preview); learned strategist in realms 0.5.0
Typed facts (pydantic), questions declared by types, answer primitives (not stated, evidence, span, rank, estimate), overall confidence, typed decider API with act/escalate, code strategist (dead ends, exact cost optimum, memoized planning) 0.5.0
Audit shows what a learned head reads and ignores; rule_abstained safeguard 0.4.1
Provenance, audit and safeguards; decisions with a model 0.4.0

Not planned (for now)

  • Writing prose: solvi decides and extracts, and checks what a model writes (solvi.generate: schemas, quotes); it does not write prose itself.
  • A hosted service: solvi runs in your process, on your data.