Serving: HTTP, MCP and System One¶
solvi serve puts a System behind an HTTP API, or behind an MCP server so that an agent calls its questions as tools.
The System is named as for solvi diff: module:attribute or file.py:attribute (a System, or a function returning one).
solvi serve myapp/decisions.py:system --store decisions.db # HTTP on 127.0.0.1:8000 (--host, --port); every answer stored
solvi serve myapp.decisions:system --mcp # an MCP server over stdio: each question is a tool
solvi serve myapp.decisions:system --decider solvi-ai/solvi-base # + POST /v1/systemone
| Endpoint | What it does |
|---|---|
POST /ask |
{"state": {...}, "questions": [...] (default: all), "store": true} → Response.to_dict() plus stored_id and trace_hash |
POST /ask/{question} |
the input state itself as the body → the same response, for that question |
POST /ask_text |
{"text": "...", "question": null, "store": true, "today": null} → a free text through ask_text: the response as for /ask plus read — the question it asks, each field with its status, value and quote [text, start, end], missing, clarify (a question asking for what is missing) and escalated |
GET /questions |
each question: its text, answer type and the JSON schema of the input state it reads |
GET /health |
solvi's version, the questions, the catalog's fingerprint, the store, the decider |
POST /v1/systemone |
the System One API answered by a solvi decider (--decider) |
The OpenAPI document (/openapi.json, /docs) is built from the same pydantic types as the rest of solvi: a question's
input schema lists the given facts its flow reads — typed by System(input_model=...), else by the types its typed readers
declare — with the ones it cannot be answered without as required (solvi.serve.question_inputs); its response schema
has each answer as its closed set (System.response_schema). The web layer does not validate the state: it goes to
System.ask as it is, so a wrong-typed field is handled as solvi handles it — the fact is missing, the answers that need
it abstain, and the response says why (safeguard type_rejected) — rather than as a 422. Unknown questions are a 404.
With --store (or a System built with storage=), every answer is saved with its whole trace; stored_id finds it
(store.get(id)) and solvi verify / replay / diff work on the store. Storing is the server's policy: a request's
"store": false is ignored unless the server was started with --allow-client-no-store
(create_app(..., allow_client_no_store=True)). Asks are served one at a time: a System
updates its measured costs and stats in place. A System with async def (or blocking=True) parts is served with
aask instead: its endpoints are async and asks run concurrently on the server's event loop
(the MCP server too).
Text in. POST /ask_text reads a message with solvi.textin.TextIn(system, decider) — --decider picks the entry
point (any decider: a checkpoint, systemone:URL#model, llm:URL#model), and the deterministic CueExtractor reads
the fields (TextIn(extractor=DeciderExtractor(decider)) uses the decider's span pointer); create_app(..., textin=TextIn(...)) or
Service(..., textin=...) sets synonyms, patterns and cues. Without a decider a text can only go to a named question
(or to the one question of a System with one), else the request is a 422. Dates without a year, two-digit years and
relative dates are read against today — the request's ("today": "2026-09-28"; the MCP tool takes it too), else the
TextIn's — which the trace records; the server never supplies its own date, so without one "paid 12 September" is
not read (the field is missing, "the year is not stated") rather than given this year. A text that does not say which
question it asks is not an error: read.question is null, read.escalated says why, the likely questions abstain, and
read.clarify asks which one is meant; a required field the text does not give is listed in read.missing and the
question abstains for lack of it — nothing is guessed.
MCP. With --mcp, each question is a tool: its input schema is the question's input state schema, and a call returns
the question's result — answer, confidence, status, why, guard, evidence, the safeguards that fired — with stored_id
and trace_hash, as JSON text and as structured content. One more tool, ask_text (solvi_ask_text if a question has
that name), takes {"text", "question"?, "today"?} and returns what POST /ask_text does, so an agent can pass a user's message
as it is. An abstention is a result, not an error; an exception is a tool
error (isError). The official mcp SDK (2.x, solvi[mcp]) serves it when installed; otherwise solvi's built-in stdio
JSON-RPC server answers initialize, ping, tools/list and tools/call (--mcp-impl sdk|builtin chooses). The two
answer alike — an unknown tool is a JSON-RPC error (-32602) in both — except for what the SDK decides itself:
arguments that are not an object are its protocol error (the built-in server returns a tool error), and a call still
running when stdin closes is not answered. For an
MCP client:
{"mcpServers": {"refunds": {"command": "solvi", "args": ["serve", "/path/to/refunds.py:system", "--mcp",
"--store", "/path/to/decisions.db"]}}}
A guard in front of an MCP server. solvi serve --guard catalog.py:guard --upstream CMD is the other way round: an
MCP proxy that checks every tool call an agent makes to another MCP server — see
Guarding an agent's tool calls.
System One. With --decider (a checkpoint folder, a Hugging Face id already in the local cache — solvi serve
never downloads one unless you add --pull, as solvi models pull would —, systemone:URL#model or module:attr;
--backend onnx|torch), the same server
answers POST /v1/systemone — the protocol solvi.systemone speaks as a client — so solvi can stand where a Jev or Kev
client points:
{"state": "I was charged twice" (or a JSON state), "model": "...",
"questions": {"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Charges", "shipping": null}},
"urgent": {"type": "noul", "instructions": "Urgent?"},
"level": {"type": "score", "instructions": "Priority?", "criteria": {"low": null, "medium": null, "high": null}}}}
→ {"model": "<--model-name, default the decider's id>", "usage": {"questions": 3, "passes": 3}, "latency_ms": 41.2,
"answers": {"team": {"type": "choice", "choice": "billing", "confidence": 0.93, "probabilities": {...}},
"urgent": {"type": "noul", "noul": 0.12},
"level": {"type": "score", "score": 0.4, "confidence": 0.7, "legend": ["low", "medium", "high"],
"probabilities": {...}}}}
noul is P(yes); a score's score is the expected level index (0 = legend[0], the lowest); criteria are the options
in order, with optional descriptions. Every option is scored ("other" / "none" included: the API has no abstain option),
and the answers carry no act / escalate signal: thresholds (act_guard and the rest) belong to the client, where
systemone(url, model) turns the probabilities back into a decider. solvi serve --decider X without a System serves
only this endpoint. Without FastAPI, solvi.serve.Service(system, decider) answers the same requests in-process
(.ask(state), .systemone(body), .tool(question, state)).
Security. The defaults are for a service on your own machine (127.0.0.1); before you expose it:
- Authentication.
SOLVI_SERVE_TOKEN=... solvi serve ...(or--token, which other local users can see in the process list) makes every HTTP request — the docs and/healthincluded — carryAuthorization: Bearer <token>; the token is compared in constant time; an empty--token ""is refused (create_app(token="")raises), and an emptySOLVI_SERVE_TOKENcounts as no token, with a warning. Without a token the server warns when it listens beyond the loopback address. For anything more (users, rate limits, TLS) put it behind a reverse proxy. MCP runs over stdio: the client that starts the process is the one that can call it. - Limits. A request body (an MCP message) is at most
--max-bodybytes (default 1 000 000: 413 above it), its JSON at most--max-depthlevels deep (default 32: 400), and a request takes at most--timeoutseconds (default 60: 504; an MCP tool error). A sync System cannot be interrupted: the ask finishes in a worker thread, and the next ask waits for the System at most--queue-timeoutseconds (default 10), then gets a 503 "busy"; at most--max-inflightrequests (default 8, a timed-out one included until its thread ends; an async System's requests count too) are running or waiting at once — more get a 503 at once, so slow asks never pile up.POST /v1/systemonetakes at most--max-questionsquestions (default 32) of at most--max-optionsoptions each (default 64): 422 above. An async System's parts get 80% of the timeout asaask's timeout (unlessSystem(timeout=)or the part sets one), so a slow part makes its questions abstain (safeguardtimeout) and the request still answers. - Errors. A refused request says what was refused. Any other failure is logged on the server with its traceback
(logger
solvi.serve); the client gets a 500 with an incident id to look it up — never an exception text, a traceback or a path. An exception inside a catalog part is not a server error: it is part of the decision (the questions that need it abstain). Its text may carry paths or data, so the answer the client gets names only its type and an incident id — in the step'serror, the alternatives tried,whyand the safeguards' details ("rule not computed: RuntimeError (incident 3f2a…)"); the full text is in the server log under that id and in the stored trace (solvi replay/verifyread it;trace_hashis the stored trace's)./healthnames the store by its file name only. - Every entry point.
POST /ask_textand the MCPask_texttool go through the same token, limits, timeout and error hiding as the questions. A malformed MCP message (a tool name that is not a string) is a JSON-RPC error, and nothing in a message stops the built-in server. The MCP proxy (--guard --upstream) bounds each client message by--max-body/--max-depthand answers a failure of its own with an incident id; an upstream server's own errors are passed on. - CORS is off: no
Access-Control-Allow-*headers, so browsers on other origins cannot read the answers.--cors https://app.example(repeatable) allows one origin. - Nothing is loaded from request data. The System and the decider are named on the command line only; a request's
modelfield is a name echoed back, and states are data. - JSON. Responses are strict JSON: a non-finite float (an escalation threshold no calibration could meet is
inf) is written as{"$float": "inf"}("-inf","nan"), as in stored records and calibration files;Response.from_jsonreads it back as the float.