Skip to content

Serving: HTTP, MCP and System One

solvi serve puts a System behind an HTTP API, or behind an MCP server so that an agent calls its questions as tools. The System is named as for solvi diff: module:attribute or file.py:attribute (a System, or a function returning one).

solvi serve myapp/decisions.py:system --store decisions.db     # HTTP on 127.0.0.1:8000 (--host, --port); every answer stored
solvi serve myapp.decisions:system --mcp                       # an MCP server over stdio: each question is a tool
solvi serve myapp.decisions:system --decider solvi-ai/solvi-base   # + POST /v1/systemone
Endpoint What it does
POST /ask {"state": {...}, "questions": [...] (default: all), "store": true} → Response.to_dict() plus stored_id and trace_hash
POST /ask/{question} the input state itself as the body → the same response, for that question
POST /ask_text {"text": "...", "question": null, "store": true, "today": null} → a free text through ask_text: the response as for /ask plus read — the question it asks, each field with its status, value and quote [text, start, end], missing, clarify (a question asking for what is missing) and escalated
GET /questions each question: its text, answer type and the JSON schema of the input state it reads
GET /health solvi's version, the questions, the catalog's fingerprint, the store, the decider
POST /v1/systemone the System One API answered by a solvi decider (--decider)

The OpenAPI document (/openapi.json, /docs) is built from the same pydantic types as the rest of solvi: a question's input schema lists the given facts its flow reads — typed by System(input_model=...), else by the types its typed readers declare — with the ones it cannot be answered without as required (solvi.serve.question_inputs); its response schema has each answer as its closed set (System.response_schema). The web layer does not validate the state: it goes to System.ask as it is, so a wrong-typed field is handled as solvi handles it — the fact is missing, the answers that need it abstain, and the response says why (safeguard type_rejected) — rather than as a 422. Unknown questions are a 404. With --store (or a System built with storage=), every answer is saved with its whole trace; stored_id finds it (store.get(id)) and solvi verify / replay / diff work on the store. Storing is the server's policy: a request's "store": false is ignored unless the server was started with --allow-client-no-store (create_app(..., allow_client_no_store=True)). Asks are served one at a time: a System updates its measured costs and stats in place. A System with async def (or blocking=True) parts is served with aask instead: its endpoints are async and asks run concurrently on the server's event loop (the MCP server too).

Text in. POST /ask_text reads a message with solvi.textin.TextIn(system, decider) — --decider picks the entry point (any decider: a checkpoint, systemone:URL#model, llm:URL#model), and the deterministic CueExtractor reads the fields (TextIn(extractor=DeciderExtractor(decider)) uses the decider's span pointer); create_app(..., textin=TextIn(...)) or Service(..., textin=...) sets synonyms, patterns and cues. Without a decider a text can only go to a named question (or to the one question of a System with one), else the request is a 422. Dates without a year, two-digit years and relative dates are read against today — the request's ("today": "2026-09-28"; the MCP tool takes it too), else the TextIn's — which the trace records; the server never supplies its own date, so without one "paid 12 September" is not read (the field is missing, "the year is not stated") rather than given this year. A text that does not say which question it asks is not an error: read.question is null, read.escalated says why, the likely questions abstain, and read.clarify asks which one is meant; a required field the text does not give is listed in read.missing and the question abstains for lack of it — nothing is guessed.

MCP. With --mcp, each question is a tool: its input schema is the question's input state schema, and a call returns the question's result — answer, confidence, status, why, guard, evidence, the safeguards that fired — with stored_id and trace_hash, as JSON text and as structured content. One more tool, ask_text (solvi_ask_text if a question has that name), takes {"text", "question"?, "today"?} and returns what POST /ask_text does, so an agent can pass a user's message as it is. An abstention is a result, not an error; an exception is a tool error (isError). The official mcp SDK (2.x, solvi[mcp]) serves it when installed; otherwise solvi's built-in stdio JSON-RPC server answers initialize, ping, tools/list and tools/call (--mcp-impl sdk|builtin chooses). The two answer alike — an unknown tool is a JSON-RPC error (-32602) in both — except for what the SDK decides itself: arguments that are not an object are its protocol error (the built-in server returns a tool error), and a call still running when stdin closes is not answered. For an MCP client:

{"mcpServers": {"refunds": {"command": "solvi", "args": ["serve", "/path/to/refunds.py:system", "--mcp",
                                                         "--store", "/path/to/decisions.db"]}}}

A guard in front of an MCP server. solvi serve --guard catalog.py:guard --upstream CMD is the other way round: an MCP proxy that checks every tool call an agent makes to another MCP server — see Guarding an agent's tool calls.

System One. With --decider (a checkpoint folder, a Hugging Face id already in the local cache — solvi serve never downloads one unless you add --pull, as solvi models pull would —, systemone:URL#model or module:attr; --backend onnx|torch), the same server answers POST /v1/systemone — the protocol solvi.systemone speaks as a client — so solvi can stand where a Jev or Kev client points:

{"state": "I was charged twice" (or a JSON state), "model": "...",
 "questions": {"team":   {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Charges", "shipping": null}},
               "urgent": {"type": "noul",   "instructions": "Urgent?"},
               "level":  {"type": "score",  "instructions": "Priority?", "criteria": {"low": null, "medium": null, "high": null}}}}
→ {"model": "<--model-name, default the decider's id>", "usage": {"questions": 3, "passes": 3}, "latency_ms": 41.2,
   "answers": {"team":   {"type": "choice", "choice": "billing", "confidence": 0.93, "probabilities": {...}},
               "urgent": {"type": "noul", "noul": 0.12},
               "level":  {"type": "score", "score": 0.4, "confidence": 0.7, "legend": ["low", "medium", "high"],
                          "probabilities": {...}}}}

noul is P(yes); a score's score is the expected level index (0 = legend[0], the lowest); criteria are the options in order, with optional descriptions. Every option is scored ("other" / "none" included: the API has no abstain option), and the answers carry no act / escalate signal: thresholds (act_guard and the rest) belong to the client, where systemone(url, model) turns the probabilities back into a decider. solvi serve --decider X without a System serves only this endpoint. Without FastAPI, solvi.serve.Service(system, decider) answers the same requests in-process (.ask(state), .systemone(body), .tool(question, state)).

Security. The defaults are for a service on your own machine (127.0.0.1); before you expose it:

  • Authentication. SOLVI_SERVE_TOKEN=... solvi serve ... (or --token, which other local users can see in the process list) makes every HTTP request — the docs and /health included — carry Authorization: Bearer <token>; the token is compared in constant time; an empty --token "" is refused (create_app(token="") raises), and an empty SOLVI_SERVE_TOKEN counts as no token, with a warning. Without a token the server warns when it listens beyond the loopback address. For anything more (users, rate limits, TLS) put it behind a reverse proxy. MCP runs over stdio: the client that starts the process is the one that can call it.
  • Limits. A request body (an MCP message) is at most --max-body bytes (default 1 000 000: 413 above it), its JSON at most --max-depth levels deep (default 32: 400), and a request takes at most --timeout seconds (default 60: 504; an MCP tool error). A sync System cannot be interrupted: the ask finishes in a worker thread, and the next ask waits for the System at most --queue-timeout seconds (default 10), then gets a 503 "busy"; at most --max-inflight requests (default 8, a timed-out one included until its thread ends; an async System's requests count too) are running or waiting at once — more get a 503 at once, so slow asks never pile up. POST /v1/systemone takes at most --max-questions questions (default 32) of at most --max-options options each (default 64): 422 above. An async System's parts get 80% of the timeout as aask's timeout (unless System(timeout=) or the part sets one), so a slow part makes its questions abstain (safeguard timeout) and the request still answers.
  • Errors. A refused request says what was refused. Any other failure is logged on the server with its traceback (logger solvi.serve); the client gets a 500 with an incident id to look it up — never an exception text, a traceback or a path. An exception inside a catalog part is not a server error: it is part of the decision (the questions that need it abstain). Its text may carry paths or data, so the answer the client gets names only its type and an incident id — in the step's error, the alternatives tried, why and the safeguards' details ("rule not computed: RuntimeError (incident 3f2a…)"); the full text is in the server log under that id and in the stored trace (solvi replay / verify read it; trace_hash is the stored trace's). /health names the store by its file name only.
  • Every entry point. POST /ask_text and the MCP ask_text tool go through the same token, limits, timeout and error hiding as the questions. A malformed MCP message (a tool name that is not a string) is a JSON-RPC error, and nothing in a message stops the built-in server. The MCP proxy (--guard --upstream) bounds each client message by --max-body / --max-depth and answers a failure of its own with an incident id; an upstream server's own errors are passed on.
  • CORS is off: no Access-Control-Allow-* headers, so browsers on other origins cannot read the answers. --cors https://app.example (repeatable) allows one origin.
  • Nothing is loaded from request data. The System and the decider are named on the command line only; a request's model field is a name echoed back, and states are data.
  • JSON. Responses are strict JSON: a non-finite float (an escalation threshold no calibration could meet is inf) is written as {"$float": "inf"} ("-inf", "nan"), as in stored records and calibration files; Response.from_json reads it back as the float.