solvi.serve¶
solvi serve: the questions over HTTP and MCP, and a decider behind the System One API.
solvi serve: a System's questions over HTTP (FastAPI) or as MCP tools, and a decider behind the System One API.
solvi serve myapp/decisions.py:system --store decisions.db # HTTP on 127.0.0.1:8000
solvi serve myapp.decisions:system --decider solvi-ai/solvi-base # + POST /v1/systemone (a cached model; --pull)
solvi serve --decider ./my-decider --model-name kev-latest # only POST /v1/systemone
solvi serve myapp.decisions:system --mcp # an MCP server over stdio
solvi serve --guard catalog.py:guard --upstream "CMD" # an MCP proxy: the guard checks every tool call
HTTP (solvi[serve]: fastapi, uvicorn):
POST /ask {"state": {...}, "questions": [names] (default: all), "store": true} → Response.to_dict() plus
"stored_id" (its id in the store, or null) and "trace_hash" (the hash at the end of its trace)
POST /ask/{question} the state itself as the body → the same response, for that question only
POST /ask_text {"text": "...", "question": null, "store": true, "today": null} → a free text through
System.ask_text: the question it asks (routed by the decider), the fields read with their quotes,
the missing ones and a clarifying question ("read"), and the answers as for /ask
GET /questions each question: its text, answer type and the JSON schema of the input state it reads
GET /health solvi's version, the catalog's fingerprint, the store, the decider
POST /v1/systemone the System One API, answered by a solvi decider (--decider): a drop-in for a Jev / Kev client
(solvi.systemone is the client side)
The OpenAPI schema (/openapi.json, /docs) comes from the same pydantic types: each question's input schema from the types of
the given facts its flow reads (System(input_model=...) fields, else the types its typed readers declare) and each response's
answers as their closed sets (System.response_schema). Inputs are not validated by the web layer: the state goes to
System.ask as it is, so a wrong-typed field is what it is in solvi — the fact is missing, the answers that need it
abstain, and the trace and the audit say why (safeguard type_rejected). With --store, every answer is stored with its
whole trace in a TraceStorage (hash-chained), and solvi verify / replay / diff work on that store — a request's
"store": false is honoured only with --allow-client-no-store.
MCP (--mcp): each question is a tool whose input schema is the question's input state schema; a call answers that
question and returns its result (answer, confidence, status, why, safeguards) with the stored id and trace hash; the tool
ask_text takes a free text and, optionally, the date it is read on (as POST /ask_text). It uses
the official mcp SDK (2.x, solvi[mcp]) when it is installed, else a built-in stdio JSON-RPC server with the subset of
the protocol that tools need (initialize, ping, tools/list, tools/call).
Security (see docs/guide.md, Serving): --token / $SOLVI_SERVE_TOKEN requires Authorization: Bearer <token> on every
HTTP request (constant-time compare); a request body / MCP message is at most --max-body bytes and --max-depth levels
of JSON; a System One request at most --max-questions questions of --max-options options; at most --max-inflight
requests run or wait at once (an async System's too), each sync one waiting at most --queue-timeout s for the
System (then 503 busy); a request
takes at most --timeout seconds (async Systems: through System.aask's part timeout, so the answer
abstains rather than the request failing); a failure the client did not cause is logged here and answered with an
incident id, never a traceback or a path — also inside an answer: a part that raised is in the response as its
exception's type and an incident id (records[].error, the alternatives tried, why, the safeguards' details), its text
only in the stored trace and the server log; CORS headers only with --cors ORIGIN; nothing is imported or loaded from
request data; --decider never downloads without --pull.
Limits
dataclass
¶
Limits(max_body: int = 1000000, max_depth: int = 32, timeout: Optional[float] = 60.0, max_questions: int = 32, max_options: int = 64, max_inflight: int = 8, queue_timeout: Optional[float] = 10.0)
What one request may be (see the module docs, Security). max_body: bytes of an HTTP body / an MCP message;
max_depth: nesting of JSON objects and arrays in it; timeout: seconds a request may take (None: no limit) — HTTP
answers 504 after it, an MCP tool call an error; an async System's parts are given 80% of it as System.aask's timeout
(unless System(timeout=) or the part sets one), so a slow part makes its questions abstain (safeguard timeout) and
the request still answers. max_questions / max_options: questions in one System One request, options (criteria) of
one of them — more is refused (422). max_inflight: requests answered at once — by a worker thread (running or
waiting for the System; a request whose thread timed out still counts until the thread ends) or, for an async
System, on the event loop — more are refused at once
with 503 "busy"; queue_timeout: seconds a request waits for the System while another is being answered (None: as
long as it takes) — then 503 "busy", instead of piling up threads behind a slow one.
RequestError ¶
Bases: Exception
A request the server refuses; its message is written by solvi (never an exception text from the catalog's code) and
is safe to return to the client. status: the HTTP status.
Busy ¶
Bases: RequestError
The server is answering as many requests as it may (Limits.max_inflight), or the System stayed busy longer than Limits.queue_timeout: try again later.
Service ¶
Service(system=None, decider=None, storage=None, model_name=None, limits=None, textin=None, allow_client_no_store=False)
What the HTTP app and the MCP server call: asks a System (one at a time: a System learns costs and counts stats
in place; a System with async parts is asked with System.aask, concurrently on the server's event loop), answers
System One requests with a decider. limits: a Limits (the request size, JSON depth and timeout).
part_timeout
property
¶
System.aask's timeout for the parts of an async System: System(timeout=) if set, else 80% of the request's.
is_async
property
¶
Does the System have parts that aask awaits? Then asks go through System.aask.
exclusive ¶
The System (or the decider) for one sync request: a slot among Limits.max_inflight (none free: Busy at once), then the lock, waited for at most Limits.queue_timeout (then Busy) — never an unbounded queue of threads.
aslot
async
¶
An async request's slot among Limits.max_inflight (shared with the sync ones; none free: Busy at once) — an async System answers requests concurrently, so this is what bounds them.
storing ¶
Whether a request is stored: always (a server with a store keeps every answer), unless the server allows clients to opt out (allow_client_no_store) and this one did.
ask ¶
→ Response.to_dict() with "stored_id" and "trace_hash". questions: names (None: all). store=False is honoured only when the server allows clients to opt out of storing (allow_client_no_store).
tool ¶
An MCP tool call: one question → its result (answer, confidence, status, why, ...), the safeguards that fired for it, "stored_id" and "trace_hash".
textin ¶
The TextIn that reads texts for this server: the one given, else TextIn(system, decider) — made once; today
(a date or ISO string, recorded in the trace) on a copy per request. Without the request's today the TextIn's
own applies (None by default): the server's date is never supplied, so a date without a year is not read ("the
year is not stated") rather than given this year.
ask_text ¶
A free text → Response.to_dict() of System.ask_text plus "read" (the question it asks, the fields read with their quotes, the missing ones, a clarifying question), "stored_id" and "trace_hash".
aask_text
async
¶
ask_text, for an async System (System.aask_text).
health ¶
{"status", "solvi", ...}: the questions, the catalog's fingerprint, the store (its file name only: no paths leave the server), the decider.
systemone ¶
A System One request (a dict, see SystemOneRequest) → the response dict, answered by the decider: choice →
the probability of each option (criteria in their order), noul → P(yes), score → the probability of each level
(criteria in order, lowest first) and the expected level index. other / none options are scored like any
other (the API has no abstain option). The answers carry no act / escalate signal: the client decides. The
request's "model" is only a name echoed back: nothing is loaded from request data.
AccessGuard ¶
ASGI middleware in front of the app (Guard up to 0.7 — a second meaning of the agent guard's name): the bearer
token (constant-time compare), the body size (Content-Length, and the bytes actually received) and the JSON depth of a
request body — refused before FastAPI parses it.
question_inputs ¶
The given facts a question's flow reads → {"properties": [fact], "required": [fact]}. Planned with every given fact
of the catalog present; a fact is required when the question cannot be answered without it (the strategist leaves it
unresolved), so the inputs of alternative producers are optional. Planned by the system's own strategist
(System(strategist=)), as ask plans.
input_model ¶
A pydantic model of the input state a question reads (extra keys allowed) — for the schema; solvi validates.
input_schema ¶
The JSON schema of the input state a question reads.
too_deep ¶
Does a JSON value nest objects / arrays deeper than max_depth? (Iterative: no recursion on hostile input.)
parse_json ¶
Bytes / text of a request → the JSON value, within the limits (BadRequest / RequestError 413 otherwise).
internal_error ¶
Log the exception being handled (with its traceback) on the server → the message for the client: an incident id, nothing of the exception (its text may carry paths, data or code).
redact ¶
A response's data for a client, without the text of any exception a part raised → d (changed in place): each exception text in the trace's records (their error, the alternatives tried) — and wherever it is repeated (why, the safeguards' details, a group's "no producer accepted") — becomes "Type (incident …)". The full texts are logged on the server under the incident id and stay in the stored trace; the response's trace_hash is the stored trace's.
trace_hash ¶
The hash at the end of a response's trace (its last record's; the input's hash when nothing ran).
create_app ¶
create_app(system=None, decider=None, storage=None, model_name=None, title=None, limits=None, token=None, cors=None, textin=None, allow_client_no_store=False)
The FastAPI app (see the module docs). system: a System; decider: a DecideModel for POST /v1/systemone and for
routing texts (POST /ask_text); storage: a TraceStorage or a path (every ask is stored); model_name: the model name
System One answers carry; limits: a Limits (request size, JSON depth, timeout); token: every request must carry
Authorization: Bearer <token> (None: no authentication); cors: the origins browsers may call it from (None: no
CORS headers at all); textin: a solvi.textin.TextIn for POST /ask_text (default: TextIn(system, decider));
allow_client_no_store: honour a request's "store": false (default: with a store, every answer is saved — the
server's policy, not the client's). A token that is empty or blank is a configuration error (ValueError).
mcp_tools ¶
The questions as MCP tools: [{"name", "description", "inputSchema"}] (a name outside [A-Za-z0-9_-] is mapped).
run_builtin ¶
A stdio MCP server without the SDK: JSON-RPC 2.0, one message per line; initialize, ping, tools/list, tools/call
(notifications are read and ignored). A message is at most svc.limits.max_body characters and max_depth deep; a
tools/call runs in a worker thread and fails after svc.limits.timeout seconds (the thread cannot be stopped: it
finishes in the background, and a later call waits for the System at most queue_timeout seconds, then is told the
server is busy). A malformed message is answered with a JSON-RPC error; nothing in a message stops the server.
sdk_available ¶
Is the official MCP SDK (2.x, handlers in the Server constructor) installed?