solvi.llm¶
Any OpenAI-compatible chat-completions server (OpenAI, OpenRouter, vLLM, llama.cpp, Ollama) as a solvi decider.
Any OpenAI-compatible chat-completions server as a decider — OpenAI, OpenRouter, vLLM, llama.cpp, Ollama, LM Studio: the LLM proposes, solvi's checks, rules and thresholds decide.
from solvi.llm import llm
model = llm("http://127.0.0.1:8080/v1", "qwen2.5-7b-instruct") # api_key="..." for a hosted service
part = model.decision("team", "Which team should handle this?", "email", {"billing": "Charges", "shipping": "Delivery"})
part.act_guard(examples, max_risk=0.10) # start with the LLM alone, calibrated
One question is one request (POST {base_url}/chat/completions, temperature 0), the input between
{"answer": <one of the options>, "probabilities": {<option>: p, ...}, "quote": "<the passage that supports it>"}
(a multi-label answer is a list and its probabilities are per option; a span answer is a passage of the text and a
confidence; "not stated" is an option only when the question allows it). The schema goes as response_format
json_schema when the server takes it, else the same contract is in the prompt and the reply is parsed. A request that
asks the model to reason (extra_body with reasoning / reasoning_effort ...) puts the contract in the prompt from
the start: a server that enforces a reply format by constrained decoding can skip the thinking altogether (some
providers behind one model name do, under json_schema and under json_object alike), and the answers are then worse.
A reply that shows no reasoning when it was asked for is marked
extra["llm"]["reasoning"] = "none", with a warning once. When the server
returns log-probabilities for the answer's tokens, the probabilities come from them (the chosen option: the product of
its tokens' probabilities; the others: the alternatives at its first token), not from the numbers the model wrote;
extra["llm"]["probabilities"] records which, and a calibration refuses examples that mix the two (two scales; see
solvi.decide.one_source).
Everything is validated: the answer is one of the options, the probabilities are numbers in [0, 1] that agree with the
answer, the quote is in the text (literally, up to typographic quotes and apostrophes, dashes, runs of whitespace and
letter case; what is recorded is the text's own spelling). A
quote that is not in the text escalates when the question asks for evidence; otherwise it is dropped (the answer stands,
extra["llm"]["quote_dropped"] records it); a span answer not in the text escalates, its passage in
extra["llm"]["rejected"]. An invalid reply, a refusal, a cut-off reply or a server that does not answer
(after retries) escalates — "model escalated: invalid LLM output — ..." — and is never turned into a guess: the
decision has no value (None), no probabilities and confidence 0. The
probabilities become the decider's logits (log p), so everything built on a DecideModel works unchanged: act_guard /
conformal / calibrate_for on your labelled examples (on the confidence: an LLM gives no act signal), fit / teach / adapt,
Cascade / Vote / Route, the audit and the trace.
The trace records the endpoint (without credentials or query), the model name and the hash of the prompt template in the
model's id and fingerprint, and per decision extra["llm"]: where the probabilities came from, the reply format used,
the model the server says answered, the quote and the tokens used. The API key is sent in the Authorization header only;
it is never in the trace, the fingerprint or an error message. The server can change the weights behind a name: calibrate
again when it does. An LLM is not replayed (its output is not reproducible bit for bit): replay checks the recorded
output instead (trust_models).
Cost and latency: every question about every input is a paid request of hundreds of tokens (the options, their descriptions and the text) and a network round trip, which a local decider does not pay; decisions are cached by (question, input) for the life of the model object. Start with it alone under act_guard; where a local decider (solvi-large) is about as strong on your stream, a Vote of the two can answer more at the same risk. "A small model first, the LLM second" is not a good default: where one model is clearly stronger, a cascade adds almost nothing over it alone (the guide's "Which combination with an LLM").
InvalidOutput ¶
Bases: ValueError
The LLM's reply broke the contract (not JSON, an answer outside the options, a quote not in the text, ...).
answer: the rejected answer when there is one (a span answer not in the text), recorded in extra["llm"]["rejected"].
LLMError ¶
Bases: RemoteError
The server refused the request itself (a wrong key, model or URL: HTTP 401, 403, 404 and other client errors) — a configuration error, raised rather than escalated. The message has the endpoint, never the key.
LLMScorer ¶
LLMScorer(base_url, model, api_key=None, *, timeout=60.0, retries=2, backoff=1.0, response_format='auto', logprobs='auto', ask='probabilities', max_tokens=None, seed=None, headers=None, extra_body=None, workers=4, opener=None, sleep=None)
Bases: RemoteClient
A scorer for DecideModel over an OpenAI-compatible POST {base_url}/chat/completions (standard library HTTP, no
dependencies; the shared client of solvi.remote). See the module docs; llm(...) builds the DecideModel.
request ¶
→ (the server's response dict, the format used); NoAnswer when the server does not answer, Refused when it refuses this request's input (after the reply-format ladder), LLMError for a wrong key, model or URL.
one ¶
One Item → the scorer's output; an invalid reply or no answer → uniform logits and escalate (never a guess).
template_hash ¶
The hash of the prompt templates and reply forms (part of every LLM decision's fingerprint).
schema ¶
The JSON schema of the reply to one question (strict: every property required, nothing else allowed).
locate ¶
A passage → (start, end) of its first occurrence in the text. Literal up to typographic quotes and apostrophes (’ ‘ “ ” as ' "), dashes (– — ‑ − as -), runs of whitespace (a new line written as a space) and, when nothing matches with the case kept, letter case ("wre54g" for "WRE54G"; ignore_case=False: not); nothing else. The offsets are into the text as given, so text[start:end] is the text's own spelling, never the passage's. None when it is not there.
llm ¶
llm(base_url, model, api_key=None, *, timeout=60.0, retries=2, backoff=1.0, response_format='auto', logprobs='auto', ask='probabilities', max_tokens=None, seed=None, headers=None, extra_body=None, workers=4, opener=None, sleep=None, max_len=None)
A DecideModel over an OpenAI-compatible chat-completions server (see the module docs).
base_url: the API root ("https://api.openai.com/v1", "https://openrouter.ai/api/v1", "http://127.0.0.1:8000/v1" for
vLLM, "http://127.0.0.1:8080/v1" for llama.cpp, "http://127.0.0.1:11434/v1" for Ollama). model: the model name the
server knows. api_key: sent as a Bearer token, never recorded. response_format: "auto" (json_schema, then json_object,
then the contract in the prompt only, as far as the server accepts — stepped down only before the first successful
request and on a 400 about the format; any other 400 / 413 / 422 escalates that question: "the LLM server refused the request"; when extra_body asks for reasoning, "auto" is the contract in the prompt only, so the model thinks before it answers — see the module docs), or one of them. logprobs: "auto" (ask for them; drop them when the server refuses; a gateway whose providers differ answers some requests from them and some from the written numbers — a calibration then refuses the mix), True, False. ask: "probabilities" (one per option) or "confidence" (one number, fewer tokens; the rest shared evenly). retries / backoff: for network errors, timeouts, 408 / 409 / 429 / 5xx. seed: sent only when set (default None: not sent; some providers reject seed 0, and at temperature 0 it rarely matters). headers: extra HTTP headers (OpenRouter's HTTP-Referer, X-Title). extra_body: server-specific request fields merged into every request's JSON — OpenRouter's provider routing and reasoning settings, vLLM's sampling extras. A field solvi sets itself (model, messages, response_format, logprobs, top_logprobs, temperature, max_tokens, seed, stream, n) is refused with ValueError, never overridden; extra_body enters the fingerprint. Pinning one OpenRouter provider, with no fallback to another:
llm("https://openrouter.ai/api/v1", "openai/gpt-oss-20b", api_key=KEY,
extra_body={"provider": {"order": ["groq"], "allow_fallbacks": False},
"reasoning": {"effort": "low"}})
workers: parallel requests for the inputs of one part.decide([...]) / calibration call (the questions of one
System.ask go one after another). opener: a replacement for urllib's urlopen
(tests, proxies); sleep: for the backoff (tests).
max_tokens: the reply's limit (default 512; 2,048 when extra_body asks for reasoning). A reasoning model's thinking
counts against it on most servers, and a reply cut off at the limit escalates ("the reply was cut off (max_tokens)"):
with reasoning on, keep the limit well above what a short answer needs, and raise the timeout too. max_len: the tokens one request reads under long="retrieve" (words
and punctuation × 1.3, the question included; default None: 512, as for a local decider). A text up to that length is sent whole; a longer
one, with long="retrieve", is read by its best sections within it — max_len=3000 reads about six times more of a
contract per request (and pays for it). Without long= the whole text is always sent. It enters the fingerprint
(through the part's long-text settings).