solvi.generate¶
Generation by an LLM — a text, or JSON validated against a pydantic model or a JSON schema, with quotes checked in a text — on solvi.llm's client settings, recorded in the trace as a model's output.
Generation by an LLM: a text, or JSON checked against a pydantic model or a JSON schema — through the same OpenAI-compatible client settings as solvi.llm, recorded in the trace as a model's output.
from solvi.generate import generator
writer = generator("http://127.0.0.1:8080/v1", "qwen2.5-7b-instruct", extra_body={"reasoning": {"effort": "low"}})
g = writer.generate("Write one SQLite query that ...") # g.value: the reply's text
g = writer.generate(messages, schema=Plan) # g.value: a Plan (pydantic), validated
g = writer.generate(messages, schema=Table, text=doc, quotes=["rows"]) # every row literally in doc
g = writer.sample(messages, k=3, temperature=0.8) # g.value: 3 outputs (None where one failed)
cat.fn(writer.part("sql", prompt, k=3)) # a catalog part: the outputs recorded, not re-run
solvi.llm asks closed questions (options, yes / no, a span); this module is for the model that writes something — a query, a plan, a JSON extraction of a table — which solvi's checks then judge. The model proposes; nothing here decides.
What is checked, and what escalates. A reply is accepted only when it is complete and well-formed: a refusal, a cut-off
reply (finish_reason "length"), an empty one, a parse function that raises, JSON that does not parse or does not
match the schema, and a quoted string that is not literally in the given text all raise solvi.llm.InvalidOutput
with the reason — never a repaired or guessed value. A server that does not answer after the retries, or refuses the
input (HTTP 400 / 413 / 422), raises Unanswered; a wrong key, model or URL raises solvi.llm.LLMError. In a catalog
any of these makes the part fail: the fact is missing, and the questions that need it abstain with the cause.
Quotes. quotes=["rows", "items.*.source"] names the strings of the value that must be copied from text: each is
looked up as written, as whole words and numbers (the rule of Claim evidence — "3" is not found in "30"). Ask the model
to copy table rows as the text writes them and parse them in code: a wrong number is then not in the text, where a
number copied into a field of its own may stand elsewhere in the text and pass. The quotes become the output's evidence
(Claim.evidence), so inside a System they are located again in the given text, recorded with their offsets and
checked on replay.
The trace. generate and sample return a Generated — a Claim whose extra["generated"] holds, per output, the
model id, the request's fingerprint (sha256 of the request body), temperature, seed, finish reason, tokens and, for a
structured or parsed reply, the reply's text. Returned from a catalog part, the value is the fact and the record keeps
that detail; writer.part(...) builds such a part with the model attached (model=), provenance proposed. An LLM's
output is not reproducible bit for bit, so replay does not call it again (replay="rerun" asks it to): it checks
the recorded output instead — the recorded reply read again through the same parse and schema must give the recorded
value, and the quotes must be in the recorded text. The API key is never recorded.
Not done here: no retries on an invalid reply (that is solvi.refine's loop: the reasons go back to the model), no response caching (put a caching proxy in front of the server), no streaming, no tool calls.
Unanswered ¶
Bases: RuntimeError
The server did not answer after the retries (network, timeout, 408 / 409 / 429 / 5xx), or refused this request's input (HTTP 400 / 413 / 422): an escalation, not a configuration error. The message has the endpoint, never the key.
Generated
dataclass
¶
Generated(value: Any, evidence: list = list(), confidence: float = 1.0, source: str | None = None, extra: dict = dict())
Bases: Claim
What a generator returns: a Claim whose value is the output (a text, a validated schema value, or a list of K
outputs with None where one failed), whose evidence is the quoted strings, and whose extra["generated"] is the
record of the request(s) — one dict, or a list for K outputs. Returned from a catalog part, the fact is the value.
Schema ¶
What a structured reply must be: a pydantic model (or any type pydantic validates: list[Row], dict[str, int], ...) or a JSON schema (a dict; the keywords listed in _SUPPORTED, else ValueError).
check ¶
Parsed JSON → the validated value; InvalidOutput with the first reason it is not.
Generator ¶
Generator(client, *, max_tokens=1024, temperature=0.0, seed=None, response_format='prompt', workers=4)
A generating model over an OpenAI-compatible server. Build it with generator(...), or share a decider's client
with Generator(llm_model.scorer) / Generator.of(llm_model). See the module docs.
of
classmethod
¶
A Generator on the client of a decider from solvi.llm.llm(...) (or an LLMScorer): the same endpoint, key, headers, extra_body, retries and timeout.
body ¶
The request body: extra_body, then the model, the messages, temperature and max_tokens (and seed when set, and the reply format when response_format asks for one).
read ¶
A reply's text → (value, quotes): parse (text → value; raising or returning None rejects the reply), then
the schema (JSON parsed from the text, or from what parse returned when it is a string), then the quotes.
InvalidOutput with the reason when any step rejects it. Replay reads a recorded reply through this again.
generate ¶
generate(messages, *, schema=None, parse=None, text=None, quotes=None, source=None, temperature=None, seed=None, max_tokens=None)
One reply → a Generated whose value is the text, parse(text), or the JSON validated by schema (a pydantic
model or type, or a JSON schema dict). text / quotes: the strings at the quotes paths of the value must be in
text as written; they become the evidence (source: the given fact holding the text, for a catalog part).
Raises InvalidOutput, Unanswered or LLMError (module docs) — an invalid reply is never repaired.
sample ¶
sample(messages, k=3, *, temperature=0.8, schema=None, parse=None, text=None, quotes=None, source=None, max_tokens=None, first_greedy=True)
k replies to the same messages → a Generated whose value is the list of k outputs, None where a reply failed
(its record says why). The first at the generator's own temperature (first_greedy=True: the reply generate gives),
the others at temperature with seeds 1 .. k-1 (so a rerun asks the same requests). Raises only when every reply
failed (the first failure's exception, with how many failed).
part ¶
part(name, prompt, *, schema=None, parse=None, k=1, temperature=0.8, text=None, quotes=None, max_tokens=None, replay='trust')
A catalog part that generates: cat.fn(writer.part("sql", prompt)). prompt: a function of facts by name (its
parameters are the part's inputs) → the messages or one string. k > 1: sample (the fact is the list of k
outputs). text: the name of the given fact the quotes must be in (one of prompt's parameters). replay: "trust"
(default: the model is not called again; the recorded output is checked) or "rerun" (call it again and compare —
only for a server that answers the same request the same way).
proposer ¶
proposer(messages, *, schema=None, parse=None, text=None, quotes=None, first=None, template=FEEDBACK)
A proposer for solvi.refine.refine: (state, rounds) → a Generated. The first round asks messages (or
messages(state) when it is a function); each later round asks the same messages followed, per earlier
round, by its proposal as the assistant's turn (left out when empty) and its feedback as the user's turn — a text
as it is, a list of reasons through template ("{reasons}": one "- reason" per line). first: another
Generator for the first round (a stronger setting, say).
GenerationPart ¶
The model behind a part made by Generator.part: the generator, the prompt function's code, the schema and the
settings — its fingerprint covers all of them, so a changed prompt or schema shows in a replay as a changed model.
check_record re-reads a recorded output on replay (solvi.runtime calls it when the model is not re-run).
check_record ¶
A recorded output, without calling the model → the reasons it is not what this part accepts: each recorded reply read again through parse and the schema must give the recorded value; a failed candidate must carry its error. (The quotes are checked in the text by the replay's grounding check, from the recorded evidence.)
json_errors ¶
Where a JSON value breaks a JSON schema (the keywords _SUPPORTED) → the first reason, or None.
quoted ¶
The strings at the quotes paths of a value, each checked to be in text as written (whole words and numbers)
→ the list of them; InvalidOutput naming the first one that is not there.
generator ¶
generator(base_url, model, api_key=None, *, max_tokens=1024, temperature=0.0, seed=None, response_format='prompt', timeout=120.0, retries=2, backoff=1.0, headers=None, extra_body=None, workers=4, opener=None, sleep=None)
A Generator over an OpenAI-compatible chat-completions server (see the module docs). The connection settings are
solvi.llm's (llm(...) takes the same: base_url, api_key — sent as a Bearer token, never recorded — timeout, retries
and backoff for network errors, timeouts and 408 / 409 / 429 / 5xx, headers, extra_body — refused when it sets a field
solvi sets: model, messages, temperature, max_tokens, seed, response_format —, opener for tests and proxies).
max_tokens: per reply (reasoning tokens count against it on reasoning models). temperature: 0 by default; sample
sets its own for the extra replies. seed: sent only when set. response_format: "prompt" (default: the contract is in
your prompt, the request carries none), "json_object" or "json_schema" (the schema's JSON schema is sent) — for a
structured reply only; whatever the server enforces, the reply is validated here.
several ¶
One reply from each of several generators (several models) to the same messages → a Generated whose value is the list of outputs, None where one failed — the candidates for solvi.agree. Raises only when every one failed.