Skip to content

solvi.generate

Generation by an LLM — a text, or JSON validated against a pydantic model or a JSON schema, with quotes checked in a text — on solvi.llm's client settings, recorded in the trace as a model's output.

Generation by an LLM: a text, or JSON checked against a pydantic model or a JSON schema — through the same OpenAI-compatible client settings as solvi.llm, recorded in the trace as a model's output.

from solvi.generate import generator
writer = generator("http://127.0.0.1:8080/v1", "qwen2.5-7b-instruct", extra_body={"reasoning": {"effort": "low"}})
g = writer.generate("Write one SQLite query that ...")            # g.value: the reply's text
g = writer.generate(messages, schema=Plan)                         # g.value: a Plan (pydantic), validated
g = writer.generate(messages, schema=Table, text=doc, quotes=["rows"])   # every row literally in doc
g = writer.sample(messages, k=3, temperature=0.8)                  # g.value: 3 outputs (None where one failed)
cat.fn(writer.part("sql", prompt, k=3))                            # a catalog part: the outputs recorded, not re-run

solvi.llm asks closed questions (options, yes / no, a span); this module is for the model that writes something — a query, a plan, a JSON extraction of a table — which solvi's checks then judge. The model proposes; nothing here decides.

What is checked, and what escalates. A reply is accepted only when it is complete and well-formed: a refusal, a cut-off reply (finish_reason "length"), an empty one, a parse function that raises, JSON that does not parse or does not match the schema, and a quoted string that is not literally in the given text all raise solvi.llm.InvalidOutput with the reason — never a repaired or guessed value. A server that does not answer after the retries, or refuses the input (HTTP 400 / 413 / 422), raises Unanswered; a wrong key, model or URL raises solvi.llm.LLMError. In a catalog any of these makes the part fail: the fact is missing, and the questions that need it abstain with the cause.

Quotes. quotes=["rows", "items.*.source"] names the strings of the value that must be copied from text: each is looked up as written, as whole words and numbers (the rule of Claim evidence — "3" is not found in "30"). Ask the model to copy table rows as the text writes them and parse them in code: a wrong number is then not in the text, where a number copied into a field of its own may stand elsewhere in the text and pass. The quotes become the output's evidence (Claim.evidence), so inside a System they are located again in the given text, recorded with their offsets and checked on replay.

The trace. generate and sample return a Generated — a Claim whose extra["generated"] holds, per output, the model id, the request's fingerprint (sha256 of the request body), temperature, seed, finish reason, tokens and, for a structured or parsed reply, the reply's text. Returned from a catalog part, the value is the fact and the record keeps that detail; writer.part(...) builds such a part with the model attached (model=), provenance proposed. An LLM's output is not reproducible bit for bit, so replay does not call it again (replay="rerun" asks it to): it checks the recorded output instead — the recorded reply read again through the same parse and schema must give the recorded value, and the quotes must be in the recorded text. The API key is never recorded.

Not done here: no retries on an invalid reply (that is solvi.refine's loop: the reasons go back to the model), no response caching (put a caching proxy in front of the server), no streaming, no tool calls.

Unanswered

Bases: RuntimeError

The server did not answer after the retries (network, timeout, 408 / 409 / 429 / 5xx), or refused this request's input (HTTP 400 / 413 / 422): an escalation, not a configuration error. The message has the endpoint, never the key.

Generated dataclass

Generated(value: Any, evidence: list = list(), confidence: float = 1.0, source: str | None = None, extra: dict = dict())

Bases: Claim

What a generator returns: a Claim whose value is the output (a text, a validated schema value, or a list of K outputs with None where one failed), whose evidence is the quoted strings, and whose extra["generated"] is the record of the request(s) — one dict, or a list for K outputs. Returned from a catalog part, the fact is the value.

meta property

meta

The record of the request: {"model", "request", "temperature", "seed", "finish", "usage", ...} (a list for K).

text property

text

The reply's text (a list for K outputs).

Schema

Schema(schema)

What a structured reply must be: a pydantic model (or any type pydantic validates: list[Row], dict[str, int], ...) or a JSON schema (a dict; the keywords listed in _SUPPORTED, else ValueError).

check

check(data)

Parsed JSON → the validated value; InvalidOutput with the first reason it is not.

Generator

Generator(client, *, max_tokens=1024, temperature=0.0, seed=None, response_format='prompt', workers=4)

A generating model over an OpenAI-compatible server. Build it with generator(...), or share a decider's client with Generator(llm_model.scorer) / Generator.of(llm_model). See the module docs.

usage property

usage

Tokens used by this client (shared with a decider built on the same client).

of classmethod

of(model, **kw)

A Generator on the client of a decider from solvi.llm.llm(...) (or an LLMScorer): the same endpoint, key, headers, extra_body, retries and timeout.

body

body(messages, schema=None, temperature=None, seed=None, max_tokens=None)

The request body: extra_body, then the model, the messages, temperature and max_tokens (and seed when set, and the reply format when response_format asks for one).

read

read(content, schema=None, parse=None, text=None, quotes=None)

A reply's text → (value, quotes): parse (text → value; raising or returning None rejects the reply), then the schema (JSON parsed from the text, or from what parse returned when it is a string), then the quotes. InvalidOutput with the reason when any step rejects it. Replay reads a recorded reply through this again.

generate

generate(messages, *, schema=None, parse=None, text=None, quotes=None, source=None, temperature=None, seed=None, max_tokens=None)

One reply → a Generated whose value is the text, parse(text), or the JSON validated by schema (a pydantic model or type, or a JSON schema dict). text / quotes: the strings at the quotes paths of the value must be in text as written; they become the evidence (source: the given fact holding the text, for a catalog part). Raises InvalidOutput, Unanswered or LLMError (module docs) — an invalid reply is never repaired.

sample

sample(messages, k=3, *, temperature=0.8, schema=None, parse=None, text=None, quotes=None, source=None, max_tokens=None, first_greedy=True)

k replies to the same messages → a Generated whose value is the list of k outputs, None where a reply failed (its record says why). The first at the generator's own temperature (first_greedy=True: the reply generate gives), the others at temperature with seeds 1 .. k-1 (so a rerun asks the same requests). Raises only when every reply failed (the first failure's exception, with how many failed).

part

part(name, prompt, *, schema=None, parse=None, k=1, temperature=0.8, text=None, quotes=None, max_tokens=None, replay='trust')

A catalog part that generates: cat.fn(writer.part("sql", prompt)). prompt: a function of facts by name (its parameters are the part's inputs) → the messages or one string. k > 1: sample (the fact is the list of k outputs). text: the name of the given fact the quotes must be in (one of prompt's parameters). replay: "trust" (default: the model is not called again; the recorded output is checked) or "rerun" (call it again and compare — only for a server that answers the same request the same way).

proposer

proposer(messages, *, schema=None, parse=None, text=None, quotes=None, first=None, template=FEEDBACK)

A proposer for solvi.refine.refine: (state, rounds) → a Generated. The first round asks messages (or messages(state) when it is a function); each later round asks the same messages followed, per earlier round, by its proposal as the assistant's turn (left out when empty) and its feedback as the user's turn — a text as it is, a list of reasons through template ("{reasons}": one "- reason" per line). first: another Generator for the first round (a stronger setting, say).

GenerationPart

GenerationPart(gen, prompt, schema, parse, k, temperature, quotes, max_tokens, rerun)

The model behind a part made by Generator.part: the generator, the prompt function's code, the schema and the settings — its fingerprint covers all of them, so a changed prompt or schema shows in a replay as a changed model. check_record re-reads a recorded output on replay (solvi.runtime calls it when the model is not re-run).

check_record

check_record(r)

A recorded output, without calling the model → the reasons it is not what this part accepts: each recorded reply read again through parse and the schema must give the recorded value; a failed candidate must carry its error. (The quotes are checked in the text by the replay's grounding check, from the recorded evidence.)

json_errors

json_errors(s, v, path='$')

Where a JSON value breaks a JSON schema (the keywords _SUPPORTED) → the first reason, or None.

quoted

quoted(value, quotes, text)

The strings at the quotes paths of a value, each checked to be in text as written (whole words and numbers) → the list of them; InvalidOutput naming the first one that is not there.

generator

generator(base_url, model, api_key=None, *, max_tokens=1024, temperature=0.0, seed=None, response_format='prompt', timeout=120.0, retries=2, backoff=1.0, headers=None, extra_body=None, workers=4, opener=None, sleep=None)

A Generator over an OpenAI-compatible chat-completions server (see the module docs). The connection settings are solvi.llm's (llm(...) takes the same: base_url, api_key — sent as a Bearer token, never recorded — timeout, retries and backoff for network errors, timeouts and 408 / 409 / 429 / 5xx, headers, extra_body — refused when it sets a field solvi sets: model, messages, temperature, max_tokens, seed, response_format —, opener for tests and proxies). max_tokens: per reply (reasoning tokens count against it on reasoning models). temperature: 0 by default; sample sets its own for the extra replies. seed: sent only when set. response_format: "prompt" (default: the contract is in your prompt, the request carries none), "json_object" or "json_schema" (the schema's JSON schema is sent) — for a structured reply only; whatever the server enforces, the reply is validated here.

several

several(generators, messages, **kw)

One reply from each of several generators (several models) to the same messages → a Generated whose value is the list of outputs, None where one failed — the candidates for solvi.agree. Raises only when every one failed.