A model that writes: generation, agreement and the re-ask loop¶
solvi.llm asks a model closed questions. When the model's output is something it writes — a SQL query, a plan, a JSON
extraction of a table — three pieces put solvi around it: solvi.generate makes the call and records it,
solvi.agree compares several candidates under a key you give, and solvi.refine runs propose → check → re-ask with
the reasons → escalate. The model proposes; the checks decide; every round is a recorded, replayable decision. None of
them makes the model better at writing: they decide what is returned without a person, and say why the rest is not.
A runnable example, with a stand-in for the model (three queries per round, other ones once the checks have spoken):
import sqlite3
from solvi import Answer, Catalog, Fail, Question, System
from solvi.agree import agree
from solvi.refine import refine
db = sqlite3.connect(":memory:")
db.executescript("CREATE TABLE orders(id, amount, status);"
"INSERT INTO orders VALUES (1, 30, 'paid'), (2, 70, 'paid'), (3, 20, 'open');")
def rows(sql): # the key: the rows a query returns, in any order
return tuple(sorted(db.execute(sql).fetchall()))
cat = Catalog()
@cat.fn
def candidates(drafts, feedback): # a stand-in for writer.part("candidates", prompt, k=3): 3 queries,
return drafts[min(len(feedback), 1)] # other ones once the checks have said something
agree(cat, "sql", "candidates", key=rows) # facts: sql, sql_agreement, sql_tally
@cat.check(hard=True, then={"answer": "no"})
def most_agree(sql_agreement) -> bool:
return sql_agreement >= 2 / 3 or Fail(f"only {sql_agreement:.0%} of the queries return the same rows")
@cat.rule("answer")
def answer(sql, most_agree) -> bool:
return True
system = System(cat, [Question("answer", "Return the query without a person?", Answer.yes_no(),
requires=["most_agree"])])
drafts = [["SELECT sum(amount) FROM orders", "SELECT sum(amount) FROM orders WHERE status = 'paid'", "SELECT 1"],
["SELECT sum(amount) FROM orders WHERE status = 'paid'", "SELECT 100", "SELECT 120"]]
run = refine(system, {"drafts": drafts}, "answer", rounds=3)
print(run.accepted, run.response.values["sql"], run.response.values["sql_agreement"])
for r in run.rounds:
print(r.index, r.accepted, r.reasons)
print(run.rounds[0].response["answer"].why)
print(run.replay(system)["ok"])
True SELECT sum(amount) FROM orders WHERE status = 'paid' 0.6666666666666666
0 False ['only 33% of the queries return the same rows']
1 True []
hard check most_agree is false: only 33% of the queries return the same rows
True
With a model the stand-in becomes one line, cat.fn(writer.part("candidates", prompt, k=3, parse=sql_block)), where
prompt(question, schema, feedback) builds the messages from the facts it names, and the rest stays as it is.
A check that says why: Fail¶
A check returns Fail("Harold is busy on Monday 13:30 - 15:30") instead of False when it can say what is wrong
(several reasons: Fail(*reasons)). It is False wherever a bool is read — rules, hard checks, plain Python (not
Fail(...) is True) — and -> bool checks keep their type. The reasons are recorded with the check
(record.extra["reasons"]), added to the answer's reason when a hard check decides ("hard check nobody_busy is false:
Harold is busy …"), shown on the check's line of the audit, and compared on replay (a check that now gives other reasons
is a mismatch). A check that returns plain False keeps working: its reason is its docstring's first line, else
"solvi.refine.failed_checks(res, question) lists the checks that are False with their reasons.
Generation: solvi.generate¶
from solvi.generate import generator
writer = generator("https://openrouter.ai/api/v1", "openai/gpt-oss-120b", api_key=KEY, max_tokens=3000,
extra_body={"reasoning": {"effort": "low"}})
g = writer.generate(messages) # g.value: the reply's text
g = writer.generate(messages, schema=Plan) # a pydantic model (or a JSON schema dict): validated
g = writer.generate(messages, parse=sql_block) # your parser: raising or None rejects the reply
g = writer.generate(messages, schema=Table, text=doc, quotes=["rows"]) # every row literally in doc
g = writer.sample(messages, k=3, temperature=0.8) # greedy first, then seeds 1, 2: a list, None if one failed
cat.fn(writer.part("plan", prompt, schema=Plan)) # a catalog part; k=3 for samples, text=/quotes= as above
The connection is solvi.llm's: generator(...) takes the same endpoint, key, headers, extra_body, retries, backoff,
timeout and opener, and Generator.of(decider) shares the client of a decider made with llm(...). A reply is
accepted only whole: a refusal, a cut-off reply (finish_reason "length" — reasoning tokens count against
max_tokens), an empty one, a parser that raises, JSON that does not parse or match the schema, and a quoted string
that is not in the text raise solvi.llm.InvalidOutput with the reason; it is never repaired. A server that does not
answer after the retries, or refuses the input (HTTP 400 / 413 / 422), raises solvi.generate.Unanswered; a wrong key,
model or URL raises LLMError. In a catalog each of these makes the part fail, and the questions that need it abstain
with the cause (… caused by sql: InvalidOutput: the reply was cut off (max_tokens)). A JSON schema is checked for its
common keywords (type, properties, required, additionalProperties, items, enum, const, bounds, lengths, anyOf); one with
a keyword it does not check ($ref, pattern, format …) is refused when it is given — pass a pydantic model then.
response_format="json_schema" also sends the schema to the server; the reply is validated here either way.
Quotes. quotes=["rows", "items.*.source"] names the strings of the value that must be copied from text. Each is
looked up as written, as whole words and numbers ("3" is not found in "30"). Ask the model to copy a table's rows as the
text writes them and parse them in code: a wrong number is then not in the text. A number copied into a field of its
own can stand elsewhere in the text and pass (below). The quotes become the output's evidence, so in a catalog they are
located again in the given fact text= names, recorded with their offsets and checked on replay.
The record. generate returns a Generated — a Claim whose extra["generated"] holds the model id, the
fingerprint of the request body, temperature, seed, finish reason, tokens and, for a structured or parsed reply, the
reply's text (one per output for sample). Returned from a part, the value is the fact and the record keeps the rest.
writer.part(...) attaches the model (GenerationPart, provenance proposed); its fingerprint covers the generator's
settings, the prompt function's code, the parser and the schema, so a changed prompt shows on replay as a changed model.
Replay does not call the model again (part(..., replay="rerun") does, for a server that answers the same request the
same way): it reads the recorded reply again through the parser and the schema and names a recorded value that does not
follow from it. The key is never recorded.
Agreement of candidates: solvi.agree¶
from solvi.agree import agree, consensus
agree(cat, "sql", "candidates", key=row_digest, prefer=returns_rows) # facts: sql, sql_agreement, sql_tally
consensus(queries, key=row_digest) # the same outside a catalog: {"index", "value", "share", "groups", ...}
Vote combines decisions over the same closed options; generated outputs have none. key(candidate) says what makes
two candidates the same — the digest of the rows a query returns, a normalized plan, a parsed number — and may read other
facts by name (def row_digest(sql, db_path)). The largest group wins (a tie: the group whose first candidate comes
first); sql is its first candidate, sql_agreement the share of all K candidates in it — a plain number fact a rule,
a head or a guarantee reads like any other signal. A candidate that is None (its generation failed), whose key raises or
is None, does not vote and still counts in K. prefer: when any candidate passes it, only those vote (rows before an
empty result). When nothing votes, sql is missing — its error lists each candidate's reason — and the share is 0.0.
sql_tally records per candidate its key or why it has none, the groups, the choice and the share; replay recomputes it
(keep the key deterministic, or cache what it computes). Candidates from several models: solvi.generate.several(
[writer_a, writer_b], messages).
The loop: solvi.refine¶
from solvi.refine import refine
run = refine(system, {"problem": text}, "accept", propose=writer.proposer(messages, schema=Plan), into="plan",
rounds=3, accept="yes", feedback=lambda r: my_wording(r.reasons))
run.accepted, run.proposal, run.escalation, run.rounds # stored rounds: r.stored_id
run.replay(system) # {"ok", "mismatches": [(round, what, why)], ...}
One round: propose(state, earlier_rounds) returns a proposal — a value, or a Generated whose record is kept in the
round — which is given to the System under into; the System is asked. accept: "checks" (default: every hard check
that governs the question was evaluated and passed, whatever the answer), an answer or a list of answers, or a function
of the Response. A round not accepted has reasons — the reasons of the failed hard checks that govern the question, in
catalog order; when the question could not be decided at all, the errors of the parts that failed ("spec: ValueError:
no slot in the plan"). feedback(round) turns them into what the proposer is told (default: the list of reasons).
The loop stops at the first accepted round or after rounds, and escalates: run.escalation is "not accepted after 3
round(s): Generator.proposer(messages, schema=…, first=…) builds the proposer: the first round asks
the messages (or first, another generator — a stronger setting for the first try); each later round appends, per
earlier round, its reply as the assistant's turn and its feedback as the user's.
Without a proposer the System generates itself: each round gives the earlier rounds' feedback as the fact
feedback_into (default "feedback", a list of reasons) and the generating part reads it — the example above.
history= continues an earlier refinement. run.to_dict() / Refinement.from_dict(d, catalog=cat) store and load it;
run.replay(system) replays every round's trace and checks that the loop did what its record says: each round's
acceptance, failed checks and causes follow from its response, the feedback is what the feedback function gives (pass a
custom one again; accept too when it was a function), each round's input holds its proposal and the feedback before
it, nothing ran after an accepted round, and the escalation matches.
What to expect. Re-asking with the violated constraints quoted lowers the share of wrong answers among those given,
and more of the hard cases go to a person instead; the model often trades one violation for another, so a re-ask is
not a fix. Typed facts extracted with schema= and quotes= are accepted or rejected by their quotes — a quote check
does not see a row left out. Agreement of several samples is not correctness: samples can agree on one reading of the
question that is not the intended one.
Not done here. No search: a loop re-asks one proposer, it does not enumerate alternatives or keep the best of two
valid ones — that is solvi.search (below). No promise that
re-asks converge. The
share of agreement is a signal; calibrate it on labelled examples before you trust a threshold. No streaming, no tool
calls, no caching of replies (put a caching proxy in front of the server).
Search over alternatives: solvi.search¶
When the candidates can be enumerated — the slots of a week, the orders of a few cities, the friends to meet — a search through the System's own checks beats asking a model to propose: run each candidate through the checks, keep the accepted ones, take the best.
from solvi import Answer, Catalog, Question, System
from solvi.refine import Fail
from solvi.search import Tree, search
cat = Catalog()
FLIGHTS = {("Oslo", "Rome"), ("Rome", "Paris"), ("Paris", "Oslo"), ("Rome", "Vienna")}
@cat.fn
def cities(problem: str) -> list: # the problem read into facts: computed once for the whole search
return problem.split(", ")
@cat.check(hard=True, then={"ok": "no"})
def direct_flights(order: list) -> bool: # false on a prefix → false on every order that starts with it
bad = [f"no flight {a} - {b}" for a, b in zip(order, order[1:]) if (a, b) not in FLIGHTS and (b, a) not in FLIGHTS]
return Fail(*bad) if bad else True
@cat.check(hard=True, then={"ok": "no"})
def every_city(cities: list, order: list) -> bool:
return sorted(order) == sorted(cities)
@cat.rule("ok")
def ok(direct_flights, every_city) -> bool:
return True
system = System(cat, [Question("ok", "A valid trip?", Answer.yes_no(), requires=["direct_flights", "every_city"])])
def orders(facts): # the space, read from the facts: one more city per step
cs = facts["cities"]
return Tree([], lambda o: [o + [c] for c in cs if c not in o], complete=lambda o: len(o) == len(cs))
run = search(system, {"problem": "Oslo, Rome, Paris, Vienna"}, "ok", orders, into="order",
prune=["direct_flights"], keep=2)
print(run)
# search ok: 28 asked, 2 accepted
# best: ['Oslo', 'Paris', 'Rome', 'Vienna']
# exact: every candidate was asked or cut — pruned by direct_flights (10); the first accepted in the space's order
# (and not the only one), given that the prune checks direct_flights stay false below a node
# rejected by: direct_flights (4)
# computed once: cities
run.response["ok"].answer, run.kept, run.replay(system)["ok"]
search(system, state, question, space, *, into=, objective=, maximize=True, prune=(), keep=1, budget=10_000,
accept="checks", store=True, hold=True):
- The space: a list or any iterable of candidates (given as the fact
into); a dict{fact: [values]}(every combination, the first fact outermost, given as those facts); aTree(root, children, complete=, bound=)walked depth first; or a function of the facts computed fromstatethat returns one of these — the space read from the problem. - Accepted: as in
refine(accept="checks": every hard check governing the question passed), and the question did not abstain — a candidate the System could not decide is never chosen.run.rejectedcounts the rejections by deciding check. - The best: with
objective(a function of the candidate, or the name of a fact the question computes) thekeepbest accepted; without one the firstkeepin the space's order, and the search stops there (keep=2says whether the first is the only one). - Cuts:
prunenames hard checks that, false on a partial node, stay false on every node below it; such a node is not expanded. A Tree'sbound(node)is the best objective any candidate below can reach; a node that cannot beat what is kept is not expanded.budgetcaps the asks. - When it is exact:
run.exactis True when the search ended by itself — every candidate was asked or cut — andrun.why_exactnames what that rests on: that the prune checks are monotone and the bound optimistic. Neither is checked by solvi; a wrong promise can cut the best candidate. When the budget stops it,exactis False and the best is the best of what was asked; when nothing is accepted,run.escalationsays why (and arefinewith a proposer can take over a space too big to search). - Facts computed once: parts that do not read the candidate (the problem read into typed facts — by rules or by a
model) run once and are held for every candidate (
run.held); a model reading the problem is called once, not per candidate. The searched asks are not stored or counted; the winner is asked again in full with the System (run.response, stored when the System stores), and if that full ask does not accept it the search escalates instead.run.to_dict()/SearchRun.from_dict(d, catalog=cat);run.replay(system)replays the winner's trace and checks it is accepted and its objective recomputes (pass a function objective again).
Where the candidates can be enumerated, prefer a search to a model's proposals: the checks decide every candidate, and every winner replays. What stays problem-specific is the space (a few lines per kind of problem) and, for an order, a walk of partial orders so the checks can judge a prefix. The price is speed: every candidate is a full ask, so a search runs far more asks than a hand-written search runs steps.
Not done here: no proposals by a model, no bisection over numbers (res.counterfactual does that), no parallel
asks, no proof of the prune and bound promises. Each candidate is a full ask with the trace hashed (see the README's
Speed table, benchmarks/ask_speed.py), so a space of millions is for code, not for this search.