Skip to content

A model that writes: generation, agreement and the re-ask loop

solvi.llm asks a model closed questions. When the model's output is something it writes — a SQL query, a plan, a JSON extraction of a table — three pieces put solvi around it: solvi.generate makes the call and records it, solvi.agree compares several candidates under a key you give, and solvi.refine runs propose → check → re-ask with the reasons → escalate. The model proposes; the checks decide; every round is a recorded, replayable decision. None of them makes the model better at writing: they decide what is returned without a person, and say why the rest is not.

A runnable example, with a stand-in for the model (three queries per round, other ones once the checks have spoken):

import sqlite3
from solvi import Answer, Catalog, Fail, Question, System
from solvi.agree import agree
from solvi.refine import refine

db = sqlite3.connect(":memory:")
db.executescript("CREATE TABLE orders(id, amount, status);"
                 "INSERT INTO orders VALUES (1, 30, 'paid'), (2, 70, 'paid'), (3, 20, 'open');")

def rows(sql):                                   # the key: the rows a query returns, in any order
    return tuple(sorted(db.execute(sql).fetchall()))

cat = Catalog()

@cat.fn
def candidates(drafts, feedback):                # a stand-in for writer.part("candidates", prompt, k=3): 3 queries,
    return drafts[min(len(feedback), 1)]         # other ones once the checks have said something

agree(cat, "sql", "candidates", key=rows)        # facts: sql, sql_agreement, sql_tally

@cat.check(hard=True, then={"answer": "no"})
def most_agree(sql_agreement) -> bool:
    return sql_agreement >= 2 / 3 or Fail(f"only {sql_agreement:.0%} of the queries return the same rows")

@cat.rule("answer")
def answer(sql, most_agree) -> bool:
    return True

system = System(cat, [Question("answer", "Return the query without a person?", Answer.yes_no(),
                               requires=["most_agree"])])
drafts = [["SELECT sum(amount) FROM orders", "SELECT sum(amount) FROM orders WHERE status = 'paid'", "SELECT 1"],
          ["SELECT sum(amount) FROM orders WHERE status = 'paid'", "SELECT 100", "SELECT 120"]]
run = refine(system, {"drafts": drafts}, "answer", rounds=3)
print(run.accepted, run.response.values["sql"], run.response.values["sql_agreement"])
for r in run.rounds:
    print(r.index, r.accepted, r.reasons)
print(run.rounds[0].response["answer"].why)
print(run.replay(system)["ok"])
True SELECT sum(amount) FROM orders WHERE status = 'paid' 0.6666666666666666
0 False ['only 33% of the queries return the same rows']
1 True []
hard check most_agree is false: only 33% of the queries return the same rows
True

With a model the stand-in becomes one line, cat.fn(writer.part("candidates", prompt, k=3, parse=sql_block)), where prompt(question, schema, feedback) builds the messages from the facts it names, and the rest stays as it is.

A check that says why: Fail

A check returns Fail("Harold is busy on Monday 13:30 - 15:30") instead of False when it can say what is wrong (several reasons: Fail(*reasons)). It is False wherever a bool is read — rules, hard checks, plain Python (not Fail(...) is True) — and -> bool checks keep their type. The reasons are recorded with the check (record.extra["reasons"]), added to the answer's reason when a hard check decides ("hard check nobody_busy is false: Harold is busy …"), shown on the check's line of the audit, and compared on replay (a check that now gives other reasons is a mismatch). A check that returns plain False keeps working: its reason is its docstring's first line, else " is false". solvi.refine.failed_checks(res, question) lists the checks that are False with their reasons.

Generation: solvi.generate

from solvi.generate import generator
writer = generator("https://openrouter.ai/api/v1", "openai/gpt-oss-120b", api_key=KEY, max_tokens=3000,
                   extra_body={"reasoning": {"effort": "low"}})
g = writer.generate(messages)                                    # g.value: the reply's text
g = writer.generate(messages, schema=Plan)                       # a pydantic model (or a JSON schema dict): validated
g = writer.generate(messages, parse=sql_block)                   # your parser: raising or None rejects the reply
g = writer.generate(messages, schema=Table, text=doc, quotes=["rows"])   # every row literally in doc
g = writer.sample(messages, k=3, temperature=0.8)                # greedy first, then seeds 1, 2: a list, None if one failed
cat.fn(writer.part("plan", prompt, schema=Plan))                 # a catalog part; k=3 for samples, text=/quotes= as above

The connection is solvi.llm's: generator(...) takes the same endpoint, key, headers, extra_body, retries, backoff, timeout and opener, and Generator.of(decider) shares the client of a decider made with llm(...). A reply is accepted only whole: a refusal, a cut-off reply (finish_reason "length" — reasoning tokens count against max_tokens), an empty one, a parser that raises, JSON that does not parse or match the schema, and a quoted string that is not in the text raise solvi.llm.InvalidOutput with the reason; it is never repaired. A server that does not answer after the retries, or refuses the input (HTTP 400 / 413 / 422), raises solvi.generate.Unanswered; a wrong key, model or URL raises LLMError. In a catalog each of these makes the part fail, and the questions that need it abstain with the cause (… caused by sql: InvalidOutput: the reply was cut off (max_tokens)). A JSON schema is checked for its common keywords (type, properties, required, additionalProperties, items, enum, const, bounds, lengths, anyOf); one with a keyword it does not check ($ref, pattern, format …) is refused when it is given — pass a pydantic model then. response_format="json_schema" also sends the schema to the server; the reply is validated here either way.

Quotes. quotes=["rows", "items.*.source"] names the strings of the value that must be copied from text. Each is looked up as written, as whole words and numbers ("3" is not found in "30"). Ask the model to copy a table's rows as the text writes them and parse them in code: a wrong number is then not in the text. A number copied into a field of its own can stand elsewhere in the text and pass (below). The quotes become the output's evidence, so in a catalog they are located again in the given fact text= names, recorded with their offsets and checked on replay.

The record. generate returns a Generated — a Claim whose extra["generated"] holds the model id, the fingerprint of the request body, temperature, seed, finish reason, tokens and, for a structured or parsed reply, the reply's text (one per output for sample). Returned from a part, the value is the fact and the record keeps the rest. writer.part(...) attaches the model (GenerationPart, provenance proposed); its fingerprint covers the generator's settings, the prompt function's code, the parser and the schema, so a changed prompt shows on replay as a changed model. Replay does not call the model again (part(..., replay="rerun") does, for a server that answers the same request the same way): it reads the recorded reply again through the parser and the schema and names a recorded value that does not follow from it. The key is never recorded.

Agreement of candidates: solvi.agree

from solvi.agree import agree, consensus
agree(cat, "sql", "candidates", key=row_digest, prefer=returns_rows)    # facts: sql, sql_agreement, sql_tally
consensus(queries, key=row_digest)              # the same outside a catalog: {"index", "value", "share", "groups", ...}

Vote combines decisions over the same closed options; generated outputs have none. key(candidate) says what makes two candidates the same — the digest of the rows a query returns, a normalized plan, a parsed number — and may read other facts by name (def row_digest(sql, db_path)). The largest group wins (a tie: the group whose first candidate comes first); sql is its first candidate, sql_agreement the share of all K candidates in it — a plain number fact a rule, a head or a guarantee reads like any other signal. A candidate that is None (its generation failed), whose key raises or is None, does not vote and still counts in K. prefer: when any candidate passes it, only those vote (rows before an empty result). When nothing votes, sql is missing — its error lists each candidate's reason — and the share is 0.0. sql_tally records per candidate its key or why it has none, the groups, the choice and the share; replay recomputes it (keep the key deterministic, or cache what it computes). Candidates from several models: solvi.generate.several( [writer_a, writer_b], messages).

The loop: solvi.refine

from solvi.refine import refine
run = refine(system, {"problem": text}, "accept", propose=writer.proposer(messages, schema=Plan), into="plan",
             rounds=3, accept="yes", feedback=lambda r: my_wording(r.reasons))
run.accepted, run.proposal, run.escalation, run.rounds        # stored rounds: r.stored_id
run.replay(system)                                             # {"ok", "mismatches": [(round, what, why)], ...}

One round: propose(state, earlier_rounds) returns a proposal — a value, or a Generated whose record is kept in the round — which is given to the System under into; the System is asked. accept: "checks" (default: every hard check that governs the question was evaluated and passed, whatever the answer), an answer or a list of answers, or a function of the Response. A round not accepted has reasons — the reasons of the failed hard checks that govern the question, in catalog order; when the question could not be decided at all, the errors of the parts that failed ("spec: ValueError: no slot in the plan"). feedback(round) turns them into what the proposer is told (default: the list of reasons). The loop stops at the first accepted round or after rounds, and escalates: run.escalation is "not accepted after 3 round(s): ". A reply the generator rejects (not JSON, outside the schema, a quote not in the text) is a round too, and its reason is fed back; a proposer that fails otherwise (the server does not answer) ends the loop with "the proposer failed: …". Generator.proposer(messages, schema=…, first=…) builds the proposer: the first round asks the messages (or first, another generator — a stronger setting for the first try); each later round appends, per earlier round, its reply as the assistant's turn and its feedback as the user's.

Without a proposer the System generates itself: each round gives the earlier rounds' feedback as the fact feedback_into (default "feedback", a list of reasons) and the generating part reads it — the example above. history= continues an earlier refinement. run.to_dict() / Refinement.from_dict(d, catalog=cat) store and load it; run.replay(system) replays every round's trace and checks that the loop did what its record says: each round's acceptance, failed checks and causes follow from its response, the feedback is what the feedback function gives (pass a custom one again; accept too when it was a function), each round's input holds its proposal and the feedback before it, nothing ran after an accepted round, and the escalation matches.

What to expect. Re-asking with the violated constraints quoted lowers the share of wrong answers among those given, and more of the hard cases go to a person instead; the model often trades one violation for another, so a re-ask is not a fix. Typed facts extracted with schema= and quotes= are accepted or rejected by their quotes — a quote check does not see a row left out. Agreement of several samples is not correctness: samples can agree on one reading of the question that is not the intended one.

Not done here. No search: a loop re-asks one proposer, it does not enumerate alternatives or keep the best of two valid ones — that is solvi.search (below). No promise that re-asks converge. The share of agreement is a signal; calibrate it on labelled examples before you trust a threshold. No streaming, no tool calls, no caching of replies (put a caching proxy in front of the server).

Search over alternatives: solvi.search

When the candidates can be enumerated — the slots of a week, the orders of a few cities, the friends to meet — a search through the System's own checks beats asking a model to propose: run each candidate through the checks, keep the accepted ones, take the best.

from solvi import Answer, Catalog, Question, System
from solvi.refine import Fail
from solvi.search import Tree, search

cat = Catalog()
FLIGHTS = {("Oslo", "Rome"), ("Rome", "Paris"), ("Paris", "Oslo"), ("Rome", "Vienna")}

@cat.fn
def cities(problem: str) -> list:                 # the problem read into facts: computed once for the whole search
    return problem.split(", ")

@cat.check(hard=True, then={"ok": "no"})
def direct_flights(order: list) -> bool:          # false on a prefix → false on every order that starts with it
    bad = [f"no flight {a} - {b}" for a, b in zip(order, order[1:]) if (a, b) not in FLIGHTS and (b, a) not in FLIGHTS]
    return Fail(*bad) if bad else True

@cat.check(hard=True, then={"ok": "no"})
def every_city(cities: list, order: list) -> bool:
    return sorted(order) == sorted(cities)

@cat.rule("ok")
def ok(direct_flights, every_city) -> bool:
    return True

system = System(cat, [Question("ok", "A valid trip?", Answer.yes_no(), requires=["direct_flights", "every_city"])])

def orders(facts):                                # the space, read from the facts: one more city per step
    cs = facts["cities"]
    return Tree([], lambda o: [o + [c] for c in cs if c not in o], complete=lambda o: len(o) == len(cs))

run = search(system, {"problem": "Oslo, Rome, Paris, Vienna"}, "ok", orders, into="order",
             prune=["direct_flights"], keep=2)
print(run)
# search ok: 28 asked, 2 accepted
#   best: ['Oslo', 'Paris', 'Rome', 'Vienna']
#   exact: every candidate was asked or cut — pruned by direct_flights (10); the first accepted in the space's order
#   (and not the only one), given that the prune checks direct_flights stay false below a node
#   rejected by: direct_flights (4)
#   computed once: cities
run.response["ok"].answer, run.kept, run.replay(system)["ok"]

search(system, state, question, space, *, into=, objective=, maximize=True, prune=(), keep=1, budget=10_000, accept="checks", store=True, hold=True):

  • The space: a list or any iterable of candidates (given as the fact into); a dict {fact: [values]} (every combination, the first fact outermost, given as those facts); a Tree(root, children, complete=, bound=) walked depth first; or a function of the facts computed from state that returns one of these — the space read from the problem.
  • Accepted: as in refine (accept="checks": every hard check governing the question passed), and the question did not abstain — a candidate the System could not decide is never chosen. run.rejected counts the rejections by deciding check.
  • The best: with objective (a function of the candidate, or the name of a fact the question computes) the keep best accepted; without one the first keep in the space's order, and the search stops there (keep=2 says whether the first is the only one).
  • Cuts: prune names hard checks that, false on a partial node, stay false on every node below it; such a node is not expanded. A Tree's bound(node) is the best objective any candidate below can reach; a node that cannot beat what is kept is not expanded. budget caps the asks.
  • When it is exact: run.exact is True when the search ended by itself — every candidate was asked or cut — and run.why_exact names what that rests on: that the prune checks are monotone and the bound optimistic. Neither is checked by solvi; a wrong promise can cut the best candidate. When the budget stops it, exact is False and the best is the best of what was asked; when nothing is accepted, run.escalation says why (and a refine with a proposer can take over a space too big to search).
  • Facts computed once: parts that do not read the candidate (the problem read into typed facts — by rules or by a model) run once and are held for every candidate (run.held); a model reading the problem is called once, not per candidate. The searched asks are not stored or counted; the winner is asked again in full with the System (run.response, stored when the System stores), and if that full ask does not accept it the search escalates instead. run.to_dict() / SearchRun.from_dict(d, catalog=cat); run.replay(system) replays the winner's trace and checks it is accepted and its objective recomputes (pass a function objective again).

Where the candidates can be enumerated, prefer a search to a model's proposals: the checks decide every candidate, and every winner replays. What stays problem-specific is the space (a few lines per kind of problem) and, for an order, a walk of partial orders so the checks can judge a prefix. The price is speed: every candidate is a full ask, so a search runs far more asks than a hand-written search runs steps.

Not done here: no proposals by a model, no bisection over numbers (res.counterfactual does that), no parallel asks, no proof of the prune and bound promises. Each candidate is a full ask with the trace hashed (see the README's Speed table, benchmarks/ask_speed.py), so a space of millions is for code, not for this search.