solvi.experimental.compile¶
A specification — a policy, a regulation, a constraint description — compiled by an LLM into catalog parts (rules, hard checks, computed facts) that cite the clauses they implement; accepted only when two independent drafts agree on generated inputs and pass tests derived from the text; a person in the loop for what the drafts dispute; versions, recompilation of what a change touches, and the decisions a change moves.
Compile a specification into catalog parts: an LLM writes the rules, hard checks and computed facts a policy, a regulation or a constraint description states; solvi accepts them only after checks that need no labels.
from solvi.experimental.compile import Inputs, Spec, compile_spec, recompile, Versions
from solvi.core.slow.generate import generator
spec = Spec(policy_text) # split into numbered clauses: spec.clauses, spec.hash
inputs = Inputs({"amount": (0, 20000), "country": ["DE", "FR", "US"]}) # what a decision reads, and its values
writer = generator(URL, "openai/gpt-oss-120b", max_tokens=24000, extra_body={"reasoning": {"effort": "medium"}})
c = compile_spec(spec, [Question("decision", "...", Answer.choice(["approve", "refuse"]))], inputs, writer)
c.accepted, c.reason, c.parts # each part: kind and the clauses it implements
system = c.system() # an ordinary System; refused when not accepted
c2 = recompile(c, spec.revise(new_text), inputs, writer) # only what the change touches; c2.changes
decision_diff(c, c2, store=store) # which stored decisions change, and the clauses why
What the writer produces. A module of plain functions: each catalog part is a function whose name is the fact it sets
and whose argument names are what it reads, and a literal PARTS dict gives each part's kind ("fn", "check", "rule"),
for a hard check its then, for a rule its question, and the clauses it implements. Clauses no part implements are
listed in NOT_NORMATIVE with the reason. The module is checked by solvi.experimental.compile.sandbox (pure functions, standard-library
imports only) and run there, in a subprocess with limits, before anything of it enters this process.
Acceptance needs no labels. Two drafts are written independently (by default two samples of one writer: the first at
temperature 0, the second at 0.7 with seed 1; pass two writers for two models). A draft is accepted when:
1. it passes the sandbox and the module contract (every part listed, every clause cited or declared not normative);
2. the two drafts give the same answer to every question on every input of the pool — inputs generated from
Inputs (with the numbers of the specification and of both drafts, ±1, as boundary values), the samples you give,
and the tests' inputs. A disagreement is shown to both writers: the input, both answers, the clauses the deciding
parts cite;
3. both pass the tests a separate call derived from the specification — it never sees the code, and each test names
the clause it checks. A test that every draft which answers it fails goes back once to the test writer,
which works the answer out again and keeps, corrects or drops it (recorded). Every input of the pool must get an
answer;
4. both match the labelled examples and the reference, when you give them.
A draft with failures is rewritten from its module and the failures, up to rounds rounds; then the compilation is not
accepted, Compiled.catalog() refuses it, and reason says why. A draft that has not run for two rounds in a row
(refused by the contract or the sandbox, or no module in the reply) is replaced by a fresh one, written from the task
again with a seed of its own and told only what the stuck one was refused for — at most fresh_drafts times, each
replacement recorded; the fresh draft meets the same checks. The record keeps the spec's hash and clauses, every
prompt and reply, the writer, the tests, and each check's outcome per round.
c = compile_spec(spec, questions, inputs, writer, review=ask_a_person) # a person resolves what drafts dispute
c.record["person"] # every question asked, every answer, the tests they became
A person in the loop (review=) is asked, within a budget, about the inputs two drafts decide differently (one input
of each of the largest kinds of disagreement) and about disputed tests (instead of the test writer's re-check). An
answer becomes a test (source "person"), never code: both drafts must pass it, and the other conditions stay; the
answers are trusted like labels (a wrong one is a wrong test). An answer saying the specification does not decide the
input is a gap, and blocks acceptance. The person sees only what the drafts dispute, not a misreading they share.
What it does not do. Agreement of two drafts is not correctness: two samples of one model can share a misreading, and the tests come from the same model. Coverage is checked by citation, not by meaning — a part may cite a clause it implements wrongly. The inputs pool decides what "agree" covers: a region of inputs nobody generates is not compared. Labels, when you have them, are the stronger check. Nothing here writes extractors from text or searches; the compiled parts read structured inputs.
Rejected ¶
Bases: RuntimeError
A compilation that was not accepted was asked for its catalog. .reason says why it was not accepted.
Spec ¶
A specification as numbered clauses (c1, c2, ...). Spec(text) splits a text (Markdown-like: paragraphs, list
items and table rows are clauses, headings are sections); Spec(clauses={id: text}) takes them as given.
spec.revise(new_text) aligns a changed text with this one: unchanged clauses keep their ids, and changes says
which clauses changed, were added or removed.
render ¶
The clauses as the writer reads them: [c3] (section) text, with marks ({id: "changed"}) in front.
revise ¶
A new version of this specification, aligned with it: a clause whose text (and section) did not change keeps its
id; a replaced clause keeps the id of the one it replaces (changed); new clauses get new ids (added); the
others are removed. new.changes = {"unchanged", "changed", "added", "removed"} (lists of ids).
Inputs ¶
Inputs(fields: dict | None = None, *, describe: str | None = None, samples=(), generate=None, n: int = 1000, seed: int = 0)
What a compiled decision reads, and where the inputs that compare two drafts come from.
fields: {name: domain} — a list of the values that occur (drawn from), a (lo, hi) tuple of numbers (a range: ints
when both ends are ints), or a type (str, dict, a pydantic model: not generated; give samples or generate).
describe: the inputs in words for the writer (default: made from fields). samples: real inputs (unlabelled),
all put in the pool. generate(rng) → one input: your generator, called n times. Without generate, n inputs are
drawn from the domains, when every field has one. Boundary values: for every number field, every number of the
specification and the drafts inside its range, -1 / exactly / +1, on a random input each.
validate ¶
Why x is not a complete input ([] — it is): a missing or unknown field, a value outside a listed domain.
pool ¶
The inputs two drafts are compared on: the samples, n generated inputs, and the boundary inputs around
numbers (the specification's and the drafts'). Deterministic for the same numbers.
Compiled
dataclass
¶
Compiled(spec: Spec, questions: list, source: str, parts: dict, not_normative: dict, accepted: bool, reason: str, record: dict = dict(), changes: dict | None = None, _ns: dict | None = None)
A compilation: the module, its parts with the clauses each implements, whether it was accepted and why, and the
record (spec, prompts and replies, writer, tests, checks per round). catalog() / system() load the module (through
solvi.experimental.compile.sandbox.load) and refuse an unaccepted compilation unless allow_unaccepted=True.
fingerprint
property
¶
The catalog's fingerprint (what a stored decision records as the catalog that decided).
clauses_of ¶
The clauses a part implements (a rule may be named answer:<question>).
part_of ¶
The PARTS name of a trace step (answer:<question> → its rule's function).
catalog ¶
→ (Catalog, questions): the module loaded in this process (sandbox.load), each hard check required by the questions it names. Raises Rejected for a compilation that was not accepted.
system ¶
An ordinary System over the compiled catalog (kw: System's own — storage, ...).
reviewer ¶
The person's recorded answers as a reviewer (compile_spec(..., review=c.reviewer()) reruns this
compilation with them): a question asked before gets the same ruling; a new one gets none (None).
save ¶
Write module.py (the code, as accepted) and compiled.json (everything) into a folder.
Dispute
dataclass
¶
Dispute(kind: str, input: dict, answers: list, clauses: list, round: int, spec: str = '', test: dict | None = None, similar: int = 1, clause_text: dict = dict())
What a person is asked during a compilation.
kind: "disagreement" — the two drafts answer input differently (answers[i] is draft i's answer, clauses[i]
the clauses its deciding parts cite; similar: how many inputs of the pool disagree the same way); "test" — every
draft that answers the test test (its clause, input, expected answer and why) fails it (answers: what they
gave). round: the round; spec: the specification's name; clause_text: the
texts of the clauses involved.
Ruling
dataclass
¶
A person's answer to a Dispute. verdict: "draft" (draft draft is right), "answer" (the right answer is
expect, {question: answer}), "neither" (the specification does not decide this input: a gap — or, with expect,
the right answer neither draft gave), "keep" / "drop" (a disputed test is right / is not a test of the
specification). A plain value is read too: an int picks a draft, a dict is an answer, "keep" / "drop" / "neither".
DecisionDiff
dataclass
¶
The decisions a new compilation changes: changed = [{"id", "question", "old", "new", "causes": [{"step",
"part", "clauses", "why"}]}]; report is solvi.core.store.diff's own report.
moved
property
¶
The changes whose answer moves (the others keep their answer and are decided another way: by another check, or forced where a rule answered).
Versions ¶
Compiled catalogs as numbered versions in a folder (v1/, v2/, ... each module.py + compiled.json).
system(n) builds a version's System; replay_all(store) replays every stored decision with the version whose
catalog fingerprint it recorded — an old decision is checked against the rules it was made with.
split_clauses ¶
A text → [(section, clause text)]: each paragraph, list item and table row is a clause (a table row is written
as header: cell; ...); headings name the section and are not clauses themselves.
numbers_in ¶
The number literals of a text or a module (thousands separators allowed).
read_module ¶
A module's PARTS, NOT_NORMATIVE (and REMOVE for a patch) → (parts, not_normative, remove, problems). The checks of
the contract: literals, kinds, every part a top-level function, a known question and an allowed answer in then, a
rule per question (not for a patch), cited clauses that exist.
coverage ¶
Clauses neither cited by a part nor declared not normative.
merge ¶
merge(old_source: str, old_parts: dict, old_nn: dict, patch_source: str, patch_parts: dict, patch_nn: dict, remove: dict, spec: Spec) -> tuple[str, dict, dict]
The current module with a patch applied → (source, parts, not_normative). Statements the patch does not name stay byte-identical and in place; replaced ones are swapped in place; new ones go before PARTS.
reference_reviewer ¶
A simulated person for experiments: reference(input) → {question: answer} (a hand-written reference). On a
disagreement it picks the draft whose answers equal the reference's, else gives the reference's answer (an answer
neither draft gave); on a disputed test it keeps the test when the reference agrees, else corrects it. An input the
reference cannot read is skipped (a test on it: dropped).
compile_spec ¶
compile_spec(spec: Spec, questions, inputs: Inputs, writer=None, *, rounds: int = 4, tests: int = 40, examples=(), reference: Callable | None = None, mem_mb: int = 1024, item_s: int = 5, fresh_drafts: int = 2, stuck_after: int = 2, review: Callable | None = None, review_budget: int | None = 20, review_per_round: int | None = 5) -> Compiled
Compile spec into catalog parts answering questions over inputs (see the module docs). writer: a
solvi.core.slow.generate Generator (or two, for two models), or a base URL (model openai/gpt-oss-120b). examples: labelled
[(input, {question: answer})] that both drafts must match; reference(input) → {question: answer}, compared on the
whole pool. fresh_drafts: how many times in all a draft that did not run (refused by the contract or the sandbox,
or no module in the reply) for stuck_after rounds in a row is replaced by a fresh draft, written from the task
again with a seed of its own (0: never). review(dispute) → Ruling: a person in the loop — asked about the drafts'
disagreements and the disputed tests (at most review_per_round questions a round, review_budget in all; None:
no limit); each answer becomes a test both drafts must pass (see Dispute, Ruling). Never raises for a bad
draft: the result says whether it was accepted.
recompile ¶
recompile(old: Compiled, spec: Spec, inputs: Inputs, writer=None, *, rounds: int = 4, tests: int = 40, examples=(), reference: Callable | None = None, mem_mb: int = 1024, item_s: int = 5, fresh_drafts: int = 2, stuck_after: int = 2, review: Callable | None = None, review_budget: int | None = 20, review_per_round: int | None = 5) -> Compiled
Recompile an accepted compilation for a revised specification (old.spec.revise(...)): the writer returns only
the parts it adds, replaces or removes — each citing a changed or added clause — and the rest stays byte-identical.
Accepted by the same checks as compile_spec (two independent patches agree, tests written for the new
specification; review= a person in the loop, as for compile_spec). The result's changes: the clauses' changes
and the parts added / replaced / removed / kept.
decision_diff ¶
Which decisions new would change: the stored ones (store, decided by old's catalog — filters as for
solvi.core.store.diff.diff) or inputs (each decided by old first, in a temporary store). Each change's causes are the steps
solvi.core.store.diff names, with the clauses their parts implement (in the new compilation, else the old one).
to_guard ¶
to_guard(compiled: Compiled, guard, tools=None, *, question: str | None = None, on_fail: str = 'deny', allow=None, name: str | None = None) -> list
Register the compiled policy with a solvi.Guard → the names of the policies added. The policies read
the compiled catalog's inputs (they must be facts the guard gives: tool_name, tool_arguments, conversation, ... or
your declared facts) and run the compiled System on them.
allow=None: each compiled hard check that names question becomes a policy, which fails with the check's reasons
when the check is false. allow="yes" (the answer that lets a call through): one policy (name, default
compiled_<question>) that asks the compiled question and fails when the answer is anything else — with the
clauses of the parts that decided (the false hard checks, else the question's rule) as the reasons — or when the
compiled policy cannot answer (it abstains: an input it cannot read).
question: the compiled question (default: the only one).