Skip to content

solvi.experimental.compile

A specification — a policy, a regulation, a constraint description — compiled by an LLM into catalog parts (rules, hard checks, computed facts) that cite the clauses they implement; accepted only when two independent drafts agree on generated inputs and pass tests derived from the text; a person in the loop for what the drafts dispute; versions, recompilation of what a change touches, and the decisions a change moves.

Compile a specification into catalog parts: an LLM writes the rules, hard checks and computed facts a policy, a regulation or a constraint description states; solvi accepts them only after checks that need no labels.

from solvi.experimental.compile import Inputs, Spec, compile_spec, recompile, Versions
from solvi.core.slow.generate import generator

spec = Spec(policy_text)                                  # split into numbered clauses: spec.clauses, spec.hash
inputs = Inputs({"amount": (0, 20000), "country": ["DE", "FR", "US"]})    # what a decision reads, and its values
writer = generator(URL, "openai/gpt-oss-120b", max_tokens=24000, extra_body={"reasoning": {"effort": "medium"}})
c = compile_spec(spec, [Question("decision", "...", Answer.choice(["approve", "refuse"]))], inputs, writer)
c.accepted, c.reason, c.parts                             # each part: kind and the clauses it implements
system = c.system()                                        # an ordinary System; refused when not accepted

c2 = recompile(c, spec.revise(new_text), inputs, writer)  # only what the change touches; c2.changes
decision_diff(c, c2, store=store)                         # which stored decisions change, and the clauses why

What the writer produces. A module of plain functions: each catalog part is a function whose name is the fact it sets and whose argument names are what it reads, and a literal PARTS dict gives each part's kind ("fn", "check", "rule"), for a hard check its then, for a rule its question, and the clauses it implements. Clauses no part implements are listed in NOT_NORMATIVE with the reason. The module is checked by solvi.experimental.compile.sandbox (pure functions, standard-library imports only) and run there, in a subprocess with limits, before anything of it enters this process.

Acceptance needs no labels. Two drafts are written independently (by default two samples of one writer: the first at temperature 0, the second at 0.7 with seed 1; pass two writers for two models). A draft is accepted when: 1. it passes the sandbox and the module contract (every part listed, every clause cited or declared not normative); 2. the two drafts give the same answer to every question on every input of the pool — inputs generated from Inputs (with the numbers of the specification and of both drafts, ±1, as boundary values), the samples you give, and the tests' inputs. A disagreement is shown to both writers: the input, both answers, the clauses the deciding parts cite; 3. both pass the tests a separate call derived from the specification — it never sees the code, and each test names the clause it checks. A test that every draft which answers it fails goes back once to the test writer, which works the answer out again and keeps, corrects or drops it (recorded). Every input of the pool must get an answer; 4. both match the labelled examples and the reference, when you give them. A draft with failures is rewritten from its module and the failures, up to rounds rounds; then the compilation is not accepted, Compiled.catalog() refuses it, and reason says why. A draft that has not run for two rounds in a row (refused by the contract or the sandbox, or no module in the reply) is replaced by a fresh one, written from the task again with a seed of its own and told only what the stuck one was refused for — at most fresh_drafts times, each replacement recorded; the fresh draft meets the same checks. The record keeps the spec's hash and clauses, every prompt and reply, the writer, the tests, and each check's outcome per round.

c = compile_spec(spec, questions, inputs, writer, review=ask_a_person)   # a person resolves what drafts dispute
c.record["person"]                                    # every question asked, every answer, the tests they became

A person in the loop (review=) is asked, within a budget, about the inputs two drafts decide differently (one input of each of the largest kinds of disagreement) and about disputed tests (instead of the test writer's re-check). An answer becomes a test (source "person"), never code: both drafts must pass it, and the other conditions stay; the answers are trusted like labels (a wrong one is a wrong test). An answer saying the specification does not decide the input is a gap, and blocks acceptance. The person sees only what the drafts dispute, not a misreading they share.

What it does not do. Agreement of two drafts is not correctness: two samples of one model can share a misreading, and the tests come from the same model. Coverage is checked by citation, not by meaning — a part may cite a clause it implements wrongly. The inputs pool decides what "agree" covers: a region of inputs nobody generates is not compared. Labels, when you have them, are the stronger check. Nothing here writes extractors from text or searches; the compiled parts read structured inputs.

Rejected

Rejected(reason)

Bases: RuntimeError

A compilation that was not accepted was asked for its catalog. .reason says why it was not accepted.

Spec

Spec(text: str | None = None, *, clauses=None, name: str = 'specification')

A specification as numbered clauses (c1, c2, ...). Spec(text) splits a text (Markdown-like: paragraphs, list items and table rows are clauses, headings are sections); Spec(clauses={id: text}) takes them as given. spec.revise(new_text) aligns a changed text with this one: unchanged clauses keep their ids, and changes says which clauses changed, were added or removed.

hash property

hash: str

sha256 of the clauses (ids, sections, texts).

render

render(marks: dict | None = None) -> str

The clauses as the writer reads them: [c3] (section) text, with marks ({id: "changed"}) in front.

revise

revise(text: str | None = None, *, clauses=None) -> 'Spec'

A new version of this specification, aligned with it: a clause whose text (and section) did not change keeps its id; a replaced clause keeps the id of the one it replaces (changed); new clauses get new ids (added); the others are removed. new.changes = {"unchanged", "changed", "added", "removed"} (lists of ids).

Inputs

Inputs(fields: dict | None = None, *, describe: str | None = None, samples=(), generate=None, n: int = 1000, seed: int = 0)

What a compiled decision reads, and where the inputs that compare two drafts come from.

fields: {name: domain} — a list of the values that occur (drawn from), a (lo, hi) tuple of numbers (a range: ints when both ends are ints), or a type (str, dict, a pydantic model: not generated; give samples or generate). describe: the inputs in words for the writer (default: made from fields). samples: real inputs (unlabelled), all put in the pool. generate(rng) → one input: your generator, called n times. Without generate, n inputs are drawn from the domains, when every field has one. Boundary values: for every number field, every number of the specification and the drafts inside its range, -1 / exactly / +1, on a random input each.

text

text() -> str

The inputs in words for the writer.

validate

validate(x) -> list[str]

Why x is not a complete input ([] — it is): a missing or unknown field, a value outside a listed domain.

pool

pool(numbers=()) -> list[dict]

The inputs two drafts are compared on: the samples, n generated inputs, and the boundary inputs around numbers (the specification's and the drafts'). Deterministic for the same numbers.

Compiled dataclass

Compiled(spec: Spec, questions: list, source: str, parts: dict, not_normative: dict, accepted: bool, reason: str, record: dict = dict(), changes: dict | None = None, _ns: dict | None = None)

A compilation: the module, its parts with the clauses each implements, whether it was accepted and why, and the record (spec, prompts and replies, writer, tests, checks per round). catalog() / system() load the module (through solvi.experimental.compile.sandbox.load) and refuse an unaccepted compilation unless allow_unaccepted=True.

fingerprint property

fingerprint: str

The catalog's fingerprint (what a stored decision records as the catalog that decided).

clauses_of

clauses_of(part: str) -> list

The clauses a part implements (a rule may be named answer:<question>).

part_of

part_of(step: str) -> str | None

The PARTS name of a trace step (answer:<question> → its rule's function).

catalog

catalog(allow_unaccepted: bool = False) -> tuple[Catalog, list]

→ (Catalog, questions): the module loaded in this process (sandbox.load), each hard check required by the questions it names. Raises Rejected for a compilation that was not accepted.

system

system(allow_unaccepted: bool = False, **kw)

An ordinary System over the compiled catalog (kw: System's own — storage, ...).

reviewer

reviewer() -> Callable

The person's recorded answers as a reviewer (compile_spec(..., review=c.reviewer()) reruns this compilation with them): a question asked before gets the same ruling; a new one gets none (None).

save

save(path) -> Path

Write module.py (the code, as accepted) and compiled.json (everything) into a folder.

Dispute dataclass

Dispute(kind: str, input: dict, answers: list, clauses: list, round: int, spec: str = '', test: dict | None = None, similar: int = 1, clause_text: dict = dict())

What a person is asked during a compilation.

kind: "disagreement" — the two drafts answer input differently (answers[i] is draft i's answer, clauses[i] the clauses its deciding parts cite; similar: how many inputs of the pool disagree the same way); "test" — every draft that answers the test test (its clause, input, expected answer and why) fails it (answers: what they gave). round: the round; spec: the specification's name; clause_text: the texts of the clauses involved.

text

text() -> str

The question as a person reads it.

Ruling dataclass

Ruling(verdict: str, draft: int | None = None, expect: dict | None = None, note: str = '')

A person's answer to a Dispute. verdict: "draft" (draft draft is right), "answer" (the right answer is expect, {question: answer}), "neither" (the specification does not decide this input: a gap — or, with expect, the right answer neither draft gave), "keep" / "drop" (a disputed test is right / is not a test of the specification). A plain value is read too: an int picks a draft, a dict is an answer, "keep" / "drop" / "neither".

DecisionDiff dataclass

DecisionDiff(changed: list, report: Any, total: int)

The decisions a new compilation changes: changed = [{"id", "question", "old", "new", "causes": [{"step", "part", "clauses", "why"}]}]; report is solvi.core.store.diff's own report.

moved property

moved: list

The changes whose answer moves (the others keep their answer and are decided another way: by another check, or forced where a rule answered).

Versions

Versions(path)

Compiled catalogs as numbered versions in a folder (v1/, v2/, ... each module.py + compiled.json). system(n) builds a version's System; replay_all(store) replays every stored decision with the version whose catalog fingerprint it recorded — an old decision is checked against the rules it was made with.

add

add(compiled: Compiled, label: str | None = None) -> int

Store an accepted compilation as the next version → its number.

get

get(n: int | None = None) -> Compiled

Version n (the latest when None).

find

find(fingerprint: str) -> int | None

The version whose catalog has this fingerprint.

replay_all

replay_all(store, **filters) -> list

Replay every stored decision against the version that made it (by the recorded catalog fingerprint) → the ones that do not replay (as TraceStorage.replay_all), plus decisions whose catalog is no version here.

split_clauses

split_clauses(text: str) -> list[tuple[str, str]]

A text → [(section, clause text)]: each paragraph, list item and table row is a clause (a table row is written as header: cell; ...); headings name the section and are not clauses themselves.

numbers_in

numbers_in(text: str) -> list[float]

The number literals of a text or a module (thousands separators allowed).

read_module

read_module(source: str, spec: Spec, questions, patch=False) -> tuple[dict, dict, dict, list[str]]

A module's PARTS, NOT_NORMATIVE (and REMOVE for a patch) → (parts, not_normative, remove, problems). The checks of the contract: literals, kinds, every part a top-level function, a known question and an allowed answer in then, a rule per question (not for a patch), cited clauses that exist.

coverage

coverage(spec: Spec, parts: dict, nn: dict) -> list[str]

Clauses neither cited by a part nor declared not normative.

merge

merge(old_source: str, old_parts: dict, old_nn: dict, patch_source: str, patch_parts: dict, patch_nn: dict, remove: dict, spec: Spec) -> tuple[str, dict, dict]

The current module with a patch applied → (source, parts, not_normative). Statements the patch does not name stay byte-identical and in place; replaced ones are swapped in place; new ones go before PARTS.

reference_reviewer

reference_reviewer(reference: Callable) -> Callable

A simulated person for experiments: reference(input) → {question: answer} (a hand-written reference). On a disagreement it picks the draft whose answers equal the reference's, else gives the reference's answer (an answer neither draft gave); on a disputed test it keeps the test when the reference agrees, else corrects it. An input the reference cannot read is skipped (a test on it: dropped).

compile_spec

compile_spec(spec: Spec, questions, inputs: Inputs, writer=None, *, rounds: int = 4, tests: int = 40, examples=(), reference: Callable | None = None, mem_mb: int = 1024, item_s: int = 5, fresh_drafts: int = 2, stuck_after: int = 2, review: Callable | None = None, review_budget: int | None = 20, review_per_round: int | None = 5) -> Compiled

Compile spec into catalog parts answering questions over inputs (see the module docs). writer: a solvi.core.slow.generate Generator (or two, for two models), or a base URL (model openai/gpt-oss-120b). examples: labelled [(input, {question: answer})] that both drafts must match; reference(input) → {question: answer}, compared on the whole pool. fresh_drafts: how many times in all a draft that did not run (refused by the contract or the sandbox, or no module in the reply) for stuck_after rounds in a row is replaced by a fresh draft, written from the task again with a seed of its own (0: never). review(dispute) → Ruling: a person in the loop — asked about the drafts' disagreements and the disputed tests (at most review_per_round questions a round, review_budget in all; None: no limit); each answer becomes a test both drafts must pass (see Dispute, Ruling). Never raises for a bad draft: the result says whether it was accepted.

recompile

recompile(old: Compiled, spec: Spec, inputs: Inputs, writer=None, *, rounds: int = 4, tests: int = 40, examples=(), reference: Callable | None = None, mem_mb: int = 1024, item_s: int = 5, fresh_drafts: int = 2, stuck_after: int = 2, review: Callable | None = None, review_budget: int | None = 20, review_per_round: int | None = 5) -> Compiled

Recompile an accepted compilation for a revised specification (old.spec.revise(...)): the writer returns only the parts it adds, replaces or removes — each citing a changed or added clause — and the rest stays byte-identical. Accepted by the same checks as compile_spec (two independent patches agree, tests written for the new specification; review= a person in the loop, as for compile_spec). The result's changes: the clauses' changes and the parts added / replaced / removed / kept.

decision_diff

decision_diff(old: Compiled, new: Compiled, *, store=None, inputs=None, **filters) -> DecisionDiff

Which decisions new would change: the stored ones (store, decided by old's catalog — filters as for solvi.core.store.diff.diff) or inputs (each decided by old first, in a temporary store). Each change's causes are the steps solvi.core.store.diff names, with the clauses their parts implement (in the new compilation, else the old one).

to_guard

to_guard(compiled: Compiled, guard, tools=None, *, question: str | None = None, on_fail: str = 'deny', allow=None, name: str | None = None) -> list

Register the compiled policy with a solvi.Guard → the names of the policies added. The policies read the compiled catalog's inputs (they must be facts the guard gives: tool_name, tool_arguments, conversation, ... or your declared facts) and run the compiled System on them.

allow=None: each compiled hard check that names question becomes a policy, which fails with the check's reasons when the check is false. allow="yes" (the answer that lets a call through): one policy (name, default compiled_<question>) that asks the compiled question and fails when the answer is anything else — with the clauses of the parts that decided (the false hard checks, else the question's rule) as the reasons — or when the compiled policy cannot answer (it abstains: an input it cannot read). question: the compiled question (default: the only one).