Skip to content

A specification compiled into the catalog: solvi.experimental.compile

Experimental. The API may change. What it promises is the procedure below, not correctness: a compiled part is as right as the drafts and the tests that agreed on it.

A policy, a regulation or a constraint description says what to decide; solvi decides with catalog parts. solvi.experimental.compile lets an LLM write those parts from the text and accepts them only after checks that need no labelled examples:

  1. the text is split into numbered clauses (Spec); every part the writer returns names the clauses it implements, and every clause is cited by a part or declared not normative with a reason;
  2. the module is pure functions checked by solvi.experimental.compile.sandbox (an ast allowlist of standard-library imports, no files, reflection or dunders) and run there — in a subprocess with memory and time limits — before anything of it enters your process;
  3. two drafts are written independently and must give the same answer to every question on every input of a pool: inputs drawn from the values you declare (Inputs), boundary values around every number of the text and of both drafts, your unlabelled samples, and the tests' inputs. Every input must get an answer: a draft that abstains or raises on one has a bug. A disagreement goes back to both writers with the input, both answers and the clauses their deciding parts cite;
  4. tests derived from the text, written by a separate call that never sees the code, each naming the clause it checks; both drafts must pass them. A test that every draft which answers it fails (at least one answers) goes back once to the test writer, which works the answer out again and keeps, corrects or drops it — recorded, since a test can be wrong as well (a draft that abstains on the test's input says nothing about the test: that goes back to the draft);
  5. labelled examples or a reference function, when you have them (examples=, reference=), as further checks.

A draft that fails is rewritten from its module and the failures, for up to rounds rounds. Acceptance is automatic when everything passes; otherwise c.accepted is False, c.reason says why, and c.system() raises Rejected.

A draft that does not run for two rounds in a row (stuck_after=2) — the module contract or the sandbox refuses it, or the reply holds no module — is replaced by a fresh draft: written from the task again, with a seed of its own, told what the stuck draft was refused for but never shown its code. At most fresh_drafts=2 replacements per compilation (0: never); each is in c.record["replaced"] (round, draft, why). The fresh draft meets every check above; it only keeps a draft stuck on the contract from blocking a partner that works.

import json

from solvi import Answer, Question
from solvi.experimental.compile import Inputs, Spec, compile_spec

POLICY = """# Shipping
- An order of 50 or more ships free; otherwise shipping costs 5.
- Orders to the world zone heavier than 30 kg are refused.
"""

MODULE = '''
def free_shipping(total):
    return total >= 50

def small_enough(zone, weight):
    return not (zone == "world" and weight > 30) or Fail(f"{weight} kg to the world zone")

def ship(free_shipping):
    return "free" if free_shipping else "paid"

PARTS = {
    "free_shipping": {"kind": "fn", "clauses": ["c1"]},
    "small_enough": {"kind": "check", "hard": True, "then": {"ship": "refused"}, "clauses": ["c2"]},
    "ship": {"kind": "rule", "question": "ship", "clauses": ["c1"]},
}
NOT_NORMATIVE = {}
'''


class StandIn:                       # a stand-in for the writer: generator(URL, "openai/gpt-oss-120b", ...)
    model_id = "stand-in"

    def fingerprint(self):
        return "stand-in"

    def generate(self, messages, parse=None, **kw):
        tests = [{"clause": "c1", "input": {"zone": "home", "total": 50, "weight": 1}, "expect": {"ship": "free"},
                  "why": "50 or more ships free"}]
        fence = "`" * 3                 # the writer answers in a fenced block
        text = (f"{fence}json\n{json.dumps(tests)}\n{fence}" if messages[-1]["content"].startswith("# Write tests")
                else f"{fence}python\n{MODULE}{fence}")
        return type("G", (), {"value": parse(text) if parse else text, "meta": {"text": text}})()


spec = Spec(POLICY)
inputs = Inputs({"zone": ["home", "world"], "total": (0, 200), "weight": (0, 50)}, n=300)
c = compile_spec(spec, [Question("ship", "Ship free, paid or refused?", Answer.choice(["free", "paid", "refused"]))],
                 inputs, StandIn())
print(c.accepted, c.reason, c.record["rounds"][0]["agreement"])
print(c.parts["small_enough"]["clauses"], spec.clauses["c2"].text)
print(c.system().ask({"zone": "world", "total": 80, "weight": 40})["ship"].answer)
True accepted in round 1 {'inputs': 318, 'disagree': 0}
['c2'] Orders to the world zone heavier than 30 kg are refused.
refused

With a model the stand-in is generator(base_url, "openai/gpt-oss-120b", max_tokens=24000, extra_body={"reasoning": {"effort": "medium"}}) (or a base URL string, which builds that). Two drafts by default are two samples of one model — the first at temperature 0, the second at 0.7 with seed 1; writer=[a, b] takes two models.

What the writer is asked for. A module of plain functions — a part's name is the fact it sets, its argument names are the inputs or facts it reads (a part that reads a name nothing gives is refused before it runs, unless other parts call it as a plain function: then it is a helper the writer listed in PARTS — it leaves PARTS, its clauses go to the parts that call it, recorded in the round's "notes"; a hard check is never treated so; and a part that reads a key of a dict-valued input by its own name — friends inside the input facts — gets an accessor part def friends(facts): return facts["friends"], added and noted, which raises (so the decision abstains) when the key is missing) — and two literal dicts: PARTS (kind fn / check / rule, for a hard check its then, for a rule its question, the clauses; a check must cite one, a fact that only reads an input or a rule giving a default may cite none) and NOT_NORMATIVE. A hard check that names a question is required in that question's flow. A check may return Fail("why"). The prompts ask for one part per quantity a clause defines, so a stored decision shows each.

The record. c.record keeps the spec's hash, the writer, every prompt and reply, the tests (and the invalid ones with why), the reviews of tests, and per round each draft's problems, which check caught it ("contract", "sandbox", "abstained", "tests", "disagreement", "labelled examples", "the reference") and the agreement. c.save(folder) writes module.py and compiled.json; Compiled.load(folder) reads them back and refuses a module that was edited.

A person in the loop: review=

Two drafts that disagree, or a test every draft fails, can stop a compilation that is nearly right: one draft misreads a clause, or the derived test is wrong. review= puts a person where the loop cannot settle it alone:

from solvi.experimental.compile import Ruling, compile_spec, reference_reviewer

def ask_a_person(d):                     # d: a Dispute
    print(d.text())                      # the input, each draft's answer and the clauses it cites
    return Ruling.pick(0)                # or Ruling.answer({"ship": "free"}), Ruling.neither("the text does not say"),
                                         # and for a disputed test Ruling.keep() / Ruling.drop()

c = compile_spec(spec, questions, inputs, writer, review=ask_a_person, review_budget=20, review_per_round=5)
c.record["person"]                       # every question asked, every answer, and the tests they became
c2 = compile_spec(spec, questions, inputs, writer, review=c.reviewer())    # rerun with the same answers

After both drafts ran in a round, the person is asked about:

  • disputed tests — a test every draft that answers it fails — instead of the test writer's own re-check: keep it, drop it, or give the right answer (the test is corrected);
  • disagreements — the inputs are grouped by both answers and the clauses the deciding parts cite, and one input of each of the largest groups is asked about (new groups first; a group answered in an earlier round that still divides the drafts is asked again, with another input). The person says which draft is right or gives the right answer, or says the specification does not decide the input.

At most review_per_round questions a round and review_budget in all (None: no limit); a reviewer that returns None skips the question. An answer becomes a test (source "person"), never code: both drafts must pass it from then on, and a test the writer derived for the same input that contradicts it is corrected. The other conditions of acceptance stay as without the person — both drafts pass every test, agree on every input of the pool and answer every one — but the person's answers can replace or drop derived tests, so they are trusted like labels: a wrong answer becomes a wrong test, and both drafts can follow it. An answer saying the specification does not decide an input (Ruling.neither without an answer) is a gap: the compilation is not accepted until the text is amended.

What the person does not see. Only what the drafts dispute. A misreading both drafts share — say, both accept only a bare "yes" as the user's confirmation ("Yes, I confirm!" refused), where the policy means any explicit yes, or both refuse a call after any earlier call, which the policy never says — gives no disagreement and passes the tests, so nobody is asked and it is accepted. (The same limit as N-version programming, whose independent versions share misreadings, and as asking questions only where sampled programs differ.) In our runs that happened on parts of a customer-service policy. Look at some decisions the drafts agree on before you rely on a compiled policy.

reference_reviewer(fn) is a simulated person for experiments: fn(input) → {question: answer}, a hand-written reference. It picks the draft equal to the reference, else gives the reference's answer; it keeps a disputed test the reference agrees with, else corrects it. recompile takes the same options.

A changed specification: recompile and the decisions it moves

spec2 = c.spec.revise(new_text)            # unchanged clauses keep their ids; spec2.changes: changed / added / removed
c2 = recompile(c, spec2, inputs, writer)    # the writer returns only the parts it adds, replaces or removes
c2.changes["parts"]                         # {"added", "replaced", "removed", "kept"}; the kept ones are byte-identical
print(decision_diff(c, c2, store=store))    # which stored decisions change, and the clauses of their causes

The writer sees the new text with its changed and added clauses marked and the removed ones listed, and the current module; it returns a patch. Each added or replaced part must cite a changed or added clause, each removed part the clause that removes it, and no part may still cite a removed clause — otherwise the patch goes back with the reasons. Then the same acceptance runs on the merged module, with tests written for the new text. decision_diff re-runs stored decisions (or inputs=, decided by the old version first) through solvi.core.store.diff and maps each cause step to the clauses its part cites — as fine as the parts are: a rule that cites every clause names every clause.

Versions and replay

from solvi.experimental.compile import Versions
versions = Versions("policy_versions")      # v1/, v2/ ... each module.py + compiled.json + version.json
n = versions.add(c2, "after the change")    # accepted compilations only
system = versions.system()                  # the latest; versions.system(1) the first
versions.replay_all(store)                  # every stored decision against the version that made it ([] — all replay)

A stored decision records the fingerprint of the catalog that made it; replay_all replays each against that version, so old decisions keep verifying after the rules changed, and a decision made by a catalog that is no version here is reported.

Into an agent guard

to_guard(c, guard, tools) registers each compiled hard check as a policy of a solvi.solutions.guard.Guard: the policy reads the compiled catalog's inputs (they must be facts the guard gives — tool_name, tool_arguments, conversation, conversation_roles or your declared facts), runs the compiled System and refuses with the clause the check implements as the reason.

A policy compiled as a question — "may this call be made?" — goes in whole with allow=:

guard = Guard(fact_names=FIELDS)                         # the facts the compiled policy reads, given with each call
to_guard(c, guard, allow="yes", name="shop_policy")      # one policy: the compiled answer must be "yes"
d = guard.check({"name": "refund_order", "arguments": {}}, facts=call_facts)        # e.g. a refund of 250
d.outcome, d.reasons    # deny, ['shop_policy: ... — [c6] Refunds over 200 go to a human: ... [deny]']

The policy asks the compiled question and refuses any other answer, naming the clauses of the parts that decided (the false hard checks, else the question's rule); an input the compiled policy cannot answer (it abstains — a fact it cannot read) is refused as well, with the reason. When the policy text changes, compile it again (recompile patches only what the change touches) and decision_diff(old, new, inputs=calls) lists which of the calls you pass move, with the clauses why — before the new guard goes live.

Not done here. Agreement is not correctness: two samples of one model can share a misreading, and the tests come from the same model — a wrong reading that both drafts and the tests share is accepted. Coverage is by citation, not by meaning. "Agree" covers the pool only: inputs nobody generates are not compared, so declare the domains and give samples of the real inputs. The parts read structured inputs: nothing here writes extractors from text, a search, or features for a head. Once loaded, a compiled module runs in your process with restricted builtins; the subprocess limits hold only during compilation (see solvi.experimental.compile.sandbox). Labels, when you have them, are the stronger check — pass them.