Skip to content

solvi.refine

A check that says why (Fail), and the loop propose → check → re-ask with the reasons → escalate, recorded and replayable.

Propose → check → re-ask with the reasons → escalate: a loop around a System whose checks judge a proposal.

from solvi.refine import Fail, refine

@cat.check(hard=True, then={"accept": "no"})
def nobody_busy(facts, slot) -> bool:
    busy = [f"{p} is busy {b}" for p, b in clashes(facts, slot)]
    return Fail(*busy) if busy else True          # False, with the reasons the proposer is told

run = refine(system, {"problem": text}, "accept", propose=writer.proposer(messages), into="slot", rounds=3)
run.accepted, run.proposal            # the first proposal the checks accepted, or None
run.escalation                        # why a person gets it: not accepted after 3 rounds, with the last reasons
run.rounds                            # every round: the proposal, its response (stored, replayable), the reasons
run.replay(system)                    # every round's trace, and that the loop did what its record says

A check gives its reason by returning Fail("...") (one or more reasons) instead of False: it is False everywhere a bool is read — rules, hard checks, plain Python — and the reasons are recorded with the check (record.extra["reasons"]), added to the answer's why when the check decides ("hard check nobody_busy is false: Harold is busy 13:30 - 15:30"), shown in the audit and compared on replay. A check that returns False keeps working: its reason is its docstring's first line, else " is false".

One round: the proposer is called with the state and the earlier rounds and returns a proposal (a value, or a solvi.generate.Generated — its record of the request is kept in the round); the proposal is given to the System under into, and the System is asked. The round is accepted when accept says so — "checks" (default): every hard check that governs the question was evaluated and passed; an answer or a list of answers ("yes"); or a function of the Response. Otherwise its reasons — what each failed hard check governing the question says, in catalog order, or, when the question could not be decided at all, the errors of the parts that failed ("spec: ValueError: no slot in the plan") — become the round's feedback (feedback(round) may reword them), and the next round's proposer sees them. The loop stops at the first accepted round or after rounds, and then escalates. A reply the generator rejects (solvi.llm InvalidOutput: not JSON, outside the schema, a quote not in the text) is a round too: its reason is fed back. A proposer that fails otherwise (the server does not answer) ends the loop with an escalation.

Without a proposer the System generates itself (a part made by Generator.part): each round gives the feedback of the earlier rounds as the fact feedback_into (a list of reasons), and the part reads it.

Not done here: no search over alternatives (a loop re-asks one proposer; it does not enumerate), no judgement of which of two accepted proposals is better, and no guarantee that a re-ask converges — measure the rounds on your own data: a model can trade one violation for another.

Fail

Fail(*reasons)

Bases: Claim

A check's False with its reasons: return Fail("Harold is busy 13:30 - 15:30"). Falsy (bool(Fail(...)) is False), so the check reads as False everywhere; the reasons are recorded with the check. Fail() with no reason is a plain False.

Failed dataclass

Failed(check: str, hard: bool, reasons: list, governs: bool = False)

A check that is False in a response: its name, whether it is hard, and its reasons.

Round dataclass

Round(index: int, proposal: Any = None, generated: Any = None, response: Any = None, accepted: bool = False, failed: list = list(), causes: list = list(), feedback: Any = None, error: str | None = None, said: str | None = None)

One round of a refinement: what was proposed, what the System said, and what goes back to the proposer.

reasons property

reasons

What is wrong with the proposal: the reasons of the failed hard checks that govern the question, else the causes.

Refinement dataclass

Refinement(question: str, rounds: list, accepted: bool, escalation: str | None, max_rounds: int, into: str | None, feedback_into: str | None, accept: Any = 'checks', feedback_fn: str = 'reasons')

The record of one refine() call. accepted, proposal (the accepted one, or None), result (the question's Result in the last round), escalation (why a person gets it), rounds; to_dict / from_dict; replay.

from_dict classmethod

from_dict(d, catalog=None)

A stored refinement back (responses restored with catalog's types, as Response.model_validate does).

replay

replay(system, accept=None, feedback=None, trust_models=False)

Re-check the whole loop: every round's trace replays under system (its own mismatches are listed), and the loop did what its record says — the round's acceptance follows from its response (accept: needed again when it was a function), the recorded failed checks and causes are those of the response, the feedback is what feedback gives for the round (with the default, the round's reasons; a custom one is checked only when passed again), each round's input carries the proposal recorded for it and the feedback of the rounds before it, nothing ran after an accepted round, and the escalation matches the outcome. → {"ok", "rounds", "mismatches": [(round, what, why)], "feedback": "checked" / "unchecked"}.

failed_checks

failed_checks(res, question=None)

The checks that are False in a response, in catalog order → [Failed]. question: only the checks in that question's flow.

causes

causes(res, question)

Why a question could not be decided: the errors of the parts that failed by themselves ("spec: ValueError: no slot in the plan"), not the steps that only lacked their inputs; else the answer's reason. [] when it was decided.

accepted

accepted(res, question, accept='checks')

Is a response's proposal accepted? accept: "checks" — every hard check that governs the question was evaluated and passed (the answer itself may still abstain, say for low confidence); an answer or a list / tuple / set of answers — the question's answer is one of them; a function of the Response → bool.

refine

refine(system, state, question, propose=None, *, into='proposal', rounds=3, accept='checks', feedback=None, feedback_into=None, history=None, store=True)

Propose → check → re-ask with the reasons → escalate (see the module docs) → a Refinement.

system: the System whose checks judge a proposal; state: the given facts; question: the question whose hard checks (or answer) decide acceptance. propose: (state, earlier rounds) → a proposal — a value or a Generated (Generator.proposer(...) builds one) — given to the System as the fact into; None: the System generates itself and reads the earlier rounds' feedback as the fact feedback_into (default "feedback"). rounds: proposals at most. accept: "checks" (default), an answer or answers, or a function of the Response. feedback: round → the text or the list of reasons the proposer is told (default: round.reasons). history: earlier rounds the proposer should see first (a Refinement's rounds, to continue it). store: whether a System with storage stores every round.