Skip to content

solvi.multi

Several models, one decision: cascade, vote and route.

Several models, one decision: a cascade, a vote and a route over decision parts (solvi.decide). The models propose; deterministic code over their proposals decides; every proposal is in the trace.

from solvi.multi import Cascade, Route, Vote
small = base.decision("team", "Which team?", "email", TEAMS)
large = big.decision("team", "Which team?", "email", TEAMS)

team = Cascade([small, large])            # ask the small model; the large one only when the small one escalates
team = Vote([small, large], rule="all")    # answer when they agree and each is sure enough; else escalate
team = Route({lambda email: len(email) > 2000: large}, default=small)     # code picks the model per input

cat.fn(team)                               # like any decision part: a fact, or `team.question(cat)` for an answer
team.act_guard(examples, max_risk=0.10)        # P(answered alone and wrong) ≤ 10%, for the combination as a whole

A combination is a catalog part like a DecisionPart: its value is one of the options, its provenance decided, its fingerprint covers every model's, and the trace records every proposal: extra["stages"] and extra["answered_by"] (cascade), extra["votes"] (vote), extra["route"] and extra["routed"] (route), and extra["calls"], the models called for this decision. The parts must answer the same question (kind and options; the task and the facts they read may differ). Combinations nest: Cascade([small, Vote([mid, large])]).

Thresholds. Uncalibrated, each part escalates by its own thresholds (act_threshold, escalate_below, min_margin). After act_guard, one threshold t applies to every part's signal (its act probability when its model gives one, else its calibrated confidence) — a one-dimensional family (scale="raw", the default); scale="rank" puts t on each part's rank among its own calibration signals instead, for models whose signals live on different scales (an LLM's confidence near 1 and an act probability spread over [0, 1]) — it can help or hurt depending on the data, so compare both. A cascade's loss is not monotone in t (a higher t can pass a question from a wrong small model to a right large one, or back), so conformal risk control runs on the loss monotonized from above — the maximum over thresholds ≥ t — which keeps the guarantee. A cascade saves cost where the small model is often sure; a vote lowers the error among the automatic answers at the price of answering less.

Combination

Combination(members, name=None, costs=None)

What Cascade, Vote and Route share: a catalog function over the union of the parts' facts that returns a Decision; the model recorded in the trace (fingerprint over every part's).

The decider protocol: a combination has every public method of a DecisionPart, with the same signature and result keys. decide / score / act_guard / calibrate_for / conformal / save_calibration / load_calibration act on the combination as a whole (one threshold shared by every part); fit / adapt / teach / reset / memory / remove_lora go to every part (a list per part, in leaves order, where the part returns one value); calls() counts the models called. What belongs to one part — adapt_lora, save_lora, load_lora (an adapter is one checkpoint's, for one question), budget, sections_k, long_key, long_input (each part reads long texts by its own long=), in_pass (a shared forward pass is for parts of one model) — raises NotImplementedError naming the part to call it on.

labels property

labels

The options as the models read them (the first part's; every part answers the same question).

adaptation property

adaptation

Every part's adaptation (adapt / fit / teach), in leaves order.

lora property

lora

Every part's LoRA adapter (or None), in leaves order.

parts property

parts

The members: DecisionParts and nested combinations.

question

question(cat, name=None, text=None, min_confidence=None, requires=None, require_evidence=False)

Make this combination a question's answer (as DecisionPart.question) → the Question.

same_question

same_question()

What the parts must agree on (kind, options, "not stated", rank k, number bins).

fingerprint

fingerprint()

A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).

text_of

text_of(facts)

The facts this combination reads, from a dict of facts (System.teach passes them back to teach).

decide

decide(text)

An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.

score

score(text)

The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.

candidates

candidates(d)

The conformal answer set of a decision (after conformal(...)), most probable first.

act_guard

act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')

Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.

scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.

groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.

The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or "shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n", "guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question" (models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part brings its own; scale says how they share one threshold). Every option after the examples is keyword-only; risk= is the 0.7 name of max_risk=.

calibrate_for

calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')

Choose the shared threshold for a target error rate among the answers the combination gives alone, on labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs; method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals, Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its fingerprint and clears its conformal set.

The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.

save_calibration

save_calibration(path)

Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.

load_calibration

load_calibration(path, groups=None, strict=True)

Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another question (strict=False loads it anyway). → self.

conformal

conformal(examples, coverage=0.9)

Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.

calls

calls()

What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's token count; usage() here was the 0.7 name of calls().

usage

usage()

Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.

teach

teach(text, correct)

One correction, absorbed by every part at once (an input, or Facts by name) → total ms.

fit

fit(examples, lam=1.0, folds=4)

Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].

adapt

adapt(texts)

Label-bias correction for every part (see DecisionPart.adapt) → [per part].

reset

reset()

Every member part's reset() (their adaptations and calibrations).

memory

memory(memory=None, **settings)

A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.

remove_lora

remove_lora()

Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.

adapt_lora

adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)

Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.

save_lora

save_lora(path)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

load_lora

load_lora(path, strict=True)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

budget

budget()

Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.

sections_k

sections_k()

Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.

long_key

long_key()

Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).

long_input

long_input(text)

Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.

in_pass

in_pass(siblings, args, names=None)

Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).

check_record

check_record(r)

A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).

Cascade

Cascade(members, name=None, costs=None)

Bases: Combination

Ask the parts in order; answer with the first whose decision does not escalate; escalate when all do (with the last part's answer as "would have answered"). The next model is asked only when the one before escalates, so a cheap model that is often sure saves the large model's calls — calls() and extra["calls"] count them; costs=[45, 137] (ms, per part) makes act_guard and usage report the expected cost; a member that is itself a combination takes its parts' costs itself (Vote([...], costs=[...])), so costs= with one raises.

labels property

labels

The options as the models read them (the first part's; every part answers the same question).

adaptation property

adaptation

Every part's adaptation (adapt / fit / teach), in leaves order.

lora property

lora

Every part's LoRA adapter (or None), in leaves order.

parts property

parts

The members: DecisionParts and nested combinations.

question

question(cat, name=None, text=None, min_confidence=None, requires=None, require_evidence=False)

Make this combination a question's answer (as DecisionPart.question) → the Question.

same_question

same_question()

What the parts must agree on (kind, options, "not stated", rank k, number bins).

fingerprint

fingerprint()

A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).

text_of

text_of(facts)

The facts this combination reads, from a dict of facts (System.teach passes them back to teach).

decide

decide(text)

An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.

score

score(text)

The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.

candidates

candidates(d)

The conformal answer set of a decision (after conformal(...)), most probable first.

act_guard

act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')

Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.

scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.

groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.

The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or "shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n", "guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question" (models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part brings its own; scale says how they share one threshold). Every option after the examples is keyword-only; risk= is the 0.7 name of max_risk=.

calibrate_for

calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')

Choose the shared threshold for a target error rate among the answers the combination gives alone, on labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs; method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals, Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its fingerprint and clears its conformal set.

The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.

save_calibration

save_calibration(path)

Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.

load_calibration

load_calibration(path, groups=None, strict=True)

Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another question (strict=False loads it anyway). → self.

conformal

conformal(examples, coverage=0.9)

Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.

calls

calls()

What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's token count; usage() here was the 0.7 name of calls().

usage

usage()

Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.

teach

teach(text, correct)

One correction, absorbed by every part at once (an input, or Facts by name) → total ms.

fit

fit(examples, lam=1.0, folds=4)

Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].

adapt

adapt(texts)

Label-bias correction for every part (see DecisionPart.adapt) → [per part].

reset

reset()

Every member part's reset() (their adaptations and calibrations).

memory

memory(memory=None, **settings)

A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.

remove_lora

remove_lora()

Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.

adapt_lora

adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)

Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.

save_lora

save_lora(path)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

load_lora

load_lora(path, strict=True)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

budget

budget()

Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.

sections_k

sections_k()

Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.

long_key

long_key()

Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).

long_input

long_input(text)

Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.

in_pass

in_pass(siblings, args, names=None)

Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).

check_record

check_record(r)

A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).

Vote

Vote(members, rule='all', name=None, costs=None)

Bases: Combination

Ask every part; answer when the rule holds — "all": every part proposes the same value, "majority": more than half do — and every agreeing part answers alone (its signal ≥ the threshold); otherwise escalate, listing the proposals. Parts of one model that can share a forward pass are asked in one pass. The probabilities are the mean of the parts'; the confidence the lowest among the agreeing parts'. Models of different families tend to disagree more usefully than a student and its teacher, which often make the same mistakes.

labels property

labels

The options as the models read them (the first part's; every part answers the same question).

adaptation property

adaptation

Every part's adaptation (adapt / fit / teach), in leaves order.

lora property

lora

Every part's LoRA adapter (or None), in leaves order.

parts property

parts

The members: DecisionParts and nested combinations.

question

question(cat, name=None, text=None, min_confidence=None, requires=None, require_evidence=False)

Make this combination a question's answer (as DecisionPart.question) → the Question.

same_question

same_question()

What the parts must agree on (kind, options, "not stated", rank k, number bins).

fingerprint

fingerprint()

A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).

text_of

text_of(facts)

The facts this combination reads, from a dict of facts (System.teach passes them back to teach).

decide

decide(text)

An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.

score

score(text)

The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.

candidates

candidates(d)

The conformal answer set of a decision (after conformal(...)), most probable first.

act_guard

act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')

Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.

scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.

groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.

The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or "shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n", "guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question" (models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part brings its own; scale says how they share one threshold). Every option after the examples is keyword-only; risk= is the 0.7 name of max_risk=.

calibrate_for

calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')

Choose the shared threshold for a target error rate among the answers the combination gives alone, on labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs; method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals, Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its fingerprint and clears its conformal set.

The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.

save_calibration

save_calibration(path)

Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.

load_calibration

load_calibration(path, groups=None, strict=True)

Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another question (strict=False loads it anyway). → self.

conformal

conformal(examples, coverage=0.9)

Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.

calls

calls()

What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's token count; usage() here was the 0.7 name of calls().

usage

usage()

Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.

teach

teach(text, correct)

One correction, absorbed by every part at once (an input, or Facts by name) → total ms.

fit

fit(examples, lam=1.0, folds=4)

Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].

adapt

adapt(texts)

Label-bias correction for every part (see DecisionPart.adapt) → [per part].

reset

reset()

Every member part's reset() (their adaptations and calibrations).

memory

memory(memory=None, **settings)

A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.

remove_lora

remove_lora()

Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.

adapt_lora

adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)

Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.

save_lora

save_lora(path)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

load_lora

load_lora(path, strict=True)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

budget

budget()

Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.

sections_k

sections_k()

Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.

long_key

long_key()

Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).

long_input

long_input(text)

Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.

in_pass

in_pass(siblings, args, names=None)

Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).

check_record

check_record(r)

A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).

Route

Route(routes, default, name=None, costs=None)

Bases: Combination

Pick one part per input by code: routes maps a predicate (a function of facts by name — its parameters are facts the route reads — returning true to take that part) or a fact name (its value is true) to a part; the first that holds picks, else default. Only the picked part's model is called; extra["route"] records which part and why, extra["routed"] its proposal. An input that is not Facts(...) (decide(text), the examples of act_guard): a state with the route's facts as keys gives them; otherwise a predicate with one parameter is called with the input itself, and a route keyed by a fact name — or a predicate of several facts — raises: a text is not the value of a fact. A fact that Facts(...) does not give raises too (it is not read as false).

labels property

labels

The options as the models read them (the first part's; every part answers the same question).

adaptation property

adaptation

Every part's adaptation (adapt / fit / teach), in leaves order.

lora property

lora

Every part's LoRA adapter (or None), in leaves order.

parts property

parts

The members: DecisionParts and nested combinations.

pick

pick(src)

The index of the member this input goes to.

question

question(cat, name=None, text=None, min_confidence=None, requires=None, require_evidence=False)

Make this combination a question's answer (as DecisionPart.question) → the Question.

same_question

same_question()

What the parts must agree on (kind, options, "not stated", rank k, number bins).

fingerprint

fingerprint()

A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).

text_of

text_of(facts)

The facts this combination reads, from a dict of facts (System.teach passes them back to teach).

decide

decide(text)

An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.

score

score(text)

The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.

candidates

candidates(d)

The conformal answer set of a decision (after conformal(...)), most probable first.

act_guard

act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')

Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.

scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.

groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.

The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or "shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n", "guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question" (models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part brings its own; scale says how they share one threshold). Every option after the examples is keyword-only; risk= is the 0.7 name of max_risk=.

calibrate_for

calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')

Choose the shared threshold for a target error rate among the answers the combination gives alone, on labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs; method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals, Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its fingerprint and clears its conformal set.

The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.

save_calibration

save_calibration(path)

Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.

load_calibration

load_calibration(path, groups=None, strict=True)

Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another question (strict=False loads it anyway). → self.

conformal

conformal(examples, coverage=0.9)

Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.

calls

calls()

What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's token count; usage() here was the 0.7 name of calls().

usage

usage()

Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.

teach

teach(text, correct)

One correction, absorbed by every part at once (an input, or Facts by name) → total ms.

fit

fit(examples, lam=1.0, folds=4)

Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].

adapt

adapt(texts)

Label-bias correction for every part (see DecisionPart.adapt) → [per part].

reset

reset()

Every member part's reset() (their adaptations and calibrations).

memory

memory(memory=None, **settings)

A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.

remove_lora

remove_lora()

Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.

adapt_lora

adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)

Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.

save_lora

save_lora(path)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

load_lora

load_lora(path, strict=True)

Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.

budget

budget()

Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.

sections_k

sections_k()

Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.

long_key

long_key()

Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).

long_input

long_input(text)

Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.

in_pass

in_pass(siblings, args, names=None)

Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).

check_record

check_record(r)

A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).