solvi.multi¶
Several models, one decision: cascade, vote and route.
Several models, one decision: a cascade, a vote and a route over decision parts (solvi.decide). The models propose; deterministic code over their proposals decides; every proposal is in the trace.
from solvi.multi import Cascade, Route, Vote
small = base.decision("team", "Which team?", "email", TEAMS)
large = big.decision("team", "Which team?", "email", TEAMS)
team = Cascade([small, large]) # ask the small model; the large one only when the small one escalates
team = Vote([small, large], rule="all") # answer when they agree and each is sure enough; else escalate
team = Route({lambda email: len(email) > 2000: large}, default=small) # code picks the model per input
cat.fn(team) # like any decision part: a fact, or `team.question(cat)` for an answer
team.act_guard(examples, max_risk=0.10) # P(answered alone and wrong) ≤ 10%, for the combination as a whole
A combination is a catalog part like a DecisionPart: its value is one of the options, its provenance decided, its
fingerprint covers every model's, and the trace records every proposal: extra["stages"] and extra["answered_by"]
(cascade), extra["votes"] (vote), extra["route"] and extra["routed"] (route), and extra["calls"], the models
called for this decision. The parts must answer the same question (kind and options; the task and the facts they read
may differ). Combinations nest: Cascade([small, Vote([mid, large])]).
Thresholds. Uncalibrated, each part escalates by its own thresholds (act_threshold, escalate_below, min_margin). After
act_guard, one threshold t applies to every part's signal (its act probability when its model gives one, else its
calibrated confidence) — a one-dimensional family (scale="raw", the default); scale="rank" puts t on each part's rank
among its own calibration signals instead, for models whose signals live on different scales (an LLM's confidence
near 1 and an act probability spread over [0, 1]) — it can help or hurt depending on the data, so compare both. A
cascade's loss is not monotone in t (a higher t can pass a question from a wrong small model to a right large one, or
back), so conformal risk control runs on the loss monotonized from above — the maximum over thresholds ≥ t — which
keeps the guarantee. A cascade saves cost where the small model is often sure; a vote lowers the error among the
automatic answers at the price of answering less.
Combination ¶
What Cascade, Vote and Route share: a catalog function over the union of the parts' facts that returns a Decision; the model recorded in the trace (fingerprint over every part's).
The decider protocol: a combination has every public method of a DecisionPart, with the same signature and result keys. decide / score / act_guard / calibrate_for / conformal / save_calibration / load_calibration act on the combination as a whole (one threshold shared by every part); fit / adapt / teach / reset / memory / remove_lora go to every part (a list per part, in leaves order, where the part returns one value); calls() counts the models called. What belongs to one part — adapt_lora, save_lora, load_lora (an adapter is one checkpoint's, for one question), budget, sections_k, long_key, long_input (each part reads long texts by its own long=), in_pass (a shared forward pass is for parts of one model) — raises NotImplementedError naming the part to call it on.
labels
property
¶
The options as the models read them (the first part's; every part answers the same question).
question ¶
Make this combination a question's answer (as DecisionPart.question) → the Question.
same_question ¶
What the parts must agree on (kind, options, "not stated", rank k, number bins).
fingerprint ¶
A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).
text_of ¶
The facts this combination reads, from a dict of facts (System.teach passes them back to teach).
decide ¶
An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.
score ¶
The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.
candidates ¶
The conformal answer set of a decision (after conformal(...)), most probable first.
act_guard ¶
act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')
Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.
scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.
groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.
The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or
"shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n",
"guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question"
(models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part
brings its own; scale says how they share one threshold). Every option after the examples is keyword-only;
risk= is the 0.7 name of max_risk=.
calibrate_for ¶
calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')
Choose the shared threshold for a target error rate among the answers the combination gives alone, on
labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's
signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration
decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs;
method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for
inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals,
Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the
target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its
fingerprint and clears its conformal set.
The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.
save_calibration ¶
Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.
load_calibration ¶
Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another
question (strict=False loads it anyway). → self.
conformal ¶
Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.
calls ¶
What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in
leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's
token count; usage() here was the 0.7 name of calls().
usage ¶
Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.
teach ¶
One correction, absorbed by every part at once (an input, or Facts by name) → total ms.
fit ¶
Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].
memory ¶
A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.
remove_lora ¶
Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.
adapt_lora ¶
adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)
Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.
save_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
load_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
budget ¶
Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.
sections_k ¶
Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.
long_key ¶
Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).
long_input ¶
Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.
in_pass ¶
Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).
check_record ¶
A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).
Cascade ¶
Bases: Combination
Ask the parts in order; answer with the first whose decision does not escalate; escalate when all do (with the
last part's answer as "would have answered"). The next model is asked only when the one before escalates, so a
cheap model that is often sure saves the large model's calls — calls() and extra["calls"] count them;
costs=[45, 137] (ms, per part) makes act_guard and usage report the expected cost; a member that is itself a
combination takes its parts' costs itself (Vote([...], costs=[...])), so costs= with one raises.
labels
property
¶
The options as the models read them (the first part's; every part answers the same question).
question ¶
Make this combination a question's answer (as DecisionPart.question) → the Question.
same_question ¶
What the parts must agree on (kind, options, "not stated", rank k, number bins).
fingerprint ¶
A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).
text_of ¶
The facts this combination reads, from a dict of facts (System.teach passes them back to teach).
decide ¶
An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.
score ¶
The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.
candidates ¶
The conformal answer set of a decision (after conformal(...)), most probable first.
act_guard ¶
act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')
Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.
scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.
groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.
The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or
"shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n",
"guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question"
(models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part
brings its own; scale says how they share one threshold). Every option after the examples is keyword-only;
risk= is the 0.7 name of max_risk=.
calibrate_for ¶
calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')
Choose the shared threshold for a target error rate among the answers the combination gives alone, on
labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's
signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration
decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs;
method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for
inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals,
Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the
target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its
fingerprint and clears its conformal set.
The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.
save_calibration ¶
Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.
load_calibration ¶
Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another
question (strict=False loads it anyway). → self.
conformal ¶
Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.
calls ¶
What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in
leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's
token count; usage() here was the 0.7 name of calls().
usage ¶
Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.
teach ¶
One correction, absorbed by every part at once (an input, or Facts by name) → total ms.
fit ¶
Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].
memory ¶
A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.
remove_lora ¶
Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.
adapt_lora ¶
adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)
Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.
save_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
load_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
budget ¶
Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.
sections_k ¶
Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.
long_key ¶
Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).
long_input ¶
Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.
in_pass ¶
Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).
check_record ¶
A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).
Vote ¶
Bases: Combination
Ask every part; answer when the rule holds — "all": every part proposes the same value, "majority": more than half do — and every agreeing part answers alone (its signal ≥ the threshold); otherwise escalate, listing the proposals. Parts of one model that can share a forward pass are asked in one pass. The probabilities are the mean of the parts'; the confidence the lowest among the agreeing parts'. Models of different families tend to disagree more usefully than a student and its teacher, which often make the same mistakes.
labels
property
¶
The options as the models read them (the first part's; every part answers the same question).
question ¶
Make this combination a question's answer (as DecisionPart.question) → the Question.
same_question ¶
What the parts must agree on (kind, options, "not stated", rank k, number bins).
fingerprint ¶
A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).
text_of ¶
The facts this combination reads, from a dict of facts (System.teach passes them back to teach).
decide ¶
An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.
score ¶
The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.
candidates ¶
The conformal answer set of a decision (after conformal(...)), most probable first.
act_guard ¶
act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')
Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.
scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.
groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.
The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or
"shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n",
"guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question"
(models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part
brings its own; scale says how they share one threshold). Every option after the examples is keyword-only;
risk= is the 0.7 name of max_risk=.
calibrate_for ¶
calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')
Choose the shared threshold for a target error rate among the answers the combination gives alone, on
labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's
signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration
decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs;
method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for
inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals,
Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the
target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its
fingerprint and clears its conformal set.
The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.
save_calibration ¶
Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.
load_calibration ¶
Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another
question (strict=False loads it anyway). → self.
conformal ¶
Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.
calls ¶
What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in
leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's
token count; usage() here was the 0.7 name of calls().
usage ¶
Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.
teach ¶
One correction, absorbed by every part at once (an input, or Facts by name) → total ms.
fit ¶
Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].
memory ¶
A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.
remove_lora ¶
Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.
adapt_lora ¶
adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)
Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.
save_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
load_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
budget ¶
Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.
sections_k ¶
Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.
long_key ¶
Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).
long_input ¶
Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.
in_pass ¶
Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).
check_record ¶
A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).
Route ¶
Bases: Combination
Pick one part per input by code: routes maps a predicate (a function of facts by name — its parameters are
facts the route reads — returning true to take that part) or a fact name (its value is true) to a part; the first
that holds picks, else default. Only the picked part's model is called; extra["route"] records which part and
why, extra["routed"] its proposal. An input that is not Facts(...) (decide(text), the examples of act_guard): a
state with the route's facts as keys gives them; otherwise a predicate with one parameter is called with the input
itself, and a route keyed by a fact name — or a predicate of several facts — raises: a text is not the value of a
fact. A fact that Facts(...) does not give raises too (it is not read as false).
labels
property
¶
The options as the models read them (the first part's; every part answers the same question).
question ¶
Make this combination a question's answer (as DecisionPart.question) → the Question.
same_question ¶
What the parts must agree on (kind, options, "not stated", rank k, number bins).
fingerprint ¶
A hash of the combination: its kind and rule, every member's fingerprint, and its calibration (threshold, guarantee, conformal set).
text_of ¶
The facts this combination reads, from a dict of facts (System.teach passes them back to teach).
decide ¶
An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.
score ¶
The probabilities the combination answers with (see decide) → {option: probability}; a list → a list.
candidates ¶
The conformal answer set of a decision (after conformal(...)), most probable first.
act_guard ¶
act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1, scale='raw')
Answer alone only as far as a guarantee allows, for the combination as a whole: on labelled examples of your stream [(input, correct)] (an input is what every part reads, or Facts(...) by name) every part is asked, and one threshold t shared by every part is chosen by conformal risk control so that P(answered alone AND wrong) ≤ risk for inputs like the examples — a share of all questions. A cascade's loss is not monotone in t, so it is monotonized from above (the maximum over the thresholds ≥ t) before the choice, which keeps the guarantee. Replaces the parts' own thresholds inside this combination (the parts themselves are not changed); changes the combination's fingerprint and clears its conformal sets (call conformal afterwards). Too few or too hard examples → everything escalates (threshold inf). → {"threshold", "answered", "error" (among the answered), "risk" (answered and wrong, on the examples), "n", "guarantee", "calls_per_question" (models called per question), "cost" (with costs=), "scale", and for a cascade "answered_by" (the share each stage answered) and "warnings" when a stage answers alone on less than 5% of the examples (the cascade is then no better than one model) or a part's signal does not separate right from wrong answers (also a UserWarning), "promise" (in words: of all inputs, not of the answered ones — "error" is not bounded)}.
scale: what the shared threshold is on. "raw" (default): the parts' signals themselves (the act probability when the model gives one, else the calibrated confidence). When the scales differ — an act probability spread over [0, 1], an LLM's confidence near 1 — one raw threshold effectively fits one model and the combination behaves like that model alone, which is often the stronger one. "rank" (opt-in): each part's signal is replaced by its rank among that part's own signals on the calibration examples (the share of them ≤ it), so every part can take part. It helps where a stage never answers on the raw scale and can lower the share answered alone elsewhere; the guarantee holds either way. Compare both on held-out calibration data. The rank uses the calibration inputs, not their labels (the guarantee then holds up to a term of order 1/n). The sorted calibration signals of each part (at most MAX_RANKS = 1024, evenly spaced by order when there are more examples) are kept in the combination and in its calibration file.
groups, min_group, delta: one shared threshold per group, as DecisionPart.act_guard(groups=...) — on the same monotonized loss, so the promise holds within every group; the group facts join the combination's inputs. Adds "groups" to the result.
The decider protocol: the signature and the keys of DecisionPart.act_guard — "signal" ("shared", or
"shared-rank" with scale="rank"; also guarantee["signal"]), "threshold", "answered", "error", "risk", "n",
"guarantee", "promise", "base_error", "must_escalate_at_least", "warnings", "groups" — plus "calls_per_question"
(models called per question; "calls" in 0.7), "cost", "scale", "answered_by". signal: "auto" only (each part
brings its own; scale says how they share one threshold). Every option after the examples is keyword-only;
risk= is the 0.7 name of max_risk=.
calibrate_for ¶
calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1, scale='raw')
Choose the shared threshold for a target error rate among the answers the combination gives alone, on
labelled examples [(input, correct)] — as DecisionPart.calibrate_for, with one threshold t for every part's
signal (on scale, as in act_guard). method="empirical": the lowest threshold at which the calibration
decisions the combination answers alone are wrong at most max_error of the time — no guarantee on new inputs;
method="ltt" (learn-then-test): the error among the answered is ≤ max_error with probability ≥ 1 − delta for
inputs like the examples (a binomial test at every threshold of a grid of at most 64 quantiles of the signals,
Bonferroni over the grid, which holds although a cascade's error is not monotone in t). No threshold reaches the
target → everything escalates (inf). Replaces the parts' own thresholds inside the combination, changes its
fingerprint and clears its conformal set.
The decider protocol: the signature and the keys of DecisionPart.calibrate_for — "signal" ("shared" or "shared-rank"), "threshold", "answered", "error", "n", "max_error", "method", "guarantee" — plus "calls_per_question" and "scale". signal: "auto" only (each part brings its own). Every option after the examples is keyword-only; error= is the 0.7 name of max_error=.
save_calibration ¶
Write the combination's calibration (the shared threshold, per group too, the guarantee, the conformal set) with the question and every member's fingerprint to a JSON file (solvi.calibfile). → path.
load_calibration ¶
Apply a file written by save_calibration / solvi calibrate; refuses one made for other members or another
question (strict=False loads it anyway). → self.
conformal ¶
Conformal answer sets for the combination's decisions (the probabilities it answers with: the answering stage's for a cascade, the mean of the parts' for a vote), from labelled examples at its current threshold — so call it after act_guard. Every decision then carries extra["candidates"]; an escalation lists them. Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.
calls ¶
What the combination cost since it was made: {"asked" (decisions), "calls" (models called, per part, in
leaves order), "calls_per_question" (models called per decision), "cost" (with costs=)}. usage is a scorer's
token count; usage() here was the 0.7 name of calls().
usage ¶
Deprecated (removed in 0.9): calls() — "per_question" is "calls_per_question" there.
teach ¶
One correction, absorbed by every part at once (an input, or Facts by name) → total ms.
fit ¶
Few-shot "S" for every part (see DecisionPart.fit) → [Adaptation].
memory ¶
A memory of corrected cases for every part (DecisionPart.memory with these settings) → [CorrectionMemory], in leaves order; False detaches every part's → None. Inside a combination a part's memory only checks (it can escalate, never answer). A memory belongs to one part, so an existing one is attached on that part (part.memory(mem)), not here — that raises.
remove_lora ¶
Every part's remove_lora() → [the removed adapter's hash or None], in leaves order. The combination's own threshold was fitted on the parts with their adapters: calibrate it again.
adapt_lora ¶
adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)
Not for a combination (raises NotImplementedError): an adapter is trained on one checkpoint's encoder for one question, and its holdout recalibrates that part's own threshold, which a combination replaces with its shared one. Train it on the part, then calibrate the combination.
save_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
load_lora ¶
Not for a combination (raises NotImplementedError): an adapter file holds one part's adapter.
budget ¶
Not for a combination (raises NotImplementedError): each part reads an input by its own model's length.
sections_k ¶
Not for a combination (raises NotImplementedError): each part retrieves by its own long= setting.
long_key ¶
Not for a combination (raises NotImplementedError): each part reads long texts by its own long= setting (every part's is in the combination's fingerprint).
long_input ¶
Not for a combination (raises NotImplementedError): each part reads a long text by its own long= setting.
in_pass ¶
Not for a combination (raises NotImplementedError): the runtime's shared forward pass is for parts of one model; the strategist never puts a combination in one (plan_batches).
check_record ¶
A recorded decision of this combination, without re-running any model: does the recorded answer follow from the recorded proposals by this combination's rule? → [reason] (empty: it does).