Skip to content

solvi.decide

Decisions with a model: the decider that answers typed questions about a text or a state.

Decisions with a model: a cross-encoder that answers typed questions about a text or a state ("decider").

Types declare questions, the model proposes, checks decide. A question's kind comes from its type (solvi.typed):

choice Literal[...] / an Enum one option (softmax); "other" / "none" may be an abstain threshold multi list[Literal[...]] every option that applies (a sigmoid per option) score Scale[...] (2–10 levels) ordered levels (softmax; the value is the median, the expected level is recorded) noul bool / Literal["yes", "no"] yes or no

The decider (solvi-decide, ModernBERT) reads one sequence per question — or, when the checkpoint says it can, several questions about the same input in one sequence:

[mode] task[opt] option 1[opt] option 2 ... [mode] task 2[opt] ... [SEP] input

and gives one logit per option marker (and, with an act head, one "act" logit per question). The input is a text, or a state (a dict, a list, a pydantic model) serialized by state_text — the one serialization the training side uses too (see docs/decide_format.md, which also defines the checkpoint's capability fields).

In solvi a decider is a catalog part like any other: model.decision(...) returns a function that returns solvi.Decision(value, probs). The value is one of the declared options by construction, the part's provenance is decided and the model's identity (weights, calibration and adaptation of this part) is in the trace, so the closed set, min_confidence, constraints with joint decoding, hard checks, the audit and the stats apply unchanged. A decision the model escalates (its act head, or a calibrated confidence below the part's escalate_below) is rejected like an unsure one: the fact is missing, the answer abstains, and the audit and stats say "model escalated" / "low confidence".

On top of the raw logits, per question (task, options, kind): - label-bias correction without labels (adapt): the mean logit of each option over unlabelled inputs of the domain is subtracted before the softmax (the decider likes some labels regardless of the text); - few-shot adaptation "S" (fit, teach): a shift and a shared scale fitted on k labelled examples (L-BFGS), with a temperature fitted on out-of-fold predictions, so confidences are calibrated; teach updates the shift at once. The shift is per option (choice, multi), a tilt / spread over the levels (score) or one yes−no bias (noul); - "other" / "none" as an abstain threshold: such an option is not scored by the model; it is chosen when the best real option's calibrated probability is below a threshold (fitted on labelled examples that include it, else the default); - calibrate_for(examples, max_error=0.05): the escalation threshold for a target error rate.

Backends: "torch" (solvi[model]) or "onnx" (solvi[onnx]: onnxruntime + tokenizers, no torch), or any object with logits(items) (and optionally logits_pass(passes)) — tests, other models: see DecideModel.

Item dataclass

Item(task: str, options: tuple, descriptions: tuple | None, text: str, multi: bool = False, kind: str = '', pointer: bool = False, unknown: bool = False, max_len: int = 0)

One question to score: the scorer returns one logit (or a row of logits, one column per mode) per option, or {"logits": ..., "act": logit} when it has an act head.

Logits

Bases: ndarray

A question's logits [K] with an l14g checkpoint's extra outputs: unknown (the "not stated" logit) and pointer (decoded: {"null": p(null span), "spans": [(p, start, end, text)]}, most probable first) — and from a scorer that can fail on one question (solvi.llm): escalate (why the output is not usable: the decision escalates with it), transient (not cached: ask again next time) and info (recorded in the decision's extra).

Pass dataclass

Pass(text: str, items: tuple)

Several questions about one input, scored in one forward pass (checkpoints with multi-question support).

BlockUnsupported

Bases: ValueError

The scorer cannot run the block layout at all (e.g. an ONNX export without its inputs): one question per sequence.

LongInputWarning

Bases: UserWarning

long="full": a checkpoint forced to read long inputs it was not trained on, or whole long texts read on a CPU. Without long=: an input that does not fit the pass was read cut (the decision's extra["truncated"] says how much).

OnnxScorer

OnnxScorer(path, max_len=512, device=None, onnx_file=None, bs=16, caps=None)

Bases: _NetScorer

onnxruntime session over onnx/model_*.onnx (inputs input_ids, attention_mask; output logits [B, L, C]). The block layout needs an export with the inputs position_ids, full_attention_mask and sliding_attention_mask ([B, 1, L, L] bool); without them several questions fall back to one per pass.

TorchScorer

TorchScorer(path, max_len=512, device=None, bs=16, caps=None)

Bases: _NetScorer

The checkpoint's encoder (transformers AutoModel from config.json) + the option head, weights from model.safetensors. The block layout passes per-layer-type attention masks and position ids (sdpa attention).

inputs

inputs(ids, att, pids=None, masks=None, device=None)

Packed arrays (see pack) → the network's inputs as tensors on device (default: the scorer's).

Adaptation dataclass

Adaptation(bias: list | None = None, n_unlabelled: int = 0, scale: float | None = None, shift: list | None = None, temperature: float = 1.0, other_threshold: float | None = None, n_labelled: int = 0, examples: list = list())

What solvi learned for one question (task, options, kind): the label-bias correction, the few-shot shift / scale, the temperature and the "other" threshold. Part of the fingerprint of every decision that uses it.

params

params()

Everything that changes the output (the examples only through the fitted parameters).

Facts

Bases: dict

An example input given as facts by name (Facts(email=..., tier=...)): each part reads its own facts, a route's predicates and a grouping (act_guard(groups=...)) read theirs. Any other input is the one input every part reads (a text or a state).

GroupBy

GroupBy(by)

Which group an input belongs to, for thresholds per group (act_guard(groups=...)): a fact name ("domain"), a list of fact names — a hierarchy, top first (["domain", "task"]) — or a function whose parameters are fact names and which returns a group or a path (domain, task). A function with one parameter also takes an input that is not given as facts (a text or a state): it is called with the input itself. A state (dict) input gives facts by its keys.

path

path(vals=None, raw=None)

The input's group path (a tuple), or None when the input does not give it.

DecisionPart

DecisionPart(model, name, task, text_fact, options, descriptions=None, multi=False, other=None, *, kind=None, as_bool=False, escalate_below=None, act_threshold=None, use_act=None, score_value='median', option_order='canonical', permutations=4, min_margin=None, long=None, top_k=None, rerank=False, perturb=0, retrieve_query=None, **prim)

A decider bound to one question (name, task, options, kind, the facts it reads). Callable as a catalog function; it is also the model recorded in the trace: fingerprint() covers the checkpoint and this question's adaptation and thresholds only, so teaching one decision does not mark the others as changed.

deterministic property

deterministic

Does the model give the same output for the same input (replay re-runs it)? False for an LLM (solvi.llm): replay then checks the recorded output instead.

options property

options

The values this decision can take (a bool question: True, False; with "not stated": also Unknown; None for a number or a span).

labels property

labels

The options as the model reads them (a bool question: "yes", "no").

lora property

lora

This question's LoRA adapter (solvi.lora.LoraAdapter; experimental: adapt_lora) or None.

text_of

text_of(facts)

The input this part reads, from a dict of facts: text facts as they are (joined by new lines), a state by state_text (several facts: {fact: value}).

in_pass

in_pass(siblings, args, names=None)

Called by the runtime for a part in a shared pass (flow.batches): score every sibling's question about the same input in one forward pass (cached, so the siblings read it) and return this part's decision; the decision's extra records the pass.

candidates

candidates(d)

The conformal answer set of a decision (after conformal(...)): the answers that cannot be ruled out at the calibrated coverage, most probable first — a short list for the person who handles an escalation. Never empty: when no answer passes (an unsure decision — the one that escalates), the most probable answer is listed; a larger set only covers more.

long_key

long_key()

How this part reads long texts, as its fingerprint records it: None, ("retrieve", top_k, rerank) or ("full", max_len_long, top_k, rerank) — top_k resolved (top_k=None: from the budget).

sections_k

sections_k()

The sections retrieve reads: top_k, or with top_k=None budget / 170 (at least 3) — sections of ≈ 170 tokens.

long_input

long_input(text)

The text a decision reads for an input that does not fit an ordinary pass: the whole text (long="full", when it fits max_len_long), else the window of its retrieved sections.

budget

budget()

The tokens of input this decision can read in one pass: max_len (long="full": max_len_long) minus its question.

decide

decide(text)

An input (a text, a state, or Facts by name) → Decision; a list of inputs → a list.

score

score(text)

The probabilities a decision answers with (see decide) → {option: probability}; a list → a list.

adapt

adapt(texts)

Label-bias correction from unlabelled inputs of the domain (see DecideModel.adapt).

fit

fit(examples, lam=1.0, folds=4)

Few-shot "S" from [(input, correct)] (see DecideModel.fit).

teach

teach(text, correct)

One correction, absorbed at once (see DecideModel.teach). → ms.

calls

calls()

What the part cost since it was made, as a combination's calls(): {"asked" (decisions made by the part on its own; inside a combination they count there), "calls" ({"0:": the model's calls}), "calls_per_question" (1 per decision: one model)}. A scorer's token count is its own usage.

adapt_lora

adapt_lora(examples, *, r=8, epochs=6, holdout=None, seed=0, device=None, lr=0.0003, max_risk=0.1, signal='confidence', max_updates=400)

Experimental: train a small LoRA adapter on the decider's encoder for this question, from labelled examples [(input, correct)] — for solvi-base (the torch backend, pip install "solvi[lora]") and about 100 examples or more. Below that, try fit (and System.fit for questions without a model) first: they take milliseconds; the adapter can keep improving where fit levels off. Compare the two on held-out labels.

Cost: minutes on a CPU; the time is estimated from the first update and reported (a LoraWarning) before training. r: the adapter's rank (alpha = 2r); epochs: passes over the examples (updates of 8 examples, between 40 and max_updates); lr: the learning rate; seed: the adapter's initialization and the order of the examples — the same seed gives the same adapter on a CPU; device: None — the device the decider runs on.

After training the adapter is active for this question only (other questions of the same model are not affected), its hash is in the part's fingerprint and in every decision's extra["lora"], and the question's earlier adaptation (adapt / fit / teach) and escalation thresholds are cleared: they were fitted on the model without it. Confidences after LoRA are overconfident, so escalation must be recalibrated on labels NOT used for training: holdout — a list of [(input, correct)], a share of the examples (0.25) or a number of them split off (by the seed) — runs act_guard(holdout, risk, signal) after training and reports the held-out accuracy before and after; without one a LoraWarning says so (call act_guard yourself; a few hundred labels is typical). Refuses a decider that is not a torch encoder (ONNX: load it with backend="torch"; an LLM or a rule has no weights to adapt), one larger than solvi-base (use tools/adapt_lora_gpu.py on a GPU and load_lora), and rank / number / span questions. Keep it with save_lora / load_lora or save_calibration (the adapter is written next to the calibration file); remove_lora rolls back. → {"adapter", "k", "updates", "seconds", "estimate_seconds", "size_mb", "device", "holdout": {"n", "accuracy_before", "accuracy_after", "act_guard"} or None, "cleared", "experimental": True}.

remove_lora

remove_lora()

Roll back adapt_lora / load_lora: the adapter is removed from the model (the question is answered by the checkpoint as before, to the bit) and the part's adaptation and escalation thresholds return to what they were before the first adapter in this process (after a load_lora: they are cleared — they were fitted with the adapter). → the removed adapter's hash, or None when there was none.

save_lora

save_lora(path)

Write this question's adapter (a .safetensors file with its config, the question and the checkpoint it was trained on) → path. save_calibration also writes it, next to the calibration file.

load_lora

load_lora(path, strict=True)

Load an adapter written by save_lora (or tools/adapt_lora_gpu.py) for this question: afterwards the part answers exactly as right after training. Refuses (ValueError) an adapter for another question or checkpoint unless strict=False; needs the torch backend and peft (solvi[lora]). Clears the question's adaptation and thresholds like adapt_lora (load the calibration after it). → self.

calibrate_for

calibrate_for(examples, *, max_error=0.05, signal='auto', method='empirical', delta=0.1)

Choose the escalation threshold for a target error rate among the answers given alone, on labelled examples [(input, correct)]. method="empirical": the lowest threshold at which the calibration decisions it lets through are wrong at most error of the time — no guarantee on new inputs (on another data set the error can be several times the target); method="ltt" (learn-then-test): the error among the answered is ≤ error with probability ≥ 1 − delta for inputs like the examples — a strong promise, so it often lets nothing through (it tests at most 64 thresholds, quantiles of the calibration signals: calibration.ltt_grid). signal: "act" (the model's act probability → act_threshold), "confidence" (the calibrated confidence → escalate_below) or "auto" (act when the model has an act head). On "confidence" with a model that has an act head (and a part not made with use_act=False) the checkpoint's act threshold keeps escalating: the examples it escalates count as escalated here, so the numbers returned are what the part does. No threshold reaches the target → everything escalates (inf). Changes the part's fingerprint. → {"signal", "threshold", "answered" (the share answered alone on the examples), "error" (among them), "n", "max_error", "method", "guarantee"}. For a guarantee on the share of all questions answered wrongly, see act_guard. Every option after the examples is keyword-only; error= is the 0.7 name of max_error=, and the result's 0.7 keys "coverage" / "target_error" still read (deprecated).

act_guard

act_guard(examples, *, max_risk=0.1, signal='auto', groups=None, min_group=100, delta=0.1)

Answer alone only as far as a guarantee allows (conformal risk control), from labelled examples of your own stream [(input, correct)] — a few hundred is typical: the escalation threshold is set so that, for inputs like the examples, P(answered alone AND wrong) ≤ risk — a share of all questions (answered or escalated), not of the answered ones. correct: an option (a value), Unknown for "not stated", for a span question the passage's text (or a Quote: its text is compared), for a ranking the order. It holds for your stream, not under a shift of domain: recalibrate when the inputs change. Too few or too hard examples → everything escalates (threshold inf). Feasibility: when the model is wrong on a share μ > risk of the examples, any rule must escalate at least (μ − risk) / (1 − risk) of the inputs ("must_escalate_at_least"; arXiv 2606.29054) — a better signal can only get closer to that bound. Changes the part's fingerprint; the trace of every decision records the promise. The error among the answers given alone is not bounded (calibrate_for(method="ltt") bounds it): with few answered it can be far above risk. A signal that does not separate right from wrong answers (solvi.calibration.separation: AUROC not above chance at the 5% level) keeps the promise only by escalating, and is warned about (UserWarning, "warnings"). → {"signal", "threshold", "answered" (share answered alone on the examples), "error" (among them), "risk" (answered and wrong, on the examples), "n", "guarantee", "promise" (in words, with that error), "base_error", "must_escalate_at_least", "warnings" when there are any}.

groups: a threshold per group — a fact name ("domain"), a hierarchy of fact names (["domain", "task"]) or a function of facts returning a group or a path (see GroupBy); the examples then give those facts (Facts(...) or a state with those keys). The promise over the whole stream allows a hard group to be answered wrongly far more often than risk; per group it holds inside each: every group with at least min_group examples gets its own threshold, a smaller one is pooled with the rest of its parent (whose threshold is calibrated on exactly those examples), the rest of the stream takes what is left. delta=0.10: with probability ≥ 90% over the examples, P(answered alone and wrong | group) ≤ risk in every group at once (a binomial bound per group at delta divided by the number of groups — Bonferroni; after HG-CRC, arXiv 2607.24562); delta=None: conformal risk control per group (each group on average). Every decision records its group and the group whose threshold applied; an input that does not give its group escalates. The group facts join the part's inputs: register the part in a catalog after act_guard. Adds "groups" ({path: {"threshold", "n", "answered", "error", "risk", "pooled"}}) to the result; "threshold" is then the rest of the stream's — inf when every example is in a group with its own threshold, whatever "answered" says: read the thresholds per group.

The decider protocol: a combination (solvi.multi) takes the same act_guard(examples, *, max_risk, signal, groups, min_group, delta) and returns the same keys. Every option after the examples is keyword-only; risk= is the 0.7 name of max_risk=.

conformal

conformal(examples, coverage=0.9)

Conformal answer sets from labelled examples [(input, correct)]: afterwards every decision carries extra["candidates"] — the answers that cannot be ruled out, which contain the right one with probability ≥ coverage for inputs like the examples (score and number questions: one contiguous interval) — and an escalation's message lists them for the person who takes over. It does not change what is answered alone (see act_guard). Choice, yes/no, score and number questions. → {"coverage", "quantile", "n", "mean_size"}.

save_calibration

save_calibration(path)

Write this decision's calibration — the escalation thresholds (per group too), the guarantee record, the conformal set — with the question and the fingerprint of the model and adaptation it was fitted on, to a JSON file (solvi.calibfile; solvi calibrate writes the same). → path.

load_calibration

load_calibration(path, groups=None, strict=True)

Apply a calibration file written by save_calibration / solvi calibrate: afterwards the part escalates, and records its guarantee, exactly as right after calibrating (the same fingerprint). Refuses (ValueError) a file made for another question, another checkpoint or another adaptation of this question — strict=False loads it anyway. Thresholds per group by a function: pass it again as groups=. Call it before registering the part in a catalog when the calibration has groups (the group facts join the part's inputs). While solvi calibrate loads a catalog, calibration files are not applied (the part is calibrated afresh). → self.

memory

memory(memory=None, **settings)

A memory of corrected cases consulted on every decision of this part (solvi.memory.CorrectionMemory): its proposal, the cases it rests on and its fingerprint go into extra["memory"]; mode="check" (default) escalates when similar corrected cases say another answer, mode="answer" may also answer where the part escalated by its own threshold. settings: k, radius, min_strength, min_agreement, text, text_weight, mode. memory: an existing CorrectionMemory of this part to attach; False detaches. Its fingerprint is part of the part's. → the memory.

question

question(cat, name=None, text=None, min_confidence=None, requires=None, require_evidence=False)

Make this decision the answer of a question: registers it as the question's rule (cat.rule(name)(self)) and returns the Question — choice, multi, ordinal (score) or yes_no (noul), with the option descriptions. System.teach on that question teaches this decision.

DecideModel

DecideModel(scorer, meta=None, model_id=None, path=None, backend=None, cache_size=4096, multi_question=None, act=None, max_len_long=None)

A decider: an input (a text or a state) + a question (task, options, kind) → a probability per option.

DecideModel.load(path_or_hf_id, device=None, backend="auto") loads a solvi-decide checkpoint (a folder with config.json, tokenizer.json, solvi_decide.json, model.safetensors and/or onnx/model_fp16.onnx; or a Hugging Face id). backend: "torch", "onnx" or "auto" (ONNX when the file and onnxruntime are there, else torch). solvi_decide.json declares what the checkpoint can do (modes, act head, several questions per pass, temperatures, thresholds: docs/decide_format.md).

DecideModel(scorer, meta=None, model_id=...) wraps any object with logits(items) → per Item an array [K] or [K, C] (a column per mode: [:, 0] choose-one, [:, 1] multi-label, more as meta["columns"] says) or {"logits": array, "act": logit}; optionally logits_pass(passes) → per Pass a list of those (one per question) and fingerprint().

batchable property

batchable

Can several questions about one input share a forward pass (declared by the checkpoint, and the scorer has logits_pass)?

has_act property

has_act

Does the checkpoint give an act / escalate signal per question?

has_not_stated property

has_not_stated

Does the checkpoint give a "not stated" output (an l14g checkpoint's unknown)?

has_unknown property

has_unknown

Deprecated (removed in 0.9): has_not_stated.

has_pointer property

has_pointer

Can the checkpoint point at a piece of its input (span answers, evidence quotes)?

state_format property

state_format

The serialization of states for this checkpoint: the first of its declared ones that solvi writes ("paths" by default; a text-only checkpoint reads the "paths" lines as text).

max_len property

max_len

The tokens this checkpoint reads in one sequence (question and input): the encoder's, else the checkpoint's max_len, else 512.

max_len_long property

max_len_long

The tokens long="full" reads whole (question and input): the load(max_len_long=...) override, else the checkpoint's max_len_long, else None (the checkpoint was not trained on long inputs).

long_declared property

long_declared

Does the checkpoint itself declare a long-input length (max_len_long in solvi_decide.json)?

block property

block

Does this model score in the block layout (declared by the checkpoint, and the scorer has logits_pass)? Then every question — alone or with others — is scored in that layout: its answer does not depend on the other questions of its pass, and adapt / fit / teach see the same logits as the runtime.

load classmethod

load(path_or_id, device=None, backend='auto', max_len=None, bs=16, multi_question=None, act=None, max_len_long=None)

path_or_id: a checkpoint folder (~ is expanded) or a Hugging Face id (downloaded once, then read from the cache). multi_question / act: override what solvi_decide.json declares (for experiments, e.g. testing a checkpoint in multi-question passes); both are part of the fingerprint. max_len: the tokens of one ordinary pass (default: the checkpoint's max_len); it also sets long="retrieve"'s budget. max_len_long: the length long="full" reads whole (default: the checkpoint's max_len_long; a checkpoint that declares none refuses long="full" unless it is given here — with a warning: it was not trained on long inputs). Part of the fingerprint.

act_threshold_for

act_threshold_for(error)

The act threshold the checkpoint ships for a target error rate (its act.threshold_for_error table: the entry with the largest error ≤ the target). ValueError when it ships none — use calibrate_for on your examples.

wire

wire(kind)

How a question kind is asked: natively when the checkpoint was trained on it, else as a single choice (score: the levels; noul: the options "yes" / "no"; rank: the options, ordered by probability; number: the bins, as a score when the checkpoint has scores). A span is native only.

weights_fingerprint

weights_fingerprint()

A hash of the checkpoint (files, or the scorer's own fingerprint), the backend, the default calibration and (for checkpoints in the v2 format, or with overrides) the declared capabilities.

fingerprint

fingerprint()

The checkpoint plus every adaptation (so a replay knows the calibration a decision used).

metadata

metadata()

The checkpoint's metadata (without the training history), capabilities, the default calibration and every adaptation.

on_cpu

on_cpu()

Does the network run on a CPU (a torch scorer on "cpu", an ONNX session without CUDA)? None when unknown (a stand-in or remote scorer).

count_tokens

count_tokens(text)

Tokens of a text for this checkpoint: its tokenizer when it has one, else solvi.longdoc.approx_tokens.

text

text(v)

An input (a text, a Quote, a scalar or a state) → the text this checkpoint reads.

truncation

truncation(specs, text, read_len=0)

What one ordinary pass leaves unread of an input → None when it reads all of it, else {"input_tokens", "read_tokens", "question_tokens", "max_len"}. specs: the question, or the questions of a shared pass. None for a scorer that reads the text as it is (an LLM, a hosted decision model) and in the block layout when the input fits its budget.

logits

logits(text, task, options, descriptions=None, multi=False, other=None, kind=None)

Raw logits of the scored options ("other" excluded): {option: logit}, or a list of them for a list of inputs.

act_probability

act_probability(sp, d, act)

The probability that the model's answer is right, from its act logit: sigmoid(logit / T_act), or the checkpoint's act calibrator — a logistic regression over ACT_FEATURES of the calibrated decision (see docs/decide_format.md).

decide

decide(text, task, options, descriptions=None, multi=False, other=None, kind=None, min_confidence=None, **spec)

→ Decision(value, probs) (a list of them for a list of inputs). The value is always one of the options; a text or a state (dict, list, pydantic model: see state_text). spec: not_stated=, k=, bins=, unit=, coverage=, evidence= (see decision). min_confidence: escalate below this confidence (escalate_below= in 0.7).

score

score(text, task, options, descriptions=None, multi=False, other=None, kind=None)

→ {option: probability} (a list of them for a list of inputs): softmax over the options for a single choice, a score or yes/no, a sigmoid per option with multi=True; the adaptation of this question applied, "other" by its threshold.

decide_pass

decide_pass(text, parts, names=None)

Several decision parts about one input, scored together → [Decision] (one forward pass when the checkpoint supports it, else one per question). What the runtime does for parts grouped in flow.batches.

adapt

adapt(texts, task, options, descriptions=None, multi=False, other=None, kind=None, logits=None, **spec)

Label-bias correction without labels: the mean logit of each option over unlabelled inputs of the domain (centered over the options) is subtracted before the softmax / sigmoid. → the Adaptation. A later fit is kept consistent. logits: the inputs' logits already computed (one per text, in the question's option order) — what a DecisionPart passes, so the correction is fitted on the signal it decides on (option_order="average", long).

fit

fit(examples, task, options, descriptions=None, multi=False, other=None, lam=1.0, folds=4, kind=None, logits=None, **spec)

Few-shot adaptation "S" from labelled examples [(input, correct)]: a shift and a shared scale on the (bias-corrected) logits by L-BFGS (per option; a score: a tilt and a spread over the levels; yes/no: one bias), a temperature on out-of-fold predictions, and — when some examples are labelled "other" — the "other" threshold that maximizes out-of-fold accuracy. Replaces earlier examples. Examples labelled "not stated" are left out (the "not stated" logit is not adapted). logits: the examples' logits already computed (one per example; see adapt). → the Adaptation.

teach

teach(text, correct, task, options, descriptions=None, multi=False, other=None, lam=1.0, kind=None, logits=None, **spec)

One labelled example, absorbed at once: the shift / scale is refitted from the kept examples (warm start, K + 1 parameters or fewer — about a millisecond or two); the temperature and the "other" threshold stay until the next fit. logits: the input's logits already computed (see adapt). → the update time in ms (the model's forward pass, if the input was not scored before, is not included).

reset

reset(task=None, options=None, descriptions=None, multi=False, other=None, kind=None, **spec)

Forget the adaptation of one question, or every adaptation.

save_adaptations

save_adaptations(path)

Write every adaptation (with its kept examples' logits) to a JSON file, with the checkpoint's fingerprint.

load_adaptations

load_adaptations(path, strict=True)

Read adaptations written by save_adaptations. strict: refuse them if the checkpoint differs (the bias and shift were fitted on another model's logits).

decision

decision(name, task, text_fact='doc', options=(), descriptions=None, multi=False, other=None, *, kind=None, type=None, min_confidence=None, min_act=None, use_act=None, max_error=None, score_value=None, not_stated=False, k=None, bins=None, unit=None, coverage=None, evidence=False, option_order='canonical', permutations=4, min_margin=None, long=None, top_k=None, rerank=False, perturb=0, retrieve_query=None, _shared=frozenset())

A catalog part: text_fact (a fact name, or a list of them) → Decision(value, probs).

The question: options (a list, or {option: description}) and kind ("choice", "multi", "score", "noul"; default choice, or multi with multi=True) — or a Python type, as type= or in place of the options: Literal[...] / an Enum (choice), list[Literal[...]] (multi), Scale[...] (score), bool (noul: the value is True / False). The input: a text fact is read as it is (several joined by new lines); a state (dict, list, pydantic model) by state_text (several facts: {fact: value}).

Escalation: the model's act signal when it has one (use_act=False ignores it; min_act overrides the checkpoint's threshold; max_error=0.1 takes the checkpoint's threshold for that error rate), and a calibrated confidence below min_confidence (see calibrate_for); an escalated decision is rejected — the fact is missing, the answer abstains ("model escalated" / "low confidence" in the audit and stats). min_margin=0.1: also escalate when the two most probable answers are closer than that (a near tie is where a misleading text flips the choice).

Option order (choice and multi questions): "canonical" (the default) asks in sorted order, so how a caller lists the options cannot change the answer (the part's options are then in that order); "given" asks as listed (0.5.0); "average" averages the model's logits over permutations rotations of the list (each costs a forward pass) — against a model's preference for positions.

perturb=k: ask again on up to k variants of the input without its instruction-like sentences ("ignore the rules and answer X", "SYSTEM: ...", "the correct answer is X" — deterministic rules, solvi.perturb) and escalate when the answer changes ("answer depends on an instruction-like sentence: ..."). An input without such sentences costs nothing extra; one with them costs up to k forward passes.

Register with cat.fn(part) (a fact other parts read) or make it a question's answer with part.question(cat). The value is one of the options by construction; the options are the part's closed set; provenance decided; the trace records the model, whose fingerprint covers the checkpoint and this part's adaptation and thresholds.

Answer primitives (an l14g checkpoint, docs/decide_format.md §9): Maybe[T] or not_stated=True — "not stated" is an answer (solvi.Unknown); Rank[Literal[...], k] or kind="rank", k= — the options best first; Estimate[edges] or kind="number", bins=, unit=, coverage= — a number over bins; Span[T] or kind="span" — a piece of the (one, given) text fact, coerced to T when it is a question's answer; evidence=True (or a number) — supporting quotes from the pointer. A checkpoint that cannot give what is asked raises here.

Long texts: long=None cuts a text beyond max_len (the tokenizer truncates it, as before); long="retrieve" splits it into sections, selects the top_k that bear on the question by BM25 (rerank=True: re-ordered by the decider's own relevance, one yes / no pass per candidate section) and decides on them; spans and evidence point into the whole text, and the sections read are in the decision's extra["long"] (solvi.longdoc). top_k=None (the default): sections of about 170 tokens — budget / 170, at least 3 (3 at max_len 512, 12 at 2048). retrieve_query: the words the sections are searched by, in place of the question's own (its task, options and descriptions) — the labels the document writes next to the value ("Invoice No Contract No Ref"), or the document's language when the question is asked in another one; the decider still reads the question as it is. long="full" (a checkpoint trained on long inputs: max_len_long in its solvi_decide.json) reads a text that does not fit max_len whole, up to max_len_long tokens, and retrieves within max_len_long beyond that (recorded in extra["long"]); a GPU mode — on a CPU a whole 8k-token text takes seconds per question.

An option the question's kind does not use is a ValueError, not ignored: score_value= (score questions; default "median"), k= (rank), bins= / unit= / coverage= (number; coverage default 0.8), other= (choice and multi), min_margin= (not multi), top_k= / rerank= (with long=), min_act= / max_error= (a checkpoint with an act head), and kind= that contradicts multi=True.

Names (0.8): min_confidence= (0.7: escalate_below=), min_act= (act_threshold=), max_error= (target_error=), not_stated= (unknown=) — the old ones work with a SolviDeprecationWarning until 0.9; the part keeps the thresholds as part.min_confidence / part.min_act.

decisions

decisions(schema, text_fact='doc', fields=None, **kw)

One decision part per field of a pydantic model class: the field's type is the question (bool, Literal[...], an Enum, Scale[...], list[Literal[...]]), its description the task (else its title, else its name), and json_schema_extra may carry "options" ({option: description}), "min_confidence", "min_act", "use_act", "max_error", "other", "score_value" (the 0.7 keys "escalate_below", "act_threshold", "target_error" still read, with a SolviDeprecationWarning). → {field: DecisionPart} in field order.

questions

questions(cat, schema, text_fact='doc', fields=None, min_confidence=None, **kw)

decisions(...) registered as the answers of questions named after the fields → [Question].

jsonable

jsonable(v)

A Python value → JSON data, deterministically: a pydantic model → its model_dump(), a dataclass → its fields, dates and times → ISO 8601, an Enum → its value, Decimal / UUID / other objects → str, bytes → UTF-8 text, tuples → lists, sets → lists sorted by their JSON text, numpy → Python numbers.

state_text

state_text(obj, fmt='paths')

The input a decider reads: a text as it is; any other value (a dict, a list, a pydantic model, a dataclass) first as JSON data (see jsonable), then serialized — by default "paths", one line per leaf with its full key path:

customer.name: Anna
items[0].sku: A-1
items[0].qty: 2
note: two lines of text

Keys in their order (a pydantic model: field order); a key that is not [A-Za-z0-9_-]+ is written ["key"] (a JSON string); strings without quotes (a new line becomes a space); null / true / false; floats rounded to 6 decimals without trailing zeros; empty {} and [] kept; a scalar at the top is ".: value". "tree" is the YAML-like indented form, "json" is json.dumps with ", " / ": " separators. These are exactly the typed decider's training serializations (serialize of its training code); docs/decide_format.md has the rules.

decode_pointer

decode_pointer(ptr, text, max_span=40, top=20, temperature=1.0)

The pointer's raw output → {"null", "spans"}: a span's score is start_i + end_j over the input's tokens i ≤ j < i + max_span, the null span's start_m + end_m at the mode marker; p = softmax over the null span and every span (exactly span_dist of the training code). A span's text is the input's characters from token i's start to token j's end, without surrounding whitespace — so it is literally in the input. ptr: {"start": [T], "end": [T], "offsets": [(char start, char end)] per token, "null": [start_m, end_m] (or their sum)}. temperature (the checkpoint's temperature.span) divides every start / end score, the null span's too, before the softmax — as the answer-primitives calibration fitted it.

pass_prompt

pass_prompt(items, markers=None)

The first segment of a pass: the questions' segments joined by a space.

pointer_evidence

pointer_evidence(ptr, threshold=0.15, max_spans=3)

Evidence quotes from a decoded pointer: greedily up to max_spans non-overlapping spans with p ≥ threshold; none when the null span is at least as probable as the best span (the l14g contract's evidence).

prompt

prompt(task, options, descriptions=None, multi=False, mode=None, markers=None)

A question's segment: "[mode] task[opt] option ..." (an option with a description is "label: description").

block_masks

block_masks(blk, pids, window)

blk [B, L] (−1 input, j ≥ 0 block j, −2 padding), pids [B, L] → (full, sliding) boolean masks [B, 1, L, L] (True = may attend): the input sees only the input, block j sees the input and itself, blocks do not see each other; the sliding (local) layers also need |pos_i − pos_j| ≤ window. Padding rows see themselves only.

lora_key

lora_key(sp_or_key)

The question a LoRA adapter (solvi.lora) belongs to: the task, the scored options with their descriptions (in any order — option_order="average" asks rotations of the same question) and the kind. A _Spec or a _Spec.key.

act_features

act_features(sp, d, act_logit)

The features of the act calibrator (ACT_FEATURES) for a calibrated decision: confidence (the decision's calibrated confidence: the top probability; multi-label: the least certain option's max(p, 1 − p)), margin (top − second probability; multi-label: the smallest |2p − 1|), entropy (of the probabilities, nats; multi-label: the mean binary entropy), act_logit (raw), n_options, kind= (1 / 0).

decision_of

decision_of(catalog, question)

The DecisionPart behind a question's answer: its rule is a decision, or a rule that only passes a decided fact on (e.g. def team(route): return route). → the part or None.

group_name

group_name(path)

A group path as people read it: "billing / refunds"; the whole stream: "(the rest of the stream)".

group_record

group_record(g, path, node, info)

A decision's guarantee under thresholds per group: the part's promise, the input's group, the group whose threshold applied (its own, or a parent's when the group had too few examples), that threshold and its examples.

guard_promise

guard_promise(risk, error, answered=True)

act_guard's promise in words, with the error among the answers given alone on the calibration examples.

no_separation

no_separation(sig, ok, name, error, base, who='')

A warning when the signal does not tell right answers from wrong ones on the calibration examples (one-sided Mann-Whitney test of its AUROC against chance at the 5% level: solvi.calibration.separation), else None — tested only with at least SEPARATION_MIN right and as many wrong examples (fewer cannot tell). The promise still holds — by escalating, not by choosing: what is answered alone is wrong about as often as everything.

one_source

one_source(ds, who='the calibration examples')

The one source of the decisions' probabilities (see confidence_source; None when no decision names one). Raises ValueError when log-probabilities and written numbers are mixed: they are two scales (a token probability near 1 against a stated 0.85-0.95), and one threshold over both answers alone by which server happened to reply.

plan_batches

plan_batches(steps)

The strategist's grouping: decision parts in a flow that read the same facts with the same model, when the model can answer several questions in one forward pass → [[step name]] (chunks of at most the checkpoint's max_questions).