Questions without a rule: fit, learn_rule, teach¶
Some answers are hard to write as a rule (a risk level, a region from a messy address). solvi offers two ways to learn
them from labeled examples: an answer head (fit) and a readable rule list (learn_rule). Both use the facts that your catalog computes, not raw text.
fit: a learned answer head, corrected instantly¶
history = [(init_state_1, "low"), (init_state_2, "high"), ...] # labeled examples
head = system.fit("risk", history)
print(head.features, head.selection) # the facts it kept, and the leave-one-out error after adding each
print(head.loo_acc) # exact leave-one-out accuracy
system.fit(question, examples, features=None, *, select=None, min_gain=0.0) fits a closed-form ridge head
(solvi.heads.FastHead): one matrix decomposition per ridge strength, so it takes milliseconds to a few seconds, and the
ridge strength is chosen by exact leave-one-out accuracy (head.loo_acc).
- Features are all facts computable from the examples'
init_statekeys, the given keys themselves included (a given number is a feature as it is, without a function around it; a given value that cannot be encoded — a long text, a dict — is left out and named inhead.dropped). Numbers are encoded as a value plus thresholds at training quantiles, booleans as +/-1, strings with at most 20 distinct values as categories, and numeric vectors (a document embedding fromLongSpanExtractor.embedder()) per dimension; when there are few of them, their pairwise products are added so middle classes and interactions can be expressed. - Without
features=, the facts are selected greedily: start from the answers' shares and add the fact that lowers the exact leave-one-out squared error most, while it lowers it by more thanmin_gain× the error of the shares (min_gain=0.0: any decrease). The selected facts become the question's flow, so later requests compute only what the head reads. The squared error is a proper score: a rare answer counts.features=[...]keeps every fact listed;select=Falsekeeps every computable fact;select=Trueselects among the facts listed. - The answer's
whylists the largest feature contributions,probsgives all class probabilities. - A head left with no feature answers the same for every input;
fitwarns when that happens and says why: no fact lowered the leave-one-out error (the facts do not tell the answers apart on these examples), or no fact could be computed from the examples' inputs — every parameter of a part is a fact it reads, one with a default value too (def fn(facts, _nm=nm)waits for a fact_nm).
Before 0.8 fit was a logistic regression whose features were chosen by cross-validated accuracy, and fit_fast was
the ridge head on every feature. Accuracy is a poor guide when one answer is most of the examples: it can keep almost
no fact. The selection by squared error keeps the facts that matter for a rare answer too. On very few examples (a few
dozen) choosing among many facts overfits: pass select=False there. benchmarks/fast_head.py compares the selection
with every fact on the example tasks — accuracy on fresh examples and the fitting time; re-run it on your own data.
fit_fast(...) still works in 0.8, with a SolviDeprecationWarning: it is fit(..., select=False); it goes in 0.9.
Its other property is online learning: system.teach(question, init_state, correct) updates the head immediately with
a rank-one Sherman–Morrison step (about 0.1–0.2 ms: benchmarks/fast_head.py, examples/10_learn_in_milliseconds.py) and returns the time in ms. Other questions, rules and hard checks
do not change, and the example still goes to the journal.
head = system.fit("suspicious", history[:10], select=False) # start small: every fact
for state, label in reviewer_corrections:
system.teach("suspicious", state, label) # each one is absorbed at once
Refits as examples accumulate. A rank-one step keeps what the first fit chose: the ridge strength, the featurizer
(number scales and the known values of each category — a value first seen later counts as none of them) and whether
pairwise products are used. Chosen on 10 examples these are often wrong for 300. So the head keeps its examples and,
each time teach doubles their number (at 20, 40, 80, 160, … after a start on 10), fits again on all of them — exactly
a fresh fit on those examples and the facts it reads (the selection is not made again: the flow stays) — then goes on
with rank-one steps. Without refits a head started on a few examples and taught up to a few hundred stays behind a
fit on all of them; with refits it catches up. The cost:
- the update that triggers a refit takes as long as a fit on that many examples, and an early refit that switches pairwise products on (hundreds of columns from few examples) is the slowest — as long as the first fit on those examples would take. The other updates are unchanged, and a refit usually drops the pairwise products a start on few examples switched on, so later updates are often faster;
- the kept examples: their fact rows, about 1 KB each for 14 plain facts (vector facts such as embeddings cost their
length), up to
refit_untilexamples (2000). Past that no refit is due, the rows are dropped and the head goes on with rank-one steps only; - a refit is a change like any update:
teachmakes it at once and it is not gated. The learning loop (System.learning) manages decision parts, not fitted heads: while it is attached withgate_teach=True,teachonly stores the correction and the head (and its refit schedule) does not move. The head depends only on its first fit and the sequence of corrections, so replaying them gives the same head (the same fingerprint); keep acopy.deepcopy(head)to go back.
fit(..., refit=None) turns it off (rank-one steps only, no examples kept); refit=1.5 refits more often.
learn_rule: a readable rule list¶
rules = system.learn_rule("zone", examples, features=["address_upper"], min_support=3, min_precision=0.8, max_rules=40)
print(rules)
This learns an ordered decision list ("if feature then answer", with a default at the end) and installs it in the
catalog as the question's rule. From then on it behaves like a hand-written rule: deterministic, explained by its inputs,
and re-checkable by replay.
facts: the computed facts to build literals from. Booleans givefact is True/False, numbersfact ≈ rounded value, strings give one literal per upper-cased word or number (in any script:"ул. Северная, 12"givesУЛ,СЕВЕРНАЯ,12), plushas number starting 'NN'for numbers of 5 or more digits (postcodes, codes). Other values are compared for equality. A name that is neither a part of the catalog nor a given fact of the examples raisesValueError, and so does an empty list of examples; nothing is installed then.- Each step adds the literal with the best smoothed precision on still-uncovered examples, with at least
min_supportexamples, while precision stays at or abovemin_precision; at mostmax_rulesrules. system.learned_rules[question]keeps theRuleList;printshows each rule with its support.
A learned rule replaces any rule previously registered for that question, and takes precedence over a trained head. Because the list is readable, you can spot rules that only memorized a few examples, or that reproduce a labeling error, and fix the labels or the catalog.
teach: record corrections¶
teach appends the correction to the journal or storage (it does nothing without one). It does not retrain by itself:
read the corrections back (system.storage.corrections(), or the {"teach": ..., "init": ..., "answer": ...} lines) and
include them in the next fit or learn_rule call. Non-JSON values in init_state are stored as JSON (dates as ISO
strings; 0.5 wrote their repr), other objects as their repr.
Each correction keeps where it came from: system.teach("risk", state, "high", label_source="outcome", by="ledger",
of=res.stored_id) — source is "human" (the default: a person corrected or confirmed the answer), "outcome" (what
really happened) or "rule" (your code rejected a model's proposal and decided instead); anything else raises
UntrustedLabel. of is the stored id of the decision it corrects; corrections() returns all three.
Learning from corrections with gates and rollback (experimental)¶
system.learning(...) turns the stored corrections into updates of the model decisions — only through gates, recorded,
and reversible. It is off until you call it, and experimental (it warns ExperimentalWarning; its API and defaults may
change):
loop = system.learning(store, gates={"honesty": "tests/honesty/core_v1.json"})
# ... the system runs; people correct escalations with system.teach(...) — now stored only, not learned at once
rep = loop.run() # labels → a proposed update → gates → promoted or undone; recorded either way
print(rep) # the rung per question and each gate's numbers
loop.versions() # [{"version", "fp", "time", "action"}]
loop.rollback(2) # any promoted version, in this process or another one on the same store
Labels come only from outside the model: the store's corrections from "human", "outcome" and "rule" sources
A correction from any other source is refused and listed (loop.labels()["rejected"]); the stored decisions — the
system's own answers — are never read as labels, so self-training is impossible by construction. Each label goes, by a
hash of its id, to "train", "calibration" or "holdout" (50 / 20 / 30% by default): a held-out label is never trained
on, in this update or any later one. While the loop is attached, System.teach only stores the correction
(gate_teach=False keeps the instant update; loop.detach() ends it).
The ladder, per question, by the number of training labels: under fit_below (50) — the decider's shift / scale
(fit); under memory_below (1000) — fit plus a memory of the corrected cases (part.memory, mode "check" unless
ladder={"memory": {"mode": "answer"}}); beyond — the adapter hook when you give one ((part, [(text, answer)]) →
a JSON-able description; an object with state(part) / restore(part, state) is rolled back too), else the memory.
The gates — the update is promoted only if every one passes; otherwise it is dropped and recorded as rejected. The
candidate is built and gated on a shadow of the system (copies of the decision parts, their thresholds and memory, and of
the model's adaptations; the checkpoint is shared), so asks that run meanwhile see the state in force, never an un-gated
candidate; the live parts change only when the update is promoted. An adapter hook runs on the shadow part and reaches
the live one through its state / restore.
| Gate | Passes when |
|---|---|
consistency |
at most max_conflict (20%) of the new training labels are contradicted by a memory of the labels already learned — a batch of wrong corrections hurts more than right ones help |
heldout |
on the held-out labels, asked through the whole system, (right − wrong answered alone) / n improves by at least min_gain (0.01), with at least min_holdout (5) labels |
honesty |
the honesty numbers (confident errors, coverage at risk, quote support) on the held-out labels — and on your own honesty set (gates={"honesty": path or cases}) — get no worse than tolerance (0.02) |
act_guard |
a part calibrated with act_guard / calibrate_for is calibrated again, with the same risk, on at least min_calibration (30) calibration labels: an old threshold says nothing about a changed signal; conformal answer sets are recalibrated on them too, or dropped (and recorded) with fewer |
size |
shadow run: the stored decisions whose input is held out for a learned question — the labels' split, per question (up to shadow_limit, 500) are asked with the current and the candidate state and compared (solvi.diff.compare); at most max_change (30%) may change |
max_change is deliberately low: the first update of a badly biased decider can move far more than 30% of the decisions
and is then rejected until you raise the limit for it (gates={"max_change": 0.8}) — a decision a person should take.
The record. Every run that proposes an update writes a record of kind "update" to the changelog (by default the
same store, hash-chained with the decisions; changelog= for another): the version it came from, the fingerprints before
and after, the rung and label counts per question, the training label ids, every gate's numbers and — for a promoted
state — the state itself (the adaptation with its examples, the thresholds and guarantee, the memory's cases). The first
run records the state it started from as a baseline version. Each decision's trace names the part's fingerprint, and
loop.version_of(fp) the version it belongs to. Limits: only questions answered by a single decision part (not a
combination), and per-group act_guard thresholds are not recalibrated (such an update is rejected).