Skip to content

Questions without a rule: fit, learn_rule, teach

Some answers are hard to write as a rule (a risk level, a region from a messy address). solvi offers two ways to learn them from labeled examples: an answer head (fit) and a readable rule list (learn_rule). Both use the facts that your catalog computes, not raw text.

fit: a learned answer head, corrected instantly

history = [(init_state_1, "low"), (init_state_2, "high"), ...]   # labeled examples
head = system.fit("risk", history)
print(head.features, head.selection)       # the facts it kept, and the leave-one-out error after adding each
print(head.loo_acc)                        # exact leave-one-out accuracy

system.fit(question, examples, features=None, *, select=None, min_gain=0.0) fits a closed-form ridge head (solvi.heads.FastHead): one matrix decomposition per ridge strength, so it takes milliseconds to a few seconds, and the ridge strength is chosen by exact leave-one-out accuracy (head.loo_acc).

  • Features are all facts computable from the examples' init_state keys, the given keys themselves included (a given number is a feature as it is, without a function around it; a given value that cannot be encoded — a long text, a dict — is left out and named in head.dropped). Numbers are encoded as a value plus thresholds at training quantiles, booleans as +/-1, strings with at most 20 distinct values as categories, and numeric vectors (a document embedding from LongSpanExtractor.embedder()) per dimension; when there are few of them, their pairwise products are added so middle classes and interactions can be expressed.
  • Without features=, the facts are selected greedily: start from the answers' shares and add the fact that lowers the exact leave-one-out squared error most, while it lowers it by more than min_gain × the error of the shares (min_gain=0.0: any decrease). The selected facts become the question's flow, so later requests compute only what the head reads. The squared error is a proper score: a rare answer counts. features=[...] keeps every fact listed; select=False keeps every computable fact; select=True selects among the facts listed.
  • The answer's why lists the largest feature contributions, probs gives all class probabilities.
  • A head left with no feature answers the same for every input; fit warns when that happens and says why: no fact lowered the leave-one-out error (the facts do not tell the answers apart on these examples), or no fact could be computed from the examples' inputs — every parameter of a part is a fact it reads, one with a default value too (def fn(facts, _nm=nm) waits for a fact _nm).

Before 0.8 fit was a logistic regression whose features were chosen by cross-validated accuracy, and fit_fast was the ridge head on every feature. Accuracy is a poor guide when one answer is most of the examples: it can keep almost no fact. The selection by squared error keeps the facts that matter for a rare answer too. On very few examples (a few dozen) choosing among many facts overfits: pass select=False there. benchmarks/fast_head.py compares the selection with every fact on the example tasks — accuracy on fresh examples and the fitting time; re-run it on your own data. fit_fast(...) still works in 0.8, with a SolviDeprecationWarning: it is fit(..., select=False); it goes in 0.9.

Its other property is online learning: system.teach(question, init_state, correct) updates the head immediately with a rank-one Sherman–Morrison step (about 0.1–0.2 ms: benchmarks/fast_head.py, examples/10_learn_in_milliseconds.py) and returns the time in ms. Other questions, rules and hard checks do not change, and the example still goes to the journal.

head = system.fit("suspicious", history[:10], select=False)     # start small: every fact
for state, label in reviewer_corrections:
    system.teach("suspicious", state, label)                     # each one is absorbed at once

Refits as examples accumulate. A rank-one step keeps what the first fit chose: the ridge strength, the featurizer (number scales and the known values of each category — a value first seen later counts as none of them) and whether pairwise products are used. Chosen on 10 examples these are often wrong for 300. So the head keeps its examples and, each time teach doubles their number (at 20, 40, 80, 160, … after a start on 10), fits again on all of them — exactly a fresh fit on those examples and the facts it reads (the selection is not made again: the flow stays) — then goes on with rank-one steps. Without refits a head started on a few examples and taught up to a few hundred stays behind a fit on all of them; with refits it catches up. The cost:

  • the update that triggers a refit takes as long as a fit on that many examples, and an early refit that switches pairwise products on (hundreds of columns from few examples) is the slowest — as long as the first fit on those examples would take. The other updates are unchanged, and a refit usually drops the pairwise products a start on few examples switched on, so later updates are often faster;
  • the kept examples: their fact rows, about 1 KB each for 14 plain facts (vector facts such as embeddings cost their length), up to refit_until examples (2000). Past that no refit is due, the rows are dropped and the head goes on with rank-one steps only;
  • a refit is a change like any update: teach makes it at once and it is not gated. The learning loop (System.learning) manages decision parts, not fitted heads: while it is attached with gate_teach=True, teach only stores the correction and the head (and its refit schedule) does not move. The head depends only on its first fit and the sequence of corrections, so replaying them gives the same head (the same fingerprint); keep a copy.deepcopy(head) to go back.

fit(..., refit=None) turns it off (rank-one steps only, no examples kept); refit=1.5 refits more often.

learn_rule: a readable rule list

rules = system.learn_rule("zone", examples, features=["address_upper"], min_support=3, min_precision=0.8, max_rules=40)
print(rules)

This learns an ordered decision list ("if feature then answer", with a default at the end) and installs it in the catalog as the question's rule. From then on it behaves like a hand-written rule: deterministic, explained by its inputs, and re-checkable by replay.

  • facts: the computed facts to build literals from. Booleans give fact is True/False, numbers fact ≈ rounded value, strings give one literal per upper-cased word or number (in any script: "ул. Северная, 12" gives УЛ, СЕВЕРНАЯ, 12), plus has number starting 'NN' for numbers of 5 or more digits (postcodes, codes). Other values are compared for equality. A name that is neither a part of the catalog nor a given fact of the examples raises ValueError, and so does an empty list of examples; nothing is installed then.
  • Each step adds the literal with the best smoothed precision on still-uncovered examples, with at least min_support examples, while precision stays at or above min_precision; at most max_rules rules.
  • system.learned_rules[question] keeps the RuleList; print shows each rule with its support.

A learned rule replaces any rule previously registered for that question, and takes precedence over a trained head. Because the list is readable, you can spot rules that only memorized a few examples, or that reproduce a labeling error, and fix the labels or the catalog.

teach: record corrections

system = System(cat, questions, storage="decisions.jsonl")
system.teach("risk", init_state, "high")

teach appends the correction to the journal or storage (it does nothing without one). It does not retrain by itself: read the corrections back (system.storage.corrections(), or the {"teach": ..., "init": ..., "answer": ...} lines) and include them in the next fit or learn_rule call. Non-JSON values in init_state are stored as JSON (dates as ISO strings; 0.5 wrote their repr), other objects as their repr.

Each correction keeps where it came from: system.teach("risk", state, "high", label_source="outcome", by="ledger", of=res.stored_id) — source is "human" (the default: a person corrected or confirmed the answer), "outcome" (what really happened) or "rule" (your code rejected a model's proposal and decided instead); anything else raises UntrustedLabel. of is the stored id of the decision it corrects; corrections() returns all three.

Learning from corrections with gates and rollback (experimental)

system.learning(...) turns the stored corrections into updates of the model decisions — only through gates, recorded, and reversible. It is off until you call it, and experimental (it warns ExperimentalWarning; its API and defaults may change):

loop = system.learning(store, gates={"honesty": "tests/honesty/core_v1.json"})
# ... the system runs; people correct escalations with system.teach(...) — now stored only, not learned at once
rep = loop.run()                  # labels → a proposed update → gates → promoted or undone; recorded either way
print(rep)                        # the rung per question and each gate's numbers
loop.versions()                   # [{"version", "fp", "time", "action"}]
loop.rollback(2)                  # any promoted version, in this process or another one on the same store

Labels come only from outside the model: the store's corrections from "human", "outcome" and "rule" sources A correction from any other source is refused and listed (loop.labels()["rejected"]); the stored decisions — the system's own answers — are never read as labels, so self-training is impossible by construction. Each label goes, by a hash of its id, to "train", "calibration" or "holdout" (50 / 20 / 30% by default): a held-out label is never trained on, in this update or any later one. While the loop is attached, System.teach only stores the correction (gate_teach=False keeps the instant update; loop.detach() ends it).

The ladder, per question, by the number of training labels: under fit_below (50) — the decider's shift / scale (fit); under memory_below (1000) — fit plus a memory of the corrected cases (part.memory, mode "check" unless ladder={"memory": {"mode": "answer"}}); beyond — the adapter hook when you give one ((part, [(text, answer)]) → a JSON-able description; an object with state(part) / restore(part, state) is rolled back too), else the memory.

The gates — the update is promoted only if every one passes; otherwise it is dropped and recorded as rejected. The candidate is built and gated on a shadow of the system (copies of the decision parts, their thresholds and memory, and of the model's adaptations; the checkpoint is shared), so asks that run meanwhile see the state in force, never an un-gated candidate; the live parts change only when the update is promoted. An adapter hook runs on the shadow part and reaches the live one through its state / restore.

Gate Passes when
consistency at most max_conflict (20%) of the new training labels are contradicted by a memory of the labels already learned — a batch of wrong corrections hurts more than right ones help
heldout on the held-out labels, asked through the whole system, (right − wrong answered alone) / n improves by at least min_gain (0.01), with at least min_holdout (5) labels
honesty the honesty numbers (confident errors, coverage at risk, quote support) on the held-out labels — and on your own honesty set (gates={"honesty": path or cases}) — get no worse than tolerance (0.02)
act_guard a part calibrated with act_guard / calibrate_for is calibrated again, with the same risk, on at least min_calibration (30) calibration labels: an old threshold says nothing about a changed signal; conformal answer sets are recalibrated on them too, or dropped (and recorded) with fewer
size shadow run: the stored decisions whose input is held out for a learned question — the labels' split, per question (up to shadow_limit, 500) are asked with the current and the candidate state and compared (solvi.diff.compare); at most max_change (30%) may change

max_change is deliberately low: the first update of a badly biased decider can move far more than 30% of the decisions and is then rejected until you raise the limit for it (gates={"max_change": 0.8}) — a decision a person should take.

The record. Every run that proposes an update writes a record of kind "update" to the changelog (by default the same store, hash-chained with the decisions; changelog= for another): the version it came from, the fingerprints before and after, the rung and label counts per question, the training label ids, every gate's numbers and — for a promoted state — the state itself (the adaptation with its examples, the thresholds and guarantee, the memory's cases). The first run records the state it started from as a baseline version. Each decision's trace names the part's fingerprint, and loop.version_of(fp) the version it belongs to. Limits: only questions answered by a single decision part (not a combination), and per-group act_guard thresholds are not recalibrated (such an update is rejected).