solvi.heads¶
The learned answer heads: FastHead (what System.fit builds: closed-form ridge, exact leave-one-out, rank-one updates) and
select_features, CandidateHead (a choice among candidates that change), and the older Head. solvi.fast up to 0.7.
Answer head for questions without a rule: features from computed_state → answer probabilities. A number is encoded as its value plus thresholds at training quantiles ("greater than threshold" steps). Multinomial logistic regression with L2 (Adam, 300 steps — milliseconds; one-vs-rest ridge could not express the middle one of ordered classes). Features are chosen by greedy forward selection on 5-fold cross-validation accuracy (gain ≥ 1 pt): the selected facts define the question's flow.
System.fit built this head before 0.8; it now builds a FastHead (below), whose selection is scored by the leave-one-out squared error (accuracy kept nothing on imbalanced questions). Head stays for code that builds one itself; its Featurizer is shared with FastHead.
FastHead, VecFeaturizer and CandidateHead were solvi.fast up to 0.7 (the module still reads, with a warning):
Fast answer head: closed-form ridge regression, learned in milliseconds and updated instantly from every correction.
Features come from computed facts (numbers with quantile steps, booleans, categories — as in Head) and from vector facts such
as a document embedding (see LongSpanExtractor.embedder). Answers are one-vs-rest ridge scores turned into probabilities.
The head keeps the inverse of (XᵀX + λI), so update is a rank-one Sherman–Morrison step: a new labeled example is absorbed in
about a millisecond without retraining and without touching the rest of the system. Leave-one-out accuracy is exact and free
(the ridge hat matrix).
A rank-one step keeps what the first fit chose: the ridge strength, the featurizer (number scales, the known values of each
category) and whether pairwise products are used. Chosen on a handful of rows they are often wrong for hundreds: a head
started on a few rows and taught one at a time ends up less accurate than one fitted on all of them. So the head keeps its examples and, each time their number doubles (refit=2.0), fits again on all of them — the same as a fresh
fit on those rows — and goes on with rank-one steps. The amortised cost stays a constant per update; the update that
triggers a refit is as slow as a fit. Past refit_until examples (2000) it stops refitting and drops the kept rows.
Head ¶
VecFeaturizer ¶
Bases: Featurizer
Featurizer plus fixed-length numeric vectors (lists / arrays), standardized per dimension.
compact ¶
One value per number (standardized), ±1 per boolean, one-hot categories; vectors left out — the base for pairs.
FastHead ¶
lam: ridge strength (None — chosen by exact leave-one-out accuracy); pairs: add pairwise products of the base features
(None — only when there are at most 40 of them, so interactions and middle classes can be expressed). refit: when
update brings the number of examples to refit times the number of the last fit, fit again on all of them (the
ridge strength, the featurizer and the pairs decision are chosen again, as given here); None — never, and no rows are
kept. refit_until: no refit beyond this many examples; the kept rows are dropped once none is due.
fingerprint ¶
A stable hash of the head's parameters (recorded with every answer it gives; changes with fit and each update).
teach ¶
Absorb one labeled example (rank-one update of the inverse; a refit on all kept examples when their number reaches the refit schedule); returns the time it took in ms.
contributions ¶
Per fact: its own weight times value (pairwise terms are split equally between the two facts).
CandidateHead ¶
A choice among candidates that change with every decision, learned from the candidates' own features.
An answer head has fixed options. An agent's step has other candidates each time — the exits of a room, the rows a search returned, the tools on offer — so there is nothing to attach a per-option weight to, and fit / teach have no key to learn under. What does carry over from step to step is what a candidate is: its distance, its kind, whether it was a dead end. This head learns "is this candidate the one to take?" over those features (a FastHead: ridge, fitted in milliseconds, taught by one correction) and chooses the candidate it scores highest. relative=True also gives each numeric feature relative to the other candidates of the step (its gap to the smallest and the largest).
head = CandidateHead(["kind", "distance", "reward", "dead_end"]).fit(steps) # steps: [(candidates, chosen index)]
i, probs = head.choose(candidates) # candidates: [{feature: value}]; probs sum to 1 over them
head.teach(candidates, 2) # one correction: the candidate that should have been taken
Labels come from a rule you are replacing, from people, or from outcomes — and an outcome label must be the criterion of the sub-goal the step served (progress towards the goal for "go on", a level gained for "train"): one global measure teaches the head to ignore every step that does not move it.
relative=True is off by default: whether the gaps help depends on the task, so compare both on held-out steps. The head learns the rule it is shown: trained on a rule's choices it reproduces the rule and can replace a model there, it does not beat the rule.
rows ¶
The rows the head reads: each candidate's features and, for numbers, the gaps to the step's min and max.
fit ¶
steps: [(candidates, chosen)] — chosen: the index of the right candidate, or a set of indexes when several are as good. The others of the step are its "no" examples.
scores ¶
The head's score of each candidate being the one to take (higher: better), in the candidates' order.
choose ¶
→ (the index of the best candidate, [probability per candidate]) — ties go to the earlier candidate.
teach ¶
One correction: this candidate (index, or a set of them) was the one to take. → ms.
select_features ¶
Greedy forward selection for a FastHead: start from no feature (the answers' shares) and add, one at a time, the
fact that lowers the exact leave-one-out squared error most; stop when none lowers it by more than min_gain × the
error of the answers' shares. → (chosen, {fact: error after adding it}, {fact that cannot be encoded: why}). A
proper score, so a rare answer counts:
accuracy, the old criterion, keeps nothing where one answer is most of the examples (a fact seldom changes the
majority answer).
Each fact is chosen with its own encoding only (no pairwise products while choosing).