Skip to content

solvi.heads

The learned answer heads: FastHead (what System.fit builds: closed-form ridge, exact leave-one-out, rank-one updates) and select_features, CandidateHead (a choice among candidates that change), and the older Head. solvi.fast up to 0.7.

Answer head for questions without a rule: features from computed_state → answer probabilities. A number is encoded as its value plus thresholds at training quantiles ("greater than threshold" steps). Multinomial logistic regression with L2 (Adam, 300 steps — milliseconds; one-vs-rest ridge could not express the middle one of ordered classes). Features are chosen by greedy forward selection on 5-fold cross-validation accuracy (gain ≥ 1 pt): the selected facts define the question's flow.

System.fit built this head before 0.8; it now builds a FastHead (below), whose selection is scored by the leave-one-out squared error (accuracy kept nothing on imbalanced questions). Head stays for code that builds one itself; its Featurizer is shared with FastHead.

FastHead, VecFeaturizer and CandidateHead were solvi.fast up to 0.7 (the module still reads, with a warning):

Fast answer head: closed-form ridge regression, learned in milliseconds and updated instantly from every correction.

Features come from computed facts (numbers with quantile steps, booleans, categories — as in Head) and from vector facts such as a document embedding (see LongSpanExtractor.embedder). Answers are one-vs-rest ridge scores turned into probabilities. The head keeps the inverse of (XᵀX + λI), so update is a rank-one Sherman–Morrison step: a new labeled example is absorbed in about a millisecond without retraining and without touching the rest of the system. Leave-one-out accuracy is exact and free (the ridge hat matrix).

A rank-one step keeps what the first fit chose: the ridge strength, the featurizer (number scales, the known values of each category) and whether pairwise products are used. Chosen on a handful of rows they are often wrong for hundreds: a head started on a few rows and taught one at a time ends up less accurate than one fitted on all of them. So the head keeps its examples and, each time their number doubles (refit=2.0), fits again on all of them — the same as a fresh fit on those rows — and goes on with rank-one steps. The amortised cost stays a constant per update; the update that triggers a refit is as slow as a fit. Past refit_until examples (2000) it stops refitting and drops the kept rows.

Head

Head(options)

fingerprint

fingerprint()

A stable hash of the head's parameters (recorded with every answer it gives).

contributions

contributions(row)

Contribution of each feature to the chosen answer (for explanations).

VecFeaturizer

VecFeaturizer()

Bases: Featurizer

Featurizer plus fixed-length numeric vectors (lists / arrays), standardized per dimension.

compact

compact(r, facts)

One value per number (standardized), ±1 per boolean, one-hot categories; vectors left out — the base for pairs.

FastHead

FastHead(options, lam=None, pairs=None, refit=2.0, refit_until=2000)

lam: ridge strength (None — chosen by exact leave-one-out accuracy); pairs: add pairwise products of the base features (None — only when there are at most 40 of them, so interactions and middle classes can be expressed). refit: when update brings the number of examples to refit times the number of the last fit, fit again on all of them (the ridge strength, the featurizer and the pairs decision are chosen again, as given here); None — never, and no rows are kept. refit_until: no refit beyond this many examples; the kept rows are dropped once none is due.

fingerprint

fingerprint()

A stable hash of the head's parameters (recorded with every answer it gives; changes with fit and each update).

update

update(row, answer)

Deprecated (removed in 0.9): teach(row, answer).

teach

teach(row, answer)

Absorb one labeled example (rank-one update of the inverse; a refit on all kept examples when their number reaches the refit schedule); returns the time it took in ms.

contributions

contributions(row)

Per fact: its own weight times value (pairwise terms are split equally between the two facts).

CandidateHead

CandidateHead(features, relative=False, pairs=None, refit=2.0)

A choice among candidates that change with every decision, learned from the candidates' own features.

An answer head has fixed options. An agent's step has other candidates each time — the exits of a room, the rows a search returned, the tools on offer — so there is nothing to attach a per-option weight to, and fit / teach have no key to learn under. What does carry over from step to step is what a candidate is: its distance, its kind, whether it was a dead end. This head learns "is this candidate the one to take?" over those features (a FastHead: ridge, fitted in milliseconds, taught by one correction) and chooses the candidate it scores highest. relative=True also gives each numeric feature relative to the other candidates of the step (its gap to the smallest and the largest).

head = CandidateHead(["kind", "distance", "reward", "dead_end"]).fit(steps)   # steps: [(candidates, chosen index)]
i, probs = head.choose(candidates)             # candidates: [{feature: value}]; probs sum to 1 over them
head.teach(candidates, 2)                      # one correction: the candidate that should have been taken

Labels come from a rule you are replacing, from people, or from outcomes — and an outcome label must be the criterion of the sub-goal the step served (progress towards the goal for "go on", a level gained for "train"): one global measure teaches the head to ignore every step that does not move it.

relative=True is off by default: whether the gaps help depends on the task, so compare both on held-out steps. The head learns the rule it is shown: trained on a rule's choices it reproduces the rule and can replace a model there, it does not beat the rule.

rows

rows(candidates)

The rows the head reads: each candidate's features and, for numbers, the gaps to the step's min and max.

fit

fit(steps)

steps: [(candidates, chosen)] — chosen: the index of the right candidate, or a set of indexes when several are as good. The others of the step are its "no" examples.

scores

scores(candidates)

The head's score of each candidate being the one to take (higher: better), in the candidates' order.

choose

choose(candidates)

→ (the index of the best candidate, [probability per candidate]) — ties go to the earlier candidate.

teach

teach(candidates, chosen)

One correction: this candidate (index, or a set of them) was the one to take. → ms.

select_features

select_features(rows, answers, options, features, min_gain=0.0)

Greedy forward selection for a FastHead: start from no feature (the answers' shares) and add, one at a time, the fact that lowers the exact leave-one-out squared error most; stop when none lowers it by more than min_gain × the error of the answers' shares. → (chosen, {fact: error after adding it}, {fact that cannot be encoded: why}). A proper score, so a rare answer counts: accuracy, the old criterion, keeps nothing where one answer is most of the examples (a fact seldom changes the majority answer). Each fact is chosen with its own encoding only (no pairwise products while choosing).