Skip to content

solvi.learning

Learning from corrections with gates and rollback (experimental): System.learning(...).

Learning from corrections, with gates and rollback (experimental): System.learning(...).

loop = system.learning(store)              # nothing happens until you run it; System.teach now only stores
rep = loop.run()                           # labels → a proposed update → gates → promoted or rejected, recorded
print(rep)                                 # what changed, each gate's numbers
loop.versions()                            # every promoted state; loop.rollback(2) restores one

Labels come only from outside the model, read from a TraceStorage: human corrections (System.teach, save_correction), known outcomes (source="outcome") and your rules' rejections (source="rule"). A correction from any other source is refused and listed in labels()["rejected"]; the stored decisions — the system's own answers — are never read as labels, so self-training is impossible by construction, not by a setting.

Each label is put, by a hash of its stored id, into "train", "calibration" or "holdout" (default 50 / 20 / 30%): a held-out label is never trained on, in this update or any later one.

The ladder, per question, by the number of training labels: fewer than fit_below (50) — the shift / scale of the decider (DecisionPart.fit); up to memory_below (1000) — fit plus a memory of the corrected cases (solvi.memory, mode "check" by default); beyond — the adapter hook when you give one (a callable (part, [(text, answer)]) → a JSON-able description; an object with state(part) / restore(part, state) is rolled back too), else the memory.

The gates (every one must pass, else the update is undone and recorded as rejected): consistency the new training labels agree with a memory of the labels already learned: at most max_conflict (20%) of them may be contradicted by close, agreeing earlier corrections (B9: wrong corrections hurt more than right ones help); heldout on the held-out labels, asked through the whole system: the share answered alone and right, minus the share answered alone and wrong, must improve by at least min_gain (0.01), with at least min_holdout (5) held-out labels; honesty the honesty numbers (solvi.honesty: confident errors, coverage at risk, quote support) on the held-out labels — and on your own honesty set when given (gates={"honesty": cases or a set file}) — must not get worse by more than tolerance (0.02); act_guard a part calibrated with act_guard / calibrate_for is recalibrated on the calibration labels (at least min_calibration, 30), with the same risk: an old threshold says nothing about a changed signal; a part's conformal answer sets are recalibrated on them too (with fewer labels they are dropped, and the gate's record says so — they no longer hold for the changed probabilities); size shadow run: the stored decisions whose input is held out for a learned question (the same split as the labels: per question, by content_key) — at most shadow_limit, 500 — are asked with the current and the candidate state and compared (solvi.diff.compare) on those questions and the unlearned ones; at most max_change (30%) of them may change — one update may not move more.

The candidate is built and gated on a shadow of the system: copies of the decision parts (their thresholds, memory, and the model's adaptations; the checkpoint itself is shared) in a shallow copy of the catalog and the System. The live parts are not touched until the update is promoted, so asks that run meanwhile (another thread, the server) see the state in force, never an un-gated candidate. An adapter hook runs on the shadow part: its effect reaches the live part through its state(part) / restore(part, state); a hook without them only keeps what it changes in objects the shadow shares with the live part (the checkpoint), and such changes are not isolated. When a gated candidate cannot be carried over to the live parts (a hook without state / restore changed the shadow part itself, so the promoted fingerprint is not the candidate's), run() puts the live parts back as they were, records the update as not promoted (gate "promotion") and raises RuntimeError.

A promoted update gets the next version number; every run that proposes an update is recorded in the changelog (a TraceStorage, by default the same store: kind "update", hash-chained with the decisions) with its gates' results, the fingerprints before and after, the labels it trained on and — for a promoted state — the state itself (adaptations, thresholds, memory), so rollback(version) restores any promoted version, in this process or another. Each decision records the fingerprint of the part that made it, which names the version (loop.version_of(fp)).

Label dataclass

Label(id: str, question: str, init: dict, answer: object, source: str, by: str | None = None, of: str | None = None, time: float | None = None, split: str = 'train')

A trusted label: a stored correction (or rule rejection) of one question.

UpdateReport dataclass

UpdateReport(action: str, version: int | None = None, current: int | None = None, questions: dict = dict(), gates: dict = dict(), fp_before: str | None = None, fp_after: str | None = None, rejected_labels: list = list(), record_id: str | None = None)

What one run of the loop did: "none" (nothing new to learn), "promoted" or "rejected".

Learning

Learning(system, storage=None, parts=None, ladder=None, gates=None, changelog=None, holdout=0.3, calibration=0.2, gate_teach=True, harvest_rules=None)

The learning loop of a System (see the module docstring). Made by System.learning(...); experimental.

current property

current

The version whose state is in force (the latest one with this fingerprint), or None when the live state is not a recorded version (before the first run, or after changing the parts by hand).

detach

detach()

Stop gating System.teach (the loop's records stay in the changelog).

labels

labels()

The trusted labels of the loop's questions → {"labels": [Label], "rejected": [(id, why)]}. Only corrections (teach records) from TRUSTED_SOURCES, with an answer the question and its decision can take.

fingerprint

fingerprint()

The fingerprint of the loop's current state: every learned part's fingerprint.

history

history()

Every record the loop wrote to its changelog, in order (dicts: "action", "version", "promoted", "gates", ...).

versions

versions()

The promoted versions → [{"version", "fp", "time", "action", "id"}] (a rollback is listed as the version it restored).

version_of

version_of(part_fp)

The versions in which some learned part had this fingerprint (as recorded in a decision's trace).

rollback

rollback(version)

Restore a promoted version's state (adaptations, thresholds, memory) and record the rollback. → the version.

run

run()

Collect the labels, propose an update by the ladder, run the gates, promote or undo it, record it. → UpdateReport. The system's answers do not change unless the update is promoted.

split_of

split_of(label_id, holdout=0.3, calibration=0.2)

A label's split by a hash of its key (content_key) — "holdout", "calibration" or "train" — the same in every run.