Skip to content

Knowledge: what a system learned, and from whom

solvi.core.knowledge is the low level of what a system knows across decisions: facts people told it, outcomes it observed, rules from a written spec, what an environment accepts and refuses, goals and gates, and the plans that failed. Every piece is plain data with a journal, so a decision can say what it was given, and a wrong item can be taken back with everything that rests on it. It never learns from the system's own unverified answers.

The high level over these pieces — solvi.Knowledge (one object for the store, the agenda, the action model and the failure memory), solvi.Agent (an agent acting in an environment on it), build(knowledge=) and Guard(knowledge=) — is on its own page: Using solvi: agents and knowledge.

The knowledge store

from solvi.core.knowledge import KnowledgeStore

ks = KnowledgeStore("decisions.jsonl")          # the journal in any TraceStorage backend (or None: in memory)
vip = ks.add("fact", {"s": "c17", "r": "tier", "o": "vip"}, source="person", by="crm")
ks.add("fact", {"s": "c17", "r": "tier", "o": "regular"}, source="outcome")   # the world changed: the old fact is refuted
res = system.ask({"ticket": text, "knowledge": ks.snapshot(about="c17")})      # given to a decision: in its trace
ks.retract(vip, by="ann", why="wrong customer")                                # with everything derived from it
changes, justification_only = ks.redecide(store, {vip}, system)                # what the retraction moves
ks.verify(); ks.report()

An item is a fact ({"s", "r", "o"}), a rule, a skill, an action or an episode; its source is "person", "outcome" (what really happened), "spec" (a written specification) or "verified" — a System 2 answer given alone under its guarantee, named by of=<stored decision> and checked in the store. Anything else is refused and the refusal is journaled: the system's own answers never become knowledge, by construction. ks.item(id) is the full record: kind, body, scope, source, who, evidence, derived_from, confidence counts (confirmations, refutations, uses, uses that were wrong, a guarantee level), status, version and supersedes, and two time ranges — when it held in the world (valid=(since, until), as you give it) and when the store learned and invalidated it (journal positions).

Contradictions resolve by rank: an observation refutes whatever it contradicts (the world changed), a higher-ranked source (outcome > person > spec > verified) refutes a lower one, a producer's new version (supersedes=) refutes the one it names. Anything else is a dispute: every side is disputed, none is given to System 1 as a fact, and one person question is opened per dispute event — the same two items disputed again after a resolution are asked again; ks.disputes() lists the open ones and ks.resolve(key, value, by=...) is the answer.

A retraction is exact: the store after retract(x) has the fingerprint of the store rebuilt from its journal as if x had never been written (ks.rebuild(skip={x})). benchmarks/knowledge/retraction.py checks it on a synthetic store of 10,000 items (facts, derivations up to depth 5, disputes, resolutions, promotions, clock expiry, drift flags): 1,000 of 1,000 random retractions exact, and verify() true at the end. Nothing is deleted: the journal keeps every entry and its chain stays whole. A retraction's cost grows with the connected component it touches (premise edges and shared keys): there, a retraction touched a component of 6,057 of the 10,000 items at the median and took 0.19 s (rebuilding the whole store: 0.79 s); retractions in small components took about 0 s (Spearman ρ between component size and time 0.88). The store has been measured up to 10,000 items, not beyond.

A decision is given ks.snapshot(...) as a fact: the active items that match (facts), the hints (hypotheses, expired or stale items, with why), the ids they rest on, the store's fingerprint and its journal position — so the trace shows what the memory said, and the decision replays. ks.redecide(store, retracted, decide) finds the stored decisions whose snapshot rested on retracted items, re-builds each snapshot as it would have been without them, re-runs the decision on the old and the new snapshot, and splits the list into "the answer changes" (for a reviewer) and "only the justification changes" (recorded). The journal can share the decisions' TraceStorage: records of kind "knowledge" in the same hash chain; store.iter() and the reports read decisions only.

Staleness: flags, expiry and rollback

flag = ks.flag(scope={"question": "intent"}, why="open-set gate: new kinds of input")   # a drift flag
ks.stale(item)          # True while a pending flag covers it: read the flag, not the status
ks.rollback(update, by="supervision", why="coverage loss")   # the previous version comes back — only as a hint
ks.end_flag(flag, by="ann", why="re-confirmed on 30 new labels")

A flag (a drift or open-set flag, reconfirm(id) for one item) makes the items it covers that were learned before it hints, not facts, until a newer observation or a person confirms each one, or the flag ends. Two rules close the holes a stream with an abrupt shift showed in the research behind this store: a rollback while a flag covers the rolled-back version (or its scope) restores the previous version only as a hint — a stale version is never restored as a fact — and staleness is read from the flags (stale, usable), never from the item's status, which a later write can set back. reconfirm_after (per kind and source, RECONFIRM_DEFAULTS; the store's clock moves with tick) expires items that weaker evidence keeps: a "verified" fact after 2,000 decisions; observations, person labels and spec facts have no clock — the next contradicting observation refutes them. A fact nobody observes again stays believed: a world where a first contact is costly needs a clock of its own.

What a flag does not do: it does not close the window between an abrupt shift and its detection. A calibrated promise ("≤ 5% wrong among the answers given alone") does not hold in the first decisions after an abrupt shift to unseen inputs, before any monitor flags it (see the output).

Write gates

KnowledgeStore(gates=[...]) takes WriteGates — admit(store, item, shadow) → Verdict(admit, reason, measured). SourceGate (the source check) always runs first; ConsistencyGate refuses an item whose premises are missing, retracted or refuted, or that would refute its own premise. A fact or an episode a gate admits is present at once; a rule, a skill or an action (behaviour-changing kinds) is promoted with every gate's verdict in the journal, and one a later gate holds stays a hypothesis. The shadow gate — measure a behaviour-changing item on a shadow of the system before it is promoted — is experimental: solvi.experimental.learning.ShadowGate.

The action model: what the environment accepts

from solvi.core.knowledge import ConservativeActionModel, Vocabulary

vocab = Vocabulary({"status": lambda state, args: state["orders"][args["order_id"]]["status"],
                    "own_method": lambda state, args: args["payment_id"] in state["user"]["methods"]},
                   hard=set())                    # names whose violation is a hard rule, never traded
model = ConservativeActionModel(vocab, store=ks)
model.observe(state, "cancel_order", args, accepted=False)          # what the environment did
p = model.predict(state, "cancel_order", args)
p.verdict, p.risk, p.support, p.reason, p.hard, p.effects            # "accept" | "refuse" | "unknown", ...
model.commit()                                                       # each action's model as an "action" item

Per action, the values each condition took in accepted calls are allowed; a refused call is explained by the conditions whose value was never accepted, and the minimal such sets are its refusal signatures. predict says "accept" when every condition holds an allowed value, "refuse" when the violated conditions contain a signature, and "unknown" otherwise — "unknown" goes to System 2 or a person. A condition that cannot be read on a state (an entity not looked up) is never allowed until an accepted call shows it, so the model says "unknown" there. The vocabulary is an explicit argument and its predicates' fingerprints are in every action item.

Measured (benchmarks/knowledge/taubench_action_model.py, τ-bench retail, no LLM): probes around the gold steps of 400 training tasks (41,626 transitions, 26,772 refused) and of the 115 test tasks (11,763 held-out transitions, 7,626 refused). With a vocabulary of typed predicates from generic templates over the tools' JSON schema (existence, status, lengths, enums, pair features of two conditions on one argument, money comparisons), refusal precision and recall over the answered held-out transitions are 1.000 and 1.000 (7,602 of 7,602 answered refusals, no false refusal), the model abstains on 30 (0.26%), and it replays all 41,626 training transitions. Without pair features it abstains on 2.6% (precision still 1.000); without money comparisons 10.7% and makes 1 false refusal (precision 0.9998); without both, 12.9% and again 1 false refusal. Read it this way:

  • vocabulary-bound: it learns only checks the vocabulary can express. That vocabulary was written by someone who had read the tools' code; the result says what a conservative learner recovers given such a vocabulary, not that it discovers the checks. When the vocabulary cannot explain a refusal, the model answers only condition vectors it has seen exactly, and those answers can be wrong (the false refusals above);
  • it learns what the environment checks: a rule the environment does not enforce — confirm with the customer, authenticate first, one customer per conversation — is never refused by it. Those come from a written policy: agenda gates or hard checks;
  • necessary conditions transfer, sufficient conditions are not promised to: conditions that held at every success by accident (an item carried, a place's name) make it over-conservative in a new world — "unknown", not wrong (the km_place arm of the toy world below);
  • effects are predicted when every accepted call with the same condition values (else of the action) had the same effect; on the τ-bench probes, the touched order's new status was exact on 2,585 of the 2,587 accepted calls it predicted an effect for (1,544 accepted calls got no effect prediction: their condition values had been seen with different effects).

Protection by default, justified risk as an option

from solvi.core.knowledge import LearnedGate, Protect, RiskBudget

policy = Protect()                                      # the default: refuse → avoid, unknown → System 2
policy = RiskBudget(max_risk_per_episode=1.0, min_gain_ratio=1.0, min_support=3)
d = policy.decide("descend", prediction, gain=0.3)      # RiskDecision: take / avoid / ask_s2, with the rationale
policy.new_episode()

gate = LearnedGate("depth", expiry=10, floor=lambda level: level + 1)
gate.failed(context=level, level=depth); gate.arrived(level, depth, failed=False); gate.end_episode()
gate.predict(level, depth)                              # a Prediction a policy decides on

Knowledge used only as hard gates removes the failures it targets and can cost a metric that rewards risk. Every prediction therefore carries a verdict and an estimate (risk, support); Protect keeps the verdict, RiskBudget takes a refused action when its expected gain is at least min_gain_ratio × its risk and the risk fits the episode's budget, sends a refusal resting on fewer than min_support cases to System 2, and never takes a hard prediction (an instant-harm or a policy rule) — in any mode. A gate learned from failures is bounded: it reads the last expiry episodes only, never falls below floor, and is reopened by evidence (recorded arrivals with a low failure rate).

On a toy dungeon (benchmarks/knowledge/risk_dungeon.py, 20 streams × 30 games; a toy with hand-set hazards, not evidence about your domain): the mean deepest level was 3.47 without knowledge (482 deaths from the hazards the knowledge targets), 5.67 with Protect (40 such deaths) and 6.65 with RiskBudget (196 such deaths; 5.8 risky actions taken per game). From the first to the last third of the streams Protect fell from 6.74 to 5.02 as learned refusals accumulated, and with a gate that never expires and has no floor from 6.70 to 2.99 — the gate tightening itself — while RiskBudget went from 6.29 to 6.94. Measure RiskBudget on your own metric before you turn it on.

The agenda: goals, gates, order

from solvi.core.knowledge import Agenda

ag = Agenda(ks)                                         # goals and gates are rule items, goal states fact items
ag.goal("authenticated", done=lambda s: s.get("user_id") is not None)
ag.goal("exchanged", done=lambda s: s.get("status") == "exchange requested", requires=["authenticated"])
ag.gate("auth_first", lambda s: s.get("user_id") is not None, blocks=["modify_*", "exchange_*"])
ag.update(state)                                        # runs the done checks; every change is journaled
ag.open(state), ag.blocked(state), ag.done()
ok, why = ag.allows(state, "exchange_delivered_order_items")   # no action past a gate
ag.override("auth_first", by="lead", why="verified by phone")  # a person, with a recorded reason
report = ag.dry_run(recorded_successes)                 # [(state, action)] that worked → how often each gate blocks them

A goal is done only when its code check says so, never because a model claims it. A gate is a hard check over the state: allows refuses an action it blocks while its check is False, and a goal it blocks is not open. Validate a gate before you make it hard: a gate written from policy text ("remind the customer to confirm they listed every item") can block the calls a customer wanted. dry_run replays recorded successful actions through the gates and reports, per gate, how many it would have blocked; a gate that would have blocked many is not ready. In the toy world below the gates are the written rules, so "0 actions past a gate" holds by construction.

Failure memory: do not repeat what just failed

from solvi.core.knowledge import FailureMemory

fm = FailureMemory(window=20, unit="steps", max_blocked=8, min_open=1)
fm.install(cat, plan="plan", then={"plan": "ask_person"})   # a hard check over the given fact "recent_failures"
res = system.ask({"situation": s, **fm.given(options=plans)})
fm.failed(plan, why="the door stayed shut")             # or fm.succeeded(plan): evidence reopens it
fm.step()

A failure blocks its plan for window steps (or episodes), then expires; at most max_blocked plans are blocked at once, and given the plans on offer at least min_open stay open — a memory fed by its own failures cannot tighten without bound. The blocked plans are a given fact, so the check is a pure function and the decision replays.

The toy world: memory, agenda and failure memory together

benchmarks/knowledge/toy_crafting.py runs the protocol of the research program's crafting-game runs on a toy world (8 places, 6 operators with the environment's own checks, 150 steps; 10 streams × 20 worlds) using only these pieces. It is not Crafter: Crafter installs, but the agent that played it there was a thousand lines of hand-written reflexes that are not part of solvi. Over the last third of each stream:

arm achievements (of 6) steps to the first stone System 1 share failed attempts per 100 steps
nothing carried 5.40 33.6 33% 16.5
the action model carried across worlds 6.00 15.6 59% 0.01
nothing carried, agenda (gates = the written rules) 5.77 15.6 48% 0
carried, agenda 5.97 15.6 59% 0
carried, the place's name in the vocabulary 5.90 24.1 50% 8.4

A proposer that does not learn (a random operator of an open goal, like a model asked again without memory) repeated a failure in the same situation within 10 steps 29.3 times per 100 steps; behind the failure memory's hard check, 0.53 times. A changed world (5 episodes in world A, then 5 in world B with the map carried): the first contradicted arrival dropped the carried map on 10 of 10 streams (18.6 claims dropped, 10.2 refuted in B on average); System 1 walked a wrong carried claim once per stream — the move that found the change — and never after the drop. Every knowledge journal verified. The toy's numbers say the pieces work together as described; how much they gain depends on the environment and the agent around them.

An agent's memory as an input: episodes

A decision replays because it depends on its recorded input only. An agent that takes many steps keeps state between them — what it tried, where it has been — and when that state lives in the harness, the decisions stop replaying, the model does not see what was already tried, and every agent writes its own loop detection. solvi.core.knowledge.episodes keeps that state as plain data that is given to each decision:

from solvi.core.knowledge.episodes import Chooser, Episode, EpisodeView, LongMemory
ep = Episode("ticket 4411")
ep.note("act", "restart the router")                       # an event
ep.progress("the customer confirmed")                      # explicit progress: the counts "since progress" start again
res = system.ask({"message": text, "episode": ep.snapshot()})

@cat.check(hard=True, then={"action": "handoff"})          # a part reads the snapshot like any fact
def not_in_a_loop(episode):
    return not EpisodeView(episode).looping(stalled=20)

EpisodeView gives the counts (since the last progress and in total), the facts board and the detectors repeated, ping_pong, stalled, revisits, and looping (stalled and one of the first two — single detectors fire on honest repetition). Chooser(model, storage=...).choose(name, task, {option: action}, context=..., rule=..., episode=ep) is the step built from these: the model proposes an option, a validator turns down what was already done without progress (and what your check refuses), the rule's option answers otherwise; chooser.replay() re-checks every stored step. LongMemory keeps outcomes across episodes — record(context, key, +1 / −1), decayed per episode — and scores(context) is given to the decision as a fact (a key that is not a string — a tuple, a dict — is kept as its JSON text, like an event's key).

Say what progress is — a sub-goal reached — and not "something changed": a wrong action changes the page too, and then erases the memory of itself. Without the episode in its input a model proposes again what has already failed; the memory keeps it from that and keeps every step replayable, but it does not make a model-driven agent better than rules a person wrote for the same task — where such rules exist, use them. A finished episode can be kept in the knowledge store as a record (ep.record(ks, outcome, stored_ids=...), see Knowledge).

A map the agent builds: worldmap

An agent that works in the same environment again — a site, an internal tool, a command line, a file tree — finds its structure anew on every task unless it keeps a map. solvi.core.knowledge.worldmap.WorldMap is written as the agent acts: every edge is a claim "(state, action) leads to state" with a status (hypothesis, confirmed), a source (seen, observed, told, human) and its evidence, and every write is an entry of a hash-chained journal. The journal is what a saved map is loaded from: load checks the chain and rebuilds the claims by replaying it (an edge edited in the file changes nothing; a broken chain raises), verify() also compares the map with its journal, and rebuild(upto=n) gives the map as it was after the first n entries.

from solvi.core.knowledge.worldmap import WorldMap
m = WorldMap("console.map.json")              # loaded when the file exists; m.save() writes it
m.see(page, "Billing", to="/billing")         # on offer here (`to` when the environment shows it, as a link does)
m.arrive(page, "Billing", "/billing")         # taken: confirmed — or refuted, whoever made the claim
m.next(page, {"/billing/refunds"})            # the action towards a target over what is known, else None
m.explore(page)                               # ... towards the nearest claim nobody has checked
m.human(page, "Reports", "/audit", note="Anna")    # a person's or a document's claim: a hypothesis like the others
m.snapshot(page, targets)                     # the part a decision needs, as a given fact

The adapter — list a state's actions, take one — is yours; the map only knows what these calls told it. A state or an action is a string, a number or a tuple of those (("room", 3)); save() and a later load keep them as they are, and anything else is refused when it is reported. Keep one map across the tasks: the gain is the map carried between tasks in a deep environment met again (a command line, a file tree, a documentation site). It does not shorten a first exploration, it does nothing where every state is one step away, and it does not choose which state a task needs. With knowledge= the map's claims are also fact items of a knowledge store, WorldMap.view(store) is the map those facts give, and m.drop(why) turns a carried map's confirmed claims back into hypotheses when the world may have changed (see Knowledge).

The world map, episodes and corrections in the store

WorldMap(knowledge=ks, scope=...) writes every claim with a destination as a "leads_to" fact (an arrival as an outcome, a document's claim as spec, a person's as person); WorldMap.view(ks, scope) is the map those facts give. m.drop(why) is what a drift flag does to a carried map: confirmed claims become hypotheses until the next arrival re-confirms or refutes them (with a store, a flag on the map's scope). Episode.record(ks, outcome, stored_ids=...) keeps a finished episode as an "episode" item — a record and evidence, never an answerer. attach(part, knowledge=ks) makes a memory of corrections whose cases are the store's correction facts: retracting one removes its case.

What is not here, because it was measured and failed or was not measured: recalibrating System 1 continuously from memory (it broke the promise when the threshold was refitted), answering from episodes under a per-decision guarantee (no threshold could be certified), facts as a growth channel on classification and matching streams (no gain over a static System 1), compiling code from examples, and the knowledge memory at 10⁵–10⁶ items (not measured).