solvi.episode¶
An agent's memory as an input of its decisions: what was tried, what did not help, what worked before.
An agent's memory as an input of its decisions: what was tried, what did not help, what worked before.
A decision in solvi depends on its recorded input and nothing else — that is what makes it replay. An agent that takes many steps keeps state between them (what it tried, where it has been), and when that state lives in the harness the decisions no longer replay, the model does not see what was already tried and offers it again, and every agent writes its own loop detection. Here the memory is plain data:
from solvi.episode import Episode, EpisodeView
ep = Episode("ticket 4411")
ep.note("act", "restart the router") # an event (kind, key)
ep.progress("the customer confirmed the fix") # something moved: the counts "since progress" start again
res = system.ask({"message": text, "episode": ep.snapshot()}) # the snapshot is a given fact: in the trace
@cat.check(hard=True, then={"action": "handoff"})
def not_going_in_circles(episode): # a catalog part reads it like any fact, and stays a pure function
return not EpisodeView(episode).stalled(8)
EpisodeView (what catalog parts use) reads a snapshot: counts of events since the last progress and in total, the
facts board, the recent events, and the loop detectors — repeated (the same action again without progress),
ping_pong (A → B → A → B ...), stalled (many steps without progress), revisits (the same state again), and
looping (stalled and one of the first two: single detectors fire on honest repetition, the combination rarely does).
A detector is a signal: answer it softly — close that option for a while, escalate — not by stopping the agent.
Chooser is the step these pieces make together: a closed list of actions, a decider proposes one, a validator turns
down what is not in the list or was already done without progress (plus your own check), a rule answers when the
proposal is turned down or the model escalates. Everything a part reads is in the input, so Chooser.replay()
re-checks every stored decision.
LongMemory carries outcomes across episodes: what led to the goal in a context and what was a dead end, with decay,
as scores that are given to the decision as a fact (not a hidden bias).
What to expect: a model that sees only the last message proposes again what has already failed; with the episode in its input and repeats turned down it stops doing that, and with the state in the input instead of the harness the stored decisions replay. The memory makes a model-driven agent sound and auditable, it does not make the model better than rules: where a hand-written script or a runbook exists, it can do as well or better. Where the bottleneck is skill rather than memory, it changes nothing.
EpisodeView ¶
A snapshot of an episode, read-only: what catalog parts use (EpisodeView(episode) on the given fact).
count ¶
How many events of this kind (with this key) since the last progress (since="total": in the whole episode).
repeated ¶
This action was already taken limit times without progress.
ping_pong ¶
A → B → A → ... for at least round_trips round trips in a row since the last progress → {A, B}, else an
empty set. Several events in a row with the same key count as one stay.
revisits ¶
How many times this state (key, an event of kind) was already met since the last progress. The key is
required: count(kind) counts every event of a kind.
looping ¶
Stalled AND (a ping-pong OR an action repeated): the combination that told real loops from honest repetition (training, walking back to heal). repeat: (kind, key, limit) — key None: any action of that kind.
Episode ¶
Bases: EpisodeView
The memory of one episode: events with their numbers, counts since the last progress and in total, a board of
facts. snapshot() is what a decision is given; keep: the recent events a snapshot carries.
note ¶
Record an event: what was done, where the agent is, what was seen (its key; a value worth deciding on is a fact: set(name, value)).
progress ¶
Something moved: the counts since the last progress start again (dead ends before it no longer count). Say what counts as progress explicitly — a sub-goal reached, an event flag — not "anything changed": a wrong action changes the page too, and then it erases the memory of itself.
snapshot ¶
The episode as a decision's input: plain JSON data, the same for the same history.
Chooser ¶
One step of an agent as a solvi decision: choose an action from a closed list, with the episode in the input.
chooser = Chooser(model, storage=store)
action, who, info = chooser.choose("next step", "What should support do next?", {"restart": "restart the router",
"replace": "send a new router"}, context=dialogue, rule="restart", episode=ep)
options: {option: action} (the option is what the model reads, the action what the episode counts). The producers
of the choice: the model (rejected when its option is not in the list, when its action was already taken
repeat_limit times without progress, when check(option, question, episode) says no, or when it escalates), then
rule — the option a rule would take (one of the options: another value raises ValueError). who: "model" | "rule" |
"only" (one option: no decision) | "abstain" (the model's option was rejected and there is no rule: the action is
None — the caller decides, e.g. asks a person); info: the option and the producers tried. A single option is
returned without a decision.
replay ¶
Re-check every decision this chooser stored → (replayed, total, the first mismatches).
LongMemory ¶
Outcomes across episodes: in this context, what led to the goal (+) and what was a dead end (−).
lm = LongMemory("memory.json", decay=0.8)
lm.begin("ticket 4411") # a new episode: older scores decay
lm.record(("customer", "c17"), "restart the router", +1, why="solved")
lm.scores(("customer", "c17")) # {"restart the router": 1.0} → a fact of the decision
lm.save()
A score is the sum of the outcomes recorded for (context, key), each decayed once per episode since, within ±cap.
Labels come from outcomes — progress, a dead end — never from what the model answered. Every item keeps the episode
it was last recorded in and why. A key that is not a string (a tuple, a dict: ("tool", "ping")) is kept as its JSON
text, like an Episode event's key, and scores gives it back in that form. The scores are given to a decision as a fact, so the trace shows what the memory
said; nothing is applied behind the decision's back.