Skip to content

solvi.episode

An agent's memory as an input of its decisions: what was tried, what did not help, what worked before.

An agent's memory as an input of its decisions: what was tried, what did not help, what worked before.

A decision in solvi depends on its recorded input and nothing else — that is what makes it replay. An agent that takes many steps keeps state between them (what it tried, where it has been), and when that state lives in the harness the decisions no longer replay, the model does not see what was already tried and offers it again, and every agent writes its own loop detection. Here the memory is plain data:

from solvi.episode import Episode, EpisodeView
ep = Episode("ticket 4411")
ep.note("act", "restart the router")              # an event (kind, key)
ep.progress("the customer confirmed the fix")      # something moved: the counts "since progress" start again
res = system.ask({"message": text, "episode": ep.snapshot()})     # the snapshot is a given fact: in the trace

@cat.check(hard=True, then={"action": "handoff"})
def not_going_in_circles(episode):                 # a catalog part reads it like any fact, and stays a pure function
    return not EpisodeView(episode).stalled(8)

EpisodeView (what catalog parts use) reads a snapshot: counts of events since the last progress and in total, the facts board, the recent events, and the loop detectors — repeated (the same action again without progress), ping_pong (A → B → A → B ...), stalled (many steps without progress), revisits (the same state again), and looping (stalled and one of the first two: single detectors fire on honest repetition, the combination rarely does). A detector is a signal: answer it softly — close that option for a while, escalate — not by stopping the agent.

Chooser is the step these pieces make together: a closed list of actions, a decider proposes one, a validator turns down what is not in the list or was already done without progress (plus your own check), a rule answers when the proposal is turned down or the model escalates. Everything a part reads is in the input, so Chooser.replay() re-checks every stored decision.

LongMemory carries outcomes across episodes: what led to the goal in a context and what was a dead end, with decay, as scores that are given to the decision as a fact (not a hidden bias).

What to expect: a model that sees only the last message proposes again what has already failed; with the episode in its input and repeats turned down it stops doing that, and with the state in the input instead of the harness the stored decisions replay. The memory makes a model-driven agent sound and auditable, it does not make the model better than rules: where a hand-written script or a runbook exists, it can do as well or better. Where the bottleneck is skill rather than memory, it changes nothing.

EpisodeView

EpisodeView(snapshot)

A snapshot of an episode, read-only: what catalog parts use (EpisodeView(episode) on the given fact).

count

count(kind, key=None, since='progress')

How many events of this kind (with this key) since the last progress (since="total": in the whole episode).

get

get(name, default=None)

A fact from the board (Episode.set).

last

last(kind=None, k=5)

The last k recent events [n, kind, key] (of one kind).

repeated

repeated(kind, key, limit=2, since='progress')

This action was already taken limit times without progress.

ping_pong

ping_pong(kind='place', round_trips=3)

A → B → A → ... for at least round_trips round trips in a row since the last progress → {A, B}, else an empty set. Several events in a row with the same key count as one stay.

stalled

stalled(steps)

No progress for at least steps events.

revisits

revisits(key, kind='state')

How many times this state (key, an event of kind) was already met since the last progress. The key is required: count(kind) counts every event of a kind.

looping

looping(stalled=150, kind='place', round_trips=3, repeat=('act', None, 8))

Stalled AND (a ping-pong OR an action repeated): the combination that told real loops from honest repetition (training, walking back to heal). repeat: (kind, key, limit) — key None: any action of that kind.

Episode

Episode(name='episode', keep=40)

Bases: EpisodeView

The memory of one episode: events with their numbers, counts since the last progress and in total, a board of facts. snapshot() is what a decision is given; keep: the recent events a snapshot carries.

note

note(kind, key)

Record an event: what was done, where the agent is, what was seen (its key; a value worth deciding on is a fact: set(name, value)).

set

set(name, value)

Put a fact on the board (overwrites).

progress

progress(reason='')

Something moved: the counts since the last progress start again (dead ends before it no longer count). Say what counts as progress explicitly — a sub-goal reached, an event flag — not "anything changed": a wrong action changes the page too, and then it erases the memory of itself.

snapshot

snapshot()

The episode as a decision's input: plain JSON data, the same for the same history.

Chooser

Chooser(model, storage=None, min_confidence=0.5, check=None, repeat_limit=1, use_act=None)

One step of an agent as a solvi decision: choose an action from a closed list, with the episode in the input.

chooser = Chooser(model, storage=store)
action, who, info = chooser.choose("next step", "What should support do next?", {"restart": "restart the router",
                                   "replace": "send a new router"}, context=dialogue, rule="restart", episode=ep)

options: {option: action} (the option is what the model reads, the action what the episode counts). The producers of the choice: the model (rejected when its option is not in the list, when its action was already taken repeat_limit times without progress, when check(option, question, episode) says no, or when it escalates), then rule — the option a rule would take (one of the options: another value raises ValueError). who: "model" | "rule" | "only" (one option: no decision) | "abstain" (the model's option was rejected and there is no rule: the action is None — the caller decides, e.g. asks a person); info: the option and the producers tried. A single option is returned without a decision.

replay

replay()

Re-check every decision this chooser stored → (replayed, total, the first mismatches).

LongMemory

LongMemory(path=None, decay=0.8, cap=5.0)

Outcomes across episodes: in this context, what led to the goal (+) and what was a dead end (−).

lm = LongMemory("memory.json", decay=0.8)
lm.begin("ticket 4411")                                   # a new episode: older scores decay
lm.record(("customer", "c17"), "restart the router", +1, why="solved")
lm.scores(("customer", "c17"))                            # {"restart the router": 1.0} → a fact of the decision
lm.save()

A score is the sum of the outcomes recorded for (context, key), each decayed once per episode since, within ±cap. Labels come from outcomes — progress, a dead end — never from what the model answered. Every item keeps the episode it was last recorded in and why. A key that is not a string (a tuple, a dict: ("tool", "ping")) is kept as its JSON text, like an Episode event's key, and scores gives it back in that form. The scores are given to a decision as a fact, so the trace shows what the memory said; nothing is applied behind the decision's back.

begin

begin(name)

A new episode: every score decays.

scores

scores(context)

{key: score} of the context, without zeros, keys sorted — what a decision is given.

forget

forget(context=None)

Remove a context's items (None: everything) → how many were removed.