Skip to content

solvi.solutions.agent

The environment agent (solvi.Agent): System 1 acts on what the knowledge predicts, System 2 searches when it cannot, the agenda's gates and the action model's refusals are hard checks, every decision is stored and replays. See Using solvi: agents and knowledge.

The environment agent: System 1 acts on what the knowledge predicts, System 2 searches when it cannot, hard checks hold in both, and what the environment shows is written back — so an environment met again needs fewer slow decisions.

import solvi
from solvi.core.knowledge import RiskBudget

km = solvi.Knowledge(vocabulary={"here": lambda s, a: s["here"], "wood": lambda s, a: s["wood"] > 0})
km.goal("wood", done=lambda s: s["wood"] > 0)
km.goal("table", done=lambda s: s["table"], requires=["wood"])
agent = solvi.Agent(env, knowledge=km)                       # protection: a predicted refusal is never taken
agent.run(seed=1, steps=150)                                 # one episode → its numbers
agent.run(seed=1, steps=150)                                 # the same world again: fewer slow decisions
agent.report(); agent.replay()                               # what it did; every decision re-checked
solvi.Agent(env, knowledge=km, risk=RiskBudget(max_risk_per_episode=1.0))   # justified risk, within a budget

env is a solvi.core.Environment (reset(seed), actions(state), step(action) → Outcome). Every step is one decision of a Dispatcher (agent.dispatcher) over System 1 (agent.system), a solvi System with one question, "action" — which of the offered actions — answered over the given facts of the step:

options one row per action the environment offers: the action model's prediction (verdict accept / refuse / unknown, risk, support, reason, hard), the risk policy's choice on it (take / avoid / ask_s2, why), the hard blocks (agenda gates, a hard refusal, the failure memory), the open goal it advances and how ("skill": the goal was done right after it before; "route": the first step of a confirmed route to where a skill worked), the expected gain, System 1's and System 2's order; goals_open the agenda's open goals; at: the state's key; knowledge: the store's journal position and hash; s2_left what is left of System 2's per-episode budget (when budget= is a number).

System 1 takes the first option that advances an open goal and that the risk policy takes (with the default Protect: the action model predicts "accept"); its hard checks — not_past_a_gate, not_avoided (a hard refusal is never traded; a refusal the policy avoids), not_a_recent_failure — are catalog checks over the chosen option. When it has no such option it abstains and the dispatcher hands the step to System 2: a SearchPath over the offered actions in an exploration order (an open goal's action first, then actions never tried at this key, then the first step towards the nearest key with untried actions, ...) through the same hard checks — an action the model is unsure of ("unknown") may be tried there, an avoided one never. s2= replaces it with a SlowPath of your own (it answers "action" with the index of an option, from the same facts). When System 2 has nothing (its budget is used up, every candidate is blocked) the agent takes a reproducible pseudo-random option the hard checks allow ("fallback"), or stops the episode when there is none.

What is written back (Knowledge.observe): the action model learns from the outcome, the failure memory blocks a refused action at this key for a while, the agenda's done checks run on the new state, a goal done right after an action makes a skill, and map facts ("leads_to", "accepts") are written on the episode's map. Knowledge carried across episodes: rules (the action model, skills) are carried everywhere; map facts are scoped to the map an episode plays on (run(seed, steps, map=), default: the seed — the same seed, the same world) and the first contradiction of a carried map fact drops the carried map (a flag: its facts become hints until confirmed again).

Risk. risk=None is Protect (the verdict is final). risk=RiskBudget(...) takes a refused action when its expected gain (gain=: a function (state, action) → number; default 1.0 for an action that advances an open goal, else 0) outweighs its estimated risk within a per-episode budget; a hard prediction, a gate and the failure memory are never traded. Every risky take is decided again for real (the budget is charged) and counted.

Every decision is stored (storage=: a path or a TraceStorage, hash-chained "dispatch" records; else kept in memory) and replayable without the environment: replay() re-checks System 1's trace, the dispatch, System 2's search record and the answer from the recorded facts.

What it does not promise: it learns what the environment shows over the vocabulary you give; rules learned in one world can be too cautious in another (sufficient conditions do not transfer); a learned memory on top of written gates added nothing in the research behind it; justified risk reduces the cost of protection, not to parity with an agent without knowledge. See the guide's "Using solvi: agents and knowledge".

Agent

Agent(env, *, knowledge, s2=None, budget=None, storage=None, risk=None, key=None, gain=None, seed=0, max_options=64)

See the module docstring.

env: an Environment. knowledge: a solvi.Knowledge (goals, gates, the action model, the failure memory, the store). s2: a SlowPath of your own (default: the search over the offered actions). budget: System 2's budget per episode — a number (System 2 decisions) or a Budget (dollars, calls, ms of a model-backed SlowPath; the dispatcher's total, reset each episode). storage: where every decision is stored (a path or a TraceStorage). risk: a RiskPolicy (None: Protect). key: state → a string key (the map's places; default: the state's canonical JSON). gain: (state, action) → expected gain for the risk policy (None → the default). seed: the tie-breaks of the exploration order and the fallback. max_options: the most actions one state may offer.

new_episode

new_episode(seed=None, map=None)

Start an episode by hand (run() does it): the map it plays on (default: the seed), the risk policy's and System 2's per-episode budgets reset, the agenda's goals open again.

run

run(seed=None, steps=100, map=None)

One episode: reset the environment with seed, act and observe for at most steps steps (or until the environment says done). → the episode's numbers: steps, decisions by System 1 / System 2 / fallback, System 1's share, refused actions, risky takes, the goals done, drift.

act

act(state)

Decide the next action in state → the action (one of env.actions(state)), or None when nothing the hard checks allow is offered. The decision is stored; call observe(outcome) after the environment's step.

observe

observe(outcome)

The environment's Outcome of the last action (or True / False): written back to the knowledge (the action model, the failure memory, the agenda, skills, map facts). → what Knowledge.observe returns.

replay

replay()

Re-check every stored decision without the environment or a model → {"decisions", "ok", "mismatches": the first few [(n, mismatches)]}.

report

report()

What the agent did: per episode (steps, System 1 / System 2 / fallback decisions, System 1's share, refused actions, risky takes, goals done, drift, why it stopped), the totals, and the knowledge's report.

s1_pick

s1_pick(options)

System 1's choice: walking the options that advance an open goal in System 1's order, the first one the risk policy takes and no hard block stops — or None (System 1 abstains: System 2's turn) when there is none, or when a goal's own skill comes first and the policy sends it to System 2 (the action model is unsure of it here).

search_order

search_order(facts)

System 2's space: the offered actions in System 2's order (none when its per-episode budget is used up).

state_key

state_key(state)

The default key of a state: its canonical JSON.