Skip to content

solvi.core.environment

The Environment protocol and its Outcome: a world acted in step by step (see Building blocks).

The Environment protocol: a world an agent acts in, one step at a time — what the environment agent of solvi 1.0 (solvi.Agent, the high level) runs on, and what an action model learns from.

from solvi.core import Environment, Outcome

class Corridor:                              # the smallest environment: walk right to the door
    def reset(self, seed=None):
        self.pos = 0
        return {"pos": 0}
    def actions(self, state):
        return ["left", "right"]
    def step(self, action):
        self.pos = max(0, self.pos + (1 if action == "right" else -1))
        return Outcome({"pos": self.pos}, accepted=True, effect={"pos": self.pos}, done=self.pos == 3)

solvi.testing.conformance.check_environment checks one (the same seed and the same actions give the same outcomes; every step returns an Outcome; a refused action leaves the state as it was).

Outcome dataclass

Outcome(state: Any, accepted: bool = True, effect: Any = None, done: bool = False, info: dict = dict())

What one step did. state: the state after it (what actions and the next decision read); accepted: did the environment take the action (False: refused — a wall, a rule, a tool that said no — and the state did not move); effect: what changed, as data the environment itself reports (an action model learns from it and is checked against it); done: the episode is over; info: anything else (a score, a message), not read by solvi.

Environment

Bases: Protocol

A world acted in step by step.

You implement: reset(seed=None) → the first state (the same seed gives the same episode); actions(state) → the actions possible in that state (a list; each one a plain value: a name, a tuple, a dict); step(action) → an Outcome(state, accepted, effect, done).

You get for free (with solvi.Agent, the environment agent of the 1.0 high level): System 1 = the action model's prediction plus the open goals of the agenda, hard checks = the agenda's gates and the action model's refusals (no action past a gate, no action predicted to be refused), System 2 (a search) when System 1 is unsure, a recorded and replayable step per decision, knowledge carried across episodes.

Stability: stable (solvi.Agent calls reset(seed), actions(state) and step(action) only).