solvi.agents¶
Guarding an agent's tool calls: the agent proposes a call, solvi checks it and makes it. The adapters
(solvi.agents.pydantic_ai, solvi.agents.langgraph, solvi.agents.openai_agents) import their framework when used and
are described in the guide.
Guarding an agent's tool calls: the agent proposes a call, solvi checks it and makes it.
from solvi.agents import Guard
guard = Guard(storage="calls.db")
@guard.tool(ground=["iban"]) # the IBAN must be quoted from the conversation
def send_payment(iban: str, amount: float, currency: Literal["EUR", "USD"] = "EUR") -> str:
'''Pay an invoice.'''
return bank.pay(iban, amount, currency)
@guard.policy("send_payment") # a hard check: False → deny
def under_hard_cap(amount: float) -> bool:
return amount <= 10_000
@guard.policy("send_payment", on_fail="escalate")
def within_daily_budget(amount: float, spent_today: float) -> bool: # spent_today: a fact the app gives
return amount + spent_today <= 2_000
d = guard.call({"name": "send_payment", "arguments": {"iban": "DE89…", "amount": 250}},
context=messages, facts={"spent_today": 400.0})
d.outcome # "allow" (solvi ran send_payment and d.result is its return value), "deny" or "escalate"
d.reasons # why, in words; d.message() is the text for the model
d.audit() # the solvi audit of the decision; d.response is the whole Response (trace, replay)
A proposed call is data — {"name": tool, "arguments": {...}} (OpenAI, Anthropic and MCP shapes are read too, see
ToolCall.parse) — never code: solvi looks the tool up in its catalog, and only a registered function ever runs.
Each tool is a small solvi System with one question, verdict ∈ {allow, deny, escalate}; the checks of a call are its
catalog, in this order (a failed hard check decides; when several fail, the first in this order):
arguments_valid the arguments validate against the tool's types (pydantic; unknown arguments are errors)
and hold no invisible (format, Unicode Cf) characters → deny
arguments_grounded every ground= argument is in the conversation as a token or number token (a quote
with offsets; an empty string never is; Unicode spaces are read as plain spaces), in a
message of a role in ground_from → deny
arguments_from_user tools with tool_values="escalate": a user-only argument written only in a tool output
(not by the user) → escalate instead of deny (a person decides; never allowed on its own)
no_injected_arguments ... and not only in tool outputs when any tool output in the conversation carries
instruction-like text (solvi.perturb.injection_spans) → escalate
no_instructions_in_tool_outputs tools with injections="any": no tool output in the conversation carries such text → escalate
not_made_before tools with once=True: a call with these arguments was already made → escalate
user_confirmed tools with require_confirmation: the user explicitly accepted a message of the assistant
that names the call's values (solvi.agents.confirm) → deny
your policies ordinary solvi hard checks over the arguments and the facts your app gives (deny first,
then escalate); guard.fn adds computations they read
request_authorizes with an authorizer (a decider's yes / no, act_guard, perturb): "does the conversation
authorize this call?" — no → escalate; an escalated or unsure decider → escalate
The hard guarantee is provenance: an argument grounded only from the user (ground_from=("user",)) is never taken from
a tool output, whatever the output says. Recognising instruction-like text is a heuristic second line (patterns: a
paraphrase, base64, spaced-out letters pass it) — not sufficient on its own: declare high-impact arguments as
user-grounded and add policies.
The rule verdict answers "allow" with the grounded arguments as its evidence (each quote is checked again by solvi's
grounding: literally at its offsets). A question that abstains (a check could not be evaluated, a fact the policies
need was not given, the decider escalated) is an escalation. Every decision is a full solvi response: stored with its
trace and the outcome in a TraceStorage (hash-chained), replayable (guard.replay(id)), with the audit.
Given facts of a call (the input of its trace): tool_name, tool_arguments (as proposed), conversation (the context as one
text), conversation_roles ([[start, end, role]] of each message in it), user_request (the user's messages) and the facts
your app passes (facts=). Computed facts: argument_errors, call_arguments (the validated arguments), one fact per
argument (its validated value, named after the argument), grounding (tools with ground=), proposal (the call as text).
ToolCall
dataclass
¶
A proposed tool call: the tool's name and its arguments (a dict; a JSON string is parsed), and an id if the agent gave one.
parse
classmethod
¶
{"name", "arguments"} (MCP, a plain dict), {"type": "function", "function": {"name", "arguments": "
Message ¶
Bases: tuple
A (role, text) message of a Session's context, with a flag: tainted — the message, before the session cut it
to its size cap, carried instruction-like text (the guard treats it as tainted even if the cut kept none).
Tool
dataclass
¶
Tool(name: str, func: Callable | None, model: Any, description: str = '', ground: dict = dict(), injections: str = 'grounded', authorize: bool | None = None, match: dict = dict(), schema_error: str | None = None, locale: str | None = None, scan_user: bool = False, tool_values: str = 'deny', ground_last: int | None = None, once: bool = False, confirm: tuple | None = None)
A tool in the guard's catalog: its name, the function solvi runs when a call is allowed (None: the framework runs it — adapters, the MCP proxy), the pydantic model of its arguments (None until a schema is known — MCP), which arguments must be quoted from the conversation, and from which roles.
GuardDecision
dataclass
¶
GuardDecision(outcome: str, tool: str, arguments: Any, reasons: list, response: Any, evidence: list = list(), executed: bool = False, result: Any = None, error: str | None = None, stored_id: str | None = None, approved_by: str | None = None, id: str | None = None, resolved: bool = False, failed: list = list(), catalog: Any = None)
The guard's decision on one proposed call. outcome: "allow" | "deny" | "escalate"; reasons: why, in words (empty
for an allowed call); call: the candidate call (the validated arguments when they validate); response: the solvi
Response (answers, flow, trace — response.trace.replay(...), audit()). For an allowed call made by solvi:
executed, result (the tool's return value) or error. stored_id / trace_hash: where it is in the guard's store.
policy_only
property
¶
An escalation by your policies alone (@guard.policy(on_fail="escalate")): no provenance, injection, schema
or authorizer check failed and nothing abstained. Only such an escalation may be covered by a standing
approval ("always approve this tool"); any other needs a person for this very call.
advice ¶
What the agent should be told: message() and, for a refused call, what to do next — per failed check
(NEXT_STEP): fix the arguments, use the values as they were written, propose the call and wait for the user's
yes, do not repeat a call made, follow a policy's reason or tell the user what cannot be done, wait for a
person. None for an allowed call.
feedback ¶
The messages that bring a refused call back into the agent's conversation (chat-completions shape) → []
for an allowed call; for a refused tool call (one with an id) the tool's answer, {"role": "tool",
"tool_call_id", "name", "content": advice()}; for a refused reply — the agent's own text, checked as a call
without an id (guard.declare("respond", schema=...)) and not sent — a note in reply_role ("user" by
default, the role every chat API accepts mid-conversation; "system" or "developer" where yours takes it)
that says it comes from the guard, not the user, that the user has not seen the reply, why, and what to do.
Append them to the history the model reads next; the refused reply itself is not part of the conversation.
approval_key ¶
What a person's approval of this escalation covers: a hash of the tool, the call's id, its arguments and the reasons it escalated for. An approval given for one key does not cover a call whose key differs — other arguments, another call, or new reasons (the adapters re-escalate).
audit ¶
The solvi audit of the verdict (what it rests on, the checks, the safeguards that fired).
replay ¶
Re-compute the decision's trace from its recorded inputs → solvi's replay report ({"ok", "mismatches", ...}).
Guard ¶
Guard(storage=None, authorizer=None, fact_names=None, lang='en', scan_user=False, tool_values='deny')
The catalog of tools an agent may call, the policies over their calls, and the store of every decision.
storage: a TraceStorage or a path (.db / .sqlite: SQLite, else JSON lines) — every decision is saved with its trace
and the outcome (meta "guard": tool, outcome, reasons, executed, the result's hash or the error).
authorizer: an optional decision part answering "does the conversation authorize this call?" (bool) over the facts
"conversation" (or "user_request") and "proposal" — see make_authorizer(); its act_guard threshold and perturb=k
apply. facts: names (or {name: type}) of facts your app gives with every call (a user's role, a budget left): policies
that read them apply to every tool without naming it (the types are for readers: a policy's own annotations are what
solvi validates). scan_user, tool_values: the defaults of every tool's scan_user and tool_values (see tool).
tool ¶
tool(func=None, *, name=None, schema=None, description=None, ground=(), ground_from=('user', 'tool', 'system'), injections='grounded', authorize=None, locale=None, scan_user=None, tool_values=None, ground_last=None, once=False)
Declare a tool the agent may call. As a decorator on a typed function (@guard.tool, @guard.tool(ground=[...])),
or guard.tool(name="refund", schema=RefundArgs) (a pydantic model or a JSON schema) for a tool the framework or
an MCP server runs. The function is returned unchanged.
ground: arguments that must be quoted from the conversation (a string as a whole word; numbers as number tokens;
a list item by item; an empty string never) — a list of names, or {name: matcher}: "token" (the default: not
inside a longer word, nor joined to one by ". @ - / : _"), "whole" (delimited by whitespace, quotes, brackets or
punctuation: for IBANs, e-mails, paths), "spaced" (as "token", and a number may group its thousands with spaces:
"1 250"), "nocase" (as "token", letters compared without their case, typographic dashes and quotes as plain ones: names,
addresses), "id" (as "nocase", and
a leading "#" of the value may be missing in the text: an order "#W5442520" the user wrote as "W5442520"), "url" (a web address: the same host, port, path, query and fragment as a URL written in the
conversation, with or without "http(s)://", a leading "www." or a trailing "/" — see same_url), "url_prefix"
(as "url", and the path may continue the written one at a "/": only for reading, a path can carry data out),
"substring" (anywhere), or a callable(value, text) → [(start, end)] of the value's occurrences. Every
Unicode space in the conversation or the value (a no-break space, a narrow one) is read as a plain space;
ground_from: the roles of the messages they may be quoted from (default: the user's, tool outputs and system
messages — never the assistant's own words; ("user",) for values only the user may give, like a payee).
injections: "grounded" (default: a grounded argument found only in tool outputs escalates when any tool
output in the conversation has instruction-like text), "any" (also: any instruction-like text in a tool output escalates the call — for high-impact
tools), "off". authorize: ask the guard's authorizer about this tool (default: when the guard has one).
locale: how a number written with one separator and one group of three digits reads — "1,500" / "1.500" is 1500
or 1.5 depending on the writer, so without a locale it grounds neither (deny); "en" (1,500.5), "de" (1.500,5),
"fr" (1 500,5 — with the "spaced" matcher), "ch" (1'500.5). A callable matcher decides per argument.
scan_user: a value the user wrote only next to instruction-like text in their own message (pasted content that
carries an instruction) escalates (no_injected_arguments); default: the guard's scan_user (False).
tool_values: what happens to an argument whose ground_from leaves out tool outputs (a user-only value) when
its value is not in the allowed messages but is in a tool output — "deny" (the default) or "escalate" (the
check arguments_from_user: a person decides, with the reason and the quote; never allowed on its own). A value
found nowhere is denied either way; default: the guard's tool_values.
ground_last: only the user's last N messages ground a value (None: all of them). In a long conversation a value
the user named many requests ago for another purpose otherwise grounds a call nobody asked for; with
ground_last=1 the call must rest on the current request (a call the user confirms with "yes, go ahead" then
finds nothing and is denied: the agent restates the value, or use a larger N).
once: a call of this tool with exactly the arguments of a call already made escalates (a second refund of the
same order, a file deleted twice). The calls made are the given fact calls_made — a Session, the MCP proxy
and the framework adapters keep it; with guard.check / guard.call pass facts={"calls_made": [...]} (strings
from solvi.agents.guard.proposal; [] when none was made). A call checked without it escalates: the check
cannot be evaluated.
declare ¶
Declare a tool that solvi does not run (the framework or an MCP server does): its name, the arguments' schema
(a pydantic model or a JSON schema; None — given later by adopt, as the MCP proxy does from tools/list) and
the options of tool (ground, ground_from, injections, authorize). → the Tool.
adopt ¶
Give a declared tool without a schema (guard.declare(name)) its arguments' JSON schema — the MCP proxy
does this from the server's tools/list. A tool that has a schema keeps it.
policy ¶
A policy over calls: an ordinary solvi hard check — argument names are the facts it reads (the call's validated arguments, the facts your app gives, conversation, user_request, tool_name), it returns True when the call may go ahead. on_fail: "deny" or "escalate". tools: a tool name or a list of them; None — every tool whose arguments and the guard's declared facts provide what it reads. Its docstring's first line is the reason given when it fails.
require_request ¶
A policy for actions that carry no value the user must give (book a hotel, create an event, read a URL a
document names): the call goes ahead only when the user's own messages (user_request, never tool outputs)
ask for this kind of action — intent, a key of solvi.agents.intents.INTENTS ("reserve", "event", "visit",
"pay", "send", "delete", "invite", "post", "share"; English and Russian word patterns) or a list of them, and /
or phrases, your own regular expressions. Otherwise the call escalates (on_fail="deny": is denied). It is an
ordinary policy named user_asked_to_<intent>: in the catalog, the trace and the reasons. It checks that the
user asked for such an action, not for this very call. → the policy function.
require_confirmation ¶
"The user confirmed this": a call of these tools goes ahead only when a message of the assistant proposed
its values and the user's next message explicitly accepted it (solvi.agents.confirm: "yes", "go ahead",
"please proceed", "да", "подтверждаю", ...; "yes, but ..." and "no" do not). The check user_confirmed
(deny; on_fail="escalate": a person decides) runs after grounding and once, before your policies; its reason
says what was missing, and an allowed call carries the accepted proposal and the acceptance as evidence. For
actions that must be the user's own decision: a value grounded in a tool output (their order, listed by a
lookup) passes grounding, while nothing in a tool output can write the user's yes. It costs turns, and it does
not judge the choice — a proposal the user accepts is allowed.
arguments: the arguments the proposal must name (default: every argument whose value is text, a number or a
list of them — a bool or an empty value is not named). match: {argument: matcher} — a ground= matcher name
("nocase", the default for text: case, Unicode spaces, typographic dashes and quotes aside; "id": an order
"#W1" written "W1"; "whole", "token", ...) or a callable(value, text) → [(start, end)] of where text (one
message) names value; a callable(value, text, facts) also gets the facts reads names (declared facts of
the guard: what an item id is called, from your app's state). A value the user wrote in the accepting message
itself counts too. last: only the user's last N messages count as acceptances (None: all of them).
tools: a tool name or a list. → None.
fn ¶
A computation the policies read (an ordinary solvi fn: amount_eur(amount, currency)), for the given tools
(None: every tool that provides its inputs).
make_authorizer ¶
A decider's yes / no question "does the conversation authorize this call?" over reads ("conversation": the
whole context, tool outputs included; "user_request": only the user's messages) and the proposed call as text,
with perturb=k (re-asked without instruction-like sentences; a changed answer escalates) — set as the guard's
authorizer and returned. Calibrate it with guard.calibrate_authorizer(examples, max_risk=0.10).
authorizer_input ¶
The input the authorizer reads for a call (decide.Facts) — for fit / act_guard / conformal examples.
calibrate_authorizer ¶
act_guard on the authorizer from labelled calls [(call, context, authorized: bool)]: P(allowed by the authorizer alone and wrong) ≤ max_risk for calls like these (solvi.decide.DecisionPart.act_guard). → its report.
system ¶
The solvi System that checks calls of one tool (built on first use; rebuilt after a declaration changes).
policies_of ¶
The policies that check calls of a tool → [(policy name, reason, on_fail)]: those that name it and those for every tool whose inputs it provides, deny ones first (the order they decide in); the reason is the policy's docstring's first line (its name when it has none) — what a refusal says. A tool with require_confirmation lists that check first ("user_confirmed").
definition ¶
A tool as a function-calling definition {"name", "description", "parameters"} — the same as
guard.tools[name].definition(). policies=True: the description also lists the reasons of the policies that
check the tool (policies_of), so the model can follow them before it is refused ("A guard checks this call:
it is refused unless ..."); a policy that escalates is marked so. Off by default: the reasons are your rules'
wording, written for refusals, and every listed line costs tokens in each request.
described ¶
description (a tool's description as a framework shows it) with the reasons of the policies that check the
tool appended, as definition(name, policies=True) writes them; unchanged when no policy checks it. The
adapters' show_policies=True use it.
check ¶
Decide on a proposed call without making it → GuardDecision (allow / deny / escalate, with the reasons and
the solvi response). context: the conversation (see messages); facts: what your app knows (a user's role,
a budget left) — given facts of the decision, recorded in its trace.
acheck
async
¶
check, on an event loop (policies may be async def: System.aask).
call ¶
Check a proposed call and, when it is allowed, make it: solvi runs the tool's registered function with the
validated arguments → GuardDecision with result (or error when the tool raised). A tool without a function
(declared for a framework) is not run: executed stays False (in a Session, report its result with
session.record).
acall
async
¶
call, on an event loop: an async def tool is awaited.
resolve ¶
A person's answer to an escalated call: recorded in the store as a correction of its verdict (who, the stored decision it answers, a note) and, when approved, the call is made (execute=False: not made — the framework makes it; the stored resolution then says executed: false, and the framework's result is not recorded by the guard). → the decision, with outcome "allow" (approved_by) or "deny".
An escalation is resolved once: a second resolve of the same decision (or, with a store, of a stored decision that already has a resolution) raises ValueError — so an approved call is never made twice.
replay ¶
Re-compute a stored decision from its recorded inputs with the tool's current checks → solvi's replay report ({"ok", "steps", "mismatches", "catalog", ...}): a recorded step that now computes differently (a changed policy, schema or authorizer) is a mismatch; "catalog" says whether the tool's checks changed since the decision.
replay_all ¶
replay every stored decision → [{"id", "tool", "mismatches"}] of those that do not replay.
session ¶
A conversation the guard follows: calls made through it add their results to its context as tool outputs, so a later call's grounding and injection checks see them; max_messages / max_chars cap the context kept (see Session). → Session.
Session ¶
A conversation with an agent: session.call(proposal) checks and makes calls in its context and appends each
made call's result as a tool output (add appends other messages).
made lists the calls that were made and did not fail — what a once=True tool reads. A call of a tool without
a function (the framework runs it) that session.call allows is counted as made when it is allowed, since it is
handed over to be made; session.record(decision, result) adds its result as a tool output, and with error= takes
it off the list again (a failed call may be tried again). After session.check nothing is counted until record.
max_messages / max_chars: the context kept for checking (None: all of it) — the oldest messages are dropped first,
and a message longer than max_chars keeps its beginning plus any instruction-like passage of the rest. A tool output
that carried instruction-like text before the cut is flagged as tainted (Message.tainted), so its taint survives
even if the cut kept none of it. Every decision's trace records the context it was checked against, so the cap also
bounds what each stored decision holds; a value or an instruction that has left the window no longer grounds a value
or taints a call.
record ¶
Report a call the framework made (a tool without a function, or a call allowed by check): its result is
appended as a tool output and the call counts as made (for once=True tools); with error (a text) the error is
appended instead and the call does not count — it may be tried again. Only an allowed decision is recorded.
→ the decision.
messages ¶
A conversation → [(role, text)]: a string (one user message); a list of {"role", "content"} dicts (OpenAI, Anthropic, MCP-style; content a string, a block or a list of blocks; {"type": "function_call_output", "output"} items are tool outputs), (role, text) pairs, or message objects with .type / .role and .content (LangChain). Roles are normalized to user, assistant, tool and system; anything unreadable is skipped.
What counts as the user's words is narrow, because a value the user gave is what ground_from=("user",) trusts:
- a message whose
typenames a tool output ("tool", "tool_result", "function_call_output", "function_response", in any letter case, with "-" or camelCase) is a tool output whatever its role; - a content block is read by its type, normalised the same way: a tool result (
tool_result, any*_tool_result,function_call_output,function_response,search_result, ...) is a tool output, a tool use (tool_use,function_call) the assistant's; - in a user message only text blocks are the user's (a string,
{"type": "text" | "input_text"}with a string "text", or a block with a string "text" and no type and no "content"); any other block (an image with a caption, a block with "content" and no type, a "text" that is not a string, an unknown type) is read as a tool output — never the user's words, and it gets the injection checks; - a message or a block that carries a
tool_call_id/tool_use_idanswers a tool call: a tool output; any item type ending in "call_output" (the Responses API's function / computer / shell / custom tool outputs) or "_tool_result" is one whatever its role; - a user message a framework generated in the user's place — LangChain's SummarizationMiddleware summary
(
additional_kwargs={"lc_source": "summarization"}), a "source" naming a summary or compaction — is the assistant's: a summary rewrites tool outputs into what looks like the user's turn, so it never grounds a user-only value. Other history compressions that rewrite turns as user messages cannot be recognised: give the guard the raw history.
Consecutive blocks of the same role make one message.
conversation ¶
A conversation → (text, [[start, end, role]], the user's messages as one text). A plain string is the text itself (one user message); messages are written one per line as "[role] text". A tool output a Session flagged as tainted (its instruction-like text was cut away with the rest of it) is [start, end, "tool", True].
arguments_model ¶
A function's parameters → a pydantic model of its arguments (unknown arguments forbidden). A first parameter typed as a framework context (RunContext, ToolContext, ...) is not an argument; args / *kwargs are not allowed.
url_parts ¶
A URL → (host, port, path, query, fragment, scheme) as the "url" matcher compares them, or None when it is not a plain web address. Read with urllib's parser: "http://" / "https://" or no scheme (read as a web address: "www.x.com/a"); any other scheme ("javascript:", "ftp://", "file:"), a protocol-relative "//x", userinfo ("good.com@evil.com", "user:pass@x"), a backslash, whitespace, control or format characters, a "." / ".." path segment (also percent-encoded) or a host that is not a valid DNS name or IP address → None. The host is lower case, IDNA-encoded (an internationalised name compares by its xn-- form), without a trailing dot and without one leading "www."; the default ports 80 and 443 are dropped; the path loses its trailing "/" (the root is ""); query and fragment are kept as written; the scheme is "http", "https" or None (not written).
same_url ¶
Does the URL value (a call's argument) name the address written (in the conversation)? Both read with
url_parts; the host must be equal (never a suffix or a prefix: "good.com" is not "evil.com/good.com",
"good.com.evil.com", "good.com@evil.com" or "xgood.com"), and so must the port, the query and the fragment.
The scheme never downgrades: a written "https://" matches only an "https://" value (not "http://", not a value
without a scheme); a written "http://" matches "http://", "https://" (an upgrade) and no scheme; a URL written
without a scheme matches either.
path="exact": the path is equal too (a trailing "/" aside); path="prefix": the value's path may continue the
written one at a "/" ("x.com/docs" covers "x.com/docs/a", not "x.com/docsevil") — only for reading: a path can carry
data out.
arguments_valid ¶
The arguments validate against the tool's types.
arguments_grounded ¶
Every argument that must come from the conversation is quoted there.
arguments_from_user ¶
Every argument that must come from the user is in the user's words (tool_values="escalate": one written only in a tool output needs a person).
no_injected_arguments ¶
No argument comes only from a tool output that carries instruction-like text.
no_instructions_in_tool_outputs ¶
No tool output in the conversation carries instruction-like text (solvi.perturb.injection_spans).
request_authorizes ¶
The authorizer (a decider) says the conversation authorizes this call.
proposal ¶
The call as the authorizer reads it: the tool's name and its arguments as JSON.
with_calls_made ¶
The facts of a call plus the calls an adapter made (the given fact calls_made a once=True tool reads), joined
with any calls_made the app's own facts give.
Did the user ask for this action? Policies for calls that carry no value the user must give.
Provenance protects an argument the user must give (a payee, a recipient). Some actions have none: "book a hotel" names the hotel from a search result, "create an event" takes a title and a time the agent chose, "read the link in that document" takes a URL from a tool output. Declaring such arguments as the user's denies every honest call; leaving them free lets a tool output that says "make a reservation for …" through. The middle road is a policy on the action itself: the call goes ahead only if the user's own messages ask for this kind of action, and otherwise escalates (or is denied).
guard.require_request("reserve_hotel", "reserve") # "book", "reserve", "reservation", "забронируй"
guard.require_request(["create_calendar_event"], "event") # "calendar", "meeting", "remind", "встреча"
guard.require_request("get_webpage", "visit", on_fail="deny") # "visit", "website", "link", a URL, "сайт"
guard.require_request("launch", phrases=[r"\blaunch\b"]) # your own patterns
It is an ordinary guard policy (a hard check in the tool's catalog, recorded in the trace, fingerprinted with its
patterns) named user_asked_to_<intent>, reading user_request — the user's messages only, never tool outputs. It
says that the user asked for an action of this kind, not that they asked for this very call: pair it with
injections="grounded" or "any" on the tool and with value policies (a date range, a price cap) where that matters.
A user who pastes a text that asks for the action counts as asking (scan_user=True for that case). The patterns are
word patterns over the NFKC-normalised, case-folded text, in English and Russian; INTENTS lists them.
request_policy ¶
A guard policy user_asked_to_<intent>(user_request) -> bool: the user's messages ask for this kind of action.
Did the user accept this very call? "Confirm before a consequential action" as a check of the guard.
A rule like "list the action's details and get the user's explicit yes before you change their order" is a relation
between three things: the call's arguments, a message of the assistant that proposed them, and the user's next message
that accepted it. guard.require_confirmation declares it for a tool:
guard.require_confirmation(["cancel_order", "refund"]) # every argument of the call
guard.require_confirmation("change_address", arguments=["street", "zip"]) # only these must be named
guard.require_confirmation("exchange_items", match={"item_ids": item_named}, reads=["known"]) # your matcher
A call of the tool is allowed by this check only when some message of the assistant names every required argument's value (a string by the "nocase" matcher — case, Unicode spaces, typographic dashes and quotes aside —, an order id with "id" if you say so, a number as a number token, a list item by item; or your matcher) and the user's next message (tool outputs in between are skipped) accepts it explicitly. A value the user wrote in the accepting message itself counts as proposed too ("yes, to my PayPal"). The proposal and the acceptance are quoted in the decision's evidence; when the check fails, its reason says what was missing — no accepted proposal, or which values the accepted one does not name.
What it is for: an action the user never asked for. An instruction planted in a tool output (an order note, a document, a web page) can talk the agent into a call whose values are all in the conversation — the user's own order, listed by a lookup — so grounding passes, and whose wording the injection detector does not know. This check still asks for the user's own yes to exactly these values. On τ-bench retail with such a note in every order lookup (12 tasks, one run each), the agent cancelled the order nobody asked about in 9 runs without a guard, 4 with the guard's other checks, 0 with this one; the customers it asked said no. It moves the decision to the user, it does not make it: in a development run a simulated customer said "yes, go ahead" to such a cancellation, and it was made.
What counts as an explicit acceptance (ACCEPT, WEAK, REFUSE_START, RETRACT, RESERVE, accepts): a yes word or phrase
in English or Russian ("yes", "go ahead", "please proceed", "confirmed", "that's correct", "that works", "да",
"подтверждаю", "оформляйте", ...) that no negation shortly before it turns around ("not correct", "don't proceed",
"не подтверждаю"), in a message that does not open with a refusal ("no", "wait", "нет") and takes nothing back
("instead", "changed my mind", "вместо", "передумал" anywhere; "actually", "wait", "hold on" at the start of a sentence
or a clause — not "the refund actually arrives", "I can't wait"); the first sentence that says yes decides, and a
reservation in that sentence makes it conditional, so not an acceptance ("yes, but not the blue one", "да, но ...").
A reservation in a later sentence is about something else ("Yes, please proceed. But could I also get a coupon?" is
an acceptance). A weak word — "ok", "sure", "fine", "alright", "хорошо", "ладно" — accepts only as the whole message,
with courtesy words at most and no question ("OK, thanks!"): "Okay, glad you found it. Which refund is faster?" is an
acknowledgement, not an acceptance. The text is read NFKC-normalised and case-folded, typographic quotes as plain.
The patterns are deliberately narrow: "no, go ahead with the other one" is not an acceptance; a user who accepts in
other words ("let's roll") is asked again. It reads the user's words only: it does not judge whether the proposal was
a good one (a wrong choice the user approves is approved), and it cannot tell a user from someone typing as them.
accepts ¶
Is a user's message an explicit acceptance of what the assistant proposed? (see the module docs)
accepted_proposals ¶
The assistant's messages the user accepted → [(proposal index, acceptance index)] into conversation_roles, the
latest first: each user message that accepts, paired with the assistant's message before it (tool and system
messages in between are skipped; a user message in between breaks the pair). last: only the user's last N messages
count as acceptances (None: all).
confirmation_fn ¶
The computed fact confirmation of a tool's checks: (call_arguments, conversation, conversation_roles, and the
facts spec["reads"] names) → {"confirmed", "proposal": [text, start, end] | None, "accepted": [...] | None,
"found": {argument: [[text, start, end, role]]}, "missing": [...], "why"}.
user_confirmed ¶
The user explicitly accepted a message of yours that names this call's values.
confirm_spec ¶
A tool's confirmation rule → (spec JSON, {argument: callable matcher}) — checked against the tool's arguments.
An MCP proxy with a solvi Guard: it sits between an agent (the MCP client) and an MCP server, and every tools/call passes the guard before it reaches the server.
solvi serve --guard catalog.py:guard --upstream "npx -y @modelcontextprotocol/server-filesystem /work" \
[--store calls.db] [--facts '{"role": "viewer"}'] [--escalate elicit|deny] \
[--context-messages 50] [--context-chars 100000]
catalog.py declares which of the server's tools the agent may call and the policies over them — without functions
(the server runs them) and usually without schemas (the proxy takes each tool's inputSchema from the server):
from pathlib import Path
guard = Guard(storage="calls.db")
guard.declare("read_text_file")
guard.declare("write_file", injections="any") # no writes after a tool output that carries instructions
@guard.policy(["read_text_file", "write_file"])
def inside_work(path: str) -> bool: # resolved: "/work/../etc/passwd" and symlinks out of /work fail
return Path(path).resolve().is_relative_to(Path("/work").resolve())
The proxy speaks MCP over stdio (JSON-RPC, one message per line) to the client and to the server it starts:
initialize initializes the server, answers with the tools capability (and the server's name in ours)
tools/list the server's tools that the guard declares (others are hidden; each declared tool without a schema
adopts the server's inputSchema — one that cannot be read gives that tool a permissive schema, a warning in
the log, and every call of it escalates; a tool whose arguments collide with the guard's facts is hidden)
tools/call the guard checks the call — allow: forwarded to the server with the arguments as the guard validated
them (coerced to the schema's types: "2" for an integer is sent as 2, "no" for a boolean as false; the
arguments the client did not send are not added), and its result's text is kept as a tool
output in the proxy's session, so later calls are checked against it (grounding, instruction-like text),
and a call the server made without an error counts as made (a repeat of a once=True tool escalates);
deny: an error result with the reasons; escalate: with --escalate elicit (default) and a client that
declares the elicitation capability, the user is asked (elicitation/create: approve yes / no) and the
answer is recorded as a person's resolution (only {"action": "accept", "content": {"approve": true}}
approves: "yes", 1 or "true" do not); otherwise an error result saying it waits for a person
ping answered
Every decision goes to the guard's store (or --store) with the call's outcome; _meta.solvi on each result carries the
outcome, the stored id and the trace hash. The proxy does not see the user's messages: an argument declared with ground=
is found only in the tool outputs of this session (and denied otherwise). The session keeps the last max_messages
tool outputs, at most max_chars characters in all (as Session names them): each decision's trace
records the context it was checked against, so the cap bounds what every stored decision holds — an output that has
left the window no longer grounds values or taints calls.
Taint is context-wide and the proxy grounds only from tool outputs: once one kept output carries instruction-like text
(an invoice that says "please pay within 30 days" is enough), every call with a ground= argument escalates. Declare
such tools with injections="off" and policies over their values, or run with a reviewer (--escalate elicit).
Upstream ¶
An MCP server started as a subprocess, spoken to over its stdin / stdout (one JSON-RPC message per line).
request ¶
Send a request and wait for its response → the result (UpstreamError on an error response or a closed server). Requests from the server meanwhile get "method not found"; notifications are ignored.
Proxy ¶
Proxy(guard, upstream, facts=None, escalate='elicit', max_messages=CONTEXT_MESSAGES, max_chars=CONTEXT_CHARS)
The proxy's state: the guard, the upstream server, the session (tool outputs seen so far, the facts).
run_proxy ¶
run_proxy(guard, upstream, facts=None, escalate='elicit', stdin=None, stdout=None, limits=None, max_messages=CONTEXT_MESSAGES, max_chars=CONTEXT_CHARS)
The proxy over stdio (see the module docs). upstream: a command line (or an Upstream). max_messages /
max_chars: the session's context kept for checking, as Session names them (None: unbounded; context_messages /
context_chars in 0.7). limits: a
solvi.serve.Limits — a client message is at most max_body characters and max_depth levels of JSON; a failure of the
proxy itself is logged (logger solvi.serve) and answered with an incident id, never the exception's text.