solvi.textin¶
Text in: a free text → the question it asks and its typed input state, every value read with a quote; deterministic parsers, dialogue updates, replay of the reading.
Text in: a free text → which question is asked (an entry point) and its typed input state, every value read with a quote.
tin = TextIn(system, decider)
read = tin.read("Please refund 1.5 million RUB for order A-10457, bought on 12 September 2026")
read.question # "request_refund" — the decider's choice among the entry points (escalates when unsure)
read.state # {"order_id": "A-10457", "amount": 1500000.0, "currency": "RUB", "purchase_date": date(2026, 9, 12)}
read.fields["amount"].quote # Quote("1.5 million", 21, 32, "request_text") — where the value was read
read.missing # required fields the text does not state → ask for them (read.clarify()), never guessed
res = system.ask_text(read) # the question answered from that state, all in one trace
Entry points are the system's questions with the typed input state each one reads (system.entry_points(), from the same
schemas as solvi serve). The model does two small things: the decider picks the entry point (a choice over the entry
points with their descriptions; below min_confidence or a near tie it escalates, and nothing is asked), and an extractor
points at the piece of text that holds each field (the deterministic CueExtractor by default; the decider's own span
pointer, DeciderExtractor, when you name it). Code does the rest: a deterministic parser per type turns the quote into the value —
numbers ("1,500", "1.5 million", "2k", "полтора миллиона"), dates ("2026-09-12", "12.09.2026", "12 September", "12 сентября";
without a year only with today= — else the field is unread: "the year is not stated"), enums by label or synonym, booleans, strings — and a quote that does not parse leaves
the field unread. A field the text does not state is "not stated"; a required one is listed in missing for a clarifying
question instead of a guess.
Provenance: a field read from the text is quoted by a model (the extractor's identity and fingerprint are recorded), never
given, so the audit counts it among the model outputs ("quoted by model"), not in the deterministic share; the choice
of entry point is decided by the decider. The records (kind "textin") are hash-chained into the answer's trace, and replay
re-checks each one: the quote is literally in the text at its offsets, the parser gives the recorded value from it (the
canonical form and the typed value rebuilt from it), and the value is the one the flow read. System.ask_text re-derives
each field from its quote before asking (rederive): a caller-built TextRead cannot carry a value its quote does not
state. Nothing is executed from the text: the "call" is data — a question name from the closed
set of entry points and typed values — that solvi checks and then asks itself.
EntryField
dataclass
¶
One input field of an entry point: its type, description and whether the question needs it.
EntryPoint
dataclass
¶
A question as an entry point: its name, text and the typed input state it reads (schema: the JSON schema).
tool ¶
As a function-calling tool: {"type": "function", "function": {"name", "description", "parameters"}}.
ParseError ¶
Bases: ValueError
A quote that does not parse as the field's type.
FieldSpec
dataclass
¶
FieldSpec(name: str, type: Any, kind: str, description: str | None = None, required: bool = False, cues: list = list(), labels: dict | None = None, values: dict | None = None, pattern: str | None = None, hints: list = list(), percent: bool = False, negatives: list = list(), stops: list = list())
What the extractor and the parser know of a field: kind (number, integer, date, bool, enum, text, or unsupported), cue words, enum labels with synonyms, a pattern.
CueExtractor ¶
A deterministic extractor: candidates of the field's type in the text (numbers, dates, enum labels and synonyms,
cue words for a yes / no, a pattern), the one nearest after a cue word of the field (its name, its description's
words, cues=) first. A string without a pattern is read only after one of the field's own cue words (its name,
cues=; never its description's words): the words after a connector ("address: ...", "address is ...") up to the end
of the clause, cut before the next "key:" of a list, another field's cue word after a separator, a new clause ("and
my ...", ", please ..."), or after an identifier followed by a separator; or an identifier right after the cue
("order A-10457"). A field named as an identifier (…_id, …_number, …_code, …_ref) takes one token with a digit in
it, or nothing. It never calls a model, but the choice of the span is still a guess, so a value it reads is recorded
as quoted by it (its identity in the trace) and counted with model outputs.
find ¶
→ [Quote] candidates, best first (confidence 1.0 near a cue, lower without one); [] when not stated.
DeciderExtractor ¶
The decider's own span pointer: for each field a span question ("What is the
FieldRead
dataclass
¶
FieldRead(name: str, status: str, value: Any = None, quote: Quote | None = None, confidence: float = 0.0, canonical: Any = None, parser: str | None = None, spec: dict | None = None, why: str | None = None, required: bool = False, model: dict | None = None, was: Any = None)
One field as read from the text. status: read | not_stated | unparsed (a quote that does not parse as the type) |
unsure (below min_field_confidence) | unsupported (a type no parser reads) | given (passed by the caller to update) |
conflict (update: a turn restates a field already read but the new quote does not parse — the old value is in was,
not in the state).
Change
dataclass
¶
A field a dialogue turn changed (TextIn.update): old and new value, and the quote of the new one.
TextRead
dataclass
¶
TextRead(text: str, question: str | None, fields: dict, route: dict, source: str = SOURCE, changes: list = list(), entry_points: list = list(), router: dict | None = None, reader: Any = None)
A text read as a call: the question (None when the entry point escalated), each field of its input state, the
route (probabilities over the entry points), missing required fields, and changes after TextIn.update.
missing
property
¶
Required fields the text does not give (not stated, unparsed, unsure, unsupported), and any field in conflict (a dialogue turn restated it in a form that does not read).
init_state ¶
The init_state System.ask reads: the text itself (under source) and the fields read from it.
clarify ¶
A clarifying question for what is missing (None when nothing is): the entry point when it escalated, else the required fields not read.
to_dict ¶
The read as JSON-ready data, values written as Response.to_dict() writes them (a date as its ISO string, a Decimal as its text, an Enum as its value).
records ¶
The trace records of this read (kind "textin"): the entry point, then one per field — fresh Records, chained by the System when it appends them.
TextIn ¶
TextIn(system, decider=None, extractor=None, *, entry_points=None, descriptions=None, synonyms=None, cues=None, patterns=None, negatives=None, percent=(), today=None, dayfirst=True, decimal=None, min_confidence=0.6, min_margin=0.1, min_field_confidence=0.5, task='Which request is this text making?', source=SOURCE)
Free text → (question, init_state) over a System's entry points.
decider: picks the entry point (a DecideModel, or any object with decide(text, task, options, descriptions=, kind="choice") → Decision); not needed with one entry point or when the question is given. extractor: finds each field's span — an object with find(text, FieldSpec) → [Quote] (best first), or a list of them tried in order; default: CueExtractor, whatever the decider (measured: benchmarks/textin_extractors.py — on the repository's texts with typed fields the cue finder read 282 of 306 stated values right and 1 wrong, solvi-base's span pointer 111 right and 11 wrong; the pointer is used only when named, DeciderExtractor(decider)). In a list, a field the first extractor does not read — nothing found, nothing that parses, or found below min_field_confidence ("unsure") — goes to the next.
entry_points: the question names to choose from (default: every question). descriptions: {question: text} for the router (default: the question's text). synonyms: {field: {label: [synonym]}} for enum fields; cues: {field: [word]} extra cue words; patterns: {field: regex} for string fields; negatives: {field: [phrase]} for yes / no fields — the phrases that mean False ("not urgent", "no rush"; without one a negated cue does not parse) (a System(input_model=...) field's json_schema_extra may carry "synonyms", "cues", "negative_cues", "pattern", "percent" too). percent: the number fields in percent ("5%" → 5; elsewhere a percentage does not parse). decimal: "." or "," — the decimal separator of the texts (default None: "1,000" is a thousand and "1.000" is ambiguous, so it does not parse). today: a date for year-less, two-digit-year and relative dates (recorded in the trace). dayfirst: 12/09 is 12 September. min_confidence / min_margin: the router escalates below this probability or when the two best entry points are closer than the margin. min_field_confidence: a span found with less confidence is "unsure" (not used). task: the routing question the decider reads. source: the init_state key the text is given under.
route ¶
→ {"question", "probs", "confidence", "candidates", "escalated", "by"}: the decider's choice among the entry points, escalated (question None) below min_confidence, on a near tie or when the decider escalates.
read ¶
A text → TextRead: the entry point (routed, or question), and each input field read with its quote.
update ¶
A dialogue turn: prev (a TextRead, or a state dict with question=) and the next message → a new TextRead over
the whole dialogue (the turns joined by a new line; quotes point into it) whose changes list the fields the turn
states anew — old value, new value, quote. Fields the turn does not state keep their value and quote; a turn that
repeats the old value next to a new one ("not A-10457 but A-10475") changes it to the new one; a turn that restates
a field in a form that does not parse makes it a conflict (in missing, asked by clarify(); the old value is
not kept as if confirmed). The entry point stays the one already chosen (an escalated read is routed again on the whole dialogue).
entry_points ¶
The questions of a system as entry points (see System.entry_points).
parse_number ¶
A quote → the number it states, as a canonical string ("1500000", "12.5"): digits with thousands separators, a decimal point or comma, a scale word ("1.5 million", "2k", "3 млн") or a number word with one ("a million", "полтора миллиона", "half a million", "two and a half million"); a currency sign or code around it is allowed. Exactly one number: two numbers in the quote are an error. A fraction the parser does not compute — "quarter of a million", "three quarters of a million", "5 and a half thousand" — is refused, never read as the number next to it. Refused as ambiguous rather than guessed: a one-letter scale apart from the number ("5 m" — metres? "5m" and "$5 m" are read), a single ".ddd" group ("1.000"; spec {"decimal": "." | ","} decides), digits grouped by plain spaces without a currency next to them ("3 100" may be two numbers; "1 500 000 руб" is read), a percentage ("5%"; spec {"percent": True}: the field is in percent, 5% → 5), and a two-digit year is the date parser's. spec {"integer": True}: the number must be whole.
parse_date ¶
A quote → the date it states, ISO ("2026-09-12"): 2026-09-12; 12.09.2026 / 12/09/26 (day first; spec {"dayfirst": False}: month first); 12 September 2026, September 12, 2026, 12 Sep, 12 сентября; today / yesterday / tomorrow. A lower-case "may" after a number, without a year and before a verb or a pronoun ("these 2 may be wrong"), is the modal verb, not the month. A date without a year, a two-digit year, or a relative date needs spec {"today": "YYYY-MM-DD"} (TextIn(today=...)): without it it is an error — for a date without a year "the year is not stated", never a year guessed. With it a date without a year is read in today's year (an assumption the caller made by passing today, not a reading: "28 December" read on 5 January is the December ahead). A two-digit year is the one within (today − 80 years, today + 20 years]: with today 2026-09-28, "85" is 1985 and "30" is 2030. Exactly one date in the quote.
parse_bool ¶
A quote → True / False: yes / no words (en, ru); a declared negative cue (spec {"negatives": [...]}: "not urgent",
"no rush") → False; a cue answered by a yes / no word ("urgent: no", "urgent = false", "is it urgent? no" → False;
"urgent: yes" → True); the quote is exactly one of the field's cues (spec {"cues": [...]}: its name and the cues=
words) → True. A cue with a negation near it, before or after ("isn't urgent", "far from urgent", "anything but
urgent", "urgent? not at all", "was urgent yesterday, not anymore", "urgent but cancelling isn't", "не срочно") is
neither: it does not parse — never True, and False only through a declared negative cue.
parse_enum ¶
A quote → one of the labels (spec {"labels": {label: [synonym, ...]}}): the quote is a label or a synonym (case, spaces, underscores and punctuation ignored), or contains exactly one of them as a whole word.
field_spec ¶
An EntryField → FieldSpec: the kind from the type, cue words from the name (plus cues), hint words from the
description (they rank candidates, never decide), enum labels from Literal values / Enum members (plus synonyms
{label: [...]}, keyed by the value or the member name). A bool field's value cues are only its name ("urgent", "is
urgent" → "urgent") and cues; negatives (or json_schema_extra "negative_cues") are phrases that mean False.
percent (or json_schema_extra "percent"): a number field in percent, so "5%" reads as 5.
rederive ¶
The fields of a TextRead re-derived from their quotes before System.ask_text trusts them: each field read must
quote the text at its offsets, its parser must be the one the field's declared type gives, and its parser arguments
(the spec: a yes / no field's cues and negatives, an enum's labels, a pattern, the decimal separator, ...) must be the
ones the field's own spec gives — rebuilt from the entry point's field by textin (else the TextIn that made the read,
else TextIn(system)); only a date's today is taken from the read. With them, the quote must parse to the recorded
canonical form and typed value. A field that does not re-derive becomes unparsed (so a required one is missing and
the question abstains or escalates); a field the entry point does not read is dropped. → the read itself when
everything holds, else a copy (the caller's TextRead is not changed).
replay_record ¶
Re-check a "textin" trace record against the recorded input → [(step, name, reason)]: the entry point is one of the recorded entry points; a field's quote is literally in the text at its offsets, the recorded parser gives the recorded value from it, and the flow read that value (the given fact equals it); a field not read was not given either.