Skip to content

Text in: from a message to a question

system.ask(state) needs a typed state. A person writes a message instead: "please refund order A-10457, I paid 1.5 million rubles on 12 September". solvi.textin turns such a text into the question it asks and that question's input state, reads every value with a quote, and leaves the decision to the catalog as before.

from solvi.textin import TextIn

eps = system.entry_points()          # the questions with the typed input state each one reads
eps[0].fields["amount"]              # EntryField(name="amount", type=float, description=..., required=True)
eps[0].tool()                        # the same as a function-calling tool: {"type": "function", "function": {...}}

tin = TextIn(system, decider, today=date(2026, 9, 28),            # fields by the cue finder (the default)
             synonyms={"currency": {"RUB": ["rubles", "руб", "₽"], "EUR": ["euro", "€"]}},
             patterns={"order_id": r"[A-Z]-\d+"})
read = tin.read("Please refund order A-10457: I paid 1.5 million rubles on 12 September.")
read.question                        # "request_refund"
read.state                           # {"order_id": "A-10457", "amount": 1500000.0, "currency": "RUB",
                                     #  "purchase_date": date(2026, 9, 12)}
read.fields["amount"].quote          # Quote("1.5 million", 36, 47, "request_text", ...)
read.missing, read.clarify()         # required fields the text does not give, and a question asking for them

res = system.ask_text(read)          # or system.ask_text(text, decider) / ask_text(text, textin=tin); aask_text is async
res["request_refund"].answer

Entry points. Every question is an entry point (or the names you pass: TextIn(..., entry_points=[...])); its input fields are the given facts its flow reads, with their types (System(input_model=...), else the types the catalog's parts declare) and whether the question needs them — the same schemas solvi serve publishes at GET /questions.

Who does what. The decider picks the entry point: one choice question over the entry points, each described by its question text (or descriptions={name: text}). Below min_confidence (0.6), on a near tie (min_margin 0.1), or when the decider's act signal escalates, nothing is chosen: read.question is None, system.ask_text runs nothing and the likely questions abstain with guard escalated, and read.clarify() asks which one is meant. The extractor points at the text of each field: by default CueExtractor — a deterministic finder of candidates of the field's type (numbers, dates, enum labels and synonyms, cue words, a pattern) nearest after a cue word (the field's name, plus cues={field: [...]}; its description's words rank candidates too), whatever the decider. The decider's own span pointer reads the fields only when you name it, extractor=DeciderExtractor(decider); any object with find(text, FieldSpec) → [Quote] works, and a list of extractors is tried in order (the trace records which one read each field). Code does the rest: a deterministic parser per type turns the quote into the value.

The default was chosen by measurement (benchmarks/textin_extractors.py, solvi-base in ONNX, one process): every text in this repository that carries typed fields — the shop requests of this section (18, English and Russian), the e-mails of examples/04_refunds.py (40), the invoices of examples/03_invoices.py (40), the tickets of gallery/11_refund_double_charge (16) and the claims of examples/16_primitives.py (3) — read field by field with the question given, against the values the repository's own hand-written code reads (306 stated values, 9 fields the text does not state):

extractor right wrong missed
extractor right wrong missed
--- --- --- ---
CueExtractor 282 1 23
solvi-base's span pointer (DeciderExtractor) 111 11 184
the pointer, then the cue finder 271 12 23
the cue finder, then the pointer 282 1 23

The last three columns are a set written for the benchmark before the cue finder's reading of strings was last changed (30 texts, 57 stated values: order ids, addresses, names, vendors and invoice numbers with no pattern, in "key: value" lists and in sentences; "wrong" counts a value read where the text states none). The pointer answers "not stated" or a confidence below min_field_confidence for most fields it is asked about (the amount and the currency of "please refund order A-10457, 1.5 million rubles, paid 12 September"); a field one extractor reads below that confidence, or not at all, is passed to the next one in the list. A string without a pattern is read after one of its own cue words (the field's name, cues= — never its description's words): what follows a connector ("address: …", "address is …") up to the end of the clause, cut before the next "key:" of a list, another field's cue word, a new clause ("and my …", ", please …") or after an identifier followed by a comma; or an identifier right after the cue ("order A-10457"). A field named as an identifier (…_id, …_number, …_code, …_ref) takes one token with a digit, or nothing. Before this, "order: A-10457, amount: 1" read the order id as "A-10457, amount: 1" and the set scored 20 right, 17 wrong, 20 missed (the cue finder then the pointer: 27, 19, 11). It is still a guess — "The vendor will be confirmed later" reads the vendor as "confirmed later" — so give an identifier its pattern (patterns= or the field's json_schema_extra={"pattern": ...}) and an enum its synonyms. Routing has no "none of these" option: a text that asks none of the questions is escalated only when the decider is unsure (min_confidence, min_margin), so a confident wrong route is possible — add an entry point for "something else" if your texts can be about anything.

Type Reads
int, float, Decimal 1500, 1,500.50, 1 500 000 руб, 12,5, 2k, 5m, $5 m, 1.5 million, 3 млн, a million, half a million, two and a half million, полтора миллиона (an int must be whole). Not guessed, so unparsed: a fraction the parser does not compute (quarter of a million, three quarters of a million, 5 and a half thousand — never read as the number next to it), 5 m / 2 b (a one-letter scale apart from the number may be a unit), 1.000 (a thousand or one? TextIn(decimal="," or ".") says), 3 100 (digits grouped by plain spaces with no currency next to them may be two numbers), 5% (unless the field is declared in percent: TextIn(percent=[field]) or json_schema_extra={"percent": True})
date 2026-09-12, 12.09.2026, 12/09/26 (dayfirst=False: month first; a two-digit year only with today=, within 80 years back and 20 ahead), 12 September 2026, September 12, 12 сентября; today / yesterday / tomorrow. A lower-case may after a number, without a year and before a verb or a pronoun ("these 2 may be wrong"), is the modal verb, not a date
Literal[...], an Enum the label (or member name), or a synonym: synonyms={field: {label: [...]}} or the field's json_schema_extra={"synonyms": ...}
bool yes / no words; the field's name or a cues= word ("urgent") → True; a phrase declared in negatives={field: [...]} (or json_schema_extra={"negative_cues": ...}) → False. Description words only rank candidates. A cue answered by a yes / no word ("Urgent: no", "urgent = false", "Is it urgent? No.") is that answer. A cue with a negation near it, before or after it in the sentence ("isn't urgent", "far from urgent", "anything but urgent", "urgent? not at all", "was urgent yesterday, not anymore", "urgent but cancelling isn't", "не срочно") is unparsed — never True, and False only through a declared negative
str the quote, trimmed; patterns={field: regex} must match it whole

A date without a year is not guessed. Without TextIn(today=...) it is not read: the field is unparsed with the reason "the year is not stated", a required one is in read.missing, and read.clarify() asks "Please tell me the purchase date (I read '12 September' but the year is not stated)." The same holds for a relative date and a two-digit year. With today= you take the assumption on: a date without a year is given today's year, recorded in the trace with today — wrong around the turn of a year ("paid 28 December" read on 5 January becomes 28 December of the new year, almost a year ahead). Where a rule compares such a date with today (a refund window), add a check that the date is not in the future, or leave today out and ask for the year. solvi serve and solvi ask --text pass a today only when the request ("today") or the command line (--today) gives one. Every field ends in one state: read, not_stated, unparsed (the quote does not parse), unsure (found with confidence below min_field_confidence, 0.5) or unsupported (no parser for the type). A required field that is not read is in read.missing: the question is asked anyway (a hard check may already decide it), and without that field it abstains — "not stated in the text: purchase_date; cannot compute: ..." — instead of guessing.

Provenance. The text itself is a given fact (init_state["request_text"]); the values read from it are not. The trace of ask_text holds, after the flow's steps, one record for the entry point (kind textin, provenance decided, the decider's identity and probabilities) and one per field (textin:<field>, provenance quoted, the quote's offsets, the parser and its arguments, the extractor's identity and fingerprint). The audit lists those fields under quoted with the model, counts them as "quoted by model" and the entry point as "decided" — not in the deterministic share — and an answer's confidence is at most the entry point's and the read fields' confidences. ask_text does not trust a TextRead it is handed: each field is re-derived from its quote (the quote at its offsets, the parser of the field's type, the typed value) with the field's own parser arguments — rebuilt from the entry point's field by textin= (else the TextIn that made the read, else a default TextIn(system)), so a read that brings its own cues ({"cues": ["banana"]}), labels or pattern does not re-derive; only a date's today may come from the read. A field that does not re-derive is unparsed — a required one is missing and the question abstains. Replay re-checks each record: the quote is literally in the text at its offsets, the recorded parser gives the recorded canonical form and the typed value rebuilt from it, and the flow read exactly that value. Even CueExtractor, which is plain code, is recorded this way: which number is "the amount" is still a guess.

A dialogue. tin.update(read, next_message) reads the next turn over the whole dialogue (turns joined by a new line; every quote points into it) and lists changes — field, old value, new value, quote. A turn that names the old value next to a new one ("the order is not A-10457 but A-10475") changes it to the new one; fields the turn does not state keep their value and quote; a field the turn restates in a form that does not parse becomes a conflict (in missing, asked by clarify()), and its old value is not kept as if confirmed; the entry point stays the one chosen (an escalated read is routed again on the whole dialogue). tin.update({"order_id": "A-1"}, text, question=...) starts from a state you already have: those fields stay given. system.ask_text(updated) answers on the whole dialogue, in one trace.

The call is data. A text can only select one of the entry points and fill typed fields through the parsers: nothing in it is executed, and the functions that run are the catalog's, planned by the strategist as for any ask.