solvi.longdoc¶
Long texts: sections split at headings, BM25 retrieval, and the window a decider reads.
Long documents: find first, then decide. A text longer than the decider reads (its max_len) is cut into sections by code (headings, then paragraphs, then sentences), a cheap lexical scorer (BM25, standard library only) — and optionally the decider's own relevance — picks the few sections that bear on the question, the decider reads only those, and every offset it gives (a span answer, an evidence quote) is mapped back into the full document.
doc = LongDocument(contract, max_tokens=200)
doc.sections # [Section(start, end, heading, index)], in document order, covering the text
top = doc.select("How many days' notice does termination require?", k=3)
win = doc.window([s for s, _ in top]) # the selected sections joined, in document order
win.to_doc(12, 30) # an offset range in the window → the same text's range in the document
part = decider.decision("notice", "Notice period for termination?", "contract", Span[str], long="retrieve", top_k=3)
With long="retrieve" a decision part does this whenever its text does not fit; the sections it read (offsets, heading,
score) are recorded in the decision's extra["long"], so they are in the trace and hashed; a full replay (models re-run)
re-checks them, since the selection is deterministic — a trusted replay (trust_models=True) verifies the recorded output
and does not re-select. Texts that fit are decided as before, with nothing recorded.
Section
dataclass
¶
doc[start:end]; heading: the heading line of the part of the document it belongs to (None before any heading).
Window
dataclass
¶
Selected sections joined in document order: text, and pieces [(window start, document start, length)].
to_doc ¶
A range of the window → the same characters' range in the document, or None when it starts or ends in a separator, or runs over sections that are not neighbours in the document. A range over neighbouring sections maps to the document's text from its start to its end (the whitespace between them as the document has it, not the window's separator).
BM25 ¶
Okapi BM25 over a list of documents' terms (k1, b as usual); score(query_terms) → one score per document.
LongDocument ¶
A text as sections that each fit max_tokens (counted by count, default approx_tokens): split at headings (a
Markdown "#", "1.", "1.2", "Article 5", "Section 3", "§ 4", a Roman numeral, an ALL-CAPS line), then at blank lines,
then at sentence ends, then at spaces; the sections cover the text in order (only whitespace between them).
select ¶
The sections that bear on query, best first → [(Section, score)]: the top k by BM25 that fit budget tokens
together (ties: the earlier section). rerank: a function [section text] → [relevance] (e.g. the decider's p(yes),
see DecisionPart) applied to the best rerank_top (default 3k) by BM25; its score orders them instead (BM25 breaks
ties). When fewer than k sections contain a query term, the rest of the k are the document's first sections
(score 0, in document order) — so k sections are read whenever the document has them, and with no match at all
they are the first k.
window ¶
The sections joined in document order → Window (offsets map back with to_doc).
approx_tokens ¶
A tokenizer-free estimate of a subword model's tokens: words and punctuation marks × 1.3 (rounded up).
terms ¶
Lowercase word terms without stop words, a plural "s" stripped (termination / terminations match).