Skip to content

solvi.longdoc

Long texts: sections split at headings, BM25 retrieval, and the window a decider reads.

Long documents: find first, then decide. A text longer than the decider reads (its max_len) is cut into sections by code (headings, then paragraphs, then sentences), a cheap lexical scorer (BM25, standard library only) — and optionally the decider's own relevance — picks the few sections that bear on the question, the decider reads only those, and every offset it gives (a span answer, an evidence quote) is mapped back into the full document.

doc = LongDocument(contract, max_tokens=200)
doc.sections                          # [Section(start, end, heading, index)], in document order, covering the text
top = doc.select("How many days' notice does termination require?", k=3)
win = doc.window([s for s, _ in top])  # the selected sections joined, in document order
win.to_doc(12, 30)                     # an offset range in the window → the same text's range in the document

part = decider.decision("notice", "Notice period for termination?", "contract", Span[str], long="retrieve", top_k=3)

With long="retrieve" a decision part does this whenever its text does not fit; the sections it read (offsets, heading, score) are recorded in the decision's extra["long"], so they are in the trace and hashed; a full replay (models re-run) re-checks them, since the selection is deterministic — a trusted replay (trust_models=True) verifies the recorded output and does not re-select. Texts that fit are decided as before, with nothing recorded.

Section dataclass

Section(start: int, end: int, heading: str | None = None, index: int = 0)

doc[start:end]; heading: the heading line of the part of the document it belongs to (None before any heading).

Window dataclass

Window(text: str, pieces: list = list(), sections: list = list())

Selected sections joined in document order: text, and pieces [(window start, document start, length)].

to_doc

to_doc(start, end)

A range of the window → the same characters' range in the document, or None when it starts or ends in a separator, or runs over sections that are not neighbours in the document. A range over neighbouring sections maps to the document's text from its start to its end (the whitespace between them as the document has it, not the window's separator).

BM25

BM25(docs, k1=1.5, b=0.75)

Okapi BM25 over a list of documents' terms (k1, b as usual); score(query_terms) → one score per document.

LongDocument

LongDocument(text, max_tokens=256, count=None)

A text as sections that each fit max_tokens (counted by count, default approx_tokens): split at headings (a Markdown "#", "1.", "1.2", "Article 5", "Section 3", "§ 4", a Roman numeral, an ALL-CAPS line), then at blank lines, then at sentence ends, then at spaces; the sections cover the text in order (only whitespace between them).

scores

scores(query)

BM25 of each section (its heading counts as part of it) for the query.

select

select(query, k=3, budget=None, rerank=None, rerank_top=None)

The sections that bear on query, best first → [(Section, score)]: the top k by BM25 that fit budget tokens together (ties: the earlier section). rerank: a function [section text] → [relevance] (e.g. the decider's p(yes), see DecisionPart) applied to the best rerank_top (default 3k) by BM25; its score orders them instead (BM25 breaks ties). When fewer than k sections contain a query term, the rest of the k are the document's first sections (score 0, in document order) — so k sections are read whenever the document has them, and with no match at all they are the first k.

window

window(sections, sep='\n\n')

The sections joined in document order → Window (offsets map back with to_doc).

approx_tokens

approx_tokens(text)

A tokenizer-free estimate of a subword model's tokens: words and punctuation marks × 1.3 (rounded up).

terms

terms(text)

Lowercase word terms without stop words, a plural "s" stripped (termination / terminations match).