solvi.extract_long¶
@extract for long, general-purpose documents: a field defined by its description (LongSpanExtractor; needs torch).
@extract for real-world, general-purpose documents (the extractor protocol it shares with solvi.extract_multi.MultiSpanExtractor: fit(items), predict(text, description), field(name, description), save / load, fingerprint()): a field is defined by its DESCRIPTION, the document may be long (windows), and the field may be absent ("no answer"). ModernBERT: input "field description [SEP] document window", two pointer heads; "no answer" is position 0 (the special token). Prediction: the best span across all windows; an answer exists if its score is above the field's threshold (tuned on held-out examples). Also works for fields unseen in training, from the description alone ("new field").
LongSpanExtractor ¶
LongSpanExtractor(model_name='answerdotai/ModernBERT-large', max_len=1024, stride=128, max_span=96, device=None)
stride: overlap between adjacent windows in tokens (as in transformers); max_span: maximum answer length in tokens.
make_examples ¶
items: [(text, description, (start, end) | None[, neg])] → training windows: every window with the answer + up to neg_per_item (or the item's own neg) windows without it.
predict ¶
→ (start, end, span score, "no answer" score) — the best span across all windows.
tune_threshold ¶
tune_threshold(name, items, grid=(0.0003, 0.001, 0.003, 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8))
Per-field "answer present" threshold from held-out examples: maximizes present/absent decision accuracy.
tune_default_threshold ¶
tune_default_threshold(items, grid=(0.0003, 0.001, 0.003, 0.01, 0.02, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5))
One threshold from held-out examples of many fields — used for fields that have no examples of their own.
save ¶
Encoder weights (bf16 safetensors), tokenizer, span head and settings into a directory.
load
classmethod
¶
From a directory written by save(), or a Hugging Face model id (downloaded once and cached).
embed ¶
Document embedding: the encoder's token states averaged over every window of the text (no field description).
embedder ¶
A catalog part: doc → embedding vector, usable as a feature of System.fit.
fingerprint ¶
A stable hash of this extractor: settings, thresholds, the span head, sampled encoder weights and the weight files it was loaded from (recorded in the trace; replay compares it with the catalog's current model).