Extracting fields from documents¶
With pip install "solvi[model]", solvi provides three ModernBERT extractors. All of them predict a start and an end
position in the text, so the extracted value is always a substring of the document with exact offsets. Each gives you
plain functions doc -> Quote to register with cat.extract.
Labels are character spans: for each training document and field, (start, end) of the value in the text, or None if
the field is absent. extract-base's model card suggests labelling about 25–100 documents per task and fine-tuning. A
GPU is recommended for training and for fast inference.
MultiSpanExtractor: all fields in one pass¶
solvi.extract_multi.MultiSpanExtractor reads a document once and has a start/end head pair per field. Use it for
documents that fit into one window (receipts, invoices, forms).
from solvi.extract_multi import MultiSpanExtractor
ex = MultiSpanExtractor(["company", "date", "total"],
model_name="answerdotai/ModernBERT-large", max_len=1024)
ex.fit(train, epochs=4, lr=3e-5, bs=8)
# train: [(text, {"company": (s, e), "date": (s, e), "total": (s, e) or None})]
ex.save("receipts-extractor") # MultiSpanExtractor.load("receipts-extractor") reads it back
ex.fit_temperature(calib_docs, lambda field, i, span: span == calib_spans[i][field]) # optional, per-field temperature
for name in ["company", "date", "total"]:
cat.extract(ex.field(name)) # registers the fact "company", ... read from init_state["doc"]
@cat.fn
def total_value(total):
return float(total.replace(",", ""))
ex.field(name)returns a function namednamewith one argumentdoc, which returnsQuote(doc[s:e], s, e, confidence=c).ex.predict(text)returns{field: (start, end, confidence)}(ex.predict(text, field)one of them). Results are cached per text, so all fields of one document cost a single forward pass. The two extractors share one protocol —fit(items),predict(text, field),field(name[, description]),save/load,fingerprint(); 0.7'sfit(docs, spans)andpredict_doc(text)still work with a deprecation warning. A labelled span past the encoded text (max_lentokens) is left out of training.fit_temperature(docs, gold_ok)picks a softmax temperature per field that minimizes log loss of "confidence vs. correct" on held-out documents;gold_ok(field, doc_index, (start, end))tells whether a prediction is correct.- The extractor always returns a span; this one does not model "field absent". Use
LongSpanExtractorwhen fields may be missing.
LongSpanExtractor: long documents, fields by description, "no answer"¶
solvi.extract_long.LongSpanExtractor is for long documents (contracts) and optional fields. The input is
[CLS] field description [SEP] window of the document [SEP], windows overlap, and position 0 means "no answer in this
window".
from solvi.extract_long import LongSpanExtractor
GOV_LAW = "the clause that says which state's or country's law governs the contract"
lx = LongSpanExtractor(max_len=1024, stride=128, max_span=96)
lx.fit([(text, GOV_LAW, span_or_none) for text, span_or_none in train], epochs=3, neg_per_item=3)
lx.tune_threshold("governing_law", [(text, GOV_LAW, span_or_none) for text, span_or_none in heldout])
cat.extract(lx.field("governing_law", GOV_LAW))
@cat.fn
def governing_state(governing_law):
for s in ["Delaware", "New York", "California"]:
if s.lower() in governing_law.lower():
return s
return "other"
strideis the overlap between windows in tokens;max_spancaps answer length in tokens.- Training uses every window that contains the answer plus up to
neg_per_itemwindows without it. One extractor can learn several fields: pass items with different descriptions. predict(text, desc)returns(start, end, span_score, no_answer_score), the best span over all windows.tune_threshold(name, items)picks the score threshold for "the field is present" that maximizes present/absent accuracy on held-out items.field(name, desc)returns a functiondoc -> Quote. When the score is below the threshold, it returnsQuote("", 0, 0), an empty value, so downstream functions should treat""as "not found" (e.g.has_tax = tax != "").- Cost grows with length: one pass per field per window, so a long contract with several fields takes many passes; use a GPU for long documents.
- Fields are specified by description, but a field that was never labeled in training is not extracted reliably from its description alone (extract-base's model card: 14% and 66% on two held-out fields for a model trained on other fields only). Label examples for every field you need. A universal extractor that handles new fields is in progress.
SpanExtractor (removed in 0.8)¶
solvi.extract_model.SpanExtractor, one field per pass with no field() helper and no save / load, had no caller and
is gone: importing SpanExtractor from solvi.extract_model still works in 0.8, warns, and gives LongSpanExtractor
(removed in 0.9). It trains on the same items, fit([(text, description, (s, e) or None), ...]); predict(text,
description) returns one (start, end, score, no_answer_score), and field(name, description) is the @extract part.
Hardware notes¶
- A GPU is recommended for training and for long documents.
- On a CPU, use the fp32 ONNX export; no int8 export is provided. If you quantize one yourself, compare its spans with the fp32 export's on your own fields before using it.