Skip to content

Extracting fields from documents

With pip install "solvi[model]", solvi provides three ModernBERT extractors. All of them predict a start and an end position in the text, so the extracted value is always a substring of the document with exact offsets. Each gives you plain functions doc -> Quote to register with cat.extract.

Labels are character spans: for each training document and field, (start, end) of the value in the text, or None if the field is absent. extract-base's model card suggests labelling about 25–100 documents per task and fine-tuning. A GPU is recommended for training and for fast inference.

MultiSpanExtractor: all fields in one pass

solvi.extract_multi.MultiSpanExtractor reads a document once and has a start/end head pair per field. Use it for documents that fit into one window (receipts, invoices, forms).

from solvi.extract_multi import MultiSpanExtractor

ex = MultiSpanExtractor(["company", "date", "total"],
                        model_name="answerdotai/ModernBERT-large", max_len=1024)
ex.fit(train, epochs=4, lr=3e-5, bs=8)
# train: [(text, {"company": (s, e), "date": (s, e), "total": (s, e) or None})]
ex.save("receipts-extractor")              # MultiSpanExtractor.load("receipts-extractor") reads it back

ex.fit_temperature(calib_docs, lambda field, i, span: span == calib_spans[i][field])   # optional, per-field temperature

for name in ["company", "date", "total"]:
    cat.extract(ex.field(name))            # registers the fact "company", ... read from init_state["doc"]

@cat.fn
def total_value(total):
    return float(total.replace(",", ""))
  • ex.field(name) returns a function named name with one argument doc, which returns Quote(doc[s:e], s, e, confidence=c).
  • ex.predict(text) returns {field: (start, end, confidence)} (ex.predict(text, field) one of them). Results are cached per text, so all fields of one document cost a single forward pass. The two extractors share one protocol — fit(items), predict(text, field), field(name[, description]), save / load, fingerprint(); 0.7's fit(docs, spans) and predict_doc(text) still work with a deprecation warning. A labelled span past the encoded text (max_len tokens) is left out of training.
  • fit_temperature(docs, gold_ok) picks a softmax temperature per field that minimizes log loss of "confidence vs. correct" on held-out documents; gold_ok(field, doc_index, (start, end)) tells whether a prediction is correct.
  • The extractor always returns a span; this one does not model "field absent". Use LongSpanExtractor when fields may be missing.

LongSpanExtractor: long documents, fields by description, "no answer"

solvi.extract_long.LongSpanExtractor is for long documents (contracts) and optional fields. The input is [CLS] field description [SEP] window of the document [SEP], windows overlap, and position 0 means "no answer in this window".

from solvi.extract_long import LongSpanExtractor

GOV_LAW = "the clause that says which state's or country's law governs the contract"

lx = LongSpanExtractor(max_len=1024, stride=128, max_span=96)
lx.fit([(text, GOV_LAW, span_or_none) for text, span_or_none in train], epochs=3, neg_per_item=3)
lx.tune_threshold("governing_law", [(text, GOV_LAW, span_or_none) for text, span_or_none in heldout])

cat.extract(lx.field("governing_law", GOV_LAW))

@cat.fn
def governing_state(governing_law):
    for s in ["Delaware", "New York", "California"]:
        if s.lower() in governing_law.lower():
            return s
    return "other"
  • stride is the overlap between windows in tokens; max_span caps answer length in tokens.
  • Training uses every window that contains the answer plus up to neg_per_item windows without it. One extractor can learn several fields: pass items with different descriptions.
  • predict(text, desc) returns (start, end, span_score, no_answer_score), the best span over all windows.
  • tune_threshold(name, items) picks the score threshold for "the field is present" that maximizes present/absent accuracy on held-out items.
  • field(name, desc) returns a function doc -> Quote. When the score is below the threshold, it returns Quote("", 0, 0), an empty value, so downstream functions should treat "" as "not found" (e.g. has_tax = tax != "").
  • Cost grows with length: one pass per field per window, so a long contract with several fields takes many passes; use a GPU for long documents.
  • Fields are specified by description, but a field that was never labeled in training is not extracted reliably from its description alone (extract-base's model card: 14% and 66% on two held-out fields for a model trained on other fields only). Label examples for every field you need. A universal extractor that handles new fields is in progress.

SpanExtractor (removed in 0.8)

solvi.extract_model.SpanExtractor, one field per pass with no field() helper and no save / load, had no caller and is gone: importing SpanExtractor from solvi.extract_model still works in 0.8, warns, and gives LongSpanExtractor (removed in 0.9). It trains on the same items, fit([(text, description, (s, e) or None), ...]); predict(text, description) returns one (start, end, score, no_answer_score), and field(name, description) is the @extract part.

Hardware notes

  • A GPU is recommended for training and for long documents.
  • On a CPU, use the fp32 ONNX export; no int8 export is provided. If you quantize one yourself, compare its spans with the fp32 export's on your own fields before using it.