Skip to content

solvi.extract_multi

A single-pass @extract with a pair of pointer heads per field (needs torch).

Single-pass @extract: ModernBERT reads the document once, with a pair of pointer heads (start / end) per field. For solvi: extractor.field(name) returns a function doc → Quote; all fields of a document come from one pass (cached by text).

The extractor protocol it shares with solvi.extract_long.LongSpanExtractor: fit(items), predict(text, field), field(name[, description]), save(path) / load(path), fingerprint(). items here are [(text, {field: (start, end) | None})]; 0.7's fit(docs, spans) and predict_doc(text) still work with a SolviDeprecationWarning.

MultiSpanExtractor

MultiSpanExtractor(fields, model_name='answerdotai/ModernBERT-large', max_len=1024, device=None)

fit

fit(items, spans=None, epochs=4, lr=3e-05, bs=8, seed=0, log=print)

items: [(text, {field: (start, end) | None})]. (0.7: fit(docs, spans) — still read, with a warning.) A span past the encoded text (max_len tokens) is left out of training.

predict_doc

predict_doc(text, max_span=64)

Deprecated (removed in 0.9): predict(text).

predict

predict(text, field=None, max_span=64)

→ {field: (start, end, confidence)} in a single pass, or one field's (start, end, confidence).

fingerprint

fingerprint()

A stable hash of this extractor: fields, settings, temperatures, sampled weights (see LongSpanExtractor).

fit_temperature

fit_temperature(docs, gold_ok, grid=(0.5, 0.75, 1.0, 1.5, 2.0, 3.0, 5.0))

Per-field temperature fitted on held-out documents: minimizes log loss of confidence vs. whether the span matched gold. gold_ok(field, doc_index, (start, end)) → bool.

save

save(path)

Encoder weights (bf16 safetensors), tokenizer, the span heads and the settings into a directory (load reads it back) — as LongSpanExtractor.save.

load classmethod

load(path, device=None)

From a directory written by save(), or a Hugging Face model id (downloaded once and cached).