solvi.aliases¶
Name matching: when a catalog's parameter names do not match its facts. accept and apply are plain code; the matcher
model (NameMatcher) is experimental: no checkpoint is published.
Name matching: when the catalog's parameter names do not match its facts —
parts written by different teams, each with its own naming style — a matcher model proposes aliases ("the parameter
INVC_AMNT is the fact invoice_total"), and deterministic code decides which to accept.
from solvi.aliases import NameMatcher, propose, accept
m = NameMatcher.load("path/to/strategist-checkpoint/matcher")
props = propose(cat, questions, init_keys, m) # [Proposal(aliases={name: fact}, score)]
got = accept(cat, questions, props, examples, probes=states, oracle=ask_person) # labelled examples + targeted questions
if got.aliases is not None:
cat2 = apply(cat, got.aliases) # parts rewired to the facts; cat2.aliases lists them
System(cat2, questions) # plans by exact names again; the trace is unchanged
Acceptance: a proposal is accepted only if its answers match every labelled example AND no neighbouring wiring
(another proposal, one alias swapped for another candidate, two same-typed aliases exchanged) also matches the examples but
answers differently on the unlabelled probes. In the active mode, while such neighbours remain, solvi picks the probe on
which most of them disagree with the proposal and asks oracle(state) for its right answers (a person labels that case);
a proposal that contradicts an answer is dropped. No neighbour left → accepted; questions used up → not accepted
("ambiguous": ask a person). The model never decides alone: a wrong alias costs coverage (no aliases accepted), not a
silent wrong wiring — as far as the examples and probes can tell wirings apart.
Accepted aliases are applied by rewiring: every part that read an aliased name now reads the fact itself (its function is
wrapped, its docstring lists the aliases), and the new catalog lists them in catalog.aliases; planning, checks, quotes into
given texts, execution, the trace and replay behave exactly as for a catalog written with one naming.
The matcher model (NameMatcher) is experimental: no checkpoint is published — NameMatcher.load reads one you trained
yourself (docs/strategist.md has the format), and examples/17 uses a stand-in. propose takes any object with the same
methods; accept and apply are plain code.
NameMatcher ¶
The link encoder: all-MiniLM-L6-v2 over texts (name, type, docstring), for arch "char" fused with a character CNN over the identifier (score = w_t·cos_text + w_c·cos_char), for arch "text" the text part alone; divided by T (0.05). Backends: "torch" (transformers) or "onnx" (onnxruntime + tokenizers).
Experimental: no checkpoint is published — it reads one you trained yourself (docs/strategist.md has the format).
split_ident ¶
An identifier → lowercase words (snake, camel, UPPER): exactly the matcher's training split.
unresolved ¶
Names that parts read but that are neither given nor a fact → {name: [(reader part, type)]}.
link_table ¶
For each unresolved name its top-k sources with log-probabilities → {name: [(source, logp)]}.
propose ¶
propose(catalog, questions, init_keys, matcher, k=4, beam=16, max_props=24, prune=4.6, init_types=None, table=None)
Joint alias proposals by beam search over the unresolved names (best first): each name takes one of its top-k sources; a part never reads one fact twice; no cycles. → [Proposal] by score.
apply ¶
A new catalog in which every part reads the aliased facts under their own names (name → fact): the result is the
catalog as if the teams had used one naming — planning, checks on computed facts, quotes into given texts, types and
the trace behave exactly as they would, and every declaration of a part (cost, timeout, blocking, validate, …) is
kept. The accepted aliases are listed in catalog.aliases (name → fact) and in the
rewired parts' docstrings.
accept ¶
accept(catalog, questions, proposals, examples, probes=(), k=5, oracle=None, active=None, check_neighbours=True, minimal=True, strategist=None, confirm=True)
Deterministic acceptance of alias proposals. examples: [(state, {question: answer})] labelled cases (the first k are used); probes: unlabelled states (targeted questions and the distinguishability check); oracle(state) → {question: answer} (a person) for the active mode; active=(k0, m): k0 labelled examples + up to m targeted questions (instead of k examples). minimal: drop the accepted aliases that change no answer on the examples and probes (they stay unresolved). confirm (active mode): spend the whole question budget even when no neighbour disagrees any more (the extra questions go where the other proposals disagree most), so a wiring is never accepted on the k0 random labels alone. strategist: plans each candidate wiring (default CostStrategist(): the deterministic plan with dead ends dropped).
match_names ¶
match_names(catalog, questions, init_keys, matcher, examples, probes=(), oracle=None, active=(3, 7), k=5, init_types=None)
Everything in one call: propose, accept (active mode when an oracle is given), apply. → (catalog, Acceptance) — the catalog with the accepted aliases, or the original one when nothing was accepted (or nothing was unresolved).