solvi.agree¶
Agreement of generated candidates under a key you give: the largest group, its share as a fact, the tally recorded.
Agreement of generated candidates under a key: K outputs (samples of one model, or one each from several models), grouped by what you say makes two of them the same, the largest group chosen, its share as a signal.
from solvi.agree import agree, consensus
c = consensus(queries, key=lambda sql: digest(run(db, sql))) # {"index", "value", "share", "groups", "keys", ...}
cat.fn(writer.part("candidates", prompt, k=3)) # three queries (solvi.generate)
agree(cat, "sql", candidates="candidates", key=row_digest) # facts: sql, sql_agreement, sql_tally
solvi.multi.Vote combines decisions over the same closed options; generated outputs have no options — two queries that
differ in text can return the same rows, two plans in different words can be the same order. key(candidate) says what
counts as the same: the digest of the rows a query returns, a normalized plan, a parsed number. The share of the
candidates in the chosen group is a plain number fact (<name>_agreement), so a rule, a learned head or a guarantee can
read it like any other signal; "all K agree" is the share 1.0.
Rules. A candidate that is None (its generation failed), whose key raises, or whose key is None does not vote; it still
counts in K, so a failed candidate lowers the share. prefer(candidate): when any candidate with a key passes it, only
those vote (rows that are not empty before an empty result, say); otherwise all with a key vote. The largest group wins;
a tie goes to the group whose first candidate comes first (with K samples, the first is usually the greedy one). The
chosen value is the first candidate of that group. Nothing voted: no value — the fact <name> is missing (the part
raises with each candidate's reason) and the questions that need it abstain, while <name>_tally and
<name>_agreement (0.0) are still there for the checks and the feedback.
The record. <name>_tally is a plain dict in the trace — per candidate its key (or why it has none) and whether it
voted, the groups, the chosen index, the share — so the audit shows how the vote went and replay recomputes it from the
recorded candidates (the key function is re-run: keep it deterministic, or cache what it computes).
Not done here: agreement is not correctness — K samples of one model often agree on the same mistake; measure the share against labels before you trust it, and put a guarantee on it rather than a hand-picked threshold.
consensus ¶
K candidates → the tally: {"index" (the chosen candidate, -1 when nothing voted), "value", "key", "share" (the chosen group's size / K), "k", "groups" ([[indexes]], largest first), "candidates" ([{"index", "key" or "why", "voted"}])}. key(candidate[, facts by name]) → anything hashable that says which candidates are the same; prefer: see the module docs. facts: the extra arguments key / prefer take by name (agree passes the catalog's facts).
agree ¶
Register in cat the agreement of the candidates in the fact candidates (a list) under key → three facts:
<tally> (default <name>_tally: the record, see consensus), <name> (the chosen candidate; missing when nothing
voted) and <share> (default <name>_agreement: the chosen group's share of all K, 0.0 when nothing voted).
key / prefer: functions of a candidate and, by name, of other facts (def key(sql, db_path)), which become inputs
of the tally part. → the names (tally, name, share).