Skip to content

The catalog

from solvi import Catalog, Quote

cat = Catalog()
Decorator Kind Returns Sets fact
@cat.fn computation any value function name
@cat.check soft check bool function name
@cat.check(hard=True, then={"question": "answer"}) hard check bool function name
@cat.extract extraction from text Quote(value, start, end, source="doc", confidence=1.0) function name
@cat.rule("question") answer rule one of the question's options (a bool is accepted for yes/no) the answer

A part's contract is its signature. You do not write a separate schema:

  • argument names are the inputs: keys of init_state or facts produced by other parts;
  • the function name is the output fact.

Part names must be unique in a catalog (a duplicate raises ValueError). A rule is registered per question; registering a second rule for the same question replaces the first. The decorators return the original function, so parts stay ordinary, testable Python. A part and a key of init_state cannot share a name: ask raises ValueError for an input key named like a part (the given value would replace the part — a hard check included), and System(input_model=Model) refuses a model with such a field when it is built. Name a part after what it computes (savings_points), not after the input it reads (def savings(savings) reads its own name).

A part may be an async def function (a database or HTTP lookup): system.aask awaits it, concurrently with the other steps; timeout= (seconds) and blocking=True (a sync function that waits: run in a worker thread) on any decorator apply under aask (see Async execution).

@cat.fn
def total(items):                           # reads "items", sets "total"
    return sum(q * p for q, p in items)

@cat.check
def big_order(total):                       # soft check: a fact like any other
    return total > 1000

@cat.check(hard=True, then={"ship": "no"})  # hard check: if False, the answer to "ship" is forced to "no"
def paid(payment_status):
    return payment_status == "paid"

@cat.rule("ship")
def ship_rule(big_order):
    return not big_order                    # bool is normalized to "yes"/"no" for yes/no questions

Soft and hard checks

  • A soft check is a boolean fact. Rules can use it, answer heads can learn from it, and failed soft checks are listed in the reason of a learned answer.
  • A hard check decides the answers it governs whenever it is in the question's flow and evaluates to False:
  • if then names the question, the answer is set to that value with status="forced" and confidence 1.0;
  • if the question lists the check in its requires (or the check has no then at all) but then does not name it, the question abstains;
  • otherwise — the check has a then for other questions and only reads this question's facts — it is an ordinary failed check for this question and is listed in the reason.

A check can say why it is False: return Fail("Harold is busy 13:30 - 15:30") — see A check that says why.

No rule and no model confidence can override a failed hard check. To make sure a hard check is always in a question's flow, list it in the question's requires. When several hard checks fail, the first one declared in the catalog decides: declare the most important first. A question's flow does not depend on which other questions are asked in the same request, and neither does its answer — except through a constraint between answers, which applies only when all its questions are asked.

A hard check that could not be evaluated (it raised, or a fact it reads is missing) never counts as passed: the questions it governs abstain. The reason names the check and, when the check could not run for lack of an input, the part further up that failed: hard check day_allowed could not be evaluated: missing inputs: violations; caused by spec: ValueError: no slot in the plan. A rule that could not run says the same (rule not computed: ...; caused by ...).

Extractors and Quote

An extractor reads text from init_state and returns a Quote:

import re

@cat.extract
def amount(doc):
    "the total amount to pay"
    m = re.search(r"Total due:\s*([0-9,]+\.[0-9]{2})", doc)
    if not m:
        raise ValueError("no total")        # a failed part becomes a missing fact; dependent answers abstain
    return Quote(float(m.group(1).replace(",", "")), m.start(1), m.end(1))
  • doc[start:end] is the supporting quote. source names the init_state key holding the text. When the Quote does not name it, it is the extractor's text: "doc" if the function reads doc, else its only argument (def refund_word(message) quotes message), else its only str-typed argument; if that is ambiguous, registration raises and asks for @cat.extract(source="..."). A Quote that names another source keeps it.
  • A quote whose offsets fall outside the source text is rejected: the fact is missing (dependent answers abstain) and the error is recorded. A quote from a model-backed part must also be literally the text at its offsets — see Grounded decisions. Hand-written code may return a value derived from the quoted text (a label, a parsed number, a date); @cat.extract(exact=True) makes it literal too.
  • confidence (default 1.0) is propagated to every answer that depends on this value (see Confidence).
  • Downstream parts receive the plain value, not the Quote.

You can write extractors by hand (regular expressions, parsers) or get them from a trained model (Extracting fields from documents).

Several producers of one fact (fallback chain)

A fact can have alternative producers — a cheap regular expression and a slower model, say. Give each its own function name and say which fact it provides:

@cat.fn(provides="total", cost=0.1, validate=lambda v: v > 0)
def total_regex(doc):
    m = re.search(r"TOTAL\s*:?\s*(\d+\.\d\d)", doc)
    return None if m is None else float(m.group(1))           # None = "not found": rejected, the next producer runs

@cat.extract(provides="total", cost=50, min_confidence=0.8)
def total_model(doc):
    return extractor.field("total", "the total amount paid")(doc)

@cat.features("total")                                      # optional: cheap features for the producer policy
def total_features(doc):
    return {"lines": doc.count("\n"), "has_total": "TOTAL" in doc.upper()}
  • The producers form a group in the catalog under the fact's name; the strategist plans it like one part (its inputs are the union of the producers' inputs). Catalogs without provides behave exactly as before.
  • A producer's output is accepted only if it is not None, a Quote lies inside its text (literally, for a model-backed producer), a Decision is among its options, it reaches min_confidence, and validate(value, [any of the producer's inputs by name]) returns true. An exception counts as a rejection. The first accepted output is used; if none is accepted the fact is missing and dependent answers abstain.
  • Order: declaration order by default. System(cat, questions, producers="learned") lets a policy choose the order per input: producers predicted to agree with the reference producer (the last declared, normally the most trusted) at least system.producer_policy.quality of the time (0.9) go first, cheapest expected cost first (cost ÷ P(accepted)). The policy learns from every run: acceptance of each producer that ran, and — when two producers ran on the same input — whether the cheaper one agreed with the reference. With probability producer_policy.explore (0.1) the remaining producers also run as a shadow (never used, only compared) so the policy keeps getting agreement labels. The policy only orders; acceptance is always the producer's own deterministic check.
  • The record of the fact says which producer was used (record.producer) and every producer that ran with its outcome (record.tried, e.g. [["total_regex", "no value"], ["total_model", "accepted"]]). Both are hashed into the chain. replay recomputes the value with the producer that was used, checks it still passes its validator, and checks that the producers tried before it are still rejected (shadow runs are not re-checked).
  • cost= (ms) is a prior; measured run times replace it (system.cost_book).