Check that extracted numbers add up
Reading a document tells you what it says. It doesn’t tell you whether what it says is right.
A scanner turns a 3 into an 8. A model reads the wrong column of a VAT summary. Both come back looking perfectly confident, because the reader genuinely did read something — and a wrong number that nobody questions is the one that gets paid.
Checks close that gap. You declare, once on the recipe, what the numbers ought to add up to:
VAT base × 21% should equal the VAT chargedEvery flow that uses that recipe now verifies the arithmetic, and tells you in plain language when a document doesn’t hold together.
Add a check to a recipe
Section titled “Add a check to a recipe”Checks live on the extraction recipe, alongside the fields it pulls out — not in an individual flow. Declare them once and every flow using that recipe inherits them.
A recipe’s checks block looks like this:
"checks": { "ok_field": "vat_consistent", "notes_field": "vat_check_notes", "rules": [ { "id": "vat_bracket_21", "label": "VAT at 21%", "assert": "vat_base_21 * 0.21 == vat_amount_21", "tolerance": { "abs": 1.0, "rel": 0.01 } } ]}Each rule takes four keys:
| Key | Required | What it does |
|---|---|---|
id | yes | A short stable name, unique within the recipe. Identifies the rule in results. |
label | no | What a person reads when it fails — “VAT at 21%”. Defaults to the id. |
assert | yes | The identity that should hold. See the grammar below. |
tolerance | no | How much rounding to forgive. Leave it out and the check is exact. |
ok_field and notes_field name the two fields the verdict is written into. They default to checks_ok and checks_notes; name them yourself when something downstream — a table column, a message template — already expects a particular name.
What you can write in a check
Section titled “What you can write in a check”The assert language is deliberately small — it’s arithmetic, not programming:
- Field names from the recipe:
vat_base_21,total_amount received_on— the day the document arrived (see below)- Numbers:
0.21,1 - Arithmetic:
+-*/, and brackets sum(…)— addition that tolerates missing values (see below)- Exactly one comparison:
==<=>=<>
Anything else is rejected the moment you save the recipe, with a message naming the rule:
check 'vat_bracket_21': 'a.b == c' contains Attribute which a check may not use;a check is arithmetic over value names, numbers, sum(…) and one comparisonThat’s deliberate. A typo you find while you’re looking at the recipe costs you seconds. The same typo discovered six weeks later, on a real invoice, costs you a document.
Checking a date against when the document arrived
Section titled “Checking a date against when the document arrived”Dates compare like numbers, counted in days: issue_date <= due_date says an
invoice can’t fall due before it was issued, and a tolerance of
{"abs": 1} forgives one day. A rule may also name received_on, the day the
document arrived:
received_on - issue_date <= 120The reading itself can’t know when the document arrived, so it skips such a rule. It is checked when the document is filed into a module’s own records — an expense document, for instance — where the arrival is known. A document issued a year before it arrived is far more often a misread year than a late invoice, and this is how that gets caught before it lands in the wrong month.
Why sum() exists
Section titled “Why sum() exists”Say you check that the VAT total matches the per-rate amounts:
vat_total == vat_amount_21 + vat_amount_12On an invoice that only uses the 21% rate, there is no 12% field at all — so that rule can never be evaluated, and the check quietly does nothing on your most common document.
sum() adds up whichever values are actually present:
vat_total == sum(vat_amount_21, vat_amount_12)Now one rule covers the single-rate and the multi-rate invoice. It only steps aside when every value it names is missing.
Tolerance: forgive rounding, not errors
Section titled “Tolerance: forgive rounding, not errors”Real accounting systems round VAT per line, so the total can sit a crown or two away from a clean base × rate. An exact check on money generates false alarms, and false alarms train people to ignore the flag.
Tolerance is abs, or rel (a fraction of the larger number being compared), whichever is more forgiving:
| Declared | Comparing | Result |
|---|---|---|
{"abs": 1.0} | 100.00 vs 100.60 | passes — within 1 |
{"rel": 0.01} | 1,000,000 vs 1,005,000 | passes — within 1% |
{"rel": 0.01} | 100 vs 105 | fails — 5% out |
| (none) | 1.000 vs 1.005 | fails — exact by default |
The rel part is what keeps one recipe honest across scales: the same 0.01 behaves sensibly on a 1,500 CZK invoice and a 1,500,000 CZK one.
Leave tolerance off when the numbers should match exactly — a count of line items, a quantity.
Three things checks will never do to you
Section titled “Three things checks will never do to you”Missing data is skipped, never failed. A rule that names a value the reader didn’t find simply doesn’t run. An invoice with no VAT summary at all — a foreign supplier under reverse charge, a supplier who isn’t VAT-registered — comes back consistent, with the note no checkable values extracted. A flag that fires on absent data is a flag people learn to ignore.
A broken rule never breaks a run. If a rule can’t be evaluated — dividing by zero, say — that one rule is reported and the rest still run. The document is not lost:
Ratio: could not be checked ('a / b == c' divides by zero)Checks don’t touch needs_review. How well we read the page and whether the document adds up are different questions with different answers. A crisp, perfectly-read invoice whose supplier did the arithmetic wrong should not look like a bad scan — those need different people and different fixes. Branch on the check field when you want to act on it.
What you get back, and where
Section titled “What you get back, and where”The verdict arrives in two places at once.
As two fields, inside the extracted fields. They look exactly like fields the reader found, so anything downstream reads them with no extra wiring — a table’s ingestion mapping picks them up like any other column, a template prints them, a Branch tests them:
{{ extract.fields.vat_consistent.value }} True / False{{ extract.fields.vat_check_notes.value }} the one-line verdictThe note is ok when everything passed, no checkable values extracted when nothing could be evaluated, and otherwise names what went wrong — with both numbers, so nobody has to reopen the PDF:
VAT at 21%: expected 282143.04 but the document shows 100000, off by 182143.04 (tolerance ±2821.43)
As a structured result, for a flow that wants the detail:
{{ extract.checks.ok }} true / false{{ extract.checks.notes }} the same one-line verdict{{ extract.checks.evaluated }} how many rules actually ran{{ extract.checks.results }} every rule: id, label, status, detail, lhs, rhs, toleranceEach rule’s status is pass, fail, skipped, or error.
Worked example: a supplier invoice
Section titled “Worked example: a supplier invoice”Here is the full set of rules the built-in invoice recipe uses. Four rules cover the whole of a Czech VAT summary:
"checks": { "ok_field": "vat_consistent", "notes_field": "vat_check_notes", "rules": [ { "id": "vat_bracket_21", "label": "VAT at 21%", "assert": "vat_base_21 * 0.21 == vat_amount_21", "tolerance": { "abs": 1.0, "rel": 0.01 } },
{ "id": "vat_bracket_12", "label": "VAT at 12%", "assert": "vat_base_12 * 0.12 == vat_amount_12", "tolerance": { "abs": 1.0, "rel": 0.01 } },
{ "id": "vat_total_sum", "label": "Total VAT", "assert": "vat_total == sum(vat_amount_21, vat_amount_12)", "tolerance": { "abs": 1.0 } },
{ "id": "header_total", "label": "Invoice total", "assert": "sum(vat_base_21, vat_base_12, vat_base_0) + sum(vat_amount_21, vat_amount_12) == total_amount", "tolerance": { "abs": 1.0, "rel": 0.01 } } ]}Run a real rent invoice through it — base 1,343,538.27 at 21%, VAT 282,143.03, total 1,625,681.30 — and you get vat_consistent = True, vat_check_notes = ok. The 12% rule skips itself, because that invoice has no reduced-rate lines.
Change the VAT to 100,000 and the same recipe says:
VAT at 21%: expected 282143.04 but the document shows 100000, off by 182143.04 (tolerance ±2821.43); Total VAT: expected 282143.03 but the document shows 100000, off by 182143.03 (tolerance ±1); Invoice total: expected 1443538.27 but the document shows 1625681.30, off by 182143.03 (tolerance ±16256.81)
Three rules caught it, each naming its own numbers.
Acting on a failed check
Section titled “Acting on a failed check”A failed check is information, not a verdict on what to do next — that’s your flow’s decision. Two patterns cover most cases:
Save it, flagged. Let the row land with vat_consistent = False and the note beside it. A flagged row somebody can review beats a lost document, and the reason travels with the record.
Route it to a person. Put a Branch on {{ extract.checks.ok }} and send the failures to Flag for review or Ask a person, so someone confirms before it’s booked.
Where to go next
Section titled “Where to go next”- Working with data: tables & imports — landing extracted documents into a table
- Keep a human in the loop — routing anything doubtful to a person
- Extract — the full field reference for the reader itself