Extract
Module · Stable
Turn an unstructured source (scanned/digital PDF, image, or free text) into typed fields with per-field confidence, driven by a reusable extraction recipe.
The single, format- and domain-agnostic reader for unstructured documents. A reusable, versioned extraction RECIPE declares the target fields (typed + required), the reader mode, and a confidence policy; the node reads the source — a PDF text-layer-first with a vision fallback, an image via OCR, or free text as-is — then runs one structured LLM pass that returns each field as {value, value_normalized, confidence, source_page, source_excerpt}. Honours the recipe’s confidence policy: a field below the floor is dropped to null (uncheckable but safe — never a confidently-wrong value), and a weak required field sets needs_review so a Branch can route to a person. A recipe may also declare CHECKS — arithmetic identities its numbers must satisfy (a VAT base times its rate equalling the VAT charged, the parts summing to the header total). Checks run after the read, with a declared rounding tolerance; a rule whose values weren’t extracted is skipped rather than failed, so a document with nothing to check comes back consistent. The verdict lands in ‘checks’ AND as two synthetic fields inside ‘fields’, so a table binding or a Branch reads it like any other value. No document standard is hardcoded — an invoice is just a recipe authored on top. A read of a STORED document under a NAMED recipe is kept, keyed on the document’s content hash plus the recipe’s key and version, so re-running a flow over documents it has already read costs nothing and returns the same answer with ‘reused’ set. A recipe’s version is found by its content, so changing the recipe re-reads everything under it — nobody bumps a number; ‘refresh’ forces one re-read of an unchanged recipe. Every read is stamped with the id of the exact recipe version it was made under (‘recipe_version_id’). Inside a Job, a read of a stored document is also kept as that Job’s READING of it (the recipe, what the recipe says the document is, the fields), listed on the document’s page in Memory beside any other Job’s reading of it. Never raises on an extraction failure; mis-configuration (no source / no recipe) does raise.
When to use
Section titled “When to use”Use when a document arrives from an upstream step — an Incoming email attachment, a Document Store output, or an actor upload — and you need its fields (invoice number, total, dates, …) as typed values WITH a confidence you can branch on. Apply the matching extraction recipe by slug; the fields dict and needs_review flag feed naturally into a Branch (confidence gate), an Ask a person review step, Create record, or any send step. When the numbers also have to ADD UP, declare the arithmetic as checks on the recipe instead of computing it in a Code step — the rates and tolerances become editable data, and every flow using that recipe gets the check.
When not to use
Section titled “When not to use”For a spreadsheet (Excel / CSV) do NOT use an LLM — read it deterministically with Read spreadsheet. (An ISDOC e-invoice, .isdoc or .isdocx, or another XML file IS fine here: Extract reads it natively as labelled text, no OCR.) For a printed leaflet with MANY price tags per page use Read PDF. For one field on a proposal PDF the Extract Proposal Field tile is more specific. To extract from audio, transcribe first with Speech to Text and run Extract on the transcript.
Inputs
Section titled “Inputs”Configured per use: file_path, ref, text, recipe_key, recipe_version, recipe, fields, mode, model, refresh.
Outputs
Section titled “Outputs”fieldsmissingneeds_reviewfilledread_outcomeoverall_confidencereaderrecipe_keyrecipe_versionrecipe_version_idai_modelchecksreusedreading_id
Auto-generated from the skill registry (load_skills()). Do not edit by hand.