Document Understanding

Reading the documents a process runs on, with a person on what the model is unsure about.

Engagement

Implementation. Scoped to a document set, with accuracy measured before and after.

Who it is for

Finance, operations and shared services teams processing documents by hand.

When people call us

Invoices, field tickets or statements arrive in dozens of layouts, and someone keys the header data every time.

What we build

01

Document taxonomy

The document types the process handles and the fields that matter in each — the design step that decides how well everything downstream works.

02

Digitization

Scanned, photographed and native files handled, including the ones that arrive as a phone picture of a ticket in a pickup truck.

03

Classification

Sorting what arrived before anything is extracted, so an invoice, a statement and a delivery note each take the right path.

04

Extraction

Machine learning where layouts vary, rules and regular expressions where they do not, and both where that is the honest answer.

05

Validation with a person in the loop

Action Center review for anything below a confidence threshold, so low-certainty results become corrections rather than errors.

06

Confidence thresholds

Set with the business against the cost of being wrong, not left at a default nobody chose.

07

Accuracy measurement

Straight-through rate, field-level accuracy and correction volume, measured before and after so the result is a number.

08

Retraining

Corrections fed back so the model improves, with someone owning the cycle rather than it ending at go-live.

From an attachment to a posted record

The model handles the confident cases. People handle the rest.

  1. 01
    Digitize

    Scans, native files and the phone photograph of a field ticket, all made readable.

  2. 02
    Classify

    Invoice, statement, delivery note — sorted before anything is extracted.

  3. 03
    Extract

    Machine learning where layouts vary, rules where they do not, against a taxonomy designed first.

  4. 04
    Post

    Straight through into the system of record, with the confidence score kept.

Below the threshold, a person validates

Action Center review for anything the model was unsure about, at a confidence threshold set against what being wrong costs. Every correction is a labelled example, fed back so accuracy improves rather than stalling at go-live.

We build the document side of an automation: the taxonomy, classification and extraction, the validation step where a person confirms what the model was unsure about, and the accuracy measurement around all of it.

The taxonomy decides the ceiling

Document types and the fields inside them are a modelling exercise before they are a technology one. Get it wrong and no amount of training data compensates; get it right and the rest is tuning.

Confidence thresholds are a business decision

How certain is certain enough depends on what being wrong costs. A misread invoice total is not a misread reference number. We set those thresholds with the people who carry the consequence, and route everything below them to a person.

Corrections are training data

Every validation a person performs is a labelled example. Feeding those back is what turns a static deployment into one that gets better — and it needs an owner, because it does not happen on its own.

What you keep

The code and the data·Your existing relationships·Approval and control·The ability to stop·The off switchWhat that means

Start with a 45-minute briefing.

No pitch. We’ll map your situation against what actually works and tell you honestly where to start.