Document Understanding
Reading the documents a process runs on, with a person on what the model is unsure about.
Implementation. Scoped to a document set, with accuracy measured before and after.
Finance, operations and shared services teams processing documents by hand.
Invoices, field tickets or statements arrive in dozens of layouts, and someone keys the header data every time.
What we build
Document taxonomy
The document types the process handles and the fields that matter in each — the design step that decides how well everything downstream works.
Digitization
Scanned, photographed and native files handled, including the ones that arrive as a phone picture of a ticket in a pickup truck.
Classification
Sorting what arrived before anything is extracted, so an invoice, a statement and a delivery note each take the right path.
Extraction
Machine learning where layouts vary, rules and regular expressions where they do not, and both where that is the honest answer.
Validation with a person in the loop
Action Center review for anything below a confidence threshold, so low-certainty results become corrections rather than errors.
Confidence thresholds
Set with the business against the cost of being wrong, not left at a default nobody chose.
Accuracy measurement
Straight-through rate, field-level accuracy and correction volume, measured before and after so the result is a number.
Retraining
Corrections fed back so the model improves, with someone owning the cycle rather than it ending at go-live.
The model handles the confident cases. People handle the rest.
- 01Digitize
Scans, native files and the phone photograph of a field ticket, all made readable.
- 02Classify
Invoice, statement, delivery note — sorted before anything is extracted.
- 03Extract
Machine learning where layouts vary, rules where they do not, against a taxonomy designed first.
- 04Post
Straight through into the system of record, with the confidence score kept.
Action Center review for anything the model was unsure about, at a confidence threshold set against what being wrong costs. Every correction is a labelled example, fed back so accuracy improves rather than stalling at go-live.
We build the document side of an automation: the taxonomy, classification and extraction, the validation step where a person confirms what the model was unsure about, and the accuracy measurement around all of it.
The taxonomy decides the ceiling
Document types and the fields inside them are a modelling exercise before they are a technology one. Get it wrong and no amount of training data compensates; get it right and the rest is tuning.
Confidence thresholds are a business decision
How certain is certain enough depends on what being wrong costs. A misread invoice total is not a misread reference number. We set those thresholds with the people who carry the consequence, and route everything below them to a person.
Corrections are training data
Every validation a person performs is a labelled example. Feeding those back is what turns a static deployment into one that gets better — and it needs an owner, because it does not happen on its own.
What you keep
The code and the data·Your existing relationships·Approval and control·The ability to stop·The off switchWhat that means
Start with a 45-minute briefing.
No pitch. We’ll map your situation against what actually works and tell you honestly where to start.