Vendor bills
Header, line items, tax, freight, terms, remit-to, PO reference, and bank details — including the multi-page ones where the total lives on page four.
AI capability
Every finance function runs on documents somebody retypes. Extraction turns them into records with a confidence score per field, and the score is what makes it usable — a system that returns eleven fields without telling you which two it is unsure about has moved the problem rather than solved it.
What it does
Header, line items, tax, freight, terms, remit-to, PO reference, and bank details — including the multi-page ones where the total lives on page four.
Photographed, crumpled, thermal, and faded. Merchant, amount, tax, date, and category, matched against the card transaction that already posted.
PDF statements from institutions with no usable feed, parsed into transactions that reconcile like any other line.
Term, value, ramps, renewal and termination clauses, and deliverables — feeding revenue recognition rather than being summarised into a notes field.
W-9s, certificates of insurance, and lien waivers captured with their expiry dates so they chase themselves before they lapse.
Not one score for the document. Eleven fields, eleven scores, and only the ones below threshold interrupt a person.
Most extraction products report one accuracy figure for a document. That is the wrong unit, because the fields are not equally consequential. Getting a vendor name slightly wrong is recoverable; getting an amount wrong by a decimal place is a payment error.
Scoring each field separately means the routing can be proportionate. A bill where every field is high-confidence except the PO reference goes to a person with that one field highlighted, taking four seconds. The same bill under a document-level score either interrupts a person for everything or nothing.
Not exotic layouts. The recurring problems are mundane: a vendor changing its invoice template mid-year, a supplier issuing statements that look like invoices, credit notes formatted identically to bills with a minus sign somewhere unobvious, and multi-page documents where the total appears on both the first and last page with different figures because one includes freight.
Each of those is handled with a specific check rather than left to general capability. The statement-versus-invoice distinction in particular is one of the highest-value ones, because paying a statement is how you pay the same invoice twice.
Extraction on a vendor you have never seen is a general-capability problem. Extraction on the forty vendors who send you ninety percent of your bills is a pattern-learning problem, and that is where accuracy climbs quickly. After a handful of documents from a given vendor, the layout is known and the confidence scores rise accordingly.
This is why the accuracy curve steepens in weeks two to six rather than arriving flat, and why we would rather show you the curve than quote a single headline figure.
Limits
Every capability page on this site carries one of these, because a feature described without its boundaries is a claim rather than a description.
Printed and typed text is solid. Handwritten amounts on a delivery note are not, and those are flagged low rather than guessed at.
It is tuned for finance documents. Feed it an engineering drawing or a legal brief and it will do a mediocre job, because it was not built for them.
A photograph of a screen showing a PDF, at an angle, in poor light, produces poor extraction. We surface that as a quality problem rather than absorbing it silently.
Questions
We will report extraction accuracy field by field against what your team actually entered.