Research

Accuracy by workflow, including the ones we are bad at

A single headline accuracy figure is the least useful number a vendor can publish, because it averages a workflow that reaches ninety-four percent with one that struggles to reach sixty-six. Here is the breakdown by workflow, with the spread across customers and the method stated.

Benchmark your own

Send fifty documents. We will report straight-through and confidently-wrong rates against your team’s actual coding.

1 / 3
n = 23 customers2025–2026Spread published, not just median
WorkflowWeek 4 → week 20Wk 20
AP · bill codingRecurring vendors dominate; climbs fastest
91%
AP · PO matchingStructured comparison, little ambiguity
94%
Bank reconciliationExact and split only; fuzzy always waits
93%
Expense codingReceipt quality is the binding constraint
84%
Vendor deduplicationGenuinely hard; we score ourselves low
69%
Multi-line allocationRequires context often not on the document
66%
Contract obligationsProposes only; judgement stays human
61%
Median across 23 customers, 2025–2026. Shaded band is the interquartile spread — the bottom quartile of customers sits well below these figures.

How to read this without being misled

Three things are worth noticing before the numbers themselves.

  • The spread matters more than the median. Vendor deduplication has a 20-point interquartile spread, which means the median of 69% describes almost nobody. A quarter of customers sit meaningfully below it.
  • The bottom three workflows are genuinely weak. Vendor deduplication, multi-line allocation, and contract obligation identification do not reach seventy percent, and we are not going to present that as a rounding issue.
  • Week 4 to week 20 is the honest window. Quoting a week-20 figure without the starting point implies immediate performance that does not exist.
A median with a twenty-point spread describes almost nobody. Publishing the spread is the difference between a benchmark and a marketing figure.

Why the weak workflows are weak

Vendor deduplication requires deciding whether two records are the same commercial entity, which is often genuinely ambiguous — a subsidiary and its parent, a company mid-rebrand, a franchisee. Human reviewers disagree with each other on these, so a ceiling below human consensus is unsurprising.

Multi-line allocation frequently depends on information not present on the document. Splitting a facilities invoice across departments needs headcount or square footage, and where no allocation rule exists the agent is guessing at something unknowable.

Contract obligation identification is a judgement under ASC 606, and the agent proposes rather than decides by design. Measuring it as accuracy is slightly unfair to it, and we include it because omitting the workflow we score worst on would make the table dishonest.

Method

  • Population. 23 customers live for at least twenty weeks between January 2025 and June 2026. Small sample, stated deliberately.
  • Metric. Straight-through rate — the share of documents reaching a posted transaction with zero human edit. Measured from the audit trail rather than reported separately.
  • Correction defined. Any human edit to a proposed action before or after posting, including account, dimension, amount, vendor, or rejection. Approval without modification is not a correction.
  • Excluded. Customers under 100 monthly documents in a workflow, because the figures are too noisy to mean anything.
What this is not

It is not independent research. We collected it from our own customers, on our own product, and we chose which workflows to publish. The mitigations are that the method is stated, the spread is shown, and the workflows we perform worst on are included. Treat it as a vendor disclosure rather than a study.

The curve

What the first twenty weeks look like.

Accounts payable for a professional-services customer at roughly 350 bills a month. The steep section is the recurring vendor base being learned; the flattening is the long tail.

0%25%50%75%100%wk 1wk 2wk 4wk 6wk 8wk 12wk 16wk 2091%

Questions

About these numbers.

Why publish the workflows you are worst at?
Because a table showing only the strong ones is an advertisement. The bottom three are where a buyer should push, and hiding them would mean nobody could trust the top three either.
Is this independently verified?
No. It is our data, from our customers, on our product. The method is published and the spread is shown, but treat it as vendor disclosure rather than research.
What will we actually get?
It depends more on your vendor concentration and coding consistency than on anything we control. A business with 40 recurring vendors climbs faster than one with 400 occasional ones. We would rather benchmark your data than have you rely on ours.
Do you report confidently-wrong rate?
Yes, per customer in the product, and it is the number we gate releases on. We have not published it as a benchmark because the denominators differ enough across customers to make a cross-customer figure misleading.
How often is this updated?
Quarterly, with the population size and period stated each time. If a number moves down, it moves down in public.

Benchmark your own documents.

Fifty real bills is enough to produce your numbers rather than relying on ours.