Research · updated August 2026

What our agents actually achieve

Every vendor in this category publishes an accuracy figure and almost none publish the spread behind it. A median tells you about the customer in the middle; the interquartile range tells you what happens if you are not that customer, and it is the number worth reading.

What would you automate?

Tell us your highest-volume workflow. We will tell you the realistic band before you commit.

1 / 3
n = 23 customers, 2025–2026Spread published, not just medianSmall sample, stated
94%PO matching, median straight-through at week 20
66%Multi-line allocation — our weakest published workflow
26ptWidest interquartile spread in the set
23Customers contributing data, all opted in
WorkflowWeek 4 → week 20Wk 20
AP · bill codingRecurring vendors dominate; climbs fastest
91%
AP · PO matchingStructured comparison, little ambiguity
94%
Bank reconciliationExact and split only; fuzzy always waits
93%
Expense codingReceipt quality is the binding constraint
84%
Vendor deduplicationGenuinely hard; we score ourselves low
69%
Multi-line allocationRequires context often not on the document
66%
Contract obligationsProposes only; judgement stays human
61%
Median across 23 customers, 2025–2026. Shaded band is the interquartile spread — the bottom quartile of customers sits well below these figures.

Straight-through rate from week 4 to week 20. Shaded band is the interquartile spread — a quarter of customers sit below its lower edge.

What we found

Six things the data says that the marketing usually does not.

These are the patterns that recur across every customer, and they are more useful for planning than any single headline figure.

The curve is steep then flat

Most of the gain arrives between week four and week twelve. After week twenty the rate moves by one or two points a quarter, which means the twenty-week figure is close to the ceiling rather than a waypoint.

The spread is the real story

Bill coding runs 91% at the median and 78% at the lower quartile. Planning on the median when your data resembles the lower quartile is how automation business cases fail.

Structure beats volume

PO matching reaches 94% because the comparison is structured. Multi-line allocation sits at 66% at ten times the volume, because the context needed is often not on the document.

Receipt quality is a hard ceiling

Expense coding tops out at 84% and the binding constraint is photograph quality, not the model. Customers who fixed capture gained more than any tuning delivered.

Vendor concentration predicts the rate

Customers whose top 50 vendors cover 80% of bills automate faster and higher. A long tail of one-off suppliers is the single strongest negative predictor we see.

Deduplication is genuinely hard

We score ourselves at 69% on vendor deduplication and publish it because the alternative is implying a solved problem. Two subsidiaries with different names and addresses should not match automatically.

Why we publish the bad numbers

A benchmark that only shows the median is a benchmark designed to be quoted. It sets an expectation half of customers will not meet, and the half that do not meet it conclude the product underperformed rather than that they were told the wrong number.

Publishing the interquartile spread costs us deals where a competitor claims a single higher figure with nothing behind it. It gains us customers whose business case survives contact with their own data, and that is the trade we would rather make.

A median with no spread is a number chosen to be quoted. Ask any vendor for the bottom quartile and watch what happens.

What determines where you land

Four factors explain most of the variance we see, and all four are knowable before you buy. Vendor concentration: how much of your volume comes from suppliers you bill repeatedly. Document quality: whether invoices arrive as structured PDFs or as photographs. Policy clarity: whether coding rules are written down or reconstructed each time. And history depth: how many prior examples exist for the agent to learn the pattern from.

A customer strong on all four typically lands above the upper quartile within twelve weeks. A customer weak on two of them will sit below the median indefinitely, and no amount of model improvement changes that — the information required simply is not present.

What we deliberately do not count

Straight-through means the transaction completed without a person touching it and was not subsequently reversed. It does not include transactions a person approved quickly, because an approval click is still a person in the loop and counting it would inflate every figure here by ten to fifteen points.

We also exclude the first three weeks of any deployment. Early rates are dominated by configuration rather than capability, and including them would make the improvement curve look more impressive than it is.

Reversal rate matters more than accuracy

An agent operating at 91% with a 0.3% reversal rate is safe to grant authority to. One operating at 94% with a 3% reversal rate is not, because reversals are errors that got through and were found later — which means the ones that were not found are also out there.

Reversal rate across the set runs between 0.1% and 0.6% by workflow. It is the figure we watch most closely and the one that triggers an authority reduction when it moves.

How we measured this

Twenty-three customers running at least one agent workflow continuously between January 2025 and June 2026, all of whom opted in to aggregate reporting. Figures are computed from production data rather than reported by customers.

Straight-through is defined as: the transaction completed with no human interaction and was not reversed within 90 days. Weeks are counted from the date authority was first granted for that workflow, not from contract start.

The sample is small and skews toward services businesses and software companies, which is what our customer base looks like. It under-represents distribution and manufacturing, and we would expect different results there. We state that rather than extrapolating.

No customer is identifiable in the aggregates, and any workflow with fewer than five contributing customers is excluded rather than published on a thin base.

Questions

Common follow-ups.

Why is the sample so small?
Because we are early and would rather publish 23 real customers than a large number we cannot substantiate. It is stated at the top of the page rather than in a footnote.
Do these numbers apply to us?
Imperfectly. The four factors that drive variance are knowable in advance, and the only benchmark that predicts your result is one run on your own historical data. We offer that before purchase.
Why exclude approval clicks?
Because a person clicking approve is still a person in the loop. Counting it would add ten to fifteen points to every figure here and would be misleading.
What is a reversal?
An automated action later undone. It is the figure we watch most closely, because reversals found imply errors not yet found.
Will these numbers improve?
The model side, yes, gradually. The data-quality side is yours rather than ours, and for several workflows it is the binding constraint.

Get the number for your data.

A few hundred historical transactions with known outcomes is enough to tell you where you would land.