Straight-through rate
The share of documents that reach a posted transaction with zero human edit. Reported per workflow, per customer, per week. This is the headline number and the one we publish.
AI governance
Any vendor can quote an accuracy figure. Almost none will tell you what it was measured against, whose data it came from, or what happens when a release makes it worse. This page is our method, including the parts that are unflattering.
Tell us your setup and we will report straight-through and confidently-wrong rates against what your team actually did.
What we measure
A single accuracy percentage hides the distinction that matters most: the difference between an agent that asks for help and one that is wrong without asking.
The share of documents that reach a posted transaction with zero human edit. Reported per workflow, per customer, per week. This is the headline number and the one we publish.
Actions taken above the confidence threshold that a person later corrected. The most important number in the system, because these are the errors nobody reviews.
How wrong a wrong answer was. A bill coded to the neighbouring account is a different problem from one coded to a different entity, and averaging them hides both.
Of the items the agent held for a person, how many genuinely needed one. Escalating everything is trivially safe and useless; this keeps caution honest.
The gate
Every prompt and model change runs against a held-out set of your own historical documents before it reaches you. Two gates, both of which must pass.
Straight-through rate is the number customers ask about, and on its own it is a dangerous target. You can raise it trivially by lowering the confidence threshold — the agent stops asking for help, more documents post automatically, and the headline metric improves while the system gets materially worse.
That is why confidently-wrong rate is gated separately and why a release that improves the first while degrading the second is blocked rather than traded off. The two numbers move in opposite directions under exactly the optimisation pressure a vendor is under, which is precisely when a hard gate earns its keep.
From your team, as a by-product of working. Every correction a controller makes to a proposed coding, every rejected match, every reassigned department becomes a labelled example specific to your chart of accounts and your conventions. Nothing needs to be annotated deliberately, which is why the corpus grows rather than stalling after the initial enthusiasm.
A held-out slice is reserved and never used for retrieval or few-shot context, so evaluation is measured against documents the system has genuinely not been tuned on. This is unglamorous discipline and it is the difference between a real number and a number that has memorised its own test.
A change that improves accuracy for a distribution business with 400 vendors may degrade it for a consultancy with 30. Because gates run against each customer’s corpus, a release can ship to one tenant and be held for another, and we would rather carry that operational complexity than average across customers and call the mean a result.
Below is a representative curve from a professional-services customer processing around 350 bills a month. Weeks one and two are the agent learning your conventions while every proposal is reviewed. The steep section is the recurring vendor base being learned. The flattening after week sixteen is the long tail — one-off vendors, unusual allocations, and genuinely ambiguous documents that a person should be looking at anyway.
This is one real customer’s accounts-payable curve, not a median and not a promise. Volume, vendor concentration, and how consistently your team has coded historically all move it. A business with 40 recurring vendors climbs faster than one with 400 occasional ones.
Questions
Fifty documents and a week is enough to see real numbers on your own chart of accounts.