Customers

What actually changed

Every figure here is measured from production data rather than reported in a testimonial, and every one shows where it sits in the distribution. A median with no spread sets an expectation half of customers will not meet.

What would this look like for you?

Send a few figures and we will tell you which of these you would plausibly see.

1 / 3
Measured, not reportedDistribution shownIncluding where it did not work

Customers

Six outcomes, with their spread.

Median across customers who ran the relevant workflow for at least two quarters, with the interquartile range in brackets.

Close length

Median fell from 13 to 7 working days (IQR 5–9). The gain concentrates in customers whose bank reconciliation was on the critical path — for those it was the single change that moved everything downstream.

Bills coded without a person

91% at week 20 (IQR 78–95%). Vendor concentration is the strongest predictor: customers whose top 50 vendors cover most volume land above the upper quartile within twelve weeks.

Bank lines matched automatically

93% (IQR 86–96%). The remaining 7% is judgement — short payments, unexplained deductions, batched wires against many invoices — and it stays with a person by design.

Days sales outstanding

Median improvement of 9 days (IQR 4–14). Almost entirely driven by earlier invoicing and by collections being worked systematically rather than by customers paying faster.

Time to consolidate a group

Median fell from 5 days to under 1 (IQR 0.5–2). This is the outcome multi-entity customers most consistently report as the reason the engagement paid for itself.

Reconciliation coverage

From a median of 3 untied balance sheet accounts to 0. Unglamorous, and the finding auditors respond to most directly.

Read the spread, not the median

The interquartile ranges above are wide, and that is the honest part. A customer at the lower quartile on bill coding is at 78% rather than 91% — still useful, and a materially different business case.

Four factors explain most of the variance, and all four are knowable before you buy. Vendor concentration, document quality, whether coding rules are written down, and how much history exists for the pattern to be learned from.

A customer strong on all four typically lands above the upper quartile within twelve weeks. A customer weak on two of them sits below the median indefinitely, and no amount of model improvement changes that, because the information required is not present in their systems.

A median describes the customer in the middle. The spread describes what happens if you are not that customer, and it is the number a business case should be built on.

Where it did not work

Two patterns account for nearly every disappointing outcome we have had, and both were visible in advance.

Data quality below the threshold

One customer’s expense coding plateaued in the low sixties because receipts arrived as photographs of crumpled paper, frequently illegible. That is not a model problem and no tuning fixed it. The honest fix was changing how receipts were captured, which was a policy decision rather than a software one.

Rules that were contested rather than undocumented

Another customer’s allocation workflow escalated constantly because four people in the business described the allocation rule differently. Automating it would have encoded one person’s version permanently, so we stopped and told them the governance decision came first.

We now screen for both during scoping and will decline the work rather than take it and underperform. That costs us engagements and it costs us fewer than the alternative does.

How these are measured

From production data, not from customer self-reporting. Straight-through means the transaction completed with no human interaction and was not reversed within 90 days — an approval click does not count, which is why these figures are ten to fifteen points below what a looser definition would produce.

Before figures are measured during the first month rather than recalled. Where a customer could not supply a reliable baseline, they are excluded from that metric rather than estimated.

The sample skews toward services businesses and software companies, which is what our customer base looks like. It under-represents distribution and manufacturing and we would expect different results there.

The only benchmark that predicts your result

Is one run on your own data. Send several hundred historical transactions with known outcomes and we will score against them and show you where it fails, before you commit to anything. A vendor unwilling to be measured that way before a purchase is telling you something.

Questions

Common follow-ups.

Are these figures typical?
They are medians with the interquartile range shown. A quarter of customers sit below the lower quartile figure, and planning against the median when your data resembles that quartile is how business cases fail.
Why exclude approval clicks?
Because a person clicking approve is still a person in the loop. Counting it would add ten to fifteen points to every figure here.
What predicts where we would land?
Vendor concentration, document quality, whether coding rules are written down, and history depth. All four are knowable before you buy.
Do you publish outcomes where it went badly?
Yes, on this page. Two patterns account for nearly all of them and we now screen for both during scoping.
Can we test this before buying?
Yes, and it is the only benchmark that predicts your result. Several hundred historical cases with known outcomes is enough.

Test it on your own transactions.

Our medians describe our customers. Only your data describes you, and we will score against it before you commit.