Research · 2026 edition

The AI in ERP Report

What finance teams actually adopted rather than what they said they would. Drawn from 23 live deployments and 41 rescue engagements between 2024 and 2026 — a small sample, stated as such, from a vendor with an obvious interest in the conclusions.

Get the full report

The written edition with methodology, per-workflow data, and the anonymised deployment profiles.

1 / 3
n = 23 deployments, 41 rescues2024–2026Vendor disclosure, not research

Six findings

The gap between intention and adoption.

Adoption stops at Level 1

Two thirds of automated workflows remain at draft-and-review a year in. Not because accuracy is inadequate, but because nobody with authority signed off on a policy having consequences.

AP is the only universal starting point

Every deployment we observed began with accounts payable, regardless of industry. It is half the manual volume and the fastest to a high straight-through rate.

Approval fatigue arrives at week nine

Approval rates climb toward 99% and time-per-item falls below four seconds within roughly two months, which is the point at which review has become a formality.

Governance questions arrive late

Security and audit review typically starts after selection rather than during. Where it starts during, the deployment is measurably smoother.

The chart of accounts is the binding constraint

Deployments on charts above 400 accounts took roughly twice as long to reach a stable straight-through rate as those under 250.

Nobody reduced headcount

Of 23 deployments, none reduced finance headcount in the first year. All redeployed capacity into close acceleration, collections, and analysis that was previously never done.

Adoption

Where deployments actually sit after a year.

The distribution is heavily weighted to Level 3 with low authority — agents initiating work and drafting, humans committing nearly all of it.

Software initiates, person governs

A bill arrives; an agent codes it and routes it.

The work starts without a person. Requires an authority model, an audit trail, and a policy engine — which is why it cannot be retrofitted onto a screen-driven system.

The distinction that matters is not how clever the model is. It is who is responsible for noticing the work exists.

Accuracy

Straight-through rate by workflow.

WorkflowWeek 4 → week 20Wk 20
AP · bill codingRecurring vendors dominate; climbs fastest
91%
AP · PO matchingStructured comparison, little ambiguity
94%
Bank reconciliationExact and split only; fuzzy always waits
93%
Expense codingReceipt quality is the binding constraint
84%
Vendor deduplicationGenuinely hard; we score ourselves low
69%
Multi-line allocationRequires context often not on the document
66%
Contract obligationsProposes only; judgement stays human
61%
Median across 23 customers, 2025–2026. Shaded band is the interquartile spread — the bottom quartile of customers sits well below these figures.

Failure

And why ERP projects stall generally.

From the 41 rescue engagements. Automation does not change this distribution — it inherits it, which is why readiness matters more than capability.

Primary causeShare of stalled projects
Requirements never settledScope kept moving because nobody could decide
31%
Data was worse than anyone knewDiscovered during load, not during diagnostic
24%
No internal owner with authorityDecisions escalated and stalled
19%
Customisation replaced process changeEvery gap closed with a script
14%
Partner capacity or turnoverTeam rotated mid-project
8%
The software genuinely could not do itRare, and usually knowable in week one
4%
n = 41 stalled or failed deployments we were brought into between 2024 and 2026. Small sample, stated deliberately.

The finding we did not expect

Nobody reduced headcount. Across 23 deployments, in the first year, not one finance team got smaller — despite most business cases including a headcount assumption and despite straight-through rates reaching the high eighties on the largest workflow.

What happened instead was consistent: close duration fell, collections finally received attention, and analysis that had been permanently deferred started happening. In four cases the team grew, because the company grew and finance stopped being the constraint.

Every business case assumed headcount reduction. None of the twenty-three delivered it, and all twenty-three considered the deployment a success.

We do not think this means the savings are illusory. It means the capacity is real and gets spent on work that was previously not being done — which is a better outcome and a harder one to put in a spreadsheet. If your business case depends on removing people, the honest thing is to say that is a management decision rather than an automatic consequence.

The Level 1 ceiling

The most actionable finding is that two thirds of workflows remain at draft-and-review a year in, and the cause is almost never accuracy. In interviews the recurring reason was that nobody was willing to be the person who authorised a policy that would post transactions automatically.

Where deployments did move to Level 2, one factor was present in almost every case: a simulation showing what the policy would have done against historical data. That is a narrow and slightly self-serving finding — we build that feature — and it was strong enough in the interviews that omitting it would be dishonest.

Method and limitations

  • Population. 23 deployments live at least twenty weeks, plus 41 rescue engagements. All are our own customers or prospects.
  • Period. January 2024 to June 2026.
  • Segment. US companies between $5M and $250M revenue, weighted toward services. Manufacturing and distribution are underrepresented.
  • Method. Product telemetry for accuracy and adoption; structured interviews for the qualitative findings.
Read this as vendor disclosure

This is not independent research. It is our customers, on our product, analysed by us, and we chose what to publish. The mitigations are that the sample size is stated everywhere, the workflows we perform worst on are included, and the findings that are inconvenient — no headcount reduction, adoption stalling at Level 1 — are the ones we led with.

Questions

About this report.

Is this independent research?
No. It is our customers, our product, our analysis, and our choice of what to publish. We state the sample everywhere and led with the findings that are inconvenient for us, which is the most we can offer short of commissioning someone else.
Can we cite it?
Yes, with attribution and a link, and please carry the sample size with the figure. A number from 23 deployments quoted as though it came from 2,300 is how research decays.
Will you publish it annually?
That is the intent, with the population growing. If a finding reverses next year it will be published as a reversal rather than quietly dropped.
Why so few manufacturing deployments?
Because we do not build manufacturing functionality, so we do not have those customers. It is a real limitation of the sample and the reason we would not generalise these findings to that segment.
Can we get the underlying data?
Aggregated and anonymised, yes, under a permissive licence. Customer-level data is not shared under any circumstances.

Get the written edition.

Full methodology, per-workflow data, and anonymised deployment profiles.