Comparison · updated August 2026
Every vendor in this category now has an AI story, and most of them are describing the same thing: a language model with read access to an existing application, summarising and suggesting. That is useful. It is also categorically different from software that can do bounded work on its own authority.
Send the AI section of a vendor proposal. We will tell you what it actually commits to.
At a glance
| AI features | erp.io | |
|---|---|---|
| What it does | Assists a person doing the work | Does bounded work and escalates the rest |
| Authority | None — a person performs every action | Granted per workflow, entity, and threshold |
| Audit record | The person’s action is logged | The agent’s reasoning, policy, and confidence are logged |
| Failure mode | A bad suggestion a person declines | A bounded action that is reversible and logged |
| Measurement | Usually none published | Accuracy, escalation, and reversal rates per workflow |
| Ceiling | Limited by how fast a person can review | Limited by the authority you grant |
| Where it breaks | Silently, when suggestions get rubber-stamped | Visibly, as a rising escalation rate |
| Data access | Whatever the application can see | Scoped as an actor with its own permissions |
| Governance | The person is the control | The policy engine is the control |
| Honest maturity | Shipping widely today | Early, including ours |
Assistive AI features are genuinely useful and there are cases where they are all you should want.
What can it do without a person clicking approve. What is recorded when it acts. What happens to a release that scores worse on your evaluation set. And what is the published accuracy per workflow, including the bottom quartile of customers. The range of answers to those four is where the category actually separates, and most proposals cannot answer the third.
We are early too, and it would be dishonest to present this as a mature category. Our benchmarks show workflows ranging from ninety-four percent straight-through down to sixty-one, and the bottom quartile of our customers sits well below the median. Those numbers are published because the spread is the honest part.
What we would defend is the architecture rather than the current accuracy. An authority model, a policy engine that every actor passes through, and an audit trail that records reasoning are the things that make the accuracy improvable safely. Retrofitting them onto a system where the model shares application credentials is considerably harder than building them first.
If you are evaluating this category, the most useful thing you can do is ask each vendor for the four answers above and compare them side by side. It separates the products faster than any demo.
Questions
What can it do unapproved, what is recorded, what blocks a bad release, and what are the published numbers.