Fewer, better decisions
The goal is not to route everything to a person. It is to route the small number of things where a person adds judgement, so that approving means something.
AI governance
Every AI vendor promises a human in the loop. Almost none design for what happens when that human faces four hundred items a week — which is that they approve everything, the control becomes theatre, and the audit trail records consent that was never really given.
Send a quarter of bills and we will show you the review volume at each authority level.
The design
The goal is not to route everything to a person. It is to route the small number of things where a person adds judgement, so that approving means something.
Every item arrives with what the agent concluded, why, what it compared against, and its confidence. Approving without that is rubber-stamping with extra steps.
What is unusual about this item relative to the last forty like it, shown before the detail. A reviewer should not have to derive the exception themselves.
Rejecting is as fast as approving and always carries a reason, which becomes a labelled example. Where rejection is slower than approval, approval becomes the default.
Amount, department, vendor, and account thresholds read from your existing approval structure, with delegation for absence and escalation for stalling.
Who, when, on what basis, under which policy version, with the item as it stood at that moment. An approval you cannot reconstruct is not evidence.
The risk people worry about with AI in finance is the agent doing something wrong unsupervised. The risk that actually materialises is subtler: the agent does everything correctly, routes all of it to a person for approval, and that person — facing several hundred items a week that have been right every time — starts approving in batches without reading.
At that point the control has inverted. The audit trail records a human approval on every transaction, which looks excellent in a controls test, and no human judgement was applied to any of them. It is worse than no approval step, because it manufactures evidence of oversight that did not occur.
Approval rate, time spent per item, and rejection rate are reported per approver. If someone is approving ninety-nine percent of items in under three seconds each, that is surfaced — not as a performance criticism, but as a signal that the routing threshold is wrong and too much is reaching them.
The correct response is almost always to raise the policy’s confidence bar so fewer, genuinely uncertain items arrive. Counter-intuitively, routing less to a person usually produces more actual review.
A reviewer given a full invoice has to work out what is unusual about it. A reviewer given "this vendor is normally coded to 6420; this one is proposed as 6310, because the line description mentions consulting" is making a decision immediately.
Presenting the difference before the detail is a small interface decision with a large effect on whether review is real. The detail is one click away for the cases where it matters.
In many systems approving is one click and rejecting requires a reason, a routing choice, and a comment. That asymmetry has a predictable consequence: under time pressure, people approve.
Rejection here is a single action with a reason picked from a short list, and each rejection becomes a labelled example that improves the agent. The person doing the reviewing is, in effect, training the thing that will bother them less next month — which is the only incentive alignment that survives contact with a busy week.
Bulk approve. There is no select-all on a review queue, at any authority level. If a queue is large enough that bulk approval feels necessary, the routing threshold is wrong and the fix is upstream.
Questions
A quarter of bills is enough to model review volume at every authority level before you commit to one.