Services

An agent for the work only you do

The standard agents cover the workflows every finance function shares. What they do not cover is the review your industry requires, the allocation your contracts specify, or the exception pattern that only appears in your data. Those are buildable, and the interesting part is how you know it works.

Which workflow?

Describe the judgement your team makes repeatedly. We will say whether an agent fits.

1 / 3
$15,000–$40,000 typicalEvaluation set before deploymentStarts at observe, always
Candidatenew prompt or modelYour corpusheld-out, labelledEval run2,400 documentsStraight-through ratemust not fallConfidently-wrong ratemust not riseShipsBlockedboth gates evaluated per customer — a change can ship for one tenant and be blocked for another

The situation

What building one actually involves

Most of the work is not the model. It is establishing what correct means and proving the thing meets it.

The decision, written down

What inputs matter, what the rules are, where judgement enters, and what the escalation cases look like. This is the longest phase and the one that determines whether the rest works.

An evaluation set first

Several hundred real historical cases with known correct outcomes, built before any agent is written. Without it there is no way to distinguish a working agent from a plausible one.

The agent itself

Built against your data model with access to the context it needs — contract, history, budget, policy — rather than to the single document in front of it.

Calibrated confidence

It has to know when it does not know. An agent that is confidently wrong ten percent of the time is worse than no agent, because the errors are the ones nobody checks.

Authority you grant

Deployed at observe, promoted only when the shadow record justifies it, and bounded by the same ceiling as every other agent — no payment release, no permission changes.

Measured after launch

Accuracy, escalation rate, and reversal rate reported monthly. An agent whose numbers deteriorate gets its authority lowered rather than defended.

The evaluation set is the project

Anybody can build something that produces plausible output. The question that matters is whether it is right, and answering it requires several hundred real cases with known correct answers — assembled before the agent exists, so the agent cannot be tuned to the test.

Building that set is genuinely tedious. It is also the only thing separating an agent you can grant authority to from one you have to check, and checking everything is the same cost as doing everything.

If you cannot say what correct looks like across three hundred historical cases, you cannot know whether the agent works. You can only know whether it sounds right.

What makes a good candidate

High volume, recurring, and judgement-based with a knowable correct answer. Reviewing purchase requests against a policy that has seventeen exceptions. Classifying claims by recovery likelihood. Checking submissions against a specification before they go out.

The pattern is a decision a competent person makes in under two minutes, several hundred times a month, using information the system already holds.

What does not work

Decisions with no verifiable correct answer, because there is nothing to evaluate against. Decisions made twice a month, where the build cost never returns. And anything where being wrong is catastrophic and unrecoverable — those we decline, because bounded authority is not a sufficient control when the bound itself is the risk.

Where accuracy honestly lands

Our published benchmarks are the right expectation. Structured comparison work reaches the low nineties. Contextual classification lands in the eighties. Anything requiring information not present in the system — because it lives in someone’s head or in an email — plateaus in the sixties regardless of how good the agent is.

We will tell you which band yours falls into during scoping. A workflow whose ceiling is sixty-five percent may still be worth automating, but you should decide that knowing the number rather than discovering it.

Where to start

How it runs

01

Scoping and feasibility

What the decision is, whether correct is knowable, and what accuracy band to expect. Free, and it sometimes concludes an agent is the wrong tool.

02

Evaluation set

Several hundred historical cases with known outcomes, assembled with your team. Tedious, unavoidable, and the reason the rest can be trusted.

03

Build and measure

Agent built, scored against the set, iterated until it clears the threshold you agreed rather than until it looks convincing in a demo.

04

Observe, then promote

Deployed watching only. Authority raised when the shadow record justifies it, and lowered instantly if the numbers move.

Questions

What people ask.

What does it cost?
$15,000 to $40,000 depending on complexity, most of which is the evaluation set rather than the agent. Hosting is included in your subscription.
How accurate will it be?
Structured comparison reaches the low nineties; contextual classification the eighties; anything needing information not in the system plateaus lower. We tell you the band during scoping.
Why does the evaluation set come first?
So the agent cannot be tuned to the test. An evaluation built after the agent measures the agent’s assumptions rather than its accuracy.
Can it approve or pay?
No. Custom agents sit under the same ceiling as standard ones: no payment release, no vendor banking, no permission grants, no period close.
What if it stops performing?
Accuracy, escalation, and reversal rates are reported monthly. Deterioration lowers authority automatically rather than triggering a defence of the agent.

Describe the decision your team repeats.

We will tell you whether an agent fits and what accuracy to realistically expect.