What you should take away
- Review minutes per interaction are the most powerful cost variable in most regulated AI workloads.
- Universal review caps the benefit at the reviewer's speed advantage, not at the model's.
- Risk-tiered review, full review where consequence is high, sampling where it is low, is where the economics work.
- Review is a control with a price; treat it as a design decision, not an afterthought.
A 2-minute human check at £25 an hour costs 83 pence. An inference call for the same interaction may cost a fraction of a penny. The review is therefore not a rounding error against the model bill, it is often two orders of magnitude larger than it, and it scales with volume just as remorselessly.
This is the single most common reason an approved AI business case fails to deliver. Not because the model underperformed, but because a control introduced late, every output reviewed before release, converted an automation programme into an assisted-drafting programme with a permanent labour line.
The arithmetic, plainly
Interactive model
What review minutes do to unit economics
Illustrative bank service queue. Change the review time applied to escalated and checked interactions and watch total cost per successful resolution against the model line.
Total cost per successful outcome
£1.20
Model inference per successful outcome
£0.08
Human escalation cost
£4.48m
AI-attempted workflow cost
£6.03m
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative assumptions, not benchmarks. The ratio between the two unit numbers is the point: review cost dwarfs inference cost in almost every regulated configuration.
Design review by consequence, not by uniform policy
| Interaction type | Consequence of error | Proportionate control |
|---|---|---|
| Balance, transaction or statement query | Low, self-evident to the client | Automated checks plus QA sampling |
| Fee or interest explanation | Medium, potential mis-statement | Retrieval citation required, sampled review |
| Goodwill or redress offer | High, financial and conduct | Hard limit in the tool, human approval above threshold |
| Vulnerability indicators detected | High, conduct and regulatory | Immediate human handover by rule |
| Complaint expression | High, statutory timelines | Routed to complaints, never auto-closed |
A control matrix like this is also a cost model: each row sets review minutes per interaction for its share of volume.
Three ways to reduce review cost without reducing control
- Constrain actions in the tools, not in the prompt: a goodwill limit enforced by the API cannot be talked around, which removes a whole class of review.
- Make outputs checkable: citations and structured fields let a reviewer verify in 20 seconds what prose takes two minutes to assess.
- Route by confidence: release high-confidence, low-consequence interactions automatically and concentrate human minutes where they change the outcome.
The goal is not less oversight. It is oversight priced deliberately, placed where consequence lives, and measured well enough that you can tell your board what it buys.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.