← Insight
Risk & controls9 min readUpdated 2026-09-18

Why Human Review Can Cost More Than the LLM

Add a reviewer to every AI output and you have not built automation. You have built a drafting tool with a labour cost attached, sometimes rightly.

Written for COO, Head of Controls, operations leadership

What you should take away

  • Review minutes per interaction are the most powerful cost variable in most regulated AI workloads.
  • Universal review caps the benefit at the reviewer's speed advantage, not at the model's.
  • Risk-tiered review, full review where consequence is high, sampling where it is low, is where the economics work.
  • Review is a control with a price; treat it as a design decision, not an afterthought.

A 2-minute human check at £25 an hour costs 83 pence. An inference call for the same interaction may cost a fraction of a penny. The review is therefore not a rounding error against the model bill, it is often two orders of magnitude larger than it, and it scales with volume just as remorselessly.

This is the single most common reason an approved AI business case fails to deliver. Not because the model underperformed, but because a control introduced late, every output reviewed before release, converted an automation programme into an assisted-drafting programme with a permanent labour line.

The arithmetic, plainly

Interactive model

What review minutes do to unit economics

Illustrative bank service queue. Change the review time applied to escalated and checked interactions and watch total cost per successful resolution against the model line.

Total cost per successful outcome

£1.20

Model inference per successful outcome

£0.08

Human escalation cost

£4.48m

AI-attempted workflow cost

£6.03m

Human review / escalation minutes6.4 min
Illustrative assumption
Share released without review75%
Illustrative assumption
Fully loaded reviewer cost£25.00/hour
Illustrative assumption

Where the cost sits

  • Human escalation£4.48m74%
    Calculated
  • Implementation, annualised£500k8%
    Calculated
  • Model inference£420k7%
    Calculated
  • Retrieval & data£180k3%
    Calculated
  • Failures & retries£180k3%
    Calculated
  • Platform, evaluation & monitoring£150k2%
    Calculated
  • Tools & APIs£120k2%
    Calculated

Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.

Illustrative assumptions, not benchmarks. The ratio between the two unit numbers is the point: review cost dwarfs inference cost in almost every regulated configuration.

Design review by consequence, not by uniform policy

Interaction typeConsequence of errorProportionate control
Balance, transaction or statement queryLow, self-evident to the clientAutomated checks plus QA sampling
Fee or interest explanationMedium, potential mis-statementRetrieval citation required, sampled review
Goodwill or redress offerHigh, financial and conductHard limit in the tool, human approval above threshold
Vulnerability indicators detectedHigh, conduct and regulatoryImmediate human handover by rule
Complaint expressionHigh, statutory timelinesRouted to complaints, never auto-closed

A control matrix like this is also a cost model: each row sets review minutes per interaction for its share of volume.

Three ways to reduce review cost without reducing control

  1. Constrain actions in the tools, not in the prompt: a goodwill limit enforced by the API cannot be talked around, which removes a whole class of review.
  2. Make outputs checkable: citations and structured fields let a reviewer verify in 20 seconds what prose takes two minutes to assess.
  3. Route by confidence: release high-confidence, low-consequence interactions automatically and concentrate human minutes where they change the outcome.

The goal is not less oversight. It is oversight priced deliberately, placed where consequence lives, and measured well enough that you can tell your board what it buys.

Apply this to your own workload

The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.