What you should take away
- Token pricing is one line in a stack of seven; in service workloads it is rarely the largest.
- Human escalation is normally the dominant cost, and it is driven by autonomous resolution rate, not by model choice.
- Regulated workloads carry cost lines consumer workloads do not: evidence retention, quality assurance sampling, model risk documentation.
- The decision metric is cost per successful outcome, compared with today's cost-to-serve for the same outcome.
Most AI business cases in financial services are built on the wrong unit of account. A vendor quotes a price per million tokens, someone multiplies it by an estimated message volume, and the resulting number is small enough that the programme gets approved without much argument. Nine months later, the run-rate bears no resemblance to the paper, and the programme is reopened at exactly the moment it needs credibility.
The gap is not a forecasting error in the token line. It is that six other cost lines were never in the model, and that the two assumptions which dominate the result, how much of the queue the agent attempts, and how much of that it finishes without a human, were stated as facts rather than as things to be validated.
The seven lines of an AI workflow
| Cost line | What it actually is | Behaviour |
|---|---|---|
| Model inference | Input and output tokens across every call in the loop | Scales with volume × context × calls per interaction |
| Retrieval and data infrastructure | Indexing, embedding, vector storage, refresh pipelines, data egress | Largely fixed, steps up with corpus and refresh frequency |
| Tools and system integration | Core banking, cards, CRM and payments calls; middleware; per-call fees | Scales with actions taken, not conversations held |
| Platform, evaluation and monitoring | Orchestration, eval harness, tracing, drift detection, logging retention | Fixed base plus a per-interaction observability cost |
| Human escalation and review | Agents, specialists, complaints handlers, QA sampling | Scales with unresolved interactions and with assurance obligations |
| Failures and retries | Repeated calls, reworked cases, remediation, goodwill | Scales with error rate; often invisible until measured |
| Implementation allocation | Build, integration, model risk work, change management, amortised | Fixed, but it belongs in cost per outcome |
Categories are structural. The values in the model below are illustrative, not benchmarks.
Illustrative annual cost stack, AI-attempted workflow
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Total modelled AI-attempted workflow cost £6.03m per year. Click any band to see its share.
Why escalation dominates
An interaction the agent fails to resolve costs twice: once in the attempt, and again in the human who finishes the job. In a regulated environment it can cost three times, because the failure that produces a complaint also produces an investigation, a response within a statutory window, and possibly redress.
That is why autonomous resolution rate, not model price, is the variable to argue about. Move the slider below and watch what a five-point change in resolution rate does, then compare it with halving the model bill.
Interactive model
Cost per successful resolution in a service workload
An illustrative UK retail bank: 1,000,000 client conversations a month, 20% already self-served, 8-minute average human handling at £25/hour fully loaded.
Total cost per successful outcome
£1.20
Model inference per successful outcome
£0.08
Human escalation cost
£4.48m
Modelled annual operating difference
£16.37m
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative assumptions only, not an industry benchmark. Arithmetic is deterministic and unit-tested; no language model is involved in the calculation.
The cost lines regulated workloads add
- Evidence and retention: every automated decision that affects a client needs a reconstructable record, including the retrieved sources and the model version.
- Quality assurance sampling: a percentage of AI-handled interactions reviewed by humans, permanently, not as a pilot-phase measure.
- Model risk governance: documentation, validation, periodic review and sign-off, cost that recurs on every material change.
- Vulnerability and complaints routing: detection and mandatory handover paths, which lower autonomous resolution by design.
- Change control: prompts, retrieval indexes and tool permissions are production changes and are governed like production changes.
None of these make the case worse than it looks, several strengthen it, because they are also costs of the current human operation. They make it different. A model that ignores them is not conservative; it is simply incomplete.
What to put in front of the committee
- Today's cost per successfully resolved interaction, for the specific queue you intend to automate.
- Modelled cost per successful autonomous resolution, with all seven lines present.
- The share of the queue the agent will attempt, and how you arrived at it.
- The resolution rate the case requires, and the evidence you have that it is achievable.
- The cost of the interactions the agent will not attempt, they do not disappear.
- Three scenarios, and the two or three assumptions the answer is most sensitive to.
That is the whole discipline. Not a forecast presented as certainty, but a model that states what must be true, and names the numbers you should go and measure before releasing capital.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.