What you should take away
- Adoption is broad and measurement is thin: most firms report AI in production, far fewer report a unit cost per outcome.
- Published research is strongest on adoption and governance, weakest on economics, which is precisely where capital decisions are made.
- Across four modelled workloads, inference is never the largest cost line; human escalation or verification is.
- Fixed cost lines mean unit economics are volume-dependent, so a pilot's numbers rarely justify or condemn a programme.
Part one, what public research actually establishes
There is now a substantial body of public material on AI in financial services from supervisors and international bodies. Read together, it supports three conclusions with reasonable confidence: adoption is widespread and accelerating; governance frameworks are being stretched rather than rebuilt; and third-party dependency is the risk supervisors return to most often.
It supports almost nothing about unit economics. Surveys ask what firms use AI for and how they govern it. They do not publish cost per resolved case, cost per research note, or measured autonomous resolution rates, because firms largely do not measure them in a comparable way, and would not disclose them if they did.
Sources for this edition, read them directly
- Bank of England & FCA: Artificial intelligence in UK financial services
The primary UK adoption and governance survey. Useful for use-case distribution and third-party reliance; not a source of cost data.
- Bank of England: Financial Stability in Focus: AI in the financial system
Financial Policy Committee analysis of system-level risks, including concentration in model and infrastructure providers.
- FCA: AI update
How existing rules, Consumer Duty, SM&CR, operational resilience, are applied to AI. Sets the control expectations that shape cost.
- Financial Stability Board, The Financial Stability Implications of Artificial Intelligence
International view on adoption patterns, vendor concentration and data quality dependencies.
- BIS: Intelligent financial system: how AI is transforming finance
Analytical framing of AI's effect on financial intermediation and supervision.
- IMF Global Financial Stability Report, chapter on AI and capital markets
Market-structure effects, particularly relevant for trading and research workloads.
Our reading of the gap: the industry has adoption evidence and governance guidance, and lacks a shared economic vocabulary. Firms cannot compare AI programmes to each other, or to the operations they replace, because there is no agreed unit. That is the gap this report exists to narrow, edition by edition.
Part two, four modelled workloads
Everything from here is a modelled scenario. The volumes, rates and cost lines are illustrative assumptions chosen to be plausible for a mid-to-large UK institution; they are not measurements, and they are not averages of anything. Every figure is computed by the same deterministic engine, and every slider is yours to move.
Workload 1, Retail client service
The most-funded workload, and the one where the economics are best understood. High volume, short interactions, a clear outcome definition, and an existing cost-to-serve baseline to compare against.
Interactive model
Retail client service, modelled scenario
1,000,000 conversations a month, 20% already self-served, 8-minute human handling at £25/hour fully loaded.
Current annual human cost
£32.00m
AI-attempted workflow cost
£6.03m
Total cost per successful outcome
£1.20
Modelled annual operating difference
£16.37m
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Modelled scenario, illustrative assumptions. Not a forecast, not a benchmark, not a guaranteed saving.
Workload 2, Onboarding and KYC review
Lower volume, longer handling, higher consequence. Automation here is document-heavy and control-heavy: the review step is mandatory for adverse findings, which caps the achievable autonomous rate by design rather than by capability.
Interactive model
Onboarding and KYC review, modelled scenario
45,000 cases a month, 35 minutes of analyst time each at £40/hour fully loaded.
Current annual human cost
£11.34m
AI-attempted workflow cost
£3.48m
Total cost per successful outcome
£17.35
Modelled annual operating difference
£5.03m
Where the cost sits
- Human escalationCalculated£1.97m57%Calculated
- Implementation, annualisedCalculated£600k17%Calculated
- Retrieval & dataCalculated£260k7%Calculated
- Model inferenceCalculated£220k6%Calculated
- Tools & APIsCalculated£180k5%Calculated
- Platform, evaluation & monitoringCalculated£160k5%Calculated
- Failures & retriesCalculated£90k3%Calculated
Largest modelled component: Human escalation (£1.97m). Outcome measured per case cleared without human review.
Modelled scenario. Third-party screening calls are a genuine per-case cost in this workload and are frequently omitted from AI business cases.
Workload 3, Investment research support
Modest volume, very large context per output, and a verification step that cannot be removed. This is the workload where inference cost is most visible as a share, and still not the largest line.
Interactive model
Investment research support, modelled scenario
8,000 requests a month, 45 minutes of analyst time each at £95/hour fully loaded.
Current annual human cost
£6.84m
AI-attempted workflow cost
£2.34m
Total cost per successful outcome
£50.69
Model inference per successful outcome
£5.64
Where the cost sits
- Human escalationCalculated£1.22m52%Calculated
- Implementation, annualisedCalculated£350k15%Calculated
- Model inferenceCalculated£260k11%Calculated
- Retrieval & dataCalculated£220k9%Calculated
- Platform, evaluation & monitoringCalculated£140k6%Calculated
- Tools & APIsCalculated£90k4%Calculated
- Failures & retriesCalculated£60k3%Calculated
Largest modelled component: Human escalation (£1.22m). Outcome measured per research output produced without deep rework.
Modelled scenario. Data licensing may prohibit parts of this workload entirely, a constraint that precedes any cost model.
Workload 4, Credit paper drafting
Low volume, high value per output, and a review step performed by expensive people. Here the case rarely rests on cost reduction at all; it rests on cycle time and consistency, with cost neutrality as the test.
Interactive model
Credit paper drafting, modelled scenario
900 papers a month, 4 hours of credit analyst time each at £70/hour fully loaded.
Current annual human cost
£3.02m
AI-attempted workflow cost
£1.20m
Total cost per successful outcome
£261.86
Modelled annual operating difference
£1.37m
Where the cost sits
- Human escalationCalculated£482k40%Calculated
- Implementation, annualisedCalculated£300k25%Calculated
- Retrieval & dataCalculated£140k12%Calculated
- Platform, evaluation & monitoringCalculated£110k9%Calculated
- Model inferenceCalculated£90k7%Calculated
- Tools & APIsCalculated£40k3%Calculated
- Failures & retriesCalculated£40k3%Calculated
Largest modelled component: Human escalation (£482k). Outcome measured per paper drafted without substantial rework.
Modelled scenario. At this volume, amortised build cost is the dominant driver of unit cost, the clearest argument for buying rather than building at small scale.
What the four scenarios have in common
- Model inference is never the largest line. Human escalation or verification is, in every one of the four.
- Unit cost is a function of volume, because four of the seven cost lines are largely fixed.
- The single most powerful lever is the rate at which work completes without a human, and it is a design and data-quality outcome, not a model-selection outcome.
- Controls reduce that rate deliberately. Any case that looks excellent without them is describing an unrunnable system.
- The benefit is credible only where a cost actually leaves a budget, or a loss is actually avoided.
What we would like to publish in the next edition
- Measured autonomous resolution rates by workload type, anonymised and ranged.
- Measured tokens per interaction in production agent loops, which would replace the weakest assumption in every model above.
- Observed escalation handling times compared with pre-AI handling times for the same queues.
- Realised versus modelled cost, twelve months after go-live.
Until then, treat every number in part two as what it is: a transparent model, built from stated assumptions, designed to show you which assumptions your own decision depends on.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.