← Insight
Flagship18 min readUpdated 2026-09-18

The State of AI Economics in Financial Services

A recurring flagship report. Edition one: what public research establishes, what it does not, and modelled scenarios for the four workloads financial institutions are funding most.

Written for Board, executive committee, investment committee

What you should take away

  • Adoption is broad and measurement is thin: most firms report AI in production, far fewer report a unit cost per outcome.
  • Published research is strongest on adoption and governance, weakest on economics, which is precisely where capital decisions are made.
  • Across four modelled workloads, inference is never the largest cost line; human escalation or verification is.
  • Fixed cost lines mean unit economics are volume-dependent, so a pilot's numbers rarely justify or condemn a programme.

Part one, what public research actually establishes

There is now a substantial body of public material on AI in financial services from supervisors and international bodies. Read together, it supports three conclusions with reasonable confidence: adoption is widespread and accelerating; governance frameworks are being stretched rather than rebuilt; and third-party dependency is the risk supervisors return to most often.

It supports almost nothing about unit economics. Surveys ask what firms use AI for and how they govern it. They do not publish cost per resolved case, cost per research note, or measured autonomous resolution rates, because firms largely do not measure them in a comparable way, and would not disclose them if they did.

Sources for this edition, read them directly

Our reading of the gap: the industry has adoption evidence and governance guidance, and lacks a shared economic vocabulary. Firms cannot compare AI programmes to each other, or to the operations they replace, because there is no agreed unit. That is the gap this report exists to narrow, edition by edition.

Part two, four modelled workloads

Everything from here is a modelled scenario. The volumes, rates and cost lines are illustrative assumptions chosen to be plausible for a mid-to-large UK institution; they are not measurements, and they are not averages of anything. Every figure is computed by the same deterministic engine, and every slider is yours to move.

Workload 1, Retail client service

The most-funded workload, and the one where the economics are best understood. High volume, short interactions, a clear outcome definition, and an existing cost-to-serve baseline to compare against.

Interactive model

Retail client service, modelled scenario

1,000,000 conversations a month, 20% already self-served, 8-minute human handling at £25/hour fully loaded.

Current annual human cost

£32.00m

AI-attempted workflow cost

£6.03m

Total cost per successful outcome

£1.20

Modelled annual operating difference

£16.37m

Share of queue attempted70%
Illustrative assumption
Autonomous resolution rate75%
Illustrative assumption
Escalated handling time6.4 min
Illustrative assumption
Model inference cost£420k/year
Illustrative assumption

Where the cost sits

  • Human escalation£4.48m74%
    Calculated
  • Implementation, annualised£500k8%
    Calculated
  • Model inference£420k7%
    Calculated
  • Retrieval & data£180k3%
    Calculated
  • Failures & retries£180k3%
    Calculated
  • Platform, evaluation & monitoring£150k2%
    Calculated
  • Tools & APIs£120k2%
    Calculated

Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.

Modelled scenario, illustrative assumptions. Not a forecast, not a benchmark, not a guaranteed saving.

Workload 2, Onboarding and KYC review

Lower volume, longer handling, higher consequence. Automation here is document-heavy and control-heavy: the review step is mandatory for adverse findings, which caps the achievable autonomous rate by design rather than by capability.

Interactive model

Onboarding and KYC review, modelled scenario

45,000 cases a month, 35 minutes of analyst time each at £40/hour fully loaded.

Current annual human cost

£11.34m

AI-attempted workflow cost

£3.48m

Total cost per successful outcome

£17.35

Modelled annual operating difference

£5.03m

Cases cleared without human review55%
Illustrative assumption
Review minutes on referred cases18.0 min
Illustrative assumption
Screening & data provider calls£180k/year
Illustrative assumption

Where the cost sits

  • Human escalation£1.97m57%
    Calculated
  • Implementation, annualised£600k17%
    Calculated
  • Retrieval & data£260k7%
    Calculated
  • Model inference£220k6%
    Calculated
  • Tools & APIs£180k5%
    Calculated
  • Platform, evaluation & monitoring£160k5%
    Calculated
  • Failures & retries£90k3%
    Calculated

Largest modelled component: Human escalation (£1.97m). Outcome measured per case cleared without human review.

Modelled scenario. Third-party screening calls are a genuine per-case cost in this workload and are frequently omitted from AI business cases.

Workload 3, Investment research support

Modest volume, very large context per output, and a verification step that cannot be removed. This is the workload where inference cost is most visible as a share, and still not the largest line.

Interactive model

Investment research support, modelled scenario

8,000 requests a month, 45 minutes of analyst time each at £95/hour fully loaded.

Current annual human cost

£6.84m

AI-attempted workflow cost

£2.34m

Total cost per successful outcome

£50.69

Model inference per successful outcome

£5.64

Outputs accepted with light review60%
Illustrative assumption
Verification minutes on the rest25.0 min
Illustrative assumption
Model inference cost£260k/year
Illustrative assumption

Where the cost sits

  • Human escalation£1.22m52%
    Calculated
  • Implementation, annualised£350k15%
    Calculated
  • Model inference£260k11%
    Calculated
  • Retrieval & data£220k9%
    Calculated
  • Platform, evaluation & monitoring£140k6%
    Calculated
  • Tools & APIs£90k4%
    Calculated
  • Failures & retries£60k3%
    Calculated

Largest modelled component: Human escalation (£1.22m). Outcome measured per research output produced without deep rework.

Modelled scenario. Data licensing may prohibit parts of this workload entirely, a constraint that precedes any cost model.

Workload 4, Credit paper drafting

Low volume, high value per output, and a review step performed by expensive people. Here the case rarely rests on cost reduction at all; it rests on cycle time and consistency, with cost neutrality as the test.

Interactive model

Credit paper drafting, modelled scenario

900 papers a month, 4 hours of credit analyst time each at £70/hour fully loaded.

Current annual human cost

£3.02m

AI-attempted workflow cost

£1.20m

Total cost per successful outcome

£261.86

Modelled annual operating difference

£1.37m

Drafts requiring only light edit50%
Illustrative assumption
Rework minutes on the rest90.0 min
Illustrative assumption
Build, amortised annually£300k/year
Illustrative assumption

Where the cost sits

  • Human escalation£482k40%
    Calculated
  • Implementation, annualised£300k25%
    Calculated
  • Retrieval & data£140k12%
    Calculated
  • Platform, evaluation & monitoring£110k9%
    Calculated
  • Model inference£90k7%
    Calculated
  • Tools & APIs£40k3%
    Calculated
  • Failures & retries£40k3%
    Calculated

Largest modelled component: Human escalation (£482k). Outcome measured per paper drafted without substantial rework.

Modelled scenario. At this volume, amortised build cost is the dominant driver of unit cost, the clearest argument for buying rather than building at small scale.

What the four scenarios have in common

  1. Model inference is never the largest line. Human escalation or verification is, in every one of the four.
  2. Unit cost is a function of volume, because four of the seven cost lines are largely fixed.
  3. The single most powerful lever is the rate at which work completes without a human, and it is a design and data-quality outcome, not a model-selection outcome.
  4. Controls reduce that rate deliberately. Any case that looks excellent without them is describing an unrunnable system.
  5. The benefit is credible only where a cost actually leaves a budget, or a loss is actually avoided.

What we would like to publish in the next edition

  • Measured autonomous resolution rates by workload type, anonymised and ranged.
  • Measured tokens per interaction in production agent loops, which would replace the weakest assumption in every model above.
  • Observed escalation handling times compared with pre-AI handling times for the same queues.
  • Realised versus modelled cost, twelve months after go-live.

Until then, treat every number in part two as what it is: a transparent model, built from stated assumptions, designed to show you which assumptions your own decision depends on.

Apply this to your own workload

The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.