What you should take away
- Scope the case to one queue with a measurable outcome, not to a capability.
- Benefit is cost-to-serve avoided plus loss avoided, not FTE 'capacity released' unless the capacity is actually removed or redeployed.
- Every material assumption carries an evidence type; the committee should see which ones are guesses.
- Stage the capital release against measured resolution rate, not against delivery milestones.
The weakest AI papers in banking share a shape: a capability ('an AI agent for client service'), a benefit expressed as FTE equivalents, a single-point cost, and a nine-figure market statistic in the appendix. They get approved, and they get reopened.
The strongest ones look like credit papers. A defined exposure, a stated base case, named assumptions with sources, a sensitivity table, and conditions under which the facility is extended or withdrawn.
1. Scope to a queue, not a capability
"AI in client service" cannot be modelled. "Disputed card transactions under £100 in the retail contact centre" can: it has a volume, an average handling time, an owner, a quality measure, and a regulatory perimeter. Pick the queue where the outcome is already counted, because that is where the before-and-after comparison is defensible.
2. Establish today's economics first
| Input | Where it comes from | Evidence type |
|---|---|---|
| Queue volume | Contact centre reporting | Measured |
| Average handling time | Workforce management system | Measured |
| Fully loaded hourly cost | Finance, including supervision, premises, attrition | Measured |
| Repeat contact rate | CRM linkage over a defined window | Measured, often missing |
| Complaint and redress rate | Complaints system | Measured |
| Quality assurance sampling cost | Operations budget | Measured |
If three of these are unavailable, that is your first finding, and a cheaper piece of work than the AI programme.
3. Model the whole proposed workflow
The modelled cost is not the AI bill. It is the AI-attempted workflow cost, inference, retrieval, tools, platform and monitoring, escalation, retries, amortised implementation, plus the cost of everything the agent will not attempt, which stays exactly where it is.
Interactive model
Base case, staged by ambition
Set the share of the queue the agent attempts and the resolution rate you believe is achievable in year one. The model shows the before-and-after operating position.
Current annual human cost
£32.00m
AI-attempted workflow cost
£6.03m
Modelled annual operating difference
£16.37m
Total cost per successful outcome
£1.20
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative defaults. A scenario output is not a forecast, and an operating difference is not a saving until the cost is actually removed from a budget.
4. Be honest about benefit realisation
- Cost avoided is real only where headcount, overtime, outsourcer volume or premises actually reduce, with a named budget line and a date.
- Capacity released is a benefit only if the released capacity is redeployed to work with a stated value.
- Loss avoided (complaints, redress, fraud losses, breach exposure) is often the larger and better-evidenced benefit in regulated queues.
- Revenue benefit from faster service should be modelled separately, with weaker confidence stated plainly.
5. Show the sensitivity, not just the answer
Which assumption should the committee interrogate?
Relative shape under the illustrative defaults used across BillingEngine. Your ranking will differ, the point is to compute it rather than assume it.
6. Stage capital against measurement
- Tranche one funds instrumentation: resolution definition, repeat-contact measurement, eval set, baseline cost per resolution.
- Tranche two funds a limited live deployment on one queue, with the resolution-rate threshold written into the approval.
- Tranche three funds scale, and only after the measured rate holds for a defined period at production volume.
- Each tranche has a stated stop condition. An AI programme without one is not a business case; it is an aspiration with a budget code.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.