What you should take away
- Permissions and limits carry more risk weight than model selection.
- Evidence, inputs, retrieved sources, model version, action taken, is the control auditors and complaints handlers rely on.
- Evaluation sets are the model-risk equivalent of back-testing, and they are a cost line.
- Every genuine control reduces autonomous resolution and therefore appears in the economics. That is a feature.
Existing model risk frameworks in banks were built for models that produce a number: a PD, a VaR, an IFRS 9 provision. Generative systems produce language and take actions, which breaks several assumptions, there is no single correct output, behaviour changes when a prompt or an index changes, and the vendor may update the underlying model beneath you.
That does not mean the discipline is absent. It means the control surface moves from the model to the workflow around it.
The five questions that carry the risk
- What actions can this system take, and what is the maximum financial consequence of any single one? If the answer is unbounded, nothing else in the framework will save you.
- What is it allowed to say, and where do those statements come from? Retrieval with citation is the difference between an answerable complaint and an unanswerable one.
- How would we reconstruct a specific interaction eleven months later? Inputs, retrieved passages, model and prompt versions, tool calls, outcome.
- How do we know quality has not drifted? A maintained evaluation set, run on every change, with a threshold that blocks release.
- What happens when it is wrong at scale? Kill switch, fallback routing, client remediation path, and a named owner.
Controls have a price, and the price is legitimate
Mandatory handover for vulnerability indicators lowers autonomous resolution. Citation requirements increase context and inference cost. Retention of full interaction evidence increases storage and platform cost. Sampling consumes reviewer minutes. A model that shows an excellent business case with none of these present is describing a system you would not be allowed to run.
Interactive model
Price the control set
Controls typically reduce autonomous resolution and raise platform cost. Set both and see whether the case still stands, a stress test worth running before the risk committee does it for you.
Total cost per successful outcome
£1.20
AI-attempted workflow cost
£6.03m
Modelled annual operating difference
£16.37m
Escalated interactions per year
1,680,000
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative assumptions. A business case that only works with controls switched off is not a business case.
What to read, from the source
Primary supervisory and public material
- Bank of England & FCA: Artificial intelligence in UK financial services
Joint survey of AI use, governance and perceived risks across UK firms. Read the governance and third-party sections closely.
- Bank of England: Financial Stability in Focus: AI in the financial system
The Financial Policy Committee's view of system-level AI risks, including concentration and model-monoculture concerns.
- FCA: AI update and approach
How the FCA maps AI onto existing rules, including Consumer Duty and senior manager accountability.
- NIST AI Risk Management Framework
A control vocabulary many institutions have adopted to structure AI governance documentation.
Read them as a business leader, not as a technologist. The recurring theme is accountability: firms are expected to know what their systems do, to evidence it, and to remain answerable for outcomes regardless of which vendor supplied the model.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.