← Insight
Architecture10 min readUpdated 2026-09-18

RAG for Financial Services Leaders

Retrieval is how a model stops guessing about your product terms. It is also a data-governance programme, a refresh pipeline and a permanent cost line.

Written for Head of AI, Head of Data, COO

What you should take away

  • Retrieval does not teach the model your business; it supplies the relevant passages at the moment of answering.
  • In regulated answers, citation is the point: an answer a supervisor cannot trace is not usable.
  • Retrieval quality, not model quality, is usually what limits autonomous resolution in banking.
  • Retrieval adds a largely fixed cost line plus a variable context cost on every single call.

A capable model knows a great deal about the world and nothing about your fee schedule, your forbearance policy, your product variants withdrawn in 2019 but still held by 40,000 clients, or the wording your complaints team is required to use. Retrieval-augmented generation is the mechanism that closes that gap: at question time, the system finds the relevant passages from your own documents and gives them to the model alongside the question.

For a financial institution, the important consequence is not accuracy in the abstract. It is traceability. A retrieval architecture produces answers with sources attached, which means a supervisor, an auditor or a complaints handler can reconstruct why a client was told what they were told.

What it costs, structurally

ComponentNatureCost driver
Ingestion and chunkingProject then recurringDocument count, format variety, PDF quality
EmbeddingPer document, per re-indexCorpus size × refresh frequency
Vector storage and searchFixed with stepsIndex size, replicas, latency target
Refresh pipelineRecurring engineeringHow often terms, rates and policies change
Access control at retrievalRecurring engineeringEntitlement model complexity
Context tokens at answer timeVariablePassages per answer × calls per interaction

Structure is general; values depend on your corpus and latency requirements.

Retrieval quality is the resolution-rate lever

When an agent fails to resolve a banking query, the cause is usually not reasoning. It is that the right passage was never retrieved, or that three contradictory versions of a policy exist and the system had no basis to prefer the current one. That is a content and governance problem wearing an AI costume, and it is why document estate hygiene often returns more than a model upgrade.

Interactive model

What retrieval investment has to buy

Retrieval spend is justified by the resolution rate it unlocks. Raise the retrieval and data line, then raise resolution, and see where the trade turns positive.

Total cost per successful outcome

£1.20

AI-attempted workflow cost

£6.03m

Human escalation cost

£4.48m

Modelled annual operating difference

£16.37m

Retrieval & data infrastructure£180k/year
Illustrative assumption
Autonomous resolution rate75%
Illustrative assumption

Where the cost sits

  • Human escalation£4.48m74%
    Calculated
  • Implementation, annualised£500k8%
    Calculated
  • Model inference£420k7%
    Calculated
  • Retrieval & data£180k3%
    Calculated
  • Failures & retries£180k3%
    Calculated
  • Platform, evaluation & monitoring£150k2%
    Calculated
  • Tools & APIs£120k2%
    Calculated

Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.

Illustrative assumptions. A retrieval programme that lifts resolution by a few points usually pays for itself many times over; one that does not is data tidying with a licence fee.

Questions to ask before funding a retrieval programme

  1. Which documents are authoritative, and who owns the answer when two disagree?
  2. How quickly must a rate, fee or policy change appear in answers, same day, or next release?
  3. How are withdrawn products and historic terms handled for clients who still hold them?
  4. What entitlement model governs retrieval, and how is it tested?
  5. How will we evidence, a year later, which version of a document produced a given answer?

If those five have owners and answers, retrieval is an engineering exercise. If they do not, the retrieval programme will discover them for you, slowly, at production cost.

Apply this to your own workload

The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.