← Insight
AI economics10 min readUpdated 2026-09-18

AI Agents in Banking: From Tokens to Cost per Customer Outcome

A client's message is a hundred tokens. What the agent processes is several thousand. Here is how to get from a vendor's price list to a number your finance function can use.

Written for Head of Client Operations, Finance Business Partner

What you should take away

  • You pay for assembled context on every call, not for the client's question.
  • Agent loops multiply calls per interaction, the single most under-modelled consumption driver.
  • Cost per outcome requires a denominator: resolutions achieved, not conversations started.
  • Blended model routing is a cost-quality lever whose value depends on measured traffic mix, not on a vendor's claim.

"Why was my card payment declined?" is six words. A bank's agent answering it may process system instructions defining tone and goodwill limits, the prior conversation, retrieved fee and FX terms, the client's recent transactions returned by a core-banking tool, a fraud-flag lookup, and then generate a compliant reply. Every one of those is billable input, on every call in the loop.

One customer request

“My card payment was declined. Why?”

The customer message is tiny. But the AI system may need much more information to answer it.

Model context so far

100 tokens

  • The visible question, plus the surrounding turn structure.

Illustrative workload only. Actual token consumption depends on architecture, model, context and workflow design.

Don’t understand tokens, RAG or agents yet? Learn AI interactively →

Calls per interaction is the multiplier everyone forgets

A chatbot makes one call. An agent runs a loop: read, decide, call a tool, read the result, decide again. A transaction-dispute interaction may take four to eight model calls, each carrying the accumulated context. Consumption is therefore volume × context size × calls per interaction, and the third term is usually estimated at one.

Illustrative consumption per interaction (7,000-token context)

1 call, simple FAQ answer~7k tokens
3 calls, balance and transaction lookup~21k tokens
6 calls, dispute triage with tools~42k tokens

Illustrative arithmetic to show the shape of the multiplier. Real consumption depends on architecture, context management and caching.

From consumption to cost per outcome

Consumption gives you a model bill. It does not give you a decision. To decide, you need the denominator that your current operation already uses: outcomes actually achieved. In service, that is a resolved client issue, not a conversation held, not a deflection, not a session.

  • Interactions attempted: the share of the queue routed to the agent.
  • Autonomous resolutions: attempts closed without human involvement, and without a subsequent repeat contact or complaint.
  • Escalations: attempts handed to a human, which carry both the AI cost already incurred and the full human handling cost.
  • Repeat contact: the quiet killer of AI service economics, a 'resolution' the client comes back about was not one.

Model routing, honestly

Routing simple traffic to a small model and hard traffic to a capable one lowers blended cost. Whether it pays depends on what share of your traffic genuinely is simple, and on whether misrouting pushes interactions into escalation, where they cost far more than the model saving.

Interactive model

Does cheaper inference or better resolution move the case?

Same illustrative bank. Change the model bill and the resolution rate independently, and watch cost per successful resolution.

Total cost per successful outcome

£1.20

Model inference per successful outcome

£0.08

AI-attempted workflow cost

£6.03m

Successful outcomes per year

5,040,000

Model inference cost£420k/year
Illustrative assumption
Autonomous resolution rate75%
Illustrative assumption

Where the cost sits

  • Human escalation£4.48m74%
    Calculated
  • Implementation, annualised£500k8%
    Calculated
  • Model inference£420k7%
    Calculated
  • Retrieval & data£180k3%
    Calculated
  • Failures & retries£180k3%
    Calculated
  • Platform, evaluation & monitoring£150k2%
    Calculated
  • Tools & APIs£120k2%
    Calculated

Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.

Illustrative assumptions. Halving inference and improving resolution are not equivalent levers, the model shows you which one your case is actually sensitive to.

The three numbers to take to finance

  1. Cost per successful resolution today, for this queue, fully loaded.
  2. Modelled cost per successful autonomous resolution, with escalation included.
  3. The resolution rate at which the two are equal, your break-even condition, stated as a target the operation can be held to.

A business case expressed that way stops being a forecast and becomes a contract with reality: here is the number we must hit, here is how we will measure it, here is when we stop.

Apply this to your own workload

The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.