What you should take away
- You pay for assembled context on every call, not for the client's question.
- Agent loops multiply calls per interaction, the single most under-modelled consumption driver.
- Cost per outcome requires a denominator: resolutions achieved, not conversations started.
- Blended model routing is a cost-quality lever whose value depends on measured traffic mix, not on a vendor's claim.
"Why was my card payment declined?" is six words. A bank's agent answering it may process system instructions defining tone and goodwill limits, the prior conversation, retrieved fee and FX terms, the client's recent transactions returned by a core-banking tool, a fraud-flag lookup, and then generate a compliant reply. Every one of those is billable input, on every call in the loop.
One customer request
“My card payment was declined. Why?”
The customer message is tiny. But the AI system may need much more information to answer it.
Model context so far
100 tokens
The visible question, plus the surrounding turn structure.
Illustrative workload only. Actual token consumption depends on architecture, model, context and workflow design.
Don’t understand tokens, RAG or agents yet? Learn AI interactively →Calls per interaction is the multiplier everyone forgets
A chatbot makes one call. An agent runs a loop: read, decide, call a tool, read the result, decide again. A transaction-dispute interaction may take four to eight model calls, each carrying the accumulated context. Consumption is therefore volume × context size × calls per interaction, and the third term is usually estimated at one.
Illustrative consumption per interaction (7,000-token context)
Illustrative arithmetic to show the shape of the multiplier. Real consumption depends on architecture, context management and caching.
From consumption to cost per outcome
Consumption gives you a model bill. It does not give you a decision. To decide, you need the denominator that your current operation already uses: outcomes actually achieved. In service, that is a resolved client issue, not a conversation held, not a deflection, not a session.
- Interactions attempted: the share of the queue routed to the agent.
- Autonomous resolutions: attempts closed without human involvement, and without a subsequent repeat contact or complaint.
- Escalations: attempts handed to a human, which carry both the AI cost already incurred and the full human handling cost.
- Repeat contact: the quiet killer of AI service economics, a 'resolution' the client comes back about was not one.
Model routing, honestly
Routing simple traffic to a small model and hard traffic to a capable one lowers blended cost. Whether it pays depends on what share of your traffic genuinely is simple, and on whether misrouting pushes interactions into escalation, where they cost far more than the model saving.
Interactive model
Does cheaper inference or better resolution move the case?
Same illustrative bank. Change the model bill and the resolution rate independently, and watch cost per successful resolution.
Total cost per successful outcome
£1.20
Model inference per successful outcome
£0.08
AI-attempted workflow cost
£6.03m
Successful outcomes per year
5,040,000
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative assumptions. Halving inference and improving resolution are not equivalent levers, the model shows you which one your case is actually sensitive to.
The three numbers to take to finance
- Cost per successful resolution today, for this queue, fully loaded.
- Modelled cost per successful autonomous resolution, with escalation included.
- The resolution rate at which the two are equal, your break-even condition, stated as a target the operation can be held to.
A business case expressed that way stops being a forecast and becomes a contract with reality: here is the number we must hit, here is how we will measure it, here is when we stop.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.