What you should take away
- Seven cost lines, two of which are volume-driven, four largely fixed, and one, escalation, driven by quality.
- Fixed lines make unit cost a function of volume: the same system is cheap at scale and expensive in a pilot.
- Year one and steady state are different businesses; model both.
- Report a range with named sensitivities, never a single number.
This is the reference template we apply in a BillingEngine assessment. It is deliberately boring: the value is in completeness and in stating how each line behaves, because behaviour, fixed, volume-driven or quality-driven, is what determines your unit cost as you scale.
The template
| Line | Behaviour | What to ask for |
|---|---|---|
| Model inference | Volume × context × calls per interaction | Measured tokens per interaction from a pilot, not a vendor estimate |
| Retrieval & data | Fixed, steps with corpus and refresh | Index size, refresh cadence, pipeline engineering days per quarter |
| Tools & integration | Per action taken | Which core systems, which per-call fees, which middleware |
| Platform, eval & monitoring | Fixed plus small per-interaction | Licence, eval runs per release, trace retention period |
| Human escalation & review | Quality-driven | Resolution rate, handling minutes, QA sampling percentage |
| Failures, retries & remediation | Error-rate driven | Measured retry rate and average remediation cost |
| Implementation, amortised | Fixed | Total build cost and the amortisation period you will defend |
Illustrative annual cost stack, AI-attempted workflow
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Total modelled AI-attempted workflow cost £6.03m per year. Under the illustrative defaults, escalation dominates and inference is a minor band.
Why the same agent has two very different unit costs
Four of the seven lines barely move with volume. That means unit cost falls steeply as volume rises, and it means a pilot's unit economics tell you almost nothing about steady state. Move the volume slider below and watch cost per outcome change without a single assumption about model quality changing.
Interactive model
Unit cost as a function of scale
Same illustrative agent, different volumes. Fixed lines are unchanged; only throughput moves.
Total cost per successful outcome
£1.20
Model inference per successful outcome
£0.08
Successful outcomes per year
5,040,000
AI-attempted workflow cost
£6.03m
Where the cost sits
- Human escalationCalculated£4.48m74%Calculated
- Implementation, annualisedCalculated£500k8%Calculated
- Model inferenceCalculated£420k7%Calculated
- Retrieval & dataCalculated£180k3%Calculated
- Failures & retriesCalculated£180k3%Calculated
- Platform, evaluation & monitoringCalculated£150k2%Calculated
- Tools & APIsCalculated£120k2%Calculated
Largest modelled component: Human escalation (£4.48m). Outcome measured per successful autonomous resolution.
Illustrative assumptions. The lesson to carry into a vendor conversation: any per-outcome price is meaningless without the volume it assumes.
Year one versus steady state
- Year one carries build, integration, model risk documentation, parallel running and a deliberately low attempt rate.
- Steady state carries a permanent operating team, periodic revalidation, and higher attempt rates on a broader queue.
- Presenting only steady state overstates the case; presenting only year one kills programmes that would have paid.
- Show both, and show the year in which cumulative position turns positive.
If you want this applied to your own workload rather than an illustrative one, the BillingEngine assessment walks the same template question by question and lets you answer "I don't know" wherever you genuinely do not.
Apply this to your own workload
The figures above are illustrative assumptions. A BillingEngine assessment replaces them with your company's numbers and shows which assumptions decide the answer.