The visible cost is the easy one

Per-call model pricing is published, predictable, and usually modest at business volumes. Budgeting from it alone is why so many of these projects come in over — the real costs are staff time and they land in someone else’s budget line.

That displacement is the whole problem. The project is approved against a software budget and paid for out of an operations team’s week, which means nobody sees the total in one place and the overrun is invisible until someone complains about workload.

1. Review time, which never reaches zero

Someone checks the output. That cost falls as trust grows but never disappears, and if reviewing takes nearly as long as doing the work then you automated the wrong step.

Model the shape rather than a single figure: near-total review in the first weeks, falling as corrections become rare, settling at a sampling rate that never reaches zero and rises again after any change to the model, the prompt or the knowledge. Budget the settled rate as a permanent line.

2. Knowledge maintenance

Business rules change, and each change has to reach the system. Budget recurring hours for whoever owns the knowledge; without them the thing decays into confidently repeating last quarter’s policy.

The load is proportional to how fast your business changes rather than to conversation volume. A stable service with settled pricing needs little; a company shipping product changes monthly needs someone genuinely on it, and that asymmetry is worth stating when the budget is set.

The model is rented. The maintenance is permanent.

3. The escalation path

Every escalation consumes a human, and escalations concentrate the hard cases. The average handling time of what reaches your team goes up even as volume goes down — a real cost that looks like a productivity regression if nobody expected it.

Warn the team before launch, and warn whoever reads their metrics. A support lead whose average handling time deteriorates without explanation will reasonably conclude the automation made things worse, when in fact it removed all the quick wins from the denominator.

4. Drift you don’t notice

Context grows, prompts accumulate, retries stack. The same quiet inflation that makes AI features slower and dearer after launch applies here, and only shows up if someone tracks cost per successful outcome rather than per call.

The two costs nobody puts in the model

Provider changes. A model deprecation or a behaviour shift means re-running the evaluation, possibly adjusting prompts, and raising the review rate for a fortnight. It is not frequent and it is not zero — budget it as an occasional certainty.

Incidents. At some point it will say something wrong to a customer, and the response consumes management attention, a correction, and a change to the evaluation set. Costing an AI employee as though that never happens is the same error as costing a hire as though they never make a mistake.

How to build the number

Take the four recurring lines as monthly hours at a loaded rate, add the model spend, and add an allowance for the two occasional costs. Then set it against the alternative you would otherwise fund — a hire, an outsourcer, a queue, or leaving the work undone.

Present it that way rather than as a technology cost. A board comparing a total against a named alternative can make a decision; a board looking at an API price with hidden staff time attached will approve it and be annoyed later — the framing that keeps the next proposal credible.

Compare against the honest alternative

AI employees usually win on bounded, high-volume, verifiable work and lose on judgment-heavy work with low volume. The arithmetic says which is which before you commit, and it is worth doing on paper first — it is considerably cheaper than discovering it in month four.