Why deflection rate misleads
Deflection counts conversations that didn’t reach a human, which includes every customer who gave up. It reliably overstates value, and it is the number vendors lead with for exactly that reason. A system can post an excellent deflection rate while quietly damaging retention.
The mechanism is simple enough to be worth stating plainly: a customer who abandons a frustrating conversation and buys elsewhere is recorded as a success. Every metric built on absence of escalation inherits that flaw, which is why the four below all measure presence of something instead.
1. Resolution, confirmed downstream
Count an interaction as resolved only if the customer didn’t come back about the same issue within a sensible window. This one change usually cuts a headline number substantially, and what remains is real.
Pick the window from your own data rather than a convention — long enough to capture a genuine repeat contact, short enough that unrelated enquiries don’t pollute it. For most businesses somewhere between a few days and a fortnight, and the right answer is visible in how your repeat contacts already cluster.
Repeat contacts are also the cheapest place to find what the AI employee is getting wrong, because the second conversation is usually explicit about what the first one missed.
2. Human minutes returned
Measure the time your team no longer spends, not the volume the AI handled. These diverge when the AI takes on work nobody was doing anyway, or when reviewing its output costs nearly as much as doing the task.
Measure the review time explicitly, as a separate line. It starts high and should fall as trust builds; if it plateaus at a level close to the original task time, you have automated the wrong step and no amount of model improvement fixes that.
If reviewing the output takes as long as doing the work, you’ve automated the wrong step.
3. Quality against your own baseline
Sample the AI’s handled conversations and score them the way you score your team’s. The bar is not perfection — it is your existing standard. Teams often discover the honest comparison is more favourable than they expected, and occasionally that it is much worse in one specific category worth pulling back.
Score a matched sample of human-handled conversations at the same time, blind if you can manage it. Without that comparison the exercise measures your reviewer’s mood against an imaginary ideal, and the resulting number persuades nobody.
4. Coverage nobody was providing
Some of the value is work that simply wasn’t happening: after-hours responses, follow-ups that were dropped, research nobody had time for. This shows up as new pipeline or faster response times rather than as saved cost, and it is frequently the largest line once measured.
Count it separately from savings. Mixing the two produces a number nobody outside the project believes, and finance will find the seam immediately — as they should.
The costs that belong on the other side
A worth-its-seat calculation is only honest if the full running cost sits opposite the benefit: review time, knowledge maintenance, the escalation load, and the model spend. Those are mostly staff hours rather than invoices, which is exactly why they get omitted.
Escalation deserves its own note. Automation removes the easy cases first, so what reaches your team is disproportionately hard — average handling time on human-handled work goes up even as volume falls, and that looks like a productivity regression if nobody predicted it.
Then compare against the real alternative
The comparison is not against a perfect system, it is against what you would otherwise do: hire, outsource, queue the work, or let it go undone. An AI employee that handles a bounded job at your existing quality standard for less than the alternative is worth its seat, and that is the whole test.
Set the review point in advance — a date, with these four numbers and the costs beside them — and agree what result would mean stopping. Deciding that while nobody is invested is the only time it can be decided cleanly, and it makes switching it off a normal outcome rather than an admission of failure.