Measure the Queue, Not the Bot
When we rolled out an AI support agent on our 5,000-ticket monthly queue, we never built a deflection dashboard. Not out of principle - we just couldn't figure out what it would tell us.
Deflection counts the conversations that didn't reach a human. It doesn't count whether anything got solved. A customer who gave up isn't a success; they're a refund waiting to happen, or a repeat ticket next week under a new subject line.
This August, three very different corners of the industry landed on the same conclusion. Automattic's CX lead argued that deflection is not a support strategy. A vendor-side column in Smart Customer Service called for resolution rate to replace it. And China's consumer-protection coverage quoted a legal expert saying AI service should be judged on resolution and escalation efficiency - with a regulator behind the sentence.
So the debate is settled? Almost. Here's the part I keep coming back to: resolution rate can be gamed too.
Close-on-silence, where no reply within 72 hours counts as resolved. Bot-confirmed resolutions, where the agent asks "did that help?" and treats a dropped session as a yes. Reopens logged as fresh tickets, so every failure becomes two successes.
Any metric scoped to the bot's slice of the queue can be flattered by choosing the slice.
What worked for us was refusing to give the AI its own scoreboard. One queue, one clock: AI-handled and human-handled tickets in the same two numbers - average resolution time and customer satisfaction, measured across everything. If the agent "resolved" something it didn't, the customer came back, and the number moved against us. There was nowhere for failure to hide.
That design is also what let us expand with confidence. We started the agent on 10% of the queue and grew it only as the whole-queue numbers held. Six months later it covered everything, and average resolution had dropped from about 24 hours to about 10. Satisfaction stayed above 95% - measured on all of it, not on the bot's best days.
I get why deflection dashboards exist. They're available on day one, and they always go up.
But the metric you hold a team to should be one your customer would recognize as real. Nobody ever wrote in to thank us for deflecting them.