Skip to main content
2026 Report

Deflection is not resolution

The AI chatbot market's favourite metric counts how many tickets never reach your team, not how many customers are actually well served. This report explains why that difference matters and which metrics you should demand instead.

By Javier Manzano August 2026 12 min read
Executive summary

What this report argues

The AI chatbot market sells deflection: the percentage of conversations that never reach a human. It is a cost metric dressed up as a quality metric, and buying it at face value has consequences for your customers, your brand and your bill.

60-80%
Promised deflection

This is the range much of the market advertises. The number tells you how many conversations never reached your team; it does not tell you how many customers got a correct answer or how many gave up along the way.

=
Abandonment counts as success

A customer who closes the chat in frustration without asking for an agent is counted exactly like one who was perfectly served: as a "deflected" conversation. The metric cannot tell resolving from discouraging.

2024
The Air Canada precedent

A Canadian tribunal ordered Air Canada to honour the refund policy its chatbot made up. Since then, answer quality has stopped being just an experience problem: it is a legal risk.

5
Missing metrics

Verified resolution, post-conversation satisfaction, escalation quality, abandonment rate and cited-source answers. Five metrics no vendor shows you in the demo, and which this report defines.

Section 1

What deflection actually measures

Deflection rate measures the percentage of conversations that end without a human agent stepping in. It was born as an internal cost metric in help desks, and the AI market adopted it as its headline sales metric.

A cost metric, not a quality metric

A ticket that never reaches your team can mean three very different things: the chatbot resolved the question, the customer gave up in frustration, or the customer went to another channel (an email, a phone call, a public review) where the problem no longer gets counted. For deflection, all three cases are identical: success.

That bias is no accident. Deflection is the metric that justifies the purchase internally ("we save X agents") and, at the same time, the one that commits the vendor to the least: it does not require proving the answers were correct, only that no human stepped in. That is why it is the number on every product homepage.

The scorekeeper's conflict of interest

Across much of the market, the party measuring the metric is the party getting paid for it. Per-resolution pricing models bill for every conversation the system classifies as "resolved" — and that classification is made by the system itself. The vendor defines the metric, measures it and invoices based on the result.

You do not need to assume bad faith to see the problem: when the criterion for "resolved" is set by whoever bills per resolution, the definition will tend to be generous. Any serious chatbot audit starts by asking who decides what counts as resolved and how it can be verified independently.

The kill-switch thought experiment

A chatbot that answered "I can't help you with that" to every question and never offered a button to talk to an agent would have a deflection rate close to 100%: not a single ticket would reach your team. By the market's favourite metric, it would be the best possible chatbot. If a metric rewards switching support off, it is not measuring the quality of support.

~100%
deflection of a bot that doesn't help
Section 2

The five metrics you should demand

None of these metrics show up in the demo, because none of them can be measured in a demo: they all require real conversations and after-the-fact review. That is exactly why they matter.

1. Verified resolution

Percentage of conversations where a human review (by sampling) confirms the answer was correct and complete. It is the honest version of deflection: instead of counting tickets that never arrived, it counts customers who were well served. It is measured by auditing a weekly sample of conversations against the actual documentation.

2. Post-conversation satisfaction

CSAT asked right at the end of the conversation with the bot, not the generic helpdesk survey. It separates two questions the market blends together: "did it solve your problem?" and "how was the experience?". Low CSAT with high deflection is the classic signature of a bot that scares customers away instead of resolving.

3. Escalation quality

Two sides: how many conversations that needed a human were escalated in time (not trapping the customer in a loop), and how many escalations were unnecessary (not interrupting your team with what the bot could resolve). A good assistant gets both directions right and hands off with full context.

4. Abandonment rate

Percentage of conversations the customer cuts off with no goodbye, no resolution and no request for an agent. It is deflection’s black hole: today, every one of these cases counts as success. Measuring it requires logging how each conversation ends, not just whether a human stepped in.

5. Cited-source answers

Percentage of answers that cite the document they come from. Without a citation, there is no way to tell whether an answer comes from your documentation or from the model’s imagination, and no way to audit errors after the fact. Since the Air Canada precedent, this metric is also a legal defence.

The synthesis: cost per real resolution

What you pay per month divided by the conversations with verified resolution (metric 1). It is the only number that lets you genuinely compare two vendors with different pricing models. You can estimate it with your own volume in our chatbot cost calculator.

Want to know what a chatbot would cost you at your real volume?

Cost calculator →
Section 3

Why the market doesn't measure this

It is not a technical oversight: every missing metric has an economic incentive behind it. Understanding them helps you negotiate with any vendor — including us.

doesn’t scale in self-service

Verified resolution costs the vendor money

Verifying resolutions requires people reviewing conversations against the client’s actual documentation. It is ongoing work that no prompt can scale away. A self-service platform with thousands of accounts cannot offer it without changing its business model, so the metric gets replaced by the bot’s own self-assessment.

?
the question nobody asks

Post-bot CSAT is scary

Asking "did we solve it for you?" right after a conversation with the bot produces uncomfortable data that does not support the sales pitch. The survey gets deferred, aggregated with the human team’s, or simply never sent. If a vendor won’t give you bot-only CSAT, ask yourself why.

+%
ambiguity adds to success

Abandonment is invisible by design

Telling "resolved" apart from "gave up" requires classifying how every conversation ends, and every unclassified conversation pads the good number. The vendor’s incentive is to leave the ambiguity in place: deflection goes up on its own.

[1]
the missing citation

Citing sources exposes the mistakes

A bot that cites the document behind each answer makes every error auditable: you can see what it invented and what it misread. A bot that doesn’t cite is unfalsifiable. The opacity is not a technical limitation (RAG with citations has been standard for years): it is a product choice.

The conclusion is not that platforms lie: it is that they measure what their business model allows them to measure. A vendor billing per conversation has no incentive to reduce conversations; one billing per resolution has none to tighten the definition of resolution; and one on a flat fee only keeps the client if the assistant genuinely works. Before you compare prices, compare incentives.

Section 4

Checklist: what to ask before you sign

Seven questions for any AI chatbot vendor, including us. The answers separate those who sell deflection from those who commit to resolution.

01

Who defines "resolved"?

Ask for the exact definition of a resolved conversation and who applies it. If the bot classifies it itself and nobody audits it, the number you’re shown is a self-assessment.

02

Can I audit conversations?

Demand full access to the conversation history and the ability to review samples with your own team. Without access to the conversations there is no verifiable metric.

03

Do answers cite their source?

Every answer should link to the document it comes from. It is the difference between being able to audit an error and discovering it when a customer (or a tribunal) suffers it.

04

How is abandonment measured?

Ask what happens to conversations the customer cuts off without resolution. If the answer is that they count as deflected, you now know how the number on the homepage was built.

05

Is there separate bot CSAT?

Bot satisfaction must be measured at the end of the conversation with the bot and reported separately from the human team’s. Aggregating them hides exactly what you want to know.

06

What happens when I double my volume?

Run the pricing numbers with your current volume and with twice that. If the bill grows in proportion to the assistant’s success, the pricing model is working against you.

07

Who maintains the assistant?

Your customers’ questions change; so does your documentation. Ask who reviews failed conversations every week and tunes the assistant: if the answer is "you", the real price includes your team’s hours.

Our position in this debate is not neutral, and it is worth saying so: the Soamee AI assistant cites its sources in every answer, includes conversation reviews with the client and works on a fee agreed upfront, precisely because we built the product around the answers we would want to receive to these seven questions. This report lays out the criteria you can use to evaluate us too.

Methodology and sources

How we compiled this report

This report is based on the public documentation of the market's leading chatbot and AI assistant platforms (product pages, technical documentation and published pricing pages), on the public literature on customer support metrics (deflection rate, CSAT, FCR) and on the publicly available February 2024 decision of the British Columbia Civil Resolution Tribunal (Canada) in Moffatt v. Air Canada.

The definitions of the five proposed metrics come from our experience putting AI assistants into production for clients across different sectors. We include no data from any specific client and no internal figures from any project: the numerical examples in the report are illustrative.

This report is published for informational purposes. Soamee develops and sells its own AI assistant, and therefore holds a position in the market analysed here; we have taken care to make every claim independently verifiable and we flag our position where relevant. If you spot an error, write to us at info@soamee.com and we will correct it.

Put us to the seven-question test

Bring us your real documentation and we'll show you the assistant answering the questions your customers actually ask, with its sources cited. No strings attached.

Deflection Is Not Resolution: AI Chatbot Metrics

Tell us your challenge. We'll propose a solution.

No commitment. Within 24 hours, you'll receive a proposal with scope, timeline and budget. No fine print.

Book a free call →