It's easy to launch a conversational AI assistant and judge it by feel — "it seems to be handling things" — without ever knowing the actual containment rate, where it's failing, or what it's costing per conversation. That gap is one of the most common reasons conversational AI projects stall or get quietly abandoned: not because the technology failed, but because nobody could point to a number proving it worked.
Good analytics turns "it seems fine" into a specific, trackable answer.
The Core Metrics
- Containment (or resolution) rate. The share of conversations the assistant handles end to end without a human. The single most-watched number, but never sufficient alone.
- Escalation rate and reason. How often conversations hand off to a person, and why — a well-timed, well-context escalation is a success; a late one after a wrong attempt is not.
- Accuracy against a test set. Performance on a curated set of real questions with known correct answers, scored before launch and re-tested periodically.
- Customer satisfaction on assisted conversations. Whether people who interacted with the assistant were actually satisfied with the outcome, not just whether the conversation technically ended.
- Cost per conversation. What each resolved conversation actually costs, especially relevant when comparing a per-conversation platform fee against a custom build.
Why Containment Rate Alone Is Misleading
A high containment rate sounds like success, but it can just as easily mean the assistant is confidently answering questions wrong and never triggering escalation — which is worse than a lower containment rate with honest handoffs. Containment always needs to be read next to accuracy; a number in isolation tells you activity happened, not that it happened well.
Building an Evaluation Set That Actually Reflects Reality
The most useful test sets come from real transcripts, tickets, or call logs — the actual questions people ask, in their actual phrasing, including the awkward and ambiguous ones — rather than a tidy list of questions an internal team imagines customers will ask. A test set built from imagined questions consistently overstates how well a system will perform once real users get to it.
Watching for Drift
Conversational AI accuracy degrades quietly when the business changes underneath it — a new product line, an updated policy, a seasonal spike in a specific question type — and the assistant's grounding data hasn't caught up. Regular review catches this before it shows up as a wave of bad answers; waiting for a complaint to surface it is the more expensive way to find out.
Turning Numbers Into Action
Dashboards only matter if someone acts on what they show. A weekly drop in accuracy on a specific topic usually means the underlying source data changed and the assistant's grounding hasn't caught up; a rising escalation rate on a particular question type often signals a gap in scope worth addressing directly rather than routing around indefinitely. Treating analytics as a feedback loop that shapes what gets fixed next, rather than a report that gets glanced at and filed away, is what separates teams that steadily improve their deployment from teams whose assistant quietly degrades over time.
Where Analytics Should Live From Day One
Analytics should be part of a deployment's initial build, not a phase-two addition once something feels off. Every deployment we build ships with conversation logs and a defined evaluation set from launch, so containment, escalation, and accuracy are tracked against agreed numbers rather than assumed. See our conversational AI overview for how measurement fits into the broader build process, and our page on conversational AI use cases for what "working well" looks like across different applications.
Frequently asked questions
What are the most important conversational AI metrics?
Containment or resolution rate (conversations closed without a human), escalation rate and why escalations happen, accuracy against a test set of real questions, customer satisfaction on assisted conversations, and cost per conversation. Together these show whether the assistant is actually working, not just whether it's being used.
What is containment rate and why does it matter most?
Containment rate is the share of conversations the assistant resolves without needing a human. It's the closest single number to "is this working," but it should always be read alongside accuracy — a high containment rate paired with wrong answers just means people are getting bad information faster.
How do you measure accuracy for a conversational AI system?
By building a test set of real questions with known correct answers, drawn from actual transcripts or tickets, and scoring the assistant against it before launch and periodically afterward. Without this, accuracy is usually just an impression rather than a number.
Should we track conversations that get escalated as failures?
Not automatically — a well-designed escalation, where the assistant correctly recognizes it should hand off and does so with context, is a success, not a failure. What's worth flagging is escalations that happen late, after the assistant already attempted and got something wrong.
How often should conversational AI analytics be reviewed?
Regularly enough to catch drift — new products, policy changes, or seasonal questions the assistant wasn't built for. Many teams review dashboards weekly and do a deeper accuracy re-test whenever the underlying data or business rules change meaningfully.
