AI & Automation

Measuring Chatbot Quality: What to Track After Launch

Launch is the start of the work on a chatbot. A small set of measures and a weekly review keep quality from drifting.

Kiaanlab Engineering Updated October 4, 2026 3 min read
Performance analytics graphs on a laptop screen

Photo by Luke Chesser on Unsplash

A chatbot is not finished on launch day. Documentation changes, customers ask new things, and the model provider updates the model. Without measurement, quality drifts and the first sign is a complaint. A handful of measures, reviewed regularly, is enough to stay ahead of that.

Resolution, not containment

The number most dashboards show is containment: the share of conversations that ended without a person. It is easy to count and misleading on its own, because a customer who gave up in frustration also counts as contained.

Resolution is the better measure: was the customer's question actually answered? Two practical ways to estimate it are a one-click question at the end of the chat, and checking whether the same customer contacted you again about the same topic within a few days.

Answer correctness

Customers do not always know when an answer is wrong, so satisfaction scores miss some errors. Take a random sample of conversations each week, thirty is enough, and have someone who knows the subject mark each answer as correct, partly correct or wrong. This is manual and it is the most reliable signal you will get.

For the wrong ones, note the cause: missing content, wrong passage retrieved, or right passage and wrong answer. Each cause has a different fix.

Questions the bot could not answer

Log every case where the bot said it did not know or handed over. Group them by topic. This list is a ready-made plan for what to write next in your documentation, ordered by how often customers need it.

Handover rate and reasons

Track how often conversations go to a person and why. A rising rate on one topic usually follows a product change that the documentation has not caught up with. A sudden rise across all topics suggests something technical, such as a failed update of the search index.

Speed and cost

Measure how long the bot takes to respond and what each conversation costs in model usage. Both can change without anyone touching the system, for example when conversations get longer or a tool starts returning more text. Set an alert on daily cost so a jump is noticed quickly.

Run a fixed test set

Keep a set of questions with known correct answers and run it automatically whenever the content, the prompt or the model changes. If the score drops, you know before customers do. Add every real failure you find in the weekly review to this set, so the same mistake cannot return unnoticed.

A weekly routine

None of this needs a large team. A workable routine takes about an hour a week.

  • Review the sample of thirty conversations for correctness.
  • Read the list of unanswered questions and choose the content to write.
  • Check resolution, handover rate and cost against the previous week.
  • Add new failures to the test set.

The routine matters more than the tooling. A spreadsheet reviewed every week beats a dashboard nobody opens.

Be careful with thumbs up and down

Feedback buttons are useful and biased. Few people click them, and those who do are often unhappy for reasons unrelated to the answer, such as not liking the policy. Use them as a pointer to conversations worth reading, not as a score.

Summary

Measure resolution and correctness, mine the unanswered questions, watch handovers and cost, and keep a test set that grows with every failure. Monitoring and review are included in the AI chatbots we build. If you have a bot in production and no clear view of how it performs, we can help you set this up.

KE

Kiaanlab Engineering

The engineers who design and build Kiaanlab's own AI and software systems, writing about what actually works in production.

Tell us what you're building.

A short call, no sales script, just an honest read on scope and timeline.

Discuss a similar project