Chat analytics that change decisions
Message volume is a vanity metric. Here are the four numbers that actually cause someone to do something differently.
Vishal ChiniwarCo-founder and CTO2026-04-081,218 words
Most chat dashboards lead with conversation count and message volume. Both go up when traffic goes up and neither tells you to do anything. They are activity measures presented as performance measures.
Here are four that pass the test: if the number moves, somebody does something differently.
1. Unanswered rate
The share of conversations where the assistant could not answer. This is the direct measure of content gap size.
If it rises, either your content stopped covering what people ask, or your traffic mix changed and a new audience is arriving with different questions. Both are actionable, and the action differs, which is what makes the metric useful. Segment it by landing page to tell them apart.
2. Question categories by distinct askers
Not message counts. Distinct people asking about each topic, ranked. This is the content roadmap, generated rather than debated.
The specific decision it drives: what gets written next. When this list is available, the content meeting is a five minute confirmation instead of an argument about priorities.
3. Handoff rate, split by reason
Total handoff rate on its own is ambiguous, because escalation is sometimes success and sometimes failure. Split by trigger and it becomes readable.
Escalations because the assistant failed twice are a quality problem. Escalations because someone asked about contract terms are a commercial opportunity. If the first category grows, fix content. If the second grows, that is good news and sales should know.
4. Post conversation behaviour
What visitors did after the conversation ended. Did they continue to a pricing page, leave, or come back later. This is the only one of the four that connects conversations to outcomes rather than measuring the conversation in isolation.
It is also the hardest to instrument honestly, and it deserves a caution: correlation here is weak evidence. People who chat are already more engaged, so they were always more likely to continue. Do not present this as a causal conversion lift unless you have actually run a controlled comparison.
What to leave out of the main view
- Total messages. Rises with traffic, decides nothing.
- Average conversation length. Longer can mean engaged or stuck, and averaging destroys the distinction.
- Satisfaction ratings with low response rates. A five percent response rate from self selected raters is not a measurement.
- Deflection rate as a headline. On a marketing site it rewards the wrong outcome.
The review that makes the numbers useful
Numbers tell you where to look. Reading ten actual conversations a month tells you what is happening. Teams that only read the dashboard develop confident theories that the transcripts would have corrected in twenty minutes.
Pick the ten from the categories where the numbers moved. That way the reading is targeted rather than random sampling.
The metrics that look important and change nothing
Total conversations. It moves with traffic, so it tells you about marketing rather than about the assistant, and it goes up when you add the widget to more pages regardless of whether that helped.
Average conversation length. Longer can mean engaged or can mean the visitor had to ask four times. The number is genuinely ambiguous and no threshold makes it less so.
Satisfaction score. Response rates are low and skewed toward people who had a strongly bad experience, so the sample is both small and unrepresentative.
Answer accuracy as a single percentage. It requires labels you do not have, and any figure produced without them is an estimate wearing a decimal point.
The four that do change decisions
Refusal rate by topic. Points directly at a page to write or fix, ranked by how many people needed it. This is the single most actionable number available.
Follow up rate by topic. Isolates answers that were correct and useless, which no accuracy measure catches and which are a large share of real dissatisfaction.
Escalations that a content fix would have prevented. Turns support volume into a content brief with an effort estimate attached.
What the visitor did next. Conversation followed by a pricing page visit is the closest thing to commercial signal available, and it is honest about being correlation rather than proof.
Making the numbers survive contact with a meeting
The failure of most analytics work is not measurement, it is that nobody acts on it, usually because the output is a dashboard rather than a decision.
Report as a ranked list of proposed changes with the evidence attached. Fourteen people asked about return postage and we refused all fourteen, so here is the paragraph to add. That is a decision someone can approve in a sentence.
Bring one verbatim quote to every review. A real question a real visitor typed does more than any chart, because it is impossible to argue with and impossible to abstract away.
And report the same four numbers every month rather than rotating through whatever looks interesting. Consistency is what makes a trend legible.
A monthly review that takes an hour
Fifteen minutes reading twenty transcripts end to end, unfiltered. This is the part that changes what you think, and it is the part that gets skipped because it is not a chart.
Fifteen minutes on the refusal list, picking the top three clusters and writing one sentence each describing the fix.
Fifteen minutes checking whether last month's fixes worked, by rerunning the questions that prompted them.
Fifteen minutes writing it up as three proposed changes with counts attached. Not a report. Three lines somebody can approve.
An hour a month sustained beats a dashboard nobody opens, and it is the only cadence that has reliably survived contact with a real team.
The number to put in front of a board
None of the four operational metrics travels well upward, because they describe content health rather than commercial outcome. The translation worth making is refusals closed.
Count the gap clusters you identified, the ones you fixed, and the drop in refusals on those topics afterwards. Nineteen refusals on returns last quarter, one this quarter, after a paragraph was added. That is a sentence a board understands and it is entirely defensible.
Resist attaching a revenue figure to it. The causal chain from an answered question to a purchase is real and weak, and a number invented to make it look strong will be the number you are asked to defend later.
Instrumenting for the questions you have not thought of yet
The most common regret with chat analytics is not measuring the wrong thing. It is discovering three months in that the raw data needed to answer a new question was never stored.
Store the whole conversation object rather than derived fields. The question as typed, the answer as sent, the passages retrieved with their scores, the model that answered, whether it refused, whether it escalated, the page, the timestamp, and a session identifier. That is a small record and it makes almost any future question answerable.
Derived fields are the trap. A pipeline that extracts topic and sentiment and discards the text has thrown away the ability to recluster later with better categories, and the categories always turn out to be wrong on the first attempt.
The session identifier matters more than it looks. Without it you can measure conversations and cannot measure what happened afterwards, which is the only bridge from a chat log to a commercial outcome.
None of this needs a data warehouse. A structured log with those fields, queryable, covers everything discussed here and can be built in an afternoon.
Terms in this articleHuman Handoff
Related reading
What a good handoff actually contains
Escalating a conversation without context makes the visitor start over. Here is the minimum payload a human needs.Sachin Aathreyaa K M2026-04-22Qualifying without an interrogation
Two questions asked at the right moment beat six asked upfront. Most qualification signal is already in the transcript.Sachin Aathreyaa K M2026-04-15
Questions
Signal
Have a question about how this works?
Access is by waitlist, demo request or public-content pilot. Ask directly and we will answer, including where it will not fit.