WhatsApp conversation metrics: which ones actually matter
Delivered, read, replied and blocked each answer a different question. How to read them together, and why a high delivery rate can be bad news.
· 6 min read
Four statuses, four different questions
Message reporting gives a business a small number of statuses, and each answers a genuinely different question. Read individually they are easy to misinterpret; read against each other they are informative.
Delivered answers whether the message reached the device. It is a technical fact about the phone number and the network, and says nothing about a person.
Read answers whether it was opened. This is the first status that involves a human, and it is about attention rather than interest — someone can open a message and dismiss it in a second.
Replied answers whether it produced action. For most business messaging this is the closest available proxy for the message having worked, since a reply requires the recipient to decide to do something.
Blocked answers whether it caused harm. It is the only status with consequences beyond the individual message, because blocks feed the quality signals that determine how many people the business can reach at all.
The common mistake is treating these as a funnel where the goal is to maximise each stage. They are diagnostics. A change in one, relative to the others, is what carries information — and the ratios between them are far more useful than any single figure viewed in isolation.
Why a high delivery rate can be bad news
Delivery rate is the most reassuring metric and the least informative, and there is a specific situation where a high one is a warning.
Delivery tells you numbers are valid and phones are on. That is worth knowing — a falling delivery rate genuinely indicates a decaying list, wrong numbers, or contacts who have left WhatsApp. But high delivery with low read is the pattern worth attention: the messages are arriving and being ignored.
That combination narrows the possibilities usefully. The list is technically fine, so the problem is not data quality. Something about the message, its timing, or the relationship is causing people not to open it. Bad timing, an audience that does not recognise the sender, or content that has stopped being relevant are the usual explanations.
It is also a leading indicator. People who ignore messages do not ignore them indefinitely; a proportion eventually block. So a business watching delivery and read together can see a block problem forming before it arrives, at the point where reducing frequency or tightening segmentation still prevents it.
A business watching delivery alone sees a healthy number all the way up to the rating drop. That is the practical case for never reporting delivery as a headline figure without read beside it.
Read rate, and what it cannot tell you
Read rate is the most-quoted engagement figure on this channel and it needs a few caveats to be used sensibly.
It depends on read receipts, which recipients can disable. A portion of any audience will register as unread having read the message perfectly well, so read rate systematically understates attention by an unknown amount. That makes it usable for comparing your own messages against each other, and unreliable for comparison against any published benchmark.
It is also inflated by the nature of the channel. A message arriving in the same application where someone talks to their family gets opened often, sometimes reflexively. A high read rate on WhatsApp therefore means less than a high open rate would elsewhere, and reading it as evidence of interest overstates what happened.
What it is genuinely good for is comparison under controlled conditions. The same message sent at different times, two versions of a message to comparable segments, or the same campaign across successive months are all fair comparisons, because the receipt-disabling population is roughly constant across them.
What it cannot do is tell you whether the message achieved anything. Read is attention, not action, and the gap between the two is where most disappointing campaigns live: high read, no replies, nothing happened.
Replies and blocks: the two that carry weight
Reply rate is the most underused metric available. It requires the recipient to decide to act, which makes it a far better signal than opening. It also has a mechanical consequence: a reply opens the customer service window, letting the business converse freely without a template.
It only works as a measure if the message actually invites a reply. A message with no reason to respond will have a low reply rate regardless of quality, so the figure is meaningful only for messages designed to prompt one. That is itself an argument for designing messages that way — a template inviting a genuine response produces both a better outcome and a usable measurement.
Block rate is the metric to treat as a constraint rather than a number to optimise. It is the direct input to the quality rating that governs reach, so it should be read relative to volume rather than as an absolute: the same number of blocks means something different against a large send than a small one.
Blocks also lag. Someone irritated by today's message may block on receiving the next one, so a campaign can look fine and damage the following one. This is why block rate should be tracked as a trend across campaigns rather than assessed per campaign, and why a rising trend justifies slowing down even when the current figure looks acceptable.
Opt-outs belong here too. They are the polite version of a block and should be counted as the same signal, not merely processed.
The numbers nobody reports but everyone needs
Several of the most useful figures on this channel do not appear in standard reporting and have to be counted deliberately.
How many inbound conversations went unanswered before the 24-hour window closed. Each is a conversation that now requires a template instead of a sentence, and the count is a direct measure of whether the cheap, flexible version of the channel is being used.
Time to first human reply, measured during open hours and excluding automated acknowledgements. Averages hide the cases that cause harm here, so the worst case and the proportion answered within an hour are more actionable than the mean.
What inbound enquiries are actually about, in categories rough enough to be sustainable. Repeated questions about price, hours or order status are not customer service volume to be absorbed; they are missing information on a website, in a catalogue, or in a confirmation template. Counting them turns a workload into a fixable list.
Cost per outcome rather than cost per message. Under per-message billing a sequence is several charges, so the meaningful figure divides total messaging cost by the number of results — orders, bookings, resolved queries — rather than by messages sent.
And the ratio of transactional to promotional volume, which is a reasonable proxy for whether a business is mostly serving customers or mostly interrupting them.
Reading them together
The habit that makes any of this useful is looking at combinations rather than individual figures, because a pattern points at a cause in a way a single number does not.
Low delivery with everything else normal is a list problem: wrong numbers, stale contacts. High delivery with low read suggests timing, relevance, or an audience that does not recognise the sender. Good read with no replies means the message was seen and gave nobody a reason to act — usually a content problem rather than an audience one. Rising blocks alongside steady read rates points at frequency: people are still opening, and enough of them have had enough.
Comparisons should be internal. Your own figures over time, and between segments, are the only fair benchmark, because published averages mix industries, business sizes and measurement definitions that may share nothing with your case.
And it is worth checking figures during a campaign rather than after it. A send that is going badly can be paused; a send that has finished can only be regretted. That means agreeing in advance what would justify stopping, since a decision made mid-campaign, under pressure to complete a send already paid for, tends to go the wrong way.
The least useful report is the one showing only the numbers that look good. Delivery and read flatter; replies and blocks tell you what happened.
Common questions
My delivery rate is high but almost nobody reads the messages. What does that mean?
That your list is technically fine and the problem is elsewhere — timing, relevance, or recipients not recognising who is messaging them. It is also an early warning, because people who ignore messages eventually block them. Catching it at this stage means reducing frequency or tightening segmentation can still prevent a rating problem.
Is a high read rate on WhatsApp a good sign?
It means less than it appears. Messages arrive in the same app people use for family conversations, so they get opened readily, sometimes reflexively. Read is attention, not action. Meanwhile read receipts can be disabled, so the figure understates attention by an unknown amount. Use it to compare your own messages with each other, not against published benchmarks.
Why should I track blocks if the number is small?
Because it is the direct input to the quality rating that decides how many people you can reach, and because it lags. Somebody irritated by today's message may block on receiving the next one, so a campaign can look clean and damage the following one. Read it relative to volume and watch the trend across campaigns rather than judging each in isolation.
Which metric should I report if I only track one?
None alone, but if forced, replies for whether messaging is working and blocks for whether it is safe. Delivery and read flatter without telling you much. Better still, count something standard reporting omits: how many inbound conversations went unanswered before the 24-hour window closed, which measures whether you are using the cheap, flexible part of the channel at all.