Why open rates became unreliable, and what to use now
Apple Mail preloads images before anyone looks, so opens include events no human caused. Which metrics survive, and how to reset your reporting.
· 5 min read
How open tracking ever worked
An email open has never been directly observable. Nothing in the mail protocols reports back that a human looked at a message, so the industry built a proxy: a transparent image, typically one pixel square, with a URL unique to each recipient. When a mail client displays the message it fetches remote images, that request reaches the sender's server, and the request is recorded as an open.
That mechanism was always approximate in both directions. A recipient reading with images disabled generated no open, so genuine reads went uncounted — historically a large share, since many clients blocked remote images by default. A recipient who left a message visible in a preview pane for a second while scrolling past it generated one.
This is worth stating plainly because it changes how the recent shift should be understood. Open rate did not go from accurate to inaccurate. It went from a rough proxy with known biases to a measure dominated by automated fetches, and the earlier number was never the clean figure people remember it as.
What Mail Privacy Protection changed
Apple Mail's Privacy Protection is a setting Apple Mail users are prompted to enable. When it is on, message content is routed through proxy servers Apple operates, and remote content — every image in the message, including any tracking pixel — is preloaded in the background when the message arrives, rather than when the person opens it.
Two signals break at once. The open event now fires whether or not a human ever looks, so an open is no longer evidence of a read. And because the fetch comes through Apple's infrastructure, the IP address it carries is Apple's rather than the recipient's, which removes the location and device inferences that open tracking used to supply as a side effect.
Apple is the largest single cause but not the only one. Corporate security gateways fetch and scan remote content in messages before delivering them, for reasons that have nothing to do with privacy and the same effect on your statistics. Any environment that pre-fetches content produces opens with no reader attached.
Why your number moved, and in which direction
What happened to a given sender's open rate depended on the share of their audience reading in Apple Mail with the setting enabled. Senders with a consumer audience heavy on iPhones and Macs typically saw reported opens rise, sometimes sharply, with no change in behaviour behind it. Senders with a mostly corporate audience on other clients saw less movement.
The more important effect is not the level but the comparability. Any comparison that spans the change is measuring two different quantities and attributing the difference to your campaigns. A year-on-year open rate chart crossing that boundary is not a trend line; it is a measurement change drawn as a trend.
This has a second-order consequence that catches teams out. If open rate is an input to something automated — a dormancy definition, a 'resend to non-openers' rule, a lead score, a segment used for a discount — then the inflation propagates silently into decisions. The dashboard looks healthier and the automation is quietly acting on a signal that no longer identifies engaged people.
Which metrics survive
Rank the alternatives by how hard they are to trigger accidentally.
Revenue and conversions are the strongest: a purchase or a booking attributed to a campaign required a person to complete something. Replies are nearly as strong and badly underused, particularly for smaller senders where a reply is a genuine conversation.
Clicks are the practical workhorse, with one caveat that is often omitted: link-scanning services and security tools follow URLs in messages too, so click data carries some machine traffic. It is far cleaner than open data, and the machine share can be reduced by ignoring clicks that arrive within a second of delivery or that hit every link in a message at once.
Unsubscribe rate and spam complaint rate remain fully meaningful, because both require deliberate human action, and both are more diagnostic of a campaign's reception than opens ever were. Delivery and bounce rates were never affected by any of this. Notably, the metrics that survive are the ones closest to a decision the reader actually made.
What you genuinely cannot recover
Some things are gone rather than replaced, and a tool that presents them confidently is misrepresenting what it knows.
You cannot know whether a specific recipient read a specific message. Not approximately, not with a correction factor — for a recipient using a proxy that preloads content, the signal does not exist. Vendors publishing 'estimated real opens' are modelling, and the model rests on assumptions about client mix that they cannot verify per recipient.
Downstream, three things need rebuilding rather than adjusting. Engagement segmentation has to be based on clicks and purchases. Automation triggered by 'did not open' is unreliable, since non-openers under this regime include people who did open. And subject-line testing scored on opens is measuring something partly independent of the subject line.
One genuinely useful consequence is that open-based tracking was always the weakest evidence in the stack. Losing it forces reporting onto measures that were more honest anyway.
Rebuilding reporting honestly
Start by marking the boundary. Pick the date your audience's client mix meaningfully shifted, note it in whatever reporting you keep, and refuse to compare across it. Resetting a baseline is less satisfying than a continuous chart and considerably more truthful.
Then choose one primary metric per campaign type, before sending rather than afterwards. A promotional campaign is judged on orders or revenue. A newsletter is judged on clicks and replies. A win-back is judged on confirmed opt-ins. Choosing afterwards means choosing whichever number looks best, which is how a campaign that achieved nothing gets reported as a success on the strength of an inflated open rate.
Keep opens if they are useful to you, but label them for what they are — an indicator of deliverability and of automated fetching, not of readership. And if you report to somebody else, say so explicitly. The alternative is that a number everybody knows is unreliable keeps being quoted as though it were not, which is the condition most email reporting is currently in.
Common questions
Should I stop tracking opens altogether?
Not necessarily, but demote them. Open data still indicates that a message was delivered and that content was fetched, which has some diagnostic value. What it cannot do is measure readership or drive segmentation, so the useful change is to stop using it as a success metric and stop feeding it into automation.
Can I identify which opens came from Apple's proxy?
Partially. Requests routed through Apple's infrastructure can often be distinguished by the network they arrive from and by their timing relative to delivery, and some platforms label them as machine opens. This lets you estimate how much of your open volume is automated, but it does not let you recover whether any particular person read the message.
Does this affect click tracking too?
Much less, and not in the same way. A click requires an action on a link rather than a passive content fetch, so it remains a strong signal. The contamination comes from link-scanning and security tools that follow URLs, which you can largely filter by ignoring clicks that fire within a second of delivery or that hit every link in one message.
How do I run an A/B test now?
Score it on the outcome you actually wanted — clicks, replies, orders — rather than on opens. Those events are less frequent, so you need more volume before a difference is distinguishable from chance. That constraint is real, and it is better to run fewer, larger tests you can trust than frequent ones scored on a polluted metric.
Related pages