Cohort analysis: what totals hide about your customers
Group customers by the month they first bought and track each group's repeat rate. Why total revenue can rise while retention quietly gets worse.
· 5 min read
The failure that totals are structurally unable to show
Total revenue can rise every month while the business gets worse in a specific, measurable way. It happens when new customers are arriving faster than existing ones are leaving. The incoming volume more than covers the loss, so the total climbs, and every figure on a conventional dashboard supports the conclusion that things are going well. Underneath, the average customer is becoming less valuable, and the business is becoming more dependent on continuous acquisition to stand still.
This is not a subtlety that better totals would reveal. A total mixes everyone together, so a customer acquired two years ago and one acquired last week contribute to the same figure with no distinction between them. Any measure computed over all customers at once inherits this: it cannot separate a change in how customers behave from a change in how many arrived. Answering the question requires keeping the groups apart, which is what a cohort analysis does and what nothing else does.
What a cohort is
A cohort is a group of customers who share a starting point — most usefully, the month in which they first bought from you. Everyone whose first purchase was in March is the March cohort, and they stay in it permanently. The cohort does not gain members later, which is the property that makes it useful: because membership is fixed, any change you observe in a cohort over time is a change in behaviour rather than a change in composition.
That single property is what totals lack. When you look at a cohort's second month, its third and its sixth, you are watching the same people, so if the share still buying falls, those specific customers stopped buying. Nothing else could account for it. And because every cohort can be measured at the same point in its own life — month three for the March group and month three for the September group — you can compare cohorts fairly even though they started at different times. That comparison is the entire payoff, and it is available from a sales list with two columns.
Building the table in a spreadsheet
You need one row per sale with a customer identifier and a date. First, for each customer, find the month of their earliest purchase; that is their cohort. Then build a grid: cohort months down the left, and across the top the months since first purchase — 0, 1, 2, 3 and onward. In each cell, count how many customers from that cohort bought in that month, then divide by the cohort's original size to get a percentage.
The result is a triangle rather than a rectangle, because recent cohorts have not lived long enough to fill the later columns. That shape is correct and worth preserving; filling those cells with zeros is the most common error in a first attempt and it makes recent cohorts look catastrophic. Read the table twice. Across a row shows how one cohort decayed over its life. Down a column compares different cohorts at the same age, which is the reading that answers whether you are improving — if month-three retention for recent cohorts is lower than for older ones, something changed, and it changed for customers acquired after a particular point. A monthly grain suits most businesses; where purchases are naturally infrequent, quarterly cohorts avoid a table that is mostly empty.
What the pattern tells you
Every cohort declines, and that is normal rather than alarming — some share of any group will not return. What matters is the shape. A curve that falls steeply and then flattens is the healthy pattern: you lose a portion quickly and retain a stable core indefinitely, and that core is the part of the business that compounds. A curve that declines steadily without flattening means there is no core, and the business will need acquisition at an increasing rate forever, because nothing accumulates.
The most valuable finding is a change between cohorts, because it is dated. If cohorts from before a certain month retain well and later ones do not, something changed around that month, and the list of candidates is short enough to investigate: a price change, a supplier or quality change, a new acquisition channel bringing different buyers, a staff departure, a competitor's arrival. Note the ambiguity honestly, though — a channel that brings poorly-retaining customers and a product that got worse produce a similar dip, and telling them apart means splitting cohorts by channel or asking people. The table locates the change in time; it does not name the cause.
Acting on an ugly trend
The first response should be to check the measurement rather than the business, because retention figures are unusually easy to get wrong. Confirm that customers are identified consistently — the same buyer recorded under two phone numbers appears as two customers who each bought once, which manufactures a retention problem out of a data problem. Confirm the window suits your buying cycle; if customers typically return every four months, a table read at month two shows a collapse that is only a schedule.
Once the trend is real, the useful move is to narrow it before acting. Split the failing cohorts by acquisition channel, by first product bought and by order value, and look for where the decline concentrates. It usually does concentrate, and that turns an unmanageable problem into a specific one: discount-acquired buyers not returning is a different issue from a particular product's buyers not returning, and they have different remedies. Then talk to people in the affected group, because the table cannot tell you why and they can. The strongest argument for doing any of this is arithmetic: a retained customer needs no acquisition spending, so a small improvement in the flat part of the curve is usually worth more than the same effort spent on new customers.
What a cohort table cannot tell you
It cannot tell you why, and it is worth being firm about that because the table's precision is persuasive. A grid of percentages looks like an explanation and is only a measurement. It also cannot distinguish a customer who has left from one who has not returned yet, and the difference is invisible until enough time passes — which means the most recent, most decision-relevant cohorts are always the least reliable ones. Any conclusion about a cohort three months old is provisional.
Small numbers compound this. A cohort of twenty customers moves four or five percentage points when one person buys or does not, so a difference between two small cohorts can be nothing at all. Look at the direction across several consecutive cohorts rather than the gap between two, and be suspicious of any cohort small enough that individual behaviour dominates. And the table is silent on customers who never bought once: the people who considered you and did not, who may be the larger group and are absent from every cohort by construction. Retention analysis describes the customers you got, not the market you addressed.
Common questions
How many customers do I need for this to be worth doing?
Enough that a single person's behaviour does not move a cohort's percentage much — a few dozen per cohort is where it starts becoming readable. Below that, use longer periods such as quarterly cohorts, and treat the direction across several cohorts as the signal rather than any individual figure.
Should I use months or quarters?
Match the grain to how often customers naturally buy. Monthly cohorts suit businesses with frequent repeat purchases; for something bought a few times a year, monthly produces a mostly empty table and quarterly is more readable. The wrong grain makes normal purchasing rhythm look like churn.
Can I run this on revenue instead of customer counts?
Yes, and it answers a related question: revenue per cohort over time shows whether retained customers spend more or less as they age. Do the count version first, because mixing the two lets a few large orders disguise the loss of many customers.
What is a good retention rate?
There is no figure that transfers between businesses, because it depends entirely on how often your product is naturally rebought. The useful comparison is your own cohorts against each other over time. A published benchmark from a different category tells you nothing actionable about yours.
Related pages