The one question worth asking your own sales data
What is different about customers who come back versus those who don't? Compare source, first product and order value — the difference is your lever.
· 5 min read
Most analysis fails because it has no question
The usual approach to business data is to look at it and hope something stands out. You open the sales sheet, sort a few columns, notice that one month was strong, and close it again none the wiser. This is not a failure of skill or software. It is what happens when analysis has no question, because data does not volunteer conclusions — it answers what it is asked, and an unasked question gets no answer however long you stare.
The fix is to start with one question specific enough to have an answer and consequential enough that the answer changes something. For most small businesses, one question outperforms every other candidate: what is different about the customers who come back compared with the customers who do not? It qualifies on both counts. It is answerable from records you already keep, and its answer points at a lever, because whatever distinguishes the two groups is something you have at least partial control over.
Why this question in particular
Repeat customers are the most valuable population in most businesses for an arithmetic reason: they cost nothing to acquire. Every rupee of acquisition spending buys a first purchase, and everything after that arrives free. So a business that understands what produces a second purchase has found the cheapest available growth, and it does not require more traffic, a bigger budget or a new channel.
The question is also unusually well-suited to small data. It does not need many customers to be informative, because it is a comparison of two groups rather than an estimate of a population value, and a clear difference between groups is visible at modest numbers. It works with the crudest possible records — a list of sales with a customer identifier, a date, a product and an amount. And its answer is nearly always actionable, because the distinguishing factor tends to be something you decide: which channel you spend on, which product you lead with, how you follow up. Compare this with the average vanity metric, which is precise, effortlessly available, and attached to no lever at all.
How to run the comparison
Start by splitting customers into two groups: those who bought more than once, and those who bought exactly once. Take care with the window — a customer who bought their first item last week has not had the chance to return, so exclude anyone whose first purchase falls inside your typical repurchase interval, or you will classify recent buyers as one-time customers and manufacture a pattern.
Then compare the two groups on the attributes you happen to have. Acquisition source is usually the most revealing. First product bought is next, and often the most immediately useful. First order value matters, and the direction is not obvious in advance — sometimes small first orders lead to loyalty because they are a low-risk trial, sometimes large ones do because they indicate a serious buyer. Then the time of year they arrived, whether they used a discount, and their location if that varies meaningfully. For each attribute, work out the proportion of each group falling into each category, not the raw counts, since the two groups are different sizes and raw counts will mislead you every time. A spreadsheet pivot table does all of this in a few minutes.
Reading the result without over-reading it
You are looking for a difference large enough to be worth acting on. If a third of customers acquired through one channel return, against a twentieth from another, that is a gap worth investigating regardless of statistical formality. If the figures are close, the honest conclusion is that this attribute does not distinguish the groups, and saying so is a real result — it removes a hypothesis and stops you from acting on noise.
The important discipline is not to convert a difference into a cause too quickly. A channel whose customers return more often may be attracting better-suited buyers, or it may be that those buyers were already familiar with you from somewhere else and the channel merely collected them. A product associated with loyalty may create the habit, or it may simply be what committed buyers choose first. Both readings imply different actions, and the data will not separate them. What it does reliably give you is a much better question to investigate, and the cheapest way to investigate is to ask several people in the group. A ten-minute conversation with five repeat customers about why they came back will usually explain the pattern the table found.
Turning the finding into something you do
A finding that produces no change was entertainment. The action depends on what the difference turns out to be, and the mapping is mostly obvious once stated. If a channel's customers return more, shift spending toward it and check that the pattern holds as volume grows, since channels frequently degrade when scaled. If one product is associated with repeat buying, consider leading with it, featuring it, or making it the natural first purchase — while remembering the causal ambiguity above.
If discount-acquired customers return much less, that is among the most valuable things you can learn, because it means discounting is buying revenue rather than customers, and the acquisition cost of a discounted first order is higher than it appears. If a season or a location explains it, that shapes when and where you spend. In every case, treat the change as a test rather than a conclusion: write down what you expect to happen and check in a few months whether the repeat rate moved. That converts a one-off analysis into something that compounds, because each round narrows what you believe about your own business.
What the comparison cannot settle
It cannot establish cause, and that limit is structural rather than a matter of getting more data. You are observing groups that differ in many ways at once, including ways you do not record, so any single attribute you find is entangled with others — the channel that brings loyal customers may also bring older customers, in a particular city, buying a particular product, and the table cannot separate those. It also says nothing about people who never bought at all, who are absent by construction and may be the larger and more interesting group.
Small numbers deserve genuine caution. With a few dozen customers per group, one or two individuals move a percentage noticeably, and a difference that looks decisive can evaporate next quarter. Look for gaps large enough that a couple of people could not explain them, and prefer patterns that persist across several periods to a single striking result. And a difference you find is a description of the past: it held for the customers you have already acquired, under the conditions that applied then. It is a hypothesis about the future, not a fact about it, which is why the next step is always a test rather than a commitment.
Common questions
How many customers do I need for this to work?
It becomes readable at a few dozen in each group, because you are comparing two groups rather than estimating a population value. Below that, the comparison is still worth doing to build the habit and the records, but treat any gap as a hypothesis to watch rather than a finding to act on.
What if I do not record who my customers are?
Then start, because this analysis is impossible without a customer identifier. A phone number attached to each sale is enough, and it is the single highest-value addition to most small-business records. Consistency matters more than completeness: the same buyer recorded two ways becomes two one-time customers.
Which attribute usually explains the most?
Acquisition source and first product bought are the two that most often show a clear gap, and both are things you control. But the answer genuinely varies between businesses, which is the reason to check yours rather than adopt someone else's conclusion.
Should I ask customers instead of analysing records?
Do both, in that order: the records tell you where to look, and the conversations tell you why. Asking without the analysis means you do not know which customers to ask; analysing without asking means you have a pattern and a guess about its cause. The combination is much stronger than either.
Related pages