An SEO audit checklist that needs no paid data
Which audit checks a page's own HTML answers and which need a licensed vendor — a checklist in three columns, with every row placed on one side.
· 6 min read
Why a useful checklist needs more than one column
Most SEO audit checklists circulating online are a single flat list: check your title tag, check your rankings, check your backlinks, check your heading structure. Read as a list of things worth knowing, it is fine. Read as a list of things you can actually go and find out this afternoon, it is misleading, because the rows are not the same kind of row. Some of them you can answer in ninety seconds with the view-source command your browser already has. Others you cannot answer at all without an account at a company that sells data, and no amount of care or effort on your part changes that.
A checklist that does not mark which is which sets a business up to feel like it failed at the audit. The owner works down the list, does well on the first few rows, hits 'check your ranking for your main keyword' and stalls — not because they lack skill but because the row requires a purchase nobody mentioned. Splitting the same rows into columns by what they actually depend on turns the same information into something completable. Column one is answerable from the page itself. Column two needs a vendor. Column three is free but needs the site owner's own permission, which is a genuinely different condition from either of the other two.
Column one: rows a single page's HTML settles
Open one page, view its source, and the following are all simply readable, no interpretation and no third party required. Is there a title tag, and how long is it. Is there a meta description. Is there exactly one h1, and does it say what the page is about. Do the heading levels run in order, or does the page jump from an h2 straight to an h4. Does every image carry an alt attribute with real words in it. Is there a canonical link, and does it point at this page or somewhere else. Is there a robots meta tag, and does it accidentally say noindex. Does the html element declare a lang. Is the page's word count plausible for the question it claims to answer.
A few more rows belong here that people rarely think of as auditable without a tool. Does any sub-resource on this https page load over plain http, which a browser will warn about or block outright. Does any link on the page say only 'click here' or 'read more', wasting the strongest hint a reader gets about the destination. Are the Open Graph tags complete, or will a link shared on WhatsApp appear with no image and whatever text the app manages to scrape. Does any structured data on the page assert a value — a price, a rating, a headline — that does not actually appear in the visible text. Every one of these is a fact about the page in front of you, which means you can re-check it the moment you fix it rather than waiting for a report to run again.
Column two: rows that need a company that sells data
Now the rows that cannot be completed from the page. Where does this page currently sit in results for a given query. How many people search that query in a month. How hard would it be to compete for it. Which queries does a competitor currently earn traffic from. Who links to this page from elsewhere on the web, and how many of them are there. Every one of those is a fact about the world outside the page, and none of them is written in the page's markup.
The reason is structural rather than a matter of trying harder. Google does not run a public position-lookup service and does not publish query volume to anyone who asks. Every difficulty score, every volume figure and every domain-strength number quoted anywhere is one company's estimate, produced from that company's own crawl and its own modelling, and priced accordingly. Backlink indexes work the same way: each one that exists is an independent crawl of the web, which is expensive enough that everyone who has built one charges for access. None of that makes the numbers worthless. It makes them sourced, and a row in an audit that reports one without naming whose estimate it is has quietly converted a private guess into an apparent fact.
Column three: free, but gated on the owner's own consent
The column most checklists collapse into one of the other two is the interesting one. Some data is genuinely free, genuinely accurate, and comes from the search engine itself — but only to the person who owns the site, and only after they prove it and grant access. Search Console is the clearest case: real impressions, real clicks, real average position by query and by page, at no cost, for a property the owner has verified. Nobody else can retrieve it on their behalf without going through that same verification and consent.
Field performance data sits nearby. Core Web Vitals measurements from real visits, and the lab measurements that accompany them, are free to query but need an API key that somebody with the right access has to create and configure first. It does not simply arrive the way an on-page check does. Both of these rows deserve to be marked differently from column two, because the blocker is not money — it is a setup step nobody has done yet. Telling a business 'this needs a credential you have not connected' is actionable. Lumping it in with 'this needs a paid subscription' is not, and it makes people give up on data they could have had for nothing.
Running the checklist on one page, start to finish
The practical shape of a first pass is narrower than most people attempt. Pick the single page that matters most commercially — for a Pune dental clinic that is usually the page about the treatment people actually search for, not the homepage — and work only column one on it. View source, read the head, note the title length and whether the description exists. Search the source for the word noindex and for the word canonical, and read what each one says. Count the h1 tags. Scan the img tags for empty alt attributes. Look at every anchor's visible text.
That pass takes well under an hour on one page and produces a list of concrete defects with no ambiguity about whether they are real, because you read them out of the page itself. Only after that list is empty is there much point moving on, because column two rows cost money and column three rows cost a setup conversation, and neither is worth starting while the page still says noindex or still has no title. Order the work by what the page can already tell you, and the unaffordable rows stop being the bottleneck they appear to be at the top of a flat list.
What to write in a row you cannot answer
The temptation, in a report or a spreadsheet, is to leave nothing in the empty cells or to put in a plausible-looking figure so the document reads as finished. Both are worse than the honest alternative, which is naming what the row is blocked on. 'Position for this query: needs a licensed provider, none connected' is a more useful sentence than a blank, and enormously more useful than a number nobody measured. It tells the next person reading the document exactly what would unblock it, and it stops a future reader treating an invented figure as a baseline to improve on.
This is also the habit worth applying to somebody else's audit before you act on it. For each number in the report, ask which column it came from. If it describes something written in your page's markup, verify it yourself in a minute. If it describes rankings, volume or links, ask which provider's estimate it is, and treat the figure as that provider's opinion rather than as a measurement. A single composite score out of 100 is worth particular suspicion, because producing one requires blending all three columns together with weights the report almost never discloses.
Common questions
Can I do a genuinely useful SEO audit with no budget at all?
Yes, for the on-page and technical half. Everything readable from a page's own markup is free to check and free to re-check the moment you fix it. What no budget cannot buy you is position data, query volume and backlink data, because those are not in your page and no search engine hands them out.
How many pages should a first audit cover?
One, chosen for commercial importance rather than convenience. A single page worked properly produces a list of defects you can act on this week; a site-wide sweep produces a spreadsheet nobody finishes. Widen it once the first page has nothing left in column one.
If a tool shows me a keyword volume for free, where did it come from?
From a paid data source somewhere, or from a model that estimates one. Google does not publish volume figures, so a free display of one means either the tool is absorbing a vendor cost elsewhere or the number is a derived guess. Neither is disqualifying, but it is a different kind of number from your own title tag length.
Is a score out of 100 a fair summary of an audit?
It is a compression that hides its own inputs. Any single score has to weigh an unfixed alt attribute against an estimated difficulty figure, and the weighting is rarely published. Separate counts of errors, warnings and notes, each traceable to the check that produced it, carry the same information without burying the provenance.
Related pages