Schema that refuses to lie
Fifteen schema types, each generated only when required properties are present and, if page text is supplied, only when the values actually appear on the page.
Works todayAvailable now
In build
The whole team
Nineteen specialists, each with a defined job and an honest status label.
See all nineteenA score out of 100 is easy to sell and impossible to act on — nobody can point at the seventeen inputs that produced 73. Khoji has no score column, and no route whose path even contains rank, position, traffic or volume, a fact a backend test checks directly against the mounted routes. What it has instead is history: every audit is stored, so the second run of a page reports exactly what changed since the first.
Why this exists
Open almost any SEO tool and the first thing it shows you is a number out of 100 — a health score, an optimisation score, a readiness score. It is a genuinely useful piece of marketing, because a single number is easy to screenshot and easier to sell against a competitor's lower one. It is also, on inspection, close to meaningless: nobody publishes the weights, the inputs shift between runs, and two audits of the same unchanged page can return different scores because the vendor tuned the model between visits. A number you cannot interrogate is not a measurement. It is a feeling wearing a decimal point.
Khoji does not have one, and the omission is load-bearing enough that a backend test checks it directly: no column in the audit table holds a score, and no route anywhere in the router even has rank, position, traffic, volume or difficulty in its path. What it has instead is a narrower, checkable claim — every audit of a page is stored, so the second run of the same URL can report exactly what changed since the first: a title edited, a canonical pointing somewhere new, a noindex tag that appeared out of nowhere. "Your pricing page went noindex on Tuesday" is a sentence that needs a before to be true, and a one-off scanner structurally has none. A stored history is a smaller promise than a score, and it is the one Khoji can actually keep.

What it refuses to do
There is no endpoint in Khoji's router whose path contains projected, guarantee or forecast, and a backend test asserts that directly against the mounted routes. A position is only ever one you checked yourself and logged, or one Google Search Console itself reported — never a number Khoji projected forward. Google states plainly that placement cannot be bought or guaranteed, so a tool that implied otherwise would be wrong, not optimistic.
Ubersuggest, Semrush and Surfer all lead with one composite number. Khoji's audit table has no score column at all — only error, warning and note counts, each one interrogable back to the specific issue that produced it. A single number with unstated inputs is exactly the kind of figure this product bans outright.
Generating a rating, a price or a name that a visitor cannot actually see on the page is refused outright, with a 400 naming what is unsupported. The same refusal rule also runs backwards over markup somebody else already wrote, which is where the damage usually already is.
Feed Khoji a paragraph of SEO advice from anywhere — an agency proposal, a freelancer's plan — and it flags a promised #1 ranking, a guaranteed traffic percentage, or a tactic named in Google's spam policies, with the reason attached to each finding.
A URL that resolves to an internal or private address is refused with a 400 before any request leaves the server — the same SSRF guard used elsewhere in the product, not a weaker copy written just for Khoji.
There is no endpoint in Khoji's router whose path contains projected, guarantee or forecast, and a backend test asserts that directly against the mounted routes. A position is only ever one you checked yourself and logged, or one Google Search Console itself reported — never a number Khoji projected forward.
How it works
Three steps, each usable the first time with nothing connected.
Post the URL. The first run is a checklist against seventeen on-page checks; every run after the first also returns what changed since the last one — a title edited, a page that went noindex, a status code that turned into a 404. Nothing is fetched twice for a comparison: the diff reads two stored rows.
Post the site's URL and Khoji reads the file, reporting per named crawler — Googlebot, Google-Extended, Bingbot, and seven AI crawlers including GPTBot and ClaudeBot — whether it is allowed in. Most robots.txt files were generated once and never read again; this is a two-second check that finds a blocked crawler nobody meant to block.
Post the schema type, the field values, and the page's own visible text. Leave out the page text and validation is skipped; supply it and any value not actually on the page is refused with a 400 naming which one.
Audit history
A one-off audit tool can only ever describe the page as it is right now, which is a checklist, not a warning system — it cannot tell you that something changed unless it happens to catch you looking at the exact moment it broke. Khoji stores every audit of a URL, so the second run onward compares against the last one and reports exactly what moved: a title edited, a canonical pointing somewhere new, a noindex tag that was not there before. Regressions are flagged separately from fixes, because the two need opposite reactions.
Robots and AI crawlers
robots.txt is usually written once, by a developer solving a different problem, and then never opened again — until a marketer notices organic traffic dropped and nobody can say why. Khoji reads it directly and reports, per named crawler, whether each one is allowed in: not just Googlebot and Bingbot, but the AI crawlers this category tends to ignore — GPTBot, ClaudeBot, and five others — because being invisible to those is a newer and quieter failure than being invisible to Google.
Schema markup
The fastest way into a structured-data manual action is a rating or a price in the markup that a visitor cannot actually see on the page — a stale review score left in a template, a price that changed and the JSON-LD did not. Khoji's generator will not emit a value the supplied page text does not contain, and the same rule now runs in reverse over markup somebody else already wrote, checking an existing page for exactly that mismatch.
What it does
No vendor, no per-lookup cost. Each card's label is derived from the backend's capability module rather than written here.
Fifteen schema types, each generated only when required properties are present and, if page text is supplied, only when the values actually appear on the page.
Works todaySeventeen on-page checks per URL, stored, with a before/after diff from the second run onward. A bounded, robots-respecting site crawl (capped at 100 pages, depth 4) now runs alongside it, closing four checks a single-page audit never could: broken internal links, redirect chains, orphan pages sitting in the sitemap but linked from nowhere, and site-wide duplicate titles.
Works todayFetching a sitemap and robots.txt, reporting every submitted URL robots.txt then blocks, and building a real sitemap.xml from one of your own crawls all run today with no credential. Whether Google actually indexed a page is a different claim — that needs the URL Inspection and Sitemaps-submit APIs behind the same Google Search Console connection google-search-console-insights uses, not yet connected on this deployment.
Needs a keyScans any block of text — an agency proposal, advice found online — for a promised ranking and for tactics named in Google's spam policies, naming which phrase triggered which finding. Saved drafts are now scored for real against a first-party checklist (keyword placement, density, length, heading structure) — a genuine, deterministic score never called a ranking score. Benchmarking that score against what is actually ranking needs a paid SERP data API this deployment doesn't have, so that route refuses honestly rather than faking a comparison.
Built, with limitsA canonical name/address/phone per workspace, with every citation entry auto-compared against it (whitespace, case and phone punctuation normalised) on every rescan. Entering what a directory currently shows is manual, by design — scraping directories one at a time is high-maintenance and prone to breaking silently. The comparison is automatic; the input is not.
Works todayNeeds no rank data at all: your own declared page-to-keyword mapping is the input, and two pages declared for the same term is the conflict. Re-checked on every read, with a primary-page recommendation from the most recent crawl's word counts when one exists — resolving is a person's call, and it never merges pages or silently suppresses the next scan.
Works todayBuilt entirely from your own crawl: a link graph (source, target, anchor text, in-degree per page) and suggestions matching an under-linked page to whichever other crawled page shares the most title words. A heuristic, not a model — no third party, no LLM, recomputed fresh from the crawl on every request.
Works todayGenerating and reading a report is complete with zero setup: it rolls up crawl issues, tracked-keyword snapshots, verified backlinks, synced Search Console totals and citation mismatches, and every section carries its own as-of date so a source that was never synced reads as stale, never as silently zero. Only scheduling delivery — sending a copy through email on its own — needs ZEPTOMAIL_API_KEY and a verified sender address this deployment hasn't set.
Needs a keyProjects, manually-entered candidate terms and shortlisting all work against your own workspace data today. Automatic expansion with real search volume, difficulty and intent needs a licensed keyword-data API (DataForSEO, Semrush, Ahrefs and similar) — no vendor is connected on this deployment, so expanding a topic refuses honestly rather than inventing a volume number.
Built, with limitsA competitor set (your domain plus up to 20 others, each validated) is stored and read for real. Generating the actual keyword-gap report needs a third-party index of competitor ranking keywords — a business cannot pull another domain's Search Console data, and no such vendor is connected, so a report comes back honestly marked blocked rather than an invented keyword list.
Built, with limitsTracked keywords, deletion and full history all work locally, and a person who checks Google themselves can log a real, dated position — genuine data, never a projection. Automated daily checking needs a paid or self-run SERP-fetch service; none is connected, so that route refuses and names the gap instead of inventing a position.
Built, with limitsA backlink your workspace already knows about is verified for real — Khoji fetches the claimed source page through the same SSRF-guarded client the audit tool uses and checks whether the target URL is still actually linked, raising new/lost alerts as that changes. This is verification, not discovery: finding backlinks nobody told Khoji about needs a paid backlink-index API, which isn't connected here.
Built, with limitsBriefs are stored per keyword and feedback on them works fully today. Generating the actual outline needs the pages currently ranking for that keyword and the real questions people ask around it — a SERP data API. Writing an outline from a language model's own recollection instead, with no vendor connected, would be exactly the ungrounded content Google's own spam-policy guidance warns against — so every brief is created honestly marked blocked rather than fabricated.
Built, with limitsLocal keywords and manually-logged pack positions (bounded to Google's own pack size, 1–20) work locally, and profile-check is a genuine first-party completeness score over fields you enter — name, address, phone, hours, category, photos — never framed as a ranking guarantee. Automated local-pack checking needs a SERP API with local-pack support, not connected here.
Built, with limitsStorage and the opportunity-extraction logic (which feature types a page can win, what format each wants) are real and exercised directly. Nothing populates the table on this deployment, though — it needs the same paid SERP data API rank tracking does, and with no vendor connected both routes honestly return empty rather than an invented opportunity.
Built, with limitsNot built yet
Listed flat, the gaps read as a product that does not work. Grouped by what actually blocks each one, they read as what they are: two free connections nobody has made, eight that need data Google does not sell to anybody, and three that need engineering.
Google publishes this data and gives it away. Both of these are built up to the boundary and waiting on a credential the owner creates in a few minutes at no cost — a Search Console property you already own, and a Google Cloud API key.
The one ranking-adjacent data source that is genuinely free — but it needs a dedicated OAuth connection (GOOGLE_SEARCH_CONSOLE_CLIENT_ID and GOOGLE_SEARCH_CONSOLE_CLIENT_SECRET, plus the workspace completing consent for a property it has already verified) this deployment does not hold yet. Once connected, real Search Analytics rows and rule-based findings — top query, zero-click-high-impression, striking-distance — come from them directly, with no model involved.
The real PageSpeed Insights v5 API, reading real-user field data via CrUX when Google has enough traffic on the URL and a lab run otherwise, labelled plainly which one it is, classified against Google's own thresholds. It answers unauthenticated, but Google throttles that far below anything a shared deployment can rely on — a free PAGESPEED_API_KEY from Google Cloud Console is needed for a usable rate limit.
Before you switch
Tell us the page you are worried about and we will show you the audit, the robots.txt check, and exactly what a second run reports.
Questions
Not automatically, and it cannot become able to by accident: There is no endpoint in Khoji's router whose path contains projected, guarantee or forecast, and a backend test asserts that directly against the mounted routes. A position is only ever one you checked yourself and logged, or one Google Search Console itself reported — never a number Khoji projected forward. You can log a position you checked yourself, with the date, and see its real history — that's genuine data, never a projection. Automated daily checking needs a licensed SERP provider this deployment does not have. Google publishes no API for rank, search volume, keyword difficulty or the link graph, to us or to anyone; every tool that shows you those numbers is reselling its own crawl, which is why we would rather name the gap than show you a number we made up.
Not for most of what Khoji does — the audit, the whole-site crawl, schema tools, robots.txt check, advice checker, citation checker, cannibalization detector and internal-linking suggestions all run with no credential. Two things are real code waiting on a connection this deployment hasn't made: a free Google Search Console OAuth connection would confirm whether Google actually indexed a page and unlock click data, and a free Google Cloud API key would unlock Core Web Vitals.
It is built specifically to avoid that. The generator refuses to emit a name, headline, price, rating value or review count that the page text you supply does not actually contain, returning a 400 naming which value is unsupported. A rating in the markup a visitor cannot see is the fastest route to a structured-data manual action, and this is the check that stops it.
It is reported, not hidden behind a tool failure. The single issue code for a fetch error carries the actual status code, and it is stored and diffed the same as any other finding — so a page that started 404ing since your last check shows up in the changes list.
Yes, within a bound: a robots.txt-respecting crawl capped at 100 pages and depth 4, run inline alongside the single-page audit. It's what finds broken internal links, redirect chains, orphan pages sitting in the sitemap but linked from nowhere, and site-wide duplicate titles — none of which a one-page audit can see. No third-party data or credential is needed for it.
Partially, and honestly about the rest. You can store a target domain plus up to 20 competitor domains today, and a gap report runs against them — but generating an actual list of keywords a competitor ranks for and you don't needs a paid third-party keyword index, since a business cannot pull Search Console data for a domain it doesn't own. With no such vendor connected, a report comes back marked blocked rather than an invented keyword list — the same honesty applies to verifying a backlink you already know about (real, works today) versus discovering ones nobody told Khoji about (needs a paid backlink index).