Does covering more topics get my page cited by AI?
We reviewed the largest public dataset on AI citation and coverage. AirOps scraped 16,851 queries and 353,799 pages from ChatGPT. We read their numbers against the retrieval literature.
Empirical studies on content structure and AI retrieval.
We reviewed the largest public dataset on AI citation and coverage. AirOps scraped 16,851 queries and 353,799 pages from ChatGPT. We read their numbers against the retrieval literature.
We expected six models to lose competence on non-English pages and built three controls to prove it: each model measured against its own English self, every failure re-run to separate a real fault from a flaky one, and the same eight pages run in English, French, and German.
We tested whether closing a concept gap beats deepening a thin mention for on-page answerability, across three natural concept states plus a controlled leave-one-out experiment, on 45 real pages with a non-Anthropic judge panel.
We tested whether the explanatory depth flagged as missing on commercial pages actually helps those pages get retrieved, by developing each flagged concept and measuring on-page answerability with a non-Anthropic judge panel.
We held the topic constant and varied the page type (product, category, blog, landing) across 25 topics and 100 real pages, ran each three times, and compared what the tool inferred, scored, and flagged across types.
We ran the pipeline on 40 pages written natively in eight non-English languages, excluding translations and English-origin global brands, then had Claude Opus 4.8 judge, concept by concept, whether each of the five structural signals genuinely applied and was read correctly.
We retrieved the top five chunks for 166 real questions across 30 real pages, then asked a cross-family panel, blind to the score, whether each chunk could answer the query on its own, and checked the panel against a human.
We ran the same explanatory content through the tool in English, French, German, Japanese and Korean, and had reviewers fluent in each language check the result. Then we separated what users see from what the analysis actually understood.
We held the information on the page fixed and split 12 topics across pages four ways, then asked whether 192 readers got their question answered at their own level and stage, across four retrieval stacks.
We developed one under-covered concept three ways, as a dedicated page and as a section on the original page, then asked AI 160 real questions to separate the page boundary from the content.
We measured ContentGrapher structural completeness on 217 pages: 135 cited in Google AI Overviews and 82 uncited organic pages ranking for the same queries. Then we checked whether the cited pages scored higher.
ContentGrapher asks who the reader is before it analyzes a page. We ran 60 pages four ways to test whether that input earns its place.
We gave four writing models ContentGrapher's structural recommendations and asked each to rewrite the same pages, then scored the rewrites against a free rewrite with no recommendations.
We tried to replace the model behind one of our scope calls with a cheaper one. A challenger looked fifteen points better, so we checked whether the difference was real.
We took 30 real pages where ContentGrapher said an idea deserves its own page, built that page, and asked 166 real search questions against a matched look-alike control.
Eight models judged concept scope on the same 49 real pages, twice each, with a five-maker review panel checking the calls.
A controlled test of structural completeness and AI retrieval across 40 third-party pages, with a decoy control arm.