ContentGrapher
ContentGrapher

Research

Empirical studies on content structure and AI retrieval.

The model language transfer studyPublished
July 20266 models · 60 production pages · 8 translated pairs · 3 controls

Does a model's competence on English survive a change of language?

We expected six models to lose competence on non-English pages and built three controls to prove it: each model measured against its own English self, every failure re-run to separate a real fault from a flaky one, and the same eight pages run in English, French, and German.

Key findingZero of five open models showed the predicted drop. Five apparent language effects, one of them our own measurement, dissolved under the controls into dropped connections, unequal page sets, serving noise, and a non-reproducing panel lean. Reliability is holding the output contract; language told us nothing.
6 modelsWithin-model controlFailure attributionNull result
Read study
The presence premium studyPublished
June 202645 pages · 398 ablations + 60 leave-one-out · cross-family panel · control-validated

Does adding a missing concept help a page get retrieved more than deepening a present one?

We tested whether closing a concept gap beats deepening a thin mention for on-page answerability, across three natural concept states plus a controlled leave-one-out experiment, on 45 real pages with a non-Anthropic judge panel.

Key findingNo presence premium in the score's flags: developing a flagged concept rarely helps (missing 17%, thin 10%, developed 8%), because ~62% of flagged-missing concepts are already answerable from the page. But a controlled experiment, removing a concept the page relied on and adding it back, lifts answerability 37% of the time versus 10% for deepening, and 82% when the removal actually broke the answer. Presence matters once the gap is real; the lever is flag precision, not the presence-versus-depth weighting.
Most 'missing' flags are already answerableNo presence premium in the flagsReal gaps matter (controlled test)Lever is flag precision, not reweighting
Read study
The depth necessity studyPublished
June 202625 topics · 100 pages · 802 ablations · cross-family panel · control-validated

Should a sales page be scored more leniently than a blog for missing depth?

We tested whether the explanatory depth flagged as missing on commercial pages actually helps those pages get retrieved, by developing each flagged concept and measuring on-page answerability with a non-Anthropic judge panel.

Key findingNo retrieval case for scoring commercial pages on a curve: the flagged out-of-place depth was about as necessary (23%) as in-scope depth (15%), the wrong direction, and a type-aware scorer would hide a genuinely useful concept in 23% of exclusions. We kept the score type-blind. Necessity read low everywhere, including the blog control, hinting that deepening already-present concepts buys little retrievability on any page type.
No format artifact to correctType-aware scoring not justifiedScore stays type-blindDeepening present concepts: low payoff
Read study
The page type studyPublished
June 202625 topics · 4 page types · 100 pages · 300 runs · cross-family panel

Does the tool expect the same things of a product page and a blog?

We held the topic constant and varied the page type (product, category, blog, landing) across 25 topics and 100 real pages, ran each three times, and compared what the tool inferred, scored, and flagged across types.

Key findingPage type almost completely changes the expected-concept set: a product page and a blog about the same topic overlap by just 3%, while re-running one page reproduces 99%. The coverage score crossed a band across the four types on every topic. There is consistent directional evidence the tool over-flags commercial pages for out-of-place depth (product and landing flagged far more than blog), but the two judging panels agreed on direction, not on the exact rate, so we report it directionally.
Type rewrites the expected setScore moves with typeDirectional, not calibratedRepeat-run noise floor
Read study
The concept validity studyPublished
June 202640 native pages · 8 languages · 5 concepts · single Opus 4.8 judge

Do our five core concepts hold for content written for non-English readers?

We ran the pipeline on 40 pages written natively in eight non-English languages, excluding translations and English-origin global brands, then had Claude Opus 4.8 judge, concept by concept, whether each of the five structural signals genuinely applied and was read correctly.

Key findingAll five concepts held in every language: 40 of 40 language-by-concept verdicts came back go, and only 15 of 200 individual judgments slipped below a clean go, none enough to change a result. The harder problem was reach. French, Chinese, Japanese, and Dutch publishers blocked the crawler on half to three-quarters of the pages we tried, and French and Chinese ran out of reachable pages before hitting five. The structural reading travels; getting to native content is the bottleneck.
All five go in 8/8Native-authoredCrawler access is the wallSingle Opus judge
Read study
The sufficiency studyPublished
June 202630 pages · 166 questions · panel + 20-pair human check

When a chunk is retrieved, can it actually answer the question?

We retrieved the top five chunks for 166 real questions across 30 real pages, then asked a cross-family panel, blind to the score, whether each chunk could answer the query on its own, and checked the panel against a human.

Key findingAI reads a slice of your page, not the whole thing, and most slices cannot stand alone. How many is judge-dependent: a model panel said 89% of retrieved sections could not answer without leaning on the rest of the page, a human nearer 60%, and they disagreed enough that the gap is part of the finding. Either way, being close to the query is what those failing sections already did. The lever is depth, one concept covered fully in one self-contained section, not a nearer keyword match.
Proximity ≠ sufficiencyJudge vs human gapCross-family panelBootstrap CI
Read study
The language studyPublished
June 202612 pages · 5 languages · 3-layer decomposition · multilingual panel

Can ContentGrapher read non-English content?

We ran the same explanatory content through the tool in English, French, German, Japanese and Korean, and had reviewers fluent in each language check the result. Then we separated what users see from what the analysis actually understood.

Key findingThe tool understands non-English content as well as English (reviewer accuracy ≥ English), but the score it reports drops 0.15 to 0.23. The cause is one scoring step that checks for concept names the analysis writes in English and cannot find in non-English text. Today, treat non-English scores as unreliable; the analysis is sound.
5 languagesBootstrap CIMultilingual panelCapable, score broken
Read study
The personalisation studyPublished
June 202612 topics · 4 architectures · 4 retrieval stacks · 192 readers

At equal information, does page architecture change what AI retrieves?

We held the information on the page fixed and split 12 topics across pages four ways, then asked whether 192 readers got their question answered at their own level and stage, across four retrieval stacks.

Key findingSplitting did not help. The boundary-split cluster did not beat a single combined page on any stack, and no architecture rescued the long tail. What moved the outcome was the reader's level: beginners were answered 65% of the time, advanced readers 19%. Splitting earns its keep when it makes you add depth, not when it re-files depth you already have.
Null result4 retrieval stacksBootstrap CILevel beats layout
Read study
The architecture studyPublished
June 202620 pages · 3 conditions · 160 questions · cross-family panel

Its own page, or a section on the page?

We developed one under-covered concept three ways, as a dedicated page and as a section on the original page, then asked AI 160 real questions to separate the page boundary from the content.

Key findingA dedicated page and a developed section on the original page were equally findable on narrow questions (both 70% vs 15% for the source). The URL boundary added nothing. The dedicated page also lost broad-topic coverage (13% vs 57%).
3 conditionsBootstrap CIHuman-calibratedB ≈ C
Read study
The AIO citation studyPublished
June 202622 queries · 217 pages · 4 metrics

Does Google AIO cite structurally complete pages?

We measured ContentGrapher structural completeness on 217 pages: 135 cited in Google AI Overviews and 82 uncited organic pages ranking for the same queries. Then we checked whether the cited pages scored higher.

Key findingNo significant difference. Coverage score gap: 0.011. Direction flips in 6 of 17 qualifying queries. AIO citation selection does not favor pages that score higher on structural completeness.
Observational corpusBootstrap CINull result
Read study
The audience studyPublished
June 202660 pages · four conditions · a counterbalanced panel

Does telling the tool who is reading change what it says?

ContentGrapher asks who the reader is before it analyzes a page. We ran 60 pages four ways to test whether that input earns its place.

Key findingThe audience changes priority emphasis, not coverage, and the knowledge level does the work: a wrong role disrupts the ranking nearly as much as the right one sharpens it. Once a position bias was removed, a blind panel could not tell the versions apart.
Four-condition designCounterbalanced panelWilcoxon tests
Read study
The translation studyPublished
June 20264 models · 3 families · 12 pages · 2 conditions

Can a model follow a structural recommendation?

We gave four writing models ContentGrapher's structural recommendations and asked each to rewrite the same pages, then scored the rewrites against a free rewrite with no recommendations.

Key findingEvery model raised the score with the recommendations, but only GPT-4.1 gained more from them than from a free rewrite. The others reached the same score by padding, whether or not they read the list.
Recommendation premium4 models, 3 familiesBootstrapped CIs
Read study
The reliability studyPublished
June 20265 challengers · a 5-model panel · 3 reruns

When two AI models disagree, is it real?

We tried to replace the model behind one of our scope calls with a cheaper one. A challenger looked fifteen points better, so we checked whether the difference was real.

Key findingThe two models overlapped less than the shipped model overlaps with itself. Most of the apparent disagreement was noise, not a better model, and the judges were never checked against a person.
Model comparisonSelf-consistencyFive-maker judge panel
Read study
The findability studyPublished
June 202630 pages · 166 questions · 3 AI reading systems

When an idea gets its own page, does AI find the answer?

We took 30 real pages where ContentGrapher said an idea deserves its own page, built that page, and asked 166 real search questions against a matched look-alike control.

Key findingWith the recommended page in place, AI found the answer to 84% of the questions. Without it, 4%. The recommended page won 164 of 164 head-to-head comparisons.
Controlled trial3 embedding modelsBootstrapped CIs
Read study
The agreement studyPublished
June 20268 models · 49 pages · 2 passes

Do AI models agree on what belongs on a page?

Eight models judged concept scope on the same 49 real pages, twice each, with a five-maker review panel checking the calls.

Key findingThe “belongs elsewhere” rate splits into three stable clusters, from 15.9% down to 1.9%, and the split does not follow the open-source versus closed divide.
8 models, 6 makersTest-retestFive-maker review panel
Read study
The decoy studyPublished
June 2026n = 40 pages

Does it matter which gap you fill?

A controlled test of structural completeness and AI retrieval across 40 third-party pages, with a decoy control arm.

Key findingOn pages with 5 or more flagged gaps, specific picks outperformed random structural addition by +11.2pp. On pages with 2 or fewer, no measurable difference.
Controlled trialChromaDB + OpenAIGPT-4o judge
Read study