What the GEO Research Actually Says About AI Search

tl;dr

Most "GEO hacks" for AI search visibility do not survive contact with realistic testing conditions. Keyword-stuffing backfires, llms.txt is explicitly ignored by Google, and even the strongest early finding (loading content with statistics and citations) loses most of its effect once researchers test it under real retrieval pipelines. What still works is unglamorous: named evidence, clean structure, and verifiable authorship.

You followed the checklist. You added an llms.txt file. You loaded your pages with statistics and named sources because an agency blog post told you that is what gets cited in ChatGPT and AI Overviews. Weeks later, your AI citation share has not moved, or it moved and then quietly reverted.

That is not a sign you did it wrong. It is a sign the checklist itself was built on a misreading of the research it cites.

This article was inspired by a closer read of the actual papers behind the current wave of "AI search optimization" advice, not the agency summaries of them. Here is the LinkedIn post that prompted this deeper dive:

Why Do Most GEO Tactics Fail Under Realistic Conditions?

Most tactics fail because they were validated on a single benchmark and then generalized past what that benchmark actually tested. C-SEO Bench, presented at NeurIPS 2025, found that most tested GEO methods are largely ineffective under realistic conditions, and several actively hurt ranking. The statistics-addition method, the single strongest lever in the original research, lowered ranking in 19 of 24 tested settings once the benchmark controlled for realistic retrieval behavior.

The deeper problem is structural, not tactical. The same study found that a document's position within the model's context window was a far stronger predictor of citation than any content-level trick. At full adoption, when every competitor applies the same optimization, gains decline toward a zero-sum, congested equilibrium. This is the part most GEO advice skips: these tactics operate in a competitive ranking system, not an absolute one. If it works because you found the trick first, it stops working the moment your category catches up.

Two more findings from adjacent studies confirm this pattern. SAGEO Arena, a February 2026 preprint, found that pure body-text optimization can actually hurt retrieval and reranking performance, while enriching structural fields (schema markup, meta descriptions, heading hierarchy) lifted retrieval by roughly 22%. A separate 2026 preprint on diagnosing and repairing citation failures found that 43% of topically relevant pages earn zero AI citations, most often because no passage on the page stands alone as a self-contained answer, a structural failure rather than a topical one.

What Did the Original GEO Research Actually Prove?

The original Generative Engine Optimization paper (Aggarwal et al., accepted to KDD 2024) tested a range of content interventions across a large, multi-domain query benchmark. Adding statistics and citing authoritative sources measurably increased visibility, by up to 40%, while keyword-stuffing failed outright. This is the one result in the entire GEO literature that is genuinely peer-reviewed and methodologically solid.

It is also the most misquoted finding in the space. A specific statistic (a claimed "+115% visibility gain for position-5 pages, -30.3% for position-1 pages") circulates constantly in agency content as if it were a new discovery. It is not new: it comes directly from this original 2024 paper, and most of the blog posts citing it never mention the source, let alone the follow-up research that narrowed its real-world applicability.

That narrowing matters. The original benchmark measured citation frequency in a controlled retrieval setup. It did not test what happens when ten competitors in the same category apply the identical tactic simultaneously, which is exactly the condition C-SEO Bench introduced a year later. The lesson is not that the original paper was wrong. It is that a lab result and a durable competitive advantage are not the same claim.

Why Doesn't llms.txt Do Anything for AI Search Visibility?

llms.txt does nothing for AI search visibility because the crawlers that matter do not read it. Google has stated multiple times, through Gary Illyes and John Mueller, that it does not use llms.txt and has no plans to, drawing an explicit comparison to the long-discredited keywords meta tag from early-2000s SEO.

That is not a hedge or a "wait and see" position; it is a direct statement from the platform with the largest share of AI-influenced search traffic.

Independent log analysis across more than 500 million AI-bot crawl events, tracked over a 90-day window, found that only a few hundred requests ever targeted the file directly. That is a negligible fraction of overall AI crawler traffic. If llms.txt has any future use, it is closer to agent-facing infrastructure for autonomous tools navigating a site than a search visibility lever.

The pattern to watch for

Every tactic in this article that turned out to be weak or false shares one trait: it was easy to implement, easy to sell as a checklist item, and never required anyone to produce new evidence, sources, or structure. That correlation is not a coincidence.

What Actually Moves AI Citation Share?

The tactics that hold up under repeated, adversarial testing are the same signals that have driven search authority for two decades: named evidence, clean structure, and verifiable authorship. None of it is a trick. It is the input quality that both classical search ranking and generative retrieval were built to reward.

The table below separates what current research actually supports from what remains marketing narrative.

LeverEvidence statusWhat it requires
Named statistics and cited sourcesEvidenced (original GEO paper, KDD 2024)Real data, attributed to a specific, checkable source
Structured schema (Article, FAQPage, HowTo)Preliminary (SAGEO Arena, 2026 preprint)Machine-readable markup, not just visual formatting
Answer-first paragraph structureEvidenced (extraction-pattern research)The direct answer in the first paragraph, not buried after preamble
Freshness and recencyEvidenced, with a caveatAI-cited content skews measurably fresher; citation likelihood drops sharply after roughly three months without a refresh
Author entity and E-E-A-T signalsEvidenced, expanded since the December 2025 core updateVerifiable credentials, consistent author attribution, primary-source claims an LLM cannot synthesize from training data alone
Keyword density and stuffingRefutedN/A, actively harmful
llms.txtRefutedN/A
Statistics-loading at full market adoptionEffect shrinks toward zeroDiminishing as a category-wide tactic

The E-E-A-T expansion is worth pausing on. Since the December 2025 core update, Google has extended E-E-A-T signal weighting beyond strictly YMYL (Your Money or Your Life) categories into general content. In practice, that means named-expert quotes, primary-source claims, and information a language model cannot simply synthesize from its training data are increasingly what separates cited content from ignored content, regardless of topic.

How Should You Rebuild Your Content Process Around This?

Rebuild around substance, not a new set of tricks. The real difference between content that gets cited and content that gets ignored is not a better GEO checklist; it is whether the piece says something specific and true, or just restates what every other page in the category already says. Language models default to the most common phrasing in their training data, so anything generic blends in by construction.

A practical framework, in order of leverage:

  1. Start from a real point of view. Before generating anything, get clear on what you actually claim and how you back it up. That is the input a generic AI writing tool never has, which is why its output reads like the category average.
  2. Ground every claim in a checkable source. Every statistic, quote, and factual claim should trace back to a source that survives scrutiny. This is both a citation-probability lever (per the original GEO research) and a brand-safety requirement: an unverified number in corporate content is a reputational liability, not just a missed optimization.
  3. Structure for extraction, not just readability. Lead with the answer, use a clear H2/H3 hierarchy, and mark up content with the appropriate schema. Large language models extract at the passage level; most citations come from content in the first third of a page.
  4. Treat freshness as a maintenance task, not a one-time step. Given the roughly three-month citation cliff, plan refresh cycles into the calendar rather than leaving evergreen content to decay silently.
  5. Measure like brand media, not performance marketing. AI citation counts fluctuate 40 to 60% month over month across most monitoring tools, and no reliable cross-channel attribution model exists yet. Track reach, topical presence, and branded-search lift, not a fabricated ROI number.

Common Questions

Does GEO replace SEO, or work alongside it?

Alongside it. SEO governs crawlability, indexing, and organic ranking; AEO and GEO layer on top for direct-answer formats and generative citation. The same underlying signals (authority, structure, content quality) feed all three, so a strong SEO foundation is a prerequisite, not an alternative.

Is llms.txt worth implementing at all?

Not for search visibility. The evidence is clear that major AI crawlers do not read it and Google has stated it never will factor into ranking or citation. It may eventually matter as infrastructure for autonomous agents navigating a site programmatically, but that is a different use case entirely from search optimization.

Why did the "+115% GEO stat" spread so widely if the research does not fully support it?

Because it originated in a real, credible paper (Aggarwal et al., KDD 2024) and got detached from its context as it moved through secondary sources. The underlying finding, that statistics and citations help, is real. The specific number and its generalization to "always works" is not what the original researchers claimed.

How much should I trust AI-visibility monitoring tool numbers?

Treat them as directional, not precise. Different tools return different citation counts for identical prompts, month-over-month swings of 40 to 60% are common, and no tool currently offers reliable, auditable attribution from citation to business outcome. Use them to spot trends, not to justify a specific ROI figure.

What is the single highest-leverage change to make first?

Get clear on your actual point of view and evidence before you generate anything. Every tactic in this article compounds on top of that and does close to nothing without it, because content with nothing specific to say still reads like the category average, however well it is structured.

Key Takeaways

  • Keyword-stuffing and llms.txt are not underperforming tactics; the research and Google's own statements show they do not work at all.
  • The strongest documented GEO lever (statistics and named citations) is real but shrinks toward zero once a whole category adopts it, because generative ranking is a zero-sum system.
  • Structural signals (schema, heading hierarchy, answer-first paragraphs) matter as much as content-level tricks, and sometimes more.
  • E-E-A-T and named authorship have expanded beyond YMYL content since the December 2025 core update and are now a general ranking and citation factor.
  • Treat AI search visibility as a brand-media metric (reach, topical presence, branded-search lift), not a performance-marketing channel with clean attribution, because that attribution does not exist yet.

This article was inspired by content originally written by Mario Ottmann. The long-form version was drafted with the assistance of Claude Code AI and subsequently reviewed and edited by the author for clarity and style.