Reading 'GEO: Generative Engine Optimization' (III): Use Research Results With Boundaries
The third reading of the KDD 2024 GEO paper covers domain differences, multi-page competition, combined strategies, its Perplexity experiment, and study limitations for a more robust GEO workflow.
Reading 'GEO: Generative Engine Optimization' (III): Use Research Results With Boundaries
A superficial reading of *GEO: Generative Engine Optimization* might produce an overly simple takeaway: add statistics, citations, and quotations to gain AI-search visibility. The latter part of the paper is more valuable because it explains why GEO is not a fixed script.
Results vary with domain, question type, source position, concurrent optimization of several pages, platform behavior, and time. Turning research into an operating practice requires retaining those boundaries.
Different domains need different evidence
The paper's label-level analysis finds that authoritative writing helps in debate, history, and science; citations help more with statement, fact, law, and government questions; statistics are particularly strong for law, government, debate, and opinion; and credible quotations suit people-and-society, explanation, and history questions.
The implication is not to deploy a single sitewide template. For enterprise procurement, cases, pricing, risk boundaries, and implementation detail may be pivotal. For an educational concept, definitions and explanatory structure matter more. For an industry trend, dated data, source attribution, and interpretation are essential.
A low-ranked source can still matter
The paper also evaluates multiple sources optimized simultaneously. With citations, the source in the fifth search position saw a 115.1% relative visibility improvement in one result table, while the first source declined 30.3%. Credible quotations and statistics also showed fifth-position gains of 99.7% and 97.9%, respectively.
This does not mean a lower-ranked page will always overtake a leading one. It means that a page supplying clear, verifiable, question-fit material can gain more explanatory space in a generated answer. SEO remains valuable because retrieval and position affect whether a page enters the source pool; GEO asks whether the page is then absorbed faithfully.
Combined work resembles real content practice
In a test-set subset, fluency plus statistics outperformed either strategy alone, and citations showed meaningful gains when combined with other methods. A credible page generally combines accurate facts, readable structure, evidence for key claims, dated and scoped data, and stated suitability limits.
Optimizing one signal in isolation can create AI-shaped copy rather than user-serving material. A page built solely from statistics may introduce trust and context problems; a polished page without evidence may be readable but replaceable.
The Perplexity experiment is a prompt to retest
In the paper's smaller Perplexity.ai validation, credible quotations performed best on Position-Adjusted Word Count and statistics performed best on Subjective Impression. Keyword reinforcement underperformed baseline on some measures.
That does not establish an eternal platform preference. It establishes a sensible operating rule: retest across platforms. ChatGPT, Gemini, Perplexity, Claude, Doubao, Tongyi Qianwen, Kimi, and DeepSeek differ in retrieval, source preference, citation style, and answer format. One platform's outcome is not a global rule.
The paper's stated limitations matter commercially
The authors note that the experiments cover only two generative engines, one a commercial system; engines and query distributions change; traditional-search ranking effects were not evaluated; and automatic labels can contain noise. Do not convert reported gains into a commercial guarantee or treat a single test as a durable conclusion.
Use GEO as continuous observation and content governance: observe, diagnose, revise, retest, and report.
A bounded implementation checklist
Build a fixed question set before editing content. Cover recommendation, comparison, buying concerns, price, industry use cases, risk boundaries, and localized needs. Select actions by domain: B2B pages need cases, integration, cost, and differences; local services need locations, coverage, qualifications, and review sources; consumer products need specifications, price, stock, service, and genuine review context.
Also record the quality of sources. Being named is only the first signal; the important question is whether official pages, media, community material, product listings, old content, and competitor sources are being mixed correctly.
GEO Radar at https://www.georadar.top can help teams conduct multi-platform visibility analysis, competitor comparison, repeatable question testing, and structured reporting. It does not guarantee a change in AI answers. It helps make changes visible and content decisions evidence-led.
The practical conclusion from the research is simple: write for real questions with credible, complete, citeable answers - not merely for an AI system.
Sources for this article
- arXiv, submitted November 16, 2023 and revised June 28, 2024, GEO: Generative Engine Optimization: https://arxiv.org/abs/2311.09735
- arXiv HTML full text, Analysis, GEO in the Wild, and Limitations: https://arxiv.org/html/2311.09735v3
- GEO project page, paper, code, and dataset: https://geo-optim.github.io/GEO/GEO
- Hugging Face, GEO-bench dataset: https://huggingface.co/datasets/GEO-optim/geo-bench