← Back to GEO Academy
Playbooks

Reading 'GEO: Generative Engine Optimization' (II): Evidence Beats Keyword Stuffing

The second reading of the KDD 2024 GEO paper examines GEO-bench, nine content methods, experimental results, and why citations, quotes, and statistics outperform keyword stuffing in its setting.

Published 07/25/2026 8 min read
GEO methodsGEO-benchAI search experimentsGenerative Engine Optimization

Reading 'GEO: Generative Engine Optimization' (II): Evidence Beats Keyword Stuffing

The first article examined how *GEO: Generative Engine Optimization* defines visibility. This second reading examines its experiments. The paper does more than propose a concept: it builds the large GEO-bench benchmark and evaluates several content-revision methods in one framework.

That shifts the discussion from guessing what AI "likes" to asking which changes improved visibility under particular metrics and conditions.

What GEO-bench contributes

The authors noted that no public query set was then designed specifically for generative engines, so they built GEO-bench with 10,000 queries: 8,000 for training, 1,000 for validation, and 1,000 for testing. Queries come from MS MARCO, ORCAS-1, Natural Questions, AllSouls, LIMA, Davinci-Debate, Perplexity.ai Discover, ELI5, and GPT-4-generated queries.

This mix includes not only conventional search queries, but questions requiring synthesis, reasoning, explanation, debate, or purchase judgment. For every query, the benchmark includes cleaned text from the top five Google results to simulate source inputs for a generative engine. GEO therefore concerns both the final response and the material the system is able to retrieve.

The nine revision methods

The paper tests: authoritative writing; quantitative statistics; query-keyword addition; reliable citations; credible quotations; simpler language; greater fluency; unique terminology; and technical terminology.

Evidence enrichment - statistics, citations, and quotations - gives an engine material with which to support an answer. Expression changes - simplicity, fluency, authority, and technical framing - chiefly affect how easily material can be understood, compressed, and restated.

Read the headline results carefully

Across GEO-bench, citations, quotations, and statistics were among the strongest methods. On Position-Adjusted Word Count, the best method reached a relative gain of roughly 41% over baseline; on Subjective Impression, the best method reached roughly 28%. The paper summarizes top-method gains as approximately 30-40% and 15-30% across the respective metric families.

Those figures are not a promise that adding numbers will reliably create a 40% increase for every site. They reflect a specific experimental setting, query set, and evaluation method. The defensible interpretation is that credible evidence can make material more useful for constructing a generated answer.

The cautionary result is equally important: conventional keyword reinforcement was weak and, on some measures, below baseline. In the paper's Perplexity.ai experiment, keyword reinforcement performed worse than baseline on Position-Adjusted Word Count. A generative engine does not simply match repeated phrases; it needs information that can support a conclusion and explain why it is warranted.

Apply the result without turning it into a formula

Do not make GEO a keyword checklist. A page still needs to establish its category, but repeated synonyms and thin automated pages do not create a credible source. Instead, add answerable evidence: capability definitions, version differences, pricing logic, cases, industry data, certifications, methods, comparisons, and FAQs.

Separate facts, opinions, and marketing language. State the basis for statistics, the conditions of a case, the scope of a feature, and the boundaries of each claim. Different query types call for different evidence: a procurement prompt may need cases and cost boundaries, while a trend prompt needs dated data and explicit attribution.

Borrow the experiment logic, not the laboratory wholesale

GEO Radar at https://www.georadar.top can help a team keep fixed recommendation, comparison, price, risk, and use-case prompts across platforms. Make one limited content change at a time - for example a case page, an FAQ clarification, a data definition, or a third-party evidence section - and then retest after an appropriate interval.

The important lesson is methodological: replace mystical rewriting with a baseline, an explicit change record, and repeatable observation.

Sources for this article