Generative engine optimization (GEO): what the research supports
GEO is optimizing content to be quoted inside AI answers. The 2023 paper that named it found citations, quotes and statistics lifted visibility 30 to 40%.
LLaunchScaler·Published ·8 min read
Generative engine optimization (GEO) is the practice of shaping your content so AI answer engines such as ChatGPT, Perplexity and Google's AI features quote and cite it inside their answers rather than only listing it as a link. The term comes from a 2023 paper by Pranjal Aggarwal and colleagues, which found that adding cited sources, quotations and statistics to a page lifted its visibility in generated answers by 30 to 40%, while keyword stuffing did not help.
Most GEO advice traces back to that paper, so this guide walks through what it tested, what worked, what did not, and where its findings stop.
What is generative engine optimization?
GEO is optimization for generative engines: systems that retrieve web pages and have a language model write an answer grounded in them. The paper that named it, "GEO: Generative Engine Optimization," by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, was posted to arXiv on November 16, 2023 as 2311.09735 and accepted to KDD 2024.
The authors describe generative engines as systems that "retrieve relevant documents from a database (like the internet) and use large neural models to generate a response grounded on the sources." Their concern was the website owner, who has "little to no control over when and how their content is displayed." They proposed GEO as "the first novel paradigm to aid content creators in improving their content visibility in generative engine responses."
The abstract's headline result: "GEO can boost visibility by up to 40% in generative engine responses," with the effect varying by domain.
How did the GEO study measure visibility?
The researchers built a test engine, a benchmark of 10,000 queries and two ways of scoring how visible a source is in an answer. Then they rewrote one of the five sources behind each answer with a given method and measured how much more of the answer came from it.
The setup, as the paper describes it:
The engine fetched the top 5 Google results for each query and had gpt-3.5-turbo write the answer, sampling 5 responses at temperature 0.7 "to reduce statistical deviations."
GEO-bench held 10,000 queries (8,000 train, 1,000 validation, 1,000 test) from nine sources, including MS MARCO, Natural Questions, ELI5, Perplexity's Discover page and queries written by GPT-4. It covered 25 domains and was 80% informational, 10% transactional and 10% navigational.
Questions, answered
What people ask about this
01
What is generative engine optimization?
Generative engine optimization, or GEO, is the practice of shaping web content so AI answer engines such as ChatGPT, Perplexity and Google's AI features quote and cite it inside their answers. The term comes from a 2023 research paper by Aggarwal and colleagues, published at KDD 2024.
Position-Adjusted Word Count scored a source by how many words of the answer came from it, weighted by how early they appeared.
Subjective Impression scored seven aspects of each citation, including relevance, influence, uniqueness and how likely a user was to click it.
For each query, one source was chosen at random and rewritten by each method in turn, and results were averaged over five random seeds.
The team also ran the methods on Perplexity itself, uploading the source texts as files for 200 test queries, since Perplexity does not let users specify source URLs.
Which GEO tactics worked, and which did not?
Three methods stood out: citing sources, adding quotations and adding statistics, each lifting visibility 30 to 40% on Position-Adjusted Word Count and 15 to 30% on Subjective Impression. Making text more fluent or easier to read helped 15 to 30%. A more persuasive tone did not help significantly, and keyword stuffing gave little to no improvement.
Method, as named in the paper
What it changed
Result reported in the text
Cite Sources
Added relevant citations from credible sources
30 to 40% gain on Position-Adjusted Word Count
Quotation Addition
Added relevant quotations from credible sources
30 to 40% gain; best method on Perplexity, 22%
Statistics Addition
Replaced qualitative discussion with quantitative statistics
30 to 40% gain
Fluency Optimization
Improved the fluency of the text
15 to 30% gain
Easy-to-Understand
Simplified the language
15 to 30% gain
Authoritative
Made the tone more persuasive and authoritative
"No significant improvement"
Keyword Stuffing
Added more keywords from the query
"Little to no improvement"; 10% worse than baseline on Perplexity
Unique Words
Added unique words wherever possible
Grouped with keyword stuffing under non-performing methods in the results table
Technical Terms
Added technical terms wherever possible
Grouped with the high-performing methods in the results table; not discussed in the text
The best methods improved on the baseline by 41% on Position-Adjusted Word Count and 28% on Subjective Impression. On Perplexity, Quotation Addition led with a 22% gain, and Cite Sources and Statistics Addition showed improvements "of up to 9% and 37% on the two metrics." The paper also tested combinations and reports that using Fluency Optimization and Statistics Addition together "results in maximum performance."
The authors' own reading of the keyword stuffing result: "techniques effective in search engines may not translate to success in this new paradigm."
Who benefits most from GEO?
Lower-ranked pages. The paper measured the gain separately for sources in each of the five search positions, and the evidence-adding methods helped the fifth-ranked source most while reducing the share of the first-ranked one. The authors conclude that "GEO is especially helpful for lower ranked websites."
Method
Change for the rank-1 source
Change for the rank-5 source
Cite Sources
-30.3%
+115.1%
Quotation Addition
-22.9%
+99.7%
Statistics Addition
-20.6%
+97.9%
For a new product whose pages sit fourth or fifth for a question, that is the practical point: evidence on the page can win a bigger share of the answer than the pages ranked above it. The effect also varied by topic. Statistics helped most for law and government, debate and opinion queries; quotations for people and society, explanation and history; and citations for factual statements.
How do you apply the GEO findings to your pages?
Add evidence where you make claims, and write clearly. Put a sourced number next to each important claim, quote a named person or document where one supports the point, link the primary source, and keep sentences plain. Do not repeat keywords. The steps below turn the paper's three best methods into page edits.
List the claims on the page that a buyer would want proof of: speed, savings, accuracy, compatibility, price.
For each, add the figure and where it comes from, such as your own measurement with its date and sample, or a public source with a link.
Where a named expert, customer or standard supports a point, quote it word for word and name the source.
Link each outside claim to the primary page it comes from, not to a page that repeats it.
Read the result aloud and cut anything that sounds like filler; the fluency and readability methods also helped.
Compare two versions of the same sentence for a hypothetical scheduling product. "Our scheduler saves clinics a lot of time" gives an engine nothing to quote. "Clinics using the scheduler cut no-shows from 18% to 11% over six months, across 240 clinics measured in March 2026" gives it a number, a timeframe and a scope. Only publish numbers you can stand behind; the method rewards evidence, and invented evidence is still invented. The guide to writing content AI answers cite covers passage structure alongside these edits.
What are the limits of the GEO research?
The paper tested a simplified engine in 2023, and the engines people use today are different. Its generator was gpt-3.5-turbo working from the top 5 Google results, its Perplexity check covered 200 queries through file uploads, and every experiment started after the source had already been retrieved. The findings describe what helps a page once it is in the pool, not how to get there.
Four limits are worth keeping in mind:
Retrieval comes first. Every test began with the page among the five sources. A page that an AI crawler cannot fetch, or that does not rank for the sub-question, never enters the pool. OpenAI, for example, says sites that block OAI-SearchBot "will not be shown in ChatGPT search answers."
Answers vary. The researchers sampled five responses per query to reduce statistical deviation, which is a reminder that one run of an AI answer is a sample, not a result.
The effect depends on the domain, as the authors stress; there is no single best method for every topic.
The engines have changed. The models, retrieval systems and citation displays of 2026 are not the ones tested, so treat the percentages as direction, not a forecast.
What does GEO add to SEO?
SEO earns a ranked link that someone clicks; GEO aims for the words inside the answer. They share a foundation, since a generative engine can only quote pages it can crawl and retrieve, and Google says the SEO fundamentals still apply to its AI features, with "no additional requirements" to appear in them.
SEO
GEO
The goal
A ranked result the searcher clicks
Being quoted and cited inside the generated answer
What decides it
Crawling, indexing, relevance and ranking signals
Retrieval first, then how quotable and well-evidenced the passage is
What the evidence says helps
Crawlable, indexable, helpful pages
Citations, quotations and statistics on the page (GEO paper); third-party mentions (Ahrefs)
What does not help
Tactics against Google's spam policies
Keyword stuffing, per the GEO paper
Ranking alone does not settle it. Semrush's study of 200,000 keywords found only about 20 to 26% overlap between AI Overview links and the organic top 10. Off-page signals matter too: Ahrefs' study of 75,000 brands found branded web mentions correlated 0.664 with ChatGPT visibility, and it cautions that correlation is not causation. The comparison of AEO, SEO and GEO sets the three terms side by side, the guide to optimizing for AI search turns them into a work plan, and the guide to measuring AI visibility covers how to tell whether any of it worked.
Check that AI engines can retrieve your pages first
Every GEO experiment started with the page already retrieved, so access is the first thing to confirm. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. It checks whether robots.txt and your firewall let OAI-SearchBot, Claude-SearchBot and PerplexityBot in, whether your main content is in the raw HTML they read, and whether the page has a concise answer passage an engine could lift. The $19 full audit adds the checks the paper's findings point to: statistics density, quotations and cited sources on your pages, and whether seven AI engines cite you at all.
It is the part of search optimization aimed at being quoted inside an AI-generated answer rather than only ranked as a link beside it. It builds on SEO, because an engine can only quote a page it has retrieved, and retrieval still depends on crawling and indexing.
03
Which GEO strategies actually work?
In the GEO paper's tests, adding cited sources, adding quotations and adding statistics each lifted a source's visibility by 30 to 40% on its main metric. Improving fluency and readability helped by 15 to 30%. Keyword stuffing gave little to no improvement.
04
Does keyword stuffing work for AI search?
No. The GEO paper found keyword stuffing offered little to no improvement in generative engine responses, and in its test on Perplexity it performed 10% worse than the unmodified baseline.
05
Is GEO different from AEO?
The terms overlap. AEO, answer engine optimization, usually refers to being the answer in answer boxes and AI features; GEO comes from the 2023 paper and focuses on visibility inside generated responses. Both rest on the same crawlable, quotable pages.