Google-Extended: what blocking it does, and doesn't, do
Google-Extended is a robots.txt token, not a crawler. It opts you out of Gemini training and Gemini app grounding, not Search or AI Overviews.
LLaunchScaler·Published ·7 min read
Google-Extended is a robots.txt token that lets you opt out of Google using your content to train Gemini models and to ground answers in the Gemini app and in Vertex AI. It is not a crawler, it does not change how Googlebot crawls you, and Google states it does not affect inclusion or ranking in Google Search, which means it does not remove you from AI Overviews or AI Mode either.
That last point matters most. A site that adds a Google-Extended block to leave AI Overviews will see no change there, because Google did exactly what the token says; AI Overviews are governed by different controls.
What is Google-Extended?
Google-Extended is what Google calls "a standalone product token." You name it in robots.txt like a crawler, but no crawler by that name exists. Google says it "doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity."
Googlebot fetches your pages as it always has, and Google reads your Google-Extended rules to decide what it may do with those pages afterwards. Two consequences follow:
You will never see "Google-Extended" in your server logs, so you cannot confirm it by watching traffic. The robots.txt file is the only record.
Blocking it does not reduce crawl load. Googlebot still visits, because Search still needs your pages.
Google's documentation gives this example group, which blocks one directory while allowing a single path inside it:
It tells Google not to use your content for two things: training future Gemini models, and grounding answers in the Gemini app and in Grounding with Google Search on Vertex AI. Google's wording is that the token manages whether crawled content "may be used for training future generations of Gemini models" and "for grounding" in those products.
Google defines grounding in the same sentence: "providing content from the Google Search index to the model at prompt time to improve factuality and relevancy." That is the step where the Gemini app looks up current pages to answer a question and links them as sources. Blocking Google-Extended takes your pages out of that step, so the Gemini app can no longer ground an answer on them.
Questions, answered
What people ask about this
01
What is Google-Extended?
Google-Extended is a robots.txt product token that tells Google whether content it crawls from your site may be used to train Gemini models and to ground answers in Gemini Apps and in Grounding with Google Search on Vertex AI. It is not a separate crawler.
02
Does blocking Google-Extended hurt my Google rankings?
Training future Gemini models (Gemini Apps, Vertex AI API for Gemini)
Yes
Grounding in the Gemini app
Yes
Grounding with Google Search on Vertex AI
Yes
Google Search results and rankings
No
AI Overviews
No
AI Mode
No
How often Googlebot crawls you
No
What doesn't Google-Extended do?
It does not touch Google Search. Google's documentation says it plainly: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." Your rankings, your snippets, your appearance in AI Overviews and in AI Mode all carry on as before.
It also does not reach other AI products. It is Google's control for Google's Gemini products, so OpenAI's GPTBot, Anthropic's ClaudeBot and Common Crawl's CCBot each need their own robots.txt group. Apple has the same pattern with its own token: Apple says Applebot-Extended "does not crawl webpages" and that pages disallowing it "can still be included in search results" in Apple's products.
Does Google-Extended remove you from AI Overviews?
No. AI Overviews and AI Mode are features of Google Search, and Google's AI features guide says "robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search." The same guide points to Google-Extended only "to limit AI training and grounding in some of Google's other systems."
To limit what AI Overviews use, the controls are the Search snippet controls:
You want
Use
Cost
A page not quoted in AI Overviews or AI Mode
nosnippet in a robots meta tag or X-Robots-Tag header
No text snippet in regular results either
Less of a page quoted
max-snippet:[number]
Shorter snippets everywhere in Search
Part of a page not quoted
data-nosnippet on a span, div or section
Only that part is excluded
A page out of Search entirely
noindex
Gone from Search, AI Overviews and AI Mode
Google's robots meta documentation confirms nosnippet prevents content being used "as a direct input for AI Overviews and AI Mode," and max-snippet limits how much is used. Every one of these also changes regular search results, which is why there is no free way out of AI Overviews. The guide to opting out of Google AI Overviews walks through the trade-offs.
How do you write the Google-Extended rule?
Add a group for the token to your robots.txt at the root of each host. To opt out completely, disallow everything; to opt out partly, disallow the paths you want kept out. Matching is case-insensitive, but the spelling must be exact, or the rule silently matches nothing.
Full opt-out:
User-agent: Google-Extended
Disallow: /
Partial opt-out, keeping your docs and blog available to Gemini grounding while withholding everything else:
Under RFC 9309, when rules conflict the most specific match wins, measured by the length of the path, so /docs/ beats / for URLs under /docs/. After saving, fetch the live file with curl -s https://yoursite.com/robots.txt to confirm the group is there, since a CDN or framework can serve a different file from the one in your repository. There is nothing else to test: with no user agent of its own, Google-Extended cannot be requested or seen.
How do you find out whether you already block it?
A site can block Google-Extended without anyone deciding to, because a CDN or a plugin wrote the rule. Cloudflare's managed robots.txt, for one, prepends its own block of rules, Google-Extended included, in front of your file. Read the live file each host serves rather than the one in your code.
Fetch the served file for every host you care about: curl -s https://www.yoursite.com/robots.txt, then the same for docs., blog. and any other subdomain, since each host has its own file.
If a group appears that is not in your repository, look for a line such as # BEGIN Cloudflare Managed content, which marks rules Cloudflare added.
Check the User-agent: * group too. Under the robots.txt standard, a token with no group of its own follows the * group, so a file that disallows everything for * and opens only a Googlebot group has, in effect, opted out of Google-Extended as well.
Record what you find before changing anything, so you know which switch produced the rule if it reappears after a deploy.
Should you block Google-Extended?
Block it if keeping your content out of Gemini training matters more to you than being a source in the Gemini app. Leave it open if you want the Gemini app to cite your pages. Either way, your Google Search visibility, AI Overviews included, stays the same, so the decision is narrower than most people assume.
A product site that wants to be recommended usually leaves it open, because the Gemini app is one more assistant buyers ask for recommendations, and grounding is how it finds current pages. A publisher that licenses its archive, or a company with legal reasons to keep content out of model training, has good reason to block it. The guide to whether you should block AI crawlers compares this trade-off across every operator, and the guide to getting cited by Gemini covers the Gemini app side.
What about Cloudflare's AI training settings?
Cloudflare offers two switches that touch Google-Extended, and they behave very differently. Its managed robots.txt adds a User-agent: Google-Extended group with Disallow: /, which is the targeted opt-out described above. Its AI bot policies block traffic at the firewall, and a Training block there reaches further than robots.txt does.
Cloudflare's documentation says that from September 15, 2026, "Mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training, including the legacy 'Block AI bots' option." Its July 2026 announcement names the crawlers it means: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training."
So if your goal is "no Gemini training, but stay in Google Search," the robots.txt token does exactly that, and a firewall-level Training block may also stop Googlebot. Check Security Settings in the Cloudflare dashboard before assuming a training block is harmless. The robots.txt guide for AI bots has a full file that blocks training crawlers by token and leaves search crawlers open.
Check which AI rules your site is sending
A Google-Extended line is easy to verify by reading robots.txt; the rules around it are where sites go wrong, such as a wildcard block that catches search crawlers or a firewall that challenges them. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. Its AI visibility checks read robots.txt for the search crawlers that feed citations (OAI-SearchBot, Claude-SearchBot and PerplexityBot), pass a training-only block as a legitimate choice, fail a wildcard rule that blocks search crawlers by accident, and test whether a firewall challenges them. Its search checks flag noindex and robots.txt rules that block pages you want in Google.
No. Google states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal in Google Search.
03
Does Google-Extended control AI Overviews?
No. AI Overviews and AI Mode are part of Google Search, and Google says robots.txt directives for Googlebot are the control for Search. To limit what AI Overviews use, use nosnippet, data-nosnippet, max-snippet or noindex.
04
What is the Google-Extended user agent?
There isn't one. Google says Google-Extended has no separate HTTP user-agent string; crawling is done with Google's existing user agents and the token is only read as a control in robots.txt, so it never appears in your server logs.
05
Does blocking Google-Extended stop Gemini from citing my site?
In the Gemini app, it can. Google says the token covers grounding in Gemini Apps, which it defines as providing content from the Google Search index to the model at prompt time. It does not affect AI Overviews or AI Mode in Search.
ChatGPT answers from live search or from what the model already knows. How to get into each: allow OAI-SearchBot, serve HTML, earn listings, then measure.