Gemini grounds its answers in Google Search, so get indexed first, keep Google-Extended allowed, and give Google clean entity signals it can match.
LLaunchScaler·Published ·8 min read
Gemini cites websites it retrieves through Google Search, so the first requirement is being indexed by Google; a page that is not in the index cannot be pulled into a Gemini answer. After that, two things decide it: the Google-Extended token in your robots.txt must allow Gemini to ground on your content, and your brand needs entity signals clear enough for Google to match your pages to the question.
The rest of this guide covers where Gemini appears inside Google Search, what Google-Extended actually switches off, why Gemini is the one assistant that sees JavaScript-rendered pages, and how to test whether it links to you.
How does Gemini choose the websites it cites?
Gemini answers from its model and from pages it retrieves at the time you ask. Google's explanation of how Gemini works calls this retrieval augmentation: when a user sends a prompt, Gemini "works to retrieve more relevant information from these external sources (e.g., Google Search)." The pages retrieved that way are the ones that can appear as sources.
In the Gemini app, those sources appear behind a Sources button at the bottom of the response or as links inside it. Google's help page is clear that this is conditional: "Not all responses include related links or sources." A response with no Sources button did not link any website, yours or anyone else's.
The developer version works the same way and shows the mechanics. In Grounding with Google Search, the model decides whether a search would improve the answer, generates and runs one or more search queries, then returns a response with citations that tie each cited span of text to a source URL. The website that gets cited is the one whose page answered one of those generated queries.
So the path into a Gemini answer runs through the Google index: indexed page, retrieved for a query Gemini generated, quoted in the answer.
Is Gemini in Google Search the same as the Gemini app?
They are related but separate. Google's AI Overviews and AI Mode run on a custom version of Gemini 2.5 inside Search, while the Gemini app is its own product, with its own help pages for computer, Android and iPhone. The controls differ, which matters when you are deciding what to block.
Questions, answered
What people ask about this
01
Does Gemini use Google Search to answer questions?
Yes. Google's own explanation of how Gemini works says it retrieves information from external sources such as Google Search, a process it calls retrieval augmentation. A page Google has not indexed cannot be retrieved that way.
02
Does blocking Google-Extended remove my site from Gemini?
Indexed and eligible for a snippet in Google Search
noindex, nosnippet, max-snippet, data-nosnippet
AI Mode
The AI Mode tab in Google Search
Same as AI Overviews
Same as AI Overviews
Gemini app
The Gemini app on the web and on phones
Indexed, and not opted out through Google-Extended
A Google-Extended disallow in robots.txt, or anything that removes the page from the index
Grounding with Google Search on Vertex AI
Apps that developers build on Gemini
Same as the Gemini app
Same as the Gemini app
If your goal is "show up when people ask Gemini inside Google Search," you are optimising for AI Overviews and AI Mode, where Google says there are no requirements beyond Search eligibility. If your goal is the Gemini app, Google-Extended becomes the deciding switch.
What does Google-Extended control for Gemini?
Google-Extended is a robots.txt product token, not a crawler. Google says it manages whether crawled content "may be used for training future generations of Gemini models" and "for grounding" in Gemini Apps and in Grounding with Google Search on Vertex AI. Google also states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal."
Google defines grounding here as "providing content from the Google Search index to the model at prompt time to improve factuality and relevancy." That is the retrieval step citations depend on. A site that disallows Google-Extended to keep its content out of Gemini training also keeps its pages out of the Gemini app's grounded answers.
Check your robots.txt for this group:
User-agent: Google-Extended
Disallow: /
If it is there and you want Gemini app citations, remove it or narrow it to the paths you want kept out, such as Disallow: /members/. Google-Extended has no user-agent string of its own, since Google crawls with its existing user agents and reads the token only as a control. You will never see "Google-Extended" in your server logs, so the robots.txt file is the only place to check. The Google-Extended guide covers the token's full scope.
Can Gemini see JavaScript-rendered content?
Yes, mostly, and that makes it the exception among AI assistants. Vercel's study of AI crawler traffic found that "none of the major AI crawlers currently render JavaScript," naming OpenAI's, Anthropic's, Meta's, ByteDance's and Perplexity's crawlers, while Gemini uses Googlebot's infrastructure, "enabling full JavaScript rendering."
Googlebot processes pages in three phases: crawling, rendering in a headless Chromium, then indexing the rendered HTML. A single-page app whose raw HTML is an empty <div id="root"></div> can therefore still reach Gemini, because Gemini retrieves from an index built on rendered pages.
Two cautions keep that from being an excuse. Rendering is queued, so Google may index a JavaScript page later than a server-rendered one. And every other assistant reads only the raw HTML, so a site that works for Gemini alone is invisible to ChatGPT, Claude and Perplexity. Google's own advice is that "server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript."
Which entity signals help Gemini describe your brand correctly?
Gemini has to connect a question to the right company, which is harder for a new or ambiguously named brand. The signals that help are the ones Google uses to disambiguate an organization: Organization structured data on the home page, one consistent name and description everywhere, and sameAs links that tie your official profiles to your site.
Google says Organization markup on the home page "can help Google better understand your organization's administrative details and disambiguate your organization in search results." It has no required properties, but these are the ones worth filling:
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Acme Scheduling",
"url": "https://www.acmescheduling.com",
"logo": "https://www.acmescheduling.com/logo.png",
"description": "Acme Scheduling is appointment software for clinics with 2 to 50 providers.",
"sameAs": [
"https://www.linkedin.com/company/acmescheduling",
"https://github.com/acmescheduling",
"https://x.com/acmescheduling"
]
}
Google asks that name match the name you use for your site name, and describes sameAs as "the URL of a page on another website with additional information about your organization." Then make the off-site copies agree. If your LinkedIn page says "scheduling platform," your directory listings say "practice management suite" and your home page says "appointment software," you have given the model three answers to "what is this company." Pick one sentence and use it on every profile you control. The Organization schema guide covers every property, and the entity SEO guide for a new brand covers the off-site side.
What stops Gemini from citing a page that is indexed?
An indexed page can still be skipped when something between the index and the answer gets in the way. The usual causes are a Google-Extended disallow, a canonical that points Google at a different URL, or an answer that exists on the page only as an image or behind a click. Check each before rewriting anything.
Cause
How to spot it
Fix
Google-Extended disallowed
A User-agent: Google-Extended group with Disallow: / in robots.txt
Remove the group, or limit it to private paths
Canonical points elsewhere
URL Inspection shows a Google-selected canonical that is not this URL
Make the canonical tag, redirects, sitemap and internal links name one URL
Answer is not text
The key fact sits in a screenshot, a PDF or a tab that loads on click
Put the answer in the HTML as a sentence
No direct answer
The page talks around the question for several paragraphs
Open the section with a 40 to 60 word answer, then give the detail
Wrong page competes
An old post outranks the page you want cited for the same question
Update the old post or redirect it to the current page
How do you test whether Gemini cites you?
Test with a fixed set of prompts every month and record what Gemini links. A single run is a sample of one; the pattern over several months is what tells you whether a change worked. Keep the prompts identical so that the only thing changing between runs is Gemini.
Write 15 to 20 prompts in three groups: category questions ("best scheduling software for small clinics"), comparison questions ("Acme vs the market leader") and brand questions ("what is Acme Scheduling").
Run each one in the Gemini app in a fresh conversation, once a month on the same day.
Open the Sources button and record every URL listed. Note whether your domain appears and which of your pages it is.
For brand questions, record whether Gemini's one-line description of your company is accurate.
Run the category questions in Google AI Mode as well and log those separately, since the two surfaces select sources differently.
The page Gemini links is as useful as the fact that it linked you. If it keeps citing a three-year-old blog post instead of your pricing page, that post is what Google retrieves for the question, and it tells you which page to update. The guide to measuring AI visibility covers how many runs make a result reliable.
Find out whether Gemini cites you, and what it says instead
The monthly log works, but it takes an afternoon each time. LaunchScaler's full audit runs the sampling once for $19, one time: it puts 10 buyer questions for your category to 7 answer engines, Gemini among them with ChatGPT, Perplexity, Claude, Copilot, Google AI Overviews and AI Mode, and reads all 70 answers to show which cite you and which competitors they name. It also checks the entity side that only the full audit reads: whether your home page carries Organization schema, and whether the one-line description each engine gives of your brand is accurate. It unlocks on your report after the free scan, which checks the robots.txt, noindex and rendering problems that stop retrieval in the first place. Run the free scan first, then open the full audit.
It stops Google using your content for Gemini training and for grounding in the Gemini app and in Grounding with Google Search on Vertex AI. It does not affect Google Search, AI Overviews or AI Mode.
03
Can Gemini read a JavaScript-rendered website?
Largely, yes. Vercel's crawler research found Gemini uses Googlebot's infrastructure, which renders JavaScript, while GPTBot, ClaudeBot and PerplexityBot do not. Server rendering is still safer, because Google says not all bots can run JavaScript.
04
Why doesn't Gemini show sources for my question?
Gemini only shows a Sources button when it provided links for that response. Google's help says not every response includes related links or sources.
05
How do I get Gemini to describe my company correctly?
Publish one plain description of what the company does and use it on your home page, your Organization structured data and every profile you control, with sameAs links tying those profiles to your site.
Allow PerplexityBot, answer questions where Perplexity looks (Reddit and niche directories), keep pages fresh and dated, and write self-contained answers.