What is Google-Extended?
Definition
Google-Extended is a robots.txt product token that lets site owners control whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and the Vertex AI API for Gemini. It is not a separate crawler. Google states that it does not affect a site's inclusion in Google Search and is not a ranking signal.
Also known as: Google Extended, Google-Extended token, Gemini training opt-out

A preference switch, not a bot
What sets Google-Extended apart from other AI bots is that it has no crawler of its own. According to Google's crawler documentation, Google-Extended does not have a separate HTTP user-agent string; crawling is done with existing Google user agents. It is a product token used only in robots.txt, and it tells Google what the content it crawls may be used for.
The practical upshot: you will never see a request from "Google-Extended" in your server logs. The rule does not change how content is collected, only where it may be used afterwards.
What it controls
Per Google's documentation, Google-Extended governs:
- Training future generations of Gemini models that power Gemini Apps and the Vertex AI API for Gemini
- Use of the content for grounding (basing answers on sources) in those products
The same page states plainly that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal.
Where AI Overviews and AI Mode fit
AI Overviews and AI Mode are features of Google Search. Google's documentation on AI features says that to limit what is shown from your pages there, you use Search's own controls: nosnippet, data-nosnippet, max-snippet or noindex. That documentation points to Google-Extended for limiting training and grounding in "some of Google's other systems". Blocking Google-Extended is therefore not the tool for keeping content out of AI Overviews.
robots.txt examples
Keep all content out of these uses:
User-agent: Google-Extended
Disallow: /Exclude a section but allow one path inside it (modeled on Google's own example):
User-agent: Google-Extended
Allow: /archive/2024-report/
Disallow: /archive/The longer, more specific rule wins, so /archive/2024-report/ stays available while the rest of the archive is excluded.
Verifying the rule
Because Google-Extended never appears in logs, you cannot watch its effect in traffic; what you can check is the file itself. The robots.txt report in Search Console shows which version of your file Google fetched most recently and any issues it hit while parsing it. Make sure the token is spelled correctly and the rule sits in the group you intended.
A costly mistake
Blocking Googlebot to limit AI use removes the site from Google Search as well. Google-Extended is the right tool for the Gemini training and grounding preference; your Search visibility depends on Googlebot and the rules for it in robots.txt. For comparison with other providers' training bots, see GPTBot and AI crawler.

