CrawlRadar
Google logo

Google · policy directive

Is Google-Extended blocked on your site?

Google-Extended is Google’s policy directive token for Gemini training and grounding, and whether it reaches you is decided by one line in your robots.txt. It is a robots.txt directive, not a crawler: no request ever arrives under this name. Disallowing it opts you out of Gemini training and grounding. It does not affect Google Search, including AI Overviews.

  • Not critical if blocked
  • Google-Extended robots.txt token

By · Published · Last updated

What is Google-Extended?

Google-Extended is a standalone robots.txt product token, not a crawler. Google crawls with its existing user agents and reads Google-Extended to decide whether that content may be used to train future Gemini models and to ground answers in Gemini Apps and Grounding with Google Search on Vertex AI. Google states that it does not affect a site’s inclusion in Google Search.

OperatorGoogle
Purposepolicy directive · Gemini training and grounding
robots.txt tokenGoogle-Extended
In the user-agentNever sent: this token is only read from robots.txt
CrawlRadar weight1 of the crawler-access weights, never flagged critical
Vendor docsGoogle: Google’s common crawlers

How to allow Google-Extended in robots.txt

Give Google-Extended its own group with Allow: /. A named group overrides the User-agent: * group, so this works even when the wildcard disallows everything:

User-agent: Google-Extended
Allow: /

How to block Google-Extended in robots.txt

Replace Allow with Disallow. To block only part of the site, disallow that path instead of /:

User-agent: Google-Extended
Disallow: /
  • Tokens are matched case-insensitively, but writing Google-Extended exactly as the vendor spells it avoids surprises with stricter parsers.
  • Check your CDN and security plugins too: a firewall rule can block Google-Extended without any line in robots.txt.

Should you block Google-Extended?

Block it if you want to opt out of Gemini training and grounding; it costs you nothing in Google Search. It is not the control for Google AI Overviews or AI Mode: those are Search features built from Googlebot’s crawl, and the way to limit how Search uses your content, AI features included, is snippet controls such as nosnippet and max-snippet.

How to verify a request is really Google-Extended

You cannot, because no request ever carries the Google-Extended name. Google fetches pages with its regular crawlers and reads Google-Extended from robots.txt as a permission. A log line claiming to be Google-Extended is not from Google.

Is Google-Extended visiting your site?

Your analytics will not tell you. AI crawlers do not run JavaScript, so a tag-based tool such as Google Analytics never records them. Google-Extended itself never visits, but Google’s crawlers do, and CrawlRadar’s collector records them from your own server. See the full AI crawler directory for the other tokens CrawlRadar checks.

Frequently asked questions

Does blocking Google-Extended remove me from AI Overviews?

No. AI Overviews are part of Google Search and are built from Googlebot’s crawl, and Google documents that Google-Extended does not affect inclusion in Search. To limit how Search shows your content, including in its AI features, use snippet controls such as nosnippet or max-snippet.

Why can I not find Google-Extended in my server logs?

Because it never visits. Google states that Google-Extended has no HTTP user-agent string of its own: Google’s regular crawlers fetch the page and the token is read from robots.txt as a permission.

See whether Google’s crawlers actually visit you

robots.txt says who may crawl; only your own request log says who did. CrawlRadar’s collector records every AI crawler visit, verifies it against the vendor’s published ranges, and shows the pages each one fetched.

14-day free trial. No credit card.

Or start with the free AI crawler checker.