Opens in a new tabSkip to content
Agent LighthouseAgent Lighthouse

    Searches the text of every published page. The evidence sources themselves are not in this index — search all of them on the trusted sources page.

    GitHub ↗
    Browse checks and page contents
    access-crawl-control/google-extended

    Google-Extended allowed

    What it checks

    Without an explicit robots.txt rule, Google-Extended may still crawl your site but has no signal that it is welcome. Adding an explicit allow rule improves your visibility in AI-powered search and ensures consistent crawler behavior.

    Why it matters

    Disallowing Google-Extended stops the site’s content being used to train Gemini models and to ground answers in Gemini Apps and Vertex AI Grounding-with-Google-Search. It has zero effect on Google Search crawling, indexing, ranking, or AI Overviews.

    Evidence

    Google-Extended allow/block state in robots.txt

    Google documents Google-Extended as ‘a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models’. Those are the models ‘that power Gemini Apps and Vertex AI API for Gemini and for grounding … in Gemini Apps and Grounding with Google Search on Vertex AI’. It also states that the token ‘does not impact a site’s inclusion in Google Search nor is it used as a ranking signal’. It is a robots.txt token only — no crawler fetches with that UA — so it is safe to block without traffic loss.

    Limits

    Critical, widely-misreported limitation: Google-Extended does not control AI Overviews or AI Mode. Google’s AI-features page states ‘robots.txt directives for Googlebot is the control for site owners to manage access to how their sites are crawled for Search’ and directs publishers to nosnippet / data-nosnippet / max-snippet / noindex for AI feature control. Any audit implying a Google-Extended disallow keeps content out of AI Overviews is wrong. Also, BuzzStream measured 92.3% citation retention among sites blocking Google-Extended — the highest of any bot studied.

    How it scores

    Google documents the token by name and states exactly what it governs: whether crawled content “may be used for training future generations of Gemini models” and for grounding in Gemini Apps and Vertex AI. That is a vendor statement about a named token, which is the grade-A bar. The grade does not extend to the claim most often attached to this token: Google-Extended does not control AI Overviews or AI Mode, and Google directs publishers to Googlebot’s own directives and to nosnippet for those. The audit reports the training and grounding effect only.

    Sources