ClaudeBot crawl access
What it checks
ClaudeBot collects web content that may contribute to Anthropic’s model training, and Anthropic states its bots honour robots.txt.
This reads the robots.txt rules that actually apply to ClaudeBot — its own group if it has one, otherwise the catch-all — and reports whether they let it fetch the site root. A named group is not required: under RFC 9309 §2.2.1 an open catch-all grants a named crawler the same access a named group would.
The legacy anthropic-ai and Claude-Web tokens are detected and reported on the result when a group names them, but they never change the status or the score.
Why it matters
ClaudeBot allow/block state in robots.txt — Disallowing ClaudeBot stops Anthropic from collecting the site’s content for potential model training; Anthropic states its bots honor robots.txt.
anthropic-ai — ‘anthropic-ai’ is a legacy/undocumented token widely copy-pasted into robots.txt boilerplate; it appears in no current Anthropic documentation, so blocking or allowing it has no vendor-confirmed consequence.
Evidence
ClaudeBot allow/block state in robots.txt
Anthropic documents ClaudeBot as ‘collecting web content that could potentially contribute to their training’ and states ‘Anthropic’s Bots respect do not crawl signals by honoring industry standard directives in robots.txt’, with IP verification at claude.com/crawling/bots.json. It is very much active in 2026, and often the highest-volume AI crawler. Cloudflare Radar had ClaudeBot and GPTBot together at nearly half of all AI crawl activity. Known Agents records 21% of top websites blocking ClaudeBot as of 2026-08-19 — the highest block rate of any Anthropic token.
anthropic-ai
Known Agents classifies anthropic-ai as an ‘Undocumented AI Agent’ — ‘Crawls websites without disclosing its purpose, collecting data for an unknown AI use case’ — while attributing it to Anthropic. Adoption is nevertheless substantial: 16% of top websites block anthropic-ai, evidence of how deeply it is embedded in circulated robots.txt templates. Claude-Web is in the same category: ‘currently unclear exactly what it’s used for, since there’s no official documentation.’
Limits
ClaudeBot allow/block state in robots.txt — Anthropic has by far the worst crawl-to-refer ratio measured by Cloudflare Radar (~50,000:1 overall, 2,500:1 in News & Publications), so allowing ClaudeBot buys essentially no referral traffic — the allow-side case is about training/corpus inclusion, not visibility. Note the canonical support URL moved from support.anthropic.com to support.claude.com; audits hard-coding the old host will 301.
anthropic-ai — Decisive negative: Anthropic’s current, canonical crawler support article names only ClaudeBot, Claude-User and Claude-SearchBot. Neither ‘anthropic-ai’ nor ‘Claude-Web’ appears anywhere on it. There is no vendor doc, no published IP range, and no Cloudflare Radar breakout for anthropic-ai. Treat its presence as harmless legacy cruft — never as evidence a site has configured Anthropic access, and never award or deduct points for it. The same applies to Claude-Web.
How it scores
ClaudeBot allow/block state in robots.txt — Anthropic’s current crawler article names ClaudeBot, and describes it as “collecting web content that could potentially contribute to their training”. It asserts that “Anthropic’s Bots respect do not crawl signals by honoring industry standard directives in robots.txt”. It publishes an IP list for verification. A named agent, a named directive and a stated behaviour is the grade-A bar.
anthropic-ai — The token is genuinely widespread — 16% of top sites carry it — but that is adoption by publishers, not consumption by an agent. It appears in no current Anthropic documentation, has no published IP range and no traffic breakout, so nothing confirms any consequence of allowing or blocking it. Wide adoption plus an unproven mechanism is grade C, which is why this signal is reported and never scored.
Sources
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic, vendor-doc (verified 2026-08-21)
- A deeper look at AI crawlers: breaking down traffic by purpose and industry — Cloudflare, article (verified 2026-08-20)
- ClaudeBot — Known Agents — Known Agents, dataset (verified 2026-08-20)
- anthropic-ai — Known Agents — Known Agents, dataset (verified 2026-08-20)