Opens in a new tabSkip to content
Agent LighthouseAgent Lighthouse

    Searches the text of every published page. The evidence sources themselves are not in this index — search all of them on the trusted sources page.

    GitHub ↗
    Browse checks and page contents
    access-crawl-control/perplexitybot

    PerplexityBot allowed

    What it checks

    Without an explicit robots.txt rule, PerplexityBot may still crawl your site but has no signal that it is welcome. Adding an explicit allow rule improves your visibility in AI-powered search and ensures consistent crawler behavior.

    Why it matters

    Allowing PerplexityBot is required for the site to be indexed and cited in Perplexity search results; disallowing it removes the site from Perplexity’s index. Perplexity states the crawl is not used for model training.

    Evidence

    PerplexityBot allow/block state in robots.txt

    Perplexity documents an explicit allow-side recommendation: ‘To ensure your site appears in search results, we recommend allowing PerplexityBot in your site’s robots.txt file’. The UA is ‘PerplexityBot/1.0; +https://perplexity.ai/perplexitybot’, and IPs at perplexity.com/perplexitybot.json, and states it is ‘not used for AI model training’. Perplexity also has by far the best crawl-to-refer ratio of the major AI operators in Cloudflare Radar’s data: 118:1 overall, and 32.7:1 in News & Publications, against OpenAI at 887:1 and Anthropic at about 50,000:1. An allow here returns more actual referral traffic per page crawled than any other AI operator.

    Limits

    Perplexity has been publicly accused of crawling from undeclared user agents and rotating IPs to evade blocks (a widely reported 2025 dispute), so a PerplexityBot disallow may not be sufficient to prevent access. Independent 2026 reporting also indicates Perplexity’s crawl-to-refer ratio has worsened (~225:1), eroding the referral argument. The ‘not used for training’ claim is vendor-asserted and unverifiable externally.

    How it scores

    Perplexity publishes an explicit allow-side recommendation: “To ensure your site appears in search results, we recommend allowing PerplexityBot in your site’s robots.txt file.” It publishes the user agent and an IP list alongside it, and states that the crawl is “not used for AI model training”. Vendor documentation of a named token and its effect is grade A. Enforcement is a separate question the audit does not overstate: Perplexity was publicly accused in 2025 of crawling from undeclared agents and rotating addresses, so a disallow here is a declaration rather than a guarantee.

    Sources