AI Crawler Atlas / domains

cbf.org: AI crawler policy in robots.txt

robots.txt as fetched on 2026-10-10: cbf.org disallows 4 of 22 AI and search crawlers from the site root: Bytespider, Claude-SearchBot, Perplexity-User, Claude-User. That includes crawlers behind AI answers (Claude-SearchBot, Perplexity-User, Claude-User).

The file refuses training crawlers and also some crawlers that power AI answers, so the site is likely to be missing from AI citations and the referral traffic that follows.

This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.

Training crawlers

Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.

CrawlerOperatorStatusHow the file treats it
GPTBotOpenAInot disallowednamed in the file
ClaudeBotAnthropicnot disallowednamed in the file
Google-ExtendedGooglenot disallowednamed in the file
CCBotCommon Crawlnot disallowednamed in the file
meta-externalagentMetanot disallowednamed in the file
BytespiderByteDancedisallowedcovered by the default rule
Applebot-ExtendedApplenot disallowednamed in the file
AmazonbotAmazonnot disallowednamed in the file
cohere-aiCoherenot disallowednamed in the file
DiffbotDiffbotnot disallowednamed in the file
omgiliWebz.ionot disallowednamed in the file
TimpibotTimpinot disallowednamed in the file

Search and citation crawlers

These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.

CrawlerOperatorStatusHow the file treats it
OAI-SearchBotOpenAInot disallowednamed in the file
PerplexityBotPerplexitynot disallowednamed in the file
Claude-SearchBotAnthropicdisallowedcovered by the default rule
DuckAssistBotDuckDuckGonot disallowednamed in the file

User-action fetchers

These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.

CrawlerOperatorStatusHow the file treats it
ChatGPT-UserOpenAInot disallowednamed in the file
Perplexity-UserPerplexitydisallowedcovered by the default rule
Claude-UserAnthropicdisallowedcovered by the default rule
meta-externalfetcherMetanot disallowednamed in the file

Classic search crawlers

Classic search crawlers, shown as the baseline.

CrawlerOperatorStatusHow the file treats it
GooglebotGooglenot disallowednamed in the file
BingbotMicrosoftnot disallowednamed in the file

How cbf.org compares

Among 65,831 readable sites in the Top 100,000, 12.9% disallow GPTBot, 10.8% ClaudeBot, 6.0% OAI-SearchBot and 7.9% PerplexityBot. Among 7,359 readable .org sites: 9.5% disallow GPTBot, 7.2% ClaudeBot.

History

First read 2026-10-10; the file has not changed since. Read the file itself.

The AI-related lines in the file

User-agent: Googlebot
User-agent: Bingbot
User-agent: GPTBot           # OpenAI (ChatGPT, Bing Copilot)
User-agent: OAI-SearchBot    # OpenAI (live search/citation)
User-agent: ChatGPT-User     # OpenAI (user-driven fetch)
User-agent: ClaudeBot        # Anthropic Claude
User-agent: anthropic-ai     # Anthropic Claude
User-agent: Google-Extended  # Google Gemini/AI Overviews
User-agent: PerplexityBot    # Perplexity AI
User-agent: Applebot-Extended
User-agent: Amazonbot        # Amazon Alexa/AI
User-agent: DuckAssistBot    # DuckDuckGo AI
User-agent: YouBot           # You.com AI
User-agent: Meta-ExternalAgent   # Meta AI
User-agent: Meta-ExternalFetcher # Meta AI 

Check it live. This page shows the file as read on 2026-10-10. Run a live check of cbf.org and get told when it changes, or get the dataset for a segment.