Which websites block AI crawlers?
We read the robots.txt of the most popular websites and worked out, for each one, which of 22 AI and search crawlers it asks to stay away: GPTBot, ClaudeBot, Google-Extended, CCBot, OAI-SearchBot, PerplexityBot and more. The census covers 77,154 sites whose file we could read (last full read 2026-10-10); 18,500 of them have a page of their own here because their file says something distinctive, and every change is kept.
Percentages are of the 65,831 sites with a robots.txt that was read; 11,323 more have no rules at all and 22,846 could not be read reliably and are not described anywhere. This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.
Look up a site
Disallow rates by popularity
| Popularity | Readable sites | GPTBot | ClaudeBot | Google-Extended | CCBot | OAI-SearchBot | PerplexityBot | Googlebot |
|---|---|---|---|---|---|---|---|---|
| Top 1,000 | 766 | 21.8% | 22.1% | 18.8% | 25.6% | 11.6% | 17.9% | 1.8% |
| Top 10,000 | 7,107 | 16.6% | 15.2% | 13.0% | 17.4% | 7.3% | 11.2% | 1.1% |
| Top 50,000 | 33,383 | 13.9% | 11.8% | 10.5% | 13.9% | 6.0% | 8.6% | 1.3% |
| Top 100,000 | 65,831 | 12.9% | 10.8% | 9.8% | 12.2% | 6.0% | 7.9% | 1.4% |
Most disallowed crawlers
- GPTBot (OpenAI, training): disallowed by 12.9% of readable sites
- CCBot (Common Crawl, training): disallowed by 12.2% of readable sites
- ClaudeBot (Anthropic, training): disallowed by 10.8% of readable sites
- Bytespider (ByteDance, training): disallowed by 10.6% of readable sites
- Google-Extended (Google, training): disallowed by 9.8% of readable sites
- meta-externalagent (Meta, training): disallowed by 9.5% of readable sites
- Amazonbot (Amazon, training): disallowed by 9.1% of readable sites
- Applebot-Extended (Apple, training): disallowed by 9.0% of readable sites
What the jobs mean
- Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot and others): disallowing keeps content out of future model training.
- Search and citation crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot): disallowing can remove a site from AI answers and the traffic that follows.
- User-action fetchers (ChatGPT-User, Perplexity-User, Claude-User): disallowing stops an assistant reading the page when a person asks.
Browse
GPTBot ClaudeBot Google-Extended CCBot meta-externalagent Bytespider Applebot-Extended Amazonbot cohere-ai Diffbot omgili Timpibot OAI-SearchBot PerplexityBot Claude-SearchBot DuckAssistBot ChatGPT-User Perplexity-User Claude-User meta-externalfetcher Googlebot Bingbot