Which websites block AI crawlers?

We read the robots.txt of the most popular websites and worked out, for each one, which of 22 AI and search crawlers it asks to stay away: GPTBot, ClaudeBot, Google-Extended, CCBot, OAI-SearchBot, PerplexityBot and more. The census covers 77,154 sites whose file we could read (last full read 2026-10-10); 18,500 of them have a page of their own here because their file says something distinctive, and every change is kept.

12.9%of readable sites disallow GPTBot
10.8%disallow ClaudeBot
6.0%disallow OAI-SearchBot (ChatGPT search)
7.9%disallow PerplexityBot

Percentages are of the 65,831 sites with a robots.txt that was read; 11,323 more have no rules at all and 22,846 could not be read reliably and are not described anywhere. This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.

Look up a site

Disallow rates by popularity

PopularityReadable sitesGPTBotClaudeBotGoogle-ExtendedCCBotOAI-SearchBotPerplexityBotGooglebot
Top 1,00076621.8%22.1%18.8%25.6%11.6%17.9%1.8%
Top 10,0007,10716.6%15.2%13.0%17.4%7.3%11.2%1.1%
Top 50,00033,38313.9%11.8%10.5%13.9%6.0%8.6%1.3%
Top 100,00065,83112.9%10.8%9.8%12.2%6.0%7.9%1.4%

Most disallowed crawlers

What the jobs mean

Browse

GPTBot ClaudeBot Google-Extended CCBot meta-externalagent Bytespider Applebot-Extended Amazonbot cohere-ai Diffbot omgili Timpibot OAI-SearchBot PerplexityBot Claude-SearchBot DuckAssistBot ChatGPT-User Perplexity-User Claude-User meta-externalfetcher Googlebot Bingbot