AI Crawler Atlas / domains

kaktus.media: AI crawler policy in robots.txt

robots.txt as fetched on 2026-10-10: kaktus.media disallows 7 of 22 AI and search crawlers from the site root: GPTBot, ClaudeBot, CCBot, meta-externalagent, Bytespider, Applebot-Extended, Amazonbot. It leaves the search and citation crawlers and the user-action fetchers open.

A protective stance that stays visible to AI answers: the file refuses training crawlers and leaves open the crawlers that let answer engines cite and link the site.

This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.

Training crawlers

Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.

CrawlerOperatorStatusHow the file treats it
GPTBotOpenAIdisallowednamed in the file
ClaudeBotAnthropicdisallowednamed in the file
Google-ExtendedGooglenot disallowedcovered by the default rule
CCBotCommon Crawldisallowednamed in the file
meta-externalagentMetadisallowednamed in the file
BytespiderByteDancedisallowednamed in the file
Applebot-ExtendedAppledisallowednamed in the file
AmazonbotAmazondisallowednamed in the file
cohere-aiCoherenot disallowedcovered by the default rule
DiffbotDiffbotnot disallowedcovered by the default rule
omgiliWebz.ionot disallowedcovered by the default rule
TimpibotTimpinot disallowedcovered by the default rule

Search and citation crawlers

These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.

CrawlerOperatorStatusHow the file treats it
OAI-SearchBotOpenAInot disallowednamed in the file
PerplexityBotPerplexitynot disallowednamed in the file
Claude-SearchBotAnthropicnot disallowednamed in the file
DuckAssistBotDuckDuckGonot disallowedcovered by the default rule

User-action fetchers

These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.

CrawlerOperatorStatusHow the file treats it
ChatGPT-UserOpenAInot disallowednamed in the file
Perplexity-UserPerplexitynot disallowednamed in the file
Claude-UserAnthropicnot disallowednamed in the file
meta-externalfetcherMetanot disallowedcovered by the default rule

Classic search crawlers

Classic search crawlers, shown as the baseline.

CrawlerOperatorStatusHow the file treats it
GooglebotGooglenot disallowednamed in the file
BingbotMicrosoftnot disallowedcovered by the default rule

How kaktus.media compares

Among 65,831 readable sites in the Top 100,000, 12.9% disallow GPTBot, 10.8% ClaudeBot, 6.0% OAI-SearchBot and 7.9% PerplexityBot.

History

First read 2026-10-10; the file has not changed since. Read the file itself.

The AI-related lines in the file

Content-Signal: ai-train=no, search=yes, ai-input=yes
User-agent: Googlebot
Content-Signal: ai-train=no, search=yes, ai-input=yes
Content-Signal: ai-train=no, search=yes, ai-input=yes
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Content-Signal: ai-train=no, search=yes, ai-input=yes
# Content-Signal выше также сообщает эту политику неизвестным новым ботам.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Bytespider
User-agent: Amazonbot
User-agent: AI2Bot
User-agent: Applebot-Extended
Content-Signal: ai-train=no, search=ye

Check it live. This page shows the file as read on 2026-10-10. Run a live check of kaktus.media and get told when it changes, or get the dataset for a segment.