AI Crawler Atlas / domains

sdsu.edu: AI crawler policy in robots.txt

robots.txt as fetched on 2026-10-10: sdsu.edu does not disallow any of the 22 AI and search crawlers we check from the site root.

The file asks nothing of these crawlers, so it is open to AI training, AI answers and live assistant fetches as far as robots.txt goes.

This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.

Training crawlers

Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.

CrawlerOperatorStatusHow the file treats it
GPTBotOpenAInot disallowedcovered by the default rule
ClaudeBotAnthropicnot disallowedcovered by the default rule
Google-ExtendedGooglenot disallowedcovered by the default rule
CCBotCommon Crawlnot disallowedcovered by the default rule
meta-externalagentMetanot disallowedcovered by the default rule
BytespiderByteDancenot disallowedcovered by the default rule
Applebot-ExtendedApplenot disallowedcovered by the default rule
AmazonbotAmazonnot disallowedcovered by the default rule
cohere-aiCoherenot disallowedcovered by the default rule
DiffbotDiffbotnot disallowedcovered by the default rule
omgiliWebz.ionot disallowedcovered by the default rule
TimpibotTimpinot disallowedcovered by the default rule

Search and citation crawlers

These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.

CrawlerOperatorStatusHow the file treats it
OAI-SearchBotOpenAInot disallowedcovered by the default rule
PerplexityBotPerplexitynot disallowedcovered by the default rule
Claude-SearchBotAnthropicnot disallowedcovered by the default rule
DuckAssistBotDuckDuckGonot disallowedcovered by the default rule

User-action fetchers

These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.

CrawlerOperatorStatusHow the file treats it
ChatGPT-UserOpenAInot disallowedcovered by the default rule
Perplexity-UserPerplexitynot disallowedcovered by the default rule
Claude-UserAnthropicnot disallowedcovered by the default rule
meta-externalfetcherMetanot disallowedcovered by the default rule

Classic search crawlers

Classic search crawlers, shown as the baseline.

CrawlerOperatorStatusHow the file treats it
GooglebotGooglenot disallowedcovered by the default rule
BingbotMicrosoftnot disallowedcovered by the default rule

How sdsu.edu compares

Among 7,107 readable sites in the Top 10,000, 16.6% disallow GPTBot, 15.2% ClaudeBot, 7.3% OAI-SearchBot and 11.2% PerplexityBot. Among 1,188 readable .edu sites: 3.5% disallow GPTBot, 2.5% ClaudeBot.

History

First read 2026-10-10; the file has not changed since. Read the file itself.

Check it live. This page shows the file as read on 2026-10-10. Run a live check of sdsu.edu and get told when it changes, or get the dataset for a segment.