AI Crawler Atlas / domains

eluniversal.com.mx: AI crawler policy in robots.txt

robots.txt as fetched on 2026-10-10: eluniversal.com.mx disallows 20 of 22 AI and search crawlers from the site root: GPTBot, ClaudeBot, Google-Extended, CCBot, meta-externalagent, Bytespider, Applebot-Extended, Amazonbot, cohere-ai, Diffbot, omgili, Timpibot, OAI-SearchBot, PerplexityBot, Claude-SearchBot, DuckAssistBot, ChatGPT-User, Perplexity-User, Claude-User, meta-externalfetcher. That includes crawlers behind AI answers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, DuckAssistBot, ChatGPT-User, Perplexity-User, Claude-User, meta-externalfetcher).

The file refuses training crawlers and also some crawlers that power AI answers, so the site is likely to be missing from AI citations and the referral traffic that follows.

This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.

Training crawlers

Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.

CrawlerOperatorStatusHow the file treats it
GPTBotOpenAIdisallowednamed in the file
ClaudeBotAnthropicdisallowednamed in the file
Google-ExtendedGoogledisallowednamed in the file
CCBotCommon Crawldisallowednamed in the file
meta-externalagentMetadisallowednamed in the file
BytespiderByteDancedisallowednamed in the file
Applebot-ExtendedAppledisallowednamed in the file
AmazonbotAmazondisallowednamed in the file
cohere-aiCoheredisallowednamed in the file
DiffbotDiffbotdisallowednamed in the file
omgiliWebz.iodisallowedcovered by the default rule
TimpibotTimpidisallowedcovered by the default rule

Search and citation crawlers

These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.

CrawlerOperatorStatusHow the file treats it
OAI-SearchBotOpenAIdisallowednamed in the file
PerplexityBotPerplexitydisallowednamed in the file
Claude-SearchBotAnthropicdisallowednamed in the file
DuckAssistBotDuckDuckGodisallowedcovered by the default rule

User-action fetchers

These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.

CrawlerOperatorStatusHow the file treats it
ChatGPT-UserOpenAIdisallowednamed in the file
Perplexity-UserPerplexitydisallowednamed in the file
Claude-UserAnthropicdisallowednamed in the file
meta-externalfetcherMetadisallowednamed in the file

Classic search crawlers

Classic search crawlers, shown as the baseline.

CrawlerOperatorStatusHow the file treats it
GooglebotGooglenot disallowednamed in the file
BingbotMicrosoftnot disallowednamed in the file

How eluniversal.com.mx compares

Among 7,107 readable sites in the Top 10,000, 16.6% disallow GPTBot, 15.2% ClaudeBot, 7.3% OAI-SearchBot and 11.2% PerplexityBot. Among 96 readable .mx sites: 12.5% disallow GPTBot, 11.5% ClaudeBot.

History

First read 2026-10-10; the file has not changed since. Read the file itself.

The AI-related lines in the file

User-agent: Googlebot
User-agent: Googlebot-News
User-agent: Googlebot-Image
User-agent: Googlebot-Video
User-agent: Googlebot-Mobile
User-agent: bingbot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: GPTBot
User-agent: ChatGPT-User
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: anthropic-ai
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: CCBot
User-agent: Bytespider
User-agent: Amazonbot
User-agent: meta-externalagent
User-agent: meta-externalfetcher
User-agent: Diffbot
User-agent: cohere-ai
User-agent: AI2Bot
User-agent: PetalBot

Check it live. This page shows the file as read on 2026-10-10. Run a live check of eluniversal.com.mx and get told when it changes, or get the dataset for a segment.