AI Crawler Atlas / domains
ignatius.com: AI crawler policy in robots.txt
robots.txt as fetched on 2026-10-10: ignatius.com does not disallow any of the 22 AI and search crawlers we check from the site root.
The file asks nothing of these crawlers, so it is open to AI training, AI answers and live assistant fetches as far as robots.txt goes.
This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.
Training crawlers
Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| GPTBot | OpenAI | not disallowed | named in the file |
| ClaudeBot | Anthropic | not disallowed | named in the file |
| Google-Extended | not disallowed | named in the file | |
| CCBot | Common Crawl | not disallowed | named in the file |
| meta-externalagent | Meta | not disallowed | named in the file |
| Bytespider | ByteDance | not disallowed | named in the file |
| Applebot-Extended | Apple | not disallowed | named in the file |
| Amazonbot | Amazon | not disallowed | named in the file |
| cohere-ai | Cohere | not disallowed | named in the file |
| Diffbot | Diffbot | not disallowed | named in the file |
| omgili | Webz.io | not disallowed | named in the file |
| Timpibot | Timpi | not disallowed | named in the file |
Search and citation crawlers
These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | not disallowed | named in the file |
| PerplexityBot | Perplexity | not disallowed | named in the file |
| Claude-SearchBot | Anthropic | not disallowed | covered by the default rule |
| DuckAssistBot | DuckDuckGo | not disallowed | covered by the default rule |
User-action fetchers
These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| ChatGPT-User | OpenAI | not disallowed | named in the file |
| Perplexity-User | Perplexity | not disallowed | covered by the default rule |
| Claude-User | Anthropic | not disallowed | covered by the default rule |
| meta-externalfetcher | Meta | not disallowed | named in the file |
Classic search crawlers
Classic search crawlers, shown as the baseline.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| Googlebot | not disallowed | covered by the default rule | |
| Bingbot | Microsoft | not disallowed | covered by the default rule |
How ignatius.com compares
Among 65,831 readable sites in the Top 100,000, 12.9% disallow GPTBot, 10.8% ClaudeBot, 6.0% OAI-SearchBot and 7.9% PerplexityBot. Among 32,014 readable .com sites: 12.5% disallow GPTBot, 10.6% ClaudeBot.
History
First read 2026-10-10; the file has not changed since. Read the file itself.
The AI-related lines in the file
User-agent: AI2Bot User-agent: Ai2Bot-Dolma User-agent: Amazonbot User-agent: Applebot-Extended User-agent: Bytespider User-agent: CCBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Diffbot User-agent: GPTBot User-agent: Google-Extended User-agent: Meta-ExternalAgent User-agent: Meta-ExternalFetcher User-agent: OAI-SearchBot User-agent: PerplexityBot User-agent: PetalBot User-agent: Timpibot User-agent: YouBot User-agent: anthropic-ai User-agent: cohere-ai User-agent: omgili User-agent: omgilibot
Check it live. This page shows the file as read on 2026-10-10. Run a live check of ignatius.com and get told when it changes, or get the dataset for a segment.