AI Crawler Atlas / domains
vocus.cc: AI crawler policy in robots.txt
robots.txt as fetched on 2026-10-10: vocus.cc disallows 13 of 22 AI and search crawlers from the site root: GPTBot, ClaudeBot, Google-Extended, CCBot, meta-externalagent, Bytespider, Applebot-Extended, OAI-SearchBot, PerplexityBot, Claude-SearchBot, ChatGPT-User, Perplexity-User, Claude-User. That includes crawlers behind AI answers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, ChatGPT-User, Perplexity-User, Claude-User).
The file refuses training crawlers and also some crawlers that power AI answers, so the site is likely to be missing from AI citations and the referral traffic that follows.
This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.
Training crawlers
Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| GPTBot | OpenAI | disallowed | named in the file |
| ClaudeBot | Anthropic | disallowed | named in the file |
| Google-Extended | disallowed | named in the file | |
| CCBot | Common Crawl | disallowed | named in the file |
| meta-externalagent | Meta | disallowed | named in the file |
| Bytespider | ByteDance | disallowed | named in the file |
| Applebot-Extended | Apple | disallowed | named in the file |
| Amazonbot | Amazon | not disallowed | covered by the default rule |
| cohere-ai | Cohere | not disallowed | covered by the default rule |
| Diffbot | Diffbot | not disallowed | covered by the default rule |
| omgili | Webz.io | not disallowed | covered by the default rule |
| Timpibot | Timpi | not disallowed | covered by the default rule |
Search and citation crawlers
These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | disallowed | named in the file |
| PerplexityBot | Perplexity | disallowed | named in the file |
| Claude-SearchBot | Anthropic | disallowed | named in the file |
| DuckAssistBot | DuckDuckGo | not disallowed | covered by the default rule |
User-action fetchers
These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| ChatGPT-User | OpenAI | disallowed | named in the file |
| Perplexity-User | Perplexity | disallowed | named in the file |
| Claude-User | Anthropic | disallowed | named in the file |
| meta-externalfetcher | Meta | not disallowed | covered by the default rule |
Classic search crawlers
Classic search crawlers, shown as the baseline.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| Googlebot | not disallowed | covered by the default rule | |
| Bingbot | Microsoft | not disallowed | covered by the default rule |
How vocus.cc compares
Among 33,383 readable sites in the Top 50,000, 13.9% disallow GPTBot, 11.8% ClaudeBot, 6.0% OAI-SearchBot and 8.6% PerplexityBot. Among 70 readable .cc sites: 8.6% disallow GPTBot, 10.0% ClaudeBot.
History
First read 2026-10-10; the file has not changed since. Read the file itself.
The AI-related lines in the file
# --- 以下三個前綴一律回 404,開放檢索是為了讓 Googlebot「看得到」那個 404 --- # __NEXT_DATA__,Googlebot 把內嵌 JSON 裡的路徑當連結抓 → 被下方 `Disallow: /` # 下方的 AI-input 專屬群組(Claude-User/ChatGPT-User/PerplexityBot 等) Content-Signal: search=yes, ai-input=yes, ai-train=no User-agent: Googlebot-Image # === AI 爬蟲:對應 Content-Signal 政策(search=yes, ai-input=yes, ai-train=no)=== # Content-Signal 是 informational header,並非所有 AI 爬蟲都會解讀。 User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Claude-User User-agent: Claude-SearchBot # 但這個 group 正是它們的目標讀者(Claude-User/ChatGPT-User/ # PerplexityBot 等 AI-input 爬蟲),只加在 `User-agent: *` 對它們無效, User-agent:
Check it live. This page shows the file as read on 2026-10-10. Run a live check of vocus.cc and get told when it changes, or get the dataset for a segment.