AI Crawler Atlas / domains
url.kr: AI crawler policy in robots.txt
robots.txt as fetched on 2026-10-10: url.kr disallows 3 of 22 AI and search crawlers from the site root: CCBot, Bytespider, Amazonbot. It leaves the search and citation crawlers and the user-action fetchers open.
A protective stance that stays visible to AI answers: the file refuses training crawlers and leaves open the crawlers that let answer engines cite and link the site.
This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.
Training crawlers
Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| GPTBot | OpenAI | not disallowed | named in the file |
| ClaudeBot | Anthropic | not disallowed | named in the file |
| Google-Extended | not disallowed | named in the file | |
| CCBot | Common Crawl | disallowed | named in the file |
| meta-externalagent | Meta | not disallowed | covered by the default rule |
| Bytespider | ByteDance | disallowed | named in the file |
| Applebot-Extended | Apple | not disallowed | named in the file |
| Amazonbot | Amazon | disallowed | named in the file |
| cohere-ai | Cohere | not disallowed | covered by the default rule |
| Diffbot | Diffbot | not disallowed | covered by the default rule |
| omgili | Webz.io | not disallowed | covered by the default rule |
| Timpibot | Timpi | not disallowed | covered by the default rule |
Search and citation crawlers
These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | not disallowed | named in the file |
| PerplexityBot | Perplexity | not disallowed | named in the file |
| Claude-SearchBot | Anthropic | not disallowed | named in the file |
| DuckAssistBot | DuckDuckGo | not disallowed | covered by the default rule |
User-action fetchers
These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| ChatGPT-User | OpenAI | not disallowed | named in the file |
| Perplexity-User | Perplexity | not disallowed | covered by the default rule |
| Claude-User | Anthropic | not disallowed | named in the file |
| meta-externalfetcher | Meta | not disallowed | covered by the default rule |
Classic search crawlers
Classic search crawlers, shown as the baseline.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| Googlebot | not disallowed | named in the file | |
| Bingbot | Microsoft | not disallowed | named in the file |
How url.kr compares
Among 33,383 readable sites in the Top 50,000, 13.9% disallow GPTBot, 11.8% ClaudeBot, 6.0% OAI-SearchBot and 8.6% PerplexityBot. Among 148 readable .kr sites: 30.4% disallow GPTBot, 25.0% ClaudeBot.
History
First read 2026-10-10; the file has not changed since. Read the file itself.
The AI-related lines in the file
User-agent: Googlebot User-agent: bingbot # 2026-09-24 갱신: 아래 6개를 추가했다. 기존 목록은 anthropic-ai, Claude-Web 처럼 # OAI-SearchBot ChatGPT 검색 결과 노출용 (GPTBot 은 학습용이라 별개다) # ClaudeBot Anthropic 크롤러 현행 이름 # Claude-User 사용자가 Claude 에서 이 페이지를 열 때 # Claude-SearchBot Claude 검색 색인 # PerplexityBot Perplexity 색인 # Applebot-Extended Apple 지능형 기능 User-agent: Google-Extended User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-User User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Applebot-Extended User-agent: anthropic-ai User-agent: CCBot User-agent: Amazonbot User-agent: Bytes
Check it live. This page shows the file as read on 2026-10-10. Run a live check of url.kr and get told when it changes, or get the dataset for a segment.