AI Crawler Atlas / domains
singaporeair.com: AI crawler policy in robots.txt
robots.txt as fetched on 2026-10-10: singaporeair.com has no robots.txt rules (the file is missing or empty), so it asks none of the 22 AI and search crawlers we check to stay away.
A missing or empty robots.txt is neither a ban nor a permission: the site asks nothing of crawlers, so well-behaved ones treat everything as allowed.
This reports what the file asks. It does not show whether a crawler can actually reach the site: many sites also block crawlers at their CDN or firewall, and some crawlers ignore robots.txt.
Training crawlers
Disallowing these keeps a site's content out of future model training. It does not remove what was already collected.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| GPTBot | OpenAI | not disallowed | no rules |
| ClaudeBot | Anthropic | not disallowed | no rules |
| Google-Extended | not disallowed | no rules | |
| CCBot | Common Crawl | not disallowed | no rules |
| meta-externalagent | Meta | not disallowed | no rules |
| Bytespider | ByteDance | not disallowed | no rules |
| Applebot-Extended | Apple | not disallowed | no rules |
| Amazonbot | Amazon | not disallowed | no rules |
| cohere-ai | Cohere | not disallowed | no rules |
| Diffbot | Diffbot | not disallowed | no rules |
| omgili | Webz.io | not disallowed | no rules |
| Timpibot | Timpi | not disallowed | no rules |
Search and citation crawlers
These feed AI answer engines that cite and link their sources. Disallowing them can take a site out of AI answers and out of the referral traffic that comes with them.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | not disallowed | no rules |
| PerplexityBot | Perplexity | not disallowed | no rules |
| Claude-SearchBot | Anthropic | not disallowed | no rules |
| DuckAssistBot | DuckDuckGo | not disallowed | no rules |
User-action fetchers
These fetch a page live when a person asks an assistant about it. Disallowing them stops the assistant reading the page for that person.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| ChatGPT-User | OpenAI | not disallowed | no rules |
| Perplexity-User | Perplexity | not disallowed | no rules |
| Claude-User | Anthropic | not disallowed | no rules |
| meta-externalfetcher | Meta | not disallowed | no rules |
Classic search crawlers
Classic search crawlers, shown as the baseline.
| Crawler | Operator | Status | How the file treats it |
|---|---|---|---|
| Googlebot | not disallowed | no rules | |
| Bingbot | Microsoft | not disallowed | no rules |
How singaporeair.com compares
Among 7,107 readable sites in the Top 10,000, 16.6% disallow GPTBot, 15.2% ClaudeBot, 7.3% OAI-SearchBot and 11.2% PerplexityBot. Among 32,014 readable .com sites: 12.5% disallow GPTBot, 10.6% ClaudeBot.
History
First read 2026-10-10. Read the file itself.
Check it live. This page shows the file as read on 2026-10-10. Run a live check of singaporeair.com and get told when it changes, or get the dataset for a segment.