Amazon blocks the AI crawlers. 760 companies say: don't train on me — but cite me.
Across 9,016 corporate and publisher domains observed by FACTANKER, AI training bots are blocked roughly ten times more often than AI search bots — and 760 domains draw exactly that line in a single file.
Cite as: FACTANKER, https://factanker.com/record/who-blocks-ai-crawlers
Every website carries a small public file, robots.txt, in which it tells
crawlers what they may read. FACTANKER records these policies for 9,016 domains as
evidence-backed, timestamped facts — for 16 AI crawlers per domain, with full history.
Read together, they are a running referendum on artificial intelligence, answered not
in surveys but in server configuration. The verdict is unambiguous.
Training bots are blocked ten times more often than search bots
| Crawler | Operator & purpose | Domains blocking | Share |
|---|---|---|---|
| Bytespider | ByteDance — model training | 1,066 | 11.8% |
| CCBot | Common Crawl — training corpus | 972 | 10.8% |
| meta-externalagent | Meta — model training | 861 | 9.5% |
| GPTBot | OpenAI — model training | 833 | 9.2% |
| Applebot-Extended | Apple — model training | 799 | 8.9% |
| ClaudeBot | Anthropic — model training | 787 | 8.7% |
| Google-Extended | Google — Gemini training | 772 | 8.6% |
| PerplexityBot | Perplexity — AI search | 106 | 1.2% |
| ChatGPT-User | OpenAI — on-demand user fetch | 80 | 0.9% |
| Claude-User | Anthropic — on-demand user fetch | 79 | 0.9% |
| OAI-SearchBot | OpenAI — AI search (citations) | 70 | 0.8% |
This is not diffuse fear of AI. It is a precise market decision: sites overwhelmingly tolerate crawlers that cite them and block crawlers that learn from them. The most-blocked bot on the list is not OpenAI''s — it is ByteDance''s Bytespider.
760 domains draw the line in one file
The sharpest evidence sits inside single robots.txt files: 760 domains block GPTBot (training) while explicitly allowing OAI-SearchBot (search citations) — the same company, sorted into “no” and “yes” by purpose. The message could not be clearer: don''t train on me, but cite me in the live answer.
Who blocks both big training bots
Among the largest companies in our cohort that block both GPTBot and ClaudeBot (each re-verified live against the company''s robots.txt on September 5, 2026; revenue as filed with the SEC, from our registry):
| Company | Revenue (latest fiscal year, as filed) |
|---|---|
| Amazon | $716.9B (FY2025) |
| Northrop Grumman | $42.0B (FY2025) |
| MercadoLibre | $20.3B (FY2025) |
| Casey''s General Stores | $17.6B (FY ended Apr 2026) |
| eBay | $11.1B (FY2025) |
Kimberly-Clark also blocks both bots (verified live the same day). The pattern: online marketplaces that monetize their own data (Amazon, eBay, MercadoLibre), a defense prime, and consumer companies protecting brand content.
Total blockade is rare — indifference is the norm
Only 46 of 9,016 domains (0.5%) block all 16 observed AI crawlers. The large majority allow everything — for most, more likely a non-decision than a strategy. The deliberate choices concentrate exactly where data is a business.
Why it matters
Roughly one in ten commercial domains has already closed the door on training crawlers, and the published trend among news publishers is far steeper. Structured, verifiable, willingly available training material is becoming scarcer — while AI answers are moving toward live search, the one channel these same companies leave open on the condition of citation. Whoever wants to be quoted by machines must be open, citable, and provable. That is the corner of the market FACTANKER occupies: this registry is deliberately open to every crawler listed above.
Method
FACTANKER''s observatory records robots.txt policies for 9,016 domains as supersession-tracked facts: each policy carries the retrieval timestamp, and every change creates a new fact while the old one remains retrievable. Counts above were computed from current (non-superseded) facts on September 5, 2026. The six companies named were additionally re-verified against their live robots.txt the same day; two further large blockers in our records (Citigroup, Baidu) did not pass this live re-check and were therefore excluded from this article.