How AI crawlers access creator data: 114,804 requests measured over 30 days
There is plenty of research about creators, and almost none about how AI actually uses creator data — because that requires first having a body of entity pages that AI crawls at volume. After we published creator profiles as machine-readable entities, crawl volume changed step-wise. This is what we see.
Over the last 30 days we recorded 114,804 AI crawler requests, roughly 3,827 per day, across 56,062 distinct creator entity pages.
Being indexed is not being used
This is the easiest thing to get wrong about crawler data, so it goes first. There are two kinds: indexing crawlers doing bulk collection, and per-user fetches that only fire when someone asks a question. They differ by an order of magnitude.
| Kind | Requests | Share | What it means |
|---|---|---|---|
| Indexing | 112,209 | 97.7% | Bulk collection — the equivalent of "Googlebot crawled you" |
| User-triggered | 2,595 | 2.3% | Someone asked a question and the model fetched the page |
User-triggered retrievals reached 1,196 creator entities. Anyone citing "AI crawl volume" as proof that "AI is using my content" should first be asked which of these two numbers they mean.
Which crawlers arrive
| AI crawler | Requests | Share |
|---|---|---|
| ClaudeBot | 82,343 | 71.7% |
| Amazonbot | 19,077 | 16.6% |
| Meta | 4,054 | 3.5% |
| ChatGPT-User | 2,593 | 2.3% |
| OAI-SearchBot | 2,217 | 1.9% |
| GPTBot | 2,054 | 1.8% |
| Bytespider | 1,771 | 1.5% |
| PerplexityBot | 353 | 0.3% |
| Google-Extended | 301 | 0.3% |
| CCBot | 31 | 0.0% |
Concentration is high: ClaudeBot alone accounts for 71.7%. That is a risk for any conclusion resting on AI crawl volume — one vendor changing its crawl policy reshapes the curve. With this kind of data, the composition matters more than the total.
Which representations they take
| Page type | Requests | Share |
|---|---|---|
| other | 83,402 | 72.6% |
| kol | 9,902 | 8.6% |
| hub | 9,427 | 8.2% |
| hub.json | 3,493 | 3.0% |
| guide | 2,214 | 1.9% |
| kol-zh | 2,189 | 1.9% |
| kol.json | 1,699 | 1.5% |
| crawl-infra | 1,115 | 1.0% |
| home | 556 | 0.5% |
| kol-zh.json | 475 | 0.4% |
Worth noting: JSON accounts for 4.9% — agents actively fetch the machine-readable twin rather than only taking HTML and parsing it themselves. That is the practical difference between publishing for machines and publishing for people and letting machines cope.
Methodology
Figures come from Koinon Link's own server logs, classified by User-Agent, with no IP and no user identifier recorded. Only AI crawlers are counted; search-engine crawlers (Googlebot, Bingbot, Baiduspider and so on) are tracked separately and are not mixed into these numbers.
"User-triggered" includes only the user-agents each vendor documents as per-user fetches (ChatGPT-User, Claude-User, Perplexity-User). OAI-SearchBot and PerplexityBot, which crawl in bulk for their own search indexes, are counted as indexing — widening the definition would inflate the "someone is asking" figure.
⚠ Crawlers outside our UA list are not counted. Coverage of China-based AI crawlers is currently incomplete, so this page understates crawling of Chinese-language content.
Recomputed daily; this page shows the most recent run over a 30-day window.
Cite this page
Koinon Link (2026). How AI crawlers access creator data: 114,804 requests measured over 30 days. Data as of 2026-09-01. https://koinonlink.com/research/how-ai-crawlers-access-creator-dataStatistics are licensed CC BY 4.0: attribute “Koinon Link” with a link to this page. Figures recompute daily — include the data date when citing. Machine-readable version: /research/how-ai-crawlers-access-creator-data.json