Verified 11 July 2026
AI crawlers directory
A practical directory of every AI bot Aiola tracks, using the same taxonomy as Aiola Analytics. AI answers are user-triggered fetches, Indexing bots build search or retrieval indexes, and Training crawlers collect data for model or dataset development.
A user-agent string is only a claim. Where a vendor publishes CIDRs, rDNS rules, or an ASN, the detail page explains how to verify the request. “No fetched ranges” means exactly that: Aiola does not substitute guessed infrastructure.
AI answers (14)
Fetched to answer a user's question now
ChatGPT-User · OpenAI
ChatGPT-User fetches a page after a person asks ChatGPT to visit or summarize it; it is not the automated training crawler.
Claude-User · Anthropic
Claude-User retrieves a page in response to a person's request inside Claude.
Perplexity-User · Perplexity
Perplexity-User visits a URL as part of a specific user action and may ignore robots.txt because the fetch is user initiated.
MistralAI-User · Mistral
MistralAI-User fetches pages when a Mistral product user requests current web content.
Grok-DeepSearch · xAI
Grok-DeepSearch fetches sources while a Grok user runs DeepSearch.
Google-NotebookLM · Google
Google-NotebookLM fetches a source URL that a NotebookLM user explicitly adds to a notebook.
Google-Read-Aloud · Google
Google-Read-Aloud fetches pages for Google's text-to-speech and read-aloud services.
Google-Agent · Google
Google-Agent identifies user-triggered requests made by Google agent products.
Copilot · Microsoft
Copilot fetches web content for a Microsoft Copilot interaction rather than performing Bing's general index crawl.
Amzn-User · Amazon
Amzn-User retrieves a page after a user action in an Amazon AI experience.
Meta-ExternalFetcher · Meta
Meta-ExternalFetcher fetches a URL after a person shares or requests it in a Meta product, including link-preview generation.
DuckAssistBot · DuckDuckGo
DuckAssistBot fetches pages in real time for DuckDuckGo's AI-assisted answers and is not used to train AI models.
Kimi-User · Moonshot AI
Kimi-User fetches a URL for a specific Kimi user request.
Qwen-User · Alibaba
Qwen-User fetches web content in response to a Qwen user request.
Indexing (21)
Crawled for AI or search indexing
OAI-SearchBot · OpenAI
OAI-SearchBot discovers and indexes pages so ChatGPT search can retrieve and cite current web results.
Claude-SearchBot · Anthropic
Claude-SearchBot builds Anthropic's search index so Claude can find current web sources.
PerplexityBot · Perplexity
PerplexityBot indexes pages for Perplexity search results and citations; Perplexity says it is not used for foundation-model training.
MistralAI-Index · Mistral
MistralAI-Index indexes web pages for Mistral's search and retrieval features.
xAI-SearchBot · xAI
xAI-SearchBot discovers pages for xAI search retrieval used by Grok.
Google-InspectionTool · Google
Google-InspectionTool fetches a URL for Search Console inspection and related Google testing tools.
Googlebot · Google
Googlebot crawls and renders pages for Google Search indexing, including surfaces that supply search results to AI features.
Applebot · Apple
Applebot indexes the web for Apple search experiences including Spotlight, Siri, and Safari.
Bingbot · Microsoft
Bingbot crawls and renders pages for the Bing index, which also supplies results to Microsoft Copilot.
Amzn-SearchBot · Amazon
Amzn-SearchBot discovers pages for Amazon search and AI retrieval experiences.
TikTokSpider · ByteDance
TikTokSpider crawls pages for TikTok search, previews, and discovery surfaces.
meta-webindexer · Meta
meta-webindexer indexes public pages for Meta's web-search and AI retrieval features.
YouBot · You.com
YouBot indexes public pages for You.com's search engine and cited AI answers.
Kimi-SearchBot · Moonshot AI
Kimi-SearchBot indexes pages for Kimi's web-search retrieval.
Baiduspider · Baidu
Baiduspider crawls and indexes pages for Baidu Search.
msnbot · Microsoft
Microsoft's legacy search crawler, still seen alongside Bingbot.
YandexBot · Yandex
Yandex's primary search crawler.
DuckDuckBot · DuckDuckGo
DuckDuckGo's search crawler.
SeznamBot · Seznam
Search crawler for the Czech portal Seznam..
PetalBot · Petal
Crawler for Huawei's Petal Search..
Qwantbot · Qwant
Search crawler for the French engine Qwant..
Training (32)
Scraped for model training or dataset development
GPTBot · OpenAI
GPTBot collects public web content that may be used to improve OpenAI's generative AI models.
ClaudeBot · Anthropic
ClaudeBot automatically crawls public pages for Anthropic model development and improvement.
anthropic-ai · Anthropic
anthropic-ai is Anthropic's legacy training-data crawler token, separate from Claude's user-triggered fetcher.
GrokBot · xAI
GrokBot is an xAI crawler identity associated with Grok's automated web collection.
xAI-Web-Crawler · xAI
xAI-Web-Crawler is a broad xAI automated web-crawler identity.
xAI-Grok · xAI
xAI-Grok is an xAI crawler token associated with Grok content collection.
xAI-Bot · xAI
xAI-Bot is a general xAI automated crawler identity seen in server logs.
Grok · xAI
Grok is the broad Grok match token retained by Aiola for xAI crawler traffic not identified by a more specific token.
Google-Extended · Google
Google-Extended is a robots.txt control token, not a standalone HTTP crawler; it controls whether crawled content may help ground or improve Gemini and Vertex AI.
Google-CloudVertexBot · Google
Google-CloudVertexBot crawls sites at their owners' request when building Vertex AI data stores.
GoogleOther · Google
GoogleOther is Google's generic crawler for product teams' one-off and R&D fetches; it is not the Search indexer (that is Googlebot) and may use separate crawl controls.
Applebot-Extended · Apple
Applebot-Extended is Apple's robots.txt opt-out token for use of Applebot-crawled material in generative foundation-model training.
Amazonbot · Amazon
Amazonbot is Amazon's general-purpose web crawler used to improve services such as Alexa and generative AI.
Bytespider · ByteDance
Bytespider collects public web content for ByteDance products and model development.
Meta-ExternalAgent · Meta
Meta-ExternalAgent automatically crawls public web content for Meta AI products and model improvement.
CCBot · Common Crawl
CCBot builds Common Crawl's open web corpus, which is used by researchers and many downstream datasets.
cohere-ai · Cohere
cohere-ai collects public web data for Cohere model training and improvement.
KimiBot · Moonshot AI
KimiBot automatically crawls public pages for Moonshot AI's model and product development.
QwenBot · Alibaba
QwenBot collects public pages for Qwen model and product development.
TongyiBot · Alibaba
TongyiBot is an Alibaba crawler identity associated with Tongyi model development.
AliyunBot · Alibaba
AliyunBot is an Alibaba Cloud crawler identity used for automated web collection.
ERNIEBot · Baidu
ERNIEBot collects public web content associated with Baidu's ERNIE model ecosystem.
YiyanBot · Baidu
YiyanBot is a Baidu crawler identity associated with the ERNIE/Yiyan generative AI service.
ChatGLM-Spider · Zhipu AI
ChatGLM-Spider collects public web content for Zhipu AI's ChatGLM model ecosystem.
DeepSeekBot · DeepSeek
DeepSeekBot is an automated crawler associated with DeepSeek model and product development.
AI2Bot · Allen AI
AI2Bot collects public web content for the Allen Institute for AI's research datasets and models.
Diffbot · Diffbot
Diffbot extracts structured knowledge from public pages for Diffbot's Knowledge Graph and crawl products.
Timpibot · Timpi
Timpibot crawls public pages for Timpi's decentralized search index.
ImagesiftBot · ImageSift
ImagesiftBot collects and analyzes public images and page context for ImageSift's image intelligence services.
Doubaobot · ByteDance
Crawler for ByteDance's Doubao assistant.
LinerBot · Liner
Crawler for the Liner AI answer engine..
QualifiedBot · Qualified
Crawler operated by Qualified for its AI sales-agent product..
Other (3)
Fetched for ads or link previews rather than answering, indexing, or training