← AI crawler directory

Common Crawl

CCBot

TrainingNo fetched rangesVerified July 2026

What it is

CCBot builds Common Crawl's open web corpus, which is used by researchers and many downstream datasets. Aiola classifies it as training because it performs automated collection rather than a user-triggered fetch.

In Aiola's taxonomy, Training means scraped for model training or dataset development.

User agent

Look for the match token inside the complete HTTP User-Agent header. Tokens can be spoofed, so use the network checks below for authentication.

Full example User-Agent
CCBot/2.0 (https://commoncrawl.org/faq/)
Case-insensitive match token
CCBot

Official IP ranges

Common Crawl has not published official IP ranges that Aiola successfully fetched for CCBot. Aiola does not list cloud-provider ranges or community guesses as if they authenticated this crawler.

Verify authenticity

Start with the source IP recorded by your trusted edge or server, not an untrusted forwarded header.

Reverse- and forward-DNS double lookup
host REQUEST_IP
# Confirm the hostname ends in an official suffix
host RETURNED_HOSTNAME
# The forward lookup must return REQUEST_IP

For production checks, test against every current prefix in the vendor feed. Re-fetch feeds regularly: a July 2026 snapshot is evidence of publication, not a permanent firewall list.

Control it with robots.txt

Common Crawl documents crawler controls for this bot family. Robots.txt can direct cooperative crawling, but it cannot prove that a request is authentic.

Allow CCBot
User-agent: CCBot
Allow: /
Block CCBot
User-agent: CCBot
Disallow: /

Read Common Crawl's official crawler documentation ↗

Track it

Aiola tracks this crawler on your site — crawls, pages, trends. See when CCBot arrives, which URLs it requests, and how activity changes over time.

Explore Aiola Analytics

FAQ

What is CCBot?

CCBot builds Common Crawl's open web corpus, which is used by researchers and many downstream datasets. Aiola classifies it as training because it performs automated collection rather than a user-triggered fetch.

What user-agent token identifies CCBot?

Match the case-insensitive token “CCBot” in the User-Agent header. A matching header alone does not authenticate the sender.

Can I verify CCBot by IP address?

No official CCBot CIDR snapshot was successfully fetched for this verification date. Do not treat an arbitrary cloud IP as proof of identity.

Can robots.txt block CCBot?

Common Crawl documents crawler controls for this bot family. Robots.txt can direct cooperative crawling, but it cannot prove that a request is authentic. Use the specific “CCBot” group when you want a rule for this identity without changing rules for every crawler.

Related