Diffbot
Diffbot
What it is
Diffbot extracts structured knowledge from public pages for Diffbot's Knowledge Graph and crawl products. Aiola classifies it as training because it performs automated collection rather than a user-triggered fetch.
In Aiola's taxonomy, Training means scraped for model training or dataset development.
User agent
Look for the match token inside the complete HTTP User-Agent header. Tokens can be spoofed, so use the network checks below for authentication.
DiffbotDiffbotOfficial IP ranges
Diffbot has not published official IP ranges that Aiola successfully fetched for Diffbot. Aiola does not list cloud-provider ranges or community guesses as if they authenticated this crawler.
Verify authenticity
Start with the source IP recorded by your trusted edge or server, not an untrusted forwarded header.
host REQUEST_IP
# Confirm the hostname ends in an official suffix
host RETURNED_HOSTNAME
# The forward lookup must return REQUEST_IPFor production checks, test against every current prefix in the vendor feed. Re-fetch feeds regularly: a July 2026 snapshot is evidence of publication, not a permanent firewall list.
Control it with robots.txt
Diffbot documents crawler controls for this bot family. Robots.txt can direct cooperative crawling, but it cannot prove that a request is authentic.
User-agent: Diffbot
Allow: /User-agent: Diffbot
Disallow: /Track it
Aiola tracks this crawler on your site — crawls, pages, trends. See when Diffbot arrives, which URLs it requests, and how activity changes over time.
Explore Aiola AnalyticsFAQ
What is Diffbot?
Diffbot extracts structured knowledge from public pages for Diffbot's Knowledge Graph and crawl products. Aiola classifies it as training because it performs automated collection rather than a user-triggered fetch.
What user-agent token identifies Diffbot?
Match the case-insensitive token “Diffbot” in the User-Agent header. A matching header alone does not authenticate the sender.
Can I verify Diffbot by IP address?
No official Diffbot CIDR snapshot was successfully fetched for this verification date. Do not treat an arbitrary cloud IP as proof of identity.
Can robots.txt block Diffbot?
Diffbot documents crawler controls for this bot family. Robots.txt can direct cooperative crawling, but it cannot prove that a request is authentic. Use the specific “Diffbot” group when you want a rule for this identity without changing rules for every crawler.