What is CCBot (Common Crawl)?
Definition
CCBot is the web crawler run by the Common Crawl Foundation, a non-profit that publishes a free, open archive of web pages. That archive is a major raw source for many large language model training datasets. CCBot identifies itself with the CCBot/2.0 user agent and follows robots.txt rules, including the Crawl-delay directive.
Also known as: CCBot, Common Crawl, Common Crawl bot, Common Crawl crawler

What Common Crawl is, and where CCBot fits
Common Crawl is a 501(c)(3) non-profit founded in 2007. It regularly crawls a large sample of the web and gives the resulting archive away at no cost to researchers, companies and individuals. CCBot is the crawler that fills that archive. According to the foundation, it is built on Apache Nutch and runs on Hadoop: crawl candidates are sorted by host and distributed across a set of crawler servers. The data is stored on Amazon S3, where it can be bulk downloaded or processed in place.
The archive is a sample, not a mirror. Common Crawl says it does not archive entire websites but a randomly selected subset of each, so a visit from CCBot does not mean every page on your site ends up in the dataset.
Why it matters for model training
CCBot does not belong to an AI lab and does not train anything itself. Its significance comes from what others do with the open data it produces. Filtered Common Crawl snapshots are a staple ingredient of large language model training sets. A well-known example: the 2020 GPT-3 paper lists a quality-filtered version of Common Crawl as the largest component of its training mix, about 410 billion tokens or 60% of the mix. The foundation itself reports that its data has been cited in more than 10,000 research papers.
That is why CCBot usually sits in the "training data" group whenever AI crawlers are discussed. There is an extra step in the chain, though: Common Crawl publishes the data, but it does not decide which model developers use it or how they filter it.
User agent and robots.txt rules
The crawler sends the user agent CCBot/2.0 (https://commoncrawl.org/faq/), and its robots.txt token is CCBot. To keep it out entirely:
User-agent: CCBot
Disallow: /If load is the concern rather than inclusion, CCBot honours Crawl-delay. This rule asks it to wait two seconds between requests to the same site:
User-agent: CCBot
Crawl-delay: 2Other documented behaviour: it reads robots.txt before fetching anything and retrieves allowed pages with HTTP GET; it speaks HTTP/1.1, and HTTP/2 over TLS only; it does not execute JavaScript or use cookies; it uses sitemaps announced in robots.txt; it backs off when a server returns 429 or 5xx; and it does not follow links marked nofollow. The JavaScript point deserves attention: pages whose content only exists after client-side rendering are archived empty or incomplete.
Telling the real CCBot from impostors
Common Crawl warns that other crawlers falsely identify themselves as CCBot, so a user agent string proves nothing on its own. The real crawler runs from dedicated IP ranges, and IPv4 requests can be checked with a forward-confirmed reverse DNS lookup: the PTR record should resolve to a host under crawl.commoncrawl.org, and that hostname should resolve back to the same address.
$ host 18.97.14.84
84.14.97.18.in-addr.arpa domain name pointer 18-97-14-84.crawl.commoncrawl.org.
$ host 18-97-14-84.crawl.commoncrawl.org
18-97-14-84.crawl.commoncrawl.org has address 18.97.14.84Reverse DNS is not yet available over IPv6, so compare those requests against the IPv4 and IPv6 ranges the foundation publishes in ccbot.json. Getting this right is a precondition for classifying bot traffic correctly during log file analysis.
What blocking does and does not achieve
- It does keep your pages out of future Common Crawl crawls, and so out of datasets later derived from them.
- It does not remove copies already in past crawls, and it has no effect on anyone who has downloaded the data. Legal removal requests follow a separate process; the foundation publishes the requests it has received in a public "Opt-Out Ledger".
- Search visibility is unaffected: Google and Bing use their own crawlers. Other training-oriented tokens, such as GPTBot or Google-Extended, need their own decisions and their own robots.txt groups.
Which training crawlers to allow is a content policy question, not a technical fault to fix. The GEO Checker lists the robots.txt status of training crawlers, CCBot included, for information only; blocking them never affects the score. The authoritative description of the crawler's behaviour is the Common Crawl FAQ.

