CrawlRadar

Common Crawl · training

Is CCBot blocked on your site?

CCBot is Common Crawl’s training token for the Common Crawl archive, and whether it reaches you is decided by one line in your robots.txt. Common Crawl’s archive feeds many models indirectly. Blocking it is a broad training opt-out with a long tail.

  • Not critical if blocked
  • CCBot robots.txt token

By · Published · Last updated

What is CCBot?

CCBot is the crawler run by Common Crawl, a non-profit that publishes a free, open archive of the web. Many AI labs and researchers train on Common Crawl snapshots, so CCBot reaches more models indirectly than any single vendor’s crawler does directly.

OperatorCommon Crawl
Purposetraining · the Common Crawl archive
robots.txt tokenCCBot
In the user-agentCCBot/2.0 (https://commoncrawl.org/faq/)
CrawlRadar weight1 of the crawler-access weights, never flagged critical
Vendor docsCommon Crawl: FAQ

How to allow CCBot in robots.txt

Give CCBot its own group with Allow: /. A named group overrides the User-agent: * group, so this works even when the wildcard disallows everything:

User-agent: CCBot
Allow: /

How to block CCBot in robots.txt

Replace Allow with Disallow. To block only part of the site, disallow that path instead of /:

User-agent: CCBot
Disallow: /
  • Tokens are matched case-insensitively, but writing CCBot exactly as the vendor spells it avoids surprises with stricter parsers.
  • Check your CDN and security plugins too: a firewall rule can block CCBot without any line in robots.txt.

Should you block CCBot?

Block CCBot if you want the broadest training opt-out you can get from one line, knowing it only affects future snapshots: pages already archived stay in past releases. It has no effect on live answer engines, which crawl with their own bots.

How to verify a request is really CCBot

There is no reliable way: Common Crawl publishes no IP ranges and no reverse-DNS scheme. A request claiming to be CCBot could be anyone, so CrawlRadar records such hits as unverified rather than crediting them to Common Crawl.

Is CCBot visiting your site?

Your analytics will not tell you. AI crawlers do not run JavaScript, so a tag-based tool such as Google Analytics never records them. CrawlRadar’s collector reads your own request path and records each CCBot visit, the page it fetched, and whether it came from Common Crawl. See the full AI crawler directory for the other tokens CrawlRadar checks.

Frequently asked questions

Does blocking CCBot remove my pages from existing AI models?

No. It keeps your pages out of future Common Crawl snapshots. Snapshots already published, and models already trained on them, are unaffected.

Is CCBot an AI company’s crawler?

No. Common Crawl is a non-profit web archive. The AI connection is that its archive is one of the most common sources of training data.

See whether CCBot actually visits you

robots.txt says who may crawl; only your own request log says who did. CrawlRadar’s collector records every AI crawler visit, verifies it against the vendor’s published ranges, and shows the pages each one fetched.

14-day free trial. No credit card.

Or start with the free AI crawler checker.