Common Crawl crawlers
Common Crawl operates 1 crawler that may reach your site. It does a specific job, so a generic blanket rule for it is often the wrong call.
| User-agent | Type | What it does |
|---|---|---|
CCBot | Training | Collects pages to train future models. |
What blocking each one costs
- Training
- Blocking costs nothing today, but means the next model has never heard of you.
robots.txt
To let every Common Crawl crawler through:
User-agent: CCBot
Allow: /To block all of them:
User-agent: CCBot
Disallow: /robots.txt is honoured voluntarily. It is the right way to state your intent, and it is not an access control — treat anything you genuinely need kept out as needing authentication.
Sources
Is Common Crawl allowed on your site?
We'll read your robots.txt and tell you which AI crawlers it lets through. No signup.