Google crawlers
Google operates 5 crawlers that may reach your site. They do different jobs, so a single blanket rule for all of them is almost always the wrong call.
| User-agent | Type | What it does |
|---|---|---|
Google-Extended | Training | Collects pages to train future models. |
Google-Agent | Search | Builds the index an assistant searches when it answers. |
Gemini-Deep-Research | Search | Builds the index an assistant searches when it answers. |
Google-CloudVertexBot | Other | Crawls for a related product rather than a public assistant. |
CloudVertexBot | Other | Crawls for a related product rather than a public assistant. |
What blocking each one costs
- Training
- Blocking costs nothing today, but means the next model has never heard of you.
- Search
- Blocking removes you from answers happening right now.
- Other
- Rarely load-bearing for AI visibility; judge it on bandwidth rather than citations.
robots.txt
To let every Google crawler through:
User-agent: Google-Extended
Allow: /
User-agent: Google-Agent
Allow: /
User-agent: Gemini-Deep-Research
Allow: /
User-agent: Google-CloudVertexBot
Allow: /
User-agent: CloudVertexBot
Allow: /To block all of them:
User-agent: Google-Extended
Disallow: /
User-agent: Google-Agent
Disallow: /
User-agent: Gemini-Deep-Research
Disallow: /
User-agent: Google-CloudVertexBot
Disallow: /
User-agent: CloudVertexBot
Disallow: /robots.txt is honoured voluntarily. It is the right way to state your intent, and it is not an access control — treat anything you genuinely need kept out as needing authentication.
Sources
Is Google allowed on your site?
We'll read your robots.txt and tell you which AI crawlers it lets through. No signup.