Meta crawlers
Meta operates 5 crawlers that may reach your site. They do different jobs, so a single blanket rule for all of them is almost always the wrong call.
| User-agent | Type | What it does |
|---|---|---|
FacebookBot | Search | Builds the index an assistant searches when it answers. |
meta-webindexer | Training | Collects pages to train future models. |
meta-externalfetcher | User-triggered | Fetches your page because someone asked about it just now. |
Meta-ExternalAgent | User-triggered | Fetches your page because someone asked about it just now. |
Manus-User | User-triggered | Fetches your page because someone asked about it just now. |
What blocking each one costs
- Search
- Blocking removes you from answers happening right now.
- Training
- Blocking costs nothing today, but means the next model has never heard of you.
- User-triggered
- Blocking means that person gets told the page could not be read.
robots.txt
To let every Meta crawler through:
User-agent: FacebookBot
Allow: /
User-agent: meta-webindexer
Allow: /
User-agent: meta-externalfetcher
Allow: /
User-agent: Meta-ExternalAgent
Allow: /
User-agent: Manus-User
Allow: /To block all of them:
User-agent: FacebookBot
Disallow: /
User-agent: meta-webindexer
Disallow: /
User-agent: meta-externalfetcher
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: Manus-User
Disallow: /robots.txt is honoured voluntarily. It is the right way to state your intent, and it is not an access control — treat anything you genuinely need kept out as needing authentication.
Sources
Is Meta allowed on your site?
We'll read your robots.txt and tell you which AI crawlers it lets through. No signup.