What is AI Crawler?
Definition
An AI crawler is a bot operated by an AI company that visits web pages and retrieves their content. AI crawlers serve three main purposes: collecting content for model training (GPTBot, ClaudeBot), building an index for AI search (OAI-SearchBot, PerplexityBot), and fetching a page on demand when a user asks (ChatGPT-User). Most can be managed with robots.txt rules.
Also known as: AI bot, AI web crawler, LLM crawler, AI user agent

Three jobs, three kinds of bot
"AI bot" is not a single thing. Most major providers run separate user agents for separate jobs, and blocking each one has a different effect:
| Type | What it does | Examples |
|---|---|---|
| Training crawler | Collects content that may be used to train future models. | GPTBot (OpenAI), ClaudeBot (Anthropic); for Google, the Google-Extended token controls this choice |
| Search / index crawler | Indexes pages that can be shown and linked as sources in AI search answers. | OAI-SearchBot, Claude-SearchBot, PerplexityBot |
| User-triggered fetcher | Reads a page at the moment a user asks a question or shares a link. | ChatGPT-User, Claude-User, Perplexity-User |
The distinction matters because blocking a provider's training bot does not block its search bot. A site that disallows GPTBot but allows OAI-SearchBot signals that it does not want its content used for training while remaining eligible for ChatGPT search.
Controlling them with robots.txt
Most of these bots state that they follow the rules written for their name in robots.txt. A setup that opts out of training while staying open to AI search might look like this:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: *
Disallow: /admin/One detail catches many people out. Under the robots.txt standard (RFC 9309), a bot that finds a group matching its own name follows only that group; the rules under User-agent: * no longer apply to it. In the example above, the search bots can also reach /admin/. Restrictions you want for every bot have to be repeated in each named group.
User-triggered fetchers play by different rules
According to the providers' own documentation, robots.txt plays a smaller role for fetches made on a user's behalf. OpenAI says that because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says disabling Claude-User prevents its system from retrieving your content in response to a user query. Check each provider's current documentation for exact behavior.
robots.txt is not access control
Following robots.txt is voluntary. The standard itself says its rules are not a form of access authorization (RFC 9309). Well-behaved bots comply; software that ignores the rules or pretends to be another bot does not. Content that truly must not be accessed needs passwords, sessions or server-side access control.
User-agent strings are trivial to fake, so do not trust a name in your server logs on its own. Providers such as OpenAI and Perplexity publish their bots' IP ranges as JSON files, which lets you verify whether a request really came from them.
Making the call
- Blocking search bots lowers your chances of being cited in that provider's AI search; blocking training bots expresses a preference that your content not be used to train future models. Decide on each separately.
- Rules written for OpenAI, Anthropic or Perplexity bots do not concern Googlebot, and Google states that Google-Extended does not affect inclusion or ranking in Google Search.
- New bots and user-agent names appear from time to time, so review your robots.txt against the providers' current lists periodically.
Doruva's GEO Checker reports whether a site's robots.txt allows access for the major AI bots.

