AI crawlers
Also known as: AI bots, LLM crawlers, AI user agents
AI crawlers are bots that companies such as OpenAI, Anthropic, Perplexity and Google use to visit websites, whether to collect training data for models, find sources for AI search, or fetch a page at a user's request.
Most AI companies run several bots for different jobs rather than just one, and they publish the names in their documentation. Broadly, they fall into three groups.
Three kinds of bots
Training crawlers collect data to train models. Examples: OpenAI’s GPTBot and Anthropic’s ClaudeBot. Common Crawl’s CCBot also gathers an open dataset that many models are trained on.
Search crawlers find and index pages to show in AI search. Examples: OAI-SearchBot (ChatGPT search), Claude-SearchBot and PerplexityBot.
User-triggered agents visit when a user asks the assistant to open a link, or when the assistant needs to read a page right then. Examples: ChatGPT-User, Claude-User and Perplexity-User. Some companies note that robots.txt rules may not always apply to these user-initiated requests.
Google works differently. Google-Extended isn’t a separate crawler but a robots.txt token. It controls whether content can be used for Gemini models, and it doesn’t affect Google Search or AI Overviews, which rely on regular Googlebot crawling.
Why it matters
The split gives site owners a choice. You can block training crawlers and still allow search crawlers, for example. A site that blocks search crawlers largely gives up its chance of being cited in AI answers. Blocks aren’t always deliberate, either; firewall and CDN bot-protection defaults can stop these crawlers too.
Example
If your server logs show no visits from OAI-SearchBot, checking robots.txt and your security settings is the first step. Our guide to AI crawlers and llms.txt has the bot list and sample configurations.