robots.txt · llms.txt

AI crawlers, robots.txt and llms.txt: which doors to keep open

AI companies run bots that do different jobs: some collect data for model training, others read your pages during live search. You should choose which ones to allow on purpose.

The most mundane reason a site does not appear in AI answers is that the bots cannot get in at all. We see this often in the sites we analyse: someone wrote a broad block into robots.txt years ago, or a CDN security setting mistook a new bot for an attacker and turned it away. The site owner usually has no idea.

In this article we go through the bots run by the major AI companies, how they differ, and how to manage them with robots.txt. At the end we look at llms.txt, a file that has been getting a lot of attention lately.

Not all bots are the same

It helps to sort AI bots into three groups:

  1. Training bots. They collect content to train future models. Blocking them keeps your content out of future training runs.
  2. Search bots. They build the AI tool’s own search index. If you block them, your site may not be shown as a source in that tool’s search answers.
  3. User-triggered fetchers. They fetch a page at the moment a user asks a question or asks the tool to open a link.

This distinction matters because “I don’t want my content used for AI training” and “I want to appear in AI answers” can both be true at once. You can block the training bot and keep the search bot open.

Who is who

Company Bot What it does robots.txt
OpenAI GPTBot Collects content for model training Respects it
OpenAI OAI-SearchBot Surfaces sites in ChatGPT search Respects it
OpenAI ChatGPT-User Fetches pages for user actions in ChatGPT Rules may not apply, as the user initiates it
Anthropic ClaudeBot Collects content for model development Respects it
Anthropic Claude-SearchBot Crawls to improve search result quality in Claude Respects it
Anthropic Claude-User Fetches pages for a user’s question Respects it
Perplexity PerplexityBot Indexes sites for Perplexity search results Respects it
Perplexity Perplexity-User Fetches pages for a user’s question Generally ignores it
Google Google-Extended Not a separate bot but a control token. Allows or disallows use for Gemini training and Gemini apps Respects it

This table is based on each company’s own documentation. Bot names and behaviour change from time to time, and rules written against an old list may no longer work. We recommend checking the current state in the sources listed at the end.

Why Google-Extended is different

It is worth stressing that Google-Extended is not a separate crawler. Google crawls with its existing bots, and Google-Extended is only a token used in robots.txt. According to Google’s documentation, it does not affect a site’s inclusion in Google Search and is not used as a ranking signal.

One more thing: Google’s AI Overviews and AI Mode are part of Google Search. To limit what is shown from your pages in these features, Google points to Search controls such as nosnippet, max-snippet or noindex, not to Google-Extended. Google-Extended covers training and use in Google’s other systems. To appear in these features, your page needs to be in Google’s index.

A robots.txt example

The example below reflects a common choice: content is kept out of model training but stays open to AI search.

# Bots that collect for model training: blocked
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

# AI search and user requests: allowed
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

# All other bots
User-agent: *
Disallow: /admin/

Sitemap: https://www.example.com/sitemap.xml

This is an example, not a template. Some businesses are fine with their content being used for training and leave everything open. Others want the exact opposite. Both are legitimate choices. What matters is that the decision is made knowingly.

Two reminders:

  • robots.txt is a request, not a lock. Google says so in its own documentation: not every crawler follows these rules. If a page must stay private, protect it with a password.
  • As the table shows, some user-triggered fetchers may not apply robots.txt rules. Because they act on a user’s request in the moment, the companies do not treat them as ordinary crawlers.

The hidden barrier: CDNs and firewalls

Your robots.txt may be spotless and the bots may still be unable to get in. CDNs such as Cloudflare and web application firewalls (WAFs) offer separate settings for AI bots. Cloudflare’s AI Crawl Control shows which AI services access your site and lets you set allow or block rules for each crawler individually. These settings sometimes switch on through a dashboard update or a single click by a colleague.

The result: robots.txt says “open”, but the bot gets a 403 at the door and leaves. We check for this separately in our analyses. You can look at:

  • Bot and AI settings in your CDN dashboard
  • Firewall rules that block by user agent
  • The response codes AI bots receive in your server logs. 200 is good, 403 or 503 is a problem.

What llms.txt is, and what it isn’t

llms.txt is a proposal for a Markdown file placed at the root of a site. Jeremy Howard put it forward in September 2024. The idea is simple: give large language models a short description of the site and a clean list of its most important pages. The file contains a title, a short summary and annotated lists of links.

What you need to know to use it sensibly:

  • It is a proposal, not a standard. llmstxt.org describes itself as a proposal. It has not been adopted by a standards body.
  • Support from the big players is unclear. Google states in its AI features documentation that you do not need new machine-readable files, AI text files or special markup to appear in those features.
  • It does not replace robots.txt. llms.txt is not a permissions file. Which bots may enter is decided by robots.txt.

So should you add one? Our view: it is easy to produce and does no harm, but don’t expect miracles. It can be useful for software companies with documentation, since AI-assisted coding tools may read such files. On an ordinary business site it sits near the bottom of the priority list. First make sure the doors are open and the site is fast and readable.

Where to start

  1. Open your site’s robots.txt (yoursite.com/robots.txt) and see what it says to AI bots.
  2. Make a separate decision for training bots and search bots, and write it into the file.
  3. Check your CDN and firewall settings for rules that affect AI bots.
  4. Confirm in your server logs that these bots actually receive a 200 response.
  5. Add llms.txt if you like, but after the first four steps.

For short definitions, see the AI crawlers, robots.txt and llms.txt glossary entries. If you’d like us to check whether your site is open to these bots, fill in the free analysis form.

Sources

  1. OpenAI: Overview of OpenAI crawlers developers.openai.com
  2. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? support.claude.com
  3. Perplexity: Perplexity Crawlers docs.perplexity.ai
  4. Google Search Central: Google's common crawlers (Google-Extended) developers.google.com
  5. The llms.txt proposal (llmstxt.org) llmstxt.org
  6. Cloudflare: AI Crawl Control developers.cloudflare.com
Free analysis

Let us see who AI recommends in your industry.

Leave your website and industry. We run the first analysis for free and send you the result.

Prefer to write directly: [email protected]

We use your details only for this analysis.