robots.txt
Also known as: Robots Exclusion Protocol, REP
robots.txt is a plain-text file at the root of a website, based on the Robots Exclusion Protocol, that tells search engine and AI crawlers which parts of the site they may crawl and which they should leave alone.
robots.txt has been in use since 1994 and became a formal standard in 2022 as RFC 9309. The file sits at example.com/robots.txt. Each rule names a bot (User-agent) and the paths it may or may not access (Allow, Disallow).
robots.txt and AI crawlers
OpenAI, Anthropic, Perplexity and Google publish the names of their AI crawlers and state that they respect robots.txt. That lets a site owner, for example, block the bot that collects training data while leaving the one that searches and cites sources open.
Two things to keep in mind:
- robots.txt is a request, not a security measure. A bot that doesn’t follow the rules can simply ignore it.
- Default security and CDN settings sometimes block AI crawlers without the site owner knowing. That alone can explain why a site never shows up in AI answers.
Example
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
These rules block OpenAI’s training crawler while allowing the bot used for ChatGPT search. Our guide to AI crawlers and llms.txt explains what each bot does.