How it works
It lives at /robots.txt and groups rules by crawler: a User-agent line (Googlebot, Bingbot, GPTBot, or * for everyone) followed by Allow and Disallow lines with paths, plus an optional Sitemap line. The format was standardised as RFC 9309 in 2022. Reputable crawlers obey it, but it is a request, not a lock: it offers no security, and anyone can read the file.
It controls crawling, not indexing. A blocked page can still appear in results, without a description, if other sites link to it; to keep a page out of search, let it be crawled and add a noindex robots meta tag or X-Robots-Tag header. AI companies often use separate crawlers for training and for search (OpenAI has GPTBot and OAI-SearchBot), so each can be allowed or blocked on its own.
robots.txt vs the alternatives
Related terms
More in Analytics, SEO and growth
Telling crawlers