llms.txt vs robots.txt: what each one is for
Two files at the root of the site, two different jobs: robots.txt controls who gets in, llms.txt indexes the content. How to configure them for the AI engines.
robots.txt and llms.txt both live at the root of the site, but they do opposite jobs. robots.txt controls who gets in: it tells each crawler what it may visit. llms.txt invites: it gives the AI engines a curated index of the content that matters. One regulates access, the other guides the reading. Confusing them costs visibility.
- robots.txt governs access; llms.txt indexes the content
- robots.txt is essential and must allow the relevant AI crawlers
- llms.txt is optional, emerging and cheap, and tends to gain weight
- Blocking AI crawlers removes you from the answers without giving back the clicks
Two files, two jobs
Both are plain text files, both live at the root of the domain, and that is where the similarities end. robots.txt has existed since the 1990s and is an access convention: it lists crawlers and tells them where they may or may not go. llms.txt was born in 2024 as a proposal and is a convention of orientation: instead of barring, it guides, pointing the models at the site's main content in a format they read easily.
The common confusion is treating them as alternatives. They are not. One decides whether the engine gets in; the other, once it is inside, helps it understand what matters. They make sense together.
robots.txt: the door
robots.txt is the first thing a crawler consults. For AI visibility, what matters is allowing the right crawlers. Each engine has its own, and it is worth distinguishing two kinds: those that collect content to train the models and those that search the live web to answer a question at that moment.
If you want to appear in the answers, the list to allow includes, among others, OpenAI's GPTBot and OAI-SearchBot, Anthropic's ClaudeBot and Claude-SearchBot, PerplexityBot, and Google-Extended. Blocking any of these removes the site from the sources that engine can read.
The mistake that costs most is blocking by default. Many sites inherit, through a plugin or a template, a robots.txt that bars AI crawlers with no conscious decision behind it. The result is being left out of the answers without even knowing why. The decision to block can be legitimate, a publisher with paid content has reasons to bar training, but it should be chosen, not suffered.
llms.txt: the index
llms.txt is a markdown file that describes the site in a structured, model-readable way: who it is, what it does, and links to the pages that matter. The llms-full.txt variant goes further and delivers the site's main text in a single read, saving the crawler the work of jumping between pages.
It is honest to say that, on its own, it works no miracles. It is an emerging standard and not every engine reads it today. But the cost of creating it is an afternoon's work, there are no contraindications, and the probability of it gaining weight is high as the engines mature. It is the kind of low-risk bet worth making early.
The table, side by side
| Dimension | robots.txt | llms.txt |
|---|---|---|
| Function | Control crawler access | Index the content for reading |
| Stance | Allow or bar | Invite and guide |
| Maturity | Established standard | Emerging proposal (2024) |
| Required for GEO | Yes, must allow AI crawlers | No, but recommended |
| Format | user-agent and allow or disallow rules | Markdown with links |
The minimum package of technical hygiene
Neither of these files works in isolation. The technical foundation the engines expect has three pieces: a robots.txt that allows the relevant AI crawlers, an llms.txt that guides them, and structured data (schema.org) on the main pages so the brand is identified without ambiguity. All three are cheap and are done once. For a worked example, destaque.ai/robots.txt and destaque.ai/llms.txt are public.
Frequently asked questions
What is the difference between llms.txt and robots.txt?
robots.txt governs access: it tells each crawler what it may or may not visit. llms.txt does the opposite of governing, it is an invitation: it gives the AI engines a curated index of the site's most important content. One regulates entry, the other guides the reading. They do not replace each other, they complement each other.
Do I need both?
A correct robots.txt is essential and must allow the relevant AI crawlers if you want to appear in the answers. llms.txt is optional and emerging, cheap to implement and with no contraindications, and tends to gain weight. For a company that wants to be found, the correct robots.txt is mandatory and llms.txt is a good low-cost bet.
Does blocking AI crawlers in robots.txt protect my content?
It blocks the reading, but at the cost of leaving the answers where your customers look. For whoever lives on being found, the trade rarely pays off. The fine-grained choice exists, deciding crawler by crawler and distinguishing training bots from search bots, and it should be a conscious decision, not a plugin's default.
Read next
- What llms.txt is and how to create yours, the file in detail.
- Six commands to find out whether the AI can read your site, to check both at once.
- Schema.org for B2B SaaS: the minimum viable, the third piece of the package.