Periscopy
Articles
Methodology8 min read

What llms.txt is and how to create yours

A practical specification of llms.txt, the markdown file at the root of the site that AI crawlers read before the HTML. Structure, rules, a real example and the mistakes to avoid.


llms.txt is a markdown file at the root of the site that describes the content in a structured, language-model-readable way. An emerging convention in 2024 to 2026, adopted by Anthropic, Mistral and several technical companies. It coexists with robots.txt and sitemap.xml, it does not replace them. This piece shows the minimum structure, the expanded llms-full.txt version, and the typical mistakes.

Update · June 2026

Google published in June 2026 that it is not necessary to create an llms.txt in order to appear in AI answers in its search. That does not contradict this article: we always described it as low-cost hygiene, not as a citation lever. It still makes sense to publish it (documentation, engines other than Google, zero cost). The full context of what changed is in Google made GEO official.

Key takeaways
  • llms.txt is plain markdown and lives at /llms.txt at the root
  • it works as a curated index, not an exhaustive sitemap
  • the llms-full.txt variant carries the expanded body for a single fetch
  • it replaces neither robots.txt (access) nor sitemap.xml (indexing)
  • implementation effort is low; the upside scenario is large

What llms.txt is

llms.txt is a public markdown file, hosted at the root of the domain (https://example.com/llms.txt), that describes the structure and purpose of the site in language optimised for models. The specification was proposed in 2024 and defined three fixed blocks: name, a one-line description, and sections of curated links grouped by topic.

The intent is to solve a concrete problem. When an AI crawler arrives at your site, it has to infer the structure through HTML, headers, menus, CSS classes. That process is fragile. With llms.txt, it receives a declarative summary, in a format the models extract with near-total fidelity.

Why it matters for GEO

Three practical reasons. First, fidelity: structured markdown has less noise than HTML rendered with JavaScript. Second, curation: the choice of what goes in is yours, instead of leaving the crawler to guess. Third, single fetch: the llms-full.txt version lets the model pick up the whole context in one call, instead of making 10 or 20 subsequent requests.

In GEO terms, this shows up in what the engines cite. The cleaner the representation of your site, the more likely the citation is to be correct: right URL, right name, right description.

Minimum structure

The official specification is simple. Title as h1, description as blockquote, and h2 sections with lists of links. A real example, adapted from the llms.txt of destaque.ai:

# destaque.ai

> Empresa de tecnologia que mede a visibilidade das marcas nas respostas de IA.
> Opera a presença de marcas digitais e locais em ChatGPT, Claude, Gemini e Perplexity.

## Serviço
- [Como trabalhamos](https://www.destaque.ai/servico)
- [Glossário GEO](https://www.destaque.ai/glossario)

## Empresa
- [Sobre](https://www.destaque.ai/sobre)
- [Contacto](https://www.destaque.ai/contacto)

## Blog
- [Posts](https://www.destaque.ai/blog)

## Recursos para IA
- [llms-full.txt](https://www.destaque.ai/llms-full.txt)
- [ai.txt](https://www.destaque.ai/ai.txt)

Note three choices. The links are absolute, not relative. The sections are grouped by intent (service, company, blog), not by page type. And there is an explicit resources-for-AI section pointing at llms-full.txt, which makes the crawler's life easier.

llms.txt vs llms-full.txt

llms.txt is the index. llms-full.txt is the index plus the expanded body of the key pages, concatenated in markdown.

In practice: the crawler starts with llms.txt to map the terrain; if it wants depth, it goes to llms-full.txt and picks up everything at once. The cost of serving llms-full.txt is low (a static route) and the saving in crawler requests is considerable.

Rule of thumb: if you already have 5 well-written key pages, it is worth generating the expanded version. If you are still building content, focus on llms.txt first.

llms.txt vs ai.txt vs robots.txt

Four files, four distinct responsibilities:

  • robots.txt governs access. It tells the crawlers (including GPTBot, ClaudeBot and the rest) what they may or may not crawl.
  • sitemap.xml governs indexing. It lists the site's URLs so crawlers discover pages.
  • llms.txt governs understanding. It explains in markdown what each block of the site is.
  • ai.txt governs the AI usage policy. It specifies which AI crawlers may index and on what terms. Analogous to robots.txt but focused on use by generative models.

All four should coexist. Each solves a different problem in the AI visibility work.

How to create yours, in five steps

The process is shorter than it looks. Typically an afternoon's work.

  1. Identify the pages that matter. Typically: home, service, about, contact, blog index. For most B2B SaaS sites, that is 5 to 10 pages.
  2. Write a descriptive sentence for each. Short, specific. "How we work: the four phases of the method" is better than "Service page".
  3. Group by intent, not by page type. Service, company, resources, blog. Not static pages and dynamic ones.
  4. Generate the file as a static route or route handler. In Next.js, a route handler at app/llms.txt/route.ts does it. In other stacks, a static file in public/.
  5. Reference it in robots.txt. A simple line: Sitemap: https://example.com/llms.txt. It is not a formal sitemap, but it gives extra discovery. Repeat for llms-full.txt.

Typical mistakes

We see the same four again and again:

  • Trying to include everything. llms.txt is curation, not inventory. 200 links with no context are worth less than 30 well described.
  • No hierarchy. A flat list with no h2 sections gives the model zero clues about how the site is organised.
  • Linking obsolete pages. Every link that returns a 404 or a redirect is noise. Audit at least quarterly.
  • Forgetting to update it. llms.txt does not maintain itself. If you change the structure of the site, update it immediately.

How to validate

There is no official validator yet. The manual process is simple:

  • curl https://example.com/llms.txt: does it return the content? Do the headers say Content-Type text/markdown or text/plain?
  • Open it in a browser. Are the line breaks correct, with no HTML escaping?
  • Run the content through a model (Claude, ChatGPT) and ask: read this llms.txt and describe the company. Does the answer make sense?
  • Validate the URLs with a status checker: do they all answer 200?

Frequently asked questions

Does llms.txt replace sitemap.xml or robots.txt?

No. They are three files with different purposes. robots.txt says who may enter; sitemap.xml says where the pages are for indexing; llms.txt says, in LLM-readable markdown, what each page is and how it is organised. They coexist.

Do AI crawlers actually read llms.txt?

Adoption is growing but still partial. Anthropic, Mistral and some large technical companies have adopted it. OpenAI and Google have not confirmed official support. Even so, the effort is low and the upside scenario is large, so it is worth putting in place.

Do I have to create llms-full.txt as well?

It is recommended, yes. llms.txt works as an index; llms-full.txt carries the expanded body in markdown, the main text of the key pages concatenated. A single fetch for the crawler to resolve everything in one call.

How big should llms.txt be?

Short. llms.txt is a curated index, not an exhaustive sitemap. For most companies, 30 to 100 lines is enough. llms-full.txt can go up to 50 to 100 thousand tokens depending on the size of the site.

Sources