llms.txt: Who Actually Reads It, and What to Put In It

The file everybody argues about

If you've been anywhere near AEO discussion in the last two years, you've watched the same fight loop: llms.txt is the new sitemap, llms.txt is the new meta keywords, publish it immediately, don't bother.

Both camps are partly right, and the reason is that they're answering different questions. "Will llms.txt get me cited in ChatGPT?" and "will llms.txt help an AI agent use my site?" have different answers. Almost every take I've read collapses them into one.

Here's the sourced version, with dates on everything, because this is a moving target and anything I write today has a shelf life.

What llms.txt actually is

It's a plain Markdown file at your site root — example.com/llms.txt — that gives an AI tool a short, structured summary of what your site is and where the important pages live. Jeremy Howard proposed it on September 3, 2024. The spec went to v2 on August 10, 2026.

Think of it as sitting between robots.txt and sitemap.xml. robots.txt says who may crawl. sitemap.xml lists every URL with no opinion about which matter. llms.txt is the curated version: here's who we are, and here are the ten pages that actually answer things.

The v2 spec is short. Exactly one element is required — an H1 with the name of the site or project. Everything else is optional:

  • A blockquote immediately after the H1, holding a one-paragraph summary.
  • Zero or more Markdown sections of prose (anything except headings).
  • Zero or more H2-delimited sections containing lists of links. Each list item needs a Markdown hyperlink, optionally followed by : and a short note.
  • By convention, an H2 named Optional marks links an agent can skip when it needs a shorter context.

V2 added two things worth knowing about. First, a formal way to point at Markdown versions of your pages — either append .md to the whole filename (/docs/tutorial.html.md) or swap the extension (/docs/tutorial.md). Second, discovery via link relations: <link rel="alternate" type="text/markdown"> for the Markdown twin of a page, and rel="describedby" pointing at the llms.txt that covers it. Both work as HTML <link> elements or as HTTP Link: headers.

That's the whole standard. You can write a compliant file in fifteen minutes.

Who actually reads it: the uncomfortable part

Let's do the bad news properly, with sources.

Ahrefs, June 15, 2026. They looked at server logs for 137,210 domains that had traffic in May 2026. 28% of them published an llms.txt. Of those files, 97% received zero requests during the month. Not "few requests." Zero.

Of the 3% that were fetched at all, here's who did the fetching:

Requester category Share of requests
SEO audit tools 21.7%
Unknown / unidentified 14.9%
General web crawlers 13.1%
Tech profiling tools 11.6%
AI agents and infrastructure 10.5%
GEO/AEO tools 5.8%
AI training crawlers 5.3%

GPTBot was 4.51% of requests. ClaudeBot was 0.80%. AI retrieval bots — the ones that fetch pages to answer a user's question right now — were 1.1% of the total.

Read that table again. The single largest consumer of llms.txt files in May 2026 was SEO audit tools checking whether the file exists. Ours is one of them. That's a genuinely funny thing to have to admit in a post like this, so I'm admitting it up front.

Ahrefs' own conclusion: "If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration."

SE Ranking, November 7, 2025. Different sample, different method: roughly 300,000 domains, Spearman correlation plus an XGBoost model with SHAP analysis, looking for any relationship between having llms.txt and being cited in LLM answers. They found none. Their model actually got more accurate when they removed llms.txt as a variable — it was adding noise, not signal. Their adoption number was 10.13%, well below Ahrefs' 28%; different samples, different dates, and I'd trust the direction of both more than either exact figure.

Google. On May 15, 2026, Google Search Central published its guide to optimizing for generative AI features. It says you don't need to create machine-readable files, AI text files, markup, or Markdown to show up in Google Search or its AI features. Gary Illyes said at Search Central Live in July 2025 that Google doesn't support llms.txt and has no plans to. John Mueller compared it on Reddit to the keywords meta tag — a thing the site owner asserts about itself, which is exactly why nobody trusts it — and when asked whether Google publishing its own llms.txt files counted as an endorsement, answered "no."

The AI labs. OpenAI, Anthropic, and Perplexity all publish llms.txt files for their own developer docs. None of them has published anything saying their assistant reads your llms.txt when answering a user. Publishing one and consuming one are different claims, and only the first has been made.

If somebody sold you an llms.txt file as the thing that gets you into ChatGPT, you were sold something nobody has demonstrated.

Who actually reads it: the part that got buried

Now the other half, which the "llms.txt is dead" takes skip.

The file isn't being consumed by answer engines. It's being consumed by agents — software doing a job on a user's behalf, right now, with a token budget.

  • Coding agents pointed at documentation. LangChain ships mcpdoc, an open-source MCP server whose entire job is to expose llms.txt files to host applications — it names Cursor, Windsurf, Claude Desktop, and Claude Code — and hand them a fetch_docs tool to read the URLs listed inside. The agent reads your index, picks the two pages it needs, and skips the rest of your site.
  • Documentation platforms. Mintlify auto-generates llms.txt and llms-full.txt for every docs site it hosts. Fern does the same. This is why developer-docs sites dominate the adoption numbers — for most of them, nobody decided anything; the platform shipped it.
  • Chrome. Lighthouse 13.3 added an Agentic Browsing category, and one of its audits is llms.txt. Google's own Lighthouse documentation says: "Without this file, agents may spend more time crawling the site to understand its high-level structure and primary content." The audit passes if the file is retrieved, fails on a server error, and is marked Not Applicable on a 404, because publishing one is still optional.
  • Cloudflare recommends llms.txt alongside per-page Markdown endpoints in its docs-for-agents guidance, on the straightforward grounds that Markdown wastes fewer tokens than HTML.

So Google Search says skip it and Google Chrome audits for it. That looks like a contradiction and isn't quite one — it's two product teams answering two questions. Search is telling you it won't move rankings or AI Overviews. Chrome is telling you an agent visiting your site will work faster if you have one. Both are true at the same time.

That's the honest frame. llms.txt is not a search-visibility play. It's an agent-ergonomics play. If your client's business involves anybody's AI agent landing on their site to look something up — and over the next couple of years, that's most businesses — a good llms.txt is a cheap courtesy to that agent. If your client wants to be cited in AI answers, the file is not the lever. Schema, answerable content, and letting the crawlers in are the levers.

What to actually put in it

Here's a complete, spec-valid file for a small local business. This is the shape our own generator produces, genericized:

# Ridgeline Plumbing & Heating

> Licensed plumbing and HVAC contractor serving Boulder County, Colorado since
> 2004. Emergency service available 24/7. Residential and light commercial.

## Services

- [Emergency plumbing repair](https://example.com/emergency-plumbing): 24/7
  dispatch, typical arrival within 90 minutes in Boulder and Longmont
- [Water heater replacement](https://example.com/water-heaters): tank and
  tankless, same-day install on most models
- [Furnace and AC service](https://example.com/hvac): maintenance contracts and
  emergency repair

## About

- [About Ridgeline](https://example.com/about): family-owned, 12 licensed
  technicians, Colorado master plumber license #PL-12345
- [Service area](https://example.com/service-area): Boulder, Longmont,
  Louisville, Lafayette, Erie

## Contact

- [Contact and booking](https://example.com/contact)

Phone: (303) 555-0142

## Optional

- [Financing options](https://example.com/financing)
- [Careers](https://example.com/careers)

Four things make that file work, and they're the same four things whether you're a plumber or a SaaS company:

  1. The H1 is the actual business name. Not "Home." Not "Welcome." The name a person would say out loud.
  2. The blockquote answers what, where, and who for, in one paragraph. If an agent reads nothing else, this is what it takes away. Write it like the answer to "so what do you do?"
  3. Every link has a note after the colon. [Water heater replacement](url) tells an agent nothing it couldn't guess from the URL. : tank and tankless, same-day install on most models is the part that gets you picked over a competitor's identical link.
  4. The low-value pages are under ## Optional. Careers and financing aren't why anyone's here. Marking them optional is a signal that you understand the format, and it lets an agent on a tight context budget skip them cleanly.

Five ways to get it wrong

The empty file. A file containing an H1 and nothing else is spec-valid and useless. It passes the presence check on every tool that looks, including ours, and helps nobody. Plenty of sites shipped one to check a box.

The template nobody filled in. I have seen live llms.txt files still containing > Optional description goes here. That's the literal placeholder from the spec's own example. Search your file for the word "Optional" outside an H2 and go look at what you find.

Treating it as a robots.txt copy. They do unrelated jobs. robots.txt controls access; llms.txt describes content. Pasting User-agent: and Disallow: lines into llms.txt produces a file that is neither. If you want to control AI crawler access, do it in robots.txt, where GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, and Bingbot are actually looking.

Letting it go stale. This is the real maintenance cost and the reason to keep the file short. Every link in there is a promise. When a page moves or dies, the file silently degrades — an agent following a dead link doesn't get an error it can reason about, it just gets less of your content. Ten good links you'll actually maintain beat sixty you won't.

Writing it for the machine instead of the reader. The temptation to stuff llms.txt with claims that don't appear on your actual pages is exactly why Mueller made the meta-keywords comparison, and exactly why nobody weights it heavily. If the file says something the site doesn't, you've built a small cloaking problem for no benefit. Keep it true to the pages it points at.

What it's worth in our scoring, and why

We score llms.txt as one of six modules in the audit, and it's worth 5 points out of 100. Four checks: the file exists and is reachable (2 points), it has an H1 (1), it has a blockquote summary (1), it has at least one section of links (1).

Five points is deliberate. Given everything above, weighting this module heavily would be dishonest — it would let a site paper over a blocked GPTBot or missing schema with a fifteen-minute file. Structured Data is worth 28 points and Content Structure is worth 20 for a reason: those are the things the evidence actually connects to citations.

But it isn't zero either, and I don't think it should be. It's cheap, it's forward-looking, the agentic use case is real and growing, and a site that's bothered to publish a good one has usually bothered with the rest too. Five points is roughly what "cheap early-mover signal with a real but narrow use case" is worth. If the agent traffic keeps climbing, that number goes up, and I'll say so here when it does.

The paid audit generates a complete, filled-in llms.txt for the site it scanned — real business name, real description, real service links pulled from the crawl — as one of three paste-ready assets in the PDF. If the site already has one, it extends it rather than overwriting: it only adds the pieces that are missing and leaves everything the owner already published alone.

The 20-minute version

If you're a consultant and a client asks whether they need llms.txt, here's the answer I'd give:

It won't get you cited in ChatGPT. Anyone telling you otherwise is guessing, and the two largest studies to date both came back empty. Publish one anyway, because it takes twenty minutes, it's genuinely useful to the agent tooling that does read it, Chrome now audits for it, and being early on a cheap convention has never cost anybody anything. Then go spend the rest of the afternoon on the things that actually move citations — schema, question-shaped headings with real answers under them, and making sure your robots.txt isn't quietly locking the AI crawlers out.

That last one is worth checking before anything else. It's the most common serious failure we see, it's almost never intentional, and it makes everything else on this list irrelevant while it's broken.

The free Grade Check scans one page and gives you a 0–100 score and a letter grade off the same engine that runs the paid reports, including whether the llms.txt is there and whether it's any good. Takes about a minute, costs nothing, and works on any URL — yours or a client's.

Run it before you write the file. You may find out the file is the least of the problems.


Every claim about third-party adoption, crawler behavior, and vendor positions in this post is dated and sourced above, and accurate as of September 2026. This is a fast-moving area — we re-verify these claims quarterly and will update or pull this post if the data changes.

Frequently Asked Questions

Does ChatGPT read llms.txt?

OpenAI has never published anything saying ChatGPT, GPTBot, or OAI-SearchBot parse llms.txt or treat it specially. In Ahrefs' May 2026 server-log study of 137,210 domains, GPTBot accounted for 4.51% of the requests to llms.txt files that were fetched at all — and 97% of published llms.txt files got no requests whatsoever that month.

Does Google use llms.txt for rankings or AI Overviews?

No. Google's May 15, 2026 Search Central guide to optimizing for generative AI features states you don't need to create machine-readable files, AI text files, markup, or Markdown to appear in Google Search or its AI features. Gary Illyes said at Search Central Live in July 2025 that Google doesn't support llms.txt and isn't planning to.

So is llms.txt worth publishing?

Yes, but for the right reason. It won't buy you AI search citations. It's cheap insurance and a real convenience for the agentic tooling that does read it — coding agents pointed at your docs, MCP servers, and Chrome's Lighthouse Agentic Browsing audit. Treat it as a 20-minute job, not a strategy.

What does a valid llms.txt file contain?

Per the llms.txt v2 spec (August 10, 2026), the only required element is an H1 with the site or project name. Optionally: a blockquote one-paragraph summary, plain Markdown prose, and H2-delimited sections containing lists of Markdown links, each optionally followed by a colon and a short note. An H2 named 'Optional' marks links an agent can safely skip.