What Is llms.txt, and Does Your Website Actually Need One?
It was pitched as "robots.txt for AI" the moment it launched. Two years of data later, that comparison looks like the single biggest reason people overestimated what this file actually does.
llms.txt is a plain Markdown file, placed at a website's root, meant to give AI tools a curated index of a site's most important pages. Jeremy Howard, co-founder of Answer.AI and fast.ai, proposed the format on September 3, 2024, and it's maintained at llmstxt.org ever since. The pitch was simple and appealing: the way robots.txt tells search crawlers what to crawl, llms.txt would tell AI models and agents what to read. Two years on, the actual data about what reads it and what it changes tells a much smaller story than the pitch did.
Writing or cleaning up a Markdown file like this? ClearMark converts Markdown to HTML and back, nothing uploaded anywhere.
The Format Itself
An llms.txt file follows a specific, minimal structure: an H1 heading with the site or brand name first, a blockquote summarizing what the site is, section headings grouping related pages, and a list of links under each section with a short description of what each one covers. It's deliberately simple — closer to a table of contents than documentation — which is part of why it spread quickly among developer tools and SaaS products: it's cheap to produce and easy to automate.
What the 2026 Data Actually Shows
Adoption climbed steadily after the 2024 launch, but usage tells a different story than adoption does. Roughly 10% of domains in a 300,000-site study have an llms.txt file in place — yet AI crawler traffic requests it in only about 0.1% of visits, meaning the overwhelming majority of crawls that could use it simply don't ask for it.
| Question | What the Data Says |
|---|---|
| Domains with llms.txt | ~10% (of 300,000 studied) |
| AI crawler requests for it | ~0.1% of crawler traffic |
| Google Search uses it? | No — confirmed by Google's John Mueller |
| Effect on AI-citation prediction models | Removing it improved model accuracy |
That last line is the most pointed finding: SE Ranking's research into what predicts whether a page gets cited by AI search tools found that including llms.txt as a factor made the predictive model less accurate — removing it from the model improved results. Rather than being a neutral factor with no effect, it was actively adding noise.
Where "Robots.txt for AI" Breaks Down
Robots.txt works because search engines built their crawlers to read and respect it — it's backed by two decades of infrastructure and convention on the crawler side. llms.txt has no such backing: no major AI company — not Google, not OpenAI, not Perplexity, not Bing — lists it as a control mechanism in their official crawler documentation. It isn't a formal web standard in the way robots.txt or sitemap.xml are; it's a convention one developer proposed that a meaningful slice of the web adopted anyway, largely on the strength of the "AI's robots.txt" framing rather than on confirmed results.
The comparison to robots.txt was never really about function — it was about borrowing legitimacy from a format everyone already trusts. Robots.txt works because every major crawler agreed, years ago, to actually read it. Nothing comparable has happened for llms.txt, and two years of adoption data suggest it mostly hasn't needed to: the traffic that would use it barely shows up.
Where It Does Seem to Help
The one validated use case sits outside search and AI-citation visibility entirely: developer-facing AI coding tools. Assistants like Cursor and GitHub Copilot retrieve a project's documentation more efficiently when a well-structured llms.txt points them directly at the right reference pages, instead of making the model crawl and infer structure from a full docs site. That's a genuinely useful, narrow win — just a different one than the "get cited by AI search" pitch that drove most of the adoption.
Should You Still Add One?
Adding an llms.txt file costs very little — it's a small Markdown file, not a structural change to the site — so there's little downside to having one. The mistake is treating it as a shortcut that substitutes for the things that actually affect whether AI tools cite a page: clear, well-structured, genuinely useful content, proper headings, and a site that's easy to crawl in the first place. An llms.txt file organizes a site that's already solid; it doesn't fix one that isn't.
The honest summary: llms.txt is a real, simple, low-cost format with one confirmed use case — helping AI coding assistants navigate developer documentation faster. The broader promise of it as "robots.txt for AI" that gets a site cited more in AI search hasn't held up: Google doesn't read it, AI crawlers barely request it, and at least one study found it actively hurt prediction accuracy when used as a ranking signal. Add it if it's easy, but don't expect it to do the heavy lifting content and site structure still have to do.
For questions or inquiries contact us at info@cleartexteditor.com