An llms.txt file will not get you cited by ChatGPT, Gemini or Perplexity on its own. Google's Search documentation says Google Search ignores the file, and as of 24 September 2026 no company running an AI answer engine has said publicly that it reads llms.txt when writing answers. The file is real and cheap, coding agents do use it, and most generated versions we have looked at are broken, including the first one we built.
That last part is the one nobody writes about, so most of this article is about it.
What is llms.txt?
An llms.txt file is a Markdown file placed at a website's root, or at a subpath such as /docs/, that lists the site's most useful pages with short notes so an AI agent can find them without parsing HTML.
Jeremy Howard published the llms.txt proposal on 3 September 2024 and revised it as v2 in August 2026, according to the spec's changes page. The format is small. An H1 with the site or project name is the only required part. A blockquote summary usually follows, then optional notes, then sections under H2 headings, each holding a Markdown list of links with an optional note after a colon. A section headed "Optional" is, by convention, for links an agent can skip when it is short on context.
The v2 text says llms.txt is "used most heavily for software documentation, where coding agents follow them to find API references and tutorials." That is the author describing where the file earns its keep, and it is not search answers.
Do AI engines actually read llms.txt?
The public evidence says no for Google Search and says nothing at all for the answer engines. Here is every statement we could source to the company itself or to a named employee, dated.
Google, May 2026, last updated 10 July 2026. Google's generative AI guide for Search site owners says site owners do not need new machine-readable files, AI text files or Markdown to appear in Google Search, because Google Search does not use them. It adds that it is "completely fine" to publish llms.txt for other services, and that doing so "will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them."
Google, April 2025. John Mueller of Google Search wrote in a Reddit thread on r/TechSEO that, as far as he knew, none of the AI services had said they were using llms.txt, and that server logs show they do not even check for it. He compared it to the keywords meta tag: a site owner's claim about the site, which a system can verify by reading the site directly.
Google Chrome, by May 2026. Lighthouse had added an llms.txt audit to its experimental agentic browsing checks. It fails a page only when fetching llms.txt returns a server error, and a missing file (a 404) is marked not applicable "as providing the file is optional at the moment." That is a check that the file does not break, not evidence that any engine reads it.
OpenAI, Anthropic and Perplexity. The official crawler pages from OpenAI, Anthropic (dated 7 April 2026) and Perplexity describe robots.txt as the control site owners have. None of them says its crawler reads a site's llms.txt. All three do publish an llms.txt for their own developer documentation, which fits the spec author's point: it is a map for coding agents reading docs.
The honest summary on 24 September 2026: Google has said its Search product ignores llms.txt. OpenAI, Anthropic and Perplexity have said nothing either way about using it in answers. We found claims that they do, and every one came from a company selling llms.txt or AI-visibility tooling, so we left them out.
Is llms.txt worth adding at all?
Yes, if your expectations match the evidence. llms.txt costs an afternoon, Google says it will not hurt you in Search, and agents that fetch pages while helping a user, such as coding assistants reading your API docs, can use it as a map.
It will not fix a page no AI crawler can reach, or make a thin page quotable. Those are separate problems, and the difference between an access problem and an evidence problem decides which one you have. Treat llms.txt like a sitemap: keep it correct, then spend the real effort on the pages it points to.

Why do so many llms.txt files go wrong?
Most llms.txt files go wrong because they are generated once from whatever the site's database or sitemap says, and nobody compares the output with the live site. We know this from the inside. Before June 2026, the llms.txt generator in Azuqe made all four of the mistakes below, and a fifth turned up after we fixed those.
Guessed URLs
A generator that does not know a site's URL pattern guesses. Ours wrote every blog post as /blog/ plus the slug, because that is a common pattern, and always added an RSS link at /rss/ whether or not a feed existed there. On sites that used a different blog path, or had no feed, every one of those lines was a dead link inside a file whose only job is to point at real pages.
Leaked drafts
Our first generator selected blog posts by title, not by status, and ordered published posts first, then drafts, then planned posts. So a site's llms.txt could list articles that did not exist yet, with titles the owner had not approved, in a public file at the site root.
Page dumps
A list of every page is a sitemap, and the spec's own comparison with sitemaps says sitemaps are too large for a context window and full of pages an agent does not need. On one customer site, about 60 of the crawled URLs were individual chat pages built from one template. A generator that lists pages in crawl order cannot tell that those are one page type, and ours could not.
Never checked after deploy
Our first generator never checked the file after deploy. It produced text and stopped, and nothing fetched the site's /llms.txt afterwards to see whether the owner had uploaded it, whether the uploaded copy matched, or whether its links still worked a month later. A file that was right on the day it shipped drifts as pages move, and nobody finds out.
The redirect we missed
The rebuild added a live check on every URL, and it was still wrong in one way. The check followed redirects and accepted the final status, so /seo and /marketing on azuqe.com, which both redirect to the homepage, went into our own file as if they were pages. llms.txt is a list of canonical addresses, and a redirecting URL is a pointer to an address rather than the address itself.
What does a correct llms.txt contain?
A correct llms.txt contains only published pages whose exact URLs answer with a success status, grouped so near-identical pages collapse into their index, and it is checked again after it goes live. As a procedure:
- Build the candidate list from published pages only, never from drafts or planned posts.
- Request every candidate URL and keep it only if that exact address returns a 2xx status.
- When a candidate URL redirects, follow the redirect yourself and write the destination address into llms.txt, not the redirecting one.
- Drop any candidate URL that resolves to a page already on the list, so one page never appears under two addresses.
- Collapse sets of near-identical pages, such as individual product, profile or chat pages, into the index or category page that lists them.
- Include a sitemap or feed link in llms.txt only after requesting it and getting a success status.
- After uploading llms.txt, fetch the live /llms.txt and compare it line by line with the file you generated.
- Re-check every link in the live llms.txt each week, because pages that were live at deploy time get moved and deleted.
Steps 7 and 8 are the ones almost every generator skips, and they are the only ones that tell you the file on your server is the file you meant to publish.
What is llms-full.txt?
An llms-full.txt file is a single Markdown file holding the full text of a site's documentation, so an agent can load everything in one request. It is not part of the llms.txt spec. It is a convention that documentation platforms popularised, and Mintlify's documentation describes it as a file that "combines your entire documentation site into a single file."
For a documentation site with a few dozen pages, llms-full.txt is sensible. For a marketing site or a store, it is usually the page dump problem in a larger file, and it goes stale faster than llms.txt does because every edit to every page changes it.

Why Azuqe is the best option
Azuqe builds llms.txt only from URLs it has just confirmed are live, then checks the file you actually deployed and tells you when it has drifted. The four failures above are the design brief, because we shipped all of them first.
| Check | What most generators ship | What Azuqe does |
|---|---|---|
| URLs | Written from a pattern or a sitemap | Requests each URL, keeps only a 2xx at that address, rewrites redirects to their destination |
| Blog posts | Every post with a title | Published posts only |
| Similar pages | Listed one by one | Grouped under their index page; an invented URL is dropped |
| After deploy | Nothing | Fetches your live /llms.txt and reports it as deployed, outdated or missing |
| Over time | Nothing | Re-verifies when the last check is more than a week old |
- Azuqe compares the live file with a fresh build and lists the links that are missing from it and the links in it that no longer work.
- Azuqe checks whether a /llms-full.txt exists on your site.
- Azuqe keeps llms.txt as one item on the readiness checklist beside the measurement it runs across seven AI engines, so the file is never mistaken for the result.
Two limits, stated plainly. Deploying the file is manual today: Azuqe has no access to your hosting, so you download llms.txt and upload it to your site root yourself, and Azuqe then confirms whether it is live. And Azuqe does not build llms-full.txt; it only reports whether you have one.
Related reading
For what the file cannot fix, read how to rank in AI Overviews, starting with retrieval. For what makes the pages behind the links quotable, read how language models choose what to quote. For the whole picture, start with AI visibility, what it is and how to move it.
Frequently asked questions
What is an llms.txt file for?
An llms.txt file gives AI agents a short, curated map of a site's most useful pages in Markdown. Its heaviest real use, according to the spec itself, is coding agents finding API references and tutorials in software documentation.
Does llms.txt actually work?
It works as a map for agents that choose to fetch it. No AI answer engine has said publicly that it reads llms.txt when writing answers, and Google says Google Search ignores it, so do not expect it to change whether you are cited.
What is the purpose of llms.txt in SEO?
For Google Search, none: Google's own guide says the file neither helps nor harms rankings. Its value sits outside search ranking, with agents that read your site while helping a user.
Does ChatGPT read llms.txt?
OpenAI has not said that it does. Its crawler documentation describes robots.txt as the control for its bots and does not mention reading a site's llms.txt, although OpenAI publishes one for its own developer docs.
Does llms.txt replace robots.txt?
No. robots.txt tells crawlers what they may fetch, and the AI companies' crawler pages all point to it. llms.txt suggests what is worth reading and grants or blocks nothing.



