llms.txt: what it is, who reads it, and whether you need one
llms.txt is a Markdown file of links for AI agents. Google does not use it, coding agents do, and Lighthouse now audits it. How to write one.
LLaunchScaler·Published ·8 min read
llms.txt is a Markdown file served at /llms.txt that gives AI agents a short summary of your site and a curated list of links to its most useful pages, so an agent can find the right page without crawling everything. Google does not use it and its John Mueller has said no AI system does, coding agents and documentation tools do read it, and Chrome's Lighthouse now audits it, so for most sites it is a ten-minute extra rather than a priority.
The rest of this guide covers the format, a worked example for a SaaS, who reads the file and what Lighthouse checks.
What is llms.txt?
llms.txt is a proposal by Jeremy Howard, published at llmstxt.org on September 3, 2024, for a file that helps language models use a website at inference time. It points agents to concise, clean versions of your key pages instead of HTML full of navigation and scripts. A second version of the proposal came out in August 2026.
The proposal's reasoning is practical: "An HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise," and context windows are "still too small for most websites in their entirety." An agent reads the small llms.txt file, picks the links it needs and fetches only those.
Version 2 made three changes worth knowing. A file can sit at the root or at any path, and "covers the URLs under its path," so /docs/llms.txt covers /docs/. Pages can offer Markdown versions at the same URL with .md added or swapped in. And sites can point to both with standard link relations: rel="alternate" type="text/markdown" for a page's Markdown version and rel="describedby" for the llms.txt file that covers it, as HTML <link> elements or an HTTP Link header.
What goes in an llms.txt file?
The spec asks for sections in a fixed order, and only the first is required. The file starts with an H1 naming the site, then a blockquote summary, then optional notes without headings, then H2 sections that each hold a Markdown list of links. Each link can carry a colon and a short description.
Questions, answered
What people ask about this
01
What is llms.txt?
llms.txt is a Markdown file at /llms.txt that gives AI agents a short summary of a site and a curated list of links to its most useful pages. Jeremy Howard proposed it at llmstxt.org in September 2024 and published a second version in August 2026.
The name of the site or product; "the only required section"
> Summary
No
One or two sentences with the key facts an agent needs to understand the rest
Notes
No
Paragraphs or lists, no headings, on how to read the links
## Section with links
No
A Markdown list of [Title](URL): description entries
## Optional
No
By convention, secondary links an agent can skip; in v2 it no longer has special mechanical meaning
The proposal's own guidelines are short: use concise, clear language, give each link a brief, informative description, avoid unexplained jargon, and test the file by asking an agent questions about your product with only the llms.txt as its starting point.
What does an llms.txt look like for a SaaS?
A SaaS file should answer the questions an agent is most likely asked about the product: what it is, what it costs, how to set it up, what changed recently. Link the pages that answer those, with a description that tells the agent what each page holds. This example is for a fictional scheduling product with docs, pricing and a changelog.
# Acme Scheduling
> Acme Scheduling is appointment software for clinics with 2 to 50 providers. It handles online booking, reminders and insurance details, and syncs with Google Calendar and Outlook.
Plans are billed per provider per month. The API and webhooks are available on the Team plan and above.
## Product
- [Pricing](https://www.acmescheduling.com/pricing.md): Plans, per-provider prices, what each plan includes, and the free trial terms
- [Features](https://www.acmescheduling.com/features.md): Booking, reminders, waitlists, intake forms and reporting
- [Security](https://www.acmescheduling.com/security.md): Data hosting, encryption and compliance details
## Docs
- [Quick start](https://docs.acmescheduling.com/quick-start.md): Set up a clinic, add providers and publish a booking page in 15 minutes
- [API reference](https://docs.acmescheduling.com/api.md): REST endpoints, authentication and rate limits
- [Webhooks](https://docs.acmescheduling.com/webhooks.md): Event types, payloads and signature verification
- [Integrations](https://docs.acmescheduling.com/integrations.md): Google Calendar, Outlook and EHR connections
## Changelog
- [Changelog](https://www.acmescheduling.com/changelog.md): Dated release notes, newest first
## Optional
- [Blog](https://www.acmescheduling.com/blog): Guides for clinic managers
- [About](https://www.acmescheduling.com/about.md): Company, team and contact details
Three choices make it useful. The summary states facts an agent can repeat, not a slogan. Each description says what the page contains, so the agent can choose without opening it. And the links point to Markdown versions where they exist, which v2 recommends, because clean text is what the agent wanted in the first place. If you do not publish .md versions, link the normal HTML pages instead; a link that 404s is worse than an HTML link.
Who reads llms.txt?
Coding agents and documentation tools read it; Google does not, and Google's own staff have said consumer chatbots did not fetch it. The proposal itself says llms.txt files "are used most heavily for software documentation, where coding agents follow them to find API references and tutorials." The table sets out what each reader has said or shown.
Reader
Uses llms.txt?
What the source says
Google Search, AI Overviews, AI Mode
No
Google: "You don't need to create new machine readable files, AI text files, or markup to appear in these features"
Consumer AI chatbots
Not as of June 2025, per Google's John Mueller
"FWIW no AI system currently uses llms.txt"; chatbots "will fetch your pages... but none of them fetch the llms.txt file"
Google as a company
Not adopting it
Gary Illyes at Search Central Live APAC, July 2025: "not a Google initiative and not seen as beneficial, or something they're looking to adopt"
Coding agents and docs platforms
Yes
llmstxt.org: used most heavily for software documentation; platforms such as Mintlify generate the file for every docs site they host
Chrome Lighthouse
Audits it
Checks the file in its Agentic Browsing category
The proposal notes that the AI labs publish llms.txt files for their own developer documentation, naming OpenAI, Anthropic and Google's Gemini docs. That is consistent with the table: the file is a documentation convention that agents working on code rely on, not a ranking signal.
Does Lighthouse check llms.txt?
Yes. Lighthouse's Agentic Browsing category, added to the default configuration in Lighthouse 13.3 in May 2026, includes an llms.txt audit. It passes a file that has an H1, contains at least one Markdown link and is at least 50 characters long. A missing file is marked not applicable, not failed.
The audit's source code spells out the logic:
If fetching /llms.txt returns a 500-range status or fails outright, the audit fails.
If it returns a 400-range status, such as 404, the audit is marked not applicable, because Google's documentation says "providing the file is optional at the moment."
If the file loads, Lighthouse checks for a line matching an H1 (# Title), for at least one link in [text](url) form, and that the file is 50 characters or longer. Each missing item is listed as an error: "File is missing a required H1 header," "File does not appear to contain any links" or "File is suspiciously short."
So a site with no llms.txt loses nothing in Lighthouse, while a site whose /llms.txt URL throws a server error fails the audit. The guide to Agentic Browsing in PageSpeed Insights covers the other audits in the category.
Do you need an llms.txt?
You do not need one to rank in Google or to appear in AI Overviews, and OpenAI, Anthropic and Perplexity say nothing about it in their crawler documentation. Publish one if developers or agents use your documentation, or simply because it costs ten minutes and a file that passes Lighthouse. Do it after the work that does affect AI visibility.
Your site
Recommendation
Developer product with API docs
Publish one, ideally with Markdown versions of doc pages; coding agents are the main readers
SaaS marketing site with a pricing page and some docs
Optional; the example above takes minutes, and it passes the Lighthouse audit
Content site or blog
Low value; Google says it does not use it, and the chat engines' crawler documentation does not mention it
Site whose content only renders with JavaScript
Fix rendering first; an llms.txt pointing at pages crawlers cannot read helps no one
The things that decide whether AI engines can use your site come first: robots.txt that allows the search crawlers, content in the server-rendered HTML, and pages that answer questions directly. The guide to optimizing for AI search covers those, and Organization schema markup is the structured-data equivalent that search engines do read.
How do you publish and test an llms.txt?
Serve it as plain text or Markdown at /llms.txt, with a 200 status and UTF-8 encoding, and make sure every link in it resolves. In September 2026, OpenAI's developers.openai.com/llms.txt was served as text/plain; charset=utf-8, and Google's Gemini API docs file as text/markdown; charset=utf-8; both work.
Save the file as llms.txt in the folder your site serves from its root, or generate it at build time from the same source as your pages.
Check the status and type: curl -sI https://yoursite.com/llms.txt, looking for 200 and a content-type of text/plain or text/markdown.
Check every link: curl -s https://yoursite.com/llms.txt | grep -oE "\(https?://[^)]+\)" | tr -d "()" | xargs -n1 curl -s -o /dev/null -w "%{http_code} %{url_effective}\n", and fix anything that is not 200.
Run Lighthouse or PageSpeed Insights and read the llms.txt audit under Agentic Browsing.
Ask an agent a few questions about your product with only the llms.txt URL as context, as the proposal suggests, and fix the descriptions it misreads.
If you also want a single file holding all of your content, that is llms-full.txt, a separate convention; the comparison of llms-full.txt and llms.txt covers when it is worth adding.
Check your AI readiness before you write llms.txt
An llms.txt file is the last step, not the first. Run the free scan on LaunchScaler: it takes a URL, needs no account and runs 156 checks across 6 of its 7 categories at no cost. It reports a missing llms.txt as a warning only, never a failure, and puts it below the checks that decide whether AI engines can read you at all: robots.txt rules for OAI-SearchBot, Claude-SearchBot and PerplexityBot, firewall challenges to those crawlers, and content missing from the raw HTML.
No. Google's John Mueller wrote in June 2025 that 'no AI system currently uses llms.txt,' Gary Illyes said in July 2025 that it is not a Google initiative, and Google's AI features guide says you don't need AI text files to appear in AI Overviews or AI Mode.
03
Who actually reads llms.txt?
Mainly coding agents and documentation tools. The proposal's author says it is used most heavily for software documentation, where coding agents follow it to find API references. Chrome's Lighthouse also audits it in its Agentic Browsing category.
04
What does an llms.txt file need to contain?
Only an H1 with the site or project name is required. The recommended structure adds a blockquote summary, optional notes, and H2 sections that each hold a Markdown list of links, with a short description after each link.
05
Does llms.txt help SEO or AI citations?
Not for Google: Google says it is not needed for its AI features, and OpenAI, Anthropic and Perplexity do not mention it in their crawler documentation. It can help agents that read documentation, and it takes minutes to write, so treat it as a low-cost extra after robots.txt and server-rendered content are right.
AI agents fail on challenges, fake buttons, unlabelled fields, hover-only menus and moving layouts. The test and the fix for each, in the order to check.
Measure AI visibility by hand: 10 buyer questions, 7 engines, 4 metrics (mention rate, citation rate, share of voice, accuracy), and how to read the noise.