vrid.ai Logo

llms.txt SEO: do you need it in 2026? The truth table

llms.txt truth table for 2026: what Google, ChatGPT, Claude, and Perplexity actually do with the file, backed by Ahrefs crawler-log data.

23 min read

Do you need llms.txt in 2026? What each AI engine actually does with it

TL;DR: No AI engine treats llms.txt as a ranking or citation signal in 2026. Google’s own Search team says it does not read the file. Ahrefs checked 137,210 domains and found that of the roughly 38,000 with a valid llms.txt, 97% got zero requests for it in a full month, and AI retrieval bots made up only 1% of the traffic that did arrive. Publish one anyway if you already have clean docs to point it at; skip it if you’d have to build the content just to fill the file.


Table of contents

  1. The one-sentence answer
  2. What llms.txt actually is
  3. What Google says, on the record
  4. What the AI engines actually do with it
  5. Inside the Ahrefs 137,210-domain study
  6. Engine-by-engine truth table
  7. Should you build one? A decision tree
  8. How to build an llms.txt file that isn’t wasted effort
  9. What llms.txt does not replace
  10. Why Cloudflare, Anthropic, and Vercel still publish one
  11. Common mistakes when people build llms.txt anyway
  12. Frequently asked questions
  13. Key takeaways

The one-sentence answer

llms.txt is a plain-Markdown file at your domain root that lists your key pages with one-line descriptions, and as of August 2026 no major AI search engine reads it often enough to move your citations or rankings, so treat it as optional low-cost infrastructure, not an SEO tactic.

That single sentence resolves the debate that has run since the file’s proposal in September 2024: is llms.txt the new sitemap.xml, or is it a dead file format that a few hundred blog posts talk each other into implementing? The evidence from actual server logs says the second one, with one narrow exception worth understanding before you decide.

What llms.txt actually is

Jeremy Howard, co-founder of Answer.AI and previously of fast.ai, proposed llms.txt on September 3, 2024. The specification lives at llmstxt.org and defines two files: /llms.txt, a Markdown index of your most important pages with a one-line description each, and an optional /llms-full.txt, which concatenates your entire site’s content into a single Markdown document.

The pitch was straightforward. Large language models burn tokens and time parsing HTML, navigation chrome, and JavaScript to find the actual content on a page. A curated Markdown index, the argument went, would let an AI system skip the noise and go straight to what matters, the same way robots.txt and sitemap.xml give crawlers a map instead of forcing them to discover everything by brute force.

It is not a technical standard ratified by any search engine or AI lab. No engine has to obey it, parse it, or even fetch it. That distinction matters more in 2026 than it did at launch, because the gap between “here is a proposal” and “here is what engines actually do” has now been measured directly, not guessed at.

What Google says, on the record

Google’s position on llms.txt has not shifted since it was first asked, and by mid-2026 the company has said it plainly enough times that there is no ambiguity left. Gary Illyes of Google Search Relations confirmed at Google Search Central Live in mid-2025 that Google does not support llms.txt and has no plans to, a position Search Engine Land reported directly from the session.

John Mueller went further and gave the comparison that stuck: he likened llms.txt to the keywords meta tag, a self-declared signal Google abandoned decades ago because site owners gamed it the moment it carried any weight, a point covered in Search Engine Journal’s writeup of his comments. A self-declared “here is what matters on my site” file has the exact same incentive problem llms.txt would create if Google ever weighted it: every site would claim everything matters.

There is a wrinkle that confused people for a few weeks in 2025. Some Google-owned properties were serving an llms.txt file of their own. Mueller clarified, as reported by Baseline Labs, that this happened because an internal content system added the file automatically and nobody removed it. Google’s own Search team neither uses it nor endorses it, on Google’s own domains.

Google’s 2026 guidance for site owners on AI features now names llms.txt directly in a mythbusting section as a tactic that does not help rank content in AI Overviews or AI Mode, closing the loop on a claim that used to get relitigated in every SEO Twitter thread. If your traffic already took a hit when Overviews rolled out, the fix is not a file Google has said three times it ignores. It is the 30/60/90-day recovery plan keyed to how much CTR you actually lost.

What the AI engines actually do with it

Google not reading the file is the easy part of this story. The harder, more interesting part is whether the AI answer engines, the ones llms.txt was built for, actually fetch it. This is not a matter of policy statements. It is measurable from server access logs, and multiple independent log studies in 2026 measured it the same way.

A 30-day server-log study of AI crawler behavior, published by DigitalApplied, and a separate multi-site log analysis covering GPTBot, ClaudeBot, and PerplexityBot found the same result: zero hits on /llms.txt or /llm.txt from any of the three major AI crawlers across the monitored period. Each bot identifies itself clearly in access logs (GPTBot/1.2, ClaudeBot/0.1, PerplexityBot/1.0), so this is not a detection problem. The bots simply do not request the file in practice, despite growing publisher adoption.

That finding lines up with what happens when a crawler visits your site for other reasons. GPTBot crawls primarily to gather training data on a scheduled basis; ChatGPT-User fetches pages live when a user’s query triggers a browse action; ClaudeBot is Anthropic’s primary crawler for both training and retrieval. None of the three, per the log studies above, treat /llms.txt as a stop on that route.

Inside the Ahrefs 137,210-domain study

The most rigorous data point on this whole debate comes from Ahrefs, and it is worth walking through in detail because it settles the adoption-versus-usage gap that most llms.txt coverage glosses over.

Ahrefs analyzed all 137,210 domains in Ahrefs Web Analytics that received traffic in May 2026, checking each domain root for an llms.txt file returning HTTP 200, then cross-referencing Ahrefs Bot Analytics for every request to /llms.txt paths across that population, split by response code and by individual user agent. The full methodology and numbers are in Ahrefs’ published study.

Three numbers from that study do the real work:

  • 28% of the studied domains publish a valid llms.txt file. More than one in four sites have already done the work, despite no major AI platform having committed to reading it.
  • Of the roughly 38,000 domains with a valid file, 97% received zero requests for it in May 2026. Only about 1,100 domains saw any traffic to the file at all.
  • Of the requests that did arrive, 96% came from bots that are not AI retrieval bots. SEO audit tools accounted for 21% of requests, unidentified bots 14%, traditional web crawlers like Googlebot 13%, and tech-profiling tools like BuiltWith 11%. AI retrieval bots tied to ChatGPT and Perplexity made up roughly 1% of the total.

Put plainly: the file is being built faster than it is being read, and the small amount of reading that does happen is mostly other SEO tools checking whether you built one, not AI engines using it to answer questions. Adoption estimates elsewhere land in a similar range depending on sample: Rankability measured 8.7% of the top 1,000 sites globally as of August 2026, while a 300,000-domain SE Ranking sample put adoption at 10.13%. The gap between “top 1,000 sites” and “all sites” tells you adoption skews toward large, well-resourced brands with documentation teams already producing Markdown, not toward the average blog.

None of this means llms.txt is fraudulent or pointless as a concept. It means the file, in its current form, is infrastructure that a small number of AI systems occasionally touch, not a lever that moves your visibility in ChatGPT or Perplexity answers the way a backlink or a structured-data fix might.

Engine-by-engine truth table

EngineReads llms.txt todayOfficial statementWhat actually moves visibility instead
Google Search / AI Overviews / AI ModeIllyes: no support, no plans. Mueller: compares it to the keywords meta tag. Named directly in Google’s 2026 mythbusting guidance.Crawlable HTML, structured data where it fits your content type, and the same signals that already drive organic rankings
ChatGPT (GPTBot, OAI-SearchBot, ChatGPT-User)Effectively ✗ in practiceNo public promise to read llms.txt as a ranking inputCrawlable content, brand mentions, and citation-worthy pages other sites and forums already link to
Claude (ClaudeBot)Effectively ✗ in practiceNo public commitment; log studies show zero fetch activity across monitored sitesClean Markdown-adjacent HTML structure and content that answers a question directly in the first few sentences
Perplexity (PerplexityBot)Marginal, ~1% of the tiny fraction of requests that occur at allNo public commitment; Ahrefs data shows AI retrieval bots (ChatGPT + Perplexity combined) at roughly 1% of llms.txt requestsRecency, source diversity, and being the kind of page Perplexity already cites heavily, including Reddit threads
Chrome Lighthouse (Agentic Browsing category, v13.3+)✓ checks for presenceGoogle’s own Lighthouse team added an “Agentic Browsing” audit category that checks whether /llms.txt exists, explicitly separate from the SEO category and Search rankingNot a ranking factor; a browser-agent readiness signal for tools like Chrome’s own agent mode
MCP-connected coding agents and doc tools✓ meaningfullyTools like Cursor, Claude Code, and Mintlify’s MCP server actively consume llms.txt-style indexes when a developer points an agent at your docsStructured, current documentation exposed through an MCP server or a clean llms.txt index

The one row that actually reads “yes” with confidence is the last one, and it is not a search engine. It is a developer tool.

Should you build one? A decision tree

flowchart TD
    A[Do you already have clean docs<br/>or a product site in Markdown-friendly form?] -->|No| B[Skip llms.txt.<br/>Building the content to fill the file<br/>costs more than the file returns.]
    A -->|Yes| C{Do developers or agents<br/>need to consume your docs<br/>via Cursor, Claude Code, or an MCP server?}
    C -->|Yes| D[Build llms.txt now.<br/>This is the one path with<br/>measured real usage.]
    C -->|No| E{Is your goal ranking in<br/>Google, ChatGPT, or Perplexity?}
    E -->|Yes| F[Do not rely on llms.txt.<br/>Fix crawlability, structured data,<br/>and citation-worthy content instead.]
    E -->|No, just future-proofing| G[Publish a minimal llms.txt<br/>as low-cost hygiene.<br/>Zero downside, near-zero current upside.]

The decision collapses to one question: do you have an AI-agent audience, meaning developers or tools that will point Cursor, Claude Code, or an MCP client at your documentation? If yes, build it properly, because that is the one place the file gets read. If your goal is search or AI-answer visibility, the file is not the lever, no matter how many “essential 2026 SEO checklist” posts list it as step one.

How to build an llms.txt file that isn’t wasted effort

If you land on “build it” from the decision tree, the file itself takes under an hour if your content is already organized. Keep it to the actual spec instead of the bloated versions circulating in generator tools.

Structure. Start with an H1 that is your site or product name, followed by a one-paragraph summary in blockquote format. Then group your key pages under H2 sections (Docs, API Reference, Guides, whatever matches your site), with each page as a Markdown link plus a one-line description of what it covers. Do not paste your entire sitemap. The value of the file, if it has any, is curation: the ten to thirty pages that actually explain what your product does, not every URL you’ve ever published.

Skip /llms-full.txt unless you have a real reason. The full-content companion file balloons in size on any site with more than a handful of pages, and nothing in the log studies above shows engines fetching it more than the index file. It adds maintenance burden (every content update means updating two files instead of one, or your /llms-full.txt goes stale and starts contradicting your live pages) for a payoff nobody has measured.

Keep it current or don’t bother. A stale llms.txt pointing to deprecated pages or dead links is worse than no file, because the rare agent that does fetch it gets bad instructions. If you cannot commit to updating it when your docs change, the file will actively mislead the one class of consumer that reads it.

Validate before you publish. Confirm the file returns a 200 status at the exact root path (https://yourdomain.com/llms.txt, not buried in a subdirectory), uses valid Markdown, and does not 404 or redirect. Several free validators exist for this; a manual curl check on the live URL works just as well.

Do not treat this as your AI visibility strategy. If llms.txt is the only line item on your AI-search checklist, run the 40-check AI visibility audit instead. The checks that actually correlate with citations, crawlability, structured data fit, content freshness, and third-party mentions, will surface the gaps a Markdown index file cannot fix.

What llms.txt does not replace

The most common mistake in how teams talk about llms.txt is treating it as a substitute for foundational technical SEO, when it addresses none of the same problems.

It does not replace robots.txt, which controls crawl permissions and blocking, a decision with real publisher-economics tradeoffs covered in the AI crawler blocking framework. It does not replace sitemap.xml, which search engines actually parse and use for discovery and indexing at scale, unlike llms.txt. It does not replace structured data, which Google has separately said is not required for AI Overviews eligibility but which multiple engines do use as a parsing aid, a nuance covered in the structured-data-for-AI-search breakdown. And it does not replace basic crawlability: if your content sits behind JavaScript rendering that bots can’t execute, or your pages return soft-404s, no Markdown index file at your root fixes that underlying problem.

llms.txt is additive, at best, to an already-solid technical foundation. Publishing one on a site with crawl errors, missing schema, or thin content is decorating a house with a broken foundation.

Why Cloudflare, Anthropic, and Vercel still publish one

If the data above is this consistently negative, why do large, sophisticated engineering organizations keep publishing llms.txt files? Cloudflare, Anthropic, Vercel, Stripe, Mintlify, and Supabase all maintain one, and the reason has nothing to do with search rankings.

These companies are optimizing for a different consumer: coding agents and MCP-connected tools, not search crawlers. Vercel’s llms.txt file for its API documentation includes contextual descriptions specifically so an agent like Claude Code or Cursor can decide which endpoints to fetch without pulling in irrelevant pages, a use case Mintlify describes directly in its own documentation about serving agent-friendly content. LangChain built mcpdoc, an MCP server that exposes llms.txt content to IDEs, giving developers direct control over what context an agent pulls from a site’s docs.

The economic logic is different from SEO. A developer using Cursor or Claude Code to build against your API is a much smaller audience than “everyone searching Google,” but a far higher-intent one; getting your docs into an agent’s context window cleanly, without the agent burning tokens parsing your marketing site’s navigation and cookie banner, is worth the file’s minimal build cost even with zero search benefit. This is the honest case for llms.txt in 2026: it’s Business-to-Agent infrastructure for developer tools, not an Answer Engine Optimization tactic, however the term has been marketed.

If your business has an API, an SDK, or technical documentation that developers integrate against, this is the one segment where llms.txt earns its build time. If your business is a blog, a services site, or ecommerce with no developer audience, the case collapses back to “skip it or publish minimal hygiene,” per the decision tree above.

Getting a workflow right end to end, from keyword research through publishing across WordPress, Ghost, Shopify, and webhook destinations, matters more for actual AI-search visibility than any single file at your root. Vrid.ai handles the keyword research and multi-channel publishing side of that workflow, but no publishing tool, including this one, changes whether GPTBot requests /llms.txt. That decision sits with OpenAI, Anthropic, and Perplexity, and as of August 2026 the log data says they mostly don’t make it.

Common mistakes when people build llms.txt anyway

Teams that decide to build the file, whether for the developer-agent use case or as pure hygiene, tend to repeat the same handful of errors.

Copy-pasting the sitemap. The most common failure mode is generating llms.txt from a sitemap export, which produces a file with hundreds of undifferentiated links and no descriptions. That defeats the entire premise. Curation is the only value proposition the format has.

Treating it as a ranking tactic and reporting on it as one. Teams that add “llms.txt published” to a monthly SEO report as a win are setting up a false causal story the first time AI-referral traffic moves for an unrelated reason. Track AI referral traffic in GA4 properly and you won’t be tempted to credit a file that Ahrefs’ own bot logs show almost nobody requests.

Letting it go stale. A file built once during a redesign sprint and never touched again becomes a liability, pointing agents at deprecated pages or old product names months later.

Skipping the fundamentals to spend time on it. The opportunity cost is the real risk. An hour spent perfecting an llms.txt file is an hour not spent on the 11 AEO mistakes that Ahrefs, Google, and crawler logs all point to as the things that actually correlate with getting cited: crawlable content, clear direct answers near the top of the page, and the kind of brand-mention signal that AI engines demonstrably do weigh.

Frequently asked questions

What is llms.txt in simple terms?

It’s a plain Markdown file at your domain root (yourdomain.com/llms.txt) listing your most important pages with one-line descriptions, proposed by Jeremy Howard of Answer.AI in September 2024 as a way for AI systems to find key content without crawling your whole site. It is not required by any search engine or AI lab; it’s an optional convention some sites choose to publish.

Does llms.txt help SEO rankings on Google?

No. Google Search Relations has said this directly and repeatedly: Gary Illyes confirmed Google does not support llms.txt, and John Mueller compared it to the abandoned keywords meta tag. Google’s 2026 AI-features guidance names llms.txt by name as a tactic that does not affect ranking in Search, AI Overviews, or AI Mode.

Does ChatGPT read llms.txt?

Not in any measurable volume as of 2026. Server-log studies tracking GPTBot, OAI-SearchBot, and ChatGPT-User found zero requests to /llms.txt across monitored sites, and Ahrefs’ 137,210-domain analysis found AI retrieval bots tied to ChatGPT and Perplexity combined made up roughly 1% of the already-small fraction of requests that llms.txt files receive at all.

Does Perplexity actually use llms.txt?

Marginally, and it’s the closest thing to a “yes” among AI engines, but the volume is tiny. Perplexity and ChatGPT together accounted for about 1% of llms.txt requests in Ahrefs’ study, versus 96% from SEO tools, unidentified bots, and traditional crawlers like Googlebot. Don’t build a strategy around a 1% share.

Does Claude use llms.txt for search results?

Independent server-log studies covering ClaudeBot found zero fetch activity on /llms.txt paths across the monitored sites. Claude’s own product team does use llms.txt-style indexes for a different purpose, feeding Claude Code and MCP-connected developer tools, which is a distinct use case from Claude’s web search or citation behavior.

Is llms.txt the same as sitemap.xml?

No. Sitemap.xml is a machine-readable index that major search engines including Google actually fetch, parse, and use for crawling and indexing decisions, a role confirmed by Google’s own documentation and years of Search Console data. llms.txt has no equivalent commitment from any engine; it’s a proposed convention, not an adopted standard.

Is llms.txt the same as robots.txt?

No, and they solve different problems. Robots.txt tells crawlers what they’re allowed to access; it’s honored by compliant bots and enforced at the protocol level for search engines. llms.txt is a content index with no access-control function and no confirmed engine that reads it as a rule.

Should a small business website build an llms.txt file?

Generally, no, unless the business sells developer tools, an API, or technical documentation that other developers integrate against using coding agents. For a typical small-business site, an hour spent on llms.txt has near-zero measured return; the same hour spent fixing crawl errors, adding relevant structured data, or improving page-load speed has a clearer, evidenced payoff.

What does Chrome Lighthouse check about llms.txt?

Chrome Lighthouse version 13.3 and later added an “Agentic Browsing” category, separate from the SEO and Performance categories, that checks for the presence of a machine-readable summary like llms.txt at the domain root. Google has been explicit that this category measures browser-agent readiness, not Search ranking eligibility; a page can fail every Agentic Browsing check and still rank normally in Search.

Who actually uses llms.txt today?

The clearest confirmed use case is developer tooling: Cursor, Claude Code, and MCP-connected agents that developers point at a company’s documentation to pull context while coding. Companies including Anthropic, Cloudflare, Vercel, Stripe, Mintlify, and Supabase publish llms.txt files specifically to support this workflow, not to chase search or AI-answer citations.

How many websites have an llms.txt file in 2026?

Estimates vary by sample. Ahrefs found 28% of 137,210 traffic-receiving domains had a valid file as of August 2026. Rankability measured 8.7% of the global top 1,000 sites as of August 2026. A separate 300,000-domain sample from SE Ranking put adoption at 10.13%. The spread reflects that adoption skews toward larger, better-resourced sites.

What percentage of llms.txt files actually get requested by anything?

In Ahrefs’ study of roughly 38,000 domains with a valid llms.txt file, 97% received zero requests for it during May 2026. Only about 1,100 domains saw any traffic to the file, and most of that traffic came from SEO audit tools and non-AI crawlers, not AI systems.

Does llms.txt help with AI Overviews or AI Mode specifically?

No. Google’s 2026 guidance names llms.txt directly as a tactic with no effect on AI Overviews or AI Mode eligibility. Ranking in those surfaces depends on the same underlying signals that drive normal organic ranking, since Google has stated good SEO and good GEO substantially overlap for its own AI features.

Should I use an llms.txt generator tool?

If you decide to build one, a generator can save formatting time, but check its output against the actual spec at llmstxt.org before publishing. Many generator tools produce bloated files that dump your full sitemap with auto-written descriptions, which defeats the curation value the format is supposed to provide and can look worse than having no file at all.

Will llms.txt matter more in the future even if it doesn’t help now?

Possibly, but nobody can state that as fact today. No major AI lab has committed to weighting it, and the closest thing to future-relevance evidence is Chrome’s experimental Lighthouse category and MCP tooling adoption, both of which are about agent interaction rather than search or answer-engine citation. Treat any claim that it “will matter soon” as speculation, not a documented roadmap item.

What should I do instead if I want AI engines to cite my site?

Focus on what the evidence actually links to citations: crawlable, unblocked content; direct answers near the top of the page; content that other sites, forums, and especially Reddit already reference; and consistent brand mentions across the web, which multiple AI engines use as a trust signal independent of backlinks.

Does having llms.txt hurt my site in any way?

There’s no documented downside to publishing a well-maintained, accurate llms.txt file; it doesn’t dilute crawl budget in any measured study and costs nothing to serve. The risk is entirely opportunity cost (time spent on it instead of higher-impact work) and the smaller risk of a stale file misleading the rare agent that does fetch it.

How is llms.txt different from an XML feed or API for AI agents?

llms.txt is intentionally simpler: static Markdown, no authentication, no query parameters, just a curated link list. An API or MCP server, by contrast, allows an agent to query for specific, current information and act on it. That’s part of why serious MCP-connected tooling (Cloudflare’s docs, Mintlify’s MCP server) has moved past static llms.txt files toward live MCP servers as the more capable agent-facing layer.

Should I put llms.txt in my technical SEO audit checklist?

Include it as a hygiene check if you like, similar to confirming your favicon or humans.txt exists, but do not weight it alongside crawlability, indexability, or structured-data checks that are proven to correlate with rankings and citations. The 40-check AI visibility audit treats it exactly this way: present or absent, not pass or fail.

What’s the single biggest misconception about llms.txt?

That it functions like a modern, AI-era sitemap.xml that search and answer engines actively parse for ranking decisions. It does not. The people who built the standard proposed a genuinely useful idea for agent-to-content routing, but the AI search engines the format was popularly marketed toward, Google, ChatGPT, Claude, and largely Perplexity too, simply do not fetch it in any volume that matters as of August 2026.

Key takeaways

  • Google’s Search team has said three separate times, through Illyes, Mueller, and its own 2026 mythbusting guidance, that llms.txt does not affect rankings or AI Overviews eligibility.
  • Ahrefs’ 137,210-domain study found 97% of the roughly 38,000 valid llms.txt files got zero requests in May 2026, with AI retrieval bots making up about 1% of the tiny remainder.
  • Server-log studies tracking GPTBot, ClaudeBot, and PerplexityBot directly found zero fetch activity on /llms.txt across monitored sites.
  • The one confirmed, meaningful use case is developer-agent tooling: Cursor, Claude Code, and MCP servers that developers point at documentation, which is why Cloudflare, Anthropic, Vercel, and Mintlify still publish the file.
  • If you have clean docs and a developer audience, build it in under an hour. If your goal is search or AI-answer visibility, spend that hour on crawlability, structured data, and citation-worthy content instead.

Run the AI visibility audit checklist against your own site before you spend another minute on llms.txt. It will tell you, with pass/fail criteria per check, whether the gap keeping you out of AI answers is something a Markdown index file was ever going to fix.

Related Posts