Skip to content
Adrythm
AI search

llms.txt

llms.txt file / LLMS.txt / llms.txt SEO

In short

llms.txt is a proposed Markdown file, placed on a website, that gives AI agents a short summary of the site and links to its most useful pages. Google says Google Search ignores these files, so adding one neither helps nor harms how a site appears in Google, AI Overviews included.

Jeremy Howard published the proposal in September 2024. The file is written in Markdown. It opens with the name of the site as a heading, which is the only required part, followed by a short summary and lists of links to its key pages.

The idea is that an agent reads this one small file first, then follows only the links it needs. The proposal says the format is used most heavily for software documentation, where coding agents use it to find API references and tutorials. It also names a business outlining its structure and policies as a fit.

Google's position is direct. Its guide to generative AI features on Google Search lists llms.txt files among the things site owners can ignore, and says Google Search itself does not use them. Creating one will neither harm nor help a site's visibility or rankings in Google Search.

Google also says it is fine to keep one for other services or systems that read the file. What remains is a case about agents from other companies. As a Google tactic, Google has ruled it out.

In practice

A proposal for AI search work lists three deliverables: an llms.txt file, content rewritten into short chunks, and new schema markup for AI. Google's guide lists the first two among the things site owners can ignore for Google Search. On the third, it says no special schema markup is needed for its AI features, though structured data still helps pages qualify for rich results.

Not the same as

robots.txt
robots.txt tells crawlers which paths they may fetch. llms.txt is a reading list for agents and carries no access rules, so it cannot keep any bot away from any page.
XML sitemap
A sitemap lists the pages a site wants search engines to find. The llms.txt proposal describes its file as a curated overview for language models instead of a full list.

Why it matters to you

Money spent on an llms.txt file as a route into Google's AI answers buys nothing there, by Google's own account. What Google says its AI features need is the ordinary groundwork: pages that are crawlable, indexed and eligible to show with a snippet, carrying content worth citing.

What to ask or check

  1. 01Which AI services that you are targeting actually read an llms.txt file, and where do they say so?
  2. 02Is the llms.txt work being sold as a Google ranking or AI Overviews deliverable?
  3. 03Are the pages the file links to indexed in Google and eligible to show with a snippet?

What people get wrong

That an llms.txt file gets a site into Google's AI Overviews. Google says Google Search ignores these files, and that a page has to be indexed and eligible to show with a snippet before it can appear in its AI features.

Red flags

  • A quote that prices an llms.txt file as a way to rank in Google or its AI Overviews.

robots.txt

A robots.txt file tells crawlers which URLs they may fetch on your site. Google is explicit that it is not a way to keep a page out of search: a disallowed URL can still be indexed if other sites link to it. The convention dates to 1994 and became RFC 9309 in 2022.

Sitemap

A sitemap is a file listing the URLs on your site so search engines can find them. Google says it ignores the priority and changefreq tags and uses lastmod only when the date is verifiably accurate. Bing calls lastmod a key freshness signal. A sitemap aids discovery and guarantees nothing.

AI Overviews

An AI Overview is the summary Google sometimes places above the results, with links to sources. Google says they are only shown when its systems judge them additive to classic Search, so they often do not trigger. There is no special file or markup that gets you into them.

GPTBot

GPTBot is the web crawler OpenAI uses to collect pages that may be used to train its AI models. You allow or block it with a line in your robots.txt file. Blocking it is a separate decision from appearing in ChatGPT search, which a different crawler handles.

Structured data

Structured data is markup that describes your page in a shared vocabulary, founded by Google, Microsoft, Yahoo and Yandex. Google is explicit about the limit: using it enables a feature to be present and does not guarantee that it will be present, even when the markup is correct.

Google-Extended

Google-Extended is a robots.txt token that controls whether content Google crawls from your site may be used to train Gemini models and ground Gemini apps. Google states it does not affect your inclusion in Search or your ranking. It is a training control, not a way to stay out of search results.

Want this explained against your own numbers?

Twenty minutes, a straight answer, and no follow-up sequence if you decide not to work with us.