On our evidence we could not detect one. We took the 50 most-cited domains in our category on 11 August 2026 and checked each for an llms.txt file. Twenty-nine publish one and twenty-one do not, and the two groups rank almost identically: mean citation rank 25.8 for adopters against 25.0 for non-adopters, where 1 is the most-cited domain of the fifty. The two most-cited domains in the entire corpus, YouTube and Reddit, publish no llms.txt at all.
That is a null result rather than a refutation, and the difference matters. We could not detect an effect. That is not the same as proving there is none, and the section at the end explains why this particular test could not have detected a small one.
GetIntel generates llms.txt files as a product feature, so the commercial incentive here runs towards telling you it works. The dataset is published with every domain, its citation count and whether it had a file, so you can check the claim rather than trust it.
What did we measure?
Our probe corpus holds 47,093 citations across 4,211 distinct domains. We took the top 50 domains by citation count, requested https://<domain>/llms.txt for each, and counted it as present only on an HTTP 200 with a text/plain or text/markdown content type. A 200 returning an HTML error page does not count, which is a mistake worth avoiding because plenty of sites serve their 404 page with a 200 status.
| group | domains | mean citations | median citations | mean rank |
|---|---|---|---|---|
| Publish llms.txt | 29 | 338.8 | 286.0 | 25.8 |
| No llms.txt | 21 | 464.0 | 318.0 | 25.0 |

Read the rank column rather than the means. Mean citations look worse for adopters, but that is two outliers doing the work: YouTube at 1,876 citations and Reddit at 1,701, neither of which publishes the file. Rank is resistant to that, and on rank the two groups are indistinguishable.

Do the most-cited sources in the category have one?
Mostly not, and that is the part that should give an llms.txt advocate pause. YouTube and Reddit are the two most-cited domains in our corpus by a wide margin. Neither publishes llms.txt. Nor do arxiv.org, linkedin.com, ahrefs.com, techradar.com or seranking.com, all of which sit in the top twenty.
Meanwhile a large share of the adopters are AI-visibility vendors: by our own classification, published as a column in the dataset so you can disagree with it, 14 of the 29 sell tools in this category. Our read is that they publish llms.txt because they write about llms.txt, and that they get cited because they are the subject matter rather than because of a text file in their web root. We did not measure why any domain was cited, so treat that as a hypothesis.
If the file were a meaningful retrieval signal, the sources engines reach for most often would be the last places you would expect to find it missing.
Do AI engines actually read llms.txt?
There is no public evidence that any major engine consumes it at retrieval time, and one has said plainly that it does not.
Google has stated publicly that it does not use llms.txt for Search or AI Overviews. OpenAI, Anthropic and Perplexity have made no commitment either way, which means adoption rests on inference rather than on anything documented. Our measurement is a different kind of evidence: it cannot see inside a retrieval pipeline, but if a file were being consumed and weighted, you would expect the domains engines cite most to be the ones publishing it, and in our corpus they are conspicuously not.
That is the honest state of it. One documented no, several silences, and no positive evidence from anyone measuring outcomes.
Why is llms.txt so widely recommended then?
Because it is cheap, legible and satisfying, and because the alternative advice is harder.
A file you can generate in a minute and tick off a checklist has obvious appeal against "earn a mention in the Reddit thread your buyers actually read". The recommendation spreads because it is easy to give, not because the evidence behind it is strong.
The evidence that does exist points the other way. SE Ranking's study across roughly 300,000 domains found no measurable citation effect, and Google has said publicly that it does not use the file. Our own measurement is a third, much smaller data point pointing in the same direction from a different angle.
Should I still publish one?
Yes, but for the right reasons and with the right expectations.
It costs almost nothing. It is a reasonable place to state canonical descriptions of your product and link your important pages, and it is genuinely useful to a human evaluator or an agent that has been pointed at your site deliberately. None of that requires it to influence retrieval.
What you should not do is treat it as a visibility lever, sequence it ahead of work that has evidence behind it, or let a vendor sell it to you as one. If your AI visibility plan has llms.txt near the top, the plan is ordered by ease rather than by effect.
We publish one. We also would not claim it has moved a single citation, and nothing in our data suggests it has.
What this test could not have detected
Fifty domains is a small field, and a null result on fifty observations only rules out a large effect. A modest one, say a few percent, would be invisible at this sample size and this test would look exactly the same.
It is also correlational and confounded in a specific direction. Adoption skews towards category vendors who would be cited on category questions regardless, which if anything works in llms.txt's favour here: the adopters had a structural advantage and still did not out-rank the others.
We measured presence, not quality. A well-maintained llms.txt and an autogenerated stub both counted the same, and it is possible the file only helps when it is genuinely good.
And this is one category on one day. Ours is a young AI-tooling niche where Reddit and YouTube dominate. A category with different source dynamics could behave differently, and the same measurement can look very different per engine, which is a reason to be careful with any single pooled number including ours.
How to test it on your own site
You do not need our data. Take the prompts your buyers ask, record which domains get cited over a few weeks, then check those domains for llms.txt yourself. It is one HTTP request per domain.
If the sources winning your category do not have the file, you have your answer for your category, which is the only one that matters to you. If they mostly do, that still does not establish cause, but it is at least a reason to look closer.
We have since run the same test on structured data and got the same answer: across the 40 most-cited pages, those carrying JSON-LD are cited no more often than those without.
The broader habit is worth more than the specific finding: before adopting a tactic because it is widely recommended, check whether the pages already winning actually do it. Most of our own published pages earn no citations at all, and we only know which ones work because we counted rather than assumed.
