Why Google Says llms.txt Will Not Help Citations

Google answered this one directly, twice, in writing. The full record of what they said, the four-part reasoning underneath it, and what actually moves citations.

RankControl9 min read
Why Google Says llms.txt Will Not Help Citations

Most of what you can learn about AI search comes from telemetry and inference, because the engines won't explain themselves. llms.txt is the rare exception. Google has answered this question directly, twice, on the record, and the two answers agree.

So this post does something unusual for the genre. It lays out what Google actually said and walks you through the reasoning underneath it. Then it clears up the one apparent contradiction and spends the rest of your time on the things Google says do matter.

The record, dated

The first answer came from John Mueller in 2025, and it was blunt enough to make the trade press. We think it has aged well. Asked about the file, he said that as far as he knew no AI service had said it uses llms.txt, and that server logs show they don't even check for it. Then he compared it to the keywords meta tag: a site owner's claim about their own site, which an engine could only trust by checking the site directly, at which point checking the site directly would be simpler.

The confusing moment came on May 5, 2026, when Chrome shipped an experimental Agentic Browsing category in Lighthouse. One of its audits checks whether an llms.txt exists and loads cleanly, so screenshots went around labeled "Google made it official." You can see why people read it that way, even though the reading was wrong.

Ten days later, on May 15, Google published its first official AI search optimization guide, and it named the file. Site owners can skip it, the guide says, because it's unnecessary for visibility in Google's generative AI features. So the team that actually builds the answers put it in writing that the file plays no role in them.

There was also a quieter answer the whole time. Google's crawler and Search Central documentation spells out robots.txt handling and sitemap ingestion in exacting detail, and it has never listed llms.txt as an input to anything. With crawl-and-rank systems, missing documentation has always been the real answer, and May 2026 just made it explicit.

The reasoning, unpacked

Google's position isn't arbitrary, and the reasoning is worth more to you than the verdict, because it applies well past this one file. You'll find four strands of it in Google's statements.

Start with the index Google already has. AI Overviews and AI Mode retrieve from Google Search, the most heavily engineered map of the web ever built. Googlebot feeds it by reading your actual pages, your sitemaps guide it, and your structured data helps it interpret what it finds. A markdown index at your site root gives that system nothing it doesn't already have.

The second strand, the heart of the keywords-meta-tag comparison, is that self-declaration can't carry trust. When sites control an input completely and engines can't verify it cheaply, it gets abused the moment it matters. That's what happened to the keywords tag, and it's why engines spent twenty years moving trust away from claims and toward observed behavior, like links, mentions and the content itself. Making llms.txt a citation input would bring back the oldest spam vector on the web under a fresh filename, and everyone at Google knows that history firsthand.

Third, everything the file offers already exists in a form Google can verify. The file's honest pitch, helping agents with limited context skim your site, is real and modest. Its SEO pitch duplicates infrastructure that's trusted precisely because it can be checked: a curated list of your important pages is a sitemap, and a machine-readable description of your content is structured data, plus the content itself.

The fourth strand, the one we'd hold onto, sits in Mueller's aside. An engine could only trust the file by checking it against your site, and once it has checked your site, the file hasn't contributed anything. Any input that costs as much to verify as doing the original work is economically dead as a signal. You can run that test on half the AI-SEO tactics on sale this year.

RANKCONTROL

Your competitors are building backlinks while you read this.

Organic outreach, social mentions, and link exchanges, with managed backlinks available as an add-on. Grow your domain authority without running the campaign yourself.

The Lighthouse contradiction, resolved in a paragraph

The one genuinely confusing data point deserves a plain answer, and the community asked it in exactly the right words: if the file has no impact, why did Google add it to Lighthouse audits?

r/SEO_LLM· u/arjun_rao7· May 21, 2026

If llms.txt has no impact on AI visibility, why did Google add it into Lighthouse audits?

↑ 4 upvotes26 comments
Via Reddit

Two Google teams were answering two different questions. Chrome's Lighthouse team builds developer tooling for an agentic-browsing future, and it shipped an experimental, optional audit that treats the file as a courtesy for agents operating sites. That audit marks a missing file "not applicable," which isn't how a requirement behaves.

The Search team, which builds the citation systems, says the file plays no role in them. Each team is right about its own area, and an experimental nod from a nearby team doesn't turn anything into a ranking input. Only screenshots stripped of context made them look like they disagree, and the fuller episode has its own post.

But what about the other engines?

Google's reasoning would matter less to you if the rest of the field worked differently, and it doesn't. ChatGPT's stack retrieves through Bing and its own index, Claude's browsing leans on Brave, and Perplexity runs its own crawl. All of them read pages, and none of them documents consuming llms.txt.

The measured behavior matches that silence. Across 137,000 domains, 97 percent of the files received zero requests in a month. The trust argument travels too, since self-declaration is as easy to game for OpenAI as for Google. The model teams spent 2026 tightening source selection toward verifiable, official content, which makes them the last people likely to adopt an unverifiable file.

Google's statements are just the most quotable chapter of a longer story, and the full adoption-versus-behavior data tells the rest with charts.

The "Google would say that" objection

You'll hear a reflexive counter, and it deserves airing: Google benefits when optimization stays pointed at Google's systems, so of course it would dismiss a rival standard. That's a fair instinct with an answer you can check.

Guidance can be self-interested and accurate at the same time, and you tell the two apart by triangulating. In this case every independent line of evidence sides with the self-interested party. The mechanism does, since answers retrieve from indexes and the file isn't in any of them. The behavior data agrees: 97 percent of files unfetched, measured by a third party with no stake in Google. Controlled tests, ours included, show null citation effects.

The most telling line is the silence from Google's competitors. OpenAI, Anthropic or Perplexity could win easy developer goodwill by announcing llms.txt support tomorrow, and none of them has. You should distrust Google's guidance when it conflicts with independent measurement. Here it matches that measurement and the behavior of Google's rivals, so the self-interest objection has nothing left to explain.

What to do with a straight answer

A direct, on-the-record dismissal is useful in a practical way: it turns a debate that keeps coming back into a two-sentence reply. The next time the file shows up in an audit, a vendor proposal or an article your executive forwarded, you can send something like this:

"Google says in its own AI search documentation that the file is unnecessary, no engine documents reading it, and measured citation effects are null; here's the link. The hours are reallocated to the four levers Google does document, and our weekly citation data will show their effect."

That reply does two jobs. It closes the topic with sources instead of opinions, and it sets a precedent that tactics get onto your roadmap through evidence, which disposes of the next three fads before anyone pitches them. Teams that handle llms.txt this way usually end up settling a bigger question, whether the roadmap runs on checklists or on measurement, and it's worth settling once. llms.txt is the easiest case you'll get, since Google has already put its answer in writing.

You're getting AI traffic. But do you know where it comes from?

RankControl credits every visit to the assistant that sent it: ChatGPT, Perplexity, Claude, Gemini, Copilot, or Grok. Full source attribution, next to your Google traffic.

What Google says does help, because that's the useful part

The same May 2026 guidance that dismissed the file is clear about what does matter. It matches how AI Overviews and AI Mode actually behave, point for point, and it comes down to four levers.

The first is being retrievable. The answers cite from what Search retrieves, and retrieval is rank-gated across the query's fan-out, so ordinary indexing, ranking and internal-linking discipline is your entry fee. That's the eligibility half of the checklist.

Once you're retrievable, your answers need to be liftable. Put direct answers where extraction can find them, shape your sections around questions, and give concrete facts instead of prose fog, because these systems quote whatever quotes cleanly. Keep your markup honest, too. Structured data should mirror the content people can see, which is the verified kind of machine-readability and the one Google actually documents consuming.

The fourth lever is corroboration. Trust comes from observed behavior, the only kind the keywords-tag lesson left standing, so the wider web has to agree your entity exists, describe it consistently and mention it independently.

Then measure the claims, Google's and ours, on your own data. With citations tracked per engine, weekly, against a fixed query set, you'll see each of those levers show its effect within weeks, while the courtesy file has never shown one anywhere.

The epitaph, reused

The keywords meta tag has been dead as a signal since before some of today's SEOs were born. Sites kept stuffing it for a decade anyway, because it was cheap and visible and it sat on every checklist. Google's message on llms.txt, delivered politely and twice, is that history is repeating and it would rather you skipped this round.

So take the win, since a straight answer is rare in this field. Spend the twenty minutes on the file only if agents genuinely roam your docs, and give it zero strategic attention either way. Put the energy you get back into the four levers Google both documents and rewards, and let your own dashboard, rather than anyone's blog post, be the referee.

RANKCONTROL

We'll show you exactly where your brand stands in AI search.

No commitment. $0 due today, cancel anytime. See how ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode talk about your brand today.

Frequently Asked Questions

Google has answered it twice, and the answers agree. John Mueller compared the file to the keywords meta tag and said server logs show AI services don't even check for it, and the May 2026 AI search guidance called it unnecessary for visibility in Google's generative AI features. The Lighthouse audit people point to comes from Chrome's experimental agent tooling rather than Search, and it marks a missing file as not applicable.

Because both are pure self-declaration, and a claim a site makes about itself can't carry trust. Sites abused the keywords tag immediately, so engines learned to ignore the claim and check the pages directly. llms.txt is the same kind of claim about your important content, and once an engine has read your site to verify it, the file hasn't added anything.

No, and it's the most mechanical of Google's reasons. AI Overviews and AI Mode pull from Google's own Search index, which Googlebot builds by crawling your actual pages, with your sitemap to guide it and your structured data to interpret them. A hand-written index at your site root duplicates something Google has run at web scale for decades, minus the verification.

Not in any way that shows up. No engine documents reading the file, and an Ahrefs analysis found 97 percent of llms.txt files got zero requests in a month. The other big engines, ChatGPT among them, also retrieve through search indexes rather than courtesy files, so the trust problem with self-declaration hits them just as hard.

Ranking comes first, because retrieval is rank-gated: a page that doesn't rank for the query or its fan-out variants never gets considered. Beyond that, Google's guidance wants answers stated directly enough to lift and structured data that honestly mirrors the page, and it wants the rest of the web to corroborate your entity. You can check all of it in your own data within weeks, and that's the test llms.txt has never passed.

RANKCONTROL

Turn AI search into a customer acquisition channel

Content that ranks on Google and gets cited by AI search engines. Published on your domain. Citations tracked weekly.

Related Articles

THE SIGNAL

Insights on AI and Google search strategy. No fluff.

Get the latest on AI citations, Google rankings, and content strategy.

No spam. Unsubscribe anytime.