llms.txt, robots.txt & AI Crawlers: The 2026 Technical Checklist for AI Visibility

Configure llms.txt, robots.txt, and AI crawler access in under 30 minutes. Copy-paste snippets for GPTBot, ClaudeBot, PerplexityBot, and GoogleOther.

RankControl10 min read
llms.txt, robots.txt & AI Crawlers: The 2026 Technical Checklist for AI Visibility

Start by opening your robots.txt and searching it for GPTBot. Among sites that mention GPTBot in that file, 83% block it completely. If you pasted a "protect your content" snippet from some 2024 blog post, there's a real chance ChatGPT can't see your site right now, and you'd only find out by looking.

The bots haven't stayed away. Cloudflare's 2025 crawler report found GPTBot traffic up 305% in a single year, which gave site owners a figure for what they'd suspected. Whether the crawlers can read what they find is the part you control.

Three files decide it. You know two of them already (robots.txt and your sitemap), and the third is llms.txt, which the SEO crowd can't stop arguing about. Every config below pastes in as is, and the job takes about 30 minutes.

Every AI crawler you need to know

Before you change a line, find out who's visiting. Some AI bots fetch pages to answer search questions, others collect text for model training, and a few do both depending on the day. That matters because a robots.txt rule names one user-agent, and that user-agent only covers one of those jobs.

CrawlerCompanyPurposerobots.txt Respected?
GPTBotOpenAITraining + search retrievalYes
ChatGPT-UserOpenAILive browsing (user-initiated)No (acts like a browser)
ClaudeBotAnthropicTraining + search retrievalYes
PerplexityBotPerplexitySearch retrievalYes
GoogleOtherGoogleAI features, GeminiYes
Google-ExtendedGoogleTraining onlyYes
BytespiderByteDanceTraining (TikTok/Doubao)Sometimes
CCBotCommon CrawlOpen training datasetYes

ChatGPT-User and Claude-User are the odd ones out. Neither follows robots.txt, since each fetches a page for a person who asked for it, the way a browser would. So your GPTBot block stops OpenAI's crawler from indexing you for search, yet someone chatting with ChatGPT can still have your page opened mid-conversation. That request comes from ChatGPT-User, a separate system with its own user-agent.

RANKCONTROL

26 content formats. Published on your domain. Matched to your brand.

Guides, comparisons, listicles, case studies, and more. RankControl generates content that gets cited by ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode.

robots.txt: what to allow and what to block

For most SaaS sites I'd let in the crawlers that feed search answers and keep out the ones that only collect training data. Start from this file:

# Allow AI search crawlers
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: GoogleOther
Allow: /

# Block training-only crawlers
User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

Three mistakes can quietly undo that file.

Hunt first for a forgotten Disallow: / sitting under User-agent: *. Any crawler without a block of its own falls under that wildcard, including AI bots so new you haven't heard of them.

Next, if GPTBot is blocked, ask why. Maybe a 2024 blog post told you to, and back then that was sensible, because training was the big worry. In 2026 GPTBot also powers ChatGPT's search feature. Block it and ChatGPT can't retrieve your pages when people search there.

Then make sure GoogleOther appears in the file at all. Google uses it for AI features such as Gemini and AI Overviews. A file that only mentions Googlebot keeps you in traditional search while Google's AI products can't see you.

Why split it this way? Cloudflare's breakdown of AI crawler traffic answers that:

PurposeShare of AI crawler traffic
Training76.2%
Search18.5%
User-initiated browsing4.1%

Training dwarfs the rest. With the file above, you stay visible in the 18.5% that's search and opt out of the 76.2% that goes to training.

What llms.txt actually does (and what it doesn't)

Mention llms.txt in an SEO group and you'll start a fight. Several practitioners have run controlled tests and measured no impact on AI citations. SE Ranking went as far as a full analysis titled "Why Brands Rely On It and Why It Doesn't Work," and the Reddit threads are harsher still.

The skeptics have a point. None of the major AI search engines has said publicly that it reads llms.txt when deciding what to rank or cite. OpenAI hasn't, and neither have Anthropic, Perplexity or Google.

I'd write one anyway, because the file does a different job from the one people expect. It was never designed to be "robots.txt for AI," and Search Engine Land's "treasure map" label is nearer the mark: you're writing a structured summary of your site that AI systems can take in fast, so an agent reading your domain has somewhere to start.

A minimal version looks like this:

# YourCompany

> Short one-line description of what your company does.

## Docs

- [Getting Started](https://yoursite.com/docs/getting-started): Setup guide for new users
- [API Reference](https://yoursite.com/docs/api): Complete API documentation
- [Pricing](https://yoursite.com/pricing): Plans and pricing details

## Blog

- [Most Important Post](https://yoursite.com/blog/key-post): Description of the post

It goes at yoursite.com/llms.txt, under 50 lines, with only your highest-value pages in it. You're curating a short list, so pasting in your whole sitemap misses the point.

One complaint we see again and again is fair: most of these files list URLs with no context at all. A file of bare links is barely more useful than your sitemap. When you write yours, describe each page in plain English, saying what it does and who it's for.

Your competitors are getting cited by AI. You're not.

Every day without citation tracking is a day your competitors pull ahead in ChatGPT, Perplexity, and Claude.

Show me who's getting cited→2-minute overview · real case-study numbers

The complete 2026 AI visibility checklist

Seven settings decide whether AI search engines can reach and read your pages, and then whether they cite you. We track all seven across our customer base, and the gap is wide: sites that get every one right see 3-4x more AI citations than sites that only handle two or three.

CheckWhy it matters
robots.txt allows GPTBot, ClaudeBot, PerplexityBot, GoogleOtherA blanket Disallow: / can override the specific allows; test with Google's robots.txt tester
No X-Robots-Tag: noindex on key pagesHTTP headers can block AI crawlers even when robots.txt is clean
llms.txt at domain rootYour 10-20 most important pages with clear descriptions, kept curated and updated quarterly
Sitemap.xml submitted and currentAI crawlers use sitemaps to discover pages; check that lastmod dates are accurate
FAQ schema on key pagesFAQPage structured data in JSON-LD; AI models parse structured data faster than unstructured prose
Content starts with direct answersAI models pull citations from the first 30% of page content
Server response time under 3 secondsAI crawlers have timeout limits, and slow pages get abandoned mid-crawl

Headers are the sneaky one, because you only see an X-Robots-Tag if you curl your key pages and read what comes back. A stale or broken sitemap is quieter still: your new content just never gets found. For the schema, our beginner's guide to getting cited walks you through the setup.

I'd look hardest at your opening paragraphs. If they tell the brand story instead of answering the user's question, you won't get cited. Speed hides problems too. If you're doing server-side rendering with heavy database queries, your technical pages may be timing out for bots while they load fine for people.

This list overlaps with your wider SEO and GEO strategy and doesn't replace it. You need a clean crawl layer, but it won't win citations alone, since content quality and topical authority still drive most of them. For the deeper technical layer, meaning how AI agents perceive your product, see our guide on making your SaaS discoverable by AI agents.

How to verify your setup

Setup took 30 minutes. Give it 15 more to prove it works, and don't skip this part.

First, pull your live robots.txt and see what it tells each AI crawler:

curl -s https://yoursite.com/robots.txt | grep -A 2 "GPTBot"
curl -s https://yoursite.com/robots.txt | grep -A 2 "ClaudeBot"
curl -s https://yoursite.com/robots.txt | grep -A 2 "PerplexityBot"

Then request one of your important pages while posing as GPTBot. You're hoping for a 200, which means the page came through:

curl -s -o /dev/null -w "%{http_code}" \
  -H "User-Agent: Mozilla/5.0 (compatible; GPTBot/1.0)" \
  https://yoursite.com/your-important-page

A 403 or 429 means your server or CDN is turning AI bots away at the infrastructure level, whatever your robots.txt says about "Allow."

Your access logs come next. Search them for GPTBot, ClaudeBot, PerplexityBot and Bytespider. Zero hits from any of them over a 30-day window means one of two things: robots.txt is keeping that bot out, or it hasn't found your site yet.

Finally, confirm llms.txt is live:

curl -s -o /dev/null -w "%{http_code}" https://yoursite.com/llms.txt

A 404 means the file isn't deployed, so check your hosting platform's static file configuration.

The CDN and WAF gotcha

This one catches more people than you'd guess. Your robots.txt is perfect, your llms.txt is deployed, and still the AI crawlers can't read your site. The blocker is something like Cloudflare's Bot Fight Mode, AWS WAF or Vercel's bot protection, stopping them in your infrastructure before the request reaches your server.

Open your CDN's bot management settings. Cloudflare's Bot Fight Mode is aggressive by default and can block legitimate AI crawlers. Your curl tests will tell you: if AI user-agents get a 403 while a normal browser user-agent gets a 200, this is almost certainly your problem.

Each provider fixes it its own way. On Cloudflare you add WAF custom rules that allow specific AI bot user-agents. Vercel users should check vercel.json for bot-blocking headers. On AWS, update your WAF rules to whitelist the GPTBot, ClaudeBot and PerplexityBot user-agent strings.

We lost two weeks to this on a client site last quarter. Everything in robots.txt looked right. The real problem sat three layers deeper in the infrastructure stack.

The part nobody talks about: monitoring

Six months after you get everything working, a developer pushes a new robots.txt rule, or your CDN vendor changes its bot protection defaults, or a WordPress plugin update overwrites your config. Your AI visibility drops to zero, and weeks can go by before anyone notices.

Setting up crawler access is a one-time job. Catching the day it breaks is the hard part, and it's why RankControl checks your crawler access daily and your citations weekly. If a robots.txt change blocks a crawler, or your server starts answering AI bots with 403s, you'll see it in the AI Crawling tab before your citations disappear.

You can monitor all of this by hand. That means curling your robots.txt once a week and reading your server logs for bot traffic, plus asking ChatGPT about your brand name yourself. Done thoroughly, it's 2-3 hours a week. Or RankControl's agents handle it automatically while you get on with building your product.

RANKCONTROL

200+ SaaS teams already track their AI citations.

They know exactly when ChatGPT mentions their brand, and when it stops. Do you?

Show me the plan→One plan · everything included

Frequently Asked Questions

They do different jobs. robots.txt is the gate, where you allow or block each crawler, and llms.txt is the map: a plain text file at your domain root that tells AI models which of your pages hold your most important content. Keep yours short and describe every page in plain English.

At the very least, let in OpenAI's GPTBot for ChatGPT and Anthropic's ClaudeBot for Claude, plus PerplexityBot and GoogleOther, which covers Google AI. Block only the training-specific crawlers such as Google-Extended or CCBot, and only if you want to stay out of training while keeping your search visibility.

Not by itself. No major AI search engine has confirmed reading it for ranking, so it won't guarantee you citations. It belongs on the wider technical checklist next to robots.txt access and structured data, along with how your content is formatted, and if you skip that whole stack you're invisible.

Fetch your robots.txt and look for Disallow rules under the AI user-agents, then curl a few pages with each bot's user-agent string. After that, search your server logs for GPTBot, ClaudeBot and PerplexityBot hits. If you see zero AI bot traffic, something is blocking them.

Blocking all of them opts you out of AI search entirely, which is a steep price to pay. I'd block selectively: let search-facing crawlers like GPTBot and PerplexityBot in and keep the training-only ones out. Your content stays out of training datasets and you stay visible in AI answers.

RANKCONTROL

Get mentioned by ChatGPT, Claude, and Perplexity

Content that ranks on Google and gets cited by AI search engines. Published on your domain. Citations tracked weekly.

Related Articles

THE SIGNAL

Insights on AI and Google search strategy. No fluff.

Get the latest on AI citations, Google rankings, and content strategy.

No spam. Unsubscribe anytime.