Start by opening your robots.txt and searching it for GPTBot. Among sites that mention GPTBot in that file, 83% block it completely. If you pasted a "protect your content" snippet from some 2024 blog post, there's a real chance ChatGPT can't see your site right now, and you'd only find out by looking.
The bots haven't stayed away. Cloudflare's 2025 crawler report found GPTBot traffic up 305% in a single year, which gave site owners a figure for what they'd suspected. Whether the crawlers can read what they find is the part you control.
Three files decide it. You know two of them already (robots.txt and your sitemap), and the third is llms.txt, which the SEO crowd can't stop arguing about. Every config below pastes in as is, and the job takes about 30 minutes.
Every AI crawler you need to know
Before you change a line, find out who's visiting. Some AI bots fetch pages to answer search questions, others collect text for model training, and a few do both depending on the day. That matters because a robots.txt rule names one user-agent, and that user-agent only covers one of those jobs.
| Crawler | Company | Purpose | robots.txt Respected? |
|---|---|---|---|
| GPTBot | OpenAI | Training + search retrieval | Yes |
| ChatGPT-User | OpenAI | Live browsing (user-initiated) | No (acts like a browser) |
| ClaudeBot | Anthropic | Training + search retrieval | Yes |
| PerplexityBot | Perplexity | Search retrieval | Yes |
| GoogleOther | AI features, Gemini | Yes | |
| Google-Extended | Training only | Yes | |
| Bytespider | ByteDance | Training (TikTok/Doubao) | Sometimes |
| CCBot | Common Crawl | Open training dataset | Yes |
ChatGPT-User and Claude-User are the odd ones out. Neither follows robots.txt, since each fetches a page for a person who asked for it, the way a browser would. So your GPTBot block stops OpenAI's crawler from indexing you for search, yet someone chatting with ChatGPT can still have your page opened mid-conversation. That request comes from ChatGPT-User, a separate system with its own user-agent.
26 content formats. Published on your domain. Matched to your brand.
Guides, comparisons, listicles, case studies, and more. RankControl generates content that gets cited by ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode.

robots.txt: what to allow and what to block
For most SaaS sites I'd let in the crawlers that feed search answers and keep out the ones that only collect training data. Start from this file:
# Allow AI search crawlers
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GoogleOther
Allow: /
# Block training-only crawlers
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Bytespider
Disallow: /
Three mistakes can quietly undo that file.
Hunt first for a forgotten Disallow: / sitting under User-agent: *. Any crawler without a block of its own falls under that wildcard, including AI bots so new you haven't heard of them.
Next, if GPTBot is blocked, ask why. Maybe a 2024 blog post told you to, and back then that was sensible, because training was the big worry. In 2026 GPTBot also powers ChatGPT's search feature. Block it and ChatGPT can't retrieve your pages when people search there.
Then make sure GoogleOther appears in the file at all. Google uses it for AI features such as Gemini and AI Overviews. A file that only mentions Googlebot keeps you in traditional search while Google's AI products can't see you.
Why split it this way? Cloudflare's breakdown of AI crawler traffic answers that:
| Purpose | Share of AI crawler traffic |
|---|---|
| Training | 76.2% |
| Search | 18.5% |
| User-initiated browsing | 4.1% |
Training dwarfs the rest. With the file above, you stay visible in the 18.5% that's search and opt out of the 76.2% that goes to training.
What llms.txt actually does (and what it doesn't)
Mention llms.txt in an SEO group and you'll start a fight. Several practitioners have run controlled tests and measured no impact on AI citations. SE Ranking went as far as a full analysis titled "Why Brands Rely On It and Why It Doesn't Work," and the Reddit threads are harsher still.
The skeptics have a point. None of the major AI search engines has said publicly that it reads llms.txt when deciding what to rank or cite. OpenAI hasn't, and neither have Anthropic, Perplexity or Google.
I'd write one anyway, because the file does a different job from the one people expect. It was never designed to be "robots.txt for AI," and Search Engine Land's "treasure map" label is nearer the mark: you're writing a structured summary of your site that AI systems can take in fast, so an agent reading your domain has somewhere to start.
A minimal version looks like this:
# YourCompany
> Short one-line description of what your company does.
## Docs
- [Getting Started](https://yoursite.com/docs/getting-started): Setup guide for new users
- [API Reference](https://yoursite.com/docs/api): Complete API documentation
- [Pricing](https://yoursite.com/pricing): Plans and pricing details
## Blog
- [Most Important Post](https://yoursite.com/blog/key-post): Description of the post
It goes at yoursite.com/llms.txt, under 50 lines, with only your highest-value pages in it. You're curating a short list, so pasting in your whole sitemap misses the point.
One complaint we see again and again is fair: most of these files list URLs with no context at all. A file of bare links is barely more useful than your sitemap. When you write yours, describe each page in plain English, saying what it does and who it's for.

Your competitors are getting cited by AI. You're not.
Every day without citation tracking is a day your competitors pull ahead in ChatGPT, Perplexity, and Claude.
The complete 2026 AI visibility checklist
Seven settings decide whether AI search engines can reach and read your pages, and then whether they cite you. We track all seven across our customer base, and the gap is wide: sites that get every one right see 3-4x more AI citations than sites that only handle two or three.
| Check | Why it matters |
|---|---|
| robots.txt allows GPTBot, ClaudeBot, PerplexityBot, GoogleOther | A blanket Disallow: / can override the specific allows; test with Google's robots.txt tester |
No X-Robots-Tag: noindex on key pages | HTTP headers can block AI crawlers even when robots.txt is clean |
| llms.txt at domain root | Your 10-20 most important pages with clear descriptions, kept curated and updated quarterly |
| Sitemap.xml submitted and current | AI crawlers use sitemaps to discover pages; check that lastmod dates are accurate |
| FAQ schema on key pages | FAQPage structured data in JSON-LD; AI models parse structured data faster than unstructured prose |
| Content starts with direct answers | AI models pull citations from the first 30% of page content |
| Server response time under 3 seconds | AI crawlers have timeout limits, and slow pages get abandoned mid-crawl |
Headers are the sneaky one, because you only see an X-Robots-Tag if you curl your key pages and read what comes back. A stale or broken sitemap is quieter still: your new content just never gets found. For the schema, our beginner's guide to getting cited walks you through the setup.
I'd look hardest at your opening paragraphs. If they tell the brand story instead of answering the user's question, you won't get cited. Speed hides problems too. If you're doing server-side rendering with heavy database queries, your technical pages may be timing out for bots while they load fine for people.
This list overlaps with your wider SEO and GEO strategy and doesn't replace it. You need a clean crawl layer, but it won't win citations alone, since content quality and topical authority still drive most of them. For the deeper technical layer, meaning how AI agents perceive your product, see our guide on making your SaaS discoverable by AI agents.
How to verify your setup
Setup took 30 minutes. Give it 15 more to prove it works, and don't skip this part.
First, pull your live robots.txt and see what it tells each AI crawler:
curl -s https://yoursite.com/robots.txt | grep -A 2 "GPTBot"
curl -s https://yoursite.com/robots.txt | grep -A 2 "ClaudeBot"
curl -s https://yoursite.com/robots.txt | grep -A 2 "PerplexityBot"
Then request one of your important pages while posing as GPTBot. You're hoping for a 200, which means the page came through:
curl -s -o /dev/null -w "%{http_code}" \
-H "User-Agent: Mozilla/5.0 (compatible; GPTBot/1.0)" \
https://yoursite.com/your-important-page
A 403 or 429 means your server or CDN is turning AI bots away at the infrastructure level, whatever your robots.txt says about "Allow."
Your access logs come next. Search them for GPTBot, ClaudeBot, PerplexityBot and Bytespider. Zero hits from any of them over a 30-day window means one of two things: robots.txt is keeping that bot out, or it hasn't found your site yet.
Finally, confirm llms.txt is live:
curl -s -o /dev/null -w "%{http_code}" https://yoursite.com/llms.txt
A 404 means the file isn't deployed, so check your hosting platform's static file configuration.
The CDN and WAF gotcha
This one catches more people than you'd guess. Your robots.txt is perfect, your llms.txt is deployed, and still the AI crawlers can't read your site. The blocker is something like Cloudflare's Bot Fight Mode, AWS WAF or Vercel's bot protection, stopping them in your infrastructure before the request reaches your server.
Open your CDN's bot management settings. Cloudflare's Bot Fight Mode is aggressive by default and can block legitimate AI crawlers. Your curl tests will tell you: if AI user-agents get a 403 while a normal browser user-agent gets a 200, this is almost certainly your problem.
Each provider fixes it its own way. On Cloudflare you add WAF custom rules that allow specific AI bot user-agents. Vercel users should check vercel.json for bot-blocking headers. On AWS, update your WAF rules to whitelist the GPTBot, ClaudeBot and PerplexityBot user-agent strings.
We lost two weeks to this on a client site last quarter. Everything in robots.txt looked right. The real problem sat three layers deeper in the infrastructure stack.
The part nobody talks about: monitoring
Six months after you get everything working, a developer pushes a new robots.txt rule, or your CDN vendor changes its bot protection defaults, or a WordPress plugin update overwrites your config. Your AI visibility drops to zero, and weeks can go by before anyone notices.
Setting up crawler access is a one-time job. Catching the day it breaks is the hard part, and it's why RankControl checks your crawler access daily and your citations weekly. If a robots.txt change blocks a crawler, or your server starts answering AI bots with 403s, you'll see it in the AI Crawling tab before your citations disappear.
You can monitor all of this by hand. That means curling your robots.txt once a week and reading your server logs for bot traffic, plus asking ChatGPT about your brand name yourself. Done thoroughly, it's 2-3 hours a week. Or RankControl's agents handle it automatically while you get on with building your product.
200+ SaaS teams already track their AI citations.
They know exactly when ChatGPT mentions their brand, and when it stops. Do you?




