AI Crawler Checker: Find Out If Your Site Blocks ChatGPT, Claude or Perplexity
Your site loads fine in every browser. Google has it indexed. Then a buyer pastes your pricing page into ChatGPT and gets told the page can't be opened, and you have no idea why.
That gap is what this AI crawler checker is for. It knocks on your homepage the way ChatGPT, Claude and Perplexity do, then the way a normal browser does, and shows you which ones got in. It also reads your robots.txt for the training crawlers, so you can see the whole picture in one place. It's free, there's no signup, and it takes about 20 seconds.
Most people find out about a block the slow way. Traffic from AI answers never starts, or it stops, and nobody can say when it changed.
What the AI crawler checker tests
Each AI company sends more than one crawler, and they don't do the same job. Some fetch pages so an AI answer can quote them today. Others collect pages to train future models. Treating them as one thing is an easy mistake, and it can cost you.
The checker sends a real request to your homepage as each of the six search and browsing crawlers below. For the six training crawlers, it reads your robots.txt rules, since blocking those is a choice many sites make on purpose.
| Crawler | Who sends it | What it does |
|---|---|---|
| OAI-SearchBot | ChatGPT search | Finds pages ChatGPT can show and link in its answers |
| ChatGPT-User | ChatGPT browsing | Opens a page when someone asks ChatGPT to read it |
| Claude-SearchBot | Claude search | Finds pages Claude can use when it searches the web |
| Claude-User | Claude browsing | Opens a page when someone asks Claude to read it |
| PerplexityBot | Perplexity search | Finds pages Perplexity can cite in its answers |
| Perplexity-User | Perplexity browsing | Opens a page a Perplexity user asked about |
| GPTBot | OpenAI training | Collects pages that may train future models |
| ClaudeBot | Anthropic training | Collects pages that may train future models |
| Google-Extended | Gemini training | Tells Google if Gemini may train on your pages |
| CCBot | Common Crawl | Builds a public web archive many AI labs train on |
| Applebot-Extended | Apple AI training | Tells Apple if its models may train on your pages |
| Meta-ExternalAgent | Meta AI training | Collects pages for Meta's AI models |
The first six are the ones that decide whether AI search can see you. If any of them gets turned away, that engine can't read your page when a buyer asks about you. The last six only affect training, which has a slower and much weaker link to whether you show up in answers.
Two places a block can hide
A crawler can be stopped at two points on its way to your page, and many free checkers only look at one of them.
The first point is your robots.txt file. It's a short text file at the root of your site that tells each crawler what it may fetch. Well-behaved AI crawlers read it and obey it. A single line meant for a training bot can catch a search bot too, which is how a lot of sites end up blocking the crawler they actually wanted.
The second point is your firewall, CDN or security plugin. These sit in front of your site and can refuse a request before your server ever sees it. Bot protection tools often have an AI bot setting, and it's easy to switch on without noticing that it covers search crawlers too. When that happens, robots.txt still says "welcome", your server logs show nothing, and the crawler gets a 403.
That's why the checker sends live requests and doesn't stop at robots.txt. It compares what each AI crawler gets back with what a normal browser gets back. When the browser loads the page and the crawler gets refused, something in front of your site is picking on that crawler by name.
Open doors are step one. RankControl fills your site with pages buyers search for, then tracks what AI search does with them.
How to read your results
Every search crawler gets one of five labels, and each one points to a different fix. Training crawlers get a sixth label of their own.
| Label | What happened | What to do |
|---|---|---|
| Allowed | The crawler got your page and robots.txt lets it in | Nothing. This is what you want |
| Blocked by robots.txt | A Disallow rule covers this crawler | Edit robots.txt (see the fixes below) |
| Likely blocked | The crawler was refused while a browser got through | Check your firewall, CDN or security plugin |
| Challenged | Every visitor, including our browser, got a challenge page | Your bot shield is on for everyone, so we can't tell |
| Not reached | Your site didn't load for us at all | Check that the site is up and not slow or geo-blocked |
| Blocked (your choice) | robots.txt blocks a training crawler | Fine if you meant it. It doesn't affect AI search |
We made one choice on purpose when we built the checker. A blocked training crawler never shows as an error, because blocking it is a fair call for a lot of sites. It shows as "Blocked (your choice)" so you can confirm it was your call, and the search results above it stay the ones that matter.
"Likely blocked" is worded with care too. The checker sends each crawler's name from our own servers. Some firewalls also check where a request comes from, so a real crawler can be treated a bit differently from our test. Read the label as a strong hint, then confirm it in your firewall's event log by searching for the crawler's name.
What we found on 88 sites that rank on page one
To see how common this is, we ran the checker on 88 sites in October 2026. They were every site, other than forums and social networks, that ranked in Google's top 10 for 12 everyday buyer searches. The searches covered software, online shops, home services, insurance and legal help, and agencies.
Seventy of the 88 sites let all six AI search crawlers in with no trouble. The other 18, about one in five, turned at least one of them away.
| What we found | Sites |
|---|---|
| All six AI search crawlers allowed | 70 |
| Likely blocked at the firewall, with robots.txt saying yes | 7 |
| Challenged every visitor, browser included | 6 |
| Blocked by a robots.txt rule | 3 |
| Didn't load for us | 2 |
The number that stood out to us was seven. On each of those sites, robots.txt allowed every AI search crawler, yet the crawlers got a 403 while our browser loaded the page fine. A checker that only reads robots.txt would have given all seven a clean pass.
Firewall blocks outnumbered robots.txt blocks more than two to one, seven sites to three.
Online shops had it worst. Six of the 16 store sites blocked or challenged AI search crawlers, and half of those were big retailers whose bot shields challenge every visitor. Software sites did better, with four problems out of 29.
One site let Claude's crawlers through but refused the ones from ChatGPT and Perplexity. Blocks aren't always all or nothing, which is why the checker lists every crawler on its own line.
Training crawlers were a quieter story. Only four of the 88 sites blocked any training crawler at all. The other 84 left every one of them alone.
How to fix each kind of block
The fix depends on where the block lives. Start with the label the checker gave each crawler.
Blocked by robots.txt
Open your robots.txt file by adding /robots.txt to your domain in a browser. Look for a Disallow line that names the crawler, or a broad rule under "User-agent: *" that catches it. Then add an allow rule for each search crawler you want in. These lines cover the three big AI search engines.
| Line in robots.txt | What it does |
|---|---|
| User-agent: OAI-SearchBot | Starts the rules for ChatGPT search |
| Allow: / | Lets it fetch every page |
| User-agent: ChatGPT-User | Starts the rules for ChatGPT browsing |
| Allow: / | Lets it fetch every page |
| User-agent: Claude-SearchBot | Starts the rules for Claude search |
| Allow: / | Lets it fetch every page |
| User-agent: PerplexityBot | Starts the rules for Perplexity search |
| Allow: / | Lets it fetch every page |
If you want to keep training crawlers out, block GPTBot, ClaudeBot and the others by name in their own groups. Don't use a catch-all rule for "AI bots", since that's the template that blocks search crawlers by mistake. On WordPress and Shopify, an SEO plugin or theme setting often writes this file for you, so make the change there or it may get overwritten.
Likely blocked
Look at whatever sits in front of your site. That might be a CDN, your host's firewall, or a WordPress security plugin. Search its settings for bot protection, AI bots or AI crawlers. Many tools have one switch that blocks every AI crawler, training and search alike.
Turn that switch off, or add exceptions for the search crawlers by name. Then open the tool's security log and search for OAI-SearchBot or PerplexityBot. If you see them being refused, you've found the rule.
Run the checker again once you've saved the change. Most CDN changes take effect within a few minutes.
Challenged
Your site shows a challenge page to every new visitor, so our test browser got stopped too. In this case we can't tell whether AI crawlers get through, but the odds aren't good, because crawlers can't solve a challenge page.
Lower the challenge level for verified bots, or turn the challenge on only for pages that need it, like checkout or login. Then check again.
Not reached
Your homepage didn't load for us at all. Check that the site is up, that it loads in under a few seconds, and that it isn't blocking whole countries. Some hosts block visits from data centers by default, and every AI crawler comes from one.
Fix the access issue once, then let RankControl watch it every day and fill your site with pages worth crawling.
Should you block training crawlers?
This is the one choice in this whole topic that's really yours. Blocking GPTBot, ClaudeBot or Google-Extended keeps your pages out of future training sets. It doesn't stop ChatGPT, Claude or Gemini from finding and linking your pages in search, because different crawlers do that job.
There's a slow cost, though. What a model already knows about you comes partly from what it read in training. If you sell a product and want AI tools to know your name, we'd leave training crawlers open. If you publish paid content, run a news site or simply don't want your work in a training set, block them by name and keep the search crawlers open.
Never block the search crawlers to make a point about training. That trade costs you visits now in exchange for very little.
Why one check isn't enough
A clean result today doesn't mean a clean result next month. Sites change all the time, and most changes that block AI crawlers happen by accident.
A security plugin updates and adds a new bot rule. Someone on the team turns on a stricter firewall mode after a spam attack. You move hosts and the new one blocks data center traffic by default. A developer copies a robots.txt template from an old project. None of these come with a warning that ChatGPT can't see you anymore.
The real problem isn't fixing it once. It's knowing when it stops working, ideally before a month of AI search visits goes missing.
RankControl checks every day whether the AI search crawlers can reach your site. When one gets blocked, a banner shows up in your dashboard, and it comes back if the status changes again. Connect the WordPress plugin or Cloudflare and the Analytics screen also shows which AI crawlers visit your pages and how often. You see the visits instead of guessing.
Access is the start, pages are the point
An open door doesn't bring anyone in by itself. AI search engines quote pages that answer the question a buyer just asked, with real numbers and plain words. If your site has a homepage, a pricing page and three old blog posts, there isn't much for a crawler to find.
That's the part RankControl does. It finds the questions buyers in your market search. Then it writes and publishes up to 100 articles a month as native posts on your own site, whether that runs on WordPress, Shopify, Webflow or most other site builders. Auto-publish is off by default, so you can read each one first.
Then it tracks 50 of your buyer questions every week across six AI engines, Google AI Mode included, next to your Google rankings and the visits AI answers send you. Traffic from Google and AI search is the first thing to show up.
Mentions in AI answers take longer, and they mostly come from other sites linking to you and talking about you. Link Control helps with that. It finds sites to pitch for each article you publish and sends the emails from your own inbox after you approve them.
You can do all of this by hand. Check your crawlers every few weeks, write two articles a week, pitch ten sites a month and track your questions in a spreadsheet. That's about 20 to 30 hours a month for a small team. Or RankControl can run it for $400 a month, with a 7-day free trial.
Run the check above, then start a free trial and let RankControl turn open doors into AI and Google traffic.