AI Crawler Checker: Find Out If Your Site Blocks ChatGPT, Claude or Perplexity

Your site loads fine in every browser. Google has it indexed. Then a buyer pastes your pricing page into ChatGPT and gets told the page can't be opened, and you have no idea why.

That gap is what this AI crawler checker is for. It knocks on your homepage the way ChatGPT, Claude and Perplexity do, then the way a normal browser does, and shows you which ones got in. It also reads your robots.txt for the training crawlers, so you can see the whole picture in one place. It's free, there's no signup, and it takes about 20 seconds.

Most people find out about a block the slow way. Traffic from AI answers never starts, or it stops, and nobody can say when it changed.

What the AI crawler checker tests

Each AI company sends more than one crawler, and they don't do the same job. Some fetch pages so an AI answer can quote them today. Others collect pages to train future models. Treating them as one thing is an easy mistake, and it can cost you.

The checker sends a real request to your homepage as each of the six search and browsing crawlers below. For the six training crawlers, it reads your robots.txt rules, since blocking those is a choice many sites make on purpose.

CrawlerWho sends itWhat it does
OAI-SearchBotChatGPT searchFinds pages ChatGPT can show and link in its answers
ChatGPT-UserChatGPT browsingOpens a page when someone asks ChatGPT to read it
Claude-SearchBotClaude searchFinds pages Claude can use when it searches the web
Claude-UserClaude browsingOpens a page when someone asks Claude to read it
PerplexityBotPerplexity searchFinds pages Perplexity can cite in its answers
Perplexity-UserPerplexity browsingOpens a page a Perplexity user asked about
GPTBotOpenAI trainingCollects pages that may train future models
ClaudeBotAnthropic trainingCollects pages that may train future models
Google-ExtendedGemini trainingTells Google if Gemini may train on your pages
CCBotCommon CrawlBuilds a public web archive many AI labs train on
Applebot-ExtendedApple AI trainingTells Apple if its models may train on your pages
Meta-ExternalAgentMeta AI trainingCollects pages for Meta's AI models

The first six are the ones that decide whether AI search can see you. If any of them gets turned away, that engine can't read your page when a buyer asks about you. The last six only affect training, which has a slower and much weaker link to whether you show up in answers.

Two places a block can hide

A crawler can be stopped at two points on its way to your page, and many free checkers only look at one of them.

The first point is your robots.txt file. It's a short text file at the root of your site that tells each crawler what it may fetch. Well-behaved AI crawlers read it and obey it. A single line meant for a training bot can catch a search bot too, which is how a lot of sites end up blocking the crawler they actually wanted.

The second point is your firewall, CDN or security plugin. These sit in front of your site and can refuse a request before your server ever sees it. Bot protection tools often have an AI bot setting, and it's easy to switch on without noticing that it covers search crawlers too. When that happens, robots.txt still says "welcome", your server logs show nothing, and the crawler gets a 403.

That's why the checker sends live requests and doesn't stop at robots.txt. It compares what each AI crawler gets back with what a normal browser gets back. When the browser loads the page and the crawler gets refused, something in front of your site is picking on that crawler by name.

Open doors are step one. RankControl fills your site with pages buyers search for, then tracks what AI search does with them.

How to read your results

Every search crawler gets one of five labels, and each one points to a different fix. Training crawlers get a sixth label of their own.

LabelWhat happenedWhat to do
AllowedThe crawler got your page and robots.txt lets it inNothing. This is what you want
Blocked by robots.txtA Disallow rule covers this crawlerEdit robots.txt (see the fixes below)
Likely blockedThe crawler was refused while a browser got throughCheck your firewall, CDN or security plugin
ChallengedEvery visitor, including our browser, got a challenge pageYour bot shield is on for everyone, so we can't tell
Not reachedYour site didn't load for us at allCheck that the site is up and not slow or geo-blocked
Blocked (your choice)robots.txt blocks a training crawlerFine if you meant it. It doesn't affect AI search

We made one choice on purpose when we built the checker. A blocked training crawler never shows as an error, because blocking it is a fair call for a lot of sites. It shows as "Blocked (your choice)" so you can confirm it was your call, and the search results above it stay the ones that matter.

"Likely blocked" is worded with care too. The checker sends each crawler's name from our own servers. Some firewalls also check where a request comes from, so a real crawler can be treated a bit differently from our test. Read the label as a strong hint, then confirm it in your firewall's event log by searching for the crawler's name.

What we found on 88 sites that rank on page one

To see how common this is, we ran the checker on 88 sites in October 2026. They were every site, other than forums and social networks, that ranked in Google's top 10 for 12 everyday buyer searches. The searches covered software, online shops, home services, insurance and legal help, and agencies.

Seventy of the 88 sites let all six AI search crawlers in with no trouble. The other 18, about one in five, turned at least one of them away.

What we foundSites
All six AI search crawlers allowed70
Likely blocked at the firewall, with robots.txt saying yes7
Challenged every visitor, browser included6
Blocked by a robots.txt rule3
Didn't load for us2

The number that stood out to us was seven. On each of those sites, robots.txt allowed every AI search crawler, yet the crawlers got a 403 while our browser loaded the page fine. A checker that only reads robots.txt would have given all seven a clean pass.

Firewall blocks outnumbered robots.txt blocks more than two to one, seven sites to three.

Online shops had it worst. Six of the 16 store sites blocked or challenged AI search crawlers, and half of those were big retailers whose bot shields challenge every visitor. Software sites did better, with four problems out of 29.

One site let Claude's crawlers through but refused the ones from ChatGPT and Perplexity. Blocks aren't always all or nothing, which is why the checker lists every crawler on its own line.

Training crawlers were a quieter story. Only four of the 88 sites blocked any training crawler at all. The other 84 left every one of them alone.

How to fix each kind of block

The fix depends on where the block lives. Start with the label the checker gave each crawler.

Blocked by robots.txt

Open your robots.txt file by adding /robots.txt to your domain in a browser. Look for a Disallow line that names the crawler, or a broad rule under "User-agent: *" that catches it. Then add an allow rule for each search crawler you want in. These lines cover the three big AI search engines.

Line in robots.txtWhat it does
User-agent: OAI-SearchBotStarts the rules for ChatGPT search
Allow: /Lets it fetch every page
User-agent: ChatGPT-UserStarts the rules for ChatGPT browsing
Allow: /Lets it fetch every page
User-agent: Claude-SearchBotStarts the rules for Claude search
Allow: /Lets it fetch every page
User-agent: PerplexityBotStarts the rules for Perplexity search
Allow: /Lets it fetch every page

If you want to keep training crawlers out, block GPTBot, ClaudeBot and the others by name in their own groups. Don't use a catch-all rule for "AI bots", since that's the template that blocks search crawlers by mistake. On WordPress and Shopify, an SEO plugin or theme setting often writes this file for you, so make the change there or it may get overwritten.

Likely blocked

Look at whatever sits in front of your site. That might be a CDN, your host's firewall, or a WordPress security plugin. Search its settings for bot protection, AI bots or AI crawlers. Many tools have one switch that blocks every AI crawler, training and search alike.

Turn that switch off, or add exceptions for the search crawlers by name. Then open the tool's security log and search for OAI-SearchBot or PerplexityBot. If you see them being refused, you've found the rule.

Run the checker again once you've saved the change. Most CDN changes take effect within a few minutes.

Challenged

Your site shows a challenge page to every new visitor, so our test browser got stopped too. In this case we can't tell whether AI crawlers get through, but the odds aren't good, because crawlers can't solve a challenge page.

Lower the challenge level for verified bots, or turn the challenge on only for pages that need it, like checkout or login. Then check again.

Not reached

Your homepage didn't load for us at all. Check that the site is up, that it loads in under a few seconds, and that it isn't blocking whole countries. Some hosts block visits from data centers by default, and every AI crawler comes from one.

Fix the access issue once, then let RankControl watch it every day and fill your site with pages worth crawling.

Should you block training crawlers?

This is the one choice in this whole topic that's really yours. Blocking GPTBot, ClaudeBot or Google-Extended keeps your pages out of future training sets. It doesn't stop ChatGPT, Claude or Gemini from finding and linking your pages in search, because different crawlers do that job.

There's a slow cost, though. What a model already knows about you comes partly from what it read in training. If you sell a product and want AI tools to know your name, we'd leave training crawlers open. If you publish paid content, run a news site or simply don't want your work in a training set, block them by name and keep the search crawlers open.

Never block the search crawlers to make a point about training. That trade costs you visits now in exchange for very little.

Why one check isn't enough

A clean result today doesn't mean a clean result next month. Sites change all the time, and most changes that block AI crawlers happen by accident.

A security plugin updates and adds a new bot rule. Someone on the team turns on a stricter firewall mode after a spam attack. You move hosts and the new one blocks data center traffic by default. A developer copies a robots.txt template from an old project. None of these come with a warning that ChatGPT can't see you anymore.

The real problem isn't fixing it once. It's knowing when it stops working, ideally before a month of AI search visits goes missing.

RankControl checks every day whether the AI search crawlers can reach your site. When one gets blocked, a banner shows up in your dashboard, and it comes back if the status changes again. Connect the WordPress plugin or Cloudflare and the Analytics screen also shows which AI crawlers visit your pages and how often. You see the visits instead of guessing.

Access is the start, pages are the point

An open door doesn't bring anyone in by itself. AI search engines quote pages that answer the question a buyer just asked, with real numbers and plain words. If your site has a homepage, a pricing page and three old blog posts, there isn't much for a crawler to find.

That's the part RankControl does. It finds the questions buyers in your market search. Then it writes and publishes up to 100 articles a month as native posts on your own site, whether that runs on WordPress, Shopify, Webflow or most other site builders. Auto-publish is off by default, so you can read each one first.

Then it tracks 50 of your buyer questions every week across six AI engines, Google AI Mode included, next to your Google rankings and the visits AI answers send you. Traffic from Google and AI search is the first thing to show up.

Mentions in AI answers take longer, and they mostly come from other sites linking to you and talking about you. Link Control helps with that. It finds sites to pitch for each article you publish and sends the emails from your own inbox after you approve them.

You can do all of this by hand. Check your crawlers every few weeks, write two articles a week, pitch ten sites a month and track your questions in a spreadsheet. That's about 20 to 30 hours a month for a small team. Or RankControl can run it for $400 a month, with a 7-day free trial.

Run the check above, then start a free trial and let RankControl turn open doors into AI and Google traffic.

Common questions

Questions, answered

How do I check if ChatGPT can access my website?

Enter your domain in the AI Crawler Checker above. It requests your homepage as ChatGPT's search and browsing crawlers, plus the crawlers from Claude and Perplexity, and compares each answer with what a normal browser gets. You see allowed, blocked or challenged for each one in about 20 seconds.

Does blocking GPTBot remove my site from ChatGPT search?

No. GPTBot only collects pages for training. ChatGPT search uses OAI-SearchBot and ChatGPT-User, so as long as those two can reach your site, ChatGPT can still find and link your pages.

Why do AI crawlers get blocked when my robots.txt allows them?

Something in front of your site, like a CDN, a host firewall or a security plugin, is refusing them before your server sees the request. In our run on 88 sites that rank on page one, seven had exactly this problem. Look for an AI bot or bot protection setting and add exceptions for the search crawlers.

Which AI crawlers should I allow?

Allow every search and browsing crawler, which means two each from ChatGPT, Claude and Perplexity. Training crawlers like GPTBot and ClaudeBot are your choice, and blocking them doesn't affect AI search.

What does Challenged mean in my results?

Your site showed a challenge page to every visitor, including our test browser, so we can't tell whether AI crawlers get through. Crawlers can't solve challenge pages, so the safe move is to relax the challenge for verified bots and check again.

How often should I check my AI crawler access?

After any change to your host, CDN, firewall, security plugin or robots.txt, and at least once a month. Most blocks start by accident during an update. RankControl checks every day and shows a dashboard banner when an AI search crawler gets blocked.

Does the checker test every page on my site?

No. It checks your homepage and reads your robots.txt rules, which covers most blocks because firewall and bot settings usually apply to the whole site. If you think one section is blocked, look in robots.txt for a Disallow rule on that path.

Is the AI Crawler Checker free?

Yes. It's free, needs no signup, and you can run it again any time. Each result gets its own page you can share with whoever runs your site.