How To Audit AI Crawlers In Server Logs

The hands-on version: where the logs live, the exact pipelines to run, how to verify bots against published IP ranges, and the seven-question worksheet.

RankControl9 min read
How To Audit AI Crawlers In Server Logs

One "AI traffic" number hides the part that matters, because every vendor's crawlers move through your site in their own way. One team worked through three months of logs, eleven million log events in all, and their conclusion was blunt: each AI crawls your site completely differently, and the differences are the findings.

r/SaaS· u/UptownOnion· Jun 11, 2026

Each AI crawls website completely differently. Here's what 3 months of 11 million event logs actually show.

Here's what we found after 3 months of tracking 11 million real crawler logs across 34 websites. It's quite fun how each AI bots have personalities, like people. GPTBot: Crawls relentlessly, all day every day and barely checks the rules. It...

↑ 8 upvotes9 comments
Via Reddit

This is the hands-on half of a pair. The first part covered what agent traffic can and can't tell you, and this one is the audit itself, with commands you can paste. You won't need new tools, only your logs and a terminal. Give the first run ninety minutes. After that, the monthly check takes about as long as a coffee.

Step 1: Find your logs

First, make sure your logs record the user agent. The standard combined format does, but some minimal configs skip it, and if yours is one of them, switch it on today so next month's audit has something to read.

If you run nginx or Apache yourself, the logs are on the server, usually at /var/log/nginx/access.log with its rotated siblings beside it. Pull 14 to 30 days.

Behind a CDN you may barely need raw files, because Cloudflare-class dashboards show verified bot analytics directly and answer half the worksheet on their own. Paid tiers export raw logs if you want more. On a managed platform like Vercel, use the log drain or export feature, since the runtime logs in the console are usually sampled.

While you're in there, settle retention. Rotated logs age out quietly, and the audits worth the most compare this quarter with the last, so decide now how far back you'll want to look.

Step 2: Take the census

Start with who's visiting. In combined-format logs the user agent is the sixth quoted field, so this ranks every kind of visitor by hits:

awk -F'"' '{print $6}' access.log | sort | uniq -c | sort -rn | head -30

That's the whole census, humans and bots together. I'd sort the AI names by job, since each job means something different for you:

JobUser agents
Index buildersOAI-SearchBot, PerplexityBot, plus classic Googlebot and bingbot
Live user-fetchersChatGPT-User, Perplexity-User, Claude's user agent
Training crawlersGPTBot, ClaudeBot, Meta's crawler, Amazonbot, Applebot, Bytespider, CCBot

One line counts them all:

grep -iE "gptbot|oai-searchbot|chatgpt-user|claudebot|claude-user|perplexitybot|perplexity-user|amazonbot|applebot|bytespider|ccbot|meta-external" access.log | awk -F'"' '{print $6}' | sort | uniq -c | sort -rn

Write the counts on the worksheet with the date. One census is only a snapshot, but repeat it every month and it becomes the trend line you read every other step against.

Step 3: Verify before you conclude

Scrapers borrow GPTBot's name constantly, hoping your firewall sees family and waves them through, so check a sample before you act on any number that surprises you.

How you check depends on the vendor. OpenAI, Anthropic and Perplexity publish the IP ranges their agents operate from, so take a few source IPs for a claimed bot and see whether they fall inside the list. Google and Bing use reverse DNS instead: resolve the IP to a hostname, confirm it ends in the vendor's domain, then resolve it back.

This lists the claimed-GPTBot source IPs by volume:

grep -i "gptbot" access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head

On Cloudflare, most of this is done for you, since its bot analytics separate verified crawlers from imposters. For me that's the best single reason to run the audit at the edge. Unverified traffic wearing a famous name is a finding of its own, and it usually deserves a rate limit rather than panic.

RANKCONTROL

Know exactly what AI says about your competitors.

RankControl's Recon Agent monitors competitor citations across ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode. See where they show up and you don't.

Step 4: Coverage, the question that matters most

If you cut everything else, keep this step: do the crawlers that feed answer engines reach the pages your revenue depends on? List your ten money pages first (pricing, comparisons, docs, integrations), then pull the path distribution for each bot that matters:

grep -i "oai-searchbot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -25

Compare the output with your ten pages. Three results come up a lot, and each has a fix. A money page missing entirely points to internal linking and your sitemap, because crawlers follow prominence. Crawl budget burned on parameter junk and pagination is a robots hygiene job.

The third is the one I'd look for first: a money page the training bots crawl but the index builders never visit. I call it a feed-only page. You're feeding the models without being retrievable for answers, which is backwards from what most sites want.

Freshness is sitting in the timestamps. Note when each index builder last visited each money page, because if you've edited one since, the engines are citing it from memory.

Step 5: Errors, bandwidth and compliance

Three shorter passes finish the audit. First, what do machines actually get served? Status codes per bot:

grep -i "claudebot" access.log | awk '{print $9}' | sort | uniq -c | sort -rn

A healthy profile is mostly 200s. Anything else usually has one of these causes:

What you seeWhat it usually means
Clusters of 403sA bot-protection rule someone set during an incident and forgot
429sYour rate limits are throttling indexers
5xx spikes on bot traffic that humans don't seeAn origin struggling under crawl bursts

Cost comes next. Summing the bytes field gives you bandwidth per bot:

grep -i "bytespider" access.log | awk '{sum+=$10} END {printf "%.2f GB\n", sum/1073741824}'

Run it for the training crawlers. Most sites find the cost trivial. Some don't (the seven-million-hit horror stories are real), and for them blocking or rate-limiting should be a budget decision made on purpose instead of a surprise.

The last pass checks your rules. If robots.txt disallows anything, grep for those bots after the date you set the rule. Documented crawlers comply and the long tail doesn't, and a bot that ignores robots.txt is a firewall problem. Before adding disallows, make sure the block actually does what you think it does, because the most common rule in the wild blocks the wrong bot for the intended goal.

Platform notes, because logs differ

Each of these setups has a wrinkle that costs first-timers an hour.

If Cloudflare or another CDN proxies your traffic, origin logs show the CDN's IPs unless you restore the real client IP from the forwarded header, and cached responses may never reach the origin at all. Use the CDN's own logs or dashboard for bot work. They sit in front of the cache and see everything, verification included.

Vercel and similar hosts sample the logs in their console, which quietly wrecks any counting. Use the log drain or export instead, or front the site with a CDN whose analytics you trust.

WordPress on shared hosting is the fiddliest. Access logs usually exist but hide in the hosting panel, often rotated daily with short retention, so download a couple of weeks before they age out. Watch for security plugins that block bots themselves, too. They make robots.txt look like it works perfectly when the plugin is the one answering.

Make it standing instrumentation

I'd do the first audit by hand, since that's how you learn what normal looks like on your site. Then freeze it into a script that writes the census, the money-page coverage check and the error and bandwidth summaries to a dated file, and let cron run it monthly.

Diff each new file against last month's and read the changes rather than the totals. That's where you'll catch a new crawler turning up or an old one going quiet, along with a key page whose coverage lapsed after a redesign.

Set exactly one alert, on the error rate served to verified index builders. That failure has a real deadline, and everything else can wait a month. I'd skip the live dashboard too. The monthly diff plus that one alert covers every decision this data feeds, and the hours a real-time bot wall would eat belong on the citation side, where the outcomes are.

The worksheet, assembled

It all fits on one dated page with seven rows:

#What goes in the row
1Which AI bots visited, by verified count
2Money-page coverage per index builder, as a simple yes-list
3When each money page was last crawled
4Error rate served to bots versus humans
5Training-crawler bandwidth, in gigabytes and dollars
6Which pages the user-fetchers hit, and how often
7Your rules versus what the bots actually do

Row six, the user fetches, deserves the closest read, because each of those hits is a live buyer question reaching for a specific page. Any row that's off has a fix: linking and sitemaps for coverage, firewall rules for imposters, origin work for errors, a budget decision for bandwidth. File each sheet next to last month's and you've got a trend line, the only version of "AI bot analytics" that holds up in front of a skeptical CTO, because the counts are verified and every finding maps to an action.

RANKCONTROL

26 content formats. Published on your domain. Matched to your brand.

Guides, comparisons, listicles, case studies, and more. RankControl generates content that gets cited by ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode.

What the audit can't see

Logs have a hard limit. They show machine attention and never outcomes, so a perfectly crawled page can go uncited forever, and an answer built from an engine's index leaves no line in your logs at all.

Think of the audit as plumbing that clears mechanical blockers between your pages and the engines. Whether an engine then picks you is another instrument's job, the weekly per-engine citation check on the queries that come before your deals.

Run the log audit monthly and the citation check weekly, and treat their disagreements as your best diagnostic. A page that's crawled but never cited needs content and trust work. When one is cited but its last crawl is stale, update it and earn a fresh fetch. If a page is missing from both, start with robots.txt.

Your competitors are getting cited by AI. You're not.

Every day without citation tracking is a day your competitors pull ahead in ChatGPT, Perplexity, and Claude.

Show me who's getting cited→2-minute overview · real case-study numbers

Frequently Asked Questions

Sort them by vendor and by job. OpenAI alone runs three: GPTBot for training, OAI-SearchBot for its search index and ChatGPT-User for live fetches when someone asks a question. Anthropic has ClaudeBot plus its user and search agents, Perplexity has PerplexityBot and Perplexity-User, and Googlebot and Bingbot cover the index layer. For the cost picture the big training crawlers count too, Meta's among them, along with Amazonbot, Applebot, Bytespider and CCBot, and each vendor documents its exact strings.

Check where the hits come from, because scrapers borrow popular bot names hoping for a free pass. OpenAI, Anthropic and Perplexity publish the IP ranges their crawlers use, so compare a sample of hits against those lists, while Googlebot and Bingbot get verified with a reverse DNS lookup instead. If you're behind Cloudflare or a similar CDN, it does the verification for you and labels the verified bots in its analytics.

Seven things, and the one I'd never skip is whether the index-building crawlers reach your money pages. Start with which bots visit at all, then look at that coverage and how fresh the last visits are. After that come the errors machines receive and the bandwidth training crawlers use, then which pages the live user-fetch agents hit and whether your robots.txt decisions are respected. Every one of them maps to a concrete fix.

Go deep once, then do it monthly on a one-page summary, with a single alert on error spikes. Bots change behavior when vendors change their infrastructure, and that doesn't happen daily, so a real-time dashboard is usually more instrumentation than you need. Comparing each month with the one before catches everything you'd act on.

Partially. A CDN dashboard like Cloudflare's gives you verified per-bot analytics with zero setup, and that covers which bots come and how much they take. It's weaker on per-path coverage of your own money pages, which is where raw logs, or your platform's log export, are still worth the effort.

RANKCONTROL

Turn AI search into a customer acquisition channel

Content that ranks on Google and gets cited by AI search engines. Published on your domain. Citations tracked weekly.

Related Articles

THE SIGNAL

Insights on AI and Google search strategy. No fluff.

Get the latest on AI citations, Google rankings, and content strategy.

No spam. Unsubscribe anytime.