One "AI traffic" number hides the part that matters, because every vendor's crawlers move through your site in their own way. One team worked through three months of logs, eleven million log events in all, and their conclusion was blunt: each AI crawls your site completely differently, and the differences are the findings.
Each AI crawls website completely differently. Here's what 3 months of 11 million event logs actually show.
Here's what we found after 3 months of tracking 11 million real crawler logs across 34 websites. It's quite fun how each AI bots have personalities, like people. GPTBot: Crawls relentlessly, all day every day and barely checks the rules. It...
This is the hands-on half of a pair. The first part covered what agent traffic can and can't tell you, and this one is the audit itself, with commands you can paste. You won't need new tools, only your logs and a terminal. Give the first run ninety minutes. After that, the monthly check takes about as long as a coffee.
Step 1: Find your logs
First, make sure your logs record the user agent. The standard combined format does, but some minimal configs skip it, and if yours is one of them, switch it on today so next month's audit has something to read.
If you run nginx or Apache yourself, the logs are on the server, usually at /var/log/nginx/access.log with its rotated siblings beside it. Pull 14 to 30 days.
Behind a CDN you may barely need raw files, because Cloudflare-class dashboards show verified bot analytics directly and answer half the worksheet on their own. Paid tiers export raw logs if you want more. On a managed platform like Vercel, use the log drain or export feature, since the runtime logs in the console are usually sampled.
While you're in there, settle retention. Rotated logs age out quietly, and the audits worth the most compare this quarter with the last, so decide now how far back you'll want to look.
Step 2: Take the census
Start with who's visiting. In combined-format logs the user agent is the sixth quoted field, so this ranks every kind of visitor by hits:
awk -F'"' '{print $6}' access.log | sort | uniq -c | sort -rn | head -30
That's the whole census, humans and bots together. I'd sort the AI names by job, since each job means something different for you:
| Job | User agents |
|---|---|
| Index builders | OAI-SearchBot, PerplexityBot, plus classic Googlebot and bingbot |
| Live user-fetchers | ChatGPT-User, Perplexity-User, Claude's user agent |
| Training crawlers | GPTBot, ClaudeBot, Meta's crawler, Amazonbot, Applebot, Bytespider, CCBot |
One line counts them all:
grep -iE "gptbot|oai-searchbot|chatgpt-user|claudebot|claude-user|perplexitybot|perplexity-user|amazonbot|applebot|bytespider|ccbot|meta-external" access.log | awk -F'"' '{print $6}' | sort | uniq -c | sort -rn
Write the counts on the worksheet with the date. One census is only a snapshot, but repeat it every month and it becomes the trend line you read every other step against.
Step 3: Verify before you conclude
Scrapers borrow GPTBot's name constantly, hoping your firewall sees family and waves them through, so check a sample before you act on any number that surprises you.
How you check depends on the vendor. OpenAI, Anthropic and Perplexity publish the IP ranges their agents operate from, so take a few source IPs for a claimed bot and see whether they fall inside the list. Google and Bing use reverse DNS instead: resolve the IP to a hostname, confirm it ends in the vendor's domain, then resolve it back.
This lists the claimed-GPTBot source IPs by volume:
grep -i "gptbot" access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head
On Cloudflare, most of this is done for you, since its bot analytics separate verified crawlers from imposters. For me that's the best single reason to run the audit at the edge. Unverified traffic wearing a famous name is a finding of its own, and it usually deserves a rate limit rather than panic.
Know exactly what AI says about your competitors.
RankControl's Recon Agent monitors competitor citations across ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode. See where they show up and you don't.

Step 4: Coverage, the question that matters most
If you cut everything else, keep this step: do the crawlers that feed answer engines reach the pages your revenue depends on? List your ten money pages first (pricing, comparisons, docs, integrations), then pull the path distribution for each bot that matters:
grep -i "oai-searchbot" access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -25
Compare the output with your ten pages. Three results come up a lot, and each has a fix. A money page missing entirely points to internal linking and your sitemap, because crawlers follow prominence. Crawl budget burned on parameter junk and pagination is a robots hygiene job.
The third is the one I'd look for first: a money page the training bots crawl but the index builders never visit. I call it a feed-only page. You're feeding the models without being retrievable for answers, which is backwards from what most sites want.
Freshness is sitting in the timestamps. Note when each index builder last visited each money page, because if you've edited one since, the engines are citing it from memory.
Step 5: Errors, bandwidth and compliance
Three shorter passes finish the audit. First, what do machines actually get served? Status codes per bot:
grep -i "claudebot" access.log | awk '{print $9}' | sort | uniq -c | sort -rn
A healthy profile is mostly 200s. Anything else usually has one of these causes:
| What you see | What it usually means |
|---|---|
| Clusters of 403s | A bot-protection rule someone set during an incident and forgot |
| 429s | Your rate limits are throttling indexers |
| 5xx spikes on bot traffic that humans don't see | An origin struggling under crawl bursts |
Cost comes next. Summing the bytes field gives you bandwidth per bot:
grep -i "bytespider" access.log | awk '{sum+=$10} END {printf "%.2f GB\n", sum/1073741824}'
Run it for the training crawlers. Most sites find the cost trivial. Some don't (the seven-million-hit horror stories are real), and for them blocking or rate-limiting should be a budget decision made on purpose instead of a surprise.
The last pass checks your rules. If robots.txt disallows anything, grep for those bots after the date you set the rule. Documented crawlers comply and the long tail doesn't, and a bot that ignores robots.txt is a firewall problem. Before adding disallows, make sure the block actually does what you think it does, because the most common rule in the wild blocks the wrong bot for the intended goal.
Platform notes, because logs differ
Each of these setups has a wrinkle that costs first-timers an hour.
If Cloudflare or another CDN proxies your traffic, origin logs show the CDN's IPs unless you restore the real client IP from the forwarded header, and cached responses may never reach the origin at all. Use the CDN's own logs or dashboard for bot work. They sit in front of the cache and see everything, verification included.
Vercel and similar hosts sample the logs in their console, which quietly wrecks any counting. Use the log drain or export instead, or front the site with a CDN whose analytics you trust.
WordPress on shared hosting is the fiddliest. Access logs usually exist but hide in the hosting panel, often rotated daily with short retention, so download a couple of weeks before they age out. Watch for security plugins that block bots themselves, too. They make robots.txt look like it works perfectly when the plugin is the one answering.
Make it standing instrumentation
I'd do the first audit by hand, since that's how you learn what normal looks like on your site. Then freeze it into a script that writes the census, the money-page coverage check and the error and bandwidth summaries to a dated file, and let cron run it monthly.
Diff each new file against last month's and read the changes rather than the totals. That's where you'll catch a new crawler turning up or an old one going quiet, along with a key page whose coverage lapsed after a redesign.
Set exactly one alert, on the error rate served to verified index builders. That failure has a real deadline, and everything else can wait a month. I'd skip the live dashboard too. The monthly diff plus that one alert covers every decision this data feeds, and the hours a real-time bot wall would eat belong on the citation side, where the outcomes are.
The worksheet, assembled
It all fits on one dated page with seven rows:
| # | What goes in the row |
|---|---|
| 1 | Which AI bots visited, by verified count |
| 2 | Money-page coverage per index builder, as a simple yes-list |
| 3 | When each money page was last crawled |
| 4 | Error rate served to bots versus humans |
| 5 | Training-crawler bandwidth, in gigabytes and dollars |
| 6 | Which pages the user-fetchers hit, and how often |
| 7 | Your rules versus what the bots actually do |
Row six, the user fetches, deserves the closest read, because each of those hits is a live buyer question reaching for a specific page. Any row that's off has a fix: linking and sitemaps for coverage, firewall rules for imposters, origin work for errors, a budget decision for bandwidth. File each sheet next to last month's and you've got a trend line, the only version of "AI bot analytics" that holds up in front of a skeptical CTO, because the counts are verified and every finding maps to an action.
26 content formats. Published on your domain. Matched to your brand.
Guides, comparisons, listicles, case studies, and more. RankControl generates content that gets cited by ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode.

What the audit can't see
Logs have a hard limit. They show machine attention and never outcomes, so a perfectly crawled page can go uncited forever, and an answer built from an engine's index leaves no line in your logs at all.
Think of the audit as plumbing that clears mechanical blockers between your pages and the engines. Whether an engine then picks you is another instrument's job, the weekly per-engine citation check on the queries that come before your deals.
Run the log audit monthly and the citation check weekly, and treat their disagreements as your best diagnostic. A page that's crawled but never cited needs content and trust work. When one is cited but its last crawl is stale, update it and earn a fresh fetch. If a page is missing from both, start with robots.txt.

Your competitors are getting cited by AI. You're not.
Every day without citation tracking is a day your competitors pull ahead in ChatGPT, Perplexity, and Claude.



