AI search doesn't show up in any one place you can look. The citations sit inside chat answers and the impressions in Google's beta reports, while your referral traffic hides in GA4's Direct bucket and the crawlers reading your site only appear in logs nobody opens. There is a real measurement layer for AI search now, but it's spread across half a dozen surfaces that don't talk to each other.
That fragmentation has a price, and Otterly.ai published a stark example of it this month. Claude drives 10.6% of their signups, while Google Analytics credits it with 0.1%, a hundred-to-one attribution gap at a company whose entire business is monitoring AI search. If they're flying that blind on one layer, what is the average SaaS dashboard missing?
The volume being mismeasured isn't a rounding error anymore. Backlinko reports that LLM-driven traffic to its site grew 800% year over year, and every engine keeps shipping new surfaces to measure. Positions won't hold still for you, either. A quarterly audit can only tell you where you were, because citation positions drift 40-60% month over month across platforms (Profound's tracking).
What follows is the stack that closes those gaps: seven tools, one per measurement layer, chosen to complement each other instead of duplicating. Prices were verified this week.
Before the list: the three jobs test
The most useful framework for judging any tool in this category, ours included, came out of a practitioner thread about which AEO tools are worth paying for. It splits the work into three jobs, which are tracking visibility, diagnosing why, and driving action. Most tools blur them together, and budgets tend to die in that blur, so the thread is worth reading before you buy anything:
Anyone here actually paying for GEO/AEO tools?
If yes, which one and is it worth it? I've been spending a lot of time thinking about AI search lately, not from a "how do I generate content with AI" angle, but from a traffic acquisition and attribution angle. The reason is simple:if AI r...
Its sharpest line, lightly paraphrased, is that a score you can't reproduce is selling opacity, not measurement. I think that skepticism is earned. Answers vary between runs, most tools run the same prompt-sampling methodology under different branding, and nobody should pay enterprise money for a number without the why underneath it.
One commenter made the skeptics' case with an analogy that's hard to forget: measuring LLM citations with prompt sampling is like wanting to know how often your friends mention you at lunch, so instead of listening at lunch, you write ten questions you think might come up and ask a magic 8-ball each one, ten times. It's harsh, and not entirely wrong. Sampling is an estimate, every vendor's estimate uses different prompts and a different cadence, and anyone claiming to tie visibility scores directly to pipeline is, in the words of another practitioner who sat through six sales pitches in a week, most likely guessing:
Sat through 6 "AI search optimization" pitches this month. They all sell a "visibility score." Nobody can explain how it's calculated. What's actually the real methodology?
Hopefully, this saves someone else some hours. Ran a small RFP last month for AI search visibility services. Mid-market B2B, trying to figure out whether we show up in ChatGPT / Perplexity / Gemini when someone asks about our category. Six...
The answer to that critique is better measurement: a frozen, representative prompt set, repeated sampling so the trends mean something, and first-party data wherever a platform offers it. The stack below was put together with exactly that skepticism in mind, using free first-party sources for the facts and paid tools only where they add diagnosis or action.
1. RankControl: the citation layer, plus the fix
This one is our product, so I'll say that first. We built it because the three-jobs problem is real, and a tracker that stops at tracking leaves you with a dashboard you feel guilty about.

RankControl checks your tracked queries every week in ChatGPT, Perplexity, Claude and Gemini, plus Grok and Google AI Mode, and it logs citations with the source-level detail that answers "why them, not us." The same platform then acts on that diagnosis. The content engine generates and publishes the pages the gaps call for, and AI traffic attribution tags the visitors who arrive, so measurement and fix run as one loop.
It costs $400/mo, everything included, which makes it the most expensive item on this list, because it's the only one doing the third job. If all you want is the tracking layer, the tools further down start at $99, and our ranked comparison of citation trackers puts those head-to-head.
2. Google Search Console: your AI Overviews ledger
The biggest measurement news of the year arrived quietly in June, when Google shipped a dedicated Generative AI performance report in Search Console. It's still rolling out in beta.

You get impressions and pages for your appearances in AI Overviews and AI Mode, as free first-party data straight from the source. That makes it the canonical record for one narrow question, which is how often Google's AI features surface your pages. The bigger AEO picture isn't in it: no unlinked brand mentions (most of them), no other engines, and nothing on why a competitor got the citation. I'd use it as the ledger and go elsewhere for strategy.
Setup takes nothing. If GSC is verified, the report appears under Performance as it rolls out to your property.
3. GA4 with a custom AI channel group: the referral layer
Out of the box, GA4 is where AI traffic goes to get mislabeled. The fix is a custom channel group that matches AI sources, along with realistic expectations about what it can't see. Google's own channel groups documentation now includes an AI assistants example, which tells you how mainstream this setup has become.

| Detail | |
|---|---|
| What to build | Custom channel group with source regex: chatgpt.com, perplexity.ai, claude.ai, gemini.google.com, copilot.microsoft.com, grok.x.com |
| What it shows | Sessions, conversions, landing pages, and engagement for AI-referred visits that carry a referrer |
| The catch | About 70.6% of AI referral traffic arrives with no referrer header and lands in Direct, per Loamly's attribution research |
| Price | Free |
GA4 has also begun rolling out a native AI Assistant channel. Practitioners report it landing unevenly, catching some assistant referrals and missing plenty, so I'd treat it as a bonus and keep the custom channel group as the workhorse, reading both numbers as floors rather than totals.
So if two-thirds of the channel is invisible, why bother? The visible third is your calibration set. Conversion rates on attributed AI traffic are consistently multiples of organic, as we covered in our AI referral stats roundup, and that tells you what the invisible portion is probably worth. It also means a swelling Direct bucket after a citation win stops being a mystery.
We'll show you exactly where your brand stands in AI search.
No commitment. $0 due today, cancel anytime. See how ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode talk about your brand today.

4. Bing Webmaster Tools: the Microsoft surface
Most teams skip this one, and it earns its slot twice over. Bing's index is what ChatGPT browses, and since February, Bing Webmaster Tools has shipped an AI Performance report in public preview that shows when your site is cited across Copilot, Bing's AI summaries and partner integrations, along with the URLs being referenced.

To be precise about the boundary, it reports Microsoft's AI surfaces and doesn't show ChatGPT's answers directly. The same submission and indexing hygiene serves both, though, since ChatGPT's live retrieval reads Bing's index. Verify the domain, submit the sitemap and check the AI Performance report monthly. It's free and takes fifteen minutes, and almost none of your competitors have done it.
5. Cloudflare Attribution Business Insights: the crawler layer
Everything above measures outputs, and this layer measures inputs: which AI bots read your site, what they take and what they send back. Cloudflare launched Attribution Business Insights on July 1. It classifies AI bot traffic by purpose (training, search, agent) and reports crawl-to-referral ratios per operator.

The crawl-to-referral ratio is the metric I'd watch. It tells you whether a given AI company is a trading partner or a strip miner, and whether your llms.txt and robots.txt decisions are working. Fair warning though: all of this sits behind Cloudflare's paid Bot Management tier, so check current pricing against your plan.
Teams not on Cloudflare have two free approximations. One is server-log analysis, and the other is Microsoft Clarity's bot analytics, which practitioners have been leaning on lately since it began surfacing non-human traffic, and even robots.txt violations, at no cost. Both are cruder views, but any answer to "what did the machines read this week" beats none.
6. Profound: brand mentions at competitive depth
For dedicated share-of-voice monitoring beyond your own tracked queries, Profound covers nine AI platforms and layers agent analytics on top. Besides ChatGPT, Perplexity and Gemini, it watches Copilot and Google AI Overviews, and it also reaches Meta AI, Grok, DeepSeek and Claude.

The $99/mo Starter plan covers ChatGPT only, so three platforms means Growth at $399/mo, and enterprise is custom, with SOC2 and SAML. That platform count matters more than it looks. Backlinko's tool analysis found 91% of cited URLs appear in just one LLM, which means a single-platform tracker sees only a sliver of your actual footprint.
The engines don't even source alike. Practitioners tracking citations across models report Perplexity leaning heavily on forums while Claude favors documentation pages; that's unverified as a study, but consistent with what our own sampling shows. Cross-model variance is the norm, and it's the whole argument for monitoring more than one engine. If Profound isn't the fit, two alternatives are worth a look: Ahrefs Brand Radar if you want social channels in the same view (pricier, from ~€358/mo), or Otterly.ai's $29 Lite tier as a budget entry.
Know exactly what AI says about your competitors.
RankControl's Recon Agent monitors competitor citations across ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode. See where they show up and you don't.

7. Semrush AI Visibility Toolkit: the prompt research layer
Every other tool here assumes you already know which prompts to track, and this last layer is where you find out. Sample the wrong twenty prompts and every score upstream turns into noise. Semrush's AI toolkit pairs prompt tracking with its search volume data, and the free AI Search Visibility Checker is a genuinely useful zero-cost starting point, since it surfaces the prompts already driving mentions for a domain.

Paid plans begin at $165/mo (annual) and include 50 daily tracked prompts and competitor gap analysis. Build your prompt library the way practitioners keep recommending, as a frozen set of 20-30 prompts that mirror how buyers actually research your category, changed rarely so the trend lines mean something. For sourcing those prompts from real buyer language, see our prompt research playbook.
What not to buy
Every dollar in this stack should buy either a fact you can verify or an action you'd otherwise do by hand, and the practitioner trenches have turned up three kinds of tool that fail that test. The first is anything selling "LLM search volume" as a hard number. The platforms don't publish prompt volumes, so every such figure is modeled, and vendors rarely say how.
Agencies and multi-product teams should also skip per-prompt or per-brand billing, because a pricing model that feels cheap at one brand compounds brutally at ten. And I'd pass on any tool that outputs a proprietary score without showing the underlying answers and cited URLs, since you can't audit that number or act on it.
Wiring it together
There are seven tools, but the workflow is one loop, run weekly:
- The prompt library (the Semrush layer) defines what you measure.
- Citation tracking (RankControl) samples those prompts across engines and flags movement, with share of voice as the headline metric.
- GSC and Bing WMT confirm the impression side on Google and Microsoft surfaces.
- GA4's AI channel group catches the referred visitors, and the Direct-bucket delta hints at the unattributed rest.
- Crawler analytics verify that the machines can still read you after every deploy.
A lean version costs $0 on the free tiers of layers 2 through 5, plus one paid monitoring tool. By hand, though, the time adds up. The free layers take a monthly hour each, but the citation sampling in step 2 is the recurring sink at 3-4 hours weekly when done manually, which makes it the layer worth automating first.
What the stack really gives you is the ability to answer, in one meeting, the question fragmented tooling can't: are we more or less visible to AI buyers than last month, why, and what are we doing about it?
26 content formats. Published on your domain. Matched to your brand.
Guides, comparisons, listicles, case studies, and more. RankControl generates content that gets cited by ChatGPT, Perplexity, Claude, Gemini, Grok, and Google AI Mode.

Start with the free floor
If this list feels overwhelming, sequence it. In week one, verify GSC and Bing WMT, build the GA4 channel group, run the free Semrush checker to draft your prompt library, and turn on whatever bot analytics your host already offers, which gets you the whole measurement floor for $0. In week two, add the paid layer that matches your bottleneck: monitoring if you're blind, action if you're stuck. Our list of AEO and GEO tools you can try covers every free checker and trial length.
You can operate all seven surfaces yourself and reconcile them in a spreadsheet every Friday. Or RankControl can run the citation layer every week, publish the content the data calls for and hand you the reconciled picture, while you attend exactly zero of the meetings where someone asks what the visibility score means.



