How to Actually Measure AI Search Visibility (And What the Tools Get Wrong)
I paid for an AI visibility platform for months. Then I cancelled it.
Not because it was bad. Because I was paying a real subscription for a number I could not tie to a single client outcome, and I could reproduce most of what mattered with GA4, Search Console, and a spreadsheet.
If you are trying to figure out whether AI search is working for your business, here is what I actually track now, what the paid tools are genuinely good for, and where every one of them quietly breaks.
What does it mean to measure AI search visibility?
Direct answer: Measuring AI search visibility means tracking two separate things: whether AI engines mention or cite your brand when people ask relevant questions, and whether that mention sends traffic and revenue to your site.
Most businesses only measure the first one. That is the mistake. A citation you cannot connect to a lead is a vanity metric with a better vocabulary.
The two things live in different places:
What you want to know | Where it lives |
Do AI engines mention us? | Prompt testing, manually or via a tool |
Do AI engines send us traffic? | GA4 |
What content are they citing? | GA4 landing pages, server logs |
Does that traffic convert? | GA4 conversions, your CRM |
What do AI visibility tools actually cost?
The category matured fast, and there is now a tool at every budget. As of mid 2026, entry-level monitoring starts around $29 a month with Otterly. Mid-market platforms run roughly $80 to $300. Enterprise platforms like Profound start near $499 and go up from there, with custom quotes into the thousands. Ahrefs and Semrush both sell AI visibility as add-ons to plans you may already have.
For a small business, the honest math is this: serious multi-engine tracking generally starts around $200 to $500 a month. That is a real line item. If you are spending that on measurement, you should be able to say what decision it changed.
For a lot of my clients, it did not change one.
What the tools get wrong
Paid platforms do genuinely useful things. But four problems show up across almost all of them, and nobody selling you a subscription leads with these.
1. The numbers are not stable
AI engines are non-deterministic. Ask the same question twice and you can get different answers with different sources. Tools sample a fixed set of prompts on a fixed schedule and report the result as if it were a rank. It is not a rank. It is a sample of a probability distribution, and a 4% week-over-week move in "share of voice" is often noise.
2. Prompt count is a pricing metric, not a business metric
Most tools price by how many prompts you track. So you end up designing your measurement program around a billing constraint instead of around your actual buyer questions. Teams routinely discover that their entry plan covers a fraction of the queries their customers actually ask.
3. Share of voice does not tell you what to fix
Knowing you appear in 12% of relevant answers is interesting. It does not tell you which page to rewrite, which entity signal is missing, or which competitor's content is getting cited instead of yours. Most tools diagnose. Very few tell you what to do next, and the ones that do cost the most.
4. Nobody is reporting a single global number that means anything
Most platforms report one worldwide figure. If you are a local service business in Chicago, a global share of voice number is close to meaningless. Country and region-level tracking exists on some tools, but usually not at entry pricing.
The free stack I use instead
None of this replaces a paid tool if you are a national brand with a real budget. But for small and mid-sized businesses, this gets you 80% of the signal for nothing.
Layer 1: Turn on the GA4 channel, then don't trust it
On May 13, 2026, GA4 added a native "AI Assistant" channel to the Default Channel Group. AI traffic finally gets separated from ordinary referral traffic with no setup.
That is real progress, and it is incomplete in two ways that matter. Google's channel does not cover every engine, and it lumps everything into one bucket. ChatGPT and Perplexity send traffic that behaves completely differently. You want them split.
Layer 2: Build the custom channel group anyway
Go to Admin, then Data display, then Channel groups. Create a channel called "AI Search" with the condition Session source matches regex:
chatgpt\.com|chat\.openai\.com|perplexity\.ai|gemini\.google\.com|copilot\.microsoft\.com|claude\.ai|deepseek\.com|grok\.com|you\.com
Two things worth knowing. GA4 applies channel groups at query time, so this covers your historical sessions too, not just traffic going forward. And GA4 derives session source from utm_source when one is present, which means ChatGPT's link tagging gets caught by this regex even when the referrer header is missing.
Check your referral source list monthly and add new domains as they appear. This list will look different in six months.
Layer 3: Accept that a large share is invisible
This is the part the dashboards do not show you. A significant portion of AI referral sessions arrive with no referrer header at all and land in Direct. Estimates vary widely, but every serious analysis I have seen puts it high enough that your real AI traffic is meaningfully larger than what GA4 reports.
Practically, this means two things. Watch Direct traffic as a directional signal, especially unexplained growth alongside flat brand search. And stop treating your AI traffic number as precise. It is a floor, not a measurement.
Layer 4: Search Console for AI Overviews, with limits
AI Overview clicks get categorized alongside standard organic search in GA4, and there is no clean way to separate them. Google launched generative AI performance reports in Search Console in June 2026, but they report impressions only, with no click, CTR, or query breakdown yet, and they are still rolling out.
So use it qualitatively. It tells you presence, not performance.
Layer 5: Log prompts manually in a spreadsheet
This is the unglamorous one, and it is the one that has changed the most client decisions.
Write down 15 to 25 questions a real buyer would ask before hiring you. Not keywords. Questions. Run them through ChatGPT, Perplexity, and Gemini once a month. Log four columns:
Were you mentioned, yes or no
Were you cited with a link
Who was cited instead
What did the answer say about your category
That last column is the one nobody tracks and the one that pays. When you read fifteen AI answers about your industry in one sitting, you learn exactly how these systems frame the buying decision, which competitors have become the default reference, and which questions have no good answer anywhere on the internet. That last group is your content roadmap.
Twenty prompts across three engines takes about 45 minutes a month. It costs nothing.
What to measure instead of share of voice
Once the tracking is in place, these are the numbers I actually report to clients:
AI-sourced sessions, by engine. Trend over time, not week to week.
Conversion rate of AI traffic versus organic. AI-referred visitors tend to arrive further along in the decision, and this comparison usually makes the case for the work better than any visibility score.
Landing pages earning AI citations. Tells you what format these engines actually pull from, so you can make more of it.
Citation rate on your buyer questions. From the spreadsheet. Move it from 3 of 20 to 9 of 20 and you have a real result.
Leads and revenue attributed to AI sessions. The only number anyone in the business actually cares about.
When is a paid tool worth it?
I am not anti-tool. I would pay for one again in three situations:
You are a national brand in a competitive category. Manual prompt logging does not scale past a certain query volume, and competitive share of voice starts to mean something when you have real competitors fighting for the same citations.
You need multi-market tracking. If visibility differs by country, doing that by hand is genuinely impractical.
You are an agency tracking many brands. The per-client economics change completely.
If you are a single business under those thresholds, put the $200 a month toward content that answers the questions your buyers are actually asking. That is what gets you cited in the first place.
Start with the question, not the dashboard
The businesses getting real results from AI search are not the ones with the best monitoring stack. They are the ones who figured out which buyer questions matter, published genuinely better answers to them, and built a site machines can read.
Measurement should tell you whether that worked. If your dashboard is not doing that, you are paying for a feeling.
Want to know where you actually stand in AI search? I run a manual AI visibility audit that maps your real buyer questions, shows you what the engines say about your category, and tells you which pages to fix first. Book a free consultation.



Comments