Someone in your market opened ChatGPT this morning and asked which company they should use.
They got three or four names and a sentence about each.
If you were on that list, you will never know. If you were not, you will never know that either.
That is the awkward part of AI search. There is no Search Console for it. Nothing lands in analytics with a source you can filter, no impressions, no click data.
The channel is invisible from the inside, which is why so many brands have a confident opinion about their AI visibility and nothing to back it up.
Across our client accounts, the first thing we do is replace that opinion with a number.
This article walks through the method we use. The first pass takes an afternoon, costs nothing, and produces a baseline you can genuinely re-run in three months.
First, One Piece of Mechanics
Ask ChatGPT the same question twice, and you will get two different answers.
That is not a bug, and it’s the reason most brand checks are worthless.
The model writes each answer token by token, sampling as it goes, rather than retrieving something it stored earlier. On top of that sit memory, custom instructions, chat history, model routing, and location.
Similarweb’s breakdown of the mechanics notes that even at temperature zero, some variability survives because of how requests get batched on the servers.
Two colleagues typing an identical question are, in practice, asking two different questions.
The variation is not even across query types. Objective factual questions wobble a few percent in wording. Subjective ones, which cover nearly every question that names a category and asks for a recommendation, diverge far more.
So a screenshot is not a measurement. It is one draw from a distribution.
Your brand appearing once might be noise. Appearing in seven runs out of ten is a finding.
What you are measuring is not a rank. It is a rate.
How often you appear across a set of prompts, run repeatedly, next to the competitors who appear instead of you.
Step 1: Build the Prompt Set You Will Track
Everything downstream depends on this step, and it is the one people rush.
You are not writing prompts about your brand. You are writing the questions a buyer asks before they have heard of you.
We usually build the set around four groups.
Category discovery. “Best [category] for [use case].” “Top [category] providers in [market].” These decide whether you make the consideration set at all.
Problem-first. How someone describes the situation before they know the product category has a name. “How do I stop [problem] when [context].”
Comparison. “[Competitor] alternatives.” “[Competitor] vs [competitor].” Your competitors’ names do more work here than yours.
Brand-direct. “What is [your brand].” “Is [your brand] any good?” These matter less for acquisition and more for accuracy. This is where you learn what the models believe about you.
Twenty to thirty prompts make a sensible first set.
Write them the way a customer types, not the way a marketer writes a keyword. Full questions, natural phrasing, including the clumsy ones.
One thing worth doing deliberately: write several versions of the same underlying question with different framing.
Ask “best X for Y” and then “is X actually worth it for Y”, and you will get two different sets of names, because the framing of the question is reflected straight back at you.
Test only the flattering phrasing, and you will get a flattering answer.
Step 2: Run Them Properly
Do the first round manually. You should see the raw answers at least once before a dashboard starts summarising them for you.
The setup matters more than the effort.
Log out, or use a temporary chat. Your account holds months of memory and custom instructions telling the model what you care about. Checking your own brand from an account that has discussed your brand two hundred times produces a flattering, useless result.
Run each prompt three times minimum. Five is better, in a fresh session each time. You are sampling a distribution, not looking something up.
Treat the platforms separately. ChatGPT, Google AI Overviews, AI Mode, Gemini, Perplexity, Copilot, and Claude behave differently and cite different sources.
Being visible in one tells you very little about the others. Ahrefs’ monthly citation tables show Reddit leading ChatGPT while YouTube dominates Perplexity, with Claude largely ignoring the community layers the others lean on.
Check by market, not only by language. If you sell in three countries, run the prompts from those locations or name the location in the prompt. Most tools report a single global number, which hides exactly the geography where you are absent.
For every run, record four things:
- whether your brand appeared;
- where in the answer it appeared;
- which competitors appeared;
- and which sources were cited.
That last one is the part people skip. It is also the part that turns a diagnosis into a plan.
Step 3: Score the Mentions
Three runs across twenty-five prompts gives you seventy-five data points.
Four numbers come out of them.
Presence rate. In what share of runs does your brand appear at all? Under 10% and you are effectively invisible. Over 50% and you are an established option in the category.
Share of voice. Of every brand named across every answer, what proportion are you? This is the number to watch over time, because it moves when competitors move.
Sentiment and accuracy. How are you described when you do appear? Correct price? Correct category? A caveat attached to your name that is not attached to anyone else’s?
Wrong facts about a brand are common in AI answers, and they are quietly expensive.
Citation sources. Which URLs do the models cite when answering questions in your category?
Write that list down. It is the most useful output of the whole exercise, and Step 5 is what to do with it.
Step 4: Decide Whether to Buy a Tool
All of the above can be done by hand.
Doing it every month, across seven platforms and three markets, cannot.
The tooling category is barely a year old and raised over $300M between mid-2025 and spring 2026, which tells you both that it is real and that it moves faster than any comparison can stay accurate.
Broadly, three tiers exist:
- entry-level monitoring from around $25–30/month, tracking a small prompt set across the major platforms daily, which is enough for one brand in one market;
- mid-market platforms in the hundreds per month, adding multi-brand tracking, sentiment, competitor comparison, and enough prompt volume for an agency;
- enterprise platforms, custom priced, adding governance, large prompt sets, and deeper source analysis.
Two distinctions are worth understanding before spending anything.
AI search monitoring and LLM monitoring are not the same product. The first tracks what users actually see in AI search products, citations and links included. The second probes what the model holds internally and returns no citation data.
Most buyers need the first and get sold the second.
And check engine coverage against where your buyers actually are. Several well-known SEO suites bolted an AI module onto an existing platform and cover fewer engines than the marketing suggests.
If a meaningful share of your category asks Claude and your tool cannot read it, your dashboard is confidently measuring the wrong thing.
In our own AI visibility work, we run a paid platform alongside Google Search Console and GA4, because no single tool answers every question. Our advice for a first baseline is simpler: start manual, prove the exercise leads to decisions, then buy the cheapest tool that covers your platforms and markets.
Step 5: Read the Citations, Because That Is Where the Work Is
Presence rate tells you where you stand.
The citation list tells you what to do about it.
Pull every source cited across your category prompts and count the domains. The pattern holds across every study of this: community platforms, video, encyclopaedic references, and “best X for Y” roundups dominate, while individual vendor blogs barely register.
In Ahrefs’ Brand Radar tables, YouTube and Reddit sit at the top for AI Overviews, with Wikipedia close behind across every engine. Ahrefs’ own guidance on the Cited Pages report notes that in the SEO category, “best type” listicles dominate the cited formats.
Now hold that list against where your marketing budget actually goes.
If you are buying guest posts on mid-tier blogs that appear nowhere in your citation list, you are buying visibility in a channel your buyer’s assistant does not read.
The plan writes itself from the gap.
Get Into the Roundups You Are Missing From
Take the “best X” articles that get cited in your category and work out how to be in them. Pitch the editor, correct an outdated entry, or pay where a listing is genuinely sold.
It is the highest-leverage single action available here, and it is why we treat brand mentions as a separate workstream from link building rather than a by-product of it.
Show Up Where the Community Is
If Reddit is cited in a third of your category’s answers, an absent brand is an absent brand.
Genuine participation, not astroturfing. The second one breaks platform rules and is easy to spot.
Fix the Facts on Your Own Site
Models quote plain, self-contained statements. Price, what the product does, who it is for, what it does not do.
If the answer to “is [your brand] expensive” is buried in a comparison table behind a JavaScript tab, it will not be retrieved.
Keep Things Fresh
Cited content skews noticeably more recent than the organic top ten, and ranking is no longer the entry ticket it once was.
Ahrefs found the share of AI Overview citations coming from pages in Google’s top 10 fell from around 76% to roughly 38% in a year.
Updating your strongest existing pages is cheap and directly relevant to this channel.
One Thing Not to Spend On
Ahrefs analysed server logs across 137,000 domains and found 97% of published llms.txt files received zero requests in the month studied. Among the few that were fetched, the most frequent readers were SEO audit tools, not AI engines.
SE Ranking’s analysis of roughly 300,000 domains found no relationship between having the file and citation frequency, and Google’s own guidance now states plainly that machine-readable files like this are not needed to appear in generative AI features.
Ship one if you have real documentation and a spare half hour. Do not ship one instead of fixing your page structure.
Step 6: Set the Baseline, Then Track It
Write down the four numbers, the date, the prompt set, and the exact conditions you ran under.
That is your baseline.
Re-run monthly while you are actively working on it, quarterly if you are not. Same prompts, same platforms, same conditions.
Changing the prompt set between runs is the quickest way to produce a chart that means nothing.
Expect noise. A move from 12% to 15% presence across seventy-five samples is not a result. A move from 12% to 34%, holding across two consecutive measurements, probably is.
Watch the trend, not the month.
What Good Looks Like
There is no benchmark to hit, because share of voice is relative to a category and every category is different.
But three things separate the brands doing this well from the ones performing it.
They measure a rate across many prompts instead of screenshotting one good answer.
They pass the citation source list to whoever buys content, links, and PR, so budget moves toward the pages that actually get retrieved.
And they treat wrong facts about their brand as urgent, because an incorrect price repeated across an entire channel does more damage than an absent mention.
The Honest Caveat
All of this measures a moving target with imperfect instruments.
Models change without notice. Answer variability means every number carries a margin of error that almost nobody in this industry quotes honestly.
And the correlation research underpinning the advice, of which Ahrefs’ study of 75,000 brands is the most-cited example, is genuinely useful and genuinely correlational.
Brands with heavy web presence get cited more. They are also usually the brands people already know. Nobody has untangled that.
Treat this as a directional instrument.
It will reliably tell you whether you are absent or present, whether you are gaining or losing against named competitors, and which sources the models lean on in your category.
That is enough to make budget decisions. It is not enough to promise anyone a position, and anyone selling a guaranteed AI ranking is selling something that does not exist.
Run the baseline.
It costs an afternoon, and it is the difference between having an opinion about AI search and having a number.



















