Most AI visibility tools give you a score out of a hundred and no way to act on it. The useful version of this exercise is different. You define a fixed panel of buying questions, run it across the assistants your buyers use, and record the whole answer each time. Then you read the citations rather than the score. This is the method we use, written out in full so you can run it yourself.
Search interest in the term is climbing fast. SE Ranking’s UK database put monthly volume for “ai visibility” at 10 in September 2025 and 170 in August 2026. The measurement problem arrived before the measurement discipline did.
What an AI visibility audit actually measures
Three different things get lumped together under one label, and they behave differently.
Mention. An assistant names your brand or your product in the body of an answer. No link required.
Citation. An assistant links to a page on your domain as a source. This can happen without a mention, and a mention can happen without it.
Referral. Someone clicks that link and lands on your site. This is the only one your own analytics can see.
Any credible AI visibility report separates the three. We have seen brands with strong mention rates and almost no citations, which means the model knows them from training rather than from retrieval. That is a fragile position and it needs a different fix from a low mention rate.
The purpose of the audit is not a score. It is a list of specific gaps, each traceable to a page, a feed field or a robots.txt line. Our applied AI work runs on that principle throughout.
Stage 1: Decide what counts as being recommended
Do this before you write a single question. It takes an hour and it prevents a useless audit.
Write down the answer you are trying to win. Not “we want AI visibility”. Something like: “when a UK fleet manager asks which supplier can fit their vehicles, we are one of the three named”. Or: “when a buyer asks for a torque wrench above 1300 Nm, our SKU is one of the products shown”.
That sentence determines everything downstream. It tells you which questions to ask, which engines matter, and which competitor set you are measuring against. Skip it and you end up tracking your own brand name, which nearly every company already wins and nobody searches.
Then name your competitor set explicitly. Five to eight brands. You will score share against them, so guessing later corrupts the baseline.
Stage 2: Build the question panel
This is the whole exercise. A weak panel produces a confident, worthless number.
We use four groups for a services business and five for anyone selling product. Our own live panel runs 29 prompts across three engines, which is a reasonable size. Under 20 prompts and one flaky answer swings the result. Over 50 and nobody reads the output.
Group 1: Brand questions
Six to eight. These test what the assistants already believe about you.
- What is [brand].
- What does [brand] do.
- Is [brand] any good.
- [Brand] reviews.
- How much does [brand] charge.
- Is [brand] good for [your core segment].
Low value on their own. High value as a control. If an assistant gets your service list or your price range wrong here, that error is propagating into every other answer.
Group 2: Category questions
Eight to twelve. The highest commercial value in the panel.
- Best [category] suppliers.
- Who can do [job] for me.
- Companies that do [category].
- Best [category] partners for [segment].
- Best [category] consultants.
These are the questions where an assistant produces a shortlist. Being absent from that shortlist is the finding that gets budget approved.
Group 3: Problem questions
Six to ten. The question a buyer asks before they know your category exists.
- Why are my products not showing up in AI search results.
- How do we get suppliers to give us complete product data.
- How much does it cost to enrich a product catalogue.
- How do we fix a messy catalogue with no internal team.
These matter more than people expect. Problem questions are answered largely from editorial content, so they are the group your published articles can actually move.
Group 4: Decision questions
Four to six. Comparisons and trade-offs.
- Should we buy a PIM or fix our product data first.
- In house versus a partner.
- Agency versus software.
- Platform A versus platform B.
Assistants answer these with structured comparisons. If your position is not documented anywhere public, the assistant will invent one for you.
Group 5: Product questions
Only if you sell the product. Eight to fifteen.
These are constrained buying questions naming a specification rather than a brand. – Cordless impact wrench with at least 1300 Nm breakaway torque. – IP66 rated junction box for outdoor use in the UK.
This group tests your feed and your attribute data, not your content. It usually fails for completely different reasons from the other four.
Stage 3: Pick the engines, the locale and the sample
Engines first. ChatGPT, Google AI Overview, Google AI Mode, Gemini and Perplexity cover most of what a UK or Australian B2B buyer touches. Pick the ones your buyers actually use rather than all of them.
Locale second, and this is where most audits go wrong. We run our UK and Australian panels as separate projects with separate locales. The same question in the same words returns a different brand set in each. If you trade in both markets and measure one, your number is fiction.
Sample third. Assistant answers are not deterministic. Ask the same question twice and you can get a different shortlist. So a single run gives you a binary and a repeated run gives you a rate. Presence rate over several runs is the honest metric. A one-off screenshot is an anecdote.
Stage 4: Capture the whole answer, not just the mention
Most tools store a yes or no. Store three things instead, for every run.
The full answer text. The cited URLs in the order they appeared. The brands named in the order they appeared.
That third item is what makes the audit actionable. In one of our own captures on 20 August 2026, an assistant named five firms and cited twelve URLs. The order of the citations did not match the order of the recommendations. The sources included supplier websites, third-party listicles and a directory. You cannot see any of that from a presence flag.
Capture the answer text verbatim too, because it tells you what the assistant thinks you do. Wrong claims about your service scope or your minimum project size are findings in their own right.
Stage 5: Score it
Five numbers, calculated per engine and then rolled up.
Presence rate. The share of runs where you are named. Calculate it per group, not just overall. Strong brand presence and weak category presence is a completely different problem from the reverse.
Mention position. Where you appear in the list of named brands. First is worth much more than fifth.
Citation share. Your URLs as a share of all cited URLs across the panel. This is the closest thing to a share of voice number and the one that moves when you publish.
Competitor presence. The same two numbers for each brand in the set you named in Stage 1.
Accuracy. The share of answers that describe you correctly. Track it. It is usually the fastest thing to fix.
Resist the urge to combine these into one index. The composite score is what makes commercial AI visibility tools feel useless. Each of the five points at a different remedy.
Stage 6: Read the source layer
This is the part almost nobody does, and it is where the work comes from.
Take every cited URL across the whole panel and classify it. Your own site. A competitor site. A directory or listing site. A third-party listicle. A forum or community. A review site. A video. A marketplace or retailer page.
Then look at the mix. We have seen panels where a third of all citations were third-party listicles that we had never heard of. That finding does not lead to writing another blog post. It leads to getting into those listicles.
We have also seen forums and video dominate the citations for product questions. That is exactly what our own test of a product buying question returned. If that is your pattern, publishing more pages on your own domain will not shift it much.
The source mix tells you which of three jobs you are actually facing. Fix the crawl and the feed. Publish content that answers the question directly. Or get onto the third-party pages the assistants already trust. Most catalogues need the first, and our product content enrichment work usually starts there.
Stage 7: Join AI visibility to your own server data
Panel data tells you what assistants say. Your own logs tell you what they do. Read both.
Search Console. Google launched generative AI performance reporting for site owners in June 2026. It covers AI Overviews and AI Mode, counted within the Web search type in the performance report. It excludes Search Labs experiments. The rollout is gradual, so not every property has it yet. The usual 1,000 row limit applies and recent data is preliminary.
Referral traffic. Assistants that link out create referrals. Segment chatgpt.com, perplexity.ai and gemini.google.com in your analytics. ChatGPT appends a utm_source parameter of its own to outbound links, which makes it easy to isolate.
Bot logs. Check your access logs and your CDN for OAI-SearchBot, GPTBot, ChatGPT-User, PerplexityBot and Google-Extended. Check robots.txt at the same time.
That last one produces more findings than everything else combined. OpenAI states that “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers”. Its guidance is to allow OAI-SearchBot and to confirm your host or CDN allows traffic from its published searchbot IP addresses. We have found bot management rules quietly blocking those crawlers on sites whose owners were paying for AI visibility consultancy. Check this in the first hour, not the fourth week.
What AI visibility measurement cannot tell you
We would rather be straight about the limits than sell a certainty nobody has.
You cannot see impressions. Google now reports them for its own AI surfaces. OpenAI, Anthropic and Perplexity do not publish anything comparable to site owners. Everything outside Google is inferred from panel sampling and referral traffic.
You cannot see personalisation. Answers vary by account history, memory and location. Your panel runs in a clean context that no real buyer has.
You cannot attribute revenue cleanly. A buyer can read a recommendation, remember the name and arrive by direct traffic a fortnight later. That path leaves no trace.
And no assistant publishes its ranking function. We covered what is and is not known about the mechanism in our piece on AI in product data. Treat any confident claim about the weights, including ours, as a hypothesis you test with a panel.
Running it as a repeatable AI visibility audit
Do the first run properly and freeze it. That panel becomes your baseline and you do not change the questions afterwards. Adding or rewording prompts resets your trend line, which is the most common way teams destroy six months of data.
Monthly is the right cadence for most catalogues. Weekly is noise. Quarterly misses the changes.
Re-run the source classification each time, not just the scores. The mix shifts faster than the presence rate, and it shifts first. When forums start displacing review sites for your category, you want to know that in the month it happens.
One person should own the panel. In our engagements that is usually whoever owns product data and content. Most of the remedies land in their backlog, not in a marketing plan. We wrote about the measurement side of that in product content performance.
Key takeaways
- Separate mention, citation and referral. They have different causes and different fixes.
- The question panel is the audit. Four groups for services, five if you sell product, 25 to 40 prompts in total.
- Run the panel per locale. The same question returns a different brand set in the UK and Australia.
- Answers are not deterministic, so measure a presence rate over repeated runs rather than a single yes or no.
- Classify every cited URL by source type. The mix tells you whether the job is crawl, content or third-party placement.
- Check robots.txt and your CDN bot rules before anything else. OpenAI says sites opted out of OAI-SearchBot will not appear in ChatGPT search answers.
- Freeze the panel after the first run. Changing questions destroys your trend.
We run this as a fixed-scope audit against a live catalogue and a named competitor set. You get the panel, the scored findings, the source mix and a remediation sequence in priority order. Most of the fixes turn out to be data problems rather than marketing ones. Thirty minutes on a call is enough to work out whether it is worth doing. Get in touch, or read how we approach applied AI for product catalogues first.