Oct 2026·9 min read

3 Checks That Tell You Whether Your AI Visibility Score Really Changed

A B2B brand went from 53% to 86% share of voice in 6 weeks with nothing changed on its website. Run the 3 checks in this article against your own last reported number and you will know within an hour whether the move was real or whether your tracker simply asked different questions.

Will Leatherman

Will Leatherman

Founder, Catalyst

TLDR

An AI visibility score is one answer out of many, so it moves even when nothing about your company does. Before you report a change, run 3 checks. Same questions and same platforms as the last reading, the move holds across 2 weeks of daily runs, and the gap is bigger than the margin of error your question count allows. A 33 point jump in a real audit failed the first check, because the two scans used different question sets. Passing one check never rescues a failed one.

On September 8th a company audited 6 weeks earlier got rescanned. The first scan put BambooHR at about 53% share of voice. The second put it at 86%. That is a 33 point increase in 6 weeks, and part of it is real, and part of it is the questions changing.

An AI visibility score is the share of buyer questions where an AI assistant mentions your brand, and it exists so a marketing team can tell whether buyers hear their name when they ask ChatGPT or Gemini for a recommendation. Most people treat that score the way they treat a bank balance. It is closer to a poll. Tracking companies have now published what happens when you ask the same question again, and the top brand changes in almost half of all reruns.

"The number moves by design. Knowing what to do with that is the skill." — Will Leatherman, Founder of Catalyst

Hold on to the 33 points. By the end of this article you will know whether it counted, and the answer is the one most people expect.

Why does your AI visibility score move when nothing on your site changed?

Because every reading is one answer out of many, and the LLM runs a fresh search underneath most of them.

SparkToro asked AI the same buying questions up to 100 times each. The same list of brands came back twice in fewer than 1 in 100 answers. Every screenshot in your deck is one draw from that distribution.

Part of the reason sits upstream of the text. Ask ChatGPT a question and it runs its own background searches before it writes anything. A repeat test of identical questions found ChatGPT reused those background searches only 11% of the time. Perplexity reused them 92.8% of the time. On ChatGPT you are close to a fresh search on every ask.

Turning the randomness off does not fix it either. Thinking Machines ran the same prompt a thousand times with sampling made deterministic and got 80 different answers. The wording moves even when the substance holds. Ahrefs measured Google's AI Overviews and found the meaning stayed stable between versions while the words kept changing, and a separate weekly tracker found 87% of prompts whose sources never moved at all still produced different text week to week.

"Most people treat their score like a bank balance." — Will Leatherman, Founder of Catalyst

A different answer is not a different verdict. One answer is noise, and the pattern across many answers is the signal.

What are the 3 counts hiding inside a single visibility score?

Three different things get blended into one number, and they move at different speeds.

CountWhat it meansHow fast it moves
NamedThe answer mentions youAbout 25% week to week in BrightEdge data
CitedYour page is one of the sources under the answer39% to 41% week to week
RecommendedThe answer tells the buyer to pick youThe least stable of the three

No single score blends these consistently, and the tools on the market each build theirs differently. SparkToro watched one brand show up in 69% to 71% of answers and come first in only 25% of them. Being named survives. Ranking wanders.

Ahrefs looked at pairs of answers where every cited page had changed and found the same top brand still came first 47.3% of the time. The sources underneath churn completely while the name on top holds.

BambooHR shows all 3 counts at once. The 86% from the opening is only its named count. It was cited from its own site twice and cited from other people's sites 465 times, and 74% of those citations came from rival pages. It was picked in 0 of 60 searches.

"Bamboo's actual competitors are writing about Bamboo more than Bamboo is getting quoted from its own site." — Will Leatherman, Founder of Catalyst

One strong count, 2 weak ones, and a blended score that hid both. If your tool reports a single number, ask which of the 3 it is really counting. The gap between being named and being picked is its own problem with its own fix, covered in why AI search names you but never recommends you.

Which count is worth watching, and why is it the answer itself?

Named and cited, because the clicks that used to pay for ranking are not arriving.

Ahrefs found that when an AI Overview appears, the number 1 organic result loses 58% of its clicks. Pew found people clicked a regular result on 8% of visits where an AI summary appeared, against 15% of visits where it did not.

"The value is moving into the answer itself, into being named and being cited." — Will Leatherman, Founder of Catalyst

AI Mode is not where that value sits yet. Two separate panels put it under half a percent of all search activity, and on consumer searches 44.9% of its citations point back to Google itself. ChatGPT is the one to watch. In June 2025, 1.6% of ChatGPT answers carried a citation. By May 2026 it was 6.8%.

Pull up the last AI visibility number you reported. Write today's date next to it, then how many questions it came from and which platforms. Split it into 3 columns for named, cited and recommended. If your tool cannot fill one in, leave it blank. That blank is the information.

Which 3 checks separate a real change from a measurement change?

Run them in order, and a failure on any one of them stops the report.

CheckThe questionThe trap it catches
1. Same inputsSame question list, same platforms as last timeYour tracker changed its prompts or its platform mix and the chart said nothing
2. Holds across runsDoes the move survive 2 weeks of daily runsOne good day reported as a trend
3. Bigger than the noiseIs the gap larger than your question count can resolveA 10 point move on 12 questions reported as decisive

The question list decides the answer more than almost anything else. Across platforms, only 14% to 29% of recommended brands overlap. The ChatGPT app and its API share 24% of the brands they name. Google's own two surfaces usually agree on what the answer says and share only 13.7% of the pages they cite. Profound, which sells daily tracking, found that which prompts sit in the set moves citation share about 10 times more than day to day drift.

Check 2 needs a floor. Researchers at the University of St. Gallen found sources cited overlapped only 34% to 42% from one day to the next, so they ran 7 runs per prompt per day over 2 to 4 weeks. Ahrefs rechecked the same AI Overviews for a month and read a different result on 70% of consecutive checks, with 45.5% of cited pages new each time.

Check 3 is arithmetic. On 12 questions asked once each, two readings have to differ by 25 to 36 points before the gap means anything. On 200 queries, gaps under 5 to 7 points are usually noise. Evertune ran a single ChatGPT prompt repeatedly and found a brand sitting around 10% visibility carried a 27 point margin of error at 5 runs, falling to about 12 points at 12 runs. Most of the numbers people screenshot sit inside their own margin of error.

Normal churn is larger than people expect. Week over week, ChatGPT replaces about 74% of the domains it cites and AI Mode replaces 56%. A cited source loses half its citations in about 4.5 weeks on Google and 3.4 weeks on ChatGPT. Three AI Mode runs on the same day matched on 9.2% of the pages they cited.

"One run is just a rumor." — Will Leatherman, Founder of Catalyst

Did the 33 point jump survive the checks?

No, and it failed on the first one.

BambooHR's two scans used different question sets, so the comparison was never valid. The 33 points cleared the bar on check 3 and that rescues nothing.

"Passing one check never rescues a failed one." — Will Leatherman, Founder of Catalyst

A daily monitor on the same brand across 3 platforms shows how the trap works. Answers per run fell from 35 to 26 while share barely moved, from 80% to 76.9%. Google's AI Overviews answered only 2 of the 12 questions on those runs, against 8 to 12 before, and nothing on the site had changed. One platform simply answered less often and the mention count fell with it.

Across 23 daily runs on 3 platforms the same brand ranged from 64.5% to 80.8% with no site changes at all. On September 30th it dropped to 55.2%. On the morning of this workshop it was back at 76.9%. Report the 30th and you report a crash that never happened. The weekly version of this discipline is laid out in how to monitor your AI search ranking every week.

What changed across the platforms in September?

Enough that any 6 week comparison spanning August and September is sitting on top of 3 kinds of change at once.

DateWhat changedWhat it does to your number
August 13Free ChatGPT users got a new default modelThe model behind the answer now depends on the account
August 14 and September 2Paid AI Mode users got 2 Gemini versions free users did notYour tracker and your buyer may be on different Gemini versions
August 24OpenAI took ChatGPT ads from 9 markets to 60, with 31 European countries liveSome of what a tracker counts on screen is now paid placement
August 31The Search Console AI report reached every siteA first baseline, not yet a trend
September 16OpenAI started testing sponsored agentsSame counting problem, new surface
September 24Search Console split web into text and multimodal, and the September spam update beganReporting definitions moved on the same day rankings did

The account level split matters more than it looks. Researchers asked for company recommendations in clean test sessions and got Tesla 93% of the time and Patagonia 91%. They then asked 800 real ChatGPT and Gemini users the same questions and saw Tesla 35% of the time and Patagonia 37%.

"If your tracker runs clean sessions, its number is the 93, but your buyers are actually closer to that 35." — Will Leatherman, Founder of Catalyst

The answers moved too. After the ChatGPT 5.6 release, listicles lost about half their share of ChatGPT citations. In the same weeks ChatGPT took in about twice as many sources per chat, up from 12.4 to 25.8, while the sources it actually cited rose by only about a quarter. A bigger pool explains part of the listicle drop. Most of it is real.

Sometimes the drop is the tool. One tracker showed Reddit at 3.83% of ChatGPT search citations through early August, then 0.5% across 4 days in mid August, and published the caveat that a data collection problem could not be ruled out. It told readers to check how many answers they had collected that week before treating it as real. That is check 1, run by the tracker on itself.

A real change looks different. When AI Overviews launched in France, Ahrefs found the most exposed sites lost 23.1% of their click-through in 9 days while a low exposure control group went up 2.3%. A spike happens once. A real change moves further than its control group.

Which 3 moves do you make once a change holds?

Cheapest first, and only one of them is new content.

  1. A day of work, for when you are named but never picked. Say what you are in the first reading of your main page, using the category and the buyer's question. Then let agents fetch that page. BambooHR sits here. Its homepage heading never said HR software, and Cloudflare bot protection was blocking AI agents from reading the page at all.
  2. A week of work, for when citations fall because a platform reweighted its sources. YouTube's share of AI Overview citations went from about 31% to 65% in 11 weeks and no site caused that. Show up where the platform moved.
  3. A month of work, for when you are cited but never named. Your page supplies the facts and nothing on it argues for you. Across 3 platforms one brand's citations rose while it was named only 1 to 5 times a run. Publish a figure nobody else can have, which is the same work described in building entity authority as a B2B startup.
"You want to say it, show up and then prove it." — Will Leatherman, Founder of Catalyst

What should you do this week?

Date every number you have already reported, including the screenshots and the slides, and write the question count next to each one.

Then run the 3 checks on the last move you reported. Same questions and platforms, enough runs, bigger than the noise.

"If it fails one, say so before someone else finds out." — Will Leatherman, Founder of Catalyst

For the next 2 weeks, run the same questions every day before you call anything a change. If you need a starting reading, the free AI visibility audit checks ChatGPT, Claude, Gemini and Perplexity, shows the share of buyer questions that mention you and which of them cite you, and takes about a minute. The longer walkthrough of setting your own question set lives in how to audit your AI search ranking in 20 minutes, and the prompts behind it are in the AEO skills pack.

The takeaway

A 33 point jump that nobody can defend costs more than a flat quarter, because the next number you report gets the same discount. Your score is one answer out of many, it hides 3 different counts, and the platforms moved 6 times in 5 weeks underneath it.

"Two weeks of the same questions is the first reading you can actually defend." — Will Leatherman, Founder of Catalyst

Lock your question list today and run it daily for 2 weeks. On the 15th day you will have the first AI visibility number you can put in front of a board without a caveat.

The Content Engineer

Enjoyed this article?

Frameworks like this, weekly. No fluff, just original research and actionable insight.

Ready to turn insight into pipeline?

We work with B2B companies that know content is the moat. Let's talk.