Ask an assistant the same question twice and you will often get two different answers, different brands, different order, different confidence. Nothing is broken. You are measuring a probabilistic system with a live retrieval feed, and it does not owe you the same answer twice.

Here is where the variance comes from, the screenshot fallacy it produces, and the measurement approach that works anyway.

The six sources of variance

SourceWhat changes between runs
Probabilistic generationThe model samples words; identical inputs legitimately produce different outputs
Retrieval varianceSearch-grounded runs fetch different pages at different moments
Silent model updatesProviders swap and tune models without notice; distributions shift overnight
Context and memoryConversation history and user profile tilt recommendations per person
Phrasing sensitivity"Best" vs "recommended" vs "what should I" pull different distributions
Provider experimentsA/B tests mean two users get two systems on the same day

Classic rank tracking assumed a deterministic scoreboard: position 4 was position 4 for everyone. Every assumption in that sentence is gone, and the measurement has to change with it.

The screenshot fallacy

The variance produces a predictable ritual: someone asks once, sees their brand first, screenshots it for the deck. Someone else asks Tuesday, sees a rival, declares the visibility lost.

Both are reading one draw from a distribution as if it were the distribution. The honest unit of AI visibility is not "our rank in the answer"; it is "our share of the answers", which only sampling can see.

One prompt, sampled 10 times: who appeared? brand A: 8 of 10 runs brand B: 6 of 10 brand C: 3 of 10 brand D: once ← the measurement A single conversation shows one row's worth of luck. The bars are what is actually true this week.
Screenshots sample the distribution once. Tracking samples it enough to draw the bars.

Measuring like a pollster

The tracking tools already work this way, and the method matters more than the vendor: multiple runs per prompt, a phrasing family per intent rather than one sacred wording, share-of-appearances as the metric, and weekly windows read as trends.

The numbers that come out are poll numbers: "present in 70 percent of runs for this prompt family, trending up four points over the month". That statement survives variance; "we are number one in ChatGPT" does not survive the next refresh.

It is the same statistical humility the leaderboard's average ranks encode: even the most-recommended sites on earth are describing their typical draw, not a fixed throne.

Stability itself is a signal

Sample enough and you notice prompts differ in how much they vary, and that variance is diagnostic.

Locked distributions, the same brands nearly every run, mark categories where model memory dominates: entrenched associations, the mention currency banked over years. Entry there is the long game.

Volatile distributions, different names each run, mark categories still decided by retrieval: the model has no settled opinion and quotes whatever it finds. Those are the winnable ones right now, because fresh quotable pages swing exactly such answers.

The operating rules

Never conclude from one run, in either direction. Track share weekly, read monthly, and alert only on sustained share shifts, the same discipline as any volatile metric. Date your measurements, because silent model updates make "compared to when" half of every finding.

And when share moves, check both doors before reacting: a retrieval-driven dip points at pages and freshness, a memory-driven one points at reputation, and the paired scorecard tells you which door the intent runs on.

The one-line takeaway: AI answers vary because sampling, retrieval, silent updates and phrasing all move between runs. Measure like a pollster: phrasing families, repeated runs, share-of-appearances, weekly trends, and read stability itself as the map of what is winnable now.