Ask an assistant the same question twice and you will often get two different answers, different brands, different order, different confidence. Nothing is broken. You are measuring a probabilistic system with a live retrieval feed, and it does not owe you the same answer twice.
Here is where the variance comes from, the screenshot fallacy it produces, and the measurement approach that works anyway.
The six sources of variance
| Source | What changes between runs |
|---|---|
| Probabilistic generation | The model samples words; identical inputs legitimately produce different outputs |
| Retrieval variance | Search-grounded runs fetch different pages at different moments |
| Silent model updates | Providers swap and tune models without notice; distributions shift overnight |
| Context and memory | Conversation history and user profile tilt recommendations per person |
| Phrasing sensitivity | "Best" vs "recommended" vs "what should I" pull different distributions |
| Provider experiments | A/B tests mean two users get two systems on the same day |
Classic rank tracking assumed a deterministic scoreboard: position 4 was position 4 for everyone. Every assumption in that sentence is gone, and the measurement has to change with it.
The screenshot fallacy
The variance produces a predictable ritual: someone asks once, sees their brand first, screenshots it for the deck. Someone else asks Tuesday, sees a rival, declares the visibility lost.
Both are reading one draw from a distribution as if it were the distribution. The honest unit of AI visibility is not "our rank in the answer"; it is "our share of the answers", which only sampling can see.
Measuring like a pollster
The tracking tools already work this way, and the method matters more than the vendor: multiple runs per prompt, a phrasing family per intent rather than one sacred wording, share-of-appearances as the metric, and weekly windows read as trends.
The numbers that come out are poll numbers: "present in 70 percent of runs for this prompt family, trending up four points over the month". That statement survives variance; "we are number one in ChatGPT" does not survive the next refresh.
It is the same statistical humility the leaderboard's average ranks encode: even the most-recommended sites on earth are describing their typical draw, not a fixed throne.
Stability itself is a signal
Sample enough and you notice prompts differ in how much they vary, and that variance is diagnostic.
Locked distributions, the same brands nearly every run, mark categories where model memory dominates: entrenched associations, the mention currency banked over years. Entry there is the long game.
Volatile distributions, different names each run, mark categories still decided by retrieval: the model has no settled opinion and quotes whatever it finds. Those are the winnable ones right now, because fresh quotable pages swing exactly such answers.
The operating rules
Never conclude from one run, in either direction. Track share weekly, read monthly, and alert only on sustained share shifts, the same discipline as any volatile metric. Date your measurements, because silent model updates make "compared to when" half of every finding.
And when share moves, check both doors before reacting: a retrieval-driven dip points at pages and freshness, a memory-driven one points at reputation, and the paired scorecard tells you which door the intent runs on.
The one-line takeaway: AI answers vary because sampling, retrieval, silent updates and phrasing all move between runs. Measure like a pollster: phrasing families, repeated runs, share-of-appearances, weekly trends, and read stability itself as the map of what is winnable now.