Why your ChatGPT answer is different from mine — and what that means for measuring it
AI assistants give different answers to the same question because they are probabilistic, personalised by your history, and — when they browse — reading a web that changes hourly. That variance is real and cannot be engineered away. It means a one-off check tells you almost nothing, and any tool claiming to show you what an assistant told a specific customer is overselling.
Three separate sources of variance
1. The models are probabilistic
A language model does not retrieve one correct answer. At each step it samples from a distribution of likely next words. Two runs of the identical prompt can legitimately produce different text, and if two businesses are close in the model's estimation, which one gets named first can flip between runs.
2. Answers are personalised
The consumer apps know things about the person asking: their location, their earlier conversations, saved memory, sometimes their account history. Two people in the same town asking the same question can get materially different lists. This is by design and it is getting more pronounced, not less.
3. The web underneath is moving
When an assistant browses, it summarises what it finds at that moment. A directory that reshuffles its ordering, a new review, a competitor publishing a guide — any of these changes the inputs. A grounded answer is a snapshot of a moving target.
What this means for measurement
If you take one answer as evidence, you will draw the wrong conclusion roughly as often as the right one. That cuts both ways: the reassuring result where you appear first is as unreliable as the alarming one where you are absent.
What is stable is the aggregate. Ask a dozen realistic buying questions across several assistants, repeat it weekly, and the noise averages out. A business named in 5% of answers in March and 40% in June has genuinely moved, even though no individual answer in either month was reliable on its own.
The useful framing: you are not measuring a fact, you are measuring a tendency. Treat the number like a weather forecast rather than a thermometer.
How to spot an honest tool
Since this variance is inherent, how a vendor talks about it tells you a great deal.
- Ask what they actually query. Almost everyone uses provider APIs, because that is the only programmatic access there is. That is fine — but it should be stated, not obscured behind "we monitor ChatGPT".
- Ask how often. A tool that checks once and shows you a number is selling you noise.
- Ask whether you can see the raw answers. If you cannot read the text a score was derived from, you cannot check it.
- Be wary of "what ChatGPT told your customer". Nobody has that. It is the clearest signal that a vendor is either confused or willing to mislead.
- Ask about share of voice. Tools built for big brands ask you to supply prompts containing your own name. If nobody searches your business by name yet, that metric will read zero forever and tell you nothing.
What we do, stated plainly
Nusan asks provider APIs a consistent set of category-and-location buying questions — never containing your business name, because the point is unprompted recall — and records whether you were named, how far into the answer, how you were described, and which sources the assistant cited. It repeats weekly so the trend is visible.
It is a directional index. It is not a transcript of anyone's conversation, and we will not describe it as one. You can read every raw answer behind every score, which is the only way you can sensibly check whether you believe us.
Common questions
Is the variance a bug?
No. These models sample from a distribution rather than looking up a fixed answer, so some variation is inherent to how they work. Personalisation and live browsing add more on top.
So is measuring it pointless?
No — but a single check is. Individual answers are noisy; the pattern across a dozen questions and several assistants, repeated on a schedule, is stable enough to be meaningful. You are measuring a trend, not a fact.
Can a tool show me what ChatGPT told my customer?
No. Those conversations are private and personalised. Any tool that implies otherwise is either confused or misleading you.
Why do API results differ from the app?
The consumer apps carry memory, personalisation, a different system prompt and live browsing. API results are a cleaner, more consistent signal — which makes them better for comparison over time, but they are a proxy rather than a transcript.
See whether the assistants name you
Three real customer questions for your trade and town, no signup, about a minute.
Run my free check