This is the measurement with which Geony proved itself, shared in full because the figures tell the story better than we do. Subject: a Dutch creative production company. Setup: sixty customer questions, each asked three times to ChatGPT, Gemini and Claude; 180 measurements in total, all with web grounding.
The outcome that decided everything
For the most important competitor in the field (an international comparison site) the spread was largest: mentioned in 36% of the Claude answers (interval 26 to 47), 5% in ChatGPT (2 to 13) and 1% in Gemini (0 to 7). The same questions, the same week, three realities. The intervals of Claude and ChatGPT do not overlap: this difference is proven, not sampling chance.
Why the assistants disagree
The explanation was in the sources. Of the 398 domains that the three assistants cited together, 316 appeared with exactly one assistant; only nine domains were cited by all three. So each assistant reads a different internet, and therefore recommends different companies. Whoever builds their AI strategy on one assistant optimises for a third of the field.
From measurement to task list
The source analysis made it concrete: the company's own domain was cited nowhere, while two comparison platforms together determined half the answer. The task list wrote itself: first make the mentions on those platforms complete and up to date (the shortest route to the answer), then make your own site citable for the questions where no one was mentioned; that turned out to be half of all the questions.
What you can take away from this
Three things. Measure across multiple assistants or do not measure at all; the differences are not noise but structure. Demand intervals with every percentage, otherwise you cannot tell real change from chance. And start with the sources, not with your own site: the answer is decided elsewhere. You can find the full method here, and we publish the ongoing figures every quarter in the AI Visibility Index.