Measuring something that answers differently every time
Ask ChatGPT the same question three times and you get three different answers. That makes AI visibility a sampling problem, and that is exactly how we treat it.
1. Every question, multiple draws
One measurement is a snapshot of a system that guesses. Geony asks every question three times per measurement round by default, per assistant separately, and stores every draw individually. On screen you therefore see dots per draw and not an averaged-away number: "2 of 3" and "67%" are the same figure, but the first shows how narrow the basis is.
2. Every percentage carries a Wilson interval
For every presence percentage we show the 95% confidence interval (Wilson score), the same statistic that medical studies use for small samples. Three measurements with one hit is "33% (6 – 79%)", and that bandwidth is not a weakness but the truth. The more measurements, the narrower the band.
Try it yourself: mentioned 36%, but how sure are you of that?
With 9 measurements the real visibility lies with 95% certainty between 14% and 67%: so wide that the figure says almost nothing. A single one-off check is this.
3. The alarm stays silent on overlap
A drop from 56% to 22% sounds alarming. But if the intervals overlap, the difference cannot be distinguished from chance, and then we send no alert. This comes straight from our own measurement history: that same drop occurred on a single day, same configuration, and was noise. A tool that alarms on that trains you to ignore alarms.
4. Assistants separately, never just the average
In our measurements the same brand appeared in 36% of Claude answers, 5% on ChatGPT and 1% on Gemini. Their average (14%) describes no assistant at all. That is why every assistant gets its own column with its own interval, and the alerting is evaluated per assistant too: a brand can disappear entirely on one assistant while the average shows nothing.
5. Sources from the metadata, not from the running text
Which sources an answer cites we read from the assistant's own structured citation data. A language model that names URLs in running text often invents them; metadata does not. Every source in Geony is therefore traceable to a real answer at a real moment.
6. What we do not do
We show no sentiment scores as long as we cannot prove them against a validated set. We report no trends on too little data; below the threshold the system would rather stay silent than guess. And we do not write "AI-optimized content" for you: we show which sources determine the answer, the work stays the work.
Within a day you'll know whether AI mentions or ignores your brand
The first measurement runs on your own prompts, across three assistants, with real figures. No demo data, no sales call first.
Prefer to talk it over first? Book a call