Data on AI visibility is unstable: generative models produce different answers each time, so a single measurement can be misleading. A new study by IQRush proposes a stopping rule to determine when ratings become reliable. In this article, we will break down how to distinguish real growth from measurement noise and what practical conclusions follow for marketers.
Why Single Measurements of AI Visibility Are Unreliable
Search engines like SearchGPT, Gemini, and Perplexity introduce randomness into every answer. The same query can yield different sources — this is a feature of the architecture. The study showed that, for example, when testing SearchGPT on the topic of running gear, Tom’s Guide received about 9.5% of citations, while Runner’s World received approximately 6.0%. The difference of 3.5 percentage points was within the margin of error, so claiming Tom’s Guide’s superiority was incorrect.
How Much Data Is Needed for a Reliable Rating
The answer consists of two conditions that must be met simultaneously. First: the order must stop changing. After collecting a sufficient number of responses, the top sites begin to stand out clearly. Second: the difference between the top sites must exceed the measurement error. If competitors are too close, the rating does not reflect real superiority.
In 30 tests across different platforms and topics, the number of responses needed to meet both conditions ranged from 33 to 94 (only responses with citations were considered). In three out of 30 cases, this was not achieved even after 125 queries — all on SearchGPT, where the top sites were too similar.
Practical Takeaways for Marketers
Rand Fishkin (SparkToro) advises: before spending money on tracking AI visibility, make sure the provider «shows their math.» The IQRush study provides a simple stopping rule to avoid relying on intuition. If after updating content you see a 3 percentage point increase in citations, this could be natural variation. Measure indicators before and after multiple times — a single measurement is indistinguishable from noise.
Different Platforms — Different Data Requirements
Gemini loads citations from the same sites within a single response, so many citations carry little new information. SearchGPT gives fewer citations per response but distributes them more widely — each response contains more independent data. The same number of responses on two platforms does not provide the same confidence: a budget sufficient for Gemini may leave you in the dark on SearchGPT.
When Data Is Insufficient: Know When to Stop
In three out of 30 tests, the top sites never clearly separated. In such cases, the right decision is to refrain from publishing a rating. A tracker that can say «insufficient data» is more valuable than one that outputs a confident order with every query.
Only leaders can be trusted: with enough responses, they pull away from the middle and tail. But even for the top 10, the typical margin of error is about five positions, and every fifth position is more than 10. Do not publish exact rankings beyond the top of the list.
Limitations of the Study
This is a preprint based on 30 tests across three platforms using questions generated by ChatGPT, not real user queries. The exact numbers do not transfer to your topics — treat them as a form of the problem, not a table of values. In one test, 125 questions yielded only 104 useful responses (17% loss), so the actual number of queries should be higher.
The method was validated internally: the early rating was compared to the final one, not to an external benchmark. However, an independent team from the University of St. Gallen (Julius Schulte, Malte Bleeker, Philipp Kaufmann) published similar results on their dataset in April, confirming that a single reading is unreliable.
The Future of AI Visibility: From Exact Numbers to Ranges
Reporting on AI visibility is moving toward a format with margins of error, as in advertising and web analytics. Until Search Console reports which clicks came from AI, the task falls on you: run the check multiple times and report a range, not a single number from a dashboard.
Frequently Asked Questions
How many times should you query AI search to get stable data?
From 33 to 94 responses with citations, depending on the platform and topic. There is no universal threshold.
Can you trust AI visibility dashboards?
Only if they show the margin of error and multiple measurements. A bare number without context is a red flag.
What to do if citations increase by 3% after a content update?
Take several measurements before and after. If the difference persists — it’s real growth. If not — it’s noise.
Which platforms require more data?
SearchGPT — due to fewer citations per response but greater independence of each response. Gemini — conversely, gives many citations, but they often repeat the same sites.
Conclusion
The main takeaway: AI visibility is not a fixed metric but a range. Use the stopping rule, check multiple times, and do not publish ratings if the top sites have not separated. Only then will you get data you can rely on. Start implementing these principles today to make your AI visibility reports reliable and useful for decision-making.
