The second draw is worth almost nothing

We asked 20 questions three times each across 34 scans to find out how much the repeats actually measure. The answer made us widen our own confidence intervals by 71%.

Last updated 21 September 2026

The question we wanted answered

We ask every question 3 times per engine, and we say so on the front of every report: a mention rate, not one lucky draw. That is a real claim about the design, and it costs 3× the engine spend.

So we asked ourselves the uncomfortable version: how much does the second draw actually buy? Three draws feel like three times the evidence. Whether they are is a measurement, not an opinion, and we had 34 scans to measure it on.

The test, which needs no assumptions

Within one scan, each question is asked 3 times. If those draws were independent observations, then for any question we could predict how often the three answers would disagree with each other — using nothing but the overall mention rate.

Across 34 scans of our own account: 1,146 question-engine groups, 3,438 answers, and an overall mention rate of 41.3%. At that rate, independence predicts that 834 of the 1,146 groups would have split their vote — some answers naming us and some not.

22 did.

What that means

98.1% of the groups gave the same answer every single time we asked. Translated into the language we use for confidence intervals: the correlation between draws of the same question is 0.974, and the design effect is 2.95. The 3,438 answers carry the information of about 1,167 independent observations — not 3,438.

Which is another way of saying what our own methodology page has said for a while, now with a number attached: the uncertainty in a visibility figure comes from how many questions you asked, not from how many times you asked them. The second draw of a question mostly re-asks it.

What we changed

The correction applies to the per-engine figures on our reports, which were computed from the number of answers as though each were independent. Under that assumption our own intervals were 71% too narrow. They are wider now.

We want to be precise about what moved and what did not. The point estimate is unchanged: a customer who read 45% still reads 45%. Only the range beside it moved, and that range was wrong the whole time. Silently shifting the estimate would have turned a correction into a data change, and the correction is not entitled to that.

Our reports now also show the effective number of observations next to the raw one — “named you in 27 of 60 answers · about 20 independent observations”. We would rather you see the pair than take our word for the correction.

Where this number comes from, and where it does not

This was measured on our own account, across mostly one site, with a fixed set of twenty questions. The intraclass correlation is a property of a question set and a market, not a universal constant — which is why the correction is estimated per scan rather than hard-coded from this figure. A different customer with a different question set could measure a lower number, and we are not going to publish 0.974 as though it applied to them.

It is also, itself, a small sample of sites. What we can say is what we measured: on this question set, on these engines, the repeats were nearly redundant.

Ask us the question this raises

If you are evaluating tools in this category, this is now a fourth question alongside the three we already suggest — and for most products it is the one nobody has an answer to:

  • How many times is each question asked, and is that reported?
  • Is there an interval on the figure, or a single value that looks precise?
  • When an engine call fails, is it counted as “not mentioned”, or excluded from the denominator and shown to you?
  • Do repeated asks of the same question count as independent observations? If a tool says it asks each prompt ten times, ask what those ten are worth. Sometimes they are worth one.

Our methodology, including the parts we know are weak.