Now booking · Q3 2026 · limited capacity

Notes on AEO · Note 2 · Jul 2026

One screenshot proves nothing.

§1

A screenshot is one run.

AI answers are probabilistic. The same question, asked the same way, can name a firm on one run and skip it on the next. A screenshot captures one run out of a distribution.

That cuts both ways. When a vendor shows you a screenshot with your competitor in it, they’re showing you one run. A screenshot with your own firm in it is weak evidence too. Neither settles the question that matters: out of all the answers buyers get, how many name you?

§2

What a real claim looks like.

A rate with a margin of error, from a fixed protocol. For example: named in 36% of 240 sampled answers, plus or minus 6 points at 95% confidence.

The rate says how often the engines cite you. The interval says how sure the measurement is. The fixed protocol, meaning the questions, engines and counting rules stay the same every time, is what makes two measurements comparable at all. Without all three, the number doesn’t tell you much more than a screenshot would.

§3

Why 240 runs.

Because the margin of error has to be smaller than the decision riding on it. Here’s what sample size does to the interval, at a measured share of about a third:

10 runs±28 ptsworse than guessing out loud
20 runs±20 ptsstill an anecdote with decimals
60 runs±12 ptsgetting somewhere, not actionable
240 runs±6 ptstight enough to act on, and to re-check

At around 240 runs the interval gets small enough to tell a firm that’s never cited apart from one that’s cited some of the time. It’s also small enough to pick up real movement between two measurements, which is the point of measuring twice.

§4

What sampling noise does.

Sampling noise is the run-to-run variation that has nothing to do with real change. A small sample can show “improvement” that is luck.

If a firm moves from 2 citations in 10 runs to 4 in 10, nothing has been demonstrated. If it moves from 0% to 18% with intervals of plus or minus 5 on both measurements, something real happened. Any provider who guarantees results should be measuring rigorously enough for the guarantee to mean something.

§5

Four questions before you pay anyone.

  1. Q1How many runs is that number based on?
  2. Q2What is the margin of error?
  3. Q3Is the protocol fixed between measurements?
  4. Q4Will you show the raw counts?

Anyone measuring properly can answer all four on the spot, because the answers are the product. If what comes back is “proprietary methodology”, treat that as a no.

Mitra holds itself to the same standard. Every audit we run reports counts, intervals and sources, the free one included: run the audit on your firm →