Growth & Strategy

Two AI-Search Startups Are Arguing About Margins of Error. The Number They're Fighting Over Isn't the Problem.

July 20, 2026

A competitor says the category's best-funded name is selling noise. The most interesting people in the argument think both sides are measuring the wrong thing.

Two AI-Search Startups Are Arguing About Margins of Error. The Number They're Fighting Over Isn't the Problem.
Credit:
powered by

Make State of Brand one of your go-to sources on Google

Google Icon
Add State of Brand on Google

Repetition matters when the unit of interest is a single exact prompt. Breadth matters when the goal is understanding a category, which is pragmatically what marketers are doing.

Ali Vaghar

Head of Data
Profound

The chief executive of a $19 million startup recently told a few thousand marketers that a $1 billion one has a data problem. Brian Stempeck runs Evertune, an AI-search measurement company, and the target of his post was Profound, the best-funded name in the same young category. Profound's sampling techniques, he wrote, carry a margin of error wide enough to make the data "unreliable for both measurement and optimization." Ask ChatGPT "what's the best SUV" once a day for a month, he argued, and the error bar on a brand's visibility score can reach 12 points. A brand that spends six figures on content and PR to lift its score 10 points, he argued, would never know what worked. It would be "chasing noise, not signal." The post pulled 462 reactions and 194 comments in two days. His statistics are sound. His description of how Profound actually measures is where the real argument starts.

That's what makes this particular LinkedIn fight worth more than the usual founder dunk. State of Brand readers already know the Great Flattening argument: AI answer engines are becoming the discovery layer, and a brand's visibility inside them is turning into a budget line. To make that case four days before Stempeck posted, we cited Profound's own study of 680 million citations. Now the company whose numbers we leaned on is being told those numbers are noise, by a competitor, in public. If you run a marketing budget, the fight underneath the fight is the one that matters. Can any of these AI-visibility scores be trusted enough to spend against? Follow the argument past the headline, and it lands on something uncomfortable for everyone selling one of these scores.

The math is the easy part

Start with Stempeck's core point, because it's a serious one. His arithmetic holds. For a brand that shows up in about 10 percent of answers, 30 runs of a prompt leave an error bar of roughly 11 points, and getting to about 6 points takes around 100 runs. That's textbook sampling. The instability underneath it is real, too. Language models don't return the same answer twice, and not because anyone set them to be creative. Researchers at Thinking Machines Lab showed last year that sampling one model a thousand times at temperature zero, the setting meant to make output deterministic, still produced 80 different completions. So repetition does buy precision, and a category selling confidence off 30 data points hasn't earned it.

Stempeck is also a competitor, which is worth saying plainly. Evertune's method is built on the exact fix he prescribes: it samples each prompt around 100 times. That doesn't make his statistics wrong. It's context a reader deserves, and it's part of why the response from Profound is worth reading closely, because Profound had already run this experiment itself.

The experiment had already been run

Eight days before Stempeck's post, a member of Profound's founding team, Josh Blyskal, published the company's test of this exact concern, built on analysis by Profound economist Jennifer Zou. They took a portfolio of 753 prompts across seven engines and ran it two ways for two weeks, once a day against ten times a day. The gap between them came out to 0.25 of a percentage point.

Here's the part the once-a-day description leaves out. Profound doesn't run one prompt over and over. It runs a portfolio of hundreds or thousands of different prompts, each roughly once a day, and reports the pooled average. "If you are tracking a real portfolio of prompts," Blyskal wrote, "the portfolio itself is already doing a lot of averaging for you." The 12-point error Stempeck cites is the error on a single prompt. It is not the error on the aggregate a marketer actually looks at, which pools hundreds of prompts and lands far tighter. Ali Vaghar, who works on data at Profound, made the point directly in a reply on Stempeck's own post that drew 76 reactions: "Repetition matters when the unit of interest is a single exact prompt. Breadth matters when the goal is understanding a category, which is pragmatically what marketers are doing."

So the real disagreement is narrower than "Profound has a data problem." Profound measures a wide net of different questions once each. Evertune measures a smaller set many times over. Those are two defensible answers to a genuine tradeoff, and the company on the receiving end had already run and published this test. Independent research lands between them. A study by the analyst Antonio Blago and the academic paper "Don't Measure Once" both recommend a blend, roughly 20 to 30 runs across 50 to 100 distinct prompts, rather than either extreme. The two approaches compare like this:

Core method:

  • Profound: Broad portfolio of different prompts, each run about once a day, pooled into one average

  • Evertune: A smaller set of prompts, each run about 100 times

Prompts per report:

  • Profound: Hundreds to thousands (753 in its published test)

  • Evertune: A smaller set, heavily repeated

Repetitions per prompt:

  • Profound: About 1 per day

  • Evertune: About 100

Real-user data:

  • Profound: 1.5 billion+ licensed prompts

  • Evertune: A 25 million-person panel

Optimizes for:

  • Profound: Category breadth

  • Evertune: Per-prompt precision

Funding to date:

  • Profound: $155M ($96M Series C at a $1B valuation)

  • Evertune: $19M

The argument nobody's actually having

But the most interesting voices in the thread didn't take either side. They think both companies are arguing about the wrong number.

The sharpest came from Daniel Cheung, who's building a measurement tool called ternith. With 753 prompts, he pointed out, the average was always going to converge; that was never in doubt. The finding that actually mattered was sitting inside Profound's own data, where changing which prompts were in the portfolio moved the result more than any amount of repetition did. "So the real uncertainty isn't sampling error," he wrote, "it's the sampling frame." A stable average also hides what's underneath it. Ten percent citation share, as Cheung put it, "might be 10% everywhere, or 100% on some prompts and zero on the rest," which are completely different situations for a marketer deciding where to spend. His closing question is the one neither dashboard answers. When a brand is absent, can the measurement tell you why?

Several people pushed on the Gallup comparison itself. Stempeck had reached for Gallup, whose polls hit a 4-point margin by surveying 1,000 respondents. But Gallup gets there by asking 1,000 different people once, not one person a thousand times, and repeating a single prompt is the second thing. "A hundred clean repeats of an unrepresentative prompt is precise and still wrong," wrote Deepshikha Dhankhar, a GEO strategist, restating the whole objection in a sentence.

Stempeck made much of this point himself in the replies. "You do need breadth of prompts," he wrote to one commenter, and "we do in fact use panel data" to decide which prompts to run. Which means the two firms agree on more than the headline lets on. Both use real user data to choose prompts. Both run many of them. The live disagreement shrinks to how many times you repeat each one, and which number you put on the slide.

What another hundred runs can't fix

Then Rand Fishkin, who has been measuring this stuff since before the category had a name, widened the problem past both of them. His worry isn't sampling at all. It's personalization. AI systems tailor answers heavily off even a little user history, and if that's true, Fishkin wrote, "I just don't know whether the same brands ever get recommended the same way to two different people," even across thousands of anonymized runs. His conclusion is the quiet alarm under the whole debate: "the whole field of AI tracking is built on this oversight, and we don't know how bad it is."

Emmanuel Dolle, who runs a tool called Bubbling, named the other thing repetition can't touch. A prompt run daily for 30 days isn't 30 independent draws. It's a time series. Models get updated, indexes shift, the web moves underneath the measurement. "That's drift, not sampling error, and no repetition count removes it," he wrote, and on a before-and-after comparison, drift is usually the larger term. Run the prompt a hundred times and you've tightened the smaller source of error while the bigger one sits untouched.

Arman Advani, Head of Partnerships at Search Atlas, sees both camps circling the same blind spot. "Both sides are really arguing about noise," he said. "Sample too few prompts and you're measuring randomness, not your brand... that's just math. But even with perfect sampling, AI answers shift day to day, so no one's handing you one clean number to chase. The real question is whether your structured data, entity facts, and content are solid enough for a model to cite you. Fix that and your visibility moves everywhere, not just on the prompts a vendor happened to sample. Measurement matters, but only if it points at something you can act on, and not just a score to report."

There's a business problem under the statistical one, and it belongs to every vendor in the category, Evertune included. Nobody yet knows what AI visibility is worth. As one growth advisor in the thread put it, the industry is "selling brands on a metric that we can't measure," with no line back to revenue. And the firm scoring your visibility is usually the same firm selling you the fix. "It's the personal trainer who sells you the supplements," wrote Steven Perlman, who builds a visibility benchmark of his own. "The one grading your progress also profits from what they tell you to buy." None of this is new. A measurement startup called Metricus published a piece in April, "Your AI Visibility Dashboard Is Grading Its Own Homework," naming Profound, Semrush, and Ahrefs, and a 2026 academic paper called "Don't Measure Once" told the field the same thing. Stempeck's post put a sharper public point on a critique the category had already been circling.

Four better questions

So what does a brand actually do with this? The answer from the people not selling a dashboard is a blend, not a winner. An independent study by the analyst Antonio Blago, running 144,000 evaluations across three models, landed on something like 20 to 30 runs spread over 50 to 100 prompts, neither extreme. The academic work points the same way. Which means the question to put to a vendor isn't "what's my score." It comes down to four harder ones:

  1. What's the confidence interval around that number?

  2. Which prompts define my market, and who picked them?

  3. How do you separate my content's effect from the model's own drift?

  4. Can you show me the work?

A vendor who can answer those is measuring something. A vendor who hands over one confident number and no error bar is selling precision it can't back.

The category will get its rigor eventually; the money flowing into it guarantees that. When it does, the edge won't belong to whoever runs the most prompts or repeats them the most times. It'll belong to the first vendor willing to tell a CMO what its number can't do. Until then, everyone's optimizing against a decimal point that doesn't know who's asking.

Common questions

Is running an AI-visibility prompt once a day enough?

For a broad portfolio of different prompts, largely yes. Profound's own test found that running prompts once a day versus ten times a day changed its citation-share reading by 0.25 points. Repetition matters mainly when you need precision on one specific prompt.

How many times should you run a prompt to measure AI visibility?

Independent research points to a blend, roughly 20 to 30 runs across 50 to 100 different prompts. About 100 repetitions tightens a single low-visibility brand's margin of error to around 6 points.

Is Profound's data reliable?

Its pooled portfolio metric is far more stable than the single-prompt critique implies. The harder, unsolved problems are prompt selection, model drift, and personalization, and those affect every vendor in the category.

Profound vs Evertune: what's the difference?

Profound measures a wide net of prompts once each, which is breadth. Evertune runs fewer prompts many times, which is per-prompt precision. Both use real-user panel data. The real gap is which number they choose to headline.

Outlever Logo

If this caught your attention, that’s not accidental.


Decoration line

The best editorial systems don’t happen by accident. Outlever builds them.

Partial view of green concentric circles with a solid green dot on the outermost circle on a light background.Concentric green circles with a single solid green dot on a dashed circle on a light background.Minimalist design with faint curved lines and scattered small green dots on a white background.

Come back for the reason it lands.


Subscribe for the kind of thinking that makes people stop, read and come back.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.