researchJuly 2026

The Measurement Problem in Generative Search: Why Topic-Level Tracking Beats Prompt-Level Checking

Most tools that claim to measure "AI visibility" track individual prompts. That method measures noise. LLMs return materially different answers to the same question depending on phrasing and run-to-run randomness — and modern AI search fans a single query into 8–12 sub-queries before synthesising an answer. The reliable signal sits one level up, at the topic.

J
James Deverick · Evolv Agency
45%
Output swings from meaning-preserving prompt rephrasings — the same question, different words, different answer (Cao et al.)
8–12
Sub-queries generated by AI Mode query fan-out before a single answer is synthesised — not the prompt you typed
25–39%
Overlap between traditional Google rankings and AI search citations — ranking #1 guarantees nothing about being cited
~%
of AI-cited pages sit outside the classic organic top ten — visibility has moved to a layer rankings don't see

The unit-of-measurement error

Point a tracking tool at a large language model, ask it a question, and note whether your brand shows up. Do it across a list of prompts and you have a dashboard. It looks like measurement. It behaves like anecdote. Three well-documented properties of these systems explain why.

Three problems with prompt-level tracking

Why a single prompt is not a stable unit of measurement

1. Stochasticity

Generative models sample from a probability distribution. Run an identical prompt twice and the output can differ — an effect that widens as the model's temperature setting rises. Any measurement that treats a single response as a fact is recording one throw of the dice and reporting it as the state of the world.

2. Prompt sensitivity

An LLM asked the "same" question in two phrasings a human would consider equivalent can return substantially different answers. Cao and colleagues recorded performance swings of up to 45% across semantically equivalent formulations. He and colleagues found formatting changes alone — nothing to do with meaning — could move results by 40%. Whether your brand appears can depend on a word you chose when writing the test prompt.

3. Query fan-out

Google's AI Mode, ChatGPT Search, Perplexity and Gemini do not run a single search against the exact words a user enters. They decompose the question into 8–12 parallel sub-queries, retrieve sources for each, and synthesise one answer. One 2026 analysis found a large majority of fan-out sub-queries changed from one search to the next. The sources cited are those that satisfied the hidden sub-queries — not the seed prompt you tested.

Stack the three together and the conclusion is unavoidable. A prompt-level check samples one stochastic response, to one arbitrary phrasing, resolved against a hidden and shifting set of retrievals. Run the same audit twice and you get two different pictures. That is not reproducible measurement — it is a screenshot with a trend line drawn on afterwards. It is also trivially cherry-pickable — a fact worth remembering the next time a tool or an agency shows you a flattering prompt result.

The topic is the correct unit

Ask a language model about a subject ten different ways and you will get ten different answers — but the mean and variance across those answers reveal a stable underlying tendency. The individual answers are noisy samples. The aggregate is the measurement.

A topic is the natural unit at which that aggregate becomes meaningful. By topic we mean a cluster of semantically related questions, phrasings and sub-queries that orbit a single subject — for an enterprise technology brand, something like "sovereign cloud for the public sector" or "VMware migration options," not one canonical sentence.

We don't build measurement on prompts at all: a unit that will not hold still cannot anchor a metric. Reading an individual answer can still be a useful qualitative spot-check — to see the phrasings buyers use or catch an inaccurate product summary — but that is investigation, not measurement. The measurement happens one level up.

Why topics, not prompts

Four reasons topic-level measurement is the correct approach

It is reproducible

Aggregate across enough phrasings and repeated runs and the random component averages out. Two topic-level measurements taken under the same method converge; two prompt-level checks do not. Reproducibility is the line between a metric and an anecdote.

It matches how the machines actually work

Fan-out already treats a user's question as a cluster of retrievals, not a string to match. Measuring at the topic aligns the metric with the mechanism the answer is actually built by, instead of against a seed query the system discarded.

It matches how buyers actually search

Nobody researches a purchase by asking one perfectly-worded question. A consideration is a session — a spread of questions, follow-ups and reformulations around a theme. The topic is the unit the buyer experiences too.

It is decision-useful

"We hold 34% share of voice across the sovereign-cloud topic, up six points quarter on quarter" is a number a board can trend and a strategy can act on. "We appeared in ChatGPT for this exact prompt last Tuesday" is not.

What topic-level measurement looks like in practice

At Evolv, generative search monitoring is topic-level by design — built into the Four-Layer Visibility Stack methodology from the ground up, not bolted onto prompt tracking.

Define the topic set

Map the cluster of subjects that correspond to real buying decisions in the category, rather than a keyword list. This is the unit everything else is measured against — broader than a keyword, narrower than a category.

Aggregate visibility across each topic

Track presence, citation share, competitor share and source patterns aggregated over the phrasings and sub-queries that make up a topic — across multiple engines — rather than logging one-off prompt responses.

Work from a topic-level exportable dataset

We use topic and source datasets via Waikay structured at the topic level, rather than screenshotting individual chat answers. The dataset, not the screenshot, is the source of truth.

Measure on a fixed cadence

The deliverable is the trend over a monitoring window, not a snapshot. A single audit, however clean, cannot distinguish signal from noise. Repeated measurement on a fixed cadence can.

Practical implications

What to do about it

For anyone buying, running or reporting generative search visibility, the practical implications are direct.

Report topic-level share of voice, not prompt appearances

If a report leads with "we showed up for this prompt," ask what the topic-level trend is. If there isn't one, you are looking at an anecdote dressed as a metric.

Build content as topic clusters, not single keyword pages

Fan-out rewards estates that answer many facets of a subject well over pages that answer one query vaguely. Comprehensive topical coverage is both a visibility lever and the thing that makes topic-level measurement possible.

Judge movement over a window, not a single audit

Insist on a cadence. One measurement cannot separate a real gain from a lucky sample. A documented before-and-after against the same competitor set is the minimum viable proof of progress.

Interrogate any "AI visibility" tool on its unit of measurement

Two questions settle it: does it aggregate at the topic level, and is the result reproducible if you run it again? If the honest answer to either is no, it is measuring noise with a professional interface.

Generative search has genuinely changed where and how brands are found. It has not repealed the basic requirement of measurement — that the thing you count should hold still long enough to count it. Prompts don't. Topics do.

Frequently Asked Questions

Does Evolv use prompt-level tracking?

Not as standard. We measure at the topic level exclusively — an individual prompt is too unstable to anchor a metric. Reading individual answers can be a useful qualitative spot-check for spotting phrasings or an inaccurate product summary, but we treat that as investigation, never as measurement or reporting.

What counts as a "topic" here?

A cluster of semantically related questions, phrasings and sub-queries around one subject that maps to a real buying decision — for example "sovereign cloud for the public sector" rather than a single exact-match query. Deliberately broader than a keyword and narrower than a category.

How is topic-level share of voice calculated?

By aggregating visibility signals — presence, citation share, competitor share and source patterns — across the phrasings and sub-queries that make up a topic, across multiple AI engines, and repeating on a fixed cadence so the trend, rather than any single response, is the output.

Does this apply to ChatGPT and Perplexity, or just Google? /

Across the board. Google AI Mode and AI Overviews, ChatGPT Search, Perplexity and Gemini all decompose a question into multiple sub-queries before synthesising an answer — so the single-prompt instability problem is universal, not a Google-specific issue.

How often should generative search visibility be measured?

On a regular, fixed cadence rather than as one-off audits. AI engines also weight recency heavily for time-sensitive topics, so a monitoring window additionally captures decay and gains that a snapshot would miss entirely.

What is the difference between a topic and a keyword cluster?

A keyword cluster is defined by lexical similarity — words that look alike. A topic is defined by the buying decision it maps to — the full cluster of questions a buyer asks during a consideration, regardless of the words they use. A topic includes semantic variants, follow-up questions, and the sub-queries fan-out generates internally.

Method and sources

This article draws on the peer-reviewed literature on LLM prompt sensitivity and output stability, on published descriptions of query fan-out from Google and from independent search analysts, and on 2025–2026 industry measurement research.

  • Google, AI Mode in Google Search: Updates from Google I/O 2025 (blog.google) — origin and description of the query fan-out technique.
  • Google Search Help, Get AI-powered responses with AI Mode in Search — confirmation that AI Mode and AI Overviews use query fan-out across subtopics and data sources.
  • Errica et al. (2024), What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering (arXiv:2406.12334).
  • Research on semantic and syntactic prompt sensitivity documenting output swings of up to ~40–45% from meaning-preserving rephrasings and formatting changes (Cao et al.; He et al.), surveyed in recent clinical-LLM stability work (arXiv, 2026).
  • Independent 2026 analyses of query fan-out behaviour and overlap between traditional rankings and AI-search citations (SparkToro Office Hours, Jan 2026; Surfer SEO, Dec 2025).
  • Digital Agency Network, Generative Engine Optimization Statistics (2026) — the GEO measurement gap as the leading operational challenge reported by agencies.

Figures attributed to third parties are cited for context. The methodology described under "What topic-level measurement looks like in practice" is Evolv's own.

Evolv Agency Research

Want this kind of intelligence for your brand?

We run GSO programmes for enterprise technology brands and produce original research as part of the engagement.

Talk to Us