Everyone in this industry has an opinion about AI visibility. Almost nobody has a measurement system. So for 90 days I ran a fixed set of buyer questions through five AI engines on a fixed schedule, logged every source each one cited, and turned it into a share-of-voice number I could actually defend.
The reason to bother is uncomfortable arithmetic. Ahrefs reported that AI search visitors made up just 0.5% of their traffic but drove 12.1% of signups roughly a 23x conversion premium over organic (Ahrefs, 2026). A channel that small and that valuable is exactly the kind you can’t afford to guess at. This post is the methodology, the mistakes, and the results.
Key Takeaways
- One check tells you nothing. Running the same ChatGPT prompt three times leaves only 2.2% of citations intact, across an 815,000 prompt-page-pair analysis (Growth Memo / AirOps, 2026).
- Most of the noise is the model, not you. In a 12,933-response variance study, 34.8% of variance came from resampling the same prompt while brand identity explained 0.7% (Żatuchin, arXiv, 2026).
- Five runs is the practical ceiling. Past the fifth repeat, extra runs cut error variance by ~0.0003; breadth beats depth (Żatuchin, arXiv, 2026).
- My 90-day result: citation share of voice moved from [EXPERIMENT DATA: X%] to [EXPERIMENT DATA: Y%] across [EXPERIMENT DATA: N] tracked questions.
- It’s worth it AI traffic converts at a large premium, so even small citation gains are disproportionately valuable.
Across [EXPERIMENT DATA: N] tracked buyer questions run five times each per engine per week, seofreelancerguru.com’s citation share of voice went from [EXPERIMENT DATA: X%] at baseline to [EXPERIMENT DATA: Y%] at day 90. Citation rate the share of runs where I appeared at all moved from [EXPERIMENT DATA: A%] to [EXPERIMENT DATA: B%].
The gains were not evenly distributed. [EXPERIMENT DATA: Engine name] moved first, at around week [EXPERIMENT DATA: N]; [EXPERIMENT DATA: Engine name] barely moved at all. That asymmetry turned out to be the most useful thing in the dataset, and I’ll come back to it.

Because AI answers are non-deterministic, a one-off spot check measures the model’s randomness rather than your visibility. This is the finding that should reshape how the whole industry reports on AI search, and it’s now well quantified.
Kevin Indig, working with AirOps across 815,000 prompt-page pairs, found that running an identical prompt three times in ChatGPT left just 2.2% of citations intact (Growth Memo, 2026):

Source: Growth Memo / AirOps, 2026. Any tool reporting a single-run result is reporting noise.
A July 2026 arXiv paper by Dmitrij Żatuchin decomposed where that noise actually comes from, across 12,933 responses covering 20 brands, 8 languages, 3 models, and 15 prompts. The result is humbling for anyone doing single-run brand checks:

Five repeats per prompt per engine after that, extra runs buy almost nothing. Żatuchin’s paper puts a number on the diminishing return: a repeat past the fifth reduces relative-error variance by only 0.0003, while adding another language cuts variance roughly 15x more per query than another repeat (arXiv, 2026). Indig’s independent practitioner recommendation lands in the same place: five reps per prompt per platform, weekly.
That convergence is what I built the protocol on. Spend your budget on more questions and more engines, not more repeats of the same question.
Fifteen buyer questions, five engines, five runs each, once a week, logged into one spreadsheet. That’s 375 observations per week and roughly 4,875 over the quarter. It takes about 90 minutes a week by hand, which is the honest cost.
Not keywords. Questions. I used the fifteen I hear on discovery calls, phrased the way people say them out loud: “who’s the best SEO consultant for a small B2B company,” “how much should I pay an SEO freelancer,” “is GEO worth it in 2026.” Half should be ones you could plausibly win; half should be aspirational, or you’ll never see movement.
Freeze this list on day one. Changing questions mid-experiment destroys comparability this is the mistake I made, and I’ll get to it.
Write these down and don’t deviate:
One row per run. The columns that earned their place:
| Column | Why it matters |
|---|---|
| Date, engine, question, run # | The identity of the observation |
| Cited domains (all) | The denominator for share of voice |
| Was my domain cited? (Y/N) | Citation rate |
| My cited URL | Tells you which page is doing the work |
| Position of my citation | Early citations carry more weight |
| Competitor domains present | Share-of-voice comparison |
| Answer snippet | Lets you see how you were framed |
That last column is the one people skip and later wish they hadn’t. Being cited as a caveat is not the same as being cited as the recommendation.
Citation rate how often you show up at all:
Citation rate = runs where your domain appears ÷ total runs × 100
Citation share of voice how much of the cited surface you own:
Citation SoV = your cited URLs ÷ all cited URLs across those runs × 100
Track both. Citation rate moves first and feels good; share of voice is the one that correlates with actually being the answer. Report them weekly with the run count attached, so a reader can see the sample behind the number.
Every run already gives you the full list of cited domains, so competitor tracking is free. Pick five competitors on day one and compute their share of voice with the identical formula. Without that comparison, a flat line is ambiguous you can’t tell whether you’re stuck or whether the whole category got squeezed by an engine update.
The pattern that mattered wasn’t the headline number it was that different engines responded to completely different work.
Three errors, all methodological, all worth stealing the fix for:
Yes, on the arithmetic AI referrals are small but convert at a large premium, so knowing which engine cites you is directly commercial. Ahrefs’ own data showed AI search visitors at 0.5% of traffic driving 12.1% of signups (Ahrefs, 2026):

Two honest caveats. That’s one company’s analytics, not a cross-industry study, and published conversion premiums range from roughly 1.3x to 27x depending on who measured and how they defined a conversion. And a 90-day series is a small sample against a system this noisy treat direction as the finding, not precision.
Want this run on your domain instead of mine? Book a Free AI-Visibility Check → and I’ll baseline your citation share across five engines using exactly this protocol.
For what to change once you can see the numbers, start with the GEO Playbook → and Reddit Youtube Off Site GEO → for the Reddit and YouTube half.
Run a frozen list of buyer questions through each AI engine on a fixed weekly schedule, five runs per question per engine, logged out with personalization off, and record every cited domain per run. Repetition is essential: an identical ChatGPT prompt run three times retains only 2.2% of its citations (Growth Memo / AirOps, 2026), so single checks measure randomness rather than visibility.
Five. A variance-components study of 12,933 LLM responses found that a repeat past the fifth reduces relative-error variance by only about 0.0003, while broadening across languages or models delivers roughly 15x more variance reduction per query (Żatuchin, arXiv, 2026). Spend additional effort on more questions and more engines rather than more repeats.
Citation share of voice is your cited URLs divided by all cited URLs across a set of tracked runs, expressed as a percentage. It differs from citation rate the share of runs where you appear at all because a rising citation rate with flat share of voice usually means answers got longer, not that you became more prominent. Track both.
Because generative engines are non-deterministic and their retrieval layer re-queries live. In a 2026 variance decomposition, resampling the same prompt accounted for 34.8% of total variance and query language 31.6%, while brand identity itself explained just 0.7% (Żatuchin, arXiv, 2026). The fix is sampling design, not a better single check.
Not to start. A spreadsheet and 90 minutes a week produces a defensible series, and building it manually teaches you what the tools are actually doing. Consider paid tooling once you’re tracking more questions or engines than you can run by hand but check how many runs per prompt any vendor uses before trusting its numbers.
Ninety days of manual tracking taught me less about tactics than about measurement. The engines are noisy enough that most published AI visibility claims mine included, until I fixed the protocol are within the margin of error. Sampling design is the entire game: freeze your questions, repeat five times, log every run, and compare by engine.
The short version:
I run this protocol continuously on seofreelancerguru.com and publish the numbers as they move. Want yours baselined? Let’s Talk →.
Sources retrieved 2026-08-10: