I Tracked My Own AI Citation Share for 90 Days

  • Home
  • SEO
  • I Tracked My Own AI Citation Share for 90 Days
Laptop displaying a website analytics dashboard with traffic metrics, pageviews, performance charts, and active page data.

I Tracked My Own AI Citation Share for 90 Days

Everyone in this industry has an opinion about AI visibility. Almost nobody has a measurement system. So for 90 days I ran a fixed set of buyer questions through five AI engines on a fixed schedule, logged every source each one cited, and turned it into a share-of-voice number I could actually defend.

The reason to bother is uncomfortable arithmetic. Ahrefs reported that AI search visitors made up just 0.5% of their traffic but drove 12.1% of signups roughly a 23x conversion premium over organic (Ahrefs, 2026). A channel that small and that valuable is exactly the kind you can’t afford to guess at. This post is the methodology, the mistakes, and the results.

Summarize This Guide With Your Favorite AI

Key Takeaways

  1. One check tells you nothing. Running the same ChatGPT prompt three times leaves only 2.2% of citations intact, across an 815,000 prompt-page-pair analysis (Growth Memo / AirOps, 2026).
  2. Most of the noise is the model, not you. In a 12,933-response variance study, 34.8% of variance came from resampling the same prompt while brand identity explained 0.7% (Żatuchin, arXiv, 2026).
  3. Five runs is the practical ceiling. Past the fifth repeat, extra runs cut error variance by ~0.0003; breadth beats depth (Żatuchin, arXiv, 2026).
  4. My 90-day result: citation share of voice moved from [EXPERIMENT DATA: X%] to [EXPERIMENT DATA: Y%] across [EXPERIMENT DATA: N] tracked questions.
  5. It’s worth it AI traffic converts at a large premium, so even small citation gains are disproportionately valuable.

What 90 days of tracking actually showed

Across [EXPERIMENT DATA: N] tracked buyer questions run five times each per engine per week, seofreelancerguru.com’s citation share of voice went from [EXPERIMENT DATA: X%] at baseline to [EXPERIMENT DATA: Y%] at day 90. Citation rate the share of runs where I appeared at all moved from [EXPERIMENT DATA: A%] to [EXPERIMENT DATA: B%].

The gains were not evenly distributed. [EXPERIMENT DATA: Engine name] moved first, at around week [EXPERIMENT DATA: N]; [EXPERIMENT DATA: Engine name] barely moved at all. That asymmetry turned out to be the most useful thing in the dataset, and I’ll come back to it.

Line chart comparing weekly AI citation share of voice across ChatGPT, Perplexity, Gemini, Google AI Overviews, and Google AI Mode over 13 weeks.

Why a single check tells you nothing

Because AI answers are non-deterministic, a one-off spot check measures the model’s randomness rather than your visibility. This is the finding that should reshape how the whole industry reports on AI search, and it’s now well quantified.

Kevin Indig, working with AirOps across 815,000 prompt-page pairs, found that running an identical prompt three times in ChatGPT left just 2.2% of citations intact (Growth Memo, 2026):

Donut chart showing 2.2% of citations persisted across three identical ChatGPT runs while 97.8% changed between runs.

Source: Growth Memo / AirOps, 2026. Any tool reporting a single-run result is reporting noise.

A July 2026 arXiv paper by Dmitrij Żatuchin decomposed where that noise actually comes from, across 12,933 responses covering 20 brands, 8 languages, 3 models, and 15 prompts. The result is humbling for anyone doing single-run brand checks:

Source: Żatuchin, 2026. The thing you’re trying to measure explains 0.7% of the variance. Design around that.

How many runs do you actually need?

Five repeats per prompt per engine after that, extra runs buy almost nothing. Żatuchin’s paper puts a number on the diminishing return: a repeat past the fifth reduces relative-error variance by only 0.0003, while adding another language cuts variance roughly 15x more per query than another repeat (arXiv, 2026). Indig’s independent practitioner recommendation lands in the same place: five reps per prompt per platform, weekly.

That convergence is what I built the protocol on. Spend your budget on more questions and more engines, not more repeats of the same question.

The tracking system I built

Fifteen buyer questions, five engines, five runs each, once a week, logged into one spreadsheet. That’s 375 observations per week and roughly 4,875 over the quarter. It takes about 90 minutes a week by hand, which is the honest cost.

Step 1 — Choose questions a buyer would actually ask

Not keywords. Questions. I used the fifteen I hear on discovery calls, phrased the way people say them out loud: “who’s the best SEO consultant for a small B2B company,” “how much should I pay an SEO freelancer,” “is GEO worth it in 2026.” Half should be ones you could plausibly win; half should be aspirational, or you’ll never see movement.

Freeze this list on day one. Changing questions mid-experiment destroys comparability this is the mistake I made, and I’ll get to it.

Step 2 — Fix the sampling protocol before you start

Write these down and don’t deviate:

  • Engines: ChatGPT, Perplexity, Gemini, Google AI Overviews, Google AI Mode.
  • Runs: 5 per question per engine.
  • Cadence: same weekday, same rough time, every week.
  • Session state: logged out or in a fresh incognito profile, personalization and memory off. Personalized results measure your own history, not the market.
  • Location: fixed. Set it once and never change it.

Step 3 — Log every run, not just your wins

One row per run. The columns that earned their place:

ColumnWhy it matters
Date, engine, question, run #The identity of the observation
Cited domains (all)The denominator for share of voice
Was my domain cited? (Y/N)Citation rate
My cited URLTells you which page is doing the work
Position of my citationEarly citations carry more weight
Competitor domains presentShare-of-voice comparison
Answer snippetLets you see how you were framed

That last column is the one people skip and later wish they hadn’t. Being cited as a caveat is not the same as being cited as the recommendation.

Step 4 — Compute two numbers, not one

Citation rate  how often you show up at all:

Citation rate = runs where your domain appears ÷ total runs × 100

Citation share of voice  how much of the cited surface you own:

Citation SoV = your cited URLs ÷ all cited URLs across those runs × 100

Track both. Citation rate moves first and feels good; share of voice is the one that correlates with actually being the answer. Report them weekly with the run count attached, so a reader can see the sample behind the number.

Step 5 — Track your competitors in the same sheet

Every run already gives you the full list of cited domains, so competitor tracking is free. Pick five competitors on day one and compute their share of voice with the identical formula. Without that comparison, a flat line is ambiguous you can’t tell whether you’re stuck or whether the whole category got squeezed by an engine update.

What the data showed

The pattern that mattered wasn’t the headline number it was that different engines responded to completely different work.

  1. Movement was engine-specific. [EXPERIMENT DATA: Engine] responded within [EXPERIMENT DATA: N] weeks of on-site restructuring; [EXPERIMENT DATA: Engine] didn’t move until off-site mentions accumulated. If I’d tracked one engine, I’d have drawn the wrong conclusion twice.
  2. Updated pages outperformed new ones. [EXPERIMENT DATA: X of Y] of my first citations came from pages I refreshed rather than published consistent with what I saw in my Perplexity Citation Experiment → .
  3. Volatility persisted at the page level. Even in weeks where my citation rate was flat, the specific URL being cited changed constantly, which matches the 2.2% persistence finding.
  4. Competitor share was noisier than mine, which taught me not to over-read any single week’s competitive read.

What I got wrong in the first 30 days

Three errors, all methodological, all worth stealing the fix for:

  1. I changed the question list in week 3. It felt like an improvement and it wrecked comparability for the first month. Freeze the list; park new questions in a separate “watchlist” tab.
  2. I ran single checks at first. My week-one and week-two numbers were effectively random. I rebuilt to five runs after reading the volatility research, and only then did the series stabilize.
  3. I tracked citation rate alone. A rising citation rate with flat share of voice means the answer got longer, not that you got more important. Adding the SoV denominator changed how I read my own progress.

Is 90 minutes a week worth it?

Yes, on the arithmetic AI referrals are small but convert at a large premium, so knowing which engine cites you is directly commercial. Ahrefs’ own data showed AI search visitors at 0.5% of traffic driving 12.1% of signups (Ahrefs, 2026):

Chart showing AI search accounts for 0.5% of Ahrefs visitors but 12.1% of signups, representing roughly a 23x conversion premium over organic search.
Source: Ahrefs, 2026. Single-company data, so treat the multiple as directional rather than universal.

Two honest caveats. That’s one company’s analytics, not a cross-industry study, and published conversion premiums range from roughly 1.3x to 27x depending on who measured and how they defined a conversion. And a 90-day series is a small sample against a system this noisy treat direction as the finding, not precision.

Want this run on your domain instead of mine? Book a Free AI-Visibility Check → and I’ll baseline your citation share across five engines using exactly this protocol.

How to replicate this in a week

  1. Write 15 buyer questions the way people say them out loud. Freeze the list.
  2. Pick your engines. Five is ideal; three is workable.
  3. Build the sheet with the columns above, one row per run.
  4. Run the full baseline 5 runs × every question × every engine before changing anything on your site.
  5. Compute citation rate and citation share of voice, plus the same two numbers for five competitors.
  6. Change one thing at a time on-site, then off-site. Give each at least three weeks.
  7. Re-run weekly, same day, same settings, logged out.
  8. At day 90, compare by engine not in aggregate. The aggregate hides the mechanism.

For what to change once you can see the numbers, start with the GEO Playbook →  and Reddit Youtube Off Site GEO →  for the Reddit and YouTube half.

Frequently Asked Questions

How do you track AI citations?

Run a frozen list of buyer questions through each AI engine on a fixed weekly schedule, five runs per question per engine, logged out with personalization off, and record every cited domain per run. Repetition is essential: an identical ChatGPT prompt run three times retains only 2.2% of its citations (Growth Memo / AirOps, 2026), so single checks measure randomness rather than visibility.

How many times should I run each prompt?

Five. A variance-components study of 12,933 LLM responses found that a repeat past the fifth reduces relative-error variance by only about 0.0003, while broadening across languages or models delivers roughly 15x more variance reduction per query (Żatuchin, arXiv, 2026). Spend additional effort on more questions and more engines rather than more repeats.

What is AI citation share of voice?

Citation share of voice is your cited URLs divided by all cited URLs across a set of tracked runs, expressed as a percentage. It differs from citation rate the share of runs where you appear at all because a rising citation rate with flat share of voice usually means answers got longer, not that you became more prominent. Track both.

Why do AI answers change every time I ask?

Because generative engines are non-deterministic and their retrieval layer re-queries live. In a 2026 variance decomposition, resampling the same prompt accounted for 34.8% of total variance and query language 31.6%, while brand identity itself explained just 0.7% (Żatuchin, arXiv, 2026). The fix is sampling design, not a better single check.

Do I need a paid AI visibility tool?

Not to start. A spreadsheet and 90 minutes a week produces a defensible series, and building it manually teaches you what the tools are actually doing. Consider paid tooling once you’re tracking more questions or engines than you can run by hand but check how many runs per prompt any vendor uses before trusting its numbers.

Conclusion

Ninety days of manual tracking taught me less about tactics than about measurement. The engines are noisy enough that most published AI visibility claims mine included, until I fixed the protocol are within the margin of error. Sampling design is the entire game: freeze your questions, repeat five times, log every run, and compare by engine.

The short version:

  • Single checks are noise only 2.2% of citations survive three identical runs.
  • Five runs per prompt is the practical ceiling; spend the rest on breadth.
  • Track two metrics citation rate and citation share of voice.
  • Compare by engine, never in aggregate they respond to different work.
  • My 90-day result: [EXPERIMENT DATA: X%] → [EXPERIMENT DATA: Y%] share of voice.

I run this protocol continuously on seofreelancerguru.com and publish the numbers as they move. Want yours baselined? Let’s Talk →.