60–100 Runs: SPIV for CMOs to Fix AI Discovery KPIs
CMOs: Measure AI discovery with three KPIs, mention, citation, and Brand Recommendation Share. Use SPIV prompts and run 60–100 samples per model.
· 11 min read
Three metrics matter most: mention rate, citation rate, and Brand Recommendation Share, each tracked separately by model rather than blended into one number. The non-negotiable rule underneath them is statistical: run the same prompts 60 to 100 times or more per platform before trusting a percentage, and structure your prompt set with a framework like SPIV instead of guessing at what buyers ask.
TL;DR:
- Running 60 to 100 identical prompt samples per platform is necessary to obtain statistically reliable mention, citation, and recommendation rates.
- Building a prompt set using the SPIV framework ensures accurate representation across segments, personas, intent stages, and variations, reducing skewed results.
- Different AI models prioritize distinct signals; for example, Gemini emphasizes structured data, while Perplexity relies on recent content, so targeting fixes should be platform-specific.
- A comprehensive KPI dashboard should break out metrics by platform, report sample sizes, and include confidence intervals to provide meaningful insights for decision-makers.
- Avoid making budget decisions based on ad-hoc, unstructured prompts; structured, repeatable measurement processes reveal actual visibility gaps and enable credible tracking over time.
Table of Contents
- What Are the Core AI Discovery KPIs?
- Why Arbitrary Prompt Lists Distort Your Numbers
- How Do You Run a Statistically Valid Measurement Workflow?
- Which Platform Fixes Actually Move the Needle?
- Building a KPI Dashboard Executives Will Actually Trust
- Why Marketing Leaders Need to Retire the Ad-Hoc Prompt Test
- Turn These KPIs Into an Automated Executive Report
- Sources
- FAQ
What Are the Core AI Discovery KPIs?
Most teams start measuring AI visibility the way they'd measure a rank tracker: type a prompt, see what comes back, screenshot it, move on. That approach breaks the moment you realize a single ChatGPT response to an identical prompt is rarely reproducible. SparkToro's analysis found the probability of getting an identical brand list twice from the same prompt is under 1%, which means a screenshot from Tuesday tells you almost nothing about Wednesday.
Mention rate is the foundation. It's calculated as mentions divided by total valid responses, times 100. Simple, but it only tells you that you appeared, not how prominently.
Citation rate narrows the lens to answers that link directly to your domain, divided by total valid responses, times 100. A brand can have high mention rate and low citation rate, meaning models talk about you from memory but don't trust you enough to cite your site as a source.
Recommendation rate applies specifically to commercial intent prompts (things like "best X for Y") and measures how often you're actively recommended rather than just referenced in passing.
Brand Recommendation Share goes further, weighting your presence against how many competitors share the same answer. Being one of eight brands mentioned in a sprawling list is worth far less than being the singular recommendation. The Rhinegold Compendium's framework for this metric adjusts for list length specifically so you can tell exclusivity apart from crowding.
Position-weighted visibility is the optional fifth layer: it credits brands mentioned first or in the opening paragraph more heavily than those buried at the bottom of a long list.
Statistical Note: Cross-model agreement on the "top" brand recommendation across ChatGPT, Gemini, and Claude sits at roughly 41.6% in large-scale mapping studies, meaning platforms disagree on category leaders more often than they agree. Report every KPI with its sample size attached, not as a bare percentage.
- Mention rate = (mentions ÷ valid responses) × 100
- Citation rate = (domain citations ÷ valid responses) × 100
- Recommendation rate = (recommendations ÷ commercial prompt runs) × 100
- Brand Recommendation Share = list-length-weighted mention share, not raw mention count
Why Arbitrary Prompt Lists Distort Your Numbers
Ask ten marketers how they built their AI visibility prompt list and you'll hear the same answer: they typed whatever came to mind, maybe added a few competitor names, and called it research. That's the single biggest reason AI discovery KPIs get misread. A prompt list built from guesswork skews toward whatever the person writing it already knows, and it almost never reflects how actual buyers phrase questions across different stages of their research.
The SPIV framework fixes this by forcing coverage across four axes: Segments (industry verticals or company sizes you sell into), Personas (the actual job titles doing the research, since a CFO and a DevOps lead ask very differently phrased questions), Intent stages (awareness, comparison, and decision, each producing structurally different prompts), and Variants (phrasing differences, geographic modifiers, and constraint language like budget or compliance requirements).
For a B2B technology brand, that might look like this in practice:
- Cross a mid-market SaaS segment with a "Head of IT" persona at the comparison stage, phrased as "top alternatives to [competitor] for mid-size teams."
- Cross the same segment with a CFO persona at decision stage, phrased around total cost of ownership and contract terms.
- Cross an enterprise segment with a security persona at awareness stage, phrased around category education rather than named vendors.
Multiply a handful of segments by a handful of personas and intent stages and you land naturally in the 20 to 50 prompt range, which is enough to catch real variation without becoming unmanageable.
Pro Tip: *Weight your SPIV cells by revenue contribution, not by prompt volume.
This structure also makes competitor benchmarking meaningful. Comparing your mention rate against a rival's only makes sense if you're both being measured against the same buyer-representative prompt set, not one team's arbitrary list against another's.
How Do You Run a Statistically Valid Measurement Workflow?
Getting from "we checked ChatGPT once" to a defensible KPI requires a repeatable process, and the sampling volume is where most teams underinvest. SparkToro's research puts the odds of an identical ranking appearing twice at roughly 1 in 1,000, so a single run tells you almost nothing about your true visibility percentage. Sixty to 100 identical runs per prompt per model gets you close to a stable baseline; volatile categories or newly launched products may need more.
Every run needs version control. Log the exact prompt text, because a single word change (adding "2026" or swapping "best" for "top") shifts results meaningfully. When you rerun a stale prompt set, note the changelog so a KPI swing can be traced to a prompt change rather than a genuine visibility shift.
Per-run data capture should include:
- Prompt text and version number
- Model and model version (GPT-4o vs. a newer release behaves differently)
- Run timestamp
- Mention yes/no and citation URL if present
- Position within the response
- Sentiment or framing context
- Raw output or screenshot for audit purposes
| Capture field | Why it matters |
|---|---|
| Prompt version | Isolates whether KPI change came from wording, not visibility |
| Model version | Different model releases produce different rankings |
| Raw output | Enables evidence-based dispute resolution and audits |
| Timestamp | Ties results to a specific measurement window |
Manual spot checks work for a handful of prompts run occasionally. Once you're running 30+ prompts across four platforms at 60+ repetitions each, manual tracking becomes a full-time job, and that's the threshold where most teams move to a dedicated platform for automated sampling and evidence storage.
Which Platform Fixes Actually Move the Needle?
A low mention rate on ChatGPT and a low citation rate on Perplexity are different problems with different fixes, and treating them the same wastes budget. Each assistant rewards a different kind of signal, and matching the fix to the specific gap is what separates teams that improve their numbers from teams that just keep remeasuring the same problem.
- ChatGPT and Bing browsing: Check basic indexability, structured data, and canonical page clarity. If a live-crawl model can't parse your page cleanly, it won't cite it, regardless of content quality.
- Perplexity: Rewards publication cadence and earned third-party citations. This model leans on live search results, so recent coverage from outlets it trusts outperforms static owned content.
- Gemini: Wired closely to Google's Knowledge Graph, so schema markup and Google Business Profile completeness carry disproportionate weight here.
- Claude: Lacks live browsing in most configurations, so it leans on durable, widely cited mentions from media and industry sources baked into its training data.
Pro Tip: If citation rate is low but mention rate is respectable, the model already "knows" you but doesn't trust a source enough to link it, so prioritize earned media over more owned-content production.
Different assistants reward genuinely different infrastructure: structured data favors Gemini, live crawl freshness favors Perplexity, and durable mentions favor Claude, as Lureon's platform analysis lays out. Fix the weakest signal first rather than spreading effort evenly across all four.

Building a KPI Dashboard Executives Will Actually Trust

A dashboard that reports one blended "AI visibility score" hides more than it reveals. Executives need to see mention rate broken out by platform, citation rate, Brand Recommendation Share, commercial recommendation rate, which platforms you're winning outright, and a measurement confidence note attached to each number.
Cadence should match how quickly each platform's underlying data changes. Perplexity and ChatGPT's browsing mode pull from live search, so weekly tracking catches real movement. Gemini shifts more slowly with Knowledge Graph updates, making bi-weekly checks sufficient. Claude, without live browsing, moves at the pace of its training updates, so monthly is often enough.
- Weekly: Perplexity and ChatGPT browsing, owned by the SEO/GEO lead
- Bi-weekly: Gemini, owned by whoever manages structured data and local signals
- Monthly: Claude, owned by comms or PR given its dependence on media mentions
Every report should carry sample size, the prompt universe used, the last run date, a changelog noting any prompt or model version changes, and a short note on recommended next actions tied to whatever moved. A framework like this pulls directly from the KPI hierarchy structure that separates eligibility, visibility, authority, and revenue-level metrics rather than flattening everything into one score.
| Executive KPI | Reported alongside |
|---|---|
| Mention rate by platform | Sample size per platform |
| Citation rate | Confidence interval |
| Brand Recommendation Share | Volatility note for small samples |
| Recommendation rate (commercial) | Prompt universe size |
| Platform wins/losses | Last run date |
Why Marketing Leaders Need to Retire the Ad-Hoc Prompt Test
The risk isn't that AI assistants get your brand wrong occasionally. It's that leadership makes budget decisions off a handful of manually run prompts that happened to look good the week someone checked. That's not measurement, it's anecdote dressed up as data, and boards are starting to ask harder questions than "did you check ChatGPT."
Structured sampling across a properly built SPIV prompt universe is what turns a screenshot into a defensible number. In our work building AuthorityLayer's AI Authority Index, brands that segmented their prompt coverage by persona and intent stage consistently found visibility gaps that a single generic prompt list never surfaced, often concentrated in one platform rather than spread evenly. Marketing leaders who adopt this now will be the ones with a credible answer when the board asks how the brand actually shows up in AI-generated research.
— Geraldine
Turn These KPIs Into an Automated Executive Report
Building an SPIV prompt universe, running 60 to 100 samples per prompt across four models, and maintaining an evidence trail by hand is a real operational lift, even for a well-staffed marketing team, which is why many turn to tools like the AI Search Audit to optimize their website for AI search engines. Authoritylayer automates that entire workflow: multi-model sampling at statistically valid volumes, per-platform KPI breakdowns, stored raw outputs for audit purposes, and AI Authority Index scoring that maps directly to the metrics covered here.
The Monthly AI Visibility Report converts this measurement work into the exact executive dashboard described above: mention rate by platform, citation rate, Brand Recommendation Share, and a documented confidence note for every figure, delivered on a cadence your board will actually read. If you want to see where your gaps sit before committing to anything, start with a free AI visibility scan to see how your brand currently shows up across ChatGPT, Gemini, Claude, and Perplexity, or request a full executive AI Visibility Report to bring structured, repeatable numbers into your next planning meeting.
Sources
- Who Owns the AI Recommendation? A Multi-Industry Empirical Map of Brand Category Ownership Across Large Language Models — arXiv
- AI recommendations change with nearly every query — SparkToro analysis reported on Search Engine Journal
- Lureon
- AI search visibility is not a monolith; platform-specific behavior requires segmented measurement. — HubSpot blog
- Brand Recommendation Share — The Rhinegold Compendium
FAQ
What Is the Difference Between Mention Rate and Citation Rate?
Mention rate counts how often your brand is referenced in an AI response, while citation rate counts only responses that include a direct link to your domain. A brand can have strong mention rate and weak citation rate if models discuss it from memory without linking a source.
How Many Times Should I Run the Same Prompt?
Run each prompt 60 to 100 times identically per model before trusting the resulting percentage, since SparkToro's research found single-run results are rarely reproducible.
What Does SPIV Stand For in AI Visibility Measurement?
SPIV stands for Segments, Personas, Intent stages, and Variants, a framework for building a representative prompt universe instead of an arbitrary list of test queries.
Should I Track One Blended AI Visibility Score?
No. Platform-specific behavior varies enough that a blended score hides which assistant is driving your results, so segmented measurement by model is the more useful approach.
How Does Brand Recommendation Share Differ From Mention Rate?
Brand Recommendation Share weights your mentions against how many competitors appear in the same answer, distinguishing an exclusive recommendation from being one name in a crowded list.
