CMOs: Turn 30–100 Prompts into a Schema for AI Assistants
CMOs: build a reproducible AI visibility schema from a 30–100 prompt panel. Run weekly multi-model benchmarks, require raw exports, and prioritize...
· 10 min read
A schema for AI assistants is not a piece of code. It's an AI visibility measurement framework: a fixed set of metrics, prompts, and models that turns "does ChatGPT mention us?" into a repeatable number you can benchmark, report to a board, and act on. Run correctly, it produces trackable citation and recommendation share across the assistants your buyers actually use. The rest of this piece covers the metrics, a reproducible protocol, a procurement checklist, and where to spend your first budget dollar.
TL;DR:
- A standardized framework measures AI assistant visibility through consistent prompts, multi-model testing, and raw output collection, ensuring reproducible benchmarks.
- Metrics like mention rate, citation rate, share of voice, and recommendation strength are essential, with share of voice only meaningful when comparing against specific competitors across the same prompts.
- Running weekly, fixed prompt sets across multiple models avoids snapshot biases, making trend analysis and decision-making more reliable and defensible.
- Disclosure of methodology, prompt libraries, and raw exports is crucial for credible benchmarking, with enhanced transparency required for budget and board-level decisions.
- AuthorityLayer offers ongoing monitoring and benchmarking services, producing comprehensive reports based on a fixed, repeatable protocol to support marketing and procurement decisions.
Table of Contents
- What Does "Schema for AI Assistants" Mean for Marketers?
- What Metrics Should a Visibility Schema Track?
- How Do You Build a Reproducible Benchmarking Protocol?
- Measurement Quality and the Procurement Checklist
- A Prioritized Playbook: Presence, Prominence, Citation
- Why Standardized Tracking Beats One-Off AI Checks
- Turn This Framework Into a Running Benchmark
- Sources
- FAQ
What Does "Schema for AI Assistants" Mean for Marketers?
For a marketing leader, this schema is the operating structure behind a simple question: when a buyer asks ChatGPT, Gemini, Claude, or Perplexity for options in your category, does your brand show up, and how does it get described? It has nothing to do with website code. It's a measurement discipline covering the surfaces where AI now sits between your brand and the buyer: ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, and Copilot.
The schema exists because those six surfaces answer the same question differently, and answers shift week to week. A usable framework turns that noise into decisions: which competitor is winning the comparison prompts, whether a new PR placement moved your citation rate, and what to tell the executive team about recommendation share heading into next quarter. Without it, you're reporting anecdotes. With it, you're reporting benchmarks.
What Metrics Should a Visibility Schema Track?
Five metrics carry the weight. Each answers a different question, and together they map the funnel from "does the model know we exist" to "does it actively endorse us."
- Mention rate: the percentage of relevant prompts where your brand appears anywhere in the response. This is the baseline: no mention, no chance at anything else.
- Citation rate: the percentage of mentions backed by a linked source. A mention with no citation is weaker than a mention that points to a comparison page or review site, since assistants disproportionately cite brands that already appear in third-party listicles.
- Share of voice and first-mention rate: how often you appear relative to competitors across the same prompt set, and how often you land first when multiple brands are listed. First position tends to correlate with higher click and conversion behavior in downstream research.
- Recommendation strength: whether the model frames you as the answer, an option among several, or a passing reference. Phrasing matters because "best choice for enterprise teams" converts differently than "some companies also offer this."
- Sentiment and competitive gap: how favorably you're portrayed versus the next-closest competitor on the same prompts.
Pro Tip: Mention rate and citation rate are directional on their own. Share of voice only becomes decision-grade once you're running the same prompts against named competitors, on the same cadence, across every assistant in your panel.
How Do You Build a Reproducible Benchmarking Protocol?
A single ChatGPT query run from your personal account is a snapshot, not a benchmark. It reflects your account's history, your phrasing, and whatever that model happened to retrieve at that moment. A real protocol removes those variables so the number means the same thing next week, and the same thing to your competitor's marketing team running their own check.
- Build a fixed prompt set of 30 to 100 prompts covering the intent types your buyers actually use: category discovery ("best tools for X"), direct comparison ("X vs Y"), and bottom-funnel validation ("is X worth it"). Cover every stage of the buyer journey, not just top-of-funnel category questions.
- Run the identical prompt set across a multi-model panel: ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, and Copilot. Two models given identical source material agree only about 64% of the time, and Perplexity's independent retrieval often produces recommendations no other assistant surfaces. A single-model check simply can't catch that spread.
- Sample multiple times per prompt and average results, since model outputs are stochastic. Weekly re-runs of the full panel are the practical cadence for 2026: daily runs add compute cost without meaningfully more signal.
- Capture the raw answer text and the full source array for every response, not just a computed score. Store the raw output so metrics can be audited and disputed later.
- Reference an existing public protocol rather than inventing one from scratch. Structures like the MAHI-100 benchmark, which defines a fixed 100-prompt protocol with a decay model for planning refresh cycles, give you a documented starting point instead of guesswork.
Measurement Quality and the Procurement Checklist
Not every "AI visibility" number carries the same weight, and IAB's guidance on measuring visibility in the AI era exists precisely because more than 20 vendors currently run opaque, incompatible methodologies that produce conflicting scores for the same brand. Before you put a number in front of your CEO, know which tier it belongs to.
Minimum disclosure should cover platform coverage (which assistants were tested), a summary of the prompt library, the logic used to attribute a mention or measure sentiment, sampling cadence, and whether raw exports are available on request.
Enhanced disclosure, the bar for anything feeding a budget decision, adds the full prompt library itself, reproducibility metrics showing consistent scores across runs, documented panel validity, and historical score versioning so you can trace a shift back to a specific week.
| Use case | Acceptable disclosure tier |
|---|---|
| Trend monitoring / internal awareness | Minimum disclosure |
| Competitive benchmarking reports | Minimum to enhanced |
| Budget allocation / board reporting | Enhanced disclosure only |
| Vendor or agency procurement | Enhanced disclosure only |
A vendor unwilling to share prompt construction or raw exports isn't being protective of trade secrets. That's a red flag, and it belongs in your contract terms as a required disclosure, not a nice-to-have. AuthorityLayer's own methodology documentation is a useful reference point for what enhanced disclosure should actually look like in practice.
A Prioritized Playbook: Presence, Prominence, Citation
Fixing everything at once is how visibility programs stall. Practitioner testing across ten industries points to a clear sequence: presence first, prominence second, citation last, because each stage depends on the one before it.
Presence comes first. Standardize your entity description, the same two or three sentences describing what you do, everywhere: your site, your Wikipedia or Crunchbase profile, your G2 listing, your press mentions. Inconsistent descriptions across sources confuse the retrieval layer models rely on. Publish comparison and alternatives content that names your actual category and your actual competitors. Tools like the Best Generative Engine Optimization Tools guide can help you see what that content should cover.
Prominence comes next. Target the third-party comparison pages and industry roundups that already rank for your buyer's research queries, and pursue PR placements on high-authority sites rather than owned-blog volume. Keep your category label identical across every property you control, since AI assistants weight consistency heavily when clustering brands into a market.
Citation is last, and hardest. This is where you win placement on the specific review and comparison sites the assistants already pull from. Around 68% of brands cited by assistants appeared on at least three independent listicle pages ranking in Google's top ten for related queries, so ranking your own domain isn't enough by itself.
Run this on a 12-week cycle: test the prompt set, extract which sources got cited, fix or pursue placement on those specific sources, then re-run the benchmark to confirm the citation rate moved. Treat drift over that window as your governance signal, not noise.

Pro Tip: Don't chase citation wins before presence is solid. A brand with inconsistent entity descriptions rarely gets cited cleanly even when it lands on the right third-party page, because the model still can't confirm what you actually do.
Why Standardized Tracking Beats One-Off AI Checks
Most marketing teams start the same way: someone opens ChatGPT on their own account, types a category question, and reports back what came up. That instinct is understandable and almost useless as evidence. Personal accounts carry chat history, regional defaults, and account-level personalization that skew results in ways no two people will replicate. It's a snapshot of one moment, one account, one phrasing.
![]()
What actually holds up under scrutiny is the opposite discipline: a fixed prompt set, run cold, across a multi-model panel, on a fixed schedule. That's the only way to see whether Tuesday's dip in mention rate is a real trend or a rounding artifact of stochastic model behavior. Reproducibility isn't a compliance checkbox. It's the difference between a number you can defend in a board meeting and a screenshot you can't.
I'd go further than most vendors are comfortable saying out loud: if a provider can't hand you the raw prompt library and the raw model outputs behind their score, you don't have a measurement. You have their opinion, dressed up as data. Standardization is what turns a marketing hunch into something procurement, legal, and the CEO can all trust at the same time.
— Geraldine
Turn This Framework Into a Running Benchmark
Building and maintaining a 30 to 100 prompt panel across six assistants by hand, on a weekly cadence, with raw export capability, is exactly the kind of operational load that breaks internal teams after month two. AuthorityLayer runs that benchmark for you: continuous monitoring across ChatGPT, Gemini, Claude, and Perplexity, competitive gap analysis against the names you're actually losing to, and prioritized recommendations mapped to the presence, prominence, and citation sequence covered above.
The platform produces the enhanced disclosure a budget decision requires, including full prompt libraries and raw exports, so you're not stuck defending a black-box score in front of your CFO. If you want a starting benchmark before you commit to anything, the Monthly AI Visibility Report gives you a real read on where your brand stands against named competitors this quarter, with the same protocol described above running underneath it.
Sources
- IAB | Measuring Visibility in the AI Era
- How to measure LLM brand visibility (with real data)
- I Asked ChatGPT, Perplexity, and Gemini to Recommend Brands in 10 Industries. Here Is Who Gets Cited and Why
FAQ
What Is a Schema for AI Assistants?
It's an AI visibility measurement framework: a fixed prompt set, a multi-model panel, and a defined set of metrics used to track how often and how favorably AI assistants mention and recommend a brand.
How Often Should You Run an AI Visibility Benchmark?
Weekly re-runs of the full prompt panel across all assistants is the practical standard for decision-grade monitoring; daily runs add cost without meaningfully more signal.
What's the Difference Between Mention Rate and Citation Rate?
Mention rate tracks whether your brand appears anywhere in an assistant's answer; citation rate tracks whether that mention is backed by a linked source, which correlates strongly with prior presence on third-party comparison pages.
Why Do ChatGPT and Gemini Give Different Brand Recommendations?
Different assistants retrieve from different sources and apply different judgment; two models given identical source material still agree only around 64% of the time, which is why single-model checks miss most of the real picture.
Can AuthorityLayer Run This Benchmark for My Brand?
Yes. AuthorityLayer runs the multi-model panel, computes mention rate, citation rate, share of voice, and recommendation strength, and delivers prioritized recommendations through the Monthly AI Visibility Report.
