Guide

AI Visibility Metrics Explained: The Full Reference

A reference guide to every AI visibility metric that matters: share of voice, mention rate, citation rate, sentiment, position, and more.

Why AI Visibility Needs Its Own Metrics

Search engine optimization built its metrics around a list: rankings, positions, click-through rate, impressions. Those metrics work because a search results page is a stable, orderable object — position 3 today is comparable to position 3 tomorrow. AI answers are not lists. ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews and Copilot generate a fresh block of text for every prompt, and that text may name your brand, name a competitor instead, name both, or name neither. There is no fixed slot to rank in.

That's why AI visibility requires a different measurement vocabulary. Instead of asking "where do I rank," the useful questions become: how often am I named, how am I named, where in the answer am I named, and does that change by engine and over time. This guide defines each metric a marketer needs to track AI visibility properly, how it's conceptually calculated, and what to actually do when the number moves.

Share of Voice, Share of Answer, and Share of Model

Share of voice is the foundational AI visibility metric. Run a fixed set of prompts relevant to your category, count every brand mention across every response, and calculate what percentage of total mentions belong to you versus competitors. If your brand appears in 40 of 200 total brand mentions across your tracked prompt set, your share of voice is 20%. It answers the core competitive question: when this category gets discussed, how much of the conversation is about you?

Share of answer narrows the lens to a single response. When an AI answer names three brands, share of answer looks at how much of that specific answer's attention — word count, detail, or emphasis — goes to each one. Two brands can both be "mentioned" in the same answer while one gets a full paragraph of explanation and the other gets a passing clause; that's a share-of-answer gap, invisible to a simple mention count.

Share of model is share of voice calculated separately for one AI engine. Your share of voice on ChatGPT and your share of voice on Perplexity are two different numbers, and both matter more than a blended average across engines (see the section on cross-engine benchmarking below). Track share of voice as your top-line competitive metric, but always look at it per engine, and use share of answer to explain why two brands with similar mention counts get treated very differently.

Mention Frequency and Mention Rate

Mention frequency is the raw count: how many times your brand was named across a tracked set of prompts and engines over a given period. It's the simplest metric to compute and the easiest to misread, because a rising count can simply mean you added more prompts to your tracking set, not that AI visibility actually improved.

Mention rate fixes that by normalizing: it's the percentage of prompt runs in which your brand appears at all, regardless of whether competitors also appear. A mention rate of 15% on a set of 100 category prompts means you showed up in 15 of them. Unlike share of voice, mention rate doesn't care about the competitive field — it only asks whether you were part of the answer.

Use mention frequency to spot volume changes worth investigating, but use mention rate as the metric you actually report on and compare over time. A mention rate that's climbing even while share of voice stays flat usually means the whole category conversation is growing faster than your competitive position — worth knowing, and a different problem than losing ground to a named competitor.

Citation Rate and Citation Depth

A mention is a name-drop; a citation is a sourced reference — a linked or attributed pointer back to a specific page on your domain. Citation rate is the percentage of relevant answers that cite your site as a source, and it matters most on engines that visibly show sources: Perplexity, Google AI Overviews, and Copilot. ChatGPT and Claude cite less consistently, but citation still happens and is worth tracking separately from mention rate.

Citation depth looks at what gets cited, not just whether. Are AI engines pulling from your homepage over and over, or are they reaching into specific product pages, comparison pages, and documentation? Depth also captures multiplicity: does a single answer cite one page on your site, or several?

The gap between mention rate and citation rate is diagnostic. High mention rate with low citation rate means the models already know your brand — likely from training data or reputation — but aren't finding your own content authoritative or crawlable enough to cite. That's a signal to fix content structure and citability, not brand awareness. Low citation depth concentrated on one page means your site has one asset doing all the work and nothing else structured well enough to surface.

Sentiment: Positive, Negative, and Neutral Mentions

Not every mention helps you. Sentiment classifies each brand mention as positive (recommended, praised, described favorably), negative (criticized, described as a poor fit, actively discouraged), or neutral (named factually — listed as an option — without an evaluative framing either way). Most AI-generated brand mentions are neutral or mildly positive; genuinely negative mentions are rarer but disproportionately important because they can talk a prospect out of considering you before they ever reach your site.

Sentiment scoring conceptually works by having a model — or a human reviewer — read the surrounding context of each mention and classify the tone, not just presence. An aggregate sentiment score (like "72% positive") is a useful summary, but it's a summary, not a diagnosis.

Always read the actual quotes behind a sentiment shift rather than acting on the score alone. A drop in positive sentiment might mean one AI engine started describing you as "better suited for small teams" when you sell to enterprise — a framing problem you can address in your own content — versus a factual error about pricing or features, which is a different fix entirely. Track sentiment delta (the change over time) alongside the raw split; a stable score hides a slow negative drift just as easily as it hides improvement.

Mention Position

When an AI answer lists multiple brands — "top tools for X include A, B, and C" — position is where your brand falls in that list. Being named first is not the same as being buried third or fourth, even though a naive mention count treats them identically. Users skim, and AI answers themselves often front-load the option the model considers most relevant, so first position correlates with being treated as the primary recommendation.

Position is harder to track consistently than presence because list order can vary between runs of the identical prompt — another reason single snapshots mislead (see the trend section below). Track position as a distribution over repeated runs, not a single observed rank.

When you're consistently mentioned but consistently buried, the useful question isn't "how do I get mentioned more" — you already are. It's "what does the brand named first have that I don't." Compare the context around the top-positioned competitor's mention against your own; the differentiator is usually specific proof points, pricing clarity, or a narrower use-case claim that the model can confidently lead with.

Prompt Coverage

Prompt coverage measures how much of your actual category's question space you're tracking, out of everything a prospect might plausibly ask an AI engine that touches your market. Most teams start by tracking a handful of obvious prompts — "best [category] tools," their own brand name — and stop there. That's a small, biased sample of a much larger set that includes comparison questions, use-case questions, budget questions, integration questions, and objection-handling questions.

There's no fixed target number, because it depends entirely on your category's breadth, but the direction to watch is coverage relative to your own prompt library growing over time, and relative to the range of ways real buyers phrase their questions. Low coverage means your other metrics are computed on a biased sample and may not represent your actual AI visibility.

Expand prompt coverage deliberately: pull real questions from sales calls, support tickets, and review sites, not just marketing guesses about what people ask. A high share of voice on five prompts you happened to choose tells you much less than a moderate share of voice across eighty prompts that reflect how buyers actually search.

  • Branded prompts: questions naming your product directly
  • Category prompts: "best tools for X," "top X software"
  • Comparison prompts: "X vs Y," "alternatives to X"
  • Use-case and buyer-journey prompts: "how do I solve [problem]"

Competitive Share of Voice and Competitor Overlap

Overall share of voice compares you against every brand that shows up in your tracked prompts, which can include irrelevant or distant players. Competitive share of voice narrows that to a defined set of named, direct competitors — the handful of companies you actually compete against for deals — and recalculates your share against just that set. It's a more actionable number because it isolates the competition that matters.

Competitor overlap measures how often you and a specific competitor are named in the same answer together. High overlap with a competitor means AI engines consistently treat you as part of the same consideration set — you're being compared, which is a reasonable place to be. Low overlap with a company you consider a direct competitor is the more interesting signal: it usually means the models aren't grouping you together at all, which often points to positioning or content that doesn't clearly stake out the same category.

Use competitive share of voice as your primary competitive scorecard, and use overlap data to check whether you're even being considered alongside the right companies before worrying about whether you're winning that comparison.

Visibility Trend Over Time and Cross-Engine Benchmarking

A single measurement of any of these metrics is close to meaningless on its own. AI answers vary run to run even for the identical prompt on the identical engine — generative responses aren't deterministic, so one query can return a different set of named brands than the same query run an hour later. Underlying models also get updated on schedules outside your control, and engines with live retrieval (Perplexity, Google AI Overviews, Copilot) pull from whatever the web looks like that day. A single snapshot captures noise as much as signal.

The fix is treating every metric above as a trend line, not a point-in-time score: repeat measurement on a consistent cadence, across a stable prompt set, and judge movement over weeks and months rather than reacting to any one run. A dip that persists across three consecutive weekly measurements is worth investigating; a dip in a single run usually isn't.

Cross-engine benchmarking matters for the same underlying reason. The same brand can have a strong share of voice on Perplexity, which leans heavily on live web content and citations, and near-zero visibility on Claude or Gemini, which weight training data and different retrieval sources differently. That's not a measurement error — it reflects real differences in how each engine sources its answers. A single blended "AI visibility score" averaged across engines hides this and can mask a serious gap on one platform behind good performance on another. Track every metric in this guide per engine, then look at the blended view only as a summary, never as the number you act on. Tools like MentioningYou exist specifically to automate this — running consistent prompt sets across ChatGPT, Claude, Gemini, Perplexity, Google AI, and Copilot on a schedule, and reporting each metric per engine and as a trend rather than a one-off score.

Frequently asked questions

What's the difference between share of voice and mention rate?

Share of voice is your percentage of all brand mentions across a tracked prompt set, so it's inherently competitive and depends on how many other brands appear. Mention rate is simpler: the percentage of prompt runs where your brand shows up at all, regardless of who else is named. A brand can have a rising mention rate while its share of voice stays flat if the whole category's total mention volume is growing just as fast.

Which AI visibility metric should I prioritize first?

Start with mention rate and prompt coverage before anything else, because every other metric is only as reliable as the prompt set it's measured against. Once you have a representative, broad prompt set tracked consistently, share of voice and citation rate become the two metrics worth watching weekly, with sentiment and position as secondary diagnostics when something changes.

How often should these metrics be measured?

At minimum weekly, on a fixed prompt set and a fixed list of engines, so you can build a trend line rather than reacting to single-run noise. Daily measurement is useful for engines with live retrieval, like Perplexity or Google AI Overviews, where content changes can shift results within a day.

Is a single blended AI visibility score across all engines useful?

It's useful as a summary but dangerous as a decision-making number, because it can average a strong result on one engine with a poor result on another and look acceptable overall. Always keep the per-engine breakdown available and treat the blended score as a headline, not a diagnosis.

What's the difference between citation rate and mention rate?

Mention rate counts every time your brand is named, whether or not the answer links or attributes anything to your site. Citation rate specifically counts answers that include a sourced reference back to your domain. A brand can have a high mention rate from general recognition and a low citation rate because its own content isn't structured or authoritative enough to be pulled in as a source.

Why does mention position matter if my brand is already being mentioned?

Being mentioned confirms the model knows about you, but position affects how much weight a reader gives that mention. A brand named first in a list is typically being framed as the leading or default option, while a brand named last or in a secondary clause is more likely to read as an afterthought, even though both technically count as a mention.