The short answer
An AI visibility score is the percentage of eligible AI-generated responses in a defined sample that mention your brand. Calculate it as responses mentioning your brand divided by all eligible responses, multiplied by 100. Report recommendation rate, citation coverage, and competitive share of voice beside it instead of blending unlike signals into one opaque score.
The phrase "AI visibility score" sounds precise. Most scores are not. A number such as 64/100 is meaningless unless a reader can see which prompts were tested, which answer engines were included, how many times each prompt ran, and what counted as a mention.
That is not a minor reporting detail. Generative answers can change across engines, runs, prompt wording, and time. Recent research argues that AI visibility should be treated as a sampled distribution rather than a fixed rank. The practical response is not to abandon measurement; it is to make the method reproducible and the uncertainty visible.
Definition worth saving
AI visibility score measures how often a brand appears in a controlled sample of AI answers. It is a sample estimate, not a universal ranking of the brand across every conversation.
The formula
Start with a score anyone can reproduce
Count each eligible response once. If the brand is named five times inside one answer, that is still one visible response. If the engine fails, refuses the query, or returns no usable answer, define whether that run is excluded before the project begins and apply the rule consistently.
Visibility score =responses mentioning your brandeligible responsesx 100
Worked example
A SaaS team tests 20 buyer intents across four AI engines. After repeats and removing failed runs, it has 120 eligible responses. The brand appears in 42.
42 / 120 x 100 = 35.0 AI visibility score
The correct conclusion is: "Our brand appeared in 35% of this defined sample." It is not: "Our brand owns 35% of ChatGPT."
Keep the supporting metrics separate
Visibility, recommendation, and citation are different events. Combining them with arbitrary weights can make the dashboard look polished while making the result harder to audit. Show the core score first, then the diagnostic metrics that explain it.
| Metric | Formula | What it answers |
|---|---|---|
| Visibility score | Responses mentioning your brand / eligible responses | How often the brand enters the answer |
| Recommendation rate | Responses recommending your brand / eligible responses | How often the answer actively selects the brand |
| Citation coverage | Responses citing your domain / eligible responses | How often owned content supports an answer |
| AI share of voice | Your brand mentions / all tracked brand mentions | Your competitive presence inside the sample |
Use your numbers
Free AI visibility score calculator
Enter counts from one reporting period. The calculator keeps the core score and diagnostic rates separate and adds a directional 95% interval, so a small sample does not pretend to be exact.
AI visibility score calculator
Use one consistent prompt set and reporting period.
Core score
Your brand appeared in 42 of 120 eligible responses.
Approximate 95% sampling interval: 27.1%-43.9%. Treat it as directional because repeated AI answers are not perfectly independent.
The interval uses the Wilson method for a proportion. AI runs can be correlated by prompt, engine, and time, so use it as a warning about sample size rather than a guarantee of statistical independence.
Measurement design
Build a sample that represents buying decisions
The denominator shapes the score. Fifty generic prompts can produce a more stable number and a less useful business signal than 20 carefully chosen buyer intents. Start from decisions your market actually makes, then represent each intent with realistic wording and constraints.
A practical starter design
- Choose 15 to 25 buyer intents across discovery, comparison, validation, and implementation.
- Include the AI engines your audience actually uses; do not average an engine into the score just because it is available.
- Run priority prompts more than once and keep the exact wording stable for trend reporting.
- Record engine, model or mode, market, date, complete answer, brands named, recommendations, citations, and cited URLs.
- Keep a stable core set for the trend and a separate discovery set for emerging prompts.
- Publish sample size and coverage beside every score.
Start with the reporting template
Prompt-level columns, classifications, source URLs, and review notes.
Why one run is not enough
The paper Don't Measure Once argues that answers vary across runs, prompts, and time. A 2026 ACL study comparing Google organic search with five generative systems also found substantial variation in retrieval footprints and stability. That means a single favorable screenshot is evidence that an outcome can occur, not evidence of a dependable visibility level.
For weekly operations, report the current period and a four-week rolling average. Investigate a movement only when it is larger than the normal week-to-week noise and appears in the relevant prompt cluster, not merely in the blended total.
Do not change the prompt set quietly. If the sample changes, publish a new baseline. Otherwise a stronger score may simply reflect easier prompts.
Diagnosis
Read the pattern, not only the headline number
There is no universal "good" AI visibility score. A score depends on category maturity, prompt difficulty, brand size, market, engine mix, and whether the sample contains informational or decision-stage questions. Benchmark against named competitors on the same sample and against your own baseline.
Low visibility + low citation
The brand and its evidence are rarely retrieved.
Next check: Check crawl access, category clarity, high-intent coverage, and third-party presence.
High visibility + low recommendation
The brand is known but not selected for the buyer's need.
Next check: Inspect positioning, fit, proof, objections, pricing, and comparison context.
High recommendation + low owned citation
Others may be defining the brand for you.
Next check: Strengthen canonical product, use-case, comparison, and evidence pages.
Owned pages cited + brand absent
Your content is useful, but the answer does not connect it to the product.
Next check: Make the entity, product category, author, and relevant offer explicit on the cited page.
Large weekly swings
The sample is too small, prompts changed, or the engine is unstable.
Next check: Freeze the test, repeat runs, segment by engine, and report a rolling average.
Visibility is more than citation count
The foundational GEO research presented at KDD 2024 explains why a traditional rank is a poor fit for generated answers: sources can appear at different positions, lengths, and levels of influence inside one response. Newer work separates citation selection from citation absorption - whether a source is selected and whether its evidence actually shapes the generated answer.
For a brand team, that leads to a clean reporting hierarchy: visibility tells you whether you entered the answer; recommendation tells you whether you won the decision; citation tells you which evidence supported the output; accuracy tells you whether the description helps or harms the buyer.
From score to work
Turn score gaps into an AEO work queue
A score is useful only when it narrows the next action. Review low-performing prompts with their complete answers and cited sources, then classify the reason before creating content.
- 01Access. Can OAI-SearchBot and other relevant crawlers fetch the page without a robots rule, authentication wall, WAF block, or noindex directive?
- 02Relevance. Does the page state the category, audience, use case, constraints, and product facts in language that directly answers the prompt?
- 03Evidence. Are claims supported by current product details, customer proof, original data, expert sources, and transparent comparisons?
- 04Corroboration. Which third-party sources are cited when competitors win, and is your brand accurately represented in those source types?
- 05Consistency. Do your site, profiles, listings, pricing, integrations, and positioning agree across the public web?
- 06Verification. After a change is discoverable, does the affected prompt cluster improve across repeated runs and more than one period?
OpenAI's publisher guidance confirms one clear technical prerequisite: public pages need to allow OAI-SearchBot to be eligible for inclusion in ChatGPT search summaries and snippets. OpenAI also says there is no way to guarantee top placement. Crawl access opens the door; it does not manufacture relevance or trust.
Measure what buyers see
From one score to the prompts behind it
BotSeen tracks buyer prompts across AI engines, shows the competitors and sources winning each answer, and turns visibility gaps into ranked actions your team can ship.
See your AI visibilityRelated guide
How to track competitor mentions in ChatGPTBuild the prompt set, preserve the answers, and turn competitor wins into actions.FAQ
AI visibility score questions
What is an AI visibility score?
An AI visibility score is the percentage of eligible AI-generated responses in a defined sample that mention your brand. A reproducible score always documents the prompt set, engines, repetitions, market, date range, and exact denominator.
How do you calculate an AI visibility score?
Divide the number of eligible responses that mention your brand by the total number of eligible responses, then multiply by 100. If a brand appears in 42 of 120 responses, its AI visibility score is 35.0.
What is a good AI visibility score?
There is no universal good score because results depend on the category, prompt set, engine mix, market, and sampling method. Compare your score with named competitors on the same sample and with your own baseline over time.
Is AI visibility score the same as AI share of voice?
No. Visibility score measures the percentage of sampled responses in which your brand appears. AI share of voice measures your brand's mentions as a percentage of all tracked brand mentions, including competitors.
How many prompts are needed to measure AI visibility?
Start with 15 to 25 buyer intents, cover several buying stages, test the engines your audience uses, and repeat important prompts. The right sample is one your team can rerun consistently; publish the count and an uncertainty range instead of presenting it as complete market coverage.
How often should AI visibility be measured?
Weekly measurement is practical for active categories, while a monthly executive report reduces noise. Keep the core prompt set stable and compare rolling periods before attributing movement to a specific change.
Primary sources
Research behind this guide
- 1. OpenAI: Publishers and Developers FAQ - crawler access, inclusion, citations, and ChatGPT referral tracking.
- 2. Aggarwal et al.: GEO - Generative Engine Optimization - the KDD 2024 framework for multi-dimensional generative-engine visibility.
- 3. Schulte et al.: Don't Measure Once - why AI visibility needs repeated measurement.
- 4. Kirsten et al.: Characterizing Web Search in the Age of Generative AI - retrieval, source diversity, and stability across generative systems.
- 5. Sielinski: Quantifying Uncertainty in AI Visibility - repeated samples, noise floors, and uncertainty in citation visibility.