# How to Evaluate an AI Visibility Vendor

> The evidence and questions needed before relying on an AI visibility tool, agency or consultant—and before placing its score among business KPIs.

> 💡 **TL;DR**
>
> - A vendor should be asked to explain what its score means, which answers it relies on, and how prompts are selected and weighted.
> - Evidence is needed that the test relates to the relevant customers. Test conditions, changes and methodological limits also need to be known.
> - A product should be bought for the use the available evidence supports. Without further validation, a benchmark score should not be treated as customer reach or commercial impact.

Suppose an agency offers to raise an AI Visibility Score from 63 to 80. The dashboard is easy to read, the competitors are familiar, and the proposal associates the improvement with reaching more potential customers.

Before discussing the target, the buyer needs to establish what changes when 63 becomes 80. More mentions in selected answers? More recommendations in buying conversations? A brand shown to a larger share of potential customers? Each claim needs different evidence.

An AI visibility tool can be worth buying without measuring all of those outcomes. The purchasing decision depends on whether the evidence it provides supports the decision it is expected to inform.

Repeating a test can make its result more stable without proving that the test represents customers. The vendor meeting should establish what the supplier can demonstrate and what must be assumed.

## What does the visibility promise mean, and how will it be measured?

The vendor should be able to define what its score measures in one sentence.

It may define the score as the proportion of AI answers in a tracked prompt set that mention a brand. That is an understandable benchmark claim. If the promised outcome is a greater share of customer attention, the connection between the definition and the promise needs to be explained.

“Directional” is not enough by itself. Directional toward what, for which audience, and with what evidence connecting score movements to changes in that audience's encounters with the brand?

The intended use should be clarified before the metric is judged. A metric for monitoring answers to selected prompts can provide useful research signals. A score used to evaluate an agency, justify a substantial budget or claim increased customer reach needs stronger evidence.

## The calculation should be examined with the answers behind it

The distinction between mentions, recommendations and prompt weighting is explored in [What Does an AI Visibility Score Measure?](/what-does-an-ai-visibility-score-measure). In a vendor meeting, the question is how those distinctions have been applied in the report being offered.

For the hypothetical score of 63, the supplier should be asked to open several answers from the report and explain how each contributed to the result.

What counts as visible? If the score is a ratio, what is divided by what? Does every answer have the same influence, or are some weighted more heavily? If competitors are part of the calculation, how does changing the comparison set affect the score?

This shows what counts as “visibility” and how those observations become a number. “Our proprietary AI analyses your authority” does not explain the calculation.

A few different answers should be inspected: one positive, one negative and one that is difficult to interpret. Does a brand named in a warning increase visibility? Does a citation to a product manual count as a recommendation? Can several brands receive credit in one answer? Are inclusion among purchase options and the brand recommended at the end of a conversation measured separately?

The vendor does not need to disclose source code to explain this. It needs to disclose enough to show which situations increase the score and how it is calculated. If commercial reasons prevent some disclosure, the vendor needs another credible basis for asking the buyer to trust the result.

## Do the test questions really reflect the customers?

[Do AI Visibility Tests Represent Real Customers?](/ai-visibility-scores-precision-validity) explains the underlying measurement issue.

The supplier should be asked where its test questions, or prompts, came from. Were they drawn from search-engine queries, observed AI questions, direct customer input, consultant selection or AI generation? These sources provide different kinds of evidence.

Relevance to a product does not show that customers ask a question often. It also says nothing about the stage of the buying decision at which it occurs or how much influence it should have on the total score.

The weighting of each question should therefore be clear. Are questions counted equally, weighted by observed frequency, or weighted by business priority? Where search volume is used, is it described as search-engine demand or as an assumption about frequency in AI conversations? The latter interpretation needs evidence connecting search volume and conversational frequency.

Coverage needs the same scrutiny. Which countries, languages, products and customer groups are included? Which buying stages are tested and which are left out? A large database does not necessarily provide enough relevant data about a particular customer group.

Some vendors use customer profiles, or personas. The vendor should explain how those profiles were developed. A detailed account of a person's age, occupation and family life does not show that the person decides like a real customer segment. A useful explanation identifies which behavioural differences the personas represent and what customer observations support them. Demographic detail or an AI-written biography is not enough.

Nor does a longer test automatically resemble a real buying journey. Follow-up questions extend a conversation; evidence is still needed that the conversation reflects customer behaviour. The explanation should distinguish observed behaviour from assumptions. The buyer need not know how the supplier built the system, but should be able to understand why the tested conversation is relevant to its customers.

## The test conditions need to be known

Naming the AI product used in a report is not enough. The conditions in which the result was obtained need to be disclosed.

Does the vendor observe answers in the consumer interface or connect to a model through an API? Which model and product mode are used? Is web search enabled? Does every prompt begin a fresh conversation, or can the model see previous messages? Is information beyond the prompt and prior messages supplied to the model?

Language and country settings need clarification as well. Translating a prompt or selecting a country in a dashboard does not by itself show which local conditions the test represents.

Some settings of an AI product may be unavailable to the vendor. Stating that clearly is different from reporting unavailable settings as if they were known. A report should distinguish known conditions, assumptions and conditions that cannot be observed.

Comparisons between this month and last also need a check on whether the conditions remained the same. The model, questions, scoring rules or prompt weights may have changed. The vendor should record such changes and consider them when explaining a trend.

Without those records, a score increase caused by a measurement change can be mistaken for campaign success. The vendor may not be able to establish every cause with certainty, but the uncertainty should be stated.

## Has the result been compared with real customer behaviour?

If a score is claimed to describe what customers actually see, evidence supporting that claim should be requested.

Has the vendor compared its score with observed customer behaviour? What was observed, for which customer group and period? How closely did score changes match behavioural changes, where did they fail, and is the evidence relevant to the buyer's market or from a materially different group?

If a commercial result such as increased sales is promised, its connection to visibility needs its own evidence. A case study in which both a score and sales increased can be informative. It does not prove that the score increase caused the sales increase; other changes may have affected sales at the same time.

The type of evidence matters. The supplier's own analysis, an independently reviewed study and a satisfied customer's comment do not carry the same weight. Adding a detailed chart does not strengthen the evidence itself.

Where available, first-party signals should also be used as a reality check: UTM-tagged referrals, server logs, CRM records or another agreed source. They will not capture every visit that began with an AI system and may have their own attribution gaps. They can nevertheless help show whether the claimed increase has any observable connection to the business. If no such comparison is possible, the score may still be useful as a controlled benchmark, but the claim should remain limited to the test conditions.

A tool can still be useful when its test result has not been compared with real customer behaviour. It may be bought to monitor selected answers, compare results under defined conditions or identify questions requiring closer research. There may nevertheless be insufficient grounds for presenting the result as a measure of customer reach.

The supplier should be asked directly about its limits: insufficient data in some markets, unobserved personalisation, or no demonstrated relationship between its score and the rate at which real customers see a brand. Clear limits make it easier to decide where the tool can be used.

![A product appears in an AI answer, a person compares two options, and a purchase is shown in three separate scenes.](/images/ai-visibility-series/separate-events.avif)

*Appearing in an answer, entering the set of options considered for purchase and being bought are separate events. Evidence of one does not show that the next has occurred.*

## Did the score improve, or did the calculation change?

Care is needed when the same vendor selects test questions, sells work intended to improve visibility and reports its success through the same score. Changing questions or calculation rules can make progress easier to show. This can happen without an intention to mislead.

It should be clear who can add or remove prompts, change their weights, revise the competitor set or alter scoring rules. Can the timing of changes be seen? When the score rises, can it be checked that the two periods used the same rules?

The same risk applies to the buyer's own team. Repeatedly adding questions on which the brand already performs well can raise the reported score without producing new insight about customers.

Before a performance target is agreed, the change in answers needs to be separated from an increase caused by a changed calculation. It also needs to be clear what evidence would be required to attribute an improvement to the consultant's work. Seeing a higher score does not explain why it rose.

## The intended use of the tool should determine the decision

A purchase decision does not require every uncertainty to be resolved. It does require knowing which gaps matter for the intended use.

If the calculation is unclear, the answers behind it cannot be inspected or measurement changes are not disclosed, that score should not be used to assess a team or vendor until the gaps are resolved. Even where a relationship to customer behaviour has been shown, the result should be interpreted only for the markets and customer groups for which it is valid.

The product may be bought for a defined purpose, tried before broader use, subject to a request for further evidence or declined. The supplier's size and the apparent sophistication of the dashboard should not determine the conclusion.

The points established in the meeting should be kept in a short record: what the supplier claims, the evidence offered, the limits of the measurement and the decisions the score is allowed to inform. When the score later appears in a board presentation, the limits stated in the sales conversation may have been forgotten. The record helps prevent the number being given a meaning it does not measure.

Regularly monitoring answers to selected questions may be a service worth buying on its own. If the score is to prove that a team's or consultant's work has exposed the brand to more customers, a stronger basis is needed. The supplier should explain what the score does and does not show, and that boundary should remain intact when the result is used.

---

Language: English
License: CC BY 4.0
License URL: https://creativecommons.org/licenses/by/4.0/
Scope: Evren Bal-authored text, unless this article expressly states otherwise.
Excluded: Third-party material, quoted excerpts, logos, and separately marked images retain their own rights.
Attribution: Credit Evren Bal, link to the canonical source and license, and indicate changes.
Source: https://evrenbal.com/evaluate-ai-visibility-vendor
