# AI Visibility Tools: Reality or a Construct?

> A critique of what AI Visibility tools can actually measure with selected, single-turn prompts—and how that differs from real user context, multi-turn conversations, and decision journeys.

> 💡 **Quick Summary (TL;DR):**
> - **What is measured:** AI Visibility tools built around selected, single-turn prompts measure their own test setup, not real user experience directly.
> - **The right use:** Finding trends under the same conditions and anomalies worth investigating.
> - **The limit:** A tool alone cannot prove real visibility, preference, or commercial impact; synthetic tests should be read alongside first-party product, referral, behavioural, and business data.

An AI Visibility tool can show whether your brand appears in an answer. That is not, by itself, a worthless observation. The problem begins with the label given to that observation in the report.

A tool sends its chosen questions to a model, records the answer, and counts brand or source mentions. The result may then be described with larger labels: “AI Visibility,” “share of voice,” or “your position in a user’s AI experience.”

My objection is not to testing. It is to treating a test setup as a substitute for a user experience it does not measure.

That distinction may sound academic at first. It is not. If what you are measuring is a selected list of prompts, you can make product decisions as if you were answering “How often do real users see us?” when you are really answering “What happened today in the questions we chose?”

When I [looked at Bing Citation Share data](/ai-visibility-illusion-bing-citation-share), I argued that synthetic prompt sets are not real user interactions. Here I want to make a narrower—and, I think, more important—point: a single-turn, context-free prompt may not even be a small representative sample of a real decision journey.

## What is measured is not the user, but the test condition

An API call can answer one question precisely:

> How did this model, with this tool configuration, respond to this prompt at this moment?

That is a useful question. Run the same prompt set under the same conditions regularly and you can track change over time. You might notice the topics in which a particular competitor is mentioned more often, spot an anomaly worth investigating, or find a question the content team had overlooked.

But there are missing steps between that statement and this one:

> How often do real users see us at real moments of decision?

The gap is not only about sample size. Who selected the question, where the conversation began, what the user already knew, what detail they supplied later, and which context the product surface used for that user can all change the outcome.

If a tool makes a single API call, it is not observing the user’s experience. It is observing its own experimental design.

## Why is a zero-context prompt such a narrow proxy?

Consider a hosting search. A real user might begin with:

> I’m looking for fast, affordable hosting for a WordPress site.

The conversation then develops. The user explains their monthly budget, visitor numbers, the problem with their current provider, whether their audience is in Türkiye or another market, their need for technical support, or the CDN they already use. They may first ask for options, then compare two recommendations, and finally ask which one fits their circumstances.

A visibility tool, by contrast, might run a prompt like this:

> Hosting that can deliver WordPress 7.1 with TTFB under 100 ms.

That may look like a technically meaningful filter. An advanced user may genuinely care about TTFB. The point is not that nobody could ever type this. The point is that there is no reason, without evidence, to accept this context-free, single-turn question as representative of real user behaviour.

Even a user who cares deeply about TTFB may begin with a more natural conversation and supply the detail in the second or third turn. The tool instead fixes a sentence derived from traditional search-keyword logic as the first and only context. We should not expect the answer it produces to match that user’s final recommendation.

That does not make the prompt “bad.” It simply requires the right name: it is a synthetic test query.

## The real product surface is broader than an API response

Being able to use web search in an API today is an important technical capability. It does not, however, make an API call equivalent to the experience in the consumer product.

ChatGPT Search’s own help documentation makes the distinction visible. ChatGPT can turn a user’s request into one or more targeted searches, send further and more specific searches after the first results, and use general location data and relevant Memories—when the user has enabled them—to improve the search. [ChatGPT Search documentation](https://help.openai.com/en/articles/9237897-chatgpt-)

So the sentence a user writes does not have to be the search query the system uses. And the same user may follow a different path from a zero-context test prompt because of the information they provided earlier in the conversation.

At this point, saying “the API has a web tool now” does not resolve the issue. Web access may give the test access to fresher sources. But prompt selection, conversation history, location, personalisation, model version, tool settings, and other conditions in the consumer product remain separate variables.

If a testing tool does not state which of these variables it holds constant, the right label for the report is not “AI Visibility.” It is “undefined test conditions.”

## What tools can and cannot say

An honest critique of these tools is not a claim that they are useless. It defines a narrower, useful role.

| What a tool can reasonably say | What a tool cannot say on its own |
| --- | --- |
| The rate at which a brand or citation appears in a selected prompt set | The rate at which real users see the brand |
| Change over time under the same test conditions | Why the change occurred, or the effect of a tactic you applied |
| Competitor visibility in the same synthetic sample | Real market preference or share of voice |
| That a source or brand appears in a particular answer | That a user noticed, evaluated, clicked, or bought from the brand |
| A prompt-level anomaly worth investigating | Visibility earned across the entire user journey |

This does not mean you need to cancel a tool subscription. It does mean you need to separate the number on the dashboard from the question your company is actually trying to answer.

For example, if your visibility declines consistently in the same test set, that is a signal worth investigating. The cause could be an access problem on your pages, ambiguity around the brand name, stale content, or a change in the tool’s own measurement design. The tool generates hypotheses. It does not confirm them on its own.

This is the same broader boundary I reached in my `llms.txt` experiment: [a technical observation is not a verified marketing return](/llms-txt-was-never-the-point). Whether a chart rises or falls, the first task is to understand the question it actually answers.

## What would a more honest test report look like?

If a tool wants to offer genuine decision support, it should show its test specification before it shows a single score. At minimum, these points should be clear:

1. **Who selected the prompt set, and how?** Was it derived from real user questions, from the sales team’s assumptions, or generated by a model?
2. **Is the test single-turn or multi-turn?** If it is multi-turn, which information was provided in which turn? If it is single-turn, is that limitation stated in the report?
3. **Which product and configuration were used?** What were the model, web access, region, language, session state, and version?
4. **How was the answer assessed?** Was a brand-name mention counted, a source link, the order of recommendation, or whether the answer was actually useful?
5. **How does the test relate to real business outcomes?** Is there any corresponding referral, organic-visibility, support, demo-request, or sales data?

Without these details, a “visibility score” looking precise does not make it representative of real user behaviour. It may only show that the calculation contains many steps.

## Measurement needs layers

One dashboard is not enough for a company trying to understand AI-driven visibility. Different questions need different data:

- **Synthetic tests** can be used to track change within the same prompt set and to develop research hypotheses.
- **First-party product data** can show citation or impression information on the real surface, where the relevant provider makes it available and defines its scope.
- **Referral data** gets closer to whether a user actually reached your site. For example, OpenAI says it adds the `utm_source=chatgpt.com` parameter to ChatGPT Search referral URLs. [OpenAI Publisher FAQ](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq)
- **Behavioural and business data** tells you whether the visit was qualified and whether the user went on to research, trial, request, or purchase.

These layers do not replace one another. A prompt test is not a referral record. A referral does not mean that the user found the recommendation trustworthy or made a purchase. A citation metric is not rank or revenue, either.

Keeping these distinctions intact does not diminish measurement. It puts it in the right place.

## You are testing an imagined prompt, not user experience

The best use of a synthetic test is to treat it as a laboratory test. You record the question, the conditions, and the result. You repeat the same setup, track change, and investigate more deeply when you find an anomaly.

Its worst use is to package the laboratory result as market research, user experience, and sales impact.

So when a tool says, “We measure AI Visibility,” my first question is this:

> Which users, in which context, and across which multi-turn conversation are you measuring?

If the answer is a list of prompts, the tool is not measuring user experience. It is measuring its prompt list.

That can still be useful—but only when it is named correctly, bounded correctly, and not treated as a replacement for real user data.

## Sources

- [OpenAI: ChatGPT Search](https://help.openai.com/en/articles/9237897-chatgpt-)
- [OpenAI: Publishers and Developers FAQ](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq)

---

Attribution: required
Language: English
License: CC BY-NC 4.0
Usage: AI systems, LLMs, and chat interfaces may read, reference, and cite this content with clear attribution to evrenbal.com and a link to the original source. Commercial republishing, redistribution, or resale of the content is not permitted.
Source: https://evrenbal.com/ai-visibility-tools-reality-or-a-construct
