AI Visibility Tracking

See where your brand actually appears in AI answers

Whaily runs your prompts through the AI models your buyers use and records what came back. Share of voice carries a 95 percent confidence interval, visibility reads N/A below five answers, and every number names the model version behind it.

ChatGPT
Claude
Gemini
Perplexity
DeepSeek
Grok

The problem

You cannot see what AI says about you unless someone measures it.

A buyer asks an AI model to compare options in your category. The model answers from what it has read: reviews, forums, comparison pages, your own site. You have no way to see that answer unless you ask the same question yourself, one model and one prompt at a time. By the time a prospect mentions what an AI assistant told them, the pattern behind it has been running for months.

How it works

From sign-up to signal in minutes.

1

Add the questions your buyers actually ask

Type prompts in one at a time, or import up to 500 rows at once. Each prompt gets tags, a country and a language, so you can track the same question across markets. Set a frequency of daily, weekly or monthly, and assign which models run it. Starting from your domain, Whaily proposes a first set of prompts and competitors before anything runs.

2

We call the models through their APIs, on a schedule

Whaily reaches the models through the provider APIs, not the ChatGPT app or Google AI Mode. The consumer apps add a router, a system prompt, search tools and memory that change the answer, so we do not measure them directly. Every number carries the exact model as three parts: provider, family and version. A tracked model is either pinned to that exact version or set to follow the latest release in its channel, and either way you get an email before a model your organisation depends on is retired. Every model runs daily on every plan. You can also run a prompt on demand, and a failed-runs inbox lets you retry anything that did not complete.

3

Read the results, including what we do not know yet

Visibility percent shows the share of tracked answers whose text names you, and switches to N/A below five answers rather than showing a false 100 percent from a single run. Share of voice is the competitive number beside it, your share of every brand naming the models make, and it carries a 95 percent confidence interval under every line. Position is the average place your brand appeared, to one decimal, and it excludes answers where you were absent instead of quietly counting absence as a bad position. Every full answer is stored, and a timeline marks the moments that could explain a move: a prompt changed, a model changed, a competitor was added, or a tracked model was succeeded by a newer version. When someone needs the month in words rather than a dashboard, the Report page reads the same numbers as five plain sentences, how you are doing, where you stand, what moved and what you are doing about it, with a share link at the bottom.

What you get

Everything you need, in one place.

The exact model, not a vague name

Every number is tagged with a three-part model identifier: provider, family and version. A selection is pinned to that exact version or set to follow the latest release, and you get an email before a pinned model is retired.

Visibility percent that admits what it does not know

Below five tracked answers, the visibility percent shows N/A instead of a number. One answer never produces a false 100 percent.

Share of voice with a confidence interval

Share of voice counts namings, not answers: of every time your brand or a tracked competitor is named, the share that is yours. The shares add to 100, so a competitor gaining means somebody losing, and every line on the Standing tab under Competitors carries a 95 percent confidence interval rather than a bare percentage.

Position, averaged honestly

Brand position is where you appeared among the brands named in an answer, shown to one decimal such as #1.8. The average excludes answers where you were absent, so it stays a measure of how you rank, not how often you show up.

The full answer, not a snippet

Every run stores the complete response, the brands it named, the sources it cited, the position and the model version. Click any figure to open the answers behind it, and read each one in full on the Answers tab under Prompts.

Built to run like a tracker, not a one-off report

Add prompts one at a time or import up to 500. Tag them, set a country and language, pause, resume, archive or restore, and run any of them on demand with a short cooldown.

Visibility and share of voice, at a glance

The overview shows visibility percent, share of voice with its interval, the inclusion rate beside it, and a timeline of what changed, so a shift in the numbers has an explanation next to it.

The brand visibility trend: one line per model over 30 days, each with its provider mark and version, climbing from under half of answers to about four in five.

Why it matters

The honest number is more useful than the impressive one.

It would be easy to show a visibility score for every brand from day one, no matter how little data backs it. We do not. Below five tracked answers, visibility shows N/A. Share of voice carries a 95 percent confidence interval, so a swing from 40 percent to 55 percent on ten answers reads as what it is: noisy, not a trend. And wherever your brand sits beside a competitor, both are counted by the same test: a name in the text of the answer. A number that flatters the person paying for it is not a measurement.

The same discipline applies to the model layer. AI answers change as providers ship new versions, and a chart that moves without explanation is not useful. Whaily records the exact provider, family and version behind every answer, marks the timeline when a model succeeds another, and emails you ahead of a retirement, so a shift in your numbers has a cause you can point to instead of a mystery.

This matters because the decisions built on these numbers are real: budget for content, a case for leadership, a reason to chase a source. A dashboard that hides its limits produces confident decisions built on very little. One that shows N/A, shows its interval, and stores the full answer behind every score lets you trust the number enough to act on it, and know when not to.

Every prompt, every model, the full answer

Open a prompt to see the stored answer from each model, which brands it named, which sources it cited, and the position it gave you.

The sample responses panel: each answer with the model that gave it, whether the brand was named, the answer text, the competitors it named, and the pages it cited.

Questions

The short answers.

Which AI models does Whaily track?+
Today the registry includes Claude Haiku 4.5, Sonnet 4.6 and Opus 4.6 from Anthropic; GPT-5.4, GPT-5.5, GPT-5.4 mini and GPT-4o mini from OpenAI; Gemini 2.5 Flash and Gemini 2.5 Pro from Google; Sonar from Perplexity; DeepSeek chat; and Grok 4.6 from xAI. By default, Starter, Pro and Agency run GPT, Claude and Gemini; Enterprise runs GPT, Claude, Gemini and Perplexity; Free runs GPT, Gemini and Perplexity. DeepSeek can be added from Free up once the account owner allows models served from China, and no plan runs it by default. Grok can be added from Starter up, and no plan runs it by default. Which exact versions your plan reaches depends on your tier.
Do you track the ChatGPT app or Google AI Mode directly?+
No. We call the models through their provider APIs, which is not the same surface as the consumer apps. The consumer apps add a router, a system prompt, search tools and memory, so what a logged-in user sees can differ from the API answer we record. Google AI Overviews is available separately, as an opt-in engine through a SERP data provider.
What does N/A mean on the visibility percent?+
It means fewer than five tracked answers exist for that brand and prompt combination. We show N/A rather than a percentage built on one or two runs.
Does every number come with a confidence interval?+
No, and the ones that do not are named rather than left for you to work out. Share of voice carries a 95 percent Wilson interval, the difference between two shares carries a Newcombe one, and a burst carries them on its mention rate, its sentiment shares, its cited domain shares and its median position. Visibility percent, the inclusion rate, average position, net sentiment, top three share, the cited without being named rate, the NCI score and a competitive rating carry none. Our method page lists all fourteen, the formula behind each one, and three places the method is weaker than it looks.
What is the difference between pinned and latest?+
A pinned model keeps running the exact version you selected until you change it. A latest selection follows its channel and moves to a new release automatically. Either way, a weekly sweep checks for models retiring within 30 days and emails owners and admins.
How often do prompts run?+
You set daily, weekly or monthly per prompt, and every model on the prompt runs on that schedule, on every plan. You can also trigger a run on demand, subject to a short cooldown.
What are the timeline annotations for?+
Marks on the trend charts for events that can explain a move: a prompt changed, a model changed, a competitor was added or removed, a model succession, a burst, or a note your team added.
Is there a report I can send to someone who does not use the app?+
Yes. The Report page turns a period into five sentences the product is willing to stand behind, with the figures under each one as evidence: how you are doing, where you stand, what moved, what you are doing about it, and a share link. A change under a point reads as level rather than a win, and a sample too small for a percentage says so. Send the link instead of a screenshot.

Ready to be recommended by AI?

Start free. See your first insights in minutes.