AI Visibility Tools: What to Look For Before You Buy

Most AI visibility tools report a number. The better question is what it could not see. Seven checks from running one, not from reading vendor pages.

AI Visibility Tools: What to Look For Before You Buy

TL;DR: Every AI visibility tool will show you a number. The number is the least interesting thing it produces. What separates a tool you can act on from a dashboard you will stop opening is whether it tells you what it could not see: which probes failed, which engines went quiet, and whether a citation is a real citation or merely a retrieval. Seven checks, drawn from operating a monitor rather than from reading vendor pages.


Contents


Why the Number Is the Wrong Thing to Compare

Tool comparisons usually start with coverage - how many engines, how many prompts, how often. Those are easy to put in a table and they are rarely what goes wrong.

What goes wrong is quieter. A monitor runs on a schedule against systems it does not control, and those systems change without notice: a model is renamed, an API deprecates a parameter, a credential expires. The same problem shapes how you measure AI visibility honestly and what to demand of any tool that reports one. A monitor’s hardest problem is not measuring - it is noticing that it has stopped measuring. Almost every failure in this category produces a result that looks exactly like a healthy one, and the dashboard cannot show you what the dashboard is wrong about.

So the checks below are all versions of one question: when this tool cannot see something, does it say so, or does it hand you a clean number anyway?


Check 1: What Happens When a Probe Fails

Ask the vendor directly: when one engine returns an error, what appears in the report?

There are three possible answers and only one of them is safe.

behaviourwhat you seewhy it matters
reports the failure“58 of 60 measured, 2 unmeasured”you can audit it
drops the rowa slightly smaller samplethe rate goes up
counts it as zeroa worse-looking scorewrong, but at least visible

The middle one is the dangerous one, and it is the most common, because dropping a failed row is the path of least resistance for whoever wrote the code. A shrinking denominator flatters the metric. If a monitor probes four engines and one dies, the surviving three still produce a rate - and that rate is higher than the true one, computed from a smaller sample, with nothing on screen to indicate it.

A citation rate reported without an unmeasured count cannot tell you it is broken. Ask for both numbers, and treat a vendor who only offers the percentage as having answered the question.

This is the single discipline that separates a usable figure from a decorative one, and it is worth reading alongside how to measure AI visibility honestly before you shortlist anything.

The general form is worth keeping beyond tool selection: any rate you rely on should report what it could not measure, because a metric with no unmeasured count fails silently and in your favour.


Check 2: Cited, or Merely Consulted

This is the distinction that separates a serious tool from a generous one.

When a generative engine answers, it typically retrieves a set of documents and then composes an answer citing some of them. Those are two different populations. Your page can be in the retrieved set and absent from the answer - the engine looked at you and did not use you.

Some tools count the first as visibility. It is a much larger number and it is not what your buyer sees. Being in the set of documents the model consulted is not the same as being named in the answer someone reads.

Ask precisely: which field does your visibility figure come from? A vendor who can name the field is measuring something specific, which is the whole basis of useful AI brand monitoring that still means something in six months. A vendor who says “we track mentions” is not answering, and the ambiguity almost always resolves in the more flattering direction.


Check 3: Is the Model Identifier Pinned

A time series only means something if the thing being measured stayed the same.

Generative engines ship new model versions continuously, and most APIs offer a floating alias that silently points at the newest one. A monitor built on a floating alias produces a chart that looks continuous across a model change that invalidates the comparison. Worse, when a model identifier is retired, a monitor pinned to it starts failing - which is fine if it says so, and invisible if it drops the rows.

Ask whether runs record the exact model identifier they used, and whether the report shows it. If two points on a chart were produced by different models, they are not two points on one series; they are one point each on two series.


Check 4: Does It Keep the Raw Answers

Summaries answer the question you thought of when you built the summary.

The most useful findings in this work come from re-reading old runs with a question you did not have at the time - which competitor started appearing, which publisher rose into the citation slot, when exactly the phrasing changed. None of that survives if the tool stores a score and discards the text.

Ask whether you can export the raw responses, and whether they are retained. Keeping the text is what turns a dashboard into ongoing brand monitoring you can interrogate later. A tool that keeps only its own summary has decided in advance what you are allowed to learn.


Check 5: Can It Separate the Four Causes of a Change

When a number moves, there are four candidates:

  1. The engine changed - a new model, or different retrieval behaviour.
  2. The sources changed - a widely-cited publisher updated an article and took the slot.
  3. Your site changed - the only one you control.
  4. Nothing changed - ordinary variance in a probabilistic system.

The first and fourth are the most common and the least actionable. A tool that reports movement without helping you separate these is generating work rather than insight, because the default interpretation of any drop is “we did something wrong” and most of the time that is false.

The practical test: does the tool let you re-run the same probe set on demand? A single run is one draw from a distribution. A change that does not survive a second sample was never a change.


Check 6: Does the Question Set Hold Still

A monitor that adds prompts as it goes cannot compare this month to last.

This sounds obvious and it is routinely violated, usually for a good reason - someone adds three new questions because a new product launched. The series breaks quietly: the denominator changed, so the rate moved, and nobody logged why.

Ask how the question set is versioned, and whether a report states which version produced it. A tool that treats its prompt list as configuration rather than as part of the measurement is going to hand you comparisons that are not comparable.


Check 7: Who Does It Say the Sources Were

The most actionable field in this entire category is not your own score. It is the list of domains the engine drew from when it answered.

That list tells you what the model considers authoritative on your topic, and it is frequently uncomfortable: for many commercial queries the sources are forums, encyclopedias, established news outlets and review sites - not vendor pages. If no company in your category appears in that list, no amount of on-site optimisation will put you there, because the model is not looking at company pages for that question.

This single field changes strategy more than any score does. It tells you whether the job is to improve your page or to become the sort of thing the sources cite - which is the actual work of generative engine optimization rather than of tooling. A tool that does not expose it is withholding the part you would actually act on.


Key Takeaways

  • Compare tools on what they admit they cannot see, not on how many engines they list. Coverage is easy to advertise; honesty about failure is not.
  • Demand an unmeasured count alongside any rate. A percentage with an invisible denominator fails silently and always in the flattering direction.
  • Make the vendor name the field. Cited in the answer and present in the retrieved set are different numbers, and only one of them is what a buyer sees.
  • A series requires a fixed model identifier and a fixed question set. Without both, a chart is a picture of your configuration changing.
  • The source list is the most actionable output. It tells you whether the work is on your page or somewhere else entirely, and it is where GEO actually starts.

If you are evaluating tools, the fastest way to sort them is to ask what appears in the report when an engine fails. The answer takes thirty seconds and tells you most of what you need to know about everything else the tool reports.

Frequently Asked Questions

What is an AI search visibility tool?

An AI search visibility tool asks generative engines the questions your buyers ask and records how your brand appears in the answers. The useful ones record four distinct things per run: whether your domain was cited as a source, whether your brand was named in the answer text, which competitors appeared, and which domains the engine actually drew from. The weak ones collapse all of that into a single score, which is easy to read and impossible to act on, because a score cannot tell you which of those four things moved.

How do you evaluate an AI visibility tool before buying it?

Ask what it does when a probe fails. A tool that cannot answer that question is reporting a number it cannot vouch for. Every monitor eventually hits a dead engine, a renamed model, or an expired credential, and the tool has three options: report the failure, drop the row silently, or count it as a zero. Only the first is honest, and the difference is invisible in the dashboard, because a silently dropped probe and a probe that found nothing produce the same clean-looking result.

Why does a citation rate need an unmeasured count?

Because a rate with no unmeasured count cannot tell you it is broken, and it fails in the direction that flatters you. If one engine stops responding and its rows quietly vanish, the denominator shrinks and the rate goes up. Nothing in the number looks wrong. A tool that reports fifty-eight of sixty probes measured, two unmeasured, is giving you something you can audit; a tool that reports a percentage alone is asking you to trust a denominator you cannot see.

What is the difference between being cited and appearing in the sources an engine consulted?

They are different fields and conflating them overstates your position. An engine may retrieve your page while composing an answer without citing it, and some tools count that retrieval as visibility. Being in the set of documents the model consulted is not the same as being named in the answer a buyer reads. Ask any vendor which field their number comes from, and treat a tool that cannot answer precisely as reporting the more flattering of the two.

How often should an AI visibility tool run its probes?

Regularly enough that a change is still attributable, and on an interval that never varies. Weekly suits most brands. What matters more than the frequency is that the question set, the engines, and the model identifiers stay fixed between runs, because a series is only a series if the thing being measured stayed the same. A tool that silently upgrades to a newer model version has ended your old series and started a new one without telling you.

Do AI visibility tools replace rank tracking?

No, they answer a different question. Rank tracking tells you where a page sits in a list of links. AI visibility tells you whether a machine repeats what you said when someone asks it a question, which can be true while you rank nowhere and false while you rank first. Most businesses need both for a while, because buyers still use both surfaces, and the two rarely move together.

About the Author

Kaxo CTO leads AI infrastructure development and autonomous agent deployment for Canadian businesses. Specializes in self-hosted AI security, multi-agent orchestration, and production automation systems. Based in Ontario, Canada.

Written by
Kaxo CTO
Last Updated: August 28, 2026
Back to Insights