In traditional SEO, we monitor keyword rankings and backlinks as proxies for visibility and authority. New tools are beginning to track AI visibility, but it's still early.
Building a tool that goes beyond brand sentiment and gives you data you can use to improve mentions and AI search visibility is harder than it sounds. Most solutions available today give you only a partial picture of your presence in AI models.
The limits of current tracking tools
I've researched more than two dozen AI search ranking tools, including Visibility Kit, SEMrush's AI Toolkit, Ahrefs Brand Radar, Profound, FalconRank, Geostar, Writesonic, ZipTie, Nightwatch, AthenaHQ, Otterly, the Hubspot AI share of voice tool, and SE Ranking. Most cost a lot, and some cost an exorbitant amount for what they do. Even proven tools like SEMrush, which we use for 100+ agency clients and internal projects, are still early in their AI tracking features.
I've used traditional rank trackers for more than 15 years. The gap between what we could track in SEO and what we can track with large language models (LLMs) is significant. My conclusion is that LLMs' inconsistent, nondeterministic results present a greater challenge than the tools' pricing or polish. Their answers are based on probabilities and can change with training data, prompt phrasing, the context they receive, and model versioning.
So what does that mean for tracking?
- A keyword has no fixed "ranking" in an LLM. Each result is a sampled snapshot, and there is no stable leaderboard.
- The wording of a prompt matters. A small change can completely alter the answer and which brands it mentions.
- You don't know the original prompt. With Google, you can target specific queries, while with LLMs you can't see what users asked. You have to guess what users might have asked and hope your brand appears.
Current LLM tracking tools can give you a rough sense of your AI search visibility. Some results may reflect citations the model hallucinated. Those citations come from the model and aren't a problem with the tracking tools themselves. Tracking tools monitor a system that builds responses dynamically and can't show you the original question.
What’s actually being measured?
Take Visibility Kit's tracking feature as an example. You can enter keywords to see how your brand performs in outputs from frontier models. It reports metrics such as:
- Average rank: Your average position across all generated outputs tied to the prompts you track
- Rank distribution: How often your brand appears in different visibility tiers
- Keyword movement: How your visibility shifts over time
The rankings come from simulated prompts. They show how your brand appears across a fixed set of test queries that approximate user behavior. They don't measure what actual users ask.
Tracking like this gives you valuable directional data, especially over time. The visibility it reports is an estimate, and you won't have access to the original prompts. You also can't tell how often users were exposed to your brand in a real-world setting.
Impressions ≠ influence
Impressions from AI tools that pass referrer data are one of the few true signals you can reliably measure. When someone clicks through to your site from an AI-generated response, that visit shows up in your analytics.
In GA4, impressions and clicks give you a simple framework:
-
Impressions = Visibility
How often your brand is referenced or cited in AI responses -
Clicks = Action
Whether those mentions bring someone to your site
If you're getting plenty of impressions and no clicks, you may have a positioning issue. Your brand appears in AI responses without being compelling enough for someone to act on. If you're getting neither, your business is likely absent from what the AI model knows.
The signal isn't perfect. It's one of the few signals that ties visibility to behavior and gives you something to act on.
Why the model matters
AI search tools can handle the same question in different ways:
- ChatGPT with browsing works differently from models that rely only on internal knowledge.
- Claude leans on Brave, ChatGPT on Bing, and Perplexity uses a blend.
- Some AI models draw on retrieved information, while others invent citations outright.
Content and links are part of your visibility in AI search. How an AI model is built, what data it can access, and how it decides to surface information also affect whether you show up.
With AI-powered tools like Perplexity, you can sometimes trace your presence by checking which source URLs the tool cites. With pure LLM outputs like ChatGPT 4o-mini without browsing, you have no such trail. You can't know why you were mentioned or left out, and the answer may change the next time you ask the same question.
What’s missing from current LLM tracking capabilities
Current LLM tracking tools can't show you what users are asking or how a model interpreted their prompts. That limit comes from how LLMs work.
The responses can change with prompt wording, timing, and model updates. A tool can record whether your brand appears in responses to test prompts. It can't tell you the deeper intent or sentiment behind a response unless the response states it outright.
Most tools look for direct citations. That works for some brands, though not all, and nuanced references or indirect mentions may be missed. Only a few platforms pass referrer data that lets you tie AI-generated impressions to on-site behavior.
The data still gives you a sense of direction. You can track how often your brand appears across test queries, see whether your visibility is rising or falling, and identify the prompts associated with your brand. You won't see real prompts, real users, or real intent.
A potential way forward
You can triangulate visibility by checking several signals together:
- Check AI responses to your questions on a schedule to see whether your business appears more or less often.
- Track visits from AI tools that share where the traffic came from.
- Check branded search queries and on-site behavior to see what users are asking.
- Improve citations and mentions on other sites that AI models might use.
We're seeing a rush of visibility dashboards from established tools like SEMrush and new VC-backed startups. Some offer overlays for search results, and some focus on simulating prompts. Others monitor AI citations across articles and knowledge panels.
For now, I use my own lightweight dashboard to track visibility across LLMs with a defined set of prompts. It shows us what's being cited, what's getting clicked, and what's trending up or down. It doesn't try to reverse-engineer the models, and it helps us observe what's happening so we can adjust accordingly.


