AI visibility tracking, also called LLM tracking, means asking AI tools such as ChatGPT, Gemini and Perplexity the questions your customers ask, over and over, and counting how often your business is named. We call that count visibility, which is how often you're named across repeated runs and several phrasings of the same question, reported with a band that shows how far it moves on its own. From June 2026 our tracker at Complete SEO ran each question up to five times per engine in one scan, and in August our own repeat study showed that those runs were measuring one sitting while the number that matters moves between sittings. We rebuilt the tracker in September around daily single runs, clusters of phrasings and a noise band.
Why one check can't tell you much
One check tells you little because the same question rarely gets the same list twice. In our three-day repeat study, we collected 12,120 ChatGPT answers to 46 questions through DataForSEO, a commercial scraper of the ChatGPT app, in August 2026, and two runs of the same local question a few minutes apart named the identical set of businesses 3.6% of the time, across 188,520 pairs. Each of the 40 local questions surfaced about 28 businesses that appeared in at least 5% of its runs, the median one appeared in 30%, and only 36 of 1,115 business-and-question pairs appeared in 95% or more of runs. For a business ChatGPT named about half the time, a block of five back-to-back runs found it one to four times in 89% of 6,624 blocks, never in 6% and all five times in 5%.

The number also moves between sittings, and before any correction a business's rate over 40 runs changed by 10 points or more from one burst to the next, 12 or 24 hours later, in 40% of 5,575 burst-to-burst changes across 1,115 business-and-question pairs. Two scans of the same untouched questions had to differ by about 28 points (28.3 after correcting for extraction errors) before the difference cleared that movement, however many runs each scan held. Part of it came from a change in the response format ChatGPT served the scraper on the first afternoon, and across the bursts after that change two scans still had to differ by about 14 points. The study covered 72 midweek hours, so its figures are a floor on how much a month can move.
A 2026 paper found the same instability in the sources three AI search engines cited for consumer-product questions, and concluded that "single-run visibility metrics provide a misleadingly precise picture of domain performance in generative search." Our guide covers why the answers change at all.
What our old tracker got wrong
Our old tracker ran each question up to five times per engine in one scan, and five runs in one sitting mostly measure that sitting. The repeat study found that runs inside a sitting behave like independent draws around that sitting's own rate, so repeats sharpen the read of that sitting and say little about the next one. For a business named in 5% to 20% of runs, the most common range, a five-run rate couldn't detect a change smaller than about 40 points. Extra repeats stopped helping after about 65 runs for those businesses, and after about 10 for one named 50% to 80% of the time, because the noise left over was the answer itself moving.
Adding phrasings to the same scan doesn't help against drift that moves every phrasing at once, the way the format change did on the study's first afternoon (our reading). If the drift is shared that way, the study's main reading says no number of phrasings in one sitting could detect a 10-point change, and its pre-registered reading needed a cluster of at least 21. Our dashboard's headline also marked the change across its chart as up or down, with no band around it. We changed the method in September 2026, moving three pilot clients to one run per phrasing per engine per day on Sept. 13 and most of our other clients on Sept. 22.
What to track
Track clusters of phrasings, one cluster for each service and place that brings in real work (our reading). A cluster is one thing a customer wants, such as a plumber for a slow drain, asked several ways and scored as one number, and its location comes from the town its phrasings name.
One phrasing measures one door into the answer. In our census of 44,988 local answers, changing only the wording took the number of businesses named from 3.4 on "list three" to 5.8 on the plain question and 10.6 on a paragraph asking for a locally owned company, across 3,915 answers of each phrasing. In one pilot client's cluster, asked once a day from Sept. 13 to 27, 2026, the locally owned phrasing named the client in 12 of its 15 ChatGPT answers and the other five phrasings in 2 or 3 answers each, 23 of 89 in all, while another of its clusters named the client in all 87 of its answers over the same 15 days (pilot data). A tracker watching only one of those six phrasings would have reported anything from 13% to 80% for the same client and the same service.
We're using four to six phrasings per cluster in the pilot, and the right count will come from the bands we're measuring. They're spread across the doors that change who ChatGPT names:
- A terse search, like "[trade] [town] [state]".
- The plain question, "Who's the best [trade] in [town]?"
- A few first-person sentences describing a real job and asking who to call, with no "best" in them.
- A request for locally owned companies.
- "List three [trade] companies in [town]."
- The specific service, when the cluster is about one.
Each phrasing names one town, so a question about two towns becomes two phrasings. Keep your own name and your competitors' names out of them, since a business named in the question gets a head start in the answer (our reading). Once a phrasing has a run, freeze its wording, because a reworded phrasing starts a new series, and mark the day the set changes so the chart shows the break.
The town has to be in the words, because a tracker has no location of its own. DataForSEO sends every question from the United States as a whole, and an engine's API sees only the text. A real customer asking for a plumber "near me" can get an answer shaped by where they are, since OpenAI says ChatGPT "may use an approximate location based on your IP address to provide relevant local results."
Build the list from the outside in. For one client, the owner's own list, a read of the website and Search Console each turned up real questions the other two missed (one client, unaudited). Questions that only ask for information, such as what a repair costs, usually name no business at all, so they go in a separate cluster that tracks whether your page gets cited (our reading).

Which engines to track, and through what
Track the engines your customers use, each on its own line, and know whether each number comes from the engine's app, a scraper of that app or the engine's API (our reading). We track ChatGPT and Gemini through DataForSEO's scrapes of their consumer apps, Perplexity through its API because DataForSEO doesn't scrape it, and Claude through its API once a week on a line of its own. The engines also disagree about local businesses, as our comparison of how each engine searches shows, so a blended score across them hides which one you're missing from (our reading).
The route changes the number. In two weekly sweeps in late July and early August 2026, we ran the same questions through OpenAI's API on a smaller model and through a scraper of the ChatGPT app, and across 309 matched pairs the app named the business we were tracking 143 times and the API 55. OpenAI documents its API's web search as a developer tool with its own settings, such as whether it "fetches live content or uses only cached/indexed results," so we don't treat the API as a stand-in for the app.
How accurate are AI visibility tracking tools?
It depends on the route a tool takes into the engine, and on ChatGPT the scraper route gets who's on the card of businesses (the short list ChatGPT pins to a map) mostly right and misses much of what surrounds it. There are three routes. A logged-in browser receives everything a customer's browser does, including data the screen doesn't show. A commercial scraper of the ChatGPT app, such as DataForSEO, which much of the industry measures with and our own tracker uses, returns a structured copy of the answer. An API is the developer version of the model, with settings of its own, and we treat it as a separate product from the app.
We compared the first two routes by asking the same 60 local questions on Sept. 8, 2026 in a logged-in ChatGPT and through DataForSEO, with three scraper runs per question.
| What the answer carried | Logged-in ChatGPT | Scraper, same questions |
|---|---|---|
| Businesses on the visible card | 215 | 208 of those 215 (96.7%) |
| Businesses behind Expand | 390 | 356 of those 390 (91.3%) |
| ChatGPT's own search queries | 60 of 60 answers | 9 of 180 responses |
| The search results behind the answer | 60 of 60 answers | 7 of 180 responses |
| Fields on each business record | 57 | 9 |
| Ads on 16 fresh questions | 10 of 16 | 0 of 16 |
The field count comes from 173 records captured on Sept. 23, and the ads from 16 questions asked that day, with the scraper run about five and a half hours after the browser. Among the fields the scraper drops from the record the app receives are where the star rating came from, whether the listing is claimed, a quote-request link and a ranking score. It also returns the visible card and the pool behind Expand as one list, so a scraper's card is the whole pool of about ten, where a customer sees only the first few before Expand, just under four on average across those 57 cards.
Deciding who was named adds its own error. In our census, DataForSEO's own business-name field found 43.1% of the businesses in a 150-answer reference set, against 92.6% for our extractor at 94.1% precision. Daily tracking shows the same gap, since one pilot client was named in 241 ChatGPT answers over the 30 days to Sept. 18 and appeared in that field on 128, and all 113 misses named the client in the text (pilot data). Our tracker counts a business as named when the answer's text names it, uses one of its aliases or cites its website, and it reports presence on the card as a separate number.
Ask any tracker these questions before you trust its numbers, ours at Complete SEO included:
- Does each engine's number come from the app, a scraper or an API?
- How many separate days sit behind a number, and how many answers?
- Is there a band, and does a change inside it show as no change?
- What decides that a business was named, and how often does it miss?
- Does "in the card" mean the few on screen or the whole pool behind Expand?
- Does it report a rank or an average position?
How often to run it
Run each phrasing once a day on each engine and read the results over a rolling 30-day window. Each day is a new sitting, so a month of daily runs gives the window 30 sittings where weekly five-run scans give it four, and only more sittings shrink the drift's share of the band (our reading).
Rotating a cluster's phrasings, one a day, means fewer questions to run each day, but over 30 days a rotated cluster gets about 30 answers per engine where running every phrasing gets 150 to 180, so its band comes out roughly 2 to 2.5 times wider (our estimate, which the pilot is testing).
Track one cluster by hand for four weeks
Write five phrasings for one cluster using the doors above, and name your town in each. Once a day, ask each one in a new temporary chat set to Unpersonalized, so what ChatGPT remembers about you stays out of the answer. Note whether each answer names your business, and leave any answer that fails to load out of the count. That takes a few minutes a day.
After two weeks, divide the answers that named you by all the answers you got, and compare that with the next two weeks as a whole, never with a single day. For a business named 20% to 50% of the time, a gap under about 15 to 18 points between two windows of about 70 answers is noise (our arithmetic, using the repeat study's drift figure). Use the check to find where you're missing.
What a report should show
A report should show visibility over a window, a band around it, and a change only when the change clears the band. The window's visibility is the answers that named you divided by all the answers in it, never an average of daily percentages, so a thin day can't count as much as a full one. Compare it with the previous window of the same length and show a change inside the band in gray, meaning the data can't tell it apart from the answers moving on their own. We hold the change back until the previous window has data on at least half its days.
The band adds two kinds of noise, the sampling error of the answers in the window and the drift between days divided by the number of days.1 In the 30 days to Sept. 18, one pilot client's visibility was 56.2% of 949 answers over 9 days of data, and a change against the previous 30 days would have had to clear about 10 points (9.92) to count (pilot data). Dividing the drift by the number of days assumes each day drifts on its own, and a platform change that lasts for weeks breaks that assumption, which is one reason we're measuring the band on our own daily data (our reading).
For a local business, also report how often you're in ChatGPT's card of local listings, and your listing's rating and review count against the card's median, which was 4.9 stars on 28 reviews across the 276,628 listings in our census. Share of voice is a different number, your share of all the business names in the answers, counted against every business the engines name, and a report should never present one as the other.
Some numbers shouldn't be goals at all. A rank or average position, the headline of an AI rank tracker, has nothing stable to describe when the same set of businesses comes back 3.6% of the time, a one-day change is a single sitting, and a total share of voice or a single composite score blends away the per-question and per-engine gaps you'd act on (our reading).
How to read the gaps
Read a low number by working out what kind of absence it is, since each kind points to a different fix (our reading). These are the questions we put to every client's data at the start of an AI search project with us:
- Does the engine know you? Ask about your business by name in a new chat. If it describes you and never names you for your trade, it knows you and passes you over, and if it can't describe you, it doesn't know you. Treat it as a sorting question, since nothing published confirms a cause.
- Which service, door or town? A total can hide a hole, and one early client's total hid 0 of 12 answers in one of its towns against 6 of 12 in another (June 2026). On a service, missing means nobody owns it yet and displaced means a competitor does, and the first is the easier one to take (our reading).
- Is something you do credited to someone else? Ask "who around here [does the thing you're known for]?" many times and count who gets the credit.
- Are you named while another site gets cited? Then your standing rests on pages you don't control, and the sites cited for your category are your list of where to be listed and what page to build (our reading).
- Is the gap in memory or in live search? Asking with search off and then on can separate the two, but on ChatGPT's local answers the card still appeared on 95.5% of 3,303 answers with the scraper's web search setting off, so "off" isn't pure memory there.
How our dashboard does it
Our dashboard at Complete SEO stores each day's counts for every phrasing and engine, so any window of 7, 30 or 90 days is a sum of small daily rows. Its tracking view leads with visibility pooled across the three daily engines, a tile for each engine and a tile for Claude marked weekly, and each tile's change against the previous window turns gray when it's inside the band. The chart plots the trailing 7-day line with the band shaded and marks events such as our switch to scraping ChatGPT on Aug. 5, 2026 and ChatGPT's format change on Aug. 24.
Under the chart, each cluster is one row that opens into its phrasings as counts such as 18 of 30, and a cell shows a change only when it clears the band. Local clients also get the card view described above, and we're rebuilding the monthly client report to compare a calendar month with the one before, band included. Clients of our AI search work see this same view.

Methods
The numbers on this page rest on our own data. Each note says where one set of them comes from.
- Repeat study. 12,120 ChatGPT answers to 46 questions, 40 of them local, collected through DataForSEO's ChatGPT scraper in seven bursts 12 hours apart from Aug. 24 to 27, 2026. Pre-registered. A business enters a question's pool when it appears in at least 5% of that question's answers. Every across-day figure is a floor on month-to-month movement.
- Census. The pre-registered analysis of 44,988 answers collected Sept. 3 to 7, 2026 through DataForSEO's ChatGPT scraper. The web-search-off arm is 3,303 answers to the plain question with the scraper's web search setting off, and each phrasing comparison uses 3,915 answers on matched trade and metro pairs. Named-business counts use our extractor, with 92.6% recall and 94.1% precision on 150 answers coded blind by two AI coders and read a third time.
- Scraper comparison. The same 60 local questions asked on Sept. 8, 2026 in a logged-in free ChatGPT account with memory off and through DataForSEO, three scraper runs per question. Audited Sept. 8 to 10, 2026.
- Logged-in captures. The 57 fields per business come from 173 business records in 16 conversations captured from a logged-in free ChatGPT account with memory off on Sept. 23, 2026. The ad count compares those 16 questions with the scraper's answers to them.
- Our tracker. Pilot and client figures come from our own tracker and were rebuilt from our database on Sept. 25, 27 and 28, 2026. The 30-day windows to Sept. 18 include weekly runs from before daily tracking began on Sept. 13, and the cluster comparison uses daily runs from Sept. 13 to 27. The app and API comparison comes from our late July and early August sweeps, and the town comparison from one client's June 2026 scan.
- Derived figures. The 13% to 80% range is our arithmetic on the pilot counts. The 15 to 18 point and 10-point thresholds use the change formula in the note below, with the repeat study's figure for movement between sittings until we've measured our own.
- OpenAI's documentation. Checked against its help center and its developer documentation on Sept. 25, 2026.
-
For visibility r over n answers on u days, a change between two windows has to clear 1.96 × √(2 × (r(1 − r)/n + σ²/u)), where σ² is the day-to-day drift. Until we've measured our own, we use the repeat study's 0.0092, a lower bound. ↩︎


