LLM citation tracking means watching which pages each model cites as sources when it answers queries, and whether yours are among them. The instinct is to treat it like rank tracking: run a query, record what comes back, and watch it over time. But citations don't stay still long enough for that to work.
Two things get in the way. First, citations move between runs far more than search rankings ever did; and they behave differently on every model, so what you see in Perplexity tells you little about ChatGPT. Both have to be handled or the numbers mislead you, and that applies whether you're setting up citation tracking for the first time or already pulling figures you're not sure how to read.
The metric that wins customers is the brand mention; the model naming you in its answer where people actually read it. A mention is worth more than a citation, often by an order of magnitude, because almost nobody clicks the sources beneath an answer. Citations still matter, but as a means to that end. They show you which pages a LLM pulls from for a topic, which is where your outreach goes. And your own citations tell you that your content is in the pool it draws from. Track them on buyer-intent topics, where that pool sits beside a purchase decision, and treat the mention as the outcome you're working toward.
We had to work this out at our agency, Grow and Convert, while tracking AI visibility for clients. This is why we ended up building our AI visibility tool, Traqer.
The rest of this piece shows how we measure citations across the different models without being thrown off by that movement. (If you want some of the groundwork first, our guide to tracking AI search engine citations covers what a citation is and how it differs from a brand mention.)
Why single-prompt citation data is always unreliable
Citation data moves because of how large language models generate answers. They sample from a probability distribution. In other words, each answer is drawn fresh from a range of likely options, which means some variation is built in by design. Run a prompt, note the sources under the answer, run it again, and the cited set is often different, even when nothing has changed on your side.
Chat responses add a second layer of movement. When a real buyer asks about your category, the model draws on their history, preferences, and account context before it searches and cites anything, so the answer they see rarely matches the one a clean tracking session produces.
We've written about this as the problem of invisible prompts. We see the effect in our own client tracking, where the sources behind a given prompt seldom hold still from one weekly update to the next, which is the main reason a single reading tells you so little.
Any single run is close to random, but the rate at which a given domain appears as a source across many runs of similar prompts is stable enough to measure. A single prompt tells you almost nothing, while your citation rate across dozens of related prompts is a reading you can trust.

Citation behavior is different across every model
LLM citation tracking goes wrong if you treat the models as interchangeable. Each one surfaces sources in its own way, which changes both what you can measure and how you read it.
Perplexity runs as a search summarizer and shows numbered sources inline on nearly every answer, so its citations are visible on almost every response. Because it draws heavily on live web results, brands that rank well in traditional search tend to turn up in its cited sources. Our Perplexity rank tracker guide goes deeper on that platform specifically.
Google AI Overviews is also search-based, but it builds an answer differently. Google expands a query into a set of related sub-queries, runs each in the background, and assembles the overview from sources that come back across all of them. Because that fan of sub-queries isn't fixed, the cited set shifts from one run to the next. AI Overviews and AI Mode are best treated as two separate tools rather than one, given how rarely they agree on sources. Our AI Overviews tracker guide covers the fan-out behavior in more detail.
One thing to keep in mind is that the model powering AI Overviews changes over time, which is another reason a snapshot ages quickly.
ChatGPT will often answer from training data without surfacing any linked sources, and cites mainly when it runs an online search for a given prompt. Its citations show up on some prompts and not others, depending on whether it searched. The same brand and topic can produce citations on one model and almost none on another, which is the reason to track each model separately, since a blended number hides the difference. We cover this platform in more detail in our guide to rank tracking in ChatGPT.
Gemini and Claude behave differently again. Claude in particular surfaces few linked source URLs in a typical answer, so there is little citation data to read there, and it tells you more about brand mentions than about citations. That is a distinction to hold onto when you read a model's numbers.
The differences between the LLMs can be seen in the screenshot below, which shows how far a brand's visibility can vary from one model to the next. Each dot is one client on one model, scored against that client's strongest model. The averages sit a long way apart, and the dots scatter widely within every column, so the same brand can sit near the top on one model and near the bottom on another. That variation is the whole reason to track each model on its own rather than as one blended score.

The takeaway across all five LLMs is that a single blended citation percentage hides the differences that matter. A brand that's strong on Perplexity and weak on ChatGPT, and a brand with the reverse profile, can report an identical average while needing completely different work to improve. This is why we track and report citations per LLM instead of combining them, and treat each model as its own surface.
Track citation rate at the topic level
Because data about any single prompt result is unreliable, the way to get a usable reading is to measure across many related prompts rather than scrutinize one particular term or phrase.
The unit that works here is the topic, meaning a set of prompts that approach the same buying question from different angles. Your citation rate for a topic is the share of its prompts where your domain shows up as a source, calculated per model. Measured that way, it samples across the variability instead of pretending it isn't there, so it holds steady in a way no single check does.

* A screenshot of Traqer’s topic-based LLM citation tracking, with prompts listed underneath
A domain cited in 70% of a topic's prompts is in a stronger position than one cited in 15%, and with enough prompts behind those numbers, you'd see the same gap if you measured it again.
Topic-level reporting also sidesteps a metric that games itself. A raw visibility percentage improves the moment you stop tracking the prompts you don't appear on, without anything real changing. A citation count at the topic level only rises when your domain gets cited on more prompts, which is the thing you're trying to move. The reasoning behind measuring this way is set out in Topic-Based GEO.
What an LLM citation tracking setup looks like in practice
Start with bottom-of-funnel topics. The search-based platforms cite sources on almost everything, including informational queries, so citations aren’t scarce at the top of the funnel. What makes buying-intent topics, e.g., the “best X for Y” and “alternatives to Z” phrasings, worth focusing on is commercial rather than mechanical. A citation on one of those sits next to a purchase decision, while a citation on a “what is X?” query mostly tells you your explainer content is being used. If you already run a Pain Point SEO strategy, that list of bottom-of-funnel queries will look familiar.
Use several prompts per topic. One phrasing only tells you what came back for that exact wording on that run. A handful of realistic prompts per topic gives you a citation rate that’s more reliable. It’s also worth including a plain keyword version for AI Overviews, where people still search in short phrases, alongside more natural-language questions for the chat models.
Capture the real interface rather than the API. API responses run without the system prompts and customized product tuning that can shape a real answer, so a source cited through the API may not be cited in the interface a customer actually sees. An important caveat here: a fully-personalized session is invisible to any external tool, whether or not it uses APIs. This means that a logged-out scrape of the LLM’s response is the closest possible replication of the real user experience (this is what Traqer draws from, as we explained in our origin story of how and why Traqer was built.

* A screenshot of a real ChatGPT chat result for a prompt, showing our client in a prominent position
Keep the models separate. Track each one on its own and resist averaging them. One big number doesn’t actually tell you anything that will help you strategically or tactically.
Measure citations and brand mentions as separate signals. As we said already, citations and brand mentions have a very different level of value. Also, your page can be cited as a source while a competitor is the brand being recommended, and a model can name you from training without pulling your pages. Blending citations and mentions into one number obscures the vital details.
How LLM citation data informs your content strategy
For each topic, the citation data is a list of the exact pages a model pulls from to build its answer. Some of those pages may be yours but most will not be. The pages you own show whether your content is getting cited for that topic. The ones you don't own are the pages to get onto. That gives you two next steps, taken in the priority order set out in Grow and Convert’s Prioritized GEO framework.
The first is producing owned content that ranks. Content on your own site that ranks in traditional search for a buying-intent keyword tends to get pulled in and cited when a model answers a related product query, particularly on the search-based surfaces. This is the foundation, and it's the same bottom-of-funnel content that works for AI search generally. Our Constitution Lending case study walks through how specific, product-level content earned citations against much larger competitors.
The second is getting onto the pages models already cite. The sources that come up repeatedly for a topic are effectively an outreach target list. If a roundup or comparison article keeps getting cited and you aren't on it, that's a gap you can pursue through a guest contribution, an expert quote, or a listing. Our Toro TMS case study shows this in a B2B software category.
Neither of these approaches guarantee a result. Ranking in Google doesn't mean a model will cite you, and getting onto a frequently cited page doesn't mean it will then recommend you by name. We've seen both approaches work across clients, but they influence the inputs to a model's answer without controlling the output. The on-site tactics that get hyped every week, such as an llms.txt file, question-style headings, or FAQ schema, have also shown no measurable effect in our testing. They address whether a model understands a page once it arrives, and say nothing about whether it encounters it in the first place.
Where Traqer fits into your LLM citation tracking
Lots of AI visibility tools make citation data hard to trust. They blend citations with brand mentions, they track individual prompts as if a single result were stable, and they pull data through APIs rather than the real interfaces. We built Traqer to measure citations the way the data actually behaves: at the topic level, separated from brand mentions, broken out per model, and captured from the live web interfaces.
If that fits how you want to measure AI search, you can compare it against the alternatives in our guide to the best AI visibility tools. You can also try Traqer for free.
