To measure whether ChatGPT and Perplexity cite your product, you run three layers of tracking. Build a panel of the prompts your buyers actually ask, run them against each engine on a schedule, and record whether you get named and what the model says about you. Compare that citation rate to your competitors to get share of voice. Then track referral traffic from AI engines separately, knowing it undercounts badly, and tie the whole thing to pipeline instead of watching a vanity score. My GEO post covered how to get cited. This one is about telling whether any of that actually landed.
How do I actually check if ChatGPT and Perplexity name my product?
Run a prompt panel: a fixed set of buyer questions you ask the engines on a repeating schedule, logging the answers each time.
Start with the questions a prospect asks right before they choose a tool. “Best CI tool for a small team,” “X vs Y for production,” “how do I do [job] without [pain].” Aim for twenty to fifty prompts that map to real buying intent in your category. These are the same sub-questions you optimized your pages against, so the panel doubles as a scorecard for that work.
For each prompt, record four things: whether you were named at all, where you appeared in the answer, which source the model cited for the claim, and how it described you. That last one matters more than people expect. A model that names you but calls you the wrong category, or repeats a two-year-old pricing detail, is a visibility problem and a positioning problem at once.
Run the panel on every engine your buyers use, because the answers diverge. ChatGPT, Perplexity, Google’s AI Overviews, and Gemini pull from different sources and weight them differently, so a prompt that names you in Perplexity can omit you entirely in AI Overviews. Test each one rather than assuming a single result generalizes.
Rerun on a schedule, weekly or monthly, from a clean session with no personalization or memory. Answers drift as models update and as new sources get published, so a single snapshot tells you almost nothing. What you are after is the trend over time.
What is share of voice in AI search, and how do I measure it?
Share of voice is the percentage of your category prompts where you get named, measured against the competitors named in the same answers.
Calculate it straight from the panel. If you run forty prompts and get cited in twelve, your citation rate is 30 percent. Now log which competitors show up across those same forty answers and how often, and you can rank the whole set. That ranking is the number worth reporting. Getting cited in a third of answers is nothing to brag about when the leader hits 80 percent, but it is a solid position when the leader is stuck at 35.
Share of voice also tells you where to aim next. The prompts where a competitor gets named and you do not are your highest-value targets: proof that the question has a citable answer and that the answer is currently someone else’s. Those gaps are a content and off-page roadmap, prioritized by real buyer intent.
How do I track referral traffic from ChatGPT and Perplexity?
Track it, but treat it as directional, because standard analytics undercount AI referrals by a wide margin.
In GA4 or your analytics of choice, build a channel or segment that catches sessions referred from chatgpt.com, perplexity.ai, gemini.google.com, and the other engines, then watch it over time. Some engines append tracking parameters to outbound links, which helps. The problem is what gets missed. AI mobile apps often pass no referrer when a user taps a link, and people constantly copy a URL out of an answer and paste it into a new tab, which strips the referrer entirely. Both behaviors dump the session into your “Direct” bucket, so your labeled AI traffic is a floor, not a real count.
The scale here is easy to misread in both directions. AI search engines sent just 0.29 percent of total US website traffic in 2026 (SE Ranking, 2026 study), so if you are waiting for AI referrals to rival Google sessions, you will wait a long time. The catch is that the number is undercounted at the source and growing fast, and the visitors who do arrive tend to be late in their research and high intent. It makes for a lousy vanity metric but a genuinely useful trend to watch.
One more nuance worth knowing: the referral traffic you can see is dominated by one engine. ChatGPT accounted for 92.4 percent of trackable large language model referral traffic across 6.77 million sessions from 166 GA4 properties (Previsible AI Traffic Study, via Search Engine Land). That concentration should shape where you spend panel and content effort first, while you keep an eye on the engines your specific buyers favor.
What tools measure AI visibility?
A whole category of AI visibility platforms now runs prompt panels for you, tracks citations across engines, and calculates share of voice. Profound, Peec AI, and Otterly are three of the more established names, and the money behind the category is real: Profound raised a $96 million Series C at a $1 billion valuation (Profound, 2025).
These tools save real time once your panel is large and you are tracking several competitors across four engines weekly. What they cannot do is decide which prompts matter for your buyers or judge whether the model’s description of you is right. That part is yours.
Here is how the three measurement layers compare, whether you run them by hand or buy a tool for them.
| Layer | What it answers | How to measure | Honest limits |
|---|---|---|---|
| Prompt panel | Are we named, and what is said about us? | Fixed prompt set, rerun on a schedule, answers logged | Answers drift; must rerun and read the descriptions, not just the yes/no |
| Share of voice | How do we rank against competitors? | Citation rate across the same panel, competitors logged per answer | Only as good as your prompt selection |
| Referral traffic | Are cited visitors reaching us? | AI-source segment in analytics, watched over time | Undercounts heavily; a floor, not a true count |
What can’t you measure yet?
Plenty, and it is worth being straight about it.
You cannot see how often your category is queried inside an engine, the way a keyword volume tool shows search demand. You are measuring your presence in answers, not the size of the audience asking. You also cannot fully attribute a closed deal to an AI citation, because the buyer who read your name in ChatGPT in March often arrives as “Direct” traffic in June and books a demo with no visible trail. And you cannot control the description: you can influence what the model says by fixing your docs, reviews, and third-party mentions, but you do not get a dial for it. Measure what is real, and resist the urge to invent precision that the data does not support.
How do I tie AI visibility to pipeline instead of a vanity metric?
Report AI visibility next to the downstream numbers, on one line, so it never floats free as a citation score nobody can connect to revenue.
Track four things together and review them on the same cadence: share of voice on your buyer prompts, AI referral sessions (with the undercount caveat noted every time), demo requests or signups that self-report AI as their source, and closed pipeline traced back to that content. Add a “how did you hear about us” field to your demo form and read the answers, because self-reported attribution is often the cleanest signal you will get when the referrer data is missing.
The point of all of this is a single question a founder can answer in a board meeting: when our buyers ask the model about our category, are we in the answer, and is that showing up in pipeline. If the honest answer is no, you know exactly which prompts to go win. If you want a partner to build the panel and run the program, that is what I do at Rare Bird Lab, and the results from a full engagement are in the case studies.