Your brand could be ranking #1 on Google, but still be invisible to AI.
Absent from conversations your customers are having with large language models (LLMs) about your category.
Or worse, showing up inaccurately, with outdated or incorrect information that’ll hurt your sales.
Without prompt tracking, you’d never know.
Prompt tracking (sometimes called LLM visibility tracking) is the practice of monitoring how your brand shows up in AI answers over time, through mentions or citations.
It’s different from traditional SEO rank tracking, which tells you where your URLs appear on search engine results pages (SERPS) for specific keywords.
With rank tracking, you ask, “How close are we to position 1?”
But LLMs don’t answer questions with a static list of 10 blue links.
They pull from massive amounts of information to generate a unique response every time, tailored to the user and the context of the conversation.
You can even ask the same question twice and get two different answers.
In this search experience, it matters less whether your brand is mentioned first, and more that it says true positive things about you to the right people — consistently, across many runs of similar prompts.
But without a system to monitor it, you’re flying blind.
Prompt tracking gives you the directional intelligence to spot AI visibility gaps — the queries you’re consistently not showing up for — and close them.
This guide shows you exactly how. You’ll walk away with a free tracking template, a step-by-step system, and two real-world expert setups you can steal.
Free template: Download our prompt tracking spreadsheet to start understanding your brand’s AI visibility across LLMs ASAP.
Why Brands Need to Track Prompts
According to a study from Orbit Media, 55% of US internet users rely on AI as their primary or frequent research tool. Thirty-two percent use it for product recommendations.
Translation: A growing share of buyers are learning about you in AI tools. Without ever visiting your website.
Overall AI visibility scores tell you whether you’re showing up. Prompt tracking tells you where and how.
It can help you understand:
The types of questions you’re showing up for and where they fall on the customer journey (ToFu, MoFu, or BoFu?)
The questions your competitors are pushing you out of (and your share of voice on important topics)
The sentiment around your brand mentions
The questions you’re getting cited for, but not recommended (sometimes called ghost ranking)
This level of data helps you spot specific trends and gaps in your AI visibility over time.
Then, you can prioritize exactly what to fix.
Take Gong, the sales call intelligence tool.
Their AI visibility score is a respectable 65.
Free tool: Get your own score using Backlinko’s free AI visibility score checker.
They show up for prompts at all three stages of the funnel.
With prompt tracking, Gong can focus on conversations most likely to drive revenue and stop spending time and money tracking ones that won’t.
If their mention rate stays low on key BoFu topics over time, that’s a signal to update on-site or third-party content.
They can also see which relevant topics competitors are owning while they’re absent.
For example, Gong’s Engage product helps with lead generation.
But Salesforce and Hubspot consistently own these prompts.
This tells Gong two things:
Their audience isn’t aware of their lead generation use case
They need more content around it to train AI platforms to mention them
Sentiment prompts like ‘Is Gong worth the price?’ and prompts that surface ghost ranking can reveal similar trends and gaps.
The important thing is to track data over time.
AI answers are non-deterministic. The same prompt can return different brands across different runs.
Regular, repeated tracking is how you get meaningful signals you can act on.
When LLM Prompt Tracking Isn’t Worth It (and When It Is)
Prompt tracking isn’t for everyone.
Plenty of businesses spend time and budget on it and still walk away with data they can’t use.
Prompt tracking is NOT worth it when:
Your audience isn’t using AI to find solutions in your category
Your site isn’t set up to be crawled by AI
You don’t publish content regularly
You’re not looking for competitive or brand narrative insights
You need a single KPI to report upward
You don’t have bandwidth to act on insights
TL;DR: Prompt tracking yields valuable insights — but only if you’re set up to act on them.
If you are set up, the returns can be significant.
Take Gong, the sales call intelligence tool. It’s a great candidate for prompt tracking.
It has a full content engine that publishes across multiple channels (blog, reports, video, media coverage, audio, and more).
They compete in a crowded category where comparison prompts are common.
And they serve a buyer (sales leaders) who increasingly uses AI to evaluate software.
(Seventy-one percent of B2B software buyers now rely on AI chatbots for product research, according to G2, up from 60% in 2025.)
On the other hand, a local HVAC company that gets all its leads from Google Business Profile (GBP) and word-of-mouth doesn’t need prompt tracking. At least not yet.
Even if AI is overlooking them, building a content engine from scratch just to fix that isn’t a realistic investment.
Building Your Prompt Set: What to Include and Why
Don’t try to track every possible prompt your customers could be using.
Instead, focus on prompts you actually care about getting mentioned or cited in.
They should map directly to your product offering, audience pain points, and moments close to purchase.
Types of Prompts to Include
Tracking these four types of prompts over time will yield the most helpful insights:
Evaluation prompts: “Best tool for x use case” and specific feature queries
Reputation prompts: “Is x product worth the price?”
Comparison prompts: “Alternatives to x product,” “x tool vs. x tool,” “best x tools”
Gap prompts: Priority topics your competitors are pushing you out of
The first three are focused on understanding how your product is being recommended in buying conversations.
The last one is about understanding your competitive landscape and where you could improve.
Pro tip: Don’t just track one prompt for each type. Looking at answers for a single prompt is just noise. Reviewing a cluster of prompts over time is a real signal.
Margaret Kapitany, Offsite SEO Lead at Hootsuite, shares how she focuses her prompt set:
The prompts worth tracking are the ones that most closely mirror how a potential buyer would actually ask their AI for help, especially close to a purchase decision. For me at Hootsuite, that means prompts that cover comparison, evaluation, and recommendation queries from a social media manager or CMO, phrased the way they’d talk to a colleague or trusted industry peer.
What that looks like in practice at Hootsuite:
Prompts worth tracking
Prompts not worth tracking
“How does Hootsuite compare to [competitor]?” (Comparison)
“What is social media management?” (Pure definition — won’t convert)
“Which social media management platforms integrate with Salesforce?” (Evaluation)
“Is Hootsuite a good company?” (Vanity — brand mention is baked in, nothing actionable)
Where to Find Prompts Worth Tracking
Find prompts wherever you normally go to learn about your audience.
To find prompts worth tracking, look at:
Keyword research: Commercial and Transactional intent queries like “best x software”
Google’s “People also ask” (PAA) boxes: Comparison and evaluation questions like “Best alternative to x” or “does x integrate with y”
Perplexity’s related questions: Similar to PAA, comparison and evaluation questions
Reddit, Quora, Facebook Groups in your industry: Repeated questions, especially ones that compare options or express frustration
Sales call transcripts: Repeated questions asked right before or during a purchase decision
Semrush prompt suggestions for your brand: Queries tied to buying decisions that you or your competitors are showing up for
Pro tip: For every question you uncover, do a quick gut check: “If someone asked an LLM this, would I want to see my brand show up? Would I be upset if it didn’t?” If the answer is “Yes, and yes,” keep it.
Let’s return to Gong as an example of how to find prompts.
Keyword research shows me the questions people are asking Google about my category, how popular they are, and the language they use.
If I use a tool like Semrush, I can filter to Commercial or Transactional intent keywords (the ones labeled “C” or “T”). And add the most popular and relevant ones to my prompt tracker.
Then, I can dig through Reddit forums my audience frequents to find repeated frustrations and buyer queries.
I would take “Sales enablement tech stack suggestions,” “AI tools for sales enablement.” I’d ignore the queries about SMBs or startups because those aren’t Gong’s target audience.
Next, I’d extract the comparison or evaluative prompts Perplexity surfaces when I prompt it with terms related to my business.
For Gong, I used “best conversational insights tool for sales” to find some good candidates:
Gong vs. Chorus (or any other competitor on this list)
Which conversational insights tool has the best ROI for enterprises?
Pro tip: When writing prompts, don’t agonize over exact wording the way you would with keywords. LLMs cluster semantically similar queries together. Track a few natural variants, like “best sales enablement software” and “top sales enablement tools.”
These sources are solid, but you’re still inferring and collecting them manually is slow.
Semrush’s AI Visibility tool gives you what none of these sources can: real LLM prompt volume data.
The exact prompts your audience runs in LLMs and how often they’re using them, ranked by frequency.
This grounds your prompt set in real demand that’s always up to date.
You can be sure you’re tracking questions your users are actually asking.
From this list for Gong, I’d choose to track prompts under “AI-Driven Sales Enablement” and “Sales Coaching and Enablement Tools,” as they’re BoFu queries related to my product.
Organize By Product or Use Case
To build your first prompt set, start small.
All you need is 20-30 prompts over 4-6 broad categories that align with your product offering or use cases.
Ensure every category includes a mix of your four types of high-value prompts and add them as tags.
For a B2B SaaS company like Asana, this prompt tracking setup could look like the following:
Project management
Task management
Workflow automation
Team reporting
Best project management software for marketing teams (Evaluation)
Best task tracking tools for cross-functional teams (Evaluation)
Best workflow automation software for ops teams (Evaluation)
Best project reporting tools for enterprise teams (Evaluation)
Asana vs. Monday.com for project management (Comparison)
Does Asana actually improve team productivity? (Reputation)
Asana vs. ClickUp for workflow automation (Comparison)
Is Asana’s reporting good enough for large teams? (Reputation)
Project management software with Slack integration (Evaluation)
Best task management software for small teams (Gap)
Asana vs. Notion for managing marketing workflows (Comparison)
Asana vs. Smartsheet for project visibility (Comparison)
Pro tip: Use branded prompts only for comparison and reputation tracking, as they can inflate your visibility score. Keep the rest of your prompt set unbranded so you actually learn where you’re getting found, not just where you’re already known.
This setup lets you easily see which categories you’re winning and losing in over time.
Or, what types of answers you need to do a better job of showing up in.
Example: Hootsuite Prompt Set
You may decide to add more tags or organize your prompts in a different way as you expand.
For example, Margaret uses multiple different tags — not just categories — for her prompt set for Hootsuite.
We’ve built out prompts in three ways:
Funnel stages, with most of our attention on conversion
Our target industries, with tailored terminology
Intent type: comparative, evaluative, integrative (e.g. “works with X tool”), and problem-led (“how do I solve Y”)
Almost all of our prompts carry multiple tags, e.g., BoFu + Healthcare + Evaluative. That tagging is what lets me slice the data later.
When leadership asks for numbers, she can report on which industries Hootsuite is most visible in, or cross-reference visibility for BoFu prompts with direct traffic trends.
Margaret’s setup shows an important lesson: However you build your prompt set, structure it so you can answer the questions you’ll want to ask later.
How to Track Prompts: Step-by-Step Process
All you need to get started with prompt tracking is a spreadsheet and 30 minutes a week.
Free template: Download our Prompt Tracking Template by Backlinko to follow along with the steps below.
Step 1. Set Up Your Tracking Sheet
With our tracker, you can log the following for each prompt:
The prompt itself
Category
Tags (e.g., type, industry)
The LLM you’re testing it in
Whether your brand was mentioned (yes/no)
Whether you were cited (yes/no)
Sentiment of the mention (positive, neutral, negative)
Competitors mentioned
The date
Customize it to whatever makes sense for your business.
Step 2. Run Each Prompt Across Multiple LLMs
Different LLMs pull from different sources, so answers will be different across all of them.
Tracking prompts from only one LLM won’t give you a full picture of your brand’s presence in AI search.
Ideally track prompts in all major LLMs, including:
ChatGPT
Gemini
Perplexity
Claude
If you need to save time, review Presenc AI’s 2026 platform demographics report to identify which LLMs your audience actually uses and focus on those.
Pro tip: Make sure you’re using a temporary chat to run your prompts. Your regular chat window will serve you an answer that takes into account everything it knows about you from past conversations. You want to track what an LLM might recommend to anyone, not just you specifically.
Step 3. Run Each Prompt at Least Twice Per Session (Optional)
Answers will vary run to run. So if you have the time, run each prompt 2-3 times per session for more reliable data.
You’ll catch when your brand shows up one out of three times (a 33% mention rate).
If you don’t have time, don’t worry. You’ll still see trends over time.
Step 4. Log Competitor Mentions, Too
Competitor data is half the value of prompt tracking.
Seeing a competitor get mentioned consistently in a category you’re performing less well in is a signal.
Maybe they published a new comparison page. Or were included in an influential report.
Look into their strategy to see what they’ve done to improve and learn from them.
Step 5. Track Weekly. Action on Data Monthly.
Regular prompt monitoring on a weekly basis is often enough to catch shifts.
It gives LLMs enough time to crawl and learn from new or updated content.
But don’t make any decisions with only one week of data.
A single week of low visibility in a category could be a fluke. Four weeks of it is a trend worth acting on.
(We’ll cover what to do when you spot one in the “How to Read and Act on Prompt Data” section below).
In this example, Asana shows up with a low citation rate for two LLMs only one week out of four.
Rushing to fix that ASAP could turn out to be a waste of time.
Step 6. Upgrade to an Automated Prompt Tracking Tool
Manual LLM visibility tracking is a great way to validate that you can actually get some useful insights from the practice.
But if you want to grow your prompt set beyond 20-30 prompts, it’s going to start taking much longer.
A tool like Semrush can help you move faster and suggest actionable opportunities based on your data.
To get started, open the Visibility Overview dashboard, enter your domain, and click “Check AI Visibility.”
You’ll see a summary of how often LLMs mention your brand, which competitors are mentioned alongside you, and a breakdown across each LLM.
Scroll down for a list of prompts where you’re already getting mentioned. Click “Opportunities” to see your gap prompts.
Click “Monitor” on the prompts you want to include in your prompt set.
Pick the LLMs you want to monitor, paste in your prompts, and hit “Start Tracking.”
You can add tags by intent, topic, or campaign so you can slice the data later.
Further reading: See how the top AI visibility tools stack up on pricing, LLM coverage, and reporting features.
Example: How an Agency Tracks Prompts for E-Commerce Brands
When you’re tracking hundreds of prompts across multiple clients, a basic spreadsheet won’t be enough.
Jonny Nastor, Founder & Head of Strategy at Digital Commerce Partners, knows this firsthand.
He built a prompt tracking map based on the theory that most buyers ask LLMs about a specific job they need done or task to accomplish. Then filter by their specific situation (a.k.a. constraints).
For example, when shopping for smart doorbells, a buyer might search “best video doorbell with no monthly subscription fees.”
The job to be done is “best doorbell.” The constraint is cost (low fees or no subscription).
He calls it a “Constraint Map.” Every intersection of job and constraint in the map becomes a prompt.
He gets ideas for constraints from keyword modifier data. In Semrush, you can see these in the Keyword Magic Tool.
He also pairs each prompt with search volume data to roughly understand its popularity and prioritize accordingly.
He runs his prompts across ChatGPT, Perplexity, and Gemini automatically through their APIs to log how each one responds.
Depending on the client’s needs, he tracks one or all of the following AI visibility metrics for each prompt:
How many times the brand was cited by LLMs
The brand’s recommendation (or mention) rate
Instances of ghost ranking
For this client, it was only citations on Bing’s AI search.
For each metric, he watches for trends over time, not drawing any conclusions based on one week of data.
He also audits content readiness with a mix of the following (again, depending on client needs):
If a page exists to address the prompt
If that page links directly to the specific products that answer the prompt
If product details (attributes) like price, dimensions, compatibility, etc., are listed on the page
If the content on the page is extractable by an LLM (e.g., no bot-blocking settings or JavaScript-heavy rendering that prevent AI crawlers from accessing the page)
This tells him exactly what to fix so AI is more likely to recommend the brand.
This system surfaced ~52,000 monthly searches with zero AI coverage for one of Jonny’s clients. And a content strategy for the next quarter.
How to Read and Act on Prompt Data
Generative AI prompt tracking data might seem hard to trust at face value.
LLMs give variable answers even in temporary chats.
Platforms push silent model updates that tweak source weighting.
And training data bias is real. One study found LLMs often favor global brands over local ones, meaning you might be invisible in your category because of the model’s defaults rather than your content.
Margaret at Hootsuite found the changing outputs of LLMs surprising when she first started prompt tracking:
It was way more chaotic than I expected. The same prompt can return a totally different brand list on ChatGPT vs Gemini vs Perplexity. “AI visibility” isn’t just a single thing to optimize for: you’re effectively running parallel strategies, and they don’t transfer cleanly, and answers will fluctuate constantly.
This variance doesn’t mean prompt tracking data is useless.
But it is the reason you need to track trends over time instead of getting hung up on moment-specific snapshots.
Here are a few meaningful signals to watch for.
Increased (or Decreased) Frequency Over Time
A consistent rise in mentions or citations usually means your content efforts are working.
A consistent drop could mean a competitor is gaining ground or a key piece of content has gone stale.
Action to take: If the trend is up, note what you published in the weeks before the lift. That’s your playbook. Keep doing it.
If it’s down for four or more consecutive weeks, pick one fix for that category:
Refresh a stale piece with updated data
Publish a new comparison page targeting the prompts you’re losing
Pitch a third-party source that’s being cited in that space
Consistent Source Inclusion
If you repeatedly see the same third-party sources cited in your LLM visibility tracking data, stop and take note.
It means the LLMs trust these sources.
And third-party sources are powerful for AI visibility. Airops found 85% of brand mentions come from third-party pages.
For example, if G2, TechRadar, and Capterra keep appearing across your evaluation prompts, those are your priority pitches.
Semrush’s AI Visibility Overview can help identify sources.
You can see your top “Cited Sources” and “Source Opportunities” (where you’re not being mentioned but your competitors are).
Action to take: Make a list of the sources that keep showing up. For each one, check if you’re already featured. If so, is your listing current and accurate?
Then, pick one you’re missing from and pitch it. A contributed piece, product review campaign, or quote request can all work.
Unless you’re correcting a mention, don’t try to pitch competitors’ sites. They’re not likely to accept.
And don’t only focus on categories you’re losing. Reinforcing visibility in categories you’re already winning is valuable too.
Movement from Mention to Citation
If you used to get mentioned and now only get cited as a source, you’ve slipped into what Jonny calls “ghost ranking” territory:
“Your content shows up in the citation panel, but the AI recommends a competitor.”
Like this example from Teva.
An increase in your “ghost ranking count” over weeks means it’s time to investigate.
Jonny knows this first hand:
Our prompt tracking showed our agency was being cited on agency-directory pages (Trustpilot, Clutch, Semrush) but ghost-ranked at an 83% rate. AI was using directory content as the source of truth, then recommending whichever agency had the densest, most-named presence.
Action to take: Find the sources showing up in the citation panel and audit your presence on each one. Fuller profiles, more reviews, and accurate product details all help convert a citation into a recommendation.
For Teva, this means pitching to be included in the cited articles by REI, backpacker.com, and Outdoor Gear Lab.
It also means updating their own pages (including the ones being cited) with more or newer details.
Then, watching to see if they get less ghost rankings over time.
Common Misreads to Watch for
When you first start generative AI prompt tracking, try to avoid getting tripped up by the following false conclusions.
“Our AI visibility dropped this week. Something’s wrong.”
One week of low visibility or mentions is likely natural variability.
Wait for four consecutive weeks before treating it as a trend worth acting on.
“Our score looks low. We need more content.”
A score can be low for multiple reasons, and the fix isn’t always new content.
Sometimes it’s getting included in more third-party sources or forum threads. Or getting more customer reviews.
If you’re just starting out, you may need to go back to your prompts and make sure they’re not too vague or top-of-funnel.
For example, a broad query like “mesh wifi” may return a definition answer rather than a list of brands.
“Our overall visibility improved after adding more prompts to our tracker. We’re doing something right.”
Adding more prompts to your prompt set will usually make it look like your AI visibility has increased. You’re getting mentioned in more prompts.
“Our overall visibility improved after adding more branded queries. We’re doing something right.”
A branded query is one that mentions your brand name. Of course you get mentioned in the answer.
That doesn’t tell you anything useful.
Limit tracking branded prompts to comparison and reputation prompts. And keep them in a separate cluster so you can filter them when measuring your overall AI visibility.
Weekly Prompt Tracking Workflow
Turning prompt tracking data into a strategy is where you prove the value.
Margaret starts by looking at topics where Hootsuite has low visibility.
Then I look at individual prompt answers to see “What does the internet think about us in this category, and how do we change that?” This usually translates into a few concrete questions for further research:
Where are the third-party listicles, comparison posts, and analyst write-ups that the model is pulling from, and are we represented accurately on them?
Do we have a first-party comparison or evaluation page that an LLM can confidently cite, written for the actual buyer in that vertical?
Do we have our customers (ex. case studies, reviews) reinforcing the same message?
From here, she can recommend actions like updating a case study or getting a mention in a third-party listicle.
Here’s a simple workflow you can use to turn your prompt monitoring routine into AI visibility gains over time.
Pro tip: Don’t expect overnight wins. LLMs take time to reflect new content. Watch for directional improvement over weeks and months. You’re not chasing a score; you’re watching whether your gaps are closing over time.
Step 1. Review prompt cluster trends weekly: What’s the visibility score for your BOFU prompts? Has it fallen for your healthcare cluster? Or a specific use case category?
Step 2. Spot recurring gaps: Note any clusters, categories, or topics that have been underperforming for four weeks or more.
Step 2a. Plan one fix per cluster: Use your content strategy brain to determine the most impactful fix.
That might be:
Publishing a new comparison page to improve a BOFU prompt cluster’s score
Updating an existing article with fresher data
Publishing a type of content you haven’t tried yet on this topic (e.g., video, podcast, social post)
Step 3. Identify recurring third-party sources: Note sources AI consistently cites across your categories. Reddit? LinkedIn? G2? YouTube creators? A trade publication?
Step 3a. Pitch one source you’re missing from: If accepted, you’ll build more off-site authority and increase your chances of being mentioned in the answers you care about.
Bonus resource: Pitching a journalist or news outlet? Use our Journalist Pitch Template, designed by PR experts, to get started quickly.
Start Winning AI Visibility with Prompt Tracking
Prompt tracking isn’t a scoreboard. It’s a compass.
The brands that get real value out of prompt tracking aren’t monitoring every possible prompt.
And they aren’t reacting to one bad week.
They focus on bottom-of-funnel prompts and follow the direction of the graph, not the dot.
Now, it’s your turn:
Build a set of 20 relevant prompts
Download our free prompt tracking template
Log and review data for 30 minutes every week
Once you’re up-and-running, dig into our complete AI optimization guide to get tips on how to fix the issues prompt tracking surfaces.
The post Prompt Tracking: How to Find (and Fix) Your AI Visibility Gaps appeared first on Backlinko.