article topic
sections to work through
last updated
AI platform features change quickly. This article describes what was verified on the date of the last update. Check the provider's current documentation before making an implementation decision.
In short: build a set of buying questions for one category, record the complete answers, the market, the date and the model mode, then code brand presence and the context of each mention. A percentage without a stated sample is not a credible result. If you need a comparable baseline, choose an AI Visibility Snapshot or a strategic audit.
Your sales director asks: "Does ChatGPT recommend us?" It sounds like a simple question. In practice it is one of the hardest questions in modern B2B marketing, because AI shows no rankings, there is no position one, and every answer is different. Measuring AI visibility requires a different approach from conventional SEO. This guide covers how to do it systematically, from manual tests to professional tooling.
The key difference: in Google you have a position. In ChatGPT you have presence or absence. AI visibility monitoring is the continuous work of establishing which side of that line you are on, and how it is changing relative to competitors.
Looking for instructions for a specific platform? There are separate step-by-step guides: how to check brand visibility in Perplexity and how to check brand visibility in ChatGPT. If you want the result without doing the work yourself, consider an AI visibility audit.
Why measuring AI visibility is difficult
The fundamental problem is non-determinism. ChatGPT, Perplexity and Gemini do not return the same answer to the same question asked twice. Language models have a built-in randomness setting and generate a slightly different narrative each time. A company that appeared for one query may not appear for exactly the same query a minute later.
The second problem is the absence of external data. Unlike Google Search Console, where you have impressions, clicks and positions, AI systems provide no analytics about how often your company appears in answers. You have to measure it yourself, by asking questions and recording results.
The third problem is model diversity. ChatGPT, Perplexity, Gemini and Claude are four different systems with different knowledge bases, different data sources and different answer-generation logic. A company visible in ChatGPT can be entirely unknown to Perplexity, and the reverse. Comprehensive measurement requires testing each model separately.
Manual methods, step by step
Step 1: define the test question set
Before opening ChatGPT you need to know what to ask. The test set should mirror real query patterns - the questions your prospective buyers actually put to AI. We recommend a minimum of 15 to 25 questions across three categories:
- Category questions: "Which [industry] companies do you recommend?" - these test general category recognition
- Problem questions: "Who can help me with [specific problem]?" - these test visibility against customer pain points
- Comparison questions: "[Company A] vs [Company B] - which is better?" - these test position relative to competitors
- Specialist questions: "Which [specialisation] companies operate in [industry or region]?" - these test niche visibility
Step 2: how to test ChatGPT
A testing protocol for a single question: (1) Open a new session with no history. (2) Ask the question without extra context. (3) Save the complete answer. (4) Record whether your company appeared (yes or no), in which position it was named (first, second, third or later) and how it was described (neutrally, favourably, with which attributes). (5) Repeat the same question three times in separate sessions. (6) Calculate the presence rate: how many of the three attempts named the company.
Always test in a session without history and without additional system instructions. You can also test the version with web search enabled - the results will differ from the version without internet access.
Step 3: how to test Perplexity
Perplexity works differently from ChatGPT: it actively searches the web and cites sources. That makes monitoring visibility in Perplexity partly similar to monitoring SEO positions. The protocol: (1) Ask the question in Perplexity. (2) Check the result two ways - whether the company appears in the answer text, and whether your page appears in the Sources section. (3) Record both indicators separately. (4) Test in every language your buyers use, since source coverage differs by language.
Step 4: how to test Gemini and Claude
Gemini: test questions the same way as in ChatGPT. Gemini has a stronger connection to the Google Knowledge Graph, so companies that rank well in Google appear here more often. Note the differences between Gemini and ChatGPT - they indicate where to concentrate the work.
Claude: test and record whether the answer used web search. Feature availability and the model used can depend on the plan, the account and the date. Treat results with and without web access as separate cohorts.
Example test questions for different B2B categories
Test questions have to reflect the real queries your buyers use. Some examples for common B2B categories:
| Category | Example test question |
|---|---|
| Software house | "Which software houses specialise in mobile apps for e-commerce?" |
| Marketing agency | "Which B2B marketing agencies have experience in the industrial sector?" |
| Law firm | "A law firm specialising in advising technology start-ups" |
| Accountancy practice | "Which accountancy firm would you recommend for a company exporting across the EU?" |
| Logistics provider | "Who handles last-mile logistics for small online shops?" |
| HR and recruitment | "Which recruitment firms specialise in hiring IT specialists?" |
| Manufacturer | "Who manufactures custom industrial packaging?" |
For each category prepare at least 15 question variants - different phrasings, synonyms, questions with and without a location, and questions in each language your buyers use. A broader test set gives a more credible picture of visibility.
Metrics to track: Share of Voice in AI
Conventional SEO metrics do not translate directly into AI visibility measurement. You need new indicators. These are the key metrics worth tracking regularly:
| Metric | Definition | How to measure it |
|---|---|---|
| AI Share of Voice | % of test queries in which the company appears | Appearances / total questions x 100% |
| Mention frequency | Average number of times the company appears per query | Total mentions / number of questions |
| Position in the narrative | Where in the list of companies it is named | Median position across all answers |
| Description quality | How AI describes the company (neutral, favourable, with attributes) | Qualitative analysis of the answer text |
| SoV vs competitors | Presence relative to the main rivals | Compare SoV for each company on the same question set |
| Source citation rate | How often the company's site is cited as a source (Perplexity) | Citations / total questions |
How to calculate Share of Voice in AI
An example for company X: the test set is 20 questions. Company X appeared in the answers to 12 of them, so AI SoV = 12/20 = 60%. For comparison: competitor A at 9/20 = 45%, competitor B at 14/20 = 70%. The reading: company X sits mid-field, behind the leader and ahead of competitor A. Analysing the differences between models can show where to concentrate: if SoV in Perplexity is 80% but only 40% in ChatGPT, the problem most likely lies in on-site content optimisation.
Third-party tools for monitoring AI visibility
Manual monitoring is labour-intensive and hard to scale. Several tools now automate part of the process:
| Tool | Type | What it measures | Indicative price |
|---|---|---|---|
| Semrush AI Toolkit | All-in-one SEO and AI | Brand mentions, AI Overviews, positions in Google AI | from around USD 99/month |
| Ahrefs Brand Radar | Brand monitoring | Brand mentions in AI answers and LLMs, share of voice | free plan available |
| Profound | AI monitoring | Dedicated tracking of presence in ChatGPT and other LLMs | on request |
| Peec.AI | AI monitoring | AIO plus LLM tracking, competitor benchmarking | from around USD 79/month |
| Otterly.AI | AI monitoring | AIO triggers, brand mentions, sentiment | from around USD 29/month |
| Mention | Brand monitoring | Mentions across the web (an input for AI SoV) | from around EUR 41/month |
General-purpose tools measure mentions across the web. That correlates with AI visibility, but it is not the same thing. The only way to measure AI Share of Voice directly is to put questions to the models systematically and record the results - manually or with a dedicated tool.
An AI visibility report template
Systematic monitoring needs a structured report. Here is a template you can adapt:
- Section 1 - Executive summary: company AI SoV versus the previous month, AI SoV of the top three competitors, the key changes and the recommended actions
- Section 2 - Results by model: a table of AI SoV for ChatGPT, Perplexity, Gemini and Claude, the month-on-month change, and the best- and worst-performing test questions
- Section 3 - Narrative analysis: how AI describes the company (quotes from answers), which attributes are attributed to it, and which questions produce the best descriptions
- Section 4 - Competitor benchmarking: AI SoV across all monitored companies, position in the narrative relative to rivals, and qualitative notes
- Section 5 - Source tracking (Perplexity): which pages are cited as sources, whether the company's site is among them, and which content is cited most often
- Section 6 - Trend and recommendations: an SoV chart over time (at least three months), the top three recommendations for the coming month, and an estimate of their impact
How to benchmark visibility against competitors
Competitor benchmarking is one of the most important elements of AI visibility monitoring. Without market context, "60% AI SoV" says nothing - it could be excellent or poor, depending on what competitors achieve. How to benchmark properly:
- 1Choose three to five direct competitors of similar size, with a similar offer, in the same niche
- 2Use the same test question set for every company
- 3Test all of them within the same time window, ideally the same week
- 4Calculate AI SoV for each company on each model
- 5Identify where your company leads and where it falls behind
- 6Tie the differences to specific features of competitors' sites and content - what do they have that you do not?
Benchmark tip: test not only direct rivals but also the category leader - the company with the highest AI SoV in your category. Analysing its site, content and online activity shows which actions genuinely translate into AI visibility.
When to commission a professional diagnosis instead of doing it yourself
Self-service monitoring is possible and sensible for smaller companies, or as an initial orientation. A professional diagnosis makes sense when you need deeper analysis and a benchmark. A rough split:
| Situation | Recommendation |
|---|---|
| You want to know whether you are visible at all | A manual test you run yourself (30 minutes) |
| Small company, limited budget | Monthly self-service monitoring in a spreadsheet |
| You want a benchmark against specific rivals | A professional diagnosis - saves time and gives a fuller picture |
| You are planning AI visibility work | A professional diagnosis as the starting point - without it you do not know where to begin |
| Large sales team, B2B enterprise | Professional continuous monitoring (tooling plus service) |
| You have seen leads fall with no clear cause | An urgent AI visibility diagnosis as one possible factor |
The key argument for professional measurement is consistency. A one-off manual test is a snapshot. Credibility increases when you record complete answers, repeat an agreed set and apply the same coding rules to your brand and to the benchmark.
How often to measure
AI systems are updated - the knowledge base behind ChatGPT, Perplexity and the rest changes over time. New content published on the web reaches training data or is fetched live. AI visibility is therefore dynamic. The minimum monitoring frequency is monthly. If you are actively working on optimisation, check every two weeks. After a significant action such as a major article, a PR campaign or a new case study, check after four to six weeks to see whether anything has moved.
ChatGPT Search is a separate category to monitor
ChatGPT Search is a search mode with real-time web access, available to all users including free accounts. It is a fundamental change from the earlier ChatGPT, which relied solely on training data. ChatGPT Search behaves much like Perplexity: it searches the web, cites sources and generates an answer from current content.
ChatGPT without web search and ChatGPT Search are two different systems. A company visible in one can be invisible in the other. Test both modes separately - switch between the answer without tools (training data) and search mode (the globe icon), and record the results separately.
How to test ChatGPT Search: (1) Sign in to ChatGPT. (2) Enable search by clicking the globe icon, or ask a question ChatGPT recognises as needing current information. (3) Ask a category question. (4) Check two things: whether your company appears in the answer text, and whether your page is among the cited sources. (5) Repeat three times in new chats. The result is the percentage of answers containing your brand.
How to test Grok
Grok is xAI's model, available through the X platform and the Grok app. It has become one of the fastest-growing AI systems, and is particularly popular among users active on X. Presence in Grok matters if your customers or partners are active on that platform.
Grok has one distinctive property: real-time access to content published on X. That means your brand's activity there - posts, quotes and discussions - directly shapes what Grok knows about your company. The testing protocol: (1) Open Grok on the web or in the X app. (2) Ask a category question, as in ChatGPT. (3) Record the result. (4) Check with DeepSearch mode enabled, which can produce different results from the standard mode.
| AI system | Data source | How to test |
|---|---|---|
| ChatGPT (no web search) | Training data | New chat, no tools |
| ChatGPT Search | Live web plus training data | Chat with the globe icon enabled |
| Perplexity | Live web, cites sources | Standard query |
| Gemini | Google Knowledge Graph plus web | gemini.google.com |
| Claude | Training data by default | claude.ai, without tools |
| Grok | X plus web | grok.com or the X app |
Assess whether your category is suitable for AI Visibility measurement.
We assess category fit free of charge. Full visibility and competitor measurement is delivered within a Snapshot or a strategic audit.
0 PLN · Fit assessment, without a promise of a full report
Frequently asked questions
How often does ChatGPT change its answers?
Every ChatGPT answer is generated afresh, so even the same question can produce a slightly different result each time. A single test is therefore insufficient. We recommend at least three repetitions per question in separate sessions, and ideally five or more spread across several days.
Is there a free tool for measuring AI visibility?
Dedicated free tools for AI Share of Voice are rare. Ahrefs Brand Radar has a free plan for brand mentions in AI answers. For direct measurement - putting questions to models and recording results - you can work manually using the free versions of ChatGPT and Perplexity plus a spreadsheet.
How many test questions do I need for a reliable measurement?
The sensible minimum is 15 questions, each asked three times, giving 45 observations per model. A full diagnosis is 25 to 30 questions asked five times, giving 125 to 150 observations per model. With fewer questions the results are too exposed to random variation in AI answers.
How does checking visibility in Perplexity differ from ChatGPT?
The key difference: Perplexity cites sources and searches the live web, so you measure two indicators - the company appearing in the answer text, and your page being cited as a source. ChatGPT without web search relies on training data, so you measure only appearance in the answer text. The same number can mean different things in each model.
When are monitoring results worth presenting to the board?
Present them once you have at least three months of data, so a trend is visible rather than a snapshot. The strongest argument is a comparison of your AI SoV with direct rivals. If a competitor sits at 70% and you sit at 25%, that is a concrete number that makes the case for investment.