article

The B2B SaaS guide to AI search prompt selection and share of voice measurement

AI share of voice measures how often your brand appears in LLM responses across buyer intent prompts tracked over 30 days. This guide shows CMOs how to build a 20 to 40 prompt baseline, calculate citation rates across ChatGPT and Claude, and connect AI referred sessions to pipeline in your CRM.

Liam Dunne
Growth marketer and B2B demand specialist with expertise in AI search optimisation - I've worked with 50+ firms, scaled some to 8-figure ARR, and managed $400k+/mo budgets.
July 30, 2026
15 mins

TL;DR:

  • B2B buyers now research vendors inside ChatGPT, Claude, and Perplexity before your sales team hears from them. Traditional keyword rankings miss this activity entirely.
  • Shift measurement from keyword rankings to prompt-based share of voice tracking across the LLMs your buyers use most.
  • Build a baseline of 50 to 100 category-relevant buyer prompts, run them across at least two LLMs for 30 days, and calculate AI SOV as (Brand Citations / Total Category Citations) x 100.
  • Tag all AI-referred sessions with UTM parameters and add a self-reported attribution field to your demo request form to capture zero-click buyers.
  • Combine UTM-tagged sessions as the floor and self-reported responses as the ceiling to produce a pipeline attribution number your board can act on.

How to measure share of voice in AI search: AI share of voice measures how often your brand appears in LLM responses across a defined set of buyer-intent prompts. Calculate it as (Brand Citations / Total Category Citations) x 100, tracked across your priority prompt set in ChatGPT, Claude, Perplexity, and Gemini. This is a directional signal, not an absolute KPI. Pair it with mention rate, citation rate, and UTM-tagged traffic to build an attribution story your board can act on.

Traditional keyword rankings no longer capture the full picture of organic search. B2B buyers increasingly use AI tools like ChatGPT and Perplexity during vendor research, and that research happens in LLM conversations where traditional analytics platforms struggle to track visibility. While tools like Google Analytics can capture some AI-referred sessions, they miss zero-click exposure, AI Overview impressions, and conversational brand mentions that don't result in immediate clicks. Your organic channel runs on two surfaces simultaneously: Google's blue-link results and AI-generated answers. This guide gives you the exact methodology to discover buyer prompts, build a defensible share of voice metric, and attribute pipeline back to AI citations.

Optimizing prompts to capture buyer intent#

Capturing invisible AI intent signals#

AI assistants make buyer research invisible by default. B2B buyers are running vendor research inside ChatGPT, Perplexity, and Claude, yet most tech brands appear in a small fraction of those responses. That gap is the visibility problem. Your Google rankings, Core Web Vitals scores, and meta descriptions tell you nothing about whether buyers are finding you inside an LLM conversation.

Capturing these signals requires tracking structured prompts, not isolated keywords. A keyword is "incident response software." The prompt a buyer actually types is "what's the best incident response tool for a mid-size engineering team?" These two phrases trigger entirely different retrieval processes, and ranking for one does not guarantee citation for the other. Liam Dunne's guide to AI search for B2B SaaS covers the mechanics in detail, and the AEO audit guide shows how to map this alignment across an existing content library.

Measuring AI share of voice via prompts#

LLMs do not return ranked lists. They synthesize a single answer by extracting semantically relevant passages from sources they judge credible and consistent. This retrieval process is known as dense passage retrieval, where the model matches meaning rather than just keywords, which outperforms traditional keyword matching approaches. Optimizing for keyword density alone does not move citation rates.

The table below contrasts the two surfaces so you can see where measurement priorities diverge.

Attribute

Traditional SEO

Generative engine optimization (GEO)

Primary goal

Blue-link ranking

Citation in AI-generated answers

Retrieval mechanism

BM25 keyword matching

Dense passage retrieval

Key ranking signals

Backlinks and site authority

Passage structure and entity clarity

Measurement metric

SERP position, impressions, CTR

Citation rate, mention rate, share of voice

Content format

Document-level optimization

Block-structured, extractable sections

For a deeper read on why these surfaces require separate playbooks, see our post on whether SEO and AEO differ and the video on why SEO is not AEO.

Mapping prompts to buyer search needs#

Prompts must mirror natural conversational language, not robotic search syntax. A buyer evaluating your category in Claude does not type "best [category] software 2026." They type "what tool does a Series B SaaS company use to handle incident response without needing a dedicated SRE team?" Aligning your content to that phrasing is what gets it extracted. The GEO audit vs SEO audit post explains where traditional audit approaches break down on this dimension.

Identifying high-intent AI search queries#

1. Segment keywords by purchase intent#

Start with your existing commercial keyword set and convert each term into a conversational prompt. A typical conversion might transform "Project management software for agencies" into "what project management tool should a 20-person digital agency use?" Commercial modifiers commonly include "best software for," "alternatives to [competitor]," and "how do I solve [specific problem]." These patterns translate directly into the queries buyers run in LLMs during active evaluation. The AI search audit guide includes a conversion template for this process.

2. Uncover buyer queries in sales calls#

Your sales call transcripts in tools like Gong or Chorus often contain the language buyers use when evaluating your category. Mine the questions prospects ask during discovery calls and competitive comparisons. "How does this compare to [competitor] for enterprise compliance?" is a prompt you can run verbatim in ChatGPT to see who gets cited. The free AI visibility audit checklist includes a section on extracting prompts from sales data.

3. Uncover buyer intent in support logs#

Post-purchase support questions often mirror pre-purchase research patterns. Customer support logs in tools like Zendesk or Intercom surface the specific problems buyers try to solve, and those problems closely align with what pre-purchase buyers ask LLMs. "How do I set up automated alerting without writing custom scripts?" is the kind of question a buyer would also ask Claude before evaluating vendors. These questions belong in your prompt library.

For a step-by-step method to extract prompts from both sources at scale, see the guide to mining sales calls and support tickets for AI prompts.

4. Track competitor citation gaps#

Commercially valuable prompts are often where competitors appear and your brand is absent. Our AI visibility tracking system ingests your prompt library, compares outputs across ChatGPT, Claude, Perplexity, and Gemini, then surfaces a ranked list of gaps where competitor presence is high and your citation rate is zero. The AEO checklist post explains how prompt-alignment gaps between your content and buyer queries create those absences.

5. Audit prompt performance in LLM chats#

Run your priority prompts manually in ChatGPT and Claude before automating. Note which brands appear early in each answer, what sources are cited, and how your brand is described when it does appear. This qualitative layer tells you whether you are cited accurately and in the right context, not just whether you appear at all.

The guide to auditing your AI prompt set for noise reduction walks through how to filter low-signal prompts before building your baseline.

Methodology for mapping high-value AI queries#

Building a scalable prompt framework#

Structure your prompt database with variables that reflect how your buyers actually search: audience segment (VP of Engineering, Head of IT, CFO), company size (SMB, mid-market, enterprise), and use case (incident response, compliance automation, developer onboarding). A 3x3x3 matrix generates 27 base prompts. Layer in comparison prompts ("X vs Y"), how-to prompts, and category prompts to build a comprehensive baseline set. This structure makes the database segmentable for reporting by buyer type.

Filtering prompts by buyer intent#

Filter out purely informational queries ("what is incident response?") with low conversion potential. Keep prompts that indicate active buying cycles: solution comparison, vendor evaluation, pricing research, and use-case validation. A quick proxy is whether your sales team would find value in a citation at that query. If not, the prompt does not belong in your baseline set.

Mapping prompts to buyer stages#

Organize your retained prompts into three tiers:

  1. Top-of-funnel: Problem identification queries ("why does our on-call process keep failing?"). These build brand awareness inside LLMs.
  2. Mid-funnel: Solution evaluation queries ("what are the main incident management platforms for DevOps teams?"). Share of voice matters most here.
  3. Bottom-of-funnel: Vendor comparison queries ("incident.io vs PagerDuty for a 50-person engineering team"). These drive direct pipeline.

Bottom-of-funnel prompts deserve the highest optimization priority because they sit closest to a purchase decision. The entity SEO guide explains how to structure brand entity information so LLMs recognize and accurately describe your product at the comparison stage.

Ranking prompts by commercial value#

Score each prompt using factors that predict commercial value: estimated query volume (use Google's search volume as a directional proxy), competitor citation density (how many competitors appear per response), and proximity to your ideal customer profile (ICP, the detailed description of your best-fit buyer). Weight bottom-of-funnel prompts more heavily because they sit closest to a conversion event. This scoring model tells you where to concentrate your CITABLE content production first.

Once scored, the guide to mapping AI prompts to a content plan shows how to translate that priority ranking into a production calendar.

How to measure share of voice across LLMs#

Defining AI search share of voice#

AI SOV is a directional metric calculated as:

AI Share of Voice = (Brand Citations / Total Category Citations) x 100

Run this formula across your full prompt set, aggregated over 30 days. If those prompts cite 80 total competitor instances and your brand accounts for 20 of them, your share of voice is 25%. Tracking tools rely on inferred data and simulated query environments rather than direct access to LLM weights, so treat this as a signal, not a guarantee. Initial citation movement typically becomes visible within one to two weeks of publishing new content. Monthly aggregated data smooths variance and produces a more reliable trend line.

The guide to measuring SOV across ChatGPT, Perplexity, and Google AI covers platform-specific tracking setup in detail.

Analyzing competitor AI search presence#

Calculate competitor SOV using the same formula. Run each competitor's brand through your prompt set and count appearances. A meaningful gap between your SOV and your top competitors' combined SOV is the benchmark you bring to a board meeting.

The AI visibility audit tools comparison covers which platforms can automate competitor tracking at scale.

Measuring your initial AI market share#

Most B2B brands appear in approximately 3% of relevant AI-generated answers regardless of their Google rankings, according to Walker Sands' B2B AI Search Visibility Benchmark of 828 enterprise companies. This reflects the core measurement gap: Google rankings do not predict AI citation rates because the retrieval mechanism differs. Ahrefs data from early 2026 shows only about 38% of AI Overview citations came from pages ranking in the top 10 for the same query, so most of what AI cites is not what ranks. The GEO audit guide walks through how to establish your starting position across all major LLMs.

Establishing your AI share of voice baseline#

Defining citation vs mention metrics#

A mention is your brand name appearing in the response text. A citation is a clickable link to your domain. These move independently. Different LLMs have different citation and mention behaviors: Claude cites external domains less predictably than Perplexity and weights first-party brand content heavily, according to Otterly AI's analysis of 379,321 Claude citations. Perplexity includes external links in a majority of responses. For pipeline attribution, citations drive trackable traffic. For brand-building inside LLMs, mentions matter too because they shape how buyers perceive your category position before they click anything.

Tracking metrics for buyer intent queries#

Use this three-step setup to establish your baseline:

  1. Define your prompt set. Finalize your prompts segmented by buyer stage, audience type, and use case.
  2. Establish a 30-day baseline. Run the prompts across ChatGPT, Claude, Perplexity, and Gemini weekly. Record brand mentions, competitor mentions, and citation links in a shared tracking sheet.
  3. Integrate with your CRM. Configure HubSpot or Salesforce to capture AI-referred sessions via UTM tagging and add a self-reported attribution field to your demo request form to capture zero-click buyers.

The step-by-step visibility audit checklist walks through each step with templates.

What a 40% citation rate indicates for ROI#

Moving from a near-zero baseline to a 40% citation rate across your priority prompt set typically takes three to four months of consistent content optimization. incident.io grew AI visibility from 38% to 64% and booked 22% more organic meetings by restructuring content for passage retrieval and building consistent third-party mentions in the communities where their buyers were already active.

Before engaging Discovered Labs, incident.io had been running homegrown LLM prompts with no clear strategy for what to optimize for or how to structure content for extraction. The full case study covers how that changed.

How to audit your brand's AI mention rates#

Measuring brand-specific AI citations#

Track two citation types separately: branded prompts (where you explicitly name your brand) and unbranded category prompts (where buyers search without a preferred vendor in mind). You win or lose invisible pipeline in the unbranded category. Our AI Visibility Tracker automates both, running your prompt set across ChatGPT, Claude, Gemini, and Perplexity and surfacing mention rate, citation rate, and SOV by prompt tier.

Defining AI citation quality standards#

Not all citations carry equal weight. Assess citation quality on three dimensions: accuracy (does the LLM describe your product correctly?), context (is it cited in a buying-relevant scenario?), and position (does it appear in the first two sentences of the answer or at the end?). Our free AEO Content Evaluator scores your content against the CITABLE framework before you publish, so you can identify which sections an LLM would extract versus skip. For a complete breakdown of why content fails to get cited and how to fix it, see the AI citation audit template.

Attributing revenue to AI search citations#

UTM parameters for AI visibility#

Tag all links embedded in your CITABLE content with a consistent UTM structure:

  • utm_source=perplexity&utm_medium=ai-assistant&utm_campaign=[category_prompt]
  • utm_source=chatgpt&utm_medium=ai-assistant&utm_campaign=[category_prompt]
  • utm_source=claude&utm_medium=ai-assistant&utm_campaign=[category_prompt]

Apply this to every internal link and landing page URL your content references. This captures buyers who click through from an AI-generated answer and lands them in your CRM with a clean, attributable source. The managed AI visibility audit guide includes UTM configuration as part of the full attribution setup.

Identifying AI leads in your CRM#

Configure HubSpot or Salesforce to capture AI-referred leads in two ways. First, set up automated lead source rules that tag any session arriving via the UTM parameters above as "AI Search." Second, add a self-reported field to your demo request and contact forms: "How did you hear about us?" with explicit options for ChatGPT, Claude, Perplexity, and "AI assistant (other)." This captures zero-click buyers who researched inside an LLM, visited your site directly, and never triggered a UTM tag. The AI search ROI guide provides the exact reporting structure your CFO will accept.

Closing the AI attribution gap#

GA4, HubSpot, and your CRM will give you three different numbers for AI-referred pipeline in any given month. That is expected, not a failure. GA4 undercounts dark social and zero-click traffic. HubSpot last-touch attribution misses multi-session paths. Build your board slide using a blended model: UTM-tagged sessions as the floor, self-reported attribution as the ceiling, with a stated caveat that the true number sits between the two.

Gladia applied this integrated approach across both organic search surfaces and achieved 7x sales-accepted leads in four months, with 93% of those AI-referred leads coming directly from LLM search. Pairing their CITABLE content operation with consistent third-party mentions across Reddit and independent publications gave LLMs the consistency signals they needed to cite Gladia over competitors. Our Reddit citation research, which analyzed 144,000 AI citations, found Reddit referenced in roughly 27% of ChatGPT's search results, a signal a keyword-only content program never builds.

Drafting your AI pipeline impact slide#

Your quarterly board slide needs four numbers: AI-referred sessions (UTM-tagged, from GA4), AI-referred MQLs (lead source field in HubSpot), AI-attributed pipeline (self-reported field plus UTM-confirmed deals in CRM), and your current SOV versus top three competitors. Add a month-on-month trend line for citation rate across your priority prompt set. This is a defensible, internally consistent story that answers the "prove every metric" objection before anyone asks it. Tom Wentworth, CMO at incident.io, on the strategic priority:

"It's clear that working on AI visibility is as important now as SEO was in the 2010s... I believe early adopters will win the Answer Engine Optimization battle, so it was important for me to find a partner who would help us see results quickly." - Tom Wentworth, CMO at incident.io

Building your master list of B2B search prompts#

Determining the ideal prompt sample size#

Start with a minimum of 50 to 100 category-relevant prompts tracked across two to three LLMs to produce a statistically meaningful SOV benchmark. Below 15 prompts, the sample is too small to produce a reliable signal. Programs requiring greater statistical robustness can expand beyond 100 prompts, but that requires dedicated measurement infrastructure beyond a manual tracking sheet.

For a comparison of AI prompt query volume against traditional search volume to help prioritize your initial set, see the AI prompt volume vs search volume guide.

Scheduling your AI share of voice checks#

Run your prompt set on a monthly cadence to align with standard marketing reporting cycles. Citation results for identical queries across major LLMs can vary significantly month-over-month, according to our citation tracking analysis. Monthly aggregated data smooths that variance and gives you a trend line you can defend. Report SOV, mention rate, and citation rate as month-on-month changes, not as point-in-time figures.

Key LLMs for tracking buyer queries#

Not all LLMs behave the same way. ChatGPT's citation presence reached roughly 6.8% of prompts by May 2026, up from about 1.6% a year earlier, according to Similarweb's 2026 Generative AI Landscape report, so a mention in ChatGPT builds brand presence without necessarily driving a session. Claude cites external domains less predictably than Perplexity and weights first-party brand content heavily, according to Otterly AI's analysis of 379,321 Claude citations, making it important for brand-narrative accuracy. Perplexity includes external links in a majority of responses, making it a high-value citation target for driving trackable traffic. Optimize your content for extraction first, which helps all models, then build third-party consistency signals to influence ChatGPT and Perplexity citation rates specifically. The video breakdown of the new way of SEO in 2026 covers how these surfaces interact in a connected content operation.

Measuring your citation rate progress#

Track month-on-month progress using the SOV formula across your fixed prompt set. The anonymized B2B SaaS client we worked with grew to 3,500+ trials in seven weeks, a 6x increase, by restructuring content for LLM passage retrieval and building consistent off-page signals. The citation rate moved because content structure changed, not because they published more volume.

Repurposing legacy keyword data#

Your existing blog content is not wasted. The CITABLE framework restructures it for LLM extraction without requiring you to start from scratch. CITABLE is our content methodology with components including: Clear entity and structure, Intent architecture, Third-party validation, Answer grounding, Block-structured for RAG (Retrieval-Augmented Generation, how LLMs pull relevant passages), Latest and consistent, and Entity graph and schema. The components are documented in full at our CITABLE framework post. Apply the framework to your highest-traffic, highest-commercial-intent pages first, because those pages already carry domain authority and are most likely to get extracted once structurally optimized. Run your top five commercial pages through the free AEO Content Evaluator to see where each one falls short on the CITABLE dimensions before committing to a full restructure.

Conclusion#

AI buyer research happens before your sales team ever hears from a prospect, and keyword rankings do not capture it. The methodology in this guide gives you the building blocks to close that gap: a prompt set that mirrors real buyer language, an SOV formula you can run today, and an attribution model that produces a defensible pipeline number. Running SEO and AI search as one connected operation is how incident.io, Gladia, and others grew organic pipeline without doubling content budgets. The same approach is available to any B2B SaaS team willing to shift from ranking reports to citation tracking.

If you want to see your current citation gaps across ChatGPT, Claude, and Perplexity, request a baseline Search Visibility Diagnostic. We analyze your prompt universe, map where competitors are winning your category queries, and show you exactly where to start. If you want to talk through fit before committing, book a call and we'll tell you honestly whether we're a fit.

FAQs#

How many prompts should we track to establish a baseline?#

Track a minimum of 50 to 100 category-relevant prompts across at least two LLMs to produce a statistically meaningful share of voice number for a board meeting. Below 15 prompts, the sample is too small to produce a reliable signal.

What is a good baseline citation rate for a B2B SaaS brand?#

Most B2B brands appear in approximately 3% of relevant AI-generated answers regardless of their Google rankings, according to Walker Sands' B2B AI Search Visibility Benchmark. A brand consistently appearing in 10% or more of its priority prompts is already performing above what we see at the start of most client engagements.

Perplexity includes external links in a majority of responses, making it a high-value citation target for driving trackable UTM-tagged sessions. ChatGPT mentions build brand presence but convert to fewer direct sessions than the same mention in Perplexity, based on Similarweb's 2026 Generative AI Landscape report, which puts ChatGPT's citation presence at roughly 6.8% of prompts as of May 2026.

Does strong Google SEO automatically improve AI citation rates?#

No. Ahrefs data from early 2026 shows only about 38% of AI Overview citations came from pages ranking in the top 10, so most of what AI cites is not what ranks. Google weighs links and site authority. LLMs weigh passage structure, entity clarity, and third-party consistency, which is a different optimization problem, as we cover in the SEO vs AEO explainer.

How do we attribute pipeline to AI citations when clicks do not always happen?#

Combine three data sources: UTM-tagged sessions from GA4, automated lead source rules in HubSpot or Salesforce, and a self-reported "How did you hear about us?" field on your demo request form. Use UTM-tagged sessions as the floor and self-reported responses as the ceiling. The AI search ROI guide walks through how to present this blended model to your CFO.

Key terms glossary#

AI share of voice: A directional metric calculated as (Brand Citations / Total Category Citations) x 100 across a defined prompt set over a 30-day window.

Citation rate: The percentage of analyzed AI search responses that include a clickable link to your domain, tracked separately from brand mentions.

Mention rate: The percentage of AI search responses that name your brand, regardless of whether they include a clickable link to your site.

Passage retrieval: The process where an LLM extracts semantically relevant text blocks from a document to synthesize an answer, rather than ranking the full page. Dense passage retrieval (where the model matches meaning rather than just keywords) outperforms traditional keyword matching approaches, which is why keyword optimization alone does not move citation rates.

SERP (Search Engine Results Page): The list of results Google or another search engine returns for a query. In this guide, SERP rankings refer to traditional blue-link positions, distinct from AI-generated answer citations.

Continue Reading

Discover more insights on AI search optimization

Jan 23, 2026

How Google AI Overviews works

Google AI Overviews does not use top-ranking organic results. Our analysis reveals a completely separate retrieval system that extracts individual passages, scores them for relevance & decides whether to cite them.

Read article