article

How to run a free AI visibility audit: step-by-step checklist for marketing teams

Run a free AI visibility audit using this step-by-step checklist for marketing teams with ChatGPT, Claude, and Perplexity. This guide delivers a defensible baseline citation rate you can show your CEO in 10 to 15 hours without agency spend or specialized tools.

Liam Dunne
Growth marketer and B2B demand specialist with expertise in AI search optimisation - I've worked with 50+ firms, scaled some to 8-figure ARR, and managed $400k+/mo budgets.
July 16, 2026
14 mins

TL;DR

  • Organic search spans three surfaces: web search (classic SEO), citations (LLM passage retrieval), and training data (offline model associations). Auditing only Google rankings misses the other two entirely.
  • A manual audit across ChatGPT, Claude, and Perplexity for 50 queries establishes a defensible baseline citation rate to show your CEO.
  • Information consistency across your site, Reddit, and third-party publications is the primary driver of LLM citations, not backlink count.
  • Manual audits don't scale past 50 queries. Continuous monitoring requires automated, API-driven infrastructure.

When your CEO forwards a screenshot of a competitor cited in ChatGPT and asks why your brand isn't there, you don't need an agency retainer to find the answer. You need a structured methodology your marketing team can execute today using free tools. This guide walks through exactly that, across all three surfaces where B2B buyers now research: web search, AI citations, and training data. For the complete audit framework this checklist is drawn from, see How to Audit Your AEO and GEO Visibility.

Structuring your internal AI visibility audit#

Treat your manual AI visibility audit as a repeatable diagnostic, not a one-time data pull. The goal: produce a baseline scorecard showing your citation rate across high-intent queries, which competitors win, and where your content fails passage retrieval.

The audit covers three components: AI engine testing (ChatGPT, Claude, Perplexity), training data analysis, and technical crawlability checks. Each surfaces different problems, and skipping one creates blind spots that lead to the wrong fixes.

Evaluating AI search, citations, and LLMs#

Organic search in 2026 spans three surfaces, and most traditional SEO tools only cover the first.

  • Web search: Discoverability via Google and Bing when humans and agents query the web.
  • Citations: LLM retrieval at answer time, where models pull specific passages from your content to build responses.
  • Training data: Brand associations baked into a model's weights before it runs a query.

We've written about how these surfaces interact in our three-surface model overview. The critical finding: per Ahrefs' early-2026 data, about 38% of AI Overview citations come from pages ranking in the top 10, so most of what AI cites is not what ranks, which means strong rankings offer no guarantee of AI visibility. You need to audit all three surfaces explicitly.

Note: AEO (Answer Engine Optimization) and GEO (Generative Engine Optimization) refer to optimizing content for AI-powered answer engines like ChatGPT, Claude, and Perplexity.

Validate data before external spend#

Running a manual audit before hiring an agency protects you from buying services you don't need. If your audit reveals minimal citation on your top commercial queries, you can focus on the specific content and structural issues rather than broader tactics. That changes what you'd spend money on.

The baseline also gives you something defensible for the CEO conversation. Instead of "we're working on our AI visibility," you can show a specific number: your citation rate across 50 queries, versus a named competitor. That data gets budget approved.

Key tools to execute your visibility audit#

You don't need paid software for a first-pass audit. The minimum free stack covers the major LLM platforms, a schema testing tool, and a spreadsheet. The biggest failure mode in DIY audits isn't the tools. It's query bias, single-shot testing, and sample sizes too small to separate signal from noise. If you're evaluating paid options instead of running this manually, see our comparison of AI visibility audit tools.

Verifying brand mentions in LLM outputs#

Use free accounts on ChatGPT and Claude to test whether your brand appears in AI-generated answers. Start a fresh session for each query to avoid session memory contaminating results. Testing in a clean browser session helps reduce the risk of previous queries influencing results.

Ben Moore, our CTO and AI researcher, recommends framing queries as a buyer would ask them, not as a search engine user. "What tools do B2B sales teams use for pipeline forecasting?" produces cleaner results than "best pipeline forecasting software." Run each query two to three times and record whether your brand appears in any version of the answer.

Benchmarking Perplexity search performance#

Perplexity surfaces live web sources when answering queries. Test the same 50 queries on Perplexity that you use on ChatGPT and Claude. Note which URLs Perplexity surfaces, not just whether your brand appears. That cited URL data tells you which specific content assets pass the retrieval test.

Measuring initial AI citation rates#

Citation rate is the percentage of target queries where your brand is cited with a clickable link. Mention rate is lower-bar: your brand name appears in the text without a link. Both matter, but citation rate is the harder and more commercially valuable signal.

To calculate manually: test 50 queries across three engines (150 total tests), count instances where your brand is cited with a link, and divide by 150. Track this baseline and compare it to your performance after optimization work.

Tracking audit data with custom templates#

Build a Google Sheet with these columns: Query, Engine, Brand Mentioned (Y/N), Brand Cited with Link (Y/N), Citation Position, Cited URL, Competitor Mentioned, Competitor Cited with Link, and Intent Category. Run each query and paste the AI's response summary into a notes column.

A spreadsheet handles 50 queries well enough to establish a baseline. Beyond that, manual tracking introduces inconsistency. Our AI Visibility Tracker monitors queries continuously across all major engines using automated retrieval that eliminates session bias entirely.

Step 1: Validate your current AI citation footprint#

The first step establishes where you currently stand across the three major LLM platforms. You're not optimizing yet. You're measuring. The output is a baseline citation scorecard tied to specific query types, so you can compare your brand to competitors and identify the highest-value gaps.

Select your high-intent buyer queries#

Prioritize commercial and comparison queries over informational ones. "Best incident response platform for mid-market SaaS" surfaces in active buying decisions. "What is incident response?" does not. Weight your query set toward comparison queries, problem-solution queries, and category queries that reflect how buyers actively research tools and solutions.

Test at least 50 queries to establish a meaningful baseline. A larger sample helps distinguish real performance patterns from random variation. Weight your queries toward intent categories closest to buying decisions.

Test queries across ChatGPT and Claude#

  1. Open a new incognito browser session for each engine.
  2. On ChatGPT, use the latest available model and disable memory if the option appears.
  3. On Claude, start a fresh conversation with no prior context.
  4. Type the query exactly as a buyer would phrase it, without your brand name in the prompt.
  5. Record the full response, noting brand position and whether a link appears.

We published a B2B SaaS AI search guide covering the query design methodology in detail if you want a full walkthrough.

Track AI citation and mention frequency#

Record every result as you go, distinguishing mentions (brand name appears) from citations (clickable link to your site). Citations carry stronger commercial weight because they drive clicks and prove your content passed the extractability test.

After 50 queries across three engines, total your mentions and citations separately. The gap between mention rate and citation rate tells you whether your content is recognized but not extractable, which is a structural content problem, not a brand awareness problem.

Analyze competitor citation frequency#

For each query where your brand is absent, record which competitors are cited instead. After 50 queries, you'll have a clear ranking of which competitors dominate AI share of voice in your category. This is the data point that moves budget conversations. If a competitor appears in the majority of your core category queries and you appear in a small fraction, the cost of inaction becomes concrete rather than theoretical.

Step 2: Analyze performance on key buyer queries#

With raw data collected, the next step is interpretation. Raw citation counts don't tell you why you're winning or losing retrieval. Analysis surfaces structural reasons: content extractability failures, passage length mismatches, or information inconsistency across sources.

Quantify AI citation and mention rates#

Apply the formula: Citation Rate = (Queries where brand is cited / Total queries tested) x 100. If you tested 50 queries across three engines (150 total tests) and your brand was cited 12 times, your citation rate is 8%.

Segment results by query intent. Informational queries ("what is", "explain", "how does") typically earn higher citation rates than comparative queries ("best", "vs", "recommend"), which produce more brand mentions but fewer clickable citations. If your informational citation rate is significantly higher than your commercial rate, you likely have a content gap at the decision stage, where buyers are actively evaluating options.

How to audit passage retrieval#

LLMs don't rank pages. They retrieve specific passages. Research by Karpukhin et al. on dense passage retrieval demonstrated that dense retrievers outperform keyword-based systems by 9 to 19 percentage points on top-20 passage retrieval, meaning a page ranking first on Google can still fail the LLM citation test if its content isn't structured for extraction.

To audit passage retrieval manually, look at the pages currently winning citations for your target queries. Check whether each cited passage begins with a direct answer rather than context-setting. Compact sections of 200 to 400 words that open with the answer perform better in passage retrieval than long, discursive sections. Our CITABLE framework provides the full architecture for structuring content this way.

Mapping user zero-click search behavior#

When a buyer asks Claude "what's the best tool for [your category]," they often get a complete answer without clicking a link. Your site receives no session, no referral, and no signal in Google Analytics 4 (GA4). Pew Research found only 1% of users click a source link directly from Google's AI Overviews, which means the vast majority of AI-influenced consideration leaves no referral trail.

Citation rate is therefore a more important leading indicator than AI referral traffic. The buyer who sees your brand cited three times across different queries is influenced regardless of whether they clicked.

Step 3: Review brand presence in training sets#

The third surface is training data: what a model knows about your brand from its pre-training, independent of real-time web search. This surface is harder to influence but important to measure, because it determines how your brand is characterized when a model answers without live retrieval.

Verify brand presence in AI models#

Test this by disabling web search in ChatGPT (toggle web search off in settings) and asking: "Tell me about [Brand Name] and what they do." Compare the response to what the model says with web search enabled. A meaningful difference between the two responses indicates a training data gap affecting how models characterize you without retrieval.

A model that knows your brand from training produces a reasonable description of your category, use case, and positioning. A model that doesn't will either hallucinate or give a generic non-answer. That gap tells you how much work you need on third-party mentions and information consistency.

Verify third-party entity mentions#

LLMs build consensus from independent sources. Our Reddit and ChatGPT citation analysis of 144,000 citations found that Reddit appeared in roughly 27% of ChatGPT's internal search slots during query processing, even though it showed up in approximately 0.35% of visible citations. A links-only view of off-page optimization misses a significant share of what shapes AI answers.

To audit third-party presence, search Google for your brand name restricted to Reddit, industry publications, G2, and Capterra. Check whether the descriptions match your current positioning. Conflicting or outdated claims undermine citation quality because research on LLM grounding shows models reward consistent claims across independent sources.

Audit structured data for accuracy#

Schema markup helps AI crawlers parse your content accurately. Check your existing structured data using schema validation tools. Look specifically for Organization, Product, FAQ, and HowTo schema. If your Organization schema contains outdated information or a description that doesn't match your current positioning, correct it before proceeding to content optimization.

Missing FAQ schema is one of the fastest wins available in a manual audit. Pages that answer buyer questions but lack FAQ structured data are harder for AI systems to extract cleanly.

Measuring your AI-sourced MQL contribution#

Attribution is the hardest part of the AI visibility problem. Most of the impact is zero-click, meaning the buyer was influenced but never produced a session your analytics tools can record. The goal is to capture what's trackable while acknowledging what isn't, then build the most defensible attribution model possible for your CFO.

Benchmarking citation rates by query intent#

Group your 50 test queries into intent categories: informational (education and awareness), navigational (brand and product searches), commercial (alternatives and versus queries), and transactional (ready-to-buy terms). Calculate a separate citation rate for each. Transactional and commercial queries should be your highest priority because a citation there reaches a buyer actively evaluating options.

Tracking share of voice vs peers#

AI Share of Voice measures how often your brand is cited as a fraction of total category citations: (Your Citations / Total Citations) x 100. Run this for your top three competitors alongside your brand. Share of Voice shifts slowly, so track it monthly rather than weekly.

Configuring GA4 for AI attribution#

  1. In GA4, navigate to Admin, then Data Display, then Channel Groups.
  2. Create a new channel group called "AI Traffic."
  3. Set the source condition to match referrers: chatgpt.com, claude.ai, gemini.google.com.
  4. Place this channel above Referral in priority order so AI visits aren't grouped into the generic referral bucket.
  5. Note that visits from Perplexity and some other AI platforms may not be captured by this channel and will remain in your Referral traffic. Our AI tracking platform flaw analysis found a significant share of AI referral sessions lack referrer information and land in Direct traffic. GA4 captures users who clicked from AI responses, not zero-click researchers influenced without visiting.

Identifying common AI audit blind spots#

Manual audits surface gaps that traditional SEO dashboards can't see, but they also introduce their own failure modes. Understanding the most common blind spots prevents building a false baseline.

Fixing attribution tracking failures#

Relying solely on GA4 referrers understates AI's impact significantly, because as Pew Research found, only 1% of users click source links from Google's AI Overviews. The fix combines technical tracking (GA4 channel groups, UTM parameters on links you control) with self-reported attribution on your forms.

Add a "How did you hear about us?" field to your demo request form. Self-reported attribution consistently captures AI mentions that technical tracking misses.

Solving dense retrieval ranking failures#

High Google rankings don't guarantee LLM citations. As Karpukhin et al. showed, dense retrievers use semantic vector matching to extract specific passages rather than evaluating documents at a page level. Content that buries its answer after several paragraphs of preamble will lose to content that opens each section with a direct answer.

The CITABLE framework addresses this directly. The "Block-structured for RAG" component (Retrieval-Augmented Generation) requires sections of 200 to 400 words with answer-first formatting so passage extractors identify and retrieve the right block. Score your existing content against CITABLE using our free AEO Content Evaluator.

Addressing AI content parsing errors#

Pages with perfect Lighthouse scores can still be invisible to ChatGPT if content loads via client-side JavaScript without server-side rendering (SSR). React, Vue, or Angular sites without SSR often deliver empty HTML to AI crawlers. Check your robots.txt to confirm you haven't blocked common AI crawler user agents.

Some teams create an llms.txt file at their domain root to signal which pages and use cases matter most to AI systems. This is an emerging proposed standard, and major AI providers have not formally confirmed they follow it, so treat it as a supplementary step rather than a guaranteed signal.

Aligning brand signals across channels#

Inconsistent information across your site and third-party sources is one of the most damaging citation blockers. If your pricing page states one figure but an old press release states another, models encounter conflicting data and reduce confidence in citing your content. Research on effective LLM grounding confirms models reward claims that appear consistently across independent sources.

Audit your G2 profile, Capterra listing, LinkedIn About section, and any press releases for factual accuracy. Product descriptions, pricing tiers, integration lists, and founding year should match your current site exactly. We walk through the full approach in this 2026 SEO strategy video.

Deciding when to hire an AI search specialist#

A manual audit is the right starting point and a natural ceiling. When you hit that ceiling, continuing to DIY wastes time that a specialist team could be spending on execution.

Citation rate below 5% after 30 days#

A consistently low citation rate on high-intent commercial queries suggests your content may be structurally failing passage retrieval. This isn't a content volume problem. It's a content architecture problem, and fixing it requires systematic restructuring using a framework like CITABLE. Our 2 million citation analysis gives detailed benchmarks on achievable citation rates by content type and query intent.

Defending against competitor AI dominance#

In one case, systematic work across all three surfaces moved a B2B SaaS brand's AI visibility from 38% to 64% and grew organic meetings booked by 22%. You can see the full approach in the incident.io case study.

Aligning attribution data for CFO approval#

A specialist partner helps you build a defensible attribution model connecting AI-sourced sessions to Salesforce pipeline. That model, combining GA4 channel groups, UTM (Urchin Tracking Module) tracking, and self-reported attribution, gives you a monthly board slide showing AI-referred sessions, MQLs (Marketing Qualified Leads), opportunities, and closed revenue with stated confidence intervals rather than false precision.

When to bring in specialist support#

Our Search Visibility Diagnostic is a fixed-cost, no-commitment engagement that bridges the gap between a manual baseline and a full optimization program. It delivers an automated AI visibility audit across ChatGPT, Claude, Perplexity, Google AI Overviews, and Gemini, an entity map and answer model for your brand, a schema and content structure audit, and articles optimized using the CITABLE framework. Full details are on our pricing page.

Metric

Manual Audit (DIY)

Automated Audit (Discovered Labs)

Query Scale

Limited to 10-50 queries

Thousands of queries monitored continuously

Engine Coverage

Manual copy-paste on web UI

API-driven tracking across all models

Data Accuracy

High risk of session bias

Clean, unbiased API retrieval

Attribution

Basic GA4 referrer tracking

Full CRM pipeline integration

Addressing internal AI audit methodology gaps#

Running this audit once is useful. Running it correctly is what makes it defensible to your CEO and board.

Defining your baseline citation rate#

Document your baseline scorecard with these fields: overall citation rate, mention rate, citation rate by intent category, and AI share of voice versus your top three competitors. Date-stamp it and set a 90-day review point where you run the same 50 queries again to compare.

Progress at 90 days should show measurable movement if you've restructured content for extractability and addressed information consistency. If citation rate is flat at the 90-day mark, the problem is structural rather than incremental, and that data point builds the case for specialist support.

HubSpot setup for AI attribution tracking#

  1. Add a custom contact property in HubSpot: "Self-Reported Source" (single-line text).
  2. Map it to a "How did you hear about us?" field on your demo request form.
  3. Set HubSpot to populate this field at contact creation to track it across the MQL lifecycle.
  4. Build a report filtering contacts where "Self-Reported Source contains ChatGPT, Claude, or Perplexity."
  5. Track that cohort through MQL, opportunity, and closed revenue.

Combined with your GA4 AI traffic channel group, this gives you two independent data points for the same attribution question, which is the most defensible model available given the zero-click gap.

Strategy for auditing multiple AI surfaces#

A complete AI visibility strategy addresses all three surfaces in parallel: web search optimization (classic SEO and content), citation optimization (CITABLE-structured content, schema, information consistency), and training data work (consistent third-party mentions building offline brand consensus).

Running this manually across 50 queries gives you a defensible business case. Scaling to thousands of queries monitored continuously with CRM (Customer Relationship Management) integration is where manual approaches reach their limit. Book a call and we'll tell you honestly whether a Search Visibility Diagnostic is the right next step, or whether structural gaps are worth fixing first with the free AEO Content Evaluator.

FAQs#

How long does a manual AI visibility audit take?#

A manual audit of 50 high-intent queries across ChatGPT, Claude, and Perplexity requires systematic testing and tracking. This includes query selection, testing across three engines with fresh sessions, and competitor tracking.

Citation rates vary widely based on content optimization, brand maturity, and category dynamics. Track your baseline and measure improvement over 90 days rather than comparing to external benchmarks.

How much does a professional AI visibility audit cost?#

Our Search Visibility Diagnostic is available at discoveredlabs.com/pricing. It includes automated tracking across all major engines, schema implementation, an entity map, and CITABLE-optimized articles.

How often should we run a manual AI visibility audit?#

Run a manual audit quarterly to track trend direction. We recommend automated, continuous tracking for production use because LLM updates are unpredictable and can shift citation patterns significantly between manual snapshots.

Does a high Google ranking guarantee LLM citations?#

No. Karpukhin et al.'s research shows dense retrievers outperform keyword-based systems by 9 to 19 percentage points on top-20 passage retrieval, and per Ahrefs' early-2026 data, about 38% of AI Overview citations come from pages ranking in the top 10, so most of what AI cites is not what ranks. Ranking and citation rate require different optimization strategies.

Key terms glossary#

Citation rate: The percentage of times your brand's website is cited with a clickable link in LLM search responses for a specific set of queries.

Mention rate: The percentage of times your brand name appears in LLM outputs, regardless of whether a clickable link is provided.

Dense Passage Retrieval (DPR): A retrieval system that uses semantic vector matching to extract specific passages of text to answer user queries, independent of keyword overlap.

Information consistency: The alignment of brand claims across multiple independent web sources, which LLMs use to validate facts before citing a brand.

Three-surface model: The framework dividing organic search into web search (classic SEO), citations (AEO/GEO passage retrieval), and training data (offline model weights and associations).

AI share of voice: Your brand's citation count divided by total citations across all brands for a set of category queries, expressed as a percentage.

CITABLE framework: Discovered Labs' 7-component methodology for structuring content for LLM passage retrieval: Clear entity and structure, Intent architecture, Third-party validation, Answer grounding, Block-structured for RAG (Retrieval-Augmented Generation), Latest and consistent, and Entity graph and schema.

Continue Reading

Discover more insights on AI search optimization

Jan 23, 2026

How Google AI Overviews works

Google AI Overviews does not use top-ranking organic results. Our analysis reveals a completely separate retrieval system that extracts individual passages, scores them for relevance & decides whether to cite them.

Read article