Most teams start optimizing for AI search before they know where they stand — which means they can't prioritize, can't prove progress, and can't tell whether they have an entity problem, a content problem, or an authority problem (the fixes are completely different). The audit below answers all of that in roughly a day of work, using nothing more expensive than the engines themselves and a spreadsheet. Run it once this week for a baseline, then on a fixed cadence so every optimization decision traces back to a measured gap.
Step 1: build a query set that mirrors real buyer questions
The audit is only as honest as its queries. Skip keyword-tool exports and build 50-100 questions the way buyers actually phrase them to an assistant — full sentences, context included, often with a 'for us' qualifier ('best AI recruiting platform for a 40-person startup'). Source them from reality: questions prospects asked on sales calls, support and onboarding tickets, community threads in your space, and the follow-up questions the engines themselves suggest. Organize the set into four buckets, because your visibility will differ sharply between them and the fixes differ too.
- Branded (10-15%): 'What is [company]?', '[company] pricing', '[company] vs [competitor]' — tests whether engines know and describe you accurately.
- Category/commercial (40%): 'best X for Y', 'top X platforms 2027', 'X alternatives' — the shortlist-forming queries where citations translate to pipeline.
- Problem/informational (30%): the questions buyers ask before they know the category exists — where authority is built.
- Comparison/decision (15-20%): 'is X worth it', 'X vs doing it in-house', 'how much does X cost' — late-stage queries with outsized influence.
Step 2: test each engine systematically, not anecdotally
One person asking ChatGPT three questions is an anecdote; an audit is the same query set run the same way across every engine that matters, with results logged before interpretation starts. Control what you can: use clean sessions (no memory, logged-out or fresh chats where possible), run each query once per engine per cycle, and capture the full answer plus its citations, not just whether you appear. Personalization and non-determinism mean individual answers wobble — which is exactly why you measure across 50-100 queries and read the aggregate, not any single response.
| Engine | How to test | What to record |
|---|---|---|
| ChatGPT (with search) | Fresh chat, memory off, note when it browses | Mentions, cited links, how it describes you, competitors named |
| Perplexity | Logged-out or clean thread | Numbered citations by position, your share vs. competitors |
| Google AI Overviews / AI Mode | Clean browser profile; note if no overview triggers | Whether an overview appears at all, cited sources, your presence |
| Gemini | Fresh conversation | Mentions and links, consistency with what AI Overviews shows |
| Copilot (optional, B2B-relevant) | Clean session | Mentions and cited links — its Bing retrieval overlaps ChatGPT's |
Step 3: log citations and share of voice against competitors
For every query-engine pair, log four fields: mentioned (your name appears in the answer), cited (you're a linked source), accurate (what it says about you is correct), and who else appears. From those, compute the three numbers the whole program will be managed against: mention rate and citation rate per bucket, description accuracy on branded queries, and share of voice — your citations divided by total vendor citations across the commercial bucket, tracked per competitor. Share of voice is the metric that makes the audit strategic: being absent from 'best X for Y' answers that name three competitors is a measurable, addressable pipeline leak, and it's the number that moves budget conversations.
- A spreadsheet is enough: rows are queries, column groups per engine, plus a competitor tally sheet. Tools can come later; the method matters more.
- Flag inaccurate descriptions as their own severity class — engines confidently misdescribing your pricing or ICP does damage invisibly.
- Record the answer text (or a screenshot link) so later cycles can diff what changed, not just whether numbers moved.
- Note which of your URLs get cited when you do appear — the pages engines already trust are your fastest levers for expansion.
Step 4: diagnose why you're absent — entity, content, or authority
The audit's real product is the diagnosis. Absence from AI answers has three root causes, and they're distinguishable from the data you just collected. Get the diagnosis wrong and you'll spend a quarter writing content to fix what is actually an entity problem.
| Diagnosis | Signature in the audit data | The fix |
|---|---|---|
| Entity problem | Engines answer branded queries wrongly, vaguely, or confuse you with others | Canonical description everywhere, Organization schema, consistent profiles, third-party corroboration |
| Content problem | Engines describe you correctly but never cite you on category/informational queries; competitors' answer-shaped pages appear instead | Answer-first restructuring, question-cluster coverage, tables and TL;DRs, FAQ schema on visible Q&A |
| Authority problem | You're occasionally cited on long-tail but never on 'best X' commercial queries; the same 2-3 competitors dominate via data and reviews | Original data, case studies, named frameworks, earned third-party mentions — the slow compounding layer |
| Technical problem (check first) | You're absent everywhere despite decent classic-search rankings | Robots.txt and CDN bot rules, server-rendered content, llms.txt, schema basics |
Step 5: turn it into a backlog, then re-run on a cadence
Convert findings into a prioritized backlog by expected impact per unit of effort: technical unblocks first (hours of work, binary payoff), entity fixes second (days, gate everything else), then content restructuring ordered by commercial-bucket gaps, then authority projects as the standing quarterly investment. Every backlog item should reference the specific queries it's meant to move, so the next audit cycle scores it. Re-run the identical query set monthly — same queries, same method, clean sessions — and review the trend quarterly against the baseline. Refresh no more than 10-20% of queries per quarter as your market shifts, keeping the core set stable so the trendline stays comparable; a query set that changes every cycle can show any result you want, which is to say none.
- 1Week 1: run the baseline audit and write the diagnosis (one page: scores per bucket, top three gaps, root causes).
- 2Weeks 2-3: clear technical and entity items — they're fast and they gate the rest.
- 3Weeks 4-12: work the content backlog against the commercial-bucket gaps; ship the first authority piece.
- 4Monthly: re-run the set, log deltas, promote or demote backlog items based on what actually moved.
- 5Quarterly: review the trend, refresh up to 20% of the query set, and re-baseline share of voice against competitors.
