How to measure AI search traffic, competitors, and factual accuracy

AI visibility has several checkpoints: a system can mention a brand, link to a page, send a visit, or contribute to a qualified opportunity. Measure those signals separately, connect them to verified records, and label every inference.

A worked scorecard, with hypothetical numbers

Suppose a fixed panel produces 12 answers from 20 runs. Three answers mention your brand. A separate factual review finds five supported claims, one contradicted claim and two unverifiable claims. These are examples for checking arithmetic, not Dardo results.

A decision aid
MetricCalculation
Answer rate12 / 20 = 60%
Mention rate among answers3 / 12 = 25%
Verified accuracy5 / (5 + 1) = 83.3%
Verifiability(5 + 1) / 8 = 75%

Keep the eight runs without an answer visible. A high accuracy rate can hide uncheckable claims, so show verifiability too. Flag consequential errors individually even when the aggregate looks good.

  1. 01Capture
  2. 02Classify
  3. 03Score
  4. 04Reconcile

Define the signal before opening a dashboard

Start by naming the event you want to count. A brand mention may be unlinked or contain a wrong fact. A citation links to a page, while a Search Console impression means that a site link was shown in a supported Search surface. A click leaves Search and a GA4 session begins when the visitor reaches the site. A form click or outbound WhatsApp, phone, or booking click shows intent, not a received inquiry. Count a lead after the form or destination confirms receipt, then use CRM stages for qualification and a won customer. These are different denominators. Keep them separate even when one report displays them together. This prevents a high mention rate becoming traffic, traffic becoming revenue, or a click becoming a sale.

Use Search Console for Google’s own AI surfaces

The current Search Console Help page calls this the Generative AI performance report. It says rollout was worldwide as of August 31, 2026, while access can still be gradual or absent when impressions are insufficient. It includes AI Overviews and AI Mode impressions, grouped by pages, countries, dates, and devices. Page rows generally use the canonical final URL; charts aggregate by property and tables depend on the dimension. Export both for a dated record; fresh data can be preliminary. Search Labs experiments are excluded. This is a useful first-party measure of links shown in Google’s generative features, but not a prompt-by-prompt competitor ledger. Its data is included in Web search, so do not add it to regular Web performance totals as an independent pool. Use it for Google’s surface, then GA4 and CRM for visits and outcomes.

Group assistant referrals carefully in GA4

GA4’s default channel group includes an AI Assistant channel for visits from sources such as ChatGPT, Gemini, DeepSeek, Copilot, and Grok. Google says it excludes AI Overviews and AI Mode, which are classified under Organic Search. Recognized assistant referrals use medium ai-assistant and campaign (ai-assistant), so start with Reports, Acquisition, and Traffic acquisition. Inspect Session default channel group and Session source / medium together, then compare sessions, engagement, and key events. For a custom grouping, open Admin, Data display, and Channel groups, create or copy a group, add an AI assistants channel, choose a source or referral URL condition with matches regex, and place it above Referral. A conservative host pattern is ^.*(chatgpt\.com|openai\.com|gemini\.google\.com|perplexity\.ai|copilot\.microsoft\.com|claude\.ai).*$; adapt it to hosts that appear in the property and review it when providers change. Avoid a broad pattern containing google, which blurs ordinary Search with assistant traffic. A missing referrer, privacy setting, or direct visit can hide the recommendation, so this measures observed referrals rather than all AI-influenced visits.

Build a landing-page exploration

To see what an observed referral did after arriving, open Explore and start a blank or free-form exploration. Add Landing page + query string as a dimension, then add Session source / medium and Session default channel group when the property makes them available. Use Sessions, Active users, Engaged sessions, Event count, and the relevant key events as metrics. Put the landing page in rows, or compare source rows with a landing-page breakdown, then filter the date range and market consistently. A second tab can isolate pages that received generate_lead or a later qualification event. Google’s landing-page guidance also supports a path exploration when you want to see the next pages after a landing page. Keep the scope visible: Traffic acquisition is session scoped, while User acquisition is user scoped, and their numbers should not be compared as though they answer the same question. Reports and explorations can also differ because processing, modeling, supported fields, and recent data windows differ. Set a reporting cutoff, retain the export, and do not interpret a last-48-hours gap as a campaign failure.

Instrument the lead lifecycle, not just clicks

Use recommended event names when they match the business process. Send generate_lead only after a person submits a form or request successfully; send qualify_lead when the team marks it as qualified; send close_convert_lead when it becomes a customer. These names require implementation and context; they do not appear because they were written in a report. Keep outbound clicks, link clicks, or a WhatsApp button as separate interaction events. Add parameters such as form identifier, landing page, service, language, or a stable inquiry reference, while avoiding personal data that Analytics should not receive. Mark the real lead event as a key event after agreeing its definition. Test with a labeled internal submission: use Tag Assistant or debug mode, follow the form, and inspect Realtime and DebugView for the event and parameters. Verify the saved record and notification in the CRM or inbox, deduplicate retries, and filter developer traffic. DebugView has limited attribution; use Acquisition reports for the most accurate view. A click is diagnostic evidence, not proof that the inquiry arrived.

Use a fixed prompt panel for competitor reporting

A repeatable panel should describe the buyer’s task instead of asking whether a brand is famous. One sample prompt is: “For a company choosing a studio for a bilingual website in [market], compare [brand or studio] with three relevant alternatives. Which providers can show evidence for naming, web design, development, SEO, language support, remote delivery, ownership, and a working inquiry path? Cite supporting pages, say when evidence is unavailable, and separate official evidence from your inference.” Keep wording, market, language, device, signed-in state, and products or surfaces constant. Save the exact prompt, response, date and time, engine and mode, locale, whether an AI feature appeared, named providers, citation URLs, and a screenshot or export where allowed. Preserve no-feature and no-mention runs. For a hypothetical panel, imagine 20 runs, 12 with a rendered answer, and 8 without. If three of the 12 mention Dardo, sampled mention rate is 3 divided by 12, or 25 percent. If those answers contain 20 total provider mentions and Dardo accounts for three, sampled mention share is 3 divided by 20, or 15 percent. These are arithmetic examples, not observed results. Report no-answer runs separately and never call this the share of all AI searches.

Score factual accuracy with a dated source of truth

Create a source-of-truth table before scoring an answer. Give each claim a date, an official URL or internal record, an owner, and a status for services, prices, office locations, named clients, credentials, languages, delivery models, and contract terms. Split an AI response into checkable claims and label each supported, contradicted, or unverifiable. Then report two ratios: verified accuracy equals supported claims divided by supported plus contradicted claims; verifiability equals supported plus contradicted claims divided by all checkable claims. A hypothetical answer with five supported claims, one contradicted claim, and two unverifiable claims has 5 divided by 6, or 83.3 percent, verified accuracy, and 6 divided by 8, or 75 percent, verifiability. It should also carry a severe-error flag for a false office, invented client, materially wrong price, or misrepresented service. Do not treat an unverifiable statement as correct merely because it sounds plausible. Record the exact wording, the source checked, the reviewer, and the correction needed. Score English and Spanish outputs independently when the claims or wording differ. These ratios are proposed editorial measures, not Google-defined scores; their value comes from consistent definitions and retained evidence.

Choose tools by evidence and set a cadence

Select a monitoring tool only after checking what it preserves. Useful criteria include repeatable prompts, country and language control, separation of AI Overviews, AI Mode, and other assistants, raw response and citation URL export, timestamps, change history, permitted privacy terms, and human review. A spreadsheet or small database can be enough for a low-volume panel; Search Console exports, GA4 reports, and CRM records add first-party evidence. A vendor dashboard may organize work, but an estimated score is not a Google impression or an observed lead. Do not endorse a tool that has not been tested against your prompts, locale, plan limits, and retention needs. Run a fixed baseline, repeat after meaningful page or source changes, and publish the denominator with every rate. Reconcile visits with GA4 and verified inquiries with the CRM. Referrer loss leaves some AI-influenced journeys unattributed, while Google’s AIO and AI Mode data cannot be isolated as a GA4 AI Assistant channel. A trustworthy report says what was observed, inferred, unavailable, and decided.

Sources & further reading

Sources behind this guide, with further detail from the original publishers.

How we publish these guides

Have a project in mind?

Let’s talk