← Back to the journal

How to measure AI citations without fooling yourself

Build a defensible prompt sample, preserve the answers and separate citations, brand mentions, factual accuracy and referral visits.

Orange threads connect three source cards through a glass prism to one answer card.
Conceptual editorial artwork · Generated with AI for FindVex

To measure AI citations, choose a stable set of relevant questions, collect answers under recorded conditions and inspect their visible source links. Report the number of observed answers, the questions you asked and the collection period alongside any percentage.

Keep four outcomes separate: a brand being mentioned, a page being cited, an answer describing the brand accurately, and a person visiting your site. None automatically proves the others.

This guide gives a small team an observation protocol. It is a proposed method, not a completed FindVex study. The examples and numbers below are hypothetical. A sample of answers cannot reveal every user’s experience or all the material a system used internally.

Decide which question the measurement should answer

Start with the decision you need to make. For example:

  • Do unbranded questions about our subject ever surface our pages?
  • Which of our pages appear as visible sources in the selected interfaces?
  • Do answers describe our offering correctly?
  • Does observed source visibility coincide with identifiable referral activity?

Josh Blyskal’s approach to measuring AI visibility emphasizes a defensible prompt panel, unbranded discovery questions and comparisons across the same panel. He distinguishes inclusion, position where an answer has meaningful ordering, visible citations and accuracy. The lesson we apply here is to define the sample before interpreting its score.

A prompt that already names your company tests recognition or accuracy. It does not test whether an unfamiliar buyer would discover you. Keep branded checks in a separate group and report them separately.

Build a small prompt panel around real tasks

For an English-language marketing publication serving US founders, the following are illustrative prompts, not evidence of how often people ask these questions:

Group Example prompt What the answer could reveal
Discovery: first steps How should a US SaaS founder choose their first SEO topics? Whether a relevant guide appears without a brand cue
Discovery: method How can a small team track whether AI answers cite its articles? Which measurement sources the interface presents
Discovery: distribution How can an AI startup find its first readers with four hours a week? Whether sources address the actual constraint
Branded accuracy What does FindVex publish, and who is it for? Whether the brand description matches its public pages

For every prompt, save the buyer task, where the wording came from and why it belongs. Customer questions or documented support conversations are stronger starting points than an unexplained list of generated phrases. Do not claim a synthetic prompt panel represents market demand.

Freeze a version of the panel for a collection round. If you add a new subject later, record the change and report its cohort separately. Otherwise a score may rise simply because the new questions are easier for your site to appear in.

A manageable pilot might use 10 unbranded prompts, 3 interfaces and 3 runs per prompt: 90 attempted observations. Ten is an operating choice in this example, not a statistically representative sample size. Keep branded accuracy checks outside those 90 observations.

Record the conditions and preserve each answer

Use the same prompt text in each repeated observation. Start a new conversation for each run, record the account state and locale, and avoid follow-ups that explicitly request your brand. A fresh conversation reduces one source of variation; it does not remove personalization or platform changes.

For each attempt, capture:

  • Date and time in UTC, interface and any visible model label.
  • Account state, language, country or locale setting, and relevant search mode.
  • Prompt ID, exact text, panel version and run number.
  • Whether a usable answer appeared; record refusals, errors and missing answers separately.
  • The full answer or a retrievable screenshot, plus every visible cited URL.
  • Brand mention, answer accuracy and the claim associated with a citation.

If the model label or web-search behavior is not visible, write unknown. Do not infer that a system browsed just because the answer seems current. Keep consumer-interface observations separate from API experiments: the interface and collection conditions differ.

In the supplied worksheet, repeat the observation_id when an answer has several cited URLs, with one row per URL. If the answer has no citations, retain one row with a blank URL. Count distinct observation IDs when calculating answer-level rates; otherwise a heavily cited answer gets counted several times.

Calculate rates with an explicit counting rule

Choose your owned domain set before counting. Decide how redirects, subdomains and URL parameters will be handled, and retain the original URLs so someone can audit those decisions.

Use three basic measures:

Measure Calculation What it describes
Brand-mention rate Usable answers mentioning the brand ÷ usable answers Presence in the sampled answer text
Owned-citation answer rate Usable answers with at least one owned-domain citation ÷ usable answers How often the sample visibly links to your pages
Owned citation share Owned-domain citation units ÷ all citation units Your share of visible sources under the stated counting rule

For this protocol, a citation unit is a distinct normalized URL within an answer. Repeated references to the same URL in that answer count once. If it appears in another answer, it counts once there too. This is a reporting convention, not a platform standard.

Here is a worked calculation using invented data:

A pilot attempts 90 observations. Six fail, leaving 84 usable answers. Twelve mention the brand; nine cite an owned page. The answers contain 120 distinct within-answer citation units, including 15 owned units.

Brand-mention rate: 12 / 84 = 14.3%.

Owned-citation answer rate: 9 / 84 = 10.7%.

Owned citation share: 15 / 120 = 12.5%.

Report the 6 / 90 failed attempts alongside these figures.

A usable answer with no sources remains in the answer-rate denominator. A failed request is not silently converted into a zero-citation answer. If you filter for a subset, publish the filter and its denominator.

Break the report down by interface and prompt group before presenting an overall number. Use the same questions and counting rules for competitor comparisons. Repeated answers to a small panel are correlated observations; the three runs do not create three independent markets or justify a universal confidence claim.

Inspect whether the citation supports the answer

Open the cited page and locate the passage that supposedly supports the claim. Record whether it supports the claim fully, partly, contradicts it, or cannot be assessed. A visible link is observable; the entire hidden retrieval process is not.

Consider this hypothetical error: an answer says “FindVex sells an AI visibility platform” and cites our About page. The citation counts as an owned citation, but the statement is inaccurate because FindVex is currently a publication.

That row should retain both facts: citation present, claim contradicted by the source. Removing it would flatter the report. Counting it as unqualified success would conceal a reader problem.

Inspect the source page for ambiguity first. Clarify an unclear description where needed, record the change and collect a later sample under the same protocol. A later corrected answer is an observation; it does not prove that your edit alone caused the correction.

Choose improvements by usefulness, not citation totals alone

Aleyda Solis’s content-prioritization framework separates reasons to visit, citation potential, brand mentions, business value, original contribution and effort. We use those distinctions to avoid giving every page the same job.

For a small publication, a useful backlog entry should name the reader’s task, missing evidence, intended outcome, effort and review date. For example:

  • A confusing reference table may need clearer definitions and a dated method so readers can check it.
  • A tutorial may need a usable worksheet that gives someone a reason to visit and apply the answer.
  • An inaccurate brand description may require clearer public information before another new article.

These are editorial hypotheses to test. A numerical prioritization score is not the probability of being cited.

Keep referrals and search reporting separate

If your analytics records a referral from an AI interface, report it as observed traffic under that attribution system. It is not the total number of citations, and a missing referrer does not prove there were no AI-assisted visits. Do not turn a sampled citation rate into an estimate of visitors.

Google says traffic from its AI features is included within the Web search type in Search Console. Those aggregate figures should not be presented as a complete inventory of AI citations. Its AI features documentation also says no special AI schema is required for inclusion.

Download the AI citation observation sheet. Start by filling in the panel and conditions, collect a small sample you can inspect, and write down the limits alongside the results. A report is useful when another person can understand how you reached it.

For the work before measurement, see our first-month SEO guide and distribution experiment for founders.


Method note: The recording fields, counting convention and numerical example are FindVex’s proposed protocol. The practitioner methods are attributed above; they are not claims that the authors reviewed or endorsed this guide. Revised September 26, 2026 to add prompt cohorts, counting rules, failure handling and an accuracy review.