← Back to the journal

Run one AI visibility experiment without changing everything

Test one content edit with fixed buyer questions, comparison pages, repeated observations, and a worksheet for deciding what to try next.

An experiment worksheet connects fixed buyer questions to one orange-highlighted CSV import edit and an unchanged comparison page. Repeated observations lead to three open choices: repeat, stop, or inconclusive.
Conceptual editorial artwork · Generated with AI for FindVex

If you rewrite a product page, add structured data, and launch a link campaign in the same week, a new AI citation tells you little about which action helped. You have something to record, but little guidance on what to repeat.

Start with one defined edit, a fixed set of buyer questions, and a decision rule written before collection begins. For a small SaaS team, the decision can be modest: whether to try the same edit on a second page.

This workflow uses a content edit and Google AI Overviews as its example. It can reveal a pattern worth testing again, but it cannot establish that the edit caused a citation gain. Apply the measurement structure separately to other search products, with their own eligibility checks and observations.

Define the change and the decision together

Choose one page that answers a real buyer question. Find a specific weakness you can correct without rebuilding it, such as missing conditions for importing data or an unclear explanation of a supported integration.

Write a hypothesis:

Adding a verified explanation of [capability] to [page] will increase how often that URL appears as a supporting source for [fixed question group], enough to justify [next small test].

Specify the passage and its replacement. “Improve the integration section” leaves too much room for interpretation. “Replace the integration paragraph with supported inputs, setup steps, and known limitations” gives you an edit you can inspect later. Treat that replacement as one package; the test will not tell you which part of it mattered.

Save the original and revised passages. Keep a log of other changes that could affect the observation window, including migrations, navigation updates, publicity, and product launches. Make necessary fixes, then record them and reassess whether the test remains interpretable.

Testing one change at a time is slow and can miss interactions between changes. It suits a team deciding whether a useful edit deserves another test.

Choose comparable questions before collecting results

Build a small question set from buyer conversations, support requests, or relevant public discussions. Keep the original wording and context beside the query you derive from it.

For a first test, six target questions and six comparison questions can keep manual collection manageable. These counts are workload choices with no statistical precision guarantee.

Target questions should concern the edited page. Comparison questions should concern unchanged pages with similar intent. A specific setup question is a poor comparison for a broad “best software” query. Use the baseline to assess whether the groups also have reasonably similar patterns of AI Overview appearances and citations. If the comparison is unsuitable, revise the design and collect a new baseline before editing.

Assign each question a designated URL before collection. Target questions can all map to the edited page; each comparison question should map to the unchanged page you expect to answer it. This prevents counting any convenient citation to your domain as a success.

Keep branded questions separate from unbranded discovery questions. Asking whether your named product supports an integration tests a different situation from asking which products support it. Use the guide to building an AI visibility prompt set that does not favor your brand to check your selection.

Freeze the wording before the baseline begins. Save new questions for a later test. Comparison questions remain imperfect controls: the edited page might begin appearing for them, too. Record that crossover because it weakens the comparison.

Establish eligibility and a repeatable baseline

For Google AI Overviews and AI Mode, supporting pages must be indexed and eligible to appear in Search with a snippet. Google requires no special schema.org markup for these features. Check those basics before treating absent citations as a content problem. Google’s AI features documentation

Choose one surface and keep it fixed. Google says AI Overviews and AI Mode may use different models and techniques and show different links. Their results should stay separate. Google’s explanation of AI search features

A practical starting schedule is three collection days per week for two baseline weeks. Run each question once per collection day. Keep the language, intended location, device type, account state, and session procedure consistent. Rotate question order so that one group does not always run first.

Save one record per scheduled observation:

  • Exact question, designated URL, collection time, product surface, and model information when visible.
  • Location, language, device, account state, and session conditions.
  • Answer or screenshot, whether an AI Overview appeared, and its visible source URLs.
  • Whether the designated URL was cited, whether the brand was mentioned, and any inaccurate product claim.
  • Collection errors or departures from the procedure.

An absent AI Overview is a valid observation. A failed page load is missing data. Write a retry policy before collection and apply it consistently; never rerun only disappointing answers. Report scheduled, valid, and missing observations for each group and period.

Set the crawl deadline and follow-up window

After the baseline, publish only the defined edit and record when it became available. Save any evidence of a subsequent crawl.

Google says crawling can take days to weeks, and requesting a crawl does not guarantee inclusion in results. A short period without movement may say little about the edit. Google’s recrawl guidance

Choose the timing rule in advance. For example, wait up to four weeks for evidence of a post-edit crawl. If that evidence arrives, start two weeks of follow-up collection on the next scheduled collection day, using the baseline schedule. If it does not arrive by the deadline, mark the test inconclusive because exposure to the edit is uncertain.

Those intervals are proposed operating rules, not Google deadlines or sufficient durations for every test. A crawl after publication also does not prove that a particular generated answer used the revised passage. Keep that uncertainty in the result, and do not extend collection until a favorable answer appears.

Score citations, accuracy, and business outcomes separately

Use one primary measure for each group:

Page citation rate = valid observations citing their designated URL ÷ all valid observations.

Count the designated URL at most once per observation. Include valid searches without an AI Overview in the denominator. Decide in advance how to handle URL variants, such as fragment links to the same page, and retain the raw links for review.

Also report how often an AI Overview appeared. This helps distinguish more opportunities for citation from more frequent inclusion when an Overview is present. If useful, calculate a separate rate among observations with an Overview and label that denominator explicitly. When none appeared, that conditional rate is unavailable.

Report brand mentions and factual accuracy separately. A citation accompanied by an incorrect capability claim warrants investigation even when the citation rate improves.

Keep referral visits and qualified actions in a separate business-outcome record. Google includes AI feature traffic within Search Console’s overall Web performance reporting. Growth in that total alone cannot establish an AI Overview traffic gain. Google’s measurement guidance

Inspect results by question and collection day as well as in aggregate. Thirty-six observations from six repeated questions do not represent thirty-six independent buyer situations. Missing runs can also change the mix: if one question has fewer valid observations, flag that imbalance before interpreting a pooled rate.

Worked example: clarify a CSV import guide

Suppose a fictional SaaS company has a CSV import guide that omits how duplicate records are handled. The team verifies the behavior and replaces one paragraph with a precise explanation and a short example. It leaves the title, schema, navigation, and promotion schedule unchanged.

Six import questions map to that guide. Six questions about comparable, unchanged export documentation form the comparison group, each mapped to its designated page. With six collection days in each period, each group has 36 scheduled observations before the edit and 36 afterward.

Assume all scheduled observations are valid. These numbers are hypothetical:

Group Before After
Target questions citing their designated URL 6/36, or 16.7% 12/36, or 33.3%
Comparison questions citing their designated URL 5/36, or 13.9% 9/36, or 25.0%

Using the underlying fractions, the target rate rose about 16.7 percentage points and the comparison rate rose about 11.1 points. The difference between those changes is roughly 5.6 points.

Citation frequency increased in both groups, which tempers the apparent target gain. The remaining difference does not prove an effect from the edit. The groups may respond differently to external changes, and repeated answers may be correlated. The table alone also leaves out Overview appearance rates, accuracy, and the distribution across questions and dates.

NIST describes completely randomized designs as randomly assigning factor levels to experimental units. This page test has no random assignment of edits to pages. Its comparison group provides context for a local decision, with limited grounds for a causal claim. NIST’s explanation of randomized designs

If improvement appears across several target questions and dates, accuracy remains acceptable, and no major concurrent change occurred, the team could repeat the edit on another suitable page. If all six additional citations came from one question across the six follow-up days, the finding would apply much more narrowly. Check that question’s relevance before extending the approach to other pages.

Copy this experiment worksheet

Complete the worksheet before the baseline. Attach the question-to-URL map and keep the saved observations beside it.

READER TASK AND CHANGE
Buyer question the page should answer better:
Edited page URL:
Exact original passage and proposed replacement:
Verified product behavior supporting the replacement:
Hypothesis:
Next decision this test will inform:

QUESTION SET AND COLLECTION
Fixed target questions and designated URL:
Fixed comparison questions and designated URLs:
Question sources, original wording, and adaptations:
Branded questions kept separate:
Product surface, market, language, and device:
Account state and session procedure:
Baseline dates and collection schedule:
Publication date:
Evidence required to start follow-up:
Crawl-evidence deadline:
Follow-up start rule, duration, and final review date:

SCORING AND DECISION
Citation definition and URL-variant rules:
Valid absence, collection error, and retry policy:
Scheduled, valid, and missing counts by group and period:
Primary citation rate and Overview appearance rate:
Accuracy checks and material errors:
Concurrent changes and comparison-group crossover:
Pattern required to repeat on another page:
Conditions that stop rollout or make the test inconclusive:
Final decision and remaining uncertainty:

A repeat rule should require improvement across multiple questions and dates, favorable movement relative to the comparison group, and no material accuracy regression. Set any numerical threshold before viewing follow-up results, based on the cost of the next decision. Meeting that threshold does not establish statistical significance.

At review, choose one action: repeat on another page, stop further rollout, or mark the result inconclusive because exposure, missing data, or concurrent changes prevent interpretation. A clearer passage can remain useful to readers even when its visibility effect is uncertain.

Your next task: define one test before editing

Choose one page with an answer you can make more precise. Save the proposed replacement, map your target and comparison questions to their URLs, and complete the worksheet’s timing and decision rules. Begin the baseline only when another teammate could follow those instructions and score the same observations.