Choose your AI visibility prompts before checking whether the answers mention your brand. Start with buyer problems, comparisons, and constraints, then keep branded questions in a separate test.
A prompt asking for alternatives to your product has already supplied its name. Even an unbranded question can favor your product if it repeats the exact feature combination on your homepage. Running either question more often cannot fix that selection problem.
A useful prompt set makes those choices inspectable: whose decisions it covers, where the questions came from, and which situations are missing. Its scores describe that set. They cannot tell you how often all potential buyers encounter your brand without evidence about what those buyers actually ask.
Define whose discovery you want to test
Start with one sentence describing the audience and decision:
We want to observe which solutions appear when US small-business operations managers ask how to reduce manual invoice follow-up.
That boundary excludes enterprise procurement, consumer budgeting, and accounting exam questions, even when they share vocabulary with your product.
Define the experience being tested, too. Record the AI product, interface, search mode, language, test date, and whether each question starts a new conversation. Include visible model information and relevant account settings where available. Results from one interface should stay attached to that interface.
Google says AI Mode and AI Overviews may use different models and techniques, producing different responses and links. Both may also issue related searches through query fan-out. Keep these experiences separate in your records. Google Search Central
Treat your initial set as a coverage benchmark: a defined collection of buyer situations you can revisit. Unless you know how common those situations are among real buyers, leave market-reach estimates out of the report.
Collect buyer situations and label their sources
Gather candidate questions from prospect conversations, lost-deal notes, customer interviews, public discussions, and relevant search queries. Keep the original wording and context alongside any test prompt you derive.
Each source has limits. Support requests come from people who already chose a product; sales conversations come from people who reached your company. Public discussions can reveal unfamiliar language, but their participants do not necessarily represent your market.
Use a consistent evidence label for each candidate:
| Label | What it means |
|---|---|
| Observed prompt | Someone supplied the actual prompt they submitted to an AI product, with permission to use it. |
| Adapted question | You turned a documented question or problem into a test prompt. |
| Hypothesis | Your team or an AI assistant proposed a plausible question without direct supporting evidence. |
A forum question is observed customer language. It remains an adapted AI prompt unless you know someone submitted that wording to an AI product.
If your source pool is thin, use the pain-point evidence sheet to collect situations and their context. Fifty paraphrases of one founder assumption still represent one assumed need.
Allocate room for problems, comparisons, and constraints
Organize candidates by the buyer’s main decision. For a small pilot, try the following allocation. Sixteen prompts is a planning choice, not a statistical minimum.
| Prompt group | Buyer task | Pilot slots |
|---|---|---|
| Problem | Understand how to solve a recurring issue | 4 |
| Category | Find tools for a known job | 4 |
| Approach comparison | Compare ways to do the work | 4 |
| Constraint | Find an option that fits a limitation | 4 |
For invoice follow-up, an approach comparison might ask whether to use accounting-software reminders or a separate receivables tool. A constraint prompt might specify a budget or an integration requirement.
These groups can overlap. Assign each question family to one primary group before selection so that a constrained tool question does not fill two slots. Record secondary attributes in the worksheet if they help you audit coverage.
Include constraints because buyers expressed them. If a constraint is still a team assumption, label it as a hypothesis. Avoid packing every product differentiator into one question, and retain relevant situations where your product would be a poor fit.
Separate questions containing your brand from those naming competitors. The first group tests known-brand evaluation; the second tests competitor-aware discovery. Report both separately from fully unbranded discovery.
Equal slots make coverage easy to inspect. They do not establish that each group represents a quarter of market demand. Report results by group without assigning unsupported market weights.
Select question families before reading the answers
Write the inclusion rules first. Each candidate must fit the target audience, express a distinct decision, and have a traceable source or an explicit hypothesis label.
Group duplicates by underlying need. Three ways of asking for cheaper invoice reminders usually belong to one question family. Choose one wording for the core benchmark and save the variants for a separate check of how wording affects answers.
When a group has more eligible families than slots, draw randomly within that group. For a simple manual draw, put each eligible family ID on an identical slip, mix the slips, and draw without replacement. Save the eligible list and selected IDs. Record why other candidates were excluded, distinguishing an eligibility defect from simply not being drawn.
If a group has fewer eligible families than planned, report the shortfall. Collect more evidence or add clearly labeled hypotheses; do not fill the slots with near-duplicates.
Random selection limits cherry-picking within your pool. It cannot recover audiences or situations that never entered the pool. AAPOR’s survey guidance provides a useful analogy: design the sample around the population of interest and explain the limits of nonprobability sampling. Here, that means documenting which buyer situations could enter your benchmark. It does not make the benchmark a probability survey. AAPOR’s survey research guidance
Freeze the selected wording before running the test. Keep a prompt when it produces no brand mention. Remove it only for a documented defect, such as falling outside the audience or duplicating another family, and log the change in a new version.
Worked example: separating branded and unbranded results
Imagine a fictional startup, LedgerNudge, that helps small service businesses follow up on unpaid invoices. Its founder begins with questions about LedgerNudge and prompts describing its precise feature combination.
The team rebuilds the library using the sixteen-slot proposal. These four illustrative entries show the kinds of questions it could include. They are hypothetical prompts, not customer statements.
| Group | Example prompt |
|---|---|
| Problem | How can a small agency follow up on overdue invoices without writing every email manually? |
| Category | What tools help a five-person service business manage invoice reminders? |
| Approach comparison | Should we use reminders in our accounting software or buy a separate follow-up tool? |
| Constraint | What options work for invoice follow-up if we cannot connect our accounting account? |
The last question belongs in the candidate pool if evidence supports that constraint, even if LedgerNudge requires an accounting connection. Removing it solely because the product cannot win would hide a relevant situation.
Suppose the team runs sixteen unbranded prompts three times each on one AI interface. All runs return completed answers. LedgerNudge appears in 12 of those 48 answers. In a separate test, it appears in all 12 answers to four branded prompts, also repeated three times.
| Test | Answers mentioning LedgerNudge | Mention rate |
|---|---|---|
| Unbranded prompts | 12 of 48 | 25% |
| Branded prompts | 12 of 12 | 100% |
| Combined | 24 of 60 | 40% |
All of these numbers are hypothetical. The combined calculation is arithmetically correct, but some prompts supplied the brand name. Calling the combined 40% an unbranded discovery rate would misdescribe the test.
An accurate report would say: “LedgerNudge appeared in 12 of 48 answers to our unbranded benchmark.” Include results by prompt group and family as well. Neither percentage measures market reach, and a mention may be a warning or rejection rather than a recommendation.
Repeat runs while keeping coverage limits visible
Multiple runs help expose answer variation. Anthropic’s guide to evaluating AI agents distinguishes a task from each trial and describes running multiple trials because model outputs vary. That supports repeated observation, but it does not establish a sample size for AI visibility tests. Anthropic’s evaluation guide
Choose a repetition schedule before starting and keep it consistent across prompts and reporting periods. Three trials, as in the example, are an exploratory budget choice with no precision guarantee.
Save each answer. Classify a brand mention, a recommendation, and a linked citation separately. Also record refusals, technical failures, and cases where no AI answer appears.
Decide how those outcomes affect the denominator before testing. If you calculate a mention rate only among completed substantive answers, show that count beside the total scheduled runs and the excluded outcomes. Do not silently rerun failures until every slot contains an answer.
Repeated answers to one prompt reveal nothing about buyer situations missing from the set. Keep the core stable for comparisons over time, and use a separate exploratory set for newly discovered questions. When the core changes materially, establish a new baseline or report the unchanged subset separately.
Once the library is fixed, use the guide to measuring AI citations for more detailed scoring.
Copy this prompt-selection worksheet
Create one record per question family. Keep shared test settings in a separate record so you can update them consistently.
QUESTION FAMILY
Family ID:
Exact test prompt:
Target buyer and decision:
Primary group: problem / category / approach comparison / constraint
Brand status: unbranded / own brand / competitor named
Source and source date:
Original wording and context:
Evidence label: observed prompt / adapted question / hypothesis
Adaptation made and reason:
Inclusion rule:
Selection outcome and reason:
Related wording variants kept outside the core:
SHARED TEST SETTINGS
Prompt-set version:
Eligible family IDs and selection method:
Selected family IDs:
AI product, interface, and search mode:
Market, language, and test dates:
Visible model information and relevant account settings:
New conversation for each run:
Repetition schedule:
Outcome definitions, denominator rule, and retry policy:
COVERAGE NOTE
Missing audiences or decisions:
Weak source categories:
Decisions supported only by hypotheses:
Unfilled pilot slots:
Before the first run, check that every selected prompt has a reason for inclusion that does not depend on its answer. Attach the coverage note to the results so readers can see what the score leaves out.
Your next task: audit ten prompts
Take ten prompts from your existing visibility test. Separate the branded questions, trace the remaining questions to their sources, and flag wording built around your feature list. Group paraphrases under a shared family ID.
Use the worksheet to identify the largest missing buyer situation. Find evidence for that situation and add it to the candidate pool before spending more on repeated runs.



