If you already pay for SEO, an AEO proposal creates a practical budget question: what additional work would you be buying?
Before signing a recurring contract, ask the agency to identify the uncovered work, show how it will establish a baseline, and explain which decision the evidence will support. If the scope is useful but the payoff is uncertain, start with a bounded pilot. Renew only when there is a specific next investment worth making.
AEO means answer engine optimization: work aimed at improving visibility in AI search experiences. An engagement might investigate inaccurate product descriptions, missing citations, or referrals that lead to qualified inquiries. Each calls for different evidence. The label alone tells you little about what you should buy.
Identify the work your SEO agreement leaves uncovered
Put the proposed statement of work beside your existing SEO agreement. For every deliverable, identify the affected pages or systems, the person responsible, the acceptance criteria, and whether you already pay someone to do it.
The disagreement over AEO pricing has a useful tension. In an April 2026 r/SEO discussion, the original poster argued that good SEO made separate AEO work unnecessary. Commenter u/RyanJones argued that additional tools and measurement time could still cost money even when the underlying practices overlapped. These are competing practitioner judgments, not evidence that either approach will work for your company. The discussion
Google’s guidance supports overlap within its own platform. It describes optimization for its generative search features as SEO and says those features rely on its core Search ranking and quality systems. Google Search does not use llms.txt files, and no special schema.org markup is required for generative AI search. That guidance concerns Google Search; it does not establish how every assistant works. Google’s generative AI optimization guide
An agency proposing internal-link repairs or clearer product documentation may be recommending useful work. The fee should buy additional capacity or cover a responsibility that nobody owns. Familiar work can deserve funding; an existing commitment should not quietly become a second charge.
Measurement can also be additional work. Selecting buyer questions, collecting answers, checking product accuracy, and investigating changes take time. Ask for setup, collection, analysis, and implementation costs to be visible in the proposal. That lets you compare hiring an agency with assigning the same work to your existing SEO provider or team.
A deliverable such as “monthly AI visibility report” leaves too much unresolved. Specify the products and questions it covers, the evidence you receive, and who acts on the findings. A dashboard without an implementation owner may leave your team paying to rediscover the same problem every month.
Inspect the baseline behind the visibility score
The baseline should precede site changes. Someone outside the agency should be able to inspect the answers and repeat the procedure, even if the resulting answers differ.
Start with questions that reflect your buyers’ actual decisions. Keep questions naming your company separate from questions where a buyer is still discovering vendors. A strong showing on prompts that already contain your brand does not establish discovery among people who have never heard of you.
For each observation, require a record of:
- The exact prompt and its reason for inclusion.
- The product, interface, search mode, and model when visible, including whether collection used a consumer interface or an API.
- Collection time, country, language, and relevant account settings, including whether the prompt began a new conversation.
- The complete answer, cited URLs, and factual errors about your product.
- Missing responses, failures, retries, and exclusions from the reported total.
Personalization limits what a baseline can represent. Google says AI Mode can use previous searches and saved activity to tailor responses when the relevant settings are enabled. A founder’s personal account therefore should not stand in for the whole market. Google’s AI Mode personalization documentation
Use consistent settings for baseline and follow-up observations. Report exploratory buyer contexts separately, and record settings you cannot control. Repeat collection across dates and retain the underlying answers. One favorable screenshot demonstrates that an answer appeared under those conditions. It does not establish how often buyers see it.
For any reported rate, require the numerator, denominator, and exclusion rule. A mention rate among completed answers should appear beside the number of scheduled observations and failed runs. Keep the core prompt set stable; if it changes, compare the unchanged subset or establish a new baseline. Otherwise, a better score could reflect easier questions.
Our guide to building an AI visibility prompt set that does not favor your brand helps you review the questions before approving a monitoring budget.
Keep mentions, citations, visits, and leads separate
Agree on definitions before the first report arrives. Otherwise, an agency can report an improving number while your team assumes it means something else.
| Measure | What to require |
|---|---|
| Brand mention | The answer names your company; record accuracy and context. |
| Citation | The answer links to a source; distinguish your domain from another site discussing you. |
| Referral visit | Analytics records a visit attributed to the relevant source. |
| Qualified lead | An inquiry meets your agreed customer-fit and buying-interest criteria. |
A citation can support a factual explanation without recommending your product. A visit can end without an inquiry. A qualified inquiry still does not prove the agency caused it.
Available platform reports can supplement the agency’s sampled answers. Google’s Generative AI performance report covers link impressions in AI Overviews and AI Mode, with dimensions including pages, countries, dates, and devices. Its documentation also lists access and impression-threshold limitations. Confirm what is available for your property before making it a required pilot input. Google’s report documentation
Microsoft’s AI Performance announcement describes citation reporting across Copilot, AI summaries in Bing, and selected partner integrations. Citation counts do not establish ranking or placement, and grounding-query data is a sample. This coverage should not be presented as visibility across every AI product. Microsoft’s explanation of AI Performance
Keep platform impressions, platform citations, and the agency’s prompt sample in separate reporting sections. Their populations and counting rules differ; a combined “AI reach” total would hide those differences.
For business outcomes, request the attribution rule and its gaps. A visitor who returns later through another channel may not retain the original source. Buyer self-reports can add context, but should remain distinguishable from recorded referrals. Unknown attribution should stay unknown.
Give the pilot a question, an owner, and a spending limit
An agency can commit to completing work and delivering inspectable evidence. Treat a guarantee that an independent platform will recommend your company as an unsupported promise.
A useful pilot could investigate whether inaccurate descriptions of an integration reflect missing or ambiguous documentation. The agency would preserve the incorrect answers, identify the relevant documentation, propose specific corrections, implement approved changes, and repeat the agreed observations. That makes the work reviewable without assuming every subsequent improvement came from those edits.
Choose an evaluation window that allows for implementation and observation, and ask the agency to explain its choice. It should not be presented as a universal deadline for search systems to respond. Track concurrent product launches, site migrations, publicity, and changes to the prompt set. Where practical, retain comparable unchanged pages to strengthen the comparison. A simple before-and-after increase still provides limited evidence of causality.
The spending limit should include agency fees, monitoring tools, content or engineering work, and your team’s review time. Start with coverage that could change a budget or content decision. More prompts are worth paying for only if their answers could affect what you do next.
If leadership wants an AI visibility initiative, agree on the question and review date before commissioning the dashboard. That gives the team a concrete commitment to evaluate while leaving room to stop.
Decide whether the next engagement earns its cost
Completing a pilot and earning a renewal are separate decisions. The agency may deliver everything promised and discover that there is little useful work left to do.
For the integration example, corrected documentation may close the task. Persistent inaccuracies might justify a narrowly scoped investigation if the agency can explain what it would examine next. If observed answers improve but qualified inquiries remain sparse, continued monitoring is a learning expense. Decide whether the remaining uncertainty is worth the proposed cost.
Before renewing, ask the agency to name the next action, the evidence supporting it, its owner, and its full cost. Revise the scope when the evidence reveals a different problem. Stop or defer when the report cannot be inspected, the work duplicates an existing commitment, or no finding would change your team’s next decision.
For your next agency call, annotate one proposal: mark every line as already covered, additional work, or unclear. Ask the agency to resolve the unclear lines and attach a sample evidence export. Add the pilot question, budget cap, review date, and renewal condition to the same document. You will have a scope you can compare and a price you can judge.



