Quick answer
An AEO prompt testing dataset is a fixed prompt library run on a schedule with full reproducibility fields, scored separately for mention, citation and accuracy, and cross-checked against Search Console's generative AI performance data. Build the dataset first; make visibility claims only where the rows support them.
- One prompt library across the buyer journey
- Score mention, citation and accuracy separately
- Say what the sample can and cannot support
Most AEO reporting starts at the wrong end. Someone asks an assistant a question, sees the brand named, screenshots it, and a visibility claim enters a deck. Run the same prompt an hour later on a different account and the answer often differs.
The fix is not a better tool. It is a dataset that records enough about each run to be comparable to the next one.
Start with a prompt library, not a keyword list
Buyers ask assistants in sentences, and the sentence changes with the stage. Build a library that covers the journey rather than a list of head terms.
| Stage | What the buyer is doing | Prompt shape |
|---|---|---|
| Problem aware | Naming the symptom | Why is my site getting impressions but no clicks |
| Solution aware | Comparing approaches | Do I need an audit or ongoing SEO work |
| Vendor aware | Shortlisting providers | Who does answer engine optimization in Vancouver |
| Evaluation | Checking a specific firm | What does this agency actually deliver |
| Objection | Looking for reasons not to buy | Is AEO worth paying for yet |
Keep the library small enough to run every cycle without shortcuts. A library you actually rerun beats a larger one you sample from.
Reproducibility fields for every row
Each run is one row. Without these fields a row is an anecdote, not data.
- Assistant and model version, exactly as reported.
- Date and time of the run.
- Region or locale setting used.
- Account state: signed in or out, memory or personalization on or off, any custom instructions.
- The prompt text, verbatim and unedited.
- The full response text, saved rather than summarized.
- Run number, so repeats of the same prompt are grouped.
- The tester's name.
Score dimensions separately
Collapsing everything into one visibility score is where AEO reporting loses its meaning. Keep four columns apart.
- Mention. Was the brand named in the answer text, and in what position within it.
- Citation. Was a page attributed as a source, and which URL.
- Accuracy. Is what the assistant said about you correct, including services, location and claims.
- Competitive context. Which other providers appeared in the same answer.
Accuracy is the dimension teams skip and the one with the clearest business consequence. A confident, wrong description of your service is worse than no mention.
A six-step test protocol
- Freeze the prompt library and the run schedule.
- Set the account state deliberately and record it, rather than testing from whatever session happens to be open.
- Run each prompt the agreed number of times per cycle.
- Save the full response before scoring anything.
- Score the four dimensions from the saved text, not from memory.
- Review the variation across runs before writing any conclusion about change.
Bridging to Search Console
Google says its generative AI performance report completed worldwide rollout on August 31, 2026. It covers AI Overviews and AI Mode link impressions broken out by page, country, date and device. It does not include a query dimension, so it cannot tell you which prompt produced which impression.
That shapes what the bridge can do. Use the report to see which pages accumulate AI-surface impressions over time, and use the prompt dataset to see how you are described and cited. Together they support statements about direction on specific pages. Neither supports a claim about share of recommendations across a market.
Honest reporting language
| Do not write | Write instead |
|---|---|
| We appear in 40% of AI recommendations | Named in 8 of 20 recorded runs of this prompt set in September |
| Our AI citation rate improved | Cited in more runs this cycle than last, from the same library and settings |
| We rank first in AI search | Appeared first in the answer text in these specific runs |
| AI traffic is up | AI-surface link impressions rose for these pages in Search Console |
Why we wrote this one
Between August 11 and September 7, 2026, Google Search Console recorded 189 impressions for "aeo services" at an average position of 15.56 with zero clicks for our own site. One conversational Vancouver query recorded 20 impressions at an average position of 1.00 in the same window. That sample is far too small to support a durable ranking claim, and we are reporting it as an observation only. These are ThinkProfits visibility observations, not market search volume.
Our AEO services page describes the work behind these tests, reporting dashboards shows how the dataset reaches a monthly pack, an SEO audit covers the crawlable foundations these answers depend on, digital marketing strategy is where the prompt library gets prioritized, and our SEO content analyzer helps keep the answer pages themselves tight.
Want an AEO measurement routine you can defend?
Book a free 30-minute consultation and we will map a prompt library and a scoring sheet to your buyer journey.
Book My Free Consultation →
