How to verify an AI visibility vendor's evidence before you sign
A practical checklist for procurement teams: how to tell real AI-platform observations from simulated or modelled data, what evidence fields to demand per run, how denominators and statuses should be defined, and the red flags that mean walk away.
Version 1.1 · Last updatedPublished · Miaowa GEO research team
AI visibility reports are easy to make look precise and hard to make verifiable. A mention rate of 27% means nothing until you know how many prompts, on which platforms, from where, how often, and how the failures were counted. This checklist is what we ask of ourselves and what you should ask of any vendor, including us.
1. Ask where the answers come from
| Source | What it proves | What to ask for |
|---|---|---|
| Web interface, controlled browser | What a user in that region and account state would have seen | Screenshot per run, region, whether the session was logged in, browser and time |
| Official or partner API | What the model returned to an API call, which can differ from the web product | Which API, which model version, whether web search was enabled, and a statement that API results are labelled as such |
| Modelled from a prompt database or clickstream | Estimated demand or aggregated behaviour, not your brand's actual answers | A clear separation between modelled numbers and observed answers in every report |
None of the three is wrong; mixing them without labels is. A report that shows a modelled prompt volume next to an observed mention rate should say which is which on the same page.
2. Demand these fields for every run
- A stable task or run ID that you can quote back to the vendor.
- Timestamp, platform, entry point (web or API), region and language.
- The prompt text and its version; a frozen prompt set with a version number.
- Answer text in full, not a summary, and a screenshot where the platform has a visual interface.
- Cited URLs with the position they appeared in.
- Status: answer, refusal, not triggered (for Google AI Overviews), failed, blocked.
3. Check the denominator
Write the formula down with the vendor. For a mention rate: which statuses count as observations, whether refusals are in the denominator, whether failed runs are excluded, and whether one prompt run three times counts as three observations or one. Two vendors can report 20% and 35% from the same answers by choosing differently. Neither is lying; only one has told you the basis.
4. Ask how failures are handled
- Are retries under the same task ID, and do they consume your prompt quota?
- Are persistent failures shown as failures, or silently dropped?
- Is there a per-cycle evidence completeness figure?
- Does the next cycle re-ask automatically, or does someone have to notice?
5. Test re-measurement
Ask for the same prompt set to be run again after a defined change, on the same platforms, from the same region, with the same account state. If the vendor cannot say what stayed constant between the two runs, the difference between them is not attributable to your change. A before-and-after that changed the prompts in between is a story, not a measurement.
6. Look at what is public before you look at the demo
- A methodology page that states prompts per platform, cadence, runs per prompt and status vocabulary with a version date.
- A capabilities page that says what is not supported (exports, API, seats) rather than only what is.
- A public sample with real task IDs, timestamps and statuses, including failed and not-triggered runs.
- Legal entity, filings, privacy policy, terms, and ideally a DPA and subprocessor list.
- A corrections policy: how the vendor handles being wrong.
7. Red flags
- Guaranteed rankings or guaranteed inclusion in AI answers. Platforms do not sell positions and do not publish ranking positions.
- Offers to seed forums, plant fake citations or manipulate the sources an assistant retrieves from. Google has stated that pursuing unearned mentions is treated as spam, and it is a reputational risk for your brand.
- Percentages with no denominator, or dashboards whose numbers cannot be traced to a run.
- Screenshots of the vendor's own dashboard presented as evidence of your brand's visibility.
- Case studies with no client authorisation, no time window and no sample size.
8. What we do and do not claim
Miaowa GEO samples six overseas platforms through their web interfaces in controlled browser environments, every 72 hours, 50 prompts per platform, once per prompt per platform, and stores answer text, screenshot, sources, time, region and language per task. Failures are excluded from denominators and shown as failures. We do not promise inclusion, rankings, traffic or revenue, and we have not been audited by an independent third party; our public sample at /en/report and our methodology page are the evidence we can offer. Apply this checklist to us as strictly as to anyone else.
- How many prompts are enough to trust a mention rate?
- Enough that one changed answer does not move the rate by several points. With 44 unbranded prompts on one platform, one answer is about 2.3 points; across three platforms it is under one point. Treat single-cycle differences below that as noise and look at the series.
- Should I run my own spot checks?
- Yes. Ask the same prompt yourself from a comparable region and compare with the vendor's screenshot. Differences are normal; systematic differences are not.
- What if the vendor refuses to share raw answers?
- Then you are buying a score, not a measurement. Decide whether a score is what you need.
How to read this guide
Competitor facts are restated from their public pages on the check date and may have changed since. Miaowa GEO figures follow the basis on the methodology page. Report an error through the corrections policy. All guides.
