Measure GEO with a fixed prompt set of 30 to 100 buyer questions, re-run monthly across ChatGPT, Claude, Gemini and Perplexity under consistent conditions, recording mention rate, competitor share and factual accuracy. Single checks prove nothing because outputs are non-deterministic.
Why the usual instruments do not work
Search Console reports impressions and clicks because search engines log a query, a result set and a click. Assistants do not expose any equivalent. There is no impression count, no position, no query report and no click data unless the user follows a link — and frequently there is no link.
So GEO measurement means building your own instrument. That sounds heavier than it is: the method is straightforward, and the discipline of running it consistently matters far more than sophistication.
Building the prompt set
The prompt set is the measurement instrument, so it needs designing once and then leaving alone. Changing it between runs destroys comparability, which is the most common mistake.
- Source real questions. Sales call recordings, inbound emails, Search Console queries and support tickets. Not keyword tools, which produce phrasing nobody uses conversationally.
- Cover the funnel. Category definition questions, comparison questions, vendor recommendation questions and problem-specific questions.
- Include recommendation prompts explicitly. ‘Which vendors should I consider for X’ is the prompt that matters most commercially and is often missing from prompt sets.
- Fix the size. Thirty is the practical minimum for signal; beyond about a hundred the manual effort per run stops being worth it.
- Freeze it. Version the set and change it only at planned intervals, noting the change in the report.
What to record per prompt
| Field | Why it matters |
|---|---|
| Mentioned (yes/no) | The headline metric — mention rate across the set |
| Position in answer | Named first carries more weight than named fifth |
| Competitors named | Share of voice against your actual rivals |
| URL cited | Whether the assistant linked a source and which page |
| Accuracy | Whether what was said about you was correct |
| Sentiment | Whether the framing was positive, neutral or negative |
Accuracy is the field clients care most about once they see it. Inaccurate descriptions are common, and in our experience they usually trace back to contradictions the company published itself rather than to external misinformation.
Handling non-determinism honestly
The same prompt asked twice can produce different answers. Results vary by user, region, session, account history and model version. This is a property of the systems, not a measurement flaw, and pretending otherwise makes reporting indefensible.
Three practices manage it. Run under consistent conditions — same account state, same region, logged out where possible. Accept that a single run is a sample rather than a fact. And report trends across months rather than point comparisons, because month-to-month noise is real and a two-point movement means nothing.
Reading assistant referral traffic
Some assistant traffic is visible in analytics. Segment GA4 referral sources for assistant domains and treat that as a floor rather than a total, because a meaningful share of assistant-influenced visits arrive with no referrer at all.
There is also a second, larger effect that is easy to miss: buyers who receive a recommendation frequently type the brand name into Google rather than clicking a link. That reports as branded organic search. A rising branded search trend alongside a rising mention rate is the strongest available evidence that GEO work is producing commercial effect.
Reporting it to people who want certainty
GEO reporting will be read by people accustomed to Search Console precision, and the honest position is that this data is directional. Say that explicitly and early rather than letting it be discovered.
What makes the reporting credible is consistency of method: same prompt set, same conditions, same cadence, changes documented. A directional measure applied rigorously is far more useful than a precise-looking number derived inconsistently — and it is defensible in a room where someone asks how you know.