How to measure GEO: citation rate, share of voice and mention
Three metrics to tell whether the GEO work is paying off: citation rate, share of voice and mention lift, with a methodology you can run in-house.
There are three metrics that matter in GEO: citation rate (how often you are mentioned), share of voice (your share against the competition) and mention lift (how it moved after an initiative). This piece describes the methodology for running each one in-house, prompt set, cadence, runs, handling variability, with no paid tools. The same three metrics we use with our clients.
- three metrics answer 90% of the question "is this working?"
- a fixed prompt set (30 to 100 prompts) is the foundation; without it there is no valid comparison
- measure across the four main engines (ChatGPT, Claude, Gemini, Perplexity) in parallel
- monthly for a baseline; weekly during an active optimisation phase
- 3 to 5 runs per prompt plus temperature zero keeps variability to a minimum
Why there are three metrics, not one
Each metric answers a different question. Citation rate answers: do we appear? Share of voice answers: do we appear compared with the competition? Mention lift answers: has it improved since we started?
In isolation, any of them lies. A citation rate of 60% looks good, until you find the competitors are at 80%. A share of voice of 40% looks good, until you find the whole market is collapsing. A positive lift looks good, until you realise the base was zero.
The three together give the picture. That is how we measure.
The foundation: the prompt set
It all starts with a fixed set of prompts that represents how real buyers research the sector. Without a stable prompt set, the metrics change because the input changed, and you lose the ability to compare between periods.
Good rules when building it:
- 30 to 100 prompts. For B2B SaaS in Portugal, 50 covers the relevant ground.
- A mix of intents. Comparison (best X for Y), evaluation (is X good for Y?), discovery (who offers X), technical (how X solves Y).
- Natural language. Do not imitate a Google query (short, keyword-heavy). Buyers in ChatGPT write paragraphs.
- In Portuguese and English. Portuguese B2B buyers switch language depending on the technical depth.
- Fix it and version it. The prompt set is treated like code: versioned, with a change log and an update date.
Metric 1: citation rate
The question: in what percentage of the set's prompts are we cited?
Formula: (prompts where we appear) divided by (total prompts in the set), times 100.
An important operational definition: being cited means the brand name is explicit in the answer. If the engine mentions a product feature without naming you, it does not count. If you appear in a list, it counts once (not per occurrence). If you appear with a URL, it counts, but track it separately from the plain mention, because the signal is different.
A typical baseline for a Portuguese B2B SaaS company with no prior GEO investment: close to 0%. After 3 to 6 months of consistent work, 30% to 60% is reasonable depending on how competitive the niche is.
Metric 2: share of voice
The question: of the category's mentions, how much is ours?
Formula: answer by answer. In each answer, count the distinct brands named, yours and your direct competitors. If yours is among them, it takes 1/n of that answer, where n is the number of brands named. At the end, average those fractions over the answers that name at least one brand.
An example with three answers. The first names you and one competitor: you take 1/2. The second names you and nine others: you take 1/10. The third names nobody: it does not enter the account at all. Share of voice is the average of 0.5 and 0.1, which is 30%.
Why this and not mentions divided by total mentions. Appearing in a pair and appearing in a list of ten are not worth the same to whoever reads the answer, and the sum of mentions treats them as if they were. The average per answer keeps the weight of each appearance, and it makes the number comparable between categories that answer in short lists and categories that answer in long ones.
To calculate it, you need a list of 5 to 15 direct competitors. Do not include brands that are not real competitors, even when they show up (AWS or Salesforce in generic answers do not count towards your share of voice in your category).
This metric is the hardest to move, and the most informative. A high citation rate with a low share of voice means you are appearing, but among many others when the user reads the answer. For B2B, where the recommendation is decisive, that is still weak.
Metric 3: mention lift
The question: how much has our citation rate risen since the baseline?
Formula: ((current citation rate) minus (baseline citation rate)) divided by (baseline citation rate), times 100.
If the baseline is 0% (a common scenario), treat it as absolute percentage points. Risen from 0 to 28%? Lift equals plus 28 points.
This is the metric for reporting progress. Take care to put it next to share of voice at the same time; otherwise the lift can be illusory (the whole market rose, the relative position stayed where it was).
The operating method
The typical monthly process:
- Clean sessions. Each prompt goes in a new chat, with no history. In some engines that means opening a private window.
- Temperature zero. Wherever the API allows it (Claude, Gemini, ChatGPT through the API). It reduces variability.
- 3 to 5 runs per prompt. Use the mean or the mode of the mentions. For SaaS with few competitors, 3 is enough; in very competitive sectors, 5.
- Four engines in parallel. ChatGPT, Claude, Gemini (preferably in AI Mode), Perplexity. Each has its own citation rate.
- Structured records. A spreadsheet or tool with columns: engine, prompt, cited (yes/no), position in the list, competitor brands mentioned, run, date.
- Monthly analysis. A month-on-month comparison of the three metrics, per engine. Trends by quarter.
What not to measure (yet)
Some metrics look attractive but add no reliable signal in 2026:
- Sentiment analysis of the mentions. The models tend to be neutral or positive. The differentiation is weak.
- Click-through on the cited URLs. The engines still do not expose reliable origin analytics (Perplexity is the partial exception). Expect 12 to 18 months.
- Volume of AI-origin traffic. Referrer headers are inconsistent. Estimating is speculation.
Sticking to the three core metrics until the others mature is the disciplined path.
Tools: by hand or paid
To get going, by hand is enough:
- The prompt set in a versioned spreadsheet.
- A manual run with four browser tabs, one per engine.
- Annotation straight into the sheet with a check (yes/no) and a list of the brands mentioned.
Coverage: 50 prompts times 4 engines times 3 runs equals 600 interactions. One person does it in half a day, once a month. It is repetitive, but it gives a faithful picture at low cost.
When paying for a tool makes sense: when the prompt set grows past 100, when the cadence has to be weekly or daily, or when the team wants to save operational time. That is where tools of the Profound, Otterly or Peec kind fit: they take the execution work, not the decision work.
Reporting internally
A one-page report. Top: the three metrics, with the month-on-month change. Middle: the prompts where we lost position (and why: competitor X rose, engine Y changed behaviour). Bottom: 1 to 3 actions for the following month.
Resist the temptation of pretty slides with 20 charts. For most teams, three well-understood numbers are worth more than complex dashboards.
Frequently asked questions
Can I measure GEO with ChatGPT alone?
You can, but it is not enough. Each engine has its own dataset and behaviour. Measuring only in ChatGPT gives a partial snapshot; ideally cover ChatGPT, Claude, Gemini and Perplexity in parallel on the same prompt set.
How often should I measure?
Monthly for a baseline. Weekly if you are running an active optimisation initiative (you want to see the curve). Daily only makes sense during the roll-out of something in an engine (a new feature launch, say); otherwise it is noise.
How many prompts in the set?
Between 30 and 100. Fewer than 30 gives a weak sample; more than 100 without real market volume behind it is overengineering. For Portuguese B2B SaaS, 50 prompts cover the ground well.
How do I handle variability between runs?
Averages of 3 to 5 runs per prompt, constant temperature (ideally 0 in the models that allow it), clean sessions (with no prior history). Variability drops but does not disappear: it is a property of the models.
Read next
- Knowledge vs augmented, why these metrics have to be measured twice, with search on and off, and what the difference reveals.
- Where the AI learns about your brand, the same ground from the authority side.