Why the old metrics stopped working

Rank tracking assumes a stable, ordered list of results for a fixed query. AI answers have none of those properties: they are generated per-conversation, vary between users, and change with the phrasing of the prompt. A single number cannot describe your position because there is no position.

This is genuinely uncomfortable for teams used to a dashboard that goes up. The honest response is not to invent a fake ranking metric, but to measure the things that actually vary — whether you appear, how you are described, and whether you are recommended.

Build the prompt set first

Everything depends on this, so it is worth an hour. Aim for roughly fifteen to thirty prompts spread across three groups: category questions a buyer asks before knowing any supplier, comparison questions asked while shortlisting, and the specific practical questions asked just before committing.

Write them as a buyer would type them, not as keywords. “Who should I use to build a booking system for a clinic” is a prompt; “clinic booking system developers” is a search term, and it produces a different and less useful answer.

Include a handful of prompts you expect to lose. A set you always appear in cannot show improvement, and the ones you lose are where the backlog comes from — the same reasoning as auditing where a competitor wins.

Track appearance rate, not position

Run the set across the engines that matter, and record how often your brand appears at all. Appearance rate over a fixed prompt set is a stable, comparable number, and it moves when your work moves.

Keep the prompt set fixed for at least a quarter. Changing prompts and celebrating an improvement is the most common way teams fool themselves in this discipline — and it is indistinguishable, in a report, from genuine progress.

Track how you are described

Appearing is not the same as appearing well. Capture the sentences the model uses about you and check them for accuracy, positioning, and completeness. A model that calls you a web design shop when you sell AI integration is a positioning failure that no amount of extra visibility fixes.

Score each description against three questions: is anything in it factually wrong, does it name what you actually sell, and would you have written it yourself. Wrong facts are urgent and have their own correction path; the other two are content problems.

Being mentioned is table stakes. Being described the way you’d describe yourself is the actual goal.

Track recommendation, separately

Being listed among five options and being the recommended option are different outcomes with different causes. Score them separately — appeared, listed among others, or recommended first. Recommendation tends to move with corroboration and specificity; mere inclusion tends to move with structure and markup.

Expect noise, and design for it

The same prompt asked twice can produce different answers, so a single check tells you almost nothing. Run each prompt at least three times in fresh sessions, record how many of those runs included you, and treat that fraction rather than a yes or no as the observation.

Only treat a change as real if it survives repetition across a period. Week-to-week wobble is the medium, not a signal, and a monthly cadence is the shortest interval that produces something worth acting on.

Keep it in a boring spreadsheet

One row per prompt per month, columns for each engine, the appearance fraction, the recommendation level, and the description captured verbatim. It is unglamorous and it is the entire measurement system — and because it is yours, it survives a change of agency.

Record what changed on the site each month in the same document. Without that, a movement six weeks later has no candidate cause and the whole exercise becomes uninterpretable.

Close the loop with the site

Every gap this measurement finds should convert into a specific page change: a question you never answered, a fact you left vague, a claim nothing corroborates. Measurement that doesn’t generate a backlog is theatre, and this discipline has enough of that already.

That backlog is also what makes the reporting defensible — a report with a ranked list of next actions is a working document rather than a monthly reassurance exercise.