Notes on · Search & AI
What I got wrong scoring 50 agencies, and what they told me back
Two agencies challenged the method behind my first study. Their evidence showed me where the scores need more care.
The short answer: my score and real citations did not line up
I scored 50 agency websites for the first Best in Show study. Then two agencies replied with evidence that made me look again. One tracked AI citations for agencies in the study. Across the seven sites in both sets of results, his citation rates did not line up with my scores. Another showed that four posts from his site scored in a higher band than the one I had published for him.
I intended the first edition as a baseline for improving the Standard in public. Those replies showed where the method needed more care and gave me evidence to test it against.
What does the Standard measure?
The Modern SEO + AI Optimisation Standard’s scorecard covers 32 weighted checks. For the August agency study, I used the 13 checks I could assess from public pages. Together, those checks carried 47 of the Standard’s 100 weight points. I converted each agency’s result on that subset to a score out of 100.
That number describes performance against the checks I used. It is not a measured probability that an AI model will cite the site. The checks examine things such as whether a page gives clear answers, identifies its authors and supports claims with sources. They help me find gaps in the material available to read and assess. They cannot, by themselves, tell me which source a model will choose for a particular answer.
An agency owner had his own data on how often models cited agencies. For the seven agencies we could compare, higher Standard scores did not go with higher citation rates. Seven sites are too few to establish a general rule, and the two measures ask different questions. Still, the disagreement matters. If I’m trying to make a useful measure of AI visibility, I need to test whether improvements in the score eventually go with improvements in citations, rather than assume they will.
I’ve also become less confident in how the score weights external coverage. A site can meet many checks on its own pages while receiving little credible attention elsewhere. My view is that positive references from trusted external sources may deserve more weight, because they give a model evidence beyond what the business says about itself. That is a reason to examine the weighting, not a new formula I can claim to have proved.
How did one sampled post skew an agency’s band?
For the first study, I assessed one recent blog post per agency. Applying the same rule to all 50 made the study manageable, but it didn’t make one post representative of a whole site. I should have given that limitation more weight.
One agency showed me three more posts to assess. Its published score was 45.7, in the Exposed band. Scoring those three other posts from the same site with the same method produced results from 50.0 to 63.8, all in Emerging. The agency also revised the originally sampled post, which then scored 59.6, also in Emerging. That is nearly 14 points between the lowest and highest of the three other posts. Which post I picked affected the band I assigned to the agency.
That is the mistake that bothers me most. The time saved by sampling one post was not worth a site-level judgement that could change when I opened another article. For the next edition, the method needs a stated rule for choosing several substantive posts and a way to show variation between them. A single score can still be useful, but readers should be able to see how much the sampled pages differ.
The same agency also showed me a hand-built file that sets out facts and relationships about its business. My checks did not look for it. The Standard includes a check for a clear source of facts about an organisation, so the discovery process needs to look beyond the few pages I initially sampled. I cannot claim from this alone that such a file earns citations. I can say the study missed evidence that was relevant to one of its own checks.
How did a testimonial become an article byline?
A separate error sat in the page reading process. The scraper credited the author of a customer testimonial as the author of an agency article. The byline check then received the wrong evidence.
The scraper now returns a reliable link for the article byline. Two agencies received incorrect byline verdicts in August, and I still need to reassess them using the corrected input. This is a narrower problem than the one-post sample, but it has the same lesson: a tidy score is only as sound as the evidence underneath it.
For the next run, I plan to give a full pass only when a byline links to an author biography page. A name beside an article alone would not meet that rule. Making the distinction explicit should make repeat assessments more consistent.
What happens after the first study?
The lessons here will go into the next version of the Standard. In November, I’ll run the next Best in Show study with the revised approach to sampling posts, finding relevant evidence and checking bylines. I’ll also score the agencies with the August method, side by side with the updated process. That gives me a fair way to compare the studies. If a score changes, the two sets of results can help separate a change in the method from a change on the site.
The August results will stay as published. The August write-up records the first run and provides the baseline for that comparison. I’ll use the updated process in later editions.
Testing the Standard in public was a way to expose it to proper scrutiny. Treating the first method as final would give it more authority than it had earned. In the words of Francis Bacon:
“Truth will sooner come out from error than from confusion.”
A published score with a stated method gives people a way to find specific faults. That is what the agencies did here. A vague claim about visibility would have left them little to examine or correct. In GEO and AI visibility work, which examines how brands appear in AI answers, that distinction matters: a useful finding should stand up to a close look at the evidence, or change when the evidence calls for it.
What should you ask before buying an AI visibility audit?
Ask what the audit actually measures. In particular, I would ask:
- Which prompts does it test, across which models and why were those prompts chosen?
- How many times does it run each prompt?
- Does it record follow-up questions and the answers to those?
- Which pages does it assess, and how does it choose them?
- Can you repeat the audit later using a method that makes the results comparable?
A single set of model answers gives you a snapshot. Answers can vary between conversations, so one run may say less than its spreadsheet suggests. Two or three audits spaced over time can tell you more about whether visibility is changing, provided the method is clear enough to repeat.
My Standard is an attempt to build a useful score that can be monitored alongside citations and organic rankings. It is still a work in progress. Publishing the first study let people test it against evidence I did not have. Their replies exposed errors, helped me fix parts of the process and gave me better questions for the next run. That is how I want the framework to improve.
Common questions
What is a GEO audit methodology?
It is the set of questions, samples and scoring rules used to assess how a site appears in AI search. A useful method also explains what it cannot measure and how the results can be checked again.
Does a high AI-readiness score mean AI will cite you?
No. A readiness score can show whether a site meets selected checks, but citation depends on more than those checks. Measure actual answers and citations separately.
How many pages should an audit sample per site?
One page is too little for a site-level judgement. The sample should cover several relevant pages, follow a stated selection rule and show how much the results vary.
Want to know what an audit can actually tell you?
I can show you what the checks cover, where the evidence is thin and how to measure change over time.