Search City: Measuring What We Can See
If search marketing was walking a lit street, in the age of AI it can feel like we’ve been plunged into darkness
I’ve been thinking a lot about AI visibility metrics recently.
Partly because clients keep asking reasonable questions like “How visible are we in AI?” and partly because every time I think I’ve found a good answer, I realise I’m falling back to measure a “less good” proxy for something else.
That isn’t entirely new I don’t think, SEO has always relied on a mixture of direct measurement and educated guesswork. What has changed is how much of the journey we can actually observe.
For most of my career, search felt like a relatively clear and “well-lit” place.
What if we Treated SEO Metrics like Lit Streets
If we rewound the clock a few years, the world of search - and how we measured it - felt comparatively straightforward.
Google crawled pages. Pages were indexed. Rankings appeared. Users clicked. Visits and conversions followed.
(We even had keyword data in Google Analytics way back as far as 2013!)
Each stage had its own tools and metrics, and whilst we didn’t have perfect visibility, we could usually see enough to understand what was happening.
Crawl logs, Search Console, rank trackers and analytics all acted as lit buildings, illuminating different parts of the street. Most of the important decision points sat somewhere (mostly) within view.
That’s not to say there weren’t blind spots. Attribution was never perfect and rankings alone never told the full story. But compared to where we find ourselves today, the route from content to customer was relatively well understood (lit).
Then More Streets Were Added
One of the most common narratives around AI search is that everything has become a “black box”.
That IS part of it I’m not convinced that’s quite right. Moreover, SEO itself was very “black boxy”!
Alongside Google, we now have ChatGPT, Claude, Perplexity and a growing number of AI-powered search experiences. Each has its own journey between content existing and an answer being delivered to a user.
Each exposes different signals whilst hiding different decisions. We are forced to more and more decisions (all with cost implications) with a worst measurement framework and an uncertain future.
Google still gives us visibility into crawling, indexing and traditional search performance. ChatGPT reveals different clues through citations, telemetry and browsing behaviour. Access logs can tell us something about crawler activity. Prompt testing can reveal patterns in answer generation.
Not All Darkness Is Equal
“The Search City” analogy started as a way of thinking about observability across different platforms, and one thing that quickly became apparent is that not all buildings are dark for the same reason.
Some parts of the journey are dark because platforms choose not to expose much information about what happens inside them. Google gives us extensive visibility into crawling, indexing and ranking, but considerably less visibility into how sources are selected and weighted for AI-generated answers.
Other parts are dark because direct observation is difficult. Understanding whether a page has been crawled is relatively straightforward. Understanding how information is retrieved, compared, weighted and synthesised by a large language model is a very different challenge entirely.
There are also parts that are dark because it may be best not to share them - for example Click Through Rate from AI sources.
Then there are parts that are dimly lit.
Access logs can reveal crawler behaviour - for some bots.
ChatGPT telemetry can expose search fan-out patterns and tool usage. Particularly the difference between models, logged-in/out states
Citation analysis can reveal how often a source appears in answers. How consensus is built and what content is most often used for that.
Prompt tracking can help us understand whether a brand is likely to surface for a particular topic - who also is present and how often this is the case.
None of these observations provide a complete picture, but they do illuminate parts of the journey that would otherwise remain hidden. More importantly, none are concrete and connected to the outcomes we crave understanding of - attribution to revenue-generating events.
This distinction matters not because the light went out, but because the city expanded faster than the infrastructure could keep up..
The history of search has largely been a process of reducing uncertainty. Search Console illuminated crawling and indexing. Analytics illuminated traffic. Rank tracking illuminated visibility.
AI search feels unfamiliar because we're back at the beginning again in so many ways.
That's also why so many of the metrics emerging around AI visibility look different from the ones we're used to. When you can't directly observe an event, you start looking for evidence that it happened.
Measuring Disturbances Instead of Events
Maybe that is it. Much of AI measurement feels less like measuring key events and more like measuring disturbances on the path to one.
Bot visits suggest a system may be paying attention to a page.
Citations suggest content may have influenced an answer.
Prompt tracking suggests whether a brand or page is likely to be surfaced.
Referral traffic suggests that a recommendation may have happened somewhere upstream.
Finally a tracked-conversion with AI-bot referral shows you the conversion data. But we KNOW this is highly-under reporting the value on its own.
None of these metrics directly expose the decision-making process itself. Instead, they provide evidence that a decision has taken place. The result is that modern search measurement increasingly resembles assembling a picture from multiple viewpoints rather than relying on a single source of truth.
The Danger of Optimising for the Lit Windows
There is, however, a risk in all of this as we naturally optimise for the things we can see - in traditional SEO that often meant rankings.
Today it might mean citations, AI referrals, bot visits or mention frequency. The problem is that visibility and importance are not the same thing. The value of each of these things likely has different weighting at different points of the journey.
We also don’t have proxies like search volume to direct our efforts as effectively.
Just because a particular window is illuminated doesn’t mean the most important decisions are happening in that room. The temptation is to build strategies around whatever metrics are available.
The reality is that some of the most consequential decisions may still be taking place inside buildings where only a handful of windows are lit.
The Next Generation of Streetlights
Despite the uncertainty, we can stop this from being a story about losing visibility. Why not make it a story about rebuilding it?
The history of search has largely been a process of illuminating previously hidden systems. Search Console did this for crawling and indexing, analytics for user behaviour and rack tracking for visibility.
Today’s citation tracking, telemetry analysis, prompt testing and bot monitoring feel like the next generation of that process.
They’re imperfect and very much incomplete. But so were many of the tools that came before them. Much of our longing for past data leaves out how flawed it was.
Every generation of search creates new blind spots. Our role as practitioners has never been to eliminate uncertainty entirely, but to (try and) reduce it.
AI search feels unfamiliar because we’re once again exploring newly built streets of the city, trying to work out where the important roads lead and which windows are worth looking through.
The city is larger than it used to be. Some buildings remain brightly lit. Others are almost completely dark. How do we understand the difference?



