← Insights

Measuring AI visibility before you try to fix it

Buyers now ask a model before they ask you. Most teams start optimising for that without ever establishing where they currently stand.

Sarthak Jain, Co-founder · · 6 min read

A B2B buyer evaluating vendors in 2026 does not open ten tabs. They ask a model, read the four names it returns, and open two of them. If you are not in the four, the rest of your funnel never runs.

That much is widely accepted now. What is not widely done is the boring part: finding out whether you are in the four today, before spending a quarter trying to get there.

Optimising without a baseline is not a strategy

Teams reach for the tactics first - schema markup, entity pages, a content push. Some of it helps. But without a starting measurement you cannot tell the difference between work that moved something and work that merely happened.

Worse, you cannot tell whether visibility was your bottleneck at all. Plenty of companies are cited perfectly well by models and still lose the deal on the site they land on. Fixing citations there is expensive motion.

What a baseline actually contains

A useful AI visibility baseline is narrower and duller than most people expect:

  • Prompt tests. A fixed set of buyer-shaped questions, run across ChatGPT, Perplexity and Google AI Overviews. Recorded verbatim, with the date and the model.
  • Who gets cited instead. Not just whether you appear, but which competitor or aggregator does. That tells you what the model considers authoritative in your category.
  • The crawl surface. Whether the pages you would want cited are reachable, structured, and unambiguous about what your company is.
  • Entity resolution. Whether the model knows your company is one company. Multiple spellings, a stale acquisition, or an abandoned product line are enough to fragment it.

Everything in that list is reproducible. That is the point - you rerun the identical tests later and the delta is the result.

Rerun the same tests, not better ones

The most common failure after a baseline is quietly improving the test. New prompts, a different model, a friendlier phrasing. The numbers go up and mean nothing.

Fix the prompt set at the start. Rerun it at launch, at 30 days, and at 90 days. If the set genuinely needs to change, version it and keep reporting the old one alongside.

Where this sits in the work

We open every engagement with this because it constrains what we are allowed to claim afterwards. If we say a rebuild improved discoverability, there is a dated before, a dated after, and the same questions asked both times.

It also occasionally ends the conversation early, which is the right outcome. If the tests say you are already being found, and the problem is what happens after the click, we would rather say so than sell a visibility project.

If this describes your situation, start with a baseline.

We document where you stand before recommending anything - so whatever we claim at 90 days has a before to be measured against.