syted

syted blog

LLM Brand Visibility: How to Measure It Yourself

Search for LLM brand visibility and you will find, almost exclusively, lists of tools. Fifteen best monitoring platforms. Fourteen tools to track brand visibility in AI search. Every one of them assumes you already know what the number is, what a good one looks like, and why yours would change.

None of those things is obvious, and one of them is genuinely strange: the number moves when you have changed nothing at all. That property is the first thing anyone measuring this needs to understand, and it is the thing the tool roundups skip.

This article is about doing the measurement, including by hand. It is written for the person who owns the site rather than the person who buys the software.

What the phrase actually means

LLM brand visibility is the share of relevant answers in which an assistant names your business. Ask ChatGPT, Claude, Perplexity or Gemini a question a buyer would ask, and count how often your name appears in the reply.

That is a different object from a search ranking, in three ways that matter.

There is no position ten. An assistant names three or four options and stops. You are in the answer or you are absent, and absence has no consolation prize sitting below the fold.

There is no single result page. Two people asking the same question in the same minute can get different answers, because these systems sample rather than look up. This is not a bug you can file.

And the answer is assembled rather than retrieved. The model is drawing on what it absorbed in training, plus whatever it fetched during that particular conversation, plus the phrasing of the question. Your page is an input, not the output.

Real demand numbers, so the topic is sized honestly. Measured with Ahrefs for the United States in August 2026, llm brand visibility gets 268 searches a month at a difficulty of 12. It is a young query, which is another way of saying nobody has settled what the metric should be.

The number moves on its own, and you must plan for it

This is the part that turns a measurement programme into a source of arguments if it is not stated up front.

Ask the same question twice and you can get two different lists. Several separate mechanisms cause this, and they stack:

  • Sampling. These models generate text probabilistically. Unless a provider pins the settings, repetition alone produces variation.
  • Retrieval. When an assistant browses, it gets whatever the underlying search returned at that moment, which is itself unstable.
  • Routing. Providers serve different model versions to different users and change them without announcement.
  • Context. The same question inside a longer conversation gets a different answer than the same question cold.
  • Personalisation. Memory and account history shape replies, so your own account is the worst possible measuring instrument.

The practical consequence is simple and unpopular: a single run tells you nothing. A brand that appears in one answer out of one is not visible, and a brand that misses one answer out of one has not fallen off. You need repetition before you have a number, which is why the honest version of this metric is a percentage over many runs rather than a yes or no.

Anyone reporting movement without a sample size is reporting noise. That includes tools, and it includes the screenshot someone sends you at eleven at night saying a competitor got named.

Building a prompt set that means something

The quality of the measurement is decided here, before any tooling.

The mistake almost everyone makes first is to ask about themselves. "What is Acme?" is a question only your existing customers ask, and the model answering it correctly proves nothing about acquisition. It measures recall, not consideration.

The prompts worth tracking are the ones a buyer would type before they know you exist. In practice they fall into four groups:

Prompt type Example What it tells you
Category "best tools for X" Whether you are in the consideration set at all
Constrained "affordable X for a two person team" Whether your positioning is understood
Comparison "X versus Y, which for Z" Whether you are a recognised alternative
Job to be done "how do I solve Z without hiring someone" Whether you own the problem, not just the category

Write thirty to fifty of these, in the words a customer would actually use. Then freeze the list. A prompt set you edit every month cannot show you a trend, because you will never know whether the number moved or the questions did.

Two further rules save a lot of pain. Include prompts you expect to lose, because a set you always win is a set that measures nothing. And keep the geography explicit where it matters: "in Austin" and no city are two different questions, and local businesses live or die on the first. That distinction plays out the same way in classic search, as we cover in DIY SEO for small business.

Running it, and what to write down

Cadence first. Weekly is enough for almost everyone. Daily produces a chart made mostly of sampling noise, and monthly is too slow to connect a change to a cause.

Run each prompt several times per assistant. Five is a reasonable floor, ten is better, and the point is not precision for its own sake: it is that below five, a single odd answer swings the whole figure.

Use clean sessions. No logged in account, no memory, no prior turns. Your own account has read your site, discussed your product, and will happily tell you what you want to hear.

For each run, record five things: the date, the assistant and any version you can see, the exact prompt, whether your name appeared, and whether a link to your domain appeared. Those last two are different events and conflating them hides the most useful signal you have.

Then compute two rates rather than one:

  • Mention rate. The share of runs in which the brand is named. This is awareness inside the model.
  • Citation rate. The share of runs in which your domain is linked or named as a source. This is your content being used.

A high mention rate with a low citation rate means the model knows you from elsewhere, probably third party pages, and is not reading yours. A low mention rate with a high citation rate means your pages are useful but your brand is not yet an entity the model recognises. Those two situations need opposite work, and a single blended score would have told you nothing about which one you are in.

What actually changes the number

Nothing here is exotic. The disappointing and liberating fact is that the inputs are the same ones that have always mattered, weighted differently.

Being named on pages you do not own. Listicles, comparison posts, forum threads, review sites. Assistants lean heavily on third party consensus for "best X" questions, because that is what the underlying corpus contains. This is the single biggest lever and it is also the slowest.

Pages that contain quotable facts. A model reproducing your claim needs a claim to reproduce. "Pricing starts at a fixed monthly fee with no per seat charge" is quotable. "Flexible pricing designed around your needs" is not, because there is no fact in it. We go through the mechanics of what gets picked up in how to rank on ChatGPT.

Consistency of the entity. Same name, same description, same category across your site, your profiles and your listings. Models resolve entities, and a brand described four different ways in four places is four weak entities instead of one strong one.

Being crawlable at all. If your content only exists after JavaScript runs, some assistants will never see it. This is the least glamorous item on the list and the one that most often turns out to be the actual problem. We covered the general version in LLM SEO.

What does not appear to work, based on what people keep trying: stuffing instructions to the model into your own pages, publishing volume without substance, and any of the "prompt injection" tricks that circulate every few months. They are also the sort of thing that ages into an embarrassment when a provider tightens its filtering.

A worked example, with the arithmetic shown

Numbers make this concrete faster than description does.

Take a forty prompt set, run five times each, on two assistants. That is four hundred runs a week, which sounds enormous and takes about an hour once the list is written, because most of it is pasting and reading.

Say your name appears in 96 of those 400 runs. Mention rate is 24 percent. Your domain is linked in 20 of them, so citation rate is 5 percent. Both numbers are low, and the gap between them is the interesting part: the model knows the brand more often than it reads the site.

Now split by prompt type instead of looking at the total. Suppose category prompts return 40 percent, comparison prompts 30 percent, and job to be done prompts 4 percent. The blended 24 percent hid the finding entirely. You are in the consideration set once someone already knows the category, and invisible to the person describing a problem in their own words. That is a content gap with a shape, and it points at a specific set of pages to write.

This is the same failure mode as reading one number for organic traffic instead of splitting it by query type, which is the habit we argue for in SEO for SaaS. Aggregates hide the actionable part.

Assistants and AI Overviews are not the same measurement

One distinction saves a lot of confused reporting.

Google's AI Overviews sit on top of a normal search result page. They are generated, but they are generated from a search that also produced the blue links underneath, and the sources tend to track what already ranks. Being cited there correlates strongly with ranking well, which means your existing Search Console data is a decent proxy.

A standalone assistant is a different animal. There is no result page underneath, the question is usually longer and more conversational, and the sources can include material that ranks nowhere. This is why a business can be well ranked and rarely recommended, or occasionally the reverse.

Measure them separately. Track AI Overview appearances alongside your ranking work, and track assistant answers with the prompt set described above. Averaging the two produces a number that describes neither.

Tools, and when doing it by hand is better

The tool market here is young and the pricing is aggressive. That is not an argument against buying one, it is an argument for knowing what you are buying.

Doing it by hand costs an hour a week for a fifty prompt set across two assistants, and it teaches you what the answers actually look like. That reading is worth more than the chart in the first months, because you will find out that the models are recommending a competitor you had never heard of, or describing your product as something it is not.

A tool earns its price when the set grows past what a person will actually run, when you need history you did not think to keep, or when several people need the same numbers. What no tool can fix is a badly chosen prompt set. Feed it self referential prompts and you get a beautiful dashboard of a meaningless number.

If you are evaluating a vendor rather than a tool, the questions worth asking are the same ones we set out in vetting an AI SEO agency: what exactly is measured, at what sample size, and what happens to the reporting when the number goes down.

The mistakes that waste the first three months

Four patterns show up repeatedly, and each one is cheap to avoid before you start rather than after.

Measuring with your own account. Memory, history and prior conversations all push the answer toward what you want. Every measurement run belongs in a clean session, and if that is inconvenient it is because the convenience was the contamination.

Changing the prompt set to look better. The temptation arrives the first week the number disappoints. A set that changes cannot produce a trend, and a trend is the only thing this measurement is good for.

Reacting to a single answer. Someone will send you a screenshot of a competitor being recommended. That screenshot is one sample, taken once, in an unknown session state. It is worth a note and not worth a strategy meeting.

Writing pages aimed at the model instead of the reader. The pages that get quoted are the ones a person would find useful, because usefulness is what made them get linked and discussed in the first place. Content written to please a parser reads like content written to please a parser, and it does not attract the third party mentions that move the number.

Setting expectations before you start

Write down what you expect, then check it later. This is the cheapest protection against fooling yourself that exists.

Reasonable expectations, stated plainly. Movement is slow, because the biggest lever is third party mentions and those accumulate over months. Results vary by category, and a crowded software market behaves nothing like a local trade. Nobody can guarantee that an assistant will recommend a given business, and any vendor who does is describing something they do not control.

What you can promise yourself is a measurement that is honest, repeatable and free, and a set of inputs that are known rather than guessed. That is a considerably better position than most companies are in.

The commercial question underneath all of this, which channel deserves the next thousand dollars, does not have a general answer either. We worked through the arithmetic for the search versus paid case in SEO versus Google Ads, and the same discipline applies here: measure first, then spend.

FAQ

How many prompts do I need for a useful measurement?

Thirty to fifty is the practical range for most businesses. Below about twenty, a single prompt swings the percentage too much to read a trend. Above a hundred, running them properly stops being something you will keep doing by hand, which matters more than the extra precision.

Which assistants should I track?

Start with the two your customers actually use, which for most western businesses means ChatGPT plus one other. Adding a third doubles the work and rarely changes a decision. Track more only once the first two have produced a trend you trust.

Why did my visibility drop when I changed nothing?

Most likely nothing dropped. Providers ship new model versions without notice, retrieval results shift, and sampling alone produces swings. This is exactly why the metric has to be a rate over many runs. If a drop persists across several weeks and several assistants, it is worth investigating. A single bad week is not.

Is LLM brand visibility replacing search rankings?

No, and framing it that way leads to bad decisions. They are different surfaces with heavy overlap in what feeds them. Most of the work that improves one improves the other, which is the practical reason not to build two separate programmes.

Can I do this without buying a tool?

Yes. A spreadsheet, a frozen prompt set and a clean browser session are enough to produce a defensible number. Buy a tool when the manual run stops happening, not before, because the discipline of reading the answers is where most of the early insight comes from.

Get cited by ChatGPT. Rank on Google.

You found this article through search. That is the whole product.

  • One researched article a day
  • Published on your own domain
  • Keywords checked against live results
Start writing

Get cited by ChatGPT. Rank on Google.

You found this article through search. That is the whole product.

  • One researched article a day
  • Published on your own domain
  • Keywords checked against live results
Start writing