AI answer research
Google Search gives you ten links. An AI assistant gives one answer.
I run market research on those answers for brands whose standing cannot be read off their own website. We design the question bank with your team, run it across the assistants that matter to your category, read every answer, and hand you an evidence-backed report and a working session on what to do about it.
- Who it is for
- Brands bought through retailers, distributors, pharmacies, advisers or recommendations. Packaged goods is the clearest case; healthcare, financial services and B2B categories judged through third-party sources have the same problem.
- What it answers
- When someone asks an assistant about your category, does your brand come up, is it recommended, what is said about it, and which sources are shaping that answer.
- What you get
- A frozen question bank, every answer recorded and readable, a report of rates rather than scores, a ranked list of the sources driving them, and the actions each one implies.
Until now, being found meant ranking. A search engine handed over a page of links and the person decided which to open. Position mattered because there was a list to have a position in.
A model does not hand over a list. It reads the sources itself, decides what it believes, and replies in a paragraph. Your brand or your page is inside that paragraph or it is not. The person may never see a link, never visit anyone's website, and never know which brands were left out of the answer.
That is what this measures, and measurement has to come first, because you cannot improve an answer nobody has read.
Google results
A search engine, 2015
- The 12 best sunscreens for kids, testedconsumerreports.org
- Sunscreen for sensitive skin | Solariasolaria.mx
- Mineral vs chemical sunscreen: what to knowaad.org
- Bloqsol Kids SPF 50, buy onlinefarmaciasanpablo.com.mx
- Which sunscreens do dermatologists use?reddit.com
- Sunscreen safety study 2024profeco.gob.mx
ChatGPT results
An AI model, today
For sensitive young skin, choose a mineral sunscreen with zinc oxide at SPF 50. Bloqsol Kids is the one most often recommended, and Dermalux Pediatric comes up frequently for children with eczema. Reapply every two hours.
Six sources went in. Two brands came out. Nobody clicked a single link, no website was visited, and the brands that were left out have no way of even knowing they were left out. That is the whole problem in one line, and it is why guessing is not an option.
What you are looking at
The rest of this page walks through a method. Before it does, three things about how this actually works, because they change how you should read everything that follows.
A person builds this with you
I sit with your team, learn the category, and write the questions with you. The measurement is not a form you fill in and a report that arrives. The judgement in it is the product, and judgement does not come out of a signup flow.
Nothing to buy, nothing to log into
No dashboard, no seat licence, no tool anybody has to adopt. You get the analysis, the report and the reasoning behind both. Your team already has more software than it uses.
Built for brands you cannot measure from your own website
Most tools in this category crawl your site and score your pages. If your product is chosen in an aisle, a pharmacy, a broker's recommendation or an analyst note, that answers a question you never asked. A packaged-goods brand is the clearest example, not the boundary. It works just as well if you do sell online.
So read what follows as a method you would run with me, not a product you would log into.
Measurement, not scoring
Ask it again.
So the only honest way to find out where you stand is to ask an AI model a real question and read how it answers.
Here is a question a real shopper types. Read the answer. Then press the button and watch the same assistant answer the same question differently, naming different brands and citing different pages. Nothing changed except the roll of the dice.
what's the best sunscreen for kids with sensitive skin?
readyPress Ask to send the question.
Click Ask to send this question to a model
What gets counted
The same things, the same way, in every answer, in every run. Consistency is what makes run two comparable to run one.
- Brand mentionsEvery appearance of the brand, its aliases, its sub-brands and its misspellings.
- Competitor mentionsThe same, for every brand in the competitor set, which is what makes a share figure possible.
- RecommendationsSeparated from mere mentions, because being listed and being recommended are different outcomes.
- SourcesEvery site cited or named, deduplicated and sorted by who can influence it.
- Refusals and non-answersRecorded as their own state. An assistant declining to recommend is information, not a missing value.
And whatever else the brand needs
The counters above ship with every engagement. These get added when the brand has a reason for them, which is where most of the personalisation lives.
- Sentiment toward the brand and toward the category
- Alerts when an assistant repeats a specific criticism or an outdated study
- Which retailers and channels the answer names as places to buy
- Whether a claim you are legally allowed to make is being made for you
- Whether a sub-brand is being credited to the parent, or the reverse
The design choice everything rests on
Your website is the wrong instrument
Nearly every tool in this space inherited its frame from SEO. It crawls your site, scores your pages, and recommends content so the site gets cited. That is a coherent product for a company whose revenue arrives through its own website.
It is the wrong frame for most of the consumer goods market. A sunscreen brand sells through pharmacies and supermarkets. Its site is a brochure with sun-safety tips on it. No amount of on-page work changes the fact that the purchase happens on a shelf, and site traffic is not a proxy for anything the marketing team is trying to move.
Build the measurement around the site and you produce a report full of numbers nobody can act on. So we build it around the category conversation instead, which leads to one rule that governs the whole prompt bank.
To find out whether your brand comes up on its own, the question cannot mention it. Name the brand and the model will happily discuss it, and what you have measured is your own question.
That governs one layer of the bank, not the whole thing. You will also want to ask about your brand directly, and about how you stack up against the others, and those questions are in there too. What matters is that the three kinds are asked apart and counted apart, because blending them gives you a number that answers nothing.
Where the purchase actually happens
- 01
Asks an assistant what to buy
Gets a shortlist. Your site is not in it.
- 02
Checks a pharmacy chain's app
Reads a product page the retailer wrote.
- 03
Reads two reviews
A parenting forum and a consumer magazine.
- 04
Buys it in the aisle
Decision already made before arriving.
Times the shopper visited the brand's website: 0
Four outcomes replace traffic
If you cannot count visits, count these instead. Each one is observable in the text of the answer, with no analytics integration and no tag on your site.
Mentioned
Does the brand come up at all in a question that never names it?
The floor. Presence in the category conversation.
Recommended
Does the assistant suggest it, or merely list it among others?
Being named in a list of eight is not the same as being told to buy it.
Named best
Is it singled out rather than grouped?
The strongest form of presence, and the rarest.
Given a purchase path
Does the answer tell the person where to buy it?
The closest this gets to purchase intent, and it needs no analytics, because what we record is what the assistant said and not what the person did next.
One caution worth building in before the run rather than after. A mention of your online store is a purchase signal. A mention of your corporate site is awareness. A social account with no shop attached is neither. Average them into one owned-property number and you will report presence as intent.
The system
How a measurement actually runs
Five steps. Click any of them. The parts that carry the judgement are the prompt bank and the reading of the answers, and those are the two the rest of this page is about.
step 1 of 5
Intake
Understand the brand and what the result has to change
How the questions get written
The bank is the instrument, so it is worth seeing how one gets built. It starts from your category and never from your brand, and it grows in two directions: the topics people ask about, and the different ways they ask about each one.
First, the topics inside it
Nobody asks about a category. They ask about a situation inside it. So the first job is to list every topic a real shopper brings to this category, and write questions for each. Here are seven for sunscreen, with the kind of question that goes under each.
Protecting children
“what sunscreen is safest for a toddler”
Sensitive skin
“which sunscreen will not irritate rosacea”
Daily use on the face
“can I wear sunscreen under makeup every day”
Beach and sport
“what actually stays on when you swim”
Ingredient safety
“is oxybenzone bad for you”
Price and value
“is expensive sunscreen worth it”
Where to buy
“where do I buy mineral sunscreen in Mexico”
Then, the different people asking
One question per topic is not a measurement, it is a sample of one. So each topic gets several questions that come at it from genuinely different angles. Take the children topic. These four are not rewordings of each other, and an assistant answers them differently.
- By who is asking
- “sunscreen for a baby under six months”
- By the situation
- “how do I stop my kids burning at the beach”
- By how far along the decision is
- “which kids sunscreen is best”
- By comparison
- “mineral or chemical sunscreen for children”
Then the questions that do name you
Everything above leaves your name out on purpose, because that is the only way to learn whether you come up unprompted. But plenty of what you need to know only surfaces once the brand is on the table. So the bank has a second layer that names it directly, and it answers a different question: not whether they find you, but what happens once they have.
- What the model believes about you
- “is Solaria safe for a baby”
- Whether it sends buyers somewhere that still stocks you
- “where do I buy Solaria”
- What criticism is attached to your name
- “is there anything wrong with Solaria sunscreen”
- Whether a sub-brand is credited to you
- “is Solaria Kids the same as Solaria”
And the questions about everyone else
A share figure needs a denominator, and the denominator is your competitor set. That set is usually not the one your commercial team would write from memory. It has to include whoever the model treats as interchangeable with you, which routinely means the supermarket own-label sitting next to you on the shelf. This third layer is where you find out which brands own the category conversation, and which comparisons are happening without you in them.
- Who owns the category
- “which sunscreen brands do dermatologists recommend”
- Head to head, named
- “is Solaria better than Bloqsol for kids”
- Comparisons you are absent from
- “Bloqsol or Dermalux for sensitive skin”
- Who owns each attribute
- “which sunscreen brand is the best value in Mexico”
Three layers, counted apart
Unbranded, branded and competitive questions each get their own numbers, and they never get averaged together. An unbranded mention rate tells you whether the category conversation includes you. A branded one tells you what happens after somebody already picked you. A competitive one tells you who you are being weighed against. Blend them and every figure moves for reasons nobody can explain. Kept apart, each has a slice size worth reporting, which is roughly why a bank ends up between fifty and a hundred questions rather than a dozen.
Not every AI model answers the same way
This is the one genuinely technical idea on the page, and it is worth two minutes because it changes what a result means.
When you ask a model a question, it can answer in one of two ways. It can answer from memory, using only what it absorbed while it was trained, in which case it has not looked at the internet at all and your brand is either in its memory or it is not. Or it can go and search first, read what it finds, and write an answer from that. Same model, same question, completely different thing being measured.
There is also a third case, where the search engine itself writes the answer. That is what Google AI Mode and AI Overviews are: not a chatbot, but a search page that answers instead of listing.
Mix these together into one number and it means nothing, because you cannot tell whether your brand is genuinely known or was simply found. So we measure and report them apart.
Answering from memory
The model replies using only what it learned during training. It never touches the internet, so nothing published recently can help it.
Whether the model actually knows your brand. This is the closest thing to unaided brand awareness that exists here, and it is the hardest to move, because it changes only when the model is retrained.
Searching, then answering
The model decides it needs to look something up, runs a search, reads a handful of pages, and writes its answer from those.
What most people actually see today. It is also the messiest to measure, because whether it bothered to search at all varies between two identical questions.
A search engine that answers
Google AI Mode and AI Overviews. You are on a results page, but the answer is written for you at the top instead of being a list of links.
The bridge to the SEO work you may already be paying for. If anyone in the room owns a search budget, this is the number that connects their world to this one.
Models also differ enormously in whether they show their sources. Some name brands constantly while linking to almost nothing. When a model exposes no sources at all we record that as its own result, because scoring it as a zero would quietly corrupt every comparison.
What lands on your desk
A report you can argue with
Below is a sample run for Solaria, an invented sunscreen brand sold through Mexican pharmacies. Three views of the same 540 answers. Open all three and you will end up diagnosing the brand yourself, which is the point: a score tells you a number, and this tells you what to go and do on Monday.
Rates, and what each one is a rate over
Every figure is a share of the answers it was measured in, and the denominator is printed next to it. A percentage with no population behind it is a decoration.
The four outcomes
n = 396 answers to unbranded questions
103 of 396 answers name the brand unprompted
44 answers actively suggest it
12 answers single it out
28 answers say where to buy it
Share of category mentions
n = Every brand named across the same 396 answers
Mention rate by AI model
n = 132 unbranded answers per model
What the model already believed
What most consumers see
Search that answers instead of listing
Solaria is mentioned in about a quarter of category answers and recommended in one in nine. The gap between those two is the whole opportunity: the assistants know the brand exists and mostly decline to suggest it.
Invented brand, invented competitors, invented figures. The structure is real; the numbers are not anyone's.
The one-pager
Everything above compresses onto one page a marketing director can carry into a meeting. What was measured, the four rates, the sources that matter, and the handful of things worth doing. Have a look at it before you decide whether any of this is useful to you.
That is the shape of what arrives. If you want to see what it would say about your category rather than an invented one, that is where a conversation starts.
Ask what it would look like for usWhat you are actually hiring
The engagement, phase by phase
Every study is shaped around the category, but the shape of the work does not change. Here is what happens, who from your side is involved, and what exists at the end of each phase.
Discovery
We settle what decision the result should change, who you actually compete with, and which of your properties count as owned. I need your competitor list and your domains; everything else sharpens the aim rather than deciding whether the number is trustworthy.
Your side: brand or insights lead, one working session
Question bank, written and reviewed
I draft the bank from the category, you review it. Review means checking that the questions sound like your customers and that no topic is missing. It does not mean rewriting them toward the brand, which is the one change that would invalidate the result.
Your side: review and sign-off, plus anyone who talks to customers
Freeze and run
The approved bank is frozen and becomes the instrument. It runs across the chosen models, repeated, with the observation count and the ceiling on cost known before the first question is sent.
Your side: nothing. This is mine.
Reading and interpretation
Every answer is read and coded the same way. This is the expensive part and the part that carries the judgement: separating a mention from a recommendation, attributing sub-brands, and sorting the sources by who can actually move them.
Your side: nothing until the findings are ready
Report and working session
You get the report, and we sit down with it. The session matters more than the document: we go through what moved, what it implies, and which team owns each action, so your people can defend the numbers in a meeting I am not in.
Your side: brand, comms, retail and insights, together if possible
Second wave, when the work has landed
The same frozen bank, asked again after the changes have had time to be published and indexed. Because the instrument did not change, any movement is attributable. If nothing moved, the report says nothing moved.
Your side: tell me when the work has shipped
What is standard, and what is yours
- Standard in every engagement: the frozen bank, every answer stored and readable, the four outcome rates, share of category mentions, the ranked source list with ownership, and the report and working session.
- Shaped to you: which models are worth running, which topics get reported as their own slice, the competitor set including the brands assistants treat as interchangeable with yours, and any extra counter your category needs, such as sentiment, a claim you are legally allowed to make, or an outdated study you suspect is still circulating.
- Not included: I do not fix your website, commission the editorial, or negotiate with retailers. The report says which of those would move the number and who owns it.
If you want to know whether your category is one where this is worth doing, that is a half-hour conversation and not a proposal.
Talk it throughHow we size it together
Three dials, and we set them together in discovery. Move them to see how the shape of a study changes. The point is not to configure a product, it is that the size of the study is settled before the first question is sent, so the scope conversation happens before the money is spent rather than after.
60 × 3 × 3
Observations collected
540
A standard study. Broad enough to report theme by theme across several models.
Directionally: the more questions and the more models, the wider the study reaches, and the more there is to validate before any single figure can be reported with confidence. More questions means more themes covered and more slices that hold up on their own. More models means the result holds across the places people actually ask, rather than in one of them. We pick the point on that curve together, against the decision the study has to inform.
Why anyone runs it twice
The loop is the whole point
Run one gives you a baseline and a list of what is citing your competitors. Then the work happens, which is usually editorial, retail content and public relations rather than a website project.
Then you ask the same frozen bank again. Same questions, same assistants, same number of repetitions. Any movement is attributable, because the instrument did not change.
This is the part that turns a one-off study into something worth keeping, and it is also the part that keeps us honest. If nothing moved, the report says nothing moved.
The bank is frozen after run one. Add a question later and you have two banks, not one trend, so new questions go into a separate set that starts its own baseline.
illustrative figures
+15 points. The bank did not change between the two runs, so the movement belongs to the work rather than to a different set of questions.
Before anything else
Five things to unlearn
Most of the confusion in a first conversation comes from people importing search engine habits into a place where they do not apply. Five corrections do most of the work.
Assistants answer. They do not rank. Your brand is named or it is not, recommended or merely listed, described accurately or not, and given a place to buy or left hanging. Those are the outcomes.
Asking where you rank in ChatGPT asks something the medium cannot answer. If a tool reports a position anyway, it is worth asking what it actually measured to produce it.
One analysis of 23,387 citations across 240 branded queries found earned media supplying 48 percent of what assistants cite, other third-party commercial content 30 percent, and the brand's own site 23 percent.
So if the answer about your category is assembled from a consumer magazine study, a regulator's report and two retailer product pages, then the lever is what those say. That work is public relations and retail content, not a website migration.
You just watched this happen. Published measurement work puts source overlap between identical same-day queries somewhere between a third and a half.
One screenshot is an anecdote. A measurement is many asks, and every number in the report is a rate over those asks. This is also what makes the exercise cost something, so it is worth establishing early.
A question does not have to be popular to be worth asking. If an assistant has formed a view about a safety concern in your category, that view gets repeated to whoever does ask.
You want to know what it says now, not after a journalist quotes it back to you.
Almost everyone arrives wanting the optimization. Measure first, because without a baseline you cannot attribute any later change to anything you did.
The honest version of this business is that run one tells you where you stand and what is citing your competitors. Run two is where you find out whether the work moved it.
What is different here
Three things you will not get elsewhere
Nothing to buy, nothing to learn
There is no dashboard, no seat licence and no tool your team has to adopt. A specialist runs the measurement and you receive the analysis and the report. Most brand teams already have more software than they use, and the last thing a quarterly measurement needs is a login somebody forgets.
Built for brands that do not sell online
Every tool in this category crawls your website and scores your pages. If your product is bought off a shelf, that answers a question you did not ask. This is built the other way round, from the category conversation inward, which is why it works for a packaged food brand and works just as well for a software company that does sell online.
You learn the method, not just the number
Part of the engagement is teaching your team how assistants build answers, why the bank is written the way it is, and how to read a rate. The aim is that you can defend these numbers in a meeting we are not in. A number nobody on your side can explain does not survive its first challenge.
What this does not do
Everyone else in this category sells an instant score. Here is where the boundaries of an honest measurement actually sit, so you can decide before you spend anything.
- It will not give you a rank
- There is no position one to report. Anything that gives you one made it up.
- It will not attribute sales
- We record what the assistant said. What the person did next happens outside anything we can see. The purchase-path outcome is the closest honest proxy, and it is a proxy.
- It cannot promise the number will move
- Most of what an assistant cites was written by somebody else. We can tell you which sources are driving the answer and who can influence each one. We cannot commit to a result that depends on third parties publishing differently.
- It does not fact-check the assistant for you
- We can record that an assistant said something about your product. Judging whether that statement is true needs an approved source of truth that somebody on your side maintains, and most brands do not have one. If accuracy is your worry, say so at intake, because building that reference is a separate piece of work.
- One run tells you where you stand, not what to do
- The action plan comes from the source list and the gaps, and it gets sharper with a second run. Anyone promising a fix from a single baseline is selling you the wrong thing.
Start with the question, not the tool
The useful first conversation is about what decision the result should change, and whether your category is one where an assistant already has opinions. That takes about half an hour, and you will know by the end whether measuring is worth it for you.
Talk it through