How To Do a Successful AI Search Audience and Prompt Research
25m 27s
AI search audience research requires a fundamental shift from traditional keyword-based methods. The key insight is that prompts must include real audience context—such as "I’m a crossfit athlete"—to produce consistent, relevant results, as models respond differently based on user profiles. Many errors stem from using generic or reverse-engineered prompts that reflect existing content rather than genuine user intent. Over-reliance on volume and pre-built prompt libraries leads to noise and self-fulfilling outcomes. Instead, researchers should prioritize a minimal, segmented, and context-rich prompt library aligned with specific audiences and business goals. Data must be gathered from multiple sources—Google Search Console, social listening, and user reviews—to form a holistic view of user behavior. Platform-specific tracking is essential due to significant drift (e.g., Google AI Overviews changes results 40% of the time), and sentiment bias must be carefully monitored. Finally, success hinges on clear ownership: prompts should belong to specific teams (e.g., marketing, CX), and decisions should be driven by actionability, not vanity metrics. Google’s own review analysis offers a valuable baseline for identifying recurring user concerns. By focusing on patterns, real constraints, and brand safety, teams can build a cost-effective, strategic AI search research framework.
Welcome to the AI search roadmap series today about how to do AI search audience research.
It seems something very trivial, but it's actually no, because of all of the implications.
And to talk about this, I have invited SEO, especially with a lot of experience,
who I know that are working in AI search optimization as well, to share with me some of the core key,
most critical, craziest, dosandons, the top errors that they are seeing out there and how to tackle
this process in a good, reasonable, cost-effective way, because we know how from tracking can.
Usually, make us to overspend as well. Before starting, I want to thank our sponsor,
you know, Web. Thank you very much for your sponsorship. Similar Web is one of those tools that I use
in my day today for not only from tracking, but also a lot of analysis ready with AI search as well.
So thank you for the sponsorship. And to start with today's edition, I want to welcome
Miriam Jaeger, Andy Chadwick and Dan Taylor. I would like to hear from you which are those
top errors mistakes that you see people doing when developing their from research. See that many
people are just extrapolating what they used to do with keyword research and many people just
importing the problems that are suggested by default by their AI trackers as well. We'd love to
hear more and your recommendations of what to do instead. I think just jumping off your point of
error later, that's probably a good place to start and it's the mindset shift of people approaching
from tracking in the same way. We have round tracking historically and I've worked with clients
and seeing pictures coming through where people are tracking 10,000 keywords at the moment because
they have a lot of variations of what a track was saying in my prompts and it's not actually
understanding that it's a different beast altogether and also the level of volatility what happens
because it's not like a linear serpent. It almost feels like if you're trying to do the fact you're
going to boil the ocean and a lot of these, as you said, prompt tools are very expensive as we are.
We're just adding more and more costs to a stacking. In my opinion, getting very little
lower out of it, we do try to take with ball of the ocean approach with prompt tracking.
Just to build on from that, especially on the costing side of ones, I wrote about it I think
about two weeks ago. I did a test where these LLM2 have context, right? All of them do JETGPT,
practically they have context about who you are and so if I was to search which I did, I do a lot
of crossfit in my spare time. So I searched, the literal prompt was can you recommend me a running
shoe? I'm doing the half marathon in September and it gave me three options and I've just a
matter of interest. I asked one of my friends who does rock climbing to ask exactly the same question
so not giving it any context. I'm doing a life marathon in September can you recommend me three
running trainers and he got three different results. One of them was the same but the top two
recommendations were different and that's because of the context of between both of us. I asked
ChatGPT to give me a reason why it told me different results of my friend and I can't remember the
exact reason. Something along the lines of as a cross trainer, a crossfit person, there'll be a lot
more tension on my calves and this so it recommended these two as my rock climbing friend. He has a lot
more tension on his heel and his Achilles so it recommended these two. So there was a real reason
why it recommended them and so the problem with people tracking prompts and the way they are,
you're hitting an API, most people are doing it by hitting an API and it's got a blank canvas,
it doesn't know who you are and so there's a study done and actually the article I wrote sites
that study and if you just use a base prompt with no context the results change either a third
direction over 25,000 runs. It was just wildly different up to 60% difference between each run
over 25,000 prompts. That dropped significantly and I can't remember the exact stats,
so don't quote me on this but it dropped to maybe like 20% when you gave it a contact, a person.
So if you append your prompts with the actual audience you're trying to target,
I am a crossfit to recommend me a trainer and keep running out through. The results seem to come
the same more often. I think we should have always as marketers been interested in who the audience
actually is not trying to appeal to everyone and LLAMs have just made that, now they're actually
showing it and making that more important when you're prompt tracking you really need to add the
context of the person who it is so you should append each prompt and you'll get different ones,
maybe you've got five different ICPs you need to append that to the prompt. That can quickly become
very expensive because then you're not just targeting one prompt, you're targeting the same prompt
with a different ICP attached to each one. You get two different ICPs that gives you two different
results, you're a visible one for not the other, there's a bit gap in your site targeting a specific
audience, so that is interesting anyway. But if you can only target one audience, just do one
audience and that's something that needs to be appended to those prompts.
>> To segment each one of your audiences for each product line service line and yes I understand
that this can become a little bit expensive quite easily if done like that, but what is key here
is the level of prioritization you have first and then on the other hand this should be prompts that
should be really representative of your actual customer journey and the real constraints and
characteristics that they search with because one of the main issues is I think that people try
to pull the more data the better, the more keywords the better and not necessarily right?
>> I have many many opinions about this because I have been dealing with nightmare so please
hang on for the topics of errors you don't want to make because I need them for you.
So number one a keyword is not a prompt please and thank you the way we used to search as humans
was conditioned by the technology, by library, by Google and now the scaffolding is falling off and
we can actually talk like humans to this thing so no a keyword is not a prompt. Okay first point.
Second point to personas unless you want to be testing if the model is good with trends don't
do that inject some context let me give you an example. If I ask these two gentlemen what a
naked shoe trend is they're not going to give me the answer that the old culture barri
mesmer are expecting okay. So some of these models are really good to catch on to these implicit
trends. Some of them are disastrous bringing some trends from the 90s or like orthopedic stuff that
doesn't belong there so please inject the context of the person when you're trying to do this.
Otherwise you're going to get a lot of noise okay. What happens if I'm one person analyzing
a bunch of prompts but that are in Japan, in China, in Germany and I only speak English you are
going to have to make some choices. This means that your data will be biased okay it will be because
it is what it is the prompt language impacts the sources that are going to show up but and this is
important. It has to fit what you're willing to do with this. So if you think your prompt library
or taxonomy is the end game and you're proud of yourself because that's a beautiful thing no that's
the first step okay. I have a few more. If you're waiting one week and then looking into the results
you're doing it wrong please wait a little bit be patient at least 30 days for goodness sake we're
professionals here okay. Tell people to chill and to wait. It's a living document so don't
don't let it die because some of these prompts are going to change which brings me to the last
thing I want to say here. If you are going to track some trends like I said you know some things
that are important validate the trend please because otherwise it doesn't necessarily deserve to be
here if it doesn't make sense. If you had outdated historical data it's not going to be representative
and last but not least when it comes to mistakes that's the mistake number one everyone makes.
What you're doing is an ideal scenario in a sanitized lab space. You were looking at the best case
scenario you could have not reality so please keep this in mind because this is giving you a
direction it's not giving you the full reality. I love that last part regarding how what you see
in any case should be directionally helpful and indicative of gaps on opportunities rather than
making you to obsess of whatever instance of these prompts where you are not showcased right.
There's a level of inconsistency that will tend to happen that is why what you are looking to
identify our patterns at a topical level rather than obsessing over individual behavior of prompts
as well we need to evangelize. I love how you have shared the actual most common errors mistake
issues that we're all facing. I would love to hear what are the actual steps on criteria that
you're using right now to create a good prompt library to understand to have a good directional
reference of the behavior of your users in AI platforms and then to be able to measure
how your brand is performing for it and based on that to identify opportunities to identify
gaps because we we shouldn't forget why this is important in the first place.
One of the things we're doing a lot of is just trying to again help clients better understand
volume and quality and actually get in marijuana because again I think for the last 10 or 15 years
we've trained everybody on more is better more keywords more traffic more I think is better
me obviously what traffic is better so we're doing things like down sampling and having a valuation
data set so we don't need to do a thousand prompts or a thousand variations from test cases we can
just do really optimize smoke tests 20 to 30 prompts across libraries daily iterations and then use
these fans are put together.
ever a Ruby end kind of release candidate set of prompts we actually
then do trash across ICPs, even down to using smaller models.
Everyone, you know, I have a lot of companies I work with in the US.
I kind of drawn into always a gradient for most expensive models for tracking.
We can use the cheaper models to run little test cases as well and
judges just to see if there's any work in it.
But then also I'm a big fan of putting together rubrics, a value of a
moment to actually understand if smaller models are producing certain key
trends across the right wider set is cheaper to track and then we take those
candidates and those prompts and actually we'll remember larger models
which are more likely to be used by consumers and then alongside that as
well, actually understanding that not all users and audiences will be using
with saying kind of LLM's.
I do a lot of working B2B farm vehicles, yeah.
And most, I mean, I do a lot of B2B upscale farmer in the US, but kind of
stuff where not direct chemicals and things, but if you want to buy
lab equipment, for example, you want to buy a million face masks or something.
A lot of people are on lockdown completely computers, lockdowns are using
Microsoft Edge, God bless them and using Bing and co-pilot for everything.
So co-pilot is more important than anything else for all the tools, but
mainly push on strategy, PT, Gemini etc.
So it's actually refining the stack to work with them as well.
That's a really good way.
I read another study.
I think it was by Rand Fishkin, the drift between the different platforms.
So I pulled out the stats here.
So it perplexes to the monthly citation drift is 40.5%.
And what I mean by that is it's the steadiest.
It gives you the same, for a given set of prompts, 60% of the time it gives you
the same answers, 40% of the time it changes.
Co-pilot 53.4, chat GPT 54, Google AI overview 60%.
So they all drift wildly at Google AI overviews being the most erratic.
And then what I did, and I pulled this last week is to look at where they're
calling from.
It's a gambit and wrote something on this as well, my co-founder.
So perplexity normally pulls from live web search traditional.
I mean, I know we're here this a lot, but the tactics change for each one.
So for Black State's closest to classic SEO and digital PR,
it's the underlying rank, which it's pulling through.
Chat GPT is more around consensus.
So it's more of those listicle type things that people are telling you to go on
Wikipedia, that kind of thing.
Google AI overviews has a 54% overlap with classic organic ranking.
So that's back to traditional SEO.
And I've got one for this, each platform actually has slightly different
quality optimization technique, but optimization type thing that you go for.
And when we, when you look in your whatever platform you use,
PKI or whatever it might be, and you blend them, that's not really the right thing
to be doing.
You should be reporting on each platform.
And as you just said, some have a different audience like your,
your developers are using this.
So this is, and this is where this pulls from.
So this is what you should be doing.
So that's just one to the build on that point there.
So you're going to open also asked or an equivalent, okay?
You're going to get in touch and make friends with the social media team.
If there's no social media team, you're going to do the research on your own, okay?
You're going to open Google search console.
You're going to open Google Analytics or an equivalent.
You're going to make best friends with the traditional marketing team.
So they give you the personas, okay?
You're going to ask what the actual business objectives are.
Once you have all of these sources, and the way I like to talk about it,
is that you have your historical sources, which is your Google search console,
your semi-rush, et cetera.
That's a historical diagnostic, you know, your internal analytics.
You're going to get your market intelligence by asking what the campaigns are,
what the personas are, et cetera.
You're going to do the latent research, which is what are people saying?
But also asked, you're going to use Google Trans, maybe you're going to look into social media
listening, and everything that you're going to find is going to help you figure out
what these prompts should be, okay?
The reason why I say this is that if you look only in one spot,
thank you, Google search console, you don't know what you don't know.
And you're not a social media expert, but oh my goodness.
Our third party website is really feeding these answers.
So you need to become one, okay?
So you have to get confident and actually speak the language of the people you're trying
to reach, right? As if you don't have the lingo, it's not going to happen at all, okay?
So once you have all of this, you're starting to understand what's going on.
You're able to structure your questions a lot more.
So this data, it's non-linear.
This means you will be going back and forth.
You will hesitate.
You will have to prioritize.
You'll have to segment it, figure out what it is, what it does, what it's supposed to do.
And there is one question at the end of this big first trial that you have.
You have to ask yourself out of all the prompts that you have.
If AI answers something really bad about me here in this question,
is it going to cost me a sale?
Yes or no?
And that can be a beginning.
So that can be a brand thing.
If it says, hey, like don't really trust them.
I have to show them because they are the main ones you asked me about or, you know,
they are discussed in this space, but don't really cost you a sale.
If it actually tells you it's not for you because there's missing information that can
cost you a sale, you pay attention.
That's the one question.
But beyond that, there's another sneaky question you should be asking yourself.
Is there sentiment bias in my prompt?
And this is important for you to know that doesn't mean it's going to be tagged in the tool.
That doesn't mean anything else.
But when you're going to analyze sentiment, you're going to come out and say, oh my God,
we have terrible sentiment.
How do I explain this to my boss?
No, you asked like, is this, can I trust it?
You already are negatively biasing the answer by asking a negative question.
So be aware of the sentiment in there as well.
Okay, this is one of the big things that I see not happening.
And I'm like, and then your surprise is negative.
You're probing for negative stuff.
Thank you very much for being in the negative sentiment.
Since this is assessed across most tools in the most simplistic way.
So for example, if you have a certain type of brand that had a very unique selling
proposition that might be perceived as something expensive, I have seen money,
a lamp, depending on the wording, right?
Like, oh, this is expensive.
So it's bad for you and they flag that sentiment to be negative.
But actually, it's not because it aligns with a brand unique selling proposition and values.
And it's just that the persona or the audience is not the correct one.
But that is why you need to really provide a context, the right context for
measurement and also something very important that you highlighted.
Unlike traditional search, AI platforms can be seen more beyond not only a
performance channel, but also a brand new channel.
And that is why there's always the share of valuable monitoring prompts, I call them.
Besides the core prompts, that these are the more stable ones used for
recurrent tracking, reflecting the highest priority journeys that should remain
consistent over time.
I also have like usually the experimental ones to test new products, markets,
constraints, audiences, emerging behavior, things like that.
But also the monitoring ones that to track reputation, brand representation,
risk compliance, because it is actually not only a retrieval system anymore,
but a recommendation engine that affects branding.
And that is why we need to track all of this additional behaviors.
But we always need to think is that you should always ask why?
Why are measuring this?
It should always connect with an action on something that is worthy to optimize for
a goal that we should pursue.
If it doesn't connect with that, then it was usually or very likely pointless.
And that is a good way to prioritize accordingly.
What you should actually measure are not for my SEOs out there, volume and
ranking.
Okay, these are the two things.
Your prompt volume is based off of certain things that are very interesting,
but they're not the same as SEO volume.
And I will let my peers talk about that all day.
But the second element is that you also have to condition the people working
with you.
I'm going to put in the prompts and we are going to prioritize them.
Okay, we don't, we can't track more.
It's going to cost us more.
So people will still want more.
You have to tell them, no, we have a system here are the priorities.
If you want to add more prompts, they better go through and be validated
and really belong to this.
So if you don't have someone that is responsible for the prompt, don't put it in.
What do I mean by this?
I have some sales questions.
Well, this is going to go to the customer experience team or the support team.
They own fixing that stuff.
They own the data that comes out and facilitating.
I'm giving them information, but they are responsible for improving this.
I can guide them, but it's their baby.
So there are some prompts that belong to marketing.
There are some prompts that don't belong to marketing and they have to be in
there because otherwise the company is not thriving, but they should be someone's
responsibility.
And as SEOs, we tend to think that we can solve all the problems.
No, that's not how it works.
We cannot work on Asylum 100%, we need to align with all these many other areas
and even expand further our alignment in it.
To start wrapping up, I would love to hear your last tip by last resource
and last action that you think we haven't yet covered and you would love to
recommend an audience to use or to leverage or to think about when doing good
from research or when establishing their representative from library.
There's not something you can probably transfer across all businesses,
but in where you have this data available, I found it extremely useful to do.
And that's to use Google's own analysis.
it gives out to the public against it for the problem side.
So if you have a business which has a lot of good Google reviews,
Google already runs a heavy salesman analysis
and call out on it, so you can go to reviews
and it will pull out for you when you're looking at saying
like lots of people might have it.
If you're looking at hotel, for example,
lots of people might mention the pool's good
but the restaurant's bad kind of thing.
You're not going blind in trying to think
what people are saying or talking about.
And you can hardly start from your own reviews.
And Google's already kind of done heavy categorization for you.
So you can pull that out and use that as a baseline
to make it actually start to put together
separate line brews of problems for your tracking and testing.
More of a mistake that I see a lot of people doing
because when you plug, you said at the beginning of the day,
when you plug your brand into one of these fancy tools,
they come up with a lot of the problems for you.
I think I sound like a record record.
My first one is as multiple times we've touched on this.
The persona's really important coming up with that persona,
which has always been marketing anyway,
but now LLM's are punishing you
if you don't really know who your persona is.
So that's a big part of it.
But the other mistake I see people make
if they go for those pre-written prompts that come up
is the way that they're generated.
They're often basically building it
on a library of content you already have.
So the tool is reverse engineering,
the piece of content, or URL,
and then coming up with a prompt
based on the stuff you already have.
So it's almost just self-fulfilling your own prophecy.
And then so you'll put them all in it
or say, oh great, you've got, as I've said,
I don't like this metric.
You should break it down per platform,
but 60% citation visibility.
You know, oh great, I don't need to do much.
Well, the way it's come up with those prompts
is because it's reverse engineered
what it should rank for anyway.
So you should be high.
So there's a track that I see a lot of people fall into as well.
- I do have a tip.
So once you have a picture of prompts,
please use the tagging system.
Please use the segments.
The reason why I say this is that you could have
a really, really good sentiment overall
and then you dive into customer experience
and it's a doozy.
So this is really important as well.
You have to think about your prompts
as not only an artifact that you're gonna put in there,
but hey, you're collecting data out there.
How are you going to analyze it?
How are you going to make it actionable?
So please be kind to yourself
and actually think about that system
because otherwise you're going to end up crying,
hanging it by hand or uploading another CSV,
thinking okay, now my times are in
and then you're actually duplicated your prompts
and it's costing you twice as much.
And this is a confession.
So yeah, don't be mean.
- I will tell you recommend everybody to start with
a minimal viable list for a single,
you're more important audience,
you're more important product or service line.
And once that you have that tackle
and you validate that is what you want to accurately assess
and understand you can take it from there.
The problem that I see is that money, money, time,
we get quickly overwhelmed
because we want to start with everything.
And of course, if you are working with any commerce brand
that has dozens of product lines, it's very challenging.
Another thing, we thankfully are getting more and more data.
First party data rather than just relying on third party.
So a big, a big thank you to the Big One Master tools team
because they have been the pioneers in this area.
They are giving us more, not only prompts,
but their topics and the intents behind them.
Again, please apply the recommendations
that we have all shared before,
rather than obsessing over their wording, et cetera,
because it's not for that.
It's to identify patterns, to understand better trends
or the trends behind the behavior.
They can be used and should be used very much
as an input to create your prompt libraries.
The same word also asks,
speaking about Mar William's cook, also wonderful tool,
can be used for this as well.
And if you have a little bit of more budget,
there are also tools like similar web as well
that show the prompts that are generating you a traffic.
But of course, these are the very,
but none of the funo that are not necessarily representative
of all of the prompts are being asked out there,
which is one of the mistakes that I see people doing, right?
Oh, because these are the ones that are generating
or are meant to be generating clicks,
I am going to only focus on this.
But we know that this is not how AI platforms work,
that you want to get cited.
And these are not the only ones that are actually representative
of the whole customer user journey, indeed.
Thank you all very much for your recommendations,
for your tips, for your insights, the mistakes.
Hopefully this helps SEOs, AI Search,
especially is out there, to advise their clients.
And she showed proof that they are not the only ones
going through all of these challenges,
but you can also follow them because they're always sharing
a lot of resources and insights about these
and all the AI search topics.
Thank you very much for your attention today.
I hope that you have found today's episode
very actionable, very useful.
This is one of the many more episodes to come
into this new AI search optimization series.
I want to thank again, our sponsor, similar web,
for today's episode.
And if you don't want to miss the other episode,
subscribe to the channel.
And follow again, our today's experts
in order to learn more about AI search optimization
in the right way.
And until the next edition, bye bye.
Podcast Summary
Key Points:
A keyword is not a prompt—AI search requires context-specific, human-like questioning that reflects real user intent and audience characteristics.
Adding audience context (e.g., "I’m a crossfit athlete") significantly improves prompt consistency and relevance, reducing volatility in AI responses by up to 60%.
Over-reliance on pre-generated or reverse-engineered prompts leads to self-fulfilling prophecies, where results mirror existing content rather than uncovering true user behavior.
Prompt tracking must be segmented by audience, product line, and platform, with prioritization based on business goals and real-world impact rather than volume alone.
AI platforms show significant drift—Google AI Overviews is the most erratic—so tracking should be done per platform to understand unique behaviors and biases.
Historical data from Google Search Console, social listening, and reviews should be integrated to build a holistic, non-linear understanding of user intent.
Sentiment bias in prompts can distort results—asking negative questions may falsely flag negative sentiment when it aligns with a brand’s value proposition.
A minimal viable prompt library for a single, high-priority audience should be validated before scaling to avoid cost overruns and data misrepresentation.
Summary:
AI search audience research requires a fundamental shift from traditional keyword-based methods. The key insight is that prompts must include real audience context—such as "I’m a crossfit athlete"—to produce consistent, relevant results, as models respond differently based on user profiles. Many errors stem from using generic or reverse-engineered prompts that reflect existing content rather than genuine user intent.
Over-reliance on volume and pre-built prompt libraries leads to noise and self-fulfilling outcomes. Instead, researchers should prioritize a minimal, segmented, and context-rich prompt library aligned with specific audiences and business goals. Data must be gathered from multiple sources—Google Search Console, social listening, and user reviews—to form a holistic view of user behavior.
, Google AI Overviews changes results 40% of the time), and sentiment bias must be carefully monitored. , marketing, CX), and decisions should be driven by actionability, not vanity metrics. Google’s own review analysis offers a valuable baseline for identifying recurring user concerns.
By focusing on patterns, real constraints, and brand safety, teams can build a cost-effective, strategic AI search research framework.
FAQs
People often treat keywords as if they were prompts, ignoring the need for context. This leads to inaccurate, noisy results that don't reflect real user behavior or intent.
Adding context (like being a crossfitter or a rock climber) helps AI generate more relevant results. Without it, results vary significantly—up to 60%—due to lack of personalization and user-specific constraints.
Tracking too many prompts without prioritization or context leads to high costs and poor returns. Many users overextend their prompt libraries, failing to focus on high-value, audience-specific journeys.
No. Each platform (like Google AI, ChatGPT, or Bing Copilot) has different behaviors and audiences. You should report and track per platform to understand true user behavior and platform-specific trends.
Start with a minimal viable list for one key audience or product line. Validate results before expanding. Only include prompts that align with business goals and are owned by a specific team (e.g., marketing, CX).
No. Pre-written prompts are often reverse-engineered from existing content, creating self-fulfilling prophecies. They may reflect current rankings rather than true user intent and should be reviewed critically.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.