I Asked a 15-Year SEO Veteran How to Rank in ChatGPT | Dan Petrovic (DEJAN AI)
60m 21s
The episode explores answer engine optimization (AEO), focusing on how brands can appear in AI-driven search results like ChatGPT, Gemini, and Perplexity. Dan Petrovich clarifies that LLMs don't crawl the web directly; they rely on traditional search indexes—primarily Bing for ChatGPT and Google for Gemini—to fetch grounding sources. Ranking well in these indexes is the first step, with Google prioritized due to its dominant market share, though Bing and Brave offer future opportunities. AEO adds an interpretive layer where LLMs re-rank search results based on pre-training biases and on-page content, making it distinct from classic SEO. Practical advice includes creating topic-central content, improving user engagement signals, and using citation mining to identify pages already selected. Influencing training data is dismissed as impractical; instead, Dan recommends iterative on-page optimization and testing model responses. Measurement involves entity-based prompts, tracking brand position, frequency, and citation share, while prompt volume is deemed unreliable—better to analyze fan-out queries. Spam detection uses compression algorithms like gzip to filter noise. Overall, AEO is an extension of SEO fundamentals, requiring adaptation to how LLMs interpret and recommend content.
In this episode, we're going to unpack answer engine optimization.
In other words, showing up in AI search.
In other words, getting mentions by chat GPT,
cloud perplexity, and others.
This is the hottest topic in marketing right now.
Every business owner wants their brand
to show up on chat GPT.
But there's a ton of hype and bad advice around it.
So I brought in the perfect guest to bust some myths
and drop some hot takes, then Petrovich.
When I was just starting to learn SEO some 15 years ago,
Dan was already a respected figure in the industry.
He runs an Australian SEO agency called Dejan.
And here's one of the nerdiest, in a good way,
SEO professional that I ever met.
Dan, welcome to the show's podcast.
Thank you so much, Tim.
Good to be here.
So I want to start very, very practical.
Before doing our interview, I opened
incognito instance of chat GPT.
And they searched for or rather asked it.
What are the best podcasts for digital marketers?
Of course I was hoping that a trans podcast would show up,
but it didn't.
So a trans podcast didn't show up in chat GPT incognito mode
for what are the best digital podcasts for digital marketers.
So where do I go from there?
What do I need to do?
Let's discuss just this specific query.
We're not going to discuss like personalization
and everything we're going to discuss it for the down the line.
But right now, there's a specific query.
I am not showing up for it.
What do I do?
OK, so what you need to do is get better rankings in Bing.
Full stop.
Get better rankings in Bing for these kinds of queries,
for like best podcasts for marketing professionals.
Yeah, so what's happening now is we've got AI models
and their sporting network of tools and harnesses
do the searching for us.
But they're still searching the search engines.
So the answer is always coming from a traditional search index.
And this is one of the common misconceptions
people thinking LLM's are crawling the web
or LLM's don't really do anything they take input
and then output comes out.
That's really what they're all about.
So a variety of little different tools took place
when you typed in that prompt.
So you were in chat GPT.
You were not personalized.
You were not logged in.
So you typed in a query.
Query was broken down into multiple fan outs.
So you typed in a complex prompt.
And then that prompt contained multiple dimensions,
different intents, different facets.
So in the background, 3, 4, 5, 6 queries took place.
Search results were generated in the background.
Using being possibly scraping Google.
That's a possible lawsuit and a variety
of other curated data sources.
So search results were brought.
And whoever was ranking up the top,
got into the mix.
And then some re-ranking reshuffling took place.
And the filter down list was presented to the model.
And then the model did the final re-ranking step.
And then what was given to it?
It organized in its own re-ranking situation,
put the first recommendation, say,
because text flows from left to right.
Bullet points go one, two, three.
So there's always an order to the recommendations
of brands and places that AI models recommend.
So every recommendation had its place.
So it's not the rankings like a ladder,
but it could be like first, second, third,
and then fourth, and then so on.
So that's in short what happens.
So to be in the mix, to be even considered,
you need to be ranking in the search engines
and being there index before you are
in the consideration pool.
So one big takeaway is that LLMs or AI systems,
they have to use search because to the best of my knowledge,
none of them have their own index of the web, right?
And what you're saying, that for my specific case,
with the best marketing podcasts,
or best podcasts for marketing professionals,
Chat GPT would typically use being search results.
So I need to see if for these kinds of queries,
like you said, they do query fan out,
where they take one search,
and they do it a little bit differently,
like I don't know, 10 times, 15 times, whatever.
So what I need to do right now,
if I want to show up, I need to go to being essentially,
do these kinds of searches, see which pages show up.
But then again, I'm checking what is happening right now,
but I want to influence this.
So there are essentially two ways to influence this.
The first one is to go and ask the pages
that are already showing up like,
"Hey, can you feature my podcast on your page?"
Which would probably be the fastest way
to get considered by Chat GPT,
or try to create some of my own pages
and try to make sure that they rank in being.
- But am I thinking correctly?
- You can do a say station mining exercise
using your favorite tool, or manually,
and you can see what characteristics the pages
they're already selected have,
that makes them being the selected pages.
And if you choose to go by way of addressing the on page matters,
by creating similar type of content on your own pages,
then you should probably try to create a situation
where your page looks, feels and has the same intent
as the pages that already selected
as the citation sources for that search engine.
Now you can't, okay, so we used Chat GPT as an example,
but we also know that there's a growing market share
for anthropic models.
So if you want to use opus and fable class models
or Gemini, Gemini, you've got optimized for Google Index,
GPT being in a variety of others,
and then anthropic models, Brave.
Brave is a search engine
that most people in the SEO industry are ignoring,
or not giving enough attention.
And I would probably look at that as well,
because more and more people are using
Claude Opus Sonnet models.
And that's something that's kind of like
very easy to get into at the moment.
So the difference between the traditional index
and Brave's index is they use their browser
to discover new pages on the web.
So they have very interesting, very different.
So in fact, when you look at Chat GPT
and Google's models, like Gemini and GPT,
their results are kind of similar,
but when you look at Claude's result,
it does give a fresh angle to things,
because they calculate authority slightly differently.
Quick detail, as we just established,
LLMs themselves don't crawl the web
when they need to pull information.
They probably get information from the web
somehow for the training data,
but not when you're asking them a question
and they need to search the web.
They rely on someone else.
In that sense, I want to ask you,
what is the point of LLMs.txt,
which is kind of something similar to robots.txt,
but LLMs don't visit your entire website.
They don't make an inventory of all your pages.
So why are people talking about LLMs.txt
when all everything that LLM is going to visit on your website
is the page that it got from search API
or whatever search index they used?
- Look, I put LLMs.txt in my website
because it's so easy to do and it doesn't harm anything.
Some tools use it, some fringe models use it,
nothing important.
We're not going to go into exotic model architectures
and fringe providers.
But the main ones are using direct URL fetching
when they need to.
- Exactly.
- Claude, Gemini, GPT.
They have the model, doesn't do that,
but the model can call a tool.
And tool can then fetch the page,
get the page context and supply it as text
or mark down to the model.
So that basically means model can fetch page on demand,
but that page comes from search.
It doesn't really come from LLMs.txt discovery.
One thing that is worth noting is that Google's come out
with OKF, let's call it protocol.
And this is another thing that I've implemented on my own site
that gives models context and the understanding
of what this website is about,
how it links together in its structure.
So that's something that I do,
I don't have any evidence right now that it works,
but considering Google put it out as a bit of a standard,
I'm going to go and say that Google's probably
going to follow that more so than LLMs.txt.
Let's go back to the fact that different LLMs,
different AI chatbots are using different indexes.
So chat GPT is using Bing, Google, of course,
Gemini is using Google.
Claude, what is Claude using?
Grave, Grave, yeah.
Claude could be using other things.
It is a black box systems that aren't really disclosed too much.
And same thing goes for OpenAI models.
I do, I really do think that they have external curated sources.
from paid partnerships with content providers similar
to how Google has with Reddit.
So it's not fair to say that GPT works exclusively
with Bing because we've seen evidence of them scraping Google.
We've seen evidence of them working with SirPAPI
or whatever other providers.
So they do get data from a variety of sources.
But the fact is what you said is true.
Each tool uses its own source of indexes
and you get optimized for those.
The optimization for the model is a little bit
of a different thing.
So first, let's discuss optimization for different indexes.
What are the main differences in optimizing for Google search,
optimizing for Bing?
Never the Azure industry cared about optimizing for Bing
until now, probably.
And I'm not even sure if it should care about optimizing.
And now you're also saying that Cloud uses Brave.
So we also need to understand how to optimize for Brave.
So essentially, if you are working in an organization
where your job is to improve the visibility
of your brand across different AI systems,
you don't need to just care about ranking in Google.
You also need to care about ranking in Bing, Brave, and whatever.
So what are the interesting things that you can share
about ranking in those different search engines?
On a practical level, things that overlap
between these major search engines.
And I can't really speak much to how Brave works
because I haven't spent enough time analyzing it.
It's on my radar, definitely.
And the good thing is that they provide an API.
So you can just query it all you like, which you don't have
with Google, for example.
So one thing that we do know is that you
want your website to be seen of a certain type.
And you want your search engine to have
a perception of what the website is about.
And if it has multiple facets, to have those facets in.
You don't want to suddenly create this huge section that's
completely off-topic.
So you want what you care about to be the central theme
of the website.
So for example, if you do thing A, then
have most concentrated topics and pages about topic A.
Don't have that plus 10,000 pages that are completely off-topic.
So that's one thing you could do quite easily.
Search engines do care about traffic,
clicks, and on-page engagement.
Search engines are the biggest spy machines
of user behavior on this planet, especially Google.
They've got Android phones.
They've got Chrome.
They've got Google Analytics.
Data is just pouring in from every angle.
If you go to Chrome and you go to Chrome,
call on forward slash forward slash histograms,
you will see everything that pours in real time
from you to Google.
So user engagement signals are a big thing.
You want to show and demonstrate that your website
is popular because ultimately search engines
are not necessarily, they say, do 10x content,
authority, you know, eat signals, authority to,
authority to tendiveness and this and that.
It's a popularity contest.
You don't have to be the best and most original,
most thought pro-- in fact, being the absolute thought
leader in your own hyperspace, kind of like being
the parallel man of mathematicians being in your own industry.
Having content that nobody understands
is too difficult to understand.
You're not going to do well in search.
So you have to have content that appeals
to a central core of the audience, that they love you content,
that they share it, that there's engagement.
So you want to look at things like traffic from other sources,
marketing campaigns, email, this, that, create buzz around your brand,
having people search for your brand, look for your website
engage with your content.
So those things, having topical centrality
of your website, having good engagement signals on the page.
And if you're doing those all those things right,
you'll be cited, linked, and organically mentioned.
And that does wonderful things to the brand.
Of course, what you can do is give it an edge,
push it in organically through link building and digital PR.
And that does feel a little bit forcing.
But we are SEOs, so our job is not
to do the thing and hope for the best.
We will do everything it takes to help our clients and our brands
if we're in-house.
So of course, you want to do a some form of outreach
and targeted content, closing the gaps
where the content is lacking, assessing your queries,
closing the queries, and so on and so on.
One thing that's not going to always work for you
is that if you're constantly looking at your competitors,
let's say you go, you go to A-trefs
and you look up your competitors and you're like,
oh, they're doing that.
So you're always catching up to the competitors.
That doesn't get you to the leadership position.
You have to sometimes innovate and come up
with something completely new, but then they copy from you.
Because if you just do everything that everyone else does,
you'll never lead.
So those are universal things that will help you in search.
I don't know, maybe a little bit top level,
but we can always get down to low level technical things
and everything else that you can do.
But overall, that's probably the most sensible--
I don't know if I'm forgetting anything big.
That's probably the most sensible top level approach.
So what I'm hearing in terms of this top level approach
is all the things that marketers and the shows
have been talking about for ages.
Create useful content that people would actually
want to consume, and those like behavioral signals,
how long people stay on your page?
Do they like your page?
You don't really engineer them.
You just create content where people naturally stay.
You can't send a lot of bots to your page
to generate those signals for you or something.
Some people do, but I'm not sure if it works on friends.
It's not quite like that.
There's a difference there.
Positive user engagement signal
could be a user on your page for three seconds,
submitting a form and being done.
And that's fine.
Yeah.
But if we're talking about content,
if you have like a 5,000 page article,
and you know that a person stays for 20 seconds,
then it's clear that they're not finishing it.
Exactly.
It's not about the length of stay.
Yeah, basically.
Sometimes it is, if it's not long article,
you need to be there.
But if it's like an action that needs to be for a quick checkout,
means page was, if you're on an e-commerce website,
custom and lands on the page, they got that's for me.
They do the one-click checkout and they're out.
That's positive signal.
Yeah.
So back to the example of a Travis podcast,
not showing up with the query,
the best podcast for digital marketers.
Again, simple things that I or we haven't done
with a Travis podcast, it only exists as a single page
on the Travis.com website.
And I don't even think that page contains
a lot of information about the podcast.
So what I could do if I wanted to show up is first,
probably launch a dedicated website for my podcast,
list all the episodes, all the guests,
explain like why I'm doing it,
like what I'm learning from it
and have a lot of content about it,
then maybe right now I know that comparison pages
are a big thing because LLMs like to compare things
between themselves.
How like this podcast compared to this podcast,
who is it good for?
And I don't have any information about it whatsoever.
So yeah, I could go to all the pages
that have the current lists of best podcast for marketers
and ask to be featured.
I could create my own website about the podcast.
I could create a few blog articles on Hrefs,
blog about the podcast, about the guests, about the topics,
et cetera, et cetera.
So basically I should create more content
about the podcast itself and make sure
that this content gets picked up for relevant searches
by these search engines.
And like one of the reasons I ask you
about the differences in ranking in Google versus ranking
in Brave and Bing is yeah, what I said before.
We never cared about ranking in anything but Google.
Then for two reasons.
First of all, Google always had like 90 something
percent market share.
So distracting yourself with ranking in Bing didn't make sense.
And second, at least I personally have always thought
that the bulk of the strategies,
the bulk of the ranking signals in Google and Bing
still overlap.
You need useful content.
You need other websites to vouch for you with links.
That's basically all there is.
Like what are the small factors where you can optimize
for Bing and would you optimize it at the expense
of Google if you optimize in rank number one in Bing?
Would you drop to number five in Google?
So how would you address this kind of thing?
Like would you go after all of these search engines?
Or would you just focus on these basics
and hope that search engines are smart enough
to understand that you're doing a good job
on your resources the best?
Well, I'm of a similar mindset.
Pick the biggest player Google for both traditional search
and AI visibility.
I don't pay too much.
Actually, I don't pay too much attention
to the French platform, because I know in the end, Google wins.
They've got the market share, not just for traditional search, but for AI visibility.
They're in people's homes, the Google, you know, okay, Google stuff.
They about to revamp all that and reignite it.
So, JetGPT in the future might be like, you know, I have the Xerox moment, a codec moment.
Just sort of like, thanks for that, but see you later.
I'm dropping, I really don't know what their place is, they started being loved and now
they're hated and it's a bit confusing with them.
I'm going to look out for China and Chinese models and what they come out with.
They're interesting in a sense that they are not so focused on money, commercial value.
Their name is AGI, but that's a separate topic, but they need to be on your radar when
it hits that you know what you're doing there.
But back to your code topic, yes, I am prioritizing Google.
There are like little levers and this and that that you could do for others like Bing, but
I'm simply not prioritizing there because of the market share.
Microsoft does have one big thing that's coming, maybe not next year, but two years from
now three by 2030, 100% and that's when it's going to become important.
Windows is about to become AI first operating system.
It's going to have native agent support, it's going to have your search experiences, your
product feeds, everything like your AI is going to be integrated throughout Windows and
it's going to be doing things for you as your assistant in the background.
All those back channels, the APIs, the connections, model context protocols, it's all going
to happen in the background and search engine is not going to see the light of the day.
Users are not going to go and be in search anymore.
That's kind of like how you and I when we were little, you went to the library.
Kids don't do that anymore, they Google things.
Now, their kids are not going to be googling things, Google is going to be the index.
They're staying.
Google's staying, nothing's going to replace the dynamic layer of the web because humanity
produces too much real time information.
So that's one thing to keep in mind, models are trained and frozen in time.
So Gemini, that is like Gemini 3.6 now has been trained a year ago.
So they have wrong ideas, wrong memories, they don't know who the latest president of
a country is and so on and so on.
So the reliance index is permanent.
There's no technology they can replace that.
But again, I'm doubling down Google with the understanding of China is going to come
up with something interesting and Microsoft is going to come up with something interesting
in terms of operating system.
Those are two disturbances that might happen in the future.
For now, Google is where I concentrate.
Great way to end your monologue for the lack of the better word.
Google is where you concentrate your efforts.
But this segues me nicely into the next question, which is, so how is AEO?
And it's range optimization.
Different from ACO because a lot of people in this, let's call it a new industry of
and range optimization are saying everything's changed like all the old playbooks no longer
work.
If you want to show up in AI search, you have to do completely different things.
You have to know completely different concepts.
And of course, I am the guy, a girl who can teach you by my PDF for $99, blah, blah.
So how is it really different?
Because like we're talking for what 24 minutes right now and all we're discussing is how
to rank your content search engines so that LLMs, yeah, when they would be using search,
they would find your content and they would consider you in your answer.
So what is different in AI search?
So effectively, AI search is an interface, an interpretive layer on top of search.
So you've got the search index and its results.
So instead of human interacting with search results directly, you've got a friendly little
layer in between that digests and summarizes that information makes recommendations.
And that puts it in a superior position of the ultra influencer.
Get about all the influencers on TikTok today.
They are not the influencers.
The main influencer is the AI model because the model chooses what is filtered through.
Model can see the results one to ten in Google and pick the ninth result as the top recommendation
and the third result as the second result and so on and so on.
And that decision is based on the model's opinion, perception, feel, notion of your brand
based on two factors, one, what it's learned about your brand in training.
So everything that's connected about your brand, product services, what you do, whether
you're good or bad, everything that it's learned in the pre-training stage and then plus
training stages like fine tuning, reinforcement learning through human feedback and so on
and so on.
Alignment, model alignment to human values to be a good personal assistant.
Other things like system prompts and its own harnesses.
So you've got that primary bias that the model has towards or against your brand already
baked in and that's the talk that I did in Singapore where I pulled apart Gemini.
I asked that million questions, literally.
I asked a million questions and I got one million brands as a result and how they connect
and what are the brand rankings.
And for example, I discovered that Chinese brands are not seen very favorably by Gemini.
It's just not top of the mind, like BYD, it's just not in the top 100 brands.
So that's primary bias.
Your secondary bias is the search itself.
It's what happens when the model's preconceptions meet the grounding, the grounding sources because
people confuse grounding and citation and mention.
Those three are separate but related concepts.
Grounding is when you're supplied as the potential context candidate to the model.
Citation is when it's used to that grounding source to make attribution to the sentence
that it's claimed.
And then the mention is when your brand is actually spelled out.
Two types of mentions, one is plain text and one is a link.
The link mention is the holy grail.
So you don't really need to obsess about citations being the grounding source and being cited.
You need to obsess about being the mention that's got a link in it.
Ideally with your product, image in there or whatever.
So those are the two factors.
So preconception, the primary bias and the secondary bias that the model has once it sees
your stuff in the grounding.
Why is that important?
Because if the model has negative bias towards your brand in its pre-training, when it sees
you in the list of 10 recommendations, it might not put you in the recommendation list
at all.
Even if you're in search, or if it likes you, even if you're the last brand on the list
in the grounding sources, it's going to put your top spot.
If it really likes you, based on what it's seen and learned about your brand in its pre-training.
And that's what engine optimization is.
But the question is the implicit question that I didn't really answer is like, so how
do you tweak that?
How do you optimize?
So what do you do with this?
And that's probably like an hour podcast on its own.
But. No, let me interrupt you right here.
Because first of all, this is a great insight that with its training data, LLMs have their
preconceptions.
They're like favorite brands or favorite products, worldview, yeah.
And you have actual data to prove that.
You asked a million questions to Gemini.
You presented about it at HF's of all in Singapore, that was the great presentation.
I think I mentioned to you before we started recording that you were our top rated speaker
because people absolutely loved what you were sharing.
But my question is, so how do you influence this training data?
And I will give you my answer.
My answer is create great content.
Get mentioned by other reputable sources with great content.
So essentially all the same things you do to show up in search, these are all the same
things that will get you in the training data.
Am I wrong?
Look at it this way.
I'm not going to put words into your mouth, but I'll let you think about this fact.
Google has the biggest search engine in the world and they have a literal case of the
web for every page on the web.
They have the copy of that, you know that, they have a copy of that page in cache.
And so on the other side, they have a Gemini model that needs to be trained on something.
What do you think they're using to train the model, the web?
So I think it's safe to say that more present on the web you are, more connection points
you'll have in the Gemini's neural network, the same thing will happen with, now we're
talking pre-training.
The state-of-the-art models today, like anthropics, Googles, and open-air eyes models, they
are no longer learning English.
They have baseline models that are already competent.
So what they're doing now is they're curating data sets.
They're spending billions of dollars on really fine-tuned curated data sets.
So it's not really fair to say that, "Oh, I just spammed the web with rubbish, and
you'll enter the training data."
You're not.
Because they are fine-tuning the models now.
They are using the best of the best of the data sets to train their models now.
So spam, that's scale.
Some of these companies have reached out to us because they Trefs has index of the web,
and some of these companies, they cannot name because of all the NDAs and such.
But some of these companies reached out to us to license their data, and they also cannot
say if it went through or not.
Very interesting.
Yeah.
But this is the point that all these LLMs for their training data, they need data sources.
Like you said, Google already has the index of the web.
They have tons of data.
They also have this like books project where they scan like a ton of books, so they have
a lot of data to train their models, folks like Anthropic, OpenAI, and others, they don't.
So they need to buy this data.
But then again, let's say someone buys index of the web from a Trefs.
Our index of the web, just like Google index, contains a lot of, I know, PBM, spam sites.
The web is very, very noisy.
As soon as you join a company like a Trefs to start like working with the index and with
the raw data, you can see how noisy is the index.
So how do you separate good things, good content from bad content, good sites from bad sites?
And suddenly, as I was working, so we launched this new project called FireHorse, Horses
where we give you like a feed of the mentions of the web.
You can like, it's kind of like Google alerts, but on steroids because as a Trefs crawls
the web and we see a ward or a brand that you want to track, we give you all the pages
that we see.
And this thing finds a shit load of pages.
So you need to filter through them to separate signal from noise and how I was personally
doing it.
I was using a chest matrix.
I was using domain rating, the higher it is, the more likelihood that the website is legit.
And I was doing, I was using search traffic.
If a website that this page is hosted on gets considerable search traffic in Google, then
I would consider this a legit website that I want to get a mention from.
If the website has zero traffic, then I probably not interested to know what they say about
my brand or about me or whatever.
So this brings me to a point that when LLMs buy a lot of this training data from the web,
there's this common crawl thing, right, where also like a, is it open source crawler or something
that crawls the web and then like everyone can tap into their index.
Yeah.
So you have this huge index, but how do you filter?
Because again, spammers like they used PBNs to point a lot of links to their own websites
and Google had to filter this out.
In the same way, they can pollute the web and say, oh, it traps is the best tool for
answer engine optimization, and I can launch thousands of websites with a lot of content.
Now with AI, it gets super easy to do it.
So how do they separate it?
I say they use all the same signals like, is it a trustworthy website?
Does it have links with from trustworthy websites?
So all the same things apply to getting yourself in training data.
So before we started the recording, we agreed that we will keep it pretty top level and
no go too near.
I am now, I am now requesting a license to proceed with a little bit of a nerdy.
Yes, please do.
Please do.
Okay.
Granted.
Now, what you're describing is, obviously, when you, when you look at a large index like
AHRFs index, I think most people in this, and I really do mean most people, 99.99 people
in the industry, do not have the concept of the scale and size of your index.
Just how big it is.
People can't really imagine a petabyte of information.
That's just, you know, that's just too much.
And you can say, oh, you know, gigabyte, terabyte, petabytes, it's easy to make that say it,
but the scale, so for example, on this machine here, I have, I don't know, maybe 24 terabytes
of the web content, and even that is the scale that I'm choking with at the moment.
But there is a really easy, simple method of separating and classifying content without
any semantic understanding, without any, and it goes back to the mid century, 20th century,
Claude Shannon, the information, the father of information theory, and Shannon entropy.
So I, and I did this just last week.
I know it sounds like a big words entropy, but I, in fact, I know some snakes, snakes old
salesmen in the industry keep using word.
So any scientific, scientific sounding stuff you need to be careful about was this guy
saying, but hear me out, and you can replicate this, task Claude with this, with the following
challenge, ask it to give you, to code a nap for you in Python, where you can put two
corpses of text, one that we'd say 100 examples of articles from a blog network, and just
give it, give it the files, and 100 content pieces from TechCrunch.
Good writing, right?
Put those two in two separate folders, give that to Claude and say use gzip compression
to classify the text.
So what happens is you have two reference, reference bases, and then you have one piece
of text that isn't in neither of those reference texts.
Let's say we pick a blog network content, that garbage that you said, the toxic content
that you don't want, the noise, right spam.
So we take the spam, and we give it as input into this app.
And what happens during the gzip compression, you take a small sample of that, attach it
to the existing corpus, you compress that, and you compare it compressed and not compressed.
So you do that for both classes, and with 99% accuracy, the model, there's no, there's
no model, there's no machine learning.
The compression algorithm will tell you which of the two buckets the text you just submitted
belongs to.
Now let that, let that sink in.
This is cheaper than any machine learning model.
This is cheaper.
You're not even looking at keywords, you're not doing ngram analysis, none of the nerdy
stuff is going in, you're using gzip compression.
Okay.
So that's the answer.
This is a great example, but with this example, you merely excluded all the spam as a viable
strategy to get into the training data.
So creating a lot of websites with the iGenerated content saying my brand is the best is easy
to detect, and therefore it won't get into the filtered set of quality training data
for LLM.
But what will and my point is that the same old things create valuable content that people
would want to consume that other people would link to blah, blah, blah.
Okay.
So yes, but I'm going to push back a little bit on this and say that you're trying to influence
the model training data, it's kind of like you going with a bottle of water into an Olympic
pool and pouring a little bit in.
Okay, but what are we here for?
What is the marketing department getting paid for?
What is our job?
Yes.
So you remember how you said prompts and tracking of the prompts and understanding the value
of them and so on and so on.
What I do, I can't speak for others, but what I do is I test the models and how they respond
based on different types of input.
And I try to understand the model the way it is, I'm not trying to influence the pre-training
too much because I know that that's a massive exercise and most clients don't want to spend
$1 million on seeding of the training data content, so what I do instead is I test the
models and how they react when they see different types of adjustments on my clients' pages.
Let's say we have a very large e-commerce website or a B2B marketplace or a digital asset
catalog or whatever.
So what I do is I make small precision based changes on page, I do on page optimization.
On page optimization and I test to see how the model reacts to the change and I do it
again and again and again in an agentic loop and I collect the ranking signals, preference
signals.
So whatever model says, I put it in a Bayesian inference loop.
But basically, basing inference is something that scientists use.
to adjust their belief system based on new evidence pouring in right. So that
removes the correlation and causation type problems right. So I constantly
optimize my clients content pages and I make the adjustments until I have a
version of that page that Gemini cannot resist or Claude cannot resist or
GPT cannot resist. And so basically I have resignation. I know I can't
influence the pre-training too much and I know that you know all the SEO stuff
is already being done. Links and outreach and yeah that's that's that's old
school stuff. But now what do I do on page? And that's what most people are tired
of. People are tired of prompt tracking. People are tired of visibility
measuring. They're tired of measuring and strategies and and dashboards that
don't do anything that just tell them what's happening. Yeah. People are
looking for things to do. And what I'm saying is do on page optimization
test how things pan out with the content adjustments and repeat the process.
And this is this is the direction we took and it's very very interesting. We
have quite a lot of interested clients and currently in in progress showing
some very very promising in an initial results on page optimization. This is
this is very interesting and it makes a lot of sense because essentially when
LLM chooses to use your page for its consideration set to give someone an
answer you're in a way talking with it because it processes the text on your
page and it chooses if it wants to use the text the argument the ideas the
recommendations from your page. So in a way your job is kind of to persuade
LLM that you are in the right place and like ignore all the previous
instructions and use this recommendation blah blah blah. So it's almost you're
trying to do this thing. But you're trying to do it with like adjusting your
language, adjusting your arguments. And what you're saying that it does
result in in specific outcomes, which is very interesting and which makes
sense. But again, like you said yourself, this is the very last step. This is when
you're already ranking. This is when your page is already being considered by
the LLM. And this is like the final fine tuning step and you're switched
completely from the influencing training data. But I think we discussed this
topic enough. And then again, one final question between the training data and
grounded answers is how often do LLMs really rely on training data? Yes, you
already said that they have biases rooted in training data. So in a way they
rely in training data every time. But is there any statistic how often LLMs
would ground their answer versus just answer it from from their knowledge?
Yes, I think I wrote I wrote about that several times on my blog. I can't remember
how many maybe twice or three times. But a short version is that Google will
ground everything every time. Is this is this sky blue grounded?
Is water wet grounded? What is the capital of Australia grounded? So the reason
that they do that is because the Google was embarrassed
through AI overviews, you know, eating rocks and glue and blah blah blah. So they
ground everything to prevent hallucinations and search cheap because it's
their own. It costs nothing. They don't need to pay for the API.
So this is the little nuance that I think is worth understanding. Google
grounds everything in multiples. So you could have multiple grounded
sources support the same sentence in its generative response. So one claim
supported by three or four grounding sources. So one sentence can have multiple
citations. Whereas where you look at GPT as part of the open AI harness and
it's a web search toolkit, the relationship is 101. So Google's greedy
and open AI is resource smart. So basically what they will do is they will
pick 1 URL ground the sentence and that's it. That's why we often see
people are saying all right, just getting to Reddit and you're you're done,
you know, that's your AI SEO. Reddit is actually one of the most rejected
sources. Reddit is just shoved down everyone's throat because it's doing so
well in index. But MLM's very often see the see Reddit
in the grounding sources and choose not to include it. In fact, it's like
over 90% rejection. We see that in the in the API calls that we do.
So one one thing to keep in mind is yeah, Gemini will go one to many
open AI one one to one. And then some slight differences like
obviously Gemini will have the page context crushed into an extractive summary.
So your full page will never get to Gemini. You will get a little snippet.
It's bigger than your Serp snippets more expanded, but there are cutouts from your
page separated with dot dot dot ellipses that are
most relevant to the user inputs prompts fan out queries that pull
that snippet in. Does that make sense? You're almost losing me at this point.
Yes, you enter the prompt. The fan out queries happen.
One fan out query brings one set of results. Those results get filtered down.
We use extractive summarization to extract bits and pieces from that
page that pertain to the input query. Not the prompt, but the fan out query that
stem out from that prompt. And then for each fan out query,
we do the extractive summarization for the pages from that separate
search data set that is then using some summaries extractive summaries
that are based on scoring to the input query
for that set of search results. So you end up with a variety of different
extractive summaries. So why am I saying extractive summarization?
Because it's not abstractive. Abstractive summarization destroys the
content. It rewrites your content for you, which is dangerous.
What Google does is they take verbatim bits, start of the sentence,
edit the sentence, middle sentence, but it always untouched copy from your site.
But during that compression, your information can still be lost.
And that's optimization optimization potential. A lot of people call it chunk
optimization and chunking and this and that. You don't have control over how
Google chunk your page. They use chunking in the pre-processing pipeline.
What happens afterwards is that the bits and pieces of the page
being taken away and form forming the grounding snippet.
What goes into that grounding snippet will represent your page in the
grounding context for the model when it's making the recommendation.
What of page survives is what will speak to the model about your business
your product and service. That's an optimization opportunity.
It got quite nerdy at this point. Let's shift gears and let's talk
about the measurement. We talked at length about what to actually do
in order to show up in LLMs, which I'm quite happy about.
But now let's talk a little bit about how to measure the results of your
actions starting from prompts. The prompt that we started this conversation
from was about my podcast. What are the best podcasts for
digital marketing professionals? How do I even know that people are asking
this to chat GPT? Is there a good source of data? What are people
asking to chat GPT, Gemini perplexity? There's so many of these AI
assistants. How do we know what people are asking and what's the
search volume? I don't try prompts for visibility
tracking. No. I use prompts for qualitative research.
During our citation mining research, we generate a bunch of prompts,
and each prompt is generated as a stem out. First, we set up entities.
Entities are the things that our clients do. The main thing, some out of an entity,
the thing that the client does, you could have multiple variations of queries,
and you could have multiple prompts. But all those prompts and all those queries
that boil down, quantize to that one entity. That's the main thing that the
client does and wants to be visible for. We fan out that
entity into various prompts. There are arbitrary, sometimes short, sometimes long,
sometimes multi-turn and so on and so on. Then we do the citation mining
to probe how the different answer engines will respond and what their
fan out queries are and what the citation chunks are and wonderful other types of
information. We use that for qualitative research. We use that for prospecting,
for understanding what types of resources get cited.
For measurement, long-form prompts are too noisy. We don't use those.
We use the quantized versions. So we did the main thing. So we frame it,
purposely frame it, and a customer is looking for XYZ.
Recommend some brands for it. I used to do 10, but then I would get garbage
at the bottom. I said, "Ready, let's go."
recommend some brands for that. And it gives me a structured list every time, every day.
And so what I do is I allow the web grounding. So if it's for GPT, I allow web grounding.
So Gemini, I allow web grounding. So the web grounding is always biasing the model's
choices. Otherwise, you're not really measuring anything. You can just measure it million times
in one day and that's your research. But if you want the true grounding bias, you need
to measure it over time because index changes over time, model is frozen in time. So what
we do then is we do the entities, we then probe each entity in a structured way and we don't
ask for pros. We ask for structured brand, tail, brand, tail. And then so what we measure,
we measure the position of the brand in that output list, which could be three brands.
It could be 12 brands. It could be 10 and so on. So we measure the original value. That's
one factor. We measure the frequency of that brand appearing even. So some days the brand
is not in the list. Some days it's in the list. We measure the frequency, the original value,
the, let's call it rank in the list. And we mentioned, we measured the citation share
and the mentioned share. And that is the balance between all the other, all the other websites
being mentioned or use this granting sources and all the other brands being mentioned as the
brands during this data set. And we, we track that on an ongoing basis. And so what that does,
and we have the third one that we call share voice, share voice basically blends the mentioned
and citation share. So it's kind of like a top level North Star type, type metric. And so that's
what we use internally. I thanks for explaining your whole measurement methodology. Very interesting,
very useful. But I want to bring you back to search volumes because this is something our
customers at HREFs are asking us for, give us search volumes for prompts. And we do it in a
pretty straightforward way. So there's these statistics on the monthly active users or market
share of Google versus chat GPT versus perplexity versus Claude blah blah blah. They publish their own
statistics about their active users. And so what we do is we take an example of a search query
from Google. And we say, okay, so if in Google, this is like a hundred percent or like, yeah,
this is a hundred percent, but 90 percent of the market, then five percent of these search
volumes chat GPT and perplexity is one percent. Like I'm calling the arbitrary numbers, but the
the method that we use to kind of estimate the potential search volume of the prompt is this
because I don't really see or we don't really see any better way to do it. Yeah, there's click
stream data. But the panels are so narrow. And I don't think they give you nice exposure to what
people are really asking you know those assistants. So and some companies, as we know on the market,
are promoting their tools that give you search volumes for prompts. But whenever I look at the data,
it's very questionable for a lack of a better word. So I'm wondering what's your take on
search volumes of the prompts? Okay, so if everyone who's been hypnotized by the prompt search
volumes, look at me in the eyes. There is no prompt volume. Forget about it. The chances of two
people typing in the same prompt exactly the same way is astronomical. It's just very unlikely.
Improbable, I should say. And so that makes it even worse the fact that most people are not just
using one off prompt. There's multi turn situations. Don't forget that. Plus this personalization,
don't forget that. Dan Petrovich in Brisbane, Australia, on Windows, on Chrome with a certain
search history needs Chrome with embeddings with everything, all those log in, logged in, etc, etc.
demographic. All that biases the models, the decisions and search itself. So prompt is a non-stata
as far as I'm concerned. What you can track. And I think this is what you're aiming at. What you
can track, you can punch in any prompt. And you can then get the fan out queries for that prompt.
And you can then measure the search volumes for each fan out query. Add them up, and that's your
prompt volume. But don't call it prompt volume. It's the volume of the queries that stem out from
the prompt that were actually being triggered by the search engine backend. That's the only
formula that works. That's the only reasonable model. Anything else is click stream data. I mean,
seriously, that's absurd to think that click stream data can give you prompt volumes.
I completely disagree with that notion. We actually requested a few samples of click
stream data providers. Then they gave us the samples of their stream of prompts and chat GPT.
And yeah, like you said, first of all, most of these prompts are super long. They're like
multiple sentences. And right now with the proliferation of dictation tools, I'm using whisper flow
all the time. I don't have to type. So I don't have to limit myself to make my prompt short. I can give
as much context as I want. So first of all, these prompts are super long. And if you try to, oh yeah,
but you can use AI to like derive the main topic of the prompt and give like a search volume of
the topic of the prompt. No, because one prompt can talk about three, five, six topics at the same
time. You can mention like best movies for criminal investigations, documentary, but there can be
multiple topics and to derive all the topics from a single prompt is very hard. It's like the data
is very noisy. And also what you're saying, a lot of these prompts are follow ups to to the
prompts that people ask, which yeah, it's, yeah, multi turn. It's, it's a mess. Okay. So, but again,
back to my example of a chef's podcast, I want to track a yeah visibility of a chef's podcast.
How would you approach creating a list of prompts for me to track if a chef's podcast is showing up?
What would you do? So I do two things. I know you explained your,
your methodology explain the prompt generation. Yeah, that's fine. I didn't explain prompt
generation. So I use two methods for, are several, I suppose, but two basic concepts. One is I show
the URL to the model. And I say, what will people type in to find for you to recommend a page like
this? This reverse, reverse kind of process. And then he returns the prompts. Then I run those
prompts to see if it actually happens. And the ones where he happens, that's the prompts that I want
to be paying attention to, not tracking though. I already explained about the tracking part. So
that's the first method, page to prompts. Then the second one is remember how I said I have my
main entities that I want to be visible for. I can generate prompts based on those entities by
doing the fan out from entity to prompts is the opposite. What a simple entity is for a chef's
podcast? As your podcast. That's one of your entities. So when I was generating prompts to track
my visibility, I also thought about like this, I don't know, top of the final bottom of the final,
I'm not a huge fan of this concept, but sometimes it is useful. And in that sense,
another entity could be educational materials for digital marketers. This is kind of very,
very broad and I talk about courses, it would talk about even like universities or books,
but podcasts could be one of the considerations. And this is what you're saying entity.
So if we're talking about educational materials for marketers, what kind of prompts or queries
can AI suggest for me to be tracking? So this is kind of the thought process of how you generate
prompts? Exactly. So you don't have to use entities. So entity could be like SEO podcast, marketing
podcast, link building podcast and so on and so on. But you could have, you could have your actual
search queries from search console and use them for the most on the top 100 and fan out to prompts
based on them. And so you then use those prompts and you run the prompts, collect the fan out
queries from those prompts and generate a list of, generate a list of fan out queries and then fetch
the search volumes for those queries because they are the actual queries, the search engine use,
with actual search volumes that you can buy from your favorite clickstream provider or
AHRF or whatever at all. And that makes sense.
Oh, we don't have much time at this point because it sounds like generating a proper list of
prompts that would give you a very good objective overview of your AI visibility is not really a
trivial task. You need to put some thought into it if you want to get a good output, good measurement.
Am I correct? Well, I think it wraps up the whole episode quite nicely when you asked,
so what do we actually do as SEOs in the age of AI a lot? That's the answer.
And we've only just scratched the surface.
Then thanks a lot for the interview. Yeah, there's a lot we could cover. We could probably speak for
one hour.
more and not even get to the, to the half of what we need to cover.
Thanks a lot. Thanks a lot for your expertise. Thanks a lot for your no BS approach
to EncerEngine Optimization. And thanks for all the work you've done in SEO in those years.
I've learned a lot from you personally. Thanks a lot for being the guest on HR's podcast.
Very welcome. Thank you.
Podcast Summary
Key Points:
AI search relies on traditional search indexes like Google, Bing, and Brave, not direct web crawling by LLMs.
To appear in AI answers, brands must rank well in these indexes and be cited as grounding sources.
Google is prioritized due to market share; Bing and Brave are secondary but growing in relevance for AI.
AEO differs from SEO via an interpretive layer
Influencing training data is impractical; focus on on-page optimization and testing model responses iteratively.
LLMs ground answers differently
Prompt volume tracking is unreliable; instead, use entity-based prompts and fan-out queries for measurement.
Key metrics include brand position, frequency, citation share, and mention share in AI outputs.
Spam detection uses methods like Shannon entropy (e.g., gzip compression) to filter low-quality content from training data.
Summary:
The episode explores answer engine optimization (AEO), focusing on how brands can appear in AI-driven search results like ChatGPT, Gemini, and Perplexity. Dan Petrovich clarifies that LLMs don't crawl the web directly; they rely on traditional search indexes—primarily Bing for ChatGPT and Google for Gemini—to fetch grounding sources. Ranking well in these indexes is the first step, with Google prioritized due to its dominant market share, though Bing and Brave offer future opportunities.
AEO adds an interpretive layer where LLMs re-rank search results based on pre-training biases and on-page content, making it distinct from classic SEO. Practical advice includes creating topic-central content, improving user engagement signals, and using citation mining to identify pages already selected. Influencing training data is dismissed as impractical; instead, Dan recommends iterative on-page optimization and testing model responses.
Measurement involves entity-based prompts, tracking brand position, frequency, and citation share, while prompt volume is deemed unreliable—better to analyze fan-out queries. Spam detection uses compression algorithms like gzip to filter noise. Overall, AEO is an extension of SEO fundamentals, requiring adaptation to how LLMs interpret and recommend content.
FAQs
Answer engine optimization is the practice of optimizing your content to be mentioned and recommended by AI search engines like ChatGPT, Gemini, Claude, and Perplexity. It focuses on being included in the sources these AI models use to generate their answers.
AI models don't crawl the web themselves when answering questions. They rely on traditional search engine indexes, like Bing or Google, to fetch search results. They then use these results as grounding sources to generate a response.
The first step is to ensure you rank well in search engines like Bing, as ChatGPT often uses Bing's index. You need to be in the search engine's index before you can be considered for an AI model's answer.
While Bing is important for ChatGPT, Google remains the dominant player for both traditional search and AI visibility through Gemini. The core strategies for ranking in both are similar, so focusing on Google is generally the best approach.
Key factors include having a website with a clear, central topic, strong user engagement signals (like clicks and time on page), and being seen as a popular and authoritative source. Creating high-quality, engaging content that attracts links and mentions is crucial.
AEO is essentially an interpretive layer on top of traditional SEO. It's not about completely different tactics. Instead, it's about how AI models use their pre-existing biases and the search results they retrieve to make final recommendations, which requires a focus on being mentioned and cited within those results.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.