Should You Block AI Bots From Your Website? | AJ Ghergich, Botify [Part 3 of 3]
18m 19s
This episode of "Found an AI" with Cassie Clark and AJ Gurgich focuses on AI bot governance, a topic often overlooked. AJ, a Botify consultant, explains that governance is not a simple binary decision to block or allow bots but a set of decisions involving marketing, IT, and legal. These teams must collaborate because marketing cares about visibility, IT about server load and costs, and legal about liability from inaccurate AI-generated content. AJ emphasizes starting with three critical questions: knowing what data trains AI models, ensuring its accuracy and currency, and having a plan for AI retrieval in training and real-time scenarios. He notes that many large brands cannot answer these yet. A good governance plan uses data to differentiate between training and live retrieval bots, especially for retailers who need to win at live retrieval moments (e.g., stock checks). In contrast, a bad plan is binary and lacks data. AJ also advises against prioritizing the llms.txt file for most brands, as it offers little value compared to other tasks. Cassie adds a nuance that Google’s new audits on agent functionality do not change this, as they focus on functionality, not discovery. The episode concludes with a call for cross-functional collaboration and starting with AJ’s three questions to address AI bot governance effectively.
Hey, real quick before we get into this, here's something that I've been noticing lately, and I'm seeing this with a lot of love. A lot of really smart marketers are jumping straight to the tools. They're on LinkedIn, they're asking, "Hey, what's the best say I've visibility tracker? Should we be using this tool or should we be using that one?" And these are all really good questions, but I'm seeing a pattern when we're reaching for the tools before we actually understand the fundamentals of how AI search works, like why some brands get cited and others don't, or what signals AI engines are actually looking for, or how citation and trust scores actually function. So I'm working to fix that. I'm in the planning stages of a free GEO course, a how-to GEO kind of thing, that walks you through the foundations before the tactics, because one to understand how it actually works, the tools start making sense. Without that foundation, you're really just buying a dashboard. The course is free, it's going to go live on YouTube and the podcast, and there's going to be a workbook to go with it, so that you can take notes and apply what you're learning. The newsletter is the first place I'm announcing it when it goes live, so sign up for the visibility report to get on the list, and you'll know the minute that it drops. Okay, to the episode. Hey, welcome back to Found an AI. I'm Cassie Clark, a fractional content strategist, an AI search optimization expert, and the host of the show where we talk about AI search, GEO, AEO, and one, all of it actually means so we don't get lost, signals to new ways of user search behavior. My cat is at my feet under my desk, and he's making a lot of noise. He's orange, and he's sneezing. He's not the only one sneezing. I am also sick. It's kicking my butt, and you can probably hear it in my voice. I'm sorry that you had to listen to that. But today we're getting into something that does not get nearly enough airtime. AI bot governance. I don't mean that in the dry compliance, he kind of way, but I mean it in the real messy question that every company is wrestling with right now, which bots get access to your site, what they're allowed to crawl, and what if the content they grab is used to train a model, or to answer customer and live retrieval? Spoiler, those are two very different problems with very different answers. This is the third part of the series with AJ Gurgich. He leads consulting in the AI team at botify. He's been having these conversations with some of the biggest brands in the world every day about what to do with these AI bots. We get into why governance foster the cracks between marketing IT and legal. The three questions that most of you most still can't answer. What a good plan looks like versus a bad one. And I take us on a little side road about the LMS.GXC file and kind of think goodness that I did at the time because Google is changing guidance on that by the minute it feels like. Let's get into it. When we say bot AI bot governance, define that for us. Yeah, I think governance means different things to different people. Obviously, I think inside of a company right now, the main question people are saying is AI governance marketing IT or like a legal problem. And then to be the obvious answer, I guess, is it's all three. And that's why it falls through the cracks. So it's really a governance is a set of decisions about what which bots can access your site, what they can crawl, whether the content that they get used is for training or for live retrieval for a customer. And everybody cares about it, but for different reasons. So like marketing cares because it affects visibility. IT cares because they have that issue where they're getting inundated, crushed with bot traffic that they have to serve. And that's a cost and an infrastructure problem versus humans. So it's like this huge load on the IT side. And then legal cares because maybe your product data showing up you an AI you didn't authorize, maybe it's hallucinating facts and now are you liable for that? It said it was good for this, but it's not. That could be a problem like whose fault is that going to be? Is it going to be open in an eye's fault or yours? So everyone cares about it, but I think it's really a combination right now of marketing IT and legal. Just wrestling with those questions. And I think that kind of leads to the obvious question of like where do you start then? Where do we start? Where do we start with that? So again, I get to, I mentioned I get talked a lot of folks, a lot of big companies, global companies, and growth companies. So this is the type of conversation I have all the time. Like how do we start? What should we do? That kind of thing. I think it's, you start with some simple questions. So as a brand, do you know what data of yours is being used to train AI models? And if you don't have, like if you have like an extremely loose answer to that, and you're not just a one woman shop, you know, like you should be able to answer that question as an organization. And then comes the next question. Is that data accurate? Is it updated? And then do you have a plan in place for what you're going to allow AI to retrieve for training and in real time? And I, like I said, talked a lot of brands. Most cannot answer that question. Some of the biggest brands in the world, at least I would say a year ago, it's getting better now could not answer those three questions. So I would ask the CMO to ask her those three questions, couldn't answer it. So I don't want anybody who is like, I can't answer those questions to panic. Like you're not, it's not uncommon. But I do think it's, you need to start thinking about that. You need to start thinking about what bots are hitting your site, how often, what information they're getting, all the way down to like, let's say you're not a technical person. Let's say you, you're like, "Ajami, I work on branding." You know, I'm a brand manager. Okay. Well, what if OpenAI has three CMOs ago mission statements about your brand and its training data? How could that create confusion about what you say your mission is today, right? What about your fonts? Have website, your colors, everything. Like think about all the confusion that's out there that it may be having from old deprecated data about your brand and your positioning versus today's positioning. Like, so this isn't everyone conversation. You can't just go, "Well, that's IT. I'll let them deal with the bots." It's like, no, it's a brand conversation, right? So everyone should get involved and then the question I think becomes, you know, who's running the show? Right. Right. I mean, I've been saying just for like AI Sir Documentization in general that this is not just a marketing thing. It is, but it touches every layer of the organization. So even with this conversation, it's still every layer of the organization. So let me ask you a question because I just was talking to a brand earlier in the week. She was asking me about the files that you may or may not need on your website. Talk about the robot.txt file and the LMS file. Do we need those? So this is the new like SEO bar fight question. It used to be root domain versus subdomain, which I won't dork out, but if I'll just say this, if your IT ever asks you if you want to put some content on a subdomain or your main domain, put it on your main. So it's better for everyone. But that used to cause all kinds of issues with inside the SEO community. This is the new version of that. The short answer is for 95% of brands, you don't need the LLMS TXT. It's not getting picked up in mass. Goals, representatives will come out and say they're not using it. Now that said, you can we track, we modify track's log files. We look at it. They do get used in hit. It's just, here's the thing. There's 10 other things you should be doing. This is the 11th. And so that's where I'm just like, when you're doing this, you're not doing something else. And I much rather you be doing something else than that. And then you got to maintain it and all that. It feels good. It's a shiny object. It is like, we're cool. We have a LLMS TXT. It's not going to do you any good. It's not going to ring the register. I wouldn't do it for 95% of brands. Okay. That one hit my head. I'm like, I have to ask it right now. But let's go back to the AI bot government plan. What does a good play look like versus what does a like a bad one look like? And when do you just call it done? Yeah. So I think a bad plan is very binary. Like we block AI bots. We don't. That kind of thing. You need to have enough data at your fingertips to know what are are good bots than what are bad bots. And then what bots--
our training bots and what bots are live retrieval. So for a retailer, a lot of folks think I need to get my data into ChatGPT. They just need to know about my brand and like my product. But you gotta realize, so while we're talking, maybe it's happening or maybe it's gonna happen an hour, like ChatGPT is gonna release their new model, 5.5, right? And so when that happens, the moment that happens, it's gonna be six months out of date, maybe a year out of date. And so my point is it doesn't know if those sneakers are in size 11, if you know, pink is in stock, it doesn't know, right? And so most retailers actually need to win at the moment of live retrieval. When a search, when an AI searches on a consumer's behalf to see if size 11 is in stock in pink at this price, what is the shipping? And so a lot of folks don't think about it that way and they're not looking at it that way. I would do it almost like a governance plan should include like user testing, but for AI agents, right? So can the AI agent at live retrieval see if it's in stock in pink in size 11 or not? Or does it have trouble and accessibility? So it's not a binary choice. You have to have data for the plan to work. And then honestly, it really should have a hub. Like I believe that AI could be a hub for all of your spokes. And so you see a lot of folks putting a chief AI officer, you could also even think about your chief customer officer becoming having AI in her purview perhaps, so that you're connecting all of these different spokes under this AI hub because you're right, it's not just marketing, it's everybody. - And absolutely everyone. So I mean, you make an interesting point about the training data and live retrieval being different. So if you're noticing that your training data or what the training data says about your brand is not correct, how are you going about fixing that? - Yeah, so it's a bit of a longer cycle obviously, right? 'Cause it's already out of the bag. - Right, it's already in the training set. The first thing I would be doing is figuring out where the heck it got the information. So I'll leave their name out of it, but a large cosmetics company on the phone with them where this is maybe a year ago going through their AI visibility and looking. And essentially the AI started falsely saying that their cosmetics were not like animal friendly. And like imagine the, it's storm that might cough, right? Like it's like, that is a no, are you kidding me? And it's like where is it getting the information? Because sometimes it's pulling the information from Reddit or a forum. So sometimes it's your own information. Obviously in this case it wasn't there information while somebody else falsely stating. So the first thing you want to do is hunt down where is it getting the information and the facts? And that's why you need everybody at the table. You need marketing, brand, SEO, you need everybody at the table because why does it have our story wrong? So is it because it can't read JavaScript since getting half our content wrong? Is it because our data is telling three different stories? Is it because it's reading some sweaty redditors post that is not accurate? And we need to talk about PR and going in and correcting the record. So all of those things would be good. And then proactively is if you can get into a mindset, especially if you're e-commerce and retail about pushing, getting as much information in your data feeds to push. So you'll see in the next coming weeks and months, all the major players will be adding agent to feeds for commerce. That's a great way to correctly record and add your data because you're pushing versus waiting for it to crawl and learn, so you're pushing your information in. So it's a complicated answer and not a perfect answer, but you've got to find the source of the rot at first before you can do anything. So you said everyone at the table is there anyone that we're missing? So we're talking about brand marketing, legal, IT, anyone else? Honestly, I would start there. I think every organization is going to be a little different. I wouldn't say that's a closed group. There's going to be-- there's going to definitely need to be more people on the loop. But that's where I see 80, 90% of people starting is with those people at the table. And then I would widen the tent from there. But yeah, you can't go too wrong if that's your starting block. OK, that was the final part of my conversation with AJ. And there are a few things that I want you to walk away with. AI bot governance isn't a binary block. The bots are don't kind of decision. It's a set of decisions that needs a marketing, brand, IT, and legal at the same table, because no single one of them owns it. Start with AJ's three questions when you bring everybody to the table. First, do you know what data of yours is training the AI models? Is that data accurate and current? And do you have a plan for what you'll allow AI to retrieve both for training and in real time? If you can't answer those yet, you're in a good company. Everyone is still trying to figure it out. But that's the work that we need to start with. Now, quick update on by allola.ms.dxcd.org. The story has moved a little since we recorded that. We recorded this a month or so ago. And then now there's new guidance out on the show. AJ's take was that for 95% of brands that file is not worth your time. And then Google two weeks ago has said, hey, you don't need this for visibility. That's Google's official line still. But they've also added an llms.dxt check to girls light house under a new agitated browsing set of audits that flag whether your site even has the file. So naturally people saw that and went, wait, I thought we didn't need this because it was announced within the same week of Google saying you don't need it. Here's a nuance with all of this. Those audits are about how ready your site is for AI agents and browser tools, not just search rankings. Google's John Mueller framed it as the difference between discovery, getting found, and functionality, meaning helping an agent do its job once or already. It's already on your page. So it doesn't really contradict AJ's point for most brands. If you're a developer docs or heavy agent traffic site, it might be worth looking into putting this on there. For everyone else is the same answer. There are 10 more important things to do first. And this helped you. This episode helped you any hit subscribe. I would love you forever. I'd love you even more if you had 30 seconds and you could just leave a review. Those reviews do help the show get in front of more marketers and founders who are trying to navigate this stuff. And if you're starting your content strategies this week and you're thinking, OK, where to start with all of this. Head over to castaclockmarketing.com to get an AI search visibility audit. I'll show you exactly where you're shown up, where you're invisible, what to do about it, why are competitors are winning. I hate saying that, but why they're winning and how you could beat them. All right, that's it for this week. I will see you in the next episode. Till then, stay visible.
Podcast Summary
Key Points:
Many marketers jump to tools without understanding AI search fundamentals, so a free GEO course is being planned to cover foundations before tactics.
AI bot governance is a cross-functional issue involving marketing, IT, and legal, and it requires decisions on which bots can access a site and how their data is used.
Most brands cannot answer three key questions
A bad governance plan is binary (block all or none), while a good one uses data to distinguish between training bots and live retrieval bots, with user testing for AI agents.
The llms.txt file is not necessary for 95% of brands; Google’s new audits focus on agent functionality, not search rankings, so other priorities come first.
Summary:
This episode of "Found an AI" with Cassie Clark and AJ Gurgich focuses on AI bot governance, a topic often overlooked. AJ, a Botify consultant, explains that governance is not a simple binary decision to block or allow bots but a set of decisions involving marketing, IT, and legal. These teams must collaborate because marketing cares about visibility, IT about server load and costs, and legal about liability from inaccurate AI-generated content.
AJ emphasizes starting with three critical questions: knowing what data trains AI models, ensuring its accuracy and currency, and having a plan for AI retrieval in training and real-time scenarios. He notes that many large brands cannot answer these yet. , stock checks).
In contrast, a bad plan is binary and lacks data. txt file for most brands, as it offers little value compared to other tasks. Cassie adds a nuance that Google’s new audits on agent functionality do not change this, as they focus on functionality, not discovery.
The episode concludes with a call for cross-functional collaboration and starting with AJ’s three questions to address AI bot governance effectively.
FAQs
AI bot governance is a set of decisions about which bots can access your site, what they can crawl, and whether the content they grab is used for training or live retrieval. It involves marketing, IT, and legal teams working together.
Brands should know what data of theirs is used to train AI models, whether that data is accurate and updated, and have a plan for what AI can retrieve for training and in real time.
For 95% of brands, an llms.txt file is not needed. It’s not widely used, and there are 10 more important things to focus on first, like improving data feeds and site accessibility.
A bad plan is binary, like simply blocking or allowing all AI bots. A good plan uses data to distinguish good bots from bad, and training bots from live retrieval bots, and includes user testing for AI agents.
First, find the source of the incorrect information, which could be from Reddit, forums, or your own site. Then, correct the record through PR and push accurate data via data feeds to AI agents.
Marketing, brand, IT, and legal teams should be at the table initially, as no single team owns it. The group can be expanded as needed.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.