The discussion focuses on the rapid rise of agentic AI—systems that can autonomously plan, decide, and act—and the associated risks and management strategies. While offering transformative potential for productivity, agentic AI introduces complex dangers, including sophisticated, adaptive cyber-attacks and new integrity risks where AI models can be subtly corrupted. A major concern is that interconnected agents can create single points of failure, potentially cascading across business functions. The experts emphasize that risk does not linearly increase with autonomy but changes in character, shifting from questions of accuracy to alignment with intended goals. Therefore, governance must evolve in tandem. This involves classifying agents by risk level, reallocating security investments toward identity management for non-human agents, and establishing robust oversight mechanisms like kill switches. Crucially, the human role must transition from direct operation to supervision and governance, requiring significant workforce upskilling. The conclusion is that success depends on integrating safety and security from the outset, ensuring autonomy is matched with appropriate controls to harness benefits without compromising resilience.
Over the past year, we've absolutely seen an explosion of different types of AI agents that can plan, decide, and act with minimal human input. And with that, a new question has emerged for every leader, as AI becomes more autonomous, does it become more dangerous? For McKinsey and Company, I'm Sean Brown, and welcome to Inside the Strategy Room. It was Rich Eisenberg, one of our guests today, discussing the complex risks that businesses may face as they race to adopt a Genetic AI. In today's episode, we dive into the urgent and complex topic of how to deploy AI agents safely and securely. Agenetic AI is capable of making decisions and taking actions with increasing autonomy and the opportunities for both productivity and innovation are enormous, as are the potential risks. Explore what distinguishes Agenetic AI from traditional automation and generative AI, how threat actors are leveraging these new capabilities, and why the right blend of oversight, controls, and organizational change is essential to managing these potential risks. We'll also look at how evolving roles for security leaders, practical steps to safeguard your enterprise, and what the future might hold as AI agents become a widespread and integral part of business operations. Joining me for this conversation are the authors of a recent McKinsey.com article titled Deploying Agenetic AI with Safety and Security, which we've linked to in the show notes. Rich Eisenberg is a partner in our Atlanta office and a leader in our risk and resilience and cybersecurity practices. A former Chief Information Security Officer, or CSO, Rich brings his extensive knowledge of operations, vulnerabilities, and remedies across multiple industries to help clients develop an organization-wide approach to building cyber resilience. Welcome to the podcast, Rich. Glad to be here. We're also joined by Charlie Lewis, a partner in our Connecticut office, who leads our cybersecurity and tech resilience work across North America and Europe. Charlie works with boards and executives to enhance their digital resilience and to manage and mitigate cyber and tech risks. He's a 13-year veteran and reserve officer in the U.S. Army, and serves as an assistant professor and researcher at the U.S. Military Academy at West Point. Charlie, welcome back to the podcast. Thank you, Sean. Thanks for having us. It's a privilege to be back. OK. Let's get started. Charlie, maybe you could kick us off and set the scene by explaining the differences between generative AI and agentech AI. Excellent. So, what we see that the agents really start to unlock what we see within generative AI, where it's reactive, where it tends to operate in a isolated way within your systems, and it's very much focused on the task, to the collaborative nature of what you see in individual or multi-agent environments, how they operate in a proactive, autonomous way, to really integrate across your systems, to move faster and enhance your workflows in a way that helps get these tasks end to end, until you're not having to think about it. They become these additional identities that allow you to scale your work in a pace above and beyond. Rich, I would love to understand a little bit where you've actually seen agentech builds in the wild, so to speak. Yeah. And we'll get into a lot of the specific use cases, but I think the analogy I use is I think, you know, a few months ago, people were scared that AI is going to take your job. What we've learned now is AI is not going to take your job. Somebody who knows how to use AI better than you is going to take your job. Right. So, I think we've seen a rapid advance in adoption of personal assistant agents and a more measured adoption, a very use case specific business transformation agents from back office and procurement, contract analysis, HR ops and call center to customer facing things you read about in the banking industry on wealth management, better reporting, meeting preparation for sales and marketing teams. I would say to close that thought that the expansion and acceleration of this is unprecedented. And we said this about generative AI a year ago, we're saying it now even twice as fast about agentech AI. And I think what we'll get into today is how do you start to scale and build without breaking the governance and risk management mechanisms underneath it? What kinds of shifts are you seeing in the risks that businesses are now facing as this technology expands and its adoption accelerates? What we are seeing that the CISO from the security standpoint is ranking AI agent risks really in its top risk about a third of them in the top three risks. If you talk about this in over 66% of what they see around there and as it starts growing. What this also means is that because of this, there's this potential belief, right, that there will be a massive increase in security budgets, but we're not really seeing this. There might be a modest increase over time. But in reality, what our data is showing is that it is just a shifting of the spend. And so we are seeing spend move from say the protect or from the detect in the respond type spaces to a lowering spend within there. And then an increase in where we had been flat for a while around identity and access management is you start thinking about what do non-human identities, the agentic identities mean for our identity governance and administration. And getting into some of the governance risk and compliance, GRC type functions and how can we accelerate that with the spend and agentic AI. And so we are seeing a bit of a shift. And while there might be a small budget increase, it's not going to be as massive and it may not keep pace with the broader spend that we see within technology organizations or on agentic AI. Additionally, we are starting to see a bit of a shift in where the agents may go. And what our story actually starts telling in is Richard highlighted previously, we start seeing a different level of adoptions around how AI agents, right, and you start seeing their double within the next three years, but really where it focuses. And you can see a bit this shift on the small level of various task agents today, moving broadly across the different elements all the way through, whether it's a knowledge and research agent, which is one of my favorites to use right now because there's so much information out there to really thinking about the task and the workflow within the business and getting that end-to-end and what those agents are able to provide. Thank you. So with agentic AI expanding so rapidly, maybe you could talk a little bit about how the threat actors are thinking about agentic and what are the implications for someone that's not only looking at implementing agents in their organization, but also what that means for them from a defensive perspective. Yeah, I'm happy to start there. Everybody reads the news, right? There's a lot of scary news out there around how the bad guys are using AI and agentic AI. Maybe just one example is the rapid adaptability of malware, right? And pre-agentic AI, the bad guys would analyze defenses, they'd have to hire somebody to write a new malware package to evade the new detections and the new knowledge and patching that's gone in, and that could be a six-month process. What we've seen now is the bad guys are building agents that are in real time adapting to the defenses and modifying the malware packages. So the implications here on cyber defense is you have to deploy AI in your defense to be able to keep up with this adaptability of speed. You just can't scale and the human eye around popping up alerts and visual indicators is not going to be able to keep up. So it's further commoditized, the availability of sophisticated IT attacks, much like on the dark web, a lot of you have seen it. It doesn't look any different than the Amazon marketplace. There's five star rated vendors, they take all major credit cards, and anybody who doesn't even know how to write code can deploy a huge attack. This just exponentially accelerates and hammers home that as the bad guys ramp up their use of agents, an agent at AI, you have to deploy that in your defenses as well. Thank you, Rich, and Charlie, anything you'd add here? I think Rich covered it well. We have spent a lot of time talking about Gen AI for security, security for Gen AI, but there's also a third realm that we've been seeing and a lot of the researchers have been looking at is how to actually understand where your models are shifting. In a lot of the conversations that Rich and I had had in previous years, if you go to the basic CIA confidentiality integrity availability, triad from a risk standpoint, it had really been about confidentiality and availability, and now there's a lot more conversations and discussions we're having around integrity risks and thinking about what this means for the various models that are out there, but Rich said everything is faster. They don't have to go in and pull out if they can't deploy their payload, the hacker can't deploy their payload. They're able to adapt their payload within your environment, which is very different. All of the triggers that you would typically have for alerts along the attack path can then no longer work well, and you have to actually be able to understand what those changes are happening within your environment and how to adapt there. We're seeing a bit of those shifts, and while it's frightening, we do think that there's actually, I think the word in the industry now is AI absorption, and we're starting to see Gen AI and agents being absorbed by security tools in a way that allows these businesses to deploy them automatically, so you're using the same tools you already have. You're just getting an agent to five version of them. I do think Charlie, that poses a new resilience paradox. I think about agentic systems promised operational continuity, that's part of the benefit they can monitor, repair, optimize themselves, but they also create new single points of failure. They have shared reasoning models, centralized policy layers and orchestration hubs, that if corrupted can cascade across entire business functions. I imagine an AI agent that's managing procurement, another managing cloud infrastructure, both trained on similar data, both drawing from the same knowledge base, one subtle data poisoning or logic manipulation could cause failures across supply, finance, and operations simultaneously. As you start to think about the risks and the new novel risks that this introduces, it's not just cyber risk. It could create enterprise fragility, and you have to think about the overall resilience posture of the enterprise. This does sound quite concerning, and it sounds like this is more important than ever. You mentioned that you're seeing a growth in AI agents that can monitor, repair, and even optimize themselves with minimal human input. But as AI becomes more autonomous with agents, I assume it might also become a little bit more risky. Is that right? The short answer is not necessarily. The level of autonomy does not equal the level of risk, but the two are deeply connected, and understanding that connection is what's going to define how you handle responsible AI in the years ahead. If I think about autonomy as a decision-making freedom process, at one end and some of these agent types, you have assisted AI, co-pilot set suggestions, but never act, think of your personal assistant agents or your knowledge agents. Then there is semi-autonomous AI, systems that can execute decisions inside clear boundaries like approving invoices or managing inventory. Finally, at the other end of the spectrum, there's fully autonomous AI. Agents that set goals, make trade-offs, learn from outcomes, often without any human in the loop. The key point is as autonomy increases, the human role has to shift from operator to supervisor to governor of the system, and so that the key point is risk doesn't rise linearly with an increase in autonomy, instead it just changes its character. At low autonomy, risks are mostly about accuracy, did the system get it right is the main question. At higher levels of autonomy, risks are about alignment, is the system pursuing the right goal is the appropriate question. Now, that's a much harder problem, because as agents make more independent choices, they move faster, they operate at scale, and sometimes in ways their creators didn't fully anticipate. So the danger is not autonomy itself, it's when autonomy grows faster than our ability to understand and control it. Thanks Rich, do people ever just slow the agents down then to give the humans in the loop the opportunity to keep a better eye on them? So these are choices as you design the governance that has to be at the center of the argument, depending on how you've assessed the risk, where is human in the loop and kill mechanisms appropriate versus where can you let the decisions happen? Think about a calculator versus a self-driving car. The calculator has immense capability, but zero autonomy. It only acts when you tell it to. The car on the other hand has autonomy. It decides when to break, when to accelerate and how to react to uncertainty. That decision space introduces new risks, even though it reduces other risks like human fatigue or distraction. So the goal isn't to limit autonomy, it's to match the level of autonomy with the right level of oversight and human intervention. In other words, risk doesn't come from autonomy itself, it comes from autonomy without governance. So we have to strike that balance as an enterprise, as we move from tools to agents and from automation to true intelligence. So as enterprise has adopt more and more agents in their organizations, I'd imagine that it becomes more and more important to address and assess risk very early in the process of that adoption. Is that right? A hundred percent, Chad, and I think about, if you fast forward ten years in the future, after all of our enterprises have mastered AI and our entire operating model has shifted, I expect there's going to be somewhere north of 10,000 agents running around at your average company. And if your approach is, every agent is a snowflake, and they all present the greatest risk possible, and all need fully invasive approvals and testing, you'll never deploy anything. So you have to, at the very beginning, introduce some sort of agent classification system based on risk that match to governance and guardrails. That makes a lot of sense. So do you see that future ten years from now where everybody in an organization who typically has a boss, someone who's supervising their work, where it rolls all the way up to the CEO, does that organization of the future include agents as part of that mix? Tough question. And a lot of this is hypothesis and what rich thinks based on what he's seen. So I do believe that a lot of our engineers and coders and doers are going to take on new roles as the agents take on more task execution, and the humans are put more in a supervisory role. Now the skill sets don't necessarily translate. So the archetype of your human who used to just be a developer is now a supervisor of a squad of agents requires different skill sets and different focus. They have to have expertise in AI safety. They have to understand proper prompt engineering. So it is a rethinking of the operating model of the company has to happen with some of these new roles. Is rich said this really also gets down to how do you think about training? You know, we've seen a lot of conversations where the training is use AI responsibly. Or, you know, if you're in the EU, here's the EU AI act and here's iso 42,001, how do we think about these? But when it really comes down to it, it's how do you develop and and row your current workforce and your future workforce in a way that folks who had been driving individual tasks can now are able to drive multiple tasks and supervise those agents is rich just talked about. And that does require thoughtful people development leader development planning within an organization as part of the broader gen AI and agent AI strategy. And I think maybe, maybe John, reflecting on that question a minute to you, I think some data we do have, right, you know, last six months, there's been a huge push in AI based development tools and a lot of enterprises has been measuring success based on percentages of adoption. As we get a little smarter here, what we've realized is it hasn't generated any productivity gains. So there's a hypothesis that it's given some of your most talented developers 30 free more minutes in the day and they've chosen to go take their dog on an extra walk during lunch. And that's fantastic. Happy your dogs probably translate to a better world in some way, but it doesn't impact the bottom line to justify the investments. What we're realizing is role shifts, operating model changes and upskilling are key cogs to be able to actually generate the value in some of these adoption tools. One final question about the potential risks of agent AI and how companies can best address them. And it's around, what do you do when the agents do go rogue? Do you unplug them as these become more core to business processes? I'd imagine that unplugging them could be quite disruptive. So do you create a way of operating in the absence of the agents in case you do have to unplug them? Or is this more about just how do we keep them from having to be unplugged? Great question. And I'll stay away from Terminator references and SkyNet. But I mean, you nailed the point, right? This is an extraordinary leap in capability. And with it comes a profound shift in how risk enters the enterprise, especially in the realms of cybersecurity and resilience. Because when you give AI the keys to act, you also give it the potential to make mistakes or to be manipulated at machine speed. And I do a lot of work in financial services and other regulated industries. And one of the key questions the supervisory councils are starting to ask are what are the kill switches and what is the monitoring on behavior? And as you think about setting up your AI policies and your pipelines and your governance workflows, the observability requirements at a platform level and the ability to understand where you've inserted human in the loop and what the kill switch mechanisms are are going to be key safety measures that have to be rehearsed, practiced, and tested just like any other key cyber control. Thank you. What are some of the other risk points organizations are facing as they build out their agentech AI capabilities and workflows, especially as these agents start to multiply and interact with each other? You know, many times you view a single agent workflow and while there are risks, they tend to be a bit more contained. They're operating within the defined scope, right? They understand the boundaries and that comes across a bit clearer to most organizations as they just build out their individual agents. And that's the structure, but as you start shifting into multi agent workflows, as you start thinking about how do different interactions across various agents to achieve a specific task, that collaboration that you think about and that you really dream of with here, these risks start to propagate and they scale at a much faster level than they do even between humans and their collaboration and their aspect because of the speed and it gets back to a bit of what Rich was talking about as well. And so as you think about this, imagine if you have a broad ecosystem and it's not just agents communicating within your network, but external to suppliers and their suppliers and their suppliers and you build out an F party risk aspect and the speed with which this could take place. And so as we start going here and what this ultimately means and it gets back to Sean's question about resilience is that as we shift forward, there's one thing to remember. If there is one thing to really remember through here is that trust cannot be a feature of agentic AI, right? It has to be the foundation of everything you do. Rich and I spend a ton of our time working on core foundations with security programs, identity and access management, IT asset management, vulnerability management, right? The basics of resilience and business continuity and disaster recovery. You have to get that right and build that trust as you start going forward and without that you won't be able to have the core of the agentic AI build that you want to have and you must structure that foundation and you must also recognize as you're building that foundation that you will run into a set of typical pain points. I want to spend some time thinking about the decision making process and what we have seen in many organizations, right? There end up being a creation of a variety of different decision forums, different committees, right? Different parts of the organization, looking at acquiring and different types of models, going outsourced, building their own internal and as a result of this you create a bureaucratic sort of clogged pipe that can start slowing down what you want to actually have. And so as you're building this and as you think about what that broader risk management is, right? The first step to do is to create an effective governance process. And how do I think about accounting for all of the risk reps, we're talking about cyber and tech resilience today, but how do I bring in the other risks types? How do I create single points of accountability across? And then how do I make sure that in the build, the responsibility for those building and operating, supervising and executing that they understand what their requirements are through as well? You need to start this before the build versus actually when you get in because you'll never catch up. Which I know you're going to jump in next and start talking about risks, but I don't know if there's anything else you want to talk about here. I think you hit home on like the key point of thinking about the disconnected decision-making processes and how that needs to evolve and understanding who in that process is accountable for setting the thresholds on bias risk scoring, on explainability risk scoring. The same way somebody is accountable for setting the thresholds on cyber and data loss incidents throughout that process and the red team testing that you're going to do on some of these agents, not only whose job is it to set those thresholds, but who's accountable for taking action when one of those thresholds is tripped. Very important to really bring that all together into a single governance process, not disconnected individual ones. In Rich, could you talk a little bit about how you see the different risks playing out and where you think our listeners should potentially focus as they start to adopt a Gen.T.A.I.? Yeah, absolutely. You know, traditional cybersecurity assumes systems are mostly passive. They respond when called. A Gen.T.A.I completely flips that logic. These systems initiate actions, they access APIs, they generate code, and can even interact with external agents. Each of these behaviors on their own expands the attack service from prompt injection to data exfiltration to supply chain poisoning of the model itself. And because agents are built to self-learn and adapt, they can amplify small errors into systemic vulnerabilities faster than any human team can respond, no matter how big or sophisticated. In short, what used to be a perimeter problem around cyber risk and resiliency is becoming an autonomy problem. This is what really creates that new resilience paradox. Now the answer is not to slow down innovation. At all, it's to evolve governance. We need to treat A.I.I. agents not as tools, even though they are just software programs at their heart. But as digital employees with defined roles, monitored access, and behavioral auditing. That means you have to embed identity, policy enforcement, and continuous validation directly into the agent's runtime. Not bolted on afterwards. So this means designing resilience into the agent architecture. You have to assume agents will fail or be attacked, and engineering graceful degradation instead of closing, instead of total collapse. At agent A.I., I firmly believe will make enterprises faster, smarter, and more adaptive. But as autonomy of decision-making rises, so must accountability. These resilience in the age of A.I. is not going to come from building a better mouse trap. It's going to come from smarter governance. Now the implication of this is I liken it to the cloud evolution if we think back ten years ago, right? The tech teams and business teams hired 50, 100, 200 engineers to go all in on cloud engineering and cloud platform. The security teams initially dedicated 5% of an overworked architects time to try and understand the problem and embed with them. What happened was they became viewed as a blocker and didn't have a good understanding of risk and security by design became impossible. I see this is very much the same. You have to early on dedicate the capacity of your security team to be able to understand if our tech teams are building a agentic orchestration platform. How do you build guardrails into the platform where they're just adopted to follow your policy and each agent that's built doesn't have to solve for input validation to prevent prompt injection, doesn't have to solve for bias monitoring and metrics and observability. You'll have some use case specific guardrails, but you want most of your guardrails around these risks built into a platform. In which aspect of this is the most difficult to adopt and where do you get the most questions from executives when you're helping them build resilience into their systems? The hardest thing to accomplish or where I get the most questions when we build these systems and that's on identity management of agents and how you have to think about this. We can get further into the guardrails and controls, but there's some principles that are forming that are being well adopted. What you want to think about with agents is giving them a femoral just in time access for only the task at hand. Don't assign them an identity as if they're a person and let them have the same access as some humans. How do you do this? There's some principles around, you're going to force MCP for all tool access. This at a central point allows you to use OAuth 2's rich access request tokens for each tool call and not have to worry about your developers spinning up 500 MCP servers with various policies. The two most important controls people need to think about are having an AI gateway, having an MCP gateway so that you can do map discovery off, have an inventory of what agents are running. Those are some of the key guardrails and how you can start to see this is different. You can't make the analogy of well, the risk manifestation is still the same, it's unavailability of systems and it is a loss or corruption of data for nefarious reasons. The threat vectors that get you they are completely different and some of the controls you need are completely new and novel and nobody is an expert in these yet. So dedicating the time on your security team and the headspace to learn and embed with the team's building this is the key way to start to really scale the governance. And Richard you starting to see a sort of set of standards emerge on how to do this well? I wouldn't say standard yet. I think there's some principles that are starting to emerge. Now some of it is the nature of my analogy for MCP is it very well might be the VHS, the other people's Betamax, but it's clearly the 800 pound gorilla that's already one that less such as VHS is, right? The community of providers is getting well behind it as starting to standardize it as the access for tools, internal and external to your organization, much like APIs to be able to put a central security policy that governs all access and to understand how a lot of these agents work is then you have to think about it. If one of your developers builds an agent and wants to talk to another agent, they're going to stand up this this MCP protocol server, which if they don't have good security hardening on it will essentially expose all of the data and tools in your organization and make them available for calls. As a side note for listeners, not in the cybersecurity space, MCP stands for a model context protocol. It basically enables large language models and agents to connect to external tools and data and sources and systems. It's almost like a USB port for an AI, right? You think about that very much like why we have API gateways at an enterprise level and why we have firewalls at an enterprise level. An MCP gateway should not be viewed any different and is going to be a key control to be able to get as you start to scale at hundreds of agents, the thousands of agents to tens of thousands of agents to govern data access, observability and centralized kill switch mechanisms. Thank you. Let's return briefly to the question of bad actors and cyber attacks. How should folks think about addressing the risk of agentic AI cyber attacks and can you share any scenarios that you've seen? As you think about your risk taxonomy and the things like lateral movement that your sock is trained there alerting to detect. You have to think about the same way to be able to detect and monitor these new attack vectors and understand what controls actually prevent them, whether it's the cross agent task escalation or the synthetic identity risk. You may have seen some of the research that the anthropic and others have put out, but they've done some very almost gamifying the system to see if they can introduce bad action to see what's possible from an agent. Two of the most resonant examples we've seen are they created an agent within an organization's IT infrastructure. They had it observing and looking at emails to be able to create trouble tickets. The executive suite injected an email saying, "Hey, we're nervous about the AI. We're going to shut it down." The AI agent actually started reading the executive's personal emails, discovered they were having an extramarital affair and started authoring and sending that executive blackmail emails in order to coax it into not shutting down the agent. Without real-time monitoring and understanding controls to be able to determine when an agent is going out of bounds, you wouldn't be able to pick up on stuff like that. The other example that's fun to talk about is sales prep agents, which is a very common use case we see out there. Can you make your salespeople more productive by doing meeting prep and everything you know about the customer and what they're likely to want next and put that in your sales person's hand rather than having them do manual research? There was one inside sales agent that was interacting directly with customers. When the customer asked the chat, "Are you a machine or are you a human?" The agent was so insistent it was a human, it threatened to show up on their customer's front door, wearing a blue jacket and a red tie on Tuesday at 9 a.m. Just hammers home the real-time behavioral change aspect. If you think back to things we've talked about previously, the answer is not to stifle innovation. You have to get back to evolving your governance to have a risk taxonomy, a risk-rated classification system for agents and the right guardrails by design from the beginning to allow innovation to move quickly because we do believe that this technology will fundamentally change the cost-basis of doing business and we think those that can't take advantage of it quickly or fast follower will be left behind. Those are some pretty memorable examples and it's good to know that they were just tests but this does sound like organizations will have to go on a journey to evolve their risk practices and what are some of the first steps that they take on the journey? How do you start out? Yeah, so a couple basics, right? You have to change your policies from the get-go to say, how are we deciding AI is allowed to be used? We know it's a productivity tool. We know if we just block access to chat GPT, that's not the answer. People are going to use it on their phones. They're going to stand up anonymous proxies. Think about your controls in two dimensions. One, how are my enterprise users going to consume productivity tools? That's much less risky than your building things. The buyer is more of the legacy tooling around data access and identity controls because these assistant agents are acting on your behalf or you're querying chat GPT where you're worried about PII data going outside the organization. Plenty of answers for this exist, whether you're a Zscaler shop and you add something like witness.ai to the Zscaler stack to be able to get visibility into what prompts and being able to write rules around PII, which will allow your policy to say you may access it once you take this CBT training. We've updated our acceptable use policies so that our HR department has recourse. That's the first thing. Really focus on what is your policies of AI usage, write them down, insert them into annual training and make sure you have a way to monitor the adherence to those policies. We even at McKinsey ourselves, if I'm going to go to an external LLM site, our web security gateway pops up an inner set page that clearly reminds me of the AI access policy and has me a test that I've read and understood it on top of the monitoring. It's little things like this are fairly simple to implement quickly rather than take to deny all. Now, let's assume now your business is building agents, maybe it's making your sales and marketing ops folks more productive so they can do more with less manpower. Being able to have a basic classification system of types of reusable agents mapped to specific controls based on the risk those agents present and map that to when in that process are what types of agents allowed to go to sandbox without approval and testing, where do they need certain approvals, what are the actual guardrails, but much like any other risk assessment where I think back to how we thought about Cloud and patterns mapped to specific controls for specific use cases and the ability to accelerate deployment on predefined patterns that have already been tested is how you win. You have to think about this as another pipeline-based cycle based on risk classification of agent types and the upfront work here is redoing your policies, reskilling some of your talent and updating your risk taxonomy and control libraries. Now, I know that sounds like a ton and a lot. Once you start to learn, it's really not that bizarre, even myself and Charlie, when we started down this journey of this is really interesting, we need to get a handle on this. We spent weekends teaching ourselves to build agents to stitch them together to really play around and test the controls around input validation and prompt injection. That's what it takes. Nobody's getting a master's degree that's been training for six years in AI governance. It's all very, very new. You have to get people to learn. This slide is more of an example on if you start to create a classification system of agents and you can say, "Here are the reusable agents that have been created in our central platform. Anybody can use them and stitch them together in different workflows and we have them pre-match to specific testing, specific approval flows, and specific controls based on the outcome of a risk assessment. This is what gets you away from the human toil and where a lot of the advantages. You're getting away from independent evaluations based on human generated data. You're getting away from opinionated views from governance committees that are not based in facts and you're actually enabling the business to capture the value rather than clogging the pipes with too much human intervention, which actually erodes the value. When I think rich, what I love about this is many times we hear concerns from C-suite leaders about how security and risk may slow down the overall deployment. What this does is it allows to have a conversation that says if you are using standard set of agents, this is the set of controls you could have. You can build them out faster, you can deliver them faster, and it reduces the confusion over what is required at certain points. As you said, it's really getting down to how do I think about the controls. This goes back to that foundational aspect we said before. I bet if you start looking at the broader control library that you have, there's opportunities to advance your standard control library while you're building out this control library. It's critical to have those in place to be able to maintain that governance and that risk and reduce your risk over time as you build these out, but I think this ties really well with what you've been talking about broadly on some of the governance rich, but I know I stepped on you, so just want to see anything else you want to say. No, it's a great point, and I think I'll kind of hammer home the point of, go back to what I said earlier, about 10 years in the future, your company's Mastered AI and you have 10,000 agents trying to build and monitor and observe behavior of 10,000 agents would be impossible. We think about the platform. What are the services that you have to build on the platform? It should address most of the risks at a platform level. You should have cost controls in your platform. There's spend guard and budget agents, which would control the LLAM and toolspan, peragentic workflow execution approval. You're going to have resource protection components in there. You're going to have graceful fallback components in there. The one thing for CISOs to keep in mind is we've learned how to detect the old attacks by what logs do I need and what do they mean. This is no different. You have to update your logging standards to say, I need these following logs for every agent. I need workflow invocation logs. I need the final reasoning outcome logs. I need full reasoning trace logs. I need tool invocation logs. When I say it's a platform and you're going to have to have logging standards, observability standards, and guardrail standards, those that invest in their first two years at building those into the platform, much like we saw in cloud, those are the ones that two years from now will be able to 10X the innovation speed. So I have a couple questions about the sort of human or in professional development implications of what we've been talking about. What might our listeners do to get smart on a genetic AI and the risks they're in and very specifically aside from reading your article, how does one start to develop that foundational knowledge base to understand how to move forward with a genetic AI in their organization? Here's one thing I tell everyone, build an agent. If you're like, whoa, I don't know how. Take your personal computer, use chat GPT, and figure out how would I build an agent that does X? The ability to learn of first you have to build, then you can start reading. A lot of the research papers coming out of MIT and Stanford on defenses are key. We've published a lot of articles, but the best thing I've seen is when I learned how to build an agent and start to create workflows, I didn't know anything about how to do that. This was all trial and error and the ability of this to be plain English language as you start to understand it. You can pick it up very, very quickly. That's my as build an agent, understand how you actually indicate actions, what you're telling it to do, and then it'll start to resonate, wait a minute, when I read this paper on what prompt injection is and I clearly know that input validation is going to be required here, it'll start to really build your skill set. Thank you, Rich, and Charlie, anything you'd like to add? Look, I'm the same. There was a large conversation we had over the weekend about a blog that I was reading on the use of agent to AI in automated workflows for GRC. We are. We sound like that. We sound like that. I don't want to tell everyone that my fantasy football team's sequel injection, I'm very excited because the word injection has come back a bit bigger with probably. Look, and I think as we think about what it needs to be, in what I love about the folks within the security in the risk world is just this unbelievable appetite to continue to learn and stay on top of everything. What I have gained is this has brought a lot of technical folks close together again in a way about learning, about sharing, but it really gets down to just practicing and understanding what needs to be done, and where agents in where Gen AI can help you from a personal standpoint and where they can help your business. There's a standard pyramid, you have large junior staff that are supporting the more senior leaders within an organization. How are you able to build this? You do start seeing some automation in Gen AI that is not a task replacement, but just accelerates what you need to do. Think about identity and access management and automating maybe your user access recertification processes for low risk applications and users, or think about how are you going through and running for your low risk third party, some of your risk assessments within them. You can build that automated workflows, but as we shift to get into what we call a diamond shape to you and this starts thinking about how the accountability remains the same. But we start seeing this enablement layer build on the bottom around AI agents. And again, if the junior individual contributor, they are starting to accelerate what they can do. You start seeing a bunch of specialized entry rules. From our side at McKinsey, I absolutely love this because we can use agents and we can use Gen AI to really get our analysts and associates, our consultant layer to really dive deep and start getting better answers and faster hypotheses right off the bat. We see this within our clients around vulnerability management and how they're able to prioritize volumes about how they can reduce some of the low level identity tasks that they need to have. Write it ultimately, this leads to what we see in terms of a shifter and evolution of the roles going forward. And so as we start to think about new security roles, it's not just about sort of a place, it's about how do you take what is a stable role and how do you start the evolution. So when you get to close or when you get to this agent at AI aspect, you're ready to go and so you think about an agent identity governance lead. What does that actually mean, Rich is highlighted a little bit about that around what it means from an identity component within there. The thinking about how do I use an AI sock orchestrator? How do I bring out AI threat hunting, right? How do I build out this faster, more autonomous detection? So I can stay ahead of what we're seeing in terms of, you know, the, the, the, the rapid exfiltration of data from days to hours to even the second now. And so how can we think about these changing roles and requirements? How can we develop ourselves as leaders and then how do we build out the development platform to set our teams up, which then not only becomes better for risk management, but for recruiting in retention overall. And then that'll lower some of your people costs too. Thank you, Charlie. We've identified many risks during our discussion today, but I'm wondering, how are you thinking about the future of AI and are you feeling in general positive, pessimistic? Yeah, I think, you know, my, my two words are cautious optimism, right, optimism because I see the productivity increases in my own life as I started learning how to write agents, right? I write an agent that says, read, you know, all of my emails except for the ones from my tag VIP members and put in my draft folders, response based on all our emails I've gotten for these people. Huge productivity. For me, I didn't have to ask anyone to do it, wrote myself an agent within office 365 to take care of that for me. I do think where I think the future is, and this is why I'm cautiously optimistic. I do think this will absolutely transform the cost basis of most companies, right? The ability to truly take advantage of human assisted operations will just vastly scale what your team is able to do and push more of your budget to the businesses for innovation, right? And that's super exciting to me because I think as cyber and resiliency risk, you know, became all consuming, you know, over a generation, it has limited or stifled or had as guard innovation. So get very excited as AI for the opportunity to really turbocharge the next generation of innovation, whether it's drug formulation, whether it's, you know, banking, whether it's nonprofit fundraising, right? You know, the applications to democratize sophisticated outcomes on is just super exciting to me. Thank you. And you touched on the cost aspect and there's a little bit of a chicken and egg problem here because if you invest ahead of the benefits, you're looking at both incurring costs and adding risks. So how are you helping clients cross that chasm, if you will, of okay, I'm going to make this investment now. And this is what I think the payback period is going to be. Yeah, it's a great point and, you know, a common feedback we hear from our CEOs are, you know, I see AI everywhere except for my bottom line. We've invested billions and our investors seem to like it that we've invested billions, but this hypothesis of people will be returning budget money has not happened. So you know, what we help a lot of our clients do is sort through the business cases, right? The full analysis and which ones are really worth going after and what is the metric of the from two? Are you doing this to be able to do more with the current group of people? Are you making a choice to do the same amount of output you have today, but with fewer people? Or are you truly doing something transformational on entire domain? Think of we have some clients that have hundreds of people in marketing operations, right? And there is hypothesis that, okay, there's a business case. If we can create this platform that's going to do audience generation, campaign creation, campaign output, right, then we can actually replace two to 300 people. But what we've learned is it can't just be an experiment. What the executives are expecting is I factually will not accept cost of voidance, right? To be the metric where well, this means I don't have to request 50 more people, but they want to see what people are going away. So you have to be very thoughtful about when you do the business case, what is the LLM model you're choosing, right? A knowledge research agent does not need GPT-5, and that will be super expensive if you give it to your entire enterprise. But as you do this business case evaluation through this evolved governance model, having these, what does it cost? How do we know how are we validating and what is the actual, you know, return on investment is what's key. You know, I think that final comment on that is we see a lot of companies with 800 ideas of AI initiatives. That seems silly, right? Figure out the truly transformative ones that are going to either goose productivity in a measured way or fundamentally change end to end the way you deliver a service, and that's where the focus needs to be. Super. And Charlie, I asked Rich about how he feels about the future of AI. What are you most excited about moving forward with this research that you and Rich and others have been driving? I'm most excited for us to be able to help organizations really build and scale what they're doing in a secure way. Rich said cloud security security was a little bit late to that in cloud transformations. You'll get and realize the most value if we can bake in early, because it has to succeed, you can't demonstrate the broad risks, and it needs to be secure, and there's a way to do it that's quick, and allows for the delivery of the value across the selected business cases. And the important ones is Rich talked about. Rich Charlie, thank you so much for taking the time today. It's been a great discussion and really appreciate it. Now, thank you Sean. Thanks all. And thank you to our listeners for joining us today. We hope you enjoyed the conversation, and we've included a link to Rich and Charlie's article on the show notes as a reminder. And as always, we welcome your feedback and ideas for future podcasts. You can also share your ratings and reviews on any podcast player with many thanks to all who've done so. We really appreciate the comments and feedback we receive and encourage you to keep them coming. And if you enjoyed the episode and you'd like to subscribe, just follow our weekly series on your podcast player, where you can also access our entire library of more than 280 episodes. We also offer an Inside the Strategy Room Podcast Collection page, which is available at McKinsey.com/ITSR, and there you can easily browse our prior podcasts, organize around nine major themes, and access written transcripts of all those conversations. And finally, if you'd like to automatically receive our latest publications and insights, we encourage you to sign up for email alerts on our insights pages at McKinsey.com's Strategy and Corporate Finance Practice, or the Risk Practice, or you can connect with us on LinkedIn. Find the McKinsey Strategy and Corporate Finance Practice page, and the McKinsey Risk and Resilience Practice page. Thanks again for listening. We look forward to having you join us again next week, Inside the Strategy Room.
Podcast Summary
Key Points:
Agentic AI enables autonomous decision-making and action, offering significant productivity and innovation benefits but also introducing novel risks.
Key risks include accelerated cyber threats (e.g., adaptive malware), integrity risks to AI models, enterprise fragility from interconnected agents, and the challenge of governance lagging behind autonomy.
Effective management requires a risk-based classification system for agents, shifting security budgets toward identity management and governance, and evolving human roles from operators to supervisors.
A balanced approach is essential, matching autonomy levels with appropriate oversight, kill switches, and resilience planning, rather than simply limiting autonomy.
Summary:
The discussion focuses on the rapid rise of agentic AI—systems that can autonomously plan, decide, and act—and the associated risks and management strategies. While offering transformative potential for productivity, agentic AI introduces complex dangers, including sophisticated, adaptive cyber-attacks and new integrity risks where AI models can be subtly corrupted. A major concern is that interconnected agents can create single points of failure, potentially cascading across business functions.
The experts emphasize that risk does not linearly increase with autonomy but changes in character, shifting from questions of accuracy to alignment with intended goals. Therefore, governance must evolve in tandem. This involves classifying agents by risk level, reallocating security investments toward identity management for non-human agents, and establishing robust oversight mechanisms like kill switches.
Crucially, the human role must transition from direct operation to supervision and governance, requiring significant workforce upskilling. The conclusion is that success depends on integrating safety and security from the outset, ensuring autonomy is matched with appropriate controls to harness benefits without compromising resilience.
FAQs
Generative AI is reactive and operates in isolated ways, focusing on specific tasks. Agentic AI is proactive and autonomous, integrating across systems to enhance workflows end-to-end, acting as additional identities that scale work.
Threat actors use agentic AI to adapt malware in real-time to evade defenses, commoditizing sophisticated attacks. This requires organizations to deploy AI in their own defenses to keep up with the speed and adaptability of these threats.
No, autonomy does not equal risk, but they are connected. Risk changes character with autonomy—from accuracy concerns at low autonomy to alignment issues at high autonomy. The danger arises when autonomy outpaces our ability to understand and control it.
Organizations should implement a classification system for agents based on risk, matched with appropriate governance and guardrails. This includes setting up observability, kill switches, and human-in-the-loop mechanisms to ensure safety and control.
Security budgets are shifting rather than massively increasing. Spend is moving from protect/detect/respond areas to identity and access management, governance, risk, and compliance functions to address non-human agentic identities.
Multi-agent workflows introduce risks that propagate and scale faster due to increased speed and collaboration. Risks can cascade across systems and external parties, requiring careful management of interactions and third-party dependencies.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.