Go back

Building Trust Into Agentic SOC Tools with Oren Saban

43m 58s

Building Trust Into Agentic SOC Tools with Oren Saban

The transcription discusses the rise of agentic SOC platforms, which now handle the majority of security investigations, from triage to threat hunting and case management. Unlike traditional SOAR, these AI-driven tools leverage unstructured data (e.g., Confluence, Slack) and adapt to organizational context, but success depends on marrying built-in expertise with local policies and exceptions. A key challenge is getting AI to accurately assess its own confidence, as LLMs struggle to say “I don’t know.” The job of security analysts is shifting from manual investigation to oversight, requiring critical thinking, agent governance, and the ability to validate AI outputs. Future skills will emphasize understanding the underlying technology (e.g., how AI works) to spot problems, while less value will be placed on repetitive tasks like documentation. As attackers also adopt AI, analysts must deepen their technical expertise to stay ahead, making continuous learning and adaptability essential.

Transcription

8024 Words, 42596 Characters

English
Hey everyone, before we start the episode, I just want to give you one quick, exciting opportunity that's coming up soon. And that is the SANS Cloud Security Exchange 2026. That's going to be taking place in San Francisco on August 17th and 18th. And aside from a bunch of amazing speakers and talks, if you go to this event in person, you will get the chance to work through two agentic sock workshops. Something that goes very, very well with the conversation in this episode. One is the Cloud Security Agentic sock workshop put on by Google Cloud. And the other one is from Microsoft called your new tier one analyst isn't human, lessons from building an autonomous sock. These are obvious interest to nearly everyone listening to this episode, so I would encourage you to check it out. Link is in the podcast description. Alright, I'm with the episode. Walking the RSA floor this year, something filled different. Not just more AI booths. There has always been plenty of those last couple of years, but for the first time, the tools I was seeing didn't feel like a first crack. Today we're actually shipping full fledged functioning products. Agentic sock platforms like a triage, investigate, threat hunt, close cases, respond, real production and dirty stuff with real customers. And that's exciting, but it also raises a question I keep coming back to. If a platform is handling 80% or 90% of your investigations, what does a good analyst actually now need to know? The job description has changed and I'm not sure everyone has caught up to that yet or even as sure what that means. Well, today on the podcast, we have one guest that does have answers and that is Orn's Bond, someone who has been thinking hard about exactly these kind of questions as well. First at Microsoft, we're working on security co-pilot and now I'm building one of these platforms of his own. We talked about how these tools actually work under the hood, but contexts they need to be useful for your team and what skills are going to matter more and less as the technology continues to mature. So with that, let's get into it. Orren, thank you so much for joining me on the Blueprint podcast. I wanted to kick off this episode with a little bit of a backstory. We met at RSA, which was filled with a bunch of really exciting security operations related tools that had been released and capabilities that are now starting to form and we had had a conversation around the implications of what that meant for security analysts, security leadership and what it's going to do to the job and the tasks and the capabilities. So I would love if we could start the conversation with a intro to you and your background on the problems you're solving and how you're working in this area. Orren, Orren, Sabana and Chief Product Officer, here at MATE, MATE Security were working on the word of AI for security and specifically for SOC. Before that, in my prior role, I was working on security for AI startups, so that was one role, but beyond, before that I was working on Microsoft XDR and working very closely with the teams, the security teams, running red team, blue team trainings. For a lot of the SOC teams around the world, sitting with them in the room and see how to make their life better later on. I worked on security and its ability to help, again, efficiency, efficacy, a lower fatigue and the help analysts do their job better and excel. I'm very, very keen and excited about this word of how we can bring AI into security and in what ways the security word will change. Yeah, so let's start there. For listeners that may have not been tracking some of the newest capabilities closely, I know when I walked around RSA, it seemed like, maybe just because I'm focused on this, but half of the show floor seemed like it was agentic security operations, capabilities and tools and something that was related to that, which is super exciting. For those that may have not seen that kind of stuff and watched some of these brand new companies and capabilities, how would you describe the new wave of tools that is coming out and what they're capable of doing compared to what we've maybe had just in the previous year or two? So a great question. I think there is also a big mess today in between, because all of the tools that are doing a lot of stuff, but eventually it comes down to some important portions of first of triage and investigation cases, investigating cases that I would say the bread and butter of anything that happens here. And from that, there's many different types of use cases from they came, in these tools, so the automatically thread hunt, building hypotheses, run top of that, hunt on top of your data, write detection, improve your detections, and so help you with the unknown unknowns, write and find more things that you couldn't do before or do it faster and increase your speed. There is some newer agentic case management capabilities, agentic reports, ability to summarize cases faster, all the documentation work that everybody hates and many of that now can be managed with AI. And of course, some others, I would say that's the core of the new capabilities. So we've had automation capabilities in the past, right? Sore platforms were big. We had initial takes on using AI with a little bit of capability built in and such in the tools, but what has the agentic kind of revolution in this space unlocked compared to what Sore has been able to do or what the more basic versions of even LLM's and AI was able to do just a short time ago? So I think if we look at a regular SOC hue, not all the hue is the true positives, right? It's combined of a little bit of true positive, maybe after 10%. And then a lot of cases which are B9 positives. So the detection was supposedly fine as just in this case for us, this is okay. Now to handle these types with automation becomes very hard because especially in enterprise, you have a lot of exceptions and tribal knowledge and things that are okay only in this case and the detection tooling are not always tuned to what our environment or don't even support such layer of tuning and then to build automation that can handle this human sense of something here looks fine just because I see the disuser have done this before and I know him and I spoke with him about it. Yeah, he's doing that and we say we'll create an exception, but we didn't do it yet. So there are so many hard problems to solve with automation that AI makes it way easier to do. And then supposedly if your SOC is configured very well, you have tens of people configuring that daily and you can achieve a state of very high level of automation, but for most of the teams and they never reach this state because it's so hard to build configure and change automation. And even on top of that, there are some layers which are very help if you don't have the data organized in the right way and it's unstructured. So to build SOAR that will utilize my confluence is something that is really hard to nearly impossible. And now AI opened us another door of utilizing a lot of unstructured knowledge that we've had in many different places and leveraged that to automate. If someone is picking up one of these tools and has all of this unstructured data, how do these platforms approach making sense of all that data? Is there a certain type of format that it works better with? Let me rephrase, if I'm a team buying one of these, do I need to convert all my documents? Do I need to consider how I'm going to be feeding that data into the system to get a good versus a bad result? So the first question that comes is even not just do I need to do that is, well, well, my confluence has documents from five years that no one ever touched in the last five years. And where are you going to do with how do you going to handle the garbage in? Because we don't always maintain our documentation or docs. That's a common problem that, still when I was at Microsoft, even with the internal Microsoft Copilot, that was a problem. You can serve for something and that's something that's just skewed. It's not true anymore. And the capability of the different tools is to take this data in, structure it in a way that you can trust it, and then leverage it. Because you don't want just to blindly leverage anything that you have in your confluence, or in your Slack, or in your Teams channels, in Teams groups, etc. You want to be mindful about what you leverage and how you do it. Does that just back up the problem to a data management problem or is there an easy way to do that? So the way that I see it, and some solutions offer that, some don't. They also offer the capability to sync the data, to manage the data in a way that you can trust. And you have different confidence level of this data. And definitely, again, as always, it all comes down to the data. So the context is a very common name now. We have our security context graph. It's exactly the capability to connect between different data sources and data types. But eventually, if the data is problematic, you can have as much context as you'd like. You might be wrong. If I was trying to predict my likelihood of success bringing one of these kinds of platforms in, is there any way to measure how much data, context, and other kind of surrounding information one of these tools is going to need in order to produce a good result? So the nice thing is that, and again, it depends on different tooling, different mechanisms to do so. But the nice thing is that you don't necessarily need to have a lot of these predefined, because some of it, again, just might, might to sense, we derived that from the work that analysts are doing. We derive insights on policies, made derives, its ability to understand how we do this in this exact situation, in this customer. And basically, a lot of the ISOCOS on bringing a lot of the expertise out of the box. So you don't need to teach, not mate, and all its competitors, how to do a fishing investigation, it will come with this knowledge. And then the interesting part, and that's where I think, if I'm putting the stuff in the customer's shoes, you need to lean on, is how do those marriage? So, great, there's a very smart agent that knows how to do a very complex cloud investigations. And then there is my own knowledge of how do I do it in my organization. And these two need to marriage and to build an intent-based type of automation, where the platform takes my knowledge, and because it's important. And the same thing, like you can't bring in a principal security third hunter. Put it in your sock tomorrow, and let him or her run on top of investigation. They won't do it right. They don't know the policies, they don't know our exceptions, they don't know our compliance and regulations, and they don't even know our architecture, maybe. So, just the same thing, you can't just take a very smart AI agent and let it investigate. It needs to know yourself, and it needs to have the capability to learn enough out of what you already have. That's one thing that I heard from literally everyone I've talked to in this space is context as key. If it doesn't have context, you're not going to get good results. If you have great context, you're probably going to get pretty amazing results. And depending on exactly how the work is broken up, there's a little bit of a dependency on that as well. So, you would mention having a smart agent. Architecturally, it seems like I've noticed that the move is to break up what needs to happen in any given investigation into a bunch of separate agents that have specialties focusing on a certain piece of that investigation, whether it's tearing apart an email, whether it's considering all the options for a hypothesis generation and evaluation and stuff like that. Why is that the approach that's taken as opposed to one agent that's maybe a little broader and gets the whole picture of what's going on? It goes down to two, I would say three things in three different factors. The first one is a list agency approach. It gives benefits both from how easy it is for the model to do the task. So, let's take an example. If I've given a model a task that a GPT 3.5 can run successfully, opus 4.7 would excel. This is something that we want because we are taking a non-deterministic intelligence like models. Putting it in cases where we expect that there is a terminism, we have double standards for what we want out of these tools. We wanted to behave in many cases like it's an automation when it's not. Then, the list agency and scoping the agents to do a very specific task without the need to hold up a lot of different contexts in its way because we know that even when there is a very large context window, all the risks to talk about it, it still creates an answer for a reduced quality of what the agents will provide. So, that's one thing. The scope it for a better quality. The second thing will be to scope it for a better security. You don't necessarily want to bring knowledge from, let's take an example around DLP. Maybe I want some agent to inspect the file that supposedly has in that leakage. But DLP is a very noisy type of alert. It creates a lot of false positives. So, definitely I want that. On the other hand, I have a problem of permissions. I don't want to open all files in the organization to all agents to utilize that. And then I need to bridge between what's handling sensitive data and exactly like we do with people, right? We don't give all permissions to everyone. We scope it to a specific time even. Maybe it's a just in time approach that I can give the agents to go to sensitive data. And eventually the last part, yeah, we do a lot of scaffolding that maybe over time might disappear as the models get smarter and maybe we'll come with it their own. A level of context management, but I think we're not there yet. With breaking it up like that and having to fight with scope and security and scaffolding and all of those things, what do you find is the most difficult part about producing a product that gets this right most of the time? All LLMs are not very good at saying I don't know. And this is something that was hard to, I'm calling it to teach you're not really teaching anything, but to tune the agents, our agents to the level of when they're not confident. And there is a lot of mechanisms that we've built in order to get their calibration of the confidence. It's not, if you ask LLMs to give it a confidence, it will do some calculation and reasoning, but it's not necessarily right. And it's not something that those tools are good at, at least not at the moment. And then you need another layer that will know how to leverage and give the context to be able to even derive what confidence is. Now this, yeah, I would say this is the hardest part to be able to get it to the point where it would be able to say I don't know what it doesn't have the data to be able to not be as sure in its answer. And the surface when there is a problems. And so, like we asked our employees to raise a flag when there is a problem by default, LLMs are not as good in that. And that's something that we worked a lot to achieve. That's probably a good segue into what I think is part of the most interesting things that we discussed at RSA is the more kind of philosophical implications of having tools and the job and career implications of having these kinds of tools for people who want to work in a sock and want to continue to work in a sock. When you have a tool that is now largely doing a lot of the work for you and then saying, here you go. Here's the answer. As you just said, right, LLMs don't often like to say like, oh, I'm not really sure about this. So what kind of new skills, what kind of new mindset does an analyst with one of these tools have to bring to the job to make sure that they're not enabling mistakes? Really good question and discussion we have with a lot of partners now. I would say not just analysts, but anyone in this world now that is working with AI. Of course, critical thinking should be the first one that I bring to my job. How are you so sure? How do I build then in mechanisms where I'm able to poke and check? So that's for a singular case, right? Am I able to look into a single case and say, and I need to say quite rapidly, does it look okay or maybe there is something off here? And this critical thinking now to ask the rest question and for the product, the ability to surface that in a way that it is for me to consume. But that's on a singular case. And now when we look on a swarm of agents working on top of all of my cases, my ability to move from a singular case validation into agents governance. And here I see very much like we call it instant promotion, right? Very much like I have a lot of employees. I need to look into those. I need to do QA of the word. Sometimes I micromanage and sometimes I know to back off where I see it's working well. And it's really behaving in the same way. And again, I think that's the role of the product to simplify for you to be able to do it in a very easy way to validate that it's working fine to know on the cases that are controversial and you need to look into for those agents to be able to raise the flag for you and say, there's a problem here. So this is the role that we're seeing that is changing in its ability itself. And then of course, when you need to hook in and utilizing that, the other thing is to be able for you yourself to know how to leverage AI to its maximum. And we really see that in software engineering. We really see the difference between engineers that know and build their system and kind of sharpen the saw. Like Stefan Koffee says, I really like this. Yeah, you need to sharpen your saw with AI. It's not always coming out of the box and just works. I do think prompting, prompt engineering is something that is not going to last, but there is going to be always the new thing of these power users who know how to utilize the ads to its maximum. And these are the one the people that I want in my soul. There's definitely a whole realm of skills related to this and not only are they new skills, but those skills themselves are continuously changing at lightning speed. What was useful this week might be completely superseded by something that's invented next week. To the best of your kind of ability to predict this kind of thing, if we look maybe a year to three years into the future, right. I write a course for sock analysts right to teach them what they need to know what they don't need to know in your opinion what skills are going to become more valuable and what skills might become less valuable in the next one to three years. So I do think you will still need to be able to look deeper. into what the AI has done and deeply understand it. Him, I can't remember who said that to be good engineer, you need to understand two layers below what you're doing. So if I'm writing C, I need to understand assembly and how it works below that. And if I'm working, we all leveled up now, to work at the management layer. So I can see results. I do need to understand still how this is working and how those are utilizing my tools. If not, I don't think you'll be able to spot problems. And then I think it will just leave you useless in case of actual incident. So we don't want to be there. That's one. As in, I don't really think you can ditch all the technical part of being a security analyst and understanding our platforms. And attackers are not giving us easy lives. And yeah, when we think about it, it's great AI is coming for helping the security teams. It's also coming to help the attacker side and with the capabilities such as Mitos or Mitos, they're not even putting into if this one or another. But we all understand there is going to be a very strong model of the other side. I mean, we are doomed if we are not doing a work faster, better, and with higher quality. So one thing is to be able to be technical and really understand how things work. The other is to leverage AI to its maximum. I know it's changing, but I think the learning muscle is the one that I would expect for each of our employees from myself, first of all. And then, yeah, I think to be able to learn and adapt quickly to new technologies, to new AI systems is second. And the third one is the critical thinking, the ability to lay down different hypothesis and challenge with the, like, challenge what you're seeing and ask the rest question and write question and enable in order to see if there is actually something in here or no. Is there any skill that you think is going to be the one that is like no longer needed and just kind of goes away now because not only is it automated, but we just don't even need to know the technicals behind it. Or do you still think we pretty much need to know like the basis of how to do things the hard way? So a few years back, right, we've had, and maybe in some organizations, still we've had a team doing endpoint, another team doing phishing, another team doing a cloud. It's now converging because it gives me, as an analyst, the ability to handle more cases, I can work with the chat a better understand them. So it did expand my coverage. I still think that you need to be very technical with your capability in the domain that you need to handle eventually. Yeah, there are some skills that I think they're going, not going to last. One of them is to be a master at query languages. I guess model is going to keep up on that. And at some point, we don't need to know much more than English in order to get the data. We need to have the standard data, but maybe not to get it. Like to query that in all different languages. We need to think more on that. That's a tough one, right? It's one of those things where, you know, I've been thinking about this now for, I guess, probably years. It's like every time I think, "Ah, AI is going to do it." I think, "Ah, AI is going to make that thing easier to do." It still comes back to about how do I verify the AI data? Well, I still have to know how to do it, right? So the query language is, I think, is probably one of the best answers to that. It's like, yes, knowing how to write a query the hard way might not be as necessary in the future because you can, if you have the right language and the right prompt, right? Do a correct conversion from whatever search you want to run. And LOMs are great at translation, right? No matter what the language is doing from. So that's probably a really good one. And it bridges the gap from both sides. So both from the product that creates the query language and from the different LLMs and tooling. So eventually, I think this gap is just not going to last. There is a change that we see that maybe we don't. a change of how the soak is going to be built, right? And the smart teams, they helped their people to grow into an L2, L3. But yeah, I think the the tierless soak is not the dream. It's coming into a function. We sit with a lot of teams. There's set of people, smart people. They can handle a lot of different cases. And you don't need a lot of tiers in the way. You just want people that can close the loop and to end. And now they have the power to do that because a lot of the groundwork could be offloaded completely. Do you think that there's going to be an issue teaching people the fundamentals? And if so, how do we approach getting people to learn how to do it? The hard way when they grow up and they go to school and maybe any formal education they have, all assumes access to AI and AI-powered tools. Is that going to be a new challenge that we have to approach? Definitely. Not just a security, but in any education. And that's why I was saying critical thinking is a tool that will be needed anywhere. And then exams are going to change. So if I would have to write an exam for an analyst today, I would give it to them with the summary of the day. The agent has done the investigation and give them 10 investigations. Which one is the one which you got wrong? Good luck. And this is hard because you get 10 items which in all of them, the agency is 100% sure this is a false positive. Now find me the one that it missed. And then you'll figure out that through working with the chat, through working with him, asking the right question, again guiding it to the right data, etc. That you need to be able to apply this critical thinking and to understand what's coming in underneath in order to pass such an exam. And I think that's going to change all across education, not just in security, but anywhere. Yeah. The skills are similar, right? It's either doing the work from scratch or not knowing the work was done correctly. You still need to know how it was done, but you're starting with the end and a claim that it's correct versus here's the evidence you come up with the conclusion, right? In a world where 9 out of 10 or 99 out of 100 or even more of these investigations are correct, are there any hints that analysts should be looking for that like, ah, this is the needle in the haystack that might be the one that I need to dive a little deeper on because what I worry about is like people get so used to it being right and they just completely pass over it like, ah, it's right again. And then they completely miss the attack. We transform alert fatigue, right? We got so used to it get it wrong. To an agent results fatigue or whatever you name it, that we will get so used to it, get it right. Now, this is where I think there is a change and a shift in the work of what I call governance and mechanisms. My employees here have made their idoo mistakes from time to time. We have mechanisms in our company to catch that. And then, yeah, there is no 100%. But I think we all should be fully honest and aware. And there is one seesaw that he told me like, if you get 80% right, I think I think you're better than myself. And now I don't think it's the case for most of the stock things we work with. Or maybe we just chose the best ones, but they're really good at what they do. But yeah, we also make mistakes. And I think the, like any part in security, we build multi-layered security. We don't create one critical point of failure. And it's going to be the same thing. If there is not enough confidence, let's bring in a judge. It's at an QA layer, both from us and from an agent that is more expensive. For example, and we can't utilize that escape. There is lots of mechanisms of how to do governance, right? And I think same thing, like was done in manufacturing. Then he wanted to do a governance that nothing, no test life being created, completely broken. And we're going to have exactly the same thing for any AI output. Yeah. Yeah, it would be interesting to see how that problem is solved across different kind of industries. And team sizes and all of that. It's still kind of an open problem in some ways. And I think there's going to be obviously plenty of mistakes made along the way. And some best practice that emerges here in the near future. But it's interesting and fun to be kind of on the cutting edge of this. One thing that also comes up a lot in class is, you know, there's a couple of paths. Like once you have this capability, clearly we're increasing the bandwidth of things that your sock can do. Right? When you have most of the investigations mostly done for you, you could either take that capability and have a smaller team and say where we needed 10, now we need three. Or you could take that capability and say, well, look at how much we can still do now with 10 people. But ultimately the question I'm trying to get to here is, do you think more teams, and maybe from your experience, do you see more teams reducing headcount after they have these things? Or are they doing more now that they have this with the same headcount? So at this moment, we see today, I don't see them completely reducing headcount yet, because there is understaffing for a lot of things that we wanted to achieve in security. And then eventually, in many cases, it goes back to the leadership approach on, I would say the company in AI in general. If there is, and I haven't seen stuff reduction only in security team. If there is a complete lay of sometimes around the globe, do they take from security team? Yes, but I didn't see it necessarily going directly because we have a. for Sokna, we can reduce the team. No, I see they're doing amazing things, increasing the bandwidth of what they can handle and making the organization more secure. So if I don't need to waste my time anymore and 80% for positives, now I can actually put that into hunting threats, into building mechanisms of prevention that are stronger than before, into improving my threat modeling and how do I handle these cases? And of supply chain attacks are coming in Monday or Thursday nowadays. So yeah, I see a shift of focusing on a high value type of work, which either agents don't yet do or wouldn't do at all in the near future. And I see the flexibility teams are creating in that exactly in these changing roles. That was one thing I was going to ask directly is like what are people doing with the extra time? So higher value work you said, is that continued further automation? Is that better detection engineering, more threat hunting? Any one area of focus that you see people shift to once they free themselves up from the grind of alert fatigue and other things they may have had in the past? Yeah, so first of all, we were limited by some set of guard rights. We had calculation of the stock capacity. This calculation has completely changed. Maybe I wanted to discover 20 more uses that put us under risk, but I couldn't even say that because my stock is 100% and 120% capacity. So no, you can't add another detection even though it seems to be an interesting and can be valuable one. So I can I see teams increasing the coverage of what they investigate in dramatically 500% and 1000% of what they used to investigate. And this is definitely one thing that we see out there. Another thing around threat hunting and the ability to put more time into and doing actual threat threat hunting for many teams was a dream. If you're not well funded, it's really hard to come up with and take in the day building a policies running them on top of your data. Again, you used to need to know how to query this data in all different sources. And now maybe that's some of it is solved. So threat hunting and I would say the last part is branding more simulations. I see that also in the rise and then great. Now we have another mechanism that closes the loop. Let's test that in all the variations where I can and then spot more holes before those become actual breaches. That matches a lot with what I've heard. There's a lot of teams out there that are understaffed and giving a higher bandwidth to take the junk out right more threat hunting. Something they'd love to do more coverage more of everything really. That's the more exciting stuff. So I'm glad to hear that that's what a lot of people are able to unlock with us. Being able to do more of that is generally assumed to be a good thing. However, how do we know that when we add these tools, we're actually increasing our security effectiveness, business outcome level. How are we looking at these with different metrics or are you measuring the efficiency and effectiveness of a sock that has these kind of tools? So this is a really good question because we're debating a lot and how to measure ourselves with their customers, of course, but how to measure the actual impact. Okay, let's say you've added another thousand alerts in a month and all of them were cleaned out. Is it better or not? I think there's some metrics that we've concluded that are good. So one of them is an early detection. Are you able to catch threats at the left side of the Myter attack matrix versus the right side of the Myter attack matrix? How many of those safe ones have zero do you have in a month? First one. Now, it's very hard to measure because there are external factors that will impact that, right? If there is NPM breached every week, then you're going to have one. At least until you know that you're not impacted. And so it's still hard to measure, but yeah, we're trying to measure how much actual impact I was seeing, the impact in the environment and how much of those that we managed to detect early. And I would say we are not fully there yet, but we would want to measure prevention. We want to measure. Okay, it's great. We don't want to get to like in a perfect world, which will never happen, but we don't want to get to detection, right? We want to make it hard enough for the attacker to ditch us and go to some other organization that this is our goal. So we would want to see are we actually preventing cases that might have been putting us at risk? So yeah, shift left and shift left square. That's how I call it. A shift left for just for investigation response to detection and to prevention. And even that prevention, right? Depending on the definition you're working with there, could mean one thing or another team by team and org by org. In my mind, prevention is often about did the attacker achieve the ultimate goal they set out to achieve, right? The red teams attackers don't win if they break in, fish someone, and then they're cut off, right? Yes, they got a password, but they didn't do anything with that password. Now there's some collateral damage there. You got to do some investigation and cleanup, but ultimately they didn't get anything, right? So was that a prevention or was that a detection? Well, a little bit of that is in the definition, but what we do know, which is from the business perspective, the thing is like nothing secret got out, right? There was no expensive data breach. There was no forensics and other kind of, yeah, we didn't have to call the SEC and report, you know, that things got out and there's no massive expensive problem that comes out of that. So yeah, I think that's probably one thing where if you can look at maybe where a sock team starts and then where they are after gaining these capabilities and looking at the delta there between one before and after, that's where people can start to maybe say like, ah, yeah, here's the actual difference that these tools are making. So that can be one approach. Exactly. Did they manage to contain it? If I catch it at the time where the click fits just that running, maybe I'm not impacted yet. And then that's why the shift left from the right side of the miter from the impact side left to the early detection prevention. And eventually it's in many cases it comes down to a numbers game, right? If you have a large top of the funnel, a lot of things are coming in, you didn't prevent anything, you were very open, any S3 bucket is open to door. Maybe you'll detect some things and be able to prevent them, but the next one maybe you missed. So we want to make sure the numbers game is being cut off at the beginning. And before we go down the path of we need to take and contain an action. Definitely. Different kind of question. These tools are ingesting a lot of malicious content, right? It could be malicious command line commands, malicious emails. And in a world where attackers know that AI is going to be reading, whatever they're using, how do you prevent your platform from becoming its own threat vector with prompt injection attacks and those sorts of things? 100%. As I said, my previous work has been working on security for AI. And there we already saw even attacks trying to direct to micro security for example, commands that are specifically directing into that and command line emails with invisible characters that are directing to their phishing and investigator. So first of all, it's here. It's not something futuristic. It may be not as common, maybe not in every attack, but it's definitely going to get there. Now, this all goes back to the security for AI practices, how do you separate the control plane and that plane? How do you do external validation? Do you have some layers of detection on things that look wrong to begin with before it even reaches the AI? Yeah, we support all of that defense evasion in a sense is going to definitely be around as well. If AI is in defense, how do I evade that? And then, yeah, I think we've had the exact same problem with the DRs, right? And now everybody looks like a developer when their leverage includes code for the DR is really hard to defer when this way or another and then, yeah, it's not a lot of them. It's just an idea and an adult type of machine learning that people try to evade. With that, have you seen a sock, like an AI enabled tool in a sock attacked in a way that's been effective or do you have any specific examples of what you've seen where, you know, if a sock is worried about this, right? Like, how can they detect it in those kinds of things? Yeah. So first of all, I would urge anyone to simulate a pry hardly. First of all, it's good. I think part of the knowledge and the skills that we said that are needed to be learned, critical thinking is also to be able to know what are the risks of how to do a prompt injection, how to jail-bake a model. I was so excited doing that for any of the of the chef models and others. Now it's, it got much harder. I think a plenty of the liberators, maybe the only one who's still really good in getting out of any model out there. But I would test it again and again to see that I'm able and build the testing capability to see that I'm actually, and again, this is something that I would expect the vendor to provide. But to see that I'm the, I know that if something like this will come, I will be able to defend by, I think the The first layer of prompt injections and these kind of things are also from the model's provider. It's a problem of AI alignment or dealing with it quite a bit. It's harder now than it used to be before. So if you have a good, they're prompt and judging mechanisms. It's harder to do so only from the data that you as an attacker create. You don't necessarily control all the data and write the comes in into the models. But it's still going to be there. Physically, there is still a way for an attacker to inject content into what AI's suck is looking at and definitely something that you want to test. So the typical red team approach, right? But now we have a new avenue that we have to worry about and potentially consider. Definitely. A question on this. So to the idea of testing and other things like that, you mentioned mythos earlier. Mythos is going to do a whole bunch of testing and all sorts of realms that this could be one of them. Certainly in a world where sock teams are enabled by these kinds of tools and there is a, let's say, publicly available mythos level LLM that attackers can use long term. What do you think is the trajectory? Are blue teams going to be pulling ahead of attackers or are attackers going to be enabled in a way that is still going to be able to overwhelm defenders? One thing I like about working in security word is that the problem space is continuously growing. It makes our job very interesting. Sometimes it's harsh, but that's the real truth. And I think a lot of people come for security for that. It's going to continuously, I believe, it's going to continuously grow. There is no table stakes or like a status quo of, okay, we're going to, we got to a point where we are detecting everything and preventing that and attackers are going to say, "Han, okay, I'm fine. I'm going to go home and find another job." That's not going to be the case, right? And if we do the math today, sock teams handle between two to three criticals every week and there's a lot of things they don't even look into. And now if you multiply the volume and giving these capabilities to the attacker by order of magnitude and you add a I generated polymorphic payload, the change hashes per victim, you add supply chain compromises like Axios and you add Blaster, it's a crow. The math gets harder. It's stopped being mathematically possible to handle human speed, current architecture. So I think you just, there is no other way from how we adopt and grow and utilize that in the right way. And then maybe we get to the equal state. And yeah, attackers have less regulations, they have less concerns and they think in many cases there is this, people say, right, as an attacker you need to be one time right. I disagree with that. You need to be a lot of times right in order not to get catch, especially if the blue team has a lot of very strong tools, but they will continuously try and will continuously try to stop them. So the cat and mouse game continues and employment is still going to be around for a while, so we don't need to worry about it, right? Yeah. Excellent. All right. Well, thank you so much for your time, Lauren. If people are looking to connect with you, ask questions or otherwise, where can we find you online? Amazing. So reach out to me LinkedIn. It's the orange, the ban, ORE and SABAN. I'm very excited to talk about any topic around this word or not. Okay. I'm a diverse person. So yeah, please reach out. All right. Fantastic. I want to watch a super fun conversation on one of the hottest topics out there. Appreciate your time and thanks for joining me on Blueprint. Thanks so much, John. It's quite meeting with you. So a few things stuck with me from that conversation. The context problem is a serious one. These tools are only as good as the organizational knowledge that you can feed them and most socks have not done that work yet. Also the analyst skill that matters most right now isn't going to be writing queries in a specific query language or detection logic. It's the ability to look at an outcome from an AI investigation that appears to be very confident and no one to question it. That's a different John that we had just a couple of years ago. And it's worth thinking about how you're going to approach that problem and whether your team is training for it. Thanks to ORE and for his time and the links to where to find them or in the show notes. If this episode was useful, share it with someone running a sock who's wrestling with these same questions and trying to figure out what to do with all of this. Thanks for listening and I'll see you on the next one.

Podcast Summary

Key Points:

  1. The SANS Cloud Security Exchange 2026 will feature agentic SOC workshops from Google Cloud and Microsoft, highlighting the shift to autonomous security operations.
  2. AI-driven agentic SOC platforms now handle 80-90% of investigations, changing the role of human analysts from manual work to oversight and governance.
  3. Context is critical for AI tools to succeed; they must integrate organizational knowledge (e.g., policies, exceptions) to avoid errors.
  4. Multi-agent architectures improve quality and security by scoping tasks, but a key challenge is getting AI to admit uncertainty and calibrate confidence.
  5. Future analysts need critical thinking, agent governance skills, and deeper technical understanding to validate AI outputs and handle evolving threats.

Summary:

The transcription discusses the rise of agentic SOC platforms, which now handle the majority of security investigations, from triage to threat hunting and case management. , Confluence, Slack) and adapt to organizational context, but success depends on marrying built-in expertise with local policies and exceptions. ” The job of security analysts is shifting from manual investigation to oversight, requiring critical thinking, agent governance, and the ability to validate AI outputs.

, how AI works) to spot problems, while less value will be placed on repetitive tasks like documentation. As attackers also adopt AI, analysts must deepen their technical expertise to stay ahead, making continuous learning and adaptability essential.

FAQs

It's an event taking place in San Francisco on August 17th and 18th, featuring speakers and workshops, including agentic SOC workshops from Google Cloud and Microsoft.

They handle complex tasks like triage, investigation, threat hunting, and case closure, using AI to manage unstructured data and tribal knowledge that traditional SOAR tools couldn't automate.

They sync and structure the data with confidence levels to ensure trust, but garbage in can lead to garbage out, so data quality is critical.

Specialized agents improve task quality by scoping context, enhance security by limiting data access, and manage scaffolding for better performance, as LLMs struggle with large contexts.

Getting LLMs to accurately say 'I don't know' and calibrate confidence, requiring additional mechanisms to flag uncertainty and prevent mistakes.

Critical thinking to validate outputs, ability to govern agent swarms like managing employees, and power-user skills to leverage AI effectively, while still understanding underlying technical details.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.