Go back

EP282: Debating Coupled or Decoupled SIEM with Alex Hurtado and Christopher Witter

48m 39s

EP282: Debating Coupled or Decoupled SIEM with Alex Hurtado and Christopher Witter

In this episode of the Clad Security Podcast, hosts Tim Peacock and Anton Juvakin moderate a debate between Alex Hortada (Director of Detection Engineering at Scanner) and Christopher Witter (from Dropbox) on the future of SIEM. The central question is whether centralized SIEM (one vendor handling storage, detection, and collection) or decoupled/federated SIEM (separating these functions across multiple tools) is superior. Alex argues for decoupled SIEM, citing flexibility, best-of-breed components, and easier integration of custom data sources like bespoke applications. She notes that storage costs can be reduced by using existing data lakes (e.g., Snowflake, Databricks). However, she acknowledges operational challenges, including managing multiple vendors and parser maintenance. Christopher counters with the benefits of centralized SIEM: simpler procurement, single support contracts, and reduced headcount needs for parsing and infrastructure. He emphasizes that many organizations, especially small to medium ones, prefer the ease of a single vendor like Microsoft, even if it is not the best technical solution. Both guests agree that hidden costs (e.g., parser upkeep for decoupled, vendor lock-in for centralized) complicate the decision. The debate also touches on how AI agents are used to argue for both sides, but ultimately, the discussion reveals that while decoupled SIEM is trendy on social media, centralized SIEM remains dominant in practice due to operational and procurement realities. The episode concludes with no clear winner, highlighting the trade-offs between flexibility and simplicity.

Transcription

9098 Words, 49016 Characters

English
[MUSIC] Hi there, welcome to the Clad Security Podcast by Google. Thanks for joining us today. Your host here today are myself, Tim Peacock, group PM on Google Cyclops, and of course, Anton Juvakin, we're a form analyst and now senior staff in Google Clads Office and CISO. You can find subscribe to this podcast wherever you get your podcasts, as well as on our YouTube channel, youtube.com/@cladsecpodcast. If you enjoy our content and want to deliver to you piping hot every Monday, please do hit that subscribe button. While you're at it, you can drop us a review on your podcasting app of choice. Anton, we have two guests today, one who we've wanted to have on this show for years now, and one from a company that I'm personally very grateful for, because I would not have become a PM without their involvement in my career. Did you know that story? I did not. So many years ago, when I wanted to become a PM, I asked my employer at the time, "Hey, can I be a PM?" And they looked at me and they said, "Well, you're not a PM, so no." I said, "That's it. How does anyone become a PM then?" And so I went out and interviewed at a bunch of places, and they all said, "Listen, kid, you're not a PM, so get lost." And that was a bit of a bummer, and then I met the security team at Dropbox, who wanted to hire me to be their PM, and the PM team said, "He's not a PM, get lost." And so the security team wrote me an offer to be a product specialist. I printed out that offer, which is very much not a PM offer, very different. I printed out the offer, I put my thumb over the word specialist, and I showed it to my boss at the time, and I said, "Look, I'm going to do a product at Dropbox, unless you let me do it here." [Laughter] I did not know the story. And so I am very grateful to Dropbox for giving me the opportunity to become a PM. Without being a PM before. And so listeners, if you've ever wondered how people become a PM, sometimes it's because you're an engineer, sometimes because you get an MBI, and there's a whole bunch of PMs out there like myself who got in by Hooker by Crook. Yeah, I mean, we should probably save the story how I became a PM briefly, for some really drunk episode, because there was a moment in my career, good number of years ago, when I literally was, I may have been a director of PM. I don't know, I want to not remember it. So I wasn't really bad PM, and I think maybe my story would be like, how do you become a PM and then prove that if you're really bad PM, you shouldn't. [Laughter] Well, listeners, I think it's very important in your career to figure out what you should and shouldn't do. And there's a lot of things I know I shouldn't do. And I bet there's things that Anton knows he shouldn't do other than just PM. What are we talking about on today's episode other than Funny Curious stories? Because I missed this one. So what did you talk about? This would be one of the longer banners because I feel like it's too much fun. But the topic of the show is also very fun. For many, I want to say decades, no, for many months, a long story, in AI years. Oh, that is decades, man. Yes, decades, exactly. People have been debating whether this new advances of AI agents are pushing SIM, you know, SIM, security information, event management, you know, from mid-2000s. Yeah, exactly. It's surprising, right? Two more centralized structure and more integrated than like, give me my SIM, my SOAR, my EDR, my random other crap, maybe my firewall, in the same tool. Well, maybe not the firewall, sorry guys. But the other one is more about like, hey, I don't actually want them. And I think SIM, can I have detection from vendor A, storage from vendor B, pipeline from vendor C. And this is sometimes called decoupled. And if you decoupled SIM, I think I made up the term. But if you also do storage that is not centralized, then it's also federated. Many, many, many that you can decoupled SIM capability, but keep storage centralized. So these are a little orthogonal. Okay. Okay. So few dimensions decoupled and federated. The point is that conventional wisdom, which I define as random guys on LinkedIn, random people on LinkedIn, seems to hint that the decoupled and decentralized SIM is kind of the future. It's the cool stuff. Everybody wants this. Well, centralized is kind of like, man, two thousands are over. Okay. But when you look at what people buy, it's still the two thousand. They buy, most decentralized. Yeah. And a relative minority, and there I say that tiny minority is actually buying decoupled. Even though the noise and LinkedIn implies that it's the thing. So are you Anton trying to tell me that social media is not real life? No, social media is real life. Who would say such a thing? Obviously. Oh, sorry, sorry. I'm jumping the gun and you're confusing there. My bad. Yeah, no, it's like, no, no. So I just wanted to have people who show up on the show and say, I'm going to argue position A, I'm going to argue position B and have them debate it without us, you and me, or just me as things happen. Sticking their nose in the debate, being kind of a neutral observer. And this is kind of what this episode was. Even though our guests, they're like super polite. And a few moments I thought one guest would skewer another. And they just didn't say they missed every chance of skewer. No, it's just above. We did not get I wanted to skewer fight. Then I wanted to fill a fight. I just wanted to fight, but it ain't the fight. Okay. It's a very gentle manly. Hold on. No, it's a gentle manly gentle personly gentle personly because one of the guests is a lady. So it can't be a gentle manly debate. It's a gentle person. What's the term gentle person? It's a gentle debate. It's a thank you. Okay, very good. Listeners, I think Anton and I have rambled quite enough. And what I hear from him is that this actually turns out to be one of our longer episodes. So let us know in the comments or on Twitter or on LinkedIn, whether you like the longer format, whether you like our longer banter, and whether you, well, missed waffles in the opener for this one. So without any further ado, let's turn things over to today's two guests. And today we have with us Alex Hortada, Director of Detection, Engineering, Scanner, and Christopher Witter, deonarly@drawblocks. Welcome to the show. Hi, thanks so much for having us. Yeah, thank you. So this episode is kind of special and not only because Tim is not here, but because we've been kind of stuck in this debate, in this metaphysical debate about the future of Sim or directions of Sim. And by the way, if you don't like the term Sim, think of it as a technology to support the dejection response efforts that isn't about the debate in the acronym. So where I'm going with this is that in recent years, there are many people came up with technologies as well as positions. That the centralized Sim, one big platform stores data in the cloud or in the frame, is no more. And we have to go with decentralized, federated, decoupled. There are many other terms that imply that it isn't just one product you buy from a vendor. You give the money, they give you storage, security, detections, dashboards, everything else. And instead, you need to do something, well, other than that. On the other hand, success of certain large vendors that actually centralized Sim, EDR, SOAR, is kind of the vote, at least to me, for the opposite side, that actually centralized this winning. And however interesting things are, they're actually even more interesting. Because when people start using AI agents, they actually make the agenteic argument to suit their original position. By the way, this is not Anton's going to talk forever podcast, but I want to make a point about agents and I know how to handle it after you guys. The agent's point is interesting because it's so surreal. People show up and say, Anton, obviously, decentralized is winning because now I can just have an agent go and get the data, whatever it is, and everything's going to be fine. And then I say, okay, if the agent goes and the date is not there, what it's going to do. And of course, the answer is it's going to give you confident answer based on no data. It literally happened to somebody. So they then say, okay, so agents are actually a vote for centralized approach. Then a whole bunch of other people are upset. So we just want to have this debate here. And Alex would kind of play for the decoupled decentralized side. And with her, I would play for the centralized side. Fair. And I'll try to be neutral even though I'm kind of abustered. So I'm almost never neutral. And I'm going to try to annoy both sides equally. Now the team's not here. It's fair. Fair. Okay. So I'm ready. Alex, why do you think that this whole decentralized decoupled sim is a good idea? Why is it not luck based sim? I think we can probably level set and start maybe with the definition of what centralized or decoupled as you coined it several years back really is. And to me, and also for some context, I've seen the sim sim in probably all the ways that it can be simmed as a verb. And there's many different ways to sim now. And this I have found to be a very super comforting way for many leaders that in particular are dealing with the two headed sim monster. They were, they find themselves having to manage multiple sims in their environment. So starting off the back there, but back to the definition. A decoupled sim is the sim that separates essentially storage from detection and from even now collection. So all of those kind of three parts live across different layers essentially. And it's kind of like you can pick your own tool to handle each layer. Whoever does it the best, some call it best of breed per se. And instead, and it's very, very contrary to doing it all under one monolithic architecture where the before the sim handled the pipeline, right, the log collection, the system's log forwarding, all of that. And it's also the same place where it's stored and it's also the same place where analytics happens. on top. So that is in my definition, my words, what I like to call is that decoupled Simflavor. That sounds interesting, but also operationally, time sounds absolutely horrible. You have to pay three bills, right? We've undersied them. So with her, do you have anything to add to this? You can support it. You can criticize it. You can redefine it if you feel like it. I mean, you know, she's not wrong. Like, I like the best of breed approach, but also you got to realize that a lot of organizations procurement is going to drive this. And like, you know, the three vendor problem, one throat to choke or one belly button, if you will. Like, I got to go to one person in one place. And it's that, you know, that procurement that typically is going to get in the way. They're like, oh, well, we're going to get this really cheap. And so the one vendor that rules them all, it's like, oh, they will throw in pieces and make it a harder decision for organizations, securities always cash strapped, not just headcount, but for tools and all that kind of stuff. So it makes it a lot easier to go like, hey, everything's in one place. We have one vendor contract, one MSA. We have one support contract. It's just, it's too easy for most organizations. Now, it'll very based on size too, because, you know, your small to medium market, when they first start needing this and they probably already have Microsoft, they have V5 licenses, and it's going to like easily, oh, well, we'll just throw in this. We'll bolt that on. The data's already there versus I need a parsing engine or something to parse it and send it into a data lake. Like, you start to eliminate all those pieces and parts, because it's just there, which for small organizations already cash strapped, human strapped, it's an easy one. Unfortunately, not the best always, but it's easy to do. So let me ask you a strange question. And I'm going to just like very temporarily argue for Alex's side. Okay. It is easier to buy. Everything you said makes sense. So, and the procurement and the corporate structure and one throw to choke, one face to scream at, all very logical. But if somebody shows up and says, you can do it in a piece by piece manner, which would be cheaper by a factor of, if they say 50, and that's a completely hypothetical number. Nobody says, nobody would sell you a SIM for 150 or the price of a SIM for the same data. But then I'm sure you would overcome those problems, right? If somebody says, it's sure there's complexity, sure there's work involved, but it is 50 times cheaper. Would you change your mind? Would you go bad? I'll give it procurement and do other things. Well, it never is. It's becoming a risk. It never is. But like I'm testing. Like if the decentralized is 50 times cheaper, but just as good, would with there's hypothetical organization, why do I have the head count for it? Right. Again, it becomes like, I need somebody to write detections, I need somebody to write parsing libraries. Most analysts maybe can write detections depending on the platform because your query language is the same as your investigative language. And so you're just like easy to set up. But if I need someone to write parsers, kind of that's sort of the line in the sand for most detection engineering versus and response teams will draw. It'll be like, I don't want to deal with that infrastructure, nonsense, and do I have the people in the skills and capabilities to do it? They become data janitors themselves. I feel like and you have to be very realistic about whether, you know, if you have the time to go into fixing all of or doing all of the mapping, all the parsers and ensuring that you have not only the data into a detection ready format, which takes a ton of time in and of itself to then get to the actual rural building part. Yeah, I do see that. I do really see that, but I would argue. I would argue the one throat to choke argument, which I know like procurement loves to see that coming, right? I want to say one support queue, but it's also like one roadmap, right? You're kind of fully committing to this vendor's roadmap and that's going to take a long time. You're just one of many, and I've been in a ha a lot like many hours in my life and they're, you know, they're really not going to listen unless you're about to not renew with them and then maybe they start listening. Wait a second. So I something unexpected happened because I thought that in my completely fake artificial example of the decouple being 50 times cheaper, with her didn't just say, well, if it's 50 times cheaper, I'll just buy it. But there was a nuanced answer. So that to me is kind of another strong vote in favor of viability of a decouple approach because presumably even if the price difference is artificially huge in my made up example, you would still kind of decide between the two choices. It is a huge price difference because it is becoming almost unrealistic to centralize all of the data into one place and still support the visibility that, you know, your organization is generating data every year. No, no, I meant to say that the license for softer. I think that Alex brought up the cost of storage and that changes the math. But if the two itself, and that's actually a next one point to add to the discussion, if the two license itself, like you're buying three pieces from three vendors and together they cost much less. What would they take it over? Come the resistance of no, no, no, I want one then there one too. That's more of a wither question maybe. Yeah, no, I mean, it's simple. It would come down to like, you know, do I have the people to operate the tooling, right? It would be hard, I would say because you think about what are the largest data sources for most organizations, right? You know, maybe it's my EDR, a 300,000 employee environment, right? And I don't have a lot of cloud or I'm moving to the cloud. Like there are still organizations out there that unbelievably are still like on the fence, still have on prem, still, you know, doing their own data centers or something. And so where is the source of all that data and pain because to Alex's point, like if I have someone who can write the parsers and I'm getting it for cheap and free, like I'm now DevOps tied to that forever. When my firewall vendor moves a field, I break the entire pipeline and I have to dedicate people to that, right? And they have to be aware of it and then they have to fix it in a timely manner and, you know, maybe backfill data or whatever. And that becomes a problem. But if my main data sources, the ones that are most, my analysts use the most are the easiest to access. It becomes much harder, I think, to make the case for it because even though it is cheaper, there's those hidden costs, but that's the hardest thing to like flush out for a procurement team like, oh, but it's going to cost me, you know, two and a half people to have them. Well, do you really need those two and a half people? Like, or is that a full time role? Well, it's half an FTE, but it takes like time to set up and like maintenance and overhead. So there are hidden costs that it's really easy to like be blind to, which have operational impacts, of course. Yes. Well, on the parser ecosystem, no, it may not always, I mean, I think we may be missing out the part that doesn't get set enough, I feel like, because in a, let's platformize, you know, large vendor, same, they are building parsers for the greatest common denominator, right? The sources that most of their customers have, right? Not these edge cases, not, you know, that custom app that your team, you know, built and have invested 10 years in and that they're not going to have a custom parser for that. And the decoupled model is one that, you know, leverages data lakes, right? And they are much easier to consume these one off, you know, be spoke, you know, data sources, if you will. I buy that I may buy this argument. I feel like it's probably a wash in terms of a large vendor may ignore a secondary data source. And you do have, in case of an in withers argument, you do have one face to scream at, you scream at it. And the face says, now we're not doing it this year. We're too busy with other parsers. So you may not get an advantage unless you pick the small vendor and you're a large corporation in which case you achieve more. But then on the other hand, in Alex's example, you deal with a special to vendor who is small, who just as parsers, like a pipeline vendor. And maybe you have more leverage with them. Is that plausible or not? And I don't know which side is is winning in this case. I think it is plausible, unfortunately, from experience, I hate to say it. I hate to give Alex points, but yes, she wins in that one. So let me also add one more dimension to a discussion because we sort of threw together decoupled and federated. And in reality, at least in as I originally thought about this, there is a decoupled sim cup ability and they're salarated storage. Yeah. So there is a way. And I feel like they're now more vendors that are both decoupled and federated in the sense that you can buy detection from one vendor, dashboards from another vendor, but storage you buy from no vendor. You rely on your own three buckets, your own storage at Sass vendor included storage or something else. Like, or data bricks here. And then some stuff with already a snowflake, we have a license. So it's kind of free in that sense. So the idea of a federated was to decrease the cost of storage. Otherwise, you pay it for storage. And when I hear this argument from intelligent, well-intentioned people like, well, you Alex, I shudder a little bit. And the reason I shudder is that if my auditor shows up and says, give me 12 months plus one day of data that your story of the PCI DSS, I can't tell them like it's over here and some data's over there and some with this snowflake and some of it. I think was an industry bucket, but I'm not sure if it's still there. That's what makes me kind of nervous in this. Is it fair with or maybe with or first and then in the edicts? I mean, having it all in one place, you know, everybody's got hot cold, warm tears. So I can archive my data and have it off in another place to rehydrate or access at a later date. I mean, that's not usually a problem. And that centralized approach then, I'm not chasing it down. Now granted, you can have documentation to say like this thing, this source is put here, this source is put there. And some organizations, because of the amount of time they have, will still run into that like when they're migrating from one giant platform to another giant platform, they're running them both at the same time because my data's tied up in this awesome platform that I pay one vendor for. And I have to now, in order to meet my obligations, start with a secondary license with another vendor so I can start like getting my historical collection there, which they're going to be running two systems like all do data's here, all data's there. So you will still run into a similar situation because unfortunately we all know we're tied into the platform. Once you're there, it's like lock in city. I think that migration, the migration story to me is a transient story, not ongoing, right? So you're right. Like when we were younger, presumably and symbols on appliance, what you do is you cut the contract to vendor A, keeping appliance in the shed, you don't pay anybody. Data is there. You'll migrate the data. You buy a new appliance from a new vendor. You use that in case somebody says, what's the data from last July? You pointed the closet and say over there. And it kind of mostly worked out. But if you're moving from SaaS to SaaS, yeah, you're paying both contracts. And also the data is. >> Hellbound popped up. >> Yeah. So we're constantly. >> Yeah. >> Because to me, this part of a debate, I think just straight goes to decoupled to federated. But to me, that's a transition stage. That's not, there are no operational advantages. That's a transition to advantages during migration, correct? >> Yeah, I like to call that like renting a brain. So with the advantage of the decoupled model, you're invested in your data strategy that stays put and then you could just bring, you know, insert your own, bring your own analytics engine on top or scheduler or orchestrator or what have you. And you can do that ideally with, you know, your EDR as well. In the case, if, you know, your EDR vendor also wants your Simbusness, like you can't really check out of there, you're kind of locked in in that case. But I even think it's still more of a case for, you know, you were mentioning the shuttering of when audits come around and when compliance comes around, you know, as the data going to be there in the decoupled model. In the case of a data lake, that's usually the case. You're not really looking for information in your Sim, right? So it's likely going to be in your S3 and they're very prepared for that since they've been having to do it since they became a public company. So I haven't found that to be a problem in my experience. Okay. What about the detection speed? I think like in, you know, prep notes, we kind of mentioned that possibly the detection speed argument should go to the centralized model because if I have the data being close to data is better, data is structured in a way that supports the detections with it. Can you make a strong, can you make a case where it's strong? No, absolutely for that for centralized in this case. What are actually going to win this argument? I'm just going to, I'm just going to. I mean, you can't find that. It's an easy one. It's like, you go to an all you can eat buffet and all the foods already there versus like, oh, why won't burgers and Jane over here or Alex is like, oh, I like Chinese and we're all going out together. Alex is going to go get Chinese. I'm going to go get burgers. We're going to sit down in the park and have our dinner. All you can eat buffet, it's already all there. So it's a lot easier to pick and choose your events, do your correlations. Again, going back to the centralized use case of, if you're using one vendor, like we use Microsoft for as the like simplest example and the behemoth in the space. Oh, I have them for identity. Events are already there. They're already parsed. Oh, I have Microsoft defender EDR. Oh, the events are already there. Oh, I have Office 365. If you think about investigations and most organizations from the corporate perspective, man, your investigations are done. Your detections are easy. It's all plain and simple. Then you'll have to worry about the cloud side of the house to do detections on things, which I mean, you're going to have you're probably going to be in Azure and then it's going to be just like that that much simpler. It's all like, ready to your fingertips. That's a tough one to like argue against unfortunate. But it's I think the only, I mean, if I were Alex, I'm going to channel my inner Alex, I would argue on just price performance argument. I would just make a price performance argument. I would say that yes, this approach is better. But imagine that if you're paying a million dollars to have it as good as you have, how much less would it be to use a decentralized approach? This may not be as good, but the bill would be dramatically less. Would you accept that? And I tried that in real life. And some people said, actually, no, we won't accept that because the quality of data we have is kind of what we need. If we would get slightly worse data at one tenth of the cost would say no. And other people said, okay, show me the numbers. So any reactions to this. I think that your Azure detections in Sentinel argument is solid. And of course, one then there are probably parts that are on logs, at least some with well and currently updates it. But what if somebody shows up and says you can get 10% worse and five times cheaper? I think what's not always brought up here either is the results of these detections. So yeah, you have them all in there. You have all your alerts. But are they good? I have a previous podcast episode and I literally called it, are your alerts really like the near your alerts suck for if we're censoring it? And most of the time the answer is yes, actually almost always it is. And I believe that sure you not only are is it cheaper, but also you're getting a better quality of an alert because you can do much more advanced analytics. You can do you can do a lot more complicated joins with SQL and with data science and data breaks or data lakes. It's much better prepared for those kind of things. And they have been, I mean, we just all know that because I mean, if you look at departments outside of security, I've always been jealous that they're so far beyond what we can do, right? Like e-commerce, finance, fintech, because it makes them money, right? And so of course they're going to use statistical analysis and cake clustering on all this stuff on how to make more money. Whereas we could probably do some of these same things to against swath logs or DNS, but we're not we haven't advanced like as fast as those other business units. And I do think there's are the produces a better quality of an alert rather. And that gets overseen because it's just because of the cheap of the pricing of it all. I think. I mean, I'd have to argue the meantime to detect there though is different. And if we're talking about data quality and data cost, what is the meantime to detect, right? Like, you know, there was one 1060 a very long time ago and vendors claim different things. But the longer it takes for those analytics to happen and for all that data and information to get in, especially also to with a decoupled system. Now you have like three things to break potentially or you know, your ingestion is taking, you know, five or six minutes to get these bulk logs ingested out of S3. Now I'm like load them into my detection engine or I'm running my scheduled SQL queries against them because I can't do real time because I'm not as cool as the centralized model. And so you end up with this problem that it may take 10 minutes or 15 minutes, you know, from soup to nuts from activity happened, not including, you know, coalescing of logs of vendors do or whatever to the actual like, here's a detection. Oh, that's like 18 minutes old now by time in analyses that are even longer or if you're running on scheduled queries and you run into like, there is a trade off that we can't just say like money is the root of all evil here and that's the thing that's most important. Unfortunately. Well, procurement would like us to believe it and we are a call center. So of course, we're not making money. So they want to keep it as spend as little as possible. Another thing I would argue in your favor, with her is the detection distance to wait. Wait, detection distance in what sense? Yeah. So you brought this up in our show notes and it's like the number of hops, right, that your data has to make before a detection fires, right? Clearly that's a lot more of that happens more hops that has to happen at the in a decoupled model, right? You have to pick it up from your pipeline player, then you have to put it into the, it has to land in an S3 even and then maybe snowflake and pick it up from there. That's already three hops right there. And then by the time a rule gets run, you know, against the disk. So there is also that part of it again. But then the offset is is that going to get you a better alert that's worth looking at? Probably yes. I actually, I think this side is very clearly in favor of centralized, but there's something else came up in this discussion that I want to pick on. I think the topic of real time or near real time. And this to me is an interesting question because ultimately when the Sim was a big database on premise, a lot of obsession was about real time. Can you do it in a second? Are the events delayed? And a lot of this was kind of sacrificed a little bit when Sim became SAS. So real time started to be understood as maybe seconds, maybe minutes, I don't know. But with the coupled and federated model when some APIs are kind of best. effort and my security component of a same called some storage and then API goes into your queue and the SaaS platform that has the data decides to give you the data in an hour. Like, how bad is it? Like, can you do real time in any kind of decoupled and federated model or is real time gone forever from the throat? I would say like it all depends on data source, right? You know, VPC flow logs are coalesced in, you know, 10 minute intervals or whatever. So it's not, it's already not real time. So, but if I can ingest it quickly in a, my, if we're saying that like if it was Azure and we're you know, to my defense, if you know, it's 10 minutes till it's coalesced and it gets dumped and then I ingested immediately, that's 12 minutes or something like that and now I write a detection. It's still not real time, but it happened in two minutes time. Like, I got the data immediately and I did the thing. The only time you're going to get anything near real time is if one you're streaming on ingest and you have a streaming detections platform where data is coming in. I immediately compare it against, you know, different libraries and information and kind of like to some analysis there and write some initial detections, but also at the same time to our defense for scheduled queries and kind of that data lake model being able to look back over, you know, if you think of adversary activity, well, I want to find like these three things that happened in a 10 minute window. Like, if that's important, that's not real time, but that's a solid detection. Yes, that I wouldn't want in real time. Like, don't give me three individual alerts because they're meaningless, but together those three are like very, very important and you think of like, you know, reconnectivity under something. You're like, oh, well, these are. Econy. Yeah, but if the third one comes to more a morning, that's a failure. If the third of the three comes to more a morning with the correct timestamp, to me, that's a failure. I can't really choke at the success because. Absolutely. I agree with the real time, like the sub-second real time is probably largely gone, but that type of like human scale real time depends on the data sources is a good answer, but practically, can you deliver detection value if you want to try to act in minutes? All this recent talk about vulnerability apocalypse implies that maybe you have to respond very quickly. If you have agents, they can also in theory respond on the minute timeframe. But if the logs show up tomorrow, your one minute response has failed. So, Alex, do you have anything that makes the decoupled and federated SIM delivery in this case? Well, let's not forget about some of the limitations of the wind detection happens at the time of ingest. Oftentimes, it's not that many. I mean, Splunk has a 200-query concurrent limit on how much you can keep doing and keep staggering on a concurrent basis. And you have to look at. I mean, and of course, that's the detection engineers job to see how you should be staggering. Your searches are orchestrating the scheduling of it all, but the limits really get you, especially for an organization like the size of Dropbox. I'm sure you have thousands of detections that are super, super priority. Are you reaching your limit, your max limit of real time detection? And that's something that you kind of have to stagger and prioritize. But in the case of thousands of detections running at the same time, that's going to be at the enterprise level. But maybe for a small organization, it might be okay. You might be okay under that 200 limit. I think it's going to depend to what your vendor is providing to. Sometimes, right, you know, EDR, if I'm only worrying about what the EDR vendor provides, then there's a lot less detections I'm going to write. I'm only ingesting the alerts. I can act faster because they hopefully find it immediately. It's not retroactive. There may be something that are retroactive, but that's going to be a problem. And you will, sure, you will hit limits depending on size. And it goes to your earlier comment about quality of alerts. Like if there's thousands, are they really working? Are you able to validate that? That's good. You know, if it never fires, does that mean we do? Yeah, people with thousands of rules make manual rules. Yeah, do we just start? Yeah, in some of the larger platforms actually have caps, which I have run into before at other places, where during the pandemic, when you couldn't get hard, where it was a SaaS platform, they're like, hey, actually you're capped right here right now because we can't get enough hard where to continue to scale our cloud operations. And if you think about that, that's just crazy. But also, it wasn't a pay per drink platform. So like, you know, the Sentinel, oh, an alert fires, that's 0.001 cents every time it fires like your warehouse costs are sunk. I buy a warehouse, I use it, I run my alerts or detections, am I data lake against it? That's like one operation. And so I don't have to worry so much as long as I keep it like utilized and it doesn't get overutilized. I think most of it to me is absolutely incredible. Yes, I feel like this is, if you pay more, you'll get better performance with the centralized system. It's kind of correlated. Like if you want faster queries by more cloud, by more hardware, whatever you're doing. So that to me, that success is probably not obvious to replicate for the decentralized. But let me give them that where we are on time. I wanted to bring up the elephant in the room, which is of course AI. Does either of you have an opinion whether the current user AI in detection response, security operations, socks, favors, the centralized or decentralized. And this is like an important question. So I should probably do a drum roll. We'll figure out whether I'll add interest can make a drum roll happen. But do you think AI agents in soccer DNR favor centralized or decentralized? Well, I don't know. Don't look at me. I would still argue they favor decentralized because the storage layer is abstracted. And oftentimes versus the monolithic approach, the storage layer behind your sim is not often reachable by anything else. And just coming out of B-sides as F, it was a while ago at this point. But coming out of B-sides, there was an amazing panel about how, well, one, I think we're on the same page about you should never DIY your sam or vibe code your sim. I think we all know that that's just not possible. But then the second thing was now customer, cross-specs, customers, they're now only considering products that have an MCP now. If their cloud can communicate with it, if their tools and their systems, the way that things are communicating is now fundamentally changing. And that is not possible in a monolithic model. I mean, unless they're coming out with MCPs, but I don't really see that happening yet. If you think about how AI likes to work context, and the less thing is it has to ask and query while the centralized model, maybe it is one MCP, or you're hitting up a series of APIs, all the data is in one place. It's probably likely to get the data in the same time frames properly correlated and formatted and served up onto on a platter. Now, with a decoupled model, it makes it one of the problems with decoupled, or potentially if you're running something more like that, where your analysts have more work, because they have to touch more tools or things potentially. I need to go here to search for this historic data. I need to go here to search for this other data. It's all centralized. They don't have to worry about that, but the problem is that the AI takes that problem away from the analyst and gives points to the decoupled model where it's like, "Oh, it's not important anymore because I've abstracted that away from my human beings," because they're just going to interact with their AI and say, "Hey, I need all the logs historical within this time frame." It goes, "Cool, I'm going here for that date. I'm going there for that date." Then they present it up on a platter like, "And here's what I think about this incident or investigation and make it super easy." I'd almost call it a draw, but if you want to go down to token burn and money, I'm winning. Fair enough. You guys are chickening out of this very critical question. I think both of you are just committing this, well, I guess Alex was a little more committed, but I didn't really hear like a really good argument, though, because what if the agent goes to some kind of this far away land for data and comes back with nothing? Somebody would have to engineer around that. It's a fair bit of engineering, with centralized system, it would be very clear. I need DHCP logs. I go, they aren't there. I tell the system that to make a decision, I need DHCP logs, they're not there. But you have to engineer the same thing for decoupled. For all sorts of failures, you have to decide what to do when you don't have the data and whether you can still do your job of an agent, in this case, without the data being available. Maybe it's available in the morning, but not in the afternoon. Maybe now it's available in three seconds and in the afternoon, it's very busy. That's what makes me nervous. With a centralized system, I say I need this data, I either get it or I don't get it and I know what to do with decoupled. It's more nuance. So you're saying it's more prone to hallucinations if it's a decoupled? Not more prone to hell. It's harder to engineer so that it doesn't make conclusions of their incomplete data. That's the one one narrow argument. But isn't that what the MCP is talking to each other supposed to figure out? Kind of yes, but also no. The thinking of the agents is you're over-eager intern who may be naive about the impact of that thing. And they just want to please you Alex. They want you to be so happy about everything that they presented to you that you're just like, "Oh, pat them on the back and send them on their way." And I think with But that's that like engineering case in environments where you're building it. Those are simple guardrails. I mean, we have them. We're using agents in our environment and we've we've run into those problems, right? Where it like created a situation or a scenario, but because we have humans involved who are looking and validating like early on in our experiments, we're able to easily say like, okay, no, you need to validate and go back and do these steps like, you know, counter your arguments, come back to me and say like, well, I think it's this. I'll prove that it isn't that. And you know, flip it around, which has been very interesting, but occasionally like, I'll say since the English language is so nuanced, it can put the same sentence in, it'll be wrong. If you rearrange the same and it's come down to seven words and a conversation I was having in the nails, like, no, no, if we take this word and move it here and take this word and move it there, like this sentence is 100% true. If we take it as the agent gave it to us, it's close to not exactly, but it is more even different. Yeah. So, and you know, when you're doing analysis or when you're allowing an agent to make decisions that could cost your organization millions of dollars in breach costs and GDPR fines and violations like, you don't play with that. Okay. Oddly enough, you both convinced me that it is decentralized versus centralized and agents doesn't actually call to either side. I was secretly believing that agents favor centralized, but now I'm shaken in this belief. So I think it depends kind of the answer rather than centralized is better. Okay. So, given where we are, tie wise. I want to get to our favorite traditional closing questions, which again, Tim usually does. So, usually we ask all guests for one tip and one piece of reading, a book or a resource or as Tim doesn't tire to remind everybody it can be anything except for my blog. So I'll repeat the same argument. So would you care to provide one tip and one piece of reading? All right. Well, as I say, I have the book that everybody should be reading. If you're into threat hunting and Mac, you can't go wrong. It is like, you know, it's sitting here. I bought everyone on my team got a copy. We're going to go through it one by one chapter by chapter and kind of learn because there is no resource like it out there. So threat hunting, maculasse by Jerem Bradley. It's awesome. Funny enough, I didn't even know it just literally sits on my desk as a reference. One tip would be if you are exploring using AI to our point, like evaluations are not a joke. Science is important. Changing one word in a prompt can impact the outcomes substantially. And so testing and validating and not just kind of like, Las Afer. Oh, we're doing AI. We're doing agents like having a prompt, having the output and having the model and keeping those consistent while you test and validate is super, super important. Like, you can experiment and play around. But if you're going to be serious, you have to be documenting this stuff and actually testing and retesting and having datasets, you know, and not making changes to get to where you need to be. Like, okay, we're not quite it. Where we want it, it's getting good, but it needs to be better. And when you're adding different things, you need to go back through, okay, we just changed the rule. Like we just gave it a new guard rail test everything again because one simple word, and I'm talking from changing the word from, I want to find malicious to, I want to find bad, will impact the outcomes. No joke. Perfect. I love it. I really appreciate it. So this is very useful. Alex. MacOS front, it's really would love to spotlight a content creator, another one that is writing a lot of research in this space too, especially on the detection engineering front is Olivia Galucci's. Oh, yeah. I've seen her speak at unprompted. She's awesome. Yeah, she is. But then the more of a tactical resource as well that I have up every day as, and I'm using as I'm building out the detection engineering library at scanner is the work from Andrew Vimbleet, library.tired/labs. It's fantastic technique research report library, and it goes beyond, I mean, your atomic point detection. This is like full, deep, like you can't find any CTI content on there because a lot of the times I use these CTI blogs out there are really just trying to sell you, you know, they're their product and just really filtering through that is really tough these days. It's almost like, do I even want to use my tokens for it? So I find these, this work by Andrew Vimbleet and it's entirely because of him and I know how he works and he does fantastic work in research and building very thorough and resilient detections. So I'll drop that for you guys in the show notes. Perfect. Awesome. That's really fun. I kind of want it maybe a little more pillow fight energy, but that actually worked out really well. So thank you for your new on stakes and we're going to have Tim listen to it and he decided who won. You wanted a pillow? Well, yeah, I mean, you guys do argue a lot on this podcast. We do. But winner, he seems like you a girl dad winner because it seems like you're a girl dad. No, no, because you're so kind. And I'll be honest, I could have argued either side. And so like it is to each their own and everybody is different, every environment is different and it'll make sense for different people. Yeah. Unfortunately, that's the way it is. Perfect. I argue that this take that I saw recently that I haven't yet to talk about anywhere. And I don't know if you saw it Anton some person, I won't say the name, but they said that detection engineering shouldn't even exist. And it only does because Sim has spelled us. Do you remember this by any chance? I have not seen it, but at the risk of extending the podcast beyond the 45 minutes, which we have swore to never do, I feel like I've seen if the take is the same that I've met people who say, I don't want to engineer detections. I just want to consume them. You're the vendor. Here's the money. Give me detections. I see. That take I've seen. And in a sense, that smaller companies is probably correct. Like they're not the engineer detections. They don't want to engineer anything. They don't have engineers. They have people. They have a little bit of money and they wouldn't define somebody who would take their money. But the threats will get detected. That's how we have a D.R. and that's how we have a lot of other tech. But to me, that's not a detection engineering failure or Simfailer. It's just that to me, there is a category in the market where the detection engineer is not for them. They don't want to engineer. Yeah. Well, I mean, they're still doing themselves a disservice, not building custom detections. And I'm so sad that Tim isn't here because I believe the analogy that was used and he's like the king, Tom Brady of analogies, was that nobody buys a car expecting it to crash and like expecting you to do all of this work that you have on your car. Otherwise, you're going to crash in the same way. Nobody should be expected to build and really fine tune and customize their car. Otherwise, you know, it's your fault. Separate debate. Let's just, I'm just going to abuse the host privilege and say, stop. It's really fun. We got to just have to stop. Wait a minute. No, we got to discuss some of the time. Thank you for being on the show. Give it a tear me apart for 45, 47 minutes of recording without him. So again, thanks for both of you. This is really fun. And I think the energy kind of went to lively, friendly energy, not a debate energy. So thanks a lot. Thank you. Thanks a lot. And now we are at time. Thank you very much for listening and of course for subscribing. Please subscribe so you can get new episodes piping hot. And if you love our content, please drop us a review on your platform of choice. You can find this podcast on YouTube, Apple Podcasts, Spotify or whatever you get to watch our podcast. Also, you can find us on our website cloud.withgoogle.com/cloudsecurity/podcast. You can argue with us on the Google Cloud Security Community site, googlecloudcommunity.com. You can follow us on xx.com/cloudsidepodcast, treat us, email us, argue with us, and if we like or hate what we hear, we can invite you to the next episode. See you in the next episode of the Cloud Security Podcast by Google. Thank you.

Podcast Summary

Key Points:

  1. The episode debates the future of SIEM (Security Information and Event Management), contrasting centralized (one vendor, integrated) versus decoupled/federated (separated storage, detection, collection) architectures.
  2. Alex Hortada argues for decoupled SIEM, emphasizing flexibility, best-of-breed tools, easier handling of custom data sources, and potential cost savings, though she acknowledges operational complexity.
  3. Christopher Witter advocates for centralized SIEM, citing procurement ease, single vendor support, reduced headcount needs, and simpler management, especially for smaller organizations.
  4. Both guests discuss hidden costs, such as parser maintenance for decoupled systems and vendor lock-in for centralized ones, with AI agents being used to support both positions.
  5. The debate highlights tension between what is "cool" on social media (decoupled) versus what organizations actually purchase (centralized).

Summary:

In this episode of the Clad Security Podcast, hosts Tim Peacock and Anton Juvakin moderate a debate between Alex Hortada (Director of Detection Engineering at Scanner) and Christopher Witter (from Dropbox) on the future of SIEM. The central question is whether centralized SIEM (one vendor handling storage, detection, and collection) or decoupled/federated SIEM (separating these functions across multiple tools) is superior. Alex argues for decoupled SIEM, citing flexibility, best-of-breed components, and easier integration of custom data sources like bespoke applications.

, Snowflake, Databricks). However, she acknowledges operational challenges, including managing multiple vendors and parser maintenance. Christopher counters with the benefits of centralized SIEM: simpler procurement, single support contracts, and reduced headcount needs for parsing and infrastructure.

He emphasizes that many organizations, especially small to medium ones, prefer the ease of a single vendor like Microsoft, even if it is not the best technical solution. , parser upkeep for decoupled, vendor lock-in for centralized) complicate the decision. The debate also touches on how AI agents are used to argue for both sides, but ultimately, the discussion reveals that while decoupled SIEM is trendy on social media, centralized SIEM remains dominant in practice due to operational and procurement realities.

The episode concludes with no clear winner, highlighting the trade-offs between flexibility and simplicity.

FAQs

It is a podcast by Google hosted by Tim Peacock and Anton Juvakin, covering security topics like SIEM and AI agents, with new episodes every Monday.

A decoupled SIEM separates storage, detection, and collection into different layers, allowing you to pick best-of-breed tools for each part rather than using one monolithic platform.

Centralized SIEM is easier to buy and manage with one vendor contract, one support queue, and one throat to choke, which simplifies procurement and operations for many organizations.

Decoupled SIEM can be cheaper and more flexible, especially for handling custom data sources, as it avoids vendor lock-in and allows use of existing storage like data lakes.

Hidden costs include needing dedicated staff to write parsers, maintain pipelines, and fix issues when data sources change, which adds operational overhead beyond the tool licenses.

Some argue AI agents support decentralized SIEM by fetching data from anywhere, but others say agents need centralized data to avoid giving confident answers based on no data.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.