Go back

Episode 2: Dane Stuckey

88m 44s

Episode 2: Dane Stuckey

In this episode, Dane Stucky discusses key aspects of detection engineering, focusing on the ADS framework he developed at Palantir, which standardizes detection documentation by requiring hypotheses, blind spots, and response steps. This approach improves detection quality, facilitates peer review, and scales team efforts. Stucky emphasizes the value of pre-built vendor detections for organizations with no existing coverage, as they provide immediate risk reduction and stability. However, he warns against relying solely on them, as they limit extensibility and long-term growth compared to raw telemetry, which allows for custom, threat-focused detections. He advocates for detection validation through CI/CD pipelines, using unit tests and atomic tests to ensure detections remain effective over time. Automation and enrichment, such as using SOAR tools to triage and enrich alerts, are critical for scaling response and reducing time to action, especially when centralized telemetry is incomplete. Stucky also discusses the trade-offs between raw telemetry and synthesized alerts, noting that the choice depends on organizational resources and skill sets. Overall, the conversation highlights the importance of balancing immediate security needs with long-term extensibility, leveraging frameworks like ADS to document and validate detections, and using automation to enhance response efficiency. The episode underscores that effective detection engineering requires continuous improvement, collaboration, and a strategic approach to tooling and processes.

Transcription

16303 Words, 89214 Characters

English
Hey everybody and thank you for joining this week's episode of Detection Challenging Paradimes. With us today is Dane Stucky and we talk about many different subjects but to just name a few are Dane impalenteers passion for open sourcing defensive capabilities, the benefits in potential issues with pre-built detections from vendors, the flexibility of saccles, and at what point should a detection engineer surface from the self-made rabbit hole during the research process of detection? Thank you for joining us today and Luke. You know what to do man, roll that intro. Dane, thanks for joining us today. How are you doing man? I'm doing great, thanks for having me guys. I hear you recently moved to Montana, how's Montana's right now? It's dope. Absolutely love it. Unfortunately not a whole lot of snow this winter but hopefully that will change. Nice and ball me in 40 degrees out. What has been some of the best activities you've done out there? Just a bunch of hiking, done any cool fishing yet? Yeah, I've got some fly fishing in over the summer. I mean I suck at it but it's the pursuit, it's a lot of fun. It did a bunch of shooting this summer and hopefully got into the cross country skiing once we get some snow in the ground. I have a story about cross country skiing. I have some friends that live in Oslo and so we went up to, went up there for a conference and on the weekend they took us up to their cabin and we went cross country skiing and the cabin was situated off the road about like 500 meters or so. I think like 11 pm at night we arrived, we parked in this little parking spot and you had a ski to get to the cabin and I had never skied before of course and not only that but it was uphill and so if anybody's ever cross country skiing uphill you have to kind of do this little fish thing to where you kind of turn your skis outward and then you kind of do one step at a time. Of course I face planted and I had by the time we got up to the house my entire beard was just completely covered in snow so that was my introduction so it was pretty fun. Nice. You decided to move in a cycling? Yeah that's yep exactly. I'm not cut out for the skiing. So name for stuff stuck in my beard. Is it? Yes I'm both scared. You're a beard. Yeah. Everybody here everybody on this call has a beard. Some are better than others I'll just say that. Every beard is better. We'll just say that. Yeah. Appreciate that. So name for those that don't know you. How'd you get into security? Tell us your background a little bit. Yeah man. I think a lot of people kind of stumbled into it when I was in high school took an interest into playing a lot of a day game stuff like that. Started using trainers in one of the where and like how do cheat kids work underneath the hood and that kind of led down the rabbit hole. And then again just matter of luck ended up getting into a digital forensics program for university and that kind of went down digital forensics incident response. And here I am you know a decade later still doing infos x stuff still enjoying it. Nice. Yeah. So Johnny I don't know if Johnny and Luke are old enough to have used something like game junior game shark but I remember when I when I first started using those those things for my super Nintendo you I just remember setting everything to ff and I don't remember I never knew why ff was what I would set it to and then like you know 20 years later I was taking a cyber security class in the Air Force and I learned about hexadecimal and I was like that's pretty cool. This looks familiar. But I was always jealous of all the game hackers we had a guy in our Air Force training that we had this computer game where you had to like go and bomb a bunch of things and this this one guy that worked with us. He was he worked for Microsoft using the Air National Guard and he was able to kind of like hack the game and they just utterly dominate everybody because all of us were just idiots trying to play by play by the rules and he had like super turned his turn his planes into like these you know fighters fighter jets of death and was just destroying everybody so that's that's pretty nifty. So Dane one of the things that we we really have leveraged from your work and like I was to we were talking before before we started recording. You know I met you all as in the Air Force probably in like 2015 time frame and so I've been following your work for for a decent amount of time. But one of the things that we've been able to use with our customers and kind of within our with our team is the ADS framework which is for those that aren't familiar. Dane and his team at Palantir released a a framework called the Allerian Detection Strategy framework and the idea was they wanted to be able to document well I'll let Dane explain what the idea was in a second but from my perspective is like let's document the detections that we're creating and provide kind of like a minimum bar for what we expect when we push a detection to production and then there is you know some some need for kind of peer review and that type of thing. And so that's been really helpful for us internally. It's been helpful for us in training students for like our detection training class and it's also been really helpful with our with our customers. Can you kind of like help us understand what the impetus for kind of that that project was from Palantir's perspective and what you guys gained from it and what you learned? Yeah definitely. Yeah so the idea behind the ADS framework is probably not novel but it came as a result of a lot of failing and the idea is like you basically want to come together with a hypothesis around what is your detection going to do how is it going to work and while you're doing that you want to outline the steps for response right some of it's working up at 2 a.m. and maybe they're not familiar with macOS or whatever detection technology and as we started building alerts in our team doesn't really scale linearly with people right so today I think we have like eight people on certain and we're responding to hundreds or thousands of alerts a day. It really became important to call out some other pieces of critical information like what is this not going to catch right what are the blind spots how does someone defeat this and the real value of this was very evident once we started forcing that bar because the quality of ADS shot up right you kind of had to create this hypothesis when you're testing the ADS you would have to prove that it would work so there's like an element of hunting and then by virtue you know with us we would require another engineer to review it to put it into productions you at the very minimum kind of cross train and get other people to provide insight and ideas to it. It's been a really scalable strategy for us it's had a lot of a lot of returns. Yeah I think there's a there's a few directions that I could go with kind of how this has become really valuable to me. Like one of the really important things we talked about this on our last episode with Olaf Hartong how he had released a project or like a website called sysmon.works and the idea that he used was he's run into situations to where the newest release of sysmon would have errors to where certain certain log events would not be captured. And so he basically created a cicd pipeline that would allow him to in an automated fashion test to make sure that the whatever version of sysmon was most recent actually produced and worked produce events and worked the way that you expected it to so that we think of that kind of like as a sensor validation make sure that your sensor is working as you expect it to. But yeah there's a really cool aspect of kind of detection validation which is let's define what it is that we're looking for right so maybe we're looking for Kerberostin right and so you say okay well I want to detect you you define kind of your hypothesis which is I want to detect Kerberostin or maybe I want to detect Kerberostin that is initiated by a process other than you know lss.exe or something along those lines and and then you you build out your detection and then you have validation steps to make sure that the situations which you intend to detect actually work in real life which I think is a really fundamental aspect it's as we go through the entire detection process we need to kind of be validating each step as we go and I think that's this is a really good kind of way to at least detect or validate the detection aspect of the kind of like pipeline. Yeah definitely and the the dream has always been not that we're there yet but to kind of build detection validation as CI CD where you know you created detection you have a rule language right whether it's spunk or elk or whatever and then create a test scenario like there's atomic you know Miter has their own it's called era but like build it a unit test and then modifications to that ADS can you know test against it and then you can periodically test your environment to see like are you catching all the ADSs that you would expect to fire right over time as you're making edits and stuff things drift. Yeah for sure for sure yeah the other thing that we run into a lot of times is kind of like what you talked about and I think this is there there's always a disconnect between the [BLANK_AUDIO] detection rule creator and the person responding to the alert that's generated from that detection rule. And so I think this has been really common in my experience with vendor provided detections because there's kind of like this proprietary nature of the detections that are being produced by vendors. And so that they won't necessarily say, this is how we arrived at this conclusion, but they'll say, hey, this thing happened. Don't worry about how we know that that happened. It just did. You should go figure it out. But I think a really valuable aspect is the response section, which is kind of, hey, when this alert fires, I'm not going to tell you every and maybe you have a different perspective of it. So I'm interested in what you think. But I'm not going to tell you everything that you should do to respond to this alert. But maybe here's the top five things that you should be looking for thinking when this when you see this alert, maybe like you should go validate that this, you know, that X happened or go and check, check wide where any network connections made from that process or, you know, whatever, whatever it may be. But you kind of give somebody a starting point for these are the things that you should first think about. And then after that, you kind of use your intuition and kind of your, your ability to investigate the situation and move forward. Yeah, I think that's apt. Right. Responses is both an art and a science and the sciences, you know, what are the discrete things you need to do to arrive at the right conclusion, but how you get there is the art form. And the reality with the modern enterprises, there's simply too many technologies for any one person to be able to effectively know everything, right? You know, one day you're going to be responding to EWS alerts. The next day, you might have a Kubernetes alert. And like, you have no idea how Kubernetes works. And so when you were relying on detection engineers to make those response actions, it's like you have to give them a quick primer and set them on the right path. And that's really worked out well for us to we're big believers in the system, obviously. So I'm curious to that point, when it comes to responses on the ADS or at least response section, have you seen it be valuable to maybe create not only a written playbook like what Jared was talking about a step-by-step basis by what the analyst needs to do to remediate that alert, but have you seen an automated fashion or maybe pull that some of that data in a triage fashion for the analyst, just kind of have that right there and ready for the analyst or maybe even have some of the investigation pieces somewhat automated for the analyst. Yeah, totally. And that's sort of the domain of security, orchestration, automation, response tools, right? So we're tools. And that's a huge part of it. And seeing the way we've kind of structured our program is we have multiple tiers of on call that each kind of have their own alerting responsibility. One of those rotations is the thing we call tertiary. And the idea behind it is like, look, we've got a lot of alerts that are running, we've got automation that needs to get fixed or improved or built to get that enrichment, to get the automated response. And so when you're on tertiary, it's a dedicated week of just pure improving the process of squashing bugs, making on call better for everyone else. And as part of that, when we have ADS, we kind of have a backlog of ideas like, oh, it'd be great if we could take this IP address and run it through gray noise and say, we'll remove anything that's considered noise. Or we want to go ahead and grab files or take actions on accounts. That's a crucial part of it if you want to scale your team non-linearly, just by not adding people and also just reducing time to response, right? Yeah, I think the automation thing is really interesting, especially for organizations that struggle to have centralized data. So there's this, the idea of detection at the enterprise scale, it's really important to have your telemetry centralized so that you could run these queries across or these analytics across the entire enterprise or at least some subset of the enterprise. And some organizations have the luxury of having tons of telemetry centralized and other organizations may not. And so then the question might be, how do we leverage or how do we kind of fill the gap that's left by the lack of centralized telemetry? And I think automation or orchestration is really, really valuable there. And so the way that I look at it is, you almost want your telemetry as far left in the pipeline as possible. I consider like the the sim or the you know, whatever your centralized log aggregator is to be the furthest left. And then if it's not there, it's like, how do I identify something that's interesting? And then how do I collect the telemetry that would be valuable and providing context? And you know, in some cases, that's going that context contextual to limitaries going to be centralized also. But in other cases, you have to, you know, maybe run a job to go and collect that information from the actual endpoint because maybe you weren't able to centralize it for whatever reason. Maybe it's price, maybe it's you know, volume, whatever it may be. Yeah, I think that's a good point. And you know, there was a discussion today in the Bloodhand Slack about EDR and how do you evaluate good EDR? Where do you start? Do you go buy a vendor product? Do you self-learn on Sysmon? And I think the thing that you're describing right is you would rather take all of the data out of an EDR sort of the raw streaming events third and your SIM and use that versus just ingesting like derived alerts or derived signal. Is that right? Yeah, so I think it depends on who you are and kind of what your situation is, right? So there's like, if it was up to Jared, right, I would, I want as much raw telemetry as possible. Now, I don't think that's necessarily the case for like all of our customers or you know, the entire industry because there's, you know, people may not, or organizations may not have the ability to invest in, you know, employee, like they may not invest in a security team at all, or they may not invest in or be able to hire the people that have the skills to deal with the telemetry like raw telemetry and build the detections or, you know, there's all kinds of variables that might might happen. But like I think in a perfect world, you would have, you would have a team that's able to leverage raw telemetry to build detections that are threat focused and or technique focused and I guess customize to your environment. That being said, there is like, and I kind of like, I kind of think of the EDR game. I know Johnny wants to say something about this, but I think of the EDR game as as there's a bunch of different vendors and they all are selling this thing called an EDR, but they there's different vendors that have kind of different focus areas and that's all based off like what their desired kind of customer base might be. And so some cuss some vendors are more tailored towards people like me that want the raw telemetry over everything else and like, yeah, I'll probably take I'll take the out of the box detections that you're going to give me or the synthesized alerts, but I, you know, I don't want those in the absence of raw telemetry, but then there's other organizations or teams that only want kind of that synthesized alert. And I think one of the problems with only having like alerts and I don't know that there's tons that are tons of vendors that are only giving you alerts at this point, but there's certainly vendors that are leaning more towards alerting than raw telemetry. I think the problem with that is that limits your ability to grow, right? So it kind of gives you a high starting point, but it doesn't give you the ability to like, it's not extensible, right? So like in computing and security in particular, extensibility is extremely important. And I think if you are only giving me, you know, the output of alerts, then I'm not able to extend the capability beyond what you give me, right? So that I think that that's a huge limitation. Yeah. I mean, you're talking about a moderate to high floor, but a relatively low ceiling, right? Yeah, exactly. Yep. Yeah. I was actually following that conversation in the Bloodhound Slack today. And there's a couple of things that you said in there, Dane, that really, really stuck out to me. And I wrote and talked about and dive deep into those conversations. So they were talking about pre-built detections inside of EDR products. And we somewhat touched on this with our conversation with Olaf. But one thing that I really like that you mentioned, I'm just going to quote you here because I think this is a pretty cool quote. As you mentioned, it's foolish to deny the benefits of pre-built detections and immediate, the immediate lift you get out of deploying good software. So the point, the way that I read that was for an organization that has zero coverage right now, having a good software in place where you have pre-built detections gives you some type of coverage. But you might not want to, you obviously don't want to stop there. Would you mind possibly going deeper into that thought that you had there? Yeah. And I think the assumption I was making, the question kind of came from the angle of like, hey, my company's just getting into EDR. There's a lot to look at, a lot of vendor marketing out there where did we start. And the reality is in a perfect world where you'd be able to go deep dive when is the vent logs, you'd be able to go deep dive. Sysmon really get an understanding of the how and the why. But it's not really like a viable organizational strategy because you're carrying a lot of risk while you're trying to figure that out. And honestly deploying stevels, Sysmon configs with no support, trying to get when is event logs flown and no surprises. It's pretty tough. And so I think a lot of security programs really need to consider the fact that you can deploy vendor software, get support accelerated, have some some blunts of stability, and get immediate risk reduction. And you don't have to commit to that vendor long term. But it buys you the space, it buys you the, it's a force multiplier that allows you to kind of step back and figure out where you need to tackle risk. And if you want to build your own telemetry rich detections in your sim. Great. Rip out the vendor tool. install something better, but you know, you just can't let the perfect be the enemy at the good. Absolutely. I really like that point. To me, it seems a lot of times that whenever an EDR vendor is deployed, and we've seen this inside of some organizations, that it's often times that the detection response process or that team really seems to lean on those detections. For example, ATP, they have a whole bunch of default alerts, a whole bunch of default incidents that can fire. There isn't many, there isn't much documentation on how those alerts fire, so sometimes they can be noisy. They do give additional coverage. But oftentimes it's easy to, I would say, lean on those detections instead of trying to figure out what those detections are covering, and then the gaps that are associated with those and then trying to fill in those gaps. Yeah, totally. And some of the detections are great. Some of them are best of breed. And if you go run some of the pre-canned scenarios for MDATP, you're going to get a lot of coverage of common, you know, miter attack techniques, but they're not customized for your environment. And some of the adversarial activity that you might encounter might be really easy to detect in your environment, if you can kind of custom tailor it, but a lot of the times it's not going to fire out of the box. It's good to get you that lift, but I think that's where exporting a lot of that raw telemetry or running customized hunts or something like that, really getting into the tools capabilities buys you that extra lift. Yeah, that's interesting point, sorry, Jared, one second. That's an interesting point that you just said that how it's not, it's not, it doesn't fit your environment completely in how miter as well. They try to correlate with miter, but that might not be a full coverage of your environment either. I think that's a good point. We talk about this a lot, how miter is a really good place to start in terms of maybe a priority list or things you might want to focus on from a detection perspective, but maybe even creating your own framework that is more tailored to what your environment is more susceptible to in terms of attacks might be a better place for your team to kind of look forward to long-term. Yeah, and just to pose a question to you all, sorry, Jared, let me just go down and try to go for a second. Have you seen any vendor products detect like cookie theft from like a Chrome or a Firefox cookie store? Is that a technique you've observed? That's not something I'm not familiar with now from a vendor like tool. Yeah, and so I take that as an example, right? Like a scenario that keeps me up at night is you've got Chrome extensions, you know, they get bought and sold all the time. Someone back doors a Chrome extension, you know, maybe you allow Chrome profiles to be synced from people's home computers, right? So you install some free proxy Chrome extension, it gets on your work machine that Chrome extension, it's bought. Someone introduces a back door maliciously updates. They start accessing cookies or they start, you know, stealing things that give you access, right? Session-based access to resources, maybe even bypassing MFA because it's a valid session. No vendor tool I'm aware of is going to catch that, right? But if you're a heavy, Samo-based company, you have a lot of Samo apps that are just sort of accessible on the internet and you don't have a lot of endpoint controls around Chrome, that is a scenario that I would be terrified of, right? So I'd much rather put investments into that type of security than I would try and detect, you know, lulbens or whatever. Sure. And that's where the extensibility portion is really nice, right? So off the top of my head, I don't know how you would necessarily detect that, but I assume that there's some file that they that it would be accessing or something like that that you might be able to track or whatever it may be, I don't know, or tracking the installation of different plugins and all that kind of stuff. But I think that another danger is of, so one of the big dangers that we see is often people will install an EDR and then they'll just leverage whatever detection. So this is like, like you said, it's really good as a baseline, right? It's a really good start. But it becomes dangerous when you think that your job is done. So I've literally had somebody ask me, hey, like we have vendor X installed in our environment, they're pretty smart. Do I need, like, why do I even need to add to that? Like, why don't I just accept what they give me and consider the job done? And it's like, well, you know, it's probably a good start. But like, I just, if I'm the person that's ultimately responsible for it, I'm not, I don't feel that there's enough transparency to the consumer of how it's working or what they're trying to achieve to comfortably say, yep, I'm good here. I don't need to do anything else because like to the point of the ADS, like you certainly don't have an ADS worth of transparency for all the different detection rules that the vendor is providing, right? You don't, like they say we're detecting credential dumping. But like the, the goal of the detection is likely more granular than that. So it might be like detecting, like detecting Mimi cats by hash. But then they say, oh, well, it falls into the attack category of credential dumping. So we detect credential dumping and you're like, yes, I mean, that's technically true, but it's not, it's not true in the sense that you're detecting all of credential dumping, right? Or like they're not helping you understand, hey, when, like I've had like a ALS, ASLR bypass, like I know what ASLR is and I know I understand what a bypass would do. But it's like when you get an alert that says there's an ASLR bypass, it's like, what am I supposed to do with this? And we've actually like the customer and I reached out to the vendor to kind of ask them, hey, when this happens, what do you do with that? And they're like, well, we don't know. It's just a data point. And it's like, okay, if I'm a firm believer that if you're a detection engineer, that's building detections for somebody else to consume and you don't know what you should do, what, like you don't know what you would do with an alert that's generated from that detection, you have no business giving that to somebody else and making it their problem. That's kind of like my perspective on detections. And so that becomes a sticking point because I don't think very many people actually people have an idea of what they want to do with a detection, but people don't think through what should I like what do you do with the alert when this is detected because the detection is not perfect in the majority of cases. Yeah, absolutely. And evasions come out, how do you evaluate your susceptibility to it? Yeah, if you can't actually analyze your assumptions and the foundation of what you're building on, it's really hard to independently evaluate the efficacy of it. But if you're just starting, then your detections can be a great first step and at least give you a breadcrime to start your investigation from. Yeah. Yeah, from that point, I'm curious your thoughts on this thing is how often have you seen it be valuable to create a detection, let's say for multiple detections for one technique just to somewhat because let's say process injection, for example, there's reflective DLL, there's DLL injections, all of which have different data points, but they all fall underneath somewhat of the same technique, right? And so how often have you seen creating a multiple detections cover that gap and really umbrella and start to help leverage the data in different ways? Yeah, I think that's a really good question. So if you're like a minefield, it's like we don't have to place really smart minds, we just have to have a lot of minds, eventually someone's going to step on one. And so there are different ways to approach the detection problem. So for an example, let's say you're worried about kernel modules getting loaded on a Linux server. So one type of detection might be look for people calling the kernel module loading commands, what you would expect. That's pretty good. It's like a very easy one to write, but it's not going to catch it if it's automatically loaded on boot. And so you take another perspective at it, you say, okay, well, what if I put all of my non-curl modules in a list and then I alert on anyone that's not part of that list, right? Sort of an allow list of things I would expect. Maybe use a different tool for it. This type of overlapping alerts provide a lot of value. For one, it means that you're viewing the same problem from different angles, right? You might catch something applying spot one of your alerts. Second, second aspect of it is if you have outages and telemetry, tooling, logging, maybe someone disables a tool, you still have a fighting chance to get some signal, right? It's really easy to disable or deter or work around 1 ADR. But if you're correlating syslog, EDR, OS query data, now you're talking about three, four, five different log sources. So I think you also get some defense and depths from that as well. Yeah, that's an interesting point. I think it's really interesting to have numerous different log sources that are kind of looking for the same thing. And you're almost taking a page out of malware's book to where it's like, I remember it used to be really common for to have multiple processes as part of your implant. So if one went down, the other one would start it back up. So you'd have to be very precise in how you take it down, I guess, to try to delete or terminate the process. Yeah, it's just a very expensive thing to do. You make the trade-off in engineering hours on adding additional redundancy or sophistication versus getting better coverage and other platforms. So it's a balancing act. But for stuff you really care about are really important techniques. Absolutely recommended. Yeah, I'm curious your thoughts too, because we talk a lot about detection pretty heavily. But I'm kind of curious your thoughts too on maybe the practice that you've seen be very valuable for, say, detection engineers to understand what to do with an alert after a fire. So that way, way whenever they either create another protection, create an ADS. Their response, excuse me, the response section is a lot more tuned for that alert and a lot more straightforward so the analyst has to do less guessing once that alert fires. Yeah, I'd be curious, what do you all do when you're doing defensive engagements and you're kind of approaching this problem? Do you focus on automation or you're focusing on manual checklists? How do you approach this problem? Yes, so I can speak to that a little bit. We've had some engagements to where we might have a client who wants to have an automation phase in sense of pulling data. Now, I personally do not believe that the investigation phase can be fully automated. I think there's a human aspect that needs to happen. I do think triage can be automated because in my eyes, triage is really just pulling back the data attributes or the data context to that particular data set of that behavior. We try to really focus. We take the detection, we understand what type of data is supposed to be coming back from that and then maybe set a base condition and this is something Jared's going to really like talking about is the base condition and then moving forward from that base condition, adding contextual pieces to that data set that'll then give the analyst understanding of like, okay, this is why the alert fired and here's maybe a jump start for this investigation and then maybe somewhat having a written playbook of things that you want the investigator to go look for the alert but also it kind of helps engage their brain because again, a lot of the things that we hear inside of the industry is, oh, it takes time if you haven't seen this before. Well, some of that can be kickstarted through that documentation and step phase. Yeah, that makes a lot of sense. And Jared, I assume you want to jump in there and talk about your schemas and. Oh, man. Yeah. Okay. Well, let me answer the question first and then I'll get into the next thing. So I like to think of it as like, I have my ideal use case of how you build detections, right? And so that's modeled around the ADS and providing response phases. And generally I would like to kind of say like the response phase, the way that I look at it is what are the, when this alert fires, given the context of the goal that was that the detection is trying to achieve. So it could be a very broad goal or it could be a very like discreet goal. Given that goal, what are the first five questions that somebody should be trying to answer to determine, you know, whether this is actually malicious or benign, I guess. However, generally what happens with like in practice is we have like our perfect scenario and then we kind of take the constraints of the customer's environment and try to fit within the constraints, right? So like there may be an organization that doesn't have tons of telemetry available in a centralized fashion, but they do have a sore tool. And so it's like, okay, well, part of the response should be dictating the additional context that's necessary by the sore that are to gather using using the sore. So our automation. Or maybe there's like, Hey, we don't have the ability to do automated investigation. And so what we do is we say, okay, well, you know, go gather this information, try to answer this question. And there'd be like a series of kind of runbooks that say there'd be like an overarching guide playbook that says, these are the things that you're trying to achieve. And then for, you know, the actual technical implementation, there'd be like a runbook that would say, okay, if you wanted, I don't know, base 64 to code PowerShell. This is how you can go about doing that. Or if you want to do, I don't know, whatever, download a file from an endpoint. This is how you would go about doing that. So that's kind of how I would approach that. Generally, it's dictated by kind of like the technology that the customer has and kind of the constraints that are inherent in that. Yeah, that's a good point. And I'll provide my perspective. So I don't just read the question with the question. So I think the goal right is to figure out, is this worthy of spending up an investigation and spending really valuable engineer hours investigating it? And so the fastest way there is to quickly get a sense of, yeah, this is really spooky. Or no, this is probably benign. And some of the strategies that you invoke there aren't even have the analyst triage or get more information, right? I think one of the most effective things that we've done, and this is all public stuff, is we have a really elaborate, tiered system of triaging alerts through Slack and chat apps and stuff like that. So it's everything from informational words. Hey, you know, someone just enrolled in a device and mem, right? The Microsoft endpoint manager. This was one of the things that you know, could catch a solarwind style thing. Hey, there's a new MFA device associated with my account. Well, as a detection, I don't want every single IR alert for someone hooking up a phone to go to an analyst, but send an automated alert to the recipient and give them the ability to escalate it if they don't recognize it. And so you can kind of start with informational dissemination. You can jump up to prompting, right? Hey, we saw this command that was run, can you confirm or deny it? Maybe get a do a push in there and you get really high fidelity signal for very, very cheap from a detection perspective. And that kind of turn to automate all that context. Kind of decentralized triage to some degrees. Like, I mean, sometimes assuming that the scenario isn't insider threat related, right? So, you know, in that case, you probably, you probably don't want to talk to the person about it and be like, Hey, are you doing this? But yeah, in a lot of cases, insider threat is not the concern. So like, today, Luke and I were looking at. So in cast, Microsoft, I don't even know what in cast stands for. It's a Microsoft cloud kind of kind of thing. They have, they have an alert that was somebody, somebody has logged into a logged in from a country that they aren't normally logged in from. And so it's like, that would be, that's not necessarily saying that they're an insider threat. It's just saying, Hey, somebody may have, you know, co-opted this, this user account and be logging in, you know, maliciously. And like part of the answer to that would be, Hey, you know, user, like person that's associated with this user account, were you in Morocco yesterday? And they're like, yep, you're like, okay, cool. Well, that solves that, right? As opposed to maybe there's not a great technological way to solve that problem in the first place. And you may have to end up going to the person anyway. And so if I have to, if my analyst has to go to the person, you might as well just ask the person to start to start. Yeah, totally. That's, and I'm sure you'll see that all the time, right? You run into organizations that are fixated on specific alerts. You know, I want to automate phishing triage. I want to automate impossible travelers. They move a thousand miles in an hour and one and now. But surely you see all the false positives and kind of the fallacies of those types of detections, even though they're really appealing when you're starting, right? Yeah. One thing I find interesting about automation in general, is kind of what Jared was talking about finding the five things that you want to look into first once in an alert fires, right? Jared and I talked about this for many hours. And since there is an assumption in blind spot thing that happens with detections, right? We talk about that a lot, but what isn't talked about a lot is the blind spots or maybe the assumptions that the analysts go through when looking at an actual alert. And so when you want to automate that triage section, it is super important to actually manually go through that alert, whether it's a whether it is a validation that your team ran or an actual alert. And you want to document those things because you might run into scenarios where you might not thought of a topic ahead. It's okay. Well, I went and I don't know. I checked this file to deceive. There's any access to it or I set a sackel on this and looked at the log to make sure there was anything there, right? And so sometimes that is easily overlooked. And whenever we are doing these automation phases, that's like the first thing that we want to do is say, okay, let's manually walk through this and then document what is the priority to us? And then maybe start to if these attributes do exist, start to change and apply that severity to that specific event and data set and not just necessarily kick it off to the side. But that way we have all those alerts there and ready for us. We have a priority list of which we want to look for as an analyst. Yeah, makes perfect sense. I mean, also for improving alerts, tuning alerts. And the whole work of a really good ADS is paradoxical, right? It's like if you have a really killer noise signal or high signal to learn noise alert and it fires once a year, how do you know it's not broken? Right? So this steps are crucial for being able to go test it and iterate on it, throw it in a thing like detection lab and look at the artifacts. So I'm a huge proponent of that for sure. There's somebody kind of wanted to just circle back to we don't have to rehash the whole thing. I just want to make sure that people listening don't kind of miss what I think is kind of a profound thing there, which is getting feedback from the users. And I don't think people understand exactly how many times they're actually subject to this and their daily lives and how for relatively, relatively low effort from the security themes perspective, they can actually gain like a lot of information. So there's a lot of companies out there that don't feel like they can do these advanced detections and like advanced alerting and stuff like that. But if they spent a couple weeks or a month or a quarter or whatever, it takes them to set up those kind of alerts that you guys have set up like, Hey, like did you do this thing? It takes a lot of the weight off of the security team to figure that out. And I know we kind of dance around the subject, but I just wanted to state it implicitly like every time you log into like feed. or your email, it does this for you. And the reason is because Facebook is not going to dedicate engineers to sit there and look at the logins for every single one of their users, because they can't, because there's too many of them, it would cost too much money. But they're still somewhat responsible for protecting your account on their platform. So a quick and easy way for them to secure your account is to make you do it, right? 'Cause it's on you to say yes or no. So I think a lot of security teams out there that kind of feel like they have too much going on to look at impossible user travel, because honestly it's not the best reward for time spent if you have two security engineers and you're at small to medium sized business. If you can make the user do it somehow, that way the ones that you do get are already validated and you know it's legit to a point 'cause the user said they didn't do it and they could be lying, but chances are they're not gonna be. They've already done the triage and a tiny bit of investigation for you and giving you a head start on that specific incident. So I think that's not a small thing that you guys have achieved at Palantir and I think it's something that could be valuable to a lot of different enterprises if they just spend a little time to set it up. - Yeah, I think just the closing remark is, Ryan Hubert and Slack really pioneered this technique and did some great blog posts about it. I think they deserve all the accolades that they have built. And we rely on user signal for a lot of things fishing. We talk about the users for the first line of defense. This is just sort of extending that something beyond just clicking report fishing, right? And you think about privileged actions, highly sensitive assets. That's super easy to get and cheap signal to rely on. - I remember that. - Yeah, so I was gonna, you made a comment and I was gonna say a little bit about the paradox of a good ADS being high signal low noise. And one of the things that I run into very often is the idea of noise or a high signal detection is associated with a low false positive detection. And I think that that is a relatively dangerous concept when it comes to detection engineering, right? So the idea is, hey, this detection isn't ready for a production because it's high false positive, whatever high fault, like everybody has a different definition for what high false positive means. And so what we're going to do is we're going to kind of filter this until we get it to an acceptable rate of false positives. But one of the problems that I have in general with that approach is that false, I don't know that this is 100% true, but I think it's almost completely true. False positives and false negatives are inversely proportional. So like as I reduce false positives, I'm increasing false negatives. And I think there's this problem that false positives are very verbose. Everybody knows that a false positive occurred and false negatives, they're silent errors, right? And so nobody, you know, you don't know that you have false negatives unless you are really searching them out. And so like what I run into very often is when people are building detections, their main focus is how do I build a detection for X without as little false positives as possible? And I think there's a mistake where it should really be, how do I build a comprehensive, or as comprehensive a detection as possible while reducing false negatives or false positives and false negatives as much as possible? - Yeah, I think an example of that would be, you know, you're looking for weirdness on a Linux server and you run data dog or you run some other host base to elementary agent. And so you say, you know what, anything that has this name or this path are just gonna filter out and that works really well until an adversary puts their binary there, right? - Correct. - Yeah, every time you reduce false positives, you're creating an innovation opportunity. It's probably a fair way to put it. - Yeah, it's kind of like anti-varish exclusion folders, right? It's great for performance, great for a lot of reasons. But enumeratable and abusable. - Yeah. - Yeah, it's a tough problem. And I think, you know, the core of that is, what is the hypothesis of what you're trying to do? Similar, I think, Sackles are a really good example. There are certain ridges that should never be queried, but the moment someone runs seatbelt or they run some sort of post-exploitation script, like those will light up like a Christmas tree. And that is a great place where, you know, you can kind of reduce the two or three processes you would expect and get a really good example. But obviously, that same hypothesis doesn't work for other types of alertness. - Sure. Yeah, so I actually was playing on talking about Sackles, but let me touch on this real quick, because I think this is what John was saying that I was interested in talking about. But I think the way, so I used to be a big, I kind of was a hater on this idea of like, building out threat scores, because I feel like threat scores are often generated in an arbitrary fashion. So it's like the way that somebody determines that something is risky is completely arbitrary and not scientific at all. However, like creating threat scores is a way to manage that false positive, false negative kind of issue, right? So the idea is, let's use, I talked about curb roasting earlier. So it's like, in order to curb roast, what is something that you must do every time you curb roast? Right? And so I think, I think, like, let's just say that the answer to that question, I think there's a couple answers potentially, but an answer to that is you must request a service ticket, right? And so that means that every time somebody requests a service ticket, they are potentially, that event, the request of a service ticket, could potentially be a curb roast, right? Or maybe put a different way, every curb roast is going to be represented by a service ticket request. And that makes me think, okay, well, the base condition for curb roasting. So like if I want to capture every single instance of curb roasting possible with false positives, like start with a broad stroke, right? The thing that I should base my detection on is a curb roast service ticket request, right? So there's event logs for that, for instance. Then what I can do is kind of what Johnny was talking about a triage, so the detection is, hey, a curb roast service ticket was requested. Well, that detection is going to be a very false positive prone, but I put that into my pipeline because if I start filtering things out, I'm potentially filtering out true positives, right? And creating a false negative. And so the next step would be, what's the technical context about that curb roast service ticket request that could be interesting to me? Well, there's like the user that requested the ticket, the account that the ticket's for, the SPN that's associated with it, so what service is this going to be used for? Maybe there's the encryption type of the ticket. There's all kinds of things. So I would want to enumerate all those things and then say, okay, well, what are the scenarios where these differences in the relationship between different contextual contexts or contextual factors causes me to increase the score. So for instance, like, I don't know, it like an easy one would be, if this is RC4 instead of AES-256 encrypted the ticket itself, then that may be something that raises the score. And so then you end up with every ticket receives, every ticket request receives a score because every ticket request is potentially a curb roasting opportunity. But maybe when you, you know, something that is known to be curb roasting is going to get a score of 100, something that is very unlikely to be curb roasting is going to get a score of zero. And then the idea would be, when you're doing in your investigation, you start with the 100s and you kind of work your way down. But the thing that I challenge people on is like, it's probably worthwhile to look at the zeros, right? Some of the, not all the zeros, because obviously there are like, ideally there would be significantly more zeros than there are 100s. But you should occasionally take a peak at the zeros because you might learn something, right? You might learn something that makes a service ticket benign, like the service ticket request benign for curb roasting. Or you might learn, hey, there's something to this attack technique that I wasn't considering when I was creating my score because I found like, I manually detected a malicious curb roast ticket service ticket request. And so that's kind of like, that's my current leaning on how the approach should go. It's like, start with a really broad scope and kind of narrow it in and built in like, you're ultimately creating a hierarchy of things on whether they're the likelihood that they are malicious, I guess, and then creating strategies to look at them differently depending on the score that's achieved. - Yeah, totally. I mean, that's a super valuable way to do it. You know, the challenge I would give you is rather than trying to score every curb roast, you know, every ticket request, why not have some honey pot accounts that should never have ticket requests ever issued for them and that lights up rather easily. So it definitely is just depending on the spirit of it, but I think for really sophisticated stuff, like new binaries observed in the network, right? A new deal has dropped on disk. Like that type of scoring I think could be super valuable for surfacing signal of interesting stuff to dig into. Yeah, I think like the honey pot example, right? So that's a valuable one. So if somebody requests-- if I have a service account that is not actually being used for anything, and somebody requests a service ticket, then that's a good indicator that somebody's doing some sort of automated thing that they probably shouldn't be doing. But I could perform Kerberostein without requesting-- you could do targeted Kerberostein. So you identify a specific account, and you request a ticket for that. And so if you're on like-- and Graham, we talked about layering defenses and all that kind of stuff. But to go hyperbolic, if your only solution is your honey pot account, then that may not be-- you still have an opportunity for false negatives, I guess. And so what I'm trying to address, I guess, is how do you limit that completely, or as much as possible, I suppose? I think the thing that is really important to take away from this is knowing your assumptions and your blind spots when you create that honey pot token. Your first assumption should be, this will only work if they request a ticket for that honey pot. And maybe they're really sophisticated, and they know exactly what they want to get. Maybe they run Kerberostein on everything. But I think something that we don't talk about enough as a defensive industry is the process of introspection. Like, how do we actually capture lessons learned from the detections? What we are seeing-- you talked about looking at the small tail of benign. A lot of people are just biased towards closing stuff. And it's like, do we actually take a step back to ask the question, what does this teach us about our environment or our alert? How do we improve our ADS with this knowledge? How do we carve out the time to do this incremental improvements? Because that's really where you hit that polished point of ADS is where they start to shine. Yeah, there's-- so John Henson's ski over at expel. He has really been-- I think I pronounced his last name correctly, so forgive me if I'm wrong. But he's done a lot of work on quality control and quality assurance. And so just to give listeners an idea, the way that I'm going to define it, and I might not be 100% correct. But quality assurance is, how do I design a process to achieve the desired objective as best as possible? And then quality control is, given the output of that process, how do I ensure that the output matches my desired output? So I build for detection engineering. I have a process by which I follow maybe the ADS framework for building detections and pushing them to production. So there's some rigor in what you document. There's rigor in how you research it. There's rigor in the peer review. There's rigor in validation. That's all built out. And then it's like, OK, well, I also need to do some testing to make sure that it's achieving the objective that I want to achieve. And then there's an implied feedback cycle as well. If you find a situation to where it's not achieving what you desired, then you should fix that, I guess. And I think we're really good in general when we're doing-- so the idea of we were talking about the funnel of fidelity, which is my model for detection and response. And the idea is, as you go left to right, you are reducing the amount of events, I guess, that are going through the funnel. And so you start with collection. You're just collecting much raw telemetry. You go to detection, which is identifying, say, every time somebody requests a Kerberov service ticket, you go to triage and you say, OK, well, I need to start racking stacking these service ticket requests to what's most likely to be malicious and what's most likely to be benign. Then I do investigation for the things that are most likely to be malicious. And then I do remediation on things that are confirmed to be malicious. And I think when we're doing investigation or remediation, we're dealing with a relatively small subset of events. And so we're really good at doing that feedback loop. But I would propose that we should be doing that feedback loop at each of those steps. So is there a feedback loop that's validating that your detection is achieving what you expect it to achieve? Is there a feedback loop after you do triage and you score out all these service ticket requests to say, hey, the service ticket requests that were marked as 0, should they have been 0 after some sort of manual analysis? And the way that John talks about it is you don't look at every 0. Otherwise, there would be no value in building out that scoring mechanism, right? But what you should do is you should pick out a representative sample set of 0s and look at them in a manual fashion and say, hey, based on my manual analysis, did this achieve the, what's the score, the proper score based on manual analysis? And I think this whole thing is assuming that your manual analysis is as good or better than your automated analysis. And I think that's probably a safe assumption, but I'm not 100% sure about that. But I generally think that each one of those phases, collection, detection, triage, investigation, and remediation should have some sort of quality control or quality assurance, quality control, and feedback aspect to them. And I think we're only really doing that with investigation and remediation in most cases. Yeah, that's a great point. And have you seen this executed in action? Are there enterprises that are taking it further left in their introspection? Yeah, I think there's-- so John gave a whole presentation, I think, a TAC con, talking about how expel does that. And so for those that don't know, expel is a managed-- I don't know if they want me to say this. Or I don't know if this is what they would consider themselves, but from my naive, not really having worked with them at Dunn, they're a high-end MSSP. So they will deploy sensors to your environment and then monitor them and report on incidents and things like that. That's my understanding. I think they're, again, a higher end version of that. And so you would assume that they have a high number of events that they're dealing with on a daily basis. And so the way that he discussed it at TAC con is that they implement that in their pipeline. So basically, just because something didn't meet the criteria that says this is 100% malicious, doesn't mean that we shouldn't be looking at it. But we can't look at everything. And so we need to-- the way that he looked at it was from the manufacturing industry, there was actually like a ISO standard or some standard that he found that was like, this is how you-- it's like a six sigma type thing. So how do you identify what a representative sample set was? So I have 10,000 events. How many events should I be responsible for looking at? Maybe I look at 50 of the 10,000. And so I take those 50 and I manually analyze those. But if the score is like 100, I probably should look at 100% of those within a short period of time. But if the score is 0, maybe I look at 1%, or I don't know how you arrive at the percentages, but looking at some zeros is better than looking at no zeros. Yeah, makes sense. One thing I wanted to touch on that you mentioned day in that I thought was super valuable was how it's often overlooked that after detection or alert fires, what can that teach us about our environment? And maybe going back with the detection, either creating a new detection or updating the current detection that we have at that time to after we better understand our organization. I'm curious if-- and not that this should or would take away this process whatsoever, because I think that's super valuable. But it seems like that-- I feel like that step could be taken while the detection is being created as well. So this is somewhat of-- behind the idea, I believe, and Jared correct me if I'm wrong, somewhat behind the idea of the abstraction maps. And so understanding the environment that we have with us and then understanding the technique that we have in diving deep as we can and trying to find a pivot point. We talk about finding pivot points, finding the place that the attacker has to perform this behavior. And that's something we can leverage as detection engineers. But I'm kind of curious in your thoughts, Dan and Jared too, and Luke as well, is at what point when diving deep into these abstractions or diving into these technologies is the juice not worth the squeeze anymore. So at what point is it just we're spending way too much time diving in deep into this? Is it at the point that there isn't any notable telemetry that we have or any notable tooling that we have to give us insight into that? Once we get to a certain level, or what? I'm curious you guys have thoughts on that. That's asking a hard hit or-- Hey, man, I got to do it right, man. Well, like the cop-out answer is, I think there's obviously an ROI, like return on investment, kind of calculation that you have to do, and like a risk tolerance evaluation. So there is inherently risk in spending time developing detections or producing alerts that people have to look at. So I think that's the argument for false positives. It's like false positives have risk because now you're triaging alerts that are not actually bad, you know, bad activity. And so you're not able to spend the time looking at other alerts that are potentially more likely to be bad. Right? So that's kind of like one of the main arguments against false positives. However, then there's the false negative risk, which is what happens if bad stuff happens, and I don't detect it. Now I think that there's like-- we talk about defense in depth, right? And so I think everybody kind of understands that from like a network security kind of perspective. But there's also this idea. of detection and depth, which is kind of what Dane was talking about with the minefield. So like, we don't necessarily have to comprehend. I talk about comprehensively detecting curb roasting, right? And you know, I'm talking about that kind of in a vacuum, but attacks don't happen in a vacuum, right? So like, you don't just get, you don't just magically appear at a point to where you're ready to curb roast without having done something before curb roasting. And then like you curb roast because you want to get credentials to do something else, right? Sometimes you're going to laterally move somewhere using those credentials and all that kind of stuff. And so, I think, I think you just have to make an evaluation of like how far can I go and technically understand this? And then what's the most likely next step, right? So if the next step of curb roasting is some sort of lateral movement, it's like how comfortable am I with my lateral movement detections? And so like the idea is, I don't have to catch everything they do. I just have to catch something they do. And then I have to trust that my investigation process is going to work the way that I hope it works, right? Which that may be a big assumption as well. Yeah. It's interesting. You know, I think if you're doing the alert development process, and I agree with all this points, right? It's, you know, if you can't get acts, maybe you can get the predecessor to it or the sequence after it. A good ADS development hypothesis means you have some flexibility to explore an idea that might not work out. And so for instance, if you're focused on, oh, you know, in this SolarWinds intrusion, we saw the actor abuse, WScript to do X, Y and Z. I'm going to go look at all the WScript on my network. Maybe it's not something you can work on because of high false positive fatigue or, you know, other reasons. But I think there's a couple of outcomes from that. One, really when you're developing the detection, you should be performing a hunt against your network for a lengthy period of time, right? Both the enumerate false positives also to validate the assumption. And then as you're implementing it or you're thinking about implementing it, I really think detection engineers should spend more time considering eliminating areas of abuse entirely, right? And so if you're looking at WScript and your network and you're investigating the indicators of it, you're kind of time-bounding that to a few days of work. Maybe the answer to your rive is actually we just need to kill this. You know, we need to implement these controls to prevent WScript abuse and just obviate the need for detections at this one specific area. And I don't think a lot of detection engineers really focus on like taking the insights that they have and implementing or partnering with other teams to implement proactive security controls, but perhaps they might be the folks who are best calibrated for it. I think there's, I was just, I'm spit-balling right now, but I think, I think a lot of times there's struggles from like the detection engineering team or just the sock in general with leveraging threat intelligence in a like a practical operational way. And I think like one of the things that may be valuable for instances, let's look at all the reported uses of Kerbero-ste. I'm just sticking with Kerbero-ste because whatever we already talked about it a little bit, but let's look at all the known use cases of Kerbero-ste. And then let's see if we could derive what the most frequently used follow-up techniques are. So there's a chain right of techniques that are being used. So it's like when people use Kerbero-ste and 90% of the time, the next thing they do is WMI lateral movement. It's like, okay, cool. Well, you know, I'm going to try to detect Kerbero-ste and like, there's tons of ways that you can do that. But like for instance, Wilshroder released a tool called Rubius, which bypassed some of the like more common Kerbero-steen detections. And so it's like, okay, well, that bypassed that. That's okay because I know that the next thing they're going to do is WMI lateral movement or use PSExec or what it like, there's probably some small set of techniques that are being used after Kerbero-steen. And so if I'm confident that I have an 80% solution for Kerbero-steen and an 80% solution for WMI lateral movement, then I probably, probably pretty good at that point. And then like, you know, we can make the managerial risk decision of like, how comfortable are we with that coverage, right? But I think an important aspect is being able to quantify that coverage in general, which is where that kind of like idea of the abstraction map comes in, which is, hey, you have tons of different tools that can achieve this objective. Let's understand how those tools are like the similarities between those tools and the differences between those tools. So like, like I said, you have like invoked Kerberos, which is a PowerShell tool that does Kerbero-steen. Well, that goes through like.NET and the API calls and all kinds of stuff. And then it makes a Kerberos TGS wreck, right? Well, then you also have Rubius, which skips PowerShell, skips the API calls and literally just handcrafts the TGS wreck. And so, you know, like being able to understand and explain that would help you quantify kind of what the coverage of your detection approach would be. And like, I think it's, this may not be true, I don't know. But I think generally it's easier to quantify the problem than or like to just like say what the problem is that it is to solve the problem with a detection, with a high signal detection, I guess. Yeah, totally. And all these circumstances, I think you're far better off as a defensive team than relying on an alert from ATA just saying, hey, we saw Kerbero-steen. You know, like, I don't know what to do with this. Or why it's firing. Yeah. It's interesting. Oh, go ahead, Jert. Yeah. So going back to the MCAS thing, Luke and I are literally working through a project, which is like, we have every Microsoft Azure solution that we possibly can have, like, Defender for Endpoint, for Identity for Office, for, you know, all this stuff. So Defender for Identity is ATA. That's, I guess, the latest and greatest name or whatever. But the idea is, like, hey, it has all these out of the box alerts. And like, we need to have a way to, you know, either automated, like, close these in an automated fashion or like, you know, provide people with basically an ADS that describes or a runbook that says, when this alert happens, these are the things that you should be doing. And it's like, once we clear those, like, because Microsoft's like, you know, not trying to be mean. But some of the times, some of the events or alerts, it'll say, like, review the logs. It's like, these are your response actions. You're like, oh, cool. Response actions. You're going to tell me what to do. And it's like, you know, review the logs. And it's like, okay, cool. Thanks. Yeah. Yeah. Yeah. Yeah. Very helpful. Thank you. Yeah. And that's, you know, if you have the team, peer review helps do that, right? It's like, you know, draw the rest of the owl. It's like a few responses. Respond to badness. It's like you're just setting yourself up for a 2 a.m. heartache. That's great. Yeah. Yeah. One thing that's interesting is you guys are talking about maybe in a detection opportunity for, um, overlying, overlaying the gap of the previous detection is understanding what might be next in that attack chain, right? So, um, I'm, it's interesting because this goes back to an idea that I had a while back. I think actually someone at Microsoft might have actually released a blog about this in terms of markup chains where it's basically where if you see an alert happen for a specific technique, what are the odds or what is the, um, what's the word I'm looking for? What's the percentage that they're going to move to X technique or Y technique, right? And see that really comes down to whether that's privilege based things or whatever it might be, but it's interesting to start to follow that attack chain in terms of detection because maybe that might be a gap coverage opportunity for the, the previous detection that you might have implemented. Yeah. Yeah, 100%. And, you know, maybe it's just me, but I, I, my, my goal when I'm writing detections, um, oftentimes is, is to find a way to make it so I don't need the detection. And like, that's one of the driving forces behind why we really invested a lot of time in the Windows Firewall project. It's like, you can spend a ton of different time and energy building different ADSs for WMI, PowerShell, remoting, um, Decom exploitation, like a lot of remitment. And it's like, if you just manage your endpoint firewalls really, really well, it's like this opportunity ceases to exist. Yeah. Yeah. Yeah. Yeah. Yeah. Workstations. Oh, man. Okay. So Andy Robbins would be your best friend right now. I might be like still in his thunder a little bit, but whatever. Yeah. So, uh, so Andy heard that the podcast was called detection, challenging paradigms. And he's like, what if we challenge the paradigm of detection itself? And it's like, okay, where are you going with this Andy? And he's like, well, yeah, like what if we just make it to where they can't do the bad things? Uh, I'm, I'm like hyperbolicing what he said a lot to make it make it like, you know, crazy. Other high we'll have him on this podcast to discuss his point to discuss this. Yeah. Yeah. So, but like, but his basic point was like, if you could do prevention well, then like, you don't need to do detection. And so I think the thing that we often forget. And I think this is like the crux of his point is detection is a stop gap to what should be prevented, right? And so there's a few reasons why it prevention isn't comprehensive, right? Prevention can't solve all the problems because maybe there's some operational need that's blocking prevention, right? Maybe you don't have like the political capital to get a prevention put in place because it's going to stop somebody else from doing whatever. Or maybe there's just technologically not a way to implement the prevention in the first that the five you probably want, it's probably valuable to have a detective control to see if that prevention doesn't work in what or fails in whatever case. But yeah, I think, I think sometimes we have more control over detective controls and so we forget, like I think you're in a unique situation where you have a lot of control over preventative controls. Like if you, like my experience with you is like if you think of something cool, like you have a lot of control over being able to implement that. And I don't think that that's necessarily true of a lot of situations to where it's like, yeah, I could have a great idea, but they're like, yeah, too bad we're not going to do that. And so then it's like, okay, well, I'm just going to stick to my detection lane and not not even worry about that because I tried to fight that fight and it didn't work. And so I'm not going to go that direction anymore. Yeah, and I think that's a fair flag. And you know, the real question is, is pushing on prevention, you want to talk about challenging paradigms, is that wearing helplessness? And at what point do we rely on maybe the wrong tool versus fighting the political battle? And that's an organizational problem. It's a company problem and it varies. But you know, I think detection engineers just broadly undersell their ability to influence change through their insights. Right? It's like you're not just detecting badness. You can help inform business risk. You can help inform preventative controls. Maybe you can't change everything. I think there is more value to be added than just awarding on adversary activity. Yeah, we're not so good at like stepping outside the technical component or technical like concepts and going into like the business, like looking at it from a business perspective. I think that's like just traditional of InfoSec people in general. That's kind of like the stereotype of an InfoSec nerd kind of thing that you're, you know, you think about all the technical parts, but you're like, I don't really care about the business aspect of it. Yeah. I look forward to the podcast to do it and that challenging that paradigm. Oh, boy. Yeah. That'll be fun. Yeah. All right. So one thing that I wanted to, we went down a rabbit hole, but I said that I wanted to talk about some sack of stuff because you kind of brought that up. But I think one of the things that you've, you've done a really good job at is or specifically in my kind of your blogging and your open source projects are, how do we, how do we enable collection for organizations, right? And so like you, you've done a lot of stuff with Windows Event Forwarding. You've talked about leveraging Sackles. Can you talk a little bit about kind of those, those projects, specifically like I, I'm really interested in, I tried to set up Windows Event Forwarding in a lab in like in the class that I took from you in 2015. And it was the biggest pain in the butt ever. And then you're like, Oh, no, it's, you know, here's how you do it. And you know, you just got to set up this crazy infrastructure. Yeah. There's a lot of good stuff. Yeah. Yeah. Yeah. I mean, like I think all of the stuff I've done has been out of my own interest and almost certainly has been standing on the shoulders of giants in almost every way like Jessica Payne at Microsoft is a huge proponent of the Windows Firewall. She did the talk at Tiga X New Zealand. They're really pioneered it as like a concrete defensive thing. And then same thing with with Waff, Windows Event Forwarding and Sackle stuff. But there are a lot of really interesting primitives in the Windows ecosystem that are just often overlooked. And my obsession right now is Sackles. Can you tell us what kind of talk about what Sackles are real quick? Yeah. So, yeah, basically you can you can tag different resources, files, registry keys, etc. With an audit, you know, access control list and you can have it generate an event log entry. And so it gets really interesting for many reasons. But one of the privitives I'm obsessed with are file reads or registry key reads, which is something that most EDR do not surface raw telemetry for. So the example we talked about earlier is I want enough any process other than Chrome ever opens up the Chrome, you know, SQL light containing cookies. Why would I have a process do that unless it wanted to extract cookies? Really tough to figure out how to get an EDR to do that. But Sackles, you can kind of go through tag this files and create an event log entry anytime there's a file open event. And yeah, go ahead. Oh, I was just going to say, yeah, so like we're all familiar with Dackels, right? So that's the discretionary access control list, which basically says if you've ever tried to open a file and it says, hey, you don't have access to this file, that's the Dackel kind of acting there. But then the thing that we don't know about because kind of by default it's not being used are Sackles, which is kind of this thing that allows you to say for this specific security object or resource file service name pipe process, whatever, whatever it may be, I want to know if a certain action is taken against it, which is, and it's extremely granular, right? So you could like get down into this idea of like the very specific file, like the SQL light database that you're talking about. And you can say like, like you mentioned, you can choose the type of access that somebody wants. Say I only want to know if somebody reads this or maybe I want to know if somebody reads or writes to this. And then it's, you could even go as far as saying, I'm interested only if this user does this or like, you know, this group or whatever. And so it's extremely powerful from that perspective, I think. Yeah, absolutely. And a much better description of what a Sackle is. Oh my gosh. But the detection permittives are really robust. And I look at a tool like Seatbelt and it does a ton of WMI enumeration, registry key enumeration. And those are like really, like especially registry keys are really great objects to put these Sackles on. Like you can do it with PowerShell and the moment something queries for what PowerShell transcription settings are set, which never really happens on a legitimate machine. You've got immediate telemetry that's worth digging into. And then the caveat to that is, you know, if you're using Windows Event Forwarding, it's really easy to get those native event logs off the box and into your SIM into a detection. Yeah. And for those that are listening that aren't familiar, Seatbelt is like a set of tools that are used for situational awareness. And so when somebody accesses a machine or like when an attacker accesses a machine, they want to know kind of like, you know, are there, is there antivirus? Are there EDR products? What are the different security settings in place? And so they'll run a tool and one open source tools case, Seatbelt is an example of that to where they'll query a bunch of these things. And so Dane's kind of perspective, I think, is, I mean, that's getting them at the very beginning of their access, right? So that's like early in their access. And so if you could capture, if you could alert on that, that may be an A, it's high signal, low noise, I think, is like in some cases in this like the PowerShell transcript perspective, but also it's really early in the attack chain. And so you have a lot of time or maybe, maybe not a lot of time, but more time to kind of react to that, I think. Yeah. Yeah. All right. Now, like one of the things that's interesting to me about Sackles is you can set them on all kinds of things kind of like we talked about, right, files, registry keys, processes. You could set them on threads. You could do all kinds of things. In my experience, I think there's certain, security objects are better for Sackles than others, because I guess it's like how ephemeral or not ephemeral is the, is the security object. So like a process, for instance, is ephemeral. And so like you may have, I don't know, let's say you want to know if somebody opens a handle. So this is built into the operating system now, but let's say it wasn't. So if you want to know, did somebody open a handle to LSAS like a read handle to the LSAS process? You can set a Sackle for that. But the problem is, is that the next time you boot up the box, you'll have an LSAS process, but it's a different object, right? And so that's an ephemeral or like a, you know, it doesn't, doesn't survive past reboot, I guess. And so setting your Sackle, you'd have to come up with some like startup script that would set the Sackle every single time or something like that. And that, I don't know how well that would work, for instance. But something like a registry key or a file, those things are permanent until somebody deletes it, I guess, but they're permanent, right? And so that's a really good place. And the Sackle is like built into the actual object itself, or I guess it depends on what the, what the object is. But like for a file, it's actually built into the like master file table entry for that, for that actual file. And so it's going to be there as long as that file exists. Yeah. I think you're spot on. You know, there are a lot of things you can do with that. For instance, you could run a PowerShell script, you could find document files, PDF files, you know, various things that could be sensitive and actual property, Sackle it. And now suddenly you can see file reads for those things. And you know, if you have to suppress some stuff that you would expect, but the moment you see Chrome or RAR or, you know, whatever, like you have some interesting telemetry to build on, is it expiltration, is it collection, et cetera. I think that area is just a TLP data loss prevention tries to tackle this problem. But I think there's a lot of really cool stuff you can do with Sackles in that instance. Yeah, really. Yeah, I guess, oh, go ahead, Jerry. Sorry. No, last last thing I have. So I think the reason why EDR might not be doing it for like reads, for example, is because those just happen all the time. And so like EDR is typically capturing like very broadly. And so it's like, I don't want to know about specific file creates. I want to know about every file create. And so then it's like, I don't want to know about specific file reads. I want to know about every file reads. So now you've just like completely indicted yourself. I guess the value of the sackle is that you have the ability to be very granular and say, "I don't care about all file reads, I care about this file reads," so you're flipping it on its head from that perspective, I guess. I think that's why they wouldn't have that type of information. Yeah, sorry, Johnny, you were not. Yeah, one thing that I really like about sackles from a data perspective is the ability to set the sackle on specific registry type of activity. I think the registry in general just gets so overly looked. Everything touches the registry, right? I think there is an extreme benefit from maybe setting a frequency of different, whether it's registry-create, set values, etc. within your current environment, within the registry, and then also identifying type of registry keys that shouldn't ever be touched unless an attacker is really wanting to talk to it. We talk about registry in terms of service creation quite a bit, but service creation happens a lot. I also think at least from a data perspective, when you start to monitor that registry, there's different things you do, whether it's rejects or whatever to really start to fine-tune, whether the new key was created or not. But sackles, in my opinion, are just really nice when you want to look at a specific key or a specific type of path in general that you don't think should ever be modified or anything of that instance. But yeah, that's one reason I really like sackles. Yeah, just kind of put the lid on the sackle. I think one of the really cool things that maybe is not as relevant now, but is a lot of EDR simply would record that like, "Hey, there was an update to this registry key or there was creation or deletion, it wouldn't capture values." The run key was changed. So, okay, well, now I have to go run a script to figure out what changed or query it somehow. Sackle actually shows you before and after values, which is pretty cool for catching things like persistence. So, one thing that has kind of come up a lot throughout the podcast here that we might want to touch on a bit is we've mentioned several projects that Palantir and you specifically have helped with that are open source projects. And that's not, I guess, it's becoming more common, I feel like, but it's not an extremely common practice in security. I know our company SpectroOps is pretty big on it. I know Palantir has a pretty big commitment to making sure that things that obviously not that are proprietary information, but things, especially security wise, that could help another company. You guys make sure that you get them out there for other people to see. Is that a passion of yours as well? Like getting stuff out there open source has Palantir been pretty good at enabling that. What's your open source life look like? Yeah. As long as this sounds like it's actually part of our mission statement, right? And the InfoSectimate Palantir's mission statement is, A, keep Palantir's secure, B, protect our customers, C, make the world safer against adversaries. And so, it is a core philosophy that we being in the position that we are with the resources we are very privileged and to the extent that we can for operational security and just time reasons, try to get as much security knowledge tooling, et cetera out there. And it's a passion that everyone on our team really carries. And I would love to do more. And we're actually committing to do more this year, but quite a lot. It's just the reality of working for our tech company. You don't have as much time as you want. But one of the projects that I'm really excited about that hopefully will be out in the next couple of weeks is an extension for Chrome Firefox, an edge called FishCatch. And this was created many moons ago. We actually rewrote it last year with some of the folks from Spectorops. And it basically is an extension that looks for credential reuse. So it kind of learns what your domain account is. You don't have a config that you can manage through GPO or like manage to your central management software. It learns what domains are corporate domains and it will do character by character password hashing to identify like half credentials been given to a domain that wasn't or corporate domain. And from that happens, it creates an event that notifies the user sense of tip so the infosect team could do something about it. But the idea is like, you know, everyone's working from home in COVID. Either on split tunnel VPN or TLS to crypt in their traffic or you just don't have the capability, but you're really worried about fishing. And so it's like, well, let's build something through it in the browser. And a lot of people could use this. So let's open source it. I think that's a really good stance for obviously it's not an option, I guess, to some places, but where you can make sure you kind of push that back to the community because I know I personally got no help from other people in the community like Bloodhound Slack or Twitter who have seen the same problem that I haven't helped me out. So I think even at the company level, it's good that companies like Palantir who have those resources and are solving, you know, the hard problems of the world making sure we put it back out there for those who can't. Yeah, we're fighting the same adversary. And so, you know, it's, I think there is some sort of moral imperative for defenders to consider. And if they can't push for it, open sourcing and contributing back to the community because the adversary has a lot of advantages and they target everyone, right? It's not just Palantir, it's everyone. I think there's this kind of just to close it up. I think there's this, when I was in the military, there was kind of this one end doubt overclassify the information, right? And so it's like kind of now what I see a lot of is like when in doubt, make, like call it proprietary. And so I think the challenge that people could take away from this if they've listed this long is when you're considering whether or not you should share something with the community, like really, really challenge yourself to say, why is this proprietary? Right? So like, yeah, maybe you don't want to share your specific detection query that you're using or the preventative control that like the list of preventative controls that you've installed in your environment, but like teaching people how to use saccles is probably not proprietary considering that, you know, it's built into the Microsoft operating system. Like the Windows operating system. And so it's like, you know, how can we, how can we make sure that we are not basically overclassifying the, the things that we're doing so that we can make sure that we're sharing as broadly as possible? Well, keeping, you know, keeping an eye idea of like, or trying to like make sure that we don't affect our security posture by oversharing, right? But like the answer to not oversharing shouldn't be, don't share it all. Right? So it's, you know, kind of take a hard look at that and kind of come up with a policy of, you know, this is what should be considered proprietary. And this is what shouldn't. And that could help, that could help your, your team or you individually kind of like share things out with the community and make everybody, you know, better. So with that, yeah, we just want to thank you, Dane, for your time. It's been really great chatting with you. We had, so we have a lot more that we want to chat with you about. So we may have to bring you back on at some other time. We just kind of kept kept flowing on stuff. So appreciate your time. I'll pass it over to Luke to kind of talk about the logistics of this episode. And then we'll go from there. Yeah, definitely. Cool. So you can find most of the main updates on Twitter, which is going to be at DCP, the podcast. And then if you want to know where you can listen to us or you want to say for available on your, your platform, we probably are DCPpodcast.com has all that information for you, all the links you could possibly need. And yeah, just thanks again for your time, Dane. We appreciate the kind of stuff you've stuff and buy. We'll be looking out for this.

Podcast Summary

Key Points:

  1. Dane Stucky discusses the Alerting Detection Strategy (ADS) framework, which emphasizes documenting detection hypotheses, blind spots, and response steps to improve quality and scalability.
  2. Pre-built detections from vendors offer immediate risk reduction and stability, especially for organizations with no existing coverage, but they limit extensibility and long-term growth compared to raw telemetry.
  3. Detection validation through CI/CD pipelines, using unit tests and atomic tests, helps ensure detections work as intended and prevents drift over time.
  4. Automation and enrichment (e.g., using SOAR tools) are crucial for scaling response efforts, reducing time to response, and filling gaps in centralized telemetry.
  5. The balance between raw telemetry and synthesized alerts depends on organizational resources and skill levels; raw telemetry provides a higher ceiling for customization, while vendor alerts offer a higher floor for immediate security.

Summary:

In this episode, Dane Stucky discusses key aspects of detection engineering, focusing on the ADS framework he developed at Palantir, which standardizes detection documentation by requiring hypotheses, blind spots, and response steps. This approach improves detection quality, facilitates peer review, and scales team efforts. Stucky emphasizes the value of pre-built vendor detections for organizations with no existing coverage, as they provide immediate risk reduction and stability.

However, he warns against relying solely on them, as they limit extensibility and long-term growth compared to raw telemetry, which allows for custom, threat-focused detections. He advocates for detection validation through CI/CD pipelines, using unit tests and atomic tests to ensure detections remain effective over time. Automation and enrichment, such as using SOAR tools to triage and enrich alerts, are critical for scaling response and reducing time to action, especially when centralized telemetry is incomplete.

Stucky also discusses the trade-offs between raw telemetry and synthesized alerts, noting that the choice depends on organizational resources and skill sets. Overall, the conversation highlights the importance of balancing immediate security needs with long-term extensibility, leveraging frameworks like ADS to document and validate detections, and using automation to enhance response efficiency. The episode underscores that effective detection engineering requires continuous improvement, collaboration, and a strategic approach to tooling and processes.

FAQs

The ADS framework is a method from Palantir for documenting detections, focusing on hypothesis creation, outlining response steps, and identifying blind spots to improve detection quality and team scalability.

It requires testing the hypothesis to prove the detection works, incorporating elements of hunting, and mandating peer review to ensure quality and cross-training before production deployment.

Pre-built detections provide immediate risk reduction and stability, especially for organizations with zero coverage, buying time to develop custom telemetry-rich detections.

Vendor detections often lack transparency on how conclusions are reached, limiting extensibility and the ability to customize or grow detection capabilities beyond what is provided.

Raw telemetry allows teams to build threat-focused, technique-focused detections tailored to their environment, offering extensibility, while derived alerts have a high floor but low ceiling for customization.

Engineers should use frameworks like ADS to define clear hypotheses and response steps, balancing depth with practical validation to prevent getting lost in unnecessary rabbit holes.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.