The podcast discusses the major CrowdStrike outage, which began when a faulty automated update caused widespread system failures, impacting airlines, hospitals, and other essential services globally. Recovery efforts were largely manual, involving technicians physically accessing devices to remove problematic files. The conversation explores the difficulty of preventing such incidents, given the reliance on security providers with kernel-level access and the interconnected nature of modern software ecosystems. Key takeaways include the need for better contingency planning, such as having clear recovery checklists and contact protocols, and the importance of staged software rollouts to mitigate risks. Additionally, the hosts note that cybercriminals quickly capitalized on the chaos with phishing campaigns, urging organizations to verify sources during crises. While improvements like enhanced testing are possible, systemic vulnerabilities and human error remain inherent challenges in centralized digital infrastructures.
[Music] Welcome to the Decifer Podcast. My name is Dennis Fisher. I'm here with YouTube star John Hammond to also do some security research on the side. John, thanks for being here, man. Hey, Dennis. Thank you so much for letting me join you. Oh, please. Yeah. Happy to have you. So, I wanted to talk to you, have you on today to talk about the CrowdStrike outage, which we've got a little bit of distance from right now. It was about 10 days ago as we're recording this. What were your initial thoughts that Friday when this all started happening and you started seeing what people were saying on Twitter and elsewhere and being like, "Okay, it looks like the CrowdStrike and update caused this giant outage." But then quickly, there were also a lot of like wacky theories coming out. What were your initial thoughts when you looked at it and then sort of figured out what was happening? Oh, my goodness. I know quite a can of worms, but I'll admit, I'm over on the West Coast in Pacific time. So, I saw some of the chatter and, you know, sort of some, the start of the fire on really Thursday night would have been 10 p.m. Pacific time when I saw just the Reddit thread start to grow with the floodgates open. People chime in, "Hey, we're seeing blue screens of death everywhere affecting our endpoints, workstation servers." And at that moment, I'm realizing this is just going to set the world on fire. Without a doubt, we've seen airplanes, of course, airlines with their downtime, of course, schools, banks, hospitals, even. I know it was everywhere. But I had that realization that, "Okay, everyone's going to be chatting about this and we need to start to fight fires the best we could." But now, as you mentioned, hey, we're a little bit distant. We're seeing the recovery. And I think, I'll admit, I guess I don't know if everyone's back in action yet. I've seen a lot of the airlines recover, had some travel this past couple of days. But so far, now we're just thinking, looking back, like, what are those lessons learned? Where do we go from here? That was the next question I was going to ask you is, you know, these outages happen a lot of times they're isolated. It might be a smaller piece of software or it might be an update that only got rolled out to a small percentage of the company's customers or something like that. And it doesn't draw this kind of attention because it doesn't have the cascading effects that this one had. Aside from, there's a couple of ways we can take this, I guess, sort of the practitioner side. If you're someone who had to deal with this, what do you think? I mean, I know you guys deal with the MSSP world. What are some of the lessons that you saw folks taking away from a really widespread severe outage like this? Well, if I may start, I guess, with the bad news and everyone knows how you, it was a manual effort to recover. It was every technician, every engineer kind of walking around and hand jamming whatever bitlock or recovery key if you need it or just getting into safe mode to try to remove some of those channel files that were real rooted the problem. And that, I'm sure, sucked without a doubt. So when folks tend to ask, hey, what could we have done to prevent this? I don't know, is there anything we can look towards in the future? I don't, I struggle with this one because it's really hard. It's not something that you could prevent all that easily. It took us by surprise. Normally, you have the conversations of network administrators and system owners saying, well, we weigh the pros and cons of having these automatic updates on for the new patch or hot fix comes out. And that's up to their own decision and their discretion. But this one coming from the cyber security provider that was completely different direction. You didn't have control over it. So a bit of a blind side, but when you're left now, okay, what can we do? What I have seen in these, this is the good news. This is the silver lining. If I may, a lot of folks are chiming in, well, you know, we took this as an opportunity to improve some of our planning, some of our just strategic planning is to look when we have a nightmare scenario like this, who are we going to call? Do we have the numbers? Do we have the contact details for everyone? Do we have a checklist? Do we have a document or procedure for how we get back in action? What are we thinking about backups? Are we thinking about recovery, etc? So those, I know, I always feel very, very fluffy, but that's the best that we can do when we're trying to fight fires like that. It's true. And I think, you know, on the enterprise side, if you're a large enterprise that has a dedicated security team, maybe you have internal IR folks, people like that, you probably had some kind of playbook you knew what to do, even if it took some manual, you know, let's get some USB drives and start running around, you know, recovering these devices. But if you're a small school district or you're, you know, a regional auto parts company or something like that, they got hit with this and your, all your stuff is outsourced, you might not even have anybody to call that is going to help you right away, you know. So that's those are the folks that end up being in a really tough situation with something like this. Without a doubt. And if I may, I know we were kind of bantering before we started to record is to just how central this is to a lot of our world, especially when we think about it for cyber security, because it's no secret. I work for a vendor. I work for a fellow cyber security provider. And, you know, for a lot of the access, for a lot of the permissions and privileges that we need and anyone and any XDR, EDR, whatever buzzword we'd like to throw in there, a lot of that means you've got to be working at the kernel, but at the low level, the root of the operating system. So you can hook these API calls and functions. And you can see the signals and telemetry that really help you make better informed decision to stop actual malware and bad hackers and elicit threat actors. But that's fragile. That's very sensitive. And obviously we've seen that topple over and choke quite a scale and severity. But that is when we're starting to have these conversations now. And I'd love to riff with on this. What does it mean to be in the kernel? Should we be there? Are Microsoft and the other architects of the operating systems we're using? Is there another direction we can go and how would that affect the cyber security vendors out in the industry? I was going to ask you this because I saw some speculation last week that Microsoft might start restricting even further who has that kernel level access for drivers because of I'm sure it's something they talk about all the time. Because I know they're very, you know, careful with who, what products actually get that level of access. And I wonder if we're going to start seeing them shrinking that pool of software and vendors that get that kind of access as either an indirect or direct result of this. Well, I have seen, I feel like both ends of the spectrum. All of men, I feel like I've seen folks saying, hey, we could take that sort of Apple walled garden approach. All of us often locked down. And maybe that is a fine approach. I guess time will tell. I've also seen folks that kind of look at that and scoff shrug it off, seeing the conversations and the speculation, but say, no shot. That's not going to happen. We're too far gone in the kernel world and it's just a necessity. And I will admit, and this is just a John opinion. This is just me personally. I think I'd lean in that direction. I am hesitant and I don't think we'll see too much of the change yet. Maybe we could think on some of those nerdy technical implementations like EBPF or some other opportunities where we work in the mix. But I don't know where we'll find a middle ground or if we will. I think you might be right because the way that I mean, Apple started that way. It's not a change that they made. You know, when the Microsoft trying to roll that back at this point, I don't see how that would work in practice. It seems, I think you're right, John, maybe too far down the road. And I don't mean to be a pessimist there. I'm not trying to sound all glad. It's just real. But, yeah, well, we'll do the very best that we can and we'll be cognizant of the holes, gaps and trying to board up the doors and lock the windows here or whatever we say. Right. But I, no, I tend to fall on your side of the fence on that too. And it's, you know, I don't know if it's good or bad. It's just where we are right now because Microsoft's, you know, I don't think they can make a large change like that. At least not any time in the near future. It's not a switch that they can just flip. They're like, sorry, guys, Colonel's close. You guys are out here. You don't have to go home, but you can't stay here. And one of the other things that this obviously sort of brought to mind for me was just the fragility and interconnectedness of the networks and software ecosystem that we have right now. You know, yes, there's millions of software apps out there and people use all kinds of different things. But in reality, there's a handful of large providers that everyone relies on or, you know, the vast majority of companies rely on in one way or another. And when something like this happens and it causes a cascading failure and large, you know, airlines hospitals, 911 systems get affected, that's when you start to see other people that aren't immersed in this sort of stick their head up and be like, oh my god, are we all just reliant on it like is this it? Like don't we have some checks and balances here? And the reality is we kind of don't. And I wonder if there's anything that we can actually do about that at this point. We might be too far down that road as well. Yeah, goodness, man, you're making me sound like a doom and gloom fella here. Not so much fear on certain things.
the end out and all, but we go back and forth on a look. Do we want to decentralize solutions? Should we have a centralized architecture? But it's no secret that they are going to be-- I don't want to say market leaders, but that might be the best way to put it. The folks that shape and cultivate what the industry looks like, Microsoft, Windows, that's just the prime candidate here. And I don't know of an alternative just because I haven't seen that reality yet. Maybe there is. Maybe there will be. I'm not sure. I can't say with any certainty. But I think inevitably, that's just kind of going to happen. And we'd hope. Again, we're all very hopeful in thinking, OK, they'll do that right. They'll take that lead and take that responsibility. And everyone else fighting in the trenches, cyber security providers amongst us will do the very same. But if I may, that's going to be run by people. That's going to be run by us. Human beings. And we're going to be fallible. And we're going to make mistakes. The same way we'll click on a stupid fishing email. We can see mistakes that are small, tiny boo-bos and some that make for a bigger blast radius. Unfortunately, we'll have to take it as it comes. But if I may, I think I'd be the first to say, this is not a poo poo on CrowdStrike. This could have happened to any provider just like this. But we'll do the very best we can to work through it. Yeah. Where is written by humans? Humans make mistakes. And these kind of things are going to happen. We haven't seen one on this scale in a very long time. But yeah, I mean, outages and mistakes happen, obviously. But these are the ones that grab everybody's attention and get people in Congress and elsewhere. You see it on entertainment tonight and things like that. And you're just like, oh my god, what is going on in this world right now? But maybe a couple of the things I wanted to ask you about is one of the things that CrowdStrike said after this all came out and sort of their post-mortem was, they did some testing on this and the general update mechanism that they used in this case. But this update was sent globally, like you said. It was an auto update to all their agents. And they said afterwards, hey, we're going to change that a little bit. We're going to start rolling this out to smaller pools of people at, you know, sort of staged rollout. Isn't that more of the norm in the way that software updates are deployed? I mean, they're not generally just sent out, like, hey, everybody here it is. Good luck. Yes. And I'll dance with that as politely as I can. Qualified, yes. Yes, we'd love to hear some of that blue green testing, maybe starting with some smaller sample pools of, hey, 100 customers or clients, maybe to start with, expand that to 1000, expand that to 10,000, expand that out as we're seeing signs of success. But I would really love a lot more of that internal testing. And I know there's the commentary that, hey, this is the process that they've used for years. This is what has supposed to have been tested and have its own bug checker. But there's a bug in the bug checker. So if there were a small little sandbox or something to even test, look, do we have a health check? Do we have a heartbeat? Is there a pulse from the systems that we've gotten our own QA workflow and process? I think you read egregiously C. Uh-oh, blue screens are deaf and masked. This is something we can get out in front of. But again, I won't pretend to know. I don't know the inside of CrowdStrike and all that they're up to. But I hope and I'm glad that there are some lessons learned. Yeah, I'm sure there are. Yeah, I mean, there's smart people there, obviously. On the customer side, if you're the one accepting these updates, there's some updates you can test yourself. You can grab those, put those on a staging machine or a test machine and test those yourself. But something like this that's pushed you, the agent is just getting it. You don't really have a way to do that, which is another part of that balance where you're like, okay, we want the updates quickly because there's security, you know, content updates to protect us. But when things go wrong, they go really wrong. And we don't have a way to kind of roll it back. You're right, I'm sorry. Did I miss, hey, some, I'm trying to think of what more to roll off of that. No, I know there wasn't really a question there. It was just my, me kind of getting my thoughts together. But I mean, for IT security teams, is there much that they can do when they have these auto-updaters installed or are they just kind of at the mercy of their providers? I got a fallback and I feel bad, but you might very well be the mercy of the providers there, especially if it's something out in production that they're already pushing a change to. Time will tell. Yeah, I know. So any other broad thoughts from this, John, things that sort of popped up for you as you've been looking at it the last week or so, that folks should keep in mind as we go forward. There is another thread that I might pull on, especially considering we know the root of the product problem, we know the crux of the issue and the crowd strike outage. But this is something that other thread actors and other adversaries will totally capitalize on and try to take advantage of spreading a lot of that uncertainty in chaos with, I don't know, maybe some rogue, malicious scam phone calls trying to say, hey, we're from CrowdStrike and we'd like to help work through this issue, but they're just masquerading and impersonating and leading into what will turn into more access to your computer and more damage to be done. So fake websites, fake domains, it have popped up, fake fishing emails, all this, and I know CISA had kind of chimed out inside of the game. We are seeing a lot of activities surrounding this. So it's something to really keep your ear to the ground and I don't want to say stay vigilant as everyone does, but go to the source, look for what CrowdStrike is up to and they will share themselves, but be especially cognizant of other ill intended actors that will take advantage of it. That's a great point. Yeah, CISA did send out an alert, I think that Friday, like pretty quickly after the whole thing has happened. Just out right off, yeah. Yeah, and CrowdStrike themselves, I think, in their initial or very quickly updated their blog posts to say, listen, make sure you're talking to an official CrowdStrike representative, be wary of people trying to take advantage of this. And we've already seen fishing campaigns and all sorts of things, like that. And I'm sure that will continue to be those. That's a great point because cybercrime groups are going to take advantage of anything like this. And even smaller things that you don't think are a big deal they use to their advantage for lures. And something like this is just like a golden opportunity for them, honestly. Yeah, that's a great point. Well, to wrap it up, hey, I don't mean to sound all doom and gloom again. I'm not wanting to be all that fun. But I actually have had some very sweet interactions with the CrowdStrike team following this, because I tried to chat with some reporters and journalists in PR. And they said, hey, thank you, John, for taking this with the grace that you could, because we know it's a crappy situation. Obviously, they're living through it, and they're the people on the other side of the screen on the keyboard just as well. So if I may just add that reminder to anyone tuning in, all the IT folks that you're chatting with, or maybe even some folks that are on the inside, that they're working through the best they can. And someone you do have to have some extra courtesy and empathy with when it's people at the end of it. Yeah, completely agree. All right, John, thanks so much for your time, and it was great to see you. Thank you, Dennis. This was a blast. Thanks again. See you. [MUSIC]
Podcast Summary
Key Points:
The CrowdStrike outage was caused by a faulty automated update that triggered widespread system crashes, affecting critical sectors like airlines, hospitals, and banks.
Recovery required manual intervention, highlighting challenges in prevention and the reliance on centralized security providers with kernel-level access.
The incident spurred discussions on improving contingency plans, testing updates more rigorously, and the risks of interconnected software ecosystems.
Cybercriminals exploited the chaos with phishing scams, emphasizing the need for vigilance and verified communication during crises.
While changes like staged rollouts may help, systemic vulnerabilities in centralized architectures and human fallibility remain ongoing concerns.
Summary:
The podcast discusses the major CrowdStrike outage, which began when a faulty automated update caused widespread system failures, impacting airlines, hospitals, and other essential services globally. Recovery efforts were largely manual, involving technicians physically accessing devices to remove problematic files. The conversation explores the difficulty of preventing such incidents, given the reliance on security providers with kernel-level access and the interconnected nature of modern software ecosystems.
Key takeaways include the need for better contingency planning, such as having clear recovery checklists and contact protocols, and the importance of staged software rollouts to mitigate risks. Additionally, the hosts note that cybercriminals quickly capitalized on the chaos with phishing campaigns, urging organizations to verify sources during crises. While improvements like enhanced testing are possible, systemic vulnerabilities and human error remain inherent challenges in centralized digital infrastructures.
FAQs
The outage began with widespread reports of blue screens of death on endpoints and servers, initially observed on platforms like Reddit and Twitter, indicating a significant system failure.
Organizations should improve their incident response planning, including maintaining updated contact lists, checklists, and recovery procedures to handle large-scale outages effectively.
The outage was severe because it involved a kernel-level update from a cybersecurity provider, affecting critical systems globally, such as airlines, hospitals, and banks, due to the interconnected nature of modern software ecosystems.
Kernel-level access allows vendors to monitor and protect systems effectively but is fragile; errors can cause widespread failures, as seen in the CrowdStrike incident, raising questions about security and stability.
While prevention is challenging, organizations can adopt staged software rollouts, enhance testing procedures, and ensure robust backup and recovery plans to mitigate the impact of such events.
Adversaries exploited the chaos with phishing emails, fake websites, and scam calls impersonating CrowdStrike, prompting alerts from agencies like CISA to warn users to verify sources carefully.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.