Guest Karen Pfeifer Talks AI for IT Operations: Turning IT into an Innovation Engine - The IT Leaders Podcast
0m 0s
In this podcast episode, Karen Pfeifer, Field Chief AI Officer at Pythian, discusses the transformative potential of AI for IT operations (AIOps). She highlights how traditional IT teams often face "alert fatigue," overwhelming ticket backlogs, and a "hero culture" that prioritizes reactive problem-solving over innovation. AIOps addresses these issues by using AI to filter noise, predict incidents before they escalate, and automate mundane tasks—shifting teams from a reactive to a proactive stance.
Pfeifer emphasizes the "capacity trap," where teams spend most of their time on maintenance, stifling innovation. AI can break this cycle by handling repetitive work, allowing engineers to focus on strategic projects. For successful adoption, she stresses the need for executive buy-in, aligning AI initiatives with business outcomes like revenue protection, and fostering a culture that trusts AI as a collaborative tool. She recommends a pragmatic approach: assess current KPIs, identify high-impact use cases, and start with controlled pilots. Ultimately, AIOps enables IT to evolve from a cost center to a driver of growth and reliability.
Welcome to the Podcast
Welcome to the IT Leaders podcast, brought to you by the Doyle Group.
I'm Kate Sosha, your host.
This podcast is where we spotlight the most influential IT leaders, diving into their strategies for success.
We're building high performing teams to fostering innovation and navigating the challenges of a rapidly evolving tech landscape.
We cover it all, whether you're a CIOCTOVP of IT or an emerging tech leader, There's something here for everyone who's passionate about technology leadership and transformation.
Thanks for tuning in and let's get started.
Meet Karen Pfeifer
I'm here today with Karen Pfeiffer, who's the field Chief AI Officer for Pythian, drawing on 20 plus years of leading enterprise scale, global data and AI organizations in 25 plus years in media and entertainment technology leadership with brands like Disney, Hulu, Comcast, NBC Universal, DIRECTV, AT&T and Odyssey.
Karen shares practical stories on how IT leaders are using AI to tackle alert fatigue, incident volume change, risk, and the capacity trap that keeps experts buried in toil instead of building their future.
Welcome, Karen.
Speaker 2
Thank you, Kate.
I'm super excited to talk with you today.
You and the Doyle Group are such amazing thought leaders and feel like that leadership challengers and curators.
So it's going to be a really great conversation.
I'm looking forward to it.
Oh.
Speaker 1
Thanks To kick us Off, you've spent more than two decades leading data and AI strategies for global enterprises and media entertainment and beyond.
Why AIOps Matters
What patterns did you see in IT operations that made you passionate about AI for IT OPS specifically?
Speaker 2
Now I, I think of this part as early experiences with large scale operations and seeing a lot of hero culture, recurring outages, lots of lots of toil that was almost celebrated and and that led towards eventually a OPS, but we had to have the generative AI revolution and now we're having a gentic AI come in.
So now it feels like that opportunity is meeting the capabilities.
And so I kept seeing incredibly talented engineers rewarded for heroically fixing the same problems at 2:00 AM instead of being given time to prevent those problems in the 1st place.
Across industries, the pattern was the same.
Fragmented tools, tribal knowledge, cultures of firefighting that really burned people out and, and that led to slowing the business down.
And as a business leader, one of your most valuable contracts is that between you and your employees, right?
That that contract of trust and not bringing them out.
And so that really pulled me into the AI OPS side, using data and AI not as another dashboard, but as a way to redesign how operations work so teams can finally get ahead of work.
Inside IT Ops Pain
When you walk into new organizations, what does a day in the life of a typical IT operations team look like?
And kind of where does that real pain lie?
Speaker 2
Yeah, that's a great Next up question for that.
So this space is more talking about alert flooding, ticket backlogs, manual scripting, change risk and firefighting and burnout.
We we see a lot of that.
And so most teams are drowning.
They're just drowning in alerts and tickets and they don't have an incident problem.
What they have, I'd like to say is an attention problem.
Where are they going to place, if they have this deck of cards for what, where they place their attention next?
Where are they going to place those bets?
Right?
And so the highly paid engineers are still writing one off scripts and doing copy paste work that an intelligent automation layer could handle in seconds for them.
You know, why would you do it any other way if you have the opportunity now?
And so the last thing I'll say about this is that the collateral damage is culture.
So people feel like they're always behind, you know, and always on call and never doing the strategic work that they were hired to do.
Yeah.
Escaping Capacity Trap
It's interesting this, you know, kind of with AI being brought into the workforce.
I've kind of seen it in two different ways.
One way is like, oh, this is just another burden on the team.
We got to figure out how we're going to handle this.
How do we incorporate this into our already, you know, strapped road map and where we have lack of capacity to implement new things?
And then this other school of thought where let's bring AI in, but let's use it to free up more capacity and to help us get closer to our goals.
And I think the capacity trap in IT is something that you often talk about.
Can you explain what that is and why it's so dangerous for organizations to innovate?
Speaker 2
I'm so glad to get to talk to you like this again today 'cause I know we, we talk every now and then about these sorts of things.
And so, yeah, you're right, the, the capacity trap, I think we talked about this before I was tied to this sort of 80 to 90% of time on run tasks versus innovation and change, right?
And so the constant firefighting, I mentioned before, there's, there's this negative impact on transformation and morale too, right?
So another way to think of of the capacity trap is where not just runtime, but also 89% of your smartest people are stuck just keeping the lights on and there's no oxygen left for innovation, right?
So it's dangerous because you keep adding more projects, more tools to try to fix this problem.
But because the operating model doesn't change, all you've really added is more load on the same exhausted team.
And then until you deliberately shift from run to change like I mentioned before, then often with AI and automation, right, you're simply not going to hit the transformation targets that the board is expecting.
It's not going.
Speaker 1
To to achieve that kind of transformational shift.
Is it really a culture thing?
Is it completely rewiring how you run your development teams and organizations?
Is it something else?
Speaker 2
I mean, I think we know the answer is really both, right.
You know, I mean, if you that's the hardest part of a leadership, right, is finding that balance, finding the balance between the the technology and, and the process side that you know, needs to be innovative on.
And then the people side, making sure that the culture is healthy, that there's a culture of not being afraid to innovate and then also not being burdened with lots of busy work.
And so I think if, if teams can really think of this AI era now as a way to get the burden of the busy work off their plates, then that's really going to be that sort of that X Factor of transformation for them.
Speaker 1
Yeah.
That burden and busy work concept is one that I'm hearing not only in development, but in sales, in product, it's kind of everywhere.
And I think to be able to really make that shift in transformation, it's going to change the way that work is because you get to do more of those fun, interesting projects and move a little bit away from some of those menial tasks, which I don't know about you, but I don't like doing those menial tasks at all.
Speaker 2
Definitely not.
Definitely not.
Defining AIOps Value
For leaders, when they hear AI OPS and think it's just more monitoring, how do you define AI for IT operations in a way that connects to business value and not just tooling?
Speaker 2
Yeah.
So I think of this as about AI that handles mundane, complex, repetitive tasks.
So that shift from reactive to predictive anticipatory operations and protecting revenue and customer trust probably most of all.
And so AI OPS is not more alerts, it's the opposite.
It's using AI to filter the noise and surface the few signals that truly matter to the business because it has the capacity and the capabilities and the authority to discover, prioritize and act on what matters.
AI does.
AI is given that authority, a gentic AI agency, right?
And so think of it as a digital teammate that handles the repetitive, complex pattern matching work.
So your human engineers can focus on decisions and design, and then strategically that becomes innovation.
That becomes moving your IT teams from being the cost center to being the innovation powerhouse that you know they can be, right?
So at the executive level, AI OPS is really about protecting revenue and reputation by spotting issues before your customers ever feel them, whether that's external customers or also internal customers, which are just as important to IT teams.
Agentic AI and Oversight
When you talk about spotting the issues, is that something where you're having a gentic AI alert you of a potential issue, or is that something where you know AI is doing a process and you've got a human in the loop approving or validating different outputs that it's receiving?
Speaker 2
I mean, it can be, it could be any of those things, right?
I mean, in an ideal situation, you would be able to, you know, go to sleep at night, 'cause humans need to sleep.
You'd have your agentic workforce take over and do maybe those, you know, those low risk tasks that you need them to do that you know, that you've tested out over and over again and you know that they can catch things and sort of let things hum.
And then maybe for the high risk things that get flagged, you know, it waits for the human in the loop in the morning when you get up or the next person on your team that's following the sun ideally to, you know, step in and, and help provide that managerial oversight.
I do really like the human loop concept as much as we can, right?
But it can't be a bottleneck either.
And so sometimes I talk about the concept of more humanity in the loop and having it be more about stepping back and having it be baked into the governance and even the security and the processes ahead of time overall.
And ensuring that we're thinking about it from an even broader equation than just did I catch that alert and was it a false positive or a false negative?
Speaker 1
An interesting new concept that I've been running into recently with my partners has been to actually utilize AI to help them in their data governance plan, to help them create a uniform list of terminology and align the organization.
Which I thought was kind of fascinating because you know, previously had always been we need to get our data structured, organized and ready for AI.
And now it's, let's give it a small subset of data and have it use itself to to get ready and to, you know, ultimately be successful and useful.
Speaker 2
That's great.
I'd love to see that.
Executive Sponsorship
When we talk about non IT folks, what role do non IT executives like the CFO and CEO play in successful AI adoption for IT operations and what do you need from them?
Speaker 2
You know, it's the most important one in my mind.
I, I always say to my technology teams and my data teams that my road map is not my road map.
It's not our map, it's the businesses road map.
Like if we're not aligned with whether it's the CEO, the CFO, the CMO, the CRO, the COO, right, CIO, of course, if we're not aligned with them and what the business needs, then why are we building this technology thing, this widget that we're going to be building or this process that we're going to be enabling and adopting.
So for me, this is really connected to financial impact, risk appetite, viewing reliability as a product and feature tied directly, revenue and brand.
So if I'm needing something from the CF OS or CEOSCXO, whatever you name it, it's that we need recognition that reliability is a product feature, not just an IT metric and that every minute of downtime has a line item in the PNL.
We need to have clarity on the risk appetite too.
So are we optimizing for absolute stability?
Are we optimizing for speed, for cost?
Because AI can dial each of those levers differently.
And so finally, when the business leaders sponsor the journey and show up to the road map conversations, then IT stops being a cost center and starts being a growth enabler.
Speaker 1
When you're having these conversations, I'm hearing revenue an important factor at the baseline for all of the conversations when you're going about your AI Rd. maps and your AI journeys and trying to identify where we want to implement things.
You know, of course we want to reduce load for developers so that they can be more strategic.
But are we looking at our financials and identifying first, hey, this is where we can make financial impacts?
Or are we going to IT 1st and saying, hey, where can we reduce load?
Is it a combination of both?
Speaker 2
Well, ideally we would have already, you know, we would have been had a process in place to have ITB surfacing these insights to authority.
That's the sort of dashboard way of looking at things.
But that also puts us in this sort of reactive mode all the time, right.
So this is about, as I mentioned before, about becoming proactive and having AI find that spark first before the fire starts and really being more predictive and prescriptive with how we're going to run map these things out reactive now.
Proactive Incident Response
What does it look like to move from a traditional reactive incident management to more of an AI enabled incident management that is genuinely proactive?
Speaker 2
I think it's concepts like mean time to identification, noise reduction, dynamic baselining and like I mentioned that anticipating the sparks before the fire.
So think about like this like traditional incident management starts when something's already broken my OPS post the start line forward.
So the moment that first anomaly appears, you know it's caught early.
And how that's done is instead of humans scanning 20 dashboards in the morning with their cup of Joe, right AI eyes on screen that we say, right AI continuously learns your normal patterns, IE dynamic baselining, and then flags the spark before it becomes the fire.
So when you compress the time it takes to identify the real issue, mean time to identification, then everything else gets faster, mean time response to recovery, and then ultimately customer communications become faster.
Data Sources for AIOps
What types of information are you giving an AI system access to so that it can ultimately learn your habits and trends and start actually identifying those incidents?
Speaker 2
That's a good question.
I mean, there's what's really amazing is that there's a lot of really great services out there now that you can use to tap into lots of different types of data.
Before, remember, we only had automated logs, for example, But now more than ever with, you know, with agentic AI and then the ability for large language models to be multimodal.
Now it can look at things like SharePoint or PDF documentation or code itself and start to find anomalies and all those data sources as well.
And so I think we, you know, you feed it as much as you can or that you're willing to within the, the safe boundaries and the guardrails and the governance that you set up and then let it do the work to do the pattern matching and then find those patterns.
Alert Noise Success Story
Can you share a concrete example of how AI reduced alert noise and improved MTTR?
What changed for the engineers on the ground and for the executives watching uptime and costs?
Speaker 2
Yes, that's a great question.
So I think about this as larger alert volumes versus smaller number of true incidents and high false positives are also conversely, you know, false negatives, which I also find to be kind of that sleeper problem that, you know, you really don't want your organization to be thinking, oh, you know, that's been resolved, that's not a problem.
And then it really was a problem.
And so there's also the gains from noise reduction and mean time to improvement there.
So in one environment, we took thousands of daily alerts and condensed those into some smaller amount of very actionable incidents.
And so the engineers were trusting what the system was telling them.
It what is what sometimes is a bigger problem is again, thinking about that contract you have between you and your employees, right?
And having there be not just alert fatigue, but also trust fatigue, right?
At least see this large amount of incidents.
And I don't know what I can trust anymore.
And you do not want to go there 'cause once that trust is broken, it's very hard to get back.
Then there's this mean time to identify the root cause for these challenges.
And so for us, that went from hours and hours of these different war room debates, middle of the night war room debates, to minutes of data-driven correlation.
And then so for the C-Suite, I think about this story as being very simple, if you have smaller major incidents and then faster recovery.
So when they did happen, then there is this very clear dollar value on minutes that we are giving back to our leadership.
And that's, that's really where that that power trust also goes between you and your leadership as well.
Speaker 1
Yeah, I think you reduce your overall risk as well, because if you're having so many incidents coming through, you're going to get fatigued.
People are going to ignore some of the incidents that come through because they're not necessarily correlating it to actually being a problem.
It's just more noise.
When you're determining kind of how you want to reduce that noise.
Are you making a matrix that is waiting, you know, core values or are you considering in other ways or you having AI help you determine what's actually important and what is how do you determine that?
Go.
No go.
Speaker 2
Well, I mean, that's, I mean that's the whole point to this, right, that we're using thinner the AI and the power of taking those models and finally attaching them to tools so they have agencies so they can start making those decisions for us.
We give them guardrails that we line in governance sessions and ensure that we are governing what those that level of tolerance is the acceptable level of tolerance, the acceptable level of error.
And we talked about this a lot in models, where models is about prediction.
Prediction is about having a certain amount of error.
And so there are some things that you can't have any error on.
And so you do want to ensure that you have not just machine watching, but also human in the loops.
And then there's other incidents where, like I mentioned earlier, there is an acceptable amount of error and it's OK for these lower level, lower risk things to let the automation handle it for you.
And so really, I mean, ideally he hated sort of a mixture of, you know, human, that human intelligence, that human general intelligence, you know, still watching over things.
Speaker 1
Yeah, yeah.
First 90 Days Plan
If you are a CIOCAIO or a CTO listening to this and you know your team is drowning in toil and ticketing cues, what are the first three concrete steps that you would take in the next 90 days to start an AI for IT OPS journey?
Speaker 2
OK.
So this is about outlining A pragmatic sequence for them to digest.
So assessing current state and KPIs, running AI for IT workshops, possibly doing some discovery and then prioritizing those high value, lower risk use cases and then safely deploying an initial AI capability into production.
So step one is baseline reality.
You need to measure your current incident volume.
So although measure mean time to recovery alert noise and then how your team actually spends its time.
So baseline step one.
Then Step 2 is you run a focused discovery or possibly even a workshop to map your biggest pain points and then identify 3:00 to 5:00 AI driven use cases that matter to the business.
Where are you going to place your cards?
Like I mentioned earlier, if you have a certain amount of cards in your deck, where are you going to place it on those things that are actually AI and actually have value?
And then Step 3 is one of those or maybe 2.
And then build a safe pilot with clear guardrails and get something into production that your teams can feel that they can use in their day-to-day work.
I really like to learn more towards not just doing something that's a proof of concept, but really doing an MVP so that there's something of value that was in production and you can test it out in sort of a safe production.
Speaker 1
Yeah.
And when you mentioned safe production, are you bringing your data over into a separate environment so that you're not altering, you know, your core data?
Like what does a safe production look like?
Speaker 2
Well, I mean, it's funny when I asked that I've been in other large organizations where there weren't sometimes the even the basics of how you handle this.
You have usually a testing, you have sandbox, you have testing, you have dev and you have prod environments and not all organizations, even some of the big ones who might be surprised to follow that discipline.
So ideally before you're getting to production, you already have run those tests, you sandbox them out, you've done tests, you've done dev and then you're in prod.
And at that point, you're, when you're in safe prod, you're basically you have all of those full alerting cycle in place, not just that the alerts are happening, but then there's that feedback loop that you've built in that is giving you feedback in real time and making sure that the model is improving.
And then you're not having model drift, process drift, data drift happening.
That's what I mean.
ITIL Quadrant Framework
You anchor your work in ITIL 4 and SRA using A4 Quadrant IT Operations capability model.
Why is that structure so important for making AI real and not just a science project?
Speaker 2
Hey, thank you.
Great question.
So let me explain the quadrant.
So there's reliability and service quality, there's incident and postmortems, there's life cycle change and then foundational governance.
And so it's all about tying those to AI use cases, tiny AI use cases, each of those.
So it's the risking your investments.
So the four quadrants give us a map again, reliability, incidence, life cycle and change and foundations and governance.
And then every AI idea has to land in one of those boxes.
So then that structure keeps us honest, right?
So we're not doing AI for the sake of AI.
Nobody wants that, nobody needs that.
And then we're still solving a specific named operational problem that leadership already recognizes.
And so it also helps with funding.
So when you say you can say this use case reduces change related outages or this one improves SL OS, then the business case essentially writes itself I.
Speaker 1
Imagine that'd be really useful in a brainstorming session because then you're not focused on, OK, let's find the perfect use case right away.
We're just throwing up all the use cases on like, I identified this problem, identified this problem.
And once they've all been put onto the quadrant map, then you can start to say what's going to make the biggest impact?
How are we going to drive the most value for shareholders, internal stakeholders, internal finance, everything like that?
A.
Speaker 2
100% right.
You're no idea when you're spilling it.
No idea is a bad idea, right?
You're just throwing off.
Some of them might be staying within the context of the business.
You're throwing them on that board and then you're doing the various overlays of is this going to increase our revenue?
Is this going to reduce churn?
Is this going to be some sort of efficiency for the business?
Right.
And then ideally you're layering another layer over that where you're looking at it from a risk side of it.
You know, what is my, what is my risk management here?
What is my security?
What is my compliance?
Do I have any, you know, security compliance, governance sets of measurements that even though it might be low reward, higher risk, you want to make sure that those aren't being negatively reduced down or impacted and reduced down because they don't have higher value.
But you have to do them because of government.
Yeah, yeah.
Balance.
Speaker 1
It's interesting on the concept, like there's no bad ideas because like there are some bad ideas, but sometimes you have to throw those out so that you can get people thinking or you know, help spark other ideas which then lead you to a better end state.
Speaker 2
In that case, Spark is a good thing.
Speaker 1
Yeah, yeah.
AIOps Workshop Playbook
You also run structured AI for IT operations workshops.
What happens in those sessions?
How do you plan and help leaders filter out AI hype and get 2 concrete Rd. maps?
Speaker 2
Yeah, I spend a fair amount of time doing this.
So it's about talking leadership and their teams through things like data literacy and fluency and ITIL quadrant exercises and KPI alignment, which you can imagine KPI alignment has taken me sometimes a whole day or more, you know, even get to which were the ones that the domains aligned on as the actual KP is much less the numerator and denominator and where the data sources were that could that could take weeks, right?
I'm not kidding.
Use case identification, risk and governance readiness assessments and then building that valuable road map of quick wins and and big bets.
And so we start by getting everyone on the same page about terms.
So often it starts with education in first parts of the day and then activation later on.
So understanding what AI OPS is, what Gen.
AI is ITEL is.
So the business and the technical leaders are speaking the same language, which it really is half the battle when you're getting everyone into the remark, half the conversation.
And then we map current pay points to the ITEL quadrants, identify those high value AI use cases from years of experience and lots of pain and victories and mixed in there.
And then stress test them against KPI's risk and data readiness, which is a big part of it as well.
I should say that also parts of workship they do are very much around data readiness for AI as well, because lots of teams that I talked to, they'll have these amazing AI ideas and then quickly realize that their data are not ready to support that.
And then finally, the last thing I'll say is that by the end, the teams walk out with a prioritize read a six month road map, maybe longer depending, and then quick wins, pilots, MVPS and then one or two strategic big bets with clear success metrics that we walk out of the session like that, that everyone's aligned and this is what our next steps are.
That's gold.
Speaker 1
Yeah, yeah.
And when you're having those sessions, who do you involve?
Are we just looking at the C-Suite level?
Are we breaking it down to VP directors?
And is it just IT business and finance?
Are we extending past that?
Speaker 2
That's a good question.
You know, it really depends.
Each organization is unique.
It depends on the size of the company.
It depends on the maturity of the company.
I've done it from everything at C-Suite level leadership and then they're S suite underneath them and then also done it with leaders of IT and a couple of other VPS and directors and analysts in the room.
It really, it really depends on, you know, where each team is.
Some teams within companies are a little bit farther ahead of the curve or want to be farther ahead of the curve.
And they felt that pain real time.
And they finally realized that just going with some new technology that someone says that shall do this thing is not the best thing to do, That it's better to come in and have their own road map put forward and then meet the business in the middle.
And ideally I would say too that when I do these sessions, often what we'll do is we'll send out, I'll send like a survey out and collect in business feedback and use cases and about our teams and, and kind of make sure that when we have the business's vision and mission and what and what their objectives are, then we're weaving in what the technology teams want to do and vice versa.
Speaker 1
Got it, got it.
So we're doing some pre work to really get a holistic picture to have the most effective conversations we can.
Speaker 2
Yes, with workshops with me, you don't, you don't get off these and I'll just comment and magically wave a magic wand and make it happen.
You get surveys, you got to show me your architecture, your data sources.
We really get into it.
So, but I think, but it's worth it and I think it'd be really fun to do that for anybody in the space who wants to really delve into AI for IT operations and gain that force multiplier of the automation in place.
Yeah.
Speaker 1
It sounds fascinating.
I'd love to sit in on a a workshop some point just to learn and see kind of what businesses are currently doing and how they're moving forward as well.
Speaker 2
Anytime key, I think we can make that happen.
Speaker 1
Oh, sick.
I'll call you after.
Lightning Round Fun
Well, we're going to kick off the lightning round of questions.
You ready to go?
Speaker 2
Lightning round.
OK, Yeah.
Speaker 1
They're really easy, I promise.
Not nothing too hard.
First one is coffee or tea.
Speaker 2
Tea, definitely.
Speaker 1
Early bird or night owl?
Speaker 2
Early Bird.
Speaker 1
Working from home or in the office.
Speaker 2
Man, I gotta do a split on that one.
Definitely I I've enjoyed working from home, but I also very much love people and it's just fun to get in there and talk with people every now and then.
Speaker 1
I agree, solve problems quickly.
It's lovely A.
Speaker 2
Little bit of a blend, yeah.
Speaker 1
Innovation or optimization?
Speaker 2
Again, blend.
I think it's so important to innovate, but if you're not innovating without optimization, then you're just like, you know, throwing stuff at the wall and hoping it'll stick.
Speaker 1
Yeah, hire for skill or hire for culture.
Speaker 2
Again, a blend.
I like to hire smart people who are just highly motivated.
And so that's part of that is culture too.
You know what?
What's their culture of success?
Yeah.
Speaker 1
Yeah, mountains or beach?
Speaker 2
Mountains.
That's easy.
Speaker 1
Oh, I was shocked.
I I thought you were going to say beach for sure living in California, but the mountains are beautiful.
Speaker 2
When you live next to the beach, you know, then you the grass is green on their side or the grass is green up in the mountains, so.
Speaker 1
Fair.
That's totally fair.
Books or movies?
Books.
And then the last one.
Dodgers, LA Kings, Lakers or 49ers?
Speaker 2
I got to say neither.
I've spent six years in Philly, so go birds since the eagle.
Speaker 1
Heck yeah.
Heck yeah, you can move the girl to California, but we can't take Philly out of her.
Speaker 2
That's right.
That's right.
Speaker 1
Well, thank you so much, Karen, for sharing your insights and experiences with us today.
Closing and Next Steps
If you'd like to connect with Karen, you can find her LinkedIn profile in our show notes.
And then Karen has also previously joined the Duo Group for a data panel focused on data foundations, quality governance and integrations, which you can find on our YouTube channel.
And a thank you to everyone who tuned into the IT Leaders Podcast.
Don't forget to subscribe, leave a review, and share this episode if you found it valuable.
Until next time, keep leading and innovating.
And thanks, Karen.
Speaker 2
Thank you, Kate, such a pleasure.
Podcast Summary
Key Points:
AIOps (AI for IT Operations) shifts IT from reactive firefighting to proactive, predictive management by reducing alert noise and automating repetitive tasks.
The "capacity trap" occurs when IT teams spend 80-90% of time on maintenance, leaving little room for innovation; AI can free up experts for strategic work.
Successful AI adoption requires executive sponsorship, alignment with business goals (e.g., revenue protection), and a cultural shift toward trust in AI as a "digital teammate."
Practical steps for starting an AIOps journey include baselining current metrics, identifying high-value use cases, and deploying pilot projects with clear guardrails.
Summary:
In this podcast episode, Karen Pfeifer, Field Chief AI Officer at Pythian, discusses the transformative potential of AI for IT operations (AIOps). She highlights how traditional IT teams often face "alert fatigue," overwhelming ticket backlogs, and a "hero culture" that prioritizes reactive problem-solving over innovation. AIOps addresses these issues by using AI to filter noise, predict incidents before they escalate, and automate mundane tasks—shifting teams from a reactive to a proactive stance.
Pfeifer emphasizes the "capacity trap," where teams spend most of their time on maintenance, stifling innovation. AI can break this cycle by handling repetitive work, allowing engineers to focus on strategic projects. For successful adoption, she stresses the need for executive buy-in, aligning AI initiatives with business outcomes like revenue protection, and fostering a culture that trusts AI as a collaborative tool. She recommends a pragmatic approach: assess current KPIs, identify high-impact use cases, and start with controlled pilots. Ultimately, AIOps enables IT to evolve from a cost center to a driver of growth and reliability.
FAQs
AIOps uses AI to handle mundane, repetitive tasks, shifting operations from reactive to predictive. It protects revenue and customer trust by filtering noise to surface critical signals, acting as a digital teammate that allows human engineers to focus on innovation.
The capacity trap occurs when 80-90% of an IT team's time is spent on routine tasks and firefighting, leaving no oxygen for innovation. It's dangerous because adding more projects or tools without changing the operating model only increases the load on exhausted teams, hindering transformation goals.
First, baseline your current state by measuring incident volume, MTTR, and team time allocation. Second, run a workshop to identify 3-5 high-value, AI-driven use cases. Third, build a safe pilot with clear guardrails and deploy an initial AI capability into production.
AIOps enables proactive incident management by using AI to continuously learn normal patterns through dynamic baselining. It flags anomalies early—the 'spark before the fire'—compressing mean time to identification and accelerating overall response and recovery.
Non-IT executives like the CFO and CEO must recognize that reliability is a product feature tied directly to revenue and brand. They need to sponsor the journey, provide clarity on risk appetite, and engage in roadmap conversations to shift IT from a cost center to a growth enabler.
AI filters thousands of daily alerts into a smaller number of actionable incidents, reducing noise and false positives. This allows engineers to trust the system, speeds up root cause identification from hours to minutes, and translates into faster recovery and clear financial value for leadership.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.