Go back

The Myth of the 10x Engineer | Charity Majors In The Engineering Room Ep. 41

64m 17s

The Myth of the 10x Engineer | Charity Majors In The Engineering Room Ep. 41

In this conversation, Charity Majors discusses key principles for building effective software teams and development practices. She emphasizes that high-performing teams are defined by their ability to collaborate, communicate clearly, and collectively own software, rather than relying on individual "10x engineers." Charity highlights the importance of fast feedback loops, observability, and continuous deployment to enable teams to learn and iterate quickly. She critiques the industry's focus on hiring top talent alone, arguing that systemic factors—like enabling frequent, safe deployments—are more critical for team growth and success. The discussion also covers the value of diversity in teams, noting that varied backgrounds and perspectives foster resilience and innovation. Charity addresses the impact of AI, suggesting it shifts the engineer's role toward problem-solving and maintaining durable code, rather than just writing code. She advocates for leadership that builds psychological safety, encourages curiosity, and supports incremental, evidence-based development. Overall, the conversation underscores a team-centric, systems-oriented approach to software engineering that prioritizes learning, collaboration, and sustainable practices.

Transcription

10004 Words, 54071 Characters

English
Welcome to the engineering room, a monthly series of long-form conversations with influential people from the software world. The engineering room series is sponsored by equal experts and I'd like to thank them for their ongoing support. So she'd like some help building some great software or are interested in finding a great place to work. Do check out their details in the description below. My guest today is a prolific writer, speaker and was an operations and database engineer and sometimes engineering manager. Currently she's the CEO and co-founder of Honeycomb who build tools that help organisations observe the behaviour of often complex distributed systems and is one of the people who is instrumental in introducing the concept of observability to our discipline. She's worked as a production engineering manager at Facebook and at Lyndon Lab worked on the infrastructure and databases that power second life. She's a co-author of database reliability engineering and observability engineering put both published by O'Reilly and she loves working in startups and the chaos of hard scaling problems. Somehow she always ends up in charge of the databases. I've not met her before today but I've always been interested in talking to her for a very long time now because mostly our views on software developments seem to align so strongly so please welcome today's guest Charity Majors, hard charity. Hi Dave, thanks for having me at CTO, not CEO. I'm sorry, I'm sorry, CTO. I want to see you for the first few years but Christian and I swatched jobs and it's much better this way. Cool, I think I'd prefer to be a CTO than a CEO. Okay so I want to quote yourself back to you because there are so many things that when I read something by you or watch one of your talks I'm nodding along and smiling all of the time because we're so well aligned sometimes I hear things I think I could have said that. So I want to read a couple of things that just tickled me and I loved so much. So high performing teams, teams that spend most of their time working on interesting novel problems that move the business materially forwards. What else can be high performing mean? That's such a great description. The team is sorry go ahead. And by the way I love you reading you're writing over the years just as much. It's kind of amazing that we haven't gotten to talk before now. It is funny yeah. So another quote, the team is the smallest viable unit of software ownership individuals don't own software. 100% completely agree with that. Let's start there. Let's talk a little bit about teams because I know that's a focus of yours. What does it take to build a great team? It's a hard question I suppose. But it is a hard question because it depends on what problem you're trying to solve to some extent. I mean I think I know Will Larson just publishes great blog posts called Management is a FAD engineering management is a FAD which I loved aggressive title but I loved it because it's true that like so much of what managers are told is like you know the way to it's actually a temporary condition based on current business realities. Like we all remember what is like in the zero interest rate days when like great managers were taught how to like spend all their time hiring right. So like to some extent it depends on the problem you're trying to solve. But I do think that there are some constants. Any comb for example we wait communication skills really highly. If someone can talk me through how they solve the problem I believe that they can write it in code. Even before AI that was true right now it's really true right. There are plenty of engineers out there who can't who can write the code but can't talk you through their process. And we think that that you know collaboration is such it's the fabric of how you learn how you progress and how you bring each other along how you both specialize and back each other up that I think that's a pretty good starting point. I again I agree so strongly with all of that that I have to confess that developers that can code but can't talk about it always make me a bit suspicious because I'm quoting Martin Fowler now but you know good code is designed for humans to read. Before I can write code that operate on a computer being able to write code that humans can really understand is the hard part and it plays into your the quotes that I mentioned earlier which is the team centeredness. If teams are the unit of software development delivery then that communication is absolutely central to be able to do a good job to not only talk about it but to write code that we can share and understand and work together on and collaborate on and all of those sorts of things. Yeah the writing of code has I want to say never almost never then the constraining factor on how fast a team can go. Almost never. It's almost always understanding it. Yes and it collectively understand and own it together. And if that was the hard part then we'd employ typists because they could talk about it. Exactly. Faster. It seems crazy to me. It's the thinking and I have some more quotes from you that I saw that you were saying the same thing. So it's the thinking that we need to optimize for not the typing and being able to deal with complex ideas and use those are really the currency of our job and as you were alluding to I think that AI has highlighted the importance of that to some degree because now the AI can do the typing for or some of the typing for us. It's probably best if we check it but yeah but being able to decompose the problem and think about it in small pieces and make progress and validate what's going on as we go all of those skills still seem completely centred doing a good job. Pretty valuable. Yeah. I wrote a little while ago about that I think software is sort of bifurcating into disposable and durable code because it's so fast and frictionless and easy that I think that like generating code and plot or whatever is just going to be if you work in knowledge you're going to have to know how to code generate right. That is a very different thing than durable code and durable code is what runs the world and the model of durable code is we make small, small, consistent changes to what is known to be good or good enough what is known to be stable right and I don't see I mean maybe at some point in the future AI will be able to tell me that code is better more than something having been in production for two years but I somewhat doubt it right and like so even the ability to generate all this disposable code it rests on this enormous iceberg of stable tested production code that you don't want to replace too fast because there's only one way to test and prod or live a lie as my teacher says you know only only prod is prod. Yeah it's the reality of the situation. Yes. So I've been revising some of your stuff for this conversation and I find too many of those things in my head now and I don't want to lose some kind of coherent picture for our listeners and viewers but one of the things that you were talking about observability in terms of the importance of closing those feedback loops so we learn seems to me so vitally important to my model of how I think about software development and the vital importance of incrementalism our ability to build step by small step on something that we know already works. So trying to establish that we know that it works and find out how to do that and then make a small change to figure out how to move forward from there seems really put I've got another quote of back to you of you to yourself so modern software development practices. So what makes good software development modern software development practices engineers own their code in production practice observability driven development testing production separate deployments from releases and continuous deployment or at least continuous delivery. Again that would that seems stuff that's writing in my wheelhouse and the stuff that I tend to talk about and I'm not quite so sure about observability driven deployment and I want to talk about that in a bit more detail later on because I think we are in complete agreement but I think that I use really quite different terminology when I talk about that but let's come back to that. Let's go back to the team thing so if we're thinking about the ways in which to try and how pair teams to grow and to become these good teams that do all these stuff where do you start? So I think that this is what really got me animated about this idea a while ago was when I saw a post on LinkedIn and it was from Coinbase and they were bragging about how they only hire the 0.1% of people that apply and I was just like oh you fuck you got it. You know like Netflix for years has been like we only hire the global top 10% and now these guys are just like yeah we only hire the top 0.1% and I'm just like do you know how we ridiculous this sounds and the more I thought about it and even I were talking earlier I find that I discover what I think by writing and I started writing a please about this and I realize that is really letting leadership off the hook. Absolutely. If you put all the emphasis on hiring the best engineers 10x engineers if you put all your emphasis on like you know world class people you look nobody is born a great software engineer. Every software engineer is built is is developed over years and years and years of working on hard problems together with other engineers. When starting from scratch and how excellent how world class engineer can become depends vitally on the systems that they are developing it. Like every time you ship code you learn you learn whether what you did worked or not good or bad if you work for you did or not. A team that gets to ship 12 times a year is going to learn a lot less over 10 years than a team that ships 200 times a day just like baseline right. And if it takes the slowest engineer at the company one day to ship one line of code it's also going to take the fastest engineer at the company one whole day to ship one line of code because they're participating in the same shared sociotechnical systems. It's the same system you know and or if you're bypassing the system that's in the bigger problem right. And so I feel like it's really like it is harder to craft sociotechnical systems that give everyone fast feedback loops that give everyone guard rails and give everyone the ability to follow their code to its conclusion right to watch you know when a tell that is when you ask software engineers how do you know if your code works or not and if they say my test pass you know that they haven't even begun to hook up. Core feedback loop of does my code work or not because you don't know if it works until it's been a production for a bit. Yeah absolutely absolutely and the thing about the you know hiring the top point 1% and so on. It makes me laugh as well. I always delighted when Google published the project Aristotle stuff where they you know they're fine. I think they started off thinking that they were going to find okay if we have five engineers that are in the top 10% and four engineers that are in the top 3% and one engineer is in the top point 1% will have a great team. I think they probably expected something like that but actually there's research that goes back to the 1960s by a chap called Bellbin who came up with a model that said where you put a bunch of top performers together you get dysfunctional teams you need a spread of different people a range of different talents and types of people to make a great team and the project Aristotle stuff found that the number one predictor of high performance in Google's teams was trust between the members of the team and very very different kind of picture to I think what they expected when they started. When a flower doesn't grow do you blame the seed do you blame the seed like there's a whole system you know. The same person might really thrive in one place and not thrive in another place and it doesn't necessarily mean that there's something wrong with either of them. This is what management is hard. There's so much context involved but I will say that like a pretty big part of leaders helping their systems grow up and mature and support these engineers involves hooking up these feedback loops of responsibility and then he gets safe to fail. Having a culture where I think the best teams I've ever seen work on are the ones where you merge a death and you just assume it's going out. By default it's going out. It will be out soon. You have to do something to stop the train to make the deploy stop. You can automate that if you can just make it automatic and if you can make it and if you can give engineers the tools to see it. To your point earlier the best teams are what. They're the ones where every single person is constantly curious, learning, pushing themselves, engage. If you have a bunch of staff plus engineers, they're all super high level. Someone's got to write the log in form. Someone's got to write the web page and if nobody is interested or excited or learning or growing they're just going to be on autopilot. That is not the picture of any gate. And also really you want your more senior people to have the humility of learning from people who are lower level than them. You want everyone to have the experience of ownership, of mastery, of teaching other people, communicating about the thing. I remember I credit my entire career sometimes to an overdeveloped sense of responsibility. When I was second life I'm like 19. Nobody wanted to do my SQL backups. So I came the fucking master of my SQL backups and I was the expert in those things. Next you feel good. Somebody's taking it seriously because they say, "Oh my God, I get to take care of the database backups. I have the keys to the kingdom." And I had to say, "I mean that is so much more motivating than being like, all right, we're all world class dude. We're just going to crush code together." Not actually what makes a team great. No, it's going back to your definition of high functioning teams is doing innovative, I can't remember your words now, but doing innovative things that move the dial for the organization. So doing new stuff. So the ability of the team to take on problems that they don't yet know how to solve. And I worked in, there's one team in my mind in particular that I worked with when we were building a financial exchange. And we just did that all of the time and it was so gorgeous. It was so exciting just to be working on new stuff. I had no idea we solved this problem. Let's get together and talk about it and try and figure it out. That's so great. So motivating. I also want to point out that there are lots of different sort of archetypes. There's no, like some people love high pressure. I am super ADHD. So I need, like I've never felt more alive than when the entire site's down and it's my response to be good. I can focus. So I've always gravitated to that. That is lots of people's nightmare. There are some people who would rather work 20 or 30 hours a week and be a consultant and be able to check out when they go. And that is fine too. I think one of the really core foundational elements of building a good high performance team is for leadership to be honest about doing some soul searching. Who are we? What does it take for us to succeed as an organization, as a business, as a company, as a team? Who is already succeeding here? Do they all look the same? That might be a problem. Because one of the aspects of diversity is that provides resiliency. There's no question in my mind that a monoculture can move really fast. A bunch of guys who all went to school together, had the same reference point, same touch point, same employment history, you see this all the time and still come back. But are the same mistakes? They started a company together. But at some point, you have to expand that circle. And everything just grinds to a halt while they try to work out who they're offending or what they're doing wrong or why people can't get along or the farther they go, the worse to crash it. I think when you point of being like, oh, this person would, we don't have any parents on the team. That's cool. We don't have some older people or younger people. And it's not like you're looking to construct a Benetan ad or pick one of each. But it's a fact that it's a plus when you bring other skills, other backgrounds, other ways of thinking to a team. You know, you'd always astonish. So I think our industry has an appalling record on diversity, on the whole. But just at the simple stage of certainly in Western cultures, the proportion of women that see it as a career these days is tiny. And I find that deeply depressing. But it's also stupid on the part of development organisations. It seems to me because half of the population of the planet are female. And they're going to have a different viewpoints, different perspective on things, different interests in how they want to use and interact with software sometimes. And I want that viewpoint as part of my team so that we've got perspective for more human beings. It seems crazy to me to miss those sorts of opportunities. And as you say, it shouldn't be a box ticking exercise or anything else. It's just a matter of. It's more like a unit test for your culture. Yeah, absolutely. It's just a matter of how do we do a better job? And we do a better job if we've got more viewpoints. And we can kind of collaborate to kind of come up with good answers. It seems to me. But yeah, crazy. Maybe that's not the place that we want to go. I don't know. You know, it's, it's, it's, it's, it's weird to me because I, when I got into tech, nobody was talking about women in tech. You know what it was? I remember my first conference was, was a, it was a, it was a usenix in like New Orleans. And it was like, never heard as far as, and I actually got into tech because I came from like a fundamentalist background. And I'm like, great. There are no women here. That's where I belong. Because I, I had deeply internalized misogyny. I didn't want any, so I'm just, I'm just like, I'm white dude, you know? Very problematic. Took me a lot of unpack it as an adult. But nowadays it's like, in some ways, it's like the future is here. It's just very unevenly distributed. There are so, there are so many people, Dave, over the past decade or two, who have made it a personal mission to mentor women and, and be your ally and their advocate and promote them. And there are some incredible places to work. I mean, no place is perfect, right? But there are a lot of really great, this is it, it's no worse to be a woman than it is to be a dude anyway, you know? The end of it and I almost feel like this doesn't get, sometimes I feel like the way we talk about the tech industry for women, if I was a young woman, I'd be like, oh God, that sounds horrible. And I don't think that's the reality. I think it's horrible in some places. I can, very much not in others. Yeah, absolutely, that's, I would agree with that completely too. So, I've seen that. I have good friends who are women in technology and they have a great time, they enjoy what they do, they're like me or not, nerdy, some of them are nerdy, some of them are not, but you know, they have a good time doing and enjoy their jobs. And you can do that. I wouldn't want to dissuade anybody. It's the best job in the world. We get paid money to send her back to the computer and solve puzzles all day. Yeah, yeah, absolutely, it's a fantastic job. It's just the relatively small proportion of women choose to do it and I think that's a shame. I agree. So, I guess this leads on to the stuff that we're talking about in terms of the team organisation. Another quote of yours is, you're talking about 10X engineers and how much you dislike that trope. And one of my favourite quotes of yours is, if you self-define as a 10X engineer then you're an asshole. I think most of us can probably agree with that. I like that. Yeah, every person I've ever met who was obsessed with being a 10X engineer, they were just like not conscious of the tech debt that they were sloughing off on to everyone else in their pursuit to be the best. Yeah, yeah. Yeah. I've met some people that I think probably qualified as 10X engineers. Oh, I have to. Which is what I think the thing probably exists, but I think it's a dumb target. It's a dumb target to aim for and I would agree with you that if it's the way that you think about yourself, think harder about yourself. Yeah. Yeah. No, I have met some. I have been in Silicon Valley my entire career and I have worked with some folks who are truly just prolific and they've been thinking about it constantly for 30 years and I will never be anywhere near that good at that. But it is an archetype, right? Some of those engineers, they actually function, they're like principal level, but they kind of function more like five senior engineers duct tape together. You know what I mean? And then there's some who are the same level who are not nearly as prolific. Like at Facebook, I remember there was a level seven engineer who wrote like five lines of code one quarter, the one of them saved the company $50 million. You know, it's like, there's room for lots of different definitions of excellence depending on what is needed by the team. Yeah. So I've got another, another, another quote of yours. Sorry, this is just me quoting you to yourself. I think it would be more fun for the audience if we really disagreed on something. Yeah, yeah, I don't know how we find that, but writing code, writing code is not the hard part. It never has been. The hard part of software is understanding it, maintaining it, extending it, scaling it, operating it, migrating it, refaturing it, crafting the right level of abstractions, instrumenting it and reasoning about it. It's the hard part. And that's what I think in one of my books, I tried to reclaim the term software engineering because I think it's got devalued. Yeah. I think it's got to mean either, I think in America in particular, it tends to just be a synonym for writing code. And that's not engineering. That's different, I think. And in other places, certainly in parts of the UK, it becomes a synonym for a bureaucrat approach to writing code. And I don't know if I can ask for right either. My view is that engineering as a discipline, broadly not just in software, is probably the height of human creativity. It's doing problems that nobody knows solving problems that nobody knows how to solve. And doing that in a way that's practical, functional and working to an economic target. It's a difficult thing to do. And it takes all of this stuff and more. And it's about this, you know, optimizing and tuning, learning and control to be able to do these things. And if it's, you know, if whatever it is that we're doing doesn't steer us in the direction of success, what we're doing is into engineering because engineering is by definition is what works best. What we know at this moment, what works best. And this is such a great description of why I don't trust any manager who thinks of himself as a shit umbrella. Because your job is not to protect your team and shield them from the realities of the business. It's to connect them with the reality of the business. Engineers are not code monkeys. Or rather, if engineering is truly the innovating engine of most modern companies, and it is the thing they have to be solving problems that move the business, they can't be geared to get monkeys, you know? Have you seen Greg or Hope's stuff about the architect elevator? I think I did, it's been a long time. Yeah, I think it's a really good analogy. Another of his ideas that I really liked was that you can kind of identify a culture in an organization by the reporting lines for the technical function. So if you're reporting to the CFO, then it's probably bureaucratic and the UCs of cross-center. And where you want to be, you know, sort of the generative end, generative culture kind of thing is as you're describing engineering as a seat at the table, it's a collaborator in driving the business forward. And you know, that's what makes a modern company successful. Yeah. It's so funny you bring this up. I'm writing the second edition of observability engineering and just last week. So the section I'm responsible for is a new section at the end, part six, on observability governance. And I've missed my deadlines a couple times now. And it's last Thursday, why it was I kind of ground a halt. And it's because addressing it just to observability engineering team, yeah, they need help. Engineers alone can't fix entire socio-technical systems. They can't approve budgets. They can't change reporting lines. They can't, you know, reallocate resources. They can't sign $50 million contracts, you know? And so I went, actually had a session with Clad, I dubbed so much, I'm like, Clad, help me reorganize all this material so that we're speaking to technical decision makers at every level from the CTO. You know, and I know they're only going to read it if it starts with the CTO and goes down to the, you know, the observability engineers. Like what do directors need to know? What if he's need to know? What is CTO? And what CTOs need to know is that probably observability is not on their top three, top five listed goals, but probably almost all of their goals are actually backed up behind shitty observability. Slow, you know, and one of my chapters talks about how if your observability team reports up to your CIO, you're treating it like a cost center, you're not treating like an investment. And one of the, when I wrote the first edition of this book, I was talking about observability in very technical terms, you know, it's about unknown, unknown, tie-cardon, now on the blah, blah, blah, blah, blah. Ever since AI came out and especially with the release of the latest door report, you know, they say over and over, they're like, AI, it's not going to fix your AI as an amplifier, right? This is not a tooling problem. This is a systems problem. They say that like 20 times in the first chapter, this is a systems problem. Every high performing engineering team is built on fast feedback loops and the foundation, the farthest upstream feedback loop of all is observability. It's a sense making capability of complex sociotechnical systems. And if it takes you an hour, a day, a week to ask or answer a simple question, I've been working on this list of almost like a quiz. How do you know if you're one of the tests that just walk around and ask about key metrics asking an engineer, do you know what, you know, what, what, this is, and if they turn to their tools, great. If they say, oh, go ask Tom, Tom knows, big problem, right? Everyone needs to be, if your engineers are, if your customers are really, if you're reporting most of your bugs instead of your engineers finding them after their bread. Big problem. You know, if your engineers are reaching for Excel spreadsheets, non-engineering tools to try and figure out any big problem, right? And like modern observability, the term has gotten so co-opted. I was talking with Rick Clark this weekend and he was like, it's just like cloud watching 20 years ago. He was telling you the story about how in 2008, IBM and peak cloud war time, right? They reclassified their Z series mainframe as cloud. And I'm, and he's like, this is what's happening today. It's too big of a problem, which too big of a budget. It's practically marketing malfeasance, not to market your thing as something to do with observability, but if it's not, it's not solving the system's problem, uniting, uniting the system's data, the application data, the user data, the business data, then it's all you washing. So I don't know you well enough to say this kind of thing to you. So I hope you'll forgive me, but you struck me as a, as a bit like me, a nerdy person. You're a technical person who enjoys the technicalities of the subject. And so I, the way that I see this is, I think this is deeply rooted in science. And so is part of a genuine engineering discipline. It's about that feedback. It's about being able to control the variables and be able to monitor the information that's necessary to be able to conduct your experiments. So you can kind of understand the result of your experiment, whatever it is. And so if you collect the data appropriately, then you can make other, other inferences from the data. It's like astronomy. We just kind of suck in all of the data and then decades later, somebody will look at data that was collected a long time ago and say, "Oh, that means this is true about black holes or whatever." And then you design a new experiment to learn something new or, you know, it's a science problem. It's a science problem. It's also a business problem. It's also a financial problem. It's, it's the, and something that I've been realizing recently is that we've been speaking so much to engineers and not speaking to engineering executives. And engineering executives are the ones that have to resolve the tensions of technology, business, and finance. And they don't have time to really most, some of them will, but most of them don't have time. And so we have to tell a story that's, you know, that has aimed at that level of attention. It's kind of going, it's kind of though part of the problem and going back to what you were talking about earlier in that engineers can't set the budgets or change the organization. All of those sorts of things, but CFOs or CEOs can't get the engineering insight to understand how to make this kind of change as well. It's about that collaboration and bringing the people together with the understanding, whatever understand they have to solve these genuinely difficult problems. It's so easy for an executive to accidentally kneecap an engineering order because so many of the things that it takes to do real, excellent, rigorous engineering actually sound kind of crazy and counterintuitive, especially if you come from financial background. Like look at so many of the problems that security has. Like it makes total sense if you're an accountant or in finance that you would never let someone admit their receipt and sign their receipt. That's just common sets. And from that, we get the rule that you can't have the same person write the code and deploy the code, right? And it's like, this is why a lot of a lot of a lot of VPs and directors that I hear from who are like, how do I'm like, look, if you're if your CEO runs a company where tech is what creates value, they have a moral obligation to at least read, accelerate the data's there. Yeah, yeah, absolutely. Absolutely. Absolutely. Yeah. But I think it goes back to where we started as well, which is the importance of communication. I think that part of the problem is on us. We need to be able, we technical people need to find better ways of communicating in ways that are amenable to understanding of less technically trained people. So if we're talking to somebody that's from a commercial background, then we should be talking in terms of the commercial impact of the decisions. If we're talking to somebody with a security background, we should be talking in terms of the security implications of the choices and not just saying, you know, dependency provenance or our technical jargon. Which is start by listening to what their problems actually are. Absolutely. And it's an important engineering skill to be able to have those come, you need people as part of these high functioning teams, at least some of them, if not all of them, to be able to have those kinds of conversations and to. And there's a certain engineering supremacism that has crept into certain tech cultures where we actually look down on our counterpart in the marketing department or the sales department and we just kind of, I think people aren't really that aware of it, you know, but the people under receiving end of that are absolutely aware of it. They're very sensitive to it as we all would be in their places, you know? And I think we really owe it to them, to ourselves, to try and drain that out of our culture and have more of a culture of curiosity. It's really hard to sell software. It's really hard to be a good, precise storyteller in marketing and not just like, you know, rely on flintsy tricks or, you know, which always backfires anyway. But I think if we respect our colleagues and if we listen to what their problems actually are and then try to speak to them in their language, we will build better solutions and they will actually trust us to do that. Absolutely. I'm sorry, I'm going to quote you again, but it's another quote that I really liked. The best engineering organisations are places where normal people can do great work. And I think that's what it's about. But normal people with the right kind of skill, so we need to breed, we need to build and encourage the cultures and the learning that helps people to grow to be able to do the most engineering organisations is that people work very hard, but a very small percentage of that hard work actually gets converted into forward progress for the business. Because all that energy is going to like, fighting with the system or like caretaking deploys or like fighting with each other or fear of their job, you know, they're expanding a lot of energy, but it's not converting efficiently into forward momentum. You know, and so like one way to try and solve this problem is just to try and hire the top 10% of geniuses in the world. And if that is your plan, I hope you plan on paying them top 10% salaries in the world, really bothers you. Or like $400,000 a year with higher best, you know, or you can try to build systems that convert that energy and effort much more efficiently. And you know, the greatest productivity tip of all is make sure you're working on the right thing. Doesn't matter how hard you bust your ass. If it's on something that never gets shipped, doesn't solve the problem, never gets used, doesn't have to follow through to make sure it actually hit, you know, and this again, this is largely not on the engineers, this is largely on engineering leadership to solve and that are rarely the ones in the firing line when they get it wrong and maybe they should be. But like, I think we need to hold our engineering leadership to a higher standard. I really think that like staff plus engineers should also be, I think that the future of engineering leadership is much more cross-pollination. I think that doing a tour, the best technologist I've ever worked with were ones that had spent some time as a manager, spent time as an IC, maybe done it a couple of times. And they just kind of wrote almost transcend the categories. Yeah, absolutely. I had an interesting conversation a few years ago now. I was talking to the CTO of ING, one of the big Dutch bands. Why not, ING? He's only clearly just went to work there. Okay, cool. So this was a few years ago now. And he was saying that the then CEO who was a traditional finance guy had been saying that the next CEO would be an engineer because that's what the bank was. It's an information. Any bank is an information company. So you better be bloody good at dealing with information and doing these sorts of things and thinking in those sorts of terms. But I don't think you have to have been an engineer to learn to think rigorously about these things. So the guy who's one of our VP of sales, as VP of sales, Mani, he's been at Honeycomb for years. He's never been an engineer. But I don't think you have to have been an engineer to understand the difference between a metric and a log line or structured data and unstructured data. You do have to give yourself permission to think about the concepts and understand them and reason about them. Not just parrot words around. But I think that a lot of non-technical people don't give themselves the permission to understand. I do agree with that. I suppose where I would slightly differ is that I think that depends on what we mean in terms of the terms. So I'm being lazy and I'm using engineering in that sense as somebody with a scientifically rational kind of training. I think that the problem is, and societal problem from my point of view, is that we don't treat science trainings as important enough. And so I think I've met lots of those people too who can be great working with engineering fans. I think it takes a certain skill in kind of rational thinking and reasoning. And that's I think having some kind of technical background, maybe not engineering, but technical background. I agree with you that science is underfunded, understudied everything. But as a serial dropout who has never graduated from anything, was a music major, I will always stand up for the fact that you don't have to have formal training to figure shit out. I think that's true too. So I don't have, I have some qualifications, but not a degree level qualifications in science. The level below in the UK, I have A levels in science. But I, science has been one of the loves of my life all the way through really, if I look back on it. And I agree with you, I don't think you need bits of paper to prove that you do, but I think you need, you need the ability to do that kind of reasoning. And I think that's open to everybody, but not everybody does it. And lots of people don't do it that should be doing it. Also true. And you would really expect and hope that anyone at the CEO or see anything level would be extremely good at it. And I would say that's fairly rare in my experience. Yeah. Certainly most of the organisations that I've seen. But yeah, I think that's true. Let's move on to talk a little bit about observability. So one of the things that interested me in this is you started talking about observability 2.0. And I wanted to talk about the difference between the two. I want to talk a little bit about the terminology stuff that we touched on earlier. So observability 1.0, you said that this was about how you operate software. And the observability 2.0 is more about how you develop software. And there's much more to it than that. But that was one kind of take that I heard you talk about. Observability was originally supposed to be something that's gotten co-opted right now. Everybody's doing it. I think anything that was built on the three pillars model where metrics go there, logs go there, tracing goes there, now profiling goes over there, errors and exceptions go over there. That's not treating your data like data. It's not, and no matter how much you try to, you can't put them back together again. And so I think those tools are very good at helping you achieve operational outcomes. Is it up? Is it down? or are there errors? Great. Which is what the metrics tools vendors talked about for years. But the only thing that they talked about. Yes. Yes. Yes. Yes. Nability was supposed to be further upstream than that. It's unified. It's just data. The irony is this was solved 20 years ago for business intelligence, for BI tools. When the columnar storage databases were first productionized. But for a long long time, it was so expensive that businesses would only put the gold treasure, the data that they most needed to run their business and their BI tools. To this day, you see engineers turning to BI tools to try and understand systems. But the cost of these storage systems and analysis systems has been following for two decades. And I think we need some language to help people understand that not everything being marketed as observability is the same. Because if it's scattered around in silos, all it can really give you is, I mean, people have been, those are very good engineers at your data dogs. When you're all at your, you know, your graphite, they're great engineers. You know, they've got way more resources in the middle. But the data model doesn't bring all this signals together in a way that lets you reason about things open endedly, exploratively. The ability to sort of slice and dice and this is interesting. Pull in the thread, zoom in, zoom out, look for patterns. You know, follow a trace like no dead ends, right? Just data and the companies who adopt this, they hook up fastly, but they feel frictions and then you play AI on top of that. And it just, it's like, it's like driving a fucking horse. You know, you're just like, psh, but if you're still trying to mangly sort of go, okay, well, my dashboard says there's a problem, you know, bunny hop over to my logs, try to find, I think there was saying that I know proof, you know, I want to trace, I'm going to hop over. It's just, you know, and I really, like, I, businesses got a business. Marketing teams got a market. But I think it's really, you know, I, Dave, this weekend I saw an observability panel staff with VPs and observability experts. And this guy said, I literally have, quote, he said, you know, our traditional observability tools, that stuff generally works quite well. Are there faults in our systems? Our traditional observability tools work quite well. But what we really care about is observing the quality of our product from every customer's perspective. Did you know there are lots of things that could interrupt your service that won't impact your lines? Did you know you could have a point in service render latency, you could have a card reader thought, you could have a mobile, did you know? And therefore, we invested in our own custom solution to measure key workflows. And my head is exploding. I'm like, fat, fat, that thing, that observability. There are traditional observability tools. That's money. Like, I don't even, what do you do at that point? You know, it's like, you even know. It's kind of my thing. So I think I talk about the same idea. But I use very different language. So I talk about it from an engineering point of view. And one of my five things, so I talk about engineering being a discipline that's kind of based on two different groups of skills. One of them is about being experts at learning. And one of them is being experts at managing complexity. So in, so in, so in, being experts at learning, we're kind of talking about iteration, feedback, incremental development, experimentation, and empirical learning. And those things in the field, the experimentation, empirical learning, and the feedback are kind of at the heart of what I think of as, as the observability part of the problem. So we want feedback on what our users make of our software. And part of the problem with that. So it's a one, one of my bits of kind of sound, bite style advice is work experimentally. So treat every change to production as an experiment. And if we're experimenting, we need to identify what we want to measure. And so what do you want to measure? And part of the complexity here is, seems to me, is that this isn't something that's specific to the products that we're building. It's worse than that. It's specific to the feature that we're building. Feature by feature, we've got to decide what the measures are that matter. And we've got to think, as engineers, we've got to think in terms of how do we manage the variables sufficiently? So if you can understand the signal that we're getting back. And all of that seems part of the problem to be able to get to it. But that seems to me is, it seems that that's that we are talking about the same thing when we talk we talk about that just in different ways. Yeah, different parts of the elephant. Yeah, absolutely. I would have said it's the same, same part of it, just different. Same elephant. Same elephant. Okay. You're talking about the engineering workflow and part. And I'm kind of talking about the tools perspective, but we're talking about the same thing. So, so, so, so, so, I think in terms of the information flow, I think we're talking about the same thing. So, so, you know, if I'm an engineer and I'm building a feature, that's I'm hoping is going to recruit more users. Then I'm going to, I want to be able to measure how many users I was recruiting and how many I'm recruiting after my feature goes live. So, that's part of the design of the feature. So, that's part of the design of the observability of the absolutely. Absolutely. So, so, so how I achieve how I do that the mechanisms that I use to collect that data and to interpret it. Maybe, maybe that's what you're talking about as a different thing. But in terms of fundamentally what it is that we're trying to do, design, I guess what I'm talking about is how do we design for observability? How do we design for observability? Yeah, this is more of a more on the socio part of the socio technical, I think, because there are a lot of different ways to do this. I feel like the gold standard is a feedback loop where engineer is instrumenting or code as they go. It goes out and they go and they look at their instrumentation that they just wrote and they ask themselves, is it doing what I expected to do? Does anything else look weird? Like, those are the two basic questions, right? Do I, you make a prediction and then does it come true? But, you know, Dave, the longer I do this, the lower my bar falls for this. I think I'm excited about the prospect of get pre-committed hooks that do like some little AI magic and auto instrument for you. I'm really excited about some stuff. We're building a honeycomb that will, you know, sort of not make developers have to go look for the data, but bring it to them and they're slack or they're cursor or something, just be like, hey, you just shipped your code. It just won't lie. Here are the key, but yeah, because I really do think that like I've been a pierced about this for so long and I just don't see the world making progress. And I now think that anything we can do to help connect that fundamental feedback loop that includes people writing the code and production systems will be a benefit to the world. So, so one of the things that one of the things that I, the ideas that I like that I don't hear very many people talking about is applying kind of SRE style thinking to this part of the problem. So, so the part of the problem with Google is that Google's big problem is the traditional uptime and all of those sorts of things. So, whenever they talk about SRE, they're nearly always talking about those sorts of measures. Yeah. And that's okay. Those, those measures are useful, but I also want to know, did I recruit more users? Did it make more money? Yes. And all of those kinds of things. So, so ideas like service level indicators and service level objectives. Thanks to those kinds of things. Ridical for organizing. Well, and like I wanted one of the things in my chapter is the SLO test. Do your SLOs actually drive behavior when your SLOs fall below a certain line? Do your product has your product team agreed that your engineers will start working on reliability things when you know because that does it have teeth or not? That is the ultimate test of whether or not your SLOs are just window dressing or they're actually they actually mean something. Yeah, yeah, yeah, absolutely. But, but also about the business function of our systems, not just the technical performance. And asking those questions of ourselves is surprisingly hard. Which is kind of the point. If we don't do that work, we don't we aren't aligned. We are going on the right thing. And that kind of closes the loop of what we've been talking about all the way through this is you know that's you know how we move the dial on our organizational performance. You know that's how we make our organization be serving our users better. You know that's how we learn what better means. So two point O observability two point O again I'm going back to you. I'm kind of back off. I'm kind of backed up. So I realized after I started talking about one point O and two point O that some of the vendors who I was calling one point O got really offended by it. And you know I live my life by the principle of calling people what they want to be called. So I I'm making an effort to not call it that instead I refer to it as the multiple pillars model and the unified story model. You know I do have one data set that brings together the braid of all the signals or you're sending every signal to a different location. I don't think I haven't heard that anyone's offended by that yet. So that's the language that I'm going with. So I'm I'm perfectly happy to adopt the language that you prefer for this. So so so when I when I look I don't think this is what you were talking about. But when I was listening to that that reminded me of one of my favorite architectural styles for big big complex systems which which is event streaming is it and and when you were talking about the log I was thinking of the log of events because one of the things that I like about that model architecturally is that it keeps the history tells the story. So one of the things that we did when we built the LMAX system was that we were able to record production and play it back. So that's a form of observability. So if we had a defect in production we could play it back until we got to the point and and have a break point in our code when we hit the point then we could carry on debugging it in the tested environment. We could run the node. Plugging and fraud. That's great. Absolutely. But but we could we could do it not in fraud. So no no I get it. Yeah. Yeah. So so so so so I think that to me from from my engineering model point of view that to me sounds a bit like trying to improve the determinism. So trying to control the variables so that your observability is more correct. So that you could otherwise you got the distribute. I think what you're referring to when you're talking about the the problem of the logs and the traces and all of those and and then not fixing up is the the the plastic distributed computing problem of information in different places and you don't have to correlate. Yeah. So how you stitch it all how you stitch the picture altogether. You can stitch it all together. Like metrics are aggregates at the time that they're written out. You know the heat maps are bucket right. The the traces are sampled. The logs might be sampled too. Are they sampled the same way? Probably not. You know you can't stitch it and and some vendors have tried and there are some in group. You know they they make it so you can predefined that this metric is derived from that trace field or something but those are all predefined and the cost multiplier is a obscene because for every pillar you're paying to store your your whole data all over again in a different format right which is why you know could the cost of observability. People are companies are now paying roughly 20 to 25% of their entire cloud bill goes to it's their second largest bill now it's everybody still climbing 30 40% your over year for the past 15 years. You know Gartner released some data they showed that like in 2009 this company a particular company is a customer there was paying $50,000 a year for the observability monitoring right and in 2024 they paid $24 million which is a 48% you know and it's like some data models are just more efficient than others. It's better to pay this for the data once you know then to pay this for the data five six seven times and then have to pay to create craft bridges between them you know but code and production tools have a very they have gravity data has gravity right yeah it's in the sub companies have a very explicit strategy of trying to attach even more gravity to their tools and it's sadly a pretty successful one so. So I think I really enjoyed the conversation but I noticed that we've been talking for over an hour so far although it's whizzed past for me for me anyway but I've got one more one last quote of you to yourself that I really liked which is the biggest obstacle between us and a better world is when we don't believe that one is actually possible. That's I like that idea a lot and I like your framing of it. I think I have a different version of that too so I think that one of the problems in software is that a large proportion of people in that working software development have never seen what a good software project looks like. Yes exactly exactly and and so they start to believe that yeah you have to be just a super team. You have to be in Silicon Valley that doesn't apply to us that's not possible for us our engineers aren't that good this requires some and it doesn't it is actually so much harder to work as a software engineer with a two or three months deploy cycle than it is to work as a software engineer with like an hour it is harder you have to contact switch and hold things in your head and gas and remember it is harder it you convert so much less energy to forward momentum. Shortening these cycles makes it so that you can hire worse engineers and have better results but people lack the confidence and they don't believe that it's possible and so they don't try. Yeah so I think we both of us feel a mission to try and help people to see the way to the better world. So I think we should probably wrap up so first of all my thanks to you I've really enjoyed our conversation today it's been a lot of fun and thank you for all your writings and talks and all of those in your books and all of those things because I've enjoyed them a lot it's been a lot of pleasure. For viewers or listeners thank you very much for getting this far and if you've enjoyed the content please hit like and if you can subscribe to future podcast if that's available on your podcast platform and thank you very much for listening hope to see you again soon. Bye bye.

Podcast Summary

Key Points:

  1. High-performing teams focus on communication, collaboration, and collective ownership of software, rather than individual "10x engineers."
  2. Fast feedback loops, observability, and continuous deployment are crucial for learning and improving software development practices.
  3. Diversity in teams enhances resilience and innovation by incorporating varied perspectives and skills.
  4. The rise of AI emphasizes the need for engineers to focus on problem-solving and durable code, not just code generation.
  5. Effective leadership involves creating systems that support growth, psychological safety, and shared responsibility.

Summary:

In this conversation, Charity Majors discusses key principles for building effective software teams and development practices. She emphasizes that high-performing teams are defined by their ability to collaborate, communicate clearly, and collectively own software, rather than relying on individual "10x engineers." Charity highlights the importance of fast feedback loops, observability, and continuous deployment to enable teams to learn and iterate quickly. She critiques the industry's focus on hiring top talent alone, arguing that systemic factors—like enabling frequent, safe deployments—are more critical for team growth and success.

The discussion also covers the value of diversity in teams, noting that varied backgrounds and perspectives foster resilience and innovation. Charity addresses the impact of AI, suggesting it shifts the engineer's role toward problem-solving and maintaining durable code, rather than just writing code. She advocates for leadership that builds psychological safety, encourages curiosity, and supports incremental, evidence-based development. Overall, the conversation underscores a team-centric, systems-oriented approach to software engineering that prioritizes learning, collaboration, and sustainable practices.

FAQs

The Engineering Room is a monthly series of long-form conversations with influential people from the software world, sponsored by Equal Experts.

Charity Majors is the CTO and co-founder of Honeycomb, a former production engineering manager at Facebook, and a co-author of books on database reliability and observability engineering.

She defines high-performing teams as those that spend most of their time working on interesting, novel problems that move the business forward, emphasizing collaboration and collective ownership.

She values communication skills highly because if someone can explain how they solve a problem, she believes they can code it, and collaboration is essential for learning and progress in teams.

She criticizes the focus on hiring only the top 0.1% of engineers, arguing that great engineers are developed through experience and that team systems and feedback loops are more important than individual talent.

She sees AI as useful for generating disposable code but emphasizes that durable, stable production code requires human oversight and incremental changes based on real-world testing.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.