Go back

Why software factories fail without humans | HumanLayer, Warp & LinearB

58m 39s

Why software factories fail without humans | HumanLayer, Warp & LinearB

The panel discussion centered on the emerging software factory model, where AI agents increasingly handle development tasks while humans shift to oversight roles. A key theme was measuring success, with panelists advocating for tracking both business outcomes, such as predictable delivery and customer value, and technical indicators like cost per effective PR merge and rework rates. They emphasized that humans still own code and PRs, recommending governance policies and auto-merge gates for low-risk changes to maintain quality and reduce reviewer overload. The conversation explored what can run autonomously, noting that human gates at critical points like spec and code reviews are features, not bugs, and that maturity varies by trust, with incident management being a particularly challenging area due to permissions and latency. The panel highlighted the need for unified observability across the SDLC, as current tools often fragment context from plans to production, creating inefficiencies. They also discussed engaging engineers by shifting their focus to product ownership, systems-level thinking, and maintaining the factory itself, which can be motivating. Overall, the discussion underscored that software factories are not about eliminating humans but about redefining their roles to maximize leverage, ensuring quality, and avoiding the compounding of "slop" that harms future development. The panelists shared practical insights from their experiences building and implementing these models, aiming to help leaders navigate this transition effectively.

Transcription

11537 Words, 60810 Characters

English
[MUSIC PLAYING] Welcome back to Dev Interrupted, brought to you by Linear B. Today, we're tackling the emerging software factory discussion in the industry. Engineering leaders everywhere are moving towards a model where AI agents lead the way, while humans move increasingly to the edges. And it probably says massive productivity gains, but brings up a critical question. How do you actually know what this kind of factory is working? Are the agents driving value or are they quietly generating a mess and hiding costs? And these are the questions that more and more are asking as we start to explore this topic in our industry. So Linear B and Dev Interrupted brought together a panel of experts to break down the context layer needed to govern these kinds of workflows, what signals you need to instrument first, and how to tie it all to impact an ROI. Here's our discussion on the agent tech software factory. Today, we're going to start by talking about the software factory is here. And maybe we're not using that language yet in our own teams, but we're certainly using many parts of it. And we're going to do intros for our amazing guests that we have here. And I pulled out some questions just from the industry about what people are asking other leaders and each other about a factory and what makes it successful. If we have some time at the end, I'd love to pull some Q&A out of the audience. So please be dropping questions. Please be chatting with each other. This is a software factory safe space for us to nerd outshare resources. So we'd love to just see lots of activity in the chat while we're blabbing away. So while we dive in here, just before we do intros, we're going to run a poll. Because I would love to know just from folks that are here, where you currently are in your software factory journey. So how much of your SDLC already runs like a factory? So you're going to get a pop up now. This is a Zoom poll. And if you can just go ahead and flag for us, just generally where you think you fall in the software factory story. And I'll ask my panelists here with me today, maybe look at the question too, and think about what your own answer might be. And we can use that to kick off with our introductions. So folks, just go ahead and answer the poll. Again, just taking a temperature check on how much of your SDLC is already running like a factory. And now that we have that running, I'm going to go ahead and jump into our introductions for the day. So I'm going to just start in the top left and hand it over to each of them to introduce themselves, tell them a little bit about what they're working on, maybe even answer the question, about 90 seconds each for each other. So I'm going to start with a look. A look would you like to introduce yourself? Of course, thanks for having me. So my name's Locke. I'm Warp's founding engineer and head of product engineering. I've been at Warp for a long time, six years through the battling agents for quite a while now. It really started with interactive agents as we all did. And now I spend all of my time building software factories and really helping companies build factories where they can show and demonstrate ROI, really help them improve it of time. Whether that's quality, cost, or some combination of both. Super slight for discussion. Amazing. And Dan, I'm going to hand it over to you next. Awesome. So I'm Dan Lines. I'm Linear B founder and COO. Before that, I was a VP of engineering. Now I spend a lot of time at Linear B. I'm working with CTOs, VP's of engineering. Obviously software factories is like the hot topic. So we talk about it a lot. And I would say there's a wide range of what I'm seeing out there right now of where everything stands with software factories, even though it's like a terminology that's probably been around for like 50, 60 years now. So thanks for having me on the panel. And I think we're going to have an awesome discussion. Let's go. Amazing. I know the results popped up. And we're going to get to those in just a moment. But Dex, I want to hand that to you to go and introduce yourself. Let me get off mute. What's up, y'all? I'm Dex. See you on co-founder of a company called Human Layer. We build a multiplayer coding agent workspace. It's got building blocks for software factories. We spend a lot of time working with teams solving hard problems in complex code bases. So hundreds of repos, hundreds of engineers, 5, 10, 15-year-old code bases. And how can we raise the floor, make sure that everybody on the team is able to leverage agents for the things that agents are good at. And as the name might imply, ensure humans are still installed at the right points of the workflow in a very high leverage way, where we can get the most out of AI while maintaining really high quality bar and maintaining commitments we've made to our users and our customers and our customers' customers. And of course, we do plenty of software factory building and experimenting and trying to push the frontier of what works and what's hype internally. Amazing. Well, as y'all can tell, we have some amazing panelists here for you. And one thing that really resonates across all of them and stood out to me is how much they really care about software, making it great and secure and at scale. And I think that's what we're going to be talking about today. But we're also going to be addressing some of the pains and confusions and uncertainties of that world. That way, we can all get on the same page. Because I can see here, just on the pole of results that have popped up, I believe all of y'all are seeing this as well. In the distribution, it looks like most folks fall around a step or two of their SDLC running without a human involved. That means that some part of it, maybe it's co-gen, maybe it's a review process. Maybe it's something in between. There's just not a human in the loop there currently. And you can see as well that there's many folks that are interested in pushing that further. They're testing how far can we take this kind of automation across the SDLC. And to those of you that are just starting to look at it, don't worry. You're nowhere near behind on this journey. In fact, I would say that you're way, way far ahead for most folks. Just because this is a really early topic, and you're just in like a bubble right now of folks that are really nerding out about it. And just the last one here is myself, Andrew Ziggler. I'm our host at Devon Terrupted. I'll also go to Market Engineer here at Linear B. So again, we just kind of recapped the idea of where everyone is on running a software factory. Are you calling it that? Do we have parts of it within our own world? And as we saw from the poll, the reality of that is, yeah. There are many parts of this that we're acknowledging now are more like a factory. So today, we're going to be diving into how do you talk about success and the impact of that factory? Because many engineering leaders, I know that y'all here in the audience are really, like, you're trying to optimize and build for the future, but you also have to answer to budgets and you have to be able to convey to leaders who are non-technical why we're doing this and how we're actually making progress over time. So with that, I would really love to start with our first question. And this is about how do you measure successful outcomes in a software factory? So I'm actually going to stop sharing so that we can see each other's faces and have a nice, lovely discussion with this. And Luke, I'd love to hand this to you first to go ahead and open us up about successful outcomes in a software factory. - For sure, well, I think you think about a factory regardless of the domain. There's kind of two things people are optimizing for, right? Whether that's code or product, building a car, building whatever. It's really like, what is your input? How much product are you making? And what is your cost per item? And I think that's true in the fact a software factory world is as much as true in any other factory. And so what I think about successful outcomes, it's really like, can you measure and demonstrate that you're increasing throughput and maintaining reducing cost per item? And what cost per item here is like, what's your cost to build product, ultimately cost per PR? And a group of it is how much product you're shipping and how many PRs you're shipping. So when I think about success, I think about those two metrics, I think there's a lot of kind of downstream metrics that are better proxies of I'm actually tracking that than maybe those two top on ones. And then also just general quality of product as well. I think the mistake that using a software factory can have if you do it poorly is you start shipping a bunch of really, really low quality code that has a bunch of rollbacks, which just doesn't work reliably. And so to me, the like, what a success it looks like is being able to automate more and more development that continues to maintain kind of high quality and gives the product an actual code that backs it. - Amazing. Dan, would you like to respond to that? I saw you come off mute. - No, yeah, of course. Like, I'm happy that we started here, right? How do you measure? That's actually the first place to start. So when I'm working with like CTOs, VPEs, usually they have a lot of pressure coming from their business. It sounds something like this. Hey, this AI thing is happening. All these other companies, maybe even our competitors, they say they're moving way faster now. What are you gonna do about it? And so I think the first thing on the measurement side, and this is what the business wants, it's all around business impact. You gotta have some type of measurement so you can go back to your business to say, "Hey, we actually are providing more value "the customer's faster." So at linear B, and you can choose what it is. But at linear B, we always start with predictable delivery. Are you delivering customer value on time more often than you were before? And then there's the capacity side, which is how much of that value are you able to deliver? And of course, we track that way that you know, you can do stories, you can do tasks, you get all the project delivery stuff. And once you get past that, 'cause that's a conversation that you're having with your business, I'll add on to what a Locke was saying. Now you can actually look at leading indicators for your factory. Now we have a framework around this. Our framework's called APEX, but I'll just like rattle some of the leading indicators off that I like to use. Cost per effective PR merge, gotta track it, right? That's showing, hey, is this AI actually effective for us or not? And I like to look at merge rate per developer. Then we look at rework rate. Then we look at assisted PR. So there's kind of all these leading indicators, okay, that we track the show, are you actually moving to a factory or a dark factory? I'm sure we'll get into that kind of stuff. And that's what we like to track, but at the end of the day, your business is gonna come to you and say, hey, is there a business impact? Is our customers getting more value? So we like to balance it with both. And that's usually like where I go and what I see is like most effective. - Dex, anything you wanna add about successful outcomes? - Yeah, I think it makes a lot of sense. I think measuring business outcomes for customers is like really, really hard. And the teams that do this well are like cleaning up. One team I know that is like very well-known for focusing a lot and doing the like product engineering thing of like, hey, we don't give developers tickets, we don't give them stories, we give them customer problems is GitLab. I used to work with a bunch of ex-GitLab people and it was very much like, hey, I don't care how many lines of code you shipped, I don't care how many PRs you ship, I don't care how many linear tickets you ship, like none of the, whatever your system is, probably GitLab issues in that case. But it's like, cool, can you measure the impact? Did you make a workflow that used to take 20 minutes? Now it takes 10 minutes. Great, that's what we care about. It's not quite the like top-level KPI of like, it's the business growing, are we closing new logos, are we making more money? But like, how do we measure customer outcomes? I think the like, there's like this whole hierarchy of like, okay, if you can't measure customer outcomes, then maybe measure like number of featureships, output throughput, if you can't measure number of PRs or whatever, for whatever reason measured lines of code, you can't measure lines of code measure, the amount of tokens shipped, but again, like the amount of tokens spent. And since we're talking about factories, I will, I will mention a book that I love from the 70s called The Goal, which was about optimizing like physical factories and basically tackling this problem that we had in every single factory in the US, is you would bring in a bunch of MBAs and you would optimize like one particular, they would each optimize like one particular station in the factory of like, hey, the thing that stamps the steel things, like I'm gonna make that as fast as possible and there was like less emphasis on the overall throughput. And I think we're seeing examples of that now in modern software factory building where people like, cool, how do we make sure we're spending as many tokens as possible? How do we make sure like the review thing is like super optimized without zooming out to the overall throughput? So I think like a lot of rediscovery of like problems we've had before and obviously we've applied this to software factories since before AI, because the term software factories as old as 1968 or something. The one thing I think is interesting in that metaphor of like software factory versus like physical factory is in a physical factory, you're making parts or you're making cars or you're making golf clubs or whatever it is. But the thing is like once the thing is done and on the truck and leaves the factory, it no longer impacts future work and when we're building software every change that comes out of the factory is compounded into an impacts every future change. So if you ship a bunch of slop six months ago that is going to or bad code or whatever, things that are hard to maintain, that is going to impact every future thing you do until you fix it. So that's I think a really important thing to highlight beyond just like the throughput today but is like how does the way you're approaching solving problems today impact your ability to solve problems in a month, in three months, in six months? So within the chat is chanting slop loop, slop loop. So I think that's a really smart call out is that if you compound the slop on the slop you're going to make more problems for yourself later and it's not having that like rigid attention to detail not just on the individual stations but on the how the whole thing is moving. So I think that's a really sharp insight. I also, I would be remiss to not call out that we've almost made it through without bringing up the phenomenon of token maxing which is just something that I feel like I've covered in every conversation I've had in the last two months. And what's fascinating is you mentioned facts that like if you can't measure so and so measure this, if you can't measure that measure this and then you end up on tokens and what we found is that like so many orgs decided to really latch onto that as the first like wrong of a ladder is a really great way that I've heard that put recently. There was this really great op ed from Luis Morales. He's the head of AI at super.com. He posted last week and he talked about exactly this about how most orgs they stepped up to like oh, we'll just measure how many tokens we use, right? And then they didn't actually step further up the ladder of like how do we then measure the successful outputs? And then from there how do we measure the successful outcomes? Like he sees it as a three wrong ladder. So I think we're all definitely circling around the same idea here. So I'm gonna go ahead and share my screen again and move us into the next question. 'Cause this one's juicy. This one is something that's constantly top of mind for me. And it's about ownership. Because in a software factory, by definition, a large part of it is that automations are creating and driving the code. And as we know, a PR up until very recently was a very deeply human process of like you wrote your code and you opened your PR and you went, you found your coworker and you slacked them or something. So please review my code. And then they said looks good to me. And everyone was happy and it got shipped, right? It was like a deeply like camaraderie experience. But now that's been so fully abstracted. And that code review has just ended up in a really weird position in the modern SDLC. And I'd love for us to talk about it. And who owns the kind of code in this world? So I just wanna open this up to whoever maybe wants to take this one first that they have a strong, strong leading thought. - Yeah, it's like, I think you're probably pressing on a pain point here. Maybe like the question behind the question, like, hey, this area of the process is maybe seeing a lot of pain as it's like super easy, the generate code and all of that. Now, okay, what's the next step in the process? Probably this PR area. I have like a data point here. This came from the Linear B benchmarking, our benchmark report that we did. So 37% of agentic PRs get merged if your organization is classified as like less mature in the ADLC. So it's like a super low number. These are PRs that are like fully created by an agent. And it presses on this pain point, what's the reason that they're not getting merged? Well, probably no one knows who owns them anymore. Kind of like what you just said, Andrew, am I responsible for who's agent is this, who's the reviewer and that type of thing? So here, I'll just try to give like some tips around what I'm seeing. As you get into the AI DLC or the ADLC, have a plan. Like again, the best CTOs or VPs that I'm talking to, they actually have a plan of intentionality of how are we gonna handle these agentic PRs. Now that humans are like less in the loop, let's say. And the other thing that I see them do is say, hey, the volume of PRs is like increasing like crazy. Like as you get more into that software factory, it's easy to create codes. There's so many more PRs. Let's not overload our developers, have a plan for that. So I can just like, I'll spit out like a few things that I see our customers doing and see where the conversation goes from there. This side on an ownership plan, maybe have it automated. Would linear be, we do it with GIT stream. Automatically assigns ownership. It could be based on like code base area, it could be based on like feature requests, what area of the product. Just make sure that you have a policy and you have governance around it, like do something. And then the second thing, I'll just say like a few things, think about workload reduction. There's so many more PRs now. So what's the plan in place to reduce the workload on the reviewer? Like obviously have an AI code review, but I would say, hey, what I'm hearing, not all PRs are created equally. There's some that are very high risk. There's some that are very low risk. There's someone somewhere in between, have governance. Identify, hey, this PR, super high risk. Let's make sure there's a human reviewer. Another PR, like documentation only change. That's not waste time. Human time is like super valuable now. So like my overall thing here is like take away from this question, make sure you have a plan. That's the number one thing. Because if you don't, PRs aren't going to get merged and all of that AI stuff you're doing. It's not going to be effective. It's going to cost you a lot of money. I think the thing that Dan is circling on is that humans are the ones who still own PRs. I just want to make that like super explicit. I think there's kind of a misconception that when you have a software factory, it's a fully dark factory. And it's going from linear issue to PR merge with no human and loop at any part. That's not what I think of as a software factory, and that's not only way of deploying a software factory. It's like in my mind, I think of factory where an agent realizes it needs a spec and needs a human to review the spec or an agent assigns a PR to a human to review is a feature not a bug. That's how you're going to get better outcomes as opposed to an agent kind of ripping and writing code that no human reviews until the very end. And then a human has to be like, this is completely wrong. And now I need to go spend a bunch of time to course grind it. I think what you think about in those terms, it gets engineers to kind of think of themselves as like develop productivity engineers who are maintaining a factory as opposed to having to get an influx of PRs. They're just sprayed on them where they have no context and what to do. And I think that's very, very exciting for engineers as opposed to like, I don't know, immoralizing. And so what I would suggest is make like figure out the right steps for human gates in your factory where you want to pull in humans. At war, like our opinion on software factories is that a spec and you can define the criteria of like ambiguity that requires a spec should pull in a human up front and kind of grill them around what are some of the ambiguity points and how can I resolve them before it writes a formal spec for a human review similar with code. And then humans are responsible for reviewing that code, but also like continuing to improve the factory when you get a PR that's really low quality. You realize code is merged. I was like terrible for whatever reason. And that obviously means like directly addressing the problem, but thinking about it more from a systems perspective, what are the skills need to write to resolve it? Was there a change in model? Like, whatever you need to do at a high level to continuously improve the factory is something that is explicitly, in my opinion, their responsibility of humans. And I think that's how you can get one more people excited about working this way. Yeah, I think I totally agree that every piece of work needs a human owner. I think one thing, DAX, the co-founder of Opencode had this post a while ago, which is like, you will have a bunch of engineers. If you have a large team, you probably have some engineers on your team that are using AI to do the same amount of work with way less effort. And they are slinging piles of slop and then just waiting for seed, like shifting the whole burden of like your entire job, which is to create good working software onto the senior engineers, which is going to burn them out really, really quickly. And the people who actually do care about the craft of software, if you allow this to go on, we'll burn out and leave and go find a place where everybody cares about creating good things. My kind of like sound bite on this is everyone who says, oh, we have too many PRs. We have to solve the too many PRs problem is like, you probably don't have too many PRs. If you have this as a pain point, you probably have too many bad PRs because a good PR is a joy to review. If something needs like less than 5% rework, that means that like, I'm reading through it and I'm like, whether I wrote the spec and the agent wrote the code or someone else wrote the spec and I reviewed it with them or just like, they did a good job of steering the agent so that the output code or even like, polishing it before they sent it to anybody else. And I'm reading through. I'm like, yep, this makes sense. This follows our patterns. This is what we agreed on. You've headed off all of the tendencies of agents to cut corners and do things that don't work very well. Like, that's great. I would review a thousand PRs like that and it's not painful. What's painful is I'm reading a bad PR and in my head, I'm thinking like, oh God, okay, this is going to have to change. And if a human is owning it, then I'm going to have to talk to them and they already put in all this time and effort to polish it, whether they read the code or not, but put in time to make sure it's working and it kind of like solves the problem on the surface. And like, that is just a huge both emotional and intellectual burden on the reviewer. And if there's a human behind it on the on the submitter as well to receive that feedback and react to it and realize like, oh, it's not down. I got to go fix it. So I don't know. The only other point I'll make there and then I'll I'll hand it back is I have talked to three CTOs here in San Francisco all working on like, you know, I think it was like a hundred plus person team, maybe like 30, 40 engineers who have found this weird interesting thing where they created a very tight gate of what can be auto merged. So it's like, okay, if it doesn't touch any of these parts of the code base and it's less than this particular size and a couple of their criteria that can kind of be determined deterministically or maybe with a small model of just doing a classifier, then it's allowed to be auto merge. If the checks pass, you don't need a review. And that actually was what they found was like, it kind of solved the bad PRs problem, but it also caused people to change their behavior when engineers realized that, okay, if I follow these patterns for like, what is an auto mergeable PR? The PR numbers went way up because people stopped just like lobbing giant piles of slop up and they're like, okay, if I can keep this under 200 lines, and I can solve it in a way where it's like, it's not, I don't need to modify the critical core systems. Then all of a sudden, my job gets like the my least favorite part of my job, which is waiting for someone else to read my thing and pick it apart, like, doesn't exist. And so there's like, we'll see how the risk and the downside there of things that make the code base harder to work in slipping through despite those rules, but it seemed to me like a really interesting lever. If you want to change people's behavior, you create the right incentives and those incentives are aligned with the incentives of the business, which is like, keep the code base high quality and avoid downtime and incidents. And also like, the incentives of the developer, which is like, okay, cool. I want to ship. I want to just keep moving and I want to produce value. Yeah, I would double down on that. Like, I would say our most, when I would consider our most advanced customers or maybe like the ones running the most autonomously, they have, I mean, policies kind of a lame, lame word, but whatever word you want to use, they have a policy. Just like you said, Dex, in place, it gives good incentives back to the development team. Hey, if you want to get your stuff merged, you have a way, like auto merge, you have a way to get there. Small PRs, low risks, not much of a call it. I don't know. PR churn or however you want to say it. It kind of gives like a path to that. Hey, I can move fast. You get good dopamine. It's get your work done. That type of thing. Like, I've seen it work super well. And the only other thing that I would say like shout out to the panelists here, I'm happy that we're able to talk about quality. I think with, you know, the terminology is slop now, but yeah, if I'm just like producing a bunch of junk, nobody wants that. The other metric that I see our customer sometimes use is PR maturity. So the PR maturity score is basically, hey, when you open up a PR, how many times did we have to go back and forth on it? Was it like in good shape, or did we actually, yeah, you open, you know, or the age and open the PR or human open the PR, and it's super immature and actually took us more time on the PR than actually like doing other stuff. I think the PR maturity metric is one that I would say is on the rise. And I think it's because of everything that we're talking about here. Let's remember quality matters after we got super excited that it's easy to build code now. You know, make sure you're on top of it. Your AI software factory is shipping more code. But is it actually delivering more value? LinearBee's new guide your software factory needs a context layer shows you how to find out. It breaks down why AI adoption alone isn't enough, how to measure effective PR yield, and why cost-perfective PR may be the metric that your executives want right now. You'll also get practical ways to connect source control issues, CI deployments, and AI signals into one feedback loop. Download the free guide from LinearBee today. Amazing. I'm going to move us into the next question here because we're already starting the two at the corners of this. I think Alok even mentioned this moment ago, and now Yolvo is chatting about it. So, you know, this really brings us to the next natural point of we talked about the ownership. Obviously, humans still own the responsibility of the code and the decisions that are happening within the factory. The responsibility is not truly getting abstracted. The ritualistic format in which we used to do it is just getting changed. But the real reality of owning that isn't any different. But once you acknowledge that the ownership from humans is definitely still in the picture, there's still an opportunity here to identify which parts of this can we create machinery and mechanics around identifying and safely operating at scale with automations, creating things like these auto merge gates, which really kind of gamify a golden path. We tell them we want, you know, it be this a number of line of codes and this kind of changes across these kinds of files always be prioritized or refactor over new code and you create these kinds of really kind of like a rubric, like a grading system basically. And by really acknowledging that there's criteria for everything that you're merging, you can start to create these unique gates for different types of security and automated reviews around things like docs that don't need a person, right? So like in this world of a software factory, how do you think as a software leader about identifying what in here can actually run in the dark? And when I say in the dark, I mean no humans involved from beginning to end, maybe you want to squeeze and squish that definition a little bit, maybe I'll let you, but if you have a strong idea on it, definitely please lead the way. Who would like to go first? I can tell you how we measure it at war at least, which is we don't think about it as dark versus not dark, just like Echo and I said earlier, to us it's more about how well the factory running and one is a human have to step in to hand hold an agent that's kind of gone to bri and I think that's a better metric to track so we call that autonomy and for a warp that's like we've got the 80% autonomy score which is for 80% of work that goes through the factory a human never has to kind of step in and commit code themselves or with like another agent to get something past the finish line. I think that's more valuable than measuring dark for the sake of dark because at the end of the day what you care about right now is like high quality product and I was just trying that as high quality code and that means to some extent you have humans in the loop at the in the right places and then as a engineering leader it's like on you to define when do you want to human the loop and like to some extent that's that's going to be flexible and like a ton of different dimensions like what's your stack a lot of our code is written in rust so it's not something that an agent can very easily test locally for example at least we don't have hot reloading a lot of like key pieces for a nation to test it but when we have a code right in the front end it's obviously like pretty high-level you're very very easy to change and so the kind of like bar for fully automated development is much slower. So again for us I think I think about it much more of like dark versus not dark and more about when does a human have to step in because I think they're human gates that are deeply valuable and I think of that as part of a good software factory deployment not a kind of failure software factory. That's so crazy I have the exact at least as far as like front end versus like rust code I have the exact opposite opinion which is like agents are I think agents are really good at well let me restate what you said and you you can tell me if I got it right which is like for certain things like rust code and and back end things it's a little bit harder to verify but like front end is a little bit lower stakes and and it's easier for agents to check is is that what you're I don't want to put words in your mouth but that was what I heard maybe I well so I think there's two two different pieces of verification there's like rust is a strongly typed language where it's very easy for an agent to verify because the compiler is incredible that that part of rust is great you get much higher quality code when an agent writes rust because it's most of rust is like written in the past six six seven years the language hasn't evolved that much and an agent can obviously iterate and validate its changes what I mean more is like I think about it is kind of like where does the code live in terms of high high and stack or lower the stack as obviously the deeper the stack it is the harder it is to change and at war like parts of our code base that are deeper in the stack are written rust and that's it's like more important that yeah it's correct okay when you something is like the high in the stack where it's like front end that is like JavaScript CSS really really cheap change it's like the bar for having an agent go ham show me a video of it working correctly with a understanding that I get changed or pretty cheap of the future is like something that's like great to me okay so I definitely read with the like the deeper in the stack you are the more critical the more core it tends to be and actually another founder I spent a lot of time with is this guy Ben sweared low who builds their building a like sandbox cloud platform thing for like really fat they're building their own like kernel drivers and stuff because the virtual memory that comes with the Linux kernel wasn't good enough for them or something super cracked guys but they they have this policy of like you basically have two zones of your code base on a fact other like principle engineers at larger teams that have come to similar conclusions is like you have your like core which is built by engineers who were good at architecting systems and understand how to make good libraries and then you have like your like sort of pragmatic part of the code base which is you know solve an actual problem you can import stuff from the core you can depend on stuff from the core and there's a little bit lower bar for the code quality for how what level of review needs to happen and like it's like way it's like the slap slap allowed not not encouraged but slapped tolerated zones of your code base and the rule that makes that work is that those like areas are not allowed to import from each other or depend on each other if you need to create shared functionality it has to be pulled out of one of those into the core and be subject to a higher standard of quality before other parts of the code base because otherwise you get this tangled web of spaghetti where all these things depend on each other and it's really hard to pull back apart but my last take that I will disagree with you on is actually like front end is one of the places where agents are probably the most unreliable because especially like in react world the render model and the way like the the render loop runs and react we've just found and talking to a lot of customers like over and over again the agents are really bad at like reasoning over like okay I have six use effects over here and effects over here in the factory the factory AI guys actually wrote a post on this as well so they banned use effect from their code base because agents are really bad at like building the mental model of this stuff and this was among three or four reasons why back in November we actually decided to take our we ran a lights off factory for like four or five months with kind of best in class like prompting and harness on like opus 4.1 and realize like oh this has become so tangled and messy the front end and other parts of it as well that it will be faster to kind of rebuild this from scratch on really good patterns and design principles then so then then to just try to iterate and evolve it to the right place so I will I will humbly disagree that front end is a safe place for slop especially if you're building I mean if you're doing charts and tables like great fine go go for it but if you're building like complex UIs we build a 90 E so there's a lot of state and like interconnectedness between all the different things in the UI so I uh that that would that would be where I would I would maybe disagree a little bit amazing what they're in the chat they're they're banning use effects they're all talking about it so I think the the react uncertainties are definitely resonating with the crew Dan anything you want to add here before we move on to the passion I guess it was a really really good question but I'll maybe take it somewhere else so this concept of like a dark factory right it comes from manufacturing all of this software factory stuff is like from manufacturing when I was doing some research I saw like back in the day manufacturing companies were bragging about if they were a dark factory what does a dark factory mean it means that you have all these machines and they can literally run without the lights on in the factory like we don't pay for lights because we don't have humans that need to like monitor it and like that's where it comes from and so I'll just tell like a quick funny story as fast as I can outside of just like software like for me I ask myself okay uh when I'm mowing the lawn mowing the lawn do I do it during the day or I do it in the dark at night trick question trick question and answer I do it at night because I have an AI robot lawn mower so that's like in dark mode for me and I think maybe like the spirit behind the question is companies are trying to get more and more in the dark because it means probably that you're more agentic and that type of thing and I can like brag about that but uh the place that I was going to take it to give you like an ops uh like a opposite answer I think it's about having night vision you still okay everything's in the dark but we can all still see why can you all still see you have the visibility you're monitoring you're able to see that we're not getting the slap the quality is still good hey there was an incident in production let me feed that context back to the agents we get an incident going on right now agents change your behavior I think it's more of like that night vision or another way to like say if you want like the night vision thing I think it sounds cool the lights are getting dimmed so let's get more and more in the dark but let's you know make sure that we're monitoring the qualities there like that type of stuff and I gave you a little intro into my my life I like I like robots and AI stuff other work I absolutely love the night vision thing actually like it is a perfect description for I think what a lot of us are working on I publish a skill like two weeks ago called slash show me which is literally like okay I don't want to read that whole diff like use an AI to distill it into very high quality like human readable like code snippets diffs diagrams whatever it is I mean I know AI code of robots have been doing things like this but like this is a thing that I think is going to become really important is like we need to develop new technologies and new ways of seeing so that we can maximize the bandwidth of understanding of like the you can you cannot outsource the thinking and like you cannot outsource the understanding if you lose touch with what what is happening in your code base whether you're like skimming the PRs or not like the the faster you can understand what's happening the better and I think there's a lot of opportunity for tools and skills and prompts and software and think things that help make that easier so I love that I love that too it's so fun and with that I think that's a great segue to move into our capstone question here that I prepare for us because really I the analogy of the night vision is really really clever because you still need to like understand everything that's moving through this but it's just like our touch points and like I said earlier like are the rituals like PRs and stuff they're all just going to transform and we're going to find new shapes they're new needs to be new tooling and new observed levels of observability and visibility into our tools. And also something that I've been learning and really embracing a lot lately is that like the true state of your product exists more in a production state now to where you need to have the constant feedbacks of observability and how things are actually behaving in the wild because it's a complex system that creates everything. And a lot of this means abandoning and changing things that we've just taken for granted or have put on top of our SDLC to make it work over the last several decades. So I want to turn this over to y'all what part of today's SDLC is least prepared in your opinion for what we're calling the factory. And how do you think we'll fix it or what do you think we'll break first and curious to get your viewpoint on the SDLC pain points right now? Okay, so I have something that maybe it's a common take, maybe it's controversial. I'm actually not sure and that's a great way to start. I think it's instant management because it's like it's really easy to have an agent observe that there's a problem and like come up with proposed solution. It's much harder to get an agent to actually fix that reliably. Like the reason that you will, there are problems with that is one permission. So what permissions are you giving an agent such that they can touch any arbitrary part of your production stack to fix it to it's like classic to put our discussion earlier. Like the way an agent works to test it change and it works in a loop, right? What does that mean when your server is down and you need to try a bunch of random crap to get it to work, right? So that's very, very scary. And three, it's like very latency sensitive, right? You get page in the middle of the night and on an order of minutes you want to get your service back up, depending on obviously like the severity of the incident. And so to me, instant management, it's like very easy to have an agent do some initial triage, give you some suggestions of next steps. And like I think that's great as kind of the first stab at a, I don't know, lights on version of instant management, if you want to call it that. To get to a part where we're like night vision boot lighting version of instant management is much, much harder, I think that's the one that's really a sticking point, especially when you think about permissions. Yeah, totally makes sense to me. Like what I'm hearing from customers, at least like on the linear B side, it's wherever they have the least amount of trust, it's like where are you most afraid now for the customers that are just, I don't know, starting out on this AI journey, usually it is in merging PRs because they're scared. Hey, all this code is AI generated. I don't know if I trust it. It could be there. Well, it could be where you're saying some of the ones that I think are a little more like mature that that I see, yeah, it's in the incident area. That's the area that I'm most afraid like it's in freaking production. Oh my God, I don't have trust there. Like I want, you know, full human time, it could, it could be there too. So it, so it kind of varies, just think about like where the, where the trust is least. That's probably the answer. The only other thing that I would say if like some solutions, I kind of said it before, but I think everyone knows now, it's all about context, right? Can I get the best context to my agents so that I trust them more so that they know everything that a human knows. And I can see what some of the customers are doing. This is the part that I mentioned before is what the agents know if there's an incident going on in production right now, let them know if there's a particular area of the code that could be at risk that's being worked on, maybe pause the merge or like change, how you're doing your risk assessment in real time. Like that's the type of stuff that I see gaining trust, I guess you can call it that. And I guess the last thing that I would repeat, do that thing where you're doing the automated policy risk assessment. If you're afraid like afraid, that's a way that to not be afraid, okay, now I know I have a mechanism in place. And then it will move your lack of trust probably to the next area, start with the PRs and then maybe the incident. So kind of on a case by case basis, what I see. So incident management is like the end of the funnel. That's not lose sight of the top of the funnel too, which is what to build. That is not necessarily, and it's obvious, right? It's actually important to call out, it's like a dark factory also means the factory is the one who's like looking at, putting it as a course because it's like crazy if you think about it. It's like looking at user feedback and deciding what to build. That's the one right now that I don't think is like well set up to feed into the factory because that requires identifying the problem, which an agent sometimes is good at, but it's really like the thinking to come up with the right product solution and then distill that into something that should just you go build. And that's something where humans are still deeply, deeply in the loop and they should be that is the thinking. I like it. I agree with everything that everybody just said. I want to call out a comment from the chat from Brian Cripe, what's up, Brian? Unified observability across STLC is missing. Don't treat all the phases as independent in different tools. I think back in a day like it was very normal for a product manager to understand the product and the UI and the surface of it and not really need to understand the code and how it shaped to make decisions, to make proposals, to say, hey, our customers have this problem. I sat with a designer and figured out how to do it. Here's a Google Doc or a Confluence page or a Figma board or whatever it is. Actually, Andrew, can I share my screen? Is it going to let me do that? Let's try it. Maybe you could screen share. You could start. Okay. The way I kind of think about the STLC and how forges work is before GitHub, and I know Linus still does this, but most projects do not email, get patches around on email lists and rely on the maintainer to assemble the build and build it and publish it. This was how we used to use distributed version control. It was just sending things around and it was very chaotic and thank God there were a lot of really smart people who could own this and be in charge of it. And we got GitHub where you have this kind of like, yes, it's a decentralized protocol, but everything is kind of streamlined in one place and it kind of flows. At any point, as a project owner, I know I can check this thing out and build my project and it is the state of truth for what is our main branch. Make sense? Oh, it makes perfect sense. I'm obsessed with this. This is like you must have been a classroom teacher in a past life. I just like to draw things and I'm bad at explaining, so we fall back to the drawings. What we have now or like, it's not like a problem. It's just like an inefficiency and an opportunity, which is like our Git patches are all centralized. But our plans and our session traces and like, there's a lot of people trying to bolt session trace like attribution onto Git history, stuff still lives in Figma, stuff still lives in Google Docs and Notion and the weirdest one, the closest one to this is like, I see all the time of people being like, here's what my quad said, paste in the Slack, ping another engineer or be like, hey, like what do you think of this? And then that engineer takes the message, paste it into their like coding agent session, gets a response, paste it back into Slack, and then the other user paste that into their coding agent session. This happens all the time and this is like the metaphor of like emailing Git patches around. And so like my take is like the tool that really enables an antique agent development across the whole SDLC will kind of combine all of these things in a very like, you know, a native way where they all have like a very specialized relationships between each other, between products specs and designs and code and prompts and agent sessions and plans. And it all kinds of comes together in one place. This is kind of my vision for how this space is going to evolve. Human layer will play some part in that, I don't know which one, but like this is I think the most interesting opportunity, like it's not a problem. We've shipped software like this for years and decades, but with AI, I think it becomes way more valuable and there's a huge opportunity to put this all in one place. I love this. And also chat loved it too. So make sure you go back and get all of your your love from everybody about the diagrams you showed us. It's like, I just throw this question in the air like, what do you think it's like broken or what do you think the current problems are? And you're just like, well, let me just get out this giant board and actually have the future mapped out so that we all just got like a great preview of where this kind of stuff is going only from the mind of decks. I love that. And we're coming up to near the end of the top of the hour, but I do want to pull a question out of the chat that I saw earlier that really resonated with me. I really loved it because so far we've been talking about the factory itself and all of the different minutiae of knowing if it's good or not and how to not stumble over yourself. But at no point have we really doubled down on talking about like, how do we get developers aligned around this idea of the software factory and get them on board and excited about what their roles are to play? I'm curious what how y'all are thinking about if we're going to take this step forward into this factory model, what do you think the role is for engineers based on the stuff that we've talked about and covered here today, but how there's still so much work to be done? All right. I'll bail you out, Andrew, from your team. No, no. I just, you know what, there, I think there's different types of developers. So you got to find like, who's passionate about what, but I'll say like the type that I was. If you told me, hey, we're going into like a full factory mode. It's going to be this dark factory, a lot of of this work is gonna be now like automated and agentic. What I would want to hear is, hey Dan, that means that you can do more product value now. I know you really like being with customers, you wish you could go end-to-end and do all the product stuff and the development stuff. Like that's what I would like to hear. I know there's like a group of engineers out there that just wanna build awesome stuff. And like that would kind of be where I would look to take them. Hey, let me get you more product responsibility, own this area end-to-end. All this other stuff's gonna be automated for you. Like go do like your innovative stuff, kind of your entrepreneurship within your role. Like I think that is kick ass. So for that group of developers, that's what I would say. - Yeah, I think the mistake is saying you wanna dark factory and not trying to think about getting alignment with your team. Like that's the first order of business. The way I would do that is really emphasize how a factory one still has humans in loop in the right place and two, less than focus on the parts or the thinking of software development that excites them, right? So as part of working the factory, you may not be writing all the code, you're still thinking at a systems level on a product and technically, right? What are the things to build? High level, what is the architecture? You may use them each and to do that, but you're the one still thinking. And that's like the most important thing that most engineers are very excited by. The other thing is like, I think that the shift of we are all dev prod engineers, we're all the ones who maintain the factory has been really motivating for a lot of the team at warp. And I would really try to get people in that mindset. So it's not me against the factory. It's I maintain the factory. I'm on the hook for improving it. And it's really exciting. It's really energizing. And I think as an engineering leader, you need to model that and maintain the factory and get other people on your team to maintain and improve the factory as well. - Yeah, and I totally agree, I love that. Yeah, I think this has been true since before AI is like, hey, if AI is slow and bad, like go fix it. Don't complain or throw stuff over the wall. Like we're all in this together. And I like your other point as well as like, there are things that like humans are uniquely good at that we haven't figured out how to RL into AI because like AI, we have not figured out how to replicate the experience the human engineers have that teaches them lessons, which is like, oh yeah, I've been up at three in the morning debugging this problem before. And I'm never allowing a PR that exhibits that problem to go into the code base again, because I don't want me or anybody on my team to get stuck with that. And we haven't figured out how to turn, put that shape of lesson into the weights of the model. And we may eventually, but right now today, like I think the number one thing is like, you are the human engineer, whether you have two years of experience or 20 years of experience, you have intuitions around things that will cause you pain in the future and cause your users pain in the future. And you owe it to your team and to your users and to your customers and to the industry to keep learning and keep applying that intuition and knowledge and keep building it. - Amazing. I'm going to go ahead and share my screen again. So I just want to give us a few moments just to go and start wrapping up here for our panel. So first, I just want to move into a slide for each of our panelists today. I would love for y'all to listen a little bit about what they're building in terms of this problem. And I also just want to let people know that we're going to be sending out resources regarding this as well. So they'll be a link to the video of everything that we've discussed as well as a guide that goes in deeper on some of this stuff 'cause I've been nerding out about the software factory stuff. Decks and many of the folks at Warp have been feeding me amazing quotes and ideas about where this is all going. So don't miss out on all of the stuff you're going to get afterwards. But first, I'd love to hand it over to Dan to go ahead and talk about linear beings and give you about two minutes here to talk about how you think linear beef fits into the story around software factories. - Well, listen, first of all, thanks for having me on the panel. If you all want to chat with me more, like feel free to hit me up on LinkedIn or however you want to find me. Happy to continue the conversation. - Yeah, I mean, would linear be, listen, come get your night vision for your software factory. That's where we start with. Show you what's going on and the end. That's where most, you know, I think software organizations are coming get that, we'll give you the night vision. And then after that, it's all about improvement. Know where you're going to improve, what you're going to improve, you probably heard me talk about it a lot, but I really like what we do with Git stream, the policies, the making sure that you can really get fully autonomous. And yeah, we'll just be pumped to have you all check out linear be and like I said, if you want to chat with me one on one, hit me up, I'm happy to do so. And that will be it for me. I don't need the two minutes, thanks Andrew. - Amazing. I'm going to hand it over to a low, talk about work factories for sure. So work factories is open infrastructure for you to build your own software factory. So what that looks like is we have a kind of opinionated default factory that lives in code. So you've been extended in any way. You can add the agents that makes sense for you. You can add skills. You can jage, because we have our model harness you want. And it's fully built with the right integrations that you need and it's built to help demonstrate ROI up front. So it has self improvement built in. So it improves over time and will send you PRs back to your own factory definition to improve your factory. And also has metrics to actually show here's how you're improving through put over time. So the high level vision is like you have full control over your factory, but you shouldn't have to spend all the time building every small piece of it. Focus on similar to all the other discussion we have. Focus on the high level pieces of what makes sense for a factory and then deploy it really quickly with warp factories. This we launched this last week. So it's really, really early. If you're interested, hit us up at warp.dev/factories or just email me [email protected]/[email protected]. We'd love to chat more about it. I'm like super, super excited about warp factories. - Amazing. I'm excited to see what he will build with it. And that's when I hand this over to you. I got some slides in here. And I also pulled some really amazing screen shots from your podcast with Vibov. I just love to give you an opportunity to talk about some of the stuff you share every week. - Oh yeah, nice. - Yeah, so again, like human layer building a multiplayer coding agent workspace, building blocks for software factories. What does the orchestration layer look like? And what does it look like to kind of build an ecosystem of building blocks that allow, you know, you can bring your own compute or you can shell it out, you can buy or build the harness, the dev environment, the orchestration layer. And just really looking at like how can we, how can we build workflows in your factory that put humans into the right place and give you the most leverage as we say. How can we save you three hours at PR review time by spending 20 minutes reviewing a spec with the coworker? And then yeah, I'm talked with another founder of Vibov, who is one of the best engineers I know, one of the best AI engineers I know, and we do a podcast every Tuesday at 1015 Pacific. There's a, there's a Luma, if you go to Luma.com/baml, you can sign up for the next one. And you can catch a slide. We just stream them live on Twitter. So x.com/dexhorthediexhorthy. And we'll be posting those there and try to help people cut through the hype and build AI systems that work. - Amazing, and we're gonna include links to all of the stuff for all of our presenters today and post stuff and afterwards. And like Dan and others said, be sure to come find us online, come pick a fight with us and link in our X or wherever about something we said today. I would love for us to continue this conversation beyond here. And if this is your first time coming into a webinar like this, we have events around this kind of stuff all the time. And if you're not reading Devin Torrupted, then you might just be one step behind. So be sure to subscribe to our weekly newsletter and podcast as well. And we get quotes from all of these amazing leaders you hear about today. But you'll also stay ahead for the future of what the software factory looks like because we're not abandoning this narrative anytime soon. So thanks for giving us your hour. We had a blast and we'll see you at the next one. Take care everybody. - Thanks, see ya. (upbeat music)

Podcast Summary

Key Points:

  1. Engineering leaders are moving toward software factory models where AI agents handle development tasks while humans oversee the edges, aiming for productivity gains.
  2. Measuring success in a software factory requires tracking both business impact, like predictable delivery and customer value, and leading indicators such as cost per effective PR merge and rework rate.
  3. Humans still own code responsibility and PRs, but policies with auto-merge gates for low-risk changes can incentivize quality and reduce reviewer burden.
  4. Identifying what can run in the dark involves defining human gates at critical points, like spec reviews and code reviews, rather than seeking fully autonomous systems.
  5. The least prepared SDLC areas include incident management due to permissions and latency, and the top of the funnel for deciding what to build, which still requires human thinking.
  6. Unifying observability and context across the entire SDLC, from plans to code to production, is a key opportunity to avoid fragmented tools and processes.
  7. Engaging engineers in factory models requires emphasizing their roles in product value, systems thinking, and maintaining the factory, not just writing code.

Summary:

The panel discussion centered on the emerging software factory model, where AI agents increasingly handle development tasks while humans shift to oversight roles. A key theme was measuring success, with panelists advocating for tracking both business outcomes, such as predictable delivery and customer value, and technical indicators like cost per effective PR merge and rework rates. They emphasized that humans still own code and PRs, recommending governance policies and auto-merge gates for low-risk changes to maintain quality and reduce reviewer overload.

The conversation explored what can run autonomously, noting that human gates at critical points like spec and code reviews are features, not bugs, and that maturity varies by trust, with incident management being a particularly challenging area due to permissions and latency. The panel highlighted the need for unified observability across the SDLC, as current tools often fragment context from plans to production, creating inefficiencies. They also discussed engaging engineers by shifting their focus to product ownership, systems-level thinking, and maintaining the factory itself, which can be motivating.

Overall, the discussion underscored that software factories are not about eliminating humans but about redefining their roles to maximize leverage, ensuring quality, and avoiding the compounding of "slop" that harms future development. The panelists shared practical insights from their experiences building and implementing these models, aiming to help leaders navigate this transition effectively.

FAQs

A software factory is a model where AI agents automate parts of the software development lifecycle (SDLC), while humans oversee and manage the process, often moving to roles that involve higher-level decision-making and factory maintenance.

Successful outcomes are measured by tracking business impact, such as predictable delivery of customer value, and leading indicators like cost per effective PR merge, merge rate per developer, and rework rate, while ensuring code quality is maintained.

Humans still own the code and are responsible for its quality. They review and oversee agent-generated PRs, and they are accountable for maintaining and improving the factory itself.

A dark factory refers to running the SDLC with minimal human intervention. However, it doesn't mean fully autonomous; it's about having 'night vision'—maintaining visibility and monitoring to ensure quality and catch issues, even as automation increases.

Incident management is considered the least prepared area because it requires high trust, quick response, and careful permission management. AI agents can help with triage, but fully automating fixes in production is challenging and risky.

Leaders should emphasize that the factory still involves humans in meaningful roles, like product thinking and systems-level design. Framing engineers as 'development productivity engineers' who maintain and improve the factory can be motivating.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.