Go back

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

51m 31s

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

Matan, CEO of Factory, shares the company's evolution from a vision of fully autonomous software development agents to a successful product, highlighting the challenges of being early in a market. He emphasizes that Factory's differentiator is model independence, ensuring enterprises aren't locked into any single model provider, allowing flexibility and cost optimization. A core philosophy is "create obsessed customers," focusing on output rather than input metrics like customer obsession. Factory's journey included a "two years in the desert" period where they were ahead of the market, leading to a pivotal moment: refunding all customers because the product wasn't good enough, a decision that built trust and credibility. The launch of the droid CLI in September 2025 was a breakthrough, meeting developers at their current skill level and leveraging improved models and behavioral readiness. Matan stresses that the team's resilience, forged through early struggles and a commitment to the mission, is a key advantage over competitors who haven't faced such adversity. The conversation underscores the importance of adaptability, trust, and long-term vision in the rapidly evolving AI software development landscape.

Transcription

10755 Words, 57993 Characters

English
Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. Like, you don't want to measure input metrics. It doesn't matter if you're customer obsessed. Like, you could be customer obsessed and they file a restraining order against you because they don't like what it is that you're doing. Like, our job is to build something so good that our customers themselves become obsessed with us. That is our job. It's like, you know, the analogy is if you're a coach of a basketball team, you don't want to tell your players before they come out there like, "Hey guys, make sure to sweat." It's like, "What?" No, like score points. Like, we need to score points. And in doing so, yeah, you're probably gonna sweat. And I think similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers. But the output is what matters. (upbeat music) We're here in the studio with Matan from Factory. This is our second time. - Yes, Matan. - With Matan. - Thanks for having me. - We're having the small and elite group of second time training data attendees. So thank you. - Oh yeah. - Matan is the co-founder and CEO of Factory, which makes droids, which are autonomous agents for the art of software development. - Yes indeed. - And Matan, we're gonna jump right in because I think you guys are a little bit of a dark horse candidate in this world of software development. It is a market that has absolutely taken off. There are folks like Cloud Code and Cognition and others who have a lead, but you guys are coming up strong. Talk about the competitive dynamics and what makes Factory special. - It's been a wild ride. We started Factory three and a half years ago now. So in April of 2023, when the world and the enterprise in particular was barely ready for GitHub Co-Pilot, let alone fully autonomous agents. And so I think the first two years it was kind of our journey in the desert is how I like to refer to it because we were focused on fully autonomous agents, but engineers weren't ready, procurement teams that the enterprise weren't ready. And so I think retrospectively, we really like honed our craft and learned a lot about how to build four developers in the enterprise, but it took a lot of time to actually come around to when they were ready to receive it. And so we're kind of now emerging much more in some of these other players like Anthropic or OpenAI who have a ton of distribution are going in and bringing their incredible tools like Cloud Code or Codex. The thing that enterprises are really caring about that we have learned through those two years is they do not want anyone to kind of be their single point of failure. They do not want anyone to kind of control their fate. And so something that really matters is model independence. Everyone learned from Cloud, where back in the Cloud days, it was like AWS or Azure being like, hey, come on in, sign this three year contract. It's going to be so cheap, we're going to subsidize it, it'll be great. And then a couple of years later, when it came time to renewal, they would 10x the contract. - Ha ha, did it gravity. - And we got you now. - Yeah, we got you. What are you going to do a two year migration to go to someone else? Like, no way. Everyone has scars from that now. And so everyone knows, look, Cloud Code is fantastic. Codex from OpenAI is fantastic. We cannot put our fate in any one of these model providers hands. Also, like, you just look at the risk profiles of the model labs versus the Cloud providers. What's the last piece of drama that came out of one of the Cloud providers? Versus like the model labs, it seems like there's kind of always some sort of chaos of internal fighting or getting in spats with the government or any other entities. And so if you're going to build this very important part of your business, you want to make sure that you're robust to any of these changes. And that's something that we've learned over those kind of initial two years is like, developers really care about things being modular. They want to know that they can customize it to what they want. They want to know that if there's a new model that comes out that's faster or cheaper or more performant, they can kind of hot swap it in. And that's, I think, one of the biggest reasons why a lot of the largest enterprises are taking the momentum that they had from a Codex or Cloud Code. And then are carrying that into factory because they get that performance from these fantastic models, but they do it without the vendor lock-in that the model labs directly run. And if I'm in the enterprise, I'm like, wait a minute, am I now just getting locked in the factory? What's he into to that? So it's a really good question because that is something that you might think of like, okay, wait, so we're just switching the lock-in point. All of the modularity that we build is such that, if at some point you wanted to say, hey, you know what, factory's not staying at the frontier anymore, whether it's like the automations that you build or the skills registry that we help you create, the work that we've done stays in your code base and any of the automations that we've created, the artifacts also live in your code base. In other words, there aren't really things that we're saying like our tribal knowledge about your org that we're keeping on our side and not giving to you. And that's part of the relationship that we have with customers is like, we similarly want to make sure we're providing the best experience possible if we help you arbitrage between different models to get cost optimization, we're giving you that optimization, we're not taking that away from you. And I think that's a really important part of the trust that we're building with these enterprises. - You and I were talking probably a couple months ago at this point and I was trying to give you credit for having the right vision for this market two, three years ago, and you responded with something along the lines of, thank you, but being two, three years early is the same as being wrong, yes. Which I thought was a wonderful response in so many ways. Can you talk about those two years in the desert, how did it feel to have this vision that turned out to be right, that nobody appreciated for a year or two? Can you just talk about that journey and what it has done to the DNA of your company? - Yeah, I mean, in the moment, it's really, really difficult because I hadn't had a job before. I dropped out of my PhD to start this company and over the course of those two years, convinced 20 of the smartest people that I've ever met to quit what it was that they were doing and join factory and join us on this mission. And these are people with families, these are people with kids who are like, dedicating years of their lives to this problem and going, you know, customer after customer and they weren't ready for agents, they didn't get it. Also, the models weren't as performant, but I think a lot of it was behavioral. And I mean, even just a fun anecdote of like, giving developers an NPS survey. If you ever are giving a developer an NPS survey, they do not like whatever it is that you're giving into them. Because like developers, they vote with their feet, they are very clear what they like and what they don't like. And if you're like, hmm, I wonder if they like it, they definitely don't. And, but during that time, I think there was a lot that we were learning. There was a lot that I myself was like, I'd never had a job before. Enterprise sales is not something that comes obvious to a physicist. But at the end of the day, it doesn't, it doesn't matter. There's no, you don't get any, you know, bonus points for being early because like who cares? Like there's no consolation prize. It's either you do the thing or you don't do the thing. And that's all that matters. And for the team is really tough. There were points where we ended up getting good at enterprise sales, but the product still wasn't good. And that's a very tricky position to be in because we ended up, you know, getting to a point where we were like just under two million in revenue and the product was not good. And there was a point in time where we realized this because if you're really good at sales, you can sign contracts. That's like, you can definitely do that. But if you're doing that in the developers don't like your product, it's like a ticking time bomb because it's eventually they're gonna turn and it's gonna be really, really bad. We realized this and we proactively gave all of those customers their money back. And I remember that was one of the most difficult decisions to make because not only is there a group of 20 people who are getting ridiculous offers from all the labs, they have these huge financial incentives to go elsewhere. There are all these other companies that are doing well and they decided to do this. And then we're gonna say, oh yeah, hey, by the way, that little bit of revenue we managed to get, we're actually gonna give it back because we don't think the product is making their developers happy. We also had to tell-- - Why did you make that decision? - We sold them on a good vision and convinced them that this is the right team to work with and that we were gonna deliver the solution for them. But we realized that the way that we had sold them on it and the product that we were delivering was not up to snuff in a way that I don't think it would hold true to one of our operating principles and one of our operating principles that I really like is create obsessed customers. This kind of flips over Bezos' thing where Bezos at Amazon, it's customer obsession. And input metrics are, you don't wanna measure input metrics, it doesn't matter if you're customer obsessed. Like you could be customer obsessed and they file a restraining order against you. 'Cause they don't like what it is that you're doing. Like our job is to build something so good that our customers themselves become obsessed with us. That is our job. It's like, you know, the analogy is, if you're a coach of a basketball team, you don't wanna tell your players before they come out there like, "Hey guys, make sure to sweat." It's like, "What?" No, like score points. We need to score points. And I think similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters. And I think that coming back to this, the product that we were delivering was not creating obsessed customers. And we wanted to make sure, like, this was a group of the smartest people I've ever met. We were getting there. Like we were getting a lot of intuition, things were starting to come together internally. Like we could see internally, we were starting to become a lot more, agent native and how we were doing things and the product was kind of scratching that itch. But we were kind of ahead of our customers and we wanted to maintain trust with our customers so that when it does hit, we can come back to them and say, Hey, guys, this is the real deal I promise. And to build that credibility, we had to say, Hey, look, you know, even though you were maybe happy to continue, we're going to give you this back and say three months from now, I think it'll be ready. Give us some time and I promise we will knock your socks off. How did your customers react when you had that conversation? Some of them were like, Oh, great. Sounds good. Because I think it wasn't something that they were obsessed with. Some of them were a little bit confused. But I think generally, it's especially enterprises. They're not used to these things. A lot of times enterprise budget, once it's gone, it's gone and no one really cares. And so some of them didn't even know if they had a mechanism by which to take back money. But you know, it's a difficult thing to tell also like investors who believe in you. Like, you know, I remember having the conversation with Sean. And Sean, obviously, he stayed really close with the company. So he was very like on the same page. It's kind of a scary thing to be like, Hey, by the way, you know, remember all those updates. And you're saying, Hey, look, the, you know, revenue is going up. It's about to go down to zero. It was a scary thing. And I think it was kind of a leap of faith of like, we see the signal internally early of like, this is the direction we need to go. We need to kind of pivot the approach on the product. But I remember that all hands where we told the whole team, it was like, oh my God, that was like one of the worst months of my life. Like I was just, because no, like, not everyone was going to say like, what the hell is this? What's going on? But it's kind of the looks on their faces where they kind of go a little bit pale. And they're like, Oh boy, like, is this just the early signs we're about to sink completely? How do you keep the team together through that? I think honestly, the only reason the team stayed together is we were so ruthless about hiring early on where it was like people that are genuinely really, really obsessed with the mission, which our mission is to bring autonomy to software engineering. And like really, really caring about that, making sure everyone was also like very clear feedback, loop says to like, this, the fate is in our hands. It's not like this is like, oh, something that I go do. It's like, we all have a part to play in, you know, making this work. And I think embracing how much it sucked was also, I think, something that was very valuable, just being honest about it, being super honest about like, yeah, this sucks. Like, oh, look, look at those competitors. The revenue is going up like crazy. Like, this is not good. Like we are in a very bad position. Like, we just had to give back all of our revenue. Like we need to really get our shit together. And in the moment, I think retrospectively, those are the moments where really the deepest bonds are made. Like if you talk to people who are like athletes or even like academics or whatever, whenever you're in the like stressful period, whether it's like cramming before finals or, you know, in intense, like, you know, we have some, some rowers on our team. And I think that's an example we always go to. Like, that's pure pain. Pure pain. It's pain. It's literally just, there is one number that quantifies your performance. It's just what is your time on your 2k or your time in that. But like embracing that is what creates those enduring bonds such that afterwards like we know what it's like to be at rock bottom. We know what it's like to lose. We know what it's like. I mean, when we first started the company, our valuation was 5 million. Like a lot of our competitors, a lot of the companies out there these days, they don't know what it's like to not be a unicorn. That's like manifestly. That is what they are day one. Whereas like we have been there kind of in those dark moments and not a single person left. Yeah. That makes us so resilient and so strong that, you know, going forward, things are going a lot better now. But they are going to be really bad times. But we have that resiliency in our DNA that I'm not sure some of these other companies do. I love that. So talk to us about what changed. And I'm curious your comment from earlier that the model is getting better is not the most important thing that happens. Because in my mind, the model thing, whether it's the most important thing that happens. Yes. Help me understand. Yeah. So a couple of things. So one is the interaction pattern that we were building for before was two ambitions. Like to your point, we were right in that what we were building for was fully autonomous agents. But it was two years too early, which makes it wrong. And fully autonomous agents require a complete change in behavior from the developer. And we were trying to do that out of the box before they were even using tools like co-pilot. It was just too much of a leap. It was too much of a step function jump. So it's an important day. September 26, 2025 was when we first put out, basically the the droid CLI and the droid CLI met developers where they were in a manner that previously these fully autonomous agents did not. And also its performance was like completely state of the art and it was model agnostic. So it could use every model that was out there. September 26 was also two years after we initially started. So the world had gotten much more used to using things like auto complete. Like by by late 2025, most engineers were using an auto complete tool. And many were starting to at the time use like a chat interface to ask an agent to go do changes like wholesale. So like the more agentic interaction. However, what we see is that like if you go back now and use in this like agentic interaction, some of these older models, they're still good. So the biggest thing that changed was developers and in particular in the enterprise, like being open-minded to this new way of working in particular, you know, developers, they've established their workflows over the last 30 years. They can be stubborn. A lot of them were like, no, no, no, like my craft could never be done by an AI tool. So a lot of it was just like understanding how to work with these tools and having the willingness to go in and try and also intuition about what are the guardrails that you need to provide in order for it to succeed. And so I think it was a combination of both of these things. The model is getting better so you need to do less in the way of providing guardrails, but also developers lowering their guard and being like, okay, you know what, let me go try and do these things. It's going to go do things I don't like. And then also there's a certain degree to which when Andre Carpathy tweets about something, then every engineer suddenly is like, okay, you know, maybe this is true. And Andre started to tweet about these agentic work early on. He wasn't as open to it. And then him being more open to it genuinely just changed some people's minds, which is funny, but that's some of the things that go into behavior changes. Like you hear it from people you trust, you start seeing it, you know, from people within your organization who are maybe a little bit more agent native, but that's that's kind of these things together is what changed that. And now we're all going to be on Slack. We might be on, we might be pushing the limits of Slack, which I think is going to be another interesting thing. But yeah. Okay. So September 25, you launched the droid CLI. You said, frontier performance soda. What does that mean for you? There's like the benchmarks, which have a very short half-life. Like anytime there's a good benchmark, it gets benchmarks within like three to six months. Yeah. At the time, I think the one that we kind of championed when we launched and kind of it ended up becoming a pretty good benchmark was terminal bench. So prior to that, the one that was kind of leading with sweet bench, which was kind of took some open source projects and some examples of issues that were then solved. The problem with that was it was very focused on like Python and like scripting or like individual file changes, whereas terminal bench was more one, it was in the terminal setting. So it was things like scheduling runs and things that were not just like changing the code file, but general software development tasks. And that was something that we ended up, you know, having really frontier performance on now it's like bench maxed to the extreme to where it's like, I think, you know, models that come out now are like 90% on it. And I think there's a very short time horizon from putting out a good benchmark to then it being kind of in the training data. What goes into building a great and is it that is a great harness and it seems like there's almost a lot of fun in the ecosystem of my harness is better than your harness. Yeah. You know, you need to own the model to have a good harness or actually you have a better harness if you don't own the model. Yeah. What's your mental model for? Yeah. You know, benchmark maxing aside. What keeps you at the frontier? Yeah. So a couple. So some general things that matter are the way you do caching. So, you know, cash tokens end up being like a tenth as expensive. And so one big piece of performance for a given harness is what what is your like rate of token caching. Another example would be how do you perform while in compression or compaction. So typically when you're dealing with a long session, you're going to exceed the context limit of the model itself. And so the harness will do some sort of, you know, summarization, compression, compaction, whatever you want to call it. And the way that you perform during that compaction is a big determining factor of how good your harness is. And tests that they do for that are like, you know, they call it needle in the haystack where you have some long thread and maybe there's one piece of information that's really important. How often will your harness preserve that through compaction? Other examples are like tool use or how does it use the environment to validate whatever work that it's doing. These are things that you can kind of have individual metrics on and that we kind of have our own internal benchmarks to measure how do the out of the box agents do versus how does factory perform. I think one thing that naively everyone believed initially was if you train the model and you build the harness, you're going to make them better together. And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better. What's the like my intuition would be model harness co design makes you better. Yes. What's the intuition for why it's actually not? It's very analogous to the idea maybe like, I don't know, 10 years ago, if you were to be like, Hey, I want to train my personal a back in like, like ML days before like GBT 3. I want to train my personal AI. I'm going to give it all of my data because I want it to know me. Turns out the answer was, "Training it on the whole internet, and it'll be so much better for you than if it were just trained on your data." So there's a sort of analog that emerges where it's, "What data is to a model? Models are to a harness." Where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular. And there are certain intricacies about different models that you can learn from and then improve different models performance in your own harness. And this was why, for example, we kind of stopped doing it because Terminal Bench got so bench maxed. But initially, when every new Opus or GPT model would come out, it would perform better on Terminal Bench in droid than it would in Cloud Code or Codex. Which is why, and this is something that I think was frustrating to, because from a lab perspective, you ideally want it so that it's better together because then that means you have to use their harness and you can't use a different one. But I think the reality is it's having that multi-model harness ends up getting kind of frontier on all those options. Is there a good example or illustration of that? Conceptually, it makes sense. Is there an easy way to illustrate it? Maybe a good example of it is like, "If you're familiar with the different behaviors of Opus and GPT 5.6 right now." I am, he's not. Opus tends to be. [laughter] I mean, loosely. I mean, to be fair, honestly, these days I'm not doing it as much either, but I will say this. Losely, Opus is kind of like that super-friendly colleague where you're like, "Hey, I want to go do these 20 tasks." And they're like, "Oh, yeah, cool. Hey, by the way, five of those tasks, I realized we didn't need to do it. Don't worry about it. I got other these done. Did it this way." It's a great time. Let's pick it up in the morning. Yeah, like, let's go get a beer afterwards and hang out whatever. Meanwhile, GPT 5.6 is like, "Absolutely. I will do every single one of those and nothing will stop me. I'm not going to sleep until this." It's like kind of very OCD and, you know, meticulous, but sometimes, you know, you want one where it's like, it actually realizes, "Hey, that list of 20 that you gave me." Actually, here's a better way of doing it anyway. You know, 5.6 is more methodical. If you build a harness for each of those, there are actually different things that that harness will then be good or bad at. So for example, one thing that, you know, typically agents will do is they'll have it to do list. If you have a task, it'll go and generate it to do list. The cloud code harness can, in some cases, or, and this is maybe less relevant now, but I think earlier, this is just a more illustrative example. Earlier, it was really strict to make sure it would stick to the to-do list because the model itself would typically wander. Meanwhile, codecs wouldn't do that because the model itself was really, really OCD about that. But if you're a user, you want to have the same experience regardless. You want to make sure if you switch to a different model, you're not going to suddenly lose track of whatever things that you are working on. So there are certain things where like, maybe in some cases, you really want robust tool use. And there are tools that you use to do these to-do list. You want really robust tool use. And you want to make sure that no matter what, if I'm a user, I want to see my to-do list there. Like, there were some cases where it would just like not have the to-do list. And so these are things that kind of improve the general performance. And the to-do list matters because you're doing some crazy migration and you don't have the to-do list. And then you're in this long session where there's compaction that might get lost in the summarization. And then now you forgot what your seventh step was and that could be one of the failure modes. That's kind of an example of how to do example. Yeah, yeah, yeah. So we talked about one type of maxing, benchmark maxing. Let's talk about token maxing. Yes. Because it feels like the world has changed a lot. We've gone from token maxing to now cost rationalization. What is theming for factory? Yeah. So maybe I'll lay this out just to so we're all on the same page of like the way that we see what's led us to this token maxing. So loosely there was like this phase one where maybe phase zero was like known believed in that then phase one, everyone believes in AI. And then boards were like, Mr. CEO, what are you doing about AI? What's your AI strategy? And Mr. CEO is like, shit, I don't know. What's our AI strategy? CTO like make sure everyone goes and uses AI. And so then phase two is, you know, CTO is like, okay, shit, we got to make sure everyone uses AI. Let's start putting it in performance reviews. Let's make public like bench or public like rankings of who's using tokens the most because everyone's stubborn. No one wants to use this stuff. They're all skeptical. And then we enter phase three, which is everyone sees these ratings. They see that it's part of their perfere views. And they're like, okay, I'm going to use AI for everything. And that's kind of phase three. It's this token maxing where people are using like opus for literally everything. Like, what's the weather in SF? Opus, tell me. I don't know. Like, there are banks that we are working with where they are spending literally hundreds of thousands of dollars a month on people asking things like literally, what is the weather? Or like, tell me about Python, like trivial questions that you could Google. People are asking opus. And the reality is this happened because we were so worried about adoption that we over corrected. We're like adoption by any means necessary. And I think that's actually, it's like a decent approach. Like, it's probably faster to do that and then curb uses or make usage more responsible. Then it is to start limited and be like, you can only use it for this thing. Because when you have people that are stubborn, first you want to just prove that it works. And then you can get kind of more mature about it. Where factory fits in, I think one of the most important things that we do is that we have the factory router, which allows you to dynamically route to different models based on the task that you're doing. So, you know, if you're asking what the weather is, you probably don't need the very frontier of human intelligence to answer that for you. Or you really do. I mean, it depends on what kind of answer you're looking for. You know, giving you like a full like down to the molecular level of what's happening. But, you know, allowing that, but also more importantly, for every enterprise, something that no one's dealing with yet, but 12 months from now is going to be the case is not everyone needs the same tokens. Having a blanket kind of token cap for every individual in some large bank, let's say, makes no sense. So, every CIO is going to need to answer for every incremental token. Where do we put it? And right now, it is super not obvious how you would do that. Like right now we're saying, oh, you know, the PMs who are like vibe coding dashboards get the same token limits as like the engineers who are building like critical infrastructure. That's probably not the best thing to do. Or similarly, you might be dealing with cobalt code bases where Opus is not the best model to use, but instead maybe some fine tuned model on that code base in particular. The point of the router is that we can kind of accommodate these different constraints, where maybe you say, you know what, this part of the org, they're just vibe coding. They can use Gemini flash this part of the org. They're doing cobalt. We fine tuned this great model to work on cobalt. Let's route to that when we're working on that part of the code base. Maybe this other part, we really care about reliability. So let's generate the code with open AI, test it with Anthropic, review it with like Gemini things like that. And we can actually take in your routing procedure instructions in natural language. So you could even say think like it's not purely deterministic. It can even be like, Hey, you know, Pat, I don't know like I don't know what he's doing. Like, give him flash. Give him flash. Like I'm a, or you know, I think we really need to avoid having them use open models because you know, whatever reason we don't like the way open models perform here. And we'll do internal benchmarking to know which models are better at which of these tasks. How close are the open models at this point? Which one's the best? GLM 5.2 is incredible. It's at the point where internally we have no token limits for our engineers and like half of our tokens are open to open models. Wow. Yeah. Because they're just faster and they're cheaper, just as performant. And I think the thing that everyone gets wrong is everyone is comparing like GLM 5.2 to the latest model like Opus 4.8 or GPT 5.6. But really they should be compared to Opus 4.7 or GPT 5.5. Because generally the open models come later and they're kind of a generation behind. And that's kind of the, the frontier models will be frontier. The question is, are the open models getting as good as like frontier minus one? And the answer is unequivocally yes, which I think is a really, really interesting outcome. It's great for consumers. And by consumers, I don't mean like individuals. I mean, the consumers of the APIs because if you're a, you know, a business that is doing in like software engineering, your job is at a very high level to solve problems. And if we can allow you to solve those problems faster and with cheaper models that are just as performant, that means you can solve more problems like that is a good thing. And it is a very good world where there is not like a monopoly on intelligence, but instead kind of a garden of intelligence that you can pick and choose, you know, when you'd like. Something that we joke about is like, you know, on this intelligence allocation thing, if you're, if you're trying to get a tutor for your daughter in algebra, you can probably find someone cheaper than Albert Einstein to be that tutor. Now it might be that she eventually goes and becomes like a leading, you know, physicist or something in which case, yeah, maybe let's, let's get Albert Einstein in there. But most likely you can get, you know, a high school student or something like that. And it's probably much more cost effective for you as well to do so. So since you guys do the model wrapping, like if you look at the, you know, if there's a pie chart that shows the complexion of models being used by your customer base today, what did it look like a few months ago? What does it look like today? What do you think it'll look like in a year? Yeah. I will caveat this with saying that right now enterprises haven't gone too opinionated yet into the routing procedures. Okay. We'll happen over the next six to 12 months, but right now they're just going from no router to router. That's kind of the first change. Then it's going to be like the exact nature of the routing at the beginning of the year. There's less than 1% of tokens. went to open models. In the first quarter, it became a single digit percent. It is now crossed into being a double digit percent of tokens. Now, percent of tokens is not always the same as percent of cost because the open tokens are cheaper, but it is pretty crazy to see the growth there. - Yeah. - That's your forecast. - My sense is that we will ask them towards vast majority being open just because it provides you more optionality and it's cheaper, but that doesn't mean there are going to be, like that's of token share, not necessarily of leverage share because maybe there are 1% of tokens that are incredibly, incredibly valuable and are like very key decision making and then the rest are more like implementation tokens or kind of lower stakes, if you will. I don't think there's going to be a world in which like, it's ever going to be 100 percent. I think the frontier of intelligence will inherently always be valuable for every business just because the stakes are going to get higher and the kind of intricacy with which you think is going to be more important, but will be better at offloading certain tasks. And this is like, you can loosely think of this already with the way orgs are structured, where in general, engineering leaders are more tenured engineers who in theory have more wisdom and each kind of minute of their brain power is higher leverage in theory. And even, you know, you can also imagine like, consider a human engineer and try mapping over the course of their day, like how much brain power they're using. And like, you know, it's probably going to be really low for a lot of it, but then there're going to be some moments where they're like, going pretty high, like they're deeply concentrating and thinking about some, you know, systems design problem or whatever, all of those low leverage moments, we want to automate away. And like we want to, like those like very high leverage moments, sometimes like, you know, we're referring to them as like the Urika moment. So the moments where they're like doing something that's very high leverage, what if those aren't just moments, but what if those are like hours at a time? Because you don't have to deal with all the other stuff. And I think that's kind of the way to think about intelligence allocation is, if you're an engineer and you're writing docs, that is such a low leverage use of your time. Like you become an expert in your craft and you used to spend hours writing docs. Like I remember, it was actually valuable. Like I remember Stripe had so much alpha for just having incredible docs. But imagine all the other stuff those incredible engineers could do if it wasn't writing documentation. Like we should live in a world where everyone can have docs as good as Stripe. And that is like strictly beneficial for everyone. And then the question is, okay, what do those really smart engineers do at their time once they don't have to do that? Maybe it's a good time to have that business model. Given that, you know, especially the rise of open-weight models, the cost differential, I imagine that means very different things for your cost structure. But very similar value delivered to customers. How do you think about business model pricing? - Yeah, this is more what our customers want to need as opposed to what we want to need. So for example, I think right now usage-based is clearly the way to go. We want to be aligned with like what they are doing and what we are doing. I think seat-based doesn't make sense at least for what we are doing. My sense is that eventually we will change to outcome-based. Now, I don't think the enterprise is ready for that. And we've learned our lesson from those first two years. We are not going to impose things, right? But my suspicion is that, you know, in the 2030s, things will probably look more like outcome-based. - Yeah. - What does that come based mean for your market? - Well, it would be the definition of an outcome. So maybe here's a way to put it. So right now we chart, we are usage-based. Like the more tokens you use, you know, the more you pay, the more we get. Now, since we are model-independent, we kind of, with our router, we are kind of pointing a token cannon at either OpenAI and Throbbic, AWS, GCP, you know, any one of these people. To a certain degree, this is like a really dumbed-down version of a marketplace. Well, right now there is a buy, the buy side is an engineer who wants a task done. And then you have the model providers who are saying like, either in benchmarks right now, they're like, we perform at this cost and this performance. And then we determine who we go to for that given task. There's a world in which, you know, if it's so important to get these tokens, they might kind of, like bid in a certain way of saying like, look, here is our cost for this task. We will get this task done at this cost no matter what, but they're pricing it such that, you know, they hope that they can make a margin there. They price it wrong, they're at a negative margin. If they price it right and win the bid, then they get the positive margin. And the way you determine if the task was successful is by some validation loops. Because no one is using these tools anymore where it's just like, right me code. Great, thank you. It's generally right me code. And here's how I know it was done well. And similarly, if you are like a model lab and you are given, here's a task, here's the validation criteria, you'll be able to say roughly how much you think you would be willing to pay to get those tokens. And you know, you want to have some margin on that. And then in that world, that's basically, that's a way that you kind of dynamically shift from usage based to outcome based. I think that there are so many questions with this and this is very much forward looking. But I think there's a lot of questions about how do you subdivide tasks, you know, divvying that up, I think is something that's not obvious. Yeah. But as these tools get better, doing things like that actually become way easier. Yeah, that's fascinating. Yeah, yeah, if you can scope a task and then create a competitive marketplace that'd be a fascinating version of the future. Yes, and as a user, it then creates an incentive to be very thorough in your validation criteria. Yeah. Because like, you know, there are stories of like, you know, you ask an agent to like fix my code and it deletes your code. It's like, you know, the solution is just giving it all to that. It's like that to the convalue up. So it's very easy. How oppression it was. Yeah. But like, so you need to make sure your tests are very thorough because technically it could hit all of your validation. Yeah, it's on the Van Tahnwoll. Go rogue. Yeah, exactly. Yeah, yeah, yeah. Yeah. That's what happens. Maybe zoom out a little bit. You names the company factory. Actually, you named it Droid before factory. That's right. But you named it factory before this concept took off. And now it feels like everybody wants to build a software factory. Where do you think we are today in terms of the building of software factories and how close are we to the ultimate vision of a software factory? Yeah. Everyone has a software factory whether they know it or not. It's just a very inefficient one. So it's kind of like, it feels like, you know, pre-industrialization where like, you know, people were manually like, you know, sowing things together or like woodworking or whatever it might be. And these things are very inefficient. Like right now, if you go to an organization that has more than 10,000 people and you're to ask about the process by which they decide and release a feature, there is like hundreds or maybe thousands of people in that process. And most likely they couldn't even draw it for you. Like there's very low likelihood that they would know what that process looks like. That is not because they think that is the right way of doing things. That is just kind of the nature of building large software as it is kind of today. But with these systems, so much tribal knowledge can be codified. So much of this stuff that typically we would require, oh, we need to ask this guru who's been here for 30 years, who has the wisdom, oh, we then need this approval and that approval. Oh, and I forgot, there was some doc that said, we always have to do this checklist. And it relies so much on kind of human behavior and like redundancy. So much of that can be automated and refocused on like, what actually moves the needle for our business? And I think this move towards software factories is a move towards, how do we figure out what are the actual inputs that determine what features we need to build? And that might be inputs from the customers inputs from the market inputs from like, you know, product leaders at the company. And let's be very clear, these are the signals, the inputs that we are taking in here. Okay, great. We have those signals. Then what is the process by which we build this? And really like mapping out the like assembly lines of how you are building software. It's really important because then you get to close the loop and say, did this actually deliver outcome for our business? Talking before about the tokenomics, if you're that CIO and you're faced with that question of where do you put every incremental token? Really, the question two years from now is going to become, where do you put every incremental dollar? And so you're going to have to be asked, do you put that incremental dollar towards headcount or towards tokens? And if tokens to wear in the org. And these are things that you can only really know when you have these kind of feedback loops that give you examples of like, hey, by the way, we made those decisions based on this data. And it did not matter at all. We added these new features and no one cared. It didn't create more attention. It didn't create more usage or whatever metrics that business is looking to optimize. And the only way to do this is like, you need kind of more rigor and more process. It almost feels like like 10 years from now, we're going to look back at this previous era of software. And it's going to feel like businesses in like ancient times where they didn't do accounting. It is like it's going to be like marketing in the day of mad men. Right. Where it's like all creative and you have no idea what's actually working. It makes notes like it's like, oh, yeah, let's ship that feature. Oh, I think it went well. Like, yeah, we had I got some metrics on that. Yeah. It's like, no, if you guys read the blog post that Jack Dorsey put out about how every company is like an AGI. Yeah. There's also this degree to which if your company is an AGI, you want to optimize the weights. Yeah. You want to figure out what nodes are doing the what things which are load bearing, which are not which need more tokens. Where do you need more nodes? And in order to do like, you don't train a model by vibes. I mean, okay, actually you kind of do it. But you don't, I guess more importantly, you don't do back prop in a model by vibes. Like you are running those actual like calculations and you are seeing when we change this node, what happens. Now you might be making bets on how to change the model by vibes, but you like you're, it's pretty like mathematical in what you were doing. Meanwhile, that companies, you know, people are determining token budgets just by shooting from the hip. People are. laying people off by shooting from the hip and just being like, "Yeah, yeah, like 20,000. " There is no way there is science to laying off 20,000 people. That is just like, "Here is a chunk and let's just see what happens." Instead, I think in these organizations, the way they can do things is much more mathematical. Like, this part of the business matters a lot and does better if we give it more tokens. It doesn't actually matter if we give it more humans. So let's give them more tokens. There might be other parts of the business where actually, giving them more tokens doesn't matter, but more people matter because if we build more relationships with our customers and deeper relationships with our customers, that matters. But these are things that we're going to need like quantitative insight on and you need a software factory to do that. Otherwise, you're just like shooting from the hip and just guessing, which won't work as well. In the limit, how much do you think people will spend on tokens versus on engineering hip-hop? It'll depend on the business. I think every business will have a balance and it just depends on like, they're just going to be like an easy example is, generally salespeople, they probably don't need that many tokens if they're good salespeople. Because generally where they provide the most alpha is like when they're in the seat face-to-face with their customers, talking about the customer's problems, understanding how they build software in our case and how we can make that more efficient, more productive. They can use tokens a little bit of like, "Oh, whatever, generate them some, you know, AID brief, take some notes, like help them with the follow up." But like, that's so minimal the number of tokens. It basically doesn't matter. Like, if you add more tokens to the sales team, it probably won't change their output. If you add more humans to the sales team, it probably will. Meanwhile, engineering teams are pretty different. We're engineering teams. Generally, it seems like you want people to own an outcome end to end, but then if you give them more tokens, they can produce a lot more. And so it seems like they're, and then there's a lot of kind of places in between of like operations, finance, marketing. These are places where are neither here nor there, where I think they're somewhere in between and it kind of depends on your business. But I think every business is going to have to ask, like, what is our core competency? Something that we see a lot in the market, or we use to see. And now they finally kind of hit reality. What we use to see is, "Oh, like, we're going to build our own software development agents." And we're like, "Okay, like, you're like a consumer logistics company. Are you sure you want to do that?" Yeah, yeah, we're going to, we have to do this. And it's like, "Okay." And then six months later, it's like, wait, actually, this is not a core competency for our business. We don't want to hire, you know, AI engineers to be doing this. Our core competency is, you know, consumer logistics. That's what we want to focus on. And I think this is an opportunity for every business to double down on their core competency and what matters for them. And then procure externally, whatever it is, that doesn't matter for them. Like a trivial example of this is like, I don't know, in the days of the early internet, you probably had to be a programmer to build a website. And like, websites generally help if you're a pizza shop because you want to have, you know, people come to your pizza shop, they want to be able to, or like, whatever, at that time, would you say it was a core competency of like a pizza shop to have engineers? Like, certainly not. Like, that is kind of a byproduct of like a brief moment in time. But then there were companies out there that help you build a website. You don't need to be technical. And then this is why we live in a world where like most pizza shops don't have an engineering department, which I think is probably a good thing. And I think, similarly, a lot of businesses have dealt with the reality of if you want to do XYZ other thing, you have to bring in people of this type of role. But I think that's been like something you had to do not because it's a core competency of the business. And allowing business is to focus and double down on the things that they're best at. I think it's going to be good for the consumers of their business. And so I think we're just going to see like a lot like ruthless refocusing on what actually matters, which is going to be cool to see on that. So, you know, every company kind of has to go through this process of reinvention, you know, 10 or 20 years ago, people talking about digital transformation. And I don't know if anybody's given it a buzz word now, but AI transformation. Something of that sort. A couple years ago, you ran into a bunch of organizations that just weren't ready to deal with autonomous agents. Things you've seen your customers start to change. And so the question is, when you look at your customers, is they kind of go up this maturity curve and sort of reinvent themselves for the future? Any good like tricks or techniques that you've seen them use to repot themselves a bit? Yeah, I mean, I think, surprisingly, like the companies that have been doing like company wide hackathons really end up doing well, it seems like relatively trivial, but like just setting aside a day for everyone in the workforce is just like build shit with AI. It really sets the tone and sets the pace. So just give me a look. I tried to force him to build the coding agent. It didn't go so well. We'll do it after this. You know, we gave it a great effort. But that's it. Like it literally just setting aside the time to like do it. And like even if it fails miserably, like it's fine. And also like the orgs that are okay with failing. Yeah, it feels like there are some who are like, we need to do it exactly right. We need to make the right decision from day one. No, like you're going to make mistakes. Everyone is going to. And the orgs who are kind of leaning into it and embracing it to a certain degree, I think are succeeding. Like one of our largest customers is EY. EY is not necessarily known to be like at the absolute frontier of AI. But I think for them, they were just like, look, this matters. We were kind of there. There have been other transit transformations that we relate to. We're not going to be late to this. Like we're just going to go in. We might mess up, but like obviously respecting like the things that you're not allowed to mess up. Sure. Put those aside. But like let's go and get our engineers to mess around and build this stuff and see where it breaks and understand what they like and what they don't like. I think that really matters a lot in the ones that we're seeing succeed. And also the ones who are like pretty bold in reinventing the processes that they put in place. And just saying like, hey, there's no sacred cows. Like let's put this side, trusting out if it doesn't work, put that sacred cow right back. And I think that's that's been kind of a determining factor there. And when it comes from within, if it comes from the board, probably not going to go well. Yeah. If it comes from within like the tech team or the ICs or the leadership, that's when we see it go better. Do you have any predictions for the most important changes that are going to happen in your space of the next 12 months? A lot of AI consumption is going up like crazy. And ever and super, super excited because there are revenues going wild. Like a lot of this is synchronous usage. In other words, like if everyone woke up sick tomorrow, like a lot of cloud code usage would be zero. Because it's all just, hey, cloud code or hey, codex or hey, droid, right? I think in 12 to 24 months, like 90% of tokens will be asynchronous tokens. So these are going to be, you know, droids on their own autonomous, we're being like, hey, here's some signal that I found from a customer. Let's go fix it or let's go create a first pass solution to this. And I think that is going to be where the real like agent native stuff begins. Because right now we're still kind of in like co pilot mode. Like if you're going to an agent and say, hey, go do this for me, it is more agentic because it's not going to come back and ask you a ton, but it's still like you are kicking it off. Like yeah, if you guys have ever been to Tesla's factories, which was one of the sources of inspiration for the name is like, it's just robotic arms everywhere going and doing stuff. Like it's not like there are people there like going and you know, attaching the widget to the thing. And this idea of like a dark factory where like the lights are off and things are just happening. That is where software development is going. That's kind of where the name came from is like, Elon was always talking about the factory is the machine that builds the machine. And that's been something that we took to heart. And I guess also that combined with his whole thing about how you're destined to become the opposite of your name. And in our case, you know, factory becomes artisanal. It's kind of a good a good flip there. So what's your most optimistic version of the future both for factory and for the world at large. So I think short term, there's going to be a lot of turbulence because I think a lot of companies have misallocated resources pretty poorly. There's been a lot of bloat. And I think the correction that's going to happen there is going to be really painful for a lot of people. And I think that's something that I think every AI CEO should really bear much more responsibility than they currently are for. And also figuring out ways to like address and kind of ameliorate in some way because this is something that's going to be very painful for a lot of people. Now, I have optimism that we can actually address that faster than we think we just need to start now in terms of addressing that. Now, longer term and why I think this is a good thing is why I don't believe at all like, you know, the BS that people are saying, oh, engineers are going the way. Generally, there is a huge number of problems in the world. A large subset of those problems can be solved with software. A small subset of those problems are currently being solved with software. And so in the short term, this means that, okay, first, there's a given problem that was over allocated engineering resources. So, okay, we need to reallocate those. Reallocate those is a very kind of cold way of saying some people are going to lose their jobs. But I think the thing that's going to happen in a longer term is we need engineers. Engineers are some of the best systems thinkers and the best problem solvers. And there are so many problems that can be solved with software that are not being solved with software. And so that means that we are going to take those engineers and have them go and solve problems that previously were not being solved. That is such a net good for the world. Because again, there are so many of these problems that we are not solving. And also, there's so many problems that we are maybe solving, but with really shitty software. And like, this is going to enable people to solve it with incredible software. And, you know, the vision for factories that we are kind of the factory that allows them to go and build this incredible software to solve these different problems. And these problems range from like, things that are trivial to, you know, like, government software, typically is not very good, whether it's like DMV or like IRS, like all that stuff is generally a pretty poor experience. We don't need to live like that. Like, we can live in, we can live in a world where all software is really fantastic. But also things like pharmaceutical research, like so much that goes into solving diseases is not just like a biology problem. A lot of it requires the best software engineers in the world. And previously, those problems haven't allocated the right dollars to attract the best engineers. But now, because of what's happening, I think we will be much more closely allocated to, these are the biggest problems. Let's get the best minds and the best problem solvers to solve that. I think it's kind of our job as an industry to do that relocation, reallocation, as quickly as possible. So it's not 10 years, but maybe like six months or a year. Wonderful. Mattan, I think the clarity and consistency of your vision over time has just always been very inspiring. And then just seeing how much you've grown as a leader and how much factories grown as a company, even since last time we did this training, did that. So it's truly all inspiring. So thank you for joining us again to share with you up to. I appreciate it a lot. Thank you. Thank you. [Music]

Podcast Summary

Key Points:

  1. Matan, co-founder and CEO of Factory, discusses the company's journey in building autonomous agents for software development, emphasizing a "dark horse" position in a competitive market.
  2. Factory prioritizes model independence, allowing enterprises to avoid vendor lock-in with model providers like OpenAI or Anthropic, and to hot-swap models for performance or cost.
  3. The company's operating principle is "create obsessed customers," an output metric, contrasting with Amazon's "customer obsession" as an input metric.
  4. Factory was initially two years early with fully autonomous agents, leading to a "journey in the desert" where they refined their product and enterprise sales approach.
  5. A key decision was refunding all customers when the product wasn't meeting developer needs, building trust and credibility for future success.
  6. The launch of the droid CLI on September 26, 2025, marked a turning point, meeting developers where they were and achieving frontier performance, aided by behavioral shifts in developer adoption.
  7. The team's resilience, forged through early struggles, is a core competitive advantage versus newer, well-funded competitors.

Summary:

Matan, CEO of Factory, shares the company's evolution from a vision of fully autonomous software development agents to a successful product, highlighting the challenges of being early in a market. He emphasizes that Factory's differentiator is model independence, ensuring enterprises aren't locked into any single model provider, allowing flexibility and cost optimization. A core philosophy is "create obsessed customers," focusing on output rather than input metrics like customer obsession.

Factory's journey included a "two years in the desert" period where they were ahead of the market, leading to a pivotal moment: refunding all customers because the product wasn't good enough, a decision that built trust and credibility. The launch of the droid CLI in September 2025 was a breakthrough, meeting developers at their current skill level and leveraging improved models and behavioral readiness. Matan stresses that the team's resilience, forged through early struggles and a commitment to the mission, is a key advantage over competitors who haven't faced such adversity.

The conversation underscores the importance of adaptability, trust, and long-term vision in the rapidly evolving AI software development landscape.

FAQs

Factory's mission is to bring autonomy to software engineering by building autonomous agents, called droids, for software development.

Factory returned revenue to customers because their product was not creating obsessed customers, violating their principle of creating obsessed customers, and they needed to maintain trust for when the product matured.

Factory's principle is to build something so good that customers become obsessed with the company, rather than just being internally customer-obsessed, which is an input metric.

Factory is model-independent, allowing hot-swapping of models, and ensures that all automations and artifacts live in the customer's codebase, preventing proprietary lock-in.

In September 2025, Factory launched the droid CLI, which met developers where they were, was model-agnostic, and had state-of-the-art performance, marking a shift from being too early to being ready for the market.

Factory was two years early with fully autonomous agents, as developers and enterprises weren't ready behaviorally or technically, making the vision wrong at the time.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.