Go back

The misaligned incentives behind AI coding agents | Russell Kaplan, Cognition

50m 16s

The misaligned incentives behind AI coding agents | Russell Kaplan, Cognition

The evolution of AI agents has moved beyond raw model size to focus on efficiency, productivity, and cost-effectiveness. At Cognition, the development of Devon as a cloud-based coding agent has advanced through a focus on real-world usability, beginning with niche applications like code refactoring and security vulnerability remediation. The team identified a critical gap in evaluating agent output—not just correctness, but whether changes would be mergeable and maintainable, leading to the creation of the "frontier code" evaluation benchmark. This new standard measures both binary correctness and nuanced quality, such as code style and future-modification safety, through collaborative work with open-source experts and internal evaluations. A major shift in agent design involves routing tasks between high-performance and cost-efficient models—such as with Devon Fusion—achieving up to 35% better price-performance while maintaining quality. This approach reflects broader industry trends where specialized, domain-tuned models outperform general-purpose ones. The rise of proactive agents, which operate autonomously in workflows, significantly increases human engineers' leverage, allowing them to act as CTOs of virtual armies of thousands of agents. As agent capabilities saturate across tasks, the focus has shifted from model choice to cost optimization and operational efficiency. Companies now face the challenge of token spending surpassing human salaries, driving demand for smarter routing and specialized AI models. Cognition’s strategy—combining agent autonomy, specialized models, and customer-aligned ROI measurement—positions agents not just as tools, but as core partners in engineering productivity. The future of agent UX lies in proactive, context-aware workflows that integrate seamlessly into organizational processes, empowering individuals to achieve enterprise-scale outcomes without traditional years of experience. This democratizes software engineering, shifting the value of expertise from deep domain knowledge to mastering agent-human collaboration.

Transcription

11233 Words, 60536 Characters

English
I started my own machine learning career at Tesla on the autopilot team. I talked to my friends at Tesla today and the bottleneck is no longer just training bigger and bigger models. It's actually running the e-vals. Today, I'm talking to Russell Kaplan, president at Cognition, the company behind Devon, an agent that went from viral demo to deploying code inside some of the most complex orgs in the world. We essentially have an evaluator agent that can take a session and say, "Was it productive or not?" It gave us the confidence to actually go to our customers and say, "We are actually going to make a $10 million productivity guarantee." He explains why and how proactive agents are giving human engineers outsized leverage. You have all these great suggestions of fixes that need to be applied. Oh yeah, that looks good. I want you to change that here. Individual developers have essentially realized I can be the CTO of an army of 10,000 agents. We get into what running agents at scale actually costs and how Cognition drives it down. There's organizations where the purpose and token spend is starting to eclipse the human salary spent. By being a little bit more clever about some of the routing, we can get about 35% better price performance with actually a slight increase in quality. And he argues there's no such thing as the best model anymore. The fable class models are really good for high precision, but we actually find that GP 5.5 and 5.5 cyber are better on recall. You have to use both. Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. You guys had a massive launch about two and a half years ago, three years ago at this point and I think you pioneered a lot of really interesting concepts, especially around UX and interacting with Devon in Slack. I think that was the first time I saw that so prominently. So I remember a lot from that launch. I remember a lot about the UX, I'm sure others do as well. What should people know about what you guys have been up to over the past two and a half years? Yeah, we launched Devon in March of 2024 was the original Devon video that went super viral. And at the time, I would say coding agents were just at the edge of possible. You kind of squint and see, okay, this is going to work at some point. And we sort of put together a first pass example of what could that look like. We bench was like 13%. I think you guys went in and tripled the previous page. Totally. It was like, yeah, we were like, oh, 13% sweet bench. This is a really exciting and it actually that by the way, that number reminds me a lot of some of our more recent e-veils and where we are now in agente coding with a new e-value we put out called frontier code, which we can talk about. And when we, you know, we kind of launched in March of 24, we had this point of view that eventually you're just going to have agents as teammates that you can delegate complete units of work to. And what we learned is that took us a few months from, okay, this is a prototype that we can kind of see the future of to this is something that's actually useful internally. I think it was June of 24. The Devon became the number one committer to Devon, which was like our first big milestone. That's still a while ago. It was a while ago, but it was like a lot of manual dog fooding and like really kind of grinding to make the internal workflows nice and then it took us a few more months to actually get this deployed in production and useful at a customer. And in, in the sort of 2024 era async cloud coding agents, they really couldn't do most of the tasks in software engineering, but there were some, you know, niches already. And I think like one of the best early ones was migrations and refactors. If you ran essentially a, you think of like a redgex plus plus style workflow across a large code base with sort of bits of intelligence sprinkled in, that actually already works reasonably well in late 2024. And so that's where we got our, you know, early product market fit. It was with bigger companies, you know, like more like enterprise companies who they just had lots of code that needed all this transformation. And why, why was that such a good fit? There's a few like technical reasons. This made sense. One is for a very large refactor or migration or, you know, ETL transformation, whatever. It's sort of worth it to put in the effort to carefully prompt engineer your cloud agents to be accurate. And so, you know, if you, if you kind of iterate on your prompt and your setup and your, you know, the context you're feeding in and you, you get it like just right and you can apply it across, you know, 10,000 modules in your code base. This is actually really high ROI and you could, it's obviously much better than just sort of a final replace, you know, string change, but it didn't have to have this like full general software intelligence that we are now delegating all sorts of coding tasks to. So that worked well even the early days, but it was still kind of niche. And then in December of 24, we launched Devon, so anyone could sign up and we had this big debate internally of what should be, you know, what should we really emphasize or what should be the focus? And I remember we did like that we just, we flew the whole company to Utah to just like lock in in December and get this out, you know, kind of before the end of the year. And what we settled on was actually slack as the primary interface. And so the entire launch video for Devon available self service, we called it internally @Devon. It was just @Devon @Devon @Devon and really trying to drive home this sort of user experience change, which is collaborating with agents more like teammates. And so from there, we got a lot more users, we got a lot more traction, you know, cloud agents have just gotten better and better since then, both as the infrastructure has gone better, as we've matured on that, as the models have gone better, as the harness has gone better. And now I think most of the work is being done by just delegating to ACE engagements. Yeah, maybe diving into that a little bit. What are some of the things that have gotten better that have allowed them to really take off? So a few things. I think people always talk about the models getting better and that's super important. But I think the kind of infrastructure maturity is always what was the first model that you felt was like really actually good enough. Yeah, it's, you know, we kept saying, oh, this is the model. This is the model. You know, this is the new model. If you like face changes, I would say I think sort of like the November 2024, like model series was one kind of, that was like a big step function where it's like, okay, this actually work a lot better. Now we did a self sort of launch, you know, in part based on that from open AI. And then, and then, you know, the anthropic model started getting really good. And we saw, I think probably like, you know, July of 25, another big step change. I think now we see, you know, with like, Fable 5, it's like new generation of models continue with step changes. And now the step changes are so high, there's actually no longer about the capabilities increases. I think we're kind of seeing the opposite trend in some sense, which more and more of the tasks in software engineering are getting intelligence saturated. So like literally all of the models can do them well. And so the sort of the sensitivity of a lot of developers has totally shifted from, wait, I need to be using the best possible model to holy cow. I am spending so much money on my coding agents. How can I, how can I be a little bit more efficient, you know, about this? And I think, I think this is like an underappreciated aspect of continuous gains and frontier intelligence, which is that, you know, for some workloads, you always want the most frontier intelligence. Especially if it's like an adversarial game between parties and, you know, your intelligence has to be higher than the counter parties intelligence, like if you're trading in competition or something. But for a lot of software we want to build, you know, there's kind of this, this saturation threshold where once it's good enough, what you, what you, what you care about is speed and cost. And I think that's actually where we see a huge amount of demand from developers and the companies we work with is, okay, this already works. How can I now optimize the speeding cost on the model side? Maybe one question. What are tasks where you think, fable type models are like where you see that jump today, like where, where would you recommend people use them? Where do you see people using them? So I think we're seeing maybe in like this most recent generation of models and fable being one example, um, particular big gains is in, first of all, there's like detection or mediation of cybersecurity vulnerabilities. And we had, there's a lot of guardrails now in the public versions of these models and so it's actually can be tricky to, to get them to help. We see, for example, that if you're trying to go just like clean up my security backlog, which by the way is a huge workload right now for a lot of the, a lot of the companies we work with, you know, the, the, the fable class models are really good for high precision. But we actually find that, you know, GP 5.5 and 5.5 cyber are better on recall. And so the best possible harness for finding and remedying security vulnerabilities right now, you have to use, you have to use both and you sort of filter down, you know, the ones that are maybe better found by the GP series with, with the fable series. I think that's one big example. Another one is that processing data from integrations, for example, from like data dog or, um, or other systems, it's, we do see like a step change in our e-bals in, in fable in particular. Is that because that data is like so large and messy that needs more intelligence. I think a lot of, yeah, it's just like there's like a volume consideration. There's sort of a multi step reasoning consideration. And I think there's just, you can tell that there's a lot of RL going on on like these very realistic data sources that have been scaled up a lot. Maybe the third category that's maybe the most noticeable is a category that we've only even gotten better at detecting and understanding recently ourselves with the work we've been doing on new e-vals for coding because to your point when we launched Devon, uh, and it was, you know, uh, in the teens on sweet bench, now sweet bench is totally saturated. And so how do we find, how do we find the next level of difficulty for evaluating models? We really, we looked for all of the e-vals and we just, we couldn't find any that could actually kind of match our sort of hazy internal intuitions of, uh, this model feels a lot better. And when we sort of dug into it and really asked ourselves, why is that? I think that core gap for e-vals that we found was around merge ability, which is, okay, this code is technically correct, but like, would you actually merge it? Like would you feel happy with this improve the quality of your code base? And there's lots of different subtle examples, um, of what that can mean. It could be stylistically following the expectations that you already have. It could mean that it's done in a way that it's easy to modify in the future or that if other modifications happen, it gracefully handles those modifications, all these little stylistic things, which we felt like we're not being captured in, uh, in the e-vals. And so we, we set out to basically make a new e-val to, to measure this better and also just have a new high watermark of difficulty, uh, for frontier coding capabilities. And then we released it recently, um, it's called frontier code. And it's once again, uh, it's, it's, it's, it's the new hardest coding e-val, uh, the frontier code diamond subset is like in the teens, uh, you know, uh, uh, you know, uh, pre-fable five and now I think failed five has, has pushed to the thirties. So we have, I wonder how many more months we have before that Evalus, I was going to say, yeah, I'm like a year, we'll be saying, oh, I remember when I first got in the teens for this one. Yeah, I'm wondering, I do wonder, like, it's going to get harder, you know, it gets harder and harder to make these Evalus. I mean, I started my own machine learning career at Tesla on the autopilot team. And this was in like 2017. And the bottleneck was very clearly, you know, GPUs and compute. And it was like, okay, how do we get more compute? We just need more compute. Yeah, I talked to my friends at Tesla today and the bottleneck is actually no longer just training bigger and bigger models. It's actually running the Evalus because for a self-driving system, the interventions are so rare that you have to do enormous volume of driving to even find any problem at all in the stack. And, you know, I think for software engineering, we're not very yet, like, we still find bugs, but we might, we might hit that threshold sooner than we think. We talk with a bunch of teams around creating Evalus for their own systems, using linksmith and other things like that. How did you guys create the frontier code? Like, what did that process look like? So it started with, okay, we need to go find the sort of peak high-taste developers who are going to have really strong both kind of stylistic quality, but also code correctness standards to help us answer the question, would you actually merge this code? Not just is this code passing tests? Is it, is it functionally correct? So that's kind of part one. And so we actually did like a, you know, a deep collaboration and recruiting campaign with a lot of the sort of leading open source authors who were, you know, really grateful for their collaboration on this. It wouldn't have been possible without them to go find, you know, what are the best known libraries in open source with, with really high-coding standards, and then collaborating closely with them to help encode their, their human intuition of what makes a PR something I would accept into these very carefully designed tests. That was part one. What, what would those tests look like? Are they OM as a judge? Are they programmatic assertions? Yeah, so multiple. So part of them are program assertions. Of course, there's like a correctness test programmatic assertions. There can be programmatic assertions on kind of stylistic elements too, like, okay, if you make this change in this pull request, you know, you have to actually use this module, even though using this other module would be equivalent right now, over time if they drift, you know, if you didn't use the correct module, then like you're going to introduce a bug in the future. So it can be asserted to just deterministically, but it requires the judgment of, you know, the human author to, to settle. And then the other big thing, which I think is still underdone by people practicing, she learning and working with agents is every single researcher on the cognition team, you know, hand contributed and reviewed and evaled the evales directly. And so this is not, this is not something you can sort of throw over the wall and say, okay, the data quality piece, that's, that's such a slog. Someone else can do that. You sort of have to work on it directly too. I'm assuming they worked with the open source authors to put together these assertions and, and they would be running on these open source code bases, I guess, like that was kind of the setup. It would be on the open source code base, and you know, we could run it in our infrastructure, but then we test and you know, to avoid contamination, we haven't published the full set of questions, we publish an example questions, but we're trying to preserve this eval for the community for as many months as possible until it gets, you know, completely saturated. But running running, yeah, essentially running the code that already exists with the patches applied by its models. How do you guys score these evales? Is it binary like zero one? Is there some numeric component, like if you have 10 assertions on one of the tests and it passes nine out of 10? How is that scored? Yeah, so we broke it out into two components to have first, there's a binary element, which is essentially yet did this pass all of the blocking constraints of how we would evaluate this PR, you know, the simple things of like, do the test pass? Are there hard, are there hard deterministic criteria that the maintainer of the code base has inputted must be met by the, by the agent who's writing this code? Do you know off top? Like, is that like five assertions, or like 50 assertions? Like, it's highly varied per task, but each task is, you know, like hundreds of hours of work of people putting into. So it's, it is like, this is not like a quick thing. It's the process of constructing a single test, you know, eval line item in a data set like this. It's, it feels like a, it's like a full project and labor of love per question, you know, where you're really trying to think holistically, what are all the things that I, as the expert maintainer, want to sort of imbue in my expectations? So that's how you get some of the binary, you know, pass, pass field metrics. But then we also, to your point, that's not enough to, to really know, okay, what I merge this code, and then also like, how would this code rank to someone else's code? Because it might be, you know, two different pieces of code you would merge, but one is like a little bit more preferable than the other. And so we also have the concept of score, which is essentially like a linearly weighted aggregation of all of the non blocking evaluation criteria. So if it's blocking evaluation criteria, we'd say, okay, you know, if it doesn't pass, you fail. But like, or this is just, you know, we're not, we're not merging it. If it's non blocking criteria, for example, there's sky, stylistic elements, a common one for us is scope. So I think one of the code smells of LLM still is like, they'll make the right change, but then they'll kind of mess with other files too that you really wish they didn't mess with. And so that's not going to impact your correctness score, but it's certainly going to impact this sort of stylistic elements. And that's going to downweat, you know, if you had unnecessary touches in other files, that's going to downweat your aggregate metric. Same there for, you know, LLM is a judge having cauristics that you impose in LLM as a judge can get added to this, um, to the, to the linear score. We also found that, um, reverse classical evaluation is, is really helpful. So commonly people say, okay, we need to make sure that after you accept this code change, the test pass, right? But we also care just as much that, um, you know, without this code change, the test should fail, right? And that, and if you make this other code change, the test should fail. So how do you, how do you impose kind of blocking constraints on both sides? But yeah, emails are really hard. You know, we spent a lot of time on this. I think, I think we got it, you know, we, I think we did it right enough that all of the current LLM's, uh, are pretty bad, are, are pretty bad at the CVL and it also more importantly to us, you know, as new models come out and we test them and we see their scores on frontier code, it's, it's roughly matching the vibes. Like, like, if, if the model is scoring really well in frontier code, then we get really excited. I like this idea of kind of like binary plus fail and then some numerical after that for the, or more explicit, I like the like blocking and not blocking. So we're building a benchmark of our own. We call it like issue bench for links and ascension, which goes through and finds issues. And there's a bunch of stuff that like, yeah, I would classify under like the non-blocking stuff. Like it's, yeah, it would be nice if it did this. It should probably do this. And then there's some other stuff that's kind of like more blocking. Does it just find this like really bad issue? Should be more of a blocking thing. So I like that kind of like dual objective position. Do you guys use harbor as a format for running evals or do you have your own internal kind of like evil? We do need harbor. We think in general, a lot of the kind of standards are quite early and immature. And I think one of the funny things, if you sort of look in the whole Devon code basis, because it was the first coding agent, there's actually a whole bunch of stuff that like there might be a standard now or a correct way of doing things. And we were just, we just rolled our own, our old implementation, you know, is funny example, funny examples. So even a basic agent things like the concept of, you know, skills.md file, like we had implemented a concept in Devon of knowledge that was like before any open source standard existed for skills. And now we've basically grafted on how can you also work with the open source standards, but it is like a recurring theme in our code base that we've gone off and sort of invented something weird. And then it becomes some some permutation. It becomes an open source standard that we incorporate back in. We went down this rabbit hole talking about models and talking about the best models. But you also mentioned kind of like cost and presumably kind of like speed and latency are becoming other issues. You guys have done a few things here. If I'm correct, you have Devon fusion. You also have your own series of models that you guys have post trained. How do you guys think about this, uh, this section of the model universe? We're kind of in this unique time in history where anyone can hire as many AI agents as they want usually the way the tools are work. So, you know, uh, me as a developer, yeah, I'm not going to, I'm not going to use the cheap, the cheap model. I want us, I want to use the best one for everything. Um, but everyone is kind of collectively making that decision. And then you sort of roll it all up and you realize, wow, like, we are spending a lot of money. You know, there's organizations where the, the per person token spend is is very rapidly approaching or even starting to eclipse the, you know, the human salary spend. And so it would be kind of crazy, you know, if, if like, the way you ran Langshane, for example, anyone could just hire a thousand people tomorrow without, you know, without talking to you. But that's kind of how we run our teams today. And this is starting to become like a real issue for us. And like, I like to think we're like, we're pretty, like, you know, we're still start up. We're a lax like AI native kind of like forward organization. But we are absolutely caring about kind of like token spend. I'm like, do you, are you guys internally kind of like worried about token spend for your own kind of like engineering teams as well? We have very high, uh, you know, compute investment in general. So I would think our internal spend while being very high per person to us, it's like really valuable dog fooding investment and everything we do. But our customers are definitely thinking about if you just extrapolate the trendline, you know, it's going to like eclipse the whole economy in on that long with the with exponential growth. And so, so, so people are asking the question, okay, well, how, how should we approach this? How should we even think about this? And I think we now, there's sort of two different trends that are happening at the same time that sometimes we get conflated. So, so one is obviously the models are getting a lot better and more marketable. What's happening is, you know, if you look at the most frontier capabilities and then maybe the set of models that are just behind them or a little bit behind them, every time we have these new generations of models, we, we both move the high water mark on the frontier, but then the set of tasks that all the models can do is growing a lot. And the distribution of tasks that people are like trying to accomplish that are any day in their everyday lives are changing a bit, but actually not that much, you know, and so it's like, okay, I'm trying to build a front end, you know, for my marketing registration page. It's like, that's not that hard if it's that, right? And so, so I think The way I think this plays out is is as more and more models do more of these things the marginal returns to investing and just like having a reasonably intelligent harness they can make sure you're using the right model for the right job in the right moment. Uh, it goes up a lot and both for individual developers who you know they want a fast answer they want a correct answer and they don't want you know they don't want an unnecessary ways or even really like. The cognitive overhead of having to decide every time you interact with the coding agent. Oh, like what model should I use for this one? I was going to ask do you let users of Devon choose what models they use for Devon desktop, which is what we rebranded windsurf relatively recently and for our CLI we do people developers like the individual control when they're working with local agents. Um, for our for our cloud agent we don't it's a you know we'll have like Devon fusion which we talked about which is the sort of. Um, frontier performance with cost optimization option will have like a more affordable agent we have you know. The like the max or ultra agent so we have this like sense of kind of tearing, but I actually think it's like a UX bug. In the fullness of time for people to have to think about this for their cloud agents and you know we saw these early days we ran some tests where we let people pick the models and then two things happen where one is sometimes people would then give a task to Devon and it wouldn't work and they would complain. Devon was so dumb here and we looked at well, why did you pick this model? But actually that's kind of more our fault I think than the users fault and then the same thing would happen where. They would run tasks like oh this was like really expensive uh well why did you use this very expensive thing and so. So I think this sort of natural equilibrium is people want the best performance for the best price without compromising on correctness right and so. Folks who are building agents I think have in some sense like a user experience responsibility as well as you know just building good price performance products to try to do as much of that optimization as possible and that's what we recently put out. Uh with Devon fusion which is the sort of next generation of our own harness designed for. Frontier level capabilities but with maximum price performance and so what we found is like by being. A little bit more clever about some of the routing and the decisions of like what you're doing with which model and and thinking about this in a very cash aware way. We can get about 35% you know better price performance with actually a slight increase in quality um and I think that's sort of where a lot of this needs to go. For the technology to be continued to adopt it the rates being adopted so talking about that like a little bit more because there's there's there's there's routing and then there's also open router launched open router fusion which is not routing it's running it on multiple models and parallel on then combining things back together. So when you guys have Devon fusion is it is it routing is it like running multiple things yeah so it's doing both so so we we put out a technical blog post that shares a little more detail on our implementation um but one core component of it is this idea of a psychic agent. And so you know what we'll do is we'll have the sort of frontier quality model. Executing on the task and then in parallel we'll have a more price performance model exit on the task and then there will be some decision making that the frontier quality model has to do. Of okay well when do I delegate to my sidekick but having them both work on the same task in parallel it lets you make sure that there's still context for both of these agents and we try to share a lot of context to and we do you know rights to the file system frequently if there's if there's things that are exploding out of context. But having basically both work in parallel then frontier model knowing okay I can pass us off and actually as the malls get smarter the most frontier models get smarter smarter they're one of the key skills we see is there like way better at delegating to like fable is very good at knowing ah this task I can delegate to a dumber model and it's going to be okay which kind of makes sense if you think of human career progression also like one of the aspects of of growing in your own career is learning how to delegate tasks and how to like do the highest leverage tasks. So we're kind of see the same technical pattern in in the models you guys have your own set of models as well we 1.6 yeah I think it's the most recent one why did you guys train that what do you see people using it for yeah this is a great one so sometimes people ask us you know why even RL or post train your models are like aren't the next models from the frontier labs just going to be better and better and better and better and definitely and we get super excited every time new frontier lab models come out that are better at their task but I think there's two important reasons. For us to be spending a lot of energy on our own RL and post training for our models so the first is that you actually can deliver frontier capabilities at any given moment in time through greater specialization and I'll give you a recent example which is we shipped the product we shipped a product relatively recent called Devin review which you guys are are using really interesting ways inside inside Langchand yes to deliver Devin review we wanted to not just have sort of a human interface for understanding deaths and kind of rocking large amounts of AI generated code quickly but also be able to run quick static analysis and lightweight kind of model driven analysis on are there bugs in this code are the security vulnerabilities in this code are the things that we should be sort of linting and automatically checking we can use some frontier models for that we can just shoot. Super models but then you get really tough tradeoffs on price performance curve you don't you don't want to have to spend tons of money just by virtue of having created a new PR and that's the exact type of problem where very specialized RL and post training can lead you to producing in extremely price performance model for the very specialized task and we know that that model has a half life that's going to expire and that's fine because by the time it expires will be very happy and we'll be working on the next specialized model for the next workload and so I think a lot of kind of building an AI startup right now is is. It's being very willing to think in these like three to six month increments of OK given the state of frontier capabilities today what is the most differentiated set of new product experiences I can build through my own specialization that's part one the second reason we work on it's really helpful is it does go back to price performance which is as more and more models are capable of doing more and more things how can we continue to deliver. The best possible price performance for our customers while it starts with if any model can do it well you know we should be able to serve our own models even faster even cheaper than anything else on the market and we see that you know sweet 1.6 actually the most popular model in Devin desktop by number of tokens consume it's about 4.6 level and we'll have more to say on on new models on new models coming out soon there's going to continue to be this like frontier of capabilities that use the most frontier models for and startups not just us I think more serve should be considering how do I. Get how do I make my own models that are specialized for my domain because more and more of the tasks in my domain are going to be doable by any model how many of these specialized models you guys have any given point time is it one for review one for coding and those are separate ones or yeah order of magnitude it ranges yeah I would say it's like single digits number of specialized models because because part of our own focus as a company is on software engineering and so you know having like the sweet 1.6 series that model series is something we can we can drop in and use at a lot of different places so. We don't need to manage like 50 or 100 different of these super specialized models but for the areas of our products that we think it's going to make the biggest impact not just on performance by the like kind of quality can be literally on the latency and speed and and you know for ability we want to we want to make sure we continue to invest and you guys optimized a lot for latency with sweet 1.6 right yeah one thing that was really fun. We're going to use we were the first kind of western firm to deploy cerebris at scale you know when when we tested their their chips we were really confused because we were like this is really good you know that it's it's it operates a unique point on the perito curve of price and throughput and you know we're getting about 95 seconds per second on our models which was many acts faster than what we could get on on GPUs for the same size of model and it was a little more expensive for us to serve. They were great partners and they enabled us to ship a product experience that basically wouldn't have been possible otherwise and now it was more popular you know now you know open house use for for spark I think more people are considering them and. But we want to constantly be basically pushing the frontier of well what's the next most possible thing that wasn't possible previously and yeah doing our own our own training helps with that. What do you think the next most possible thing that isn't possible right now will be as far as like model serving I mean I'm really excited about you know the sort of. Continue generations of very very specialized inference as the architectures that we use for transformers have gotten a little bit more standardized more predictable you know the people designing chips. You can make a lot more specific assumptions that enable really big really big meaningful technical gains I'll give you one final example from when I was a Tesla. So Tesla we built our own chips in house for inference on autopilot and it was a big part of why autopilot could work as performantly as it did. And remember we had I was like a deep learning researcher was one of my jobs was like interfacing with the Silicon design team to make sure that the next generations of silicon supported our needs. And you know a lot of the bottleneck in chip design it's actually thermals it's like how hot can you run without melting the chip that's the amount of power you can put in and then even within the amount of power you you put in a lot of that power is consumed by memory movement as opposed to a supposed to flop. And so we're trying to be really efficient of okay like what are the things that we don't actually need on this chip that we can take away so we can really run at max throughput for the stuff we did care about. And we had a big debate one day on whether the chip actually needed to support division because they're like well you know if we don't have to support division we can make this way faster and if you look at that time it was in a commercial network the layers in a convolutional network at inference time none of them actually needed division you know for normalization we were using batch formalization which has a division step but you know you could actually fuse the batch norm weights and biases into the into the condo weights so I think we're going to have crazy levels of specialization because there's going to be so much inference demand and that's going to unlock just like super exciting levels of performance and so when we you know we started off talking about all these things that kind of like made. coding agents and Devon like way better over the past two in half years, all this infrastructure stuff at the chip layer, I'm presuming that's part of it. And there's probably a lot more stuff to go in that capacity. Every layer of the sack, we try to think about every single layer of the stack. And one of the layers I'm most excited about right now, it's how deeply can we integrate with the rest of the team and company, you know, organizational context and way of doing things. So I'll tell you like the biggest internal difference I feel using, you know, Dev and a cognition versus a few months ago is we've gone really hard on this concept of automations. So how can you wire up your agents so that they're proactive, not reactive. And you know, in steady state, it's probably going to be the case that most of the engineering work at a company, it's done proactively, autonomously by default by your agents. And humans are still going to be making the decisions, but they're really going to be driving the new bets. And well, what are the sort of change my company trajectory, you know, technical directions to take on versus you got some user feedback, there's a bug report, you know, there's some crash needs to investigate all that should be handled proactively by agents. And so, you know, we have Dev and wired up to a bunch of our slack channels triaging every message saying, is that something I should chime in on? Is it something I should investigate? And the sort of role of being a human engineer, it's just getting like crazy leverage. And so you have like all these great suggestions of fixes that need to be applied. Oh, yeah, that looks good. I want you to change that here, and in particular, having the, having the agent have the context of who's responsible for what in, you know, in the code base in the company. So it knows, oh, you know, Harrison, you should review this change or I should review that change. It's really, it's really quite exciting. I'd love to hear you guys think about UX, because you started off, I think, and made kind of like the slack coding experience, like I think, I think of you guys when I think of that. You mentioned you, you guys teamed up with Windsurf, you also have Dev and CLI, you're now talking about kind of like a different type of UX where you're not kicking things off, but it's coming to you. How do you see like the UX of coding agents over time and in the future evolving like where will humans be in the loop and what will working with these things look like? Yeah. So I think individual developers working with coding agents are going to ask themselves the question of how do I create self-driving software? How can I do sort of, you know, fix costs upfront work for continuous productivity benefits and gains? Well, the term software factory means to do, I think some people use that to me. It's sort of the wrong term, because when I think of a factory, I think of like the same thing being produced, you know, again and again versus the whole magic of coding agents in this like proactive shift of engineering, is that it's the opposite of that. It's bespoke. It's fully custom and with all the right context of exactly what's going on. So I personally have never been a fan of the term, but generally this idea that the baseline engineering and technical function of your company is going to just be getting better and better by default because it knows what objectives it's optimizing for. It knows the input context coming in, it knows how to react to that and it can learn from your feedback over time. I think it's super powerful. And we see this within the companies we work with too, where individual developers have essentially realized I can be the CTO of an army of 10,000 agents. And maybe the most acute use case that's just growing like wildfire right now for us in this regard is vulnerability remediation because you know, there's sort of this generational shift in the threat surface to a lot of the especially large firms given how good these models are getting at cybersecurity, both offense and defense. So there's kind of a lot of urgency right now on, okay, let's go patch all the vulnerabilities that we have, but how do you actually do that? Well, you know, a developer probably has to sit down and think, okay, how can I really leverage Debin to go find all the different issues and then remediate them and mass. So you can have one person sitting off, you know, like an enterprise wide API call that's doing the equivalent of thousands of engineers worth of work. How do you do that? Right. So for security specifically, there's enough nuance that we've actually shipped pretty like vertical specific futures and product services in security. We announced something recently called Debin security swarm. One way to think about is agentic map reduce for finding and then fixing security vulnerabilities. One of the original advantages of Debin is that it was the first cloud agent. So every, every session is running on its own micro VM. And because of that, you can actually safely reproduce potential security vulnerabilities and replicate them and then validate that they're fixed. And so there's a sort of a scanning phase where you have to figure out, okay, of my massive code base that doesn't fit in the context window of any one LLM. What are the sub parts that are actually vulnerable and you have to sort of shard it out. And then once you find those parts, how do you fix them, combine them and aggregate them in a way that is technically correct and actually patching the issues? But that paradigm of one engineer is now empowered to take on the whole company, make something better. I think it's really exciting. It's like, each person, there's nothing stopping you from just having way more output than you could have had only recently. And so that's a really big change and it's like changing the skills of the skills that are most valuable, I think, of being an engineer, thinking much more big picture of, okay, like what are the most important things that we have to get done? How can I do this massive side of changes to get them done? And then leave behind a system that is self-improvement. Who has those skills? And kind of what I'm getting at is like, do you have to have a lot of years of expertise to understand how the whole system fits together? And if so, how will people who are just graduating now get that expertise in order to do this new skill, which is now maybe the only skill that matters? Yeah, sometimes people ask like, oh, what's the future of being like a new grad engineer? Because how are you even going to get the experience if like the coding agents are so good at new grad engineer level tasks? But I think the flip side is everyone has barely any years of experience when it comes to working with agents. So in some sense, we're kind of all starting at the same level. And there's a new skill ladder to climb, which is how can you work with agents really well? And the best way to do that is to just try and just learn and do stuff, consume the resources and the training and the reading materials, but actually just go build stuff and then see what's working and not working. And in particular, keep your sort of finger on the pulse of the current limits of the frontier of capabilities and constantly reassess that. I think one of the biggest mistakes people make when you're trying to learn a new way of working is you sort of try something and say, okay, that didn't work. I guess that doesn't work. But as we know, in AI every two months or three months, you have to continually retest that because these systems are getting so much better. I actually say that what I found both internally at cognition and just working with Debin is you actually have more relative advantage as a new grad or as a recent engineer because onboarding has never been easier. You can ask as many questions as you want, including the silly ones to an agent that will not judge you, it will help you understand exactly what's going on. It will point you to the right parts of the code base to say, oh yeah, like this is actually where this is implemented. This is where it's done. And so we're seeing the opposite where people who would never have called themselves engineers not that long ago are steering the creation of software in ways that are totally crazy. You know, we started focused on professional software engineers and that's still I think the main users of Debin. But so many people are now using it to make things that we didn't think possible even though you know, that wasn't our original focus. You mentioned a few of these like specialized versions of Debin, so Debin review, Debin and Security Swarm. I think there's also this idea that coding agents are general purpose and can be used for everything. How do you guys think about what should be like a special version of Debin versus just, hey, just go use normal Debin for that. Yeah. So we debate this internally a lot. I think one of the original kind of like founding theses of the company was that, yeah, if you solve code, you can kind of solve a lot of other things. We are seeing this, I think, play out in real time, both in terms of the popularity of coding agents and in, you know, the revenue numbers of other people working on code and really the whole adoption profile of code. We don't want to over engineer something that then the next version of Debin is going to make, you know, completely outdated and useless. At the same time, part of the role we serve in the ecosystem is to be the independent agent lab. So how do we take the best of every underlying model and then go really solve our user's problems and to end like in the very specific details of what they're working on. So we're still super focused on software engineering as the core of what we do. But more and more things touch software engineering and so, you know, one of the reasons we did this like specialization push on security is that we just looked at the distribution of how people were using Debin and realized, hey, that's like one of the most popular uses of Debin. So how do we double down on that? And I think in general, like my sort of personal philosophy on product development is, you want to be allocating, kind of some portion of your time to fixing the frictions and the paper cuts that your user experience. And then you also want to be allocating some portion of your time on doubling down on positive surprises. What are the ways that people are using Debin that we didn't even think of? And then how can we double down on that to make their experience even better? And a lot of times doubling down is, hey, let's like try to just improve general agent capabilities. But other times it's, okay, what are the more specific interfaces that we should build to make it work well? You mentioned being an agent lab, there's also these model labs, which you are competing with at times, but also using a lot of their models at other times. How do you think about that and what does that landscape look like over the past few years? Yeah, so we both like have really deep collaborations with the model labs and, you know, and use each other's stuff and frontier models are a really important part of Debin's delivery. And then there's definitely, I think model labs not just encoding, but in all of the application domains are starting to get into the application layer too. I think for us, we've learned a few things, you know, working in that space. One is that there is actually a pretty interesting structural differentiation of where we set in the ecosystem versus any one model lab, which is whether you're an individual developer [BLANK_AUDIO] the CIO of a large enterprise. The only thing you can be really sure about is that the underlying models are going to keep changing and it's pretty hard to predict in advance who's going to have the best model, who's going to have the best price performance model and so on and feedback we pretty consistently get from the individual users and the leaders and the CTOs that we work with is that it's really nice to have some sort of decoupling between the agents and the models so that as the models keep getting better you know stuck on you know the wrong tooling because now your model is not isn't the longer the best you you want to be in a situation where every day when you wake up you're very happy to hear that new model has come out that's even better than before regardless of who came out. I think the second big lesson is just focus on the last mile of complexity and problem solving that you know is most important for the customers. So we were pretty early in building you know for-deployed engineering team for example today our for-deployed engineering team it's we have more for-deployed engineers than non-for-deployed engineers at cognition and so you know by the numbers we're actually mostly working with customers because the core product is is Devon working on itself most of the time. What does for-deployed engineering mean for you guys because I feel like everyone's got for-deployed engineers and sometimes they mean different things. So for us it means folks who are very technical but also good at understanding customer and business problems and we can go on-site with customers to help solve big outcomes. The way we work with the large enterprise it's very different from how someone signs up on devon.ai and uses the product it's often pointed at a specific objective to go solve. It's not just about here is the tooling you're using but it's hey you know we have got to move onto this new system by the end of the year or we're going to have a lot of problems and our current schedule has this taking four years. How do we do this you know much faster and better and part of I think the appeal for individual engineers to be for-deployed right now is that you're kind of flexing a lot more muscles as as an individual engineer where you're you're still in the technical weeds of what's going on but you can be that that CTO of an army of a thousand agents and like work on your own skills to just sort of master this new way of working and that's like really valuable for us but also really valuable for for our customers what's the right profile of someone who wants to be a forward deployed engineer at cognition. We have had a lot of success with ex-founders in particular so people who are very comfortable in ambiguous kind of problem areas and you know can do the technical parts but also can understand customer pain points talk to them and and really have very high internal agency to recognize oh yeah that's that's the problem we should go solve which is another big thing we've learned in the process of deploying Devon more and more widely you know you can use coding agents for almost anything and so one of the big questions is well what's the most impactful thing or set of things to get started on and helping steer that is I think one of the things that makes like a great forward deployed engineer so great we talked about this a little bit earlier but I think costs for coding agents are becoming something that a lot of companies are paying attention to and the other part of that is kind of like the return that you get for everything that you're spending on coding agents and quantifying ROI is really hard I think you guys have done maybe one of the best jobs at or at least maybe even just one of the only people I've seen actually try to take this on in a general way you wrote a blog post on like estimating productivity and you have some productivity guarantee could you talk about how you guys think about that yeah we think about this all the time so first on the productivity side what's the problem that the problem right now is that a lot of people in the industry have a strong incentive to get customers to token max you know let's build the leaderboards of usage and kind of had things go up and the incentives are actually quite fundamentally miss a lot I think and at some point you know the bill comes due and I think we're starting to see that now at a lot of customers where they are they have been token maxing they've been using a lot and now suddenly they're spending in the tens or hundreds of millions of dollars or even more in some cases a year and then the the leadership is asking what what do we get for this so there's kind of this reset happening going back also to your earlier question of like where do we sit in the ecosystem as the independent agent lab and then one of the side effects of being independent is that structurally we're the most incentive aligned with customers to not just focus on oh you got to drive usage of this model but really focus on well what's going to be the value you're going to get and let's not even work on the stuff that's going to be low value so that was the end of one thing we've focused on early on and we've always tried to work closely directly with customers to just help them estimate and then quantify their own ROI of using of using Devon and an AI generally but part of that's a manual process it's literally sitting down and saying okay well what are the most important things for the company right now and how do we you know how could we move the needle together if we worked on them closely what we were looking to do is answer a sort of a more scalable question which is how could we make this happen in an automated way so everyone can benefit whether you know you're a big enterprise or an individual developer using Devon how can everyone benefit from that and so that immediately imposes one design constraint which is okay how can you automatically estimate the productivity or ROI value and I think our conclusion was at least today twenty twenty six it's very hard to automatically estimate ROI because you you don't have the full business context of oh you know this is going to drive your you know your revenue profits up this much or this new product is worth worth this much to you but we could do one level below that which is productive engineering output so there's a difference between I'm using an agent I'm getting a lot of tokens out a lot of code out and oh this was actually valuable time savings for for me and so to measure that we actually worked with a bunch of our of our customers to collect a data set of a bunch of Devon sessions where we scored them together with the customers say okay first of all was this actually productive or not and some rules you can do automatically so for example if you create a Devon session and that creates a PR and never gets emerged we call that an unproductive session now in reality like maybe it was still actually helpful for you but we want to conserve it so we say okay that's just blanket unproductive session if you merge a PR production then we say okay that's productive if it didn't result in any PR because sometimes people are using Devon for data analysis or other workloads then we do an ML driven classification based on kind of a manually tagged data set in collaboration with customers so the result of all that is we essentially have a new evaluator agent that can take a session and say was it productive or not part one part two is if it was productive how many hours have worked at that actually save you and again manual data collection grind to figure out that question where we surveyed a bunch of folks who were closely with them to see hey you're still using AI maybe even if not devon how so like what's the delta there and so we collected this data set of essentially engineering hours estimates for their equivalent Devon session and by combining these things together we can now automatically apply for everything you run from Devon a score of was it productive or not and if it was productive how many hours did this save you and that's really interesting visibility for individuals for teams and for companies and when we ran the numbers we realized that as we were hoping people are saving a lot of time with Devon but they were saving so much time that it gave us the confidence to actually financially underwrite it and go to our customers and say we are actually going to make a $10 million productivity guarantee which is if you're paying for Devon and people are in fact just wasting usage and they're not getting value you know if you end up paying us more than the like engineering hours of value that you would expect in dollar terms we're just going to refund the difference up to $10 million and when we debated this internally I got a lot of pushback actually because some people were saying wait this is like a big risk we're signing up for a potential liability like what if someone just like makes an API call and they just run Devon in a loop and it totally burns you know millions of dollars and the reality is yeah like if you did that then we would we would be on the hook and we would lose money so let's go add some guard rails in the product to prevent you from being able to do that is like how do we align our incentives with the customers incentives so we now have really good cost controls we have the ability for admins to say like this is the overall budget I want to allocate help me be smart about like who should get you know more compute who's still learning and it needs kind of a tighter a tighter leash and I think it's it's all about like align the incentives with customers so that they're actually getting value for what they're using I was going to ask about that last thing you mentioned which basically like how do people use this info and do they do they use it at a team level do they use it at an individual level I think intuitively inside lane chain if you asked our VP of engineer I think it probably say that more senior engineers he trusts to use these things more and more junior engineers it's easy to mess up how you use these things it's easy to accidentally spend stuff but I'm curious yeah like how do people use these insights I think that my favorite way and one of the most popular ways of using these insights it's really for learning and like professional development because you can look at you know two teams you can say oh wow this team is actually you know really economically efficient with their use they're getting tons of a productive output for what they're putting in and this team is kind of blowing up their costs a little bit well it's not like they're maliciously intending to do that they're just maybe using the tools differently and maybe they don't realize that they're doing some things that are really inefficient so we've tried to then now that we have these dashboards and probably you can see oh yeah how is my usage compared to you know other folks and where do I sit and like what can I be doing better we're trying to push more of that coaching and that like professional development sort of into the product itself where you can like learn from that definitely definitely tell you Harrison that was a bad prompt you know like like you like this is like way under specified you know you should fix this like here's here's some tips on how to be better and so we're trying to put more of that training into the product so people can kind of continue to level up with the capabilities because it's actually really hard for anyone to keep up with how fast everything is moving and things that you didn't expect would be possible a few months ago or possible now or things that you're doing today as like a workaround for limitations actually you shouldn't do tomorrow so we've tried to move as much of it in practice possible and I think people are actually using the dashboards the most is like literally seeing okay team by team individual individual how how can I learn how to be more productive more efficient and just really master this new way of working. If you liked this episode, leave a review, and subscribe. Send feedback or questions to [email protected]. We want to hear from you.

Podcast Summary

Key Points:

  1. The bottleneck in AI development has shifted from training larger models to running efficient evaluations, particularly for assessing agent productivity.
  2. Cognition’s Devon agent uses proactive, context-aware automation to evaluate code sessions, enabling a $10 million productivity guarantee by measuring real-world impact.
  3. Specialized models like GP 5.5 and Fable 5 excel in different domains—Fable for high precision and GP for recall—demonstrating that no single model is universally optimal.

Summary:

The evolution of AI agents has moved beyond raw model size to focus on efficiency, productivity, and cost-effectiveness. At Cognition, the development of Devon as a cloud-based coding agent has advanced through a focus on real-world usability, beginning with niche applications like code refactoring and security vulnerability remediation. The team identified a critical gap in evaluating agent output—not just correctness, but whether changes would be mergeable and maintainable, leading to the creation of the "frontier code" evaluation benchmark.

This new standard measures both binary correctness and nuanced quality, such as code style and future-modification safety, through collaborative work with open-source experts and internal evaluations. A major shift in agent design involves routing tasks between high-performance and cost-efficient models—such as with Devon Fusion—achieving up to 35% better price-performance while maintaining quality. This approach reflects broader industry trends where specialized, domain-tuned models outperform general-purpose ones.

The rise of proactive agents, which operate autonomously in workflows, significantly increases human engineers' leverage, allowing them to act as CTOs of virtual armies of thousands of agents. As agent capabilities saturate across tasks, the focus has shifted from model choice to cost optimization and operational efficiency. Companies now face the challenge of token spending surpassing human salaries, driving demand for smarter routing and specialized AI models.

Cognition’s strategy—combining agent autonomy, specialized models, and customer-aligned ROI measurement—positions agents not just as tools, but as core partners in engineering productivity. The future of agent UX lies in proactive, context-aware workflows that integrate seamlessly into organizational processes, empowering individuals to achieve enterprise-scale outcomes without traditional years of experience. This democratizes software engineering, shifting the value of expertise from deep domain knowledge to mastering agent-human collaboration.

FAQs

Frontier code is a new, high-difficulty coding evaluation benchmark that measures whether code changes would actually be merged by a human maintainer. It was developed to capture subtle aspects of code quality like style, maintainability, and future compatibility that existing evaluations missed.

Devon agents automate routine tasks like refactoring, migration, and vulnerability remediation, enabling developers to focus on higher-value work. They can process large codebases efficiently and provide actionable suggestions that improve code quality and reduce manual effort.

Specialized models like GP 5.5 and Fable 5 outperform general models in specific domains—such as security vulnerability detection or data processing—due to superior recall and precision. For many tasks, cost, speed, and efficiency matter more than raw model capability.

Devon Fusion runs multiple models in parallel—using a frontier-quality model for accuracy and a cost-efficient one for speed—then combines results. This approach delivers up to 35% better price-performance with only a slight quality trade-off.

Proactive agents automatically detect issues, suggest fixes, and triage tasks in real time, shifting engineering work from reactive to autonomous. This allows human engineers to focus on strategic decisions, effectively giving them 'CTO-level' leverage over thousands of agents.

Companies assess ROI by measuring whether code sessions result in merged PRs (productivity) or not. A new evaluator agent automatically classifies sessions as productive or unproductive, helping quantify time savings and engineering output.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.