Go back

AI researchers debate how close we are to recursive self-improvement

97m 1s

AI researchers debate how close we are to recursive self-improvement

The absence of a radical AI-driven transformation in 2036 is not due to external shocks but rather technical limitations rooted in the difficulty of true generalization and persistent bottlenecks in AI's ability to reason, self-improve, or discover new paradigms. Despite massive scaling and advances in training methods like reinforcement learning, models still struggle with complex, non-stationary real-world tasks that require dynamic, context-sensitive reasoning—such as law, business, or long-horizon problem solving. These challenges stem from the difficulty of transferring knowledge across domains, the lack of efficient self-improvement loops, and the inherent gap between narrow benchmark performance and real-world adaptability. The field's progress is increasingly driven by domain-specific training and simulation-based environments, which, while effective, fail to capture the full complexity of human-like interaction and learning. Crucially, AI still lacks the ability to autonomously define or optimize its own objectives—making alignment and human oversight essential. Even with advanced simulation, real-world deployment data remains underutilized due to poor reward function design and low sample efficiency. However, early signs of progress—such as online learning in tools like Composer or Cursor—demonstrate that models are beginning to learn from real user interactions. The most plausible path forward involves a combination of simulated environments, human-in-the-loop feedback, and incremental improvements in model efficiency, rather than a sudden intelligence explosion. Ultimately, the key bottleneck may not be compute or training, but the fundamental difficulty of creating AI that can generalize beyond pre-defined, easily verifiable tasks into the messy, evolving reality of human work and innovation.

Transcription

20575 Words, 112617 Characters

English
Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk, and who will also happen to be at somewhat open-ish labs and companies, so you guys can actually say things on the record. I'm Jordan Byron-Millich, who is the CTO of Zyfra, which is developing open-source models. John Schulman, who is the Chief Scientist at Thinking Machines, previously the co-founder of OpenAI, led the RLHF work that led to CHGPT, and Charlie O'Neil, who is head of model training at base 10. The first question I have, we're in 2036. It's been 10 years. We don't have like, billions of crazy super-intelligences that are running around that are like radically transformed the world. What is the most likely reason that doesn't end up being the case? Other than sort of exogenous political shocks, or like there's a war, or they banned AI or something, but what is the most likely technical reason that we don't, like 2036 isn't like a crazy alien super-intelligence world. I mean, my reason would just be, it's got to be the sort of, there's been a classic thing, almost like Marvel X paradox, right? We think of it the AI, but if it can do this, it's going to be amazing. If it can solve these hot-mouth problems, if it can win a chess plumber, and then it solves these things, and then it's not that impactful, obviously, it's somewhat impactful, but not everything. If somehow that continues, and there's never the true spark of generalization that occurs, I think that could lead to like the AI's just being like extremely good at kind of everything that people like put into a bent rock, went into an environment, but like there's still some persistent, like, some to ill, which is somehow blocking everything. I think this is kind of unlikely. I think we do actually see this kind of generalization even from our island practice already, but like, if it is just like ridiculously hard to like generalize metal-earning, plus like we don't solve container-earning, it's just like super-hard and impossible. Like this would be my like default scenario in that case. Yeah, I agree with that. Humans have a lot of advantages over models now, and each time a new model comes out, it'll sort of, it'll catch up in some of these areas, but like you end up getting bottlenecked by the places where the model is weaker, and where it has worse judgment, or the models can't check themselves well enough. Yeah, so there's this cycle that keeps repeating where people think where a new model comes out, and people are blown away, and they're like, this is it, this is the, this is a GI, but then they use it a bit, and then it starts to feel dumb after a month or so. So that cycle just might keep going, and it's hard to predict how many times it's going to repeat. And like right now, you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering, that even if the model can write way more code than a person, it doesn't make you like a hundred times more productive. But yeah, so maybe, maybe they're just more of these cycles than we would, than we would expect. For me, it's like a question of how far off like this global optimum of a learner you could have on a chip is like the transformer plus like RL, basically like the current recipe. So like, I think people imagine that even once, like once you have a, an agent which is better than all humans at AI research, even if it's like 0.1% better than all humans, then the fact that you can run like, you know, hundreds of thousands, if not millions of these in parallel, you can run them much faster, like she was going to speed up, that's going to outweigh every other like bottleneck, and like your eventually just going to like hit this like very fast takeoff with recursive self-improvement. I could imagine that if we continue along the trajectory that we're currently on, with that paradigm where you know, it's basically just like self-attention RL, scaling up our own environments. The, I guess like, if you think about what happened with Moore's law, right, like we had this very like nice straight line and that held for a really, really long time, but there were so many like discrete like discontinuities and innovations that had to happen to keep that scaling law going, and the same thing has kind of happened with LMS, like we had this pre-training like scaling law, and then that was kind of like, you know, hitting the diminishing returns, and then we came up with like RL and solved that, and then we got this new like, you know, diminishing returns curve to hit that made it keep looking like a straight line going up, and so like if it requires another one of those discontinuities to solve, like I'm not sure that like the current method of like training LMS with these RL environments, even like RSI targeted RL environments, would be able to discover that discontinuity, and if not like, we're probably going to hit this like asymptotic like curve where like, sorry, but discontinuity will be harder than anything that's come since 2012. If we have the ads for that, that we kind of have the ability to implement it, but like maybe there's the distinguish, we should distinguish between the discontinuity which adds to the current paradigm. Again, it's like cumulative, like there's some thing beyond the RL that we have to discover, and maybe they're capable of like, you know, connecting the dots in that straight line, or like, but like again, how far off the global optimal are we? Do we have to go back and throw out like, you know, gradient descent and like neural nets in general? And I don't think like if you continue to scale up the current paradigm and LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away. Yeah, the only hope really is if deep learning just can't get us to an AI which is at least can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or like, I don't know, maybe humans would also never have discovered the the next learning architecture, but do the extent humans could have discovered it eventually. But it just seems like, I don't know, if you just look at the progress that's happens in 2012 till now, and you just continue that on, I mean, I know it's been powered by huge amounts of compute scaling and so forth, but it would be weird if like, it just didn't get to the point where it could like dominate humans at least in R&D, especially over the next few years, there's going to be Ryan Greenblub was on the podcast recently, and he made this point that you could imagine as the AI is getting more and more capable in our capable of making progress on simulations which incentivize getting better at not only AI R&D, but generally at science. So this is a thing that all the labs targeting, many startups are targeting, or another intuition pump is if you look at the ELO score of chess bots since the 80s, there's just like a very linear increase in ELO over time, but there's this huge discontinuity as they cross the human range of human experts always win against AI's to like human experts never win against AI's, as this linear increase in ELO happened, and you could think, I grew with your point that so far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world, but that just because like they're slowly rising in ELO relative to humans. Yeah, I mean, I agree it would be very, I mean, the only way for this to not happen is if like, as you said, somehow asymptote just like just before basically because we're already pretty close in my opinion to like where we'll start crossing like the human ELO score, and so we'll need to ask them to before that, and like that's the only way, you know, in this scenario you post, we're like somehow we're sitting here in 2035 and like everything is normal for this to happen, I think. I mean, the only other way is like there's like some dramatic like regulation on AI, is like this is kind of what I see as like the most likely way for this scenario to happen, actually, rather than the technical thing. Yeah, I think there's different kinds of research, there's like research where it's like the order research style where the objective is already specified very clearly, and you're optimizing that objective, and I think everyone is picturing like if we continue along this path of like, you know, making pre-training, let's go down, making our environment, let's go up, that's going to lead to like improvement, but like, yeah, maybe what Ryan is talking about is like this much more open-ended type of science, which is required for like paradigm shifts, where we can't specify the objective, and the AI is definitely not able to specify that objective either. Like, we have to be really, really careful about how we specify objectives, right? Any of these things. And my view point is that like the nature of the breakthroughs that happened since 2012 is that we have found like in 2012, people weren't saying, or I'm assuming, I don't know, you guys were there, or at least Johnny were there, but I was in promiscuous. Actually, John, I'm curious, or you're like sort of wisdom of the ages of yeah, or wisdom of like being in the trenches way back when, but presumably a big breakthrough was realizing that the next token prediction is the, like, you wouldn't have thought that NanoGPT speed run is the thing to be optimizing for in 2014. But now that we have come to this new paradigm, that's you wouldn't think to do a speed run on that and have you guys get really good at that. But maybe there's like a next inner loop to optimize that the AI's wouldn't anticipate. And there's an outer loop of like revenue or something that eventually should be strong, but it's a very slow outer loop. Yeah, in fact, I remember in the early open AI days having the intuition that actually just do like minimizing log loss wasn't going to get you to intelligence because like the important bits are accounting for such a small fraction of the loss that like it was going to be overwhelmed by noise. So just training a language model on next token prediction just wasn't going to learn the interesting things you wanted to learn. And we needed to craft better objectives that would put more emphasis on the important things. And like you can make all sorts of arguments for this and you could say, oh humans probably don't learn how to like, we don't learn how to model everything in our environment. We can't most like people can't create a photorealistic reproduction of some kind of scene they've looked at. So there must be, we must need a better objective but then it turned out that it just worked anyway. And as you're pointing out, the inner loop, even in current AI research of like post training benchmarks or whatever, it doesn't necessarily translate into what users like. Oh yeah, I mean, the whole field relies a lot on generalization and it's very hard to predict when you're going to get generalization or when you're going to get some kind of out of distribution generalization. So we know that if you train on the task you care about, you're going to do better. But like the most important advances are often or types of generalization that we have no right to expect. So, for example, from just pre-training on this very naive next token prediction objective to various tasks of interest where some very, never-quire understanding of the input in some deep way or learning some skill from pre-training that's like very rare and not very heavily represented, and then also generalization from these verifiable tasks to less verifiable ones. This is also a type of generalization that there's no reason out priori to expect it. So, this is an interesting question because one intuition pump that you could have or why you would see some sort of singularity very rapidly without even scaling up the inputs to your progress that are not just AI labor is that before every single experiment you run, that's like a seven-figure experiment. You spend an equivalent amount of compute on AI labor, and so you just have automated versions of you guys spending a century thinking about what is the optimal experiment to run, do you like small-scale oblations, developing literally like a century's worth of theory so going back even before deep learning. And before you decide what experiment to run, doing extremely optimal setting up of the experiment, then you do a century of thinking after the experiment is over where you're like analyzing what happened and what the next experiment to run is. Well, I think if you think hard enough, you probably could have expected some of these things beforehand. Like there is probably some very clever way to do a small-scale experiment that will let you build the theory that then will generalize to the large-scale experiment. So I would expect that like we're nowhere near the ceiling of how well you can do research. And I would imagine a future where AI is doing a lot of analysis and theory building, spending a comparable amount of compute to the amount that you're spending on the experiments themselves doing various kinds of analysis and building a theory around what we've seen so far. I think there's really concrete examples of this when the objective is well specified. So again, all thinking can do is update your posterior based on the bits that you've gotten since you formed your prior. You can't gain any new bits from just thinking. But when the objective is well specified and there is this data sitting around, I imagine there will be this big speed up in the current paradigm we're in. And good example, this is like, you know, if you've got an AI to think about the Kaplan scaling walls, an AI at this point would have noticed that they've just taken these intermediate checkpoints and didn't account for like the annealing. And so like, this is wrong. And like, that would have caught that. Like, years earlier, we would have made like, you know, progress like would have cut off a euro to a progress just from that like observation from an AI. And like, again, once the objective is well specified which is like lower pre-training loss or whatever. Like there's many, many good examples where if you just thought about it a bit more, you would have been able to like cut down a significant on things that you've done to like muP and like how learning rate scales with like model size and like realizing the model width is important in that as well. Like I feel like you can, you can really back out a lot of these things and cut off like a lot of low hang fruits. I would imagine like a 10 time speed up if our thing is just like maximize the objective we're currently on. But I don't see that. Oh, that generalizes at all to, you know, if you can't with the right objective in the first place. So like just thinking doesn't necessarily buy either right objective in the first place. I mean, yeah, I think this is really the key question to like any kind of like, very rapid, our size like from current AIs is like how well can AIs generalize like learning their own objectives? Because have any kind of like self-propelling automated loop you need the AI to like propose objectives, optimize them, figure that out, propose a new objective. And like have this like not go off the rails at like any point for like a long long time. To come back to more of X-products, there might be like a case of more of X-products where like we think this kind of like autonomy and sort of like being like self encapsulated so we can you know think of what we should do ourselves and then go do it and like have this loop is like super easy because we always do this. And like obviously evolution needs to create creatures that can like survive on them by themselves like long period of time. And like this just might be something that for some reason it's like really hard for the AI in the same way that like no commotion stuff is really hard. But it's like math is super easy despite being super hard fast. And doesn't the timer as an increasing suggested that? Yeah, exactly. I mean, this is another possibility which like but I agree like there's no obviously evidence for this like in fact the fact that you know our agents now like super persistent and it's crazy to do this. It's kind of evidence against this. Yeah. Like this would be potentially like one of the reasons why like we just don't get this like immediate takeoff is like if this is hard. If you look back from 2012 till now or maybe from when you started doing your research till now what part of all the innovations that have happened since that time including purely engineering ones including purely conceptual ones what seems like the thing that is the thing that would be the last things humans would have to do before the AI is totally automatically earned. Probably just like it's just like asking the right questions. Like if you can get the AI to like do any experiment but like you need to decide what experiments to do and like well now I think AI is not very good at this compared to coding experiments at all. Yeah. Like whenever we talk about research they propose like a bunch of like miscellaneous things which are like very very tiny steps. Or even going from like you know deep minds approach like we're going to solve intelligence by learning to play games to the super human level. That's going to be the approach to like one random research like Radford being like I'm going to try and just predict the next token of a very widespread flip data. And then even once Radford to discover that right like it took a while before people decided to scale it up because we had to come with the idea of scaling rules and the fact that like you could very reliably predict these things. Yeah. I would say that the last job for humans or the role for humans that will last the longest is like defining the objective and like deciding what we actually want. Yeah. So like in that vein something like deciding how the assistance should behave or what it means to be helpful or what's like the objective when we're doing are all from human feedback is one such thing. And then like then later like defining like constitutions and model specs is another one. And I think even if the AI's can do all the technical work we'll have to still do a lot of that and decide what we actually want. Yeah. Alignment is the final job. Yeah. Alignment is sort of the answer but it's also alignment itself can be kind of decomposed into like specification of the objective or figuring out what the right objective should be. And then like actually like achieving or optimizing the objective you've defined. And I think the first one is not going to go away any time soon. And like if I think about like a post trading team and why you need a lot of people to be on the team. It's just because there are a lot of different like areas where you have to figure out like how the model should behave. And like there's no way of like there's no way of well it would be very hard to automate the whole thing just because someone has to think about how should the model behave in this area. Jane Street started using Antithesis to test her software in early 2025. And they were so impressed by the product that they decided to invest in the company. I recently caught up with Ron Minsky who co-leads Jane Street's tech group to ask about how Antithesis actually plugs in. The thing that I think is most impressive about Antithesis is we started using it in the team that was building higher insurance software and being really careful. And nonetheless it was able to shake out bugs that were otherwise going to be really hard to find. And that's important both because it helps make those systems more reliable but also because it helps the teams that build it to just move faster. This matters more and more as code production is increasingly automated. I think in general as we've been using agents more and more, the key problem that you run into is the verification bottleneck. Just the time it takes from people to look at code and figure out is that actually something you want to accept in your production software? And tools that make testing better are just incredibly helpful there. They just ease the verification bottleneck and make it possible for you to get more stuff done and move faster because you can have more confidence that the code generated by the agent is actually not introducing new problems. To see how Antithesis fits into your development process, go to antithesis.com/thorkech What is the story for why there isn't huge consolidation in model providers? There's just so many things that point to centralization here. If you step back over the course of years, is there something that is going to prevent that? I think distillation is the main thing that fights against the centralizing force because basically anything that can be learned through RRL can be distilled very easily because it's a small number of bits. It's something that you can learn from a small amount of data. If you can get trajectories from the model that show a behavior, you can easily distill it. I think distillation is one of the things that fight centralization. There is also a possibility that there will be companies specific models. It will be possible to learn from deployment and have a company continually improving its own model. Such a system could be provided by the current oh, gothfully of model providers or some other currently smaller company. But I think that will change the game a bit. I also want to point out that continual learning doesn't stop distillation. Even if your model is improving every day, could be distilling it every day, so it's like, the loop's good just operated at the same pace. - Right, that makes sense. - Okay, so copying model behavior. You, I guess you need to know yourself what the right distribution to prompt is in order to get like the relevant model behavior. - Yeah, for just distilling with supervised learning, the prompt distribution is extremely important, so it's very non-trivial to distill a model, even if you have full access to it and have the cut, the chain of thought and everything. Yeah, it's non-trivial to distill all of the useful capabilities from it, because you need to prompt the model with something, and you need to prompt it with like, realistic prompts, you need to have a really wide distribution of realistic prompts. So yeah, one thing that's been coming out recently is some of the, some of the Chinese companies are probably using these router services, which are designed to allow people in China to use the US Frontier models, which would otherwise be blocked in China, but they're all these router or proxy services that allow people in China to use these models mostly for coding, and these router services are collecting and selling some of the data. So I think this is like a very useful data set for distillation, because it gives you the perfect prompt distribution. - I think this is one of those things where AI has helped a lot here. Like, if you actually look at like, you know, the Frontier pipelines, they say like the Chinese models that they actually put in their papers, it's a lot of like, humans or like, they get seed prompts from somewhere, which is some combination of humans, this kind of data, and then they like synthesize a vast coverage from those seed prompts, using their existing models or like the other Frontier models. And so it's like, you can automate like an awful lot of this like, prompt distribution gathering and like, environment creation. It's just like, humans need to provide like, increasingly fuel amounts of bits, it's like the models get better. - Right, oh, everybody, it still seems you're bottlenecked by like having a service which has users or users are going through. So like, not necessarily, I mean like, yeah, that's obviously very helpful. But like, theoretically, you can just think about like what users want to like. - No, but it's a lot of tasks. - The whole point is that we don't, the user says, make me an application like this. Oh, that didn't work. I actually want you to make this new feature. But actually, let's step back and do this other thing in capturing that whole trace is the, or to the extent you could have done that anyways, then you just have like RSI anyway. - Yeah, I mean, like ultimately, like, if you have this like fully automated loop, that is basically RSI, right? Like the AI is deciding, the data is deciding, the training that that is the loop. But yeah, I mean, like, it depends how much human information you need. Like, at some point, if you're just like, I want traces that look like this, you prompt that to the model, the model will be able to like, come up with like, a pretty good approximation. - But what if you want to do like, make me really good politician, and then that's like, anticipate Denovo? - I was like, yeah. - How would a discussion like the Senate halls go, or something? - Yeah, I mean, like, there's a lot of things which I-- - Well, ironically, this is actually, I think, easier for the distillers than the Frontier Labs, right? 'Cause the distillers just like, I want to put good politician, they go to the like Frontier model. The Frontier model already knows how to be a politician, so it just like generates those traces. Whereas like, if you actually want to build the first model that does this, you have to like, actually somehow like get data on like what politicians do every day and like build that. So it's actually much easier to like, say like, I want something like this, and then like get like, the idea to produce like a billion variations, then to like actually create the thing like this to begin with. - I think you can actually make a really concrete prediction based off like this observation that the Chinese types have this route of data. So like, I think the thing that just did this originally was I was saying, isn't it weird how Sonnet 5 and Opus 5 are like, like almost objectively worse models than like GLM 5.3, Kimi K3, even though they've had access to like not only distillation, but logic distillation from like mythos. And so the counter here was that like, okay, the Frontier distribution really, really matters. Like, you need to see what users are doing so that you can distill like kind of these behaviors and things in. I think the prediction from this is that the Frontier labs don't necessarily have much of an advantage if at all in our environments now. Because yes, like user distribution matters for like general like behavior and so on, but like the best measure of a capability is the very, very hard our environments you've made at the Frontier. And so if you have access to those our environments as anthropic and you have access to logic distillation and you've still made a worse model, then maybe like the real world of point matters more than the environment. That's really interesting. So, but they had to incentivize those capabilities in the first place in Fable or the Frontier model. And so it's weird that they can't incentivize them again or like with a smaller model or something. Maybe like maybe we're just in this weird, like uncanny value where you know, like actually trying to copy that Frontier model too much. Like the student-teacher gap, whatever it is, is just like too large. And like I think maybe people made this point with Opus is it's like, the difference between Opus 4.6 and Opus 5 is that Opus 5 really feels like it's got this like AI as a judge checking every possible thing it's done. And it's like, that's why it uses so many tokens that like tries to think about all these things. But it doesn't necessarily have the big model smell of Fable to know when to like stop doing that or like when's a good path to go down, or whatever. - The reach exceeds the grass, yeah. - Yeah, it offers a slightly different hypothesis. So I would say there are a couple of different axes for the environments you can create. And like one of them is difficulty and the other is realism. It's sort of easy to create, or it's comparatively easy to create a lot of difficult environments like that are just like involved like doing a much more complicated task or doing something that requires a lot more cleverness. And you could say this is like the bench maxing distribution 'cause a lot of the most prominent bench marks just involve doing some very hard puzzle like task that's easy to verify. And then there's sort of like the realism axis where you want the model to be good in the realistic coding agent setting where there's like multiple back and forth through the human and there's like multiple objectives. And like I'd say like the people, like the labs who are crafting the model behavior for the first time need to push in both directions and to get good model behavior. You need to really push on the realism axis and have like rubrics or some kind of human feedback that's informing the reward function you use there. But I think when if you try to do distillation naively you end up just sort of matching the teacher on the bench maxing distribution. And but if you don't have enough of the environments that really exercise the capabilities in these like trickier realistic settings then you're not gonna get those into your student model. And I think maybe one thing that's happening is the big models generalize better from the like the tricky narrow tasks to these sort of more realistic tasks. So if you have a really good like realistic prop distribution for distillation you can match the big model really well. But if you only have this like this distribution of easily verifiable tasks then you can match the big model on all the benchmarks. But you do work on this broader distribution. So you might that might even explain something about the smaller anthropic models like Sonic 5 though it's hard to predict exactly what they're doing to post-training those models. It could also be that they're always changing their post-training stack and they just made they just got a few things wrong in some of these models. So they like I don't know they turned up like something too high and created some quirks that people really don't like. So it's like really easy to screw up post-training in some way that doesn't show up in benchmarks. - I mean we're just one of the sort of very basic point is just like the front AI labs pile their data from their data companies. And like the Chinese can also just buy the same data from data companies. - And they are. - And like they are exactly there's a lot of people like you know being annoyed about this but like if they have exactly the same data and like they can buy that they can also distill. It's like it means it's quite easy to like keep up really. - Yeah, yeah, yeah. Okay the other question I had is how the first models that are capable of automating AI R&D will actually be trained? 'Cause there's a toy version which is this thing that Ryan was talking about which is you just have GPT-8 try to build GPT-3 size models that are really good at like inner loop type challenges of beating video games that require continual learning or just getting to a certain loss with like the least amount of compute, et cetera. But John I think you had an interesting point that maybe that's not the way it actually will happen in practice. So I'd be curious about, yeah, by the point which you have as it are actually capable of automating AI R&D how are they probably trained? - Yeah I think we will probably do some combination of learning from human feedback to absorb like the researchers' taste and just like creating a lot of practice environments which involve like doing multi-step research projects. So I think people will in practice to some combination of those two things and just each iteration like patch whatever seems to be most broken in the last iteration. So like researchers will be using the AI's a lot and we'll notice that they have some consistent weaknesses and then those things will either be patched by collecting human feedback or like creating environments. - Yeah, maybe you saw what I think about this is like how much of the lineage we roll back and then let self play from there. Like I think in the limit like you're picturing like you know just giving them like a GPU and maybe neural nets or something and saying like okay figure out how to train a model to like do this particular task. Like the way it currently works is like we go up to the very like edge of the lineage and say okay like here are the bugs like you know anthropocas found in their training stack in the last few months will tell us into environments like you need to train to get better on the frontier. And so you obviously lost. in all the previous history of the language, but you could imagine a world in which you roll back to like, you know, before G-R-P-O or something, and then you have environments which like trying to get it to discover like the best will form to like R-L models on and then maybe you roll further and further back, but I think we will be still so compute bottleneck that like people will just keep like staying at the frontier and like diffing essentially the bugs and whatever improvements they found since the last model version turning those into training environments. It was also really good for having non-steel like new data between the model generations. It's just, again, this is basically a continual learning within the AI lab of distilling the last three months of AI research progress through environments and like R-L-JF's type stuff back into the model itself. And it is distilling right and that's maybe why some of us feel like it's asymptotic is like you're always like just trying to get the last three months of progress and that progress is being contributed to by as of course, but it also still has humans in the loop and it feels like, you know, you're just constantly inching closer and closer to what the human researchers are like finding capable of doing. I mean, the one thing I will say that was like obviously if you're just distilling on like trajectories, you can never go, but environments can go quite a far way above what a human can do. It's very easy to design environment that like no human can solve, but the AI can always be still trying to solve it. And so that would be the path to like go ahead of just like what the human AI research is. Do you have like an example of like, in terms of RSI or like, like, you know, kind of treat a certain, but doing it even faster than a human speed runner? Yeah. I mean, I feel like in AI research, especially is very easy to define like goals, which like, you know, you could say like the loss needs to be like one point three or something. And like no human can get that, you know, now, but like that's a very extreme measurable via Bible task. And if the AI gets that, then great. Right. And I'm building like a hundred million parameter model that beats Minecraft. That's maybe too easy. But like be thick and much more complicated game or something. Isn't it crazy that a hundred million parameter models will be at Minecraft. We call it that too easy. Like I mentioned, so that like five years ago, I would say a lot of research is not exactly like that, though, where it's like hill climbing on a well-defined goal. It's sort of more like, here's an intuition we have about some way models should be better. And then we also have some idea for an algorithm that seems to go a little bit in this direction. So let's come up with a task that is sort of designed to show signs of life on this approach. And like, see if we get some, get those signs of life. And then if we do, we can make successively more realistic versions of the task. Right. It's like a lot more guided by intuition. And then the inner loop is to elicit the, or make tests for that intuition, rather than like the test itself leading to the insight. Right. Like you're not directly optimizing for the eventual objective you care about or the practical like production objective, it's sort of you're, you're relaxing your objective a little bit. You're saying, yeah, let's relax on the realism axis a little bit and find some methods that actually work. And then like, then try to get back to realism later after the method matures a little bit. And then there's also like more, there's research that's more oriented towards explaining things and like developing a theory or sort of, yeah, often we don't have like mathematical theories in machine learning that are that predictive. But we have like a lot of like more informal theories for what's going on. Yeah. I mean, like presumably the models will be trained on like some combination of all of these tasks and like some will be very easily verifiable. Some will be like LMS judge or like just ask the human like does this look reasonable. And then you will set the hope would be that like these would all generalize to like these much sort of hard sort of more vague fuzzy kind of tasks. And like it probably will doesn't extend whether it generalize enough that like we could the loop can become like self-sealing without humans being in the loop at all. It's like unclear. Yeah. Yeah. Maybe taking a step back. Here's what I, here's what it seems to me that the plan for AI research going forward is. And you tell me if you think it's going to work or if you agree with this characterization. So the bet is that we will scale up RL VR training across millions of diverse environments, across hundreds of different kinds of domains. And what will emerge at the other end is an agent which has like learned these basic skills or less than basic skills around being persistent, being able to triage information in context, eventually having like end to end optimization of working with other agents and things like that. And such an agent will be very sample efficient within the context. You know, you've done research on how you actually scale up in context learning to make it like arbitrarily long, but you keep scaling it up. And so what comes out the other end will something will be something that it basically functions like a drop in remote worker or over the course of a week or a month. First of all, do you agree that is the bet the labs are making? And second, is it, is that enough? Like basically learning how to learn within the simulac or within a data center. And then getting deployed into the real world. But not actually like learning from real world deployment, only learning these meta skills from the simulated environments in the data center. Yeah, I think it's now hard to separate out like how much of the labs effort is going towards like direct RSI versus like making generally intelligent models that they can continue to deploy, to collect revenue to fund the next big training run. Yeah. I think for the latter, like yes, that's probably just the bet they're making. Like, and it's very clear like the pattern of like where these environments are going over the last few years. I mean, like anthropics lineage of environments is like a very clear example of this. Like, you know, first of all, we just focus on coding and like we're going to get really, really good at that. And then the task horizon that we've got from coding, which is probably the lowest hanging fruit in terms of like data available on the internet to create environments like their own internal stuff that they can turn into environments. Then we're going to generalize we're going to go after finance next and like literally like just so much Excel data and all that sort of stuff in there in the other training. And then, you know, it's PowerPoints, it's like this long tail of like the working economy and like that seemed to work really well and like a lot of the other labs and thing even the open source labs have now realized that that was the correct bet. But what is the implication from that when I had Dario on the podcast, the thing I asked him was, if you truly expect models, which will be human like in their ability to learn on the job, why would you try to beacon all these skills of like working with PowerPoint or something? Wouldn't you just expect the model to be able to pick that up on wireless deployed? And so yeah, there's multiple different explanations. One is just that this is, we expect models to get there soon, but they're not there yet. So why not amortize these skills into the model training? Another is that we're not concentrated on making it really good at widely deployed work. We just wanted really good at RSI and this is just like a way for us to like get revenue so that we can port back into a model that is actually like really good at doing RSI development and then like what's the singularity happens? The thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, I don't know if you have takes on like what, how much it can strew why there is so much task-specific knowledge in these models, if the path is like this kind of generalization? Yeah, I mean, if the models were good enough at learning in context and in theory, you wouldn't be, you wouldn't need to train them on finance, they would just be able to figure out, read all the books on the fly and figure out how to, how to do everything in the appropriate jurisdiction. Yeah, and you could argue that you need to do a lot of this domain-specific training just to make the more efficient. So even if they were smart enough to figure this out on the fly, you still might want to do a bunch of RL and bake all these intuitions into the weights so the model would be more efficient at runtime. Yeah. Yeah, I'd say in practice, it does seem like model providers are going domain by domain and trying to strengthen the models in the highest value domain, and I'd say that that's one of the answers to why the models have gotten so much better. It's just because the model providers have covered a lot of the high value domains and the most common types of skills. I mean, I think another thing is just like, it's not that expensive to do both at the same time, right? Because like the models are massive, they can easily afford in terms of their parameters to like learn everything. Yeah. And like there is likely some transfer and sort of even just, even if like finance is not specifically like the information is important for like RSI, just the general like meta-learning of like how to figure out what's important, how to have taste, how to like do long horizon work is potentially generalize but I'm like, there's not that much RSI like data in the world as well. Like it's kind of hard to generate and like that requires a lot of effort. So like if you can advertise in this other data, get some transfer form it, you already have massive computing, massive parameters based on like why not do that as well as like obviously the direct like commercial intent of like selling a model. That makes sense. Yeah, I'll add that, I mean there's one question about whether this current paradigm of doing like seem to real will be the dominant one forever. So basically you, you look at what the real world tasks are like and then you try to create a bunch of environments that can be simulated like in the data center and you can do RL on them. And I think obviously this has been very successful, successful but it has a lot of weaknesses because a lot of things are just kind of hard to simulate, especially if they involve like interacting with a bunch of humans in real time. Yeah, so there's some question about like, like whether seem to real will be the dominant framework forever. And I think some tool has to be the donor framework, well like sample efficiency is kind of low because like right now you need like you know thousands of thousands of interactions with the humans. No human is going to sit there and like deal with this basically being the loop of RL training. And so like we kind of have to simulate that now to like get the samples you need. But like obviously if sample efficiency improves a lot, you'd expect learning from deployment to like become like a much bigger part of it. But there are also other things you could do like you can learn off policy so you can not take all the traces and even without resimulating everything you can potentially learn something from them. Jane Street just launched a new competition, and it's their most ambitious one yet. Design a protocol emulator ASIC. Basically, if you have a chip that you want to test, you can connect it to this ASIC, and then this ASIC will simulate realistic traffic. That way, you can see how the chip responds without having to plug it into a live system. Jane Street is looking for flexible, general-purpose designs, not single protocol emulators. When I was chatting with them, they suggested that I start off by trying to implement what are apparently three very common protocols, UART, SBI, and I2C. Jane Street also mentioned that they hoped that more ambitious designs will also tackle low-speed USB and Ethernet, and any other protocols that flex your chip's specific architecture. Importantly, your design should be reprogrammable, rather than smashing a bunch of specific protocols onto a chip. If a new protocol comes out after your ASIC is taped out, your chip still needs to be able to handle it. How exactly does it is up to you? But there is one hard constraint. Your design must target an open source 130 in the animator process node. That's because Jane Street will pay to tape out the most novel submissions and send the physical copies to the winners. The competition is open till January 18th, 2027, and working in teams is highly encouraged. Go to JaneStreet.com/doorcache to download the template code and get started. I want to ask more about this, because it's sort of weird that you have 50% of compute that's spent on inference that is not directly helping the model become better. One of the key advantages you'd expect eventually digital minds to have is, unlike a human who gets to have 50 years of real world experience, a model will get to-- through all its instances, we'll get to experience millions of years of deployment across all kinds of economically relevant work in the economy. And right now, that data is just not in a meaningful sense helping the model get better. It just seems so obvious that eventually, model should be able to learn from this data. And once they do, you would have something that almost feels like a widely deployed intelligence explosion because the model is assimilating so much information across all these deployed instances. But when do you expect this kind of high-mind kind of crazy shit to be start happening? I think broadly, at a very basic level, this is already happening, just in the next generation of models. So right now, you can always take your deployment data and put this in the pre-train or the mid-train of future models, especially if you do like some kind of filtering or some kind of judgment or annotation or recent synthesization of that. How much do you think that explains the generation or the generation improvement? I think it explains quite a bit. I mean, especially-- I mean, this is, I don't know whether the labs do this because theoretically, they claim not to train on people's data. But the Chinese are 100% due. And they definitely get this advantage, both obviously deploying-- this is basically what distillation is. They take out the models. They get some of their deployment data. They get some fraction of that by pinging the model. And then they train their next generation of models on it. And they can suddenly do it on their own models as well. There's no reason not to whatsoever. I completely agree with you. So I think if you zoom out far enough, this is definitely happening. You're picturing this-- and we're all picturing this. This is what continual learning-- the Holy Grail is-- is this very, very organic live loop of an individual model, getting an experience and live updating on the spot and learning from that. And a lot of things break when you zoom into that level of granularity. But the big labs are doing this, the close models are doing this. There's also early signs of life of people using open-source models doing this in a much faster cadence. So a good example is probably like composer. You have some sort of model, and you are able to-- or like Harvey's doing the same thing with legal agents. It is getting very specific environments from the data that you have for that particular task and things that users are complaining about. And all the feedback that you're somehow extracting from your specific deployments and a lot of these companies have the advantage over the big labs and that they can use this data really, really well. And then they will create environments. They will do a big post-trained of Kimi K3. They will go deploy it. They might do some online learning as well. Composer did online, basically, reinforce for a long time. So yeah, there's still a human in the loop. There's still a human saying, OK, this is the signals we care about. Here's how we're going to create environments from the data that we have. And there's still a longer cadence than maybe the one that you're thinking of. But it really is happening. And eventually, that loop will become faster and faster. I mean, the composer thing is interesting because this is where the model, like in cursor, people press tab or they don't press tab on the next completion of the model suggests. And based on that, every single day, composer gets better at predicting the next. So that was the old tab model. They actually did the same thing for the actual, not just the tab model, but the actual generative model. It's interesting. And it was-- it's hard because when you do online reinforcement learning, you don't have groups, right? You just have one user saying one thing. And then you get one roll out. And so you have a big variance reduction problem. Yeah. And like, cursors kind of fuzzy answer to this was like, oh, you know, we have very good heuristics, which are able to estimate, like, how much better than average, like this response was or how much worse than average this response was. And then they would do like this big reinforced update. And then their solution to like whether I got worse or not was like, if it improved on cursor bench, they would deploy the new model like every 5Ls. And if it didn't, they would like throw that version out. Interesting. Yeah, I think your biggest problem is actually just not knowing what the reward function should be from natural data. And if you use some kind of superficial signal, like did they accept the code, the edit, you might, that might get reward hacked in some way. But it seems like a bigger issue with the SIM2real thing, where the longer and longer horizon tasks get, the harder they are to simulate within the data center, right? It seems to me already, potentially, even in coding, we're getting to it with a point where there's not some year long coding task that doesn't eventually require you to talk to a client or interact with the company or interact with the users. And if you think about the gamut of things, we would want AI to be capable of that. You want, eventually, super intelligent should be able to run a business, or start a new business and make it profitable, or have a profitable day trading in the market, or win a court case. And these are all things which are very hard to simulate in a data center. An inherent part of the learning there is interacting with the real world. And so maybe they made a learn how to get better at these things from the transfer between SIM2real. But alternatively, maybe you do need weight updates from these kinds of interactions in order to get better at them. And then if that is the case, if the transfer isn't strong enough and you do need weight updates, then the fact that the models are quite simple and efficient is maybe a deeper problem. And the reason I'm curious about this, I feel like by default, I don't see how you don't get some kind of crazy recursive self-improvement within the next 10 years. But the one reason why that might not happen is in terms of weight updates, the sample efficiency of weight updates, they just seem way far behind humans. Like plausibly million fold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from cold start to finishing training. And so this is all to say, first of all, is there going to be a good transfer between simulations and extremely long horizon, really complicated real shit that we want the AIS to do in the role world? And if not, does that really mean that the lack of sample efficiency in these models comes to bite us? I think maybe the way I'd break down the two types of tasks in which models get good and models will still continue to struggle is whether the task is cumulative, or you have this non-stationary distribution you have to keep learning and relegating a bunch of stuff. So maybe an example of a cumulative task might be RSI. It's theoretically possible to maybe have less than a million token, like Python file, which like from scratch, trains a model that is capable of recursive self-improvement. And like every discovery that you make is kind of a line and send that you hold. If it's true that for RSI, we don't need to discover a new attention variant or whatever. Once you've discovered attention, and then once you've discovered mixture of experts, and once you've discovered your IPO that you just add that to the training stack, and that's there. And like a good example of this is 5.6-all training, 5.6-terror, whichever one, opening, I told you to train. It didn't have to go back and discover attention. It basically probably would have called a bunch of scripts, which is pre-training.sation, post-training.sation, just did that. So that's an example of a cumulative task. I think the real world and the reason people are thinking so much about continuing learning is it's not really a cumulative task. Imagine in a law firm, you have an agent acting as a legal associate. That's a very non-stationary distribution. You have to be able to fit in your context, all the relationships between all the important people like company, which are also changing all the time. You have all these implicit ways about how things are done, where to find information, et cetera, and that's not as clean of an example of a cumulative task like RSI is. So I think that there will be this breakdown between tiles. But if the labs realize that, and they do believe that RSI is cumulative in the sense that we don't need to go back and discover some brand new architecture or whatever, then maybe more and more effort and compute gets focused on that versus the-- It's so unfortunate that RSI happened to be easier than that early call. Yeah, I don't know if you guys have thought on this. Yeah, I would say there's models-- today's models are weaker than humans in a lot of different ways. And some of them might have to do with sample efficiency in a certain regime, where I mean, in some regimes, models are very sample efficient, like learning in context. But then there might be some medium length regime where they're less sample efficient, because humans can do some kind of weight update more efficiently than models. So I think like being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that, for example, having lower diversity of thought than humans, or being bad at certain kinds of long horizon judgments. I mean, I think a lot of what people call taste is something about. behavior that works in the long run, and that people have realized it works in the long run, not everything, but like some aspect of tastes, like especially for something like software engineering, like I think a lot of taste is like, what are the systems that are gonna be maintainable and work well in the long run of this project? So yeah, I think the weaknesses of humans, which limit RSI along with other things, are there's a variety of them, and some of them are related to sample efficiency and some of them aren't. - Maybe an interesting thought experiment is like, if you were able to give a model, like a context window of, I don't know, a trillion tokens, or whatever you would have needed to fit in, like your experience prior to like, let's say RLHF, and like it's got all that experience in the context window and it has the same sample efficiency and in context learning ability as it does at a million tokens. Like do you think taste is then solved? Like would it be able to like make the same judgments that you did, or is there like something fundamentally missing apart from just a longer context window with the same sample efficiency? - Yeah, I mean, it would have to be trained to learn from that context. So I'm not sure, yeah, either it would have to be trained to learn the right update to make from that context, or we'd have to generalize. - So you don't have to think you could just like dump it all in, like your whole like life, like research experience. - I mean, like you still need the data to train at long context, right? Like even if you could theoretically get like a trillion context, you would need a trillion lengths of data to train it. Like right now, you have like to take context and it sounds dumb to something if you had that. - Yeah, very, I think, yes. I mean, this really just comes down to the question of like how meta-learnable is taste from like shorter horizon episodes? And like I feel like there's no obvious reason it's super long 'cause like humans somehow developed taste with not having many long episodes. Like we don't live to be like 10,000. So we have like, you know, we develop pretty quickly, right? And so like, you know, if you think about like, even like in a PhD, the difference between like, a first-day PhD student and like a final like postdoc or something, that's like five years maybe. And they've only done like maybe like 10, 530 research projects and total but somehow they developed taste quite quickly from like a relatively short succession of like small things. And so like theoretically it's possible to develop it like that. The AI obviously will have vastly more experience in which to develop taste to like meta-learn it. And then it's like how well does that generalize to like really long horizon things? As I think the question, which I think is really unsolved at this point, like we don't know. - Going back to this question, eventually there should be a regime where AI's were learning a ton from each individual instance of deployment that they have. Well, currently you could say there's a meta-fuzzy process by wish models doing for deployment but I feel like it's a very weak, very weak feedback loop. Do you see this around the horizon where there's just like high-mind kind of learning that's very rapid? And if so, how exactly does it happen? - Actually, I would say that around will we get a high-mind that learns from all of its deployment experience? I mean a big part of that is actually about incentives rather than being a technical question. So like companies aren't gonna wanna have the model provider learn from all of their deployment because that might just reduce the advantage of their business. I think that maybe the economics of this will pressure not necessarily wait updates to one big, like common shared model but like kind of like modules they get sub-dened. So a very obvious example, this is a law, but it might be something else like, you know, there's been a lot of work to try and fit like an arbitrary context length into a fixed size. Like there's all the linear attention stuff and all that sort of stuff and like cartridges which are essentially KVK's just trained to be very, very compressed KVK's just to fit in a lot of information. That's another example of like, you know, something that like companies may be willing to sign up for if that's get subbed into the model and it's not like actually changing the based on the line model itself. So like there's many different versions of like learning from your data in real time and like the latter ones are not really helping the big labs 'cause they are just these modules. But I think the like the economic pressure will force like the labs to go down that path first before they can embark on this, like, you know. - Which economic pressure there? 'Cause I feel like even if you have like a bunch of cartridges or dollars or whatnot, you can still just like take all these traces and just like to still just dump this to the pre-training of like your next generation of models. - Yes. So it may be a more indirect form of learning that the big labs are getting and that's obviously still really valuable to them. But I can't imagine a world of which we start off with, like, you know, we're going to just like directly train this one big model on like all the exact dollar. - No, I think it will definitely like go through stages 'cause I mean, this is assuming it's just like one discontinuous event where it's like suddenly we fix like wait updates continuously and like in practice I think it's meant to be like, the cartridges and stuff allow you specialize in deployment. Then you generate traces, you put that in your model. Like three months later you come out with a model which is better with this stuff. You specialize to get any like household data to get. And then eventually we'll just like make this leap faster and faster. So I said like every three months we release a model. Now it's like every week and then every like day and then every hour. And at which point we basically have always dissolved it. - Yeah, and I think this is a good point as well because you asked like kind of how far off the current paradigm we are from being able to do this. I think like we've done a bit of research to this and people have done a lot of research. Like at a really large scale like when you wash out enough noise and you have large enough patches like this outer loop process of like putting data into mid-training and creating our own environments. Like it does work in like some sort of continuous learning regime. But the problem is like when you zoom in close enough at like a micro level, it's like, I've got one model and I'm trying to update it again for like a law firm or something and I'm trying to do that very continuously like with a relatively small amount of data. Like all the methods kind of break down a bit. So like if I SMT the model on just like you know successful traces off policy on policy, like eventually in the very iterative regime, like when you're doing like you know hundreds of these micro updates, you see both catastrophic forgetting, you see forgetting of like previous information I've learned on the top of the base model that was much earlier on and I see degradation of general like use, general capabilities. You know on policy distillation seems to like push this horizon out a little bit but it's still eventually succumb to the same thing and RL is not very good at like, it is good at like getting capabilities in but it's not as good as getting like knowledge in and like just this very explicit knowledge of like, okay, like this person does this law firm and like this is a very specific process we find and you have to pour in a lot of compute to create the right environments to get the knowledge in itself around. - Do you think that the fundamental issue here, why you get worse at any of these other skills or there's forgetting and stuff? Do you think it's fundamentally an issue of capacity or it's an issue of techniques? - A little bit of both I think like SMT and even like distillate like on policy distillation can be like way too destructive. Like the reason RL is so nice is because like, yeah, it changes a very, very small amount of about the model and there's like a lot of evidence for why this is the case. And so like it kind of just like tweaks it in this very, very, very small like loss value to like get it into the right point, but that also then limits what you can do with RL, like how much you can actually change the model. - Worcestershire, you're saying like the reason this is the winner-take-all potentially is that it's just like very hard to distill that much information into the base model without ruining something in an iterative issue. - Like it's easy to distill into like a different base model. Like this is where I think it's mostly technique. It's not like, it's definitely not like just like there isn't capacity. Like if you had some model you know with all this data and you take like literally the same size model and pre-trained it from scratch with like all of this stuff in mid-training, it will be better. And I think that's a lot of what's happening today. And so it's very much like there's, you know, a bottom like the stops is from just keeping training the same model forever versus just like getting all the data from the old model and like training any model from scratch. And this is exactly just like saying like some combination like plasticity and like catastrophic for getting in that like the, you know, if you just naively train on like non-stationary data because you're adding new data as you go, basically this is messing with the data distribution. So like the old stuff is just forgotten. And we don't really have good methods to like stop that from happening. - And maybe at the, in the limit, you're like just bottleneck by retraining the model from scratch with all this information. - Yes. - Which of course is like very expensive. Like training model from scratch is expensive. - But you're going to do that anyways. - And so. - Not necessarily. I mean like maybe eventually if you have continued learning you're never training any model. You just like just have a model and it keeps learning and it keeps expanding. Right? - But there might be like some deep technical reason why that's very difficult because of these like- - I mean that's the question. - That's the question. - I think we have pushed back like how much from scratch we need to do, like it is definitely possible now to take like the pre-trained base and like do very good mid training on top of that like kind of continuously plus some RL from like different checkpoints that are later on in the training and like that's looking more like continuing learning about certainly not the case of like, you know, take the most recent model, apply a couple of very small updates and like iteratively like never lose. - So everybody's in this like, I'm a bit confused because isn't this literally what happens during training? Or during post-training or something, you just have like your model that's already gone through so much training. And then you like distill some fork that's been further RL or something. Isn't that literally what happens? And like why is- - It's still at a large enough scale, I think you're washing out a lot of like the noise. And like you're not just focused on one distribution which is where I'm said it's like, you know, that is now a very, if you're just like focusing on one toss, right? Like that's not- - I mean in the eventual regime, you'd be doing, I don't know, there's billions of deployed instances. You're like doing, you're learning from all of them at once. And so hopefully there's some washing out of noise and stuff from that, right? Maybe that's scale, yeah. - Yeah, I mean, I think like definitely is a sort of thing. Like you can do continual mid-training for like a long time and you can like well back to a checkpoint give any mid-training data. But at the same time like you can't do this like indefinitely. Like if you just keep continuing mid-training the same base forever, it just like get, it does, it sort of asymptotes at some point. Like you can't just learn new stuff in that base. And this is why people end up training new bases. Like otherwise you would just keep me training the same base forever. - Whenever I finish recording an interview, I immediately brain dump all my thoughts into Slack. Things like, what was the most interesting? And what should get cut? This ensures that my editors have all the contact they need to start editing the episode. But it's not like these brain dumps have any clear timestamps and my unended recordings are many hours long. They can take a ton of editor time to even find the exact moments that I was referencing. So we decided to try adding a rock-bought producer to our chat. And now whenever one of my editors puts a rough comment on it. of the episode, Groffbot opens a transcript on its own computer and starts working, usually before I even see the message. It takes the notes that I dropped in a slack and it highlights the relevance snippets in the transcript. It also uses a big case file that I've compiled with all my parapherensis, so it can suggest potential items. And when it's done, it sends me its top clip candidates so that I can review everything from my phone. This doesn't work really well. Being able to send informal messages like I'm texting my editor and then having the transcript immediately reflect my preferences has just been so helpful. Try Groffbot yourself at x.ai/bot. OK, let's talk a bit about data now. So I'm generally interested in this question of how much of AI progress is just explained by data progress. It doesn't mean it will be necessarily hard to automate, but this is ever a question. So is there some data distribution which if you trained current architectures on would result in a super intelligence that totally dominates human experts across every single field? Are we talking about pre-training plus post-training data like environments as well? I think the existence of this is obvious. It's just whether we can create the right environment. In the trivial case, we could just train it to output the Python file, which trains the actual super intelligence. Just have to memorize it in the way it's like, yes, there's probably a ladder of our environments that is possible to construct such that you would get AI researcher, which is at least as good as a human researcher, but the effort to climb each successive wrong grows kind of exponentially and that's going to be the two things that you have to trade off against as to whether how fast we're going to hit that final wrong where it's where it's better. I think that's fairly clear. And I think we're still relatively early in our environment creation. There's a lot of asymmetries that we exploit in order to create good environments. So one of the asymmetries which we've talked about before is there's environments where it's easier to go backwards and forwards. And what I mean by that is it's very easy to define this complex data generating process, and this is this kind of latent variable you keep hidden from the model. You can generate arbitrarily complex environments, and the model has to do a lot of irreducible, token spend, and irreducible work to figure out what that data editing process was. There's asymmetries in terms of you can inject information from the world, like anthropic finds a bug through tens of thousands of human and LMLs combined and turn that into a very very neat environment which is a single LML because they're radically fine within a few million tokens. So there's all these asymmetries which we're cherry picking and we're counting on this task horizon generalization. But I think again there's just going to hit diminishing returns at some point. At some point it's diminishing returns and how hard it is to create these environments in the first place, coming up with them because you can't necessarily just have these really, these processes where it's easier to go backwards and forwards, you actually have to sit down and construct something that looks with humans along enough time horizon. It's going to be a really complex task to create, and then it's also going to be the compute and time, but the next for the agent to actually do those tasks. So I think you're just going to start seeing this curve to fly up now. I saw something about how someone fine-tuned the talky model, which is only trained on data up to 1930 on this modern coding agent data, and it did better than Claude III Opus on sweet bench. So this model that has no knowledge of code whatsoever can be fine-tuned on a modern amount of data and behave better as a coding agent than this much larger pre-trained model is pretty crazy. It kind of shows you that once you have an example of the right expert behavior, it's actually surprisingly easy to copy that into a relatively weak model. But a counter example to that kind of is that there was a paper recently where they trained it up to like fifth grade maths, and also primary school English and stuff, so it was a decent language model, and they tried to RL it to do late high school and college maths, and the gap was just too large, they couldn't get it to climb at all, but if you did successive wrongs of like you know you're seven maths and then you're eight maths, and so on, like you could obviously climb to T12, so like again it's just like what is the distance between the wrongs on those letters, and how hard is it to create? Yeah, and this just comes back to like the RL signal problem, like RL is not very good at like exploring right now, and so if the model comes like in 128 RL, it's very likely to get signal to like progress, and this is why like in RL we need like curriculum, whereas like in pre-training we don't, because like it's that's not a problem for pre-training at all. And again pre-training, data is different to post-training data, and I imagine as we continue on like humans will be involved less than less, but that doesn't change the fact that you're bottlenecked on like how much signal you can extract from the real world, so like there's a lot of signal in the world, and that's true, like you know there's people doing like spreadsheet tasks, there's people doing like legal tasks and all this sort of stuff, but you know the capability frontier of where the models are at now, like how many bits in the world are actually like really relevant to like improving the models' capabilities, like you know how many new maths problems are being solved, that like there just be on the reach or grasp of the current models, like how many coding problems are being created or solved that are beyond the reach of the current models, and like I think that's why the dimension returns kicks in, because like even the world as a whole is not giving you the bits that are useful for tipping you into the next like base end of capability. Yeah, I totally agree with this, it's like really a question of like where the signal is coming from, and so like the signal doesn't you know, when pre-training of the signal is like already in common core, right, like for the tasks that you care about in pre-training, the problem is, there's not just like getting signal at all, it's like filtering out all the noise that exists, and that's quite a noticeable process, but like as the models get better as we end, and to mid-training and post-training, the signal just like doesn't exist anywhere in the original data we have, like no amount of filtering will like get this, you know, there's no like hidden proof of like the Millennium Prize problem sitting in common core, we can just like filter until we see it, right, and so like at that point you have to get bits some other way either from humans like directly like asking them to like write out their reasoning, or like by like creating environment where like humans decide like what environment should be created, what the objectives of these environment are, or like you know, some kind of like training on like the human data that exists in deployment, like you have to get the bits from somewhere. Yeah, yeah. There's a question of how much of the progress in pre-training is being driven by data. Yeah. I did this investigation with Jerry Hahn who's a student at Princeton, where we basically trained all the recipes from 2019 till now, pairwise with all the data sets from 2019 to now. So you say in like GPT-2 on the newest data set like ultra fine web and you train Delphi, which is the newest training recipe or the open source training recipe on like the pile or some old data set and you do like the whole grid and you see the getting to some level of capabilities, how much less compute does it take across this grid and you see that the data seems to explain like 9x of a computer efficiency gain, but the architecture improvements explains like a 3x computer efficiency gain at a very small scale. And so to the extent that that is true at large scale, that most of the pre-training computer efficiency gains are coming from better data. How much can that continue, like can you keep just filtering data more and more, I'm building more and more synthetic data until, yeah, I do have a sense of how much this kind of retraining progress can continue. I think my, my prior is that like again the low-hanging fruit is like somewhat exhausted with like, we got the internet as this big block and like there's, it's not like the internet is necessarily like growing at the same like race, all the useful stuff in the internet is growing at the same rate. So like we've probably got like a bunch of like 0.1 set loss drops to go, but like not to definitely not as many as have currently occurred. Yeah. But like that's also really interesting. They're like, you know, you find this like block cumulative like 27 times improvement across both. I think like it was epoch of someone who estimated like three times a year since 2019, which would imply something like, you know, three to the seven like over 2000, like times improvement. So like where's that missing, you know, 100 times or whatever coming from, like that probably gives you a good signal of like how much this is like post-training. I think the explanation has to be that a lot of the computer efficiency gains are scale dependent and we're starting an extremely small scale. And that raises the question of do the data, computer efficiency gains or the algorithmic computer efficiency gains have more of scale dependence. I don't know if you guys are probably around that. We just don't have enough computer investigating that question. I mean like just naively, right, like the theoretically the scale dependence of the architecture is like fairly well known. Yeah. And like you can fit a straight line to it, whereas like I would have no idea how to do that for like combining pre-training and post-training data and mid-training data. I mean, I feel like data is actually more like more important with scale. Like I feel like architecture is kind of like a one-time like, you know, an architect, I feel like combining like saying just like an x percent efficiency gains kind of misleading because like one architecture does is like let you reach like a qualitatively new regime which you couldn't reach with the odd architecture. And then within that regime obviously the data is like the prime we think determining it. But like you know, if we say didn't have like even like GQA, we're doing like full attentional day, we wouldn't be able to do like a million, we like predicts expenses to a million context. And like because of that we couldn't, we could never use the data which is like actually a million context. And so we couldn't get these capabilities. Even if like if you just do a naive, like how much does this do at like two K context where the architecture isn't unlocking anything then like the data, you know, that will look much more important than in some sense it is, right? It's unclear to me that these things are like really just like multiplicative gains in this way. Hmm. I see. So then what does this take over the scale dependence of data? Some in on scale dependence, I think like a lot of the like mid-training and post-training data we have now is like actually gets better with scale because like a lot of it like the very long context horizon of stuff really requires like big models to be able to like make use of them. Yeah. And like this is not, you know, if you try and train like your 100 million parameter model on like three bench traces, it's not going to get anywhere. Like it's not going to show you the same kind of improvement that you would get if you train like an actual sensible size. is model on it. Yeah. And it's hard as well now because so many of the architecture changes, you look at like Kimi, for instance, or DeepSeek, they're doing these architectural modifications with not just like dropping the pre-training loss in mind, but like, for instance how the models are going to be used in the real world, so like, yeah, the inference efficiency, like having some form of compressed attention in the DeepSeek models is not necessarily geared around, you know, this is fundamentally like a pre-training agreement, it's just like, okay, we're considering how the models are going to be used. Right, right, right. So one question I'm curious about to understand the future is how parameter scaling will go as we're getting into more of a RL heavy regime. Like I don't know how fast historically, you can look at sort of open source architectures and see how fast parameters have been scaling and maybe it's like roughly 2x every year for frontier open source models. And to the extent that like even frontier close source models have like 100 B or 200 B active parameters. Do you think that like keeps 2x in year over year or another we're going to RL regime where you also want to conserve compute on a rollout and also maybe there is like a threshold of the fact where you have enough capacity and at the point increasing parameters arbitrarily doesn't matter as much. Do you guys have a sense of in 2030 how many active parameters will a frontier model have? Yeah, I think for the next few years we're going to be like like because we have so focused on doing longer and longer horizon rollouts for RL where like inference efficiency matters a lot. And it feels like the models aren't necessarily saturated on their ability to that where the bottleneck is still the environments. And so we might see like a little bit of plateau like I have a feeling that you know like mythos and and the GPD models are much smaller than like you know the tentrally in parameter range that the people are talking about even just naively comparing the open source models you can probably back out that conclusion. So yeah probably for the next few years I wouldn't imagine a huge growth in the number of parameters. But again like there's so many different different things to trade off here like you just slide the size of your model based on like how much pre-training data you have and then like the difficulty of the RL environments that you've got to train on and you ideally want to like get to the optimal point where you know you can get like a decent pass at one or something on like the hardest environments you have and like it wouldn't make sense so like make a bigger model pass there because then you're just paying like much more inference what you need to. So there's a lot of inputs this like depends on how quickly you know like McCore and then in house these these guys can scale up the complexity of the RL environments they're all right. I would expect the models to keep getting bigger just because people are scaling up compute and the GPUs are getting bigger but I would say exactly how much they get bigger depends a bit on the scaling laws in non-obvious ways. So one thing is that I think like data efficiency is going to be a bigger driver than compute efficiency of like the exact architectures people use. Now that we're getting to the regime where we're sort of running low on like high quality pre-training data so that might affect how sparse you want to make the model. And then I also think we don't understand sparsity that well and it's like parameters are a different resource than active parameters but it's and like sparsity has definitely increased a bit but it's not clear that it's going to keep increasing without bound there might be some kind of sweet spot. There's an argument that sparsity should make data efficiency worse because you might have to learn the same thing on multiple experts so that's debatable. So I think we don't I don't think we have a good enough theory of scaling laws that we really understand why sparsity is helping and how much it'll help and if that'll like plateau at some point at a certain level of sparsity. So can you spell out exactly what the implication of data efficiency would be on, so it sounds like you'd say well it shouldn't be less sparsity but what are the other implications on grammar scaling? I guess just that the scaling law you're not necessarily looking for the most you're not trying to optimize compute efficiency so you have all your choices you can make on the architecture and each of these gives you a different scaling law and then like traditionally you would look at some kind of envelope based on compute so you would look at performance versus compute and take the envelope of like the best models but like if if we're making that decision based on data so it's like yeah we're sort of assuming we can spend a lot of compute and like we're sort of data is on our x-axis instead of in set of compute then we just get a different set of optima or a different set of models that are on that frontier. And I also don't think that we've necessarily like you know doubled like the size of the models every year for the last few years like the people have been training like one training parameter models for at least a few years like there was even an open source one called Falcon Bay like Liam from Periodic Labs like I think puts yesterday on Twitter about how like an early experiment at opening I was like training a one trillion parameter model that was very very sparse. Yeah that's what they did before opening I that was like before I could go to the switch range for you. So like it was like you know very very good at like knowledge but terrible at reasoning because it was so sparse and so like yeah it feels like we've been playing in this like a hundred billion to you know up to two trillion parameter range for like at least a little bit and like it certainly hasn't been this is like nice linear increase. Yeah I mean I feel like this two things so as Charlie was saying like inference efficiency is super important for our own and so like this will really push down active parameters quite a lot. And then I think the total parameters really depends a lot on the hardware as well. So like you really need to get like very high memory bandwidth and like the VRAM size to like actually be able to serve like multi trillion parameter models. And so like you know right now you know people still are still using a lot of like H100s and stuff. And so as everyone moves to GPs and then very ribbons will get like more actual like the ability to scale and like actually serve and like do like large RL inputs are like different at larger scales. The data question I think is interesting because naively like larger models are much more sample efficient in like the actual data points. And so like even if you're like not saturating the model it's still better to go bigger because like the models with larger models generalize better and like get to a better loss for the same amount of data. And so right now I think we kind of have a lot of data and like that's not the constraint well than computing. So we're having like small models which are like very insufficient. But if computers are along the bottom like it might come back to larger models which are like undisaturated but like they have this generalization ability because they're much larger. If you just look at like the basic central scaling law and you just maximize out parameters. Yeah. It actually decreases the amount of data you need to get to the same loss very little. Yes. Infinity on parameters the amount of data you need I think goes on less than 10x just because the nature of like the power. But we're now on the way to much data side of the Chinchilla laws right. So right now we over train more of the Chinchilla and so we could easily get back to a point to which as we're running our data we move back to like the Chinchilla optimal points. We've unlike a bit on the overtraining that you know and the training model side. But surely like even with these new chips that come online and stuff like we're just going to be so compute bottlenecks for the next few years that that won't necessarily be like. So this depends on like the ratio you have like training an inference compute really. It's like if you're super well-liked on data not on compute you should go bigger. If you're super well-liked on compute you should always go smaller and then like yeah. But you can also use computer generates synthetic data so it's like one of these very hard things to predict. Yeah. I think part of the reason it took people so long to figure out the scaling laws in the first place was that if you don't get all these things right then you don't get such a clean relationship and like the beautiful straight lines on graphs like hide a lot of complexity and how you have to make sure like to scale every hyperparameter the right way or like parameterize your optimizer in a way that scales and where you don't have to change your hyperparameters as you change the model size. And bugs have their own clean scaling laws as well right like you know like West Kaplan forgetting the cosine and the only thing or like even just like considering embedding parameters I think and so that messed up the estimate at smaller models because embedding parameters are a decent size of the model a bit on RL so I feel like a year ago a lot of people were making this argument that RL will not be super successful at scaling for models. I think Johnny wrote a research paper where you were pointing out that models learn one bit per episode when you are all basically learned did I get the answer right or did I get it wrong. And I wrote some blockbuster earlier this year I was like it's even worse than that because in the pass rate is low and the model is very unlikely to get the answer right it's the learns almost almost nothing at all from an RL episode. But we I look at the models today and they seem pretty smart and it seems to be the result of scaling up RL bearing you had a post I think a few weeks ago where you're trying to explain what's going on but why has RL been more successful than one would have naively thought. I mean so I think the success of RL comes down to a bunch of different things so first I think what is slightly underestimated is actually the mid-training so an awful lot of like what we see is success of RL actually comes from like very very good mid-training data which is basically where we're like essentially doing pre-training but aren't like synthetic reasoning data and like the kind of environments that like get the model warm started for RL. And so this actually takes the model like almost like 80% of the way to like the final RL checkpoint often. And then what RL does on top of that is it does like a lot of you know essentially tweaking to the policy. And so this is one of the reasons why it doesn't need like as many bits as you would naively think it doesn't have to learn all of these behaviors from scratch. It needs just like a few bits from these episodes which you do get and then the other thing that I really point out in my blog is that these bits are actually extremely high signal compared to like regular like pre-training which is why you need RL at all versus just like SFT on like successful reasoning traces. Maybe because it's exactly the bits about how to get the answer right. Well there's two things so yes one is exactly the bits about how to get the answer right but like this is not exactly how you think of it because in SFT you have a trace right you have like see a bunch of math reasoning and then the answer at the end. The bit is still there like you still SFT on the answer token. So that bit is still there what's important is the object ignore all the other bits. So in SFT you like have like you know to try and match like the exact reasoning token to the model producers. So you're essentially getting like too many bits about like the exact way this other model you're training on reasons. For RL you only get the one bit and that means that like this that signal is not drowned out in the noise of like all the other bits the model has. And so that's what really like it's really super dramatic like increase into the signals noise ratio during training, which is why like RL is like so dramatically efficient in terms of steps. I don't know if you guys have thought that. Yeah I like there's been so much debate about like what RL does to the model versus like you know mid training or SFT or whatever. And like you know everyone talks about how you know past one will go up but past 256 will go down like very rare correct reasoning traces will be like down weighted and kind of like outweighed by gradient signal from like easier kind of reasoning traces. And I think the simple like way to view RL now is that if you have a large enough like a large enough amount of compute to sample a large enough group size such that your probability of getting a bunch of correct answers is like past some like not insignificant probability then like it will be up weighted and like to to to balance point basically mid training and you know more pre-training like the the past one the starting point for RL like scales in the log number of pre-training tokens. It can answer very basic questions. I guess that answer makes sense and maybe there's empirical research which shows that this is what's happening. But then I just look at the models themselves and I don't know what's happened. Yeah maybe you can give me a sense of what is the basis of the AI progress over the last year but if it is yeah maybe it's just up waiting the policies which we're going to do the correct thinking anyways but it just seems like qualitatively the models have gotten so much more capable and anyways maybe there's nothing to there's no inherent contradiction there but how do we square square like the relatively small impact this take would imply that RL would have from the actual qualitative capabilities the models seem to be gaining. So like one thing I want to point out here is that like it doesn't necessarily imply that RL has a small like effect right even if you have a few bits and like you only change the parameters a small amount like the actual impact on like function space the model ends like the input to upper mapping can still be like super dramatic you know even if it's like even like one bit can change like your function space a lot in the it can like will have like half the hypothesis space which is huge. So like I don't think it's necessary the case it's like small amounts of bits small amounts of RL once you're starting from a really good point means that like you don't have dramatic impacts and behavior at least like not necessarily. I think it comes down to two things I think the first thing is that everyone was hoping that RL would like generalize this reasoning across like all these different domains and I think we necessarily got this like horizontal generalization like just training on math doesn't necessarily make you the greatest code like you do have to do RL on on code environments. I think what we did get though is like horizon generalization like the models just learned how to use more tokens for for longer and still make progress on some sort of task and so like you can train on environments where they get longer longer longer and then put them into a completely new environment and yes like they may have generalized the reasoning patterns which allow them to do well in that environment but they've at least generalized the ability to like continue on that task for longer which is correlated with like success something like there's a paper called edge bench which show that the rate at which models can work for longer is like doubling every three months so so that's a clear evidence of generalization and I think like the the final way to think about is like in pre-training there's this idea of like quanta so you have this very smooth like pre-training lost code and when you actually look at what's happening in the model like the model is learning all these like very discrete like tasks and there's like all these like emergent points but there's like kind of a phase transition like it didn't have induction heads now it has induction heads and there's like tens of thousands millions probably like hundreds of millions of these things and you average them all together and you got this very like smooth lost code I think like to us to an extent like a similar thing is happening happening for RL like there is this very slow out of loop as Bera mentioned of you know we will train a model and then RL and then like the next kind of model iteration of training we will dump a bunch of these synthetic reasoning traces into the mid-training data like we're kind of hitting all these quanta for all these different tasks and like on an individual task level it may look like a phase transition and like you're suddenly going from like a 0.5% pass rate to a 90% pass rate on like a particular like finance task or Excel task whatever but you average all these things together and plus the horizon generalization you kind of would go up wow we've got like qualitatively better models yeah I mean I think a lot of this as well it's just like I think RL does generalize a bit like suddenly you get like some transfer between like math and code or like puzzles and math and this kind of stuff also just like the the share amount to environment so I think the people are targeting it's just like vastly greater so like you know before when you try to do you know some task which like you do in your daily life like two years ago like the labs wouldn't really care about this they wouldn't like train them on for it and now like it's just so much broader right they have a lot of environments targeting this specific thing early in the conversation we're talking about RL in the context of causing this entropy collapse or pretend just you know concentrating probability on solutions the base model we're already done and causing relatively sparse updates in the policy but when I think like when I think about I think there's also another story about RL which is going back to the Atari games and then alpha go coming up with move 37 the super creative move that because it was never initialized on human data it can like think in ways that humans are not even thinking and come up with extremely creative solutions yeah do you ever sense on when we should expect or if we should expect RL on LLM's to result in things like move 37 just extreme creativity even beyond human creativity because like there's just a noble they noble initialization of intelligence I mean so a couple of things here like first off I think that the alpha goes using MCTS which obviously does like more exploration and like stuff than regular policy gradients but I kind of also think that like RL doesn't necessarily like reduce the creativity and like I mean even if we look I think this is obviously qualitative but if we look at like the the open air hugging phase incident like these models were coming up with like multiple zero days at a time to like break out of the sandbox and like this is clearly like some level of like move 37 creativity I think already which we just get from just like the general generalization properties of the RLM's like I don't think it's definitely the case of like RL's like totally destroying their like entropy yeah especially along horizons yeah I mean one thing that people call creativity is just solving hard search problems so and so that's like like move 37's obviously an example of that or like writing some kind of poem that satisfies a ton of different constraints so that's something AI is obviously going to be extremely good at if if trained for it then there's another way in which the models like the diversity of their outputs is a lot lower after RL and they sort of develop these ticks and like even though the models seem like they're good at writing when you do some kind of like distributional analysis you find that like they're reusing certain themes like all the time and they're using the same character names all the time so there's actually it's not like you're getting the same kind of diversity that you get when you like from human authors you're sort of getting one really good like style so I think that like that kind of diversity has definitely been like cut down by RL a lot and in fact not oh yes since we were talking about distillation earlier that's sort of something yeah one thing that's happening is that so many people are just stilling mostly from Claude that like like all the open weight models for the same way as Claude and use the same like have the same ticks so this seems kind of concerning to me that we're having this like this monoculture emerge again I don't think this is like fundamental to RL it's like a method though and same with distillation like even with distillation like you're just training on the data it's like just because your data is not like super bored that doesn't mean like the training method itself is somehow wrong it's like a problem with the data and I think a lot of for instance like the RL like entropy collapses basically due to like exploitation of fairly simple like verifiers when you don't have like a huge diversity environments because like for instance like the writing I think the writing is presumably graded by some judge and like the judge has some specific ticks and like the model is learning to award hack the judge and that's why like it collapses but like this is really a problem with the judge is not a problem with like RL in general okay super rapid fire predictions about the future so I want timelines on the following couple questions by when do we have models which you can here's what the it feels like to a user you basically hire them as a drop-in-the-mote worker for all kinds of white collar work not just coding but I don't know video editing law paralegal etc like it's like literally an actual remote worker but like full computer use with like literally a month of seamless learning and operation and executing on like complex projects and it required interacting with other people etc etc everything a human worker can do over a month if you like mandated to use like a browser or whatever other like these that again the firm setting up the information to be like programmatically accessible like maybe a couple years but if it's not like browser based like you can send Slack messages they can do all this stuff but still probably say around here yeah I mean I would say maybe like for the like full generality maybe like three years but I think to try this point we'll end up with like a lot of people like making their organizations easier for the AI's to use and so you get like 18 90% of the way there before that so I mean the thing that's the death between one year and three years ago is just literally like like that I think there's going to be like a long title like miscellaneous stuff which like some human can do which like we'll take them all it's like quite a while to do yeah they give me everything you have sort of computer your stuff or like basic cognitive capabilities I mean I think I think this really comes down to questions like, how quickly can we solve this kind of like online learning and like whether we can like get like 80, 90% of the way they would like compaction and like writing files to yourself and stuff. And like that's my big uncertainty, I really don't know. - And another, like maybe an example something that wouldn't be good at is like, you know, if I have to like yell at someone to get something at work or like really push someone to get something done. Like the model isn't just going to do that. It's just going to be too nice. Yeah, yeah. - I'd say there's a wide variation in quality of human remote workers. So if you try to hire someone like off of off work to do a software engineering project, there's going to be a huge variation. It's like often quite hard to get them to do, like to do a good job or like pay attention to all the feedback you're getting. And like, I would guess that in some cases it like, it'll be worse. Like the pre-AI version of this was worse than what you can get now from existing AI. So I think it might end up being a little complicated 'cause maybe to some extent we already have this, like for some, like not so high quality of work, but then like, then it's obviously like we're not, yeah, we're not matching human level in certain like higher quality, like forms of work. So, but I basically agree with Charlie and Baron that maybe, yeah, we'll, yeah, we'll have some version of this in a year. So that's like, okay, and it will be able to do, maybe we'll have that form factor and it'll be able to do some things really well, some things not so well and things will be improving from there. - Like we shift the goalposts based on the very long tail all the time. Like, I think it feels like he uses example before of like doing your taxes or something. Like this year I literally just like told Codex to like go get everything I need to do and send it to the accountant. And like there was massive list of stuff it had to use computers to click through and like download some stuff. And I didn't know, it was like fine, it was perfect. So like, I don't know, a lot of this stuff, it can already do. - Yeah, okay. Give you 10X total productivity uplift. - Basically, if it takes you a year to make a breakthrough now, you make a breakthrough every month. - I think I would just refuse to give you a scaler on this. Like we might already be past that in some like types of work. Like, let's say you're just trying to prove, you're trying to do like certain types of math. For you as AI researchers, trying to make like advance the state of AI research. Because how much are like AI researchers sped up? - Yeah. - We're uplifted. - Somewhere between five and 10 years. - Oh really? Okay, that's far away. - Well, do you think it's longer than like full-general of remote worker? - Yeah. - Interesting. - I think you're probably, I think I'm realizing you probably have very different definitions as a fully general remote worker. I could have specified that. - Yeah, this is true. 'Cause I mean like yeah, because obviously like an AI research it can be a remote worker and so like. - Yeah, no, I'm pitching like, you know, normal white color work over the period of a month. Yeah. I think it starts to diverge a little bit past a month. - A very competent white color worker. - Yeah, but not necessarily like a super creative researcher. - I would say like two years. - Two years? - Yeah. - 10x, okay. How are you, Bernard? - I can kind of see that actually. 'Cause like it really is just like, right now it's already like definitely more than 10x of like coding stuff. And so it's like, if it can do it even like one or two loops of like experimental feedback, that would actually be massive already. So 10x uplift of AI researchers within two years. If you just plug it into like a very naive model of like AI progress and how much is coming from AI researchers and they're like, there's like a 10x increase in their productivity. Yeah, you have like radically accelerated pace of AI progress starting two years from now. - Yeah, I mean, I think like this will mean that AI progress doesn't get bottlenecked on like, AI research and ability to run like small experiments that gets bottlenecked on other things. - Of course, of course. But it just like happens 10x faster. - So sure, yeah. This is a huge deal. And that also like helps the next thing which makes it gives you 100x speed up, happens sooner, et cetera. - Yeah, I'm happy to stick with a longer on that one. - And what's like the crux? Like my capacity to absorb information and make like Bayesian ultimate decision on the next experiments, makes sense. - Yeah, I mean, I'm assuming that like you can delegate some of this to the AI. So like AI is becoming decent at like deciding you know, it's run this experiment, it's got this result, it runs like the next experiment. And then if it can run like two or three experiments and run without like crashing, then like that is actually a big update in like uplift. - And okay, final question. And AI which is, which dominates top human experts across every single field of kind of work that can be done over a computer. So not only a research, but all cognitive work. And not just like short horizon work, but like literally if it takes like three years or something, the AI will still do better than humans. - So this is basically just like ASI. - Okay. - I would say like three or four years. - The fuck. (laughing) I mean that doesn't seem wrong. - I would say like AI is obviously being more, getting more attention. So it's like one of the harder things, but it's like a lot of energy is being put into it. And it's also like not one of the hardest things for AI 'cause it's like involves a lot of code and math which models are really good at. Maybe for things that involve like 3D and like spatial stuff and physical stuff, I think that will take a little longer. So especially if it's not like, yeah, if it's like mechanical engineering or something and it's not getting like the most attention right now, that might take a little longer. - But it also just includes fields where there is relatively little data because of the nature of the field. And it has to like learn that data on the fly. So for example, it has to become a superhuman at like being an engineer at TSMC or something. - Oh yeah, so you would have to assume that like the onboarding, yeah, you can give the AI the same onboarding material. And yeah, then there's some like something has to be solved about like sort of longer horizon learning or yeah. - Until five to 10? - It means yeah, you think automating AI research is like ASI complete or something? - Yeah, I think so. Yeah, I think there's so many things in the world, which like even if you have some sort of memory system external to the model. And even if like context length goes a little bit, like there are just fundamentally things like, even if you could research the information or write notes yourself, like you'd need more than amount of the context fields there. - Yeah, I mean, I kind of agree in like the five-year range, at least for like the stuff that like labs are focusing on, but I think like there's gonna be a long tail of stuff, which like, yeah, I could theoretically go out and learn about, but like no one has bothered to do it. Like the computer's been allocated, so that's it. That might take longer for like literally everything on the human experts. - Sorry, but by this I also included like the ability to learn as fast as a human, a new domain. - I mean, I think that's not necessarily necessary actually, 'cause like the AI will have vast degree to experience than like any human. - Right. - Thanks so much for doing this, yes. I feel like this was a great format for getting different experts to disagree and debate and discuss things together. It was very productive. - Cool. - Thanks for having us.

Podcast Summary

Key Points:

  1. The most likely reason AI hasn’t led to a radical transformation by 2036 is the persistent difficulty in achieving true generalization, where models fail to transfer knowledge across diverse, real-world tasks despite strong performance on narrow benchmarks.
  2. Current AI progress is bottlenecked by limitations in self-referential reasoning, objective specification, and the ability to discover new scientific paradigms—highlighting that scaling alone may not lead to superintelligence.
  3. Human judgment and alignment remain critical for defining AI objectives, with the final role of humans likely being to specify what "good" means, rather than just optimizing for technical performance.

Summary:

The absence of a radical AI-driven transformation in 2036 is not due to external shocks but rather technical limitations rooted in the difficulty of true generalization and persistent bottlenecks in AI's ability to reason, self-improve, or discover new paradigms. Despite massive scaling and advances in training methods like reinforcement learning, models still struggle with complex, non-stationary real-world tasks that require dynamic, context-sensitive reasoning—such as law, business, or long-horizon problem solving. These challenges stem from the difficulty of transferring knowledge across domains, the lack of efficient self-improvement loops, and the inherent gap between narrow benchmark performance and real-world adaptability.

The field's progress is increasingly driven by domain-specific training and simulation-based environments, which, while effective, fail to capture the full complexity of human-like interaction and learning. Crucially, AI still lacks the ability to autonomously define or optimize its own objectives—making alignment and human oversight essential. Even with advanced simulation, real-world deployment data remains underutilized due to poor reward function design and low sample efficiency.

However, early signs of progress—such as online learning in tools like Composer or Cursor—demonstrate that models are beginning to learn from real user interactions. The most plausible path forward involves a combination of simulated environments, human-in-the-loop feedback, and incremental improvements in model efficiency, rather than a sudden intelligence explosion. Ultimately, the key bottleneck may not be compute or training, but the fundamental difficulty of creating AI that can generalize beyond pre-defined, easily verifiable tasks into the messy, evolving reality of human work and innovation.

FAQs

The most likely technical reason is that AI still faces persistent bottlenecks in generalization, especially in complex, open-ended tasks. Current models excel at narrow tasks but struggle with real-world, dynamic environments requiring deep understanding and adaptability, which limits their ability to achieve true recursive self-improvement.

Generalization—the ability to apply knowledge to new, unseen situations—is critical. AI models often fail to generalize beyond training data, especially in tasks requiring human-like intuition, such as social interaction or long-term planning, which hinders breakthroughs in scientific and engineering research.

AI struggles to define or discover novel, high-value research objectives on its own. Humans still hold the key role of specifying what to research, including ethical boundaries, goals, and real-world constraints, which AI currently cannot autonomously determine or optimize.

Distillation allows smaller models to learn behaviors from larger, more advanced models by using training trajectories. This reduces reliance on a few dominant providers and enables widespread, decentralized model development, especially in regions with limited access to leading AI infrastructure.

Yes, early signs show that real-world deployment data is already being used to train and improve models. Companies like Jane Street and Composer use user feedback to create targeted training environments, demonstrating that real-world interaction contributes meaningfully to AI learning, though full self-improvement remains distant.

Tasks involving human interaction, dynamic environments, or evolving social norms are extremely hard to simulate. Real-world interactions require adaptability, context awareness, and long-term relationships, which current simulation environments cannot fully replicate.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.