Go back

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

53m 43s

Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

Rich Sutton, a pioneer in reinforcement learning, argues that the AI field has become "weird" in its focus on human knowledge and data scaling, rather than on continual, experience-driven learning that naturally evolves over time. He asserts that true intelligence lies in systems that learn from their environment continuously, not in static models trained once and then frozen. His seminal work, "The Bitter Lesson," emphasizes that long-term progress depends on algorithms that scale with computation, not on human-curated knowledge. This view is challenged by the rise of large language models (LLMs), which he sees as both a positive—enabling massive scaling through data—and a negative—eventually hitting limits due to finite human knowledge. Synthetic data generation, while promising, remains bottlenecked by human expertise and fails to capture the infinite complexity of the real world. The "big world hypothesis" posits that the world is infinitely complex, making it impossible for any single system to learn everything, thus requiring multiple agents learning from their own experiences. Oak Lab is founded on this vision, aiming to develop self-improving, continual learners through novel algorithms like continual backpropagation and generate-and-test mechanisms. These systems would learn abstractions from experience, plan with internal models, and avoid catastrophic forgetting. Unlike current LLMs that stop learning after training, these agents would evolve over time. The team prioritizes small, focused growth and deep alignment, aiming to build a scalable, self-consistent AI design that can adapt to diverse tasks—from physical robotics to abstract reasoning—without depending on human input. While ambitious, Sutton believes such systems are within reach in the next five to ten years due to advances in computation and efficiency. He stresses that AI must evolve beyond language mastery to achieve true fluid intelligence, where learning, abstraction, and self-improvement are unified in a single, adaptive framework.

Transcription

9599 Words, 51670 Characters

English
People think I have a radical point of view sometimes. They start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And I mean that like, you know, it's just the recent times people are thinking weird. Before there was all this AI craziness, you talk about, you wouldn't have to say continual learning because it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. We always act and we learn. That's just the normal way of thinking. I'm not weird. The field is weird. The field, they need to call it continual learning. It's just learning. We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You had the key students in the field, folks like Dave Silver. You wrote the essay, The Bitter Lesson, that I believe is the Bible of the field. And you have just been one of the greats in propelling the field forward. So thank you for taking the time to join us today. Rich is joined by Kuram Javed, his co-founder and former student from the University of Michigan. He's a professor at the University of Michigan. The two of you have set off to found Oak Lab. I'm very excited to talk to you about that today. So for today's session, we're going to start talking about The Bitter Lesson, the state of the world as we know it today, whether LLMs will get us there or not. And then we're going to transition to start talking about your research agenda and your plan for Oak. Rich, maybe take us back. I was going to start with The Bitter Lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy. What gave you the conviction to do that? What else are you going to do? We're trying to figure out the mind and learning is a central part of the mind and having a goal is a central part of the mind, central part of intelligence. Yeah. So I was just doubling down on what I was always thinking. Did people think you were crazy at the time? It was a winter. It was an AI winter. What year was this? It was in 2003. And it's kind of crazy, actually, the truth, because I was really sick. I was actually dying of cancer in 2003. But I wasn't quite dead. I've been trying for a number of years and I wasn't dead. I was in another remission. And so I said, well, I'm not dying. I haven't succeeded in dying. So I might as well, you know, it's going on long enough. I might as well just try to get another job. And so I went to Alberta and started teaching there. And then in the end, I didn't die. It's kind of amazing. It's like that. I'm joking about it now, but it was quite serious. And it's an even more poignant question. Why did I continue to work in Alberta? Why did I continue to work in Canada? Why did I continue to work on this research stuff when I was, you know, I only had a few months? I would always keep reminded what I think it's Benjamin Franklin is supposed to have said that, you know, if you ever wonder why someone is doing something, it's almost always one of two things. It's either habit or vanity. Okay. So I think it was probably true. Maybe it was a habit to just kept doing what I always was doing, or maybe it was a vanity. I don't know. I think it was more like habit because I was dying. Wow. Divine intervention. Yeah. It's always been easy for me to be very determined. And I'm going to go even longer on this answer. They start questions saying what I'm thinking is so different from everyone else. If you look back, what people thought about the mind for, you know, even just a decade, you'll find the kind of thoughts that, you know, learning is important. You've got to have a goal. And, you know, perception is important. We are low level beings. We are generating actions and perceiving data at a fast speed. And yet we have to think at higher levels. And, you know, go back a few, before there was all this AI craziness, you talk about, you wouldn't have to say continual learning because it wouldn't make any sense to talk about learning that wasn't continual. You know, it's not a special phase. I'm not weird. The field is weird. The field they need to call it continual learning. It's just learning. I'm not weird. Everybody else is. That's a good, though, to live by. We're going to have to send out a next post about that. We're very happy that you lived on. The field is happy that you lived on. And thank you for pushing the frontier of AI. I'm really happy. Thank you for pushing the frontier of AI. I'm sure really happy. And thank you for all of that. And you've been able to sort of educate a lot of students who pushed the frontier as well. How did you pick them? How did you, over the last 20, 30 years? Oh, well, you are giving me opportunities to be humble. I like to be humble and point out how all these great decisions are just happen. And that's the way I feel about students. I don't feel that I choose them very well. Sometimes I'm lucky, sometimes I'm unlucky. I don't feel I'm particularly good at picking my students. I'm looking at Kerm now. I think sometimes you end up with really great ones. David Silver picked me. How is it that I got you, Kerm? I finished my master's, not with you. And I was planning to join industry. And then we were collaborating on a project, which also just started organically. There was something I worked on that Rich was in a meeting. Then they mentioned that I worked on it. So I got pulled into it. We started collaborating. It went really well. I felt so happy with that collaboration. Rich also felt really good about it. And then six months down the road, we had made some progress. And it just made sense to convert that into a thesis proposal. So at no point did I apply. At no point did I ask, should you be my PhD advisor? We worked together. Then we decided this would be a pretty good thesis. And then after that, I applied for the PhD. Life works in unexpected ways. Take us to 2019. You wrote The Bitter Lesson, which has become the modern tome. 2019 was a funny time to be writing that piece because ImageNet was 2009. AlphaGo was 2015. What caused you in 2019 to reflect and to write that? Because it was before the current kind of scaling paradigm around large language models had taken off, but it was after deep learning had really proven itself. Well, it was a long time coming. As The Bitter Lesson expresses, it's something that you can observe for a long time, for many decades. And it's definitely at least as much due to the round of symbolic AI, which I lived through. It's all about not getting distracted by trying to put in your human knowledge and just paying attention to what the problem needs and how you can scale with computation. I know I wrote versions of it at least a year before. And I gave talks. I gave a talk a year before. And it wasn't a particular response to the moment. It was a particular response to my long experience. Different people trying to think in different ways about how you can make smart systems. What is the essence of The Bitter Lesson? It may be the phrase that I hear used the most in my meetings these days is Bitter Lesson Pilled. Is it not Bitter Lesson Pilled? I would imagine given the popularity of the phrase, it's probably been tortured and misused in different ways that you didn't originally intend it. So what is the essence of it? And where do you think people go wrong in their attempt to understand it? Yeah, you're making me think about X now. And I recently made a post where I tried to do the Bitter Lesson in 26 words. It goes something like, don't be distracted by human knowledge as AI traditionally has been many times. Instead, focus on learning methods that will scale with computation like search and like learning. So it's really all about focusing on algorithms and improvements. It's not saying you don't need fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with computation. Rather than scaling with data. Rather than scaling with human input. Yeah. And then the question, if I can anticipate, yeah, what about large language models? Are they consistent or inconsistent with your essay? Yeah. And I've thought about this and I think there's another X post about it. But the conclusion is that it's both a positive example and a negative example of the bitter lesson. First, large language models enabled enormous scaling with computation. And you could just drink in the Internet and scale so much. So it was a way of getting a much more capable system just by methods of scale. Then after that, as you go on further, it eventually gets limited by that information. The Internet is finite and it's hard to get more examples. And the world is big and the world is massively bigger than everything we stored on the Internet. And so in the end, it seems like it could be, I guess that would be a positive example of when we relied too much on human knowledge. And it eventually holds us back. Can I just push on this a little bit? Yeah. It seems like a lot of what the Foundation Model Labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human Internet. Is synthetic data generation kind of as part of this LLM scaling paradigm, is that a general method that leverages computation? No. Why? That's just a big mistake. Well, it's such a big, it's such a, maybe it's the next, the next big lesson. It's been floating around Alberta for five or 10 years. Okay. And we call it the big world perspective or the big world hypothesis. Kuram, who eventually wrote it up as a paper, there's a little paper called the big world hypothesis. The big world is that the world is infinitely big. There are infinitely many things. There are infinitely many things to learn, and you can have people generating these synthetic data sets, but they would always be more things to learn. And because of that, if you could just learn from experience, you could remove the humans from the loop, then you would have systems that can do everything because, you know, the world is big. There are many tasks that we want them to do, and they would be able to do anything by learning from their own experience. Going back to the synthetic data question, too, who decides what's a good synthetic data and what's the bad synthetic data? Yeah. Because I can write a program that can output a lot of synthetic data, which would hurt programs. Right now, I would say humans decide, and that's the bottleneck where, okay, you can have humans deciding how to generate these data sets, but you need human experts who know what's a good data set and what's a bad data set for that approach to scale. So it is bottlenecked by humans. Doesn't my loss curve decide, like, how much better did I get with this data set versus that data set? Right. But if all the engineers, OpenAI, Anthropic, or all the big new labs, the engineers went on vacation. Who would generate the synthetic data? That's the question. It doesn't come from agents' experience. It's not something that the agent is generating itself. Some human has to decide what is the right synthetic data to generate. And that requires human expertise. So, for example, if you want a system to do something very challenging from a physics point of view, maybe you want a drone that flies with echolocation, like a bat, for example, what's the right synthetic data for that? I think you would need to hire domain experts. To go figure out what is the right data and generate it. And then maybe you would be able to learn from that. But the domain expert has to exist first. So we are bottlenecked by human expertise at that point. But you can have infinite synthetic worlds. The existing world is finite. But let's go back to the echolocation thing, right? That's what I want. I want a drone that can localize itself and move with echolocation. That's my goal. The robot that is a robot is generating its own experience. So it could totally learn from its own experience, but it wouldn't be able to. It doesn't matter how much synthetic data you generate. It doesn't matter if you generate synthetic data that captures 50 different universes. It will not allow you to do that task without humans figuring it out first. At first, just say the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, because there's going to be a little program that will generate the synthetic data, it'll be a small world. So, for example, what's important to me is what's going on in your mind right now. Okay? And you're saying, why don't I get some synthetic data to tell me what's going on in other people's minds? No, there's no way we can have synthetic data for other people's minds. And other people's minds matter to us. You know, like. I talked to you guys about investing today, so I care what's going on in your minds. But how can I get synthetic data on such a thing? Really, you can't even get synthetic data on anything. You can't get synthetic data on how the drone is going to interact with its environment, in the physical world, in the friction, and where in the motors of this robot. The world is infinitely complex. And any simulation of it is, like, microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agent. And this is obvious because the world contains many other agents. So, because the world is massively complex, there's no way you can do anything that might claim to be optimal or perfect. You're going to be imperfect, and you have to have approximations. And those approximations will be severe. And so, because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason, we have to continue learning because we'll encounter some particular part of this immense world, and we'll have to learn an approximation that's tuned to the part of the world we're in, not to all the other parts that we're not in. Yeah. I'm going to push on this one more time. And sorry, I'm being argumentative for the sake of being argumentative. I'm trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim. And then they have to do some sort of post-training, I guess, to make sure they work in the real world. But that it's been a very effective pipeline. Yeah. So, I think the important question to ask here is, how many engineers were involved in building that simulation? And are we ready to say that the only problem worth solving are those where we can hire a large team of engineers? So, I think the only problem we have is that we can hire a large team of engineers to first make a simulation. And I'm sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a sim-to-real gap that was not acceptable, then they fixed it. So, there is this human in the loop fixing the simulation, like they're getting feedback from the real world, humans, and then they're fixing the simulation. Why can't we just remove the human and let the agent do it itself? And then when it actually drives, again, something unexpected will happen. And that's when you actually want to learn from experience. Yeah. Because there's just so much more data that's going to come from experience than there possibly can be from humans curating, creating data. Yeah. And I think there is obviously value in learning from simulation. And there is a way of doing it. The agent can learn a model from its own experience. And when the agent learns it, it's much better because if the model is incorrect, it can fix it by continuing learning. If the humans are making a simulator, then the model only gets updated when the humans figure out that something's wrong. So, yes, planning is important. The agents should learn from simulators, but simulators, they make themselves. Okay. I want to move to another part of the bitter lesson, removing human knowledge. From your essay, "Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain. But the only thing that matters in the long run is the leveraging of computation." And if I, you know, your former student, Dave, with AlphaGo and AlphaZero, for me, that was an example of a triumph. A triumph of removing human priors. Did that result surprise you? I guess, why or why not? Of course, it made me very happy and made me, you know, feel vindicated. You know, it could have gone either way. It wasn't that because prior knowledge can help. You know, there's nothing wrong with prior knowledge. You know, and I say this right at the very beginning, right at the very beginning of the bitter lesson, I say, there's no reason why there has to be a conflict between prior knowledge and then learning knowledge. You know, you can put some prior in there and then start learning. There's no reason in principle why these have to be opposed. In fact, they're all about knowledge. You know, life is gaining knowledge and having knowledge. And why are these, how somehow, you know, nature and nurture became enemies. But really, you know, prior learning is what you already, and then you learn more. And it's, they should be friends. But as I say at the beginning of the bitter lesson, in practice, they have been enemies. In practice, people who had an affection for existing human knowledge ended up, you know, wanting that to win. And so they wanted to minimize or dismiss learning. And so now I'm sure your sense of me is that I'm someone who loves learning and wants to dismiss prior knowledge. But, you know, I'm really someone who's interested in the mind. And the mind is, you have prior knowledge and then you get more. And then once you've gotten more, then that becomes your prior knowledge as you get more and more and more. And then. these two things work together. I end up appearing to be someone who's interested in learning primarily because all the rest of the world is talking about all you need is enough knowledge. You don't need to learn. You know, large language models are, oh, we're going to put all this knowledge into the system. And the large language model will not learn when it runs. You know, it's talking to people, it's interacting. It is absolutely, the weights never change. So, you know, I am not the weird one. It's you guys that are the weird one that think that that's possible that you could possibly, you know, they claim they can make a PhD level experience and expertise out of something that doesn't learn at all anymore. You know, so, you know, I'm not the weird one. So your recommendation is let the algorithms run for a much, much longer period of time before feeding it. Continually learn. Before you feed it data or prior data. So drip or drip the prior data along the way. Prior knowledge. So both are important. But in the long run, you've got to gain and structure the gaining of new knowledge. And that's what all that matters in the long run. And as you are doing this, yeah, there'll be some that you got previously. Like how it work if we, you know, look into the future when we have intelligent robots. Will we have them all learn from scratch or will we like copy them and ask them to keep learning from wherever they are? I mean, they'll be digital and it'll be easy to copy them. And so instead of having like this huge thing where we're spending zillions of dollars to retrain them from the internet, we'll just copy the agent and keep learning from there. And so in some sense, the prior knowledge will be, should be dismissed because you're just going to copy it from the previous robot. So why don't you just describe for us what you think a machine or a computer that learns from experience looks like? Well, it could look like a robot. It could, but it also could live entirely on the internet. You could like, for example, routing of packets through the internet and do that in a way that's sensitive to experience and becomes better over time. Or you can interact, be the user interface that's interacting with people, like on your phone. Or on your computer and it becomes better over time. You know, like an intelligent assistant, you know, has to become better over time, has to know what you want. Would your contention be that the current paradigm of, you know, the popular LLM-based assistants, would your contention be that these are not experiential learners or continual learners? And if so, what is the fundamental gap? Are you serious? I mean, obviously. They learn memories about me. They're, you know, they're doing some in-context learning. Their weights never change. And by the way, is a small number of the weights changing sufficient, or do you need all the weights to be changing? Well, so think of all the structuring and generation of new concepts that went into creating the large language models. All that is the weight learning. And you want to be able to continue doing that. You don't want that. You don't want that. You don't want that. You want you don't want that to happen just once. Is another way of saying is we do too much pre-training and post-training before we launch the models, and they don't learn after that. The only point that the big disagreement is we don't let them learn after that. Yeah. We don't let them learn after that. We can do as much pre-training as we want. That's okay. Post-training is fine. But then when I'm using the model, I cannot. It stops learning. You can give it more context. You can change the state of the model by giving it more context. And so it has already learned that, if the state is different, if the state says something new, then it will use that to make the next prediction. But the model is not learning. Cursors, tab, autocomplete model. It does get updated based on. Those models. Those weights change. Those are two examples. Cursors, tab, and I think the composer, they were also updating. Those are two examples of continual learning. But it can be much better. So the way they do it, as far as I understand, is a lot of people are using tab. They collect all this data. So coming from millions of people using tab, they collect all this data. So coming from millions of users or thousands of users, and then they do one update of the policy from this batch data. So this could work, but imagine I want to teach this model something specific. I don't want to fight with 100,000 other people about what they want to teach their models. I want to teach my model something very specific, and I want to do it to my version of the model. I don't care about the shared knowledge that the model has coming from other people. And so it's a very inefficient way of doing it. It seems like the way that this is currently done is that there's fundamental skills, maybe, that are learned in the weights that are common to everybody. And then there's personalization that happens in the form of context, right? Is that not the right mental model for how learning should work? Should all the contexts live in the weights themselves? So context can be in the state, too. It could be both. But you still need to be able to update the weights. So if I give you an example, some really good use studies are with HIPAA. I'm going to talk about HIPAA, and I'm going to talk about how HIPAA works. HIPAA works for people with disabilities. When a human goes through something that changes their mind or some sensors, you can see them adapt. So for example, we have proprioception. We have internal sensors that tell us how the body's positioned, and we use this for walking. There are cases where people lose this ability completely, and then they can't walk at all, because that is literally the foundation of their walking policies. It is ingrained in the brain. But then over the course of two, three years, they can learn to walk again by looking at their feet. So visual feedback. I'm going to talk about how HIPAA works, and then I'm going to talk about visual feedback. So brain is insanely plastic in the sense that it can learn a lot of things, something that has been true for 20 years. When it stops being true, it can go and update that and get rid of that. And that is the capability I think that's extremely useful we would want in our systems. What is there for us to learn from how human babies or animals learn? And how much inspiration do you take from that? Well, we take a lot of inspiration. We don't take it as a requirement that the AI has to behave like the natural system, like babies or people or animals, but it's a source of inspiration. Inspiration, but not constraint from animal learning. Consistent with the bitter lesson. Where do you think we should most seek to draw inspiration from the way that biological learning works that is not present in today's systems? I feel like I'm just giving opinions now, but they're just obvious opinions. So I think it's apparent that no animal learns by supervised learning because we don't get examples of how our muscles should twitch, and that's our output. But all of school is supervised learning. I know, absolutely not. But even if it was, school is like a tiny fraction of what we learn. Like we learn to see, and we learn to walk, and we learn how the world works. But even in school, no one tells us how we should twitch our muscles. The knowledge skills I acquire were from supervised learning in school. So I don't want to say that learning from others, transmission from others is not important. It's extremely important, and language is extremely important. But what are we, what are we missing? You know, there is, there is no supervised learning. There's no targets that are given to us. You know, you, you hear the right answer is, you know, where, where is, what's the capital of France? And we know the answer is Paris. Okay. But no one tells me how I should pronounce Paris. You say the answer is Paris, and I listen to you, and I hear your words, and you know, I will make some other, uh, muscle motions to produce the answer Paris. It's not literally supervised learning. Um, anyway, yeah. So I, I think it's really true. I mean, well, anyway, the first thing is the school is, is irrelevant. Like, you know, squirrels don't go to school and, and, and learn that. Animals don't learn that way. It's, and school is a very special thing that even we didn't have up until, you know, I don't know, a few hundred years ago. But it's not part of, it's not part of the essence of intelligence. And it's a distraction to think of that as your primary example of learning is this thing, which we didn't do as animals. I wish you had been around to tell my parents that. They forced me to have, get, go to school and deal with all the structure. The thing is like squirrels are wonderful at jumping off trees, but squirrels can't prove math theorems. And if I want to learn how to prove a math theorem, I go to school. Yeah. Uh, they also don't have, uh, DVDs and iPods. You know, there are a lot of things, they can do things that we can't do. Um, but math theorems, uh, yeah. And they don't play chess, you know, sort of like more of X paradox. They're, uh, they're these advanced things that we think of as really intelligent, but, uh, they're sort of easy for computers to do as opposed to all these regular things that are hard, like moving and seeing with attention and everything. Um, I think supervised learning is a, is a good thing. You know, just mentions, I like to think, look for obvious things. No one tells us how to twitch our muscles by giving us examples because they couldn't possibly because we have had to twitch our muscles we've had to figure that out yeah and their answer would be wrong right so if i moved my mouth and my tongue and my vocal cords exactly the same way that rich does to pronounce paris i'm sure a very different sound would come out so in some sense rich or no one knows the right way of producing a sound with my body only i know that yeah it seems to me that many of the most i guess the most raw like sensory motor capabilities especially related to movement in the physical world i agree with you that that seems something that's inherently learned from experience it seems to me though that there are higher levels of abstraction that bring us closer to you know what makes humans great and much of that doesn't live in this low level of sensory motor learning does your world model i guess span sensory motor learning all the way up yeah that's the ambition absolutely and squirrels by the way can do some enormously abstract things what's the coolest thing a squirrel can do well it can always get into your bird feeder no matter what obstacles you put in the way you know it can find new ways to jump and climb and okay and do lots of things calculate trajectories pretty well animals are pretty good at understanding the physical world without the mental calculations that we think we are doing when we think about launching ourselves into space breaking a fall they can do it in real time the right way to prevent injuries okay fair enough i think it's just a question of degree between and i like to think that animals other animals are are very close to humans i think it's hubristic to try to emphasize what we do differently you know how we're different from animals it's better to see the commonalities and i think we are just a question of degree it's degree and of course society and culture give us big advantages language gives us big advantages can i just push on some of this yeah good because i want to back up sonia so i believe animals and children learn from experience and do incredible things learning from experience and when my son was two or three or four i'm like wow this is really interesting that they're my son can learn these things without nobody really teaching them how to do these things but at the same time what i'm trying to do is i'm trying to teach them how to do these things what i'm trying to do is i'm trying to teach them how to do these things what makes human uniquely human to be able to go to outer space build a rocket those are not things that are learned 100 from experience because before you launch the rocket you actually have to abstract thinking through it in a way that is not learned from quote-unquote experience because you don't know if it's going to work or not you have to imagine it how do you we teach a machine to imagine things that were not available before that's probably the thing that we're trying to like push on because that we're not quite understanding that we're absolutely going to agree with you there you have to be able to plan you have to be able to imagine yeah would you say that humans a thousand years ago before they had done all most of the things that we're talking about were they as intelligent if for example someone from that era was exposed to this new culture would they be able to get the same skills and start doing useful things even over the last 10 000 years i don't think the human brain has evolved that much fundamentally the same machine fundamentally the same machine but we've built up 10 000 years of knowledge yes and i get to learn 10 000 years of knowledge by going to school through supervised learning right and i get all that much much faster than trying to learn through experience so right so i think you're like totally right so we we would want our systems to learn from experience and part of their experience would be getting exposed to our culture and then learning from about our culture they should learn from that that's all good but let's talk about when someone goes and does a paradigm shifting thing so everyone gives the example of einstein but i think there are many examples learning is that too like looking at learning thing versus uh programming things so when these paradigm shifts happen i would say it's a human who has accumulated all this knowledge and then from their experience they're building new abstractions they're planning with them and then they're discovering new knowledge and that skill of of coming up with new abstractions and then learning what mars and planning with them that problem is that skill is totally missing in our current systems and you can expose this at the edge of human knowledge but you can also study this problem at the sensory motor screen level so we're not arguing with the with the principle we need to form abstractions we can reason at a high level you guys are coming close to doing that thing that i said we should never do which is argue is prior knowledge important or gaining knowledge is important you know that's what you guys just said you said it's you're still going to have to learn things and you're saying oh i can get things from my culture and from prior knowledge but these should not fight for each other on the exact thing around paradigm shifts how do we create a machine that understands when to shift the paradigm yeah i think through its experience right so it would have to through its own experience it can't rely on human knowledge because we're assuming the humans see one paradigm and we want a different way of looking at things and so through its experience it has to find something that is better maybe it generalizes better and makes better predictions maybe it's better in some other ways but it has to be through its own experience the big challenge that we don't see in our field the ability we don't see in our field yet is the ability to learn a model and then plan with a model we can do the the math things we can do alpha go because the games we know the model we know how the moves work and in math we know what the operators are we you know we know lean will take us from one state of knowledge to the state of the proof to the next state but if we have to learn the models there are no i'm going to say it there's probably maybe a weird example but a counter example but i can see there's no instances of learning the model and then planning with the model in our field at least not with uh like self-discovered abstractions so there are people who say i'm just going to learn a model of what happens in the next second or next millisecond but that's not how our models work our models are more abstract our models are uh quite different so one of the things i like about what you're doing here is you're not just sitting around pontificating or lamenting the state of the world as it is you're very action oriented that's why you started the company so let's let's start talking about that a bit in 2022 rich you laid out a very specific 12-point plan the alberta plan for ai research maybe tell us about that so the alberta plan came about because we did have general ideas but we also needed to convert them into smaller chunks so the 12 steps are the attempt to crystallize particular chunks there's a very important early step step two is to convert them into smaller chunks and so the 12 steps are the attempt to uh and which is continual deep learning and we think that one is like almost the most important because it unlocks everything else if you could do continual deep learning you could then continually update your model of the world and then if you knew how to do the abstraction rights and like the second half of of the the steps are all about how to get the abstractions right so and not only my abstractions right what i mean but i don't mean get the right abstractions because no one can say what the right abstractions are that depends on the worlds you're in your agent would have to learn the correct distractions for whatever world is in and so you know if you maybe those are the two key things you have to find the right abstractions and then you have to do able to continual deep learning i think that a lot of the people in the field realize that we need models we need to plan with them but the abstractions tell us what the model should be conditioned on so what should you what should the model predict what are you going to do and then something is going to happen and more importantly where would that come from so i really like the example of elite athletes if you ask elite athletes about how they do certain things they would have weird niche terminologies for doing very specific things they were like you know i do this thing and they would have a name for it if they communicate sometimes they don't even have a name for it if they're just doing it alone so how did they come up with those abstractions that's in some sense a crucial thing that's missing that the people in the field don't even know about so i think that's something that we need to think about and i think that's something that we need to think about i think it's important to think about and i think it's important to think about and i think it's important to think about and i think it's important to think about what we need to do and what we need to be able to do because if i wanted to do call it naive updating of weights based on user interaction i can do that today right and so what in your opinion is the biggest thing that we're missing to kind of get to continual deep learning yeah so it's absolutely an algorithmic gap you you can do the naive thing but then you'll see all sorts of problems so for example if you say i'm going to take one sample and then i'm going to update my whole model with that one sample you will run into this problem that now all of the previous knowledge in the model it's impacted negatively and the way currently we get around this is exactly what cursor does they don't use one example they use a large batch coming from a lot of users so in use cases where you can have that you can do continual learning but most use cases you don't have that most use cases you have a single stream of data and then if you apply it to the naive thing it just completely destroys your prior knowledge in a very um destructive way away catastrophic forgetting yeah that is so yeah but it's totally curable you have to have the right algorithm yeah exactly what is well you know first you need to do what we call step size optimization and it means every weight in your network has to have a separate step size so some will move fast some will move slow and you will we you will have to meta learn these step sizes for each weight most of your network will be have have weights that have tiny step sizes so then when you train on a new example they don't get destroyed it happens just to the right places and then secondly you have to use some form of generate and test which is in feature space so you come up with new features or new units and and without following gradients because gradients is a very slow process you only move into direction if you know it's the helpful one and that's always going to be very slow and doesn't give you a path to grow more and more complex and and to have sustained learnings you need to have something that just proposes a bunch of new units and and and and then goes from there i guess so there is a specific thing i can say that make at least concrete which is to say we have this algorithm called continual backprop we used published in in nature and we're going to talk a little bit more about that in a little bit more detail but i'm going to talk a little bit more about that in a little bit more detail but i'm going to talk a little bit more about that in a little bit more detail but i'm going to talk a little bit more about that in a little bit more detail but every but you also plant new seeds of units that are newly initialized with random weights backprop only has random weights at the beginning of time and then as you go on all that randomness all that variety from the randomness gets used up and with continual backprop you keep injecting a bit of randomness a bit of generate and test a bit of generate and then the the operation of backprop is the tester so you can see that the backprop is the tester so you can see that the backprop is you need you need that and and if you put those together really well i think you'll have a new generation of massively superior uh continual deep learning and that's what we hope to do in the next couple years wonderful do you think that these algorithms can be applied to the current state of affairs with people scaling llms and trying to get them to do continual learning without catastrophic forgetting yeah absolutely i think it's uh so i don't think that you could take an existing model and say i'm going to just start updating it with these algorithms because these algorithms meta learn how to learn so really you have to say i'm going to learn from scratch so let's say learn a new foundation model but i'm going to learn with this these new algorithms these new algorithms in addition to learning the knowledge they're also going to learn how to learn future things so they're learning two things at the same time and then then i think you would be able to learn new things without catastrophic forgetting is that the most radical thing and that you're trying to do in your company in terms of from the current state of affairs to try to do these two things at the same time most radical thing is i think that's this goes back to i'm not crazy everyone else is crazy yeah that's perhaps not totally radical there was a point in like 2016 to 2018 where a lot of people were exploring these ideas uh quite a bit they were doing it in a much more limited setting so they would say we have a distribution of problems and then in this specific case we'll do it whereas we want to do it from a single stream of experience so our method should be more generally applicable so i think many people have explored this no one has explored this in the general setting where the resulting algorithm would be applicable everywhere so what would be the most radical thing that your company your new company is trying to do that other people are not doing what's the most ambitious thing remember i don't think i'm weird so i don't want to say all right let's what's the most ambitious ambitious thing i think is to try to have the full spectrum of knowledge both about the tiny things and about the big things you know like thinking about how you take an airplane from one city to another that's a a very big you know it's it's more it's like your your uh space flight example but it's just kind of more commonsensical to think about because we all many of us take airplanes and all of us use abstractions and all kinds of things and we all our life and even even the squirrels use abstractions so to have that spectrum of knowledge from the small to the big and to treat it in a uniform way and to be able to help have it self uh maintaining you know the big question is always you have your knowledge-based system and what keeps the knowledge in it correct well what keeps the knowledge correct in a large language model is well people did a lot of post-training and and they they made it sure it was correct and then they freeze it after that so that's what keeps it correct but really our minds we are we're always changing things and yet something keeps it organized and coherent and and and settling back into a good place rather than drifting off into crazy land that is an i think our our biggest ambition to have a mind that is self-consistent and and can keep training itself and make it coherent i love that can i ask it almost seems that it's it's such a ambitious vision and the idea that all these things can be unified into a single mind is so ambitious it's within reach i think it's within reach it's here it's 2026 yeah and our computers are so fast you know is it is it so ambitious that it's out of reach uh or do we have already inklings of how all the steps can be done and we i i think we have a vision and inklings um i don't think it's i don't think it's inappropriate your vision involves a trillion parameter model with 20 watts that seems pretty ambitious that is ambitious uh in some sense with current technology i would say it's also impossible like just storing a trillion parameters in memory would probably use more than 20 watts of energy with current memory technologies but we are really thinking of okay things are getting better computation is getting cheaper it is getting more energy efficient so where would be would be in five to ten years and i think five to ten years with the right algorithms and we can totally be in a world where this would be possible so five to ten years is two orders of magnitude of moore's law it's a standard improvement if we double every 18 months 10 years would be two orders of magnitude and so for kerm's statement to be plausible then today you should be able to do it for for what 20 watts two orders of magnitude to 2 000 watts if you can do it with 2 000 watts today yeah then in 10 years you'll be able to do it for 20 watts i think you can do it for 2 000 watts you have i think lots of people at research labs that have access to way more than that yeah i think we can can be more efficient than that even now with the right algorithm if we can be more efficient than that then why aren't we it's not like people just want to spend all their money and spend all our money sometimes it seems like sometimes it seems like they want to yeah i think that's how they show their their real men by using lots of energy at least when i look at different research groups i don't even see anyone believing in it's possible and i think if you don't believe in it you're just not going to work on the technical problems and work on the technical problems and work on the technical work through them is it that it's not possible or it's that there's so much waste in the system like one which one is it like is is there someone who knows how to do it efficiently yeah and then there's 10 times the number of people in the same lab doing all these other things and so nine out of 10 people are wasting in some sense i the way i think about it is that we are stuck in a local you know so if we want to move towards this new kind of algorithms it is almost impossible that things will not get worse before they get better so when we start exploring these new directions you're not going to get state-of-the-art performance from day one but it is because it is a different paradigm but that path leads to similar performance at a higher energy and these big labs they are so locked into a product that they it is not possible for them to pursue a path where things get worse first because their current paradigm allows them to keep scaling and this new paradigm they have to take a bet and they and they have to figure out some some technical things that are difficult that we have thought about it for many years we know people who have thought about these things for many years and when i talk to them it makes sense that it's doable but you need to think about those challenges for a long period of time so if everything goes right with oak what happens with the company what do you what kind of company are you building uh if everything goes right we uh implement the architecture we can have genuine uh continual learning and we can form abstractions so that we can do planning and reasoning and and we so sort of like true intelligence and then you know it's hard to imagine just exactly what will happen by then but i think humans will become irrelevant i i don't think that's true at all i i we don't either i think the world becomes exciting and even more exciting and interesting and at the same time it's going to be exciting for humans but in particular i think there are the you have to wonder about the large language models they might be at risk. When this eventually happens, you know, I'm sure they'll get a good run. They've already had a good run. You know, they've been very successful. And let me say just for clarity that large language models are an amazing scientific breakthrough, a breakthrough in the skillful use of language by neural networks. It's totally unanticipated. You know, it was always a holdout for symbolic methods in language, and they have totally changed how that's thought about now. Yeah, so it's a big breakthrough. So it's frustrating to me that we have to, you know, just celebrate that we've made this great progress in the subset of the problem of AI and enjoy that. Instead, it has to pretend to be all of AI. All of intelligence is not fluid, capable use of language. There's so much more. It's an important. You know, it's like 20% or a quarter of intelligence. There's more. We're not done. Yeah. If everything goes right, are you imagining that there's a single mind that can do everything from learn how to swing from tree branches to make a spaceship to, you know, all these various things we've talked about today? Is it a single mind? And is it a single set of weights that can do all these things? Or is it. It's a single design. Okay. There'll be many different. There'll be many minds. Okay. So it's a single design that reacts to different environments. And different versions of that mind will learn different things because they have different experience. This sort of goes back to the big world hypothesis that there are infinitely many things to learn. So one system cannot learn infinitely many things. And like, I think Rich already mentioned this, but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot learn infinitely many things. And like, I think Rich already mentioned this, but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot model each other because they're equally complex. So a single system would never be able to get to a point where it can learn everything. It would always be multiple systems that are learning from their own experience. You guys are hiring? What kind of people are you looking for? We are hiring. The initial team, most of it, we already have in our mind. So these are people who have thought about these ideas for a long time. And we are going to take a slightly different approach. Because this is a different paradigm. It doesn't make sense to become large very quickly, because in some sense, everyone we hire has to come to see what we see. And not everyone sees that. So we're going to start small, slowly grow to maybe a handful or two or three, and then go from there. We want to be super aligned. So that we can be very productive working together and scaling the progress. Absolutely. Very, very cool. Wonderful. I loved this conversation. Thank you for taking the time to share what you're up to. You are an unusually deep thinker about where reinforcement learning and algorithmic design will go. And it was a true pleasure to get to explore it together with you today. So thank you. Thank you very much. Thank you. It's our pleasure. Thank you. Thank you. Thank you. Thank you.

Podcast Summary

Key Points:

  1. Rich Sutton believes that reinforcement learning is fundamentally about continual learning, and that the field’s current focus on scaling with data rather than computation is misguided—what truly matters in the long run is learning methods that scale with computational power.
  2. The core of "The Bitter Lesson" is that AI systems should prioritize learning algorithms over human knowledge, enabling systems to scale with computation instead of relying on vast datasets or pre-encoded human expertise, though both are valuable in different contexts.
  3. Oak Lab aims to develop a new generation of AI systems capable of continual deep learning through self-improving models, abstractions, and experience-based learning—moving beyond current LLMs that freeze after training and fail to learn from user interaction.

Summary:

Rich Sutton, a pioneer in reinforcement learning, argues that the AI field has become "weird" in its focus on human knowledge and data scaling, rather than on continual, experience-driven learning that naturally evolves over time. He asserts that true intelligence lies in systems that learn from their environment continuously, not in static models trained once and then frozen. His seminal work, "The Bitter Lesson," emphasizes that long-term progress depends on algorithms that scale with computation, not on human-curated knowledge.

This view is challenged by the rise of large language models (LLMs), which he sees as both a positive—enabling massive scaling through data—and a negative—eventually hitting limits due to finite human knowledge. Synthetic data generation, while promising, remains bottlenecked by human expertise and fails to capture the infinite complexity of the real world. The "big world hypothesis" posits that the world is infinitely complex, making it impossible for any single system to learn everything, thus requiring multiple agents learning from their own experiences.

Oak Lab is founded on this vision, aiming to develop self-improving, continual learners through novel algorithms like continual backpropagation and generate-and-test mechanisms. These systems would learn abstractions from experience, plan with internal models, and avoid catastrophic forgetting. Unlike current LLMs that stop learning after training, these agents would evolve over time.

The team prioritizes small, focused growth and deep alignment, aiming to build a scalable, self-consistent AI design that can adapt to diverse tasks—from physical robotics to abstract reasoning—without depending on human input. While ambitious, Sutton believes such systems are within reach in the next five to ten years due to advances in computation and efficiency. He stresses that AI must evolve beyond language mastery to achieve true fluid intelligence, where learning, abstraction, and self-improvement are unified in a single, adaptive framework.

FAQs

The essence of the Bitter Lesson is that AI systems should focus on methods that scale with computation, like search and learning, rather than relying on human knowledge or scaling with data. It emphasizes that in the long run, the ability to continuously learn is more important than pre-loaded human knowledge.

LLMs are both a positive and negative example of the Bitter Lesson. They enabled massive scaling with computation by consuming vast amounts of internet data, but eventually hit limits due to finite data and lack of real-world experience, showing that over-reliance on human knowledge can hinder true progress.

The Big World Hypothesis proposes that the world is infinitely complex and contains infinitely many things to learn. It suggests that true intelligent agents should learn from their own experience, not synthetic or curated data, because no simulation can capture the full complexity of reality.

Human knowledge is valuable but not sufficient. The long-term success of AI depends on continuous learning from experience. The field often overemphasizes human prior knowledge, while neglecting the importance of algorithmic progress and lifelong learning.

Continual deep learning refers to AI systems that continuously update their models based on real-world experience without catastrophic forgetting. It requires specialized algorithms, such as step-size optimization and generate-and-test methods, to maintain prior knowledge while learning new information.

Current LLM assistants do not learn during interaction—only memorize context. Their weights remain fixed, meaning they don’t adapt or evolve. True learning requires updating model weights over time based on experience, which current systems lack.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.