Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again
53m 43s
Rich Sutton, a pioneer in reinforcement learning, argues that the AI field has become "weird" in its focus on human knowledge and data scaling, rather than on continual, experience-driven learning that naturally evolves over time. He asserts that true intelligence lies in systems that learn from their environment continuously, not in static models trained once and then frozen. His seminal work, "The Bitter Lesson," emphasizes that long-term progress depends on algorithms that scale with computation, not on human-curated knowledge. This view is challenged by the rise of large language models (LLMs), which he sees as both a positive—enabling massive scaling through data—and a negative—eventually hitting limits due to finite human knowledge. Synthetic data generation, while promising, remains bottlenecked by human expertise and fails to capture the infinite complexity of the real world. The "big world hypothesis" posits that the world is infinitely complex, making it impossible for any single system to learn everything, thus requiring multiple agents learning from their own experiences. Oak Lab is founded on this vision, aiming to develop self-improving, continual learners through novel algorithms like continual backpropagation and generate-and-test mechanisms. These systems would learn abstractions from experience, plan with internal models, and avoid catastrophic forgetting. Unlike current LLMs that stop learning after training, these agents would evolve over time. The team prioritizes small, focused growth and deep alignment, aiming to build a scalable, self-consistent AI design that can adapt to diverse tasks—from physical robotics to abstract reasoning—without depending on human input. While ambitious, Sutton believes such systems are within reach in the next five to ten years due to advances in computation and efficiency. He stresses that AI must evolve beyond language mastery to achieve true fluid intelligence, where learning, abstraction, and self-improvement are unified in a single, adaptive framework.
People think I have a radical point of view sometimes. They start questions saying how
what I'm thinking is so different from everyone else. But I don't see it that way at all.
I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking
a bit weird. And I mean that like, you know, it's just the recent times people are thinking weird.
Before there was all this AI craziness, you talk about, you wouldn't have to say continual
learning because it wouldn't make any sense to talk about learning that wasn't continual.
All learning is continual. We always act and we learn. That's just the normal way of thinking.
I'm not weird. The field is weird. The field, they need to call it continual learning. It's just learning.
We are honored to have the great Rich Sutton with us here today.
Rich, you invented reinforcement learning. You wrote the seminal textbook. You had the
key students in the field, folks like Dave Silver. You wrote the essay, The Bitter Lesson,
that I believe is the Bible of the field. And you have just been one of the greats
in propelling the field forward. So thank you for taking the time to join us today.
Rich is joined by Kuram Javed, his co-founder and former student from the University of Michigan.
He's a professor at the University of Michigan. The two of you have set off to found Oak Lab. I'm very excited to talk to you about that today.
So for today's session, we're going to start talking about The Bitter Lesson,
the state of the world as we know it today, whether LLMs will get us there or not.
And then we're going to transition to start talking about your research agenda and your
plan for Oak. Rich, maybe take us back. I was going to start with The Bitter Lesson,
but I actually want to start earlier than that. Decades ago, you decided to dedicate your
career to reinforcement learning, to deep reinforcement learning in particular, and
you established the University of Alberta as a bastion of that back when I think the field was
very much in its infancy. What gave you the conviction to do that? What else are you going to
do? We're trying to figure out the mind and learning is a central part of the mind and having
a goal is a central part of the mind, central part of intelligence. Yeah. So I was just doubling
down on what I was always thinking.
Did people think you were crazy at the time?
It was a winter. It was an AI winter.
What year was this?
It was in 2003. And it's kind of crazy, actually, the truth, because I was really sick. I was
actually dying of cancer in 2003. But I wasn't quite dead. I've been trying for a number of years
and I wasn't dead. I was in another remission.
And so I said, well, I'm not dying. I haven't succeeded in dying. So I might as well, you know,
it's going on long enough. I might as well just try to get another job. And so I went to Alberta
and started teaching there. And then in the end, I didn't die. It's kind of amazing. It's like that.
I'm joking about it now, but it was quite serious. And it's an even more poignant question. Why did I
continue to work in Alberta? Why did I continue to work in
Canada? Why did I continue to work on this research stuff when I was, you know, I only had a few
months? I would always keep reminded what I think it's Benjamin Franklin is supposed to have said
that, you know, if you ever wonder why someone is doing something, it's almost always one of two
things. It's either habit or vanity. Okay. So I think it was probably true. Maybe it was a habit
to just kept doing what I always was doing, or maybe it was a vanity. I don't know. I think it
was more like habit because I was dying.
Wow. Divine intervention.
Yeah. It's always been easy for me to be very determined. And I'm going to go even longer on
this answer. They start questions saying
what I'm thinking is so different from everyone else. If you
look back, what people thought about the mind for, you know, even just a decade, you'll find
the kind of thoughts that, you know, learning is important. You've got to have a goal.
And, you know, perception is important. We are low level beings. We are generating actions and
perceiving data at a fast speed. And yet we have to think at higher levels. And,
you know, go back a few, before there was all this AI craziness, you talk about, you wouldn't
have to say continual learning because it wouldn't make any sense to talk about learning that wasn't
continual. You know, it's not a special phase. I'm not weird. The field is weird. The field they need
to call it continual learning. It's just learning. I'm not weird. Everybody else is. That's a good,
though, to live by. We're going to have to send out a next post about that.
We're very happy that you lived on. The field is happy that you lived on. And thank you for
pushing the frontier of AI. I'm really happy. Thank you for pushing the frontier of AI. I'm
sure really happy. And thank you for all of that. And you've been able to sort of educate a lot of
students who pushed the frontier as well. How did you pick them? How did you, over the last 20,
30 years? Oh, well, you are giving me opportunities to be humble.
I like to be humble and point out how all these great decisions are just happen. And that's the
way I feel about students. I don't feel that I choose them very well. Sometimes I'm lucky,
sometimes I'm unlucky. I don't feel I'm particularly good at picking my students.
I'm looking at Kerm now. I think sometimes you end up with really great ones.
David Silver picked me. How is it that I got you, Kerm?
I finished my master's, not with you. And I was planning to join industry. And then
we were collaborating on a project, which also just started organically. There was
something I worked on that Rich was in a meeting. Then they mentioned that I worked on it. So I got
pulled into it. We started collaborating. It went really well. I felt so happy with that
collaboration. Rich also felt really good about it. And then six months down the road,
we had made some progress. And it just made sense to convert that into
a thesis proposal. So at no point did I apply. At no point did I ask, should you be my PhD advisor?
We worked together. Then we decided this would be a pretty good thesis. And then after that,
I applied for the PhD. Life works in unexpected ways.
Take us to 2019. You wrote The Bitter Lesson, which has become the modern tome.
2019 was a funny time to be writing that piece because ImageNet was 2009. AlphaGo was 2015.
What caused you in 2019 to reflect and to write that? Because it was before the current
kind of scaling paradigm around large language models had taken off, but it was after deep
learning had really proven itself. Well, it was a long time coming. As The Bitter Lesson
expresses, it's something that you can observe for a long time, for many decades. And it's
definitely at least as much due to the round of symbolic AI, which I lived through. It's all about
not getting distracted by trying to put in your human knowledge and just paying attention to what
the problem needs and how you can scale with computation. I know I wrote versions of it at
least a year before. And I gave talks. I gave a talk a year before. And it wasn't a particular
response to the moment. It was a particular response to my long experience. Different people
trying to think in different
ways about how you can make smart systems. What is the essence of The Bitter Lesson?
It may be the phrase that I hear used the most in my meetings these days is Bitter Lesson
Pilled. Is it not Bitter Lesson Pilled? I would imagine given the popularity of the phrase,
it's probably been tortured and misused in different ways that you didn't originally
intend it. So what is the essence of it? And where do you think people go wrong in their
attempt to understand it?
Yeah, you're making me think about X now. And I recently made a post where I tried to do the
Bitter Lesson in 26 words. It goes something like, don't be distracted by human knowledge
as AI traditionally has been many times. Instead, focus on learning methods that will scale with
computation like search and like learning. So it's really
all about focusing on algorithms and improvements. It's not saying you don't need
fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with
computation. Rather than scaling with data.
Rather than scaling with human input. Yeah. And then the question, if I can anticipate,
yeah, what about large language models?
Are they consistent or inconsistent with your essay?
Yeah. And I've thought about this and
I think there's another X post about it.
But the conclusion is that it's both a positive example and a negative example of the bitter lesson.
First, large language models enabled enormous scaling with computation.
And you could just drink in the Internet and scale so much.
So it was a way of getting a much more capable system just by methods of scale.
Then after that, as you go on further, it eventually gets limited by that information.
The Internet is finite and it's hard to get more examples.
And the world is big and the world is massively bigger than everything we stored on the Internet.
And so in the end, it seems like it could be, I guess that would be a positive example of when we relied too much on human knowledge.
And it eventually holds us back.
Can I just push on this a little bit?
Yeah.
It seems like a lot of what the Foundation Model Labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human Internet.
Is synthetic data generation kind of as part of this LLM scaling paradigm, is that a general method that leverages computation?
No.
Why?
That's just a big mistake.
Well, it's such a big, it's such a, maybe it's the next, the next big lesson.
It's been floating around Alberta for five or 10 years.
Okay.
And we call it the big world perspective or the big world hypothesis.
Kuram, who eventually wrote it up as a paper, there's a little paper called the big world hypothesis.
The big world is that the world is infinitely big.
There are infinitely many things.
There are infinitely many things to learn, and you can have people generating these synthetic data sets, but they would always be more things to learn.
And because of that, if you could just learn from experience, you could remove the humans from the loop, then you would have systems that can do everything because, you know, the world is big.
There are many tasks that we want them to do, and they would be able to do anything by learning from their own experience.
Going back to the synthetic data question, too, who decides what's a good synthetic data and what's the bad synthetic data?
Yeah.
Because I can write a program that can output a lot of synthetic data, which would hurt programs.
Right now, I would say humans decide, and that's the bottleneck where, okay, you can have humans deciding how to generate these data sets, but you need human experts who know what's a good data set and what's a bad data set for that approach to scale.
So it is bottlenecked by humans.
Doesn't my loss curve decide, like, how much better did I get with this data set versus that data set?
Right.
But if all the engineers, OpenAI, Anthropic, or all the big new labs, the engineers went on vacation.
Who would generate the synthetic data?
That's the question.
It doesn't come from agents' experience.
It's not something that the agent is generating itself.
Some human has to decide what is the right synthetic data to generate.
And that requires human expertise.
So, for example, if you want a system to do something very challenging from a physics point of view, maybe you want a drone that flies with echolocation, like a bat, for example, what's the right synthetic data for that?
I think you would need to hire domain experts.
To go figure out what is the right data and generate it.
And then maybe you would be able to learn from that.
But the domain expert has to exist first.
So we are bottlenecked by human expertise at that point.
But you can have infinite synthetic worlds.
The existing world is finite.
But let's go back to the echolocation thing, right?
That's what I want.
I want a drone that can localize itself and move with echolocation.
That's my goal.
The robot that is a robot is generating its own experience.
So it could totally learn from its own experience, but it wouldn't be able to.
It doesn't matter how much synthetic data you generate.
It doesn't matter if you generate synthetic data that captures 50 different universes.
It will not allow you to do that task without humans figuring it out first.
At first, just say the synthetic data is wrong.
I mean, it won't be correct.
It'll be a synthetic world.
It won't be the real world.
And it will matter.
The world is incredibly complex.
If you write a little program, because there's going to be a little program that will generate the synthetic data,
it'll be a small world.
So, for example, what's important to me is what's going on in your mind right now.
Okay?
And you're saying, why don't I get some synthetic data to tell me what's going on in other people's minds?
No, there's no way we can have synthetic data for other people's minds.
And other people's minds matter to us.
You know, like. I talked to you guys about investing today, so I care what's going on in your minds.
But how can I get synthetic data on such a thing?
Really, you can't even get synthetic data on anything.
You can't get synthetic data on how the drone is going to interact with its environment,
in the physical world, in the friction, and where in the motors of this robot.
The world is infinitely complex.
And any simulation of it is, like, microscopic.
The big world hypothesis, let's say what it is,
is that the world is massively more complex than your mind, than any agent.
And this is obvious because the world contains many other agents.
So, because the world is massively complex,
there's no way you can do anything that might claim to be optimal or perfect.
You're going to be imperfect, and you have to have approximations.
And those approximations will be severe.
And so, because of that, that is ultimately the reason why we have to continue learning.
If you want to think of it as a reason, we have to continue learning because we'll encounter
some particular part of this immense world, and we'll have to learn an approximation that's
tuned to the part of the world we're in, not to all the other parts that we're not in.
Yeah.
I'm going to push on this one more time.
And sorry, I'm being argumentative for the sake of being argumentative.
I'm trying to understand.
My understanding is that the newest cohort of self-driving car companies,
many of them were primarily trained in sim.
And then they have to do some sort of post-training, I guess,
to make sure they work in the real world.
But that it's been a very effective pipeline.
Yeah.
So, I think the important question to ask here is,
how many engineers were involved in building that simulation?
And are we ready to say that the only problem worth solving are those where we can hire
a large team of engineers?
So, I think the only problem we have is that we can hire a large team of engineers to first
make a simulation.
And I'm sure they had to do multiple iterations where they made the simulation, they learned
in it, they realized there was a sim-to-real gap that was not acceptable, then they fixed
it.
So, there is this human in the loop fixing the simulation, like they're getting feedback
from the real world, humans, and then they're fixing the simulation.
Why can't we just remove the human and let the agent do it itself?
And then when it actually drives, again, something unexpected will happen.
And that's when you actually want to learn from experience.
Yeah.
Because there's just so much more data that's going to come from experience than there possibly
can be from humans curating, creating data.
Yeah.
And I think there is obviously value in learning from simulation.
And there is a way of doing it.
The agent can learn a model from its own experience.
And when the agent learns it, it's much better because if the model is incorrect, it can
fix it by continuing learning.
If the humans are making a simulator, then the model only gets updated when the humans
figure out that something's wrong.
So, yes, planning is important.
The agents should learn from simulators, but simulators, they make themselves.
Okay.
I want to move to another part of the bitter lesson, removing human knowledge.
From your essay, "Seeking an improvement that makes a difference in the shorter term,
researchers seek to leverage their human knowledge of the domain.
But the only thing that matters in the long run is the leveraging of computation."
And if I, you know, your former student, Dave, with AlphaGo and AlphaZero, for me, that was
an example of a triumph.
A triumph of removing human priors.
Did that result surprise you?
I guess, why or why not?
Of course, it made me very happy and made me, you know, feel vindicated.
You know, it could have gone either way.
It wasn't that because prior knowledge can help.
You know, there's nothing wrong with prior knowledge.
You know, and I say this right at the very beginning, right at the very beginning of
the bitter lesson, I say, there's no reason why there has to be a conflict between prior
knowledge and then learning knowledge.
You know, you can put some prior in there and then start learning.
There's no reason in principle why these have to be opposed.
In fact, they're all about knowledge.
You know, life is gaining knowledge and having knowledge.
And why are these, how somehow, you know, nature and nurture became enemies.
But really, you know, prior learning is what you already, and then you learn more.
And it's, they should be friends.
But as I say at the beginning of the bitter lesson, in practice, they have been enemies.
In practice, people who had an affection for existing human knowledge ended up, you
know, wanting that to win.
And so they wanted to minimize or dismiss learning.
And so now I'm sure your sense of me is that I'm someone who loves learning and wants to
dismiss prior knowledge.
But, you know, I'm really someone who's interested in the mind.
And the mind is, you have prior knowledge and then you get more.
And then once you've gotten more, then that becomes your prior knowledge as you get more
and more and more.
And then.
these two things work together. I end up appearing to be someone who's interested in learning
primarily because all the rest of the world is talking about all you need is enough knowledge.
You don't need to learn. You know, large language models are, oh, we're going to put all this
knowledge into the system. And the large language model will not learn when it runs. You know,
it's talking to people, it's interacting. It is absolutely, the weights never change.
So, you know, I am not the weird one. It's you guys that are the weird one that think that that's
possible that you could possibly, you know, they claim they can make a PhD level experience and
expertise out of something that doesn't learn at all anymore. You know, so, you know, I'm not the
weird one. So your recommendation is let the algorithms run for a much, much longer period
of time before feeding it. Continually learn. Before you feed it data or prior data.
So drip or drip the prior data along the way. Prior knowledge.
So both are important. But in the long run, you've got to gain and structure the gaining of new
knowledge. And that's what all that matters in the long run. And as you are doing this, yeah,
there'll be some that you got previously. Like how it work if we, you know, look into the future
when we have intelligent robots.
Will we have them all learn from scratch or will we like copy them and ask them to keep learning
from wherever they are? I mean, they'll be digital and it'll be easy to copy them. And so instead of
having like this huge thing where we're spending zillions of dollars to retrain them from the
internet, we'll just copy the agent and keep learning from there. And so in some sense,
the prior knowledge will be,
should be dismissed because you're just going to copy it from the previous robot.
So why don't you just describe for us what you think a machine or a computer that
learns from experience looks like?
Well, it could look like a robot. It could, but it also could live entirely on the internet.
You could like, for example, routing of packets through the internet and do that in a way that's
sensitive to experience and becomes better over time. Or you can interact, be the user interface
that's interacting with people, like on your phone.
Or on your computer and it becomes better over time. You know, like an intelligent assistant,
you know, has to become better over time, has to know what you want.
Would your contention be that the current paradigm of, you know, the popular LLM-based assistants,
would your contention be that these are not experiential learners or continual learners?
And if so, what is the fundamental gap?
Are you serious?
I mean, obviously. They learn memories about me. They're, you know, they're doing some in-context learning.
Their weights never change.
And by the way, is a small number of the weights changing sufficient, or do you need
all the weights to be changing?
Well, so think of all the structuring and generation of new concepts that went into
creating the large language models. All that is the weight learning. And you want to be able to
continue doing that. You don't want that. You don't want that. You don't want that. You want
you don't want that to happen just once.
Is another way of saying is we do too much pre-training and post-training before we
launch the models, and they don't learn after that.
The only point that the big disagreement is we don't let them learn after that.
Yeah. We don't let them learn after that.
We can do as much pre-training as we want. That's okay. Post-training is fine.
But then when I'm using the model, I cannot. It stops learning. You can give it more context.
You can change the state of the model by giving it more context. And so it has already learned that,
if the state is different, if the state says something new, then it will use that to make
the next prediction. But the model is not learning.
Cursors, tab, autocomplete model. It does get updated based on. Those models. Those weights change.
Those are two examples. Cursors, tab, and I think the composer,
they were also updating. Those are two examples of continual learning. But it can be much better.
So the way they do it, as far as I understand, is a lot of people are using tab. They collect
all this data. So coming from millions of people using tab, they collect all this data. So coming
from millions of users or thousands of users, and then they do one update of the policy from
this batch data. So this could work, but imagine I want to teach this model something specific.
I don't want to fight with 100,000 other people about what they want to teach their models.
I want to teach my model something very specific, and I want to do it
to my version of the model. I don't care about the shared knowledge that the model has coming
from other people. And so it's a very inefficient way of doing it. It seems like the way that this
is currently done is that there's fundamental skills, maybe, that are learned in the weights
that are common to everybody. And then there's personalization that happens in the form of
context, right? Is that not the right mental model for how learning should work? Should all the
contexts live in the weights themselves? So context can be in the state, too. It could
be both. But you still need to be able to update the weights. So if I give you an example, some
really good use studies are with HIPAA. I'm going to talk about HIPAA, and I'm going to talk about
how HIPAA works. HIPAA works for people with disabilities. When a human goes through something
that changes their mind or some sensors, you can see them adapt. So for example, we have
proprioception. We have internal sensors that tell us how the body's positioned, and we use this for
walking. There are cases where people lose this ability completely, and then they can't walk at
all, because that is literally the foundation of their walking policies. It is ingrained in the
brain. But then over the course of two, three years, they can learn to walk again by looking
at their feet. So visual feedback. I'm going to talk about how HIPAA works, and then I'm going to
talk about visual feedback. So brain is insanely plastic in the sense that it can learn a lot of
things, something that has been true for 20 years. When it stops being true, it can go and update
that and get rid of that. And that is the capability I think that's extremely useful
we would want in our systems. What is there for us to learn from how human babies or animals
learn? And how much inspiration do you take from that? Well, we take a lot of inspiration. We don't
take it as a requirement that the AI has to behave like the natural system, like babies or people or
animals, but it's a source of inspiration. Inspiration, but not constraint from animal
learning. Consistent with the bitter lesson. Where do you think we should most seek to draw inspiration
from the way that biological learning works that is not present in today's systems? I feel like I'm
just giving opinions now, but they're just obvious opinions. So I think it's apparent that no animal
learns by supervised learning because we don't get examples of how our muscles should twitch,
and that's our output. But all of school is supervised learning.
I know, absolutely not. But even if it was, school is like a tiny fraction of what we learn. Like we
learn to see, and we learn to walk, and we learn how the world works. But even in school, no one
tells us how we should twitch our muscles. The knowledge skills I acquire were from
supervised learning in school. So I don't want to say that learning from others, transmission from
others is not important. It's extremely important, and language is extremely important. But what are
we, what are we missing? You know, there is, there is no supervised learning. There's no targets that
are given to us. You know, you, you hear the right answer is, you know, where, where is, what's the
capital of France? And we know the answer is Paris. Okay. But no one tells me how I should pronounce
Paris. You say the answer is Paris, and I listen to you, and I hear your words, and you know, I will
make some other, uh, muscle motions to produce the answer Paris. It's not literally supervised
learning. Um, anyway, yeah. So I, I think it's really true. I mean, well, anyway, the first thing
is the school is, is irrelevant. Like, you know, squirrels don't go to school and, and, and learn
that. Animals don't learn that way. It's, and school is a very special thing that even we didn't
have up until, you know, I don't know, a few hundred years ago. But it's not part of, it's not
part of the essence of intelligence. And it's a distraction to think of that as your primary
example of learning is this thing, which we didn't do as animals. I wish you had been around to tell
my parents that. They forced me to have, get, go to school and deal with all the structure.
The thing is like squirrels are wonderful at jumping off trees, but squirrels can't prove
math theorems. And if I want to learn how to prove a math theorem, I go to school.
Yeah. Uh, they also don't have, uh, DVDs and iPods. You know, there are a lot of things,
they can do things that we can't do. Um, but math theorems, uh, yeah. And they don't play chess,
you know, sort of like more of X paradox. They're, uh, they're these advanced things that we think of
as really intelligent, but, uh, they're sort of easy for computers to do as opposed to all these
regular things that are hard, like moving and seeing with attention and everything.
Um, I think supervised learning is a, is a good thing. You know, just mentions,
I like to think, look for obvious things. No one tells us how to twitch our muscles by
giving us examples because they couldn't possibly because we have had to twitch our muscles we've
had to figure that out yeah and their answer would be wrong right so if i moved my mouth and my tongue
and my vocal cords exactly the same way that rich does to pronounce paris i'm sure a very different
sound would come out so in some sense rich or no one knows the right way of producing a sound with
my body only i know that yeah it seems to me that many of the most i guess the most raw like sensory
motor capabilities especially related to movement in the physical world i agree with you that that
seems something that's inherently learned from experience it seems to me though that there are
higher levels of abstraction that bring us closer to you know what makes humans great and much of
that doesn't live in this low level of sensory motor learning does your world model i guess span
sensory motor learning all the way up yeah that's the ambition absolutely
and squirrels by the way can do some enormously abstract things what's the coolest thing a squirrel
can do well it can always get into your bird feeder no matter what obstacles you put in the way
you know it can find new ways to jump and climb and okay and do lots of things calculate
trajectories pretty well animals are pretty good at understanding the physical world
without the mental calculations that we think we are doing when we think about launching ourselves
into space breaking a fall they can do it in real time the right way to prevent injuries okay fair
enough i think it's just a question of degree between and i like to think that animals other
animals are are very close to humans i think it's hubristic to try to emphasize what we do
differently you know how we're different from animals it's better to see the commonalities
and i think we are just a question of degree it's degree and of course society and culture give us
big advantages language gives us big advantages can i just push on some of this yeah good because
i want to back up sonia so i believe animals and children learn from experience and do incredible
things learning from experience and when my son was two or three or four i'm like wow this is
really interesting that they're my son can learn these things without nobody really teaching them
how to do these things but at the same time what i'm trying to do is i'm trying to teach them how
to do these things what i'm trying to do is i'm trying to teach them how to do these things what
makes human uniquely human to be able to go to outer space build a rocket those are not things
that are learned 100 from experience because before you launch the rocket you actually have
to abstract thinking through it in a way that is not learned from quote-unquote experience
because you don't know if it's going to work or not you have to imagine it how do you we teach
a machine to imagine things
that were not available before that's probably the thing that we're trying to like push on
because that we're not quite understanding that we're absolutely going to agree with you there
you have to be able to plan you have to be able to imagine yeah would you say that humans a thousand
years ago before they had done all most of the things that we're talking about were they as
intelligent if for example someone from that era was exposed to this new culture would they be able
to get the same skills and start doing useful things
even over the last 10 000 years i don't think the human brain has evolved that much
fundamentally the same machine fundamentally the same machine but we've built up 10 000 years of
knowledge yes and i get to learn 10 000 years of knowledge by going to school through supervised
learning right and i get all that much much faster than trying to learn through experience
so right so i think you're like totally right so we we would want our systems to learn from
experience and part of their experience would be getting exposed to our culture and then
learning from about our culture they should learn from that that's all good but let's talk about
when someone goes and does a paradigm shifting thing so everyone gives the example of einstein
but i think there are many examples learning is that too like looking at learning thing versus
uh programming things so when these paradigm shifts happen i would say it's a human who has
accumulated all this knowledge and then from their experience they're building new abstractions they're
planning with them and then they're discovering new knowledge and that skill of of coming up
with new abstractions and then learning what mars and planning with them that problem is that skill
is totally missing in our current systems and you can expose this at the edge of human knowledge but
you can also study this problem at the sensory motor screen level so we're not arguing with the
with the principle we need to form abstractions we can reason at a high level you guys are coming
close to doing that thing that i said we should never do which is argue is prior knowledge important
or gaining knowledge is important you know that's what you guys just said you said it's you're still
going to have to learn things and you're saying oh i can get things from my culture and from prior
knowledge but these should not fight for each other on the exact thing around paradigm shifts
how do we create a machine that understands when to shift the paradigm yeah i think through its
experience right so it would have to through its own experience it can't rely on human knowledge
because we're assuming the humans see one paradigm
and we want a different way of looking at things and so through its experience it has to find
something that is better maybe it generalizes better and makes better predictions maybe it's
better in some other ways but it has to be through its own experience the big challenge that we don't
see in our field the ability we don't see in our field yet is the ability to learn a model and then
plan with a model we can do the the math things we can do alpha go because the games we know
the model we know how the moves work and in math we know what the operators are
we you know we know lean will take us from one state of knowledge to the state of the proof to
the next state but if we have to learn the models there are no i'm going to say it there's probably
maybe a weird example but a counter example but i can see there's no instances of learning the
model and then planning with the model in our field at least not with uh like
self-discovered abstractions so there are people who say i'm just going to learn a model of what
happens in the next second or next millisecond but that's not how our models work our models
are more abstract our models are uh quite different so one of the things i like about
what you're doing here is you're not just sitting around pontificating or lamenting
the state of the world as it is you're very action oriented that's why you started the company
so let's let's start talking about that a bit in 2022 rich you laid out a very specific 12-point
plan the alberta plan for ai research maybe tell us about that so the alberta plan came about
because we did have general ideas but we also needed to convert them into smaller chunks
so the 12 steps are the attempt to crystallize particular chunks there's a very important early
step step two is to convert them into smaller chunks and so the 12 steps are the attempt to
uh and which is continual deep learning and we think that one is like almost the most important
because it unlocks everything else if you could do continual deep learning you could then
continually update your model of the world and then if you knew how to do the abstraction rights
and like the second half of of the the steps are all about how to get the abstractions right
so and not only my abstractions right what i mean but i don't mean
get the right abstractions because no one can say what the right abstractions are that depends
on the worlds you're in your agent would have to learn the correct distractions for whatever world
is in and so you know if you maybe those are the two key things you have to find the right
abstractions and then you have to do able to continual deep learning i think that a lot of
the people in the field realize that we need models we need to plan with them but the abstractions
tell us what the model should be conditioned on so what should you what should the model predict
what are you going to do and then something is going to happen and more importantly where would
that come from so i really like the example of elite athletes if you ask elite athletes about
how they do certain things they would have weird niche terminologies for doing very specific things
they were like you know i do this thing and they would have a name for it if they communicate
sometimes they don't even have a name for it if they're just doing it alone so how did they come
up with those abstractions that's in some sense a crucial thing that's missing that the
people in the field don't even know about so i think that's something that we need to think about
and i think that's something that we need to think about i think it's important to think about
and i think it's important to think about and i think it's important to think about
and i think it's important to think about what we need to do and what we need to be able to do
because if i wanted to do call it naive updating of weights based on user interaction i can do that
today right and so what in your opinion is the biggest thing that we're missing to kind of get
to continual deep learning yeah so it's absolutely an algorithmic gap you
you can do the naive thing but then you'll see all sorts of problems so for example if you say
i'm going to take one sample and then i'm going to update my whole model with that one sample
you will run into this problem that now all of the previous knowledge in the model it's impacted
negatively and the way currently we get around this is exactly what cursor does they don't use
one example they use a large batch coming from a lot of users so in use cases where you can have
that you can do continual learning but most use cases you don't have that most use cases you
have a single stream of data and then if you apply it to the naive thing it just
completely destroys your prior knowledge in a very um destructive way
away catastrophic forgetting yeah that is so yeah but it's totally curable you have to have the
right algorithm yeah exactly what is well you know first you need to do what we call step size
optimization and it means every weight in your network has to have a separate step size so
some will move fast some will move slow and you will we you will have to meta learn these step
sizes for each weight most of your network will be have have weights that have tiny step sizes
so then when you train on a new example they don't get destroyed it happens just to the right places
and then secondly you have to use some form of generate and test
which is in feature space so you come up with new features or new units and and without following
gradients because gradients is a very slow process you only move into
direction if you know it's the helpful one and that's always going to be very slow and doesn't
give you a path to grow more and more complex and and to have sustained learnings you need to have
something that just proposes a bunch of new units and and and and then goes from there i guess so
there is a specific thing i can say that make at least concrete which is to say we have this
algorithm called continual backprop we used published in in nature and we're going to
talk a little bit more about that in a little bit more detail but i'm going to talk a little bit more
about that in a little bit more detail but i'm going to talk a little bit more about that in a
little bit more detail but i'm going to talk a little bit more about that in a little bit more detail
but every but you also plant new seeds of units that are newly initialized with random weights
backprop only has random weights at the beginning of time and then as you go on all that randomness
all that variety from the randomness gets used up and with continual backprop you keep injecting a
bit of randomness a bit of generate and test a bit of generate and then the the operation of backprop
is the tester so you can see that the backprop is the tester so you can see that the backprop is
you need you need that and and if you put those together really well i think you'll have a new
generation of massively superior uh continual deep learning and that's what we hope to do in
the next couple years wonderful do you think that these algorithms can be applied to the current
state of affairs with people scaling llms and trying to get them to do continual learning
without catastrophic forgetting yeah absolutely i think it's uh so i don't think that you could
take an existing model and say i'm going to just start updating it with these algorithms because
these algorithms meta learn how to learn so really you have to say i'm going to learn from scratch so
let's say learn a new foundation model but i'm going to learn with this these new algorithms
these new algorithms in addition to learning the knowledge they're also going to learn
how to learn future things so they're learning two things at the same time
and then then i think you would be able to learn new things without catastrophic
forgetting is that the most radical thing and that you're trying to do in your company in terms of
from the current state of affairs to try to do these two things at the same time most radical
thing is i think that's this goes back to i'm not crazy everyone else is crazy yeah that's perhaps
not totally radical there was a point in like 2016 to 2018 where a lot of people were exploring
these ideas uh quite a bit they were doing it in a much more limited setting so they would say
we have a distribution of problems and then in this specific case we'll do it whereas we want
to do it from a single stream of experience so our method should be more generally applicable
so i think many people have explored this no one has explored this in the general setting
where the resulting algorithm would be applicable everywhere so what would be the most radical thing
that your company your new company is trying to do that other people are not doing what's the most
ambitious thing remember i don't think i'm weird so i don't want to say all right let's what's
the most ambitious ambitious thing i think is to try to have the full spectrum of knowledge
both about the tiny things and about the big things you know like thinking about how you take
an airplane from one city to another that's a a very big you know it's it's more it's like your
your uh space flight example but it's just kind of more commonsensical to think about because we
all many of us take airplanes and all of us use abstractions and all kinds of things and we all
our life and even even the squirrels use abstractions so to have that spectrum of
knowledge from the small to the big and to treat it in a uniform way and to be able to help have it
self uh maintaining you know the big question is always you have your knowledge-based system
and what keeps the knowledge in it correct well what keeps the knowledge correct in a large
language model is well people
did a lot of post-training and and they they made it sure it was correct and then they freeze it
after that so that's what keeps it correct but really our minds we are we're always changing
things and yet something keeps it organized and coherent and and and settling back into a good
place rather than drifting off into crazy land that is an i think our our biggest ambition to
have a mind that is self-consistent and and can keep training itself and make it
coherent i love that can i ask it almost seems that it's it's such a ambitious vision and the
idea that all these things can be unified into a single mind is so ambitious it's within reach
i think it's within reach it's here it's 2026 yeah and our computers are so fast
you know is it is it so ambitious that it's out of reach
uh or do we have already inklings of how all the steps can be done and we i i think we have
a vision and inklings um i don't think it's i don't think it's inappropriate your vision
involves a trillion parameter model with 20 watts that seems pretty ambitious that is ambitious uh
in some sense with current technology i would say it's also impossible like just storing a trillion
parameters in memory would probably use more than 20 watts of energy with current memory technologies
but we are really
thinking of okay things are getting better computation is getting cheaper it is getting
more energy efficient so where would be would be in five to ten years and i think five to ten years
with the right algorithms and we can totally be in a world where this would be possible
so five to ten years is two orders of magnitude of moore's law
it's a standard improvement if we double every 18 months 10 years would be two orders of
magnitude and so for kerm's statement to be plausible then today you should be able to do it
for for what 20 watts two orders of magnitude to 2 000 watts if you can do it with 2 000 watts today
yeah then in 10 years you'll be able to do it for 20 watts i think you can do it for 2 000 watts you
have i think lots of people at research labs that have access to way more than that yeah i think we
can can be more efficient than that
even now with the right algorithm if we can be more efficient than that then why aren't we it's
not like people just want to spend all their money and spend all our money sometimes it seems like
sometimes it seems like they want to yeah i think that's how they show their their real men
by using lots of energy at least when i look at different research groups i don't even see anyone
believing in it's possible and i think if you don't believe in it you're just not going to work
on the technical problems and work on the technical problems and work on the technical
work through them is it that it's not possible or it's that there's so much waste in the system
like one which one is it like is is there someone who knows how to do it efficiently yeah and then
there's 10 times the number of people in the same lab doing all these other things and so nine out
of 10 people are wasting in some sense i the way i think about it is that we are stuck in a local
you know so if we want to move towards this new kind of algorithms it is almost impossible that
things will not get worse before they get better so when we start exploring these new directions
you're not going to get state-of-the-art performance from day one but it is because
it is a different paradigm but that path leads to similar performance at a higher energy and
these big labs they are so locked into a product that they it is not possible for them to pursue
a path where things get worse first because their current paradigm allows them to keep scaling
and this new paradigm they have to take a bet and they and they have to figure out some some
technical things that are difficult that we have thought about it for many years we know people who
have thought about these things for many years and when i talk to them it makes sense that it's
doable but you need to think about those challenges for a long period of time so if everything goes
right with oak what happens with the company what do you what kind of company are you building
uh if everything goes right we
uh implement the architecture we can have genuine uh continual learning and we can
form abstractions so that we can do planning and reasoning and and we so sort of like true
intelligence and then you know it's hard to imagine just exactly what will happen by then
but i think humans will become irrelevant i i don't think that's true at all i i we don't either
i think the world becomes exciting and even more exciting and interesting and at
the same time it's going to be exciting for humans but in particular i think there are the you have
to wonder about the large language models they
might be at risk. When this eventually happens, you know, I'm sure they'll get a good run. They've
already had a good run. You know, they've been very successful. And let me say just for clarity
that large language models are an amazing scientific breakthrough, a breakthrough in the
skillful use of language by neural networks. It's totally unanticipated. You know, it was always a
holdout for symbolic methods in language, and they have totally changed how that's thought about now.
Yeah, so it's a big breakthrough. So it's frustrating to me that we have to, you know,
just celebrate that we've made this great progress in the subset of the problem of AI and
enjoy that. Instead, it has to pretend to be all of AI. All of intelligence is not fluid, capable
use of language. There's so much more. It's an important. You know, it's like 20% or a quarter of intelligence. There's more. We're not done.
Yeah. If everything goes right, are you imagining that there's a single mind
that can do everything from learn how to swing from tree branches to make a spaceship to,
you know, all these various things we've talked about today? Is it a single mind? And is it
a single set of weights that can do all these things? Or is it. It's a single design.
Okay.
There'll be many different. There'll be many minds.
Okay. So it's a single design that reacts to different environments.
And different versions of that mind will learn different things because they have different
experience. This sort of goes back to the big world hypothesis that there are infinitely many
things to learn. So one system cannot learn infinitely many things. And like, I think Rich
already mentioned this, but if you have two of these systems, if you have two of the largest
systems in the world, then it is trivial that they cannot learn infinitely many things. And like, I think Rich already mentioned this, but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot model each other
because they're equally complex. So a single system would never be able to get to a point where it can learn
everything. It would always be multiple systems that are learning from their own experience.
You guys are hiring? What kind of people are you looking for?
We are hiring. The initial team, most of it, we already have in our mind. So these are people
who have thought about these ideas for a long time. And we are going to take a slightly
different approach.
Because this is a different paradigm. It doesn't make sense to become large very quickly, because in some sense, everyone we hire has to come to see what we see. And not everyone sees that. So we're going to start small, slowly grow to maybe a handful or two or three, and then go from there.
We want to be super aligned.
So that we can be very productive working together and scaling the progress.
Absolutely.
Very, very cool.
Wonderful. I loved this conversation. Thank you for taking the time to share what you're up to.
You are an unusually deep thinker about where reinforcement learning and algorithmic design
will go. And it was a true pleasure to get to explore it together with you today. So thank you.
Thank you very much.
Thank you.
It's our pleasure.
Thank you.
Thank you.
Thank you.
Thank you.
Podcast Summary
Key Points:
Rich Sutton believes that reinforcement learning is fundamentally about continual learning, and that the field’s current focus on scaling with data rather than computation is misguided—what truly matters in the long run is learning methods that scale with computational power.
The core of "The Bitter Lesson" is that AI systems should prioritize learning algorithms over human knowledge, enabling systems to scale with computation instead of relying on vast datasets or pre-encoded human expertise, though both are valuable in different contexts.
Oak Lab aims to develop a new generation of AI systems capable of continual deep learning through self-improving models, abstractions, and experience-based learning—moving beyond current LLMs that freeze after training and fail to learn from user interaction.
Summary:
Rich Sutton, a pioneer in reinforcement learning, argues that the AI field has become "weird" in its focus on human knowledge and data scaling, rather than on continual, experience-driven learning that naturally evolves over time. He asserts that true intelligence lies in systems that learn from their environment continuously, not in static models trained once and then frozen. His seminal work, "The Bitter Lesson," emphasizes that long-term progress depends on algorithms that scale with computation, not on human-curated knowledge.
This view is challenged by the rise of large language models (LLMs), which he sees as both a positive—enabling massive scaling through data—and a negative—eventually hitting limits due to finite human knowledge. Synthetic data generation, while promising, remains bottlenecked by human expertise and fails to capture the infinite complexity of the real world. The "big world hypothesis" posits that the world is infinitely complex, making it impossible for any single system to learn everything, thus requiring multiple agents learning from their own experiences.
Oak Lab is founded on this vision, aiming to develop self-improving, continual learners through novel algorithms like continual backpropagation and generate-and-test mechanisms. These systems would learn abstractions from experience, plan with internal models, and avoid catastrophic forgetting. Unlike current LLMs that stop learning after training, these agents would evolve over time.
The team prioritizes small, focused growth and deep alignment, aiming to build a scalable, self-consistent AI design that can adapt to diverse tasks—from physical robotics to abstract reasoning—without depending on human input. While ambitious, Sutton believes such systems are within reach in the next five to ten years due to advances in computation and efficiency. He stresses that AI must evolve beyond language mastery to achieve true fluid intelligence, where learning, abstraction, and self-improvement are unified in a single, adaptive framework.
FAQs
The essence of the Bitter Lesson is that AI systems should focus on methods that scale with computation, like search and learning, rather than relying on human knowledge or scaling with data. It emphasizes that in the long run, the ability to continuously learn is more important than pre-loaded human knowledge.
LLMs are both a positive and negative example of the Bitter Lesson. They enabled massive scaling with computation by consuming vast amounts of internet data, but eventually hit limits due to finite data and lack of real-world experience, showing that over-reliance on human knowledge can hinder true progress.
The Big World Hypothesis proposes that the world is infinitely complex and contains infinitely many things to learn. It suggests that true intelligent agents should learn from their own experience, not synthetic or curated data, because no simulation can capture the full complexity of reality.
Human knowledge is valuable but not sufficient. The long-term success of AI depends on continuous learning from experience. The field often overemphasizes human prior knowledge, while neglecting the importance of algorithmic progress and lifelong learning.
Continual deep learning refers to AI systems that continuously update their models based on real-world experience without catastrophic forgetting. It requires specialized algorithms, such as step-size optimization and generate-and-test methods, to maintain prior knowledge while learning new information.
Current LLM assistants do not learn during interaction—only memorize context. Their weights remain fixed, meaning they don’t adapt or evolve. True learning requires updating model weights over time based on experience, which current systems lack.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.