Ep 86: Yann LeCun on Leaving Meta, Breaking The LLM Paradigm, & Why Hinton is Wrong
81m 56s
Yann LeCun, a Turing Award winner and AI pioneer, discusses his departure from Meta and his new startup, AMI, which focuses on world models and JEPA architectures. He clarifies that LLMs are valuable for language tasks but insufficient for achieving human-level or animal-like intelligence, as they lack the ability to predict consequences of actions or plan through search. LeCun explains that world models, which predict outcomes at abstract representation levels rather than pixels, are essential for intelligent agents to anticipate and plan actions. He contrasts this with generative approaches, which he deems failures for representation learning, citing JEPA’s success over VAEs and MAEs. At Meta, he faced tension between exploratory research and the company’s pivot toward LLM-driven products, leading him to spin out AMI when results matured. He criticizes current robotics methods like VLA models, which rely on imitation learning and require enormous datasets, whereas world models could enable zero-shot generalization, as seen in humans learning to drive in hours. LeCun emphasizes the need for data-efficient systems, arguing that synthetic data from video models won’t solve fundamental gaps. He remains optimistic about JEPA’s potential to revolutionize AI for real-world applications, though he acknowledges significant challenges ahead in scaling these architectures.
You're one of the godfathers of AI, which you're kind of view of the path of progress here. Five years, complete world domination. The best way to get breakthrough research is you hire the best people, you get the fuck out of the way. I'm part of my French. You shared the Turing Award with two others. When did your views start diverging? In 2023. How do you know it was time to leave meta? It sounds like you were thinking through some of these things over a period of time. There is a big misconception about my role, my relation to Alex, and how AI was run at beta. What's like one thing you've changed your mind on in the last year? I mean, the whole idea of. Jan LeCoon is one of the godfathers of AI. He's an absolute legend in the field. Someone I've admired for a long time. And so it was such a treat to get him on unsupervised learning. He's been a noted skeptic of LLMs in many ways. And so we dug into what LLMs can do, what they can't do, some of the limitations he sees and why he ultimately decided to pursue different architecture. And we also talked about his time at meta, the things he's proud of and setting up fair, how the last few years proceeded, and what ultimately led him to spin out and start his own company, a me. I think it's just fascinating to get Jan's thoughts on everything happening in the AI ecosystem today, this tension between basic research and then pushing LLMs forward and how that's happening in a bunch of organizations today, as well as his thoughts on just where the whole space is headed. He's just an absolute giant in the field. And when I started this podcast, I hope we get guests like him. So it's just such a treat. I think folks will really enjoy hearing the conversation we had without further ado. Here's Jan. [MUSIC PLAYING] Jan, this is such a pleasure. You're one of the godfathers of AI. I feel like when I started doing this podcast years ago, I was really hoping we might one day get someone like you on. You know, I don't like that term because I even knew Jersey when you were a godfather in New Jersey. It doesn't mean the same thing. Very fair, very fair. Obviously, you're bet on neural nets when everyone doubted them. It's legendary. And I feel like today you're making a similar bet in many ways against LLMs and the kind of predominant generative architectures that so many believe in. You've recently started a new company behind this theme. And so our goal today in the conversation is to leave our listeners with a lot for more information about me, what you're doing there, some of your work at tapestry, why you think the rest of the field is pointed in the wrong direction around some of these generative models, and then also just get your reflections on the way the field's unfolded your time at meta and all that. So modest goals for a single podcast episode. A theory would be great to start with me, because the company feels like the clear statement of your technical thesis going forward. And so you recently launched the company, focused on world models and scaling the Jetbar architecture, which you obviously pioneered over at meta. And so I wonder if you could talk a little bit about the origins of that architecture and the extent to which you drew inspiration from the human brain and the way that works. So first of all, I want to say, there's nothing wrong with LLMs. In the sense of LLMs are the basis for a lot of very useful AI products that all of us use, including me. They're great. OK. For what they do. They're just not a path towards human level or human intelligence or even animal-like intelligence. So that's my claim. OK. I'm not saying it at times I'm useless. Right? I'm just saying they're not a path towards-- I mean, help build some of the first major open source ones. Right, absolutely. So what is-- I mean, so I mean, it really stands for Advanced Machine Intelligence. And the kind of subtitle, the model, if you want, is AI for the real world. So basically, a lot of AI techniques that people know about today are good for language manipulation. Either human language or computer code or mathematics or legalese, which barely qualifies as human language. Unfortunately, a lot of human language used for it. Right, sadly. Language is very special in a way. And it's particularly well-suited for the type of architectures that have been so successful recently, the large language models, GBT style architectures. But what about the real world? What about understanding the physical world? Turns out, reality is way more complicated than language. Because it's high dimensional, it's continuous, it's noisy, it's messy. And training a system to understand the real world is much, much harder. So that's really what we're after. That's what I've been after for most of my career and really kind of working on an accelerated fashion over the last five, six years or so and making significant progress over the last two years. And so it made sense to really do a startup around it and sort of go into a high gear when pushing that. And it became clear by the end of last year that META was really not the right place for that. So she's why I left and started Aminux. I think it's an interesting trend that we're seeing across the board right where it feels like there's many folks spinning out of either some of the large companies or research labs that have a particular direction of research they're excited about. And you'd have said an interesting vantage point of this from your time at fair, this almost tension that exists between go, pursue, as many different research directions as possible in these companies versus, hey, something's really working. This is the thing that we're gonna sell for the next six, 12 months, like go focus on that. You know, I'm curious your thoughts on that and what you've kind of seen in the industry at large. - Well, it's a strange trade off. There's really two modes of parenting, right? There's a lot of exploratory research, a lot of research directions, right? And sometimes something kind of seems to work and you need to push it further. And it's not research anymore. I mean, the people working on it are researchers or they're called researchers at least in the press, but really it's becoming more engineering and pushing for products, right? So that happened a number of times at meta because of things that were started at fair. Such a thing happened in, you know, early 2023, essentially when, you know, Lama, which was developed at fair, Lama one was very promising. And meta created a whole organization, GENERI to turn it into something real and a series of products and produce Lama two, Lama three, Lama four, which was a bit of disappointment. And because, you know, Mark Zuckerberg was disappointed by it, he kind of rebooted the entire organization, reorganized it and hired new people, et cetera. But what also happened over the last year is that basically the company met, I realized that that fallen behind a little bit. And so that kind of refocus the strategy on trying to catch up with the industry. And the sad side effect of it is that a lot of the exploratory research was basically not given high priority anymore. I mean, it didn't concern the stuff I was working on, all the JEPA and Wal-Marrals, because Mark himself and do buzzwords with the CTO and a bunch of other people in the company were really interested in that project, you know, he believed in the long term impact. But the rest of the company was just, you know, totally entirely focused on LLMs. And made it clear to me that meta was really not the right place to push for that project anymore. And then we started to have good results. And so it was clear that, you know, we had to kind of make that transition between research and actually kind of developing the technology is going to get up and building products out of it. And we realized also that most of the applications were probably foretanks that meta was not particularly interested in. A lot of applications of the kind of stuff that we've been working on is in industry, like manufacturing industry and stuff like that. Obviously, you're kind of pursuing world models and in that broader world. And I think there's other people that have come at the world model pace from a more like generative approach. And so I think you've got folks-- you know, you've got the Google folks in Genie and the video models. You've got folks building VLAs on the robotics side. You've got Feifei and kind of like the 3D spatial models. As you think about kind of the body of evidence that got you excited about the JEPA models and how you kind of compare them to what the generative folks have done, you know, where do you think we are today in terms of like comparing these architectures and approaches? OK, so one model is quickly becoming a buzzword. Yeah, right now, right? So we're doing research, but also in industry, to some extent. And then there are two factions, if you want. I'm not going to talk about VLA, because VLA is clearly now being seen as not going anywhere. Like it's really not working. So VLA is, you know, vision language action models, right? So basically, use the LLM technology to train the system to produce actions for like controlling or robot to something like this, right? So you have vision in, language in, action out, maybe language out too. And that's pretty much now seen as a failure, not being reliable enough, requiring too much training data, you know, things like that. OK, then there is world models. OK, so what is a world model? A world model at a regional level is something that allows an agentic system to anticipate the consequences of its own actions. OK, predict the consequences of its own actions. For my point of view, I cannot imagine how you can even think of building an agentic system without that system having the ability to predict the consequences of its actions. I mean, that's pretty essential, right? When we act in the world, we have this ability. And when we take an action without thinking about the consequences, we are taking a big risk and very often, you know, other people think we are an idiot. We're plenty of examples on the international political scene - Oh, yeah, the more people who have come here.
No ability to predict the consequences of their actions. So that's the one model. That's all it is, right? Ability to predict the consequences of your own actions. If you have this ability, then you can plan a sequence of actions to accomplish a task to, you know, satisfy a goal. And you do this by planning, reasoning, by a process of search and optimization. You don't do this by predicting one action after the other, author aggressively, like a VLLA we do. You do this by searching for a sequence of actions that will accomplish the task you set for yourself. So the blueprint for this is completely different from what LLMs can do at the moment. LLMs do not have the ability to predict the consequences of their actions and they do not have any planning abilities. Because inference is by predicting the next token, right? It's not by search. So right there, you have the two characteristics that I think are essential for intelligent behavior. Ability to predict consequences of your actions. And second, ability to plan by optimization by search. Find a good sequence of actions that will produce the correct outcome. And then there is a third characteristic which is how you predict the consequences of your actions. So if I have a water bottle in front of me, I realize some people would just listen to this and not have the picture. So I have an open, uncapped water bottle in front of me. If I push at the bottom, it's going to slide on the table. If I push near the top, it's probably going to flip. We can't predict exactly how the bottle will fall in wish direction. We can't exactly predict how it's going to slide, you know, how the water will spill. You know, whether the table is tilted in one way and the water will, you know, kind of flow in one direction or another. There's no way we can predict this at a pixel level. And so our mental model of the world predicts, but at a naps track level of representation. So as you were working on this architecture, it was a lot of it inspired by the human brain. I mean, obviously, like the way you're articulating things is exactly how we do things. Right. Or at least by cognitive science, right? Whether you can sort of translate this into an oral architecture and things like this, that's a big gap there. Okay. So that, you know, certainly cognitive science was a bit of a motivation or, you know, what psychology is called system two, which is this idea of the way you behave in sort of deliberate reflective behavior, is that you do imagine predict the consequences of your actions and you plan accordingly. Contrary to system one, when you just act, you know, reactively and instinctively. So yeah, there is an inspiration, but also there is a lot of empirical evidence that you don't want to generate pixels. Okay. I've been, I've been really interested in that problem of learning models of the world by prediction for a very long time. And then had in the beginning about five years ago, realizing that all of the architectures that have been successful to learn representations of images and videos are non-generative architectures. And all the generative ones, basically have been failures, right? So VAU, right, virtual auto encoders or auto encoders more generally. It's kind of a natural way to think about like learning abstract representations of inputs, right? So you put an image at the input of a neural net and then you try to just reproduce the inputs on the output. Now with a big neural net. Now if you just do it this way, you know, on that will not do anything interesting. We just learned the identity function. Completely uninteresting. It doesn't work. If you try to VAU to learn representations of images, you get something, but it's really not that great. Same with sparse auto encoders. Then you have another set of techniques. And it's kind of derivative of something called denosing auto encoder. Mass auto encoder is a version of this. Bert is a version of this for an LP. So you take the image, you corrupt it in some way, and then you're trying this big neural net to recover the original image. There's a huge project that fair on this called M.A.E. master auto encoder. It was very disappointing. A lot of computation and not really great satisfying result. Semitaneously, some of the same people working on M.A.E. and some other people in Paris and in New York were working on other techniques using non-generative architecture, joint embedding architecture. So take an image, corrupt it in some way, and then run the two images to encoders, and then try to predict the representation of the original image from the representation of the code. So there were representation of the corrupted one. That's JEPA. JEPA means joint embedding predictive architecture. So you have one encoder that makes an observation. Another encoder that makes a different observation. We try to predict the representation of the first one from the second one with a predictor. Those techniques turned out to work much better for representing images and video. It was like Dino, Dino V1, V2, V3 project that is still going on at Ferrin Paris. Projects like IJEPA and then VJEPA and then before that there were like SIMCM and MOCO and all kinds of different techniques. Mostly for meta, there was a bunch of others from other groups. But that turned out to be a much better way of learning representations of images than predicting pixels. So it just clicked in my mind, but not just mine. That this was the way to go when predicting pixels was kind of a losing proposition. It feels like there's all these robotics demos that are released from some of the model companies that are feel increasingly impressive and maybe seem to resemble things like planning and reasoning when they maybe haven't seen a room or a specific version of a task before and are still able to execute that task. What would you say to our listeners? I guess that sort of that stuff would feel like, "It feels like we're trending toward some real progress with some of the general approaches." Well, there is real progress and some of those demos are really impressive. But they are trained with enormous amounts of data collected either from teleoperation or from just human action with things you hold in your hand and look like. grippers that you collect the data for that or just tracking hands and fingers of a person and then translating this into commands for robot. So those things are trained with imitation learning mostly. And a little bit with reinforcement learning to find human in mostly insemination. So the issue with this is that you need a lot of data to train those systems through imitation. And it becomes expensive and it's a little brittle in the sense that you need to collect lots of data for every task you want the robot to solve. Whereas if the system had a wild model that allowed you to predict the outcome of an action, it would just plan an action to solve a new task without actually having to be trained to accomplish this task. So the degree of generalization you would get with a wild model based system is much, much larger. What is spectrum of tasks with less training data that would be required than a system with imitation learning and. No doubt those approaches require more data. And I guess this question of generalization really is a big question. Some folks have shown some results around getting better at task A helps with task B, but that obviously feels like the still the big on-answered question around those architectures. I mean you get this synergy between tasks. So the more tasks that you train the system to solve, the more tasks is going to be able to acquire with some amount of data. Regardless of whether to take with technique you use. But the hope with wild models is that the system can solve new task zero shot, which humans are completely capable of doing. And many animals as well. So that's really the hope like solving a lot more problems with either a small amount of training data or no training data at all. And just a little bit of maybe RL style of fine tuning. How is it that a 17 year old can launch a drive in like a dozen hours or maybe 20 hours? We have millions of hours of training data of people driving cars. We still don't have level 5 serving cars. So imitation learning obviously doesn't work even for just the task of a 20 year driving. Yeah, I guess a race between the ability to develop some of those capabilities which may take time and lots of data versus this kind of architecture. I feel like there's this dream of using video models to just generate like tons of synthetic data for simulation. And you know, even if it's not perfect these video models from a physics perspective, it's like helpful enough to improve robotics and the underlying physical world. What have you made of some of those approaches? Obviously, I think a video has been focused there. Google seems to be going down that road. I'm sort of asking again the question, you know, why can 17 year old launch a drive in 20 hours? You don't need millions of hours of demonstration. You don't need synthetic data. You don't need any of that. So, you know, I want a system that can learn as fast as that. If we crack that, then we don't need, you know, generated data, right? I mean, we might need to train the system in simulation, but not with the same amount of data.
of time or trials as current systems require. It's really a question of data efficiency. - I was interviewed Jerry Tork on the podcast. He was at OpenAI and spun out to start his own lab. And you could sense a similar tension where I think he actually might even agree that if you continued scaling or all the way we're scaling, you continue getting very impressive results. But I think he felt, God, there's just gotta be some way more efficient way to do this. And it's interesting, it's an interesting tension because you could imagine if you're OpenAI and you know something is gonna continue, like you could continue scaling it and it will keep getting better. There's not a ton of incentive necessarily from a business perspective to do something more data efficient. - Right, there's no incentive for the other companies to do anything different either because they're all chasing the same. Like they can't afford to fall behind the others, right? So they all work on the same thing. And it's a bit of this sort of, you know, kind of herd behavior and mostly you can see a convadiate where everybody is digging the same trench. - Yeah. - And so I purposely set up the headquarters of Amiravs in Paris. (laughing) The American office being in New York, not the convadiate. - It's really interesting because I think it points to attention that it exists in the broader ecosystem today where you could imagine the other side being, sure, maybe there are more data efficient methods out there but like almost who cares because we can keep scaling what we have to better and better results. And then obviously I think from both, you know, new things you can accomplish from these models as well as just the joy of being a researcher and finding these new things, I get why there's such an attraction to these other architectures as well. - And it's a bet. But, you know, we're pretty confident because we have results already. - Yeah. - And as you think about like the kind of, the initial space that you're most excited about for the AME technology, like what gets, you know, where do you think, you know, the technology goes and what are you most excited about? - Well, you know, AI for the real world. Like, you know, where is your domestic robot? Where is your level of self-driving car? - Yeah. - Where is, and that's, - Why am I gonna get a domestic robot? I'm excited about this. - Well, so this is several years down the line. Okay, despite the fact that there is like, huge number of companies building robots, none of those companies actually has any idea how to make them smart enough to be useful, right? - Or trust it around with a baby in the house or something or something. - Certainly not that. But even for like, you know, it's relatively now a manufacturing task, right? You know, I mean, none of them really knows how to do this reliably other than, you know, for imitation learning for a small number of tasks. So, how do we make those things useful? So that's kind of a relatively long term objective. Shorter term, there is a huge amount of applications in industry where you need to have a system and intelligence system that has the ability of, you know, predicting what's gonna happen if I change this control variable on this complex system, be it a jet engine, a chemical plant, a power plant, some manufacturing line, a patient, a human cell, right? Those are systems that are sufficiently complex that you can't model their behavior with a small number of equations, right? So the traditional way of modeling does not work. And what you need to do is train a neural net deep learning system to, you know, model the dynamics of that system from data. And what you get at the end is a phenomenological model of that process of that system. And if it's action condition, then you get basically a one model of that system that allows you to control it optimally for whatever purpose you have. And I think the number of applications of this in industry is mind-boggling. - Where do you think we'll be with, you know, jet models over the next couple of years? Are there like, you know, milestones you'd point to or like what's your kind of view of the path of progress here? - Okay, a couple of years is a little short. Like five years, complete world domination. Essentially. (laughing) Okay, so somewhere between on the path to world domination in five years. - I mean, this is kind of a joke obviously, but this is a quote from Linus Starvalds, right? You know, when people are asking what's your goal with Linux, is a total world domination. Yeah, it actually managed to do that. - Yeah, very fair. - The first approximation every computer in the world runs Linux. So that's kind of a joke. But in the end, I think this is the blueprint for intelligent systems of the future. There still be a small place for an LMS, you know, for like a language interface basically. But what we're designing are systems that are capable of thinking. They may not be capable of talking or listening initially, but they all do the thinking. And then you can add the talking and listening on top of that. - I'm sure you and the team are eagerly working to kind of get the early proof points of this. And obviously you've already had some in the work you've done. How do you think about the interim steps of what you'll be able to show on that path to five year world domination? - Well, so I think, you know, within a year or so, we'll have, I think a general methodology to train hierarchical one models on, you know, a very wide variety of modalities. We know we can do it with Java and video, with some techniques that we're not completely happy with because they have some shortcomings. But we have sort of small scale demonstration over methodology that we think is really what we want. So we need to scale that one up and get it to the similar level of performance as the other techniques that are not as satisfying if you want on things like video, but also on other types of datasets that we would get from industry partners. Okay, so we'll have demonstrations that we can train world models, perhaps action-conditioned world models that allow us to plan for a number of different use cases. Some of them would be robotics. Some of them will be industrial process control of various types. Maybe some of them in health, healthcare as well 'cause we have partners in that domain. And that should be within a year or two 18 months. And then we push this methodology and those models into those use cases with partners, some of which are investors already company. And gain experience on how to kind of essentially build a some-white universal one model if you want. - I mean, we've obviously had this experience before of kind of making this really contrarian bet on neural nets and being certainly proven abundantly right in the history books. I guess as you think about this bet, which I think if you talk to the majority of people, maybe at the kind of edge of various parts of AI, maybe would say is contrarian today. In what timeframe do you think it will become a parent like, you know, this was right? - I think it'll happen faster than expected perhaps because I mean, you can see that world model is already becoming a buzzword, right? At least at the research level and is starting to kind of permeate into the industry. - Yeah. - And you know, realising like VLA is suck and you know, LLMs don't work for real world data. Industry is realises already. Certainly on the user side. And I think because of the importance of the robotics industry, you know, a lot of people are kind of trying to figure out like how do we get there? How do you get how you make those robots useful? So I think the realization that you need a change of paradigm is happening as we speak. It will become completely obvious to people by early 2027, I think. - Yeah. - Now that doesn't mean we'll have a solution by then. We hope we will, but you know, we'll see. - I guess you know, switching gears to the LLMs side, you mentioned some of this work you're doing with tapestry, which I think would be really interesting for our listeners. And so maybe to speak to that a little bit. - Okay, so this is kind of a little bit orthogonal to Ami Labs. - As if that wasn't enough to keep you busy. - Well, it's kind of an idea of been forming over the last three years or so. It's the fact that people increasingly use AI assistance for various things, right? I mean, you see a decrease in the use of traditional search engines and you just ask a question to your favorite AI assistant. And you know, if the plan that met on others are developing of having smart devices like smart glasses and stuff like that, it's realized basically you'd just be talking to your AI assistant by voice with to your smart glasses or maybe some other smart device. And so all of your information diet will be mediated by AI assistance. And if you are someone in the world, let's say outside the US or China and you have an AI assistant and that AI assistant was built in California or, you know, Beijing or Shanghai or a shenzhen. It's not good for you. Like you may speak an language that those systems really haven't been trained to handle particularly well. You may have a culture that is not particularly well understood by people in Silicon Valley and China. Not well represented by the training data that is probably capable on the internet. You may have a value system that is absolutely not represented.
by people building those models. And certainly, you almost certainly have political opinions that are absolutely not represented by the handful of AI systems you might be able to get from the West Coast tech companies or from Chinese companies. So what is the solution to this? Like how do you serve a farmer in India or even a philosopher in France or Germany? And what you need is a platform which basically is an open free foundation model, LLM style that is fine-tunable by anyone to cater to the interest of people speaking a particular language, having a particular culture, having particular value systems, political biases, creeds, whatever it is. And so what you need is a wide diversity of AI systems. There's a lot of countries around the world that are neither the US nor China who absolutely want some level of sovereignty for AI, not just for the industry, but also for the citizen. They don't want the citizen to get brainwashed by Chinese model or a California model, actually. And so they want sovereignty. How do you get that? So the way you get an open platform like this to get to the frontier is you just train it on more and higher quality data than the proprietary systems. If you talk to people in India, in France, in Vietnam, in Switzerland, in Korea, Japan, Kazakhstan, everyone wants basically sovereignty. And you tell them, like, you guys have been training your model locally. You don't have to share your data. So that's the cultural aspect of tapestry. You would have international contributors to tapestry, contributing to training a global model that would basically constitute a repository of all the world knowledge and culture, if you want. But the contributors would contribute data and computing resources. But they would preserve the control on their data. They would not have to share their data with the other contributors. What they would contribute is parameter vectors. So it would be kind of federated learning style thing, where you have a bunch of data centers. They get the parameter vector from the global consensus of a model. Think of it as an average of all the parameter vectors of all the contributors. So all the contributors periodically tell everyone else through maybe a central server, here is my parameter vector. What is yours? And so you exchange parameter vectors like this. And a local worker, basically whatever it updates its parameter vector, it tries to also make it as close as possible to the global consensus vector. So as the training of this kind of progresses, all those parameter vectors converge towards a consensus model, essentially, which is kind of a repository of all human knowledge. Now you have an open model that is as good as if it had been trained all the data in the world. And now you can fine tune it for your own purpose, your own political cultural and linguistic biases, whatever you want, or centers of interest. And I think there is a natural force for this to happen, because most countries that are not the US, nor China, want sovereignty, but also because AI is fast becoming a platform. And there is a natural tendency for platforms to become open. That's what happened with Linux. And that's what happened with the software infrastructure of the internet or the wireless network. It's all open source. It was proprietary initially. But that was all wiped out. It's a really clever way to get around-- it would seem that this trend of decreasing open source. And obviously, I think there's been many fears that as the close source models get better, they'll be held back, and they'll be used to train the next generation. And they'll be this almost like a scape scenario for close source models where they get so much better than their open source counterparts. So remember who the big players of the internet infrastructure were in 1996? Sun macrosystems, HP, Dell, and a few others. So Sun macrosystem was showing you Solaris with their proprietary hardware, HP with HP UX. They were claiming, you know, Unix is so much more reliable than Windows. You're not going to run a web server on Windows. Dell was doing this with Windows NT, but like, who is running Windows NT now as a web server? All of this was totally wiped out by Linux, like the entire internet runs on Linux, even Azure, right? Even Microsoft in Linux. So basically, OpenAI, Anthropic, et cetera, today are the Sun macrosystem and HP UX of yesterday. Yeah, I mean, I guess it implicit in that is obviously, I think your view of the limitations of what, like, you know, these models can only get so good, and so it will be possible over time for the open source folks to catch up. They've already run out of data, right? I mean, the open, open-early-available, public-available data, text data is already all used. I mean, there's not more of it, right? So what those companies are doing is licensing commercial copyrighted data or training on synthetic data. And I guess I'm curious, because obviously, there's been some impressive results in the last few years that they have been able to drive post these large-scale returnings. IMO gold, the meter task-arisen benchmark keeps going up. That's-- OK, that's very interesting. Now, think about those two domains, right? Mathematics and code. Those are two domains where the language itself is the substrate of reasoning. It's not the only substrate of reasoning. But a lot of-- when you do mathematics, right, the formal way on the piece of paper, not the intuitive stuff. But the human-eupelate language, right? And elements are really good at this. So proving theorems and stuff like that, that's what elements are really good at. They're not so good at the sort of coming up with good concepts and definitions and things like that. It's more like, here is a problem solving. They're problem solving. Mathematics is not just problem solving, right? Most of it is actually a creative act that those things don't do. And same for code. So LLMs are good programmers. They're not software architects. They're not computer scientists, right? But they can program for us. So they're not in a state where they can just replace humans. It changes the world of humans. So humans now kind of go one level up in the abstraction hierarchy. Our world is to decide where to build, but building it, you can get help from LLMs. But OK, the important point is that LLMs are particularly successful at domains where the language itself is the substrator reasoning, not for anything else. Yeah. What would an LLM need to do to convince you otherwise? So like a zero-shot agentic system, right? The urban agentic system gives it a new problem. It's not been trained to solve that particular problem. It doesn't have a script for it. Is it going to be able to accomplish this task that it's never been trained to solve? And unless this system has the ability of predicting the consequences of its actions, and then using that for planning, it's not going to be able to do it. And you're not going to do this with an LLM. You're going to do this perhaps with a significantly augmented LLM that is capable of search and planning, blah, blah, blah. And currently, LMs that do math and code actually do this. Yeah. Right? Because they search for sequences of tokens that actually accomplish a particular task, and they can run the code or verify that the proof is correct or whatever. So you have a way of checking whether something that's produced is correct. But that's not a very efficient way of doing planning. And it only works in domains where this type of search can be performed in token space. What I'm talking about with JEPA is you don't do this in token space. You do this in Astra-X thoughts space. And I'm sure some people, Listeen, might think, well, hey, even if it's inefficient and it works, and it works at things that are done in token space, are still a large part of the economy. I mean, if it works, it's fine. I mean, there's, again, there's nothing wrong with using it for what they're good at. It's just not a pass towards you, whatever they are. You're missing, you know, in a huge domain. You seem like, hey, it's going to tap out before it can become a software architect, whereas I'm sure it's not going to tap out. It's just going to have a limited--
you know, ability to be deployed for like an, it's going to become like increasingly difficult to kind of deploy it for an increasingly large number, you know, of use cases because you're going to have to collect tons of training data for each of those use cases. And there's a basically you're not going to be able to make those systems completely reliable, you know, without hallucinations or or dangerous stuff or, etc. Unless your systems have the ability to predict the consequences of their actions, which means they're going to have to have explicit world models. Yeah, so it's the bed against you to the 100% accuracy and then also the generalization across different tasks. Right. I guess you know one thing that's so interesting about the way the field is developed is obviously you share the the Turing Award with two others and I feel like they seem much more convinced of like maybe the power potential threats or safety risks of elements over time. I'm wondering like when did your views start diverging in 2023 and what like drove that in your mind? I didn't change my mind. They changed their mind. And just about the same time it was basically GPT-4. I mean Jeff basically had, was not connected to any of that. He was never really interested in elements and discovered GPT-4, you know, 2023 when he came out. And basically he had a piffini and said, oh my god, those systems, you know, are really close to human intelligence and they have possibly they have subjective experience. And he did a quick calculation saying like, okay, the human cortex is about 16 billion neurons. If you want to do something like back prop, okay, the brain doesn't do back prop directly. But if you do something like back prop, like some sort of you know, graded estimation for some sort of objective function, you probably need like a network of a few neurons to kind of reproduce the functionality of a virtual neuron in a neural net. And so he said like, let's assume, you know, maybe you need a circuit of 10 actual neurons to reproduce what a back prop neuron does. Then all of a sudden your cortex is only 1.6 billion neurons. Oh my god, GPT-4 is really close to this. Okay, so maybe it's that smart, you know, it's going to get as far as humans. I do not believe in this claim at all. This is kind of, you know, Jeff's way of saying, okay, basically I can retire. I can declare victory, you know, I search for the learning algorithm of the cortex on my career. Maybe I didn't discover what it really was. But back prop seems to be like a good substitute for it. It works really well. And so maybe that's what we need. So I can retire. And go around the world and give talks about, you know, the potential promises and dangers of VAI. That's basically what, you know, I think what is intellectual kind of trajectory has been. He's much less vocal about the potential dangers now than he was a year or two ago. He kind of realized it's probably a way to design to the intelligence system. So first of all, he probably, you know, he realized that current LLM's are not that smart, first of all. And second, that is probably a need for a few breakthroughs, like conceptual breakthroughs before we get to human-like intelligence. And third, that the blueprint of those systems would be quite different from LLM's. And we probably have a way of, you know, making them controllable and things like that. I've been saying this for years, but, okay, he sort of discovered this recently. Yeah, same kind of, there's a similar thing with the Ashwai. I think what they are both worried about is the ability of society and the political system to make sure that the benefits of VAI will be maximized. And VAI would not, you know, just profit, you know, make, rich people, even richer. And, you know, accentuate inequalities and, you know, cause major catastrophes because of bad usage. Okay, this is not like the, the Duma scenario of VAI ticking over the world. It's more bad uses, uses. What seems possible with the LLM's of today? Which is a danger, but, you know, I don't think it's as apocalyptic as, you know, what some people have claimed. It is certainly not as apocalyptic as what even anthropic has claimed. And let's try to kind of lobby governments into, you know, scaring governments into kind of regulating AI because of that. I don't, I don't subscribe to this at all. They seem to genuinely believe it. I think they genuinely believe it, but also I think there is, you know, some kind of commercial good commercial reasons for them to believe that and to kind of, you know, you know, Renoir, some people and government to thinking their systems are dangerous. And it sounds like, you know, with these other architectures, do you think they're, because obviously it doesn't, you know, as maybe barracies you are on LLM's being the end state of everything, you know, you have some pretty ambitious timeline too for these new architectures. And so, yeah, it doesn't seem like you think we're particularly far away from, from some very compelling capabilities. How do you think about, I guess, the safety around, you know, if it ends, if these breakthroughs end up coming from newer architectures and whether that should make us rest easier or not? I'm going to say something that's again, my B controversial. And certainly my, some of my colleagues at MIT are, didn't like me saying this, but I think LLM's are intrinsically unsafe. I don't think they can be made reliable and safe. They cannot be made reliable because you can't stop them from hallucinating. And if there are agentique, you cannot guarantee they're not going to take an action that, you know, they didn't predict the outcome of that. I mean, it's a surprise you they can do these like 15 hour coding tests given the concerns around reliability. The coding is something where you can actually verify that, you know, the code that you generate, you know, satisfy your specification. But, but not everything is coding. And there are examples of, you know, coding agents like wiping up your, your I drive, or doing stupid things, right? That makes you lose a lot of money or data or whatever. So I think, I think, you know, LLM's in the current forms are intrinsically unsafe because they cannot predict the consequences of the actions. Because the way the task that they accomplish is determined is, is subject to their training. You know, you give them a prompt. And then they will accomplish a task that corresponds to that prompt only to the extent that their training has conditioned them to actually do the right task corresponding to this prompt. But it is no like, you know, hardwired constraint that will force them to accomplish this task and then, you know, predict that the task would be accomplished properly. Yeah, I mean, I think famously in the early days, right? They would, you'd ask them a question and they'd keep asking the keep asking the question, right? Right. Right. Or I mean, also they don't have common sense, right? So I mean, there's the joke that was circulating like a month ago of, you know, I need to wash my car and, you know, the car wash is 100 yards from my house. Should I walk? I tried it again, like maybe two weeks ago. They all say, yes, you should walk except Gemini. So you think so? So they're training on your video of having done, having given that speech before. There was not my video because I didn't come up with this example. Whoever came up with it. Yeah, right. But there are a few instances where I said, like, you know, an LLM can do this. And then six months later it was kept people doing it. And it's simply because, you know, as soon as people watch the podcast of me saying, "Eleven's can do this, they of course type it into chat GPT." So not because part of the training set. Right. And now, of course, you know, the next version has that, you know, that thing in the fine tuning set. And of course, you can answer the question, but it's not because it's, it becomes smart all of a sudden. It's just because it was, it's physically trained with that question. So, LLMs are intrinsically unsafe. I don't think there is a new way to fix that in the current paradigm. And what I've been proposing is the architecture I've been talking about is objective driven AI. So basically, you give an objective to an AI system, which is accomplish this task. Now, how does the system knows it will accomplish this task? It has a word model and it predicts, you know, the outcome of a sequence of actions it imagines taking. And if this outcome satisfies a cost function that, you know, describes to what extent the task has been accomplished or not accomplished. Then that system, if, if the way that system works is by, optimisation, finding a sequence of actions that accomplishes this task, minimise this cost according to its one model, it can do nothing else. Yeah. Okay. And of course, there's many things that can go wrong there. In particular, the cost function might be inaccurate. It could be that the cost function you think is actually measuring to what extent the task has been accomplished, but perhaps it's not accurate. Okay. The one model might be inaccurate. So the prediction that the system makes is actually not the right one. So it's prediction of what was going to happen as a consequence of its action wasn't right. Okay. So the system can still make mistakes.
but it can predict the consequences of its actions to some extent, which is, I think, indispensable for any agentic system. Now, you can add to that system is not just a cost function that guarantees a task it's been accomplished, but you can also add a bunch of other objective functions, other cost functions, or even constraints that are safety constraints. They say, okay, you know, don't hurt anybody on the way, right? And you cannot specify this at an abstract level, but you can have, you know, all-level objective functions that put together will guarantee that the system will not be dangerous. And the system cannot violate those things. By construction, it will have to satisfy those conditions. Not the case for an LLM. The LLM can always escape. There's a gap between your training error and test error. There's always going to be a prompt where the system is going to do really stupid things. To talk to one specific space around LLM's, like, you know, I think you're obviously really excited about a me and healthcare. And I think, if people have been using LLM's and healthcare for all sorts of things. And so I'm curious how you think about, like, the set of things where LLM's are just not going to work in healthcare, and you need, like, a model that understands the world better. - So, I mean, designing a course of treatment for a chronic disease, for example, or even a non-corny disease, for a particular patient, which may not completely fit into, you know, templates that you've observed before. But if you have a good mental model of the dynamics of the physiology of the patient, and you might design a course of treatment that will actually bring the patient to a good state. When I'm seeing a patient, it can be a cell. Okay, how do you tell a stem cell to turn into a碰queous bit of cell that produces insulin? Okay, you have a patient with type 1 diabetes. And, you know, they have, you know, the immune system basically, you know, kind of eats up their own bit of cells right into the immune. How do you keep making bit of cells? You know, can you send a message? Do you have a model of its overhuman cell that will allow you to figure out what sequence of message do you need to send to a stem cell so that it turns into a bit of a bit of a cell? - The less LLM pilled camp, and the LLM pilled camp talk past each other, but it's like, I think it's actually very possible that both, what LLM's can do, which is maybe scaling what a top doctor, the treatment you get at like the top doctor or the top place, scaling that around the world, like unbelievable potential impact of that, right? If you're able to do that. And then, you know, I think what you're talking about, which is certainly still on the come for a lot of these things is, okay, well, even better than the top doctor, like how do you go do that? - But it's more than just a top doctor, right? Because, I mean, what the LLM can do well is, you know, it can sort of regurgitate knowledge that you can read in books mostly. But if medicine was only kind of about accumulating declarative language, that declarative knowledge that exists in books, you can be a doctor by just reading books. And you can be a doctor by reading books. You have to do, you know, residency and, you know, actually kind of listen to the heart and like press on the belly and things like that to, you know, diagnose, have been decided, so whatever it is. - Yeah, it's interesting. I would be very curious to see whether LLM's themselves can provide like, you know, top quality healthcare globally. Love the check back in on that one. It seems like there's pretty close. You know, I definitely want to hit on your time at Meta, because you spent over a decade building, like one of the most respected research labs in the world. You know, obviously you recently left, does you reflect back on the time there? What do you think you got like most right and most wrong in your time running fair? - So the thing we got right is, you know, building a top research lab that really sort of innovated, produced a lot of the sort of basic methods and science and tools like PyTorch that are useful to the entire industry, right? I mean, the entire industry is built on the right research basically, except for a few people at Google. And I think a culture of, you know, openness and kind of, you know, scientific process, which I think is necessary for breakthrough innovation. - Yeah. - Because, you know, there is a lot of, there's a whole chain of innovation, right? You have Blue Sky Research, new concepts, a lot of that takes place in universities. Some of that takes place in advanced research labs in industry, which can be counted on the fingers of one hand. You know, Google is a good one. You know, fair was a good one. Hopefully it will still be, I'm not sure. And, you know, a few others. Then you have, okay, this is a good idea. Like, let's push it forward and see if it can be made useful, but still at the research level, in a sense of, we're not going to fool ourselves. We're not going to try to just, you know, find a solution that just works for this problem. We're going to see if this technique that we imagine or we picked up from other people in the community can actually be pushed and be made practical, not as a product, but like we can show that it beats some recordums, you know, some task or benchmark. And then the next stage is for the company that hosts the research lab to say, okay, now we're going to push the button, devote a, you know, big engineering effort to that vision and then push it forward. That is where a lot of projects fail. That's where a lot of companies kind of fail to pick up. Meta was actually pretty good at this, okay, but far from perfect. It was not like, you know, textbook example of how you do it wrong, like, you know, is there a spark, like totally missing out on, you know, green interface, you know, mouse and windowing systems, right? Meta was, you know, kind of missed a few steps, essentially. And it's partly just organizational, it's partly because you need an organization that is pretty close to research, but not completely a product organization to take the relay of, you know, pushing a technology a little further, not making product with a three-month deadline, but like, you know, pushing things. And we had that at one point, at Facebook and Meta, and then we lost it. And Fair was basically isolated within the company, had lots of ideas that nobody picked up on. And then in 2023, the Gen AI organization was created by basically taking about 60 or 70 scientists and engineers from Fair, right, initially, and then it built up. But then it wasn't there so much short-term pressure that basically that organization, Gen AI didn't have time to talk to Fair. And so instead of being at the forefront and innovating in LLM, Gen AI basically had to focus on short-term things and became very conservative. And so there was a gap, basically, in Pino Smith's match between research and. - Is that kind of what happened with Lama 4? - Yeah. Well, even with, you know, Lama 3, starting with Lama 3. So Lama 1 was a small project within Fair, 2022, early 2023. It's three, Gen AI was created. The Lama people were basically moved to Gen AI. They started working on Lama 2 and then much of them realized, like, I could do a startup. So that was the Gen A system of Mistral, okay? Two of the authors of Lama 1, basically created Mistral with another guy from Google. And, you know, a few people kind of left and sort of did other things. This is not a kind of a happy time at Meta for various reasons. And so there were, you know, a bunch of people kind of left. And then the Gen AI organization which kind of took over Lama 2 to some extent and Lama 3 and 4 was under so much short-term pressure that they became very conservative. And, you know, it's a combination of. Well, this is a parody of the groups, but like pressure from the leadership. And, I mean, there's many ways things can go wrong and you can't blame anyone in particular. But, yeah, that's kind of what happened. - I mean, it feels like a lot of these organizations, obviously, are under short-term pressure right now because there's just a kind of incredible race going on. And so I'm curious, like, obviously, this, you know, fair setup you had and kind of there's a similar one, you know, like Google for many years. And certainly, many researchers running around open AI and thropped trying many different things. Do you think, like, that is still possible going forward or like, is the only, you know, it's one of the only paths to leave and do your own company or, you know, are there still places within the industry that you think have this like original ethos affair even amidst the race that is the race dynamics that are happening? - I think there are a few places within Google research and deep-mind that where people actually do research. But, increasingly, the industry has become more kind of closed, right? I mean, Google certainly climbed up and, you know, meta and fair even is kind of going a bit in the same direction, there are restrictions on publication now, like more restrictions. And so, it's sort of less appealing for people who really wanna kind of do breakthrough research and, you know, they don't get as much resources. If they do something that is relevant in medium term, they are told not to talk about it. And so, it's not, you know, it's not a good ethos affair, I think, for breakthrough. It's not conducive. You know, basically, the gap,
best way to get breakthrough research of the type that you know you were getting in the early days of fair and you know at Bell Labs in the good days and zero spark is you hire the best people and those are people who have good nose to know where to work on, what projects to kind of attack. You give them the means to succeed and you get the fuck out of the way. All right, pardon my French. Yeah, I mean I'm curious like what impact that then ends up having on the broader research community. So obviously one of the legacies of fair is you trained you know so many researchers right and like they're all throughout the ecosystem and it feels like now the maybe equivalence of those people that came in younger than their careers at fair you know they're joining these these labs with maybe shorter term priorities and focus and I guess I'm wondering like you know in this current ecosystem where it feels like a lot of you know people getting into the field are thrust much more into these like short term dynamics does that change anything about the way the ecosystem evolves. Well I mean the people who tend to want to work with me are generally people who you know sufficiently crazy to do it first of all. Very fair and or you know kind of subscribe to the whole idea that in academia and during your PhD you should work on the next generation of a system you shouldn't work on the current generation. Like if you work on LLM in academia now it's incredibly boring at least to me it's boring it's basically kind of studying how and why LLMs work and explaining why the work or what the limitations are. It's like descriptive science it's really not you know kind of creative very creative like I don't find that particularly interesting it's useful. Yeah. And you know if you really want to kind of show how to do new things with LLMs you're not going to have the GPUs you need for that. So like forget that like don't work on the LLM if you're doing a PhD like there's no point you cannot contribute. How do you know it was time to leave that sounds like it was you know you were thinking through some of these things over a period of time you know was there a moment that a crystallized or it was a combination of things right. So first of all you have to understand a lot of people have like completely wrong idea about what my role at Facebook and Metat was. So I joined in late 2013 really kind of started early 2014. The first four and a half years I was director of fair so I built the fair organization set up the culture higher the key people and sort of managed it. And after four and a half years I stepped down from that role for a number of reasons and I became chief AI scientist. Okay so the reason is you know I was basically getting close to turning 60 first of all 58 and I just don't want to do management. Okay I mean I was ready to do it for a while to get the organization started but I'm just not good at it. It's not the thing I'm more like a you know scientific or technical visionary and engineering scientist so other people are much better at management than they're here. So I basically stepped down with two other people, Joie Pinot and Antoine Bard basically took over the directorship of fair and I became chief AI scientist so I was reporting to the CTO and and you know had walls of basically restarting a research project that I thought was necessary because the ambition of fair was always to build intelligent systems right and I thought you know I put my own research in in parentheses what I was running fair I just didn't have the time and I thought it was important to basically kind of design the architecture of like human level you know human like AI systems. And you know I had come up with the concept that this was going to be based on self-supervised running and on you know prediction from sensory signals like video things like that I mean this is these are all the ideas and and one models I actually gave a keynote at NERIBS in 2016 where I said like this is the way AI research would go like world models predict you know consequences of your actions and plan and I said like you know RL is not the thing that will take us there because it's too inefficient. Supervised running has shown its limits and so the future is self-supervised running and one models so how do we do self-supervised running and one models and and I started a few projects on this with like a few avenues that didn't pan out some projects on video prediction and stuff like that and and then came up with this concept that you could train self-supervised running from video but you have to train the system to make prediction in representation space so that's the idea of JEPA and if you have JEPA you can turn into a world model by making it action condition and then you can use it for planning. So I had this idea around 2020 and in 2022 I wrote a long vision paper so I said I'm just gonna write a paper with my entire vision okay spill all my secrets like I don't care but maybe that will rally a bunch of people to to that vision and boy did it work because not only did I rally you know a bunch of students who kind of came working with me yet in where you were in Paris because they wanted to work on this but also a whole team at at fair we said like this sounds great like that's what we want to work on and the angel Alpino said well maybe this should be like a major mission of of of fair we called it advanced mission intelligence that was the internal name of the park interesting okay I'm actually leave with it and now it's the name of the company and you know Mark Zuckerberg you know kind of kind of read that paper and you know what it was about and subscribe to the project and and to both words the CTO also and Mark Schreffer the previous CTO Chris Cox who was my my direct manager she product officer also love the idea so like you know there's a lot of support in the leadership about this project that we internally called me and and you know and and and you started really kind of working for video but then you know company kind of refocused all of its effort on that alarm despite support from Mark and Andrew buzz you know all the layers below like didn't see the point I think and so politically it sort of became a little difficult the applications as a as a said of Japan one model are there are applications in like you know wearable agents and stuff like that but and robotics but but meta chose to get rid of its entire robotics AI group that was led by Jitterna Malik who is not Amazon and so you know clearly it sort of wasn't the right environment anymore most of the applications were in industry that meta had no interest in fair was increasingly getting pressure to kind of basically help MSL with atlams so yeah you know it may clearly it may clear and and that you know thought ramming worked really well with investors too because when I had to raise money for amy everybody knew my story and you anybody knew you know many investors know staff at various vcs that read my paper and or at least into my talks and had bought my story that were realizing you know lm's and limitations and you know what kind of yeah interested by the idea of like building the next generation AI systems I guess was like the scale acquisition like part of this catalyst of like the pure lm focused internally yeah definitely I mean I there's probably some you know other reasons to it I think you know maybe I don't have any sort of inside your formation to comment on this but it's possible that Marx sees in Alex kind of a potential successor to himself like a younger version of himself yeah I feel like like a lot of the popular narrative are you know in the media has been like oh like you know when Alice comes in it then gets harder to run like a research organization you know I don't know if that to the things that you felt that or well okay so here's a big misconception about my role my relation to Alex and how AI was run at MITA I had zero technical contribution to lama like none whatsoever my one contribution to lama was to argue for open sourcing lama too because it was a big internal debate whether we should open source like the legal department was against it the policy department was kind of against it the comms department was for it all the engineering side was for it like bars was for it so there were like enormous internal discussions at a very high level you know 40 people from Marzikovar down every week for two hours for months so so really it was you know kind of a big debate internally and I really really we pushed argued for the fact that you know and bars also was very vocal about it that the you know
safety risks were basically overblown. The opportunities to create an industry were extremely strong. And that we were gonna jumpstart the AI industry by open source single-arm at two. And in fact, that's exactly what happened. So, but I had zero contribution to Lama, positive or negative. Like I didn't do anything to stop it or slow it down or anything. There was a lot of people working on LLMs within Fair and it was fine. And it was anything against it. Okay. Other than saying, this is not a path to human-only traditions, but it's fine. It's useful. You know, same thing for speech recognition or translation. Right. So, and particularly since 2018, when I stepped down from being director of Fair, I didn't have any direct influence on what people were working on. Other than, you know, basically publishing my vision and then rallying people around my project, but you know, they were working with me because they wanted, not because I was their boss. I wasn't telling them to work with me. And so, so I had no positive or negative influence on LLM. Okay. Within, within meta. And I had some influence on the strategy, but it was more like the long term and like how you maintain a research lab and things like this. And in the last year, you know, I mean, starting maybe early 24 and certainly in 25, the way Fair was kind of the direction in which it was moved and managed, basically did not correspond to what I thought was necessary to preserve, you know, innovation, research and breakthrough and preserve the good people. Like a lot of good people have left. Or he. Yeah. And I guess a lot of, you know, you probably was harder to get people to work on the stuff you were working on internally and then I'm sure there's pressure for yourself to work on a lot of the seldom stuff. Yeah. Yeah. No, but a lot of other people also have left. Yeah. No, it's fascinating. And one thing I'm struck by throughout our whole conversation is I feel like you've like had a remarkably consistent point of view, like, you know, in the space, like for a long time and you can go back to, you know, a bunch of the earlier talks you referenced, you know, obviously it is a fast moving space and a ton of interesting things have happened in the last year. I mean, the whole idea of what we used to call unsupervised running that we now call self-supervised running, you know, until about 2003, the whole idea of unsupervised pre-training where you get a good representation for the input data. And then you either fine tune the model with a little bit of supervised label data. And it sort of gives us, you know, some evidence that this whole technique could work. I try to apply this to video because ultimately, what I wanted to do is train the system to understand how the world works by just watching the world go by, right? I mean, that's the basic idea. And sort of started to argue for this in the sort of, you know, early 2010s. Did some some work on simple video prediction. We didn't have GPUs, OK? And then sort of doing this more seriously about after the creation of Fair by doing pixel level video prediction, realizing that wasn't working. But then arguing for self-supervised running, OK? This is an idea of like training a system generically, not to solve a task, but to basically just predict. And then using the representation that is learned this way as input to a downstream task that you can train supervised or reinforcement or whatever. So that was a bit of the topic of my second half of my keynote at Nibb's in 2016. It was to call Nibb's at the time of course, in 2016. And then I kept kind of, you know, kind of pushing for the salient try to kind of discover some methods to get that to work. And what surprised me is that that became incredibly successful, but not for video, for language. And Nibb's basically are a blendingly successful example of self-supervised running. No, that they are. Well, I feel like that's almost like the perfect note to end on. But I want to make sure to leave the last word to you. I feel like there's-- I mean, all our listeners are very familiar with you. But I want to at least give you the mic to point them to anything that you think they should check out with some of the new stuff you're doing, or I don't know. Your work you want to point to. The mic is yours. OK, let me tell you one thing. An LLM works because when you have a sequence of discrete symbols, making predictions is easy. Because it's only a front-end number of possible symbols in your language. 100,000 possible tokens or something like that, right? And you can have your neural net produce a probability distribution over all possible tokens. And then you can sample from that distribution, shift the token into the input, and then produce the next token, and you can do autocracy prediction. So that's a special case. If you have the real world, you can't use the generated model. So now you have to train a system that learns that we're a presentation and makes prediction in the representation space. There's a big issue with this, which I didn't think until about five years ago that was easily solvable, even though I invented one technique to solve it decades before that. And it's the problem that if you take two inputs, let's say the initial segment of a video and the continuation of that video, or you take one image and a corrupted version of it, you run them both through an encoder, and you train a predictor to predict the representation of one from the representation of the other. It's a very simple solution where the system basically predicts a constant representation, and now the prediction problem becomes trivial. That's called a collapse, representation collapse. So the big question of self-supervised learning for JEPA, for those joint-term-editing architectures, is how do you prevent collapse? The solution that I came up with many years ago in 1993 is contrastive learning. So basically you have examples of things that should be predictable from one another, and an example of things that should not be predictable from one another. It turns out this method works, but it doesn't scale with dimension. It doesn't scale very well. There's another technique that was actually invented by Jeffington and Sue Becker in the late '80s, I'm sorry, where you have those two networks, and you try to maximize the mutual information between them. You got Schmiduber, is mad at me because you also came up with a version of this in 1992, and it says that's JEPA. It's not JEPA. It's just another way of preventing collapse of a joint-term-editing architecture, OK? Which is fine, but it's a particular way of doing it, which I don't think is particularly good. So now you have this JEPA architecture. You have to come up with a good way of preventing collapse. And there is a couple ways-- so as I already said, contrastive methods, I think, is not a good approach. There's another set of methods that are called distillation methods. And they do prevent collapse. We don't know why. So a good example of that is Dino or Dino. That's a joint-term-editing method using the distillation method. Basically, one of the encoders trains the other one, is used as a teacher for the other encoder. And the encoder that is being trained-- you do backprop to it. The one that is not being trained, you don't do backprop. But you share the weight with the other one, with some explanation moving average. It's a collection of recipes. It was a paper from Deepline about it, it could be a URL. It was to wrap you on the latent, which uses this trick. That trick is derived from some intuition from reinforcement learning. And somehow it prevents collapse. But we don't know why. There's a few theoretical papers on it that explain why it possibly might work in some simple cases. But it's not a satisfactory function. The cost function, you think you're minimizing. You're not actually minimizing. And so you can't monitor. It actually goes up when you train. It makes-- so we don't like this method. But it works. And some of the models we've trained, large scale video representation learning system, Vijepa, Vijepa 2, Vijepa 2.1, they train using this method. Ajepa also. But we're moving away from this. And now we have a few papers that came out recently on a specific regularizer to prevent this collapse, which basically tries to maximize the information content coming out of the encoder. So it's in the same family as the Becker and Hinton from '89 and the Schmiedu-Börer 1992. And a bunch of others since then. And to some extent, also, quantitative techniques, also it's not-- although it's not-- simple, contrastive. And then the question is, how do you measure information content? How do you maximize the information content coming out of a neural net? And the problem is, if you want to maximize the quantity, you either need to be able to measure it or you need to have a lower bound on it. Information content, we only have upper bounds. We cannot measure it. We can only come up with up bounds. And so we take an upper bound and we cross our fingers. OK. Ajepa kind of works. So the latest one is called sigreg, that means sketch isotropic Gaussian irregularization. We had a previous one called VCreg or Vickreg, variance in variance covariance regularization.
And the Cigarette stuff is really cool. So this is some work by Randall Bellister-Iroh, who was opposed to document me as a system professor at Brown now. And it basically consists in forcing the distribution of variables coming out of the encoder to be John Gaussian, essentially, sort of maximizing formation, if you want. It's just a very different way of doing it than, you know, what Yogan, Shmi-Duber, and Sue Becker, and Jeff Intern, what are we doing? And so this is super promising in my opinion. And we have variations of it, one that we can produce for our presentations, another one that can produce isotropic representations, but not necessarily Gaussians. And we have a paper with Randall and a student at Mila, Luca, Maes, that where we train a world model with this, it's a small scale, but we think it's super promising. So if you want to read one paper with that paper, it's the world model, L-E, world model. Awesome, we'll definitely link to it too. Yeah, I'm not responsible for the name. Randall picked up today. Amazing. Well, John, seriously, thank you so much. It is such a privilege to get to spend the last bit of time with you. And yes, really appreciate coming on the podcast. Thanks for having me, though. It's fun. I'm Jacob Efron, and this has been on supervised learning. A podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a night-to-week ends project in addition to my day job as an investor at Red Point. But our ability to get these incredible guests on really comes from folks like you, subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so please consider doing that. And thank you so much for your support and listening. We'll see you next episode.
Podcast Summary
Key Points:
Yann LeCun, a Turing Award winner and AI pioneer, argues that LLMs are useful but not a path to human-level or animal-like intelligence.
He founded a new company, AMI (Advanced Machine Intelligence), focused on world models and JEPA architectures for understanding the physical world.
LeCun left Meta because the company prioritized LLM-driven product development over exploratory research, despite support from Mark Zuckerberg for his long-term projects.
World models enable agents to predict consequences of actions and plan via search, unlike LLMs which rely on autoregressive token prediction without planning.
JEPA architectures, which predict representations rather than pixels, outperform generative models like VAEs and MAEs for learning visual representations.
LeCun criticizes imitation-based robotics (e.g., VLA models) for requiring massive data and lacking generalization, whereas world models could enable zero-shot task solving.
He highlights data efficiency gaps, noting humans can learn tasks like driving in hours, unlike current AI systems needing millions of examples.
Summary:
Yann LeCun, a Turing Award winner and AI pioneer, discusses his departure from Meta and his new startup, AMI, which focuses on world models and JEPA architectures. He clarifies that LLMs are valuable for language tasks but insufficient for achieving human-level or animal-like intelligence, as they lack the ability to predict consequences of actions or plan through search. LeCun explains that world models, which predict outcomes at abstract representation levels rather than pixels, are essential for intelligent agents to anticipate and plan actions.
He contrasts this with generative approaches, which he deems failures for representation learning, citing JEPA’s success over VAEs and MAEs. At Meta, he faced tension between exploratory research and the company’s pivot toward LLM-driven products, leading him to spin out AMI when results matured. He criticizes current robotics methods like VLA models, which rely on imitation learning and require enormous datasets, whereas world models could enable zero-shot generalization, as seen in humans learning to drive in hours.
LeCun emphasizes the need for data-efficient systems, arguing that synthetic data from video models won’t solve fundamental gaps. He remains optimistic about JEPA’s potential to revolutionize AI for real-world applications, though he acknowledges significant challenges ahead in scaling these architectures.
FAQs
AMI stands for Advanced Machine Intelligence and focuses on world models and scaling the JEPA architecture for AI in the real world, aiming to understand the physical world beyond language.
He says LLMs are useful for language tasks but lack the ability to predict consequences of actions and plan by search, which are essential for intelligent behavior in the real world.
A world model is a system that allows an agent to anticipate the consequences of its own actions, enabling planning and reasoning through search and optimization.
JEPA, or Joint Embedding Predictive Architecture, is a non-generative approach that predicts representations of inputs after corruption, which has proven more effective for learning image and video representations than predicting pixels.
He left because Meta became overly focused on LLMs and product engineering, deprioritizing exploratory research, and his projects on world models and JEPA needed a place to transition into product development.
Imitation learning requires enormous amounts of data for each task and is brittle, whereas a world model-based system could solve new tasks zero-shot with minimal data, like a teenager learning to drive in 20 hours.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.