Go back

Inside Pathway’s Brain-Like AI: Zuzanna Stamirowska on Continual Learning, Memory & Real-Time Reasoning

48m 50s

Inside Pathway’s Brain-Like AI: Zuzanna Stamirowska on Continual Learning, Memory & Real-Time Reasoning

The transcription discusses the development of a new AI model named Baby Dragon Hatchling (BDH) by Pathway, which aims to address the issue of memory in AI systems by integrating time-related features. The model draws inspiration from brain functions, particularly employing Hebbian learning principles to establish efficient connections between artificial neurons. The conversation delves into the unique approach of the BDH model, emphasizing continual learning, long-horizon reasoning, and adaptation over time to new data. By simulating brain-like functions through artificial neurons and synapses, the BDH model strives to achieve lifelong learning capabilities and maintain context effectively. The discussion reveals a focus on creating an efficient and powerful AI system that can adapt to new information and learn over time, drawing parallels between the model's design and the structure and functionality of the human brain.

Transcription

8882 Words, 47609 Characters

We believe we're on a faster way to AGI. Whenever two neurons were interested by something, the connection between them becomes stronger, and this is memory. We actually saw the emergence of just this kind of brain appearing. We can actually glue two separately trained models together, and they become one. Well, I remember we all rushed into the office, then I see the brain, and we was like, whoa. Welcome, humans, to the neuron. A.I. explain podcast. I'm Corey Knowles, and joined as always by my partner in crime here, Grant Harvey. How are you, Grant? Doing good, doing good. Thanks for having me. Of course, of course, quite the guest. You come every time. You just never stop showing up, right? I just got to do something different, you know, about a switch or not. Well, we have an incredibly fascinating guest today that we're both really excited about. Grant, you want to tell him about her? So we have invited Zizana Stema Roska, CEO of Pathway, one of the boldest challenges to the reigning transformer-based A.I. paradigm. And today, we dig into what live A.I. really means, why Pathway is banking on it, and whether this could be the next major architectural leap in A.I. Zizana, welcome to the neuron. We're so excited. Hello, hi, Corey. Hi, Grant. I mean, thank you so much for having me. Great to see you, guys. Well, I guess to get started, one of the first things that kind of stood out to us was, how did you go from studying at a French school for politicians to complexity science in A.I.? So have you guys seen that movie, The Beautiful Mind? Yes, I love it. There's this scene where, actually, he gets an ugly prize and they're like, oh, those people who bring him pens, right? And I remember my dad always cried at that scene. He was like, for him, it was like the most beautiful, romantic thing that he did. It was so funny. I mean, a big guy always crying. And then, I mean, actually, when I was studying at that, he was positive that that's cool for, you know, French kind of presidents or whatever. I actually went to Stockholm School of Economics and I got the chance to take a course in Game Theory and I think, oh, Game Theory, well, this sounds cool. You know, having seen John Nash in the movie and all of that. And I actually took a course in Game Theory and I remember I was sitting there, coming from a very different background than the other students in a way. And I just saw all the results of the games without doing the math. Wow. And actually, the guy who was teaching it was sitting on the Nobel Prize committee. So it was just an amazing course. It was just so beautiful. I became obsessed with it. I understood that, okay. That was like, I felt like fish and water. Finally, as if somebody, you know, finally showed me the real thing I should be doing that just felt so natural. And I said, okay, this is it. I mean, there is nothing else I can, or I should be doing in my life. At the same time, I was training in the kind of management consulting 'cause this is what folks do at Stockholm School of Economics. So I mean, I got a lot of exposure to all of this. But then I knew, okay, okay, how do I make it happen? And I guess I was like lucky enough, you know, to actually, I have met John Nash once. So that was, yeah. That's awesome. That is kind of cool. Well, what was the context or how did that happen? And there was a conference in Lisbon, and he was actually a speaker there. Oh, wow. Yeah. And then, and then I actually went to, I had an option actually to go to a called "Politea King" for like my masters, et cetera, and yeah. And then, so I did my master's specializing in Game Theory on Graphs. And Game Theory on Graphs actually very quickly evolves into complexity science. Once you do it, I mean, we have you know, small particles, big structures. It's more interesting, more fun if the structure keeps on changing. And then you try to play a game on like an infinitely changing structure that keeps on growing. I mean, this sounds tricky. It is. What, look, it's hard. Yeah. But of course, false, you know, for like a pretty long time where we're trying to crack it and kind of just bring it to some more universal levels of math, especially in particle physics and this sort of stuff. I mean, at the end of the day, you had small particles bumping, like doing something between them, right? Sometimes in space, bumping into each other, sometimes having connections like in the graph and like between neurons and you kind of send things over. And then this gives rise to like small folks doing something, you know, give rise to a society or like a big phenomenon or intelligence. I mean, you name it, but somehow as you get to the math, it starts to look somewhat similar, right? I mean, I think physicists and mathematicians would kill me for it to say that, you know, it's just all the same, same. But somehow you get those intuitions in some sort of like toolkits that help you to kind of put things at levels of abstraction that make it way less complex, you know, things that all the time start to look like just a sphere. And I just say that it's like, you know, in infant dimensions, but who cares? So it's like this here. So this kind of, yeah, doing a lot of graphs and then playing games on it and venting games, like this is just, you know, I was getting a kick from it. So I did my studies at the Colbert Dicking. This is like, very funny. It's also a military school, like engineering, military. This is where like for physicists, Plancaja, for example, like what? So, you know, a pretty big physics name, it's like that double prices go there pretty often. Yeah, and then actually did research there as well. And then, yeah, then I worked at the Institute of Complex Systems of Paris, where yeah, we're kind of, you know, trying to figure out how we get to global phenomena from very local small interactions between things. - And then how did you go there to founding pathway? - So you see, when you have a system, like the one that's growing and changing, right, you kind of need to have a notion of time. If things are to evolve and emerge, they kind of need time. If you look at, well, IT systems in general, then intelligence and kind of AI right now, especially, they're kind of deprived of the notion of time. - Right. - And this bugged us. So it's not just me, you know, I mean, I was fortunate enough to over the years seem to meet wonderful people like our CSO, Adrian Kosobskin, the guy had a piece, he had 20 rights. Like quantum physicists, mathematician, and the fear of the computer scientists, you know, top level, the guy is just crazy. Jan, who, you know, was at Google Brain. And then all of a sudden, all those guys, you know, are kind of jumping off the cliff, dropping 10 years to do this thing with me. This was, that's awesome. - Pretty, pretty cool. - So was time the core insight or problem you were trying to solve to the three of you? - Right now, all the models that we see that are out there are built on one type of architecture, one type of technology. And that was an absolute kind of algorithmic breakthrough. And this is a transformer. So the transformer was like fundamentally built for language. Funnily enough, one of the quarters of transformer actually was like the first check in pathway. But this, this technology is by definition deprived of the notion of time and memory. So pathway right now is building the first both transformer frontier model, which is tackling this fundamental problem of lack of memory in AI. Memory is linked to time, of course, because you remember, you need to remember things over time. You need to remember how you were thinking, how you were solving something, for example. You need to remember to see consequences, right? - Mm-hmm. - You need to remember to stay coherent while problem solving. And the more you know, I, the longer you can stay focused on a task, I mean, this means memory, that kind of course time, right? And we kind of know right now there is this lab called meter that actually kind of measures the benchmarks, the equivalence of, okay, like how the level of human tasks, let's say that the LLEMs can do with, let's say, 50% success rate. And right now, the left of those tasks is at like two hours, 17 minutes for GPT-5. So I mean, if you were like, we could say that kind LLEMs are kind of, we're living their groundhog day every day. So they are, they don't have memory as such. The way it works is that they're trained once with a lot of a lot of a lot of a lot of data, right? To the point that by now we know, we've exhausted all the data readily available on the internet for training. This is where they get their power from, because these are like fundamentally language models. - Yeah. - So they actually managed to produce something new that they didn't necessarily see in the training data, very explicitly, right? From having so many kind of samples of data. And then everything is like a relationship to everything else, right? - Yeah, exactly. So they've seen so much, and these are language models and they don't have memory. - And I guess what they're calling memory now is essentially something that's getting pushed into the system prompt each time you run a query. Is that correct? - Yes, correct. So right now when you have an LLEM and you kind of started, it's trained once, we have these like huge models and they'll tell you how many parameters, right? And the bigger the better, that need that would be usual, the kind of the vibe that they would be giving. And then every time you use it, it's as if the model was waking up and always using the same brain that was set during training. Training is very expensive because of data and compute because you need to produce a huge model. It really works better, of course. You ask your question. You add some context like I do your private document, like some whatever you're asking about, right? Then you get your answer. But then you relift the groundhog day every time. And you have a sort of memory, query, exactly, kind of as you said. But it's more like, as if you were leaving sticky notes for yourself from the data that you know you wouldn't remember. - Yeah, or the movie "Memento" where he's tattooing it on his arm. - Yeah, exactly, exactly. So it's like, you know, there's a big difference between having a library of knowledge versus having internalized it. - Yeah. - Such that you actually created framework space on this, that you can kind of adapt to new situations. - Okay, so this is a very big difference of having like evolving memory, contextualized evolving memory, that's your own, right? - Yeah. - And having a brain, let's say, which is set once. - So before we get into what you're doing differently than transformers based on that thesis or that problem set, let's say, do you think that there's a plateau on the meter chart where only transformer based language models can only get to like X hours of consistent tasks, like, do you think that there's like, and what hour mark would that be? Or day mark, what's that be? - So I guess right now we're kinda, you know, pushing, in fact it's not through LLMs per se that we are pushing and you would even have folks from OpenAI saying some things that the way forward is through reasoning, right? - Yeah, nice. - So how far can we get with reasoning? So I wouldn't, so reasoning is reasoning, it's even less related to transformer per se. - Sure. - So there, I wouldn't like put the burden necessarily, but just, I mean, given the math, like, the memory is not there. So it's difficult and kind of somehow tires some to actually try to trick transformer into having memory. So like, what I like to talk about is like epicycles. So you guys, like, before we had, you know, Copernicus and the proper theory of solar system, people were observing the moon. And to make sense of the observations are trying to kind of design some sort of orbit that would be maybe like this because that was the only way that they could explain the observations, right? It was like cumbersome. Pretty ugly if you think about this, but then every time they got a bit better, they were getting, you know, like, well, champagne or, you know, like, they would party. And the thing is, well, sometimes you see to swap things kind of around and then the orbit is actually just a little bit, right? It kind of looks good. - It starts to make sense when you switch the perspective. - Yeah, things kind of start to fall into place. So I mean, yes, we believe that there are just, you know, some things that we need to grow back to. I mean, transformer opens, like, it's an amazing, absolutely amazing innovation, which opened the entire market and actually it's done two things. One is a technological innovation, right? A scientific and technological innovation. Second, with the go to market that happened, it's managed to tickle the imaginations of everybody. - Yeah, yeah. - And this is huge for a scientific innovation, just think about this. - Oh, yeah. - But we are still early in this kind of AI market shift. So so far, I mean, 0.7% of GDP was spent on this AI technological shift. If you compare it to other such shifts that, you know, like, in the past century, I mean, just the telecom in the '90s took over 2% of GDP to be accomplished. - Wow. - And I'd say that probably AI is more fundamental, right? - Yeah. - So we're too. - So we're super early. And what transformers most likely not, I mean, as far as we would say, it's not the ultimate technology to get us all the way through it, and yeah, we need something else. But there's a lot to be done. - All right, so we have to ask about the name, Baby Dragon Hatchling, how did you come up with it? And I have a theory, but I want to hear the actual explanation and I'll tell you my theory. - This is like Dragon Hatchling. So in the paper, like specifically, it's Dragon Hatchling and the abbreviation is BDH. And of course, the question is where does the B come from? - Dragon Hatchling, precise, coming from theory projects, the color of magic. - I love it. - Okay. - Cool. I have not read that one actually. I need to. - I strongly recommend it. And I mean, this book, this book is where fly, at least 10 business books, I think, that you could find out there for so many different reasons, at dragons being probably even a smaller one. But I mean, specific in the color of magic, there are dragons that appear more the more you think about them. And it was actually very funny for, I mean, we found it funny for reasoning models. Because we literally had to reason about reasoning to build the reasoning architecture. Actually. - So the dragon started to appear? - Yeah, yeah, exactly. And well, it's like what you see publicly, right? It's an architecture in the paper. So this is a hatchling. And this is it. - I love it. - And then yeah, we do get some questions about why BDH and the truth is, and everybody tries to put something that would naturally fit, like the B, at the very simple truth is that, I mean, I just thought that, you know, A.I.D.3, like three-letter acronyms. - I agree. - They are also, you're not wrong. - Easy to pronounce and it worked. But I had one person, a physicist who came to our office, and like he said, listen, I read your entire paper, I read everything. And because I read it, I still, I think I still need to read the appendix, because I still don't know where the B is coming from. And you had to explain it to yourself. Like, well played, well played. Oh, I love it. - Is it, is it, so is the B because it's a small version, and you're gonna grow it? - So to be perfectly honest, it's like the most truthful explanation of A.I.D.3, just the three-letter acronyms. - Oh, got it. - The B per say comes from the fact that, like the model working on, like the working name is Baby Dragon. - Yeah. - So architecture is Dragon Hatchling, and then you have Baby Dragon, because it's already, you know, somehow grown. - Gotcha. - And it was just inherited. - Right. - I could imagine the B because our internal name was, was Baby Dragon. And then, you know, we do have some dragons flying around the lab. So, I love it. - We even have a random name generation for dragons. - Oh, that's amazing. - Like Indricoric, like how, how nerdy are we getting here? Like. - Ah, no, no, no, no. I, like we literally have an LLM, because we have versions whenever you have, you have versions, you're getting you to no testimonial, like one thing against the other and stuff. So we have, we literally have a random dragon names generator. - That's cool. - I love it. Yeah, my theory was that if this is truly continual learning, it's kind of a dragon in a sense, because it could be very powerful and dangerous, if we're not careful. But I imagine it's more like a dragon in Game of Thrones, where they're controllable. So I guess the question is, you know, one, we'd love to know how it works. And two, you know, if it is continual learning, how do you control something that's potentially powerful? - You actually hit something here as well, because part of why we venture towards dragons at all is that while this is a mythical creature, right? - Yeah. - That nobody believed it could exist. This is where, okay, we're in the business of building dragons. So I think it was a realization we had very early on. And indeed, we were talking about continual learning in time, long horizon reasoning and adaptation over time, right? - Right. - To new data, new learnings, et cetera. So the way it works is actually a bit like a brain that works on silicon. So a bit like a brain, I mean, there's a concept, this is very nerdy. The concept called a Hebbian learning. And like a very simple principles of how the brain works. So I guess all you all can imagine that the brain has neurons, which are like little cells, right? And then connections between neurons, which are synapses. And this is kind of it. So we have kind of dots or like neurons. And then links between the neurons. And this is like a big network that we have in our heads. And this is the very simple model, because you know, somebody who's like a neuroscientist or a biologist would say, well, yeah, but you have all those chemical reactions all this and that. I mean, here we're really just looking at the very stripped down very basic, very simple, you know, structure. Like we know that birds fly and they have wings, right? We don't need to know the how they really like move how many bones they have in them and stuff. Now it's just literally something to fly and to have like the surface of wings. And there we go. So we're looking at those kind of little neurons, little like almost like particles, that entities, that and they have legs between each other. And they, they're actually passing signals between them. - Oh, wow. - So this is, this is why you have connections. You have the structure, this structure. We know has to be dramatically efficient. - Right. - Why? Because well, our heads are somewhat limited in space. We walk on to feet and we kind of fall over. So our brains kind of get larger. - Yeah. - So it has to be very efficient. We know it is very efficient in terms of power. But it does offer this kind of capabilities of lifelong learning. - What are? - Keeping, keeping kind of very like infinite context pretty much. So we know that there exists a physical system that is capable of doing those kind of dragon-like things. - Yeah. - So I think it's not, it's not fully impossible. So this, this, we know. Question is how to make it work and especially how to make it work on the hardware that we have right now. - Right, yeah. - And like, yeah, this is, you always have to work with the hardware that you have with the materials that are possible whenever we see big technological shifts. I mean, it's usually some sort of inflection points where many things come together. I mean, so much compute with this algorithm. All of the sudden, this is, you know, gives us a boom. So what we did is we looked a little bit at Transformer and thought like, okay, what is it remissing to, like, from the brain, right? Like to get closer to the brain. And then, and then, yeah, that was actually Adrian, you know, our, our chief scientific officer who went, who went out this journey, like, literally, like, with very strong conviction that it has to be local interactions. Looking at the brain, we have to have those, like, small particles. And our model, like, our architecture, BDH, the way it works is that you really have small neurons. Neurons are connected whenever you're having a new bit of information, like, as folks call tokens, but you have some new information coming in. Only the neurons that are interested in connected light up. - Okay. - Mm-hmm. - So, when you're and gets information, passes it on to its neighbors, those with whom he's connected. Not everybody, not everybody lights up, just the neighbors. If, if they care enough about this thing, they light up as well. So, this is the principle of neurons, you know, who are connected and they be fired together. - Does that vary based on how connected it is? For example, the idea being that some information would be more tightly tied to point X? - Yeah. - Okay. - Yeah, and actually, but this structure emerges naturally. We don't set it. - Wow. - It just comes from data. It emerges, we actually saw in our lab, there was just, that this was an amazing moment. We actually saw the emerges of just this kind of brain appearing. - Wow. - Just like, whoa. I remember we all, it was like at late in the evening, and we all just rushed into into the office, like, because we were doing something else and then Adrian just called us, "Hey, look at this." And then I see the brain, and we was like, whoa. And I remember that kind of scary. - Well, I almost have chills thinking about that, like, it's just like learning that on its own, you know, - I remember, I immediately texted like, you know, like my brothers and kind of investors, and I look at this. - Oh my God. - Yeah. - That's awesome. - It was huge. I mean, you know, emergence also for complexities scientists, right? Emergence is what kind of love to see. This is a spontaneous order. You think that something's just so random doing the gods know what, but you can strip it down to such very simple fundamental rules that, yeah, you will see this like larger order appearing. And this is what we got, like, the structure of the brain somehow appearing naturally from those, like, very local, honestly, message passing between neurons. Like, as we do it on social networks, for example, we say something to our friends, right? Imagine this rumor spreading dynamics. This is how kind of learning works here. So whenever you're actually, whenever two neurons were interested by something, the connection between them becomes stronger. And this is memory. - Yeah. - That's right. - That's right. - Because that sort of like how the hippocampus works, right? Where it's like, I'm gonna do a terrible job. - Why are you point this, you use it more. It becomes stronger. I mean, this is, this is the principle. And it's only positive activation. So there's no like positive and negative. It's only positive it gets stronger. Something is not used, you know, over time it will start fading. But journey speaking, the connections that we're useful become stronger. And then I mean, this structure is, you know, it's actually very efficient because it's like a brain. So it's computationally efficient. It distributes nicely. It gives so many nice properties that unlock a number of things, you know, that then for us, in front of the engineers, then point in terms of how it scales, how it distributes, how you can run it on many machines, et cetera, et cetera. But it's like a scale free, so this is sorry, super geeky, but it is a scale free graph structure. So point is, even if we go beyond the scales that we've seen in data and tests, we scientifically know how it will behave. It's very different from the transformer, at least as we see it now, because for transform it hasn't been studied, it would probably be difficult to study. For this, because of, we know how the emergence works, we know that, I mean, yeah, it's scale free, same laws we'll be holding, you know, above, above what we've seen in tests and kind of data until. - Does that mean that it's also more interpretable at some level, like you can kind of understand what it's gonna do or now? - Yes, in the way. So specifically, we do see very precise the neural activity. And we, because we see the neurons when they care about something, right? We just see that. - Yeah. - Right, right. - So like for LLAMS right now, for transformers, I've also trying to build MRI machines to scan the brain. Or as we sort of have a CCTV inside of the brain. - I was just gonna ask you about that. I remembered you making a CCTV analogy when we talked recently. - What does it look, what does it look like? Like how, like, is it just about the numbers? - We see the neurons that fire up when they fire up on something. So in the paper, we're showing, you know, we're showing things that fire like synopsis and neurons that fire up for the notion of currency. And you have like a dollar one, right? You fire up, firing up. Of course there is, so two scientists who may be listening as an understanding system, there's compression of information while learning. So, you know, it's not always super clear. Like some concepts may be fairly large and you'll have like a lot of things fired up, for example. - Yeah, right. - But dreams be we see this neural activity. And we see even more, we see neurons getting bored. So when you keep on repeating something to them, we just see their activity, maybe whatever. Just going down and learning, like, think about this, especially as you go get older. I mean, kids get, learn very quickly, but also because everything is new to them. Everything is a surprise. It's worth learning. With us, we get older, like, things, you know, we stop thinking about many things. We stop noticing them because they're just so obvious. I mean, our connections are strong enough, let's say, for, for, you know, for that one thing, like eating soap or whatever. We shouldn't be eating soap, bad idea. So there's this element of surprise that actually somehow shows that something is valuable and worth remembering. So it was actually for us very funny to see this surprise effect, literally, on your activity. - So will it, in the same way that the brain over time, if there are areas that are not being used, that they can weaken? Will the same thing happen in a model? And if that's going to be used for a repeat task? - No, no. So I mean, yeah, actually, it says, so we would be getting some sort of like fading connections that they're not used for, for very long, but this is, this is more a topic of, okay, how to, how to transfer also to long-term memory, right? - Yeah. - So, yes, because there are some things that, again, it's not to work like a database, right? - Okay. - Because a database, then in deployment, this is something you plug in. - Right. - To, if you want to store absolutely everything forever, right? This is less of a problem. But for reasoning and having, let's say this, your space to explore when you reason, kinda. You wanna build it in such a way that you have the most relevant and connect compact structures. - Well, then I guess what I want to know now is, so you know, you've proven this BDH works at GBT2 scale last I read with 1 billion parameters, is that correct? What's the path to scaling it to say 100 billion parameters, what needs to happen to get there or, or to larger? - Of course, first of all, we do it there are no, no reasons actually for it not to scale and like scaling loads are inherited from the transformer. But there's also no big need to scale. This is not the game of scaling of more parameters and more data because this is currently not where the value is to come from. The value is to come from faster learning, how to solve problems that haven't been seen in the training data. I like this is what we, this is where we want to get to. And actually, if we can show better learning out of smaller data, well, this is the kind of value that we want to prove. So actually, I hope that very quickly, we'll be more looking at models that are very small but capable of producing results comparable to the big ones. - Love that, that's awesome. - We're not looking at scale and root for scaling. We're looking at this getting better at puzzle solving and reasoning and hopefully in a general way as possible to get it closer to the way that humans reason work and ultimately innovate. 'Cause if you look at the real innovator, like the best ones, I know, 'cause I kind of have them on the team, right? It's not about seeing what's there, but it's seeing what's not there and what could be there. - Let's say, for example, it's learning something really complicated, like where it's, I mean, you, your complexity scientist, you know more about this than I did, but something like really complicated. Does it at some point run out of brain power or how, like, 'cause I'm used to thinking of parameters as this thing that's like, oh, this is like, it can retain a lot more information. Like, and if this thing is just continuously running, at what point does it reach its limit of what it can think of or do we just not know that? - We don't know, but I don't think those limits work in this way. I don't think, I mean, right now, we have to do chapter number of humans that's like, the problem number that you may get there, you may imagine models where you would be adding them. This is a right, you could have a little that's growing technically. But it's like, right now in Transformers, it's not the reasoning power doesn't come so much from this size per say. - Okay. - When you think about this, we actually do have a lot of compute power because of how much place we have in the brain, because of this structure. So these models, but if you look like the brain now, I would need to check again, but I mean, for the synaptic connections, the brain are in the trillions. - Right. - And I think it could be, yeah, like 1,000, when some folks would say it could be, I think 1,000 trillions or something like this, of synaptic connections in the brain. This gives you a lot of memory and a very efficient structure. So effectively, you get, you get to something that operation works like infinite context. I mean, to be super scientifically precise, he has context and BDH is limited by the size of your brain. So the number of neurons and then connections between them. But this network structure allows you to encode so much. - Well, even the human brain fits in a serial bowl. Like, no, it's really not that big, considering the amount of information it could store and its speed and its capacity, I would think. - Exactly. And then when you think about this, you keep your memory close to the core. Actually, exactly about the core. So you don't need to do lookups for technical people. You don't spend your energy on all of that. You don't need additional compute for this. So it becomes very efficient from this point of view. You have a very direct lead there. And yeah, it's like in the memory on the chip. - Wow. - Yeah. And then second thing is that you don't fire up the entire model every time, but you fire up all need those guys who are connected and who care. - Right, that makes sense. - In some of those cases, yeah. You don't fire up the entire brain. I guess when you fire up the entire, I don't know how much brain people are using, but I mean, small percentage, right? - Yeah. - At any time. But that's exactly the point. Why would you be using the full thing? Not for every task, right? I don't need that to pick up a soda or to answer the door and say hello. There's definitely lightweight tasks for sure. - Exactly. You may have some that are more specialized in something. Cool thing, however. Like with our heads, what we're not capable of doing. So we can't exactly glue two braids together. It would be very difficult. And with those models right now, since they scale only with a number of neurons. So there's just this one dimension. It's something we're showing in the paper. We can actually glue two separately trained models together fairly easily. And they become one. - Ooh. - So we showed in the paper, we have like one model trained in one language, yet we want to train in another language. We just put them together, even without leaving them to train a bit. They actually start producing sentences that mix up the two languages pretty well. And then of course, the value will be allowing them to train together a bit more. Such that they create really like a one unity 'cause they also create links between each other. But yeah, this gluing is a bit like Lego blocks that you can put together. So you could imagine a model trained in finance and the other one in legal, you know, in an enterprise. And then you put them together. You get this like super, well, I don't know what the person with, if you know, good legal and finance partner would be doing, maybe we should. You should eventually venture there. Right. I guess I'm some superpowers that I mean, we would definitely want. It's like, you know, as I'm having Adrian, like it's a different story when you have, like, if I were to have like a physicist, a mathematician and a computer scientist on a team. Yeah. Sure, cool, but still communication, different intuitions and all of that. Versus having one person who kind of has all of this in one, you know, efficient structure, which is the brain. Something I'd like to ask in this little shift gears a little bit, but I know you already have some early adopters in people who are working with this, like I understand NATO, the French Postal Service, and maybe Formula One, is that correct? So these are like, we have history of actually working with pretty amazing accounts. Wow. But as of now, they don't have the models deployed, they have the dragons nest. Okay. You wish. So I mean, you have to understand that there is, if you are to bring life intelligence, you'd have life data as well. They need the way to connect. And you need your dragon to feel cozy, you know, instant environment. And so those guys, I mean, as as of now, they are using like layers of technology that we have built for the dragon nest to make sure that we actually can feed data at low latency very efficiently and actually do cool things that are necessary once you put such models and life intelligence in production. Yeah. This is like, of course, cases are pretty cool, right? I mean, sometimes I joke that, I mean, how are we exactly getting all the coolest customers? Like, do we have a cool, cool factor in kind of the generation or is there a lot of cool there? It's the dragon factor. I can't think of any one with more data than what Formula One collects. It's unreal. Our formula is very cool. And definitely the cool factor. Hi, it's really cool, right? And for them, it's like, from the strategy of the race, right, you have so many things that could go wrong. And every car is like a prototype, it can blow up. It's like, wow, so this is, of course, NATO. I can't talk about this freely. I mean, this is, this is very fundamental. I imagine that there is, if you played any sort of, you know, war game, like, I know, Harpoon or whatever, I just imagine getting real-time information from the field and then kind of, you know, informing the strategy and stuff. So this is, this is all cool. And then, you know, I know boys and girls like buses, ships, all of that. And I mean, you know, we have these complex systems of moving parts. And like, this is how the world works really, right? We have a lot of moving, living, things that are interconnected. And idea would like to predict what's happening out of it. And somehow see the patterns in this chaos. And hopefully, it'd be able to control it. And yeah, for this, you kind of need a number of technologies that come together and we do have a nest. That's the point. Our baby dragons have nests. That's awesome. What's the, what's the roadmap then? You know, are we, are we thinking we're going to be Lego blocking a bunch of different models together? I mean, but as far as the company goes, like, where do you see this going in production? Yeah. So actually, just like a week ago, we announced a partnership with NVIDIA and AWS at Reinvent. Yes, how good. So point is, the moment we're ready, it's going to be available to AWS customers kind of right the way all built on the infrastructure and kind of made in such a way that it will be easy to plug in and test it and adopt it. So this is, this is one. Awesome. I mean, this should happen, you know, sometime next year. So, so this is it. And then I mean, we have our own roadmap. I mean, we believe we're on the faster way to EGI and, you know, working towards getting their as fast as possible. Yeah. And I keep being all the focus in this. On a philosophical level, let's, let's go a different turn. Do you see this approach as, and you just mentioned AGI, AGI as a step toward general intelligence as a whole? What are any, are there any safeguards or things along those lines that you're thinking? You might need, as it begins to think more like a brain, operate like a brain. First of all, I think it's important to say that we're looking at reasoning and we see reasoning as a primary function of intelligence. Ella Lenzas, we've seen them for like literally language models, right? As we, as we've seen them for chat code use cases, summarization, some sort of like, search, I mean, these are great use cases and brought us here, but they are, they're not the primary function of intelligence. And we would say the reasoning as the primary function of intelligence. This is not just us, actually, I think by now everybody from all the labs would tell you the reasoning is, is this? So we're looking at reasoning, then with reasoning, capability of solving and, you know, inventing very tough problems. So the North Star is to get to an innovator who sees what's not there, as opposed to see what was there and recomposing, right? Yeah. So that would be the true generalization. And then you have generalization over time. So going towards solving problems and environments and over, like, periods of time that we're not tested, that we're not seen in data, et cetera. In terms of safety, I think I'm not, I'm not yet there to give very definite answers. But I'd say that part of what we're doing is actually gaining a way better scientific understanding of how the models are working and why. For us, mapping, like, very precisely mapping and understanding this passage from micro interactions and having something that we know we called the equations of reasoning, to having escape, restructure of known laws that govern it. This is kind of important, it's important for safety. Then some discussions that we have, let's say, more internally are around getting to provable risk levels, how such system will behave, and if it will be kind of do what I mean machine or not. Because the risks that we're looking at that are more controllable, maybe for us, is making sure that these models will not just, just by themselves, won't venture into doing something just completely silly because they're hallucinating, right? Yeah. So off-policy and it will be just ridiculous because they would stop working like a predictable human. Because if you think about this, if you're hiring someone, you observe them for a couple of hours. - Right? - Yeah. - Max, yeah. You see kind of see their credentials, you know how they were kind of taught, at Princeton probably, same curriculum. (laughing) - You know, you get them, and you give them tasks. And you safely assume that they will perform in a certain way, right? - Yes. - Without getting completely crazy in blowing up the planet Earth. - Yeah. - And at the wall. And this is a first assumption, and this is something we should get with AI, right? - Agreed. - So at least mathematically, we should have some sort of, you know, comfort and that, okay, we know how those Princeton graduates (laughing) function, and then, however, we still don't control the objectives, right? - Yeah. - So somebody puts AI to a bad objective. I mean, this is something that, I mean, as it started by four now, I mean, I have no influence over. - Yeah. - From. - And this is definitely something that, I mean, as we get closer, you know, to AGI, we'll have to be resolved at many levels. And I'm sure we, you know, we won't be the only voice to, to take part in this discussion. I hope, you know, we'll actually have something to say. - How do you prevent it from learning something you don't want it to learn? And maybe you're not at the point where you can do that yet. - Well, no, this is actually not too bad. So I think the easiest is that you can roll back to a checkpoint. - Nice. - So this is, this is fairly simple. - And you have observability that other, that traditional LLMs don't have, essentially, with your, with your CCTV? - Knowing that physicists were, I mean, pointed, it looks like an, like a cascade. Like a, it matches like an epidemic spreading in a graph, right, so like, you have a system and have an epidemic spreading. This is how information is spreading, right? So if you have a very small cascade, very small epidemics, you probably don't care because it wasn't relevant enough for the model. - Yeah. - Potentially. And, but if it's small, you can still reverse it. You can literally reverse it. - You could quarantine it as well, yeah. - But if it's big, this is, this is when, you know, also the information is very relevant. And you get to a point when you, like, you wouldn't be even able to say where this came from. If you were, if you were, you know, somewhere in this graph kind of looking around, you wouldn't be able to say where it came from. I mean, then it's non-reversible per se, just from kind of physics point of view. So, so, yeah, but you, you're all back to checkpoint and you could just checkpoint your kind of model over time. So you know, if you, if you infuse something, you didn't want to have the data and the model, you take it out. - What is the, one AI capability that you're most excited to unlock with BDH that isn't possible now? - With the current architecture. - Yeah. Yeah, with not, with transformers, specifically, yeah. - But this is very easy, like, generalization. - Yeah. - But, I mean, for, - The true generalization. - To a generalization. - It can kind of generalize, but. - No, true generalization. - Yeah. - So, so let's get the getting to innovator level. I mean, why, I mean, I think there is a path towards space exploration for us. And. - As in like your models in space? - No, I mean, although I do know a guy who's kind of working in TPUs in space, this is so cool. - That is cool. - That is cool. - Yeah. But. - You mean the space more abstractly? - Yeah, I mean, TPUs in space, like, you know, where we'll be building the centers, hey, in space. - That's a big discussion lately. - Yeah, it's a very serious thing, right? And I think timelines are pretty short. It's like, and within two years day, the day is when they will have actually TPUs in space. - Wow. - So, yeah, but I think like for, you know, to just like make space travel actually work, I will need so many scientific unlocks, and especially on the energy front. So there's like a couple of very fundamental problems in technology and science that we need to crack. And I see AI, in fact, you know, not as an end in itself, but as a crucial tool that will help us lift a number of obstacles that will then allow, you know. - Yeah. - I don't know, civilization to point out. So when I take two friends, you know, you can compare this transition to maybe the moment when people started agriculture. You know, it's like before we're kind of like, hurting and moving from places to places and then kind of humanity settled down. And then I was able to build culture and civilization. So like AI, the moment we become, you know, everywhere in this powerful will allow us to build civilization to point out. - You know, and we spend a lot of time talking about the next five years, the next 10 years, the next 20. And we forget that like, time's gonna go a long way. What's gonna happen? 100 years, a thousand years, 5,000 years. Where will humanity be as a result of what we're experiencing right now at that point? And that's, I can ramble on that for days. - Well, look at the pace. Yeah, sure. Anyway, look at the pace, right? - Yeah. - Just, I remember this time last year, folks asked me for example, about, oh, so what will happen in AI next year? And I was saying, well, reasoning models are going to be everywhere and they're gonna be the main thing. And people by then didn't exactly know what reasoning models were. - Yeah, correct. - They look like you're crazy when you say it, don't they? - Yeah, I mean, I'm happy my predictions were, but the thing is just that this space, right now the speed of transition, let's say of shift is like unprecedented. In hardware, you know, in a railway, that was a huge one that changed in the world. - Yeah. - I mean, you have huge infrastructural projects that they take just so much time to be built, right? And so I mean, it's like so, so slow. Whereas here, it's like from month to month, landscape is different. I mean, you know what better than anyone? Because, you know, your job is to regain an understanding that's translated, right? - Yeah. - So hard. - Yeah. - So many times, there's something that we think of as, as being way in the past. And then we go and we look at something and we're like, well, that was, that was eight weeks. That was eight weeks to go. - I didn't have those two months ago, you know. - That's right. I mean, at my PC defense, I had a guy from United Nations who was there. And he actually commented, commented and gave me, gave me a book saying, entitled that every time I found the meaning of life, they changed it. - So true. Well, Zana, thank you so much for joining us today. We were both so excited to meet you and look forward to watching your tech grow for many, many years to come. If, if our viewers wanna learn more, how do they go find Pathway? - Yeah, so I mean, go to our, please go to our website at pathway.com and then you can connect with me on Twitter and LinkedIn and then, yeah. I mean, these follow our research papers. So this is what we do, probably also gonna start some sort of, like, blog to get some of these messages kind of closer and maybe take them out of the paper because paper is pretty long and deep, but yeah, this is it. And guys, thank you so much, it was so much fun. - Excellent. - Thank you for having me. - Well, everyone that wraps up another episode, if you haven't yet, please take just a moment to like and subscribe today's video that way we can continue bringing you more cutting-edge interviews from some of the most fascinating people at the frontier of AI today. Also, don't forget to check out the Neuron newsletter that started all of this and joined 600,000 or so other people who read it every morning. Thank you again to Zuzanna and Pathway and that's it for today. Farewell for now, humans. (upbeat music)

Podcast Summary

Key Points:

  1. The conversation revolves around the development of a new AI model called Baby Dragon Hatchling (BDH).
  2. Pathway is building a model that addresses the lack of memory in AI by incorporating time-related features.
  3. The model is inspired by brain functions, using Hebbian learning principles to create efficient connections between artificial neurons.

Summary:

The transcription discusses the development of a new AI model named Baby Dragon Hatchling (BDH) by Pathway, which aims to address the issue of memory in AI systems by integrating time-related features. The model draws inspiration from brain functions, particularly employing Hebbian learning principles to establish efficient connections between artificial neurons. The conversation delves into the unique approach of the BDH model, emphasizing continual learning, long-horizon reasoning, and adaptation over time to new data.

By simulating brain-like functions through artificial neurons and synapses, the BDH model strives to achieve lifelong learning capabilities and maintain context effectively. The discussion reveals a focus on creating an efficient and powerful AI system that can adapt to new information and learn over time, drawing parallels between the model's design and the structure and functionality of the human brain.

FAQs

Pathway is addressing the fundamental issue of lack of memory in AI models, which is crucial for tasks requiring temporal context and coherence.

Pathway is developing the first transformer model that integrates memory, allowing for contextualized evolving memory that adapts to new situations, unlike traditional transformer models.

The name 'Baby Dragon Hatchling' was chosen to symbolize continual learning and long horizon reasoning, akin to building powerful and adaptable 'dragons' in AI.

Pathway's model is inspired by Hebbian learning, a concept from neuroscience that mirrors the brain's structure with neurons and synapses, enabling efficient lifelong learning and retaining infinite context.

Memory in AI models allows for tasks such as remembering past solutions, maintaining coherence, and focusing on a task for extended periods, enhancing overall performance and adaptability.

Pathway's model features interconnected entities similar to neurons and synapses in the brain, enabling the passing of signals and efficient lifelong learning capabilities.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.