Go back

Why the Brain Computes 1,000,000x More Efficiently Than A GPU: Unconventional AI's Naveen Rao

0m 0s

Why the Brain Computes 1,000,000x More Efficiently Than A GPU: Unconventional AI's Naveen Rao

Naveen Rao argues that current AI compute faces an imminent energy crisis, as global electricity capacity cannot sustain the exponential growth of AI inference and training. He notes that while the human brain uses only 20 watts, today’s AI systems consume gigawatts, and within a few years, we will hit a hard energy wall. The root cause is the 80-year-old digital abstraction—von Neumann architecture with floating-point math—which was never designed for intelligence. This paradigm burns most energy moving data between memory and processor, with minimal efficiency gains in recent years. Rao points to biology as an existence proof: a squirrel’s brain runs on milliwatts, yet outperforms supercomputers in complex tasks like jumping between branches. The key difference is that brains use non-linear dynamics and stochastic interactions, not deterministic matrix math. His company, Unconventional AI, is building a chip based on coupled oscillators with trainable couplings, where computation emerges from the physics of the system itself rather than explicit memory access. This allows state and function to overlap, eliminating the energy cost of data movement. A demo showed a generative model running on such dynamics, morphing between image classes without explicit programming. Rao concludes that by building these systems, we can approach the thermodynamic limits of computation—three orders of magnitude more efficient than today—and finally understand how brains work by recreating them synthetically.

Transcription

2926 Words, 15946 Characters

English
Why Current AI Compute Faces an Energy Crisis We're going to jump into our next set of Frontier talks. The 1st is which from Naveen Rao. Naveen is a pioneer in the AI space. He did his PhD in neuroscience. Then he started one of the first ever AI chip companies way before it was cool. He's actually the person who ran Mosaic ML, one of the first AI training companies, and then built all of Databricks AI. He left that amazing role to do something new and redefine the future of computing, bringing cool back to neuroscience. Naveen Kamano. Speaker 2 All right. Afternoon everyone yeah, super excited to be here. I'm Naveen Rao and CEO of unconventional AI. So we are unconventional because maybe it's actually the wrong word to use because I think we don't have to change the name to conventional. It's a it's actually a great time to be a start up as as Boris Boris alluded to like like having no baggage actually I think is a true competitive advantage. We can do things so much faster than traditional, you know, sort of chip companies and full stack companies can do. And I think that's what's very exciting. We can get to tape outs in months instead of years, things like this. So anyway, let's get get started. I put that up there. I can guarantee you some of you are like trying to prove me wrong. All of a sudden you're like, Oh my God, what's you know, I can hear the I can hear the years turning like that, that that can't be true. Well, let me, let me explain my logic a little bit and, and, and maybe the definition of what I call ASI. So really I think we need to get to much greater amount of compute efficiency. And when I say compute efficiency, I don't mean algorithmic compute efficiency or data efficiency. There's lots of people working on these problems. I actually mean the fundamental substrate, actually how I do information processing at the like physics level. We chose a path of making computers work the way they did for a lot of reasons about 80 years ago. If you think about it like in the tech industry, how many things have existed for 80 years? Like not a lot. The digital abstraction, sort of floating point numbers, those were around from from the 1940s. And for a machine that was built from a completely different substrate and for a machine that was built for a completely different purpose. And now we're building machines for intelligence. So let's think through this a little bit. So part of this is that AI will make us more efficient. I mean, we've talked about coding and, you know, running 1000 agents on your phone, all this kind of stuff, right? It does make us more efficient. But at some point we start to get to a place where what, what, what does efficiency really mean? If you think about energy, actual energy, maybe you're actually not more efficient. And we're kind of butting up against those limitations of the physical world. Now, you know, already today we're we're using many gigawatts for AI inference and training. And we're going to get to a point within the next couple of years, this is not ten years away. This is 234 years where we just don't have any more energy in the world for AI. Then this topic starts to become very, very important right now. Comparing Brain Efficiency to GPU's Thermodynamic Limits You can kind of look at it like, well, electrical energy versus like food. These are two energy sources for intelligence. And you know, there's no real limitation on the on the electrical side, but that's going to hit a pretty solid wall very soon. And you know, we're talking about going to space, we're talking about building fusion reactors. Great, let's do all those things. But still these fundamental physics apply. So if we think through this a little bit, we have about 8 billion people in the world. Our brains use about 20 watts each. It's only 160 gigawatts. It's the entirety of humanity is 160 gigawatts. So just put that into comparison. We have about 9000 gigawatts of capacity in the world today. the US has about 1000. And you know, this runs everything. This is like heating in your home and you know, all the things, electric cars, all that kind of stuff. But if we said we, you know, maybe you got 50% more of this and we say, OK, great, now we got like, you know, 4000 plus gigawatts. But the problem is our current paradigm of compute is just vastly more inefficient. So a computer, I'm, I'm making some numbers up here. I mean, I could say like if I'm running inference per token, I can come up with numbers, but you know, fully loaded, like the amount of energy that goes into inference, building the model, running the model, all of that, let's call it a GW, something like that. Or, you know, definitely in the MW range, but humans are on the order of of 20 watts. And you could also argue that evolution over 4 billion years is created what we are. But the reality is today our constraint becomes how quickly can we get learning to happen? How quickly can we build intelligence on a given amount of energy? So if we want to build this future where we have lots of intelligence in the world and we're automating all kinds of things that we want to really be more efficient from an energy standpoint, we're we're going to need a lot more, we're going to need a lot more watts. Or we can think about building a vastly more power efficient computer. And that's where we come in. So I love this curve here. And I think most people really haven't really haven't thought about this because you just sort of assume the computer is a computer. And we haven't really questioned that. Again, this is the unconventional part is like, let's break that apart. The, the, the assumptions we made 80 years ago are actually not quite valid anymore. We just choose to keep building on them because I can build a product in two years, I can make something that I can sell in two years. But we're kind of taking a different tact, like let's go back to those first principles and see if we can build something much, much better. So there is a thermodynamic limit to intelligence for what? OK, that means you just can't do any better. There's there's something called the Landauer principle, which some of you may know, which basically suggests how much compute could happen within a certain amount of energy. So there is a physical reality that we can't, we can't get past. And that's sort of this asymptote here. Now biology is somewhere up here. It's actually pretty darn efficient for a billion years of evolution have created something that is actually very efficient. However, it's not at the asymptote yet. There's probably an order of two of magnitude between those two. We're actually here, by the way, we're down here. And I think limits of 2D lithography, that's what chips are built on today is call it somewhere down here. And I think with focused effort we can get to the point where we're pushing the limit of that. And so we'll have something that's, this is this, by the way, is something like 3 orders of magnitude from where we are. It is very far away from where we where we could be in terms of energy efficiency. And so really that's what we're focused on today. Unlocking Efficient Compute Through Biology's Non-Linear Dynamics So how do we do it? I mean, yeah, this is great. Make something more power efficient and wonderful, right? But the reality is we it's not so simple. We can't be thinking about the computer in exactly the same way as we have been. It's not about a machine that runs off of runs matrix math. That's been the simple way to move forward. NVIDIA, of course, has owned that market and continue to push the envelope. But if you look at the power efficiency numbers, actual power efficiency of delivering an FP8 flop, for instance, it's not that much better. Costs have gotten better because manufacturing has gotten better, Our ability to package has gotten better. But actual energy per flop with memory access has not gotten better. It's very, very incremental now. So I am a a neuroscientist. I was a computer architect for 10 years before that. So I've been thinking about this problem for a long time, actually on the order of 30 years. So it's a very exciting time for me personally. And you know, biology really does provide an existence proof. I mean, you can argue that, OK, the tokens per second out of a human are lower than the machine, but the intelligence is higher. We still haven't gotten to the point with these gigawatts that we're throwing at it that we're we're rivaling a human's intelligence in terms of discovery. We'll get there. We're going to get there in a very short amount of time, but it's going to come at the cost of a lot of energy. So what I think is most interesting here is actually not just that brains are human brains are 20 watts, but that the the waters kind of scales with the with the weight. A macaque monkey's brain is probably less than a Watt. And actually you see this all through the mammalian world, not also the insect world. Like you have very complex behavior for milliwatts. Just for reference, your phone in your pocket is about 1 Watt. So you know a squirrel jumping from branch to branch is running on a less than 10 milliwatts. That's one 100th of your phone. We can't actually do this perfectly. Like, you know, you know, squirrels jumping 10 feet across from branch to branch within wind and all that stuff. We can't do that with a much, much larger computer. So biology still created something quite amazing. And and I just, I just don't feel like there's been an appreciation for that. So I'm just, you know, just just just a little reminder there. So now great. We we see this kind of phenomenologically like biology is efficient, can do, can do amazing things, but how does it actually work? We don't really know. I will be honest as a computer scientist and a neuroscientist, but there are some ideas that we can harvest from neuroscience and one of them is that the brain is dynamic. It does not use matrix math to do compute. It uses what's called non linear dynamics to do compute. What this means is that there's a time varying interaction between neurons and that's actually where the compute lies. So can we extract that and actually apply it to to synthetic circuits Maybe. They don't do floating point math, they don't do matrix math. They do something that can be characterized as such, but it's actually much richer than that because of these non linear dynamics and they're stochastic brains. Compute is not a strict 1 and 0 in a digital computer. If we're off by a one or a 0, the whole system falls apart. So really not computers. So I'm going to try to go through this quickly. This is the thing called a Kuramoto synchronization. So if you look at a bunch of oscillators here and they're kind of rigidly coupled to each other on this plank, you know, you'll see over time that they start off in any state and then they actually synchronize. This is an example of a, of a, of a contracting or, or, or converging dynamical system. So no matter how you start it, it converges and it's only based upon the coupling between them. Well, you can generalize this to something that has kind of a flexible coupling, call it a trainable coupling between those those things. And then it can have all kinds of interesting dynamics. It can move through the state space of dynamics in many, many different ways. So if you generalize this, you can actually think about an Electro as an electronic circuit. You can say I have a bunch of oscillators and they have a fabric on which they're coupled. And now when I when I make this fabric trainable, I can actually see something that's it starts to look a little bit more like the dynamics of the brain. It actually has non linearities and they interact with each other in very interesting ways. It's actually very rich and represent a lot of information. Prototype Chip, Demo, and the New Computing Paradigm This is actual chip that we're going to be building this summer. So we went from basically no team in January to A to a full prototype in six months. And that's because of AI. So this is what's probably cool about not having baggage is you can do things in completely different ways. And the way you compete with something like this is the traditional way would be basically loop over some sort of linearized time. This is how we do things in a in a Von Neumann machine. We write state out, we retrieve it, we operate on it, we write it back. So we keep going back and forth. Turns out that's what burns most of the energy in the existing computing system with something with non linear dynamics. I actually just say, here's the initial state, kick it and let it run. So the, the the physics themselves basically do this computation and it doesn't sort of the the state is an implicit, it's not an explicit rights. So in some ways you can think about it, if you take anything from this talk that we use the time, the time access of the physics to do computing and existing competing constructs do not. And so the question then is, can I train this? And the answer is yes, I can actually steer the system into multiple, multiple different things. In fact, we're sort of traced out in state space, A unconventional logo. That's the idea here. We can train in a few different ways. So yes, we can train these systems and steer them into basically any arbitrary set of trajectories. And can we compute, can we connect it to AI problems like image generation? So actually I have a better version of this I can go to in a just a really quick demo here. Let's go to the next one for the demo. Yeah. So basically what you see here is something that's running on a on a model of dynamics that was trained on these different images. So basically I can say, OK, I have to use cats. I think Andrew Ng is here, so homage to him. But we can do anything. But basically this is a pretty simple generative model. And I can basically say like, OK, at time t = 1, I'm going to, I'm going to backfrop an error from randomness to, to a particular image class. And after that point, we let the, we let the system just run naturally. And you, what you'll find is that it actually has clumped its representation into places that are meaningful. They're no longer just random pixels, but they're actually pixel pixels that make different kind of machine or different kind of animals or whatever. So for horses, I start off as random and at t = 1 you should see it kind of converge into horse like things. And then over time you'll actually see it sort of morph between those. So it's already learned in the state space that it can move between these different things. So let's go and move out of this. So this is really the emergence of something new. So CPUs actually do very fast single threaded things the best. Even today. It's faster than a GPU. And really what you're doing is this kind of Von Neumann machine where you're moving in and out of memory and cache and doing operations. GPU basically did this with multiple operands at once. So we move a bunch of operands from memory, do some stuff to it, write it back computing memory. Like Grok sort of did the same thing, but just did it on chip. It's kind of a a more fine grained version of this. And what we're talking about is doing something in a dynamical system. The state and the the function are overlapped with the physics themselves. So you now have no separation between state and computation. And you know, computer efficiency goes goes up. Of course, Galaxy brainness goes up. And this is truly non von Neumann. So with that, I'm just going to leave you with this quote. It's been something I've guided my entire life by. And I'm I'm really excited about this time because I've been thinking about this problem for 30 years. And we're at this point where I think we can actually start to understand how brains work because now we can build them. Thank you.

Podcast Summary

Key Points:

  1. Current AI compute faces an energy crisis, with AI already consuming gigawatts and projected to hit global energy limits within 2–4 years.
  2. The human brain operates on ~20 watts, while AI systems require vastly more energy for comparable intelligence, highlighting a huge efficiency gap.
  3. Traditional digital computing (based on 80-year-old von Neumann architecture) is fundamentally inefficient for AI, with energy per flop improving very slowly.
  4. Biology uses non-linear dynamics and stochastic processes, not matrix math, to compute efficiently at milliwatt levels (e.g., squirrel brain <10 mW).
  5. Unconventional AI is building a prototype chip using coupled oscillators and trainable dynamics, achieving compute without separate memory/state storage.
  6. The new paradigm leverages time as a computational axis, enabling implicit state evolution and drastically lower energy use.

Summary:

Naveen Rao argues that current AI compute faces an imminent energy crisis, as global electricity capacity cannot sustain the exponential growth of AI inference and training. He notes that while the human brain uses only 20 watts, today’s AI systems consume gigawatts, and within a few years, we will hit a hard energy wall. The root cause is the 80-year-old digital abstraction—von Neumann architecture with floating-point math—which was never designed for intelligence. This paradigm burns most energy moving data between memory and processor, with minimal efficiency gains in recent years.

Rao points to biology as an existence proof: a squirrel’s brain runs on milliwatts, yet outperforms supercomputers in complex tasks like jumping between branches. The key difference is that brains use non-linear dynamics and stochastic interactions, not deterministic matrix math. His company, Unconventional AI, is building a chip based on coupled oscillators with trainable couplings, where computation emerges from the physics of the system itself rather than explicit memory access. This allows state and function to overlap, eliminating the energy cost of data movement. A demo showed a generative model running on such dynamics, morphing between image classes without explicit programming. Rao concludes that by building these systems, we can approach the thermodynamic limits of computation—three orders of magnitude more efficient than today—and finally understand how brains work by recreating them synthetically.

FAQs

Non-linear dynamics refers to time-varying interactions between neurons where computation emerges from their evolving relationships, unlike matrix math which performs fixed operations on stored data. This allows the brain to process information implicitly without explicit state writes, saving energy.

They leveraged AI tools to accelerate design and tape-out processes, avoiding the baggage of traditional chip companies that take years. Starting from no team in January, they reached a full prototype by summer.

The Kuramoto model shows how coupled oscillators naturally synchronize over time based on their couplings. By making those couplings trainable, the system can be steered through state space to perform computations, mimicking brain-like dynamics.

Yes, the speaker demonstrated a generative model trained to produce images of cats and horses. Starting from random noise, the system converges into meaningful pixel patterns representing those animals, and can morph between them over time.

A squirrel operates on less than 10 milliwatts (0.01 watts) while performing complex tasks like jumping between branches, whereas a smartphone uses about 1 watt. The squirrel achieves far more sophisticated behavior with 100 times less power.

The Landauer principle defines the thermodynamic minimum energy required for a computation. Biology is within about one to two orders of magnitude of this limit, while current digital computers are about three orders of magnitude less efficient.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.