Go back

Vitaly Vanchurin: The Universe Is a Neural Network That Learns

118m 32s

Vitaly Vanchurin: The Universe Is a Neural Network That Learns

The discussion centers on modeling the universe as a neural network, where the learning process itself generates physical laws. Professor Vitaly Vanchurin proposes that the universe's dynamics can be described through neural network frameworks, with optimization algorithms like stochastic gradient descent and Adam naturally producing phenomena such as curved spacetime, which enhances learning efficiency. This approach leads to the emergence of equations like those from general relativity and quantum mechanics. Importantly, he clarifies this as a mathematical model rather than a claim about ultimate reality. The conversation explores implications for consciousness, observer physics, and the unification of fundamental theories, positioning learning and optimization as core principles that differentiate this model from traditional physics, which typically lacks such inherent objective functions and learning trajectories.

Transcription

18703 Words, 104669 Characters

English
The universe is self-tuning itself. It likes to be observed. And so observers emerge not because there's carefully chosen constants of nature, but because if they were not carefully chosen, then they would be learned to evolve towards being carefully chosen. Five years ago, an unintuitive and startling result was dropped like a bombshell. Professor Vitaly Venturin of Cosmology found a way to model the universe as a neural network where the learning dynamics are the physics. This has huge implications for what the cosmos is, what you are, and potentially what consciousness is, and its relationship to everything. As you'll see in this conversation, this is not another way of saying that you can use neural networks to simulate general relativity or the standard model. That's been done. Instead, the professor shows that the universe's own learning is the physics. What happens is gravity falls out. The Dirac equation falls out, Klein Gordon falls out. The algorithm behind most modern AI, copacetically named the atom optimizer, implicitly carries a curved metric on parameter space. The presence of the curved space is essential. Essential for convergence. Space time curvature is actually there precisely because it makes the universe as learning efficient. This conversation spans natural selection at the subatomic scale. The Boltzmann brain paradox, Carl Friston's free energy principle, consciousness as learning efficiency, and the ramshackle state of observer physics, which Venturin argues demands a three-way unification of quantum mechanics, general relativity, and observers. My name is Kurt Gimongele, and on this channel, I interview researchers about their theories of reality with rigor and technical depth, even at the risk of limiting the audience, because this slow, meticulous, candid approach is superior to a fast, flashy, potentially misleading approach. The universe is a black box, but today, Venturin opens it. We'll definitely have a part two, so leave your questions in the comment section. There's plenty more to explore. Enjoy today's episode of "Theories of Everything." Professor, you claim the universe is literally a neural net, so not that it's a useful model, it literally is ontologically, justify yourself, young man. Okay, not so young anymore, but I'll try to do my best. Now, when you're saying that, I claim the universe is a neural network, and not just immortal. If I did say that at some point, I want to take this back. As a physicist, I am not, as a theoretical physicist, or as a physicist, I'm not allowed to say what the universe actually is. What I'm allowed to say is what is a good way to model it, because at the end of the day, I cannot really know or check or test, prove or disprove, whether this is how the universe works. But I can test and check whether any given model, mathematical model, is good for modeling, you know, sometimes a phenomenon. So if I did say that at some point, or somebody misinterpreted, now I always talking about what the good model of describing it. And at that point, yeah, I have to say it looks like it's a promising candidate. It's not a final, there's no final verdict yet. But it's a promising candidate that should be explored, whether it is a good way, convenient way, compact way of describing phenomena in the universe using neural networks. And then perhaps how exactly I want to I want to do it, you know, we can discuss later, but I just want to, you know, open all my cards and say, I would never claim that this is how the universe works. Now, what if we get to philosophical questions, of course, we can say, what if, right, what does it mean? This is really the universe, you know, how the universe is, what kind of philosophical conclusion we can reach out of that. But as a physicist with my physicist head on, I can only say this is a interesting good model and it works remarkably well in the places where I wouldn't expect it to work well. People who know some things about neural nets know that their universal function approximators. So why would it be surprising that neural nets can satisfy functions, which the universe is described by functions, like it would be more surprising if you had the counter claim, the universe cannot be modeled by a neural net. And then the next question, okay, and actually now I can actually put my finger and exactly what I mean, right, it is true, right, why should we be surprised, you know, neural networks are universal approximators, why should be surprised that you can use neural networks well trained neural networks to, you know, reproduce the dynamics that we have that's not so obvious, but at least for classical systems, yeah, I would say why we should be surprised, but the difference that I'm proposing is that I'm not only saying that the trained network is good at describing a given function or a given dynamics, but actually the process of train, the process of learning is a part of dynamics. So it is different, right, so now I it's not for me allowed enough to just show you here it is, there's a train network it describes well harmonica solider, no, because to train it, I've used trainable variables, I use some kind of learning algorithms and and that dynamics is not just didn't disappear, it must be the it must be part of me telling you why harmonica solider can be described. By neural network so if we remove learning it's almost trivial state if we say no, no, no, let's talk about the entire dynamical system of neural network that comes with learning dynamics that comes with activation dynamics and can that thing together combine thing can that. So it's a useful model for for describing the universe. So I just took the most popular stochastic gradient descent and let's just see it looks like the one that use little resources and still doesn't make an amazing thing for me was interesting that even a very simple algorithm can produce behaviors that you know we physicists don't have tools to describe. Just because it's a learning dinar so so that was like five years ago and that was that was all and and it was already some some very spectacular result would come out of this. With the collaborators we've showed that you know quantum behavior can emerge we can discuss all that later now now I've learned more I knew about this but now I actually understood more that it isn't just stochastic gradient descent that is. Interesting and important for modeling the phenomena but there are other well known algorithms that anybody working machine learning know about and use them daily. Adam Adam optimization optimizer this is one example and and it works for many many problems it works much better it has much better learning efficiency you know you train your model and loss function goes down much faster right. So so I wanted to understand more recently why is that the case what's what's the physical reason what is the physical reason why it works better but but now it isn't just a cast a gradian and sent but it's a whole class of learning algorithms we call it a caverian gradient descent caverian comes from the physics definition of caverians that we can again discuss. And those algorithms Adam like algorithms and their generalizations give something something I haven't again expected there's like always study something machine learning you don't expect and you get it and so and in particular the emergence of curved curved space space and space time comes naturally if you are actually thinking about. And so the presence of metric there you know whether the people in machine learning community know that there is a curves metric or don't know about this. And so the first of the curved space and and is essential essential for convergence essential for all going to be to be efficient and so so yes and I originally it was the classic gradient descent but there's new things come up like did every month that I I learned about. And so the audience of this podcast are researchers and computer science but also researchers and physics and math and and that's more on the hardcore stem side but then there's a large swath of artists and miscellaneous lay people it's quite interesting because it's one lump here that's quite hardcore and then another lump on the soft core let's call side. and it's an honor. interesting that the overlap is small. It just goes extremely nitty gritty PhD level or much more, much more layman wondering about ontology and philosophy and so forth. I forgot to mention there's also researchers in philosophy. Okay. To those who are not computer scientists, a neural net is what? What is the minimum someone needs to know? Right. So I mentioned it earlier, but let me just discuss it one more time. The neural networks comes with this one feature that even I as a physicist wouldn't know as a research in physics wouldn't know. And that's the learning dynamics and that's the dynamics that is is taking some function that machine learning researchers call loss function. Some cost function, you can call it. Now function that you're trying to optimize. What is it that you don't have to be a scientist or a researcher to understand that there is some kind of optimization. There is something that you optimize. And so the neural network dynamics comes with that something that goal, you can say, you know, what's the goal of the system? What is it trying to do? You know, it's trying to speak English language without mistakes or texture or it's trying to do speech to text recognition. Whatever it is trying to do, this is the difference. The presence of that objective function, this is the loss function, something that is essential and that is, well, I'll stress once again, it isn't something where I use to in physics. And so and and it is something that machine learning people are used in machine learning research, but not us. So I had to kind of try to get all of the nice experimental results in America's obtained from machine learning, try to, you know, use our toolbox that we use in physics and try to understand it. But yes, we are trying to tell this story to people who are not, you know, running models every day or writing equations every day. Then that's the difference. That's the difference. So you have you have a system that has this one boring dynamics that we know about before activation dynamics. There's some state that keeps changing according to some some law. And then there is this learning dynamics that there is some objective function that the system is trying to optimize. So the presence of those, you know, two things is essential. Now for neural net. Okay. Well, firstly, optimization physicists do know about it. If they're doing any minimization of the Lagrangian or or extremal point of the Lagrangian. So is there something particular about the way that the technique of optimization from neural nets compared to other optimizations that's well suited for describing the fundamental laws and such generality that you've been able to find out? Good point. Yeah. So we do use variation, variational principle, right? So we study the extrema of Lagrangians, right? Action actually, right? So so we are interested to find certain solutions. We take this beast, which is called action, to vary it with respect to degrees of freedom. And we are interested in its minimum maxima. Now what is new here is that you are not only interested in minimum maxima. You are interested in the entire, you know, trajectory from whatever you started with to whatever complicated state you're going to get. And that complicated state will certainly satisfy some variational principle in some sense. And and that's where you will kind of because of that, see the emergence of some kind of classical like behavior. But but even out of this equilibrium, right? There is this whole evolution that takes you learning optimization evolution to reach the the the the the the the the the the the the the the the the the the that is present in optimizing machine learning systems and isn't present in physics. Now I have to correct a little bit because you know, right now physicists adapt machine learning to solve lots of problems as a tool. Okay. Not as a model of physics, but as a tool, right? Yes. So let's say you have a very complex quantum, many body system. You're trying to find its ground state. You're doing some difficult problem. And so of course, you're going to use all the tools there are all the computational tools there are including using machine learn. Now, when I'm saying that a physicist aren't used to optimization as a model of of of a system that they're trying to say, but as a tool, absolutely, this is this is a great tool. And it's been used by physicist and you use by system more and more now. What is the input into this neural net? Right. Okay. So if we are talking about the neural network of the as a model of the entire universe, let's say, then there is that all there is. This is the status. You know, you describe the state of all neurons, describe the state of all connection weights. And that's the state of the system. That's your input. This is your initial state. So in physicists, we actually have a very nice set up of modeling everything. We say, well, you need two things. You need to know the state and how to host. Right. So now, quite a mechanics again, putting aside, those are the only two things you need. So the state or the input of this neural network is the state of all neurons. And then and the state of all of the trainable, trainable and non trainable variables. And then they evolve according to the one side activation dynamics and the other side. So this state. So this setup that physicists came up with, there's actually mathematicians have a much more general setup, dynamical system setup. Right. Then they don't even bother whether the dynamics is Hamiltonian or, you know, satisfies some kind of constraints. There's like energy like function. They don't care about that. So what I'm talking about, it is a dynamical system. But it isn't a dynamical system in a sense where you restrict yourself to classical hematonic like dynamics. In traditional physics, the input may be the state at time zero. And then the output may be the state of time T in neural nets. Let me just talk about an image classifier. So an image, let's say you give an image of a dog and it's just a square image. And maybe it's 20 by 20 pixels and so 400 pixels. Then you have 400 numbers as your input. I mean, if it's a grayscale and then at the end, you want to know, is it a dog? Is it a cat? Is it a flower? What have you? So however many categories you have here is your output. So what is the input on this side? Is it the whole state of the whole universe? And then the outputs, the whole state of the whole universe again. Very good. What is it? No, no, but very good question again. This is, so now you have your, well, it's tip will keep switching hats. So now you put your washing. And learning had on and said, OK, here's a common understanding what you're talking about. And I'm saying no. So in this sense, the entire network with the input and the output. Before you even started propagating your image through and figuring out whether it's a cat or a dog, this whole thing, the state of all of all of the degrees of freedom is the state of the system, not just input, the whole thing. Now in the case of them, of the cats and dogs classifier, it happened to be that there is in your problem, there is a clear distinction. What is your calling input and what is it that you're calling out? Right. So there is a kind of flow of information in that in this direction. But this is just because you set up your network that way. You didn't have to do. You could have used the recurrent networks. You could have used a lot more complicated loss functions. So for example, in this case, your loss function would be, well, did I get it right? Is it a dog or zero or one at the end? But it's, if it's just you're talking about restricted class or machine learning problems. And we, in this case, the information really flows in one way. Now if we have the entire network, entire universe, the scribe or neural network, it may happen then at some place, there is like only left going wave or right going wave. Where information only goes, goes one direction. But that's because of there, I can finish all conditions that you set up, not because your network cannot start up with some other states. And so imagine in your example with a cat's a dog, imagine that the zero and one that you got at the end, you look it back to the input, right? And then now it may not do something that you wanted to do, but it will run. You will get this, you know, pixel change, changing one of the pixel in what you call input. And then, and then, and then going through. So in this case, still the whole thing is is is input at previous time step. So it probably better call it the state of the system in the previous time step. And then once, you know, one step of activation took place in the next time step. And then another step of activation to the third time step and so forth. And that's kind of time evolution of the activation dynamics. And then there is learning, right? So then they have to upgrade your weights, which is, which is again, and you can just keep going, keep activating and and learning. When I'm wrestling with a guest argument about say the hard problems. Consciousness or quantum foundations, I refuse to let even a centilla of confusion remain unexamined. Clawed is my thinking partner here. Actually, they just released something major, which is Clawed Opus 4.6, a state of the art model. Clawed is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow, thinks with you, not for you, whether you're debugging code at midnight or strategizing your next business move. Clawed extends your thinking to tackle problems that matter to you. I use Clawed actually live right here during this interview with Eva Miranda. That's actually a feature called artifacts and none of the other LLM providers have something that even comes close to rivaling it. Clawed handles interalia, technical philosophy, mathematical rigor and deep research synthesis, all without producing slavently reasoning. The responses are decorous, precise, well structured, never sick of phantoms, unlike some other models. And it doesn't just hand me the answers. The way that I've prompted it is that it helps me think through problems. Ready to tackle bigger problems? Get started with Clawed today at clod.ai/theoriesofeverything. That's clod.ai/theoriesof everything and check out Clawed Pro, which includes access to all of the features mentioned in today's episode. Some of the key equations in physics are general relativity, so Einstein's field equations or Dirac or Klein Gordon. I know that you're not able to with your words say how you derive them exactly in such a way that is rigorous, but we can of course point to your papers and lectures on screen right now. But either way, can you just walk us through as much as you can with your words as to what you started as your input and how were you able to get these as outputs? Sure. So let's start with the field theory. So we'll know very well that the standard model standard model of high energy particles, high energy physics is where we'll describe by the you know, collection of fields. And so if you want to get that physics out of your framework, mathematical form work, you want to show how you know fields fields will emerge, how we would get fields out of it. Now it's it's a it's a difficult task. So let me just put it right away. And it's not something that I can say, well, here it is. I get you know, quarts, three generations, I get everything and it's simple and I can like write one paper and go home. No, no, it's not not even close. It took years to get Dirac equation out of it. Okay. So Klein Gordon was easier. Hamiltonian mechanics was easier. Getting getting fermions, getting Dirac equation, it turned out to be to be a difficult task. So just since we talked about this direction of information flow, it turns out that for the direct field, there's some tensor factor, something in your neural network setup has to have an anti-symmetry in it. So it has to be anti-symmetric. And so if you put that in, if you put this constraint, now, why would you put this constraint? I don't know, is that something that this constraint was learned because of that some kind of microscopic optimization optimization, algorithm that's running great. Can I show it? No, what I can show if I assume a certain constraint, if I only take into account certain trainable degrees of freedom, that's essential, so we cannot throw away trainable and certain non-trainable, then the dynamics resembles, you know, lattice field theory, where, you know, individual nodes would be like neurons and they would have some very precisely defined like connections to each other. It's not like, you know, any connections, we would do the trick. So as I said, getting Klein Gordon scalar field equations was easier. It's like kind of more generic, getting something like, Dirac is harder. And when I'm not there, I'm not ready to write down the standard more Lagrangian say, well, here it is. So that for the field theory. Now, the other part is Einstein equation. Once again, telling you that I have finalized my understanding how Einstein equation emerged from this framework would be a lie. This is not true. But what I do know, I do know how to get emergent space curves space from it. I also know how to get emergent space time from it. That's again, I mean, it's like, that's a subtle difference that most people wouldn't pick up on. Okay. So expand on that please. Sure. Sure. So, so, you know, the space you can, you can probably understand by showing like surface over Apple, right? And or chip potato chip. And you'll say, well, it looks like two dimensional surface. And, you know, since we are three dimensional beings, was easy to look and say, well, yeah, there is it's curved. It's not something flat. I cannot put it on the on my table, which is flat. And the same for the potato chip, if I take it, if I say to chip, which has negative curvature, Apple has positive current, curvature. If I put it on the table, it wouldn't be lying down. So that's kind of our understanding as three dimensional creatures of what curvature looks like. Now, this concept can be generalized to 3D. Now, I cannot now actually, you know, draw it or my hands because I am in 3D dimensions, but, but we know the tricks. We know the tricks, how to do this calculations, how to imagine, we even know how to draw three dimensional objects on two dimensional pieces of paper. So, so it's not so surprising that we are able to carry out calculations in 3D. And so when I say in the 3D curvature, I mean, the three dimensional space, which is curved, and that turns out to be actually not, you know, some feature of this theory, it is, should be a feature of any theory of everything. So if your theory doesn't produce in some limit emergence of the curved space, then you are against Einstein and, and, and of course, this is the one of the most beautiful theories that we have. And, and we cannot just throw it out of our considerations. Okay, so that's that's three dimensional space. Now, for the space time, that again involves a little bit of of, of, of, of, if you want to understand it correctly, you have to write equations, but since the audience is, you know, by module distributions, we should try to explain what space time means even in, in, in that sense. So what turns out to be that, the space time when your space, when you're talking about that, you, you have to tell how you measure distance. So what do you mean by distances between two points? If you have that definition of how to measure distances, which has a specific requirement, that, then you know, what kind of space you're doing. And, and then the apparatus for that we call it metric, tens, not very important. And space time comes with very, very strange at first, you tell to any student that this is how distances should be measured, and they will question why this looks bizarre. It turns out that to measure distances in the Euclidean space or it's just space, you take like x square plus y square and take a square root of this Pythagorean theorem. Well, it turns out that if you're working in space time, this isn't true. You should not be adding the two squares and taking a square root, we should be subtracting. So one of those squares, which corresponds to time coordinates has to be subtracted. And because of this stupid science difference, there is a huge difference between space and, and, and the space time. And so it took, you know, some time to get the curved space, but if you cannot get a space time, then again, your theory is not an agreement with observation. And we do observe a curved space time, you know, my background is in a cosmology and the space time is important there. I hope I wasn't too technical because no, no, no, no, and I have a technical question. You're absolutely right. Actually, I love that you said that because this is always true. There is those who actually know, know the terminology and would appreciate me speaking more like in the using the physicist or machine learning terms and that those who don't and you don't want to board anyone of those and so the aim of this podcast is to aim toward researchers to ward postdocs and graduate level PhDs and professors and so forth. And that the advantage of this podcast or the niche part of it, the difference in it is that it's as if for that other distribution of people, they finally get to peer into what it looks like when, when professors are talking I'm not a professor, but you get the idea. So, okay, you mentioned lattice field theory and lattice field theory has a problem with with Fermion doubling. So I'm curious if anything about your approach helps solve that problem. No, and we're not there yet. Not even close to actually fitting the lettuce like field theory, I shouldn't say the lettuce field theory because it isn't, but it is let us in a sense of how their weight matrix is arranged. So you have a lettuce. Now, this is actually why I'm not happy with this particular model of off. of how fields emerge. There is now another one, another approach, which I wasn't able to take as far as getting thermions out of this, but the approach is that it's closer to propagals as opposed to fields. Now we do know the fields were better and then particles, kind of only a good description of certain limits, right? But so speaking of that, second approach is neurons or some networks, they behave like particles in the emergent space and that emergent space is the space of actually trainable variables. And I already mentioned Adam like algorithm and so machine learning people would say, okay, now I know that Adam comes with metric and there is curvature. Now, but from the point of view of the physicist, it's small like there is second approach again. As I said, the theory is not final. And so you take all approaches, you can and you're just trying to say, okay, well, what can I say? Can I get thermos? And in this second approach, it is as if what you have, we have some kind of subnetwork of neurons that are doing their usual business activation and learning dynamics, but their motion is considered in the space of trainable variables. And that space does not have any lattice structure. It is just a completely continuous space and then you do have places in that space where no states are occupied like vacuum. And even if there are, you know, once in a while certain neurons appear to have such and such configurations of the trainable variables, this is not a field. It's a kind of discretized, more like similar to particle that I said, but also strings, right? Strings are assumed not to actually be fields in the sense of occupying. They're like one dimensional objects. So I think it's probably a good idea to say, like, you know, fields that work extremely well, three dimensional objects plus one time, strings, one dimensional objects plus time and those neurons in this second picture are like one dimensional objects plus time. Zero dimensional, zero dimensional, but that. - Okay, I think we'll be super useful for people. If on screen right now, the video editor will place in what a neural net looks like in terms of, we're then giving tutorials on what neural nets are. And then I think what's useful would be for you to say what your theory is not saying. So for instance, in the beginning I said, why at all is this surprising if a neural net can approximate any function? You're like, well, but that's not word. We're not saying that. Okay, something else you're not saying. And again, referencing this image is you're not saying that each one of these nodes is somehow space discretized. - Exactly. - Because there are other causal set models and causal dynamical triangulation models and other discretized forms. Okay, you're not saying that. Neither are you saying that this is a hypergraph model like a wolf or model. Okay, so when you start to talk about this with your colleagues, what else do they think you're saying? But you're like, no, no, no, no, that's not what I'm saying. It's this. So the first thing you identified right away and you absorbed the right, that's what people think and I'm saying, no, learning must be there and you're absolutely right. Now the second thing is I am in the superposition of saying and not saying it. So I'm saying there is two possibilities and both of them are being explored. One possibility that it is like lattice space, whether it is square lattice or some other lattice where it has triangulation, some hypergraph like model, which is and that is a possibility. And that is a possibility that I'm exploring. It comes naturally because neural networks are networks. You can easily get a graph of to this. Now in a distinction from the models, other models where you have this network or graph like structure, is that I am constrained to how this network will evolve. I'm not able to just say, look, you know, you have a graph, now I want it to form a torus or I wanted to be flat. I'm not able to just impose rules without saying where they can't fail. So where they can, from is for me to specify actually the one most important object in this entire theory. Like you know, in physics we have one object that kind of describes the entire theory, it's your reaction or Lagrange and you give it Hamiltonian and you are done. So here you have to specify loss function. And loss function is a very strict object. It's not, it's not, I cannot write, you know, it's a scalar, right? So you have to pay attention to that. And so if you want to use this hypergraph like structure and you want to see how it evolves and you know that experiments suggest that you have to form such and such approximated geometry, you have to go back and say, all right, what loss function would give you that? And that kind of puts it in. So this is one approach, which again, I say and don't say because in this approach, I do say that in this alternative approach, it's like two types of neural network theory, if you wish, right? Type one is that yes, you discretize it and you work with it. Type two, your space is the space of trainable variables and things evolve in that space. And there are pros and cons of both approaches. And there's a second approach. In the first approach, your space is discrete. There is nothing between the nodes. In the second approach, your space is continuous. Your continuous trainable variables, you know, they were not continuous. You wouldn't be able to use gradient descent or add them whatever. And so, and so, and you try and you try both things. And you try and you see that the one approach helps you. And it's very similar to what we do particles versus, you know, let us feel theory, which comes with its own both come with problems. But yeah, I don't want to say that, I don't say that, but I say that in addition to that, I also, you know, I'm going to investigate this other possibility that actually recently proved to work better in a sense because the, the curvature, curvature emerges not because I've assembled my graph in a certain way, but because it is an algorithm, which is more efficient. So, so the curvature is a way for the system to learn faster, not because of, so it's not going to do one direct way of saying where the curvature, where the geometry comes from. I do know that in machine learning, leadership people are using atom, they're not using the terminology of the curved space of trainable variables, which emerge from learning. It's again, it's not something to specify ahead of time. It's emerges as an efficient algorithm. But this is what I'm saying here. So it's not the curvature of the loss landscape that corresponds to the curved space time of our universe. Oh, absolutely no. No, it's, it's like you were saying, I think it's a very useful analogy when you're talking about loss function is to think about the grunge. So it's not the curvature of the Lagrangian landscape. Yes, okay. That gives you the curvature of a space. No, it is the degrees of freedom in the Lagrangian, which we call metric, which describes a space which is curved. So, so yeah, so there is there is a big, big disc, and same here. Now, maybe this is a good point to, to, since I'm drawing this connection between Lagrangian and loss function, originally, I was associating the loss function with more of the energy. And because, because, you know, there were like stick, stick, a thermodynamic description where a canonical ensemble naturally would emerge from that picture. Later, I understood that adding a kinetic like term to the Lagrangian, to, to loss function actually makes learning in very certain situations better. So like, you know, you had one more term, which isn't maybe the term that you are trying to optimize, but once you add it to the loss function, and once the loss function uses this term to, to, to, it learns fast. So, so, so there is, there is this little bit of, of a new twist. And if you want to minimize something, you may actually use it with, with a conditional term. And so in this case, I think it's a very good analogy to think about loss function as, as, as, as a, as Lagrangian, although they're different objects, of course. When I'm wrestling with a guest's argument about, say, the hard problem of consciousness or quantum foundations, I refuse to let even a centilla of confusion remain unexamined. Claude is my thinking partner here. Actually, they just released something major, which is Claude Opus 4.6, a state of the art model. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow, thinks with you, not for you, whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle problems that matter to you. I use Claude actually live right here during this interview with Eva Miranda. That's actually a feature called artifacts, and none of the other LLM providers have something that even comes close to rivaling it. Claude handles interalia, technical philosophy, mathematical rigor, and deep research synthesis, all without producing slavinely reasoning. The responses are decorous, precise, well structured, never sick of fantic, unlike some other models, and it doesn't just hand me the answers, the way that I've prompted it is that it helps me think through problems. Ready to tackle larger problems? Sign up for Claude today and get 50% off Claude Pro when you use my link, Claude.ai/theories of everything, all one word. Earlier you said that the quantum dynamics were extremely difficult, non-trivial, or what I have you. Walk us through the insight when you were studying machine learning, why were you even studying machine learning? You mentioned cosmology, I don't know about its connections there for you in your particular use case. Anyhow, walk us through you as a circle six years ago or so. Okay, so six years ago, I was on sabbatical leave, so when you're on sabbatical you can do whatever you want. I finished the project that I was interested in at that time and had to do with certain dualities of quantum mechanical, strongly coupled systems that I thought would be a good candidate, also describing curved spaces and quantum gravity aspects of that. And then I had time and so I attended many, many talks by machine learning people who would present nice slides, nice results and no formulas, no formulas, no equations apart from something like stochastic gradient descent or something that kind of looks trivial. And I knew that neural networks and they would always say, well, it's like a black box, black box meaning it works, we don't really know what. So I had time, I have a few months and I said, okay, why not just try to open this black box? Because universe is also black box, nobody in the beginning told us that this is the standard model and somehow we came up with the tools of Lagrangian Hamiltonian mechanics to actually understand why it works. Maybe not understand why this particular Lagrangian, but understand at least how to model it. So that was my motivation, taking it and it has nothing to do with quantum mechanics and so I, at this point, I see the system with many degrees of freedom. They evolve according to learning and activation dynamics. So I knew that some of the physics will be relevant but not all of it and because it would be more complex. If you see a system with many degrees of freedom, your first reaction will maybe you can, you can discuss ensembles, statistical ensembles and maybe in some limit, you can understand how the system behaves and what we call emerging regime. So, so something like, can we have a thermodynamics of machine learning? So something along this line, I knew it was different because of the long this time. So that I figured out sooner but can we have a certain thermodynamic description which can be verified? No, is there a notion of a temperature? Is there a notion of entropy? Is there the first second law of thermodynamics? Would they still hold or do they have to be metafide? So that was kind of, you know, you have a toolbox that you think should be the first you try to model the system and that direction I went. So as I said, quantum mechanics was not on the horizon but then I saw that because of the learning dynamics, it is not a system just doesn't go to a boring canonical ensemble distribution and stays there. It has a very interesting behavior even in equilibrium because of the presence of those two different dynamics, activation and learning. And so the idea was to set up some kind of variation of principle that might be describes it beyond thermodynamics. So first, you know, thermodynamic, are there any microscopic objects like temperature, entropy that we described? But maybe we can go beyond it. Since this equilibrium, I call it learning equilibrium, is kind of boiling and then you know, things fall out of the equilibrium and go back and it's kind of not. So getting really, you know, zoom in and say, well, I don't want to just calculate temperature, pressure, rolling, whatever you usually do. Although you have to still define all those things in machine learning system and I can you zoom in and say, I'll pay attention to, let's say, only trainable thereness. I will still integrate out and kind of course, grain over non-training but but pay attention to train them. Now the reason for that was that we know that they're not trainable, they like flip very fast, you know, you put your input image and get caught some tags or cats or dogs are out in your example, right, zeroes or one. So the, so the activation goes fast and the learning goes slowly and then you calculate a loss function and gradually propagate changes. So I knew there was two, two skills and if anything we know in physics is, that's what you should do, you should integrate out and remove the irrelevant information and keep on the relevant. So that's what I did. So I integrated out this and say, okay, how does this system behave and it turned out that behavior of this trainable variables. If you assume, you have to assume certain principle for how entropy changes. So I assume maximum entropy production extreme, station entropy production principle. But if you assume that the equations derived from that are the mandolin equations. Now mandolin equations, again, those who know, no, but those who don't know, it's close to quantum but it isn't quantum. So there is still this step of quantum is a theory where you, what propagates is not the probability but the square root probability, square root of probability and that's the relevant degree of freedom and that comes with this quantum phase complex numbers. So we all know that, you know, and so if you only pay attention to the mandolin equation as you remember, this isn't quantum yet, but it gives you hope. So maybe you can actually understand why this complex phase would emerge. What is the physical meaning of the complex phase? Not exactly quantum, but maybe as an emergent quantum, again. And then I collaborated with Katsunelsson and he pointed, a corrected point about that we need the discreetness of the phase for this to work. We don't have to call it a phase, but something has to be discrete, something you change discreetly and you'll lose function, but say it doesn't change, right? Much. Or the dynamics doesn't change. So that's what the meaning of the complex phase you're rotated by two pi and you come to the very same point. Right. And this within the system, it's like having age bar, you know, having age bar, having something that is like without this, it isn't quite quantum yet, although in certain regimes your system. And so it took this little extension to the original derivation of mandolin equation, like almost classical quantum equations. And that came from very interesting picture suggestion that we made that actually can explain it to artists to anyone. So you have your system, a learning system and yes, you pay attention to trainable variables and then yes, they follow almost a Schrodinger equation. But for you to for a Schrodinger equation, the system has to have access to a bath, to a reservoir of neurons that it can borrow. It's like, you know, external resources, you run your machine learning system, but you say, well, if you need, here is few more, you know, neurons you can use, you can plug in. If you don't need it, just give it back. And so if you have this access to the system in physics, we call it grand canonical ensemble, it moved from canonical ensemble to the grand. So if you do have that, if you like in your algorithm, you you provide that option. Whether it is an option that you provided by, you know, actually programming it this way, or whether it's an emergent phenomenon, because you know, there is an emergent phenomenon that certain neurons stop working and start working, stop working and start. If you do that, then it turns out that you do get Schrodinger equation. And in some limits, again, it's not, it's not an exact Schrodinger equation, which means that in certain limits, it should be violated, but with that little twist with this space of neurons that you can kind of hire. Like you hire to do some work and then once you don't need it, you put it back. With that, your dynamics effectively becomes a linear, because Schrodinger is linear and and described by Schrodinger equation. And then everything fell into places. And at that point, it was actually more than just a proposal for some abstract theory that in certain limits can describe mandolin plank equation, which isn't quite, but it is actually, you could see that, well, maybe this quantum… and this can emerge from completely classical system, in the sense of there's no, it's not the quantum machine learning. It's a classical machine learning, but with this little twist, it requires this quantum like behavior. - I do wanna get to consciousness. Before we get to that, I have a sort of a silly question. So in machine learning nowadays, there's a huge field of interpretability. So people wanna peer into the black box. And it makes sense to some degree, when it comes to LLMs, because there's semantics underneath of what LLMs are trying to capture. So you're trying to understand what are the LLMs doing. But when it comes to the universe, more abstractly here in your model, is our universe interpretable? - Right, so let me start with machine learning. Then move on to physics, and then maybe to more general answer the question. So in the machine learning, the fact that we need to interpret how machine learning system work would happen inside is essential. And I already said that physicists had dealt with this problem. And our tool was, yeah, model the dynamics of this system, but then model it as a Hamiltonian system, or as a Lagrangian system. So that was our way to dig into the black box. Now, with LLMs like models, or with any other machine learning systems, you start with designing certain architecture. So architecture can be written also in a certain way that is interpretable and essential. No exactly which blocks do what? And that would be similar to writing not just something that produces results, but something that produces results. And you know why it produces results because there is a term in the loss function or in the Lagrangian. And if you remove that term, then it would produce different results. Maybe your LLM would not work. Your chat GPT would break. And so in this sense, this is a way to model and understand how the interiors work. So maybe this is the term that is responsible for math being right in their large language models. This is a term that is responsible for good translation between certain languages. This is a term that is responsible. And then you can kind of concentrate on the term and say, okay, well, I don't like my LLM model is not producing, you know, not doing good calculations of tensors, which it isn't. So if you look and try to kind of do tensor calculus, all models I tried, they all, you know, at some point start producing garbage results. And so you have to kind of locate it. So there's certain tests they can don't know how to do it. Maybe just the term that you have to tune in your loss function in Lagrangian in the architecture. Now the same thing, we, and this is an interpretability. How do we interpret? Why is it, you know, certain things work and certain things don't work? Or what is it inside of the state of the neural network that was responsible for one thing? And how to do alignment, right? What is it there that I can say to my neural network and then in some sense it's will align with my interest, with my loss function? - Yeah, so we know how to do it in physics. I think this tool box in physics can be useful to enhance our interpretability, interpret how the LLM work, how machine learning systems work. Now, and coming back to the physics, well, we can also should be able to dig deeper than just writing down a symmetry group and saying, okay, well, this is the standard model, this is it. You know, we can start asking question, why is it that the case? You know, what is it? You know, because in this description, the field theory is on everything we observe in the microscopic levels, I imagine phenomena. They would come from some microscopic loss function where just to throw one example, you know, maybe each neuron wants to minimize entropy or maximize certain local loss. So if you want to minimize entropy, it's not a good idea to connect to all neurons because then it will be chaotic, every state, every next time step will be chaotic, but not connecting to anything was also good. So maybe neuron individually and a microscopic level will try to find some. And then there'll be some kind of RG flow where this microscopic loss function would give rise to more microscopic behavior. And then we would see say, okay, well, at this level, it is described as a standard model, but the flow doesn't stop there. And then you go, okay, biology level. And a lot of work that we've done, again, we may touch upon this during discussion of consciousness, but there again, I mean, it doesn't mean that at a biological level, it should be the standard model, the governance, the correct description. So if you can identify how the interiors work, you can kind of RG flow it and understand how things work. Now, and for the more general audience, this is just a way of saying, we are describing the universe around us and depending on the scale on which discuss the different languages are appropriate, more or less. So in a very microscopic level, the language is maybe neural networks. I want to say on the bigger level, maybe the right language is a field theory. Interesting. Now even bigger, genotype or phenotype, and then even bigger. And so it should not be surprised in this approach that on each level, there is just a different language, which is correct for describing certain, on each scale, on each energy, different configurations, there'll be different languages, which are more appropriate. That's kind of what we are saying here. And it shouldn't stop here. You know, once we move on to gravity, cosmological scales. Yes, we are trying right now to use the language of field theories to describe cosmology, inflation, gravity, but things don't quite work. On a cosmological level, there is this dark energy problem, there is dark matter problem. And so maybe once we adjust the language and start by describing those scales using different language, different, modeling, then you may have a more agreement with the experiments that we have now. So would now be a good time to talk about your second law of learning? Sure, sure, sure. I mean, I've touched upon this. So I think it's important to emphasize that just like a second law of thermodynamics, we really like it, but we should understand that it doesn't work. It works all the time until it doesn't. So we should not take any microscopic laws should be taken with a grain of salt. So because, but nevertheless, it is useful for many, many different calculations. And so second law of thermodynamics tells that the entropy should grow. And then if you really take it literally, then you will have hard time explaining the emergence of life. You'll have hard time explaining many biological phenomena if you want to be honest. Now if you want to wait a moment, why would you have a difficult time explaining the origins of life? Because it's always that global entropy increases, but you're going to have local. Right. But if you're only talking about global entropy, there is just one equation. There's one number. It's not very interesting. What's interesting is what, and okay, it's increases with decreases doesn't matter. What's really happens locally? Because in the thermodynamics, it says, you have any subsystem, big enough, and the locally entropy will grow. And so instead of just having one number in the entire universe that you've observed that grows, what's more interesting is to pay attention to what happens in the different subsystems and then explain how they entropy. And how you define the entropy is also, because if you talk about gravity, it's very important to think about how you actually define entropy. This is now we have a space that has curved. Now if gravity pulls things together, wouldn't it mean that the entropy decreases, right? So you have, before they would be distributed everywhere and they pulled together, now, now, you will say it decreases, but then I say, well, no, no, no, no, I'll define entropy in a different way. I'll dash the problem. So that it doesn't. But again, if you only pay attention to one number that you want so badly to increase, then OK, let it be. I'm saying that the usefulness of second law of thermodynamics that can be applied for many subsystems, very successful. And then when it-- but there are certain things that with this, you will have hard time explaining. And as a cosmologist, I have to say, we have one beautiful theory in cosmology, called the theory of cosmic inflation, that is kind of hopeless. It doesn't know what to do if you try to assign probabilities for observers to emerge. And if you use the usual classical mechanical or physical approaches and trying to calculate, what's the probability of a certain observer to emerge? Is it easier for us to go through inflation, and galaxy formation, and biology formation, or it's just easier for us to form Boltzmann brain? that are floating in the empty space with the memory of us, thinking that we are on the meeting right now. And you will see that in this very successful otherwise models of cosmology, you will see that very often you get the answer that, yeah, all this correct. What we think is correct is just gives you a lower probability. It's higher probability if we're just too nucleate out of nothing. And that's called the Boltzmann brain problem. So no, it's. First of all, once you enter gravity territory, a cosmology, you have to be careful about what you call entropy, and then there is this problem of defining probabilities for us observing what we observe. And people employ different ideas, one of which is on traffic principle. Well, let's put that in mind. So maybe we should not be just a random point in space because if we were, we probably would not be observing what we're observing now. Maybe we should only look at the places that are actually two-in-four-life, smart enough observers who can ask those questions on traffic principle. To be clear, the anthropic principle is different from an entropic principle, which is a good point, very good point. Actually, it's a lot of confusion about that, and partially, because entropy is used in the post-conquering, but you absolutely right. Antropic meaning with an A starts with an entropic principle. And that's people. Very large fraction of physicists don't like the principle they disregarded as an entropic. Now in cosmology, that's what's kind of only game in town. Now, until at least Molling proposed his national selection approach, which is, I think, closely related to what I'm trying to say. Interesting. And with the one-in-twist, I'm actually giving a mechanism. Yes, yes. And that mechanism of how instead of a universe being fine-tuned for life, it is self-tuned for life. So you start with whatever you want, because universe is learning. So it consists of learning subsystems, and the learning subsystems, and they all try to learn what around them in the most stupid way. Because of that, you're not tuning anything. It is self-tuning itself. And so this guy is giving you a physical mechanism of how ideas that were. And actually, at least, knowing to tell me that this idea came before him, they were philosophy. But there's always. Any idea you describe, there's always some philosophy for in the past, who said the same. Sure. So those ideas were, of course, but here, you can actually say what the mechanism is for the statistic place. And then in the comments section, they'll be, oh, and this philosopher is predated by the Vedic texts. Yeah, there's always someone before. Yeah, but I don't think there is competition here. Just with every new time you rediscover something, what you're trying to do is trying to make it more rigorous using the tools you currently have. So, you know, there were no machine learning systems 100 years ago. There were no neural network dynamics back then. Yes, exactly. Now we have that. And so those are the tools. Can you use those tools? Well, speaking about the tools, there's one little problem that kind of maybe unique, maybe not. Many times physicists came to realization that new tools are needed. Einstein is a great example. You know, who would have thought that curved spaces or even geometry is important for modeling the universe. Nobody. But Einstein came around and said, no, this is a mathematics unit. You need a differential geometry to describe. So, but at least at that time, there was already a body of work where you can just take this framework with neural network station is a little bit different. We have so much experiments, so many experiments. Every time you are amazed by what, what, what, what neural does. It's an experiment. And not so much of the actual theory of you being able to actually predict ahead of time. And yes, this architecture will work. And no, that will not work for such and such reasons. So that's kind of the theory is a bit behind here, which is like perfect program playground for the theories because I can, you know, set up experiment. Write down my theory and then test it experimentally numerically right away. So that's that's right. Okay, but that's. It was a diversion from. I want to get to where you said that the universe likes to be observed. We're going to get to that and we're going to get to consciousness. But before we do, I recall you saying that natural selection operates on the level of subatomic particles. Am I shaky in my recollection? Not from this conversation, but somewhere else. No, no, no, you absolutely, it may have not been in this conversation that I said that I certainly wrote about it in the papers. And yeah, and so, so this this natural selection like way by natural selection. I mean, the more useful configurations of networks survive because they help loss function to be minimized better and that and the other configuration will not survive. So in this sense, natural selection, those architectures that are useful for learning will stay. And those who are not useful for learning will be removed because because they are not they are loss function is not as low as should be so in this sense. Yeah, but it doesn't work just on the level of particles and entry particles. Remember again about this analogy of the right language on different scales. So if you want to talk about this at the level of particles, yes, you would say particles are such, you know, the way they are is because they underwent the theory series of natural selection. And then the natural selection of their scales and figured out that this is the states, this is the state of their neural networks that describe them that are allowing their loss function to be the smallest but this this is can be up this argument this. Or I call it more learning argument, right, so you're trying to minimize something on can it can be applied to any scales can be applied to scales of biology, which we usually do when we say natural selection. We usually think about the scales of organisms, right. You know organisms again some kind of configurations there more maybe fluid configurations because that not no two organisms are like but so maybe particles right so we yeah they look very similar like all electrons look very similar but we don't know maybe there is some tiny difference and then in the way we are trying to understand the tiny difference. Maybe it's already went through this very long period of natural selection and then that is the value that's what now electron mass should be and there's nothing else I can do. I see I see at this point. I imagine if you could put a mark on an electron and and a boss on or and distinguish them then the spin statistics there wouldn't apply anymore and we would see some effects of that. Right so the last function would be it's just not convenient for not to have fermions and and again we kind of understand that I think one good example I can give that maybe a very general audience will understand the machine learning people will definitely understand cars cars and and you know moving in the traffic and they are trying to sell driving cars. It's always you were 10 years from now where you know every all the cars are self driving maybe sooner maybe later I don't know and and then all of them are driving and they're all trying to optimize their loss function. But to optimize their loss function they have to do some calculations about the environment they have to scan scan the environment you know find something plug it into maybe their network and then network will say turn left turn right. And so and and at this time that information that it collects each of the cars collects we actually call in in the physics of electrodynamics we call that that bosonic field. So it's electromagnetic field or you know green function of other you know electrons that I scan around propagates to me and that gives me relevant information for what me to do as an electron or cars can around for other cars get that information and say okay that's the relevant information to me for me to make left right turn or accelerate so in this sense cars are like fermions they are advanced enough to be able to process. Not the set of the entire universe around them because they are tiny but relevant information for them to optimize their life and if they would be doing something else they would not behave as a as an electron and that would create some kind of unstable behavior and this whole system would would not work as a shoot so just like. So just like. you know, all the cars converge to some, all self-driving cars are using the same software, just because that's useful, you know, electrons kind of using all the same software or how to navigate an electric net field. So this is a analogy, maybe helps to think about electrons as, as, as, you know, self-driving particles that have already established what is the, what is the, the best, maybe they haven't established it exactly and so there's still internal degrees of freedom, you know, there's spin and there are other things so, you know, in some, so one circumstance I will be doing this and that, that, and the same for the cars, right? So you will have different cars driving in UK and US, right? Because a left, you know, left side, right, right side driving. So there's state of the self-driving software and the cars would still have to be different, have to, have to agree, yeah, but, but other than that, a lot of similarities in describing those, those cars would be and, and the same for, so if it's useful, this is, this is a correct analogy. There's a structural similarity between the cosmos if you zoom out far enough and neurons and some people use this to suggest there's a cosmic brain. Now, I want to talk about what you're not saying. Are you not saying that or are you saying that? No, that's a lot of the notes. I'm sorry, if I interrupted you, but, no, I'm not saying that. Okay, it seems to me like this comports with your theory. So it would seem like you'd be like, oh, that's great evidence for my theories. Maybe, maybe not. Yeah. So both of the things that you said are true. So first of all, I'm not saying that because I haven't, um, uh, confirmed, no, no, visually I've confirmed. They look similar. Right. We're all done that. There are papers of people who actually tried to do statistical analysis, which is the right thing to do and statistically showed that there are certain things that are similar. There is a well-known, um, critical, like phenomenon, uh, where you have some kind of scale invariance that is observed in the cosmic web and observed in the biological networks. So now I haven't mapped out exactly the dynamics of the galaxy's formation and how all this would come around. I've only done calculations suggesting that self-organized criticality or critical state is something that you should expect to see in the learning system. And it's a good thing for the learning system to have criticality. So there is this interesting, this calculation that tells you yes, criticality is good. And then you can say it and say, okay, once it's good, I should, doesn't this confirm the observed criticality in the brain activity or the observed criticality that we see? Yes. But, but, but this is indirect evidence. So maybe this is actually, we are talking about this cosmological-like scale where, um, the system performing very slowly, maybe some kind of very complex calculations. Or maybe we're just saying it slowly, but actually doing some important learning task. So yeah, so because of the, I do not say that, usually I show those pictures when I give public talks, but I do not say that I've done enough rigorous calculations for the structure formations to say that this is. I know how to do those calculations, but there's just only 24 hours a day. Yes, you have a rare quality where you will assert something and then say, and here are, here's why it's either not fully the case or I don't believe it or here's the limitation of my own model. Here's the counter evidence. I haven't seen that in almost any of the people that I interview. Okay, so, so, uh, I also hated that when I was a student because even when you're a student and you come to the class and then they tell you something that they've been taught, uh, and they take it without actually trying to question it. I think this is a horrible quantity. We physicists actually can do better. Um, I think, I don't remember who exactly said that, but, um, we should be doubting everything. So we should be doubting our own models, our own calculations, calculations that other people had done are some, even if 100 people come to you and say, um, that general relativity is wrong. It doesn't mean it's wrong. And we know all that that story when they're 100, you know, physicists wrote a letter to Einstein signing that saying that generalities is wrong. And he's reply was brilliant. I mean, again, if you don't need a hundred, if you have a point or show me a point and I'll, I will consider it. So, so yeah, I think this is the quality that you absolutely must, um, you have to do show all the good and bad things about, because I've thought about this. I mean, I, of course, I try to answer. I'm not, I don't have all the answers. So I think this is more honest and correct way of doing this. And we should be doing it not just with new theories. There are a lot of, um, uh, problems with existing theories. Classical, uh, theories are not as pure as they thought. There are divergences and things we don't, uh, we don't fully understand and we should be telling people and students about them when, when we discuss all this. So, so, uh, don't put anything under the rock. That's kind of my approach, because sooner or later, the smarter people will find what's under the rock and, and, and that's certainly very, very important. With that out of the way, and thank you for that by the way. Let's get to consciousness. Is the universe itself in your model with the universe being the neural net? Conscious? Okay. Uh, very good question. And of course, I get this question all the time. Now, um, here is my, maybe a longer answer, but I think I need, need to say that. Um, you come with the mathematical framework, new mathematical framework, which is very rich, which is the, the, the, the, the relies on neural networks and the learning dynamics. And you're trying to use that to describe some, some phenomenon, this phenomena may be physical phenomena like we talked about, or the phenomena that people are discussing in other branch of science, like consciousness. So they already have a term for something and you bring a toolbox. You don't, in this toolbox that that I have, or learning, there is nothing, uh, that I would call consciousness, but I'm trying to use it to describe what people mean by consciousness. And I can, I have, I can have many attempts. So I may suggest something and they will say, well, this looks like I know not a good definition of consciousness because, because here's the system that we all agreed, a hundred of us agreed that is conscious, but your system, uh, your, your definition tells it it's not. Then okay, then either I say, well, uh, maybe you should adjust your notion of what consciousness or maybe I should adjust my definition of kind, and, and, and both ways are funny. Now my definition of consciousness comes from how I understand it, how I can build in within framework that I understand mathematical frame. And then this mathematical framework, uh, you know, system undergoes learning dynamics and there are three macroscopic things that are directly related to learning that I can calculate. So one of the things is how fast system adapts to their new data set to the new environment. How fast it learns. So this is the, I can actually calculate it is the decay rate of the loss function. That sounds to me more like intelligence than consciousness. Right. And, and, well, okay, so, but, but then I say, uh, hold on the intelligence because I have comment about that as well. So, so, so, so, and then I'll say, I want to, as a hypothesis, call the risk rate of decay, how fast I want to call it consciousness. Maybe it will be wrong, but I will call it, uh, because I come with a new, um, new framework in this framework, I can calculate it. But I say right away, there are two more things that are also macroscopic and some people may relate it to, uh, things related to consciousness, uh, but I would relate it to intelligence. And I would say actually three things contribute to intelligence. If you judge a learning system, how well it behaves, then there are three quantities you have to calculate. One is how fast things learns. And I call it consciousness. Maybe you would call it intelligence. Second thing is, uh, how low does the loss function go? I seem to be told if I had infinite time, how will it, will go, yeah, it may be long fast, but then just hot stuff. So I would say, okay, that's another ingredient that is important, uh, you know, how low is, is a loss function. Uh, and the third one is once it reaches this asymptotic loss, it's not going to stay there. It's going to be fluctuate because that's what learning activation dynamics. You just don't stop. You, you never stop. You're always in this learning equilibrium. And sometimes lost cause a little bit up, down, up, down. So you always fluctuate. So how big are the fluctuations? So then I have a learning system and I can calculate three things. How fast you learn, how good you learn, if you had infinite time, and how stable what you learn. And so I would say that because of that, those three things, I actually describe what I think is intelligence. Not just one IQ number, three things. Uh, and then you, you can have, you know, different people. Some people learn very fast, but, uh, they, they, they stop, they kind of, they are not learning more and more. Their loss function is, uh, halted. And other people may take very long time to learn and then evaluate. Eventually they end up knowing all of the differential geometry and you know quantum field theory what's not and and then the third type of people who maybe also learn fast and maybe they know all of the advanced mathematics, but they're very unstable until they keep repeating it they you know keep forgetting and then opening your book gain will forget stuff right to learn something I forget the things I wrote in my papers right there are many papers I have to look up so so there is some degree of how my knowledge my loss function is actually fluctuate so I would put those three things I would put those three things I said this that's intelligence at least three so it's at least three maybe more because actually there are more because you know it's a stochastic variable there are statistical moments first second third time so I'm simplifying things you can describe this fluctuations as with the infinite number of parameters but at least those three things are very important and I think when you look at different systems you can actually say with all different people you can say students right you can say well yeah this is has a very good learning efficiency I would say he's more conscious and this has a very bad learning and he's less conscious but again this is just a definition if somebody tells me that even with your framework I suggest it has to be you know a square root of two times the first number of two I'm okay with that so as long as my you know I declare what I mean by and then then I'm happy I understand the first two but the last one about stability why does that have anything to do with your intelligence when it could be the universe is changing that doesn't seem to have anything to do with you yeah so andro assumption if universe doesn't change at all so so it kind of or changes are it's always changing okay so it's always changing there you know this is what processing data set you your data set you get always different images of cats and dogs you look around and every day there is like new shape of trees and leaves arranged themselves so you always have that but I say let's integrate that so it will be just some state statistical state of the universe so no major events there is no major nothing hits the earth and creates a sequence of earthquakes and there's no nuclear wars and nothing major so I'm more or less in this learning equilibrium so if that's the case if nothing major change maybe a good example would be you know you take a bacteria as an as an observer place in some kind of controlled environment where you kind of keep maybe the temperature the same the same mod of light or so so if you work with this ensemble then the last one should will still change will still change because of you know stochastic grade in descent now you know sometimes I or bacteria is the light appears on the left then then it moves to the left maybe it appears to the right or in opposite direction moves to the right so there are some some changes and it always updates its new it's trainable variables to to if why would you do that because if they are state statistical change I would be able to adapt so kind of that your ability to adapt is actually you know backfires on you and it creates more more fluctuations we have to come up and people actually know about that if you kind of set the learning rate to be smaller like in some algorithm and it will go to a stable very stable minimum but the you know it will not be as good minimum so it's not so this fluctuations are not should not be treated as is a bug it's actually a feature to get out of the local equilibrium and so that happens all the time now am I correct in saying that you said at some point that we need to unify not to not quantum theory and general relativity but quantum theory general relativity and observers okay so most physicists tend to think of observers is coming from the physics something emergent why do you think that we have to unify these three at the same time right so and most physicists will tell you biologist somehow will you know emerge from all once I have strength here you know completely done and quantum gravity quantize I call it wishful thinking there is no evidence for that other than we think we kind of you know putting things in what one way is saying putting under the rug we are just saying if something is complex all yeah yeah but if I do long enough calculations if I have long enough time that's how it's going to work and I don't for example one example most of us is convinced that quantum mechanics has plays no role for how you know for consciousness how how we function how brain works right so yeah makes sense microscopic objects why would quantum mechanics but we have no you know proof for that we have no and I think it is more wishful thinking again related to how the second law thermodynamics is a wishful thinking that it has to be really working so I don't think so but the other answer to that is that the fact that observers are very special and should how be understood I think realized by most physicists who pay attention to two important problems in physics one important problem is the measurement problem so every single physicist who actually seriously thought about foundations quantum mechanics not not the person who is just doing shut up and calculate type of things and like following the manual of the but who is trying to understand that the study will necessarily realize that there is a measurement problem measurement problem is about this third postulate of quantum mechanics that is very new because in classical physics will be needed state and how to walls he needs state how to walls and how does observed so that's one and so and then you kind of have to say something that maybe there is something additional to quantum mechanics that you have to describe maybe it is an observed observer may play a special role and and if that's the case if you realize that quantum mechanics is incomplete and the measurement seems to be playing a special role then then you are stuck with was trying to describe now in another problem in physics that comes around and also has to do with the other problem is cosmology it's called that measure problem essentially more or less different problem but more more or less the same complication coming from observers so if you're trying to assign probabilities to different observations and cosmology that should be the right probably we discussed the both some of brain problems it's it's a part of it so so you have to specify the rules you have to describe how to deal with observes and in both cases we actually think about this complexity comes from the fact that we are trying to put observer into the system so when observers outside of the system well know what to do we know you know there is Hamiltonian how it was right but once you put observer into the system so this is you know for for more general audience people know about shorting your cat problem so if you put something and then the wingers for the end problem and start putting observers inside things start to break and and is it important that the observer is conscious right so so at this point no there's something fishy about observers but we don't have a model of observers so once you say okay let's some people claim yes consciousness is important and and they have their own definition i say it's important to model observer so if you want to put it inside in the system then you really want to model how it behaves and model not just saying well maybe some kind of emergence phenomenon biology will happen and maybe some kind of wave function collapse will happen this isn't going to work if you try to do the calculation so so my answer is yes observers are very important and that's why you really have to describe them if you want to do calculations even separately in quantum mechanics and gravity and more so and maybe because those two problems persistent cosmology where is the gravity and quantum mechanics where there is where there is the measurement problem maybe the solution is actually try to understand how observers work and then once you understand that the both theories may some kind of be unified and their observer would be models as well because it seems to be the problem with all of the with both both theories which again you know you can certainly put it under the rock you can stop to ignore it maybe elephant is the room in the room is is every bed knows it's there but we are trying to look the opposite way and I'm as I said again I see the I see the Sullivan say well it's it's there at what point do you imagine observers entered into physics is it at the plank epoch is it prior so in this model everything is conscious there are observers everywhere every subsystem is an observer some of them are efficient observers they have efficient architecture so there was function you know falls down some of them are stable observers they've already reached the very low value of the last function and some of them are not stable and they've also always fluctuate out of so any subsystem system because the entire building blocks in this model neurons and they come with the trainable and non-trainable variables because of that everything is loading so everything in the sense is observed. Just some observers are capable of doing and asking perhaps more complex questions than others. Although we don't know, right? Maybe inside of electron there is whole complex neural network that already solved the problem of quantum gravity and just looking at us and laughing saying, "Oh, guys, I mean it's simple." Yes. Maybe, maybe. We don't know about that. And this model, as a model, we started with this. I'm not saying this is how it works, but as a model that could very well be. As far as I understand, there's a number that you can associate with consciousness. How conscious is this subsystem? But consciousness to us is far more than just a number. We care about how conscious is someone. Most of the time when it comes to health, are they alive? Should we remove the plug and are they going to wake up? But our consciousness or conscious of so much. So in your model, do you have any qualia? Right. So I wasn't one of the FQXI conference and there was like a heated discussion. Every time consciousness is discussed, any of their physicists and non-physists in the room, it's always a heated discussion. And so should we call consciousness the person who is actually conscious in the sense of, you know, talking and interacting with you, another conscious observer? Or should we call consciousness something? So is it like a discreet? You know, I talk to you then, I call you conscious. I, I, a new reply, you know, I talk to a dog and here applies its consciousness. So yeah, absolutely. You can then say consciousness will be defined as a coupling, a strength of the coupling of the organism with, you know, sound weight or, you know, light or some, some, some, some electromagnetic phenomena. You could do that. You can do that. And that would be your definition of consciousness. Maybe it's better than mine, right? And then we would not be arguing, you say, okay, look, this is a person, he's not conscious. But then there'll be people who say, yes, I have such and such, you know, a friend or a relative who is in karma, but he is conscious. And he would say, no, the fact that the person is in karma and isn't interacting with you the way you wanted to interact, it doesn't mean that he's not interacting with you some other way. And actually, I've mentioned, I mean, I have to say this speculative idea because, you know, maybe philosophers will love it or maybe not. But we discuss for quantum mechanics, you need this bath of neurons. Otherwise, it just doesn't behave like, so it could be this bath of neurons is always there, but it's not in our physical space. It's in the hidden space, what I call it. And if it's there, then nothing stops for a person who's not interacting with you in the physical space to interact with the hidden space, okay? Maybe that's what you do in your dreams or maybe, you know, people are interacting through this hidden space all the time and people do claim that. So there's a lot of people who claim that have this special abilities, right, to interact. And I think we don't, we don't, we don't take it seriously, physicists, I think for two reasons, we don't have a good enough framework to modeling this. And the second reason, we don't have a controlled enough experiments to do that. But I think we should not be disregarding that when we become equipped with a better mathematical model and with better experiments. So, so yeah, I wouldn't, I wouldn't like this definition where the person is conscious only if he can hear or tell reply back. But that would, you know, chat GPT would be conscious of conscious, according to that definition because it certainly replies when you're right. So again, it's, there's a lot of discussion, maybe not very important about definitions, but we need to do this. We need to define terms before we can make statements and that's just not time to do that. What distinguishes between trainable and hidden variables? Do physical entities correspond to some or even mental entities correspond to one, not the other or what? Right. So, so the, the, the hidden variables in this case are like hidden neurons and their states are described by non-trainable variables, something that, but, but, but all of the non-trainable variables in your network are connected by trainable variables, they're called weights. So, so kind of, you kind of like just draw a sharp line and say here is the, you know, trainable and here's not trainable, very much like in physics, you cannot say here's a traumatic wave and here are electrons. They're copied, they're all communicating through, through each other. Now, the, the, the difference over hidden non-trainable variables and a physical non-trainable variables is that the one that a physical, they've organized themselves in the three-dimensional structures and they've kind of discovered their effectiveness of using three-dimensional space for exchanging information and minimizing their loss function. The hidden space at this point, they can have arbitrary connections to each other. There is, it's, it's, it's not, maybe a good idea to think about initially you have a soup of neurons, everything is connected to everything and it's all hidden in a sense that none of, no physical space had yet emerged. And then there is like, you know, bubble from the bank and then there's certain number of those neurons figured out that they can learn a lot more like phase transition and can minimize the microscopic loss function if they arranged themselves in the three-dimensional space. So if that phase transition took place, then you really have, you still have the, the hidden variables, which, which you can always hire if you need to do calculations, but they need not be present in your physical space. They need not interact with, with, with a classical degrees of freedom. They still interact by providing you with this quantumness, but they need not be, you know, directly observable and coupled. So that's, that's the model for them. So and it is a correct to call them hidden variables because, you know, hidden variables is one of those. It's called interpretation quantum mechanics, but I guess it's an attempt to actually make quantum mechanics more mysterious, less mysterious, trying to actually, it comes with its own problems and, and we can certainly talk about them, but at least it doesn't put, it tries to put less stuff under the rock. We, we physicists keep doing that, like, keep lying without saying that we are lying in a good sense. We're not doing it intentionally. We're simplifying things, right? But one of those things is that we are, we can say something like, I believe in the manual world interpretation. And, but once you corner the person, he will admit that it's not as clear and, they're uproar. So, yeah. If the universe is learning, what is it learning toward? Right. So, so if the universe is learning and there is nothing but the universe, so that's it. That's all there is. It's, it's an also, unsupervised learning. And so the only thing it can learn, every subsystem can learn the rest of the universe. So you put in arbitrary boundary, me and the rest. So I'm as a subsystem, the only thing I can learn, I can, I can learn, try to learn about myself if I'm on cash, like, you know, and comma, maybe I can do that. But, but that's what I, I will be interested to do. And that will actually help me also to survive the more I learn it. So now we're moving to the biology level where we've done a lot of work trying to understand how this exactly works. But basically, you know, organism has to learn its environment, model its environment. In order to better predict how our environment will behave. And then once it able to predict it with more likely to survive. So this is what you have to do in order to survive. This is very actually. So first of all, it's similar, of course, it's to national selection, but it's similar to the phenomenological model that Carl Fiston is constructing where he's trying to say, okay, well, let's define some phenomenological function that maybe our ability to predict the state of the environment is what, and what I'm adding to this story is that, yes, that's great. And that's actually a dig deeper and give a microscopic interpretation of that. So it's like, you know, you can have thermodynamics, but then you can have a derivation of a thermodynamics from statistical mechanics. So I'm kind of saying, well, you can also derive it through statistical mechanics of how neural networks work. So that's the idea. So then coming back to your question, every sum system is learning the rest, right? And we are not different. We are as ourselves are not different that each cell is learning how to fit best into the organism and so that you optimize its own laws function. And the society is indifferent. It's only, it's still learning, but on different scales. And so languages, how we describe it, changing, you know, addressing to the physicist audience, there is our. you flow, there's the loss function changes as you start renormalizing and generalizing the concept of neurons. So in a small scale can be fundamental neurons, on the biggest scale it can be subnetworks like particles, then you can have something like cells, people, civilizations and societies like that. If I recall correctly, your second law of learning is that learning efficiencies proportional to the lapthacian of the free energy. Is that correct? So in very, very, it's just like in, in, in, in standard physics, we can derive thermodynamics in a very, very simple limits. We can unfortunately with physicists, only good in doing Gaussian integrals and, and doing calculations for a very simplified system. And for those systems, where you can simplify, you can do those calculations, you can, you can show this related to a passing of the free energy, where the free energy is actually defined microscopically as, you know, you start with energy like function, which is a loss function. Okay, and then from that, you, you define free energy. You do not start it from the phenomenology, start from Microsoft. But in this case, in that particular limit, that was the answer in more complex systems, where we're not dealing with Gaussian's. And in the critical systems that we discuss, we're not dealing with Gaussian. So there's, you know, there's many, many, many scales are important. And those limits think so much more complex. And you cannot really give this formula and say it's exact. Nice to do it. Either perturbative. So, so, so yeah, in this limit, Laplace in all the free energy was important. And more general, it may be, you know, have lots of lots of corrections. Or as you know, in the perturbation theory, it may happen that it's not just correction, but they're dominating everything. And then, yes, yes, your zero thought is wrong. It's a non-pertorbative limit. And then, right. And the answer is complete. There are, and there are reasons to believe that the system is in this sense, non-pertorbative because of the criticality that we have drove. And so there's symmetry breaking transitions take place. And there's lots of complex things that, of course, I won't be able to talk about here. But yeah, analytical calculations are hard. But I guess without them, we are not going to understand, you know, what's actually happening, what's the relevant language of describing different phenomena. Does Karl first and free energy, is it an independent claim from, from yours? Does it emerge from your framework? No, I completely agree with him. So, so, so there has to be a free energy in this, in this set up. It's like, it's a phenomenological way of saying that there is a function that you will be, you know, minimizing, optimizing. And that's right. I mean, they're, they're, you started with their beginning, how come classical mechanics also optimizes? Yeah, but it only deals with the, you know, very close when you're close to the equilibrium in the sense of the learning dynamics. And the same with the, with the, so, so he has a phenomenological model that is very intuitive and very nice. And it, it is describes that such a function must exist. It's a kind of existence. Now, he doesn't start with the learning theory, but he starts with his understanding of how, you know, organism behavior makes sense, makes total sense. Now, I'll give you an example how we can, he's his free energy may be corrected to be much better. So, for example, you can have an organism that isn't interested in predicting environment, but is interested in quantizing gravity. Okay. So, so for that organ, organism, he will spend all his resources, maybe locked in the room with no windows trying to quantize gravity, writing equations. So this organism will have its own free energy. Now, it will not be the one that tries to, you know, predict how the environment doesn't matter how, maybe a little bit, I mean, I'm going to make sure that I survive. So, so on the on the higher levels, the free energy can be can be different for different organs. The question is whether you can derive them always from the microscopic dynamics. So, when you can do this RG flow and actually starting from some microscopic loss function assumption derived. And it's an open question. The one thing I can really add to free energy principle that Carl Fistens advocating is that we can model it using this trainable and non-trainable variables and think about what you get once the non-trainable integrated out and you pay attention to a handful of trainable variables. And then the system becomes something you can calculate and then you can model and now we can model it phenomenologically. So, if you have a controlled enough experiment, you don't care like what microscopic physics give rise to this advantage. You can just calculate by seeing how the system behaves to changing environment. For example, you say, I want to interested how the system behaves on the sound or on the light or on the temperature. And then you just model it as a function of those parameters. And that may kint you to actually how such a system would emerge from some microscopic. So, I don't think it's a contradiction to what Carl Fistens says. I'm saying that we can do dig deeper. If we assume that there's this learning dynamics happening on all scales. What's a piece of advice that you found inspirational that you keep coming back to? Advice that somebody gave me. You could be also that you read in your book. It could be from an advisor. It could be from a movie. Something you found that helped you. Yeah. Well, one advice I said, and that was very, very useful to me and actually maybe not something I would advise student to do, but it worked for me is to doubt everything. So, so, so I do not trust anything that is relevant for your work until you try to do, try to do as much calculation yourself as much understanding yourself. Now, why is this a bad advice? Is because it may not be optimal for a student who is trying to get professor position, tenure position, whatever, you know, if you are going to be doubting everything and doing all calculations yourself, you may just not you will rubbish, publish or rubbish. So, so that's that's a bad advice, but it was it is something I couldn't I would I couldn't not to do so once I figure out that there is no problem that I cannot solve myself. I said, OK, I'll be doing that. And of course, I haven't done all the calculations. There are. The archive is full of the calculations. I have done, but as much as I can. And so doubting is is I think one advice that that and doubting and also and that advice, of course, that I know about more concrete advice that was given by my advisor, Alex Lincoln. So he you know, I would come every day with some new idea and he gave me advice that somebody gave to him. And then I don't know how long it was. And the device was you come up with some idea, some theory of some equations. And then and then the next day, you should try to criticize it as much as it can. So like you flip up. And like, you know, try to act as if you are opponent to that idea, have this. Yes. Or even days of the mind. It's very helpful. Like a really objective look at what you have done and say, no, no, no, I don't like that you've done it. I'll try devil said that can't right? So I'll try to disprove it. And and find all of the problems with it. And that's why when I'm talking to you and saying, like, I know why it's may or may not works. So I'm trying to to to sell you, you know, and use car and without telling you that, you know, something is broken because that wouldn't be fair. I wouldn't feel feel right. And I would be confident about the coalition that I had done. So yeah, constantly, like, flipping with this, why is this wrong? Okay, one day is come up with ideas, do calculations next day try to criticize as much as you can. And maybe like last statement here, now the judge G, G, G, G PT is horrible in doing that, you know, or other, like, attend to agree, everything you say. And so so so my advice is push it into if my advice not to use charge GPT fork correcting your work without you verifying it. So no, using it for correct work is fine, using it for suggesting ideas is fine. Like it's it's an excellent tool. We just don't know how to use it yet. We are completely we have students of LLMS once we're long how to use it, but never never trust or verify we say in Russian, you should always verify it. And you should to the point that you redo the same calculations many times because honestly, how many times we make mistakes when we do calculations? Well, I do mistakes all the time, you know, we all do mistakes. And so so you should keep questioning that. And so it's related to the doubt and then but doubt even your own ideas, I guess that's what that's mind box. That was given to me. Professor, thank you so much for spending two hours with me was two hours. Yes, it went by like that. So quick, but space time don't exist anyhow. So at least not fundamentally right. Okay. Sure, that was fun. That was a lot of fun. I appreciate it. I have it on your podcast. It's, it was very nice talking, very nice questions, by the way. So I have so many more. - Let me just full disclosure. In addition to providing information, it was experiment for me, because from the time I conjectured that the world is near-all-ment work, every time I talk to a person, I conduct an experiment. How this person reacts to what I say. And so you've been a great, a great opponent, a great person to talk to, to actually- - A great guinea pig. - Yeah, to actually, so you have been experimenting with you. Whether you know it or not. And so, so, so, and at some point, when I will be constructing, not just theory of biology, which we've done, but a theory of psychology, I might be using some of this discussion as an experimental evidence of certain psychological phenomena. - All right, I'll take that as a compliment. - It is, it is. No, no, it was really great. I mean, it's exceptionally, I'm very happy with the questions that was very good. Thank you. I'm honest with telling this that, I've given interviews to a podcast when the people were not equipped at all with with any of physics, or the lingua, or the, and, and they just were pushing their own worldviews without trying to understand what, what I am trying to say. And that was really a torture for me, because, you know, the point of the podcast, or you, I think the point was, is, is, is to try to do both, try to understand what I'm trying to say, and then try to point me in the right direction. So that's what I appreciate. And I've had good experiences with people who actually, you know, done their homework, and then, you know, so it was obvious, you, you, you have a physics degree. So that helps a lot, because at least certain things, I may say, between the lines, and you, you know, put me back and say, okay, I'll clarify that more often. So that was very useful. And so I think that, that was, thank you for that. So it was great, great experience. - All right. Okay, take care, sir. And I'm sure we'll talk again. The audience is gonna love you. I guarantee. Hi there, Kurt here. If you'd like more content from theories of everything and the very best listening experience, then be sure to check out my substack at KurtGyMungle.org. Some of the top perks are that every week, you get brand new episodes ahead of time. You also get bonus written content exclusively for our members that's C-U-R-T-J-A-I-M-U-N-G-A-L.org. You can also just search my name and the word substack on Google. Since I started that substack, it somehow already became number two in the science category. Now, substack for those who are unfamiliar, is like a newsletter. One that's beautifully formatted, there's zero spam. This is the best place to follow the content of this channel that isn't anywhere else. It's not on YouTube, it's not on Patreon. It's exclusive to the substack. It's free. There are ways for you to support me on substack if you want and you'll get special bonuses if you do. Several people ask me like, "Hey, Kurt, you've spoken to so many people in the fields of theoretical physics, a philosophy of consciousness. What are your thoughts, man?" Well, while I remain impartial in interviews, this substack is a way to peer into my present deliberations on these topics. And it's the perfect way to support me directly. KurtJayMungle.org or search KurtJayMungle substack on Google. Oh, and I've received several messages, emails, and comments from professors and researchers saying that they recommend theories of everything to their students. That's fantastic. If you're a professor or a lecturer or what have you and there's a particular standout episode that students can benefit from or your friends, please do share. And of course, a huge thank you to our advertising sponsor, the economist. Visit economist.com/toe. To get a massive discount on their annual subscription. I subscribe to the economist and you'll love it as well. Toe is actually the only podcast that they currently partner with. So it's a huge honor for me. And for you, you're getting an exclusive discount. That's economist.com/toe. And finally, you should know this podcast is on iTunes. It's on Spotify. It's on all the audio platforms. All you have to do is type in theories of everything and you'll find it. I know my last name is complicated. So maybe you don't want to type in JayMungle, but you can type in theories of everything and you'll find it. Personally, I gain from rewatching lectures and podcasts. I also read in the comment that to listeners also gain from replaying. So how about instead you relisten on one of those platforms like iTunes, Spotify, Google podcasts, whatever podcast catcher you use, I'm there with you. Thank you for listening. The economist covers math, physics, philosophy, and AI in a manner that shows how different countries perceive developments and how they impact markets. They recently published a piece on China's neutrino detector. They cover extending life via mitochondrial transplants, creating an entirely new field of medicine. But it's also not just science. They analyze culture. They analyze finance, economics, business, international affairs across every region. I'm particularly liking their new insider feature. It was just launched this month. It gives you, it gives me a front row access to the economist's internal editorial debates where senior editors argue through the news with world leaders and policymakers and twice weekly long format shows, basically an extremely high quality podcast. Something else you should know about is that if you go to their app, they not only have daily articles, but they also have long form podcasts with their editors and writers. This is also available online. Whether it's scientific innovation or shifting global politics, the economist provides comprehensive coverage beyond headlines. As a toll listener, you get a special discount. Head over to economist.com/toe to subscribe. That's economist.com/toe for your discount.

Podcast Summary

Key Points:

  1. The universe may be effectively modeled as a neural network, where the learning dynamics themselves constitute the physics, leading to the emergence of fundamental laws like gravity and quantum equations.
  2. This approach differs from using neural networks as mere simulation tools; it treats the universe's inherent learning process as fundamental, with optimization algorithms (like Adam) naturally giving rise to curved spacetime for efficient convergence.
  3. The model suggests a unification of quantum mechanics, general relativity, and observers, framing natural selection and consciousness in terms of learning efficiency and information optimization.
  4. The professor clarifies he is proposing a promising mathematical model, not making an ontological claim about the universe's true nature, emphasizing the role of objective functions and learning trajectories absent in traditional physics.

Summary:

The discussion centers on modeling the universe as a neural network, where the learning process itself generates physical laws. Professor Vitaly Vanchurin proposes that the universe's dynamics can be described through neural network frameworks, with optimization algorithms like stochastic gradient descent and Adam naturally producing phenomena such as curved spacetime, which enhances learning efficiency. This approach leads to the emergence of equations like those from general relativity and quantum mechanics.

Importantly, he clarifies this as a mathematical model rather than a claim about ultimate reality. The conversation explores implications for consciousness, observer physics, and the unification of fundamental theories, positioning learning and optimization as core principles that differentiate this model from traditional physics, which typically lacks such inherent objective functions and learning trajectories.

FAQs

The universe can be modeled as a neural network where the learning dynamics themselves constitute the physics, rather than just using neural networks to simulate known physical laws.

It proposes that the universe's own learning process is the fundamental physics, with phenomena like gravity and quantum equations emerging naturally from this framework, not just as simulation outputs.

Optimization, through a loss function, is central as it represents a goal-directed dynamics absent in traditional physics, driving the system's evolution toward efficient learning states.

The Adam optimizer and similar algorithms inherently involve a curved metric on parameter space, which is essential for convergence and mirrors the emergence of spacetime curvature in the universe.

It means the universe's constants and laws evolve through a learning process to become 'carefully chosen,' allowing observers to emerge, rather than being fixed from the start.

It argues for a unification of quantum mechanics, general relativity, and observers, suggesting observers emerge because the universe learns to be observable through its self-tuning dynamics.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.