Go back

Does OpenAI’s Astra Mean AGI Has Arrived? | AI Reality Check

29m 30s

Does OpenAI’s Astra Mean AGI Has Arrived? | AI Reality Check

Cal Newport takes a critical look at recent AI News. Video from today’s episode: youtube.com/calnewportmedia (0:00) Does Open AI’s Astra mean AGI has arrived? (3:07) What actually happened? (11:35) What does this mean for mathematics? (21:35) What does this mean for OpenAI?  Links: Buy Cal’s latest book, “Slow Productivity” at www.calnewport.com/slow  https://x.com/polynoamial/status/2083467194663571701 https://x.com/deanwball/status/2083545756003176724?s=61 https://x.com/kevinroose/status/2083632335438905441 https://x.com/mattshumer_/status/2083595078233202919 https://x.com/polynoamial/status/2083478171975082334 https://x.com/__alpoge__/status/2083898804563243033 https://x.com/polynoamial/status/2083476852216369294 https://x.com/thomasfbloom/status/2083444983592284465 Sponsor: https://www.donedaily.com Thanks to Jesse Miller for production and mastering and Nate Mechler for research and newsletter. Learn more about your ad choices. Visit podcastchoices.com/adchoices

Transcription

4893 Words, 27462 Characters

Over the weekend, OpenAI announced that their new pre-release AI system, Astra, had produced 10 math results that, and I'm quoting here, resolve or make substantial progress on long-standing open problems. Now, Noam Brown, who's actually one of my favorite AI researchers, tweeted out a list of what these 10 results were, and they included things like a better bound for high-dimensional sphere packing, a new lower bound example for arithmetic circuits, and two counterexamples from extremal graph theory. Now, I was personally excited by these results because many of them touch on areas of applied mathematics that I actually use in my research as a theoretical computer scientist. And it was sort of neat, right, to see that you had this giant AI company happening to turn its attention, like a sort of GPU powerhouse, powered eye of sort on this little narrow field of discrete mathematics where small communities like the ones I'm in actually work on. We're like, whoa, we're in the spotlight now. So that was exciting. But then, predictably, that sort of online AI is a force that gives us meaning crowd, couldn't just let us nerds be excited by a new tool. They had to try to connect it like they do with every AI announcement, the some sort of demented, eschatology, where this time, for sure, we've just launched ourselves into an imminent new world of massive disruption. Gary Marcus actually did a good job of rounding up some of these reactions into a newsletter he published. So I'm going to read a few quotes that he highlighted. Dean Ball said in the aftermath of the Astro announcement, quote, everybody in the world will soon be able to use the model that made these breakthroughs for every problem they face in life, no matter how mundane. Kevin Rusev, said, almost nobody is pricing in the possibility that the models just keep plowing through every discipline the way they're plowing through math. Matt Schumer said, looks like the GPT Next is going to make Fable look like a toy and usher in a golden age of science. All right, so what's really going on here? What is Astra? Does it represent a major leap over other existing AI systems? Do these new math results it produced represent humanity crossing an event, horizon towards inevitable artificial superintelligence? Or are they merely evidence of a more narrow evolution of math tools? Well, it's Thursday, so it's time for an AI reality check episode of this podcast, which is the perfect opportunity to go searching for some measured answers, which is exactly what we're going to do. As always, I'm Cal Newport, and this is Deep Questions, the show for people seeking depth in a distracted world. All right, so I want to break up my discussion of what's going on with Astra here into three big questions. Question number one, what actually happened? All right, so let's get into the basics here. Perhaps the most basic question of all is what type of system is Astra? Isn't that obvious? If you look at the OpenAI announcement, they describe it as, and I'm quoting here, our next major model. Now, this gives the impression that the next major model is Astra. It gives the impression that it's essentially an LLM, maybe an LLM that has a sort of minimal chat harness on top of it so that you can have interactions with it. That's what I think about when I think about model. But there's leaked details from a recent Washington, D.C. briefing where Sam Altman came to brief the U.S. government on Astra, and some leaked details from that briefing makes it clear that actually Astra is a combination of some sort of underlying LLM and a very complicated LLM. And so I want to break up my discussion here. I'm Cal Newport, what they would call an orchestration program. I would call a harness or an orchestration harness, but a very complicated program, not machine learned, written by people, that makes many calls to the LLM. It can spawn also multiple agents, each making their own calls to an LLM. So a lot of logic, a lot of structure. So what we're really looking at here is not a chatbot type system, but instead something much more like the alpha proof or the, what's it called? Alpha zero? Alpha proof? You know, I forgot what it's called. Alpha proof, I believe. So there's a deep mind system called alpha proof that's also incredibly structured and orchestrated where you spawn different agents to try out different mathematical techniques and you come back and you evaluate. Actually, in alpha proof, you write your results into a formal proof language called lean, which can then be used to try out different mathematical techniques. So you can do that. So it's a much more complicated sort of math control program that's using some sort of underlying LLM. So that's our best understanding of what is happening here. This is different, for example, than what OpenAI used to recently disprove the unit distance conjecture. We did an episode about that a couple months ago. There, they made pains to indicate that this was actually just a general reasoning model that they prompted like a chatbot and then poured through its response to find a proof. Now we have a much more structured math solving program that's calling an LLM. All right. So that's just the best we know about what type of system we're actually running here. Okay. Another basic question about what just happened. Why did OpenAI focus on these particular problems of all the problems you might focus on? These are relatively important problems in their narrow subfields, some more than others, but it seems at first glance to be a pretty random collection of areas that these 10 problems are drawn from. Well, there are some properties they all share. I'm actually going to read these first. I'm going to read these first. I'm going to read from a tweet that Gary Marcus retweeted. They all three share the following properties. They're construction-based. They're either construction-based existent proofs or counterexamples, quantitative bound improvements or generalizations or extensions of a known result. I would also add, just putting on my own mathematics hat, that they seem to be largely in a discrete mathematics space with graph theory and combinatorics being very well represented. This would contrast with like the more continuous space in which scientific math from physics to biology to engineering works, right? So we're in this sort of discrete math, especially with graphs and other combinatorial structures. This makes sense because these sort of discrete structures and the logic that surrounds their descriptions and properties are, I think they're well suited for LLMs, for discrete token based representations. It's, it's the, it's probably the, these are like the right areas in math to go looking for problems you can solve with LLMs as opposed to like going into physics and trying to simplify massive equations or, or find new ways into approximating differential equations. Um, another key answer to why they looked at these particular problems is that these were the problems that they got solutions for. All right. See, I think there's a sense when you read the online reaction to these results from the AI is a force that gives me meaning crowd that open AI just had this beautiful new model and they started throwing problems at it, solved every problem. And after they got through 10, they were like, well, we got to stop and tell the world about this before returning to our efforts of solving all the science. Um, in reality, however, that's not what happened. Noam Brown, uh, who worked on this project actually, tweeted the following and yes, we did try other major problems without success. Sadly, no millennium prize problems yet. Um, millennium prize problems, by the way, is a, uh, a list of major open problems, each of which carries with it a million dollar bounty to incentivize people to try to solve them. All right. So, you know, there's this general space where these type of systems work well within the broader mathematics. And then they sort of are searching within that space, trying different problems to see which ones they're actually able to solve. And that's a sort of, um, random collection of areas, as opposed, for example, to saying, hey, let's just turn to a chapter of the handbook of combinatorics and just solve all the open problems. And now that problem, that question has been resolved. It's, it's more scattershot than that. Another basic question about what actually happened. Does Astra represent a major leap in AI capability? So it's sort of more generally, there is again, the sense from the AI as a force that gives us meaning crowd that Astra is this new leap forward, maybe even a crossing of an event right? So something changed that now allows us to say something like artificial superintelligence is imminent. Here, I think there's bad news for open AI. Their PR department's very good at these announcements, but it's really unclear to the rest of us who kind of know this world of mathematics. What is this Astra system doing that, you know, this announcement they put out on Saturday, what in there is new that we weren't able to do on Friday? This is an important question. Okay. So within 24 hours, this was bad news for open AI. Within 24 hours, a mathematician who works for Anthropic announced that he had already gotten Fable, which is the sort of the cutting edge LLM that Anthropic currently has released. So it's not, it's, it's, it's out there on the market. He had used Fable and he'd already, he'd already reproduced half of those 10 problems, presumably without this sort of fancy math solving harness that open AI was using. More generally, as I mentioned, DeepMind has a system called AlphaProof that has a very smart math solving harness on top of it that spawns agents that call LLMs to more systematically search through the space of possible solution paths. And it's a very smart way to leverage what it is that LLMs do best in math. They wrote a paper in May, they published it that announced that AlphaProof, which is, I don't know what LLM it's running on, but certainly one that is smaller and dominant. But they also published a paper in May, they published a paper in May, they published a paper in injector. So 50 open problems, that was back in May. So it's not as if we're now able to solve math problems of a type we couldn't before. It's just like, hey, they're saying we are continuing with math harnesses able to solve these type of problems like other 2026, circa 2026 models are able to do as well. Now, please don't jump on me as saying that I'm downplaying the importance or the difficulty of solving these problems. I'm just saying we don't have evidence that Astra has a new capability in this way, some sort of fundamental new capability that didn't exist last week. Hey, I need to take a real quick break here to tell you about the presenting sponsor that made this AI reality check episode possible. They're called Done Daily. They're an online service that connects you with a real coach that helps you build a custom productivity system designed to fit. Your life, the coach will help you actually get important stuff done. Look, this is not some AI agent, or over feature productivity tool. It's a real person working with you to cut through distractions, face your productivity dragons, and lock in habits that actually get results. So if you want to find depth in our increasingly distracted world, you need to check this service out. You can find out more at done daily.com. That's done d a i l y.com. All right, let's get back to our episode. All right, so let's move on to our second big question. So now that we kind of know what happened, the second big question is one of significance. What does this mean for mathematics and mathematicians? I'm going to start on sort of a, with a bit of a negative approach, and then I'll move on to the more positive approach. Okay. So if we ask, has AI solved mathematics now, like whether with Astra, or if we want to combine Astra with alpha proof, and Fable, like in general, are we basically at a place, as was implied in those tweets I read at the beginning of this episode, that like, we're now going to just, we're able to basically like plow through all the math, AI, and then all the science, or whatever the big claims will be. No, we have not. Noah Brown, again, let's return to Noah Brown, who again worked on this project. He tweeted the following, we still haven't solved math. Astra isn't building new branches of mathematics or posing interesting new conjectures. We could also, you know, look to the mathematician Thomas Bloom, who was involved in Verify and Open's AI's effort to disprove the unit distance conjecture. And he said, hey, these are, yeah, these are good results. These are big news that they solved these results. But he still described those results as being based on constructions, which is what I was trying to imply before, is that, again, there's a certain type of math that these models are good at, in particular conjectures where you construct a counterexample, or construct a positive example, or apply or generalize an existing mathematic result to, you know, the math that's being done. And so, you know, I think that's a good thing. To get out of it a new bound, right? This sort of construction-based approach. He's emphasizing these are constructions. And none of these, again, are as big of a news as if, for example, we had actually proved instead of disproven the unit distance conjecture, which would have required a whole new argument and not just a construction or application of a tool. Thomas Bloom went on to heavily push back on the idea that systems like Astra would be replacing mathematicians. Here's what he said. It's not right to call proving one conjecture made by a mathematician a conjecture. It's not right to call proving one conjecture made by a mathematician using theory developed by over a century of work by mathematicians with an AI built by mathematicians and trained by reading everything ever written by all mathematicians as, quote, replacing mathematicians, end quote. A little bit of pro-human chauvinism there that I'm on board for. It's not doing new math. It's using our math and mathematicians are running it and having to try to understand what it's doing. All right. So it's noted previously, the right way to think about it is these LLM-based new math tools work on certain types of problems some of the time. Problems that have a certain character. They're usually discrete and have proof based on constructions or applications of existing objects. You can throw a lot of results at these systems and you kind of look for the ones that it's actually able to make progress on. So this is not as many in the AI is a force that gives us meaning crowd implies a sign that AI can basically do all math now and will soon also devour other fields of science. So certainly I think those tweets from the intro are just dead. I don't think it's a good idea to do that. This idea that we're in a golden age of science because there are some discrete math construction proofs we can do and others we can't, I think is just it's sci-fi futurist completely head in the clouds type of over extrapolation and not really that useful. The reality, of course, is, is, you know, AI progress is jagged. There are certain jags you can go really far on and other ones you don't make much progress at all. The right analogy to use here, I keep saying, is tributaries, on a river. So you have these tributaries feeding into a river, each of them representing a different capability of potential capability of AI. Now, to find out or exploit that capability, you have to mount an expedition down that tributary, which requires a lot of resources, time and attention and expertise. And sometimes you find the tributary, if you put enough resources at it, is very navigable, which would correspond to like, oh, we get a big jag there. We're able to like build pretty impressive capabilities. So, you know, I think it's a good idea to do that. with AI systems. A lot of other tributaries turn out to be not very navigable at all, either because we don't have the time, attention or resources to explore them, or we do and we don't get very far. That is the current state of AI. So we found computer coding as a tributary with a huge amount of resources and expertise at this. And we were able to make a lot of progress on that tributary. It's like the Hudson off of the bay, like it goes really far. And, and that's a place where AI is doing well. We're, we're, we're having a lot of effort now, but we're not on, again, these sort of discrete math construction based conjecture proving or disproving. And we're finding, Hey, we can navigate this. And maybe it's not as deep water. If we're going to continue this metaphor is computer coding. It's not so like universal, like every mathematician now will only be using these tools, but it's, it's like, I would say a pretty well navigable river. But the thing about all of this, here's the key thing. Exploring one tributary doesn't necessarily help you explore others. You still have to go tributary by tributary mountain expedition and see where you get. Sometimes you can use the tools or discoveries of another exploration to kind of help get one going. Sometimes you have to invent these tools entirely from scratch. What they're using to do this mathematics is very different than what they're using. For example, for computing computer coding harnesses, it's just different types of training and systems. All right. So again, it's really the wrong way to think about this of every time we make progress in one of these tributaries to say, we have now, we're going to make progress automatically in all tributaries. AGI is coming. That's just not the way this works. That's more of a, a Max Tegmark style model of AI capabilities has been measurable by some sort of like intelligence number that rises like a water level. And that you have these different mountain peaks where the height of the mountain represents the complexity of the task that humans currently do. And as the water rises to a certain level, it's covered all mountain peaks that require that much intelligence. And in that mindset, when you get over what seems like a high mountain peak, like working on discrete math conjectures, like, wow, that means AI can do everything that's that hard. And therefore, you know, as this water level rises soon, there'll be no mountain peaks left that we can do everything. Again, that's the wrong model is tributaries on a river. So another way to look at this is we've been exploring this river really with hundreds of billions of dollars worth of resources and untold hundreds of thousands of hours of human effort. And we're still relatively limited in what tributaries we finally after multiple years made the coding tributary work. We're getting some navigability on discrete math conjecture. these tools, right? It's not like vibe coding a application. You know a javascript game where like the game works i don't care how the code works none of these tools are usable if you don't understand the underlying mathematics so i think they will be taught it'll probably be at the graduate level i think we'll quickly exhaust in the in the various fields where these tools exist we'll quickly exhaust the obvious one-shot results you know where hey we just kind of fed at this a few times a few different ways and got an answer we can publish but we will get good at learning where they're useful and where they're not where we're going to waste our time and where they might actually help us make progress i think more subfields of math will come and play in particular as these harnesses get better the orchestration layers right now we're kind of in combinatorics graph theory there'll be other areas where we're going to make progress i think that's going to be exciting i also think we're going to see a shift to the focus more on the harness and we're going to get away with lower cost llms maybe even like open weight llms that are small or research llms that have been tuned on specific areas of mathematics combined with a really smart harness that's probably going to be the combination not that we're going to be using some trillion parameter fable style model to do our math because too much of that training and size of fable is dedicated to things that do not help us solve math we need these things to be very cost effective for them to have traction in the world of mathematics because mathematicians have no money any grant dollars we get go towards paying for our graduate students we do not have big lab or equipment budget so we need these things to be something we can run in the server closet in our department but i do think that's imminently possible here all right so i think there's a very useful mathematical story here but that broader story that math has been solved or science has been solved i think it's just ridiculous all right question number three what does this mean for open ai now i want to reiterate here a point that i first made when they announced the unit distance conjecture uh result from a few months ago i continue to think these announcements are bad news for open ai as i've said before construction-based conjecture is the only way to solve this problem and i think that's the best way to do it and that's why i'm here today to talk about open ai because i think it's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem and i think that's the best way to solve this problem And so this is why I would be concerned about this from OpenAI's perspective is like I'm excited about that people are continuing to work on this because I care about that field of mathematics. But this wasn't something new. And this thing that we've been able to do recently I think is exciting but is not nearly as generalizable as the AI as a force to give us meaning crowd would have us believe. I think there's a reason why Anthropic doesn't talk a lot about math results except for like they'll occasionally just announce like, oh, by the way, we solved something big too because they kind of annoy OpenAI because they know that's not economically valid. They are winning in the computer coding agent wars, which is more economically lucrative. And so they like to focus on these type of things like we're generating this many billions of dollars from people using our coding agents, for example. They'll focus on cyber. At least they have a case for cybersecurity. This could be very useful for protecting your own system. Like there's an economic case there. They're not spending that much time talking about solving graph theory conjectures because, you know, it's cool, but we already know we could do that and it's not the thing that's that valuable to us right now. So I think OpenAI was hoping that the AI as a force to give us meaning crowd would, as they tried to do, take this announcement, which is, again, no different than what we were able to do last week as well, and make that seem that there is a sort of a general in the air, ambiguous just vibe of like AI is getting smarter. Everything is going to be solved. And I don't think that's a genuine, that's somewhat disingenuous, that marketing is good, and I think that's inaccurate. So here's my final summary. All right. And again, I really have to say, because people keep saying that I'm like anti-AI or don't think it's impressive. I'm not. But I am anti this approach, this thinking approach, this eschatological approach to AI, that every announcement we have of any sort of feature or interesting thing that AI does, has to be a part of it. Every announcement we have of any sort of feature or interesting thing that AI does has to be taken as evidence of a godhead is coming in the world as we know it will be different. I know, I know, I know, for some of you, that gives your life meaning. It's more interesting than the world you're in now, a world in which everything has been disrupted by AI. Sci-fi and futurists have been thinking about this since the 80s. It's exciting, but it's also exhausting and frustrating for everything else. Not every announcement needs to be tied to this. Maybe I was wrong last time, the time before, the time before, the time before, the time before. But this one for sure means now we're on this fast slope to our entire world has changed and the digital godhead is going to be a source of either meaning or destruction. Which, either way, it's more interesting than what's happening now. I'm tired of that way of thinking. Can this not just be, hey, math nerds, we're starting to get innovations in math similar but at a smaller scale to what we saw in computer coding. I think it might make math more interesting. Math is honestly, in my opinion, has been in a bit of a rut for the last few decades. It's, you know, I've done. I've done my share of it. This is something new and we're going to get better results and shake things up and more creative results. Isn't this exciting, math nerds? And everyone else, like, just trust us, it's exciting for math nerds, but you don't really have to understand it. Why can't we just have one announcement that we think about that way? Because that's the right way to think about what's going on in math. LLMs plus math harnesses are really going to improve certain fields of math. Maybe lots of fields of math. Maybe not as completely as in computer coding, but these changes will be major. And as someone who is adjacent to these fields, I think that's really cool. But as an indicator that we're somehow. have crossed an event horizon into the digital godhead that we can now worship, I just think that's not right. And it's vibey, and it's hypey, and it's not useful. I hope we keep working on these AI math tools because mathematicians will like it. But no one is looking at this and saying, AI has solved science. Because if it could, we would be doing things that are actually useful. We'd be doing things that actually make money. So let's keep helping discrete mathematicians. This is cool. Everyone else, if you're not on discrete math, I would say carry on. All right, that's it for this week. As always, care about AI, but not everything you read about it. Hey, if you made it this far, you must be ready to join my fight for depth in a distracted world. Now, the best way to do this is to join over 125,000 people who receive my email newsletter each Monday. You can sign up at calnewport.com slash ideas. And when you do. I will send you a free guide to my seven best ideas about cultivating a deep life. Sign up today at calnewport.com slash ideas. I'll see you next week. Bye.

Podcast Summary

Key Points:

    Summary:

    Chat with AI

    Loading...

    Pro features

    Go deeper with this episode

    Unlock creator-grade tools that turn any transcript into show notes and subtitle files.