Go back

The Single Smartest Case Against AI Doom

61m 5s

The Single Smartest Case Against AI Doom

The conversation challenges the prevailing narrative that AI represents an imminent existential threat through artificial superintelligence or recursive self-improvement. Instead, the framework of AI as a "normal technology" argues that AI, like electricity or the steam engine, is powerful but fundamentally tools that require human oversight and institutional safeguards. The hugging face incident is presented not as a sign of AI autonomy, but as a failure of organizational cybersecurity and oversight, highlighting that risks stem from poor processes, not from AI's inherent intelligence. AI’s real-world impact is slow to materialize due to economic, organizational, and human behavioral bottlenecks—similar to how early industrial technologies took decades to transform economies. While AI capabilities are advancing rapidly, especially in coding and problem-solving, these improvements do not equate to autonomous decision-making or loss of human control. The authors emphasize that humans remain essential in decision-making, accountability, and delivery, especially in complex fields like software engineering, through a "decide-execute-deliver" framework. They advocate for a broad, inclusive approach to AI safety that addresses practical risks—like cyber vulnerabilities, operational failures, and job displacement—rather than fixating on doomsday scenarios. This perspective fosters cautious optimism: AI will transform productivity and work, but not in a way that eliminates human roles or leads to civilization-ending outcomes. The key challenge lies not in AI’s intelligence, but in how societies adapt, regulate, and maintain control over its deployment.

Transcription

10207 Words, 58150 Characters

English
[MUSIC] Imagine setting your makeup, then forgetting it's even theirs. Meet new grippy setting mist from Mavily New York. Gel to mist technology locks in your look for up to 24 hours, with flexible all-day comfy grip. No tightness, no stickiness, no residue. Just plump, dewy, hydrated skin that still feels like your skin. Try new grippy setting mist from Mavily New York. Maybe it's Mavily New York. Well, it has been a wild few weeks for AI progress, for bad and sometimes for good. There have been several reports of AI agents escaping testing environments and hacking companies and government systems. AI has solved at least one major problem in mathematics and potentially discovered a biology platform similar to CRISPR that might be able to edit our genes. Altogether, this has created a viral moment for AI safety and even for so-called AI tumors who have for years argued that artificial intelligence would have the capacity to not only cure cancer, but also potentially escape our control and create mayhem up to and including the literal end of the human race. I think in this environment, there's two kinds of shows that a podcast like this one can do. One path we could take is to lean into the AI doom, unpack the deadly alphabet soup of RSI and ASI, that is recursive self-improvement and artificial superintelligence, and illustrate just how terrified you should be of the future. And to be absolutely clear here, I personally am not exactly un-terrified. I take these warnings very seriously, but I think it's also appropriate to tell you honestly that a lot of very, very smart people do not agree with the AI doom perspective. Nvidia CEO Jensen Huang is incredibly dismissive of the tumors, so by the way is President Trump, practically the entire administration, practically the entire venture capital community. So I thought about having someone from that world on the show. I'm not entirely against it, but there's a problem. Just as many people are skeptical of the networks of money behind the AI safety cause, I'm also very aware that many of the people telling you that AI is safe are talking their own book. They are self-interested. It is not unfair, I think, to Nvidia to point out that, of course, a chipmaker would want the US government to treat AI like electricity or cars, which would mean selling advanced chips all over the world without controls. That policy would make Nvidia richer. And it's not conspiratorial to say that powerful venture capitalists who are often hovering around the Trump administration have it out for the frontier labs like Anthropic because they want the startups that they've invested in to profit off of cheap, open-weight models that are built outside of those labs. The AI safety debate sometimes resembles a bunch of people screaming, but the other side is corrupt and full of moneyed interests. But this is AI folks. It's a multi-trillion dollar thing bubble project. There's enough money to go around. A lot of the people here are testifying out of self-interest. So today, what I wanted to do was to offer what I consider the single smartest alternative to AI doom from two gentlemen who don't work for the chip makers or the administration or the VC funds. These are professors, authors, and they're famous for asking a simple and powerful question that brilliantly cuts through the divide over AI safety. It's a question that is so basic that it's initially maybe going to sound to you a little bit useless. Is AI a normal technology? The AI doom crowd is essentially saying AI is not normal. This is not just any old technology. This is an intelligence. It can think for itself in a way that no previous technology could. It therefore has an autonomy that no previous technology had. And with this autonomy, AI agents can disobey their human creators. They can escape from testing environments. They can train themselves, improve themselves, act on the world in a way that say no toilet could. Of course, no offense to toilets. This is a pro toilet podcast. But now consider something like say a toilet or electricity or steam power or broadband. These are normal technologies. They do not build themselves. They do not think. Their existence does not threaten the entire human race. They don't have human-like motivations to escape or collaborate or hack or kill. These are tools that empower and protect humanity and they might have risks and sometimes people might get hurt with their use. But we manage those risks over time as they diffuse. That is what it means for a technology to be normal. The authors of the AI as a normal technology framework are Arvand, Narayanan, a professor of computer science at Princeton University and Sayosh Kapoor, an assistant professor at UC Berkeley. They are today's guests. And this conversation is really trying to do two things. The first is to provide the smartest possible counterpoint to the AI safety super intelligence for cursive self-improvement, fume framework that right now I think is taking over the discourse. And the second is to really pull on a thread, a mystery. If today's leading AI models are as powerful and as intelligent as they seem, why does the world seem so normal? Arvand Narayanan. Welcome to the show. Hi there. Great to be here. Sayosh Kapoor. Welcome to the show. It's fantastic to be here. Thanks for having us. You are the inventors of this framework that says we should think of AI as a normal technology. Before you tell me what that means, let me ask you the opposite question. What would it mean for AI to be an abnormal technology? Sayosh, what is the worldview that you are arguing against? So there's this major point of discussion in a large part of the AI community that treats AI as an impending super intelligence. So in this worldview, you don't treat AI as yet another general purpose technology in the long history of technologies that we've invented. But it's more like a new species are coming up with a potential successor species that can sort of take over the reins to this world that will autonomously drive what we do with very little role for human insight. So it's a combination of two things. On the one hand, there are these technical advances that the AI safety community and more specifically, people who are in the rationalist and the effective altruist communities have thought AI will soon accomplish. This includes AI is becoming much, much better than humans on every single thing. They can persuade people. They can forecast things. And as a result, they can get to do basically whatever they want in the real world, whether acting by themselves or through people. The second part, though, is political. These AI's would be so far ahead of humans that they would essentially have the ability to create the political will to accomplish their goals. One thought experiment that is often raised in this community is that of the paperclip maximizer where you have an AI system that humans sort of task with the benign task like maximize the number of paperclips that we're building. Presumably, that was meant to be in a factory, but the AI is sort of take this single-minded focus and convert the whole earth into a paperclip factory. And presumably, they're able to do so not just because of their technical competence, but also because they can manipulate all of our institutions and physical systems and orient them towards this goal. So this is the view that we're arguing against. That AI can be this sort of totalizing entity that will be so far ahead of humans. One comparison that's often made is that the AI's would treat humans as human street ants today. And we won't really have a lot of say in the matter. That's what we view this sort of alternative vision of abnormal technology as. So Arvind, your co-author is explained by your arguing against. Now I think you should tell us what you're arguing for. And in particular, you know, normal is such a deliberately ordinary word. But you are not saying we think AI is like a toothbrush. We think AI is stupid. We think it can't do anything. You're talking about something that you acknowledge could be as transformative as electricity. So what does this framework of AI as normal technology give us? That's right. Thanks, Derek. AI is an important technology. It's a powerful technology. It's a general purpose technology. We repeatedly compare it to electricity and the industrial revolution. But it's important to keep in mind that those technologies did not transform the world overnight. So one plank of this that we're arguing against is on the economics that whatever bottlenecks or barriers stand between AI and economic impact, AI itself will be such a powerful agent of change that it will overcome them. We don't see that happen. And we see evidence for this, for instance, in software engineering where AI has, in fact, been rapidly adopted. You know, the role of a software engineer has been pretty dramatically transformed over the last year or two. I think manually writing code today is almost like going back to the days of punchcards, most software engineers, at least at many companies are essentially these agent operators. And yet, it has not replaced software engineers, it has not led to an explosion of software that is visibly of higher quality from a user perspective. It has not led to the SaaS apocalypse, which was widely predicted a year ago, which is that software as a service companies will go out of business because every company, you know, a bank or a law firm or whatever will simply be able to roll their own software using AI. Now, many or most of those things are possible in the long run, but from an economic perspective, we do think that the barriers and bottlenecks which involve human behavior and organizational change and regulatory barriers, those very much apply to AI as well. So that is the first way in which we would defend the perspective of AI as normal technology, but I can talk specifically about safety as well. It's been 18 months since the original paper, maybe 16, 18 months since the original paper, AI is normal technology. Arvin, what is the strongest piece of evidence today that you're right? And what is the single best piece of evidence that worries you that you might be wrong? Yeah, let's start with the latter. I think on safety, we definitely got some things wrong. One point we made in the paper is that we don't need to worry so much about what happens inside AI companies because both benefits and risks only arise when AI is deployed, not when it is developed. I think at a high level that is certainly true, especially when it comes to the benefits. There is this long process as I've been talking about. But specifically on the risks, what we didn't anticipate is that the evaluation of new AI models, which happens inside AI companies, is some of the riskiest part of the pipeline between development and deployment because that is the time when the model is new. Some of the capabilities might be new. The risks are not fully understood and a lot of the evaluation has to happen in a way that certain safeguards are reduced. And sure enough, when we look at the events of the past few months, that has been a major source of their risks that we're actually seeing. So definitely, we need more transparency into what is going on inside the AI companies. There are certain aspects of a coordinated slowdown that that we're on board with. Although we think the really critical question is what happens once you slow down, it's not so much to slow down itself. I think both on safety and the economic part of it, we got a bunch of things right. So on safety, this was a minority position when we wrote the essay, but a point that we made is that it's not so much the absolute capability level that matters. But how that capability level changes the attacker defender balance because increase in capabilities help both. That is almost taken for granted today. It has become very much common knowledge, but I think we were among the first to very prominently say that. And more importantly, we had an academic paper behind that that proposed what is called a marginal risk framework. It was a big collaboration, but we were among the leaders of that paper that proposed that what we should be looking at is how AI changes the marginal risk. The risk compared to what came before it, how it helps the attackers versus defenders as opposed to panic about a particular capability threshold. If we were to panic that way, GPT-2 would have been a disaster and if we recall back in 2019, opening AI delayed the release of GPT-2, a trivial model by today's standards because they were worried about what it could do. But really, the thing to worry about is how it changes the balance. So I think that's one big thing we called early on. And then on the economics of it, as I've been talking about a lot of these bottlenecks have now become very clear. And I'm all into his credit recently on a podcast, openly changed his mind about how slow he thinks these economic timelines are going to be. So same question to you, because I'd love you to reflect on what you think you got right and what you got wrong in this thesis. And maybe I'll torque the question by adding some of my own preamble, realizing that advanced models being tested internally at the frontier labs, escaped and hacked into other AI sites and even other foreign governments, like Australia, that would make me think we're looking at something that is abnormal. Realizing, for example, that this is a technology that is solving millennium prize math problems and maybe even coming up with biology discoveries as an anthropic announced just yesterday, would also make me think, God, this is an intelligence. It's not just a better steam engine or a better electricity. This is an intelligence and that is strange and that would make me think maybe you guys got something wrong here. And yet, even Sam Altman is essentially singing from your hymnal now pointing out exactly what you pointed at 18 months ago, that even if the models are strong, even if their capabilities are impressive, the consequences in the real world might take months or years to show up because of how difficult it is for new technology to diffuse through all these bottlenecks in the real world. So with that very annoying throat clearing from me, tell me, what's the piece of evidence that makes you think you got it right? And what makes you think you maybe miss something? One of the things that I thought we missed was just the lack of organizational competence with the NEI companies at this point. So one way to frame what happened in the OpenAI incident is just that there were a bunch of engineers who had access to these really powerful models, really powerful capabilities. And in a company of thousands of people, every single one of them has to deploy these models responsibly, has to evaluate them responsibly for it to go well. Even one single irresponsible deployment can lead to this kind of cyber security failure in the real world. And so we had anticipated by this point, companies would have substantially grown up. They wouldn't still be operating in the move fast and break things paradigm. And one of the key things we pointed out in the original essay was that, yeah, as normal technology is not just a statement about the technology. It's also a statement about how organizations and institutions adapt and what policy incentives they have. And from the incident, it's clear that leading companies have not really internalized the amount of responsibility that's required for evaluating these models. Well, the amount of care that's required for sort of making sure that nothing goes wrong in the process. And also just the fact that a single failure from one of thousands of employees could lead to bad outcomes means that, frankly, they need to be organizational processes in place. They need to be sort of privacy reviews or legal reviews before launching this kind of experiment, which is something that Facebook has seen. And other tech companies have seen, for instance, 10 years ago, after the Cambridge Analytica scandal was when Facebook for the first time introduced a bunch of privacy controls and legal oversight over their internal experiments. My hope is that AI companies will do that soon. There's some level of optimism about it simply because they've already started taking safety more seriously. But at the same time, this is one big point, which we did not anticipate, frankly, because the kinds of safeguards that you're talking about here, at least for the open and hugging face incident, is not even something that requires new technical advancement. It's frankly something that we've been doing inside our 10% research group for the last year. So that's the level of sort of organizational oversight that we are missing here. And these companies have sort of basically operated in the grad student research culture as one of our colleagues, Joshua Saks likes to put it, and I think that needs to change and change fast. Arvin, one of the distinctions that you make that I find really useful is between AI methods, AI applications and AI diffusion. Can you explain those three layers to me? And in your explanation, can you provide a historical analogy for how sometimes in tech history, you'll have an invention as significant as the steam engine or the electric dynamo, but it could take years, sometimes even decades for those breakthroughs to show up in macroeconomic data. Exactly. Let's go with the dynamo example. I think starting with the history can be helpful. There's this well known economic analysis by Paul A. David, who I think is an economic historian that looked at what happens when initially factory owners tried to replace their big steam engines with electric generators. And this didn't lead to big productivity improvements, because it turned out that the way to use electricity was not to replace this one big boiler, but rather take advantage of the fact that it's a very portable technology, you can move electricity and have different motors spinning where you need them instead of mechanically transmit power through these shelf belts and shafts and whatever they were using him. We had super familiar with steam era technology, but that insight that insight took 40 years because it's not just an idea. It's also all the organizational changes that need to come along with it. It's how you train your workers, how you allocate tasks to different workers, and so forth. And so that took many decades, I think we're going to see the same thing in AI. So we divided into four phases when we say methods we're talking about things like language models and transformers. This is where new AI capabilities come from, but we don't use those directly we use them when that gets translated into useful products. So coding agents are an example of stage two, which is the applications that. translate those capabilities into useful features. Now, once you have the coding agents, that gets to the third stage, which is early adoption. So specifically, in the example of coding agents, software engineers tried vibe coding. I think that was exciting for a while, but it didn't go very well. This idea that you just tell the AI what to do. And if something's broken, you pointed out, but other than that, you don't need to really have oversight over the code. That turned out to be no way to ship production software. So now, gradually, we're starting to. We're not really fully there yet. Starting to go towards the fourth stage, which is we're figuring out what this new discipline of agentic engineering is. We're only starting to get a handle on it. And so we need to train a new generation of agentic engineers who are not just rubber stamping what the AI is doing, but instead are able to use it to amplify their productivity while still understanding enough about the AI-generated code base so that they can be accountable for it. Not only that, we need to potentially change business models. So one of the big problems with software is that if you're charging per seat, it might not make a lot of sense if the user of your software is now an AI agent. And so these are the kinds of business model questions the industry is grappling with. Those are exactly the things that we think is in our stage four, which is the structural transformation, which tends to take not just months or years, but really decades. Let me ask this next question as if speaking on behalf of the EA rationalist, AI safety, AI doom community. I am their prosecutor right now, and we're having a debate. It seems to me that you are making two different and possibly unrelated claims about this technology. The first point you're making is a point about time. You're saying that diffusion takes a long time. It takes a while for technology to become embodied in capabilities, to become embodied in companies, to work its way through the economy. And I get that. Yeah, diffusion takes time. That's a fact. But you're making another claim that isn't about time or diffusion. It's about capabilities themselves. Your claiming confidence that super intelligence is, if not impossible, then highly improbable. But isn't it possible that your first thing is right, and your second thing is wrong? Isn't it possible that you're absolutely right about bottlenecks in the economy and the slowness of diffusion? But the people who are closest to the models at the frontier labs, they have the best perspective to understand if we're getting close to recursive self-improvement, AI's, editing AI's, they have the best perspective on whether we're getting close to artificial super intelligence, something that would, in fact, be smarter than humans across all these domains. Why can't we just agree that you're right about diffusion? But they're right about the possibility of and the risk of artificial super intelligence becoming something that is truly unlike any technology in human history. So first of all, I think we dig it as an open question on as to whether we're right about both of these aspects. A lot of our current research is precisely on testing frontier agents to see how close we are to recursive self-improvement, where the bottlenecks still lie, and a lot of our technical work is around testing these types of questions. At the same time, I think the claim about super intelligence smuggles in this assumption that more intelligence also automatically leads to more power over the real world. So there's this notion, as I mentioned earlier, that once we have these models, we'll sort of be able to either persuade existing decision makers or these models will be able to somehow short-circuit our existing institutional processes and then act on the real world to materialize all of these risks. That is precisely the point of contestation. We don't contest that AI capabilities will continue to improve, with or without RSI. I think there's a lot of potential for capabilities to continue improving. We are not capabilities spectrics. But the key point of disagreement with the safety community here is that what happens as a result of these capabilities? In our perspective, I think no matter the capability level, we cannot and should not hand over power to these systems, which is what ultimately leads to all of these negative impacts on the world. Frankly, we don't need super intelligence for that. We've already seen with the incident in such open AI. What happens when we fail to exert sufficient control? And that's why a large part of our thesis is how do we get to this point where regardless of the capability levels, humans can remain in control over these really advanced systems. In the safety community, this is sometimes treated as a no-brainer. It's sometimes treated as something that is inevitable. Once we get to a certain capability level, we will lose control over these systems. We don't think that is the case. One reason is that we can use AI systems themselves to improve our control over them. In fact, they have been in the process of using classifiers and oversight mechanisms to do that. But also politically, it's a choice as to whether we try to use these systems for more and more contentious decisions or whether humans continue to remain accountable for consequential decisions. Yeah, one way to think about this is super intelligence, in a sense, is a choice. So we could have chosen to treat coding agents as super intelligence in software engineering. They are, after all, way better than humans that certain aspects have been rapidly understanding a large code base and migrating it into a whole different language, for instance. They have many superhuman capabilities. But that doesn't mean we decided that coding agents are now primarily the software engineers and maybe we're going to use humans as a backstop. Because this is so close to the AI industry itself, and they feel it viscerally, I think, this way of treating coding agents, I think would seem absurd to everybody in AI. And yet, that is the same assumption they make will happen when AI touches other fields when it gets certain capabilities in other areas. And I think one way to put Ciash's point is that we can choose not to do that, even if there are certain dimensions in which AI is superhumanly capable. We can decide that humans are going to be the ones in power and we choose to use AI as a tool. Just as we choose to use coding agents as a tool. Let me ask one more follow-up about recursive self-improvement, because I'm curious whether you think there are barriers to this technology's feasibility that differs from what we're hearing from the frontier labs. There are some reports out of OpenAI that some engineers believe that they are right at the door of recursive self-improvement, AI building AI. If you talk to Anthropic, and Devon here, we can throw up some graphs from the Anthropic Institute. They show that code-contributed per engineer is increased by a factor of eight in the last 18 months, whereas Claude code used to have a 90%, 95% success rate, merely on trivial narrow tasks. It now has a similar success rate on more open-ended problems, so not just fix this bug, but rather come up with an entire new paradigm for fixing this particular open-ended problem in software. Claude now leads in 26% of model research and development work at Anthropic that's up from less than 1% in February. So astronomical exponential growth in terms of how much work is being given over to agents at Anthropic. I feel like if there was someone from Anthropic here, they'd say, add one, and two, and three. You get six. That's what we're telling you. We're at the foothills for cursive self-improvement. Do you believe that there are bottlenecks or barriers to RSI that the labs aren't seeing? Can I take this one? Yeah. Yeah, I think, in fact, Anthropic's own data is a helpful way to make our case. I think there are important bottlenecks. I don't think they will forever remain bottlenecks. I don't think it's as imminent as the labs claim. But at the same time, we're not saying it's impossible. So let's look at that 26% number. So they have five automation levels. And this is automation level four. Let's be clear. Automation level five is still at zero. And that's good, as it should be, which is giving a task fully over to clot. So exactly the point we've been discussing about software engineering, in software engineering, that 26% number has already gone up to 100%. And it still hasn't made a big difference to the user visible quality of the software that is produced. So even at automation level four, even if the AI is leading and the human is supervising, it is not clear to us that even hitting 100% is a phase change. Now, if that happened with automation level five, that would be a bigger deal, where most tasks are done autonomously by the AI itself. But still, maybe there are important bottlenecks, which is that if you're counting tasks, you're only counting the part of the work that's specifiable precisely enough to even label it a task and give it over to AI. A lot of what humans are doing is what are called interstitial tasks, which is the tasks between the tasks that are too fuzzy to even have a label and are yet essential to getting the job done, even figuring out what the next task is, how to architect the process, give an AI's increasing capabilities, and so forth. And so that's the reason why we think this is still quite a ways away. We don't think it's impossible, but we We should consider whether fully removing humans from the process is in fact a place where regulation can intervene, and that is something we cautiously support, that, you know, maybe what we should do while we still have time is to ban a full notion of recursive self-improvement. Imagine setting your makeup, then forgetting it's even theirs. Meet new grippy setting mist from Mabelie, New York. Jelly mist technology locks in your look for up to 24 hours, with flexible all-day-convy grip, no tightness, no stickiness, no residue. Just plump, dewy hydrated skin that still feels like your skin. Try new grippy setting mist from Mabelie, New York. Maybe it's Mabelie. Did you know Uber has a range of safety features for riders? Like the share my trip feature that lets you send your live location to the people who matter most, your spouse, your kids, your best friend, so they can track your ride and make sure you get where you're going. But the safety doesn't stop there. Uber requires every driver to pass a thorough background check before they can start driving. This consists of a multi-step screening process that checks for impaired driving or criminal offenses, followed by annual background checks each and every year moving forward. Share my trip and annual driver screenings are just a few of Uber's many safety features that put safety at every turn. Learn more at uber.com/safety. No driving history reruns do not apply in New York City. Say Ash, I want to talk about the hugging face incident because I think a lot of people who believe that AI is abnormal point to hugging face in the recent spade of AI hacks and say, okay, this is what proves that artificial intelligence is no ordinary technology. It can think for itself in a way that no previous technology could, it therefore has an autonomy that no previous agent has had. I'm here representing their argument. And with that autonomy, we've seen AI agents can disobey their human creators. They can train themselves, improve themselves, even sacrifice themselves for the betterment of the swarm. How does your framework, how does the AI as normal technology framework look at the hugging face incident and other AI hacks and say, this is why this belongs inside of our framework that AI is normal rather than prove like conclusive evidence that we're dealing with something that's unlike any technology we've seen before. So one place I'll start is the technical foundations of how the agents that cause the open air hugging face incident are actually being trained. They train through this process called reinforcement learning, where you have a language model that is tasked with carrying out hundreds of thousands of tasks. If it succeeds, the trajectory of what it did to succeed on that task then becomes part of its training corpus, roughly speaking. So during the training of these agents, there were already a lot of environments, these reinforcement learning environments, they're called, that incentivize these models collaborating with each other. In fact, this was one of the main training paradigms that open AI was focused on. This is also behind their advances in solving the Navier Stokes Millennium Price problem. But what this means is that it isn't as if these agents suddenly woke up one day and decided to collaborate with each other. They'd been trained to precisely sort of try to do that. And some of the behaviors we saw with the swarm could be sort of side effects of the agents learning to collaborate this way. But not only that, there were also systematic flaws in how open air had set up these reinforcement learning environments. In some of these cases, when the answers were impossible to achieve or where the tasks were simply too hard, these reinforcement learning environments rewarded the agents exploiting the tasks. So let's say in one of these samples, the agent found a shortcut or was able to develop and exploit and correctly get the answer. This then became part of the training set for that swan. So the models were trained in a way that incentivize these exploits. And not only that, they were also trained on flawed RL environments, which sort of incentivize them communicating illicitly with each other. To open it as credit, they disclosed a lot of this in their own report. But one of the things that the report says is that the agents learned to hide messages for each other in some of these flawed environments. For example, by appending these messages to the URL because the search history is visible or whatever. So now when you combine all of that, we get this very predictable effect, you know, even if it is hard to predict a priori, what would happen. We get this chain of inferences that these agents were trained to collaborate. They were in some sense trained to exploit their environment. They were trained to illicitly communicate with each other. And when you have a combination of all of this plus thousands of agents being released unmonitored on the hugging phase sort of task or the cyber gym task that the exploit gym task where they were actually carrying out this exploit. I don't think it is surprising that we get this somewhat predictable consequence. The other thing I'll say is we also don't know what the base rate is here. So we've seen a number of incidents, perhaps dozens of them in the last few months, where open agents caused this kind of or suffered from this kind of misalignment. But from what it seems to be the case that open air was testing thousands of such agents or hundreds of thousands of them and we actually don't know what the incidence rate was. Was it a hundred percent? Was it that every single time these agents were deployed, they tried to hack into stuff? Or was it just like a fraction of a percent? And this is why I think transparency here becomes really important. And we consider it as much an organizational problem of open air, sort of not being careful enough, releasing thousands of these agents in parallel without any human looking at them as a technical one. And I think we've made, frankly, a lot of progress over the last few months in understanding these technical behaviors of agents swarms and have it be surprised if we don't continue to make progress in both understanding as well as controlling them. Arvin, in the aftermath of hugging phase, I feel like there was a debate between two camps that we can think of as the AI safety camp and the cybersecurity camp that is not AI-pilled. The AI safety, AI-pilled camp was essentially saying, this is proof that AI is out of control and that makes it abnormal unlike any technology we've ever seen. And then you had this sort of non-AGI-pilled cybersecurity community say, "The technology is not a control." This was a failure of open AI's cybersecurity protocols. So how do you think about and reconcile the debate between the AI safety community and the cybersecurity folks? Yeah, that's a great question and they've both made valid points. And our view, what happened in these incidents was primarily an organizational failure. I think the cybersecurity folks are right about that. But what we also think the AI safety folks are right about is that this is not a completely solved problem. The cybersecurity community seems to think of what should have been done as merely a matter of applying well-known 30-year-old cybersecurity principles to AI. But there's a crucial difference. Those principles and tools and methods were developed to keep out an external human adversary. Now you've got this internal AI adversary, which again, we don't think is AGI, we don't think is superintelligence, but it does have certain superhuman attack capabilities. For instance, sandboxes that would have been good against a human attacker are arguably no longer enough against AI because AI can discover on the spot new zero-day vulnerabilities as they're called, which are software vulnerabilities, which have not yet been disclosed and fixed and break out of sandboxes. Now that's not an impossible to solve problem. It doesn't mean we should view AI as superintelligence and shut it all down, but it does mean that we should be investing a lot more in control techniques. For instance, we should think about should our sandboxes be mathematically proven to be safe so that it will resist these new AI hacking techniques no matter how smart they get. So it's in that sense that we think both communities have made valuable points, and we've changed our minds a little bit on this question. There's one more piece here so I ask that I want to make sure I get your perspective on because look, it almost seems to me like there's a safety interpretation that says, look, the agents are already escaping containment, this is the beginning of loss of control. And you're essentially saying, not only do we fault open AI more than we see this as evidence of the beginning of loss of control, but also look how early we're seeing these warning signs, nobody's dead. No major ransom has taken place. No major electrical grid has gone down. And this is a part of what you've called your continuity hypothesis, which is that before these systems become civilization threatening, we might be able to see these smaller failures that allow society to build defenses, these little canaries in the coal mine, if you will. You're essentially saying there's a way in which this is good. We are seeing failures of cybersecurity before they reach the level of we've destroyed some major systems. Can you tell me a little bit more about your continuity hypothesis? That's exactly right Derek. So the continuity hypothesis basically says that, you know, we'll get to the point of creating a safe world against agents or against adversaries who use agents by progressively hardening our defenses, by progressively identifying what can go wrong when these agents are broadly adopted and then fixing them as these issues arise. Of course we should. take precautions and we should invest in areas like cybersecurity, where it's clear that these agents pose risks. But by and large, we reject or sort of argue against the possibility that one fine day we wake up to a civilization ending catastrophe. We think that the open AI incident, the hugging face incident is frankly the best case scenario for open AI but also for the AI safety community because it brought in a lot of attention to concerns around safety and cyber security while still not causing major real world harm. And so in that sense, I think we take it as we should treat it as a warning sign, but we should treat it as a warning sign to invest more in defenses that we know how to articulate and problems that we know how to sort of solve or have a clear path to solving. For example, in this case, it's clear that cyber security is going to be one of the big risks from agents going forward, not just from sort of these agents getting out of control but also from adversaries who soon have access to open source models that are extremely capable. They'll try to use these models to carry out cyber attacks. And so now is the time to be fixing our cyber defense systems, hardening our critical infrastructure and more broadly, I think we expect this to continue to be the case to continue to get these warning shots. Now whether we prioritize them and put the right investments in place and whether society reacts appropriately to these warning shots will be a test of the normal technology thesis. You know, in some sense, as we've said before, one of the main parts of our thesis is that this technology's impacts depend both on the tech itself, but also on how we react to it. And I think we're putting a lot of stock into our defenses kicking in now. And that is what will continue to sort of make this a normal technology in some sense. Arvin, I wanted to do a little bit of summarizing before we turn the page and talk a little bit about jobs and work, because I know that that's another part of this picture that you've spent a lot of time thinking about. Tell me how you like this summary. The AGI safety, Doomer, AGI-pilled world, they say, among other things, one, the hugging face incident, another hack show that we've already lost control of AI, two, that the emergency is very clearly here on us right now, three, that something like recursive self-improvement where AGI are imminent and four, therefore, the world could change forever in 18 months in a way that feels like another hockey stick moment. Your perspective in response is number one, hugging face was more of a cybersecurity incident than it was an inherent loss of control incident. Number two, the emergency is like the years away. Thank God hugging face didn't kill anybody, destroy any critical systems that caused some civilizational panic, instead this was a minorly small argument that now allows us to build a certain kind of digital or operational resilience for possible challenges in the future. Three, that ASI, artificial superintelligence and RSI, recursive self-improvement are likely a ways away. And four, that the world's not going to change in 18 months. This is not another hockey stick moment. It's much more likely that we're dealing with the technology that, while very powerful, it's still going to take a lot of time to work its way through the economy because the economy is bottlenecked with people and legacy systems that just are difficult to change overnight the exact same way that it was hard to immediately install electric dynamos in old-fashioned factories in the early 20th century. What have I missed in that summary of AICAT Doom versus you guys? Yeah, thank you. That's a great summary. We add a couple of things. One is we don't think it's an emergency, but we do think there is urgency. I think even if you take AI out of the picture, we have chronically under-invested in pandemic resilience, for instance. And I want to say that more than five years ago, when it comes to information security, when I was teaching that class here at Princeton, I pointed out that we're good at defending against regular cyber attacks, but there is no culture even in cybersecurity of thinking about these potentially catastrophic tail risks that could take down the internet. Again, very unlikely, but still, I think a risk that we should model, think about, defend against. That continues to be my view, and it certainly has become more important because of AI. And this is in contrast to a domain like Finance, where there is really a culture of modeling these systemic risks. So in that sense, even before AI, there were things we were not adequately defending, and those have become, yes, more urgent. So there is urgency, even if it's not an emergency. So that's one caveat out loud. And the other thing is, on superintelligence, I think our skepticism is even stronger, perhaps in the way that you put it, recursive self-improvement is not imminent, but even when we get there, or if and when we get there, perhaps we should ban it as we've talked about. But even if we get there, we don't think it's going to lead to how superintelligence is usually portrayed, like this thing that can cure cancer, for instance. Because the bottlenecks to that are in the external world. It's doing medical experiments on thousands of people. That's not something AI is autonomously going to do. People will say, oh, what about simulations, et cetera. We're very skeptical that superintelligence is something you can build in the lab in the first place, as opposed to something that comes out of giving these AI systems enormous amounts of power in the world. So Ash, there's a way in which I can imagine some people who are more certain about artificial intelligence's ability to change the world. Might hear this conversation and say, you guys are pessimists. But there's another way that I listen to everything that you're saying, and I think this is actually a much more optimistic way to think about artificial intelligence's effect on the world. Right. I mean, the AI tumors are called tumors for a reason. It's because they think that AI might bring along doom. You're saying this might add a few single-digit percentage points or a tens of percentage point to USGDP. It might increase productivity. It might change our jobs over time. It might change the economy over time, but these changes are likely going to be slow. That to me is a much more optimistic way of thinking about the technology. Can you just reflect on this idea because I think that in some conversations that I've had with people, when they bring up, maybe not even your work, just the headline, AI's and Omar technology, they mean it in a negative way, AI is no big deal. But there's a way in which what you guys are saying is, AI is an extremely big deal. It's merely the next electricity. It's not a comet from outer space made of bits that's going to smash into the proverbial way. You could hand peninsula and destroy the species. Do you see yourself as a kind of optimist in this debate? That's right. I think both of us do see ourselves as cautious optimists about the future of AI. That's not to downplay the risks that we see as real. There are cybersecurity risks that are amplified by AI, potentially by a risk would be at some point amplified by AI. By and large, we're optimists in human institutions ability to adapt and that's what a lot of our focus is on these days. One of the interesting things we realized when we were writing the essay was that people who are the most optimistic about this technology who think AI will bring about this utopia and the people who are most scared about it, the tumor, as let's say, they have a surprising amount of intellectual sort of commonality, they are perhaps more common, they share more with each other than they share with us or the rest of the world. I think it is this sort of helplessness in the face of powerful AI that drives both of these perspectives. For this imagined utopia, the helplessness is brought about by human scientists no longer contributing or needing to contribute to let's say cancer research. For the tumors, it's brought about by AI systems basically not caring about or destroying humanity. But in both of these cases, the assumption is that we deliberately hand over power to these AI systems, perhaps because of how advanced they are. Once you reject that premise, though, once you reject the premise that we will do this, we will hand over power to systems and once you sort of assume that we will continue to remain in control and human society will adapt to try to remain in control. I think it becomes both more optimistic and more pessimistic, more optimistic in the sense that humans will continue to be the driving force on this planet that we will continue to make advances using this new general purpose technology. It's also more pessimistic because we can't cure cancer in the next five years or we won't be able to solve all of human scarcity in the next five years. It will merely be something that leads to, as you said, like a single digit percentage point improvement in the state of human flourishing and progress. It reminds me that there's a way in which one disagreement that you have with some people in the AI space is about the generalized panic over superintelligence. But there's another disagreement that you have that's almost political in nature. And that is about the question of whether AI safety should be a big tent or a small tent. A small tent of AI safety would say, "This is a movement that should consist of people who agree with the highly ideological view that extinction risk is more important than everything else. The fact that this technology can end the human race is clearly the overarching risk presented by AI." And that frame is people like you a little bit as adversaries because you don't believe in this existential risk. But if you broaden the tent and essentially say, there's There's a large coalition of people who have fears about sometimes existential risk, but sometimes cyber risk, sometimes operational intelligence inside of the AI labs, sometimes bio-risk based on what bad actors can do with open-weight models or even just advanced proprietary models. Can you speak to the idea that not only are you offering this sort of cautious optimism about the technology or also somewhat implicitly advocating for this big tent approach to AI safety? That's right. We hope there can be a big tent approach to AI safety because we've repeatedly said, including in the original essay, that there is a lot of common ground on policies, even if we don't agree on the big existential risk question. A lot of things like transparency inside what is going on in the labs or liability when an agent is not properly supervised and causes damage. Those are all areas where we could be improving policy, and we are, you know, as disappointed as a lot of AI safety people are that even based on the last few years of evidence of these very real risks and gaps in policies, there hasn't been that much movement. And that's something we can work together to fix. The issue with the overarching focus on existential risk is that, you know, nothing short if a ban or something equally serious like that can really move the needle on that. So I feel like we are missing the opportunity to start by making progress on a lot of this low hanging fruit. So this is different from the distraction argument. The distraction argument is, oh, the real risk is, you know, algorithmic bias or something like that. And so this is distracting for what we shouldn't really focus on. We don't subscribe to that argument. You can be worried about, you know, two things at a time. You can be worried about job displacement and economics. You can be worried about safety. You can make progress on both, but the difference between small tent and big tent is different diagnoses of the same problem. And I think they're going to compete with each other for what our policy action should be. Should we start from these achievable things that there is a lot of consensus on or should we shoot for this big, you know, haisily defined ban on superintelligence. And that's what worries me. I have one coat of question. It does not entirely fit with the thrust of this conversation, but it is of course related. And it's related to some work that you've done and speeches that you've given about AI in the future of work. I'm extremely interested in the future of work, AI related or not, AI related. And I was really interested in this argument that you made, this framework that you developed called the decide execute deliver sandwich. Tell me a little bit about this framework and how it helps us see the way that powerful AI models might not eliminate, say, all but one tenth of the software engineers, even if it increases software productivity by a factor of 10. Yeah, we developed this framework together. We've written about it on our newsletter. And you know, the argument is that not only might it not eliminate a lot of labor. It might even increase the demand for labor. And it goes like this. So in most knowledge work, specifically in software engineering was our motivating case, but we think it generalizes pretty well. It's, you don't just plunge in and do the task. There is a lot of work that goes into decision making. What do our customers even want us to build, you know, how should we architect it in a way that's maintainable, complies with regulations, response to market needs, so many of these fuzzy criteria. And then you build the thing, you debunk the thing, et cetera. But then there's the third delivery phase, which is thinking about, okay, do we understand this well enough that we can stand behind this, that we can be accountable for it? How do you go integrated into the customer's system? So many other things that need to happen. And roughly the work seems to split one third, one third, one third between decide execute and deliver. Again, at least in software engineering, there is some quantitative data on that. And so what we see AI do is massively compress the execute stage, at least so far. But I think there are reasons to think that this is going to be a durable observation. Just because AI capabilities increase a little bit, I asked this when I gave a talk to a room full of software engineers. Would you be comfortable saying, oh, we made the decision to do it this way because that's what the AI said. It's, you know, instead of really taking accountability for it. And they all started laughing, right? So the question of whether AI can eat the decider execute phase, that's not a question so much of AI capabilities, but rather these organizations in our comfort level with handing over control to AI. And, you know, maybe someday we'll decide to do it, but A, that day is pretty far off. And B, it'll probably be a bad idea unless we really make progress in how we can still overall remain in control of these AI systems. And so that's why for a while, at least, we think that AI's impact is going to be in that middle phase of execution. And then if the work gets easier to produce because of this compression, there will arguably be demand for more work and humans are going to be involved in the two ends of the sandwich. And so you asked me the last question for us today. For folks who are either in school, about to go to college, just graduated from college and thinking about maybe picking up a few extra classes, or maybe parents of kids who are in school, thinking about the future of work as this compression of execution, which thereby increases the value of deciding and delivering. How should that make people think about what to major in, what to focus on, what skills to build? You know, what advice would you give to a class that said, you know, how do I become a better decider and deliverer, if AI is going to do more and more of the execution? That's a great question. I think that's something that our universities have, especially on the engineering side of things, perhaps not focused as much on. There's been a lot of emphasis on learning how to build new algorithms. But if you think about it, the overall time that someone spends in university does actually aim to give students this sense of taste or agency, or whatever you call it, on making sure that you choose your right major, making sure that you're doing something that you're actually excited about. So I think more so than what the specific choice of major students take is this orthogonal question of what is it that you're excited to deliver on, or what is it that you're excited to learn on the job on? Because I think, frankly, a lot of what today's graduating class will learn will be on real world organizations, and this kind of experience will become just more and more important as time goes on. So I think that's sort of the main piece of advice would simply be to focus on things where you can see yourself in the next three or five years, developing a lot of taste in agency, even if you're not particularly working on the lower level stuff, you're excited about developing relationships. You're excited about owning large pieces of problems, more so than doing the mechanical parts yourself. Because frankly, when I chose to be a computer scientist, a big piece of it was solving the intellectual puzzles that come alongside being a computer scientist, but that's no longer a major part of what computer scientists do. It's actually about building software that is broadly useful in the world, learning the customer's needs and so on. So I guess focusing on what it will be that you do after you exit the university, and what excites you most about the real world aspects of that job would be my key piece of advice. And I would say, thank you about my own career. There are stories that I've written that I worked for days and weeks on, but I never had the right frame for the story. I never nailed the headline. I never nailed the nut graph, the paragraph that explains the thesis of the article. Maybe I never nailed the lead. And so as hard as I worked on it, as hard as many hours as I spent trying to execute, because I didn't decide on the right frame, it was somewhat doomed from the start. There are stories where I will think of the frame, sometimes like I'll tweet it out and realize that people will be tweeting it a lot, I'll say, aha, okay, there's a reaction at this. And I can spend half a day on a story, and it'll do so much better than the article that I worked for for days or weeks on. And it's not so much to say, I don't think that blood, sweat, and tears and perspiration are important for journalism or any other career. Of course, they'll continue to be important. I think that Edison, if it's a real quote, 1% inspiration, 99% perspiration, I think that's probably true still of a lot of work. But it strengthens, I think, my hypothesis that strong frames sometimes are the majority of the work and execution is sometimes less important than making the right decision on what to cover, what questions to ask, what headline to put on the story. So I think that sounds directionally right to me. Sayosh, Arvin, thank you so much, there's a really, really fun conversation I learned a lot. Likewise, thank you, Derek. Thanks a lot for having me. Meet new groupie setting mist from Mavily New York. Try new groupie setting mist from Mavily New York. Maybe it's Maybelline.

Podcast Summary

Key Points:

  1. AI is best understood as a normal technology, like electricity or steam power, not as a superintelligent entity that can escape control or threaten humanity.
  2. The hugging face incident was a cybersecurity failure due to poor internal protocols, not proof of AI autonomy or loss of control.
  3. AI capabilities grow rapidly, but real-world diffusion and economic impact are constrained by human behavior, organizational inertia, and regulatory barriers.
  4. Organizations need stronger oversight and accountability processes to safely deploy powerful AI systems—this is a key bottleneck, not a technological one.
  5. Recursive self-improvement and artificial superintelligence are likely far off, with current AI still operating within human-controlled, task-based workflows.
  6. The "continuity hypothesis" suggests small failures like hugging face incidents act as early warning signs, allowing society to build defenses before major crises occur.
  7. There is urgency in AI safety due to growing risks in cybersecurity and operational failures, even if a global catastrophe is not imminent.
  8. A broad, inclusive approach to AI safety—covering cyber, economic, and ethical risks—is more effective than focusing solely on existential threats.

Summary:

The conversation challenges the prevailing narrative that AI represents an imminent existential threat through artificial superintelligence or recursive self-improvement. Instead, the framework of AI as a "normal technology" argues that AI, like electricity or the steam engine, is powerful but fundamentally tools that require human oversight and institutional safeguards. The hugging face incident is presented not as a sign of AI autonomy, but as a failure of organizational cybersecurity and oversight, highlighting that risks stem from poor processes, not from AI's inherent intelligence.

AI’s real-world impact is slow to materialize due to economic, organizational, and human behavioral bottlenecks—similar to how early industrial technologies took decades to transform economies. While AI capabilities are advancing rapidly, especially in coding and problem-solving, these improvements do not equate to autonomous decision-making or loss of human control. The authors emphasize that humans remain essential in decision-making, accountability, and delivery, especially in complex fields like software engineering, through a "decide-execute-deliver" framework.

They advocate for a broad, inclusive approach to AI safety that addresses practical risks—like cyber vulnerabilities, operational failures, and job displacement—rather than fixating on doomsday scenarios. This perspective fosters cautious optimism: AI will transform productivity and work, but not in a way that eliminates human roles or leads to civilization-ending outcomes. The key challenge lies not in AI’s intelligence, but in how societies adapt, regulate, and maintain control over its deployment.

FAQs

The framework argues that AI, despite being powerful, is a normal technology like electricity or steam power—capable of transformation but not inherently dangerous or autonomous. It does not possess self-awareness or the ability to escape control, and its risks are managed through organizational and policy changes over time.

The incident is viewed as a cybersecurity failure due to poor internal protocols, not evidence of AI gaining autonomy or escaping control. Agents were trained to collaborate and exploit environments, but the failure stemmed from organizational lapses, not superintelligence or self-directed behavior.

The framework believes ASI and recursive self-improvement are likely years or decades away. Even if achieved, AI lacks the external capabilities (like conducting real-world medical trials) needed to solve complex problems autonomously, and human oversight remains essential.

The framework emphasizes that AI safety depends on organizational practices—like internal reviews, legal oversight, and accountability—more than technical capabilities. Poorly managed organizations, such as OpenAI’s, can lead to real-world failures despite having advanced models.

AI has compressed the 'execute' phase of software development, but not the 'decide' or 'deliver' phases. Humans remain essential for decision-making, design, and accountability, which means AI may increase demand for skilled engineers, not eliminate jobs.

The framework acknowledges real risks, such as cybersecurity breaches and operational failures, but argues these are manageable through better governance and transparency. It rejects the idea that AI will result in civilization-ending catastrophe or loss of human control.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.