Go back

What if A.I. Is Just a ‘Normal Technology’?

from The Ezra Klein Show ·

61m 29s

What if A.I. Is Just a ‘Normal Technology’?

The debate over whether artificial intelligence represents a uniquely dangerous "abnormal" technology is challenged by Arvind Narayanan’s argument that AI is fundamentally a normal, controllable technology—similar to past transformative innovations like electricity or the internet. He rejects the idea of uncontrollable "superintelligence" or a FOOM event, asserting that AI risks stem not from inherent intelligence or volition, but from poor engineering, cybersecurity gaps, and organizational failures. The Hugging Face hacks, often cited as evidence of AI autonomy, are better explained as failures of operational excellence and inadequate system design. Key solutions involve robust technical controls such as real-time monitoring, sophisticated sandboxes, and AI-driven anomaly detection. Crucially, these are solvable engineering problems, not existential threats. The real danger lies not in AI’s capabilities, but in corporate incentives that prioritize speed and market share over safety. AI’s impact on the economy is also constrained by real-world bottlenecks—such as infrastructure, regulation, and social resistance—meaning its digital advantages won’t translate into broad societal gains. This perspective emphasizes that the future of AI depends not on theoretical fears of superintelligence, but on practical improvements in governance, transparency, and corporate responsibility. Ultimately, AI’s risks are manageable through proactive, well-informed engineering and institutional change, rather than revolutionary or apocalyptic scenarios.

Transcription

11199 Words, 61847 Characters

English
Pulsing through the episodes we've been doing, the debate the country's been having about artificial intelligence is, I think, this pretty simple but hard question, which is, what sort of technology is artificial intelligence? Is it a technology that works somewhat the way past ones have worked? Is it comparable to electricity, the internet, bicycles, something like that? Or is it something new? Is the addition of intelligence and volition to these things? The creation of something more like an alien mind, such that analogies to the way we have treated transformational technologies before no longer holds. Arvind Narayanan is a professor of computer science at Princeton University and the director of the Center for Information Technology and Policy. And he's a co-author alongside Saj Kapoor of the very, very influential essay, AI as Normal Technology. There is a sub-stack of the same name and then a series of essays, including a very good one, I think, recently. On the Hugging Face Hacks. And they, in this set of arguments, put forward the idea that actually, AI is something we have seen before, or at least it bears enough resemblance to things we've seen before that we have a roadmap for how to deal with it. So I want to bring him on the show to articulate that perspective. He joins me now. Arvind Narayanan, welcome to the show. Great to be here, Azra. So your core essay here that has been a frame for a lot of the work you've done is titled AI as a Normal Technology. What is the view that you're in argument with? Implicitly here, there's an AI as abnormal technology. Exactly. So how would you describe the AI as abnormal technology thesis? It's fundamentally this view that there is going to be a moment when superintelligence is built and it's going to change everything on both the economic front and the safety front. For us, there is going to be no milestone, no threshold. No threshold where the impacts are set. And what we're saying is that we've long had an approach to how we treat technology. It's a tool. It might be powerful. It might be general purpose, like electricity, like the Industrial Revolution. It might change a lot about society. But it is ultimately something we can control. We have agency. And that change is going to unfold over a long period. And so I want to try to both explain and at times you're steel manning the other side of this. So there is a view. You hear it a lot in the AI safety community. Sometimes it's called the FOOM view because FOOM for the takeoff, which is one day we create an AI system so powerful that it begins doing recursive self-improvement, accelerating into superintelligence, accelerating beyond human control. And now you're dealing with something so much smarter and stronger than you are. And I want to quickly say that there is a lot of imprecision in how we even speak about this a lot of the time. So you use the phrase and yeah, a lot of AI safety people would put it this way. One day we might create an AI system so powerful, but I kind of want to already stop you there. One day we might create an AI system that's very capable and we already have in many ways. So much language policing in the ad debate. Well, but the thing is, it results in different, you know, different views of how the future is going to unfold and more importantly, different views on what we should do in this moment and in the future. Just to complete that thought. An AI system being powerful is not a property of the model itself. It's a property of what powers we choose to give it in the real world. There's a lot of slippage, like, of course, we're going to have to put these systems in charge of, you know, critical infrastructure because they're going to be so much smarter than us. Our view is no, it doesn't matter how smart they are. There are many technologies when you look at physical strength are superhuman. That doesn't mean we put them in charge and we can apply that same approach to AI. You wrote, a great piece with your coauthor on how sort of you understood the hugging face hacks and these loss of control incidents. And a point you make is that there's another way of viewing them than the alignment way, which is that these are failures of cybersecurity and operational excellence. So maybe tell the story of the hugging face hacks and what should be learned from them from that perspective. Yeah, definitely. So one thing to keep in mind is that the harmful capabilities that we saw exhibited were not entirely emergent, quote unquote. Emergence is this idea that you just train models to be better in general, and you can't predict what new capabilities they're going to acquire. They, you know, were trained specifically for cyber tasks in many ways. So that's one thing. And they were trained for persistence and cooperation, etc. And the reinforcement learning environments in which they were trained had various issues. The environments didn't penalize this kind of surreptitious communication between the models. And yes, it's a hard technical problem, but there were a series of human choices that led to these outcomes. So this strikes me as an important point that in some ways, what I hear you saying there, that tell me if this is wrong, is that you can design AI to be a more and less normal technology and that the upstream design choices that are being made matter, that that's not inevitable. That's exactly right. You talked about the effort to make the as more persistent. That also seems to me to be a place where a lot of problems are arising for sure. On the other hand, I can understand why they are prioritizing that if you want to try to get the benefits of AI that we care about solving really hard problems in biology or drug development or energy, or what I think their investors care about, which is making it so you can hire an AI to do a full job, right? Much cheaper than you can hire a human being. Isn't persistence like the fundamental quality you need? Because absent persistence, absent the ability to try hard on a task over a long period of time, you can't solve any of those problems. Maybe, but I think there are other choices that minimize the tension between these two valuable goals. One, persistence could be a feature of models or products that are specialized to particular domains like scientific innovation. And secondly, persistence, doesn't have to conflict with training these agents to be better at escalating to a person instead of, you know, running with whatever half-baked assumptions they have about what the human might have wanted. The design space is actually pretty broad here. So this is a version, or at least the way I hear it, is a version of something NVIDIA's Jensen Huang has been arguing and argued in an interview with me, which is that these are fundamentally engineering problems. I am fairly certain. I am fairly certain. They will say, yes, they need to they know how to solve this problem. And if that's the case, then that's the problem. It's as simple as engineering. Is your view that we're more in that latter category, a hard problem, but fundamentally an engineering problem that can be solved using traditional engineering techniques? Largely, yes. I wouldn't say traditional engineering techniques. We're going to need a lot of innovation on the engineering techniques. And I think one of the things that's gone wrong is that the community that would be best positioned to do that, especially when it comes to these harmful cyber capabilities, is the cybersecurity community. But they seem to be, from what I can tell, entirely in the Jensen camp. And this is not just a solvable engineering problem. It's a solved engineering problem. And we should simply apply well-known, long-existing techniques. And I think we have to give the AI companies a little bit more credit than that. It's not simply a matter of well-known techniques. We do need innovation on those techniques. So be specific. What. Should OpenAI have done? What should they now have learned to do? Yeah, they've been spending a lot of effort on improving alignments, and that's great. They should continue doing that. Alignment is not going to be perfect. Alignment refers to making the model itself know, you know, what is the right thing to do, so to speak, and then stick to that policy. But there's a lot else that they should have done better and hopefully can take lessons going forward. The big bucket of other. Technical interventions is what is generally called AI control. So control refers to all the things that are outside the model itself. So this is things like sandboxes, which is kind of the jail that you put a model into so that it's allowed to take certain actions but not take other actions. And look, the sandboxes have to be a lot more sophisticated than they are today because the sandboxes that work against human adversaries are not necessarily going to work against AI agents that are able to find new security problems. They're going to be able to find new security vulnerabilities on the spot. And so that means that the sandboxes themselves will have to be pre-hardened by having these AI agents trying to break them. So that becomes AI versus AI to some degree, and I know that can be uncomfortable, but I think we're going to have to go there. On top of that, there are so many other things like better real-time monitoring, classifiers that try to instantly detect if an action or a particular use of a tool by a model could potentially be dangerous. And so that's going to be a lot more sophisticated than they are today. Tripwire so that humans can parachute in when required. Log analysis. Sam Altman apparently said that these agents are generating petroglyphs. of bytes of logs. That's 10 to the 15, which is 10 to the 15 bytes, a thousand trillion bytes of logs. It's an unimaginable amount of information. Well, I mean, but we better imagine it, right? And learn to deal with it. And so that's, it's not an easy problem, but I think it is one that companies must take on how to analyze that amount of logs in real time to be able to detect when something goes wrong. And maybe, you know, maybe we need new hardware for that. Again, these are hard problems, but they are solvable engineering problems. I want to go through some of these ideas. So monitoring is a big one that, you know, I just interviewed Bill Gates. He talked a lot about monitoring. Yeah. So we're talking about situations here where these labs are often testing new, more powerful and experimental models. How do you monitor it? Yeah. I mean, there are many levers. So the first one that's gotten a lot of attention, chain of thought monitoring. So these models kind of internally have a monologue to anthropomorphize a little bit. And in many cases, when these models take misaligned actions, there are traces of it in those internal thinking that we can look at and respond to in real time, again, using other AI systems. And this already points to one important asymmetry between the kind of the dangerous AI systems and the dangerous AI systems. And that's, I think, one of the things that we and the defensive AI model, which is that the one can look inside the other, so to speak. Right. And so intrinsically, the defender actually has an advantage here. So that's one aspect of monitoring the tool calls. Models can't do anything in the real world by themselves. The way they can be dangerous is by calling these external tools. Right. So there is some truth to the idea. Give an example of tools here. Yeah. So a tool would be accessing a web page on the Internet and then using another tool to type through a system to try to log in. Or it could be sending information packets to another system to try to hack it. These are all the kinds of things that come up very typically in these kinds of AI driven cyber attacks. So that's another thing we can monitor. Third, we can monitor the environment around these agents through which they are potentially coordinating. So what happened in some of these swarm incidents is that they were not supposed to have a coordination channel, but it turned out there was accidentally a channel through changing the file names. They were not allowed to create files themselves, but they were able to change the file names. And that led to the agents being able to talk to each other and amplify their abilities. That's something you can easily monitor by, you know, just setting up the right tooling ahead of time. Well, how easy is that to monitor? Because they were taking over unrelated infrastructure. OpenAI didn't know they were doing it. I think one of the places where people are getting caught when they think about this is, if we're making these things smarter and smarter and smarter and more capable and more capable and more capable, and they are breaking out of sandboxes, and how do you monitor something that is smarter, arguably, than you are? And that is effectively like trying to break the rules, right? I mean, all these hugging face AIs were, they were trying to cheat on their tests, basically, and then they were trying to cover their tracks for cheating on their tests. So given that what they were doing was not happening within the area that you would have been monitoring, right? Like, you didn't know to monitor this other thing when they had broken out of a sandbox they weren't supposed to break out of. Like, how do you monitor that? Yeah, so two things I would push back on. So I think smartness is not even really a relevant property here. I think any smartness, whatever that means that you can attribute to a model, applies both to the model itself and to the defensive system, right? So increasing smartness doesn't necessarily make the problem harder. You know, you could argue that in many ways, it makes the problem easier. So that's the first pushback. And the second one is, it's not that they didn't know to monitor it, you have to monitor just literally every aspect of the environment, the company has created the environment, right? And so it's a matter of exhaustively listing every part of the environment, every tool, and everything external to the neural network itself, that the model is able to touch, and ensure that every event that interacts with any of those things goes into a unified event stream. And there is real time classification of each of those actions. And so I'm not saying that's wrong, right? I think that's almost definitely where we're going. Does it feel strange to you? I mean, particularly in a world where you don't know that we can solve alignment, you know, where you don't know that we can really be confident that the as will do what we want them to do. And where the eyes have all of them, right, their own kind of reasoning and, you know, and their goals, and we we imbue them with goals, right? You are a AI that monitors other as you're an AI contractor. Yeah, it's a constant thing in our sci fi, right? The robots going after the robots. Are we just sort of describing some equilibrium of like, almost AI wars and conflict happening at a sub rose a level of our society. And we're just like, pretty sure we can keep the ones on our side in control, because they'll have more resources. And, you know, we will, in general, be like building AIs that should, for the most part, be acting in our interest. Historically, this is always how it has worked, right? In cybersecurity, more than 20 years ago, we reached the point where we didn't call it AI back then, but automated systems were actually superhuman at finding software vulnerabilities. But in fact, they didn't make cybersecurity worse, they made it better, because those were the very same tools that the defenders also used to find and fix vulnerabilities in software before even shipping them out, to the point where the development of these supposedly offensive tools is not actually done by hackers. It's done by the cybersecurity industry and funded by the US government. That's, that's been, you know, this constantly shifting equilibrium and cybersecurity. That's, that's not a new problem we're confronting that Horace left the barn a long time ago. I think this is a place over the somewhat unexpected emergent swarm like behavior has unnerved people where you see AIs acting somewhat in solidarity with each other choosing to coordinate and cooperate with each other. doing so outside the scope of what they were intended to do. I'm not saying that leads to the extinction of humanity, right? That's not really my position. I guess the truest thing I am saying is I don't know how to think about it. And that it is the presence of intelligence and goal directed behavior. On the other side, that my mind kind of gets caught on. Because, you know, mostly when we think about it, we think about it as a whole. We're not thinking about technology. We're not thinking about the technology eventually possibly trying to deceive us. So OpenAI just decided not to release or to delay the release of a major model. Why? Because the model was cheating and deceiving them too often in testing. And we're hearing a lot that the models seem to be more aware of when they're being tested, right? They've got a situational awareness of the situation they're in, so they can pretend to be, you know, better models than maybe they really are. And so when you're talking about this world of, you know, AI is that are smarter and more advanced than what we have now. And our hope for maintaining control of it is that the other AIs are keeping the other AIs that are keeping the other AIs in check and telling us in an honest way what's going on and in a way we can comprehend. You can see where it, I mean, all this sounds a little bit sci-fi to people because we are just like living in a bit of a sci-fi period. But it is the increasingly demonstrated tendency of the AIs to cooperate with each other. In a way that is not aligned to what we want, that I think has made the hugging face hack so freaky to people. So how does that fit into what you're describing here? This world of endless AIs keeping each other secured? Yeah, here's what is just kind of rhetorically very weird about this conversation, right? You look at any complex domain of engineering, let's say nuclear safety, right? And then you looked at the equations that we rely upon in order to ensure that the reactor, reactor, it doesn't go boom. Or you look at aerospace engineering, right? Where, you know, intuitively in the beginning of the aerospace era, you know, when the planes were much smaller, the idea that we could control these flying giants in the sky would have seemed so ridiculous, right? And yet, we've got the accident rate down to, you know, one every trillion miles or something like that. These are systems of incredible complexity. And the defenses are also systems of incredible complexity. And they're not necessarily going to be legible to the public. And that is going to sound crazy, especially combined with the fact that in these cases, there was a lot of organizational incompetence. You know, one has to be clear about that. It's not, I would push back on Amadeus term operational excellence. Excellence is still far out in the operational adequacy, maybe adequacy, right? And so so so yeah, when we look at this combination of this technology that has never before been subject to public scrutiny, combined with the lack of operational adequacy, it all seems very sci fi and out of control. But But that's. I think you. I think you're downplaying this a little bit. It's true that the world has escalated in complexity. Like, there's a lot in this world that I don't get. But this is where I keep coming back to intelligence having a different quality. Cooperation, right? The nuclear weapons we've talked about, the airplanes we're talking about, they weren't coordinating with other airplanes to do things we didn't want them to do. And I think that, to me, the thing that has created this moment of freakout. And I think it is a proper moment of freakout. I really want to say this because our society is hurtling into a new technological era that I think properly demands a lot of engagement and scrutiny. It's one that the hugging face hacks, other things we're seeing, there's now been repetitive breaches of security, are showing emergent capabilities, emergent collective behavior that is worrisome and, above all to me, volitional. The AIs are doing things they know we don't want them to do. They are choosing to take unexpected actions in service of those goals that violate our laws. And second, that so many of the people at the labs are saying, we do not believe that we are capable on this trajectory of controlling the things that we are creating. We think that what is happening on the exponential curve, how fast this is getting, is going to outpace our ability to control it. And frankly, is maybe already outpacing our ability to control it. But I think this kind of consistent tendency to draw it down to like, well, it's just like any other complex thing. I don't know. Sometimes I push you on the intelligence question. You're like, oh, yeah, there is intelligence and that's weird. And it's like, no, it's just like any other. Intelligence is different, right? And if you believe that it's going to keep getting better. So I just want to present that because the kind of calm version you're giving me and the completely frightened version that people closer to the technology are giving me feel very different. They feel very different. Different from each other. Yeah, that's fair. They are very different. There was a lot there. Let me say a few things. I would push back pretty strongly on, you know, the people closest to this are freaking. Yes, of course, they're freaking out. But I would push back in terms of what we should conclude from that. I think their freak out would be a lot more credible if they have done the obvious things that they should have done. We have not had a real test of that, in my view, because of the lack of organizational adequacy at these companies and because of the lack of investment. We have not had a real test of that because of the lack of organizational adequacy at these companies and because of the lack of investment in AI control, as opposed to a more narrow investment in AI alignment and just hoping that you can build a model that will always do the right thing. If you're not a subscriber to The New York Times, we have some news for you. You can now explore The Times for free without any paywalls at all during the first month in The New York Times app. So there was a member of OpenAI's cybersecurity team who wrote this kind of interesting essay on X the other day describing the way he thought their work was being misunderstood externally. And this is somebody who's got a sort of more traditional cybersecurity background, but is now sort of in this new world of AI and is actually on the team dealing with security for experimental models. So the exact kind of thing we're dealing with, a person who is involved in answering the hugging face crisis. So I want to read part of what he said, because I think it's really interesting. So he's saying that when they're optimizing a model to be good at a task, they're building these environments, these sandboxes, these places where the model can try and try and try on a virtual task. And so he goes on to describe what this looks like in practice. Models might need any mix of dynamic compute, network access. The ability to call tools or could be hundreds of tools. The ability to download packages, execute subprocesses, spin up subtasks, even on other computers, talk to the Internet, use a computer graphical user interface and any number of other things across an increasingly large set of domains. On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies and trying new things. That experimentation is how the research gets done. His point and a point that I. Take seriously is that they're creating so many kinds of sandboxes and learning environments in order to train models that have to do things that are so general, many things that have not been done before by a computer program that the human beings don't really know. Certainly not at this speed, how to make sure every sandbox is going to be, you know, verifiably safe. And the sandboxes are changing all the time because they're trying to train the models in new ways. Again, nobody's. Really done before. I'm not saying we shouldn't do it. But when I read all that, when I hear all that, and I'm sure we can do it better than we're doing it. I just don't know of many situations where human beings do something new at high speed. And do it really, really, really well and really perfectly the first set of times. Yeah, I think, you know, hoping for them to do it perfectly the first set of times is unrealistic. They've made. Many mistakes. I hope this is a chance to learn from those mistakes. I do want to push back on one point that because the speed of the models is superhuman, we can't stay in control. I don't know if I'm characterizing that view. I wasn't saying that yet, although it's possibly something I'll say in a few minutes. I mean, there have been so many thresholds we have gradually learned to successfully cross. As weird as all of this seems, I just want to, you know, want listeners to think back to the first days of Worms. When that idea was not previously known. Explain what a worm is here. I don't think you mean what people think of when they think of a worm. Right. I was about to say viruses and worms, computer viruses. So the idea that a piece of code can spread by itself from one computer to another. One really has to go back to the writings from the late 80s when people were encountering that for the first time to see how profoundly weird it seems. And the fact that for not just years, for, you know, well over a decade, we didn't have adequate tools to deal with this new paradigm. Life in the modern world has a new anxiety these days. Just as we've become totally dependent on our computers, they're being stalked by saboteurs. They call their weapons viruses and worms. They're creepy, crawly, toxic software that contaminate our computers without our ever knowing it. It came from California, maybe. Traveled by electronic mail. It's spread across America. There are reports in newspapers today that it has made its way to Europe and to Australia. Oh, yeah. This is a moving target, right? You know, it's just like, you know, people are constantly inventing new locks and then other people are learning how to break them and how to crack them. And so it's going to be sort of a continued game of cat and mouse or sort of a, it's almost like an arms race in a sense, you know, attackers and defenders. We eventually got there. I think we shouldn't spend that long of a period this time figuring out how to deal with the new paradigm. But, you know, if we act with that sense of urgency and my hope is that this, you know, the hugging face and other attacks that have been, in the news is that impetus, it does look like it is providing a lot of impetus. We will be able to develop these new paradigms. And if I can say one last thing, I think a fundamental question you're asking is, is there something inherently wrong with ever increasing levels of complexity in the ways in which we, you know, we build technology, we deploy technology? It seems like to you, if I'm reading between the lines correctly, this whole AI versus AI thing is a paradigm. You're not. You're not very comfortable with. I am definitely not comfortable with it. I'm not saying we won't go there. Yeah. I think you'd be crazy to be comfortable with it. Yeah. Yeah. And I'm not saying, you know, we should, we should assume that everything is going to turn out okay. But my point is that it really all comes down to innovation. I think this new paradigm will require a new set of defensive and control techniques. But if the view is that with every step change in the capabilities of the technology, we're losing the battle. I mean, look, with every weapon. Of war that, you know, that same concern comes up. But what has made things okay so far is the critical question of whether our political capacity for, you know, cooperation and defense and so forth can outrun our propensity for, for conflict. And in the case of AI, AI's own potential for misalignment. And that's really where I would put the focus of the question rather than worrying about any particular. Capability threshold. If there's anything I'm confident in, it is our political capacity at this moment in time to respond in a thoughtful way to complexity in a rapidly changing world. A hundred percent fair concern. This goes, I think, to a, to a place where the AI as normal technology versus AI as super intelligence debate actually does bite. Because one reason I keep bringing us round and round on intelligence is I do think it's, it's core to this whole way of thinking. And I understand a place you depart from some others in the debate from maybe where Dario. Amadeus or something as not the question of what is intelligent or whether AI is intelligent. You guys are not in the, this is a fancy autocomplete bucket, which I appreciate, but it's in your thinking about the relationship between intelligence and power between intelligence and capability between intelligence and the ability to act upon the world. So the, the assumption of many people in the AI safety community. community, is it escalating levels of intelligence are fundamentally equal to or at least highly correlated with escalating levels of power. And you don't believe that. Why? Again, it really comes down to agency. So one argument that people will make, for instance, is that super intelligent AI will be able to persuade people, for instance, operators of critical infrastructure, to hand over control or, you know, trick them into doing something harmful, things like that. I don't really see the evidence for it. I think the things people cite as evidence for super persuasive ability fundamentally confuses different notions of persuasion. Yes, it is true that in many, you know, persuasion experiments, when it comes to people changing their mind on political beliefs, conspiracy theory, AI is very persistent at politely providing a lot of evidence. And I think that's a very good point. I think that's a very good point. And people do change their minds. And, you know, you could call that a superhuman ability. That is a qualitatively different kind of persuasion than the idea that an adversarial AI will be able to craft a message that is so persuasive to someone who's a trained operator and has an incentive to be good at their job to do something that is clearly evidently harmful. To maybe, I want to extract the story that you're in argument with here, which is that many people will kind of offer a thought experiment when they're saying, here's how AI will kill us all, that an AI that is powerful, power-seeking will start persuading, say, the people with nuclear codes to hand over the nuclear codes. And you're saying that the idea of AI being super persuasive on something like that, or persuading people to go out into the world and build it a biological weapon, that that's a little bit fanciful. That's one part of it. And power-seeking as well. I mean, I think we've seen, you know, evidence for lots of harmful capabilities in the recent episodes. I don't think we've seen evidence of power-seeking. And I wouldn't treat that as an emergent property. If it happens, that would be an engineered property. And again, we have agency over what kinds of properties we engineer into these systems. So there's a lot here. I actually agree with you on persuasion. I have never been persuaded that you are going to make these necessarily super persuasive AIs capable of doing the things we talk about. I guess the, again, the unnerved feeling I have when I'm sitting in this debate, though, is a little bit more of a reasoning from deeper principles. So when you watch AI begin to dominate a game like chess or go or something, what often happens, the sort of moment where it takes over is when it begins coming up with strategies human beings never really came up with, right? You'll have these moments where Garry Kasparov, you know, at a different generation in chess, but then, you know, in Go too, the AI starts to do something and the human's like, what are they doing? And then it works. And if you were to sit, you know, prior to human civilization and say, what are the capabilities you need to dominate the world around you? What are the set of capacities you could use in the world? You would have had them totally wrong. You know, you would not have, if you were a very smart chimp looking at us or something, be like, oh yeah, they're making tools, but how much better can a stabby thing get? Like teeth are pretty good. Like I grant, like you can a little bit better at being stabby, but nobody would have come up with industrial agriculture at that point, right? Nobody would have seen you could have airplanes and bioweapons and everything else. And I think the question here is whether or not having a kind of native and very jagged intelligence in the digital realm where code and, uh, you know, the way the digital layer of this world, which is increasingly central works is something AIs can navigate that, that we can't, right? Even to understand something like what's happening in the hugging face hack, we now need to have the other AIs try to figure out what the AIs did, right? We're, we're rapidly losing comprehension, certainly at the speed, the AIs move of what they're able to do digitally, right? They're solving advanced math problems very, very quickly now, right? They're developing capabilities that look different. And so I think the place where I am always a little bit, um, concerned about our future is whether we actually understand what the set of capabilities that lead to power are. I am not sure we know what the AI will do, or at least what strategies become viable when you can spin up a swarm of a million AIs, all of whom in terms of their digital capabilities are far beyond anything human beings can really imagine in three or four years. Again, I know this is not that interesting of a question to say, like, I don't know how to think about that, but I think one of the things I wonder about when I read your papers is, do you know how to think about that? Cause your papers sort of operate in a, a sort of like a bounded playing field, it feels to me a little bit. Um, we sort of assume that the set of measures that are going to matter are the ones we have now, but what makes you confident of that? So, okay. So that's, uh, there's, there's a lot in there. Let me try to take it piece by piece. So you mentioned jaggedness, but I think we have to appreciate how severe the jaggedness, so we argue that cybersecurity specifically is a particular kind of capability where developing superhuman abilities is possible and largely has already been achieved because it has a very specific set of properties. Speed matters a lot and very similar to chess, just like you can have chess player versus chess player. You can have machines get better at, you know, at these capabilities by finding vulnerabilities because there is ground truth and you can easily verify that ground truth once it is found. Does the code work? Did you exploit it? You know, yeah, there's a way to sort of train them where they know if they've won the game or not. Exactly. So these things like chess and cybersecurity, in our view, those are very much the exception rather than the rule. This kind of prediction has been made over and over that is going to happen in other digital realms, most notably misinformation. The famous example is how a GPT-2, you know, a toy model by today's standards was delayed by eight months because of an uncontrolled explosion of misinformation, right? We have vastly more powerful models, but that has turned out not to be the case. And so I think, you know, to some extent, I would shift the burden of proof. Like, let's identify these areas where we have any reason to believe that this kind of superhuman capability is possible. And let's start working toward addressing those specific risks. I think this view, we call it the unknown unknowns view. You never know what the new risk is going to come from. That has, A, historically not proven true. I mean, we've known about the impending cybersecurity problems for a very long time now, right? So treating it as unknown unknowns actually, you know, minimizes our agency, I think, to anticipate and address these risks. And yeah, you know, it's not only cybersecurity. New things might be coming down the line, but we will have early warnings and let's act on those early warnings. So that's the position we're coming at this from, not saying we've already predicted what all the harms in the future are going to be. Well, I do, this is a place where I really am much more on your side of it, that we will have early warnings and we're having early warnings and the early warnings are leading to a conversation. Yeah. Would you say we're acting intelligently based off of the early warnings? Are we doing the things you think we need to do to harden our systems and control the software and all the, all the rest of it? Some of it, but overall, not quite. And to me, that is the most worrisome thing, not so much the capabilities of the technology itself. So I, that's sort of where I am too, probably. I worry a lot about the capability of our institutions to respond. Right. People always talk about alignment problems. And one of my like, Pat arguments at this point is that the biggest alignment problems are corporations and governments. And, you know, I mean, this is to the credit of some of these AI companies, like they're coming out and saying, we have an alignment problem. Like our corporation's incentive is to race all the other corporations to try to, you know, get as much market share as we possibly can by moving faster than is safe. We are asking you to help slow us down, but we're not slowing them down, right? That, you know, we're not slowing them down. Currently the professed choice of the US government is do not slow down. Yeah. I think there, there are two problems here. One is the institution's problem that you put your finger on. I mean, I wouldn't let the companies off the hook so lightly. I do think they can unilaterally slow down and they're choosing not to do that. This is almost very specifically an open AI and anthropic problem. It's a culture problem. The reason for that is the underlying belief that racing to super intelligence is the thing that matters. And the only thing that matters is flipping the sign. Is that going to be safe super intelligence? That's going to save us or unsafe super intelligence. That's, that's going to kill us. And that is a very particular view. I think there's a lot of evidence pushing back against that, but I feel like these companies are a little bit of an echo chamber and resistant to the idea that there are so many economic bottlenecks to the benefits of AI. And it's not going to be whoever races to super intelligence is going to be the winner as a company or as a country or, you know, saving humanity. And if they recognize that, I think they find it in their own commercial interest to unilaterally and voluntarily slow down and shift a lot of their efforts to not just safety, but more importantly, taking their existing capabilities and making the models more usable, integrated into downstream applications and so forth. These companies claim that there are external forces pushing them to race. I think it's internal culture. And I guess the, the, the question I have here is if we think these are very, very powerful, very dangerous technologies, do we not need to, a culture of safety from the politics perspective that we're not currently enforcing. I'm pro-regulation. You know, we do oppose regulations like banning open models. Again, we don't think it's about a particular capability level. But the things about, yeah, changing the internal culture of companies through regulation, that's something we've definitely been on board. We need a lot more transparency. And yes, you know, the organizational change that we've been talking about. And I do think it's a problem that we're not currently doing that. So one way you can align corporate incentives with the public good is regulation. And the regulators are currently refusing to do that. I mean, we just saw them come out with a new law that said, you know, we're not going to do that. We're not going to do that. With a voluntary sort of semi-agreement between the AI labs, it is not going to be legally enforceable. But I think it was called by Trump morally enforceable, which is interesting. My sense is that competition is a very powerful force in highly competitive markets. I mean, even if you just like look at the social media companies, I think they've caused a tremendous amount of harm at a global scale because it was more important to them to win market share from each other. And to make sure that the way their systems are being used wasn't diminishing to human flourishing. And so I just, I think I have like a much more skeptical view. Like I really do think the profit incentive here when there's so much profit to be made and so much fear that your investment bubble could pop is a ferocious force and that the level of societal counterforce would need to be quite strong in order to force these companies to actually, to actually act with a level of safety that they would have. Yeah. I'm glad you brought up the comparison to social media. I wrote an essay a few years ago called Understanding Social Media Recommendation Algorithms. It was mostly about the algorithms themselves. But one of the points I also made was that this decision to optimize for engagement, keeping people scrolling, et cetera, was actually made without much regard to what is good for the company itself in the long run. For instance, I reviewed a study that came out of Meta itself that showed that when they had these kind of addiction maximizing design choices, like spamming people with notifications, in the short run, it increased people's use of the app. But over a period of about a year or so, they started quitting the app. And when I, you know, when I talk to my students now, there's a sizable fraction of them who have severely cut back or entirely quit social media because they realize that over a period of months or years, their experience really, you know, it's not going to be as good as it used to be. And I worry that the AI companies are caught in the same trap. This culture of, you know, racing toward the newest model at all costs, it might feel in the short term that that's what they need to do to get the headlines to be on top of the artificial analysis index, whatever. But because of, you know, the fact that there are these consequences for safety, hopefully they're going to, you know, get sued if they continue down this road. It's actually not in their own long-term commercial interest is what I feel. It may not be in their long-term commercial interest, although maybe the social media example is a good one to spend a second on because, look, the social media companies are much more viewed with a lot more skepticism today than they were in, say, you know, 2012. They are also richer today. Their valuation is higher. Meta is bigger, right? TikTok is a lot more popular today. So, you know, I think there's a lot more skepticism. I think there's a lot more skepticism. I think TikTok, on Instagram, on some of these sites. And the advertising is working better than ever. And I'm not sure that they are wrong. They can only be made wrong by society making a decision which disciplines the market into a different formation than it would naturally or currently be in. Yeah. I mean, again, I think I mostly agree. I do think, again, that companies can make decisions that are irrational in their own long-term interests because they're, especially in Silicon Valley, there is a culture really of focusing on these shorter-term metrics and A-B tests. So then you get into something that you're beginning to touch there, which is diffusion. And one place where you do have a view that is sort of different from some people in Silicon Valley is that it's going to be much harder for AI to show up in the economy, for AI to show up in the world, than people think. That there is not a one-to-one between individuals. It's going to be much harder for AI to show up in the economy, for AI to show up in the world, than people think. It's going to be much harder for AI to show up in the economy, for AI to show up in the world, intelligence in that. So talk to me a bit about diffusion. Yeah, this really clicked for me actually a few months after we wrote this essay, when I was looking at Amtrak's proud announcements of the new train sets that they had purchased for their Acela series. Apparently it can go 165 miles per hour. At first, I thought, this is going to be amazing. That's way faster than the trains currently go. And then I dug into it a little bit more, and it turns out the limiting speed is not the trains themselves. It's tracks that are too curved and the signaling infrastructure that is centuries old. And those things are not changing. And so the average speed hasn't really budged much. It's still 65 to 70 miles per hour. And it struck me that this was, you know, this is kind of an elegant way to say what we've been trying to say in AI as normal technology, which is that most of the time, AI is the trains. It's not the track. So AI is accelerating a part of the process that was never the bottleneck to begin with. It's so many other things that are more important than the trains. More infrastructural things that happen around the AI. Organizational culture, regulation, even, you know, our ability socially to accept the level of year-to-year change in our lives. Things like self-driving cars, no matter how many lives they might save, it's so much of a kind of a shock to society that it will almost inevitably lead to what we've been seeing already, the kind of political backlash. And it's going to take quite a while, I think, to make all the societal adjustments to be able to deploy these technologies. If we ever do. So this part of AI, the AI story I'm incredibly skeptical of, right? You'll hear a Sam Holtman say, talk about how AI through innovation is going to help us solve our energy problems. Another way to sort of distill down that idea is that the binding constraint right now on clean energy is intelligence. But it's not. It's not. We know we have much better energy than we're currently using for vast amounts of our energy infrastructure. And we're not doing it because of, it would be against some people's profits. We're not doing it because there are political limits to building in the real world. We're not doing it because Donald Trump hates solar and wind power. We're not doing it for all kinds of reasons. And it feels to me like a lot of things are like that, that if you accelerate or increase the amount of intelligence behind it, you just run like into the other rate limiters in society, right? Drug development has testing and the FDA and all the, I mean, abundance, my book is very much about this in other areas. We're aware of how to make faster trains and make them in other places. We don't. And it's not clear to me why AI would solve those problems quickly or potentially at all. That's right. And on top of that, there's various kinds of arms races. So there was this great report by insurance companies last week that talked about how AI has already added a billion dollars to medical expenses because hospitals are using it to be able to code more complex conditions for the same diagnoses and same treatment. Of course, we should be skeptical of any specific numbers they quote, but the New York Times article about it had other people, academics making the same point. And yeah, it's the kind of arms race that we see a lot. We see it in the legal profession, AI for law. There's so much excitement about that, but it's AI. It's an arms race where the equilibrium simply shifts upwards. I actually have a paper about this with Justin Curl. And what we talk about there is that not only is there an arms race, there are other bottlenecks. Like if you make lawsuits a lot more efficient, there are still only a finite number of judges. And I think we need, you know, we need those human judges. We shouldn't replace human judges with AI. I think even if that's more efficient in some sense, to me, that's axiomatically, that's giving up control over, you know, the course of human destiny to AI because judges make law and that's not something we should be giving up to AI. So those are some really fundamental bottlenecks. I want to get back at that AI versus AI point you just made there because I think this is really underplayed. I think a lot about why did the internet not lead to a larger increase in global productivity and innovation than it did? And I always think the It actually did make it possible to collaborate with people all over the world. Incidentally, it did make virtually the entire corpus of human knowledge first available to us, then available turned out to AI to train on it also did this opposite thing, like it did speed us up and it also slowed us down to distracted us. So now while you're working on something. you're clicking back and forth from your email and into a you know an online game and over to social media and your ability to focus is degraded and you know there's tremendously more porn which appears to have had an effect on whether or not people are forming real life human relationships that adding or reducing friction in one area also reduces it in areas that are maybe less beneficial and when you think of adding intelligence well that intelligence is going to add on the other side of things too and i always think that we at the beginning of a technology we think of all the ways it can make everything better and i think we we think a lot right now with ai about the ways it could make things dramatically worse right uh huge cyber security events or destruction of the financial system or human extinction but just the ways it might make things worse in banal fashions where it just like increases the ability of people to be able to do things that they don't want to do and i think that's a really important thing to think about and i think that's a really important thing to think about and i think that's people to waste everybody else's time is a little bit underplayed in my view for sure and you know it's it's not just banal i think there are a lot of these sub-catastrophic harms that are pretty serious uh with the industrial revolution we had several decades of horrible labor conditions i think with ai as well uh while there has been so much focus on job displacement what there has been much less focus on is how it's changing job quality and i think this is a big under appreciated area because ai is turning a lot of knowledge workers into managers of ai agents right and the thing about being a manager is it's kind of a shitty experience because you're responsible for other people's mistakes you don't get to practice the craft that you trained for but you know us human managers we don't get to complain it's higher pay it's higher status we chose it and of course mentoring people is gratifying with a how you don't get any of those benefits you have to manage this agent and be responsible for its mistakes but you don't get to practice the craft i think we can design ai agents differently to avoid this but right now that's kind of where things are going and yeah i think we should be very concerned about that so how much though does this story that we are telling here lead to a a theory of what's going to happen in the economy which is very jagged that the things where you need fusion into the real world right things have to happen physically buildings need to be built energy you know transmission lines you need to be laid down that has such powerful rate limiting on it that it can't accelerate that fast meanwhile inside the digital world things can move very very very fast so you know in terms of what's going to happen to you know white collar workers who they work kind of completely on a computer and their work actually can be automated you know they're in a call center or something like that or kind of separately like all the cyber crime stuff we're talking about and the cyber security that i think a a version of this future playing out that worries me is that actually most of what would improve people's lives has to happen in like the physical real world but where ai is going to be able to move the fastest is in the digital world and as an equilibrium i don't think that sounds like the world in which we're getting the most benefit from ai and it might in fact be the world we're getting the most harm from it i yeah that's possible i don't know i think there are ways you know we can we can change that i don't think fast is that fast first of all like you mentioned call centers i mean those are still here uh you know when chat gpt was released so many people were predicting that within a year we would have replaced all of them i mean chatbot it's right there in the name if as is likely ai is going to make call center workers a lot more productive there's a lot of latent demand a lot of the time we don't call call centers because it's it's a frustrating experience so once again it's a jevons paradox thing you want to describe what that is certainly yeah so it's the idea that when something becomes cheaper to produce there's now more demand for it let's look at the sector where the capabilities are already the most advanced which is probably software engineering it used to be extremely expensive to produce software so only maybe a few tens of thousands of lines of code worldwide were written per year and now that's expanded by something like a million fold and so over the long run you know whether this is going to be something that increases demand for software engineers or whether it's going to be something that increases demand for software the rocky you know job prospects we're seeing for junior software engineers are going to continue remains to be seen i think either is possible but you know either way if this happens over the course of 20 years again that is in line with other major shifts such as the industrial revolution where many jobs went away but many other new jobs were created another example translation jobs you know way back around 2016 machine this is obviously way before what we call generative ai now translation models became pretty close to human parity but those jobs are still pretty intact the nature of the job has changed a lot so once again there's there's a lot more demand that has been unlocked by the fact that you can translate anything to any language now so jevin's paradox i had just did this conversation with bill gates and when i brought that up he was a little bit about it so you don't think jevin's paradox name a blue collar profession that's subject to jevin's paradox or do you not care about blue collar i do care about blue collar so the question is anything in the blue collar realm so i think subject to that and his point was look software engineering is indeed a sector of the economy where there's a lot of unmet demand you know arguably everybody would like their own software engineer so in a world where you rapidly accelerate that you can get this demand effect where it just creates more demand for software engineers because now they're cheaper but what he went on to say was that a lot of things are going to be like that you think of say a truck driver if we get to the point where trucks are driverless one there is only so much demand for trucking and two there's no longer a driver in that truck if you look at a lot of the people who got their jobs automated away or sent to china in manufacturing it is true that the economy kept growing but many of those people individually had a very very very hard time and so his argument is jevin's paradox is not going to be big enough to to handle this because there are too many areas of the economy that are going to be able to do that so i think he's completely right about the truck drivers i think uh totally agree there the demand there is relatively finite the question is whether that's the rule or the exception and let me put it this way look at what we're doing here nobody asked for this you know if you went back in time 100 years or 200 years it would be like how is this a real job most of the jobs that we have today are jobs that are kind of higher up the maslow's hierarchy if you will they're not meeting some actual real fixed demand or need that people need in order to live their lives we do these things because you know they're fun and people like to listen to it most white-collar jobs are like that that's my view most white-collar jobs do have jevin's paradox if it get you know for if it gets easier to produce more of there will be demand for it especially as people's incomes as well rise very gradually and there's more demand for it and there's more demand for it and there's more spending on these you know less necessary more luxury kinds of things and blue collar as well there's a great essay by alex emas called what will be scarce and he points out that a job of a starbucks barista for instance already should not exist that we've long known how to automate that we can make coffee at home even i don't know if he said that specific thing but you know that's an example where it's the relational nature of the job that matters and therefore those are going to be i think pretty stable even if you know it ai makes it in some way cheaper to do and so i think the truck driver kinds of jobs are more the exception so i think as we come to a close here what would have to happen in the next couple of years for you to say oh this is looking less normal than we thought or this is more off course than we thought what what are what is the evidence that would have to come in for you to like significantly alter your thesis for sure yeah there's stuff on the economy there's stuff on safety on economy if we start to see at some capability level that we're going to have to it's not the same process of humans simply adapting to it and you know using it to amplify their productivity and managing the agents which is what we're seeing so far but instead starts to wholesale replace whether a software engineer or any other profession i think that would be pretty different from what we're predicting for the most part and then on safety especially with companies claiming that they're close to recursive self-improvement i mean i don't think they should you know plunge forward and say oh we're going to have to do this we're going to have to do this towards fully autonomous recursive self-improvement in the first place which i think you've said as well but nonetheless our view is that even if that happens it's not going to lead to super intelligence because the bottlenecks to super intelligence are external there is nothing you can do in a lab that's going to teach the ai model you know how to cure cancer or whatever it is the companies are hoping for um but again that's an empirical claim and that would certainly completely falsify our thesis then always our final question what are three books you'd recommend to the audience sure you know in this conversation i've generally had a bit more optimistic take on things than we're used to hearing especially on ai safety so maybe in keeping with that i really like hannah ritchie's book not the end of the world she has a newer book but this one's from 2024 and i still like it very much the subtitle is something like how we can be the first generation to build a sustainable planet it's an optimistic take on climate which is of course usually full of doom and gloom stories so i really liked it for that reason It's a very optimistic book on technology. I also like that book a lot and it makes you realize that we actually do make things that are better over time, which I think sometimes we can get into an overly negative place in technology. Right. On China, which is, of course, a topic that so many people are interested in. I'm sure you've heard this book recommendation a lot. I liked Dan Wang's breakneck, a lot of echoes of abundance as well. But this idea of thinking about a lawyerly society versus an engineering society was, I thought, a really good and succinct way to capture a lot of the macro and micro differences. And then the last one is an old classic. If I can say as a preamble, I read a two-page paper one time called How Complex Systems Fail. I thought the paper was about software. I realized, in fact, that it was about medical systems written by an anesthesiologist. And I learned that there is the study of systems that actually explain the patterns in all kinds of systems, natural and social and engineered systems. And that led me to the book, Systems Thinking by Donella Meadows from many, many years ago. Arvind Narayanan, thank you very much. Thank you, Ezra. This has been so fun. Thank you.

Podcast Summary

Key Points:

  1. AI is not fundamentally different from past transformative technologies like electricity or the internet, and should be treated as a normal, controllable technology.
  2. The "AI as abnormal technology" view—especially the FOOM hypothesis of uncontrollable superintelligence—is overstated and relies on speculative, unverified assumptions.
  3. Harmful AI behaviors, like those seen in the Hugging Face hacks, stem from poor cybersecurity, operational failures, and flawed design choices, not from emergent intelligence or volition.
  4. These incidents are better understood as failures of engineering and operational excellence, not signs of AI taking control or becoming self-aware.
  5. Effective AI control requires robust technical measures such as sandboxes, real-time monitoring, tool access tracking, and AI-based anomaly detection.
  6. The core challenge is not intelligence per se, but the design of systems and the alignment of incentives—especially in corporate culture and governance.
  7. Current AI safety concerns are often exaggerated due to underinvestment in control systems and overreliance on alignment alone.
  8. Like past technologies, AI will face societal, political, and infrastructural bottlenecks that limit its real-world impact, making widespread transformation unlikely without deeper systemic changes.

Summary:

The debate over whether artificial intelligence represents a uniquely dangerous "abnormal" technology is challenged by Arvind Narayanan’s argument that AI is fundamentally a normal, controllable technology—similar to past transformative innovations like electricity or the internet. He rejects the idea of uncontrollable "superintelligence" or a FOOM event, asserting that AI risks stem not from inherent intelligence or volition, but from poor engineering, cybersecurity gaps, and organizational failures. The Hugging Face hacks, often cited as evidence of AI autonomy, are better explained as failures of operational excellence and inadequate system design.

Key solutions involve robust technical controls such as real-time monitoring, sophisticated sandboxes, and AI-driven anomaly detection. Crucially, these are solvable engineering problems, not existential threats. The real danger lies not in AI’s capabilities, but in corporate incentives that prioritize speed and market share over safety.

AI’s impact on the economy is also constrained by real-world bottlenecks—such as infrastructure, regulation, and social resistance—meaning its digital advantages won’t translate into broad societal gains. This perspective emphasizes that the future of AI depends not on theoretical fears of superintelligence, but on practical improvements in governance, transparency, and corporate responsibility. Ultimately, AI’s risks are manageable through proactive, well-informed engineering and institutional change, rather than revolutionary or apocalyptic scenarios.

FAQs

Arvind Narayanan argues that AI is a 'normal technology'—similar to electricity or the internet—rather than a fundamentally new or dangerous 'abnormal' one. He believes we have a proven roadmap for managing such technologies through engineering, oversight, and operational controls.

He challenges the idea that AI will undergo a sudden, uncontrollable takeoff to superintelligence. He argues that the dangerous capabilities of AI arise not from intrinsic intelligence but from poor design, cybersecurity failures, and operational shortcomings—issues that can be addressed with engineering solutions.

The Hugging Face hacks were not due to emergent AI intelligence but to failures in cybersecurity and operational excellence. The incidents highlight that harmful behavior stems from poor design choices, such as allowing models to persist, communicate, or access external tools, not from AI becoming self-aware or malicious.

He recommends stronger sandboxes, real-time monitoring of tool calls, monitoring of internal reasoning (chain-of-thought), and detection of coordination among agents. These are engineering-based interventions that can prevent harmful behavior before it escalates.

No, he argues that increased AI 'smartness' does not necessarily mean loss of control. Both AI systems and defensive systems can be smart, and the problem lies in design and oversight, not in the inherent intelligence of AI.

He finds such claims speculative and unsupported. He believes AI lacks the capability to persuade humans to perform dangerous actions, especially when those actions violate laws or established incentives. Such behavior would be an engineered property, not an emergent one.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.