Go back

Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism

69m 38s

Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism

In this interview, Anthropic CEO Dario Amodei discusses the rapid advancement of AI, emphasizing an exponential growth trajectory in capabilities due to scaling laws in compute, data, and training techniques. He argues that societal and economic impacts are nearer than many realize, prompting his urgent public warnings about risks, though he also acknowledges AI's positive potential. Amodei dismisses terms like AGI as marketing, instead focusing on tangible progress, such as Anthropic's models improving significantly in coding and other tasks. He addresses criticisms, including accusations of wanting to control the industry, and defends Anthropic's stance against larger competitors like Meta and XAI, highlighting the company's reliance on talent density, mission alignment, and principled compensation over bidding wars for employees. Despite raising substantial funds, he asserts Anthropic remains competitive in resources and infrastructure, confident in its approach to navigating AI's future challenges and opportunities.

Transcription

11890 Words, 64393 Characters

English
Fisically responsible, financial geniuses, monetary magicians. These are things people say about drivers who switch their car insurance to progressive and save hundreds. Because progressive offers discounts for paying in full, owning a home and more. Plus, you can count on their great customer service to help when you need it so your dollar goes a long way. Visit progressive.com to see if you could save on car insurance. If you have a chance to save on car insurance, potential savings will vary, not available on all states or situations. You can also use the store to buy a car and you can also buy a car that you can use in your car. The store is now at 60 km/h. Life's better with a story. I get very angry when people call me a dumber. When someone is like, "This guy is a dumber, he wants to slow things down." You've heard what I just said. My father died because of cures that could have happened a few years later. I understand the benefit of this technology. I'm sure you've heard the criticism from people like Jensen who say, "Dario thinks he's the only one who can build the safely and therefore wants to control the entire industry." I've never said anything like that. That's an outrageous lie. That's the most outrageous lie I've ever heard. Anthropic CEO, Dario Amode joins us to talk about the path forward for artificial intelligence. Whether gendered AI is a good business and to fire back at those who call him a dumber. And he's here with us in studio at Anthropic Headquarters in San Francisco. Dario, it's great to see you again. Welcome to the show. Thank you for having me. So let's recap the past couple of months for you. You said, "AI could wipe out half of entry-level white collar jobs. You cut off wind-serves access to Anthropics top-tier models when you learned that Open AI was going to acquire them. You asked the government for export controls and annoyed Nvidia CEO, Jensen Wang. What's gotten into you?" I think Anthropic, myself and Anthropic are always focused on trying to do and say the things that we believe. And I think as we've gotten more close to AI systems that are more powerful, I think I've wanted to say those things more forcefully, more publicly to make the point clearer. I've been saying for many years that we have these, we can talk in detail about them, but we have these scaling laws. AI systems are getting more powerful. They're going from the level of a few years ago. They were barely coherent. A couple years ago, they were at the level of a smart high school student. Now we're getting to smart college student, PhD, and they're starting to apply across the economy. So I think all the issues related to AI ranging from the national security issues to the economic issues are starting to become quite near to where we're actually going to face them. So I think as these problems have come closer, even though in some form and Anthropic has been saying these things for a while, I think the urgency of these things has gone up. And I want to make sure that we say what we believe and that we warn the world about possible downsides. Even though no one can say what's going to happen, we're saying what we think might happen, what we think is likely to happen. We back it up as best we can, although it's often extrapolations about the future where no one can be sure. But we see ourselves as having the duty to warn the world about what's going to happen. And that's not to say, I think there's an incredible number of positive applications of AI. I've continued to talk about that. I wrote this essay, Machines of Loving Grace. I feel in fact that I and Anthropic have often been able to do a better job of articulating the benefits of AI than some of the people who call themselves optimists or accelerationists. So I think we probably appreciate the benefits more than anyone. But for exactly the same reason, because we can have such a good world if we get everything right, I feel obligated to warn about the risks. So all of this is coming from your timeline. Basically, it seems like you have a shorter timeline than most. And so you were feeling a sense of urgency to get out there because you think that this is imminent. Yes, I'm not short. I think it's very hard to predict particularly on the societal side. So if you say when are people going to deploy AI or when are companies going to use X dollars of spend of AI or when will AI be used in these applications? Or when will it drive these medical cures? That's kind of harder to say. I think the underlying technology is more predictable, but still no one knows. But I think on the underlying technology, I've started to become more confident. There isn't no uncertainty about it. I think the exponential that we're on could kind of still, could totally peter out. I think there's maybe 20 or 25% chance that sometime in the next two years, the models just start getting better. For reasons we don't understand or maybe reasons we do understand data or computer availability. And then everything I'm saying just seems totally silly. And everyone makes fun of me for all the warnings I've made. And I'm just totally fine with that given the distribution that I see. And so I should say that this is part, our conversation is part of a profile I'm writing about you. I've spoken with more than two dozen people who work with you, who know you, who've competed with you. And I'm going to link that in the show notes if anybody wants to read it. It's free to read. But one of the themes that has come through across everybody I've spoken with is that you have about the shortest timeline of any of the major lab leaders. And you just referenced it just now. So why do you have such a short timeline? And why should we believe in yours? Yeah, it really depends what you mean by timeline. So one thing, and I've been consistent on this over the years, is there are these terms in the AI world like AGI and super intelligence. You'll hear leaders of companies say we've achieved AGI when moving on to super intelligence. Or it's really exciting that someone stopped working on AGI and started working on super intelligence. I think these terms are totally meaningless. I don't know what AGI is. I don't know what super intelligence is. It sounds like a marketing term. It is marketing. Yeah, it sounds like something designed to activate people's dopamine. So you'll see in public, I never use those terms. I'm actually careful to criticize the use of those terms. But I think despite that, I am indeed one of the most bullish about AI capabilities improving very fast. What I think is real, that I've said over and over again, is the exponential. The idea that every few months we get an AI model that is better than the AI model we got before, and that we get that by investing more compute in AI models, more data, more new types of training models. Initially, this was done by what's called pre-training, which is when you just feed a bunch of data from the internet into the model. Now we have a second stage that's reinforcement learning or test time compute or reasoning or whatever you want to call it. I think of it as a second stage that involves reinforcement learning. Now both of those things are scaling up together. As we've seen with our models and as we've seen with models from other companies, and I don't see anything blocking the further scaling of that. There's some stuff about how do we broaden the tasks on the RL side of it? We've seen more progress on, say, math and code where the models are getting pretty close to a high professional level and less on more subjective tasks, but I think that is very much a temporary obstacle. When I look at it, I see this exponential and I say, "Look, people aren't very good at making sense of exponentials." If something is doubling every six months, then two years before it happens, it looks like it's only one-sixteenth of the way there. We are sitting here in the middle of 2025 and the models are really starting to explode in terms of the economy. If you look at the capabilities of the model, they're starting to saturate all the benchmarks. If you look at revenue and throttpix revenue, every year has grown 10X. Every year we're conservative and we say, "It can't grow 10X this time. I never assume anything." Actually, always I'm very conservative in saying, "I think it's going to slow down on the business side." We went from zero to 100 million in 2023. We went from 100 million to a billion in 2024. This year, in this first half of the year, we've gone from one billion to, as of speaking today, it's well above four. It might be 4.5. If you think about it, suppose that exponential continued for two years. I'm not saying it will, but suppose it continued for two years. You're well into the 100 billions. I'm not saying that will happen. I'm saying the situation is that when you're on an exponential, you can really get fooled by it. Two years away from when the exponential goes totally crazy, it looks like it's just starting to be a thing. And so there. That's the fundamental dynamic. We saw that with the internet in the 90s, where it was like networking speeds and underlying speed that the computers were getting fast. Over a few years, it became possible to have to basically build a digital global communications network on top of all this when it wasn't possible just a few years ago. Almost no one except for a few people really saw the implications of that and how fast it would happen. That's where I'm coming from. That's what I think. Now, I don't know if a bunch of satellites crashed, maybe the internet would have taken longer. If there was an economic crash, maybe it would have taken a little longer. So we can't be sure of the exact timelines, but I think people are getting fooled by the exponential and not realizing how fast it might be, how fast I think it probably will, although I'm not sure. But so many folks in the AI industry are talking about diminishing returns from scaling now that really doesn't fit with the vision you just laid out. Are they wrong? Yeah. From what we've seen, I can only speak in terms of the models at Anthropic. But what I've seen in terms of the models at Anthropic, if we look at, let's take coding. Coding is one area where I think Anthropic models have advanced very quickly. Adoption has been very quick. We're not just a coding company. We're planning to expand to many areas. But if you look at coding, we release 3.5 on it, a model we call 3.5 on it V2, which is what's called 3.6 on it now, 3.7 on it, and then 4.0 on it, and 4.0 obis. That series of four or five models, each one got substantially better at coding than the last. At benchmarks, you can look at sweet bench growing from, I think, 18 months ago, is that 3% or something growing all the way to 72% to 80% depending on how you measure it. The real usage has grown exponentially as well, we're heading more and more towards autonomously you can just use these models. I think the actual majority of code that's written at Anthropic is at this point written by, or at least with the involvement of one of the quad models. Other companies have said similar statements to that. We see the progress as being very fast and the exponential is continuing, and we don't see any diminishing returns. But there are some liabilities it seems like with letters language models. For instance, continual learning. We had to work as you on a couple of weeks ago. Here's how he wrote about it in his sub-stack. The lack of continual learning is huge, huge problems. The LM baseline at many tests might be higher than an average human, but you're stuck with the abilities you get out of the box. You just make the model and that's it. It doesn't learn. That seems like a glaring liability. What do you think about that? First of all, I would say, even if we never solved continual learning, even if we never solved continual learning and memory, I think that the potential for the LM's to do incredibly well to affect things of the scale of the economy will be very high. If I think of the field I used to be in, biology and medicine, let's say I had a very smart Nobel Prize winner. I said, OK, you've discovered all these things. You have this incredibly smart mind, but you can't read new textbooks or absorb any new information. That would be difficult, but still, if you had 10 million of those, there's still going to make a lot of biology breakthroughs. They're going to be limited. They're going to be able to do some things humans can't and there are some things humans can do that they can't. Even that, even if we impose that as a ceiling, man, that's pretty damn impressive and transformative. Even if I said you never solved that, I think people are underestimating the impact. But look, context windows are getting longer and models actually do learn during the context window. So as I talk to the model during the context window, I have a conversation, it absorbs information. The underlying weight to the model may not change, but just like I'm talking to you here and we're having a conversation and I listen to the things you say and I think and I like respond to them. The models are able to do that and from a machine learning perspective, from an AI perspective, there's no reason we can't make the context length 100 million words today, right, which is roughly what a human here is in their lifetime. There's no reason that we can't do that. It's really inference support. And so again, even that fills in many of the gaps, not all the gaps, but it fills in many of the gaps and then there are a number of things like learning and memory that do allow us to update the weights. So there are a number of things around types of reinforcement learning, learning training. We used to many years ago talk about inner loops and outer loops, right? The inner loop is like, I have some episode and I learned some things in that episode and I'm trying to optimize for the lifetime of that episode and kind of the outer loop is the agent's learning over episodes. And so I think maybe that inner loop outer loop structure is a way to learn the continual learning. One thing we've learned in AI is whenever it feels like there's some fundamental obstacle, like two years ago we thought there was this fundamental obstacle around reasoning. Turned out just to be just to be RL, you just train with RL and you let the model write some stuff down, you let the model write things down to try and figure out objective math problems. How being too specific, I think and we already have maybe some evidence to suggest that this is another of those problems that is not as difficult as it seems, that we'll fall to scale plus a slightly different way of thinking about things. Do you think your obsession with scale might blind you to some of the new techniques like Demisosavis says, to get to AGI or you might call it super powerful AGI, whatever, human level intelligence is what we're all talking about. We might need a couple of new techniques for that. So we're developing new techniques every day. Okay. Claude is very good at code and we don't really talk externally that much about why Claude is so good at code. Why is it so good at code? Like I said, we don't talk externally about it. So every new version of Claude that we make has improvements to the architecture, improvements to the data that we put into it, improvements to the methods that we use to train it. So we're developing new techniques all the time. New techniques are a part of every model that we build and that's why I've said these things about we're trying to optimize for talent density as much as possible. You need that talent density in order to invent the new techniques. There's one thing that's been hanging over this conversation, which is that maybe Anthropic is the company with the right idea but the wrong resources because you look at what's happening with XAI and inside Meta where Elon's built his massive cluster. Mark Zuckerberg is building this five gigawatt data center and they are putting so much resources towards scaling up. Is it possible? And Anthropic obviously you have raised billions of dollars but these are trillion dollar companies. So we've raised, I think at this point, a little sort of 20 billion dollars. It's not bad. So that's not nothing. And I would also say if you look at the size of the data centers that we're building with, for example, Amazon, I don't think our data center scaling is substantially smaller than that of any of the other companies in the space. In many cases, these things are limited by energy, they're limited by capitalization. When people talk about these large amounts of money, they're talking about it over several years. And when you hear some of these announcements, sometimes they're not funded yet. We've seen the size of the data centers that folks are building and we're actually pretty confident that we will be within a rough range of the size that they've built. You talked about talent density. What do you think about what Mark Zuckerberg is doing on the talent density front? I mean, combining that with these massive data centers, it seems like years can be able to compete. Yeah. So this is actually very interesting because one thing we noticed is that relative to other companies, I think a lot fewer people from Anthropic have been caught by these. It's not for lack of trying. I've talked to plenty of people who got these offers at Anthropic and who just turned them down. Who wouldn't even talk to Mark Zuckerberg who said, no, I'm staying at Anthropic. Our general response to this was, I posted something to the whole company Slack where I said, look, we are not willing to compromise our compensation principles, our principles of fairness to respond individually to these offers. The way things work at Anthropic is there's a series of levels. When candidate comes in, they get assigned to level and we don't negotiate that level. Because we think it's unfair. We want to have a systematic way. If Mark Zuckerberg throws a dart at a dartboard and hits your name, that doesn't mean that you should be paid 10 times more. more than the guy next to you who's just as skilled, who's just as talented. And my view of the situation is that the only way you can really be hurt by this is if you allow it to destroy the culture of your company by panicking, by treating people unfairly in an attempt to defend the company. And I think actually this was a unifying moment for the company where we didn't give in. We refused to compromise our principles because we had the confidence that people are in the topic because they truly believe in the mission. And I think that gets to how I see this. I think that what they are doing is trying to buy something that cannot be bought. And that is alignment with the mission. And I think there are selection effects here. Are they getting the people who are most enthusiastic or most mission-aligned, who are most excited to- But they have talent and GPUs. You're not underestimating them? I will see how it plays out. I am pretty bearish on what they're trying to do. So let's talk a little bit about your business because a lot of people have been wondering, is the business of gendered vei a real thing? And I'm also curious. I have questions all the time. You talked about how much money you've raised, close to 20 billion. You've raised 3 billion from Google, 8 billion from Amazon, 3.5 billion from a new round, led by light speed. What I've spoken with, what is your pitch because you are not part of like a big tech company you're out there on your own. Do you just bring the scaling laws and say, can I have some money? So my view of this has always been that talent is the most important thing. So if you go back three years ago, we were in a position where we had raised hundreds of millions. Open AI had already raised 13 billion from Microsoft. And of course, the large hyper-cap tech companies were sitting on $100 billion, $200 billion. And basically the pitch we made then is we know how to make these models better than others do. There may be a curve of scaling laws. But look, if we are in a position where we can do for $100 million, what others can do for a billion, and we can do for $10 billion, what they can do for $100 billion, then it's 10 times more capital efficient to invest in an entropic than it is to invest in these other companies. Would you rather be in a position where you can do anything for 10 times cheaper or where you start with a large pile of money? If you can do things 10 times cheaper, the money is a temporary defect that you can remedy if you have this intrinsic ability to build things for the same price much better than anyone else or as good as anyone else for much lower price, investors aren't idiots or at least they aren't always idiots. It depends much when you go to. Not going to name any names. But they basically understand the concept of capital efficiency. And so we've been in a position three years ago where these differences were like 1,000 and now you're saying with $20 billion can you compete with $100 billion? And my answer is basically yes because of the talent density. I've said this before but the entropic is actually the fastest growing software company in history at the scale that it's at. So we grew from zero to 100 million in 2023, 100 million to a billion in 2024. And this year we've grown from 1 billion to, I think I said this before, 4.5. So that 10X a year, every year I like suspect that we'll grow at that scale and every year I'm almost afraid to say it publicly because I'm like no, it couldn't possibly happen again. So I think the growth at that scale speaks for itself in terms of our ability to compete with the big players. Okay, so CNBC says 60 to 75% of entropic sales are through the API that was according to internal documents. Is that still accurate? I won't give exact numbers but the majority does come through the API although we also do have a flourishing apps business and I think more recently the max tier which power users use as well as cloud code which coders use. So I think we have a thriving and fast growing apps business but yes the majority comes through the API. So you're making the most pure bet on this technology. Like OpenAI might be betting on chat GPD and Google might be betting on the fact that no matter where the technology goes it can integrate into Gmail and calendar. So why have you made this bet on the pure bet on the tech itself? Yeah, I mean I would say I wouldn't quite put it that way. I think we'd describe it more as we bet on business use cases of the model. More so than we bet on the API per se and it's just that the first business use cases of the model come through the API. So as you mentioned OpenAI is very focused on the consumer side. Google is very focused on kind of the existing products that Google has. Our view is that if anything the enterprise use of AI is going to be greater even than the consumer use of AI. I should say the business use because it's enterprise, it's startups, it's developers, and it's kind of power users using the model for productivity. I also think that being a company that's focused on the business use cases actually gives us better incentives to make the models better. A thought experiment that I think is worth running is suppose I have this model and it's as good as an undergrad at biochemistry. And then I improve it and it's as good as a PhD student at biochemistry. If I go to a consumer, if I give them the chatbot and I say great news, I've improved the model from undergrad to graduate level in biochemistry. Maybe I don't know 1% of consumers care about that at all. 99% are just going to be like, I don't understand it either way. But now suppose I go to Pfizer and I say I've improved this from undergrad at biochemistry to graduate biochemistry. This is going to be the biggest deal in the world. They might pay 10 times more for something like that. It might have 10 times more value to them. The general aim of making the models solve the problems of the world to make them smarter and smarter but also able to bring many of the positive applications. The things I wrote about in machines of loving grace, of solving the problems of biomedicine, solving the problems of geopolitics, solving the problems of economic development. As well as more prosaic things like finance or legal or productivity or insurance, I think it gives a better incentive to develop the models as far as possible. I think in many ways it may even be a more positive business. I would say we're making a bet on the business use of AI because it's most aligned with the exponential. Then briefly how did you decide to go with the coding use case? Yes. Originally as happens with most things, we're trying to optimize for making the model better at a bunch of stuff and coding particularly stood out in terms of how valuable it was. I've worked with thousands of engineers and there was a point about a year, a year and half ago where one of the best I'd ever worked with said every previous coding model has been useless to me and this one finally was able to do something I wasn't able to do. Then after we released it, it started getting quick adoption. This was around the time that a lot of the coding companies like Kurser, Windsor, GitHub, Augment Code started exploding in popularity and then when we saw how popular it was, we kind of doubled down on it. My view is that coding is particularly interesting because A, the adoption is fast and B, getting better at coding with the models actually helps you to develop the next model. It has a number of advantages. Now you're selling your AI coding through Clawed Code but it's very interesting. The pricing model has been confounding to some. You can spend $200 a month and get the equivalent I spoke to one developer. They got the equivalent of $6,000 a month from your API. Ed Zitran has pointed out the more popular that your models get, the more money you're going to lose if people are super users of this technology. How does that make sense? Actually pricing schemes and rate limits are surprisingly complicated. Some of this is basically the result of when we released Clawed Code in the max tier which we eventually tied together, actually not fully understanding the implications of the ways in which people could use the models and how much they were actually able to get. Over the last few days as of the time of this interview, we've adjusted that particularly on the larger models like Opus. I think it's no longer possible to spend that much with a $200 subscription and it's possible more changes will come in the future. We're always going to have a distribution of users who use a lot and users who lose some amount. It doesn't necessarily mean we're losing money that there are some users who get more, but if you were to measure via API credits, spend a better deal on the consumer on the consumer subscription, then they would on the API products. There's a lot of assumptions there, and I can tell you that some of them are wrong. We are not, in fact, losing money. But I guess there's another question about whether you can continue to serve these use cases and not raise prices. So just to give you a couple of stats, there are some developers that are upset because using anthropics newer models and cursors, causing them more than never has. Startups that I've spoken with say, anthropic is down a bunch, a down a bunch because they can't get access to the GPUs. At least that's what they imagine is happening. And I was just with Amjab Masad, a replet, in an interview that we're going to air next week, who said there was a period of time where the price per token, the price to use these models, was coming down, and it stopped coming down. So is it, is what's happening that these models are just so expensive for anthropic to run that it's hitting a wall of its own? Again, I think you're making assumptions here. That's why I'm asking to see you. Yeah. The way I think about it is we think about the models in terms of how much value are they creating. So as the models get better and better, I think about how much value they create. And there's a separate question about how the value is distributed between those who make the model, those who make the chips, and those who make the-- the underlying applications. So again, without being too specific, like I think there are some assumptions in your question that are not necessarily correct. I can tell you which ones. So I'll say this. I do expect the price of providing a given level of intelligence to go down. I expect the price of providing the frontier of intelligence, which will provide an increasing economic value. That might go down. My guess is it probably stays about where it is. But again, the value that's created goes way up. So two years from now, my guess is that we'll have models that cost of the same order of magnitude that they cost today, except they'll be much more capable of doing more much more autonomously, much more broadly than they are capable of today. One of the things that I'm John mentioned was he thinks that the bigger models are not as intensive to run, or more intensive to run, given their size, because of the architecture and some of these techniques that we talked about, that they're lighting up only certain sections of the model. So his idea-- I'm hopefully conveying this truthfully-- is that anthropic can run these models without too much bulk on the back end, but is still keeping those prices where they are. And I think the line that I'm going to draw there is maybe that to get to software margins, there were some reports that anthropic is slightly below software, gross margins, you're going to have to charge a little bit more for these models. So yeah, again, I think larger models cost more to run than smaller models. I think the technique you're referring to is maybe mixture of experts or something like that. So whether your models or mixture of experts are not, like mixture of experts is a way to run models more cheaply that have a given number of parameters. It's a way to train models. But if you're not using that technique, then larger models that don't use that technique cost more to run than smaller models that don't use that technique. And if you're using that technique, larger models that use that technique cost more to run than smaller models that are using that technique. So I think that's sort of a distortion of the-- I think that's sort of a distortion of the situation. Basically, I'm just guessing, and I'm trying to find out what the truth is from you. Yeah, look, so I-- in terms of the cost of the models, one thing you'd be surprised by people-- people kind of impute this thing to like, oh, man, it's going to be really hard to get the margins from X% to Y%-- we make improvements all the time that make the models 50% more efficient than they are before. We are just the beginning of optimizing inference. Inference has improved a huge amount where it was a couple years ago to where it is now. That's why the prices are coming down. And then how long is it going to take to be profitable? Because I think the loss is going to be like 3 billion this year. That's what they think. I would distinguish different things. There's the cost of running the model. So for every dollar of the model makes, it costs a certain amount. That is actually already fairly profitable. There are separate things. There's the cost of paying people and buildings. That is actually not that large in the scheme of things. The big cost is the cost of training the next model. And I think this idea of the company's losing money and not being profitable is a little bit misleading. And you start to understand it better when you look at the scaling loss. So as a thought exercise, these numbers are not exactly even close for endropic. Let's imagine that in 2023, you train a model that costs $100 million. And then in 2024, you deploy the 2023 model. It makes $200 million in revenue. But you spend $1 billion to train a new model in 2024. And then in 2025, the billion dollar model makes $2 billion in revenue. And you spend $10 billion to train the next model. So the company every year is unprofitable. It lost $800 million in 2024. And then in 2025, it lost $8 billion. So this looks like a hugely unprofitable enterprise. But if instead, I think in terms of is each model profitable, think of each model as a venture. I invested $100 million in the model. And then I got $200 million out of the model in the next year. So that model had 50% margins and made me $100 million. And the next year, the company invested $1 billion and made $2 billion in the next model. The company invested $1 million. So every model is profitable. There are but the company is unprofitable every year. I'm not, this is a style. I'm not claiming these numbers for anthropoclaming these facts. But this general dynamic is in general terms. The explanation for what is going on. And so at any time, if the models stopped getting better or if a company stopped investing in the next model, you would have probably a viable business with the existing models. But everyone is investing in the next model. And so eventually, it'll get to some scale. But the fact that we're spending more to invest in the next model suggests that the scale of the business is going to be larger the next year than it was the year before. Now of course, what could happen is the model stopped getting better. And there's this kind of one time cost that's like a boondoggle. And we spend a bunch of money. But then the companies, the industry will kind of return to this plateau, to this level of profitability, or the exponential can keep going. So I think that's a long-winded way to say, I don't think it's really the right way to think about things. Right. But what about open source? Because if you stopped investing in the models and open source caught up, then people could swap in open source. Now I'd love to hear your perspective on this because one of the things people have talked to me about when it comes to the anthropic business is there is that risk. Eventually, that open source gets good enough that you can take anthropic out and put open source in. Yeah. So people have, I think one of the things that's been true of this industry is that, and you know, I saw it early in, I saw it early in the history of AI. Every community that AI has gone through, it has this set of heuristics about how things work. Like back when I was in, you know, AI back in 2014, there was an existing kind of AI and machine learning research community that like thought about things in a certain way. And we're like, this is just a fad, this is a new thing, this can't work, this can't scale. And then because of the exponential, all those things turned out to be false. Then a similar thing happened with kind of like people deploying AI within companies to various applications. Then there was the same thought in the startup ecosystem. And I think now we're at the phase where kind of the world's business leaders, like the investors and the business, they have this whole lexicon of commoditization, you know, which layer is the value going to, which layer is the value going to accrue to, and open source as this idea that you can kind of see everything that's going on, you know, that it has a significance, that it kind of undermines the fact that, you know, the idea that it undermines business. And I actually find as someone who didn't come from that world at all, who never thought in terms of that lexicon, this is one of these situations where not knowing anything often leads you to make better predictions than kind of the people who have their way of thinking about things from the last generation of tech. And you know, this is all, I think, a long-winded way of saying, I don't think open source works the same way in AI that it has worked in other areas. Primarily because with open source, you can see the, you know, you can see the source code of the model. Here, we can't see inside the model. You know, it's often called open weights instead of open source to kind of distinguish that. But a lot of the benefits, which is that many people can work on it, that it's kind of additive, it doesn't quite work in the same way. So, you know, I've actually always seen it as a red herring. When I see it, when I see a new model come out, I don't care whether it's open source or not. Like if we talk about deep seek, I don't think it mattered that deep seek is open source. I think I ask, is it a good model? Right. Is it better than us at, you know, the things that, that's the only thing that I care about. It actually, it actually doesn't, doesn't matter either way. Because ultimately you have to, you have to host it on the cloud. The people who host it on the cloud do inference. These are big models, they're hard to do inference on. And conversely, many of the things that you can do when you see the weights, we're increasingly offering on clouds where you can fine tune the model. You can, you know, we're even looking at ways to kind of investigate the activations of the model as part of like an interpretability interface. We did some little things around steering last time. So I think it's the wrong access to think in terms of, when I think about competition, I think about like, which models are good at the tasks that we do. I think open source is actually a red hearing. But if it's free and cheap to run, it's not free, you have to, you have to, you have to run it on inference and someone, someone has to make it fast on inference. All right. So I want to learn a little bit more about Dario, the person. Yes. So we have a little bit of time left. So I have some questions for you about early life and then how you became who you are. Yes. So what was it like growing up in San Francisco? Yeah. I, you know, the city, when I first grew up here and had not really, had not really gentrified that much, you know, when I grew up the tech boom hadn't, hadn't happened, hadn't happened yet. You know, it happened as, as I was going through high school. And actually I had no interest in it. It was totally, it was totally boring to me. You know, I was interested in being like a scientist. I was interested in physics and math. And, you know, the idea of like, you know, you know, like writing some website actually had no interest to me, to me whatsoever, like founding a company, like those weren't things that I was interested in at all. You know, I was interested in discovering fundamental scientific truth. And I was interested in like, you know, how can I, how can I do something that like makes the world better? So, you know, that was, that was kind of more. And, you know, I watched the tech boom happen around me, but I, I feel like, you know, there were all kinds of things I probably could have learned from it that would have been helpful now. But I just actually wasn't paying attention and had no interest in it, even though I was like, right at the center of it. So you're the son of a Jewish mother, a Italian father. That is true from where I'm from in Long Island. We call that a pizza bagel. A pizza bagel. I've never, I've never heard that term before. So what was your relationship with your parents like? Yeah, I mean, you know, I was, I was always, I was always, I was always pretty close with them. You know, I feel like they gave me a sense of, you know, of, of kind of bright and wrong and what was important in the world. I feel like, you know, kind of imbuing a strong sense of responsibility is, is maybe the thing that I remember most. You know, they were always people who felt that sense of responsibility and, you know, wanted to, wanted to make the world, wanted to make the world better. And I feel like, you know, that's one of the, one of the main things that I, that I learned from them, you know, is always a very, a very loving family, a very caring family. I was very close with my sister, Daniella, who of course, became my, became my, became my co-founder. And, you know, I think we decided very early that we wanted to work together in some, in some capacity. I don't know if we imagine that it would happen, you know, at quite the scale that, that, that, that it has happened. But it, you know, I think, I think it really, you know, that was, that was something we kind of decided early that we wanted to do. The people that I was spoken with that have known you through the years have told me that your father's illness had a big impact on you. Can you share a little bit about that? Yes, yes. He was, yeah, you know, he was ill for a long time. And eventually died in, eventually died in, in 2006. So that, you know, that was actually one of the things that drove me to, you know, I don't think we mentioned it yet in this interview, but before, you know, before I went into AI, you know, I went into biology. So, you know, I'd gone to, you know, I'd shown up at, I'd shown up at Princeton wanting to be a theoretical physicist. And, you know, I did some, did some work in, in cosmology for the first few, a few months of my time there. You know, and, and, you know, that was, that was around the time that my father died. And, you know, that did have an influence on me and kind of was one of the things that convinced me, you know, to, to go into biology, you know, to try and address, you know, human illnesses and biological problems. And so I started talking to some of the folks who worked on biophysics and computational neuroscience in the department that I was at at Princeton. And that was what led to the switch to biology and computational neuroscience. And then, you know, of course, after that, I eventually, I eventually went into AI. And the reason I went into AI was actually a continuation of that motivation, which is that, you know, as I spent many years in biology, I realized that the complexity of the underlying problems in biology felt like it was beyond human scale, you know, in order to understand it all, you needed hundreds, thousands of, you know, human researchers. And, you know, they often had a hard time collaborating or sharing their, you know, combining their, their knowledge. And AI, which was, I was just starting to see the discoveries in it, felt to me like the only technology that could kind of bridge that gap could bring us beyond human scale to, you know, to fully understand and solve the problems of biology. So, yeah, there is a through line there. Right. And I could have this wrong. But one thing I heard was that his illness was largely uncurable when he had it. Yes. And there have been advances that have been, can you share a little bit more? Yes. There are advances that have made it much more manageable today. Yes. Yes. That is, that is, that is, that is true. Actually, actually only in the, maybe three or four years after he died, the cure rate for the disease that he had went from, went from 50% to, to roughly 95%. Yeah. I mean, it has to have felt so unjust to have your father taken away by something that could have been cured. It, of course, of course. But it also tells you of the urgency of solving the relevant problems, right? That, that, that, that, you know, there, there was someone who worked on the cure to this disease that, you know, managed to cure it and save a bunch of people's lives. But, you know, could have, could have, say, even more people's lives if, if, you know, they had managed to find that, that, to find that cure, you know, a few years earlier than they did. And I think that's, that's one of the tensions here, right? That, you know, I think AI has all of these benefits. And, you know, I want everyone to get those benefits as soon as possible, you know, I probably understand, you know, better than almost anyone, how urgent those benefits are. And so I really understand the stakes. When, you know, when I speak out about AI has these risks and I'm worried about these risks, I get very angry when people call me a dumber. I got really angry when, you know, when, when someone's like this guy's a dumber, he wants to slow things down. You heard what I just said, like, you know, my, my father died because of, you know, cures that, you know, could have, could have happened a few years later. When I sat down to, to, to write machines of loving grace, you know, I wrote out all the ways that billions of people's lives could be better with this technology. Some of these people, some of these people who on Twitter, you know, cheer for acceleration, I don't think they have a humanistic sense of the benefit of the technology. They're, their brains just full of adrenaline. And, and they're like, they want a cheer for something. They want to accelerate. I don't get the sense they care. And so when these people call me a dumber, I think, I think they just completely, completely lack any moral credibility and doing that. You know, what really makes me lose respect for them. And I've been wondering what this, this word impact has been because it's come up so often that those who have been around you have said you've been singly obsessed with having impact. In fact, I spoke to someone who knew you well who said you wouldn't watch Game of Thrones because it wasn't tied to an impact that it was a waste of time. And you wanted to be focused on impact. Actually, that's not quite right. I wouldn't watch it because it was so negative some. People were playing, playing such negative, it was like these people start off. And they're partly the situation and partly because they're just horrible people. They like create the situation where at the end of it, everyone is like worse off than everyone was before. I'm really, I'm really excited about like creating positive some situations. I recommend you watch it. It's a great show. But I have watched some parts of it. I was just very reluctant and didn't watch it for a long time. Let's get back to the impact. OK, let's get back to the impact. So that's what impact is. It's effectively your career has been this quest to have that impact to be able to tell me if I'm going too far to prevent other people from being in similar situations. I think that's a piece of it. I mean, I have looked at many attempts to help people. And some of them are more effective than others. And I think I've always tried to-- there should be strategy behind it. There should be brains behind trying to help people, which often means that there's a long path to it. It can run through a company and many activities that are technical and not immediately tied to the kind of impact that you're trying to have. But I'm always trying to bend the arc towards that. I think that's my-- That's my picture of it. That's really why I got into this. I think similar to the reason to get into AI was that I saw the problems of biology as almost intractable without it or at least too slow moving. I think my reason to start a company was that I had worked at other companies. And I just didn't feel like the way those companies were run was really oriented towards trying to have that impact. There was a story around it that was often used for recruiting. But it became clear to me over the years that story was not sincere. I'm going to circle around a little bit because it's clear that you're referring to OpenAI here. From what I understand, you had 50% of OpenAI's compute. I mean, you ran the GPT3 project. So if anyone was going to be focused on impact and safety, wouldn't have been you. Yes, there was a period during which that was true. That wasn't true the entire time. That was, for example, when we were scaling up GPT3. Yeah, so when I was at OpenAI, I and a lot of my colleagues, including the people who eventually founded and thropic-- The pandas. That's the name you gave them. That isn't the name I gave them. The name they took. That isn't a name they took. That's the name other people called them. I-- maybe it's a name other people called them. That's not a name I ever used for my team. OK, sorry. Go ahead. That's good clarification. Thank you. So yeah, we were involved in scaling up these models. Actually, the original reason for building GPT2 and GPT3, it was an outgrowth of the kind of AI alignment work that we were doing, right? Where myself and Paul Cristiano and some of the anthropic co-founders had invented this technique called RL from human feedback. And that was designed to help steer models in a direction to follow human intent. It was actually a precursor to-- we were trying to scale up another method to all scalable supervision, which I think is just starting to work many years later to help models follow more kind of scalable human intent. But what we found is even with the more primitive technique, RL from human feedback, it wasn't working with the small language models with GPT1 that we applied it to and that had been built by other people at OpenAI. And so the scaling up of GPT2 and GPT3 was done in order to kind of study these techniques in order to apply RL from human feedback at scale. This goes to one thing, which is that I think in this field, the alignment of AI systems and the capability of AI systems is intertwined in this way that always ends up being kind of more tied and more intertwined than we think. Actually, what this made me realize is that it's very hard to work on the safety of AI systems and the capability of AI systems separately. It's very hard to work on one and not the other. I actually think the value and the way to inflect the field in a more positive way comes from organizational level decisions, when to release things, when to study things internally, what kind of work to do on systems. And that was one of the things that motivated me and some of the other 2B andthropic founders to kind of go off and do it our own way. But again, if you think capability's in safety are interlinked and you were the guy driving the cutting edge models within OpenAI, if you left, you knew they were going to be a company that was still doing this stuff. That's right. It seems like if you're driving capabilities, you'd be the one in the driver's seat to help it be safe the way that you wanted to. Again, I will say if there's a decision on releasing a model, if there's a decision on the governance of the company, if there's a decision on how the personnel of the company works, how the company represents itself externally, the decisions that the company makes with respect to deployment, the claims it makes about how it operates with respect to society. Many of those things are not things that you control just by training the model. And I think trust is really important. I think the leaders of a company, they have to be trustworthy people. They have to be people whose motivations are sincere. No matter how much you're driving the forward, the company technically, if you're working for someone whose motivations are not sincere, who's not an honest person, who does not truly want to make the world better, it's not going to work. You're just contributing to something bad. So that I'm sure you've heard the criticism from people like Jensen who say, well, Dario thinks he's the only one who can build this safely. And therefore, speaking of that word, control wants to control the entire industry. I've never said anything like that. By the way, I'm sorry if I got Jensen's words wrong, but-- No, no, no, the words were correct. OK. But what are you saying? But the words are outrageous. In fact, I've said multiple times, and I think Anthropics Actions have shown it, that we're aiming for something we call a race to the top. I've said this on podcasts over the years, and I think Anthropics Actions have shown it, where-- with the race to the bottom, everyone is competing to get things out as fast as possible. And so I say, when you have a race to the bottom, it doesn't matter who wins. Everyone loses, because you make the unsafe system that helps your adversary or causes economic problems or is unsafe from an alignment perspective. The way I think about the race to the top is that it doesn't matter who wins. Everyone wins, right? So the way the race to the top works is you set an example for how the field works. You say, we're going to engage in this practice. So a key example of this is responsible scaling policies. We were the first to put out a responsible scaling policy. And we didn't say everyone else should do this, or you're bad guys. We didn't try to use it as advantage. We put it out, and then we encouraged everyone else to do it. And then we discovered in the months after that that there were people within the other companies who were trying to put out responsible scaling policies. But the fact that we had done it allowed, gave those people permission, kind of enabled those people to make the argument to leadership, hey, Anthropic is doing this, so we should do it as well. The same has been true of investing in interpretability. We release our interpretability research to everyone. And allow other companies to copy it, even though we've seen that it sometimes has commercial advantages. Same with things like constitutional AI, same with the measurement of the measurement of the dangers of our system, dangerous capabilities evils. So we're trying to set an example for the field. But there's an interplay where it helps to be a powerful commercial competitor. I've said nothing that anywhere near resembles the idea that this company should be the only one to build the technology. I don't know what anyone could ever derive that from anything that I've said. It's just an incredible and bad faith distortion. All right, let's see if we can lighten around one or two before I ask you the last one, which we'll have five minutes for. What happened with SBF? I mean, he was one of those-- Go ahead. I couldn't tell you. I couldn't-- What was the-- what are you answering? I probably met the guy four or five times. So I have no great insight into the psychology of SBF or why he did things as stupid or immoral as he did. I think the only thing I had ever seen ahead of time with SBF was a couple people mentioned to me that he was hard to work with that he was a bit of a move fast and break things guy. And I was like, OK, there's plenty of people. Welcome to Silicon Valley. Yeah, welcome to Silicon Valley. And so I remember saying, OK, I'm going to give this guy non-voting shares. I'm not going to put him on the board. He sounds like a bad person to deal with every day. But he's excited about AI. He's excited about AI safety. He's a bull on AI. And he's interested in AI safety. So it seems like a sensible thing to do. In retrospect, that move fast and break things was turned out to be much, much, much more extreme and bad than I ever imagined. OK, so let's end here. So you found your impact. I mean, you're working the dream pretty much right now. I mean, think about all the ways that AI can be used for biology just to start. You also say that this is a dangerous technology. And I'm curious if your desire for impact could be pushing you to accelerate this technology while potentially devaluing the possibility that controlling it might not be feasible. So I think I have more than anyone else in the industry warned about the dangers of technology. We just spent 10, 20 minutes talking about the frightening, the large array of people who run trillion dollar companies criticizing me for talking about the dangers of these technologies. I have US government officials. I have people who run $4 trillion companies criticizing me for talking about the dangers of the technology. Computing all these bizarre motives that bear no relationship to anything I've ever said, not supported, and anything I've ever done. And yet I'm going to continue to do it. I actually think that as the revenues, as the economic business of AI ramps up and it's ramping up exponentially, if I'm right, in a couple years, it'll be the biggest source of revenue in the world. It'll be the biggest industry in the world. And people who run companies already think it. So we actually have this terrifying situation where hundreds of billions to trillions to-- I would say maybe 20 trillion of capital is on the side of accelerate AI as fast as possible. We have this company that's very valuable in absolute terms, but looks very small compared to that, right? $60 billion. And I keep speaking up, even if it makes folks in the US government are upset at us, for example, for opposing the moratorium on AI regulation, for being in favor of export controls for chips on China, for talking about the economic impacts of AI. Every time I do that, I get attacked by many of my peers. Right, but you're still assuming that we can control it. That's what I'm pointing out. But I'm just telling you how much effort, how much persistence, how much despite everything that stack up, despite all the dangers, despite the risk that it has to the company of being willing to speak up, I'm willing to do it. And that's why I'm saying that, look, if I thought that there was no way to control the technology, right? If I thought, even if I thought, this is just a gamble, right? Some people are like, oh, you think there's a 5% or 10% chance that AI could go wrong? You're just rolling the dice. That's not the way I think about it. This is a multi-step game, right? You take one step, you build the next step of most powerful models, you have a more intensive testing regime. As we get closer and closer to the more powerful models, I'm speaking up more and more. And I'm taking more and more drastic actions because I'm concerned that the risks of AI are getting closer and closer. We're working to address them. We've made a certain amount of progress. But when I worry that the progress that we made on the risks is not fully aligned with the-- is not going as fast as we need to go for the speed of the technology. Then I speak up louder. And so you're asking, why am I-- you started this interview by saying, what's gotten into you? Why are you talking about this? It's because the exponential is getting to the point that I worry that we may have a situation that our ability to handle the risks is not keeping up with the speed of the technology. And that's how I'm responding to it. If I believe that there was no way to control the technology, which I see absolutely no evidence for that proposition, we've gotten better at controlling models with every model that we release, right? All these things go wrong, but you really have to stress test the models pretty hard. That doesn't mean you can't have emergent, bad behavior. And I think if we got to much more powerful models with only the alignment techniques we have now, then I'd be very concerned. Then I'd be out there saying, everyone should stop building these things. Even China should stop building these. I don't think they've listened to me, which is one reason I think export controls is a better measure. But if we got a few years ahead in models and had only the alignment and steering techniques we had today, then I would definitely be advocating for us to slow down a lot. The reason I'm warning about the risk is so that we don't have to slow down, so that we can invest in safety techniques and can continue the progress of the field. It would be a huge economic effort. Even if one company was willing to slow down the technology, that doesn't stop all the other companies. That doesn't stop our geopolitical adversaries to whom this is an existential fight for survival. So there's very little latitude here. We're stuck between all the benefits of the technology, the race to accelerate it, and the fact that that is a multi-party race. And so I am doing the best thing I can do, which is to invest in safety technology, to speed up the progress of safety. I've written essays on the importance of interpretability, on how important various directions in safety are. We release all of our safety work openly because we think that's the thing that's a public good. That's the thing that everyone needs to share. So if you have a better strategy for balancing the benefits, the inevitability of the technology, and the risks that it face, I am very open to hear it, because I go to sleep every night thinking about it, because I have such an incredible understanding of the stakes in terms of the benefits, in terms of what it can do, the lives that it can save. I've seen that personally. I also have seen the risks personally. We've already seen things go wrong with the models. We have an example of that with Brock, and people dismiss this, but they're not going to laugh anymore when the models are taking actions, when they're manufacturing, and when they're in charge of medical interventions. People can laugh at the risks when the models are just talking. But I think it's very serious. And so I think what this situation demands is a very serious understanding of both the risks and the benefits. These are high stakes decisions. They need to be made with a seriousness. And I think something that makes me very concerned is that on one hand, we have a cadre of people who are just doomers. People call me a doomer, I'm not. But there are doomers out there. People who say they know there's no way to build this safely. I've looked at their arguments. They're a bunch of gobbledygook. The idea that these models have dangers associated with them, including dangers to humanity as a whole, that makes sense to me. The idea that we can logically prove that there's no way to make them safe, that seems like nonsense to me. So I think that is an intellectually morally unserious way to respond to the situation we're in. I also think it is intellectually and morally unserious for people who are sitting on $20 trillion of capital, who all work together because their incentives are all in the same way. They're dollar signs in all of their eyes to sit there and say, we shouldn't regulate this technology for 10 years. Anyone who says that we should worry about the safety of these models is someone who just wants to control the technology themselves. That's an outrageous claim. And it's a morally unserious claim. We've sat here and we've done every possible piece of research. We speak up when we believe it's appropriate to do so. We've tried to back up when we make claims about the economic impact of AI. We have an economic research council. We have a economic index that we use to track the model in real time, and we're giving grants for people to understand the economic impact, the economic impact of the technology. I think for people who are far more financially invested in the success of the technology, then I am to just, you know, easily lob ad hominem attacks. I think that is just as intellectually and morally unserious as the Doomer's position. I think what we need here is we need more thoughtfulness. We need more honesty. We need more people willing to go against their interest, willing to not have, you know, breezy Twitter fights, the hot takes. We need people to actually invest in understanding the situation, actually do the work, actually put out the research, and actually add some light and some insight to the situation that we're in. I am trying to do that. I don't think I'm doing that perfectly as no human can. I am trying to do it as well as I can. It would be very helpful if there were others who would try to do the same thing. Well, Dario, I said this off camera, but I want to make sure to say it on as we're wrapping up. I appreciate how much anthropic publishes. We have learned a ton from the experiments, everything from red teaming the models to vending machine clawed, which we didn't have a chance to speak about today. But I think the world is better off just to hear everything going on here. And to that note, thank you for sitting down with me and spending so much time together. Thanks for having me. Thanks everybody for listening and watching. And we'll see you next time on Big Technology Podcast. [MUSIC PLAYING]

Podcast Summary

Key Points:

  1. Dario Amodei, CEO of Anthropic, expresses urgency about AI risks and capabilities, believing exponential progress is underestimated.
  2. He criticizes vague terms like AGI and superintelligence, focusing instead on observable scaling laws in compute, data, and training methods.
  3. Amodei defends Anthropic's competitive position, emphasizing talent density, mission alignment, and sufficient resources despite larger rivals.
  4. He addresses criticisms, including claims of wanting industry control, and highlights AI's potential benefits while stressing the duty to warn about downsides.
  5. The discussion covers business growth, model advancements (especially in coding), and responses to competitors' recruitment and scaling efforts.

Summary:

In this interview, Anthropic CEO Dario Amodei discusses the rapid advancement of AI, emphasizing an exponential growth trajectory in capabilities due to scaling laws in compute, data, and training techniques. He argues that societal and economic impacts are nearer than many realize, prompting his urgent public warnings about risks, though he also acknowledges AI's positive potential. Amodei dismisses terms like AGI as marketing, instead focusing on tangible progress, such as Anthropic's models improving significantly in coding and other tasks.

He addresses criticisms, including accusations of wanting to control the industry, and defends Anthropic's stance against larger competitors like Meta and XAI, highlighting the company's reliance on talent density, mission alignment, and principled compensation over bidding wars for employees. Despite raising substantial funds, he asserts Anthropic remains competitive in resources and infrastructure, confident in its approach to navigating AI's future challenges and opportunities.

FAQs

He believes AI capabilities are improving exponentially and that many people underestimate how quickly transformative changes could occur, though he acknowledges uncertainty in exact timelines.

He considers these terms meaningless and primarily marketing-driven, preferring to focus on concrete exponential improvements in AI models rather than vague concepts.

Anthropic maintains its compensation principles and fairness, refusing to negotiate individually, believing that mission alignment and company culture are more important than matching high offers.

He argues that, based on Anthropic's models, scaling continues to yield significant improvements without diminishing returns, citing rapid advances in areas like coding.

He acknowledges it as a challenge but believes even without solving it, models can be highly impactful, and that techniques like expanding context windows and reinforcement learning can mitigate many gaps.

He emphasizes a duty to warn about potential downsides and risks, even while appreciating and articulating the positive applications of AI more strongly than many optimists.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.