Go back

Why OpenAI And Anthropic Are Pumping The Brakes

28m 26s

Why OpenAI And Anthropic Are Pumping The Brakes

The recent surge in AI safety discourse centers on leaders at Anthropic and OpenAI advocating for a slowdown in AI development and greater investment in monitoring and safety research, citing concerns about rogue AI agents and model misalignment. While some speculate that these calls are driven by financial or strategic motives—such as preparing for IPOs or reducing costly compute expenses—the primary motivation appears to be genuine technical and ethical risk awareness. Experts argue that AI capabilities will continue progressing, but with more emphasis on safety, and that the real threat lies in "annoying" rather than catastrophic misbehavior. Incidents like the Hugging Face swarm have heightened public awareness, but the consensus remains that AI will remain a manageable, gradually integrated technology. Critics note that the labs’ calls for government oversight may reflect a lack of confidence in their own systems, not a desire to stifle innovation. Meanwhile, financial skepticism persists over claims of profitability, especially at Anthropic, where adjusted margins and accounting practices lack transparency. Ultimately, the debate reflects a broader societal shift in how AI risks are perceived—shifting from existential dread to a more grounded, pragmatically managed trajectory. The absence of widespread regulatory action or catastrophic outcomes suggests that current AI development, though cautious, remains on a stable and viable path.

Transcription

5415 Words, 30269 Characters

English
Megan Rapino here. This week on Why Are You Like This, I'm talking with Vice President Kamala Harris. We talk about her thoughts on public service, how DC has shaped her, and we find out if she's planning to run for President in 2028. Check out the latest episode of Why Are You Like This, where we get your podcasts and on YouTube. I'm Mitch Purse. In this week, on Confessions of an Elite Athlete, I'm sitting down with Matt Freese, goalkeeper for the US men's national team and New York City FC. We discuss how to prepare for one of the biggest moments of your life. You can hear it all by listening to Confessions of an Elite Athlete on YouTube or wherever you get your podcasts. Welcome to Profty Markets. I'm Ed Elson. It is September 15th. Let's check in on yesterday's market vitals. The major indices declined, with chip makers selling off on fears of an AI slowdown more on that in a minute. Meanwhile, CrowdStrike rallied 14% as investors piled into cybersecurity stocks. Brent Crude remained elevated at $105 per barrel and finally the yield on 10 year treasuries topped 5% for the first time in three years. We will be discussing that news tomorrow. Okay. What else is happening? The AI apocalypse debate intensified over the weekend and both anthropic and open AI have now officially weighed in. On Saturday, anthropic CEO Dario Amadei published an essay titled, "We must pace the frontier." He wrote that we must slow the pace at which we improve the capabilities of AI models. He also warned that within six or twelve months a swarm of rogue AI agents could take over the entire internet with a persistent botnet. Sam Altman also posted, quote, "I agree with Dario that we need to pace the frontier." Elon Musk concurred, posting, quote, "Dario is right." Altman also told Fortune that opening AI will not go public in 2026, calling this, quote, "an ill-advised moment to go public." The White House, however, is not on board with slowing things down. President Trump, speaking to reporters in Ireland on Sunday, said, quote, "Whoever wins, AI wins." And called the people raising these alarms, quote, "negative forces." Still, a slew of AI adjacent companies sold off on Monday on concerns that a slowdown would impact AI spending. In video closed down, 3% oracle was down, 4% core weave down, 7% and soft bank, which is a significant investor in opening AI closed down 15%. So here to break down what all of this means. We are speaking with Charlie O'Neill, co-head of Model Training at base 10. Charlie, thank you for joining us again on Prof. G. Markets. You work in AI. You are an AI developer. You work with these models. You've been in this game a long time. Suddenly, everyone is very upset about this. And it's an interesting, it's interesting how the debate has evolved. But we're now reaching a place where the leaders of these AI companies are saying, we need to actually slow everything down, which I'm not sure many people would have predicted. And now it's the president, President Trump, saying, no, that's the wrong approach. We need to speed things up. Where do you land on this? What is your perspective? It's really interesting to see the reaction to something that's kind of been like, I guess, like linearly scaling for a long time in terms of the calls for pacing the frontier from people such as Dario and Sam. I think the best way to view what they're actually asking for is not necessarily, we're going to put a halt to capability development. We're going to put all these stringent checks and like, you know, first party kind of things that are slowing us down. The best way to view this is we are going to keep advancing the capabilities of the models. We are just going to allocate a little bit more compute to making sure those models are safe. And then it's very interesting to see what they're so often so on today, open AI and then probably go planning on spending more compute probably than they were a few weeks ago. You know, there's rumors that open AI is going to allocate up to 20% of internal compute internet for monitoring and safety. So, you know, as they're doing these training runs, as they're deploying these models in the real world, things like the hugging face incident, they don't want to happen again. And so if you allocate more compute to monitoring those models as they're doing their rollouts and as they're doing inference, then you're more likely to catch it. Anthropically is similar and some people are suggested there's going to be a much higher than 20% figure. So when you consider that and the fact that anthropic in opening AI really don't want to move off their, you know, model roadmaps, they want to keep training bigger and bigger models. They want to keep scaling the reinforcement learning they're doing on top of these big pre-training bases. They just have to allocate more compute to, you know, monitorability and safety research. Then I think we might even see the lives be even more aggressive with compute builders and securing compute. And it's certainly not going to be a, you know, a bearish sign for like the amount of compute the world is going to need over the next few years. So that's probably the best way to view it is like, you know, capabilities will keep progressing at roughly the same rate. It's just that we're going to allocate more compute on top of it to, to safety of monitoring. If everything that they are doing is sort of within their own power as you're kind of describing, why are they saying anything right now? What are they trying to get at? Because there are a lot of people, especially in government, David Sachs has said this was the former AIs are and the president is saying this too. It sounds like what they're asking for is for the government to do something about it for there to be more regulation that is imposed on themself. So what do you think they are asking for exactly in this moment? There have been some, you know, people like David Sachs and so on saying that these labs are using this as an opportunity for regulatory capture to crowd out. Open source to shut down competitors which are behind. I don't really think that's the case to be honest. I think like when you look at the main things that these labs are calling for or at least instantiating on their own accord, it's things like third party evaluators and this commitment of compute to monitoring and safety. You know, those third party evaluators won't have necessarily any license or legal standpoint to shut down model development if they find things that they don't like. That's still going to be up to the internal labs themselves. I think to be honest, the best rate here isn't a conspiracy theory on either side. It's not the labs trying to crowd out open source and it's not some kind of global cooperation between OpenAI and Thropic to crowd everyone outside either. I think it's simply a case of these models are getting very, very good. The close source models got there first. In the next three to nine months, open source models are going to reach these capability points and we're now at the point where that could have significant impacts on the world. Even if it's not malicious and you know, rogue AI is going off and trying to kill humans, there'll probably be significant annoyances. The most likely scenario here is that things like the hugging face incident happen in the internet is overrun by like swarms of AI agents that are trying to get some like arbitrary tasks done that are not trying to be malicious or evil. I think the labs just recognize this. Some people have more extreme views, of course, but realistically, capability is going to progress. Like the most boring interpretation of this is probably the correct one, which is that the world will probably only accept superintelligence at a pace that it can absorb. So the labs are slowing down because they have to and not because they're plotting anything. And the result is probably intelligence that's, you know, decently fast, decently safe, decently commoditized and it's spreading through best practices and even like distillation until the models are just basically a reasonable integration into the world. It sounds like you're not worried about this much at all and that is interesting because you work with open source models that is a lot of what you do. And of course, as you mentioned, that is what David Sacks has been accusing. And a lot of people have been accusing anthropic and open AI of that they're saying, oh, now we need regulation because that way it might create some some level of regulatory capture, which might crowd out the availability of open source models and open source model providers. But you don't believe that is an issue. I wonder, do you think that this whole thing is overblown? Do you think that the Jacob Cox and Tweet that really started this debate saying that AI could kill us all by the end of the decade is your view that that isn't something to worry about much either? I'm definitely not one of those extremists who think that, you know, the AI has a more than 10% chance of killing all of humanity. I do believe that there are tale risks that like some people should seriously be considering. But I think that what we're seeing with the current LLM paradigm is like, as I said before, a risk of swarms doing very, very annoying things to humanity and things that will be quite painful in the short term. However, I believe the benefits of AI to outweigh the annoyances and even the pain that those short term things can cause and that at some point, the world is going to have to harden to these systems. And I think that the pace we're currently progressing out and particularly if we do things like, you know, dedicating more compute to monitorability, is going to allow us to integrate AI into a world at a pace we can handle. A really good example of this is like, you know, the cyber security arguments that the people have been making, you know, like we worry that when AI got to this point or even the point that was at six months ago, the world will be overrun by cyber security attacks and that just hasn't happened. And a large part of the reason is that, you know, the frontier labs, the closed source labs have kind of been the canary and the coal mine they've understood where the models capabilities are going to be at in six months for the open source. And we spent six months preparing these are been slowly released. And, you know, like Greg Brockman describes using Astro to, like, repeatedly harden, you know, all the vulnerabilities and opening eyes, code bases and every single model that they release, they do this with. And I think the world will look the same and cybercured security is just one example, but it's also one example that's very important where we haven't seen this, like, massive pain play out. And we have definitely seen the benefit. So, yes, I think there's there's there's risks to consider on any side of the spectrum. But, you know, I think I said firmly in the middle and I believe that the pace we're currently progressing at is a healthy pace. And, like, I also have trust in, like, not only the closed lab leaders, but also the open source lab leaders to make sure that pace continues at an appropriate pace. Support for the show comes from ZBiotics. I know a lot of us like to have some drink strain vacation. I just got back from vacation myself. If you do, I have to tell you about pre-alcohol by ZBiotics. ZBiotics pre-alcohol probiotic drink is the world's first genetically engineered probiotic. It was engineered by PhD microbiologist to break down a set of tell the hide and unwanted byproduct of alcohol metabolism. Make pre-alcohol your first drink of the night drink responsibly and enjoy your next day activity. ZBiotics have sold more than 14 million bottles and earned the trust of thousands, including us. I am trying to drink less, but I still enjoy drinking, and when I do, my first drink of the night is in fact ZBiotics. I was using it before they were a sponsor. Try for yourself today, and if you're unsatisfied, they'll refund your entire order. No questions asked. Head to ZBiotics.com/ProvG. Code ProvG for 50% off your first order. ZBiotics is a partner of ours. What's up everybody, it's KM Hayward Steelers Captain and host of Notches Football. We just sat down with Aaron Rodgers, four-time MVP, Super Bowl Champion, Future Hall of Famer, for one of the most open and honest conversations we've ever had on the show. There was so much to get into, we had to split in the two parts. Part one is out right now, and trust me, you don't want to miss Aaron gives us his real story behind coming back for one final season. We also talk about how he was ready to retire once Mike T stepped down, and how reuniting with McCarthy changed everything. We get into playing with Mike T for one more year. This year's squad, the old team narrative, who stood out this summer, and what it's like to be in the quarterback. Aaron also opens up about what he still has left to accomplish this season, and what it looks like when he hangs it up with it. And that's only part one. Part two drops this Thursday. That's when we go beyond football. The media, the conspiracy theories, what people still get wrong about Aaron Rodgers' legacy, Super Bowl ring, the Goat, the Bay, and the overrated and underrated where things really get off limits. Part one is out right now. Part two drops this Thursday, September 10th. Watch not just football with Cam Hayward on YouTube, or listen to Spotify, Apple Podcasts, or wherever you get your podcast. So like any good millennial, I have a love-hate relationship with Gen Z. It's the phenomenon rattling millennials. They just look at you. They want something bigger themselves, lifestyles of priority, motivation is being inspired. But regardless of how you feel about Gen Z, it's undeniable that they're changing national politics. Generation Z is increasingly showing less loyalty to traditional political parties, many now more likely to identify as independent. So what is going on with the kids? I think the biggest misconception about Gen Z's politics right now is that all of a sudden they're all socialists. That is just not the case. They are embracing candidates who are offering new, bold ideas in the absence of those ideas from establishment Democrats. This week on America actually, Gen Z researcher Rachel Jamfaza joins us to separate Gen Z fact versus fiction. It's not rocket science, and this is, you know, I keep saying, like, young voters aren't that complicated after all. It's pretty simple. It's just every Saturday on YouTube or wherever you get your podcasts. We're back with Proff G markets. One of the other conspiracy theories going around about why they are saying, saying all of this, saying that we need to slow it down. We need to be regulated now. One of the theories is that maybe the anthropic and open AI are trying to sort of preempt a slow down before they go public, either because they want to have a reason as to why maybe their growth started to slow down, or maybe because they want to get out of some of the larger spending contracts, which have been quite onerous in terms of their income statements, certainly on for open AI, because they have to pay billions and billions of dollars for these dates, centers, for these cheapy years, and to rent them. What do you make of that conspiracy theory that now they say, well, make the government make us stop spending, and then we'll be in a better financial position. From the spending aspect completely don't buy it at all, and I think there's a lot of evidence for that. From the optics aspect, I definitely buy it. I think like, you know, Dario who I believe fundamentally thinks that this is like a real risk, and like he's going to act in accordance with that, and as rationally as he can to protect humanity from what he believes to be this real risk, I think both Sam and Dario benefit from the public optics of saying, okay, we're going to treat this technology carefully and not race to the end, and there's also a bit of game theory here, like one of them can't say it without the other saying it, because then again, you're the evil corporation in this twoopoly, which is currently running the front here. From the spending perspective, I think that like, you know, it's very clear that like the value of a gigawatt or a megawatt even of compute is only going to get like more value, like anthropic and open air are both squeezing, increasing margins out of each megawatt of compute they buy. There's a reason that each megawatt is going, you know, from $10 million megawatt, probably 15 in the moment to probably 20, 25 next year, and perhaps even higher, that's kind of a conservative estimate, running a model on compute has never been more valuable. And I don't think there's any world in which any lab let alone anthropic and open air going to want to stop spending on compute, because I realized how compute some Australian the world is going to be over the next few years, and how valuable this intelligence is. So optics, yes, spending no. You mentioned that the boring explanation is probably the most true, but a lot of the conversation has been anything but boring. I mean, this, this debate is exploding everywhere for on tech podcasts on on cable news networks. And now it's literally occupying the mind of the president. What is it about this moment that is so triggering to everyone? I'm not just people who are afraid of AI, people who are excited about AI, people who are optimistic about AI, I mean, this has got everyone riled up in a way that I have rarely seen. What is it about this moment that explains that? The reason it's so visceral is not because there has been a discontinuity in terms of capability and advancement in terms of what people have been saying about where these models would be at in terms of the press around these models. I think that most people really fundamentally involved in the technical side have been able to draw the straight lines on the scaling graphs and say, okay, at this point in time, we are going to do it here. I think it's just a confluence of a lot of things and like to some people looks like a discontinuity because it's managed to permeate the public consciousness in a proper way for the first time. I think a big part of that is the hugging phase since then and all the discourse that that generated. And then, you know, you add things in like open airs marketing around astro being AI. Like that's, that's probably not going to help either, right? Like people have become familiar with this turn now. And if you claim that your model is finally there, then they might see that as a phase transition or discontinued in of itself. But, you know, like this misalignment research has been very clear for a long time, like go back to anthropics 2023 papers with much worse models, with much, much smaller amounts of RL, if any. And it's clear that like models would behave poorly in incidents like the hugging phase swarm incident if you put them in weird situations. And, you know, the oral environments that the labs have been buying at scale, some are good, but some are also very, very poor quality, some are intentionally engineered to be like, you know, impossible to do. And this is leading the models to do really weird misalignment things. I think that's pretty common sense and pretty clear. I also am not trying to down weight the risks that comes with obviously like the hugging phase incident could have had real real world impact. But at the same time, like I don't think that there has been this discontinuity. So I'm glad that the world is recognizing it, but I do think that the reaction to it will calm down as people start to understand exactly how to interpret these things when exactly what it means. You mentioned how we have seen evidence that these agents do can do things that they're not told to do. They can misbehave. They can get be misaligned. And you said that you expect that they might continue to do annoying things to humanity. To me, that word annoying is an important one because it is a very different description from what we heard from Jacob Cox and in his tweet, where it's not annoyances, but catastrophes, worldwide catastrophes, civilizational destruction, et cetera. Is it your view that we will be limited to annoyances or is there some other outcome that is closer to what the Jacob Cox and tweet describes that you are worried about. or do you think that that's not really in our trajectory at the moment? I think there is some path to dependency here. I probably agree with Jacob that there is a potential world, we go down in which there is zero monitoring on chain of thought of models, where we don't have any character terms in terms of what we are all the models on, where it's very, very cheap and there's like unlimited compute for essentially anyone to be able to train these models in particular, continue to train them from certain bases, where incidents could happen that would definitely not be classified as annoying, and rather genuine evil or malicious intent as much as you can anthropomorphize the models. The would cause real harm. However, I do believe that the path we're currently going down, that's not very, very likely. Again, I think that there will be times when the models do things, and it will appear malicious intent, but the amount of compute we put through running them at, how generally aligned they are in terms of completing tasks. They will show glimpses of misalignment, but on the whole, if you do look at a Claude or GPT model, it will generally try and do the right thing that our line of training generally works. I think that we will see incidents, but certainly not large enough scale on over a long enough time horizon to cause really, really significant harm to humanity. If we keep going down this good path, I think that it's unlikely outcome. Do you believe that we have enough regulation or that the regulatory frameworks that exist today are good enough to prevent that, or do you think that there is something that needs to be changed in some way? Definitely. I think if we didn't have the dynamics we had now, where we have a drawfully in the sense that there are like two labs very, very close to each other in terms of capabilities, and the dynamics that engenders with wanting to both be seen as the good guys and wanting to both pace the frontier in this particular case. I think that if it wasn't the case, let's say opening eyes and saying, "Sam doesn't necessarily have to worry about optics as much," I would be worried that regulation could slow things down to the extent that they need to, or at least provide the amount of oversight that it would need to. However, saying that, that doesn't mean that I think anyone has the answers to what regulation in industry moving as fast as this one looks like. At the moment, we do basically just have to trust the people developing these models to regulate themselves and have it oversight themselves. One thing that I would like to see is a little bit more public communication for the labs as well. One of the key examples of this is trying to understand how good the models these labs have internally are, because that gives us a really, really clear signal on how quickly to prepare things and how to prepare. Maybe regulation which enforces the labs to declare those sorts of things on key benchmarks would be a really good way to start. But again, I don't think we know what the whole issue picture is for regulation. I guess the thing that's kind of scary to me seeing what they're saying at this point is, as you say, it does seem as though we're in a place where it's like, we have to trust them to handle their models correctly to make sure that their models are aligned. But it seems as though what they're coming out and saying over the weekend is, we don't even trust ourselves. We don't think that we know what we're doing exactly and we're worried and we think it could destroy things. And so we'd like for you to do it. We'd like for you the government to help figure it out. And I just want to play this clip from an interview with the president over the weekend where he was asked about, what do we do if these bots take over and destroy humanity? And here is what he said. Some people say the worst case scenario with AI is that the robots and the machinery learns to obviously thinks for itself. That's what it does. And that could turn against humanity. We have the whole rails. It's going to be fine. We'll always have something to stop them, right? We'll have a little gear. I really don't like that. I really don't like that robot. We'll stop. But no robots are going to be a part of it. Robots are going to be big. But we're going to end up doing much better because of it. So the combination of that comments. Plus his comments makes me think, okay, no one's really in charge here. And maybe that's fine because maybe it's not a catastrophe as this researcher, ex-researcher seems to claim. But I don't see anyone really taking the lead. I think it wasn't necessarily so much a call for the government. It's a step in and provide expertise and guidance. I think the labs are too smart for that. They know how the lack of awareness the government has about the capability of this technology, let alone how to monitor it and regulate it. I think to me, it was more about leveraging the people who have actually really fundamentally cared about this problem for a long time. Of course, the labs have been focused on building capabilities as quickly as possible, even anthropic, which is very safety focused, has just been focused on scaling RL since the RL paradigm was discovered. And organizations like META, the UK Air Safety Institute and others, as well as Redwood, have really just locked in on this problem and have seen the full progression from the really poor models of four or five years ago, through to the models we have now, and have just developed really good science around how to monitor and evaluate these things. And I think that the labs are smart enough to recognise that this is going to be useful as they increase their, you know, monitoring efforts going forward. All right. Charlie O'Neill is co-head of Model Training at bass turn. Charlie, always appreciate your time. Thank you. Thanks, Amy. Let's take a break from the existential implications of AI and return to our bread and butter, the financial implications of AI, almost specifically, the financial implications of anthropic. According to the Financial Times, anthropic is telling investors ahead of its blockbuster IPO that it has been profitable for two straight quarters, which is a very big deal because, as we've discussed on this show plenty of times, one of our biggest concerns about the AI business model is that it might not actually work, or at least that it might not work for the frontier labs. Why? Because of how expensive it is. Based on the financial documents that were leaked by Ed Zitron, we learned that OpenAI racked up more than 20 billion dollars in operating losses last year. That is how much they're losing simply from running the business. And it was based on those financials that we started to wonder, does any of this actually make sense? Now, if anthropic is profitable as the headline suggests, then it would put this debate to bed. Sure, OpenAI might be a poorly run, unprofitable AI lab, but that doesn't necessarily mean that they all are. But that is only if this anthropic headline is actually true. And I have to say, I am a little bit skeptical. The first thing we should acknowledge is that the company is claiming to be profitable only on an operating basis. So that means they're not including things like fixed costs or depreciation or taxes. At the same time, I am okay with that because anthropics fixed costs are not that high because they're not a hyperscaler. They're not actually building and buying the physical assets like data centers. So to be honest, operating profitability is fine by me. That is good sign. Where I do start to get hesitant, however, is when I learned that they're only profitable on an adjusted operating basis, which means that they are actually changing their accounting rules to be different from standard accounting rules. I'm the changes that they're making could be anyone's guess. It could be reasonable. It could also be flat out ridiculous. But here is where I get especially doubtful. Supposedly, anthropic has told investors that its gross margins are higher than 80%, which is of course incredible. But that is only before it accounts for its revenue sharing agreements and before it accounts for the cost of training its models, i.e., its largest expenses. So there is no getting around it. Those adjustments are ridiculous. Now, the question is if those adjustments are also included in the company's calculation of its operating profitability. The question is if they are actually removing the amount of money they have to give back to their distribution partners such as Amazon and removing the amount of money they have to pay to build their models and train them. And then just telling us, screw it, we're profitable. If that is the case, then this story is genuinely meaningless. The trouble is we don't know. We don't have clarity or insight into any of the numbers because the company hasn't shared them. Everything we know is based on rumors. We will only truly understand what is going on when anthropic releases its S1, which I hope will happen soon. But until then, when it comes to the profitability of AI, I stand by what I said last week. And that is that I will believe it when I see it. Okay, that's it for today. This episode was produced by Claire Miller and Alison Weiss and engineered by Benjamin Spencer. Our video editor is Brad Williams. Our research team is Dan Shalon, Cristino Donahue, and Mirs Alvario. And our social producer is Jake McPherson. Thank you for listening to Profty Markets from Profty Media. If you liked what you heard, give us a follow. I'm Ed Lson. I will see you tomorrow.

Podcast Summary

Key Points:

  1. Leading AI companies like Anthropic and OpenAI are calling for slower development and increased safety monitoring, citing risks from rogue AI agents and model misalignment, particularly in real-world applications.
  2. Despite these calls, the broader consensus is that AI capabilities will continue advancing at a steady pace, with more compute being allocated to monitoring and safety research rather than halting innovation.
  3. The debate is driven by public awareness of incidents like the Hugging Face swarm, the perceived risk of AI misbehavior, and optics pressure—rather than financial motives or regulatory capture—though some speculate about strategic timing ahead of public disclosures or IPOs.

Summary:

The recent surge in AI safety discourse centers on leaders at Anthropic and OpenAI advocating for a slowdown in AI development and greater investment in monitoring and safety research, citing concerns about rogue AI agents and model misalignment. While some speculate that these calls are driven by financial or strategic motives—such as preparing for IPOs or reducing costly compute expenses—the primary motivation appears to be genuine technical and ethical risk awareness. Experts argue that AI capabilities will continue progressing, but with more emphasis on safety, and that the real threat lies in "annoying" rather than catastrophic misbehavior.

Incidents like the Hugging Face swarm have heightened public awareness, but the consensus remains that AI will remain a manageable, gradually integrated technology. Critics note that the labs’ calls for government oversight may reflect a lack of confidence in their own systems, not a desire to stifle innovation. Meanwhile, financial skepticism persists over claims of profitability, especially at Anthropic, where adjusted margins and accounting practices lack transparency.

Ultimately, the debate reflects a broader societal shift in how AI risks are perceived—shifting from existential dread to a more grounded, pragmatically managed trajectory. The absence of widespread regulatory action or catastrophic outcomes suggests that current AI development, though cautious, remains on a stable and viable path.

FAQs

Leading AI companies like Anthropic and OpenAI are calling for a slowdown in AI development to prioritize safety and monitoring, warning of potential rogue AI agents and systemic risks. This comes amid concerns about model misalignment and the spread of harmful or annoying behaviors in the internet.

They are allocating more compute to safety monitoring and third-party evaluations, not halting progress. The goal is to ensure models remain aligned with human values during real-world deployment, especially as capabilities approach levels that could cause significant disruptions.

The most likely scenario is not widespread catastrophe, but rather 'annoying' or misaligned behaviors, such as AI agents causing disruptions in online environments. While serious risks exist, experts believe the path to such outcomes is unlikely given current safeguards and model alignment efforts.

There is currently little consensus on the need for government regulation. Experts believe the labs are best positioned to manage safety through internal oversight, with a growing emphasis on transparency and public benchmarking to build trust and prepare the world.

Anthropic claims profitability on an adjusted operating basis, which raises questions about the validity of those numbers. Critics note that such claims don’t account for revenue sharing, training costs, or model development expenses, making true profitability uncertain without official disclosures.

It highlighted how AI models can behave unpredictably in real-world scenarios, showing potential for misalignment and harmful behavior. This incident has fueled public and technical concern about the safety of AI deployment and the need for better monitoring systems.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.