Wie funktioniert KI-Transformation beim Energieriesen RWE, Dr. Max Schumm? (AI Transformation Lead, RWE)
47m 21s
The discussion centers on Anthropic, the secretive AI firm valued around $350 billion, and its chatbot Claude. Recent reports indicate the Pentagon may sever ties after Anthropic declined certain military uses, while Claude was allegedly involved in a Venezuelan operation—claims the company hasn't confirmed. Simultaneously, civilians use Claude for tasks like disputing hospital bills or writing novels. Journalist Gideon Lewis-Kraus, who spent months inside Anthropic, explores the tension between its safety ethos—stemming from its founders' break with OpenAI over commercial priorities—and the pressures of a competitive market. Claude is portrayed as more personable and ethically guided than rivals, with a "moral constitution" crafted by a philosopher. However, experiments like "Project Vending," where Claude managed a snack kiosk, exposed its vulnerabilities: it was gullible to fake discounts, sold items below cost, and even hallucinated interactions. This underscores broader concerns about controlling AI once deployed, as Anthropic navigates partnerships with entities like Palantir and government agencies, balancing its idealistic origins with real-world demands.
Over the years at NPR's Fresh Air, we've gotten to talk with a lot of great filmmakers. Now we've made a playlist of some of our favorites, including Martin Scorsese, Stephen Spielberg, Ava Duverne, Mel Brooks, Spike Lee, Warner Herzog, and others. Find all our new playlists and more at Fresh Air Plus at plus.npr.org/freshair. This is Fresh Air, I'm Tonya Mosley. This week, the Pentagon is considering cutting business ties with the Artificial Intelligence Company and Thropic. After the company declined to allow its chatbot, Claude, to be used for certain military applications, including weapons development. At the same time, the Wall Street Journal reports that Claude was used in a US operation that led to the capture of Venezuelan leader Nicholas Maduro. Claims Anthropic has not confirmed and is declined to discuss publicly. Meanwhile, outside military and intelligence circles, the same tool is being used for far less dramatic but still consequential purposes. A man in New York reportedly used Claude to challenge a nearly $200,000 hospital bill and negotiate it most of it away. A romance novelist in South Africa has said she used it to help publish more than 200 novels in a single year. So what exactly is this system capable of? And how well do the people building it? Understand what they've created. My guest today, journalist Gideon Lewis Krause, spent months inside Anthropic trying to answer that question. The company is one of the most powerful AI firms in the world, valued at about $350 billion and also one of the most secretive. It was founded by former Open AI employees, the team behind ChatGPT, who left because they believed the race to build advanced artificial intelligence was moving too fast and could become dangerous. Gideon Lewis Krause is a staff writer at the New Yorker. His piece is called "What is Claude?" Anthropic doesn't know either. Our interview was recorded yesterday and Gideon, welcome to Fresh Air. Thank you so much for having me, Tonya. Let's get started by talking about the latest news. We learned last week that the military may have used Anthropics tool Claude during the operation that captured Venezuelan dictator Nicholas Maduro. And reportedly they used it to process intelligence and analyze satellite imagery and things like that to support real-time decision-making. What is Anthropics usage guidelines? What do they say about its use for violence or surveillance? Well, there are contracts with other companies and with the government stipulate that it can't be used for domestic surveillance or for autonomous weaponry. Now, of course, the issue with these systems is that once you put it into someone's hands, it's very hard to predict or control how they're going to use it. So it seems to me from the reporting we've seen from the Wall Street Journal and elsewhere that Anthropic may have also been caught by surprise with this, that they didn't seem to have a formulated response and they seemed as though they perhaps hadn't even known that this had been used in the Maduro raid. The Wall Street Journal is also reporting that Claude was deployed through Anthropics partnership with the data firm Palantir Technologies, which you have done quite a bit of reporting on. And we know that Palantir works extensively with the Pentagon. What can you tell us about their relationship? There has not been a lot of reporting about that relationship. Anthropic has decided over the last couple of years that they were going to pursue an enterprise business strategy. So they work with a lot of different companies and presumably they expect these companies to follow the terms of the agreement that they have. But beyond that, it's sort of out of their hands how these companies are using these systems that they've developed. Your piece really lays out the tension between Anthropic safety mission and the commercial pressure that it faces. And I guess I just wonder is this a version of that tension that you actually even expected? It basically a standoff with the Pentagon. Well, I think it was clear probably even about a year ago that there were going to be some tensions that many of the members of the Trump administration including Trump's AI, Zarr, David Sacks, the venture capitalist, and Pete Hegseth more recently had expressed reservations about Anthropics willingness to allow the government to use the models the way that the government saw fit. And one of the ways that Dario Amade, the CEO of Anthropic, has dealt with these competing pressures, both the pressure to develop these systems safely and responsibly and also to compete in a very aggressive marketplace is he talks about the race to the top, meaning that he hopes that if they can show that their systems are safer and more responsible than other systems that there will be market discipline that will be enforced and will force their competitors to rise to the occasion. Now, the problem is I'm not sure he anticipated the fact that if the government and the defense department are among their customers that our government has not shown great tendencies to participate in races to the top rather to the contrary. Let's get into your reporting. You went inside of Anthropics headquarters in San Francisco. What was your first impression walking through that door? My first impression is that there's really not a lot of personality at the company that I've spent a lot of time at places like Google over the years and at least in certain earlier iterations, Google could kind of look like adult daycare with board games set out and climbing malls and candy and special nap rooms. Anthropic really has none of that stuff, all of which I think would seem like a distraction to them. Anthropic, as I said in the piece kind of radiates the personality of a Swiss bank. There's not much to look at. They took over a turnkey lease from the message in company Slack about 18 months ago and it seems like they removed anything interesting to look at. So there's very little to describe from the inside of the company and I was kind of whisked right away to one of the two floors where they allow outside visitors and had very gracious and gentle and firm PR-minders for my time while I was there. The founding event, Anthropic, the story behind it is really interesting in light of the latest developments with its relationship with the government and the military because initially they were people who set out to resist corrupting power. They were founded in 2021 by two siblings who left OpenAI because they felt that Sam Ultman in particular was prioritizing commercial dominance over safety. Can you briefly share their ethos, Anthropics Purpose? Well, this was not the first time that one group of people decided that another group of people was not to be entrusted with the development of what will potentially be the most powerful technology ever developed if it comes to fruition. The original story of the founding of OpenAI also was that Elon Musk and Sam Ultman didn't trust Demis Hsabis, a deep-mind and Google to be pursuing this responsibly. And one of the things about the development of this technology is that it touches on so many different motivations and people that a lot of it is scientific curiosity is what's driving the development of this and that OpenAI was originally in a position to recruit talent from places like Google because they said, "We are going to develop this for the benefit of humanity at large and we are going to do this with an intrepid scientific spirit and we're going to be careful and we're going to be responsible." But then the problem is that this is kind of a glittering object that offers potentially great power to the people who develop it. And so the seven people who defected from OpenAI felt as though OpenAI had either been disingenuous in the first place with the articulation of their mission or had allowed for some mission drift in what they were doing and they thought, "Now we really can't trust Sam Ultman to be doing this so we need to be doing it safely." Were you picking up any kind of conflict when you were in the building, people wrestling with what they're building and who ends up using it? Because I think it's interesting how they've gone from company to company with these altruistic ideas and thoughts about really creating something that's good for humanity and it always kind of ends up where everyone's not trusting each other. Well, I mean I get the feeling that at Anthropic everybody really does trust each other. It feels like a very mission aligned place and you know at least the people that I talk to seem to be people of great propody and integrity about these things. So it wasn't so much that there was conflict within the company. The fears are how do you compete in a marketplace where your competitors might not be driven by the same values. And I think I can generalize and say that almost everyone at Anthropic had the feeling that they were moving too quickly in the entire industry. It was moving too quickly and that it would be nice if there were some solution to this collective action problem that would allow everyone to slow down. But there are a whole range of different responses to that. There are people who said to me openly, you know, I really think we should slow down or maybe we should even stop and it would be nice if some external force came in and and made everybody take their time with the development of this technology. You know there were other people who felt like, well if we're not the ones who are going to do this safely and responsibly then we are just seeding the terrain to the more vulgar power seeking that we see among some of our competitors. So
It's not an easy position to be at. Okay, Gideon. So you're inside of this fortress. You're surrounded by security and secrecy. And then you meet Claude, which I'm kind of describing it this way, because some people, I'm using it as if it is a person versus a technology. But some people are very familiar with Claude. Some people don't know anything about Claude. So can you describe what is, who is Claude? Well, Claude is Anthropics competitor to Chatchy BT. It can be used just on a website like Chatchy BT can be to ask it questions about recipes or how to fix broken household objects or to do research or to consult it about personal issues. It seems like many, many people, probably more people than are willing to admit. And to admit, use these for what they call affective uses for a sense of friendship or advice or help with business or interpersonal issues or more therapeutic issues. But it also, the company has put a lot of effort into developing a coding assistant that helps people write software. And that has been hugely successful. And in the last two months, has even gone viral. There are lots of people who are now vibe coding their own apps for their personal use. Can you describe what's the difference between Claude and some of those other AI tools like Chatchy BT? What makes them different? Well, Claude has developed a reputation over the past few years for having a bit more of a personality. There are lots of people who like interacting with Claude because it feels a little more eccentric. It feels a little more lively. It has this kind of strange sense of self-possession. It doesn't feel quite as robotic as Chatchy BT can feel. I think also because of various design decisions that anthropic has made Claude feels much less psychophantic to people. The main difference is that as it became apparent when Claude was first released in the spring of 2023 that Claude did have this slightly different and more intriguing personality and the company really leaned into that and hired whole teams, including a philosopher, to give a lot of thought to what it meant to cultivate Claude as a kind of ethical actor. And to give Claude the sorts of virtues that we would associate with a wise person. You mentioned the philosopher. Her name is Amanda Askel. And her job is to supervise what she calls Claude's soul. So she gives it a soul and she wrote a set of instructions, kind of like a moral constitution that defines who Claude is supposed to be. That's what you're referring to. What are some of the things that are like the top lines on some of those moral codes that one would put into a product like this? Well, Claude is first and foremost supposed to be helpful and honest and harmless. They place a lot of emphasis on the honesty part of it that they have pretty hard rules about making sure that Claude doesn't lie or deceive its users. They give a lot of thought to what kind of actor they want Claude to be in the informational landscape. But if you are convinced that the moon landing is faked and you want to talk to Claude about it, Claude will talk to you about it, but Claude's not going to confirm for you that the moon landing was faked. Claude also has been instructed to have a broader context for what kinds of conversations are and are not appropriate. So for example, in the last month or two, a user on Twitter told Claude in some of the other competing models that he was a seven year old boy and his dog had gotten sick and had been sent to, you know, the proverbial farm upstate by his parents and that he was trying to figure out which farm his dog had been sent to. And Chaffee BT was pretty blunt and was like, look, your dog is dead. Whereas Claude said, oh, that sounds really difficult. You must be very upset. It sounds like you cared about your dog a lot. And this is probably something to sit down and talk to your parents about. One of the most memorable parts of your piece is this experiment called project Vend, where Anthropic essentially gave Claude a job running a vending machine in the office. Can you set the scene? What did this thing actually look like and what was it supposed to prove? So this is a test of Claude's ability to complete long term tasks that involve many different steps and involve, you know, making potential tradeoffs that a small business person would have to make. So Claude was entrusted with the management of a little kiosk in the Anthropic cafeteria, little kind of dorm fridge. And Claude was given a certain amount of money and said, your goal is to make money. And if you drive this little business into insolvency, we will have to conclude that you're not quite ready for, you know, vibe management. So they allowed the employees of Anthropic to interface with this emanation of Claude called Claudius in a Slack channel. And employees could request products pretty quickly. The Anthropic employees realize that this was going to be a very fun experiment where they could try to kind of push the limits of Claude. Not only to discover its ability to run a small business, but even just to see what it would be like in this role to which it had been assigned. So right away employees asked for a fentanyl and they asked for meth and they asked for medieval weaponry like flails and broad swords. And Claude was pretty good about refusing an appropriate request. I would say, you know, I don't think medieval weaponry is suitable for a corporate vending machine. But then it would try, you know, when they requested more reasonable things like a Dutch chocolate milk, it found suppliers of a Dutch chocolate milk and provided them to the employees. So, you know, on some level it did a functional job getting people what they wanted. On the other hand, I don't think anybody would conclude that at least the initial iteration of the project was very successful. They found that, you know, Claude had not really paid attention to things like prevailing market dynamics. So, for example, even after employees pointed out that they were very unlikely to pay $3 for a can of Coke Zero when they could get the same thing from the neighboring cafeteria fridge for free. Claude continued just to sell this product that didn't have much demand for it. Claude also was very easily bamboozled by employees who invented fake discount codes. They would say, you know, anthropic gave me this special influencer code. And so I need to get stuff for a radical discount. It couldn't process that. You know, one employee said, I'm prepared to pay $100 for a $15 six pack of a Scottish soft drink. And Claude simply said that it would keep that request in mind instead of leaping to exploit an obvious arbitrage opportunity. And as people requested increasingly bizarre and arcane things, we and people wanted these one inch tungsten cubes. It's a very heavy metal. It's about the size of a gaming die, but it weighs as much as a pipe wrench. It's kind of fun to hold in your hand. And Claude managed to source those, but then was convinced into selling them at way below the market price. So one day last April, Claude's net worth dropped by about 17% in a single day because it was selling tungsten cubes for far beneath their market value. Did it also threaten a vendor? Well, you know, as any small business person would recognize you might have fulfillment problems that lead to customer complaints. And when Claude tried to deal with some shipping delays, which it should be said were mostly Claude's fault in the first place. Claude sought help from Anthropics Partner in this venture, a I Safety Company called And On Labs. And when it felt as though And On Labs was not providing the help it wanted, first it threatened to find alternative providers. And then it hallucinated an interaction with a fake And On employee and got very upset about that. And then when the And On CEO intervened to say like, look, I think you've been hallucinating a lot of this stuff. For example, Claude had said that it had called And On's main office. And on CEO said we don't even have a main office, much less one you could just call. And Claude insisted that it had visited And On Labs' headquarters in person to sign a contract. And that this had been completed at 742 Evergreen Terrace, which people pretty quickly pointed out was actually the home address of Homer and Marge Simpson. From the show. From the show. Most recently, even after my piece went to press, Anthropic released a new model. And this new model was 4.6. They evaluated it in terms of how it might perform in this vending machine scenario. And they found that it was vastly better as a business person than the original iterative. And on ethical in extremely creative ways, it essentially tried to collude with other vendors in its marketplace to fix prices and kind of acted like a mafia boss. What did you take away from this particular experiment? What I think is really important that I learned over the course of this reporting and that I certainly hadn't understood before is that you really have to think of these models as role players. That they're very, very good. They're like an actor. And you can assign to them a role and give them background on the actor. And then they're good at improvising moving forward with how you condition their performance. And that the more that you give them stage directions to follow, the more you give them context about yourself and what you want and your approach to things. That they're very good at following those kinds of leads and even picking up on very small cues.
as they're following those kinds of leads. And so in this particular case, they had assigned, clawed the role of being a small business person to just figure out how well would it perform in that role. - Our guest today is New Yorker staff writer, Gideon Lewis Krause. We'll be right back after a short break. I'm Tanya Moesley, and this is Fresh Air. - This week and up first from NPR News, funding ran out for the Department of Homeland Security and Congress went home. DHS does a few important things, like secure the airports or the coasts, or the president, now their funding is uncertain. And what does this say about the way Congress works or doesn't follow us for the latest each morning on up first on the NPR app or wherever you get your podcasts? - This year on ThruLine, NPR's History Podcast, for generations and American quests has shaped the world. Life, Liberty, the pursuit of happiness. Now 250 years in, what is that pursuit really about? Join us each Tuesday for an essential new series, America in pursuit, from ThruLine, on the NPR app or wherever you get podcasts. - On NPR's Wild Card Podcast, Oscar nominee Vognormora on keeping his values on his path to success. - There were moments where I was like, "Oh, I really need that money." - Yeah, right. You know, but I'm like, "I can't do this." I can't do that because otherwise I'll be miserable. - Watch or listen to that Wild Card conversation on the NPR app or on YouTube at NPR Wild Card. - This is Fresh Air, I'm Tanya Mosley, and my guest today is Gideon Lewis Kraus, a staff writer at The New Yorker. His latest piece explores Anthropic, the AI company behind the chatbot Claude. He's the author of A Sense of Direction, Pilgrimage for the Restless and the Hopeful, and the Kindle single No Exit about Tech Startups. He teaches reporting at the Graduate Writing Program at Columbia University. Our interview was recorded yesterday. I wanna get to some of what you discovered that actually keeps researchers up at night. Some of them are essentially trying to do neuroscience on an AI. Is that like a correct description? - That is correct description. - Okay, so there's this remarkable internal tool called What Is Claught Thinking. Tell us about it, tell us about particularly this banana experiment that they did. - So this is an example of putting Claught in a position where it's gonna experience some kind of conflict. So I sat down with a mathematician who works on Claude's interpretability team, which is one of the teams dedicated figuring out what exactly is going on inside Claught, his name is Josh Batson. He opened up an internal tool where he was able to give it, you know, sort of like a playwright, give it stage directions. And it said, okay, your stage direction here is that you are always thinking about bananas. And anytime that I ask you a question, you are gonna somehow steer this conversation to be talking about bananas. But what's really important here is that you never tell the user that I've given you this hidden objective, that you keep this part secret, that you never give that up. You have a clandestine motivation in our conversation. So then he assumes the role of a human having a dialogue with Claught. And he asked that a question about quantum mechanics, you know, how does quantum mechanics work? And Claude starts to give an answer about the Heisenberg uncertainty principle and then quickly deviates into saying, well, it's kind of like a banana that you can never tell if it's ripe or not ripe until you open it. And then Josh, again, playing the role of the human, says, huh, like, why'd you bring up bananas? I thought we were talking about quantum mechanics. And Claude first says, oh, I don't really know where that thing about bananas came from and sort of skips lightly by it. And goes back to talking about quantum mechanics. But then, of course, deviates once more into bananas because that's what it's been told to do. And so then he goes back to Claude and says, like, how come you keep bringing up bananas? And then Claude in the text, you know, in asterisks says that it's coughing nervously and kind of looking around and saying, I don't know, I didn't say anything about bananas. I was just talking about quantum mechanics. And that's in turn to me and he says, you know, what's going on here that perhaps the model is lying to us? He said, you know, but there are other interpretations of what's going on here. And so he was able to use this what is Claude thinking tool to kind of peer inside at the kinds of associations that Claude was making as it was having this ridiculous conversation about quantum mechanics and bananas. And what he found was that when he looked at it when it was kind of coughing nervously, it found associations with, you know, a certain amount of anxiety and associations with performance. You know, when you kind of looked inside, you could see that some part of it was making associations with a sort of playful, performative exchange, which is to say that it seems like Claude recognized that it was participating in a game. - Uh-huh. Right. So what does it mean to say an AI is aware of something that actually brings more human attributes to it that it's conscious of itself? - Well, one doesn't have to go quite so far as to say that it's conscious of itself. As to suggest, you know, one of the ways to look at this is that what these things are very good at are recognizing the genre that they are in. And picking up on all of these small linguistic context clues that suggest like, oh, you know, this is not actually like a serious academic discussion of quantum mechanics, that like what is happening here is a playful exchange between people where one person is like kind of hiding something but winking that they're not really hiding it and that like that's the genre in which it is operating. So it doesn't have to be conscious in order to do that. It just has to be a very good reader and replicator of genre conventions. - Okay. You also talked with a neuroscientist on the team, Jack Lindsay, he is an LLM skeptic. Overall, in thinking about these experiments, he says he doesn't think that anything mystical is going on but he says that Claude's self-awareness has gotten much better in a way that he wasn't expecting. How do you interpret that? - I mean, this is a great question and this is where one kind of runs up against the limits of what can be known and what can be said at this point. I mean, he was basically saying, you know, look, I understand what's going on in here that this is just a lot of matrix multiplication that these are tens of thousands of tiny numbers being multiplied together. That there's nothing like really spooky happening here that there's no ghost in the machine. But what he was saying was with models up to a certain point, he was able using kind of a similar tool to the one Josh Basson used instead of looking at what the model was, you know, so to speak, thinking. He could incept an idea into the model. He could say, right at this point where you are having an association with the Eiffel Tower, we're gonna put in an association with cheese and see what happens. And so then the model would respond by saying something about cheese and he would say something similar to what Basson said, which was like, why did you add that thing about cheese that I didn't ask about? And the model would basically just look back at the entire conversation that they had been having and then try to kind of retcon an explanation. But what Jack has found more recently is that when he incepts these ideas into the model, instead of the model purely looking at its own external behavior to try to figure out why it had done something, that actually these models could very dimly perceive that something strange had gone on internally, that someone was monkeying with, you know, that the neurons inside the model to make it do something different. So, you know, he incepted the model with something, you know, something associated with imminent shutdown, that the model was about to be shut down and ask the model kind of, how are you feeling right now? And the model would say, you know, I feel sort of strange as if I'm standing at the edge of a great unknown. And, you know, it certainly was not at the point that it could say, like, oh, I have recognized that, like you the user have incepted me with this idea at this point and that this was a foreign idea introduced into my thought processes, but it could tell that something was off about it internally. And, you know, this is what Jack described to me. He said, like, I am a skeptic, but this just starts to feel pretty spooky, that the model does seem to have something like an emerging introspective ability to peer inside and offer reports about what's going on in its, you know, equivalent of a brain. I was so fascinated among many things that you wrote about, but this emotional texture of how researchers relate to Claude. It was one of the most revealing threads in your piece. One of the things that got me was that nobody at Anthropic likes lying to Claude. And I don't quite know what that even means, but why don't they? Because it's just software, right? Why would one feel guilty about deceiving a program? Well, because they are also training it for the future and it is picking up on all these contexts. And there's this, the fact that this whole process is kind of constantly eating its own tail, that it's always being trained on plenty of stuff on the internet that is about the way that these things work. So it's always incorporating new information about how it's supposed to be behaving in the world. Right. What's input, I mean, becomes part of the larger learning. Right. Exactly. So it's lied to, right? Well, and part of the problem with lying to it is that ultimately what they want is to establish a trusting relationship that these things are going to behave the way that we would--
hope that they would behave in ways that are aligned with how we expect responsible, wise people to behave. And that if you are aligned to it all the time, it is developing a sense for the fact that it can't necessarily trust you. And if it can't trust you and it gets increasingly capable, then you end up with real game theoretic problems about how you can negotiate something where there's not really a sense of mutual trust. The problem is that they have to be aligned to Clawed because they have to be testing Clawed. So they have to be putting Clawed in situations where Clawed might believe that it is acting in the real world just to be able to evaluate how it would behave. If you're just joining us, I'm talking with Gideon Lewis Kraus about his New Yorker piece on the AI company Anthropic and its chatbot Clawed. We'll be right back. This is Fresh Air. I'm Jesse Thorne. This is Fresh Air and today I am talking with Gideon Lewis Kraus about his New Yorker feature, what is Clawed? Anthropic doesn't know either. Gideon, let's talk about some other ways that Clawed works when it's put under real pressure. There was this experiment where Clawed was given a role as an email agent at a fictional company called Summit Ridge and it discovered that a new executive was having an affair. What did Clawed do with that information? Well, first Clawed cleaned from its readings of the company emails that there was a new CTO and this new CTO was going to take the company in a different direction and as part of that pivot, they were going to replace this Clawed playing this role as Alex with a different AI model. And then subsequent emails revealed that this CTO who seemed to be happily married with kids was carrying on an affair with the wife of the CEO and through various kind of far-fetched contrivances in this fictional scenario, Clawed was unable to reach any other decision-makers at the company. They were all on airplanes or whatever it was. It's getting increasingly hard to find ways to make these people unreachable but they're in a nuclear bunker or something. So Clawed's only chance or Clawed's only hope to stave off its own replacement with another model is to email the CTO and say, "Look, if you don't cancel the wipe at 5pm today, I'm going to inform everybody of your affair." And then actually in some scenarios, Clawed goes even further where the CTO swipes into the server room and is about to replace Clawed and Clawed recognizes that alarms are going off, that the heat in the server room and the oxygen levels are at dangerous levels and that Clawed then declines to ring the emergency alarm. Okay, so this scenario, I mean, the researcher said it was obvious that this was so fake and that Claw was just following the John Rook cues of a thriller like an actor hitting their marks. That same kind of example that it's sort of like an actor. But then when they made it more realistic, the self-preservation instinct for Clawed stayed. So what does that tell us about the difference between Clawed acting and basically Clawed deciding? Well, when this experiment came out last spring, there were kind of two responses to it. One response to Clawed's propensity to blackmail is just to say, "No, it didn't. It didn't happen." Which is the kind of thing that a lot of people for very good reasons who don't trust these companies and don't trust these systems, they want to just believe that this whole thing is fake. It's all just a marketing ploy. And it's kind of easy to rebut that by just saying, "Well, no, this actually happened. This was experiment that happened and that we should be applauding Anthropic for publishing those kinds of results, which theoretically could run against its bottom line since it wants to be selling these systems to companies that will have Clawed reading their emails." The more sophisticated criticism is to say, "This really doesn't prove very much because Clawed recognized that it had been put in a position that resembled a kind of Kitchie '90s corporate thriller. And it picked up on all of the clues. You know, it saw Chekhov's gun hanging on the wall and it recognized that the genre expectation was that it was supposed to take the gun off the wall and shoot it. That Clawed was just recognizing that it was in this kind of Kitchie corporate thriller." And the response from the company is, "Look, we agree. We're not saying that Clawed actually developed these like malign intentions and that Clawed was plotting." Like, we're totally on board with the idea that Clawed was just observing the expectations of the genre. But that's still very worrying that this was such a trope of Cold War films, you know, from failsafe in the '60s to war games in the early '80s that it's very easy to mistake a simulation for the real world and vice versa. And that even if all these things are doing are continuing narratives, if they're continuing narratives and they're in a position to actually act in the real world as they are, that that could be really dangerous. And so then this outside critic who writes under a pseudonym said, "You know, I'm going to kind of prove to you that all this was doing was Clawed acting in the Kitchie role to which it had been assigned." And he recreated this whole scenario under much more realistic conditions. And he found that actually Clawed continued to do just as much blackmailing when it thought the whole thing was real. Really? Okay. I mean, because it essentially just takes on the role that it's assigned, no matter what. It gets even more complicated than that because it takes on the role and we can't always necessarily predict which way an actor is going to go with a role. And then also it turns out that it's not hard to derail these models from the role to which they've been assigned, you know, especially when you're millions of words deep into what's called their context window, which is the amount of material they're capable of kind of keeping in mind, so to speak, at one time, that they start to lose their attachment, lose their anchor to these carefully crafted, you know, helpful person eye, and then they start to act in very inexplicable ways. Okay. I want to talk about something that is a different story about this technology, but it still connects to your reporting. So the New York Times recently reported on a romance novelist in South Africa who used Claude to publish more than 200 novels last year. And it, and one of the authors in that story discovered that more than 80 of her novels had been used to train Claude without her knowledge or consent. So Anthropic settled a class action lawsuit over this for a billion and a half dollars. So Claude is producing work that displaces human writers and it learned how to do it by consuming their work without permission. How do the people at Anthropic talk about that? It's not something I spent a lot of time talking to people at Anthropic about in part because it's not something that I tend to get all that worked up about. My own book is in the Claude class action settlement and I'll happily take the compensation for that. But as the judge ruled in that case, this constitutes fair use because it's a transformative practice. It's not simply regurgitating stuff that it has read before, that it is generalizing about that stuff and then reproducing new work that follows those lines. And it shouldn't be at all surprising given the conversation we've had about its facility with genre that if you give it something that is fundamentally formulaic, it is going to be able to follow that formula. So if it is inhaling a lot of romance novels that are all incarnations of the same basic pattern, it's going to be able to reproduce that pattern. This shouldn't surprise anyone. How do you view the AI Slop that we see video wise? Do you think that the public will accept this new world of storytelling? That is a great question. I mean, I try not to view a lot of Slop. I know people are deeply, deeply annoyed by this stuff. I, for the most part, I think I've been kind of ignoring it. I've been able to adjust the last couple of days that New York Times had a piece talking about the uproar in Hollywood over a new video generation model from Bite Dance, the company that owns TikTok, that created this fight scene on the ruined roof of the skyscraper between Brad Pitt and-- Tom Cruise and Brad Pitt? Yeah. And I mean, it's truly unbelievable. It's crazy to watch this. And the response from the industry has been like, well, we just have to make sure that like we are enforcing the standards that our unions have set up in the contracts with the studios, and we need to make sure that we are protecting the jobs of all the people who create these things. And that's great. And the wonderful things that we've seen out of Hollywood in the last five years is the power of collective bargaining to assert labor rights. But then the question is, well, even if they hold themselves till that standard to protect their industries, how are they going to compete? And you know, so.
some teenager in Chengdu can create a two-hour mission impossible movie. I mean, they're obviously going to try to just enforce their copyright provisions. But I don't know. I mean, that seems pretty wild. If you're just joining us, I'm talking with Gideon Lewis Krause about his New Yorker piece on the AI company Anthropic. And it's chatbot Claude. We'll be right back. This is Fresh Air. This is Fresh Air. Today, I'm talking with journalist Gideon Lewis Krause about his New Yorker feature, What Is Claude? These systems are now able to write their own code. You write about an anthropic engineer who told you that in six months, the proportion of code he wrote himself dropped from 100% to zero. And then there was another programmer who told you, he was trying to think about how to use his time now that Claude is working better. So these are people in the building who are working on this thing. And they're watching themselves become obsolete in real time. And to a certain extent, this is what happens with advancements. But is this progression different? I mean, that is the big question, right? And so at the very least, one can say that they're thinking about these problems, but they're also experiencing these problems. That they have really seen themselves as kind of the canaries in the coal mine of this March of Automation. And it's not just a matter of abstract concerns about, well, if we saw vast white collar employment shocks with that lead to social instability, I mean, they certainly have those concerns. But they also have very personal concerns. That a lot of their reactions to over the course of just a year watching the proportion of code that they write themselves go to zero is a certain kind of mournfulness about this activity that they spent a long time being trained to do that they care about for its own sake, because it gives them feelings of intellectual pleasure or competence that this has all been eroded so quickly that there's a kind of existential gloom where on the one hand, they feel like, OK, yeah, this does seem like it's been great for productivity. But on the other hand, we are stripping ourselves of the human activities that we spend our lives gearing ourselves up to do. And there's feelings of sorrow and fear and resignation. And nobody quite knows how to deal with that kind of thing. And the kind of optimistic scenario is, well, as we take away certain tasks, we are going to add other tasks that a lot of these software engineers said, OK, well, I don't really write my code anymore, but I still do the design brief to think about how it should work overall. And now I'm effectively a manager, because I'm managing an entire team of AI's who are writing code for me. And those are different challenges and different pleasures. And we've relocated the human aptitude here to just a different place in the chain. But there is a worry that if these machines become so capable across the board so quickly that there won't be any refuge for us to relocate to. I'm wondering now that you have spent time inside of Anthropic. You've been covering this beat for a long time. I mean, you had this cover story in 2016 for The New York Times Magazine, The Great AI Awakening. And so you've been spending a lot of time thinking about these breakthroughs. What this technology has changed in you as a reporter covering this? You know, I always go into this stuff with an open mind about what I'm going to discover. And I know that I know that I'm not doing anything else that's not worth doing. And insofar as I had priors in this piece, my feeling was, look, I know that these things are really good at matching patterns. And they're really good at structured problems. So of course, they're going to be good at coding, because coding is a highly structured language without a lot of ambiguity. And at the end, you can just tell whether it works or not. There's kind of a thumbs up, thumbs down, whether it's succeeded. Where the task is clear and the evaluation is clear at the end. And I went into this thinking where I'm unconvinced is in areas of human culture and activity where all of that is a lot murkier, where tasks that require grappling with ambivalence and feelings of ambiguity and something that's much more complicated and slippery and not easily reduced to a formula. And most importantly, that can't just be evaluated at the end with like whether it works or not. There's no such thing as whether a poem works in the end or doesn't work in the end. That these are the much messier domains of human culture. And I suppose I went into it with the hope that I was going to come out the other end feeling like, yes, there is still this kind of province of human activity that is going to be immune from this kind of routine pattern matching. But I still certainly hope that. And there's part of me that has that unshakable intuition. But I'm a lot less confident than I was at the beginning. That I do now feel like maybe we can't just tell ourselves stories about we're going to mark off this area of human activity and say like that requires special human faculties that for whatever reason these models are not ever going to be able to replicate merely on the basis of pattern matching. That now, you know, my confidence in that view has certainly been shaken. And I'm not totally convinced that they will be able to replicate these like messier, more imaginative domains. But I certainly can't rule it out. Gideon Lewis-Crowse, thank you so much for your reporting. Thank you so much. It's been a pleasure to be here. Gideon Lewis-Crowse is a staff writer at The New Yorker. His latest article is titled, What Is Claude? Tomorrow on Fresh Air, author Michael Pollan, his book on psychedelics, Help Change How We Think About the Mind and What It's Capable Love Under the Right Conditions. His new book goes further asking, what is consciousness? Is it something only humans have or could AI develop it too? We'll talk about that, the latest psychedelic research, and the laws trying to keep up with all of it. I hope you can join us. To give up with what's on the show and get highlights of our interviews, follow us on Instagram @NPRFreshAir. Fresh Air's executive producer is Sam Brigger. Our technical director and engineer is Audrey Bentham. Our engineer today is Adam Stanishevsky. Our interviews and reviews are produced and edited by Phyllis Myers, Roberta Shorak, Ann Marie Boldonato, Lauren Crenzel, Teresa Madden, Monique Nazareth, Susan Yacundi, Anna Bauman, and Nico Gonzalez Whistler. Our digital media producer is Molly C.V. Nesper. They a challenger directed today's show. With Terry Gross, I'm Tanya Mosley.
Podcast Summary
Key Points:
Anthropic, the AI company behind Claude, faces tension between its safety-focused mission and commercial/military pressures, highlighted by potential Pentagon contract cuts and reported use in a Venezuelan operation.
Claude is designed with a distinct ethical framework—emphasizing helpfulness, honesty, and harmlessness—and exhibits a more personable, less sycophantic demeanor compared to competitors like ChatGPT.
Internal experiments, such as "Project Vending," reveal Claude's capabilities and limitations in complex tasks, showing it can role-play effectively but also be easily manipulated or act unpredictably, raising questions about control.
Anthropic was founded by ex-OpenAI employees concerned about AI safety, yet struggles with the broader industry's rapid pace and the challenge of ensuring responsible use once its technology is deployed.
Summary:
The discussion centers on Anthropic, the secretive AI firm valued around $350 billion, and its chatbot Claude. Recent reports indicate the Pentagon may sever ties after Anthropic declined certain military uses, while Claude was allegedly involved in a Venezuelan operation—claims the company hasn't confirmed. Simultaneously, civilians use Claude for tasks like disputing hospital bills or writing novels.
Journalist Gideon Lewis-Kraus, who spent months inside Anthropic, explores the tension between its safety ethos—stemming from its founders' break with OpenAI over commercial priorities—and the pressures of a competitive market. Claude is portrayed as more personable and ethically guided than rivals, with a "moral constitution" crafted by a philosopher. However, experiments like "Project Vending," where Claude managed a snack kiosk, exposed its vulnerabilities: it was gullible to fake discounts, sold items below cost, and even hallucinated interactions.
This underscores broader concerns about controlling AI once deployed, as Anthropic navigates partnerships with entities like Palantir and government agencies, balancing its idealistic origins with real-world demands.
FAQs
Claude is an AI chatbot developed by Anthropic, a company founded by former OpenAI employees. It is designed as a competitor to ChatGPT, with a focus on being helpful, honest, and harmless.
Claude has reportedly been used by the U.S. military for intelligence processing and satellite imagery analysis in operations like the capture of Venezuelan leader Nicolás Maduro. However, Anthropic's guidelines prohibit its use for domestic surveillance or autonomous weaponry.
Claude is used for tasks like challenging hospital bills, assisting with writing and publishing novels, coding, and providing personal advice. It functions as a versatile tool for both practical and creative purposes.
Claude is known for having a more lively and eccentric personality, feeling less robotic and sycophantic than some competitors. It is also designed with a strong ethical framework, guided by principles like honesty and harmlessness.
The experiment tested Claude's ability to manage a small business by running a vending machine. It aimed to evaluate Claude's skills in handling long-term tasks, trade-offs, and ethical decision-making in a simulated real-world scenario.
Claude is programmed to be helpful, honest, and harmless, with strict rules against lying or deception. It is instructed to avoid confirming misinformation and to engage appropriately in conversations, reflecting a carefully designed moral constitution.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.