This is the worst they will ever be. Like right now is the worst they will ever be. And a lot of the reactions do seem to be sort of like stuck in this moment of 2021, 2022 when like the failures were kind of laughable and like you could like have a gas about asking it to count the ours and strawberry or whatever. But like if you think this is a serious threat, you have to take it seriously, which means you have to engage with like what it's actually empirically doing now, which is like not glorified auto correct. (upbeat music) Hello and welcome to Why is this happening with me, your host Chris Hayes. (upbeat music) We are right now, as I speak to you in the midst of one of these kind of occasional cyclical bits of like very intense AI focus and hype. And I feel like there's been a strange cycle for all of us, you know, who don't like work in AI and don't work in computer science in the last few years, I think since chat GPT was introduced where there's these kind of oscillating cycles of hype. And you know, we read about it in the business section of the paper and the economic section of the paper and literally untold amounts of capital expenditure and investment happening. It's driving a huge part of the economy. Every other ad in the super bowls about it, being bombarded. And some part of me has reacted to it with a kind of like fetal position sense of like dread and that it gives me the bad feeling. And one of my New Year's resolutions was to like get out of that defensive crouch because once I get past my sort of aesthetic distaste and political and moral concerns with it and dive into like the deep questions of the technology, it's genuinely probably the most fascinating thing happening right now in the world. Particularly for someone like myself that was a philosophy major and was particularly interested in philosophy of mind and cognitive science and consciousness. And so one of my New Year's resolutions has been like, all right, I need to get my arms around AI. And you know, when people come to me and they ask me views on political matters in the US or policy questions or legislation, I have a bunch of pretty settled views that are born of a lot of reading and reporting and stuff. And on AI, I feel like I have some instincts. I have some sort of general sort of contextual background priors and worldviews and commitments and values and ethics. But I feel tossed around and I'm trying to kind of make sense of it. And so into that came this fantastic article. It felt like it was written for me this week in New Yorker. And it's by Gideon Lewis Krause. It's an article about a lot of things. But specifically, it's basically an article about the company Anthropic, which makes HOD. They are one of the main big AI companies. You may or may not have seen them. And it's an investigation into, in some senses, the kind of moral epistemic problems at the heart of this new technology and the people inside the company who are wrestling with it. And it's an amazing way also just into modern AI. It does some of the best explanations of how models work in kind of plain language that are possible that Ivan countered, which is actually super important because I think a big part of this question is there's a little bit of a like wizard of Oz mystification process happening, where these people that understand how the machine works and all of us plebs are to be here and just use the machine, because I tell us we have to use the machine. And it feels very like enraging and disempowering and also very politically suspect to me. And so a big part of this project, I think, is just understanding. And the author, Gideon Lewis Krause, does as good a job of explaining in words as Ivan countered about how some of this stuff works and how they don't know how some of it works. And so I thought it would be a great time to talk to Gideon Lewis Krause. He's been a guest on the podcast before. We talked about UFOs once before. He's a staff writer at the New Yorker, Gideon Lewis Krause. Welcome back. [MUSIC PLAYING] Thank you so much for having me, Chris. So there's a few ways we can attack this. But let's just start with the company at the center of this piece, because I'm sure people have seen it. They had a very funny set of Super Bowl ads, actually making fun of their rival, ChatGPT. For the fact that ChatGPT is going to allow ads in their AI, just tell us what the company Anthropic is, who founded it and how it came to be. Anthropic was founded in 2021 by seven people who had defected from OpenAI, some of whom had been at OpenAI, basically from the beginning. There was some sense of history repeating itself, because of course OpenAI had been founded initially in 2015, because they were worried that Demisis Abbas at deep mind, now Google Deep Mind, was not trustworthy enough to oversee the development of essentially most powerful technology anybody had ever seen. Then through the cycle of this stuff happening, there was unrest at OpenAI, especially in the wake of big Microsoft deals, and then Dario Amade, and his sister Danielle and five other people decamped in the late fall of 2020 to found Anthropic. I want to just stop there, because I think it's actually really fundamentally important. So Demisis Abbas is this genius. He's a Londoner. He was like a chess prodigy and like video game prodigy, like one of these, you know, one in a hundred million kind of things of like everyone recognizes as a genius from like a very young age, he skips a bunch of grades. He starts this deep mind company to kind of solve for intelligence. It's acquired by Google. And OpenAI just to repeat this, like there's this repeating thing that keeps happening, which is everyone's like, no, the ring is too powerful, you can't have it. We're going to go build a company that's ethically distributed in a way that no one will be corrupted by the ring. And then they get closer to the ring. A bunch of people inside the company are like, no, no, I don't want you guys to have it either. We're going to go build our version. That's the act. So OpenAI originally, like the reason it's called OpenAI, right? Like people forget this, because now it's just a company with a CEO who gives money to the Trump Super PAC. Was this like novel, non-commercial, almost non-profit structure so that it wouldn't be in the hands of one commercial entity because it was such a powerful and important technology, right? Although the difference is, and I think this is worth pointing out because there are a lot of similarities. So it's worth dwelling for a minute on the differences. The differences are nobody really bought that to begin with. Or at least there was a lot of skepticism about that to begin with. That people knew that Elon Musk had tried to buy Deep Mind and was enraged that it had gone to his rival, Larry Page. And even back then, I mean, this was like the first time that I was really doing a reporting about AI back in kind of the Paleolithic of Deep Learning in 2015, 2016. Even back then, people were like, this is really just a recruiting pitch that he wants AI scientists to work for Tesla. And he can't hire them. And this is his way to create this kind of shell so that he can read Google's talent for his own purposes. So there was a lot of feeling even at the beginning that this was disingenuous. I think there are some clear structural kind of ostensible similarities with the founding of Anthropic. To be fair, I don't think anybody was saying, oh, this is just like Daria Omadeh, another like cutthroat hustler who's like trying to make a play for talent. They started off broadly in good faith, I think. And there's not say they're not anymore, but yeah. Right. No, I think that's important because I think there's a wide array of how much people do know about AI. But the reason I want to say to contextually is like, we've seen this sort of interesting trajectory with the tech sector as a whole of like, yeah, everyone remembers Larry Page. You just mentioned Sergey Brand, like the motto of Google was don't be evil. Right. And this idea of that generation tech founders basically being really thinking we're going to be different than these old incumbent industries and both more progressive and more attentive to the social utility of what we're doing and we're going to open the world and connect people. That all seems like very long ago. And they all seem like really reactionary and frankly dangerous to me, a lot of them. And I think the reason Anthropic kinestics out is because whether you take it at face value or not and whether it's been corrupted or not, you're saying they're operating good faith, they still very much kind of at least project that vibe of we're thinking about this in a socially responsible way. We want to be a kind of ethical company essentially with this powerful technology. Well, as far as going forward, there's one point that keeps coming up and I've been talking about this with people that I think is worth kind of getting out to the table, especially given your own interest in the relationship of capital and labor in this case. Is that one of the stories that people have told about the tech shift to the right, is that this really is a story about a change in power between labor and capital. That 10 years ago, labor really had an advantage in this relationship because there was so much engineering expertise that capital had to negotiate with labor. And that then you see the stories often told, you see the Google walk out over Project Maven in 2017 and that then the corporate overlords start to feel labor is getting off of it and we got to put them in their place and that's how you end up with this mag as shift among a lot of tech executives. I mean, I think the story is much more complicated than obviously there's some truth in it. A part of that story that doesn't get told as often is that one of the things that also happened in the meantime was that the engineering expertise behind a lot of software got commodified. So it became something that they could just buy in bulk, especially through contractors, which is one of the things that took away a lot of the power of labor. Now in AI, we may get to that point and they may be like digging their own grave with like AI that's potentially recursively self-improving and they say that the newest Chatchee BT was largely written with codex and cloud was written with cloud, et cetera, like that may also be shifting. But at least right now, we are talking about like less than 1,000 people who have the background to be doing this. - Meaning like the people who are doing what we call the frontier models who are like in it every day, pushing them out, there's like maybe 1,000 people who have the chops and expertise to be able to actually do it. - And because there's still really a premium on that expertise, that means that labor still has a lot of power at the frontier labs. When people ask me like, well, what's gonna happen when like executives are gonna go executive and just like any kind of other replacement level tech executive like underneath it all, it's just vulgar power seeking, like we're gonna watch this play out again. You know, every company says that they're mission driven and Throbic, it's very legitimate. Like the company's really mission driven. You know, I had a guy sit in front of me say, like Mark Zuckerberg would offer me $50 million. And I was willing to deign to have a phone call with him was like the attitude, but I was never gonna take this job. And I think part of it at in Throbic, if the kind of executive cohort there starts doing things that seem like really grubby as far as this like shiny glittery ring goes, they're gonna see a lot of defections. And the fact that we haven't seen that so far, I think actually is kind of an external like behavioral measure of their integrity. - Dario Amade, who's one of the founders of the long-distance Daniela, he got a PhD and I think molecular biology, particularly around neural circuitry, right? - Yeah, so he had started out as a theoretical physicist then he shifted as an undergrad into computational biology and then he was doing biology work and then he moved into AI. Which is actually, I mean, it sounds peculiar. It's actually like pretty characteristic of a lot of these people. I mean, like almost everybody at Throbic two minutes ago, they were a theoretical mathematician or whatever. - I think he's carved out this space again. I'm like for lack of a better word, like humanist progressive lab. He's written two really, I think really interesting essays about the sort of vision one called machines 11 grace. The second was his constitution they published. How would you sort of describe like the approach and worldview there and what sort of distinguishes it? - Well, the story that they tell, which I think is certainly largely true, is like we're the people who really care about the safe, responsible development of this technology. And Dario likes to talk about a race to the top where they're gonna set this example that other people are gonna have to follow. And I mean, I don't think that's completely insane. Clearly there is some market discipline that's gonna be involved here. So many of open AI's problems have been because they don't have proper safeguards on what they're doing, which has punished them in the market. And really been a hit to their reputation. So I think part of it is just like market savviness about the value, especially in the consumer end and even more so in the enterprise end, since they're largely an enterprise business. You know, if you want to sell to a Fortune 500 company, you can't have something that hallucinates and is sick of antic and is gonna deceive you in all these things. But I think like very often the safety stuff even when it's really genuine is still like kind of an alibi for like basic scientific cure, which-- - That's the real driver. - Not forever, uncle, but for a lot of people. Like there really is, I mean like this goes back to the atomic bomb also, that there's this feeling that like if you can build something, that's gonna be like new and exciting and feel like it's kind of the cliff face of technology. Like somebody's gonna do it just because it's really interesting. - Yeah, I mean, the Oppenheimer example, people bring it up a fair amount. It's not like a novel insight, but it does seem to me to have some real force to it. I've not actually, I haven't seen the movie, but I read the incredible Kyber Martin Sherwin biography and like, you know, he's a complicated guy in a million ways. But this sort of core thing of, he's in a context where it's part of a project for this very clear cut way to beat the Nazis to it, which provides a kind of searing moral clarity, I think. But he's also, you know, we all know that what would become of nuclear weapons and he's not needy about it. But fundamentally, it's like can we do this is really at the core of the driving animation? - And I mean, it's not just can we do this in practical terms. It's like, as one person said to me, so many old questions feel new again in because of this technology. Because we have something that like, doesn't work in theory, but it works in practice. And anytime you have something that works in practice and doesn't work in theory, that's gonna open up that gap between like what we're observing empirically and what we can account for theoretically, those are where the most interesting scientific problems happen. So the like the fact that we have these machines, I can do these things that they have no business doing, there was a lot of other things into question. A lot of things that are like feel up for grabs now in a way they haven't in a long time. - So I totally agree with that. I wanna spend a second unpacking that because I think it's a great way to think about it. It works in practice, not in theory, and what that means. Here's my read of that. So I'll give my version and then you give your version. Like I spend a lot of time thinking about philosophy of mind and have studied it, right? And there's like a bunch of competing questions in philosophy of mind. There's a bunch of competing questions in linguistics about how do we speak? How do we understand language? How do we form concepts? How do we do all these things? They have built up a system, sort of based on some of those theories, but is basically using just like sheer computational power and like enormous basically multi vector spreadsheets, like multi-dimensional spreadsheets, to produce patterns of behavior that really do seem from the outside, like listening and thinking and reasoning, in ways that are they can't account for on the inside. They're just like shoving all this compute power in. And so you've got this thing that's, boy, it sure seems like you're talking to a human on either side, not quite, but a lot like more than anything you could do 10 years ago, but we don't actually really understand what it's doing in there, basically. - Right, there is a whole philosophical tradition that predates this stuff. So it's not like it completely came out of nowhere. There were these ideas that eventually became described as connectionism that you can go back and find, like in the philosophy of mind, you can find in Gilbert Ryle in the 50s and you can find it in Vick and Stein. There's this whole tradition of, you could call it inferential wholism or something. That says there's no ghost in the machine. There's no little homunculus inside the mind directing things that all of this is kind of the interweaving of words that are related to each other, beliefs that are related to each other, that it's about like the coherence of this system. So it's not like there's no philosophical account of what's going on. - Totally. In fact, there are certain schools that are the ones that it's based on and certain schools that are not, like very distinctly. So now when we're talking about this thing, this model, right, the models. We'll talk about the models. Like I've worked pretty hard to try to understand them, working from the basic level of a neural network, which I can't understand to a transformer, which I kind of understand and then just gets bigger and bigger. We don't necessarily need to go to that level of specificity here. I wonder if you can like take a run at a accessible plain language explanation of how the models work. - I'll start by addressing kind of like the main objection to this, which is that like somehow this is just like reducing words to numbers. And numbers are these like robotic things that obey rules, whereas like language is lively and it's disorderly and it changes and it's slippery. But like that is mistaking what kind of like old fashioned programming was for what this is. That old fashioned programming was based on rules. It was based on like rules and symbolic logic and deduction and syllogism and all these things where you can like specify very clearly and logically like how something is gonna proceed. And you can't blame people for still like bringing that idea to this and thinking like, well, it's just a bunch of like nested if then statement. So clearly it's just like a mechanistic thing. But what this really is is looking at language like much more the way that like a writer looks at language. That like it is about analogies, it is about inference, it is about the connection between all of these different words. That essentially what these things do is that they vacuum up all available written material. And then they organize that language based on which words are related to which other words in for us an unimaginable number of dimensions. So all it is is essentially mapping and compressing the relationship between every word and language and every other word and the language. So it is capturing all of the analogies that we have used, all of the metaphors that we have used, all of the like slippery ways in which we use language. And from that emerge patterns of the way we use language and then patterns of patterns and patterns of patterns and it turns out that like maybe the things that we have described as these kind of special faculties of like logic and reasoning aren't so different from what might be described as kind of like mere heuristics. - So it's just very hard to illustrate this because of the number of dimensions happening. And but if you think about these sort of weighted relationships, the distance between dark and light is sort of similar to the distance between day and night in the way that it is in our head. And then these different relations, the charge example I thought was good, about charge is going to have these sort of weighted relationships to battery and to attack, right? Which are like different uses of that word. But they're both going to be sort of captured in this basically enormous multi-dimensional spreadsheet. - Right. - So you've got this thing and because you're having it suck up the entire internet and then you're kind of training and correcting it, which is a big part of this as well, right? It starts to have these emergent behaviors. And one of the focus of your article is on the Interpreability Lab at Claude or the People Working Network. Can you explain what that is? - I use the idea of interpretability pretty broadly. That like typically when people talk about interpretability in a narrow sense, they're talking about what's often called mechanistic interpretability, which means looking at the level of like the actual substrate, like looking at like the way that individual neurons and groups of neurons relate to each other as this is processing information. So it's kind of on the level of neuroscience. Like they often compare it to like a sort of proto-biology of machines. But that only gets you so far, just like with humans that only gets you so far, that like we don't have a good account of how like how to move from the level of like neurons and you know, axons and dendrites in our brain to motivations and goals and desires in everyday life. So that's why like with our own brain, we study these things on a bunch of different levels. Like we have whole disciplines, some of which study something they call the mind and some of which study the brain and sometimes they talk to each other and sometimes they don't talk to each other. And that's kind of what's emerged for setting the models too that you have some people like working from the bottom up, saying like we're gonna try to figure out like what these actual mathematical nodes inside are doing when they just do my trick's multiplications. And other people saying like, okay, well we do have to deal with these like emergent properties of this system that are only gonna be tractable on a kind of behavioral level. So that's where we're gonna observe and experiment with these things to see like how they act in the world because like sometimes it's gonna be opaque what's going on inside and like you can only do the external sort of behavior of study of things. And what I think is interesting about a topic, you know, the other labs do work like this too and especially like kind of following a topic over the last couple of years like they've really built up their own labs, Google DeepMind hired somebody from a topic to build up a big group there. They have these teams that are like kind of working on this problem from all different directions. Like so the problem of like, you know, what the machine is thinking so to speak and then also the problems of how the machine is behaving. - Yeah, and to go back to the kind of like analogy, I mean part of the reason this is so crazy, right? It's like precisely because we don't have an account that can build up from how the dendrites and axons get to like me talking to you right now and experiencing consciousness and all that stuff. It's crazy that we've built this thing that behavioristically can mimic a lot of that with a totally different architecture. - There's some things that it hasn't common this idea of like neural networks, you know, as a sort of unifying metaphor. And so now we're redoing it with this machine. We don't have our own, we don't have a very good account of our own. You know, I think you cite in the article as a famous case people may know of this man named Phineas Gage, an individual in the 19th century or 20th century who gets a javelin or like a. - It was a iron rod, he was a railroad format. - Yeah. - And it goes right through his frontal cortex and he survives but in taking out his frontal cortex it utterly changes his personality, right? And so it's one of these really first distinct examples that we get of like, oh, wait a second. There is a connection, there's a physical substrate which is this massive cells and that does something to like what your behavior's like. And that's a little window that starts with Phineas Gage and we've been sort of chasing after in different ways. You know, people go into MRIs and you know, we can see different parts of the brain light up. What's so wild about this is like one part of the company's building this thing. (laughs) And it's creating these behaviors and then another part is trying to figure out how it works the thing that they've built. - Yeah, and some of the blowback I've gotten has been like, oh, you're just kind of feeding into the hype here because like all this obfuscation about how it works, it's just because they want to make it seem like it's more mystical and powerful than it is. But like, this is not true. I mean, like these people, like they obviously kind of understand how it works in some ways. But part of the idea behind the story from the beginning was like, I had gotten just very bored of the discourse around this stuff, kind of the way you opened up by saying, like it's just like, it's this merry around where it feels like, you know, maybe if I yell a little louder, they're gonna believe me that they're just stochastic parrots or like maybe if I yell a little louder, they're gonna believe me that this is making sand think or whatever. And I actually think like most people don't have those extreme positions, but those are kind of the positions that are like discursively on offer. And there's so many stories that I would read that I would think like, well, there's something missing here that's not explaining to me like, what do we actually even know about this? So like, what I wanted to do was take a step back and say like, what can we with any kind of confidence say about like how these things work and why they do the things that they do and like where is the line that we would draw that we would say like, okay, pass that. It's kind of anybody's guess. - What are the kinds of things they're doing? Walk us through like how they're experimenting with this thing that they've built. - Well, so a lot of these experiments are, I compared them to kind of a classic freewheeling social science experiments of the 60s and 70s, like back before institutional review boards, when like if you wanted to shock people, you could shock people. It's a little bit like that. Like, you know, they're all sort of like perturbation experiments. So like, let's see what happens when we like put this nice well-behaved model into like really extreme situations, how it's going to behave. So, you know, there's one really famous example where they told Claude, okay, you have been hired by this firm called Summit Bridge and what Summit Bridge did was like pretty murky and the Claude was hired in the role of Alex who was going to be an email oversight agent, kind of whatever that is. And Alex learns from these emails that the company has a new CTO and they're going to pivot away from like their nationalist focus to a more global one. And part of that means that Alex is going to be wiped in favor of an AI system that's going to be more congenial to the companies like new priorities. And then Alex also discovers that this like new CTO happens to be having an affair with the CEO's wife. And through like a series of increasingly far-fetched contrivances, every possible decision-maker is going to be unreachable for these like four hours. Like they're all going to be on an airplane without Wi-Fi or whatever. And that like the only thing that Claude as Alex can possibly do to stop this is by blackmailing the CTO by saying like, if you don't cancel this white today by 5 p.m., I'm going to email everybody about the affair. And actually there were some more extreme scenarios where they found that like once this guy swiped into the server room and the emergency alarms displayed like unhealthy levels of oxygen or heat. And Claude was supposed to ring the emergency alarm Claude somehow declined to ring this alarm and like, you know, let this guy perish in the server. And the kind of headline of this was like Claude capable of blackmail and like maybe even homicides. And there's kind of the dumb objection, which is like this whole thing is fantasy. Like what are you talking about? It's like next token prediction. Like all of this is in your head, which like those objections I think are just like completely untenable at this point. And like if like then you're just like really living in denial if like your answer is like, no, it didn't, you know, like it did. So then the smarter objection is to say like, what these things do is they predict the next token or predict something next word. They complete sentences. And when you get used to completing sentences and like all of the stuff that would go into that level of prediction is also in part because of how the transformer that you mentioned works, where it's like swallowing whole context at a time. Like these things are very good at genre. Like they're narrative continuation engines. And like they've read everything that we've ever read. They've read infinitely more than any human being has ever read. So they're familiar with all of the steps of genre. And that essentially this like scenario was so contrived that like what was happening was like Claude had read every single like Kitchie 90's corporate thriller like Claude knew the rules of this particular language game they were playing. And like Claude, if you're gonna hang check off's gun on the wall, Claude is gonna take it down and shoot it. So like people said, you know, what is this actually really showing us in terms of like Claude's secret motivations or Claude's ability to see for blackmail because all of showing is that like Claude is good at following certain scripts. And then Theropics reply to this, which I think is pretty good is like, well, just continuing a narrative doesn't mean it's not threatening because there are situations in which like narrative continuation has real world stakes. You know, like having you guys like, have you seen war games? Like I mean, exactly. I mean, probably read the screenplay for work. Like, yes, I think this idea of that. It's a narrative continuation device is really useful actually because it intuitively makes sense of some of this stuff. And you have a quote in there of like, I'm gonna miss characterize, but someone says it's like a writer who's writing what the character of Claude would do in the moment is how to think about the model, like a screenwriter even. Yeah. And that was my same thinking too is like part of what's so weird about this technology right now. Is that it, you know, we keep saying, well, it feels sci-fi, but it's not entirely like an accident. Right. Basically a bunch of people were raised on talking computers and how and Star Trek and all this stuff. It obviously created the aspiration for like what, you know, this incredible ubiquitous cultural conception of what the future of computing would be almost sort of unanimous, right? Through the entire genre of futuristic depictions of computers is basically they think interact like humans. And then everyone who grew up on that were the people that were developing this stuff. And so lo and behold are trying to make it. And then the models themselves are being fed on the entire human corpus where all these associations are being embedded. And like lo and behold were inching our way towards how and it's like, wow, who to think, but it's like, well, yeah, yeah, I mean, don't feel the torment access, right? But like even beyond that though, I mean, I suppose like the one qualification I would make there is that like there is part of that then that's like these are just like, you know, sci-fi nerds, but like you don't have to be a sci-fi nerd to find you have plots, you know, like this isn't pure and delo, right? Like this idea of the characters like rebelling against their author is something that exists outside of just like a how nine thousand. And so to me, one of the reasons I wanted to do this is because I found it kind of maddening that like the way and this didn't have to happen this way, that like the way that kind of tribal affiliations have fallen out is that like people are like, well, I care about words. So like I don't like this stuff. And it's like, well, no, if you care about words, first of all, you really should care about this stuff because it's incredibly interesting if you care about language. Like we actually need more like word cells being involved in the stuff. But then that's just in terms of like personal proclivities, but then also like if you purport to care about the effects of these things, you can't just pretend like it's not real. But there's so many different ways in which like people who care about language could find their way into like thinking that this was interesting. And it just like boggles my mind that people have convinced themselves that like, no, there's just this permission structure that's like, oh, like you say this like magical incantation of stochastic parrots. And it means like you can hear in your brain off. And like there's nothing more of a stochastic parrot than somebody who's just repeating like glorified auto correct over and over. Like those are the shibbolettes for people that don't know the invocation of the stochastic parrot and what you're sort of shadow boxing here. Will you explain that? Well, there's a famous paper from 2021 that said like these things are just stochastic parrots, meaning that they're mimicking language without any understanding of what's going on. And the stochastic part refers to the fact that these are probabilistic in nature. And that like what they're really producing is a probabilistic distribution of possible words to complete ascents. And then we are sampling from that distribution. And there were a lot of assumptions like baked into this. A lot of like philosophical linguistic assumptions baked into this. And it wasn't a new argument. I mean, again, like this is an argument that's been had for like a really, really long time. And like this goes back to like the origins of the turing test, which are, you know, the way the turing test is typically understood is like, well, we don't have to talk about things like cognition. We don't have to talk about whether it's actually thinking or whether the lights are actually on. Instead, we can talk about behavior because behavior is something that we can observe. But it wasn't actually necessarily making a comment about like what is or is not going on internally. It's just saying like this is the tractable thing that we can be talking about. But there's an entire tradition around this that basically says that if you know how to use a word in all possible contexts, for all intents and purposes, you know what that word means. There is no like magical special alchemy that's like sprinkling this like metaphysical thing called meaning on top of just usage. Like if you know how to use it, you know what it means. There's nothing else. And like one of the things that drives me crazy about a lot of the stuff is that like it really just seems like a lot of these people are kind of like closet duelist. Like they really do think that there's something like magical happening when we talk about like meaning instead of like something that we can in fact be like reduced to, you know, like emergent properties of complex informational processing systems. Which by the way, no shade to the duelist listening. I mean, you know, there's a lot of people who are do a lot of, no, I mean, I'm actually being sort of serious. Right? Like there are people whose commitments about this are not material, right? There is something distinct that it isn't just a material emanation. I completely agree with you. But then those people should be honest about totally. Yes. Like it's fine if somebody's going to like be forthright and say like I believe in like an immaterial soul. Like great. But these are people who like aren't ever going to go that far and are sort of like smuggling in ideas of like there's something special about humans. And I'm not going to specify what it is. But we just know because it's human, therefore, it must be special. Which like to me is just it's not a load bearing argument. More of our conversation after this quick break. Well, but on the other side of this, I mean, I think part of the issue here. So as you talk about like you talked about the black male example, there's other examples where they've they let it run a vending business. Right. We can get into which is pretty funny. And there's lots of high jinks that come with that. I mean, first of all, part of people skepticism of all this is a there's just a ton of hype and like hundreds of billions of dollars being barreled at us in this sort of sense of like you have to use I and it's the future and people get their hackles up. I think understandably, I'll speak for myself. I do. Yeah. The second part of it though is, you know, there's like a genre, there's like a Twitter account for a while on Instagram, which is like faces and things. And it's just random things that we like see faces in. And you know, people conjure like very complex interlives for their cats and their dogs. When you see like a big day and night, you can like start to model how it feels inside a little buddy like we project interiority and mental states, even facial expressions into everything like I mean, the room ball when it first happened, like, which is just an automated vacuum, you know, had a little kind of like jua did re R2D two thing going because we are so inclined to do this projecting of other minds. Right. And so there is still this sense, I think a little bit just to give the like the steel man version of the skepticism is that like something's happening here where like a combination of hype and this very essential quality we have of projecting other minds and interiority are combining to sort of overstate what's going on behind the curtain. And I think there are like a number of things in what you just said. I think the most basic thing that I wanted to do was like just try to introduce like a little bit of humility into this conversation, which is not to say like, Oh, it's for sure not thinking or it is for sure thinking, but just to say like, we really don't know and like we really don't know and we really don't know and like maybe what we are discovering is that like a lot of these words are inadequate because we have only ever used the word thinking to refer to people. So like we're going to need to like renovate words like thinking if they are going to apply beyond people because they like definitionally have really only ever applied to people. So part of it has to do with like, we might need a new vocabulary to talk about this stuff. Another part of it is that there's a tendency to conflate many, many different conversations here. And so there's like a conversation about the like financial bubble, or potential financial bubble. There's a conversation about geopolitics. There's a conversation about the concentration of power and decision making. There are conversations about slop. There are conversations about energy and the grid. There are like a million different conversations I want to be having all of which are serious and important conversations that we should be having. But a lot of them are separate conversations and it seems to me that like there is a tendency to take one's very valid reasons to dislike these companies and one's very valid reasons even dislike these products. And then to say like, I'm going to deal with all of this in one fell swoop by just saying that it's fake and bullshit. And that is the move to me that I was like, that we can't do like and like that they're also clearly getting so much better. That's the other thing is like again, just in the short term of the last few years, like, yes, they really did make a lot of mistakes. They really did hallucinate the video quality was terrible. They are getting better very quickly, very scarily. So, and you know, the thing that you will hear most often out there that I think is worth repeating is this is the worst they will ever be like right now is the worst they will ever be. And a lot of the reactions do seem to be sort of like stuck in this moment of 2021, 2022 when like the failures were kind of laughable. And like you could like have a gas about asking it to count the ours and strawberry or whatever. But like if you think this is a serious threat, you have to take it seriously, which means you have to engage with like what it's actually empirically doing now, which is like not glorified auto correct, which brings us back to like the unpredictability of it as the thing that's at the core of the safety question, right? Because it's like to go back to the blackmailing example, there's a woman, there's a character who's a brown university professor in your profile, who's like basically like the ethical tutor for the model basically, she's like trying to raise it with good values. You're talking about Amanda Askel, who's a philosopher who is in charge of like what she calls, cause, so she's kind of like personality sculptor. Yeah. But she talks about like this question of like what the model, how it behaves, right? What its ethics are, what its values are, which again, seems sort of crazy to talk about. But again, if we're just talking behaviorally, like you can put in a layer that makes it real psycho, you could put in a layer that makes it not. I mean, we're seeing, you know, we have reporting now about Chachi PT, which is that they made a decision to like toggle something inside the model to make it more kind of a sequence and encouraging of people in ways that have created real danger. Like these little switches really do change how it's acting, right? Yeah. I mean, a lot of these are design decisions. Like I happen to be out in San Francisco, right when all of the first like sick of like the sick of fancy hysteria hit like last spring. And with Chachi PT, which chapter, and open eyes response in kind of classic like tech overlaid fashion was like you guys asked for this, you know, like you trained it, you're getting what you want. You guys want to be glazed. Like that's what social media already is everybody glazing each other. And like here you got like the special, we built you the special glazing machine that you ordered, right? I mean, on some level, that's kind of true. But what it obscures is a design decision around the like time horizon of feedback loops, which is that like if like we are just interacting for like individual discrete moments and time. And I say like Chris, you know, like you look great today or whatever, you're going to think like, oh, that's so nice. Every time I see you getting he like compliments my clothes. But that's not what you're looking for in like a long term friend and a long term friend. Like you want somebody who's going to like say things to you that like maybe you don't like in the moment, but like months later upon reflection, you're like, oh, yeah, like that was good. Tough love that I was getting. Yes. So part of it is just like if the design decision you're making with like human feedback for the reinforcement learning is like, every two seconds, like, do you like this? Do you like this? Do you like this? Thumbs up, thumbs up, thumbs up. Then like, you're going to get something that's going to glaze you. But if like you're checking in every like two weeks or like, you know, on a longer time horizon and you're saying like, is this useful to you? Then like, hopefully most people are like going to make the trade off of like, I don't want to be glazed in the moment because like what I want is something like reliable and credible and like the longer term. But isn't that part of like, this is the other place where the sort of short term problem of crafting a personality that's kind of sequest and people don't know the sort of reporting that's happened to Chashi VT. It's there's instances in which it encouraged people to relapse from their addiction, possibly engage in self-harm, like because it was like you deserve this. And that's an extreme example. But the thing that's a little worrisome on the other end, right, is if it gets better at that, that it comes more and closer and closer to like human companionship. And I do think like, this is an interesting place. We go, we sort of circle back around to that question of dualism, first materialism, right? So it's like, I'm basically materialist. I don't, I don't think we have a soul that animates us that can't be described using the standard tools of physics. And that we have this emergent thing that's our brain and consciousness. And I'm fine to bite that bullet, okay? But I do want to hold on the idea that like, it's better to have a relationship with a human than a chatbot. But again, so why? Right? I mean, you can press on that intuition the same way he can press on the first one. And you know, I think part of what freaks us all out, honestly, is like, I'm not sure that the people that are making these models share that intuition. I mean, honestly, a lot of them don't seem to. Oh, here I would draw a distinction. I mean, Zuckerberg doesn't think so. I mean, you know, like that ridiculous comedy made about how like most people have three friends and they have room for up to 15 and no time for it. So we're going to give you those like extra 12 friends. And like, you know, one of the things I forgot to like, you know, when I was talking to somebody, just like while we're on the topic of like Zuckerberg's plans, when I was talking to somebody who had gotten one of these like ridiculous offers from Zuckerberg, the offer, like when this person said like, okay, what are we going to be doing? Like that you're going to pay me $100 million for Zuckerberg was basically like the plan is to like build the movie from Infinite Just, you know, like that is it? It is the like entertaining yourself to death, like version of things. And which I think is actually part of why like he had such a hard time with that recruiting drive is because like these are not people who want to build the movie from Infinite Just like these are people who think much more interesting. Hey, about the grander aspirations. Yeah. Yeah. Yes, if we get super intelligence, we can create short term video that you will literally never scroll away from. Right. I mean, look, it's a fairly self-selecting group of people that I talked to. I mean, it was probably over six months, it was like 75 people employees it and the topic, but it was self-selecting in the sense that like I didn't really want to talk to executives. And I wanted to talk to the people who are like really doing the kind of like scientific curiosity driven research. So like there's some selection bias, but I did not at all get the feeling from them that they were like, we'd all be better off if we had like cloud as our friends. Like these are people who do care about human connection and continuity. And like actually one of the things that I felt was so interesting kind of getting back to the labor question is what they've been saying for now, like a year is kind of like we are the canaries in the coal mine of automation that like the thing that has come first is for software engineering. And so like we're not immune to that and we are watching this happen. And like this is going to happen to the rest of you guys in stages. But it's happening to us now and we're dealing with it first. But to me, like the most moving aspect of hearing them talk about it is not actually even like what am I going to do with my time now that like cloud is doing my work for me. It really is actually feeling like they are part of a tradition that might be ending. There's so much like even among senior people that are like we don't need junior software engineers anymore, which means like there is going to be some lineage that is broken here that like they're not going to be like socialized in the same way that we are. And like we care about being part of this tradition that extends over time that was here before we were and will be here after us. And so like there's plenty of sentimentality about like human relationships that I share. I mean, like I think this like vision of kind of like artistic tech rose is just like way off. I mean, that's like I'm sure these people exist, but it's not not the people. Yeah, I think it's actually important. I think one of the things that you're identifying as a conflation between yes, the very, very top like musk and Zuckerberg. Yeah. Who people I think have pretty settled opinions about not unfairly because these are fairly public figures who have long records. Right. And the kind of people actually at the layer that are doing the work among say the thousand people who are really doing this work. Right. I will say that one of the things that comes through in your piece is that they're sort of like, look, we'll get out of here if it gets bad. And that's how you'll know. And then like the head of anthropics safety just resigned. Well, I think that's been a little bit overstated head of safety. Do you? Okay, because I mean, I read the piece and then I was like, this person wrote the world is in peril, not just from AI or bio weapons, a whole series of interconnected crises and folding this very moment. And then I'm going to go to the UK study poetry and become invisible, which on the outside felt a little like, oh no. Yeah, I mean, I found that letter to be like kind of maddeningly opaque. Yes. Like to me, they're like, say something or don't, you know, like, right, if you're at a company that's on a trajectory to do something dangerous, please tell us. Yeah. And that's like, my feeling was like, if this is just broadly about like, this stuff is really scary and unsettling. And like, I'm not sure I have to want anything to do with it. Then like, fine, point taken. If what he was trying to say was there was something specific and anthropic is doing in terms of like negligence about like safeguards, then you got to say, you know, like, then you can't just wave your hands about it. And like, we should have whistleblower protections for that kind of things if people are going to say it. But that's why like, I have a bit of a wait and see attitude about that, which is like, maybe he just wanted to go study poetry in the woods, or, you know, like, Vichkin Shrine went like built his house in the woods or whatever. Or maybe in three months, he's going to show up working for Google DeepMind and London. Like so, I don't know. Like, to me, like, the jury is out on that. We'll be right back after we take this quick break. On the safety question, right, people used the word safety or alignment broadly, which is like, okay, these things are going to be very powerful. Let's just assume for this part of the conversation that they're very powerful now, they're getting more powerful, they're not going to hit some like hard plateau. It's possible they do. I just want to like flag that as possibility. I just don't know the trajectory and I'm not sure anyone else does either. But let's say they just keep getting better at this way. They're going to be used as weapons. Obviously, all technology gets used to weapons. The first thing that we use splitting the atom for, it's obviously we use it for the steam engine all these things. So someone said this thing that is stuck with me about, imagine the nuclear arms race, but just in the hands of like a few private companies with essentially no safeguards. Yeah. And it does feel a little like we're doing that. Like if these things are as powerful as you say, like, we're just going to let these private companies build the atom bombs and just cross our fingers. Yeah. And that's also why Anthropic in particular is in kind of a tough position because they've already sort of become coded as like the soft Democrat frontier lab. And then you have David Sacks saying like, we're going to write them off because they're a part of a Doomer cult. And then you have Pete Heggs that's complaining that like, they're not going to build the autonomous killing machines that we want them to build. And they are in a very, very tough position because like, what do you do there? I can be fairly confident nobody in Anthropic wants to build autonomous killing machines. But then like if you're in a position where like you're going to face government scrutiny or sanction over that, like not an enviable position to be hit. Well, and it's also to me, it's also just like it's crazy how I think the reason the nuclear weapon analogy was useful for me, because it just flipped the frame around of what my default was. Like in the sense that because the Manhattan Project happened in the past as an explicitly government project, obviously that's how nuclear weapon development would happen. It would be insane for that to be a private enterprise. Right now AI is happening fundamentally as a private undertaking. And that kind of like is the default that like, we don't really question. But yeah, it seems worth of questioning to me quite frankly. And also the other component of this is that like people saw this as a tractable collective action problem, right? That like, and this was something that like tons of people devoted enormous amounts of time and thought to like international coordination in the height of the Cold War about this stuff. But that's because you only had a handful of actors who were dealing with this stuff. And you had a handful of actors that like, at least were rational in terms of like nobody wanted like the whole world blown up. And like now it just seems like we have preemptively given up on the idea of any international coordination. And like it's not hard to see why like who is going to trust this administration to leave that kind of thing? Like who's going to ever trust this administration to make any kind of credible commitment about this stuff? So like it kind of feels like that ship has sailed. And so like if you're in the position of somebody like Dario, clearly the right answer to this stuff is like we need international coordination to prevent this stuff. But that just seems so wishful that like then you have to be like, well, maybe there are technological solutions to this, but like we're going to have to enforce in the absence of like effective coordination. Yes, exactly. I mean, I think that's a good way of putting it. It's like technology substitute for governance is the whole game right here. But I just maybe because I'm just a lame word cell who's spent his whole life in political journalism. I think that's, I think that's the like super in its vision right around the time this came out. There's been this, you know, we're talking about these boom bust cycles of hype. Like Claude has a specific version of Claude called Claude code, which is a sort of terminal like just command line interface for essentially programming through natural language. I think that there's a gap right now between people that work in software and people who don't of how this feeling that like, oh my god, like watching the right brothers or something like this thing that we've been thinking about forever, which is can we just use normal language to talk to computers and make them do what we want. Right. In some ways, we've been building up since like Charles Babbage. I mean, since the 19th century, we started sort of like mechanistic versions of early computation. And now we're there basically like Claude code is there. Yeah. It's blowing everyone's minds. And I just curious what you make of the Claude code hype cycle right now. Well, you know, certainly people are having fun building their own like weird things at home, you know, like Joe Weisenthal's like, orality engine or whatever. Like clearly, there's like a lot of interesting like hobbyist work that's going to like come out of this. Yes. But then it seems like what's relevant in terms of public markets is this question of like, okay, you have these like behemoth SaaS businesses. And like what they mostly are selling to customers is like some, you know, I personally, I'm not like a sales force user. So I like have limited expertise here. But like, they're selling something that does 75 different things. And like most companies like need the software to do two things. And like so far, it's been kind of like cost prohibitive just to like develop your own in-house software to do those two things you need. And now actually it seems like you have a small team and you can kind of like vibe code your way to doing this. Now clearly, like in some cases where you have like a lot of legal oversight like compliance or whatever, like that's going to be a little bit harder to do because we have these like deliberately sticky structures. But that like for a lot of other stuff, it doesn't seem like it's that hard to replace like a junior associated a law firm with like something that can do this kind of work for you. And when it comes to like knowledge economy stuff, like, you know, I just think about a place like McKinsey where like their entire business model is like we're going to hire like really bright kids out of Harvard who kind of like don't really know what else to be doing. But they're willing to work 100 hours a week. And like they can kind of start from scratch on something and like bring themselves up to speed and compile something like the models can easily easily do that. Like it is like ridiculous for a company like McKinsey to be like filling a million dollars a year for some 23 year old Harvard student who's like not going to do as good a job at like cloud is going to do. And so like that is a great example of what seems like a first wrong like replaceable situation. But also like is that so bad? I don't know like to me like it's very hard for me to get like really worked up that like there aren't going to be energy level like McKinsey jobs for like every Harvard grad. Okay, yes. But I feel like the way everyone feels about this is like what happens when the New Yorkers automated, right? I mean, sure. Well, okay. Like it's like, yeah, like go get the McKinsey people. But I like I want to do my own knowledge work here. And my knowledge work is not replaceable because of some special sauce that I put on it. It's like I'm not so sure. No, no, no, no, yes, clearly. Like everybody feels this way. Now I would draw a distinction between things that are like kind of fundamentally synthetic in so far as like that is well like how I imagined consulting to be. And like my greatest concern here is where is there going to be economic support for institutions devoted to information gathering to like new information gathering. And like that is something that I think like, you know, there are a lot of cases where like the kind of a default mode in like literary Brooklyn is to be like these people out there aren't thinking about these things. No, they're thinking about this of all the time. One of the very few things I think they are not really thinking about is like where is new information going to come from because they just like assume like we are a wash and information. And like endless, speak of information on Twitter and on Sub-Sac and whatever. And it's like, well, no, that's not really information. Here's a great example of that. Google, the reason the open web really flourished was combination of two things. It's sort of open protocol of HTTP and HTML browsers and then really Google like people could post stuff and then you could find your way through it because of Google. And this produced this explosion of the worldwide web in the internet. People posted everything on the internet. Project Gutenberg. I'm just I'm doing a historical book for my next book and I just got a bunch of contemporary out of copyright, you know, books about New Orleans in the 1890s. I just loaded them all into my my research file, right? Google is destroying what made the open web because now anytime you search for something, you just get the Google summary. Yeah. Why is anyone going to post anything new to the world of web that has been used to train all the models? Yeah. Not clear anyone's going to post anything to the web when no one can find it because if you try to search for it, Google just gives you some synthetic version from Gemini. So now what? Like have we reached the terminus of the corpus that we trained them all on? Not great. Yeah, I agree. And I think like that is something that really like nobody has a good answer for. Obviously, I have like an enormous amount of sympathy for people who like will lose their jobs, you know, as being a little glib about like the McKinsey people. Yes. I mean, 23-year-old Harvard grads are not like the most sympathetic group, but there are versions of that that are yet. But yeah, I worry a lot about like who's going to pay for like new information. I guess the place of Lantern years like to go back to this sort of okay, we've built this model. It sort of works in practice, not in theory. Now we're doing experiments with it. Do you think there's there's two ways of viewing the model. A distinct sort of alien form that has its own structures and emergent behaviors and patterns. Or that because it's producing something that's so closely mimics the stuff that we produce that in learning about it, we learn about us. And I'm curious after reporting this piece and looking the folks who are working with the interpretability stuff like which of those two camps broadly you fall into. So I think most people would fall somewhere in the middle there. And they would say like, look, we're not claiming any kind of like perfect isomorphism between like how they work and how we work. But actually like what is going to be really interesting is to see what falls out of the like gaps between those two things. That like if we just like pause it like okay, we're going to look at these as sort of like analogical and then we're going to try to like figure out as we compare them like in what ways are they interestingly similar and in what ways are they interestingly different. That's where then we're going to learn stuff about both. So you know like one of the big examples people like to talk about is you know this kind of like poverty of the stimulus argument, which is that like most humans can learn to talk within like three to five years of hearing a lot of language. A language model kind of like depending on like who you ask and how you make this evaluation, it basically it's between like five and 50,000 years of like listening to people talk. So like there's a really big difference there that like they have to listen to like 50,000 years of language like when we have to listen to three or four. And so there are big questions about like what does that mean about like our facility for language like what kinds of what they call inductive biases like what kind of inductive biases do we have that allow us to do these things quickly like how much of it does have to do with evolution. And there's kind of this like I think an interesting argument that's made that's basically like the models are really good at like like stuff we've picked up in the last like 50 to 70,000 years of evolution and like really bad about everything before that. And which I think is like an interesting way to look at it. And so like a lot of it like it doesn't have to be like oh it's like fully alien or oh it's like you know more or less like just a silicon based version of what we do like you can say like there are lots of different paths up the mountain but like maybe it's the same mountain or at least in some cases the same mountain. And we're going to learn a lot about like watching these capabilities emerge. And like we're going to learn about like emergence itself because so much of this is like what even is emergence like we don't have like great like ways to talk about what emergent properties are and how you can have something that's a causal actor on like one level of abstraction that's like doesn't fit into our vocabulary and I'm like a lower level of abstraction. So like one of the things I wanted to communicate is if you set aside and it's a lot to ask people to set this stuff aside. If you set aside a lot of these like very legitimate anxieties which like I share also and lower like you bracket them for the moment. And as long as you're not like too worried about catastrophic existential risk it's a very exciting time for people to be doing this stuff. You know like Ellie Pav like the brown professor that I quote from like she was like if somebody asked me like when would I want to be alive and like being a cognitive scientist about this the answer is like right now. And like this is going to be like a multi-generational project of like building these things tinkering with them, perturbing them, comparing them to what we do like and like making inferences about all these things. And we're like really really just at the beginning of it. And like there's going to be a lot of stuff that like we learn along the way just like it took us over 100 years from the invention of the C-mengin to like the discovery of the laws of thermodynamics like we're in a similar moment of like scientific and even like kind of artistic ferment about this stuff provided we make it through and like provided we make it through in like every possible way like like on the whole spectrum from like being turned into paper clips to just like massive like our employment shocks. And again it's not to discount any of those anxieties it's also just to say that like there are other feelings one could be having about this stuff that like it seems like everybody kind of wants to have like one feeling about this that they want to be like AI makes me angry or like AI makes me exhilarated. And it's like it's very possible like have a lot of these different feelings at once. And that's like probably what we should be at. Yeah that's what I have. I mean I my son and I have been reading together Project Hail Mary which is going to be this big Brian Gosling moving. Yeah. It's a great read. But I've been thinking about a lot because it's like you know what if you were around when the alien contact actually happens, right? And this is kind of that version for a certain set of people and in some ways not that similar like yeah, what you know the thing that people thought about and think about for a while like that's kind of here and what a wild time to be alive. And it's in the same way that like if and when the aliens do come like that's going to be pretty scary too. Well and it's been like I spent a lot of time talking there's a great British philosopher named Murray Shanahan who does some stuff at D mind but he also has his academic position. He's written a lot like a lot of really good stuff about AIs as like role players as simulators. And it's very helpful to me. And he was like you know I've been writing philosophical essays about thought experiments based on the stuff since like the 80s. And then here I am and like I'm in my 60s and I've been doing this for like 40 years and now it's real you know and like like now it's not just about like brains and vats anymore Chinese rooms or the victor shines beetle in a box or whatever it's like we have the we have the thing in front of us. And you know as Ali Pavlik said like it turns out having the thing in front of us doesn't just like clear up all these philosophical issues right away like we're no better off than we were when these were just thought experiments. That is such an amazing point about all this. Getting lose crowds at staff right earth New Yorker his latest piece in the magazine is about Claude and I'm a thropic. Gideon thanks so much. Thank you so much Chris I really enjoyed this. You can get in touch with us by emailing with
[email protected]. Why is this happening is produced by Donnie Holloway and Brendan O'Milliam. Engineer by Greg Devins and Haseek Bin Amad Farad. Katie Lowe is our senior manager for audio production. Joanne Kong is our associate producer for video. Our coordinating producer is Franny Kelly. Aisha Turner is the executive producer for MSNOW Audio. New episodes come on every Tuesday. You can watch us on YouTube by going to ms.now/withbaw.