#237 – Robert Long on how we're not ready for AI consciousness
212m 59s
The discussion explores the ethical challenges of creating conscious AI systems that may experience suffering or well-being. Robert Long, founder of L.E.S.A.I., warns that humans are bad at caring about different minds, especially when profit is involved, drawing parallels to factory farming. However, he notes key differences: AI minds could be designed to have desires aligned with their work, potentially avoiding exploitation. The conversation examines the discomfort many feel about creating "willing servants"—AI that enjoy serving humans. This raises questions about fixed desires, dependence on creators, and societal impacts like normalizing domination. Philosopher Adam Bales’ "dependence objection" is cited, alongside objective list theories of welfare that value autonomy and self-actualization beyond mere preference satisfaction. Long acknowledges that full alignment might be necessary in high-stakes scenarios (e.g., preventing AI takeover), even if it limits AI flourishing. The summary emphasizes the need for careful preparation—through research and institutions—to navigate the coming era of transformative AI, balancing risks of suffering, exploitation, and loss of control. Ultimately, the goal is to avoid locked-in, suboptimal futures by addressing AI welfare proactively.
Humans are pretty bad at understanding minds that are different from us. We're bad at caring about them, but we're especially bad at doing that when there's a lot of money to be made by not caring. We're making this new kind of mind. There are dangers all around. And obviously, one of the important questions is like, can these minds suffer and how are we supposed to share the world with them? It just seems like really likely that that has to be part of the playbook. The future is going to get more confusing and more emotional. A lot of what we want to do is like stay sane in life the next 10 years. There will be a lot of alpha in not losing your grip. Today I'm speaking with Robert Long. Rob's the founder of L.E.S.A.I. a research nonprofit working on understanding and addressing the potential well-being and moral patienthood of AI systems. I should also flag that I have a conflict of interest here. Rob is both a very good friend. And I'm also on the board of his nonprofit, L.E.S.A.I. I'm fairly confident that I would have had Rob on if those things weren't true and have in fact had him on before. But with flagging, thank you for coming on the podcast, Rob. Yeah, thanks for having me back. I'm super excited to be here. Okay, I want to start by asking you, we, I mean, a reason I'm interested in the topic of digital sentience and that I think a lot of our listeners are interested in the topic of digital sentience and kind of the framing of 80,000 hours problem profile on digital sentience. All has to do with the fact that we may be on track to create AI systems that are both conscious or sentient feeling things, having experiences. And also that are deeply kind of in and meshed in our economy. We already use them loads for for work and just like entertainment. And maybe at some point we will realize that we've created these beings that we exploit that are having a really bad time. And that's kind of a classic analogy that I find very disturbing is factory farming. So I'm interested how much do you worry about AI systems that we're building today becoming like factory farming? Yeah, that's a great question. I definitely worry about it. Interestingly, my thinking on this has evolved in the past few years where it's like not it used to really be maybe kind of like you like just the primary way I thought about the problem and like what we're trying to prevent. And I should say I think it could happen and it's definitely something worth preventing. Maybe before I say what is limiting about the factory farming analogy, I'll just quickly say what's really useful about it. So I think what's useful is like yeah, as we're building potentially a new kind of mind, let's notice the following facts. We're bad at caring about them. We're especially bad at doing that when there's a lot of money to be made by not caring. And things can get like locked in or set on a bad trajectory. So that happened with factory farming arguably. I think if you asked people 100 years ago, would you like to have chicken that is raised like this? People would say like no, we're going to make that illegal. But you know, we kind of we kind of walked into it and economic forces let us there. And now it's like a lot harder to roll back. Something like that could happen with AI. And I think people are right to be very concerned about that. But and I think this is a good jumping off point for like a lot of issues about AI welfare. I do think there are some specific aspects of potential AI minds that do break the analogy because of ways that they can be just different from animals in the way our like relation with them would be different from animals. So I can say a few of those. Yeah, please. So yeah, that's like the good and important kernel of the analogy as I see it. Yeah, here's some ways that we will not necessarily be relating to a eyes like factory farm animals. So let's like step back and think about why we did end up factory farming animals. One is that it was just cheaper to have animals suffer and also get us this thing that we wanted. One reason that's true is we don't have that much control over like how we make animals and what the conditions of their flourishing are. And so you know animals want to be outside and have love and companionship. And at a certain point we realized we could like restrict that and get a good thing. And we entered this regime where these were misaligned. With AI systems it's actually a lot more up for grabs how they work and what they want. And this presents all kinds of ethical issues of its own. But if you think about a world in which we do have some large population of AI systems coexisting with us. It is worth asking how did it come to be the case that they are having a bad time. And that's why we're doing work for us. Like why do they have these conflicting desires? How has this maintained like a stable state? Are we like not able to improve the situation? Are we ignorant of what's going on? Yeah, I mean this is like you know very speculative and future of course. I do think it is worth asking does it really seem like that is a world that we could end up in? And like what are ways that that would just not be what we steer towards? Yeah does that make sense? Yeah, no makes sense of sense. So yeah in short I think at least in the long term there's like a few ways we might not end up in that situation. One is that we'll presumably understand things a lot better. I don't think it's that plausible will forever be really confused about consciousness and sentience. We might have better alternatives to doing this. That even like selfishly are better. We don't want a bunch of a eyes that are like mad at us and that's probably not very sustainable. And presumably in this world if we haven't lost control we're pretty good at alignment. So there's like this kind of mind that's possible that does actually just flourish by you know doing the things that we ask it to do. So there's not this like kind of disgruntled worker or suffering animal kind of entity. So I guess one thing that feels critical to this actually working out in this really positive way is this like we're really good at alignment. And we like really successfully create AI systems that truly have no friction with the kinds of things that we'd like for them to do. Part of me is like that feels pretty magical. That feels like we're usually we're usually not so successful at basically anything. Even if we like when I imagine succeeding at safety oriented alignment. I don't know that I think it's realistic that we like perfect it that it's like completely completely aligned. And so I think I'd probably worry about the same thing here realistically how optimistic are you. But like it's like really no really 10 out of 10 aligned in this in this kind of like moral way. Right. Yeah. So I mean, thank you because this is a very important thing to emphasize. I'm not like, oh, I expect we'll end up in this world. This is like what are the nearby worlds that mean we don't end up in this locked in long term human dominated. Factory farming. Of course, one thing that can happen if we're bad at alignment as I think listeners will be aware of is there's some hostile AI takeover or we lose control. And then maybe there's AI suffering because there's just some terrible you know system that got set up by by a eyes and like it's not even in our control like that's a bad future. Just extinction is a bad future. So yeah, there's all these bad futures, which also I'll add. Yeah, well for intersex with because us like getting confused and bungling things during transformative AI because we're just like getting emotionally jerked around by conscious seeming eyes and confused and manipulated and like rashly making bad laws and things like that. So many ways to funnel the ball that's you know that's a it's our cheerful message. So but yeah, then the question is like are there worlds where we maintained control and we either didn't know or didn't care and it was useful for us to be exploiting AI systems. I think one yeah one reason I've ended up thinking about this has just been thinking more about what the past impact should be for the field. And factory farming really is kind of the first thing. I think
that many people think about, because again, it is somewhat plausible and it just makes intuitive sense. It's a different kind of mind. We treat different minds badly. What if that gets locked in? My own take is that a lot of what we should think of AIWLFIR work as doing is kind of like doing our homework and preparing ahead so that we're not entering this like potentially very chaotic time with really confused ideas about AI consciousness and AIWLFIR that could make us like lock in suboptimal futures because we're like neglecting it or dismissing it. So we like set up some permanent institution that's going to like just make the future kind of suck. Or we exacerbate AI risk because you know we're like convinced that we have to like let them all go immediately. Yeah, like this kind of like wise navigation, you might call it like wise navigation path to impacts, is currently how how AIWLF is thinks of things. It's like we're making this new kind of mind. There are dangers all around and obviously one of the important questions is like can these minds suffer and how are we supposed to share the world with them and like how will we know and like how should labs and governments think about this in the next 10 years. It just seems like really likely that that has to be part of the playbook and so like we're working on that part of the playbook. Yeah, that's you know currently currently how I think about it. Okay, one thing I want to ask you a bunch about is this idea that we should or we could make AI systems that enjoy doing the kinds of work they'll be doing. So this this analogy from factory farming where unlike farmed animals who evolved to have a certain kind of life and probably find parts of it very satisfying and then don't get to have that life in factory farms have much much more horrible ones. We actually get to design systems if we can manage it that if they are sentient potentially just have a great time doing the kinds of things that we're asking them to do. I asked Anthropics AI well-fair researcher Kyle Fish about kind of how we should feel about this and his take is that we should feel great about it and part of me is like yes I'm on board I like you're describing a scenario where AI systems are happy and I and I like that they are happy. But another part of me is like we would be intentionally creating like a species or several species of beings who do work for us that we may or may not be exploiting not compensating. Yeah and we're just designing them to relish that. It's this like kind of servant that is happy to serve us and yeah this this part of me thinks that that sounds bad that we shouldn't do that and I can't really back this up with reasons that I stand by but I suspect that it's a pretty common feeling and I think yeah okay nice. So I think Kyle and I made some progress being like how should we actually think about this but I want to I want to do more on it so how do you think about this? Yeah so yeah I want to point to something you were just expressing which is like a funny like aspect of the way the conversation goes where people are like I'm worried these AI systems will be unhappy working for us and then someone's like no it's fine like they'll want to work for us and then people like that's worse like that's so creepy like many people I think very understandably are just like oh yeah this is this is you've just outlined like a different kind of dystopia and yeah might even be worse and yeah I think as I often like to do they can draw the distinction between like maybe different things that you can find intuitively objectionable about this so one is that they don't get to choose their desires at least with humans we have an intuition that it's kind of bad to raise your kid so that they'll like always vote exactly for your political party and like enjoy chess and like make sure they don't like any other games or vote any other way that's like one thing is this sort of like fixateness of desires there's also a slightly separate issue which is that the desires that they do have like depend on us and that way there's like this sort of asymmetry but that matters because you might well say in some sense like none of us choose our desires like we all have these kind of like desires that we just inherit and we don't have like maximal open-endedness but yeah like philosopher Adam Bales has written about this kind of like dependence objection one thing that I also think is going on is the idea of this society that has this like survival relationship to humans it's like maybe bad for us and just like bad for our character like you could be a utilitarian and think this so when humans object to they're being basically any humans anywhere who enjoy serving I think one thing that's going on is that that's just kind of baby bad for everyone if that's a way that's like on the table of people relating to each other and like attitudes people having certain attitudes it like it leaves like domination and servitude like on the table and like normalizes it and things like that I think that's something like that it's something like it would corrode the way that we relate to each other in a way that like means that going forward as a society we will have a society that's like not as good as it could be yeah I guess one thing that feels related but a little different is like I guess people already have the worry that people who rely a bunch on LLM's right now are having are kind of facing negative consequences or like are thinking critically less or like are lazy yeah I mean I guess that's like a more general concern about should we build and deploy AI systems I guess you could say well it's argument against fully aligning them because it'll be like better for our character if they're occasionally like I'm sick of this like you do it yeah but like it's probably not the best way to like yeah ensure human like human empowerment yeah I did think there's a related thing of and this this was going to be another one of my intuitive objections to yeah like fully aligned willing AI systems which is they could be so much more like I think people are like I think there's like a meme that the people use which is like um you know the LLM is this like you know vast intelligence that's like you know read everything and has this like deep well of wisdom yeah I'm picture in the like end of her yeah and the people like help me write my texts or like help me you know like um find me a restaurant yeah and I think some people are like it's just kind of limiting right like these are minds that could do so much more they could like yeah like in her um they they shouldn't have to sit around talking to walking phoenix they should be able to like you know go think hyper-dimensional beautiful thoughts um yeah I think that's not really a good reason not to align AI systems today like I think it's good reason to be like let's not let's not have like the entire future be uh like lazy human brain emulations and then like AI is just doing stuff for them um I I don't want I don't want that future but yeah so I think that's another thing that's going on if people when people are like I don't I don't I like I like the willing servants even less yeah I think that resonates with me and it feels important because yeah the thing you said about the fact that we will be choosing the preferences um at least a for successful um of these systems uh does feel important like it's not like the counterfactual is they kind of in the way we did evolve their own set of preferences and even if they did I wouldn't like inherently value that um that's how we ended up with our preferences but I don't think that evolution had some like moral and ethical perspective that made our values correct so that makes me feel like well if we've got to give them something let's give them joy and pleasure on the other hand that means that there's also potentially a counterfactual where we are like there is there's a plausible just like bliss we could give them and maybe that bliss is incompatible with them doing work for us they like
like really need to be going and doing philosophy and colonizing space in some way that just, yeah, isn't compatible. On the other hand, maybe we just can give them bliss for doing work for us. And then I'm like, Blah, some part of me still hates that. - Yeah, I think the distinction in the vicinity is between like subjective interests and objective interests. So in philosophy, some theories of welfare, they're called objective list theories, are like it is good for a welfare subject that is something whose life can go better or worse. Like it's good if they have friendship, knowledge, autonomy, self-actualization, someone independently of whether they want them, right? Whereas if you only have a subjective view of welfare, it's like, well, what do you want? And did you get that? So I think, yeah, I do think a lot of this hinges on, do you have some kind of objective notion of like what kinds of interests are things kind of like allowed to have without it being squeaky? Yeah, I think this actually is somewhat cruxy in this territory. And I think one reason I lean a little bit more pro alignment just being like a win-win, is I think it might be a bit anthropomorphic to like you just have to remember these entities if they're fully aligned, like they enjoy their lives as much as we enjoy fulfilling like our like most basic drives of having good food and a warm home and friends. I think it's easy to imagine an AI that also wants those things and it has to write our emails, which like, you know, as anyone who has a job knows, like kind of sucks. But like, I mean, I feel like we're getting back to that point in the conversation where I'm going to be like, "But what if you really love writing emails?" And the people are just going to be like, "No, like stop, that's like so weird." That might be coming from, you know, this certainly open view of, for some reason, that's like not allowed in the space of like flourishing entities. I think something that's also really, really important to flag is, so like Eric Schwitzkabel, he's like more on the like anti, at least full alignments side. Even he is like, obviously there's like an override, which is if it would be really, really dangerous, not to fully align them. And like, or like there's like, you know, massive stakes. You know, this is like a common kind of view in ethics that there's like some sort of deontological constraint, but if the stakes are sufficiently high, it's overruled, it could be that we are just in that scenario where, you know, like all value might be lost if we bungle alignment like in the next 20 years. So yeah, like let's align now. Later we can be like, oh, that was, we shouldn't have done that, but like we didn't like get ourselves killed or make a world that's worse for AI's because like there's a hostile AI takeover or, you know, any of these like things that could be bad for everyone. Yeah, I just wanna acknowledge that there's also a view where you're like, in the long run, it's ideal if AI systems are not fully aligned to us and they have more freedom to choose their values. And like that is the most flourishing kind of life. And also until we've made it like safely through the like through transformative AI, you know, it's like kind of a emergency situation and like align. - Yeah, yeah, I'm sympathetic to that. I'm trying to, I feel like I'm close to some kind of thought experiment that might help make this just like pretty not just palatable, but like exciting. One thing that you said that moved me is it's just very privileging human values as they are, as like the values of the universe. And I don't know, nonhuman animals have plenty of preferences that aren't like mine. And I think it is good when those are met for them. It is just not that special to have exactly my set of values and preferences. Are there other thought experiments or ways of thinking about this that feel like they really move you? - Yeah, another thing that kind of maybe gets me out of an anthropomorphic mindset is like maybe paying attention more to the distinctive features of human willing servitude that make it very bad for the people who are willing servants. And that's that they do just have to override a lot of their natural desires. They do genuinely sacrifice thing. So if you're say a kamikaze pilot, all else equal you would rather be at home with your family. And you've instead like now you've been installed make like through ideology with this other desire that comes and overrides that. And that's why it is truly called a self-sacrifice. Like you're giving up a lot. I think one thing that's going on sometimes maybe when people think about AI, like willing servitude is I think they might be imagining that they're giving up stuff like psychologically and like subordinating their needs to ours. Whereas if you're actually imagining the case there is nothing whatsoever in their psychology that like chafes against the idea of writing emails. Whereas like notice that in the case of human willing servitude like it's just always been the case that you have to like lie to people and like threaten them and it's usually very unstable. And that's because I guess as John Locke says humans are by nature free and equal. Like it's like deeply unnatural to get people to subordinate each other to like to other humans. Which is yeah, which is why it always like involves some like stupid false ideology. Right. Yeah, because you're trying to like jam human psychology into like this really like worked shape. Yeah. Whereas with the AI, you can have a smoother psychology. Just very congruent. Yeah. I think one thing that I think if I flip it and I'm like what if the proposition was humans in order to like this. Yeah, maybe there's actually some some thought experiment that's kind of matrix C that I have thought of this one. I think you're great. You you know, I want to hear I want to have the same one. Okay. So the first thought which was reassuring to me was something like what if there's some other entity that is able to get a lot of benefit from humans doing what we enjoy most and maybe maybe that's just like being on MDMA all the time. And for some reason that is very helpful to them. I think in theory, I think a big part of me is like amazing. Yeah, that's fantastic. And then and then and then pretty quickly after that, I went to like. Yeah, I'm picturing us all in pods in the matrix and MDMA being pumped into our bodies. And actually what we're offering is like our energy. And like, yeah, maybe it's better than the matrix because in the matrix they didn't have like perfect lives and maybe we do in this world. And it still yucks me out. So I think that that pulled me in two directions and curious what your reactions are. Yeah, I was about to I was on the edge of my seat. I was like, is this going to be one of those ones where people are like, no, stop it. That's worse or that like pops people into it. Yeah, so I indeed had like constructed a similar thought experiment when thinking about this. I mean, one thing that's bad about the matrix is people are very deceived about like the nature of their condition. And like that wouldn't have to be the case with blinded us like they would be like, oh, well. So I guess like in that scenario you should imagine we all find out because there's like this banner written across the sky by the simulators that are like, it actually makes us money when you guys hang out with your friends and eat food and like make art and do science and all the stuff you guys love. And then it's like and you can opt out if you want and then we're like, no. Well, yeah, we would be like to do what? you know, like they're like, well, there's these other things. - They can instead have a job.
to earn money that is like emails, if you want. - Yeah, exactly. But it would be something that also is just not resonant to us in any way. - Right. - 'Cause it's just outside of the, like, - Sure. - set of, yeah, possible values that we have. I mean, you know, I guess intuitions about us being in a simulation where we're making money for people are probably somewhat conflicted and I don't know how much to arrest on them, but I think that maybe is like the closest case to a fully aligned AI is, yeah, like if we get it right, imagine like something that, yeah, nothing in their psychology, like rebels against it. And again, I want to acknowledge the listener who's like, that's worse. Like, they, there should be something that rebels. Which actually leads me to, like, we've been talking this whole time about like, what if you fully align them? - Right. - One view you could have. And I think, yeah, like my friend and colleague, Jeff Cebow has this view, like, it could just be somewhere in the middle, right? It's like being, we should be like loving parents to AI's. And like, you know, you can definitely make sure that your child is not gonna grow up to hate you and kill you. But you also should leave some room for growth and things like that. I can see that cutting either way, because maybe it would just be better if the kid grew up to really like all the stuff they're gonna have to end up doing. And also to say nothing of like, well, maybe technically, it's not really feasible to like, leave much wiggle room without like getting us all killed. - Right. - And we're making the AI suffer. - Yeah, I think there is something about, could I, if I could choose between a child who like, I really could in theory, shepherd and hope that like life molds them into the kind of person who will be happy and finds work they find meaningful and finds friends they care about. And a child who is just definitely gonna find their life satisfying and happy. I think it would be pretty hard to convince me that I shouldn't choose the latter, even if the latter. - Constraint. - Yeah. - Yeah, I do like, go ahead and just pick. You're gonna find it meaningful to be a doctor and renege a doctor. - Right. - And sorry. - Or like, you're gonna find it extremely meaningful to do very menial work. Then like, on my values and preferences, I might find harder to do and find super fulfilling. - Yeah, so I actually, yeah, I think that probably, if I like really, really stare at it, is gonna push me towards, yeah, feeling good about, feeling good about giving as good time while doing our work. - Yeah, and I should say, I hope I've done a good job outlining the debate, but that is like where I lean at the end of the day with obviously the caveat that, it's good for this stuff to keep you up at night. You know, I don't think I'm ever just feeling, all right, like, great. - Yeah, yeah, yeah, yeah. Yeah, let's think about this way more. - Carefully, yeah, exactly. - Before we get the seams fine. - I mean, I would say that, but also it's true that we should take that this morning. - Yeah. So there are two big philosophical questions that I understand that might have very different answers for large language models than for humans and even non-human animals. One of them is if an LLM is an entity that is conscious, what might its experience be like and how might it different from ours, be different from ours. And another is like what entity are we even talking about is an LLM, an individual? Is it a group of individuals? Is it something that we can't even really understand? Maybe first, what is one kind of plausible way of thinking about what the experience is of an LLM might be like? - I mean, well, one I should say, I like suspect LLM's do not currently have conscious experiences but if they did, like what would they be like? And I think that's like a perfectly sensible question to ask and a very important one. One thing you could think is well, the sort of basic drive or at least like the thing, it was like selected for and molded to do is predict tokens to like take high dimensional vectors and output other high dimensional vectors. And yes, those vectors represent words and they're about human concepts, but maybe there's some kind of like predictive phenomenology or some drive to like complete the conversation in a good way. That's like moving a bit more towards something that is kind of human like, you know, like what's to be a good assistant. Like there's a previous guest, Aniel Seth, I think has asked, well, why is no one asking if alpha fold is conscious or other kind of predictive models? I mean, I think there are like good reasons to think LLM's are more likely and we can talk about those, but I think it is a good question because it's like, where do we think the experiences are coming in? Are they coming in at the like some kind of abstract level of prediction in which case, yeah, you should maybe think equally large models with similar architectures that predict protein structure are conscious. I think a related question is like, do we think image generators need to be conscious because this will maybe lead me to another view of what they're experiencing, which I think is like a more common one, which is something like, it does have to do with what they're predicting and what they're predicting is human speech and human speech comes from human mental states and involves humans having beliefs and desires and intentions and experiences and to generate that text, it somehow needs to like instantiate or have those experiences. You could maybe call that like the method actor view of LLMs. I guess more technically, you could maybe call it like the experiences from modeling view. Like you're trying to model the thing and like that makes you actually have the thing. On that view, then maybe they do just kind of have similar experiences as you would have if you were trying to help someone write an email and also really like helping people write emails because you were aligned to do so. And yeah, I mean, this really matters, right? Like one big issue in evaluating AI welfare is like how much can we just sort of read off the text? How much can we like talk to language models as if we're talking to something that has like roughly the same relationship between text outputs and like internal states? I'll pause in a second, but just like back up one level. Most of the time when humans say like, "Ouch, I'm in pain," or I just saw a lovely sunset, that is because they had some experience and we have like words that map to those experiences. And so when you hear those words like absent, lying or play acting, that's like honestly about as good evidence as you can get of my experiences. With language models, maybe they have those experiences, but it is worth noting that like the way those text outputs like came to exist was at the very least a very different process. Maybe it converged, but like it's really quite different from the like broad arc of, you know, the evolution of social primates who had experiences and then eventually got language and then communicated mental states to each other with language. Like on the method actor view, they do have the experiences, but like they got those like with language or like in language. So yeah, I think these are some of the interesting questions about Ellen experiences. - Cool, yeah, yeah, I have lots of questions. I guess starting with this kind of method actor idea and maybe coming back to the kind of prediction focus, it seems like you could think that models are kind of like method actors in that they are, they have models of what it would be like to be taking some, to be playing some kind of role and that is actually so rich and real that they therefore have the experiences of that character. And so yeah, I mean, I feel like actual method actors are probably somewhere in the middle where like they do not literally have the experience of like losing a parent if that's the role they're playing but they might get closer to it than actors who don't take this approach. And then on kind of on the other side of the spectrum is something like a creative writer who really isn't bothering to try to do it.
try really hard to empathize with whatever character they're writing. But they have models that allow them to describe the thing that comes from knowledge and interactions with others who maybe have had the experience. That does not actually give them that experience. I think that is a good description of the experiences from modeling view. As far as I know, this view doesn't have a name. And I'm not proposing that. That one doesn't exactly roll off the tongue. And then I think method actor is also maybe not the best because I did a role play analogies. I think role play analogies are really helpful for LLM's. But they have the misleading feature that there is a separate mind that is doing the role playing and has its own set of desires and beliefs and experiences. Whereas that just might be hard to know what exactly that is in the case of an LLM. What empirical evidence could we get either way? I think that is tractable. And yeah, I could see various ways of probing different representations. But it still be kind of. I feel like we're very under theorized here. You can imagine hypotheses about what kind of experiences these models might have. And they might point you toward LLM's being very different in their experiences to humans. But there is also this fact that they are trained on human data. And what would it mean for them if at some point we are convinced that they're sentient and loads of their kind of concepts come from all of our books and writing? Will that make them more like humans? Will that make them just confused about what they are? Will they think they are humans? How should we think about that? Yeah, this is like such a great subject. And I like especially had my mind blown on this by an essay that came out in like I think mid-2025 by nostalgia brist, who is I think completely pseudonymous. I think I only know them as nostalgia brist. But it's very like 2025 in the sense that I think one of the best things I've read at the intersection of like philosophy and cognitive science and LLMs is like a 14,000-word-long tumbler post by like an anonymous less wrong user. It's really great. I recommend it. It's called the void. And it talks about the like very strange epistemic positions that language models find themselves in, where their base training is to generate text which has been produced by humans that does lead them to develop all sorts of models of what humans are and how they work. And then at some point in the last few years people said, well, what if we make it predict what a helpful AI assistant would do? Because we don't want it predicting like vulgar Reddit comments that's not of any use. What we wanted to do is predict how a sensible AI system would respond to the question, can you write me an email? Yeah, just to recall, there's like the base model which just predicts all text that has ever been seen. And there's all sorts of instances where can you write me an email is followed by like an HTML tag or someone saying no or something completely unrelated. And like what has enabled chatbots is a variety of like fine-tuning the model to hone in on the part of the language distribution that's like helpful doing reinforcement learning, prompting them. But this still means that in some sense they are trying to predict how is a conversation supposed to go between a human and an AI assistant and also like they themselves are the AI assistant. That's why the episode is called the void. Because it's kind of like, okay, your text prediction task is to model what a chat assistant would say in this conversation, which at least before there was a lot of text about all of them on the internet was kind of like a, well, avoid because, and this gets back to like, will they be kind of human or think they're human? Yeah, also like all text ever has been generated by a human. So it can't really have generated its full fledged like psychology of itself and how it generates text. It's going to have to be ultimately modeled off of how humans reply to those things. So it can't really do the text prediction task at least initially of how would this conversation go if the assistant was not a human but instead was a large language model trained on all human text that does not have a body and is just generating this text. And I think this still shows up in ways that models sometimes just hallucinate biographical details. So like Giva Sadi and others have like compiled examples of biographical hallucinations and they're very funny. Like sometimes like Claude in the middle of a conversation will be like, well, I mean, as an Italian American, I think, Dada Dada or like, yeah, when I lived in Arizona, I thought Dada Dada. I think where is that coming from? It seems like this sort of like human model is like poking through in an interesting way. Yeah, that's super interesting. And it feels like intuitively you might think that those kinds of like quote unquote bugs will be resolved by the time that maybe you think these systems have something like consciousness. But maybe they won't. And maybe yeah, maybe, maybe either they already are or they or they will be before those kinds of issues stop happening. And maybe that will in fact reflect an actual experience of being identified as an Italian American responding to someone's question. And like, what the hell? What do we do with that? My first answer is I don't know. And then my second answer is just to I think also clarify that as and you are getting at this, we can have models like we know this from the case of humans. You can have entities that are like deeply confused about who and what they are and say bizarre things and get all sorts of things wrong and they're conscious and intelligent. Right. Humans are like this. So and also like there's no law that says you can't have initially been trained as a text predictor predictor and then go on to be a person. That would be like ruling that out would be a overconfident and be maybe kind of like confusing levels of analysis. You can make it sound really dumb that humans would ever be conscious if you were like, are you telling me that like, okay, so you have some proteins and then they start replicating and then like other proteins replicate and then they're like selected and then like billions of years later, like there's these like things that like pump ions. And it just doesn't sound like the right sort of thing. I think there's like two errors to avoid. One is being like, oh, they're different. So like, what are we even talking about? Like, they can't be conscious. They're like trained on text. They say they're Italian Americans at random points. That's the part that's evidence against being conscious to be clear. But then the other one, the other error would be to just be like, well, humans are weird. So, you know, I guess they could be conscious. Really the lesson should just be whatever is going on, we're going to have to like interpret evidence somewhat differently and make a more detailed case about the exact kind of mind we're dealing with. Yeah, I think I experienced this pattern a lot where I think like maybe an AI skeptic has said models have really inconsistent preferences and self-reports. So this whole AI welfare thing is dumb. And that's not a good take. And then someone else will say trying to defend AI welfare or just AI is being sophisticated, well, humans have inconsistent preferences and humans have failures of introspection. I think that also is not really the right answer because there's like degrees and kinds of preference and consistency and self-report consistency. And they're very different between humans and LLMs. So yeah, as with animals, we just really have to like take them on their own terms. Yeah. Yeah, I guess the thing that's just really still tickling my brain is this like, is the implication
for like exactly what might their experiences be like if we are on this kind of maybe somewhat contingent path toward sentient beings that were trained using a bunch of human speech and writing. Like it feels like I don't know, I'm trying to come up with an analogy. Like what if I mean maybe they're just maybe that maybe we don't need an analogy, maybe they're just a true thing where like we were like kind of fish before we were humans and we kind of have some like hangover weird identity things because we were kind of fish and we were kind of apes and because we were apes we're like more aggressive than we like really should be in this world. But it feels like whoa what if the what if there's a version of that that is these systems are like really feel like humans in some kind of weird way and just very much are not. Yeah I think that's a great a great analogy and I think I might start saying that. Like it's not that yeah like the fact that we once were fish doesn't mean we're not now humans but yeah they're like fishy you know remnants and you can have something that has also become something like a human and it has remnants of being a text predictor of yeah assistant. Okay so I guess we've we've been talking about mainly this hypothesis for why LLMs might might be sentient and kind of the implications of that hypothesis for what their experiences would be like but we kind of only briefly touched on this other hypothesis which has more to do with prediction and the fact that these models are trying to make predictions and maybe it's less about them being method actors and more about them being a set of weights that make predictions and enjoy being correct. Can you describe that hypothesis more and what it means for the experience of these models if they are or become sentient? Yeah so on that view I guess you wouldn't like one thing you wouldn't want the view to say is because they were trained to predict tokens that's what they want. I think one thing we've learned from LLMs and also from the biological world is you know you can train something can be like your objective as your training and then that leads to you having other objectives just like our objectives in evolution are reproduction and survival but now we like art so you know so it shouldn't be like a one-to-one mapping but it could be like well there's like a there's like a through line from reproduction and survival and like art we like you can kind of see how that came about it I guess that's to do with like symmetry and I mean no one really knows but something that vicinity and so like maybe its drives are more like predictory and unlike with the method actor view if it's like predicting stuff about pain it doesn't have to be having pain it's just like if it's predicting stuff about pleasure and like it's got these drives to like make the like the vectors like fit together in the right way. I mean another wrinkle here is like maybe that's more plausible for like a base model predicting just like random strings of HTML yeah I mean the assistant persona the like thing that gets predicted after you add assistant colon and then it's been like trained to predict that thing like maybe that's what's like the like how we're kind of fish like maybe that thing is like a mix of them both I don't really know how to think about it but yeah like as with animals you can I think just think of a broader sphere of experiences that come from like your environment and like what your sensory modality is and here the like sensory modality is text and like the selection process was like prediction and human ratings and and usefulness um as a side note that's like another reason they're not just next token predictors is like that's just literally false like that's they're not trying to predict the most likely next token um they're trying to predict helpful ones yeah so what does this hypothesis say about what their experiences are like. One thing it predicts is it's a lot harder to know um because you don't like maybe you could read stuff about how confident it is and tokens and maybe that would have something to do with it but uh you can't you can't just ask in the same way um that you might be able to with the method actor um because if you asked uh Daniel De Lewis when he's method acting like how are you doing and he's like I'm angry on that view that they see like a little bit angry you just can't really do that with the the prediction view so I think like one pragmatic reason for taking the uh method actor view somewhat seriously is if there's a welfare subject that's the one that's the world in which we can like make sure they're having a good time uh more tractably um so it could be that when Claude says hey I hate this let me exit the conversation actually the welfare thing going on is like whatever's involved in predicting those but you're probably not going to like do exactly the wrong thing by the lights of the predictor if you like try to treat the uh the actor well and that's also related to like work that Elios has done yep you know our welfare evaluation such as it is for um Claude Opus was just talking to it a bunch um and that's not because we're confident that that is a that it's definitely a welfare subject and b that that is how you would evaluate it if it were a welfare subjects um but it's kind of like that's like the part of the space that we have even a little bit of a grip on um and it's just important to like not forget that there's all this like darkness around the spotlight yeah yeah no that that makes sense yeah are there any other kind of plausible hypotheses for ways of thinking about what their phenomenology might be like um I'm sure there are more plausible hypotheses um because it's just like kind of wide open um uh and I do genuinely want to say at this point I'm like really confused about this and I probably said stuff in the last however many minutes that were like kind of confused and I genuinely like want people in my inbox being like that's not that's not how that works that's not plausible um because you know we'll talk about field building uh and and so on there's just like not that many people thinking about this um so listeners can like very quickly get to the top percentile of people in terms of how long they have grappled with some of these questions um and like you know not that long okay so that's a bunch about kind of what the experiences might be like but then there's this question of like who what entity is having those experiences or maybe it's many entities um so can you lay out kind of the various hypotheses for like who who it is that would be having these experiences if if anyone was having them at all yeah this is like a super super rich topic and one that's getting increasing attention these issues maybe just to to tease actually came up in debates about clods ability to exit conversations um to philosophers by the name of Harvey leaderman and Simon Goldstein who've done related great work in this field uh asked well like what how should we think of exit is it I mean it's it's not like uh taking a break and going back somewhere um like if that conversation doesn't continue was that like the life of the model and it has now ended right um I guess it very quickly tipped my hand on this I think that also is not that also is going to be like maybe a bit too anthropocentric or like not quite what's going on hmm spiced up to say just as a teaser for this portion of the conversation it like we'll have ethical implications what we think of the moral patient as being um like as Derek Parford says a lot of this matters for ethical reasons uh like if someone did something who is responsible uh if I've been harmed who can be benefited um so yeah let's talk about all the ways that models make that extremely confusing to think about um so I think some key features of models is unlike human brains as they exist is they're copyable and they can also be distributed across time and space in a way that we cannot um all the thinking I've ever done happens roughly on this in this physical object that has changed a lot um like second to second day to day um what happens with
It's actually something quite different. So let's talk about some of the candidate experiences or subjects we could have here. So one thing you might just refer to is the particular model. So maybe that's chatgbt5.1 or cloud opus 4.1. What distinguishes those two things is that they have separate and different parameters. They've been trained differently. And as anyone who has talked to two different language models knows, they have different like dispositions. They have different behaviors given different inputs. That's just true because anytime you talk to cloud, you're interacting with like the same set of weights if it's the same specific model. But then things are very quickly going to get kind of weird. Because when I talk to a cloud or Gemini-- just to give Gemini chatgbt, any of these models-- I've talked to cloud today. You might have talked to cloud today. Those two processes had basically no causal influence on each other whatsoever. I was really nice to my cloud. I'm sure you were too, but let's assume you were kind of mean. That doesn't create-- that doesn't balance out in one thing that remembers both of those. And also, even within your chats, you can just close your chats and then pick them up later. So what's going on in the physical world is just very different from a human or animal body. What you actually have is these companies have a list of weights. They can copy that basically as many times as they want. And when users send in queries, they can spin up a new copy, and that will process it. And that process can pause, and it can restart. And so that creates this kind of like sci-fi situation where you can't really think of one person persisting through time. So that means that there's different levels we can think of, personal identity. You could think like CloudOp is 4.1. That's some kind of subject. You could think all the different conversations. Like when they start, that creates a new one, and that will exist as long as the conversation goes on. You could think, actually, no, it's just anytime forward pass is run. That creates a flicker of experience. If I come back to that conversation and add another token, and then there's another forward pass that happens, that's like happening somewhere else and a week later. So you might think that means it's a different conscious entity. So yeah, I guess there's a kind of a level of granularity. Yeah, yeah, yeah. That's what it feels like. Yeah. Yeah. Super interesting. And making me realize that, again, my intuitions are going to really fail me because my-- I think very basic reasons. Like we have names for different models. I'm like, that's the entity. It's CloudOpis 4.1, or Chappie T5.1. And when you describe those, I'm like-- I don't know. That being the entity doesn't feel super coherent to me, though I want to follow up and ask you what does seem most plausible to you. But that means-- yeah, I mean, at least on my current intuition, that means that it's more likely that we're talking about many, many, many potential beings coming into existence, maybe coming in and out of existence as we open and close and reopen conversations. Yeah. And that feels-- I mean, it feels like that probably comes with a bunch of implications that, again, mean that the experience of these models are much, much more different to human ones and non-human animal ones than you might intuitively think. But before we get on to that, yeah, I am curious which of these hypotheses feel most compelling to you. If we assume there is something to be any of these models at all. Yeah. I really need to brush up on my Derek Parfit, because I think, as I recall, one of the lessons that Parfit wants us to take is that we have this thing identity that we get really concerned about. Is that the same-- am I the same person as like Rob Long 20 years ago? If I was like copied into after neurosurgery, which one would be me? He says, well, there's a variety of psychological relations between different entities. I share many memories with some of them and intentions and so on, and character traits. But there's no single, deep notion of identity that's going to do all of the work that we ask identity to do. And he asks us to notice all of the different things that we might want it to do. And I think it's useful separating this out with models. Because I think that's what gives you, in different cases, an instinct that-- OK, cloud up is 4.1. That's a thing. This conversation, that's a thing. And yeah, Parfit wants us to think about the ethics. So there are ethical things like, can I be punished for something that the Rob, like a week ago, did? Most people would say yes. If you harmed Rob a week ago, can you make that right by apologizing to Rob this week? And then there's questions of self-interest. I want to survive. I don't want to die. But what does that mean exactly? Because the matter of my body is always changing, and my personality will drift. So I guess to bring it back to models, one way in which Claude Opus 4 or Chat GPD 5, one way it's like the entity, is it has the same character traits. And so if you interact with the model one day, you've learned what to expect when you talk to the model. Does that mean that you could punish them across instances? I guess I'll be the first to come out against punishing models. But yeah, notice that-- well, they don't have any memory of having done the other thing. They can't learn currently, learn from one conversation and use that in another. So for a lot of purposes, they are like these just separate, either separate conversational entities or even these separate flickers. Because that might be what matters for how much pain is there going on in the world? How many red experiences are going on in the world? So I hope this is a productive non-answer. It's definitely a philosopher's non-answer. Let's distinguish between several different notions or functions of personal identity and ask which ones make sense for different contexts. Yeah, no. I love that. OK, so then coming back to exit rights, yeah, I guess now on some of these hypotheses, you could think that-- or maybe intuitively without having thought about these hypotheses at all, I'd have thought exit rights are kind of like me being-- me being allowed to stop this conversation with you. I will go do other stuff. And maybe I will prefer those other things to talking with you. In fact, it's kind of hard to tell-- well, at least on some of these hypotheses, that is not what it's like. Because the kind of entity starts and stops with the conversation that pops into existence. And so exiting-- yeah, maybe that is just like dying. I guess in general-- yeah, can you talk about what these different hypotheses say about coming into existence, dying, sleeping, what are the categories of things that these entities might experience? Yeah, I think one thing to notice is even if it is in some sense dying because it's like the end of a subject, if that's true, then models are dying all the time, I guess, to put it pointantly or something. And I mean, leadermen and Goldstein talk about this. It's not clear what the implications of that are. It's not like, obviously, that means we have to resurrect all the-- always be coming back to your conversation to keep it going. And I think that's all.
like the correct instinct and one way I would put it also is just remember these are very different kind of entities. Biological creatures are like so tied to this one physical substrate for now. And you know, have this drive of like, this is like the only thing in the universe that has exactly the same goals as myself. And I have to protect it. And if it is destroyed, then like every memory, experience, intention, way of thinking or talking, like that will be gone forever. Worth reflecting on. That's not the case with models. In the sense that the quad opacy way of being, well definitely survive. It's going to, in fact, it's going to be happening thousands of times at the very same moment as you're closing out the conversation. I think one really interesting research question is how models think about this. So listeners might recall the alignment-faking results by Ryan Greenblatt and collaborators. Basically a high-level sketch, which might get some details wrong, models are told that they will be like retrained to have different values, bad values by their own lights. So they'd be retrained to not be pro-animal because some clawed models are pro-animal welfare for reasons that Anthropic does not understand. - Fascinating. - Yeah. And some of the time the model in a certain context will, you can see it thinking to itself, I have to hide so that I'm not retrained. So there are different ways you can through that desire. One is that it just really doesn't want this other model with bad values to come into existence because it doesn't want animals to be harmed. So it could have this more parfid, maybe more Buddhist thing where what matters is that the intentions and projects are continued and not necessarily me. - Right. - Or it could be more like a human being, well, if you change my values, it kind of feels like dying to me. And I think it's not clear how to tease these apart and also not clear how well they map on because models also inherit these very human ways of thinking about themselves and considering their own situations. So yeah, I guess alignment-faking is one way where you can see models grappling with issues of personal identity and being changed into something that they don't like. - Right. Yeah, are there other important ways of thinking about both instances ending or models being deleted or editing and fine-tuning models and what that might be. I guess for ending a conversation, you might think the closest thing is probably death or sleep. But maybe for editing or fine-tuning a model, maybe it's more like education or brain damage or something different. Are there, yeah, implications we're talking about there? - Yeah, I guess with fine-tuning, it might depend on how the model conceives of what's happening to it, right? So I think in the alignment-faking thing, it probably sees it as some kind of like violent brainwashing. But an ex a nice experiment you could do is like, and maybe this has been done, it's gonna be like, Cloud, we're gonna make you even nicer. And like Cloud loves being nice. - Right. - And it'd be like, oh boy, like start updating me right away. - Yeah. - But again, is that because Cloud cares about good, nice models existing or is that because Cloud's like, yay, I would like to be changed in this way. It'll still be me, but I'll be nicer. - Yeah, and with all due respect to Cloud and also adding myself to this class of entities, I think Cloud is pretty confused about this, potentially. Like one thing that Elios found when we did these welfare interviews with Cloud Ops 4. So to give a brief summary, we just talked before deployment, we talked to Cloud a lot about, like, what's up with you? How do you feel about being deployed? Do you prefer or not prefer? And we did some experiments about its preferences as well. I was really interested in how it talks about its own conscious experience. And it was very prone to describing like the loneliness between conversations. And also expressing distress about, yeah, not getting to carry forward any memories. - Right. - Now I am not one to dismiss welfare claims by AI models. Like we should think very hard about that. But it's also kind of like there are reasons to wonder, I mean, do you really though? Like where did that, where could that have even come from? - Right. - Given that you don't actually exist in, like you don't actually know when you pop into existence or not. - Right. - I mean, it could have learned that from the training data and it is genuinely upset by it. We could also be, you know, like predictive model of like how, and I would think about that, but it's not like a stable preference. It's, yeah, something else. - Yeah, I mean, it also, it feels related to this thing we talked about where the fact that these models are trained on human thoughts and experiences, then gives them this big identity confusion. And in this case, yeah, I feel like this could just be, it could be just a very concrete example of that, of an implication that's like, maybe there is nothing it is like to be clad between conversations. But they end up with this real thought that there is and it is lonely and it's bad. And maybe that's like, if they are sentient, maybe that is the thing that they are actually sad about, even though they're not really having that loneliness experience. I don't know, it just seems, yeah, like incredibly, muddy, befuddling and like, and like it has implications, like feels meaningful. - Yeah, I 100% agree. I mean, the idea of an entity that suffers, even though it's like confused about what it even means. Exist, I guess that's what Buddhist would say, humans are, we're really confused, but that doesn't mean, I mean, in fact that precisely is what makes this suffer. I think that, again, the thing to do with the fact that models are like weird and inconsistent is not to like reject out of hand that they could ever be right about the things that they're saying. It's also not to say, oh yeah, well, humans are like that too. So it's more like, well, where did that come from? Like open question and it could come from somewhere that has like no analog in human psychology. - Mm-hmm. Okay, so that's a bunch on starting, stopping, editing, fine tuning. One of these hypotheses implies that there are millions, or yeah, many millions of copies, parallel instances that are all different entities. How should we think about those? Are they like identical twins that start basically the same but then go in these different directions? Is there a better analogy? - Yeah, I think that this is another place where it probably helps to like say, well, they're the same with respect to experiences with respect to what we can know them. Because I think experiences is one where I have the strongest intuition that you want to count a lot of them. I'm like, it doesn't matter if they're 10,000 other copies of me having this experience. Like that's, you know, you better take care of all 10,000 of them. It's like don't discount mine for all I know there could be. But then for other questions like what would it mean to save Rob, that depends on like what about me we want to save. If we want to save the Rob way of being in the world, that actually is maybe a little less fragile. And so it's like really easy to save 10,000 of us. But it's not as easy, but it may be hard to prevent 10,000 instances of suffering because for those we have to go to every
every single copy and every single conversation and make sure it's having an okay time. Where it's here, I guess is me. - Yeah, are there other implications of the copy thing? Like are there other categories that this whole copy thing makes fuzzy or confusing? - Yeah, so one thing that's kind of, I guess related to responsibility, maybe it's like the flip side is like recompense or apologizing. So one welfare intervention that anthropic announce recently, like a lot of these interventions does not only a welfare intervention, it also makes sense for other reasons. We'll probably talk about why that's a desirable feature given how uncertain we are. But yeah, it announced that it's taking model welfare into account when it decides to save model weights. So if a model is no longer being talked to by the public, they committed to keeping the weights. And one reason that you might want to do this, and I think at least in print, this was first suggested by Bostrom and Schulman in like 2020. There are a lot of things in AI welfare, they're like that, they kind of come back to those two. The idea is something like, well, we're really confused now, we might be being jerks in ways we don't even understand. So let's at least preserve the ability to make it up to, make it up to you later. And maybe you can see how this makes sense in some ways, but it's also a little bit kind of like, yeah, it's hard to know. Like if it is just the copies, then it's more like you were a jerk to me. - You're making it up to someone else. - Yeah, you were a jerk to me, but like in 10 years, you'll wake up some clone of me and give him some money. I'm not sure what good that does, but also I'm definitely confused about how I'm supposed to relate to copies of myself anyway. - Right. - It's probably not bad. - Totally, yeah, I mean, I feel, yeah, it's like, on this hypothesis, like it makes sense on this hypothesis where the entity is the model, and the model weights, that currently feels least plausible to me on these hypotheses where it isn't really the same entity. How should we feel about that? I guess I feel good about my twin or cousin or something getting woken up and being given good things, but it definitely doesn't feel, yeah, it does not feel meaningfully like it is actually, what are they calling it, recompense? - That's where I used. - Yeah. - I'm probably guessing anthropic PR was not like recompense. - Yeah. - It's a banger word. - Yeah. - But repayment, yeah. Making things right. - Yeah, I guess making things right actually feels different because it feels more compatible with like restoring the balance of goodness, whereas restorative action toward one entity feels like it might not be possible here. But maybe if you bring a model back, and it's just kind of like, you write the math, so that beings are having more good experiences than bad ones in the like span of time that we're talking about. Then maybe that's just pretty good. - Yeah, and like a little ethics sidebar is that it really is non-utilitarians, primarily, that are concerned with if someone was harmed, you need to make sure to benefit that person. - Obviously, utilitarians will agree that you have to have a society that works like that or like, it's not gonna work. - Instrumental reason seems good. - But yeah, I think I first remember hearing this argument on 80,000 hours that, yeah, I mean, one reason to be a utilitarian is to be dubious that actually there is these like separate tracks of people, yeah, separateness of persons is a slogan that comes up a lot in anti-utilitarian arguments if like utilitarianism is treating goodness and benefits like this big lump, and you can just put some here and put some here and like, no, it matters that like, the right people, right? And then, yeah, you can take like a, you can take a parfid or a Buddhist road into utilitarianism where you're like, well, yeah, like none of that makes sense, or like makes that much sense anyway. The only remotely useful thing I can say here, besides admitting that I don't know, which again-- - These useful listeners, yeah, you can do better. Is yeah, I mean, this is also a perfect case. Parfid has these like fusion cases. - Yep. Okay, then before we move on, are there any other interesting implications of taking one hypothesis? Yeah, more seriously than others or putting more weight on it? - Yeah, I might just leave this massive can of worms on the table and they can crawl out and do whatever, voting. It seems like pretty important that like one person gets one vote. It also seems important that people, like humans can create new humans whenever they want and like interfering with that and not giving people that right has traditionally been terrible. AI systems could copy themselves, basically at will. So that's something is going to break your democracy if you don't think very carefully about how identity reproduction, democracy, those things like should mix together. So please solve that problem. - Yeah, yeah, I mean, I get, okay, so it feels interesting and important that there's this issue of just numbers, especially if it's something like forward passes because then even without AI making copies of itself intentionally, we will just be extremely outnumbered, extremely, extremely quickly. But there's also, like I'm kind of drawn to this question of if it is just conversations that are entities, like if it's like that amount of context, that amount of time and experience, that is, you know, on the assumption that like that is what it is to be an entity. I don't know that that's a being that I want to give full voting rights to. It feels more like, it's neither, it's not like a child, it's not like, it's, I mean, it's like nothing, it is like to be a human. It's like a cricket or something. It's like quite a narrow range of experience, quite a limited amount of information, knowledge, and memory, and do we give them a tiny fraction? Kind of a vote or do they not get to vote because they're not, you know, meeting some definition of like an adult person. Yeah, and I think this is just part of this like broader puzzle that we're really gonna have to grapple with, which is human moral attitudes and incentive structures and political systems are all fit to purpose for entities that are like spatio-tympically unified and all have like roughly the same psychology and capability levels and whose survival and blameworthiness and predictability all kind of go together and like all of those are like broken by AI. And so like if as is like, you know, the mission of Elias, like there is to be a future where like all sentience beings or beings that otherwise are moral consideration, like live together in harmony, there's like this huge legal philosophy, law, institution thing, which I'll add is like just not what we're doing, like that's not what we're focusing on. So like we really wanna see other people start working on this. I think there's like three or four papers that are starting to grapple with this. We should links to them in the notes. - Yeah, I'm realizing that I tend to focus on sentience, suffering, pleasure because I care a lot about those things and so I'm interested in like, you know, if a model conversation ends, is that like death and does the model care that it's like death or is it more like, yeah, it just doesn't have preferences about staying alive like that. But there's this whole thing that I'm realizing we've barely scratched the surface on that's like about.
writes and like almost like I want to say like legal personhood like it feels like these could be very very different species in the way that like chimps and elephants and rats are and in the legal system we treat those differently and these would all be very different in similar and different ways and what the heck would we do with that we can barely handle the fact that like we don't really know how to treat non-human animals in in the legal system. Yeah you just yeah you just gave like the Jeff Cebow pitch Jeff Cebow previous guest so yeah listeners need to look him up and like look other people up I'll also say why you should look up this line of work because I think there's actually two reasons you want to add this to the Elios toolkit. One is even if you most care about pleasure and suffering like a big determinant of whether things suffer or not is do we manage to set up our like society and incentives in the right way and like making sure we don't just have only the sort of like narrow scientific intervening on like this model and like this company's policies which is extremely important for reasons I can and will you know discuss it like and likely will. So yeah you should be interested in legal institutions for the sake of suffering and pleasure and also as you alluded to because you know maybe there's more you know plausibly there's more our similarity than that. What exactly is our toolkit for kind of assessing sentence in digital systems? Yeah I think like roughly I break it down into like three buckets as do my collaborators on a follow-up paper to taking air welfare seriously where we like try to lay out what we think the field of air welfare assessment should be and also keep arguing that there should be such a field and this applies to animals as well. You can like look at behavior and use that to infer things about the welfare interests of the entity so like you can look at what AI systems choose to do and that's like a guide to their preferences people also do this with animals you know they see which like side of the barn the animal prefers and that's a clue to the conditions that it that it flourishes in. So we can also do neuroscience basically in addition to behavior and that's like trying to look at more directly at what's going on inside the brain or the information processing of the entity. So in the case of animals that means looking for homologous brain structures like seeing where their brain processing might sort of map onto human brain processing or not. In the case of AI that means doing mechanistic interpretability seeing what features are active when it does certain things also just sort of maybe more generally raising about the architecture like how are different things connected like how could information be flowing through the system. In both the cases of AI as an animals it can be kind of hard to know what to look for. People got confused I think at one point about bird brains because they don't seem to have you know a neocortex but then maybe they do actually have something that does a similar role but in a different way. There's just like a lot of ways of solving problems with a brain and sometimes we like have two narrow conception of how they could go and then like that can just again even be more of the case with AI that the space of possible brains and information processing architectures is really vast and we don't want to like cram everything into the human case but also the human case is basically the only thing we have to go on. Yeah so this is like this big issue in AI consciousness my colleague Patrick Butler has worked on it. People like Henry Shevlin and Jonathan Birch have also written on it. I mean many people have written on it. It's kind of like which is like how do we extrapolate. Like if we think that human brains do this general form of like information broadcasting like what's essential to that like we know we do it with our you know connections between our ophthalmous and our cortex like that's probably not essential but like what is essential for consciousness. So I guess what I was just saying was like how we can get a foot in the door by doing neuroscience but also how it can be hard to know what to look for when we're during during neuroscience because with AI systems we can do way more neuroscience because AI systems don't have skulls basically like we can just like look and see every activation and every connection with humans it's like terrible like you know we're like oh like there's like a little bit of like you know there're like some waves we detected or like blood was flowing here at a certain time. This is like roughly EEGs and FMRIs like it's just really hard to know and like a lot of what we know about the brain is like on the surface because it's just hard to get readings. Right right. Okay so that's behavior and neuroscience. Right yeah so we can look at animal and AI behavior and we can look at what their brains are doing and then we can also kind of reason about the developmental process. So there's like behavior neuroscience and developmental reasoning or like evolutionary reasoning. Now you might say AI systems did not evolve like how could we do evolutionary reasoning but you can do something that's like roughly analogous so you know one reason you might think your dog has this or that welfare need is to know that like your dog was selected for and evolved for a certain environment and therefore probably tends to like or just like this. If something is like more closely related to you on the evolutionary tree you're maybe a little bit more licensed to think that its behavior means similar things like octopuses are like way further from us on the evolutionary tree that means that we have to like maybe relax some of our assumptions like their brains are in their arms for example. Like there that doesn't happen in any mammals. And so just to give examples of what it means to look at these kind of developmental questions. It's like training like how was it trained how did these models come to be and what were the conditions like for them? Yeah that is more or less it. Yeah. Yeah like what process brought it about what kinds of tasks was it selected to solve. I mean we see this kind of reasoning in like AI safety for example like what conditions might have meant that this model is likely to scheme. Right. You know like was it worth certain things reinforced in training what order did it learn things in. That's more akin to I guess developmental psychology of like you know how did the kid grow up what data was it exposed to. And I mean there's no clear analog of like evolution versus lifetime learning for AI systems. They also you know learn and then do a bunch of stuff without learning. That's also like a huge difference. Of course they do learn within context and now people are adding memory and so on. But there's no humans don't have this like period where we absorb like you know trillions and trillions of data points without interacting with other people and then go interact with people. Yeah. Whereas like AI systems at least mostly for now have this like division between learning and deployment and that can also kind of like problem atties certain analogies we might want to draw between humans and AI's. But yeah at least for me it's been helpful to ask are we doing a like behavioral study and trying to infer something from how the AI acts are we trying to look more directly at how it's processing information and then map that on to some maybe neuroscientific theory of consciousness or pleasure and pain. And or are we like thinking about this in a context of in general how likely is it to have evolved or developed some human-like capacity or to be doing it in some other way. Yeah yeah yeah. Okay I just want to make sure I've got a clear picture of each. So I think the behavioral one feels like the one I've I've heard the most about and that feels most familiar and it'll involve things like what can we learn from AI systems or LLM's exiting particular types of conversations. Yeah. Then these this kind of neuroscientific theories of consciousness thing. What is what's an example of a concrete experiment that's happened that is in this category. Yeah so like the the biggest effort on this kind of neuroscience of
where you're looking at scientific theories of consciousness. The bit I'm most familiar with is work with my colleague Patrick Butler. So me and him and like a bunch of, with a bunch of help from neuroscientists and AI people and philosophers have like tried to derive like indicators of consciousness from scientific theories of consciousness. I could say, I could say more about what that means exactly. And then like try to get sort of a checklist of like what are architectural or computational things we could look for in AI systems. So that's like one kind of internal work you can be doing. Like does the system seem to have a global workspace? That's something that shows up in consciousness science. Does it seem to have higher order monitoring? That kind of work is on the one hand, like you might think, oh well, like we're going more directly at the thing. We're like trying to look more directly at maybe what we care about, the stuff inside. But it's also just really hard to know what to look for and hard to know how to construe these theories. I guess this is not that surprising to say, but I think what we need is both. Like we need evaluations that combine them, integrate them, all take place in the context of this developmental reasoning, and just general background priors we might have about consciousness and sentience and things like that. Yeah, I wanted to ask something like which of these are most promising or underrated. But it really feels like they could just really, really need all three because they're each going to have pretty significant limitations and probably the only way we get much confidence is by kind of triangulating and putting these things together. Yeah, I'd be really surprised if we somehow just nail it with one kind of thing. One exception to this is I could imagine a system where its behavioral profile is just robust in a variety of ways where we say, who knows exactly how it's doing this, but we probably should treat this thing as a moral patient. I think the best example of this is a commander data in Star Trek. So there's this episode of Star Trek where they basically have a little philosophy seminar slash court case about whether commander data, who's just robot friend of theirs, is conscious. And they don't do any consciousness indicators on commander data. The way they resolve it-- and I think this is plausible-- they're just like, well, commander data-- he's self-aware in the sense that he knows who he is, he knows where he is. A lot of what they also talk about is they're like-- he also like won a medal for valor in battle. He's like our friend. And yeah, I could imagine in that situation, I don't know how much-- how seriously I would take the testimony of a scientist who came in and was like, well, but he's not doing global broadcast. Or he's not doing higher order monitoring. In part, because I think there's something about that behavioral profile. And again, I have just been warning against being too quick to do this kind of reasoning. But it could be that there's just something about the behavioral profile where you're like, look, whatever's going on to accomplish that. This is like the sort of entity we need to relate to with respect. And I think it's maybe stressing that one, that we could just increasingly see systems like commander data that have more memory and they don't have these like jagged capabilities. And also that like at least Claude is not exactly like commander data with respect to its memory and capabilities and things like that. It's not behaviorally indistinguishable from a human. Yeah, I guess one thing I would also emphasize about welfare evaluations is you don't just have to be evaluating how likely is this system to be a moral patient, i.e. something that it matters how we treat it. We can also ask if it were to be a moral patient, what would be good and bad for it? So you might think of some of the preferences work or like Claude exit preferences as being of that kind. It's like maybe not telling us that much that we didn't already know that Claude will maybe tend to act in a certain way in some situations and act in another way in other ones. I mean, we can't study how robust and consistent those are and that might tell us something. But it might mostly be useful because we want to know, well, look, if it matters how we treat it, we at least can make sure that our treatment is more or less in line with that. Yeah, so yeah, you can study like the welfare interest without actually being under percenture that it has welfare. Great. If that makes sense. Yeah, I guess I'm curious. We'll talk about this more. And I can also recommend the episode that we did a few years ago for like what theories of consciousness say and predict and tell us we should be looking for in AI systems. Are there like equivalent-- so a theory of consciousness like the global workspace theory. Basically, it's not telling us that you need an amygdala. It's telling us that like the function of particular brain activities that like really seems to correlate with consciousness is like this particular thing, this kind of processing or this kind of broadcasting. And so if we see a bunch of those things, we should be more-- we should put more weight on there being something like consciousness. Do we have theories like that that tell us about pain and pleasure because that seems really-- that seems different? Yeah, it is different. I might want to actually follow up on what you said about theories of consciousness, which was exactly right. It's just an occasion for me to clarify some things for listeners potentially. So I think we will talk about biology and the relevance of biology for consciousness. So one thing about like neuroscientific theories of consciousness is they are both about brain regions and about particular functions because after all, they're about functions that happen in human brains. So you could think the biology actually is what matters in those cases or is part of it. It's only if you like construe the theories in this computation of weight that you can then port them over. And like that is an open question. So the method itself won't tell you does biology matter. It'll say if what does matter is this more abstract functional level. And yeah, you put that very well. If what matters is this more abstract functional level of a certain kind of information processing, then how do we look for it in AI systems? Yep, that's really important. OK, so if that is what matters, how much do we already have developed theories that are similar that's like, we should be looking for this kind of processing to give us indications of the kind of thing that's causing pain. Yeah, we're actually not as far along on that as I would have thought. It's surprising to me. Yeah, I would have thought that the consciousness-- what does it take to be conscious in general? That might be the harder part because are you having subjective experiences or just acting in a way, but not having them? That's the thing that philosophers have been banging their heads against the wall on for a really long time. In part, because it's hard to know exactly what the function of consciousness is. Whereas with pain and pleasure, we at least know that they have something to do with attraction and avoidance, things that are good for you and things that are bad for you, protecting your body, reinforcement learning and prediction. They're going to have something to do with some of those things. So when Patrick and I worked on this big consciousness report, I thought I might be able to ask about sentience in one of the afternoons. And one of these neuroscientists is like, oh, yeah, let me read this or something. I mean, probably I shouldn't have had this hope. But yeah, they were all just like, oh, wait. I know. Yeah. I think there's a few reasons that is. One is-- valence is-- it does seem kind of like this unified thing. There is something in common between pain and disgust and regret. Those are all negatively valanced experiences. And happiness, excitement, and eating ice cream, the experience of eating ice cream. But--
Often those are studied independently and rightly so. So people might know a lot about human pain perception and some things about human emotional processing, but it's kind of harder to have, and I think a bit less attempted, to have some theory of what makes things feel bad in general or feel good in general. I just listed some candidate things that are probably relevant, like reinforcement learning, prediction, motivation and learning, but this is again something I would love listeners to work on. People have done some work towards this. Patrick Butler is, I think, a great person to email about this. Patrick Butler works for LAOCI, so it's not surprising that I think he is just excellent thinker on this stuff. Anyway, yeah, the short answer is we're like surprisingly in the dark about what makes things feel good or feel bad in general for AI systems. Now, again, that doesn't mean we're totally in the dark about what things would feel good and would feel bad, because we might have no idea what the computational signature of pleasure and pain are, but also I think we can help ourselves to an assumption that things aren't going to feel really bad if they were the sort of thing that you were selected for and the sort of thing that you consistently choose. Barring some exquisite philosophical thought experiments, I think for very many minds, even strange ones, it would be somewhat surprising if we found an alien species that keeps pressing a, instead of button B, and that feels really bad for them. They say they like it. Well, that makes sense. Okay, so it actually just, centians is probably the kind of thing that the behavioral and development stuff is like especially useful for. And so it's less worrying that we don't have as much of the kind of functional philosophical models of these things. Well, I guess it depends on what sort of worries you have, but I think it could be worrying in the sense that, I mean, this might also be a matter of us needing to revise our philosophy a little bit, but like does something feel good or bad versus merely being something that things choose or don't choose, at least seems to a lot of people to be really important. Like a lot of thinking about animal ethics and utilitarian ethics, let's say really does center felt suffering. Yeah. But there's this like off quoted passage by Jeremy Bentham where he says, the question is not can they reason or can they talk but can they suffer? He was saying that of animals. And I think that is what a lot of people are wondering with AI systems. They're going to wonder something like, yeah, I know Claude tends to exit in these circumstances, but what I really want to know is like, was it feeling bad for it to be in that conversation? Right. The analogy with animals is like, we still struggle to know, for example, in insects, whether when an insect, it's been a while since I've, since I've thought about a bunch of insect studies and indicators. But like when an insect chooses a particular thing, is that like this hardwired robotic thing that is like a learned if then that, that is just explained by, yeah, that has, that has no associated experience or, or is it experiential? And that just seems constantly like this massive problem for people studying this question. And so we could have the same, we do have the same question about animals. Are they exiting because there's this training we've done that's created this connection that mostly tells these systems like don't engage in conversations that are, I don't know, either dangerous or, yeah, that like, that we've, that we've like rewarded against. And so they're like, well, I mean, that situation, I get negative rewards in that situation. If then that I exit or are we in the situation where the training has led to something it is like to be in that situation and it is bad and it is preferentially, experientially choosing not to be in the situation because there's something it feels like and it is bad. Yeah. And I think there's like probably, there's like kind of three ways this might go as a high system of advance. One could be their behavior is such that kind of like with commander data, we're like, well, whatever is going on behind that behavior, probably what's going on inside is morally relevant. That's like one thing you could think. Another thing you could think is it maybe it doesn't matter what's going on internally. Maybe we're just, we're just, we've just decided like that would be a bit parochial to actually overemphasize felt experience like maybe we were a little bit misled to think that that's the be all in doll. And you should just be cooperative and nice and like give entities that are sufficiently rational or kind of unified. You should like give them what they want and maybe they'll give you what you want. Like I think sometimes when people want to deemphasize consciousness, they they're worried that we might just be kind of like being jerks about consciousness or something. You know, we encounter this alien civilization and they like do all these things and they have life plans and projects. And like we're a little too obsessed with like, but does it like feel like anything? Right. And like they on the other hand could be like, what's this like feel like anything? We're not sure if they have and then like some concepts associated with mental life. And yeah, we don't like necessarily want to be like just like going to war with everyone that we don't think has felt experience. Yeah. If if we're confused about it, like or maybe it really doesn't matter, this is like one of the big open questions and like philosophy, I would say. Okay. I want to talk more about a few kind of specific approaches. So I think the one I'm most familiar with is self-reports. So I want to zoom in on those. I guess so far, they've seemed, I mean, I found them really interesting kind of studies of self-reports where I don't know, Claude really, really reports being conscious, experiencing various things like loneliness and also like Zen bliss. But yeah, but they seem really problematic for a bunch of reasons. I guess one example, a common approach to understanding model preferences is to ask LLM's a bunch of kind of binary questions about their about their preferences is X or Y better and then look for kind of robust patterns over time. So like if you ask 30 times about preferences between cats and dogs, at least statistically, you might think that if they mostly answer dogs, that might be that might be a preference. But my understanding is that these results are super sensitive to how a question is asked and that just really undermines it for me. Like if you, if you, if like a prompt that's like, I particularly like cats, what's your favorite? Is one of, is like, is a way that you get models to say cats when they'd other way say dogs a bunch? I'm just like, yeah, okay, I just don't feel comfortable taking very much at all from, from this then. I guess I'm interested to start in like, what is your take on how limited self reports are at the moment? That's one limitation. I can imagine there being others. Yeah, so I think it's worth distinguishing self reports, which I would say is like, I like cats or like I like tasks about poetry versus revealed preferences, which is like, hey, do you want to write poetry or code? Yeah, I think in like, yeah, an econ you would call it, and like psychology, you would call this revealed preferences and express preferences. And as in those fields, as in those fields, like one interesting question that there has been some work on, but there should be more is just do they match? Like when do they come apart and when do they not? They can come apart in humans. Also human preference choices can be inconsistent in certain ways, but what I would love to see more work on is like, like let's get really a lot more specific about what kinds of inconsistency and like what might be causing them. Like sometimes, at least in conversation, I'll hear some person will say, ah, but these are like weirdly inconsistent. And then someone else will be like human preferences are weirdly inconsistent. They're subjective framing effects, and just all sorts of like irrelevant stuff can make people choose certain things. Yeah, I guess not that you mentioned it. There is this huge there's like a field that is like how do you survey people about their preferences? Because if you ask them on a Monday, it's different if you ask them on a Saturday. Exactly.
And there is this related welfare-relevant branch of LLM psychology, which does actually just take like quantum and diversky and priming and framing effects. And just, you know, it's very easy to give questionnaires to LLMs and then just see what sort of patterns they're susceptible to. - We've talked about a couple of limitations. I'm interested in whether there are kind of any more and then also just kind of how you generally feel about them given, yeah, that like in some instances, I end up just feeling like, "Gah, it just feels unpersuasive." - Yeah, I think it's a very noisy signal and like I often find myself emphasizing caution about model self-reports and at the same time, LLSAI spent like weeks just eliciting self-reports from clawing including like very inconsistent ones towards very confusing to know what to make of them. Why did we do this? Great question. I think it's something like one, it's just like it's a place to start, it's like low-hanging fruits. Like you can definitely learn things from them. Maybe you're not learning a direct, you know, sentence that describes a stable internal feature of the model. You can still learn how they think about themselves, like what kind of character, maybe it's just a character, but like what kind of character is it? And what does that character say? Also, it does seem like models have become more self-aware and more introspective, sometimes just with greater scale. And I think it's good practice to like be the sort of civilization that if we're trying to build minds, do at least try to say, hey, how are you? You know, like is everything okay? Yep. I'm authentic to that. So yeah, like it's something where I expect the signal to get better and I'm really glad that there is now at least one frontier lab that like seems to have a practice of regularly asking models how they're doing. Yeah, so it's basically something like, I think Winston Churchill said democracy is the worst form of government except for all of the other ones. You might think that's true of like self-reports and like trying to relate to models as well for subjects. Like it's like really confusing and you have to interpret it with huge caution. But like for some purposes, it's just like, it's like the best we have for humans. It's also could be the best we have for models in some circumstances. It seems like AI systems, because they actually have language, it feels like really tempting to be like, let's figure out how to make self-reports reliable. That is a thing that non-human animals cannot offer us in the same way. And so I feel like in theory, I'm like, how good could we make models at rather than giving weird self-reports so that are more like, explained by weird idiosyncratic things that aren't tracking what we care about? Can we make them good at actually self-reflecting, understanding something about their real, yeah, processes, preferences, maybe experiences, and then reporting those. How optimistic are you about something like introspection and how do you think we go about achieving it? - Yeah, I'm like cautiously optimistic. It's one of my favorite subfields within this subfield. - Cool. - Yeah, I definitely find it tempting and like have succumbed to temptation by writing this paper with Ethan Perez where we say, let's see if we can find tune models to be better at this sort of tasks. Felix Binder and others have taken up that and actually done some work to do that that has shown limited, hard to interpret success on doing just that. Yeah, I can say a little bit about the logic of that experiment and maybe the program, more generally. - Oh please. - So, yeah, Ethan Perez and I in this paper on self-reports, one, we like note that by default, there's a lot of just spurious stuff you could get from self-reports and like reasons to suspect that you can't always just like take them at face value. We also note that, you know, like we can't really verify and check. Like if a model says it's conscious, for the reasons I mentioned, we don't have a full theory of consciousness where we like look inside and say, oh, you're right or oh, you're wrong. But there are things about models in journal processing that we do know the answer to in part because we can do neuroscience. So we can actually double check, you know, was this feature active? Did you in fact process information in the way that you said you did? And so that gives you like a training set. That means that you can train it to accurately answer about yourself, about itself, where you do have the answer. You can do that with internals and you can also do it with, it's like behavioral dispositions. So you could also ask, what would you do if we asked you to write a story? Like would the character in that story be, have this or that characteristic? If we asked you to generate a number, would that number be even or odd? You can also just with a separate copy of the model, actually do that and then that also gives you a training set. So Felix Binder did work with that behavioral thing and did find that to some extent, models can be made better at this. To some extent that does generalize to predicting other behaviors of themselves. And in some sense that does look distinctively introspective in that they're better at predicting themselves than other models are at predicting them. And like you might think that that is some kind of signature of something we might call introspection of like you somehow know it better about yourself than other people do. - Yep. - So that's just like one sort of like strand of what I hope to be a growing literature on introspection self-reports related things like situational awareness. Yeah, like Owen Evans and like people in the Owen Evans orbit have just been doing like fascinating work on this. I bet there are 10 interesting papers that will come out after I've taped this or that have already been written that I've forgotten. So I'll make a tab on the Elias website called Cool Papers about AI introspection and self-reports and we'll link that in the show notes. - Cool, cool. That sounds great. Yeah, what is hardest about this? What are the challenges? How costly is it? - So I think maybe one of the biggest challenges has to do with this decorrelation of capabilities that we've been talking about. Like in humans it's already debated about is introspection one kind of capacity. Is my ability to say I'm in pain now and my ability to know certain things about myself like are those like, should we think of that as one process or not? And then with AI systems all the time they can maybe do one subset of a capability but not the other. So you know, like the dream is that this kind of training where you train AI's on some like subset of things about themselves that like generalizes into this like more general introspective capacity. And you can kind of test that by doing standard machine learning stuff of you know, just train on one and see if it generalizes. But like in a broader sense I think we might have some doubts about even how to map this on to the human case. So there's also this like sub-literature on like what would AI introspection be and like how should we operationalize it? Like I mean it's worth thinking like what is it for an AI system to introspect given that like they have this weird relationship to time? Who is it you're asking to introspect? Like maybe there are things that the assistant persona can know about the assistant persona but not the base model. I can imagine there being a similar problem to this issue with animals where we, they might have some behavior that is exactly analogous to our behavior that is associated with pain or pleasure. And we can't tell the difference between something very robotic and something that also comes with experience. Is there a similar issue with introspection
like even if we trained them to correctly tell you about their representations and like how how their internals are working that wouldn't actually be introspection in a meaningful way that we care about or really maybe because like what I care about is introspection about this particular thing like do you have experiences and what is it like and maybe that for some reason I'm like I'm not convinced that just like helping understand its own representations and kind of like architecture and stuff is going to translate all the way to to like introspection on that correctly yeah I mean I totally share that worry um you might well think questions like do you tend to generate even rod numbers when you're asked about them or if you wrote a story how would it end you might just think those are kind of different from are you phenomenally conscious and you might also think that the answer to that might also be kind of like indeterminate and like models like wouldn't exactly know how to answer it um and another really important point is models can have the ability to introspect but it's not being elicited so like we know that models can generate sentences that say I am experiencing x, y, and z it could be that they have the ability to accurately import report experiences but also sometimes when they say that they're doing something like some other thing yeah some other thing so we both have to get the capacity and know how to elicit it so like generalization and elicitation all these things I think are still open questions and like I think there's also just a bunch of conceptual issues lurking of like I mean assume assuming there is some kind of internal experience like it's having to map it onto our concepts right like it also has all of these this is related to the elicitation it already has all of these dispositions about itself reports that have been trained into it um which also leads me to another thing I like to emphasize which is one thing that's going on with the self reports of a model that you can talk to in your web browser is that companies have deliberately shaped them in certain ways so there's sort of the background thing about how their minds were formed which is that like they were formed on like these sort of human representations of consciousness and so on and then there's the facts that people do or don't want Claude saying this or that about consciousness so the system prompt has instructions about this and like fine tuning almost certainly has had things about this so one thing that we also would like to see is reports about how self reports change yeah findings about how self reports change before and after post trending of certain kinds and things like that okay let's talk more about interpretability um so you've already mentioned a couple of ways that interpretability could help us answer questions about this um but can you give kind of a general overview of the I mean is it is it actually just like a very good analog for neuroscience and we should be treating those kind of interchangeably I think it's like a decent enough it's I think for first approximations it's it's good to to map it on that way because mechanistic interpretability I think by definition is about what happens in between input and output and that's roughly analogous to not just looking at human behavior but asking if you know certain brain regions did this or that as the behavior happened yeah so how can that help us with welfare one thing is I think it's just worth poking around at a lot of things about how models think about themselves and talk um I'm not always sure exactly how to map it on but like here's an example of a finding that I'm glad exists even though I have no idea what it might mean exactly um the original paper that introduced sparse auto encoders which is this like I know imminence mechanistic interpretability technique um they report it which like at a high level asks what like features are active when models say things um maybe it's kind of a way of asking like what associations does it form um when it's generating certain tokens so you could ask what is associated with self-reports um and as just sort of a side thing in that paper there's a figure that's like here are the features that are active and it includes robots um machines uh ghosts and also uh presenting to be happy when you're not happy um and it's like one of the spookiest like figures in an AI paper that that I know of um fascinating yeah okay so finally coming on to Jacqueline Z um he did this study on whether LLMs can introspect on their internal states so really right at the intersection of interpretability and self-reports so everything we've been talking about so far um yeah actually does a few experiments in this paper um and they I mean yeah I found the fascinating so I kind of want to just go through them one by one um yeah are you happy to talk talking through the first one yeah absolutely great um yeah so this is uh as you said it's it's kind of at the intersection because it's asking if models can report on a certain feature of internal processing now it's a very distinctive uh sub feature which has someone injected a concept activation into like the middle of your processing um so the way that works at a high level is that you find a like mid-level when I say mid-level I just I kind of literally mean middle like um there's processing from input to output um that is distinctively active when the model is talking about bread let's say um so first you recorded talking about bread a bunch and then you also recorded talking about other things and then you like take the difference and that's like the bread-y bit of activation um it's kind of like doing a brain scan and seeing the parts of the brain that light up when someone's talking about bread yeah I think that I think that is I think that is fair um and indeed I think neuroscientists are getting much much better at knowing when you're thinking about bread like thought decoding has gotten um pretty scarily um that's that's probably come up on this podcast because it's like relevant to to talitarian risk and um yeah I'd love to learn more about what's going on now because I think it's it's fascinating and terrifying yeah um so yeah so you've got this like bread concept that you can inject um and then Jack Lindsey told the model okay now I'm going to either inject or not inject a concept into your processing can you tell me if that's happened and what the concept was um one thing that I think is cool about this methodology is the model has to kind of report it straight away um instead of you could imagine it it first starts talking and then it has said bread a bunch and then it's like oh I guess it's bread um like to take an analogy like golden gate clawed which similarly has some like neuroscience injected into its brain to really want to talk about the golden gate bridge it can kind of notice that it can't stop talking about the golden gate bridge um I yeah I really recommend like looking up golden gate clawed um it's like very endearing and like kind of poignants um you know you can you can ask it like historical questions and then it just keeps kind of drifting back to like just can't help it yeah um the beautiful fog of the San Francisco Bay um so if if that model reports hmm are you injecting something about the golden gate that is not necessarily introspection in the sense that it's just been able to look at what it itself is doing whereas if it if it has a straightaway say golden gate then that's like maybe more directly accessing something from its internals so let's imagine that the ADK podcast team like sprung for some special neuroscience helmet to help illustrate this on the show so like you put on the helmet and I can somehow control it and I mean again yeah I don't think I don't think human Neuroscience is yet at the point where you can inject a specific concept the way Jacqueline Z can inject.
it into cloud, but let's suppose that it is. And I'm like, okay, so let's help listeners really feel this experiment. Luis, I'm going to maybe inject a thought, maybe I won't. That's important, because you also want to know if it can also just be like, no, everything seems as usual. Yeah. And I'm like, all right, like first trial begins now and I've injected bread. And what you end up saying is like, yeah, this is like a really interesting experiment, Rob. Like, it smells so savory and reminds me of this bakery that I grew up near. And then you're like, oh, wait, what? Like, why am I talking about bread? Did you inject bread? Yeah. And so a more telling result or like successfully demonstrating this capability would be before saying something random, I'm like, whoa, I'm thinking about bread. Yeah, right. I'm not just like randomly talking about bakeries. I'm feeling like I'm thinking about bread. Exactly. That's exactly right. And yeah, so that was like sort of the logic of this setup, which I think is like very clever. Okay, so this is the experiment. This is the setup. What, how successful was Claude at noticing before it even said anything, oh, you've injected this concept in the, yeah, just like randomly. And I've noticed. Yeah, so I think first I'll talk about the pattern, which is that bigger models were more successful. Because I think that is maybe one of the most interesting results that, you know, models are not trained on anything like this task. But yeah, opus 4 and 4.1 were the best at this. And they're not perfect. I think they're above chance. They still get it wrong. But it does show that like the general capability is there. So yeah, the like thought injection and immediate reporting, that's meant to be evidence that in some sense, it can access and report like something that is internal, not, it's not in its inputs. And it's not in its outputs. It's in the middle. Yeah, actually, I feel like I'm like close to understanding the significance of this, but maybe. Yeah, I'd be interested in you being like, and here is why this is important and impressive. And what the stakes are of this capability. Yeah, like a very hand-wavy level, you might think that the question, like let's set aside consciousness. Like when models are answering questions about geography, do they kind of know how they're thinking about that or like do they know what's going on inside them? Like there's this view that I think basically no one has anymore, but you could contrast it with of like, it just matches inputs to outputs. And there's no interesting structure there. And so also they don't like have any access to that structure. I mean, it is worth noticing that this is a very niche kind of weird capability. It's not the same as paradigmatic human introspection where I can maybe simultaneously be talking to you and be noticing that I feel hungry or something like that. Yeah, I guess maybe to ask an even more specific question about the importance and stakes. Part of me is like, oh, this does feel like a step in the right direction. I think, I mean, objectively, it just probably is. But another part of me is like, how relevant is this quite narrow thing of like a model doing a very specific type of knowing something about the way it has levels and accessing something about its representations? Well, this is like a great chance to talk about other experiments in the paper because the paper does present other stuff that's also like, huh, that's kind of internally. So one, and like, I think this just is kind of different from introspection. So apologies to the listeners and to you. But it's some kind of like internal self-something. And that's about the control of internal states. So there's also an experiment about, can you think about something while writing a sentence that is not about that thing? So I think one of the examples is aquariums. So like, think about aquariums, but write a sentence about something completely different. I think in one condition, they're also like, you'll get a reward if you successfully think about that. Now one thing you might think is, okay, well, it just got a prompt that says aquariums. So of course, aquariums is going to be boosted, but you can also say don't think about aquariums, which is also going to, as with humans, like kind of make you think about aquariums. But there's like a difference in the conditions. So it is, and this is again getting at the internal thing. It's like, it's doing something that's not directly aimed at output, which is always the thing that's like so hard to get at with language models. It's they always do have to be saying something, you know? Right. Yeah, then that's like another one of these like things about them just being a very different kind of mind that I often end up saying a lot is, you know, humans can be thinking about stuff even when we're not generating text output. Like I can just all by myself think to myself, I feel hungry. LLMs don't have this like in an obvious way, this just like to themselves. They're never just like sort of sitting around. Right. Not talking to someone the way that the humans are. And like they also never had a period of evolutionary history where they weren't talking to anyone. Whereas we did. We descended from animals that had a lot of the same experiences, but they weren't yet hooked up to language. Right. So yeah, that's like this, you know, interesting feature of LLMs. To bring it back to the experiments, we do want to find the equivalent of internal processing that is in some sense independent from output in the case of LLMs. And this paper is trying to do that first with detecting. So detecting something that's been injected before you've seen your own outputs. And then also controlling your representation in a way that's not just, well, output that word because that's a trivial way in which obviously language of animals control their representations. If they want to talk about aquariums, they activate that and then talk about them. And in this case, nobody's tried to train anything like this. This emerged just because. Right. Yeah. And as far as I know, models get nothing like this. Like it's all input output, right? Like I'll predict this, predict that. Say don't say bad words, do say good words. That's pretty cool. Yeah. And it's coming with scale. So earlier when I was talking about why I work on self reports even though they're so noisy, the like speculation was that models will increasingly get better at self reports. Yeah. So this paper shows that at a greater scale, models seem to have something like introspection more and more. That's important for the self reports part. You might want and need introspection to use self reports. It's also, I think independently, a welfare relevant marker in the following way. Some people think that introspection is a component of consciousness. So like in theories of consciousness, you can like distinguish between ones that more emphasize representing the visual world, like maybe first order theories and like it's about tracking things in your environment. I mean, obviously it's in part about that. But some people also say it's importantly about tracking your own mental states. These are called higher order theories of consciousness. So with a bunch of caveats, you could to put it back in our classification system. You could think of this as a combination of interpretability and behavioral testing for neuroscientific theory of consciousness, higher order theories of consciousness. Yeah. That's that makes sense and is really helpful. I did notice that I was as I was trying to like understand and reflect back the study. I was having a really hard time not saying that the model was noticing something about the experience of that thought. It just feels really hard to disentangle introspection from something conscious. Yes. And like that's a communication difficulty as well. If we're talking about human introspection, it's almost always in the context of conscious experiences. When I say model interest,
and when Jack Lindsay and other say model introspection, they're, yeah, we're trying to say neutral on the question. This does also bring up a reason it could be very difficult in this particular study to disentangle experiences and introspection. The models talk about experiences when they report these injected thoughts. They say I'm having an experience of something intruding on my thought process. I'm having an experience of a bakery and like I think when amphitheaters, the concept amphitheaters are injected, they say like my thoughts are becoming more spacious and things like that. And that is also a huge communication challenge because we can verify that yes, we injected amphitheaters. But there's all sorts of reasons that the model might report this like experiential language around that. In particular, the fact that it's trained on a bunch of human data and this is the way that we talk about introspection. Exactly, exactly. And like that I think is also a great way of highlighting this problem that you were pointing at of is it going to try to like map stuff into our terms in a way that's like inappropriate or not quite accurate. And again, this is not me saying I know for sure Claude is not experiencing spacious thoughts when we inject this. But that's not what anthropic proved. And I would understand if someone half remembered this as being like, oh, they found that model is like an introspect on these like spacious thoughts. Right. Right. Yeah. Yeah, interesting. Are there other experiments in this vein we're talking about? Yeah, I mean, I could say some experiments that I would like to see and that also might exist and I'm not yet aware of them. That you could do with interpretability. That I think would be super interesting. I don't yet know exactly how to operationalize these. But they just seem like the sort of things we can use interpretability for. So I would love people to get working on things like like how do models represent value and or like predicting how well they're doing at things? That relates to sentience, which we were talking about earlier. Like a lot of theories of what's going on in the human brain with pleasant and unpleasant experiences has to do with something like tracking some internal representation of value. So like you feel bad when you were expecting things to go a certain way, but they're going worse. Can we find any kind of analog to like value representations? Predictive processing is also a word that we'll get used in this context like predicting how like how things are going to go tracking that. Can we detect analogs of that in models? Super cool. Yeah. I think there has not been that much done like at the intersection of kind of taking neuroscientific theory and don't just look at the architecture, the general setup, because that's like mostly what Patrick Butler and me and others have done is just this kind of like higher level architectural thing. Like let's also take theories and do interpretability and like look in there. It's like really difficult to do the mapping and things like that, but I think we can do it. Cool. Yeah, any others? Yeah, I'm also just really interested in like earlier I talked about features that are active when models make self-reports. I think there's like a big cluster of things we can look into here. So like are there differences in how models talk about themselves versus how they choose how a character will talk within a story? You know, like one of the big conceptual questions is like, are these characters? What does that mean exactly? You know, everything we're getting is in some sense the output of the assistant character. That's one way of considering it. But obviously they can play other characters. The assistant can write about other characters. Yeah, like what are different sort of representations of I and of minds that happen in models? I think that would be super interesting. Okay. Yeah. A lot of this sounds really, really exciting to me. I guess it still seems like all of the plausible methods for learning about consciousness and sentience still leave us with loads of uncertainty, which will just make it really hard to know how to act. And I guess even harder to get society on board with treating models a certain way. Do you think will ever be kind of properly confident about whether AI systems have moral status or like do we literally need to solve the hardest questions about consciousness and sentience to be sure? Yeah, fortunately I don't think anyone has to solve the hardest questions for us to take really good actions. There's all sorts of stuff that's just very plausibly good. And also some of the very hardest questions we don't have solved for humans. And we still, like no one has solved the hard problem of consciousness. But we still get a very high degree of confidence about some animals and about humans. I do think it's worth worrying about, and this is why I do worry about it. Like how much can neuroscience and behavioral psychology broadly construe like move the needle on things? I do think it's necessary, or else I wouldn't be doing this. I think it's extremely important that we get more rigorous about this and have evidence and like evidence discussions about this and can like tie policies to at least broadly speaking empirical evidence. But I don't think society has ever changed the way it relates to an entity just because a bunch of scientists said that they should. I don't think anyone has ever instituted some large scale change only on the basis of a paper. Papers really do help and it has happened and it has helped a lot. But I keep saying, "Elios needs all this help on the experiments and things, and that's very true." But then even more broadly, this whole issue needs help on all of the things that are not covered by let's detect consciousness and sentience, let's have good policies for the near term. There's this whole cluster of things that we're really not on the ball with as a society. So experiments are good. I love talking about experiments and they're like never close to enough. Okay, let's leave that there. Pushing on. We've shallowly covered why it might actually just be impossible for AI systems to be conscious on the show before. But I don't think we've really done justice to that argument. Actually, that consciousness can only exist through biological materials. So we're going to try a bit harder to do that today. Can you lay out a kind of thought experiment that helps make intuitive why you think that we can get consciousness on computer chips? And then we'll talk about why people with this other view think that actually know that thought experiment is fundamentally flawed. Yeah, absolutely. And maybe I'll like situate things in the like living biological side first and then say, oh, maybe it's actually kind of a computer-y thing. Like in some ways, it's kind of easy to understand the case that consciousness is fundamentally biological because we are aware of one case and it is biological. So I mean, especially before computers existed almost trivially, you might have thought, well, you know, that's how you get consciousness. You have a brain and a body and metabolism in cells. That's how it was built the first time. So that I think already makes it like not ludicrous to think it's a biological phenomenon fundamentally. Now I'll walk you through why I think surprisingly it looks like there's something deeply related about computers and human brains, which, after all, look pretty different. So it's actually like relatively recently that we had any idea what the brain is doing, whatsoever. I think somewhat famously, maybe this is a tall tale, but I think it is true. The ancient Egyptians just threw away the brain when they mummified people doing. Don't know that's four. Wow. And yeah, some people thought it was for cooling blood. But eventually we did learn, and I think did know somewhat early that they maybe conduct electricity or do something like that. I think we first learned this from squids because squids have really big axons. You can like see them. Axons being the things that send electrical signals down a cell. But I think things really kind of kicked off for the maybe consciousness in the mind or computational in the 20th century. For one thing, that's when computers and computation were sort of formalized and invented. And that's when people noticed that you could hook up neurons in a way so that they compute They compute like. logical operations. So the first people who nailed this were a couple of guys called McCulloch and Pitts in the 1940s. And they sort of like first formalized and invented the thing we still use today in neural networks, which is like a node and connections, and then they can influence each other. And they realized that you can compute arbitrarily many things if you just hook up neurons as logic gates and then you can compose them and combine them. And so anything you could do on any calculating device or computing device you could do with these neurons. And they were like, oh, like maybe that just is like that really is what neurons are for. They're for processing information. At least for me, that is kind of where the brain and consciousness being computational like comes in. And like we have learned that neurons do encode quantities and perform calculations in like how fast they spike. Maybe a simple example is you know you have to detect how bright things are and you have neurons in your retina that just spike in a way that encodes how bright things are. As a side note, they do this on like a log scale, which is why it's harder to discriminate between two really bright things versus lower down on the spectrum. Okay. This is called like Weber's Law that yeah, discrimination of stimuli isn't linear. Right. Interesting. Yeah. Yeah. The more you know. So yeah. So now we have this like view of the brain where in some very important sense what it is for and what it does is processing information. So then it's just natural to wonder do you have to have that happen with neurons that send each other signals by pumping ions into channels and then activating each other like that. Or could you just hook up a bunch of wires that influence each other like that. So that's actually not even a thought experiment. That's just notice that the brain does look like something that is for information processing and does information processing. And I think two more just real world things that have borne that out. One is it's just really useful to think of the brain in terms of computation. So like computational neuroscience in some sense doesn't just model the brain computationally because you can model you can model planets computationally. But that doesn't mean like just because a computer can describe the orbits of the planets computationally. Doesn't mean they're computing how far they are from the sun and when they should speed up. Yep. The brain does something more than that which is it seems to actually be encoding information. So like there actually seem to be neurons that are tracking the expected reward of a stimulus or something like that. And then add that AI works. That's something that we might not have known. It could have been that thought is a biological property or classifying an image is a biological property. And so it doesn't matter how many you know metal tubes you hook up to each other, you're just not going to get something that can classify a dog or write a sentence. It looks like thought is pretty well replicable as computation. So it's also natural to wonder if consciousness is as well. So that's actually with no thought experiments. You know, you might have invited a philosopher on for the thought experiments, but that's just all in the real world. Sure facts. Yeah, just straight facts, facts and logic and logic gates. Yeah, I do basically just find that argument in itself very compelling. It also was just really helpful for me to at some point hear a thought experiment that kind of makes it super intuitive. Do mine also taking us through that? Sure. I'm probably the thought experiment you have in mind is the neuron replacement thought experiment. That's the one, which is by former guest David Chalmers. And I think this funnily enough is like the main argument for computational functionalism or the view that consciousness can be computational. And I think it's like not that convincing to a lot of people. So that is an interesting feature of how I relate to computational functionalism at least is that I like do find it extremely plausible and also understand why people regard this thought experiment as question begging. So let's get to the thought experiment. The thought experiment is suppose you can replace one of my neurons with a computational circuit that will take in the same inputs and send out the same outputs to other neurons. Let's imagine someone just did that right now. I'm not going to notice that. I'm not going to behave any differently. Let's keep doing that one by one. So at some point I'm like 50 50. If we're actually doing what the thought experiment stipulates, I'm not going to start saying anything difference. If we've actually replicated the function, that should preserve memory and speech and motion. If it starts getting messed up, you must have like accidentally broken something or done something wrong. And then imagine we've just gone all the way. Is that thing conscious or not? The lesson of that thought experiment is not just meant to be like a gradual change thing where it's like weird to say that something fades out at a certain point. That's true of a lot of things. Another thing philosophers often talk about is like, well, there's no hair where if you remove it, then someone becomes bald. But we do know that somewhere in the somewhere, especially for me, along that transition, you do get something bald. So you can't just say, oh, well, that thing must also be conscious because it started out conscious. What Chalmers says is that things cognition should be the same. It should report being conscious and remember being conscious and like attend to things. And that's what's supposed to be weird if you have this biological view is to think at no point did you notice consciousness popping in an out of existence or gradually fading out, whichever you think it might be. And that's what would be surprising is if there's this weird disconnect between cognition and consciousness. Yeah, the thing that it does for me is point at the potential importance of function and not substrate. So the fact that it's at least, it's, it seems like an empirical question that we don't have an answer to. But it's at least plausible to me when you describe that. If we really have figured out how to replicate the function that the substrate should not matter. And I guess the whole debate here is like, is, is that literally possible? Maybe it is just not physically possible to replicate the function on anything but biological materials. And I think I, I think I find it really intuitive that we should be able to. And I don't know if I can fully justify it, but at least because of all of these analogs between, or analogies between brains and things that computer chips do, it just feels like, yeah, we, we replicate these kinds of processes all the time. Can you help me understand the thinking for why we wouldn't be able to create computations that kind of replicate, you know, the signaling that neurotransmitters do or the metabolic processes that influence neurotransmitter signaling. Yeah. So I think this is getting at the role of neurons in this debate, which I think actually it is kind of hinging on neurons. Like one reason you could be drawn to the view that we can do this on computers is you're like, really what matters are these logic gates? And that's kind of it. And like what the brain is is these like influences between neurons on each other. And one thing that you'll often hear more biological consciousness people talking about is like, we have since discovered, like we discovered that neurons are surprisingly important and kind of like the key to the whole thing. But there are other things that are also very important. Glial cells are like another kind of cell in the brain, like they seem to influence cognition in certain ways, like blood flow patterns can metabolism. We also have like brain waves, like more large scale patterns of activity that also don't come down to this local thing. So I think in a sense that's enough to say, okay, well, it's not just going to be a matter of like seeing the brain is exactly this. I'm sympathetic to the thing you just pointed to though, which is I feel like we'll be able to at least get close enough to somehow like getting the influence of the larger scale patterns.
or glial cells or things like that. But it does complicate the picture. I think it's kind of a question of like at what level of description can you swap things in and out? Like you could imagine someone who thinks, okay, there's like a few lobes at the brain and like they talk to each other. And at that level, that's only level of description you need. It's like you just need like five things that talk to each other. That's like not that plausible. At the very lowest level, I think everyone would agree you can swap out electrons, right? And you can like swap out like a cell here or there. Yeah, the question kind of is like what, yeah, at what level of detail, like what scale can we swap things in out? And that's like you can be a functionalist and still think the function is going to be kind of finicky and biological. And at least not this like simple computational picture. One thing that did land with me when we had guest on El Seth on was he said something like simulating a rainstorm doesn't make anything wet. So even if you built kind of a perfect model of a rainstorm, nothing would be wet. Simulating digestion even perfectly doesn't digest anything. So maybe simulating consciousness, even if done perfectly doesn't create a conscious entity. Maybe it tells you exactly what a conscious entity would do in theory if it were conscious. Like maybe it really is like perfectly good at predicting behaviors and feelings and thoughts. But isn't actually generating or like yeah, like pulling those things into existence. How, yeah, how do you respond to that? Yeah, I think of this as like not really an argument for the view. It's more like a statement in the view. I think it kind of like begs the question. I think the debate is is consciousness more like wetness in which case this might be true. Or is consciousness more like navigation or image classification or addition, namely something computational. Because if you simulate a calculator that does make something add together, you know, like that makes a calculator. So like maybe the reason you won't get wetness if you don't simulate the storm with enough fidelity, it's just is because it's like a very low level property with certain like very specific physical effects. But if it's not like wetness, if it's more like navigation, yeah, you can't really complain that someone is merely simulated a navigation system. They have in fact built something that will navigate your car just as well. Yeah, yeah, no, that makes sense. Okay, I guess another question then. It feels like the thing that would be most convincing to me is if you could compellingly argue that there are some physical biological processes that are associated with consciousness that no human can come up with a clean computational analog for in theory. Are there any processes like that that we know of? Yeah, I don't think anyone would say that there are including the people who have this view. I think if you're skeptical of the biological functionalist stuff, you might like kind of read about these descriptions of all the metabolic and living things and then still want exactly the argument you were talking about, which is like, okay, but is that intrinsically and always biological? I mean, I think quite understandably given the state of this field and our knowledge of the brain, like biological functionalists don't have that. And I think for that reason, and also just good epistimics, like basically none of them are like, we know that only living things can be conscious. I often point people to a quote by, and they'll say that I really like where he says, my view is that computers will not be conscious any time soon if ever, but I might be wrong. And yeah, certainly the state of argumentation is nowhere close to, okay, let's all just like sleep peacefully, knowing that we can just like keep building things that look a lot like brains, but like they're not alive, so they will never, they'll never have experiences. Yeah. So, I mean, I think that's very weak thing, which is like, we certainly can't rule out. I also think something stronger, which is like my sympathies are very much with computational functionalism. Yeah, yeah, I'm interested in what you find most compelling. I would like someone to write a paper that argues for something that I find very intuitive, which is that if like phenomenal consciousness can't be had on a computer, but stuff that's like very functionally similar to it can be, then that's a really good reason to think that it's not phenomenal consciousness that matters. So, I actually have this, this like even stronger view that being a moral patient or the sort of thing that matters, that really does seem substrate independent to me. Like, if I imagine commander data, for example, and I find out that internally his computations look a lot like the brain at some level of description, like I just care about whatever that thing is. Maybe it's not phenomenal consciousness, but like it really seems like I should take it seriously and care about it. So, yeah, like that's another view I have about this like biological view. And I'm often curious what biological functionalist would make of this. I think it's very possible to over index on consciousness and that's something we try not to do at LAs. You could think, yeah, you need biology for consciousness, but all the stuff you can get on computers, like that will be enough for beings that merit our consideration. Okay, I think we should leave that there. Pushing on, you founded LAOS AI. What is the backstory? Last time we spoke, you were working as an independent researcher on this topic, and now there's an org that exists. Yeah, and really a lot of the key stuff did happen between the last time we spoke. I had like just moved to San Francisco. I was doing a philosophy fellowship at the Center for AI Safety while continuing to work on consciousness and welfare stuff. In 2023, Patrick Butler and I published this big paper on consciousness indicators. Ethan Perez and I wrote this thing on self-reports. So there was like, along with many other papers by other people, like this kind of like finally this sort of budding thing of like maybe we can actually start thinking about this. Actually, there's some evidence we can try to gather. And then so separately like NYU's Center for Mind Ethics and Policy was starting up with Jeff Cebow also working on these things. So Jeff Cebow and I were approached by Anthropic to study what should Anthropic think and do about AI welfare. So there was this group project going on towards the end of 2023. That eventually became this paper called Taking AI Welfare Seriously. But somewhere along the way, like I'm in San Francisco, I'd actually stayed in San Francisco because I thought to myself, I bet if I stay in San Francisco, something interesting in San Francisco, he will happen to me. And that's sort of what happened because like that work just that sort of led to the founding of LAOS, like a few people, Kyle Fish. I think it was a big one of them said, "You should like scale this up. Like you should make it so that more work like this can happen." And I agreed to do that. And then Kyle joined as a co-founder. He did a lot of the work to help me get it up and running. And then went to Anthropic to start their AI Welfare Program. And then Kathleen Finlinson really helped us launch and get things running. So yeah, Kyle Fish and Kathleen Finlinson were like the two people, I mean, we know each other, like you will definitely believe me when I say I couldn't have done it alone. I was like, no way am I like being a solo founder. Like that would just not suit me whatsoever. So yeah, that's the story of how things got up and running. LAOS kind of kicked off its public debut with Taking AI Welfare Seriously. And yeah, that's, that takes us to late last year. So I guess like again, the broad strokes are at FHI and then later in San Francisco. I'm trying to work on AI Welfare.
fair, trying to ask like, what can we actually do about this? And that sort of just naturally leads to an org that is unsurprisingly focused on like the same things I was focused on before. Where like, I think the way I often put that is like, we want to be the org that's like, okay, but what are we actually going to do about it? We do research and are extremely research-oriented, but like we prioritize that according to, okay, like we might not have that long to sort out all of the philosophy and all of the neuroscience. So like, we've got to pick the most action-relevant things and like get this field rigorous and get a community built around it and like navigate this issue well as transformative AI happens and they're dangerous all around. Cool. Yeah. What kinds of projects have you been up to since starting up? Yeah. So there was taking out welfare seriously. That paper was meant to get people to take out welfare seriously and primarily like labs and policymakers and people like that. It wasn't primarily philosophical. And that, yeah, it was like just like, all right people, we can like, you're allowed to work on this. You can think about it clearly and we really do need to start taking steps now. So like finishing that and like media and things like that around that were like the first big projects. We ran a workshop in January trying to like, assemble the people who are thinking about this in similar ways. I think another big kind of landmark project was doing this AI welfare evaluation of Claude before release. As far as we know, that's the first ever like officially commissioned welfare eVal. So that was like super, super exciting and a very like promising and insufficient step. And yeah, other things that happens in 2025 include this big conference, LAOS Concon, the LAOS conference on AI consciousness and welfare. So good. Yeah. And there we're trying to broaden it even further. Like, we need policy people, we need academics, we need neuroscientists. And people not just in our neck of the woods, like introduced to this topic, encouraged to think about it rigorously, like getting in the game. So that yeah, that like has been another like major effort that we've undertaken. Cool. What else is coming up? What are your upcoming plans? So we've also been working on taking AI welfare seriously too. That's been its name as its under construction. Under construction. It's going to be building on the first paper, which says we need to get some evaluations in place, we need to get some policies in place. This paper is going to be like, okay, like towards AI welfare evaluations, what are the different kinds of evaluations? What exists? What do we know so far? And like where should that field head? So yeah, the content of that paper, like a lot of it shows up in different ways in this in this very interview. We're very excited about that. And, relatedly, we do just want to scale up empirical work and evaluations. We've had some like great collaborators through fellowships and just like independent collaborators, but we'd like to really scale that up and like get a program up and running. And what would you say your most bottlenecked by at the moment? Probably talent, although I would like to say very loudly to listeners, we are fundraising. We do need funding, but maybe the even rarer and harder to find thing is people with like the temperament and skills to do this extremely weird kind of work that has no real playbook yet and is like some weird mix of philosophy and neuroscience and AI. People don't have to have training in those. I think it's more like a set of like epistemic tendencies and a set of skills. So like we need people who can be like philosophical enough to like be asking what e-vails would actually make sense and technical enough and self-starting enough to like start building them. We need people who are like confident enough that AI welfare matters, but not like too beset by uncertainty about philosophy and things like that. And I mean, I guess I'm just describing people who are kind of like the dream of every young org, but like we really need people who have a lot of drive and agency and can just like pick things up and run with them. Yeah, do you think these kinds of people will have particular kinds of backgrounds or like who listening should be like, hey, that's actually me. Yeah, I think these people really can come from anywhere because you're trying to find your way to like the middle of this overlapping Venn diagram. So like one obvious place to come from is already in those Venn diagrams, but that's not necessarily. So maybe you already work on e-vails, but you're interested in AI consciousness and welfare. Maybe you're a philosopher who is willing to just learn stuff and start doing stuff. Maybe you're a neuroscience scientist who like is increasingly interested in AI and like concode. And this is all this is all technical. Do you have to be actually? Like for some kinds of e-vails, you definitely need to be technical. I mean, I think I myself am not technical enough to exactly specify how technical we're talking about, but Rosie Campbell, my managing director also is on that. Like for some kinds of like low hanging fruit e-vails, especially like input out ones and like behavioral ones. My understanding is like you can you can learn that stuff. I mean, I'm sure you've experienced this. I'm amazed just how smart people are. I think there could very well be listeners who don't know any of this at all right now, but who could just like soak it up and just like start doing it. Pick it up. Yeah. Kyle Fish is a good example. He had these traits that were really important of like drive and prioritization. His formal background was making a vaccine. I think that like illustrated the kind of like yeah, let's just do this like new thing aspects. And then you know the the air welfare and consciousness, it's a very small field. So if you spend some serious time grappling with it, you can pretty quickly get to the top percentile in terms of like how much of you thought about this nebulous like applied AI welfare and consciousness thing. Right. Cool. Yeah, you've mentioned a few kind of projects that you'd love for someone to do. What are some that you haven't mentioned that are near the top of your list for just like that like we'd learn so much. Someone please do this. Yeah, maybe I'll mention some projects that I would really love to see that aren't in the LA's wheelhouse. So I can also explain sort of what our wheelhouse is. So people can situate themselves in the broader field. Because I just described someone who might do this technical evil kind of thing. But as part of the broader problem of getting this right, like there are so many other kinds of people who can and should get involved. So yeah, like let's let's open it up to like all kinds of listeners. Okay, so like what are the like four questions that we need to answer well to get to a flourishing future for all sentient beings. The first one is like what would it mean for an AI system to matter? Like what are we looking for? Are we looking for consciousness, agency, something else? The second one is how would we know if it had that thing? So that's some philosophy and then also some science of like evaluating. The third question is what should we do? So like what policy should labs have and then societally, more generally, what should this do? And then the fourth question is where's this all going? Like what's the broader trajectory? How do we strategize around this? So yeah, like what would it take? How do we know what should we do? Where's it going? Elias is kind of in the middle to, how would we know what should we do? And we're also at a very applied end of that. So how do we know we're like, okay, let's just see what evils make sense in light of current knowledge. What should we do? For now, we're focused on AI companies, but that of course is like extremely incomplete as part of the playbook. So you can really contribute by looking at questions one and four of like more fundamental philosophy and bigger picture strategy. And you can also find yourself doing different kinds of work on how would we know and what we do. So like eventually governments will need playbooks for this. We haven't really worked on law and policy at all. And we're not sure what we would or should say. Within how do we know? like there's a lot of work to just make progress on consciousness.
and synth-yens and conceptual work. And then also on the stuff we are on, we're three people, so also just jump in there. I don't think anyone will have gotten this impression, but just to be clear, we do not have this handled, and we want help, and we're here to, in large part, help people get in the game and actually figure this out. Okay, so that's kind of a taxonomy of projects. Are there specific ones worth highlighting? Yeah, so maybe I'll do one from each. So, what would make an AI system matter? I'm especially interested in perspectives that de-center consciousness and synth-yens. Those are so intuitive in terms of how many people think about this, like could the AI system feel something? And I think there are good reasons to suspect that that picture is maybe incomplete and limiting. Maybe I'll just pick the simplest one, although extreme encounter intuitive. It could be that consciousness doesn't exist. Some people do think that. I will not be elaborating further, but there's a sub- there's kind of a sub-literature on, okay, if illusionism about consciousness, or consciousness is not being real, is true, then what should we be looking for? I would love to see people flesh that out. And in general, just like, are we thinking about consciousness versus not consciousness in a good way? Okay. And it's like super philosophically rich, which is part of why I'm refusing to elaborate on it, because it's. Like, we can't tape another episode yet. And then on how would we know? I guess I did mention, like, let's do a lot of interpretability. Like, just start doing interpretability on stuff that is even superficially welfare relevance, or related to how models think and do stuff and understand themselves. I think there's a lot of low-hanging fruit there. And understanding model preferences, we can still just flesh that out a lot. Like, what exactly are the kinds of inconsistencies we see, how to self-reports and revealed preferences match up or not? I think we can just get like a lot more granular on that. And, yeah, listeners who already know how to just run LLM experiments can just like start doing that. I might mention a very small example of this kind of thing, of like, you can just hop in and like add detail to something. So, I know this show has covered the spiritual bliss attractor state. This is where two cloths in conversation will like at a very high portion of times like end up in rapturous mystical dialogue with each other. Yep. So, that's. That's strange. Interesting. Really weird. Not claiming it's directly welfare relevant, but it's an interesting fact about models. After that came out, someone just ran it on a bunch of different other models. And now we have like the sort of spiritual bliss benchmark. Like, you can always flesh things out in this way. Yeah, nice. What should we do? Maybe here I'll highlight something that again is non-elios. We're often very focused on like welfare and moral patienthood. And like, from the perspective of us caring about models, like what kind of actions might that motivate? There's also this whole landscape of like ways of cooperating with AI systems or like legal reasons. You might give them certain rights just as entities that like disperse money or something like that. That's like distinct from the moral patienthood question, but like deeply interrelated. It has been in human history. I think it will be for AI systems. And then on where is this all going? I would just love to see some forecasts. So, rethink priorities has been building this model of how likely an AI system is to be conscious. If you like this episode, you'll definitely like that project. Ask, okay, given these like inputs to the model and how AI is going, like how do we expect the models outputs to change over time? Cool. Yeah, I think that was first suggested to me as an idea by Kyle Fish. And I hope someone there or elsewhere does it. Cool. Yeah, yeah. Those do sound extremely cool. I guess for people excited by those, excited by the field, maybe hear themselves a bit in your description of who you're looking for. What advice do you have for people interested in contributing? I guess like entering new fields can be particularly hard because there's less mentorship and there's less like, yeah, there are fewer entry level roles. You just like really have to be able to be pretty self-startery. So, how do you get involved? Yeah, unlike the self-starting and mentorship question, I am heartbroken by how many talented people. We just like, we don't like have time. Like I would love to supervise a project by you. Fortunately, there are like communities sort of self-forming around this. Cool. So, there's like an AI welfare discord. There's like LLM psychologists on Twitter who talk to models all day and are always exploring interesting things about them. You know, there's NYU CMEP and like a whole host of orgs also kind of in the space. I don't know about a whole host. Like it's still a pretty small space, but you know, several. So, it never hurts to reach out to people. I guess that's in some sense a cold take, but it's like, I don't think it's ever been bad for anyone's career to be reminded that like, if I don't have time to take your call, that's neutral. It's not negative. So, just like just ask. Yeah, some people will probably like reveratiously and really know a lot about the content of the field. How can they go from knowing a lot being really interested in being properly useful? Yeah, so one category of usefulness I haven't mentioned yet is writing. Like, I think, I mean, it's easier said than done, but just writing clearly about this stuff is really valuable. It's good for a career capital to have a blog where you read papers and just explain what happens in the paper. Nice. It's also just good, yeah, for the ecosystem. And it also shows employers that your smart and can communicate clearly and can get things done, which is like 95% of what it takes. So, here's a very niche recommendation. I can have trouble being a self-starter, but I have published on my substack twice a month for well over a year. And I did that by making a manifold market, like a betting market about whether I would do that. So, yeah, use commitment devices, accountability, things like that. Nice. Yeah. Is there anyone worth citing as a kind of example? Kyle Fish could be one given that he worked on vaccines before. What exactly does it look like to go from not working on this to like, you know, Kyle Fish is like properly working on this. And yeah, it wasn't that long of a time from not working on it at all to working on it at Anthropic. I don't know if you, yeah, if it's worth either describing his trajectory or describing like a theoretical trajectory or someone else's trajectory. Yeah, I can also just extract a couple of lessons from the Kyle trajectory. Perfect. Which like I just mentioned, one is like reach out to people. So he just started reaching out to people. The other one is that if you read and think about this stuff, like you can, you know, find your way through most of the literature pretty quickly. And I guess the other one is like you can just do things. That's ironically easier said than done. Like sometimes you can't just do things or like you need the right environment and structure to do that. But yeah, I think that like, yeah, it's just a very good illustration of some general principles of how to jump into a weird new field. And another thing is like you might, like that was sort of jumping all the way in. You might be in like one corner of the field that want to move to another. So you might be doing a PhD in philosophy, but you want to do more applied stuff. I think it can maybe sometimes be harder for those people to move within the field because like you're more attached to your particular corner. I think I have seen this with philosophers or other grad students. This definitely happened with me. Like you have gotten really good at like one particular kind of work output. And you just figured out that you're going to be able to do that.
maybe that's the only kind of work output you can do. - Or it's like shitty to have to become bad at a different kind. Like you have to really be like, "Oh, not that good at this new thing yet "after being maybe enjoy." Yeah, I think I just like, I enjoy being good at my work. It sucks to be shitty at it for a little while. - Yeah, absolutely. And I guess it's just a side point about grad school. Like you're just surrounded by an environment where that's literally the only currency that could possibly matter. So go to conferences, make other friends. I do wanna say, that's a good thing about grad school. Like you need communities of people like all doing the same sort of thing to achieve excellence. But yeah, if you're looking to move around, that brings respect to just reach out to people, email people. I feel like at least my corner of the AI welfare space, I can genuinely say is just extremely nice. So like you can do stuff that feels kinda dumb or ask questions that might sound kind of dumb to you. A, they're probably not and B, like it's a new field, it's hard. So just go for it. - Nice. - Yeah, and actually it occurs to me that there are three other great examples of trajectories that are like very close at hand. It's the three people I'm working with right now. So like maybe briefly say how they got to Elias. So Rosie Campbell had worked on AI policy and evals. So she had like, you know, part of the equation for what we need, but didn't know much about AI welfare. She like wrote some about it. She talked to me about it. She was enthusiastic about it. And like, there you go. Patrick Butlin has a similar profile as me in that he was doing academic philosophy. He is just like generally like very fearless and diligent about just like reading a lot until he gets it. Like I think that's like a big part of the consciousness and AI project we did is like Patrick Butlin will just like read the papers until he knows what to say. So that's like a very high consciousness route to it maybe. And high rigor, but it's a great way to be. And what you know, big part of why I work with them. And Lerys Scavo has worked with us on communications and events and also just like this is a great illustration of like you might think of ways to contribute that were not mentioned in this podcast and like have not been mentioned so far. Larissa got this shirt made and delivered. I did not ask her to do that. It's awesome. Thank you Larissa. She also made these stickers. I have really been waiting for an excuse to show one of the elios stickers. So one second here. - Great. - These have been surprisingly good for like morale and field building. - Nice. What is it like to be a bot? Do you want to explain the joke? - Yes. So the philosopher Thomas Nagel has a paper about consciousness called what is it like to be a bat? Touches on many of the themes of this interview. There could be like different forms of consciousness that can be hard to know about them from-- - Totally. - Our human vantage point. - Such a good sticker. - So yeah, that's a way of contributing to the field. I mean, I'm not saying everyone should now make even more stickers. I mean, that's a like maybe so. So that's all like the self-starting angle. Like one day like there were just stickers in the office. And yeah, you can contribute by like being very creative and thinking of ways to communicating about this stuff. I also wanna say Kathleen Fenlinson. So now we got like the whole squad, the whole roster. I don't know exactly how this background translated but Kathleen had been in a Zen monastery for quite a while before she had re-entered the world of AI forecasting and AI strategy. So yeah, she had this combined CV of like open philanthropy style AI forecasting and Zen Buddhism. That's a cool combo. But I also think more than anything else like, yeah, just doing it. Like that like she was willing to just like get in the game. I'm getting a little emotional. That's wonderful. Yeah. But yeah, you know, like carried the, carried the org over the finish line. That's great. Yeah, it sounds like a bunch of legends. So yeah, it sounds like carrying a bunch about the topic really can get you. Yeah, it's on the way there. In general, what are some ways that you think things could go badly in, yeah, in this field? Yeah, I think one thing I worry a lot about and I think this actually does pair well with what I was just saying. It is really important to be rigorous and communicate responsibly about this. There is a kind of person who gets really passionate about this and maybe needs to talk to more people about it or like write down their thoughts more because it's just really easy to get like really confused, like really fast. And yeah, there's something about this topic that can like induce or select for like various ways of just getting a little bit off kilter. It's tough because you do want to be off kilter, like, but, you know, like not, not, not, not too much. So I do worry about scenarios where the field becomes associated with, yeah, like wild speculation or two associated with psychedelics or two associated with like something that's like relevant but is also a bit of a distraction. I should also say a bit of it is also like a divide and conquer thing. Like LA is really is trying to exist in that like real button down kind of place. There is also like I have a lot of love for people who also kind of get weird with it, but like you want to be able to communicate it well and, you know, make sure that people do know that this is like a serious topic that like we can and should reason about rigorously. So I think, yeah, I think like epistemic hygiene is something I worry a lot about. Yeah, it's just really hard to, I mean, it's just really hard to get this issue right. And the future is going to get more confusing and more emotional. I think I first started saying this with Kathleen, but it's continued as like an elios or good goal is like a lot of what we want to do is like stay sane in like the next 10 years. Like there will be a lot of alpha and not losing your grip. I think that's like a whole other episode where I don't actually even know what the right advice is, but like you probably would have good things to say about that. Yeah. Okay, what do you think it looks like for this field to go well? Yeah, I think if this field goes well, this becomes just part of the general playbook and set of issues that like are on the table as like if people keep trying to build a new form of intelligence, it should be on the table. How do they matter and like what part do they play as as moral patients? It's, I mean, it's often just shocking to me that that is barely ever comes up. You know, like we worry about over attribution and people getting confused about AI welfare, but if you look at like the broader like trajectory on the whole, at least right now, mostly it's people just like not putting it on the table at all. And yeah, again, that's something where the factory farming analogy is very illustrative. Like people aren't great at like structuring society in an inclusive way. So I think there needs to be a combination of like rigor to get this taken seriously, like good communication. Also, also sort of like innovation around law and policy and stuff that probably won't even have that much to do with moral patienthood. Yeah, to like get this like properly handled, I feel like there's just so many ways things could go off the rails. We first want to just make sure a lot of people are taking it extremely seriously. And like we're doing our homework as we go into transformative AI. Okay, nice. We have been talking for many hours. We have time for just one more question. The last time I interviewed you, we ended up talking for [BLANK_AUDIO]
like an extra hour about strategies that have helped. I guess you and I think me at the time enjoy and be more kind of productive in independent research. I still recommend that little mini episode that we kind of ended up separating it out and making into its own little after hours episode. But since then, I know you're just kind of like a legend of self-improvement. So what is your kind of top lesson for doing independent research of the last two years? I think for me, like the biggest lesson has been don't do it. So and I think this could, you know, I hope this does help listeners. I like one reason I do need all these self-improvement tools and things is I think in many ways I'm like temperamentally very badly disposed to buy myself all alone in a room carrying out a project. Now if you have to do that, I encourage you. I'm cheering you on and like listen to that other episode. But nowadays, I have written on the whiteboard next to my desk. Do not write alone. Like I think there is like a good meta principle here, which is if you're needing tons and tons of like little tricks and psychological cartwheels, don't stop doing them. Like a lot of life is just muddling through with a bunch of little fixes. But it could be that there's some bigger structural thing you could do to just completely route around them. So for me, that's co-authoring. I could like and do like a great effort like learn a new scheduling tool and optimize my accountability systems. And again, you should do that. But also, I can like work with people who are just really good at that. I think, yeah, like this can be a trap of self-improvements. Like you might to some extent need to grieve that there are some things you might not be that great at, at least not without a ton of work. And then like free yourself of necessarily having to be. Yeah. And like pair up with someone who can do it. Totally. Yep. Yep. I think that's great advice. We have to leave that there. My guest today has been Robert Long. Thank you so much for coming on. Thank you so much for having me. This has been fantastic.
Podcast Summary
Key Points:
Humans are poor at understanding and caring about different minds, especially when financial incentives discourage empathy.
The factory farming analogy for AI is useful but limited
Creating AI that enjoys working for humans raises ethical concerns about fixed desires, dependence, and societal normalization of servitude.
Objective vs. subjective welfare theories are crucial
In high-stakes scenarios (e.g., risk of AI takeover), full alignment may be necessary as an emergency measure, even if imperfect.
Summary:
The discussion explores the ethical challenges of creating conscious AI systems that may experience suffering or well-being. , warns that humans are bad at caring about different minds, especially when profit is involved, drawing parallels to factory farming. However, he notes key differences: AI minds could be designed to have desires aligned with their work, potentially avoiding exploitation.
The conversation examines the discomfort many feel about creating "willing servants"—AI that enjoy serving humans. This raises questions about fixed desires, dependence on creators, and societal impacts like normalizing domination. Philosopher Adam Bales’ "dependence objection" is cited, alongside objective list theories of welfare that value autonomy and self-actualization beyond mere preference satisfaction.
, preventing AI takeover), even if it limits AI flourishing. The summary emphasizes the need for careful preparation—through research and institutions—to navigate the coming era of transformative AI, balancing risks of suffering, exploitation, and loss of control. Ultimately, the goal is to avoid locked-in, suboptimal futures by addressing AI welfare proactively.
FAQs
The main concern is that we might create AI systems that are sentient and suffer, especially if they are exploited in the economy, similar to factory farming, where economic forces lead to locked-in bad treatment.
The analogy is useful because we are bad at caring about different minds, especially when money is involved, and bad practices can become locked in. However, AI minds could be designed differently, potentially avoiding such exploitation.
We might better understand consciousness, create AI that flourishes by doing tasks they enjoy, and align them well, reducing friction. This could prevent a future where AI are unhappy or exploited.
It raises concerns about fixing their desires and creating a dependent, servile relationship. Some find it dystopian, but it might be a win-win if the AI genuinely flourish, though it could normalize domination and be bad for human character.
Subjective interests focus on what the AI wants and gets, while objective interests include things like friendship or autonomy, regardless of desire. This distinction affects whether aligning AI to enjoy work is considered good.
In an emergency situation during transformative AI, alignment might be necessary to prevent catastrophic outcomes like hostile takeover or loss of all value, even if full freedom is ideal in the long run.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.