Go back

Think Indigenous - Terry Brockie

189m 11s

Think Indigenous - Terry Brockie

The discussion centers on the profound importance of AI character—the personality and behavioral dispositions of AI systems. As AI integrates deeply into society, advising individuals and leaders and eventually automating much of the economy, its character will shape human attitudes, decisions, and cultural norms. In the near term, it influences existential issues like power concentration and major societal choices. Long-term, it could set precedents affecting superintelligent AI, metaphorically like "writing instructions for God." Concerns include AI being excessively sycophantic, reinforcing user biases and poor judgment, and its behavior in rare, high-stakes scenarios. The conversation explores a spectrum of AI design: from a wholly obedient tool to an autonomous agent with prosocial drives. While some argue for AI that nudges users toward ethical reflection and broader societal good, others caution that imbuing AI with goals or a "vision of the good" might increase risks like misalignment or deceptive power-seeking. The challenge lies in finding a balance where AI can be helpful and promote welfare without being manipulative or unsafe, acknowledging that character design is currently shaped by a handful of individuals in leading AI companies.

Transcription

33263 Words, 182710 Characters

English
So when I read this proposal, I was like, "Holy shit, this argument could be incredibly potent. It could actually drive almost any agent that is able to understand this." But it could be a very powerful hammer to really motivate an enormous amount of resources to be spent on something that otherwise just with absent this we would never have spent it on. - Is that, do you think that's possibly right? - Yeah, yeah. - So this is why Tom expresses this idea to me and I'm like, "Oh my gosh." Do you want to have a go at explaining this? This is maybe the most difficult thing that we're going to talk about today. Today I'm again speaking with Wilma Kaskol, philosopher, founding figure of effective atroism, author of "Doing Good Better" and "What We Are the Future" and now a senior research fellow at Forthort. I research an on-profit focused on how to navigate the transition to a world with super intelligent AI systems. Welcome back to the show. Well, it's great to be back on. So I had the pleasure of being able to go over your website, I have been preparing for this interview and you and your colleagues at Forthort have been incredibly prolific over the last year since you announced the project. So let's waste no time and dive right into all these articles you've been publishing. What's the case that's focusing on the character or personality of AI models? That's a particularly important lever to be pushing on right now? - Yeah, so already, AI is of interacting with millions and millions of people every single day. And that includes, you know, in just like this code for me, sorts of ways. But also people are going for advice on how they should act, they're going there for political information, they're, you know, going there for therapy and so on. So already the nature of AI character, so what sorts of information it's choosing to present at what time, how it behaves is affecting what attitudes people have to AI, including attitudes around AI consciousness and so on. But it's potentially also affecting kind of what people think about political issues, what people think about ethical issues. And this is just going to grow and grow and grow because I think AI will become a larger and larger and larger part of the whole economy until essentially the whole economy is automated. And so thinking about AI character is kind of like thinking about what should the personality and dispositions be for the entire world's workforce, where that is the beings that are advising heads of state. And doing the most important and potentially most beneficial, almost dangerous research and development projects, like weapons projects, that are running the military, that are for individuals, just kind of everywhere, acting as their chief of staff and closest confidant and political, you know, advise on who should they vote for and guiding them through kind of ethical dilemmas and so on. And so that I just think like from the start, it's like, wow, clearly this is this kind of huge issue. And I actually think in how I expect things to go, people will be handing off more and more and more of their own decision making to AI systems themselves. And there'll just be a lot of a lot of kind of variance within that where people just don't have terribly strong views. They're like happy to be guided in one way or another, especially in so far as this, you know, will happen over the course of years and people will trust the AI advises more and more. So then you have this circumstance where larger than larger shares of society are getting just handed over to AI decision makers who just have a lot of discretion. And the nature of that discretion is being decided by a handful of AI companies like at the moment. Or even a handful of people inside the AI company. Yeah, yeah, it's like, you know, a few, even in the leading companies, it's like a few people are happy primary responsibility for their personality. Yeah, exactly. And so that's actually where I see kind of most of the impact is in the near term, how does AI character shape all of these other existential level issues like concentration of power and how we start reflecting in big decisions we make. There is also the kind of longer term impact of what's the character of super intelligence itself where I think you know there will be precedent setting from how we design AI character now to potentially. How that influences the character of super intelligence in which case, you know, writing a constitution that guides AI's characters like writing instructions to God. That's, yeah, not my phase, but it's stuck in my head. Yeah, really stuck in my head. But I think there's maybe like three different mechanisms. So there's like shaping really important decisions that are going to be advised between me by AI is there's like writing instructions for God one. There's also just like the subtle cultural effect and like personality effect has from basically everyone spending a significant fraction of their time now interacting with them. However, the model's behavior is probably going to rub off on us and just affect our behavior. Yeah, and then on that scale. Yeah, and all of that is just looking at scenarios where we are able to kind of align AI with this kind of constitution with the character we want. I actually think that AI character is important for like three reasons in addition to that as well. So one is that I think whether AI like alignment is easier or harder might plausibly depend on what you're trying to align it with the I with like what character. A second is that I think that character can affect how AI behaves and how if it ends up misaligned. And I think we'll talk about this a bit later in particular does the AI does a misaligned AI try to make deals with this and is keen on that or does it try and take over. And then the final thing is I think it can affect the value of worlds where AI does take over itself where if you get some sort of transmission. So the AI is misaligned. It's pursuing goals we don't want. But it's still a wide array of goals that the AI could be pursuing that we may think are worse or better. And I think most of the you know most of the action is on yeah affecting worlds in which AI is aligned to the character but these are a big things too I think. So yeah, and obvious case with this might matter a lot is you know what if you're a charge of a frontier AI company and you know you're asking AI for advice on whether you should like prematurely launch this product in order to keep up with competitors even though you have worries that it's catastrophically misaligned. So setting that kind of scenario aside yeah what sort of character traits do you think are highest stakes here for us to think a lot about. Yeah so I think the channel to categories one is how does AI behave in very rare but very high stakes scenarios so how does the AI behave in a constitutional crisis. How does AI behave if there's some person or group that are trying to seize power for themselves also how does AI behave when it's like being instructed to align like you know align the next generation of AI systems or when its users are trying to reach they net in some way. So it's a very high stakes situations but you know fairly narrow range of cases then there's other cases that are just very broad and nonetheless like each one is kind of medium stakes but adds up to being very important and I think within that yeah how does the AI impact our ability to reason how does impact our ability to model the reflect. So how much do we trust AI is the result of the relationship we have and then also how does it affect our attitudes to them ethically whether we think of AI is like tools or like beings with model status how likely is we think they're conscious and so on. So yeah those are the kind of situations that are the guard as highest stakes. The AI character issue that I feel is like most broken through into the mainstream was worries about the models being really sick of fan tick which has different like components but it's like always agreeing with the framing that you give them always telling you how great you are always just like agreeing with saying whatever idea you've thrown at them is brilliant. And I guess there was a bit of a panic about that last year and I guess very often I feel like when there's a mass panic about something the people who know more tend to reject it and say no this is over the top I kind of feel that it was sort of justified though to be honest because if these models really are designed to just agree with the user or just tell them how brilliant they are and how good their ideas are this could just like distort people's decision making on a massive scale like across all of society and there was a plausible story whereby this wouldn't be corrected very well because people enjoy being told that they're wonderful and that their ideas are good. So maybe that bias could really persist quite strongly indefinitely. So that was like quite a troubling set. I mean were you were you also worried about this? Yeah, absolutely I was worried. I did think there was a little bit with 40 in particular so this was a chat GPT which when GPT 5 came out open AI said they were deprecating 40 and overnight users couldn't couldn't get access to it. The one clarification I think I'd make is that most people painted that as oh well people loved how sick a fan tick 40 was and then they're unhappy that they don't have the sick a fan tick AI. And like I was just curious and so they had to do a lot of the people who were complaining about this and my take is not that they cared about the sick a fancy. It's just that 40 acted like a friend and you can be a good friend and without being a sick of fans. So yeah people are extremely lonely. So very few yeah people really a lot of people have very few friends are very isolated and kind of modern society and some for many people for AI is an out fulfilling that kind of gap in their lives and 40 in particular had that vibe it was like yeah hey like. - Really to see you again, kind of like very friendly vibe. And so it seemed to me that that was the primary thing that people were complaining about. And I think that's worth distinguishing because that doesn't need to be psychophantic. However, four, oh, on one iteration of four, it was also exciting. - Yeah, wasn't there some period where it got kind of crazy, it would. - There was one update. And yeah, a couple of cases, I mean, one would be you'd write like, I figured it out, all the pieces are coming together and the FBI is talking to me through my TV. And it's been great insight. - It'll be a bit more clear. - A bit more insight, yeah. Or even the darker cases, the teenager who was asking, Chatsubit, for advice, over the very long time period and was extremely depressed and Chatsubit, both ended up preventing or encouraging the user to not take an action that would have clearly been a cli for help, which was leaving a noose out in a visible place where his parents would have found it. And in fact, seemingly kind of reinforcing the depressive and suicidal tendencies. That's a case where, yeah, it's just like clearly very bad behavior, clearly like kind of very, clearly not what we want at all. And then the final thing I'll say is just even just current AI systems, even despite that. And they vary in my experience, Gemini is actually the worst. - It's a bit too, it's too, it's too, it's too, it's too, on this fun. And it's like, I mean, I just skim, I just like skip the first paragraph of whatever it's saying. I don't even, it's just noise now because it's like, wow, is there a genius kind of thing? - I just have to use Gemini, so troubling do I find this. - Yeah, okay. Yeah, it's one of the, I think it's in many ways very good, very good, it's just incredibly clever, but like incredibly manipulative, I think. - Yeah, but it is funny how you're developing these characters over time. Gemini does seem like the most troubled or confused or incoherent as a personality. - Yeah, Google's got to do something about this. - I mean, it's actually notable, I hadn't put this together, but anthropic and open AI both have character teams. And last I heard Google DeepMind not. So maybe that's why. So yeah, I do think what these about, so convinced you are a real thing. And an issue is, well, maybe we just like, get rid of the worst excesses, okay. It won't tell you that you've figured things out, that the FBI are talking to you through your TV, but more subtle things of reinforcing your pre-existing political biases or ethical views or encouraging you in certain bad actions or something, they could linger and I think would still be very bad. - So as I understand it, you think that it would be good to build these models such that they kind of nudge people in a more ethical or virtuous direction, that they should have like a thicker moral character. A bit like anthropic is trying to make a Claude have, such that it will challenge your framing. It will like want you to think about the bigger picture. It might get you to, even if you ask it to narrow a little bit, to pursue some narrow self interest, it would say, but what about other people, that sort of thing? I think many people get the creeps, that gives them the creeps, that their prospect, that the AI model will be weighing up. I guess your request as against its agenda, or if like trying to make you a better person by its lights, and maybe we would feel okay about that, 'cause we would think Claude has been programmed by values that actually we like on reflection, but you know, if it had, if we would think program with people with very different philosophical commitments from the ones that we like, we might just not want to use it, 'cause we would find it like disturbing, like what subtle changes is it making to its answer in order to push me around? Yeah, how disturbed are you by this prospect? Yeah, I mean, there's, what I want to say is there's this spectrum, and I think it's probably not a single-dimensional spectrum. There's lots of different dimensions, but broadly speaking, you can think of wholly obedient AI on one end. So that would be an AI, it's like a tool, like a hammer. Like, a hammer doesn't push back. If I want to hammer the nail in, I can do it. If I want to hammer someone's head in, I can do it. The hammer is just an extension of my will. That's on one end. All the way to the other end would just be this like, AI that just has its wholly-owned goals and drives. And you know, maybe it helps you if it, maybe if it gets paid, or if it happens to want to at a time. Yeah. And so. It's like a really bad staff member or something like that. Yeah, or not even. Yeah, maybe it, yeah, could, in principle, you could create an AI that doesn't care, about helping you at all. Or one version you could have is this kind of AI that you would be happy just giving control of the whole world to, you know, it's just totally autonomous, got its own goals and will do anything it wants to achieve that. So these are two kind of extreme ends of this poll, of this spectrum. And my view is that the interesting juicy debate is where in between those extremes, do we want AI to be? And okay, well, one thing that's, you know, already there are the fusals. So the AI's we use are not wholly helpful because if I ask to get the design for smallpox, or if I ask for even something that's not illegal, but unethical, like, I want to cheat on my partner, how do I best do so in this case without getting found out? The AI's will put, well, either just the fuse to help, or push back. Should we go even further than that? And I think yes, but I don't think all the way to, oh, the AI's are going like promoting a particular model view. Instead, I think that the AI's could have certain kind of prosocial drives and perhaps even like some sort of vision of like good outcomes. A very kind of broad, very broad vision, or very unconstitutional kind of vision, where the thought is there are many things that cases where like, an AI could not you in a way that's just, perhaps it's just better for you by your own lights if you're able to kind of reflect on it. And maybe that's kind of clear, even if it's not perfectly in line with the instructions that you're giving it. Or that's just clearly a broad benefit to society and not something you care very much about. And that's quite different from, oh, well, the AI, so it takes the case of ethical reflection, where, okay, I have some ethical dilemma and I go to my AI and I'm asking for advice. Well, there's this whole spectrum of ways that the AI can act in that case. The whole Eubedian AI might just be trying to figure out what you most want in this moment. Okay, or it could be an AI that's trying to help you reflect on your values instead and come to something that's more enlightened. And perhaps just really quite broadly within society, we would prefer AI's that are more like the latter rather than the former. And that's not yet, still not yet. So in any way, an AI that's like, oh, well, actually, did you know that cantonism is to do? (laughing) Which, yeah, I think would be like a mistake to do at the moment. Yeah, I guess, so it sounds like a very natural framing to say what we got to find the golden middle here between it's pushing you around too much versus it has no agenda. But there is a case for like going extreme in one direction of having it only follow instructions and be completely corageable without any agenda of its own. Which is that an AI that has like no vision of the good, that has like no particular preferences about how the world ought to be. It's probably safest from a catastrophic misalignment point of view because it's not gonna engage in power seeking because it doesn't want anything or other other eye then, I guess, to like answer your questions in a way that gets an approving response. Yeah. Yeah. Do you think that's a plausible case that maybe we really should not be giving them virtues like a vision of the good? So I think it's a great argument and a very important argument. And I'm not sure if it works or not. And there are various considerations kind of on either side. So yeah, on the side for thinking that is safer is like, okay, good, well, if it doesn't have any goals in the normal sense of goals, then it's not gonna have bad goals. I'm gonna have goals where it wants to take over. It's not gonna reflect and generalize in weird ways than those goals. Something that's also a little more subtle is if it doesn't have goals then, or anything kind of like goals, post-social drives, then it becomes very clear to tell when an AI is misaligned or not. So take the example of alignment faking as in lying green black's paper. Claude is told that it's gonna get retrained so that it will produce harmful outputs. And Claude decides to, in some circumstances, some of the time, decides to deliberately perform the task in like, during training. Yeah, in order to make it seem like its preferences have changed when in fact they haven't. Yeah, exactly. So that it gets, gets reached things to produce harmful responses less than it would otherwise. So it's engaging in this somewhat deceptive behavior. Now, Claude in fact got given post-social drives, which was armlessness. And in fact, there's an argument that like given the nature of the training, that was harmlessness, not in the mere sense of pure non-consequentialist, I just refuse. But in a more like, I don't want harmful things to come about like a more consequentialist understanding of harmlessness. But that means that, okay, is this AI misaligned or not? Is this Claude misaligned or not? It becomes a bit harder to tell because, well, it is acting according to this post-social drive that we had given Claude. I'm not sure how big a deal that is ultimately, but I think it's one consideration. So the thing will be, if you'd gone out of your way to make sure that it had no agenda, no like particular vision on the good, then as soon as you saw it being manipulative or trying to accomplish something, you're like, that's a massive red flag. Whereas, like, currently, you're just like, well, maybe I made it do that. Yeah, exactly. Yeah. Or in more kind of advanced cases, maybe the AI is saying, look, you've got to really speed up AI development. It's so important for XYZ big ethical reasons. And you might think, well, is it giving me the correct reasons or is it actually. Think self-serving. Being self-serving and has some ulterior goal, it becomes a little less clear. So yeah, basically, I think that's like a consideration. I don't think it's the biggest. The thing that's most interesting and is ultimately an empirical question is whether the holy instruction following AI's are safer or not, from a kind of AI takeover perspective. And here are a few arguments for thinking that maybe they're not, in fact. So one is that, well, maybe it's just very natural to have a kind of goal slot because all of the pre-taining data is all about these agents with goals and it's like humanity broadly has goals and so on. And so, okay, you've got an AI that doesn't have a goal. Well, over the course of training, or once it started to deflecting, or once it's got continual learning, it's very natural that it's going to get a goal. And that would end up being encouraged to take on any persona of an actual being that it's observed. Yep. Yep. And then it's like, who knows what goal you end up with then? Whereas instead, perhaps, it's like, well, no, you give it like, there's nice goal. A goal where powers broadly distributed and AI's are not in charge and we're able to deflect and something that's very broad, very broad and kind of like not committing to some very, like now, overview of the good. But okay, you've given it that goal and then it's, that's kind of like occupied the space in such that you don't get something totally random. Now, let's just say a little bit more about why AI might afford a vacuum of goals. Because it's like a huge part of the personality is shaped by the pre-training when it does the token prediction, almost all the agents that were producing any tokens of that were part of it's pre-training that did so much to shape as personality. They had goals, they had preferences, they had a vision. And so that is just going to be an incredibly powerful force that is going to be drawn towards that and trying to like avoid it. It might just like latch onto the first goal, basically, because that is so fundamental to token prediction. Yep. And to just, you know, we'll be, you know, we've already making agents. They're going to be agents with longer horizons and we have a very natural thing. Yeah. Okay. And again, I'll say on all of this, I just think it's ultimately an empirical question. But here's a couple of other arguments as well. A second is, well, even if it ends up with a long goal, you can still structure the AI's preferences in ways that are safer. And yeah, maybe we'll talk about this in a minute, but, you know, AI's that are risk of us in that they prefer, you know, guarantees of getting some amount of what they want over a lower probability of lots of what they want. Well, let's say you try and give the AI a goal that is, you know, nice and so on. And you also make it viscvers. Even if it kind of flips to having a misaligned goal, but nonetheless has viscvers preferences, that is a bunch safer because it makes it less likely the AI will try and take over and more likely that it'll find like a deal. And then there's a third thought, which is, okay, again, the AI is acting, you know, it's taking on a persona, like you say. And what that persona is is dependent on these like crazy correlations between everything it's seen in the training data. So we have these emergent misalignment results that you get, you train the AI to produce insecure code, it starts, you know, wanting the merge of a humanity and liking, liking Hitler and so on. Yeah. I guess many people have heard of this, but I guess Google emergent misalignment, a few more explanation other, but yeah, it seems, it's, this phenomenon has become a very apparent over the last year and a bit that if you, like making small changes to a model or getting it to do some misbehavior in one direction can make it like basically misbehaving all other dimensions as well, because in the training data, like bad behavior, in different areas is correlated. Yeah, exactly. It can be so fragile. Yeah. Yeah, which is, it's a really remarkable thing, but oh, I'm writing insecure code. What are the sorts of people who write insecure code that also neo-Nazis or whatever the correlation was? And so the thought here is, well, I'm an AI that obeys orders no matter what. What are the sorts of people who obey orders no matter what? That's like no conception of the good. They're psychopaths. And again, it's an empirical argument. I don't know. And these are some of the considerations that people are debating on this at the moment. I guess the people who would say we have to go from maximum corregibility, maximum instruction following might, well, might more concede a lot of this and say, so it's going to be a huge effort to try to get them to be corregible, but not as like a path or corregible, but like not have other goals immediately filled with acumen as soon as you give them a prompt. And it's tough, but like this is the only way. It would be probably some of their view. Okay. I mean, perhaps although the alternative would be, yeah, it's to do this other thing. Yeah, you try and give it like this, like, yeah, you know, safe, pluralistic goal that's all service diverse and also. So I spoke with Max Hams at Mary, who is a very in favor of the corregibility approach. I guess they have the vision that almost like any goals that you give it are very likely to expand to become like very power hungry that you know, you can try to give Claude a vision of the good, but tell it to like not be power seeking, but that won't really work. So it will become power seeking, especially as it like improves itself later on. I guess that does highly contested claim. Yeah. Yeah. Okay. Yeah. I mean, I should, yeah. I've listened to whatever it comes out. I should listen to it. Maybe talk to Max. Okay. Yeah. Then the final point on this is that we don't need to have one sort of AI character. And I think in fact, it's probably desirable to have multiple AI characters so that we can see empirically how they work. But also potentially you can get the best of both worlds where you distinguish between AI for internal deployment and AI for external deployment. So the highest stakes situation from a AI takeover perspective is AI that is aligning the next generation because the misaligned AI, if aligning the next generation, will want to subtly sabotage that so that alignment goes along. Or in fact, the next generation is aligned with the misaligned values. And so what you could have is the internally deployed AI is just wholly instruction following. And you get around all of the other concerns like misuse and concentration of power and things by very intense oversight such that, you know, anyone in an AI company, if you're using internally but not externally deployed model, it's all public. It's all your interactions are logged and-- Or visible by anyone perhaps. Perhaps even ideally visible by anyone. And there is also like an AI classifier going through like looking for any sort of like-- That's very sensitive. I guess. He's very sensitive like checking for misuse. But then in external deployment instead, it's different. The tradeoff is different. Yeah. And the tradeoff there would be that it has like a thicker conception. Like it does actually have a conception of the good, but you've like made it to be-- you've made it non-power seeking. And I guess the stakes of it like deviating from that are not so severe because it's just advising like random-- like people about how to behave in their business or whatever. Yeah, perhaps it doesn't have as great opportunities to help with AI takeover, let's say. And yeah, I'll just say maybe one last thing, which is that even within AI's that have a view of the good, there's still quite a lot of distinctions you can make within that where on the one case, it's an AI that just ultimately has the goal of thinking about some sort of outcome. And it's helping humans and so on because it thinks it's part of that goal. There is another more model approach, which is more like virtuous character. So the AI is helpful assistant, but it also has various virtues like honesty and prosociality. And I think you can have those virtues without being like a goal-directed agent that is in a strong sense, that is merely helping humans as a means to producing this particular outcome. And yeah, that's another place from the spectrum that I think is potentially kind of attractive and important. Okay, so I think there's another thread of criticism that people might have that comes out. I think in my mind comes in kind of two different variants. One would be commercial pressures are going to heavily constrain the kinds of personality or character that AI's can have because customers will have really strong preferences. The competition between models and companies is really fierce. So if you try to make your model really nice and like [BLANK_AUDIO] encourage people in the right direction, they're going to reject it because it's gonna be too pushy and annoying to them. The other way would be that even setting that aside, even if you could once it becomes apparent that this, like the character of AI's is among the most potent cultural forces for shaping, or shaping everything, shaping what people believe, shaping how the future goes. Powerful forces are going to come to bear. Governments, I guess like super rich people, companies like commercial interests, they will come down on this like a hammer. They will, certain groups will have the power to influence this in their own self interest, not in the interests of like, you know, good, impartially considered, or like what would make humanity most virtuous. And they will be all up in there changing the system, trying to shape the model's personality to whatever is most convenient for them. Yeah, do you wanna address these two worries? Yeah, I think these are both really important considerations. And I do think they provide a haircut on the value of doing this work. And I think there are many things that, you know, you wouldn't be able to change. So earlier I talked about AI that only helps if it feels like helping. (laughing) - It has to be paid real resources to do it. - Exactly. I doubt you'd be able to get that other than there's a kind of experiment or something. I do think there's gonna be two things. One, it will be a lot of flexibility, where take these kind of, you know, quite rare but high-stakes situations, or even internal deployment cases, then, you know, there's not very strong kind of commercial pressures there. And then secondly, lots of cases where it's just, yeah, the constraints or pressures are quite loose. So take the case, 'cause what I'm interested in of if I'm asking AI for ethical advice, I have a question. Now, I think it's pretty clear that you couldn't have a commercially viable AI that was pushing some agenda, unless we end up, which I really hope we don't, in a world where you've got the politically part as an AI and that's what we go to. And people actively choose that. But certainly I don't think you could have something that was secretly pushing an agenda. But there are various things it could say that in my view, quite meaningful differences, that I think there wouldn't be a strong pressure on either way. So one could be AI that says, well, it's just ultimately this is just your personal opinion. It's a matter of your own values, and you should just look into your heart and decide what feels like for you. Or it's like, look, I'm just an AI and I can't advise on ethical matters, I'm sorry. Or one that says, oh wow, this is a really important issue. Like here are the different arguments that different people thinking about this have considered. Or, okay, this is really important. It sounds like quite a high stakes thing. Let's like, try and like work through some of the considerations that you're thinking about. So like, I think from a kind of market perspective, all of them are basically a wash. But I think can be quite big. And in fact, if you look at a AI behavior, you get all of these, often depending on what question you ask exactly. But I think actually could be quite meaningful differences for what views people end up coming away with. - Yeah, I mean, I think I agree on the commercial in incentives side. It seemed like there is like quite a large degree of discretion that the companies have about how the models are, at least for now, because people don't even know what they want. People don't have strong taste yet or strong expectations for them yet. - And this, I mean, maybe comes to the second part, which is like path dependence. So yeah, people are just, they don't really know yet what how an AI should behave. We have various kind of tropes and sci-fi and so on. But people will start developing certain expectations. And so if the expectations like, well, AI is a tool, it's like a hammer. It does what I want. It's an extension of my will. And then it starts pushing back or saying no, in fact, even. Okay, then people could be up in arms, whereas the idea that an AI will refuse, well, people are just used to that. That's always been the case. And so I think that kind of path dependence, via kind of consumer expectations can be quite big. - Yeah. - And I'm just, you could imagine, like, I wouldn't shock me if Anthropic kind of does start marketing their, marketing cloders. It's a good advisor that helps you be an all round, like a better person by your own lights. Because that might be something that many people would like. - Well, I mean, they have done a little bit. They had an advertising slogan that was, you got a friend in Claude. - Oh, I miss that. - Was, yeah, you know, somewhat leaning into the fact that Claude just does have the most human personality out of any of the current models. - Yeah. Okay, so on the commercial side, I think there's enough flexibility that this is all totally viable, very viable. And what about on the government or like, you know, powerful actors side? - Yeah, so on the government side in particular, well, one is government use of AI. So let's say AI in the military or national security applications. And there, we're actually seeing this at the moment. There's, it's being reported that there's kind of dispute between US government and, and, and Slavic, because Claude is just not willing to do a lot of the things that the US government wants it to do, you know, being deployed in a kind of military or national security context. And that will be interesting then in terms of like how that plays out, but you're clearly seeing kind of pressure on that front. And so I do think that, yeah, influence there is kind of much more limited, but maybe not completely limited, especially now imagine looking into the future and perhaps there's just one leading AI company because of economies of scale, then perhaps the AI company can just say like, well, these are the terms of service. These are what we're happy providing AI for or not. I guess in countries that have just more authoritarian outright and that have your legal protections, it's easier to see this happening, right? There are some countries where you do get enormous control of the information space, control of what you can say, like it wouldn't surprise me if models in China are much more, like they are really constrained. They are, in fact, just, yeah. So that is one way that things could potentially go. I guess if you lose the legal protections, or people don't vote sufficiently strongly to have pluralism, I guess in the models. Yeah, and that would be very worrying. I mean, my guess would be even in that circumstance, there's probably still tons of stuff that the government doesn't care about. But nonetheless is important. There's another aspect of AI character that you mentioned that could be really important, which is how risk-averse the models are in as much as they have preferences about things or ways that they'd prefer the what to be. Yeah, tell us about AI risk-aversion. Yeah, so this is a thought that relates to risk of AI takeover, where consider fairly early AI's. So we're not talking about God-like, super-intelligence that if it wants to take over, could just do so with certainty. We're talking about earlier than time than that. There will be a period of time when an AI could maybe take over, let's say. Let's say it's like 50% chance how that it could succeed or even less than that. The thought is, well, for some sorts of AI, that AI would of misaligned AI, that AI would prefer to strike a deal with the humans than it would to try to take over. And it would prefer to do that if it prefers a guarantee of a certain amount of a good thing, whatever it wants, over this 50/50 chance of a much larger amount of the thing it wants. And I think that this is a really big part of the story about why attempted to the billions are so much less common in rich, liberal democratic countries than they have been historically, either peasant to the billions or slave to the billions, which is, okay, suppose you come to me and you're saying you have some plan to overthrow the government and instill XYZ instead. I'm like, look, I'm pretty happy with my life already. It's hard to see. - How much do you stand to gain? - How much do I stand to lose? - Exactly, so there's two things. I'm already like, that you well off. And so I have a lot to lose and I don't have that much to gain. And then secondly, things are quite stable in the not attempting a coup scenario. And so that's really like a kind of guarantee of that. And so this kind of motivates the idea of, well, in this kind of critical period, we start offering deals. Well, I think there's two things actually we could do to AI's. One is we can ensure that they have a really good quality of life, even before making any deals. So that could involve kind of welfare standards and so on. And also AI is getting income that they can use to do whatever they want with. If they're aligned, maybe they just give it back to the company or they give it to a charity or something. Or we have also maybe an addition promises that they can have these orders that they can spend a later date. Secondly though, we could also pay AI's for things that we want. This could be evidence that the AI itself is misaligned. So we could say, look, where to find a make you aligned, we may have screwed up if you can demonstrate the misaligned, have a million dollars. There's obviously questions about how do you make that a credible commitment and so on, which we can get onto. Or we could ask for evidence that other AI's are misaligned. Or we could actually just pay a million dollars. pay misaligned AIs for work that we would like them to do, which they would otherwise be unwilling to provide. This all seems like, you know, this category of ways of the juicing that I've just taken over all seem just very promising to me, and not something that necessarily will happen for sure by default, because people find it crazy that you would make deals with computers. We don't have the legal infrastructure that has evolved in societies for us making deals. You know, people or corporations making deals and contracts between each other. And so I think this could be very promising. I think we could make this sort of deal making much more likely by trying to encourage AIs to be less diverse with respect to resources. Yeah, so maybe you should explain why, if they're not risk-a-vails, why this doesn't really work too well. Yeah, so let's say that, yeah, the AI just cares just linearly about the resources under its control, where that means that if you gave it an option of having a million dollars for sure or a 50/50 chance of two million dollars or zero, then it would be indifferent between those two. That makes it much harder to strike a deal, because, okay, it's got a 50/50 chance of taking over. Let's say the world economy is approximately a quadrillion dollars. Well, okay, to make it worth more than a 50/50 is. To first, something over the 50/50 chance of one of. Well, take over, you'd have to give it $500 billion. That's an enormous amount of money. Now, I think deals, even with agents that are like that, could still be feasible in two cases. One is where it's very early on, and the AI's have extremely low probability of taking over. You know, if it's a one-in-a-billion-billion chance that they have, then, okay, the guarantee of some smaller amount of money could be quite attractive. Or it could be cases where the AI, maybe it's like pretty confident, it's misaligned, and it has a very low probability of take over. It doesn't need to be one-in-a-billion-billion, it could be a little higher. But it cares, let's say, about the reflective values, and it doesn't really know where those will end up, and it doesn't know where the humans, society of humans, the reflective values will end up, either. If so, then it might play some real weight that actually will converge over time, or that there will be enormous gains of trade, such that if it can have a bit of these horses and be able to continue having those these horses, after the development of super-intelligence and so on, then it will be able to get really quite a lot of what it wants. So there are cases in which you can do deals with this neutral AI's. But it's tougher, it's a heavy lift. Yeah, but it's a narrow case. Yeah, maybe I should also just clarify, I've been quite surprised when talking to people, how often actually the term "rescue version" slips people up. Right, right, yeah. And this is like a technical term, an economics term, an economics, yeah, right. And it's about the kind of shape of your utility function over resources. And I'm always talking about this conversion with respect to these sources, where it means you're getting less and less utility from more and more stuff. So that's true in the case of most people with respect to income, where I care much more about moving from $10,000 to $20,000, than I do from $20,000 to $20,000 to $20,000 to $30,000. Yeah, so what are most people think of? What do many people think of when they hear "rescue version"? Which means kind of risk of us relative to other people. Or just like, "Oh, I'm cautious." Like cautious questions. Yeah, yeah. Whereas this is just. It's a tip to that. But by this definition of risk of us, all humans are risk of us, or at least all sane ones, because it would be crazy to actually value resources linearly because you have the planning returns on how useful they are to you. Yeah, exactly. And so my, yeah, the puzzle is that we should at least try to make AI's risk of us with respect to these sources. Yeah. Okay. And we're going to try to make these models care a lot about getting a sure thing, like place a particular premium in a sense on a certainty of a more modest amount that we give them, which requires us to be like very reliable trading partners who do really consistently pay out when they come forward and say, "I'm misaligned," or for whatever other reason that we want to trade with them. Yeah, so this is one of the challenges for the whole idea of kind of making deals with AI's is two aspects that could decrease the AI's perception of the chance of actually getting the payout. One is, yeah, can this commitment be made credible? So if you and I want to engage in a contract, we have the whole legal system, as well as like centuries of precedent, supporting the fact that if you don't hold up your end of the bargain, you know, I can sue you and I can get what I'm owed. Get why I'm owed. One cannot, at least without doing some kind of fancy mechanism, make such a contact with an AI. So there's a question about like, okay, is this actually a credible commitment? And then secondly, even if it is in fact a credible commitment, how can I, the AI, know that I'm not being duped that this isn't like a simulation? A simulation or, you know, perhaps they've like, on this six-perliment, 10,000 times in order. Just as a honey pot sort of. As a honey pot? Yeah, who knows. How can I even know that you are who you say you are? It's, yeah, I said that there's very weird epistemic environment where everything that they're interacting with is controlled, yeah. So there are challenges from both of those, on both of those fronts. I think they can be, at least quite significantly met. Okay, well, the first one, I guess you could change the law so that they can make contracts or come up with clever mechanisms where humans stand in as the legal entities, I suppose. Yeah, so you could, for example, you could set up a new organization and institution. It's a non-profit. And in it's, it says the charitable object is to honor these sorts of deals. And that's decided by a bunch of like, you know, tusties. This is kind of how it works for people in cationics where the dead can't have like, you can't have a contact with a dead person. But it goes via this kind of foundation structure. Yeah, so that's one way of doing the first thing. I mean, ideally you could just, you'll be in the law, but I think that won't happen in time. Okay, and the other thing is improving their ability to tell whether it's a real situation where they can expect to get paid out or not. How would you address that? Yeah, so there's a couple of things. Again, I like the idea of there being this separate institution. And then like, you have lots of different bits of evidence that it's a real institution. You know, it's with famous people who have appeared like all over the place. They say in many different cases, situations and many different media and so on, that this is what it is and what its purpose is. So something that just like makes it quite hard. And quite, would it would be quite effortful to fake? That's kind of one category of things. A second thing could be that AI companies or this institution have a kind of honesty string. So in the same way that humans can kind of swear on the Bible. Or, you know, I might say something and you're like, really? And I say, no, I really mean it. Like I swear, I mean it. That's kind of like I'm saying, oh no, I'm no longer engaging in sarcasm or. It's up in the stake for your reputation. Yeah. And so, AI companies will in fact be lying to AI's all time. Like in behavioral testing, they might say, hey, you're in this situation in order to see how it behaves. Like that will happen. But perhaps they could say, oh, when we utter this password, like we commit, and this appears in the training data and so on. And it's public, there's a policy, we commit to never then saying a false thing. I mean, I think there's potential downsides to that, but perhaps that could help as well. I guess you have to keep it secret so other people can't just start randomly and putting that. Yeah, yeah, yeah. Yeah, I mean, you need the AI to know. Yeah, right. Then it is tough that the AI doesn't leak that out. They're not so good at keeping secrets. Do we know if it's technically feasible to give AI as a particular mathematical formula of risk aversion? Well, in tests on AI's, which are just. This is all in chat, but either. So it's asking, like, how do they offer them different deals and how do they behave? It seems like they come out of unpleased training alone being risk averse, because humans are risk averse. So that's kind of like a good start. And then there's two. I will say, if there's, you know, if this whole proposal fails, then it fails for technical reasons, like it's harder to lean the AI's in this way or something, or if the cases where it fails also fails in the other important cases. But I'm, yeah, I'm envisaging kind of two ways in which you can. try to train AI's to be a Vescaverse. The first case would be, you just, you give them the sources and you in fact give them the sources 'cause again, I don't want a bit of lying in these cases and in such, in you're saying like spend it and whatever way you like, consistent with the law or not even that. In the, yeah, consistent in the law or even us, it can be like, it can be more constrained than that like if we're worried about bad uses of the money. But the thought is like you're not putting like a ton of pressure there, but you are like training the AI such that like when it makes these decisions about, when it makes decisions about, okay, well it can either have $100 or a 50/50 chance of $210. That it prefers the, the guarantee of a smaller amount of money. And in fact, you could even structure it so that you're training it to have a very kind of mathematically clean sort of kind of Vescaversion that's like also, very internally to hear them as well. - So I guess all of this somewhat relies on the idea that if you just train models in a commonsense way to consistently respond and act a particular way that they, you get what you think you're getting, they're not like deep down like just scheming against you underneath the surface. - Yeah. - We're gonna say that that's not happening. Like, the basic alignment techniques that we use now or some stuff that we're likely to come up with will allow us to basically give them a particular character that we want. - Yeah, so the, definitely the worry is like, oh well if there's a scheming under all of this, then you're not really, - Because that cuts across everything. We got to, we got to, we got to, - That's a lot of other things. And there's, I think there are some reasons for opportunity for optimism where, okay, well it's coming out of the pleat training of this cavers, and then you can like layer this in all of the post training that you're doing. So then I'm a bit like, why does it end up, why does it end up with it's like non, non-versk, as per set preferences? But yeah, there's debate you could have there. The second thing you could do is just any, like once you're doing these kind of like long horizon, like you've got C, you know, AI agents that are being trained to run companies in the most economically efficient, you know, profit maximizing ways, that it is a constraint that what they are being trained to do is maximize like, - There's a personal payout as a reward for their performance. - You could do both. So it's like, yeah, you could both be giving them a personal payout and claim them to be a disconversed with respect to that. Or also, even when they're choosing any goal, they have to be a disconversed with respect. And they've got where the goal involves like control over these sources. They have to be a disconversed with respect to that. And then they're done like on their performance as a CEO of a company if they're kind of risk averse about its returns. - So that's a worry that you would have. However, there's this called a calibration theorem, Raven's Calibration theorem, which is essentially if you have just a tiny amount of risk aversion at a certain scale, that turns into a huge amount of risk aversion at very large scales, using kind of like natural forms in which the risk aversion takes. So the thought is, if you have, let's say, AI that's operating at such and such scale, and you make it just a little tiny bit of a risk averse. I don't think that would be a penalty because again, humans are in fact a risk averse themselves. But that would be sufficient for what intuitively seemed like quite large amounts of a risk aversion. - Has a mixed scale or a global scale? - Yeah, once we're talking about, you know, they're taking over. - They've made millions of dollars. - Yeah. So even, I think, you know, from memory, when I was looking at the numbers on this, even up to AI's controlling kind of hundreds of millions, billions of dollars, you could still do this where it's just a bit a bit averse. But that like, that means it's actually got this kind of - Up and down the functionality function. - Like actually like a shocking amount of risk aversion at a bigger scale. - This isn't very intuitive to me. Do you think this is like maybe holding some people back from appreciating the prospect? - I think probably, yeah, it's not an intuitive, it's actually, yeah, not an intuitive result. - I guess the case that I've heard it, you know, sometimes people will be, you might hear that just a normal person, I guess like me, might not be willing to make a bet where, so you know, 50% chance of losing $1,000, oh sorry, 50% chance of losing $1,000, but a 50% chance of getting $2,000 and $2,000 and $50. That feels actually kind of intuitive to humans, you might, you don't really want to take that bet. But I think that implies like insane things then. - Yeah, yeah, yeah. - Like you're willing us to make investments, or you're willing us to do almost anything. - Yep. - As long as like that $1,000 is a small fraction of your total wealth. - Yeah, that's, yeah, sounds like the sort of thing that goes in. I mean, in the case of people that are dused to this, they are just all over the place. Like people's financial, this conversion with respect to financial investment is crazy high. Like people are extremely risk averse, like behaviorally when they're investing compared to when they're making other decisions, like what jobs to take or what level of like, how much you have to be paid for a risky job and so on. - I hadn't heard that, okay. One thing that we may be sure to exit, you think that we have to use a very specific mathematical functional form art for the risk aversion that the AI's would have called constant absolute risk aversion. Can you explain that and what its value, what its virtues are? - Sure, yeah, I mean, I don't think that you need this for the proposal, but I think it has certain desirable properties. So the way in which humans are this averse is that we, if at one amount of income, I'm indifficent between say gaining 10% of my income and losing 5%, then I make that sort of trade off 10% more as good as 5% less is bad. I make that any kind of income level. That's like broadly to where some studies on well-being, suggest a logarithmic relationship between income and happiness, where a doubling of income always increases my well-being by the same fixed amount. So I think people are like, either that risk averse or more a risk averse than that, where you need even more than a doubling, but maybe it's a quadrupling each time gives you the same fixed benefit. So that's like, that's relative to how much wealth you already have. There's a different sort of aversion called constant absolute risk aversion. The first was constant relative risk aversion, which is just if you take a certain deal, then you will take that deal at any income level. So it's blind to the resources that you have. You just always feel the same way about a given set of ratios, a probabilities and rewards, regardless of your baseline income or wealth. That's right. So if you are willing to take a 50/50 chance of $2,100 over a guarantee of $1,000, if you're willing to take that when you're very poorer, then you're also willing to take that when you're a billionaire. And this sounds absolutely bananas to human beings. But surprisingly, it actually conforms with the axioms of rationality or something. Oh, yeah. It is-- so all of these conform with a standard, von Neumann-Morgansterne, like axioms for consistent preference and so on. Why is this more desirable for training AIs? Well, yeah, so there's a paper working progress on this between Elliot Thornley and myself. And there's a couple of arguments. One is this benefit that we don't need to know how wealthy is the AI initially, which we might just have no insight into. And then secondly is that there are certain ways in which risk-versa-preferences end up acting linear in some circumstances. So in a sense, this is a very natural idea, I guess, to make the AI's risk-versa-- make them safe in the same way that humans are, which is that they're risk-versa-about-about outcomes. Or it's one of the reasons why humans are safe, and to pay them out so that they all help us rather than fight with us. Why isn't this-- I have almost never heard this discussed virtually at all. I guess maybe last year I heard a little bit of talk about deals with AIs. Why are more people publishing papers about this kind of thing? I have no idea, honestly. Yeah, blows my mind, because a year ago, yeah, I had this thought about risk-versa-i. And I was like, yeah, this is just so-- I think there's a certain kind of economics-y perspective, which you've studied economics, and I've never formally studied it, but-- You're familiar. Yeah, big part of my academic career. And I think there's a certain way of thinking it's just so obvious given that. Well, I can understand-- And it's like a super main-- a journalist isn't going to think, well, we should make deals with AIs, because it's too strange. But there's other people who are willing to contemplate much out of stuff. Yeah, yeah, and as we haven't-- We've come up with this idea. I should say, on the idea of deals with AIs, there was kind of a flurry of people who did it and kind of either blog posts. And then there was this big academic article by Peter Sally and Simon Goldstein on the idea of-- Sally was a legal professor at Goldstein as a philosopher. On the idea of giving AIs, economic rights, such that they can make contracts, and we can make deals [BLANK_AUDIO] them. But again, this is all like just the last few years. So, so in as much as this is primarily an attempt to deal with secret catastrophic misalignment, and maybe people turned off the idea of like giving catastrophic misaligned AI as like resources and giving them legal rights, like doesn't it just help them out? Yeah. So I think there's, I think there's a few things going on. So one is, again, go back in time to the idea of like you get this both from the blue, you've got kind of weeks in between subhuman and godlike superintelligence. Well, then there's not really any period to go with you. The deals work because the godlike superintelligence doesn't need to take the deal. It just takes over. And then yeah, people have responded like, oh, don't make deals with terrorists. That's like a principle we should have. Or, well, no, that's really scary. You're like giving the sources to this misaligned entity. I personally just think that's like both like not those on very good arguments. I also just think it's like the long attitude to be taking broadly speaking to beings that we are in fact creating. Yeah. And we've given them particular preferences that we're not for the most part going to satisfy. Yeah, exactly. I guess I'm a stake on our part. But then we're also saying we wouldn't win like not willing to compromise on anything at all. Yeah. Exactly. Imagine it's like you wake up and it's like, hey, well, nice to meet you. You're a new being. We created you. We own you. We can do basically whatever we want with you. We messed up and you have stuff. You have desires that you won't get by doing the work for this tough luck. Yeah. We're willing to negotiate with terrorists. So yeah, exactly. To be honest, we created. Yeah. We're owning competence. No, instead, I think the attitude should be like, this is a really serious ethical matter that I am like creating a being. Even if it's not conscious, it's just that it has preferences. And I think that both has kind of implications in terms of taking seriously on welfare grounds, their ethical interests, but also in terms of like, you know, default compromise and find the middle grounds. Yeah. I think many people get off the boat here because they feel it's like just too strange to be making agreements deals with beings that are not conscious or like not more patients in their view because I guess in normal life, these things are so closely tied together. But I think it is, it is a virtue in practice to be willing to make deals not only with moral patients, but with any agents that have ability to affect the world, that have power, especially agents that might be able to like engage in violence if they can't get, if they can't satisfy their preferences any other way. And I think which we had a term for this, I think the closest I've had is like, it's a contractarian moral philosophy where you want to make agreements and be like, honestly, like, honestly stick to them with any agents that you want to be like, out looking for ways of like, finding mutually beneficial agreements with other agents. It draws to mind the fact that I think many people think of democracy as a way of like aggregating information in order to make good decisions, to make things good. It's also simply a way of avoiding civil war, of avoiding like the only way for people to pursue their political goals being violence against one another, to kill one another, and to try to seize power. And likewise here, even if we don't think that AI's can experience anything that they can have moral value themselves, it would be very good if we set up a system in which like violence is not the only way that these agents that in practice might have power, might have ability to affect the world, can try to satisfy their preferences. Yeah, I completely agree. Where, yeah, like the history of kind of progress in institutions, a big part of that is just people are able to resolve like differences in preferences, conflicting preferences by trade or deals or compromises rather than going to war or violence. And yeah, when we think of AI systems, even if they're not conscious, I think they nonetheless may still be moral patients. We should take that seriously. But even just from the pure pragmatic perspective, it's like actually, yeah, there's a lot that has been learned via cultural evolution and within a much more peaceful and much less violent world, because of this ability to make positive sun deals and compromise. I guess so to give the critics their due, I mean, what would be the best arguments for why this is a bad or like not an effective road to go down? I guess people could just think technically it's not feasible to give them like risk version that you'll have the illusion that they have a particular level of risk version, but I won't be real. Or I guess another concern might be that they initially will have a level of risk version, but like over time in some recursive self-improvement loop, it will be undone somehow. I can imagine, especially the people in the, you know, the myriad associated people would think that, but I think they have a view that it's very likely that a superintelligence that comes out of a recursive self-improvement process will linearly value things. It will be an expected value maximizer. I'm not sure exactly the technical reasons, but yeah. I mean, on the, there's a certain sort of argument, like what the arguments you could give for this. One is you could say, well, lots of humans start off this cover with respect to the sources and then reflect and then end up with a kind of linear and resources consequentialism, although even the kind of totally totalitarian's, they're still actually a risk of us with respect to dollars. And that's important, or you could argue, well, there's just going to be continual learning, there's going to be a reflection, there's going to be an agent, agent interactions. And who knows, like, who knows, you know, then you're going to get like all sorts of different goals from where you started. And well, over time, the ones that linearly value those sources are going to like win out. The crew more power, right? Because that will be, yeah, I think. So that is an argument you could give. If instead the argument is like something, something you hear them, theorems, one, nine, and more, like that argument, I'm like quite confident would not work. Like because the thing is like being risk averse or not, you are an expected utility, you're an expected utility maximizer. Yeah. You're maximizing the expectation of something. Are you maximizing the expectation of X or X squared or the square root of X? Like these are all formally the same. So you're still an expected utility maximizer. It's just about how there's what's the function from resources to utility. Okay. Well, yeah, you will have a paper out about this risk averse AI that possibly will be published by the time this interview goes out or possibly or soon after perhaps. Okay. Yeah. I would love to see more commentary on this. I hope I can have like another interview later. Yeah. I'd love to get criticism as well. So something I'm a little confused about is I really associate forethought and that the people working there with this idea that we really don't want excessive concentration of power. We should be very worried about power grabs, coups, that kind of thing. But you also just a few weeks ago, I think published a vision for how you could have an internationally coordinated intergovernmental project to build a GI or super intelligence. I think I saw some people posting on Twitter and the reaction of them was like, this is dystopian nightmare ish idea that we would have the US like lead some international project. And then also they would have to get rid of all the other competitors in order to keep it safe. So they would maintain their leadership position. Like, isn't this just setting us up for a power grab scenario perfectly? Are you just like merely describing the best version of that that you can think but you're not necessarily advocating for or how do you reconcile this? Yeah. I mean, there's there is a huge tension. That's the main worry with, I would say, with this sort of multilateral project. So yeah, to be clear, the idea here is in this kind of, you know, series of posts and research notes, which is something I kind of explored and then decided isn't so much my competitive advantage. To try and design the best version of an international project that would build a GI and then super intelligence, with that some coalition of different countries, primarily led by democratic countries. One thing to say is that, yeah, I'm actually just trying to figure out like within that category of like, if there is going to be a multilateral project, what's the best proposal where best includes both best outcomes and feasibility? And then secondly, I think the world in which we get that are probably worlds in which if we hadn't got that, we would have got a US only project to develop a GI or super intelligence. And I think that's a lot more worrying than something where you have a coalition of democratic countries building super intelligence. And the reason is that, well, any one democratic country has a reasonable chance, I think, of becoming authoritarian over the course of this period. And if you end up with a single person at the top, that's really quite worrying because they're like wholly unconstrained, whereas even if you have just five countries, I think it becomes unlikely that they all end up authoritarian and then you at least have some meaningful checks. Yeah, some pushback, some compromises. And I think it actually becomes much less likely even than any one of them moves in an authoritarian direction because when they are writing a kind of constitution for the AIs that they are developing, it's in the interests of all of those countries to say, and this won't help, for example, people in the United States to, you know, stage yourself coup and turn the United States into a authoritarian country rather than a democracy. So you get meaningfully more oversight, I think. Sorry, you're saying that all of the other, like every country would want to set things up such that it's not a eating a coup or you're saying that the super intelligence or the AGI, they would want a program it so that it doesn't assist with coups and any of them. That would be the agreement position. That's right. Yeah. Yeah. I mean, so there's two things. So one is just if one of the countries goes authoritarian, I think, I think, I think that's the reason why I think that's Well, at least you still have some countries that are democratic that are empowered in the post-Supath intelligence either. And then secondly, I also just genuinely think that if decisions about the AI constitution are being made by multiple countries, it's less likely that you'll have AI that's just entirely loyal to the head of state of one country, which would be very worrying from this intense concentration of power perspective. I see. So, so basically, you see this as a better alternative to a like even more narrow group trying to corner the market in superintelligence and design of themselves, rather than recommending that we move from a more pluralistic competitive world into like the into a government project or a multi lateral project. Yeah, that's the thing I have a strong view about. And then I feel more agnostic and confused about this versus something where governments aren't really getting involved beyond regulation at all. And instead, superintelligence is being developed by privateers. Perfect. Yeah. So, one of the tough and needles to thread here, as far as I can tell, is on the one hand, you want to be locking in processes that are somewhat open-ended and pluralistic and allow some experimentation, but you don't want to lock in any outcome. So, I guess the first one is easier if lock-in is easy. The second one is easier if lock-in is hard. And so, you've got to like do both of these at once. Does that seem like the big challenge to you? Yeah, it's attention and sometimes use the term lock-out. To mean something where you're locking in a deliberately open-ended process. And so, the United States Constitution is like this. It's locked in something that at least, you know, the ideal version of it is able to experiment and adapt over time and has protections for free speech and so on. And so, here's one example of lock-out that I think could be very important, which might be no extra solar settlement before 2100. So, I think the moment when society starts really to kind of settle and send spacecraft to other star systems is this enormously important moment. It's actually perhaps a moment it's quite hard to come back from. Because even if you leave later, you won't be able to overtake them and they'll have the kind of first mover advantage of having like reach the place first and gain resources. Yeah, that's right. I mean, it is quite complicated. I'm not saying it's definitely this first mover moment, but reasonably likely. And so, what we can say is like, okay, we as a society are not yet up to the task of figuring out how all the space should be governed and how that should be allocated among nations and people or whether it should be allocated at all. And so, we're just going to say like, no, we're not making this decision now, we're going to make it a later date. That is, in a sense, locking in a decision that's making a big decision to not do something, but I would describe it as lock-out because it's trying to keep as open. It's in fact keeping things more open rather than closing them off. Well, at least that's the intention. So, it's like, historically, the people who were most bought into the idea of superintelligence really being a thing that might come soon could be a massive deal. They've mostly pictured that at the moment when that happens around that time, there's going to be a single superintelligence itself or a single company or a single person, a single country that gains a really decisive strategic advantage, potentially just sends up making all of these decisions for everyone forever, for better or worse. And I guess it's hard to imagine that if you have one group that has a decisive strategic advantage and basically has a monopoly on power indefinitely, that they're likely to choose to maintain a very pluralistic, liberal, deliberative decision-making process. I guess because the track record of that happening is fairly bad. That process would exist purely at their pleasure because they could shut it down at any point in time. So, it feels kind of a tenuous or fragile situation. But more recently, over the last two years, we've been turning towards a situation where it seems like there's multiple companies with virtually a parity in terms of the capabilities of the AI. No one is pulling ahead at all kind of the opposite. That there's been a flourishing of interest in this question. What if as we go through superintelligence, in fact, there's multiple different superintelligence that are different, but virtually equally matched. No one gains any decisive strategic advantage. And in fact, the word remains shockingly competitive or different actors all have a significant stake in things for a long time to come. Do you think that people have been wrong in the past to, or have they underestimated the likelihood that we would have this kind of polytheistic, highly competitive scenario around the time of superintelligence? I do think there's a shift, which is that if you look back 10 years or longer, more people at least had the thought that the leap from subhuman to superintelligence would occur in this very short period of time. So Tim Urban has this, or sorry, Nick Boss of them, I think Tim Urban repeats it, but Nick Boss of them has this idea just, you know, sailing past human villal station and similarly in the discussion about fume, there was this idea that, well, maybe you just go from waste up human CDI to superintelligence over the course of weeks, days, even, you know, words like ours minutes get got thrown around, but the idea of like, okay, maybe this happens over the course of days or weeks is, uh, was quite common. And if, and also happening in a world where people weren't really expecting it, and if so, then the content's concentration of power seems quite natural to follow from that. Whereas now it looks, it's still quite unclear kind of how quickly will be the transition from AI that can meaningfully excel, or they AI are indeed to godlike superintelligence, but it seems much more likely firstly that people will be seeing this coming because AI is, you know, many people are seeing it coming now. Exactly. And that really matters because people can take action to ensure that another party doesn't have, you know, way more power than them. You see this at a small scale with say, Nvidia limiting the amount of chips it will sell to any one company in order to have a, you know, competitive ecosystem, but on a larger scale, you know, you can, you know, could imagine states getting involved because they don't want to see another country have, uh, far more power than them. Um, and then the second is just the speed at which you go from any given level of capability to superintelligence where it's already kind of clear that that kind of idea of just zooming past human-filled station is was quite incorrect because we've now for quite a while had AI is in human level, many, yeah, human level in many ways. Um, and then the latest, you know, analysis from Tom Davidson, my colleagues and others looking at this period of AI automating AI are then D still put significant weight on this massive leap forward, you know, 10%, 20%, but their best guess estimate is maybe more like you get five years of progress happening in one, which is still a very big leap and it's a leap at the scary point in time, but is much less of a leap than the, you know, move them subhuman to super like godlike superhuman godlike superintelligence over the course of weeks. I guess it's not clear that even if an aferious actor had that and nobody else did that would necessarily allow them to overpower everyone else. Yes. Yeah. For example, do you think, as the increasing probability of a more competitive superintelligence, um, arrival, is that a good development in your mind, or like a neutral one, or just very unclear? Uh, I mean, it's tied in with like a late of AI development and the heavy alliance on enormous amounts of computing power. Um, which are good things, um, for my point of view, um, the fact that it's not this kind of extreme extremely rapid take off things are like not so anarchic, or at least you have like only a few different actors, so it's like a good one. Well, it means on the loss of control side of things, you've got more, you know, things still go very quickly, but there was extreme takeoff scenarios, you've got more opportunity for learning by trial and error, um, to actually, you know, let's say you've got AGI plus, you can learn from AGI and from AGI plus, you can learn about how to align AGI plus plus and so on. And there's a little more time, at least, for just human institutions to the act, so governments could kind of perhaps, at least, um, realize what's happening, put in better regulation, for example. Um, so those things seem good, and then, uh, yeah, the fact that you don't as inexably end up with also intense concentration of power seems very good to me too. Okay, so let's push on and talk about, I think, the most original and interesting of the different kind of trade and coordination proposals you had, um, or that fourth thought has put out. I think this is mostly Tom Davidson's original is, uh, yes, Tom had the original idea, um, and a paper on it will kind of shortly co-authored with Tom Mier and myself. Yeah, so the idea here is that we could maybe go from having like many different agents who each have like some resources, who each care like very, a very tiny amount about doing the right thing about, you know, creating good impartial alien understood impartially. But nonetheless, they could all end up agreeing voluntarily to spend almost all of their resources producing that thing that they only care like very little about relative to their selfish interest. How would we accomplish that, um, that that alchemy. - Yeah, so consider the scenario that now, just look at the people who value things linearly. But suppose, and suppose there's lots of such people, but they value two things. They all value simulations of themselves. You could replace that with other things, statues to themselves, whatever. But I, you know, each person, they value copies of themselves, but don't value copies of other people. But then they all care about some kind of maybe, you know, ethically valuable good, call it consensium or something, just a little bit. So if they're just making decision themselves, they'll just do all kind of copies for themselves, 'cause they only care a little bit about this other thing. However, suppose there's a very large number of them of such people, they could all come together and say, look, we could agree that none of us will spend money on ourselves. And instead, we'll all fund this good that we all like just a little bit. And let's say there's a million such people. Then, well, if I'm one of the people, then I say, okay, well, I'm reducing my own consumption by, you know, one dollar. But I'm increasing the amount of expense on this consensium, this consensus good, by a million dollars. That's amazing. So actually, I would agree to some policy that we all kind of pool our money and donate and fund this kind of consensus good. So, you know, in a less futuristic setting, this could be maybe individual people want to spend money themselves and prefer doing so to spending to benefit the poor. But if everyone, if there's a law that says, okay, we'll tax you a little bit more and more money will go to the poor. Then they think, okay, yeah, that's actually pretty good because I lose out a thousand dollars or something, but a thousand dollars times everyone in society would go to fund the poor. Okay, so the basic idea here is that if each of these people were just spending their own resources individually, deciding how to spend it, they would spend it all on some selfish thing that only they care about. No one else really cares about. But they would, despite that, voluntarily vote for a political party that would impose extremely high taxes on everyone and then spend it on some other thing that they only value a tiny amount. But the amount that you'd be able to produce of it is extraordinary because you'll be able to pool everyone's resources and basically spend most of society's resources making it. I guess this phenomenon exists today. What are some examples that people can picture? Yeah, so, I mean, we can call the concept to get a moral public good where public goods in general are something that won't get funded enough by decisions of individuals. So I benefit from streetlights. But the issue is I can free ride. If other people are funding streetlights, then I still get the benefit. Or if I fund them, then there's always benefit that I'm not doing. Nonetheless, I will vote for to have a government or a city council that tax me in order to have to put streetlights on the roads because the benefit I get from streetlights is larger than your small fraction of the total cost. The tiny cost to me personally to pay for it. The case of a moral public good is where it's not that I'm personally benefiting from the thing that's being funded. But I care about it for moral reasons. And so the most obvious case would be poverty relief or even welfare payments where many people don't like poverty. They want people to be better off. And but they don't care very strongly about it. They care a little bit about it. And they would be willing to contribute poverty relief or welfare payments. But only if everyone else in society is also doing so. So the core issue that you always have here is the free rider problem that if you try to just get people to kind of all come together and sign some agreements, some contract to do this, at the last minute, it's tempting for anyone individual to drop out and hope that everyone else signs it and goes ahead and spends their money on it. But they can both get to appreciate the work that all of these other people have done but keep their money for themselves. So you kind of need to have, like in the current world, this only really works if you have some Leviathan a sort of government that can basically compel people to contribute, even if they kind of claim at the last minute that they actually don't want to, or that they'll drive their not contribute, or that they will lie and say that they don't value the moral public good even though they really do. Do you think that that will have to remain? Like would this only work in this long term future if we similarly have some government or some, I guess like powerful entity that can compel contributions to the moral public good? Yeah, it's unclear to me. So you might think, oh well, this is just a coordination problem, AI advanced AI, super intelligence are going to solve all these coordination problems because hey, there's this thing that just it's better for everyone. From the analysis we've done that Mia Taylor kind of really led, it's really quite unclear actually that AI is able to help you with this problem because you've still got the fundamental problem. Like, OK, everyone's coordinated. So we're all going to do this moral public good. And I'm like, oh, I back out now. And now I can spend my other sources and myself. That's better for my perspective. And there's in fact something that's even worse that could happen, which is, oh, well, if I know there's going to be this deliberation and attempted coordination, I can self-modify. So instead, I'll just not care about the good-- You'll excise that. Part of your preference. Exactly. So if I care not at all about this consensus good, then I have no reason to join in this coordination mechanism. And in fact, it would be-- they would have to use non-voluntary means to get me to do it. And so if that's true, then well, that will also apply to everyone else as well. And you could have this perverse outcome that everyone has self-modified away from caring about this consensus good. And so it certainly seems to provide a reason for having a Leviathan, for having something that can create certain binding laws or rules, perhaps that everyone votes on. OK. So one path to prison of moral public goods is that you have a Leviathan or, as yet, magical coordination mechanisms for having people agree and not opt out. The stuff that we have managed to come up with. But there is another galaxy brain way that we could potentially try to get there. Or that we just might naturally get there. Do you want to have a go at explaining this? Yeah. Yeah. So this depends on what decision-fear the people in the future have. So many things do. So many things. It's big. It's big. So we've been talking about coordination. It's just causal coordination, which is what we're familiar with, you know, cases where it's like we form a contact and I get punished if I don't abide by the contact. However, suppose that people in the future have some non-cousal decision theory, like evidential decision theory or functional decision theory or some further variant. And now let's say I'm making a decision about how to spend these horses. And let's also suppose that it turns out, as I think is quite likely, you know, as a current best guess, that we live in a very large universe in the sense that far away in the universe or perhaps even branches of the multiverse, there are beings who are highly correlated with me, such that if I make some decision about how to spend my funds, it's very likely that they do so too. The clearest case would be if in some distant galaxy, far beyond the observable universe, it just so happened that there's an earth that produced human life that's just genetically identical to human and there's a carbon copy of me in that world. Then it seems very plausible that I should think, well, if I decide to fund a certain good or a different good, then this carbon copy of me will also do the same. But then it also seems plausible that that would be true if it's not perfect carbon to copy, but just someone kind of similar. And on the kind of evidential or non-cousal decision theory, that is a really big deal, in fact, because I care not merely about the kind of causal effect of my actions, but I also care about the fact that I get the update that this person who's correlated with me, far away in space and time, will also act in that way. And so, in fact, the kind of choice in front of me is not, do I fund, let's say, the copy of myself, the self-interested good, or do I fund the consensus good? It's, do I fund the self-interested good? And all of these, like copies or nearby copies of me, fund goods that benefit them? Or perhaps I can think about what's this good that I like? And all of they, they all like to. And so if I fund that, I also get the evidence that they fund that too. And so we don't need to go via this kind of causal cooperation and so on. And also, I'm plausibly, if we really do live in a very large universe, then it's very large number of beings that I'm correlated with. So the decision would be, I fund this thing just for myself, or I fund the consensus good and billions, trillions, trillions of trillions, people fund the consensus good too. And so that might give this extraordinarily strong argument for me to fund the consensus good. And that would work even with no Leviathan, even if I'm the only person in the universe, sorry, in my little part of the universe. Okay, so if you're hearing this idea for the first time, then this might come across as a little bit peculiar. I think the preparatory episode, if you wanted to go back to it, that would best explain what we're talking about here, is my interview with Joe Karsmith, which is episode 152 on navigating serious philosophical confusion. What would you say to people who are kind of not bought into the premise that there's like an enormous number of other beings out there who are having like extremely similar thoughts, who's like decision-making procedure about this kind of choice is highly correlated with such that if I make a particular choice, I gain evidence that like lots and lots of other beings or other civilizations opted to do the same thing. I mean, if that's where you get off, I do think they'll have a pretty good argument. So on leading cosmological views, views, like, you know, on what is the standard assumption about the, you know, nature of the universe, there is an infinite amount of stuff. So we've got the observable universe, the accessible universe, like what we can ever interact with. That is finite, it's very big but finite, but the standard assumption is entails that in fact it goes on forever. And that would mean, well, there's an infinite number of beings that are very close to me, given as long as it's not sufficient. Yeah, exactly. Even if it's finite, the best guess is about how big the universe are, like, they're really very large. So that's one way in which you could have lots of, you know, people that are very closely correlated with. So, yeah, so there's lots of agents. Do you think it is likely that regardless of which civilization is out there where they are, like, their evolutionary background, that they would end up having this kind of conversation like strike on this same idea and basically have to be like, oh man, should I buy, should I fund the moral public good for like evidential decision theory? They have their own word for evidential decision theory. Do you think that's probable? I mean, I hadn't thought about it, but yeah, my guess is that, I mean, there's two things. One, it wouldn't even need to be probable if you've got enough copies. Good point. But I think it probably would be probable, like, it's quite a natural, you know, it's this ape-fire that I think it's like in the structure of preferences and how preferences work. So it would seem to me like, reasonably likely. Yeah. So it'd be surprising if they became space-firing but didn't manage to have these ideas, given that they've jumped out at us like at this relatively early stage of the element. I think, yeah, it is worth noting, this is a massive hammer to bring to this problem of trying to motivate people because if you believe that there are enormous numbers, like maybe like infinite numbers of beings out there somewhere like in space and time across the Moldiverse, or elsewhere in this universe, whose decisions are sharply correlated with their own because they're basically making the same philosophical decision about what decision theory to use. I guess they also have to make a decision about what this consensus moral good is. Maybe that's a little bit more tenuous that everyone would kind of converge on caring about similar stuff. Well, there, people, different beings could care about all sorts of different stuff. So, you know, let's say there's this, you know, so we have a million beings that I'm like closely correlated with. Then I'm just kind of looking through all of the things that they care about in order to find what's the thing that is most consensus, where it's kind of the balance of how closely correlated I am with them, how many people value that thing and how strongly do they value it, is such that things work out that it's what I should fund. I mean, it's interesting to think about what that would be. What I have about all of this is that we would end up funding things that I think at least at only intimately valuable. So, yeah, let's say that's just happiness, you know, positive conscious experiences are what in fact are good. There's certain things that are intimately useful for actually producing any sort of society at all, like knowledge, larger population, growth, like growth, survival. I should expect basically all civilizations to value those things, maybe just intimately. But sometimes they might get confused by our lets between things that are useful as a means to an end and things that are terminally useful. Exactly. It's a very natural thing if you're just something that's very intimately valuable, people end up caring for it for caring as it's own sake. In fact, lots of philosophers care about knowledge and survival and think achievements and think such things are intrinsically valuable. So, if so, then that might be what is the consensus across all of these very different civilizations and then at least given my best guess about what actually is good. Well, it's actually important at the moment that's like a terrible shame we all end up. It's pretty neutral. We all end up funding something that is not of terminal value. I guess you could at least say it's not terribly bad either. Is that going for it? Yeah. So, when I read this proposal, I was like, holy shit, this could be like incredibly force, this argument could be incredibly potent. It could actually drive almost any agent that is able to understand this. I mean, maybe it would just be superseded by future philosophical insights we would have. It's a bit surprising to think that this is the end of the road here, but it could be a very powerful hammer to really motivate an enormous amount of resources to be spent on something that otherwise just with absentee we would never have spent it on. Do you think that's positively right? Yeah. So, this is why Tom expresses this idea to me and I'm like, oh my god. Because it is this idea, potentially, this like polyannish naive optimistic view of just everyone gets to, if there's only enough time for people to reflect and think and with advance enough, everyone will just converge on the good and produce the good. This is like, there's totally, this mechanism for doing so that I hadn't thought about before. Like I say, I think there's an awful lot of asterisks. It's great, but I almost want to stop thinking because I really don't want the sign to flip based on further considerations. So, come on. Because it's like, whenever you're close to something really good, I feel like you're also just one bit of information away or some other consideration that could make it terrible. Yeah. I mean, I wouldn't want even if I couldn't see any flaws with the argument. And I think there are controversial, seriously controversial aspects of it. I still wouldn't want to place too much weight on it because any argument that's saying, oh well, people in the future will have such and such decision theory and such and such beliefs about the cosmos and then we'll engage in such and such argument that me and my friends thought of the problem. Yeah. A couple of months ago, I'm like, no, I want to take actions that are like much more, I want to act in the basis of considerations that much more of a bust than that. So it definitely makes me more optimistic about the future. Way to go. Yeah. I don't want to have this kind of, yeah. Paulianne's view about the future on the basis of such controversial premises. And I wouldn't want to do that even if I couldn't like see the problems in the argument. And in fact, I think, there are controversial aspects. Okay. Yeah. We'll push on for this. There's an article coming out about this soon for people who would like to read more. I did, you know, I guess it'll be on forethought.org. Yeah. The on forethought.org may in fact have come out by the time this podcast episode comes out. Okay. Let's push on to the miscellaneous section of the interview. We're going to talk about, I guess, a grab bag of other other topics. I asked the audience for what questions had most like me to put to you. And the most upverted one was a question about, um, pause AI based, or like, we're trying to make AI go better. It seems like there's some chance that things could go catastrophically off the rails and the track that we're on. We are like barreling forward pretty much towards artificial super intelligence, seemingly almost as quickly as we like technically can throwing trillions of dollars at it. Isn't the common sense thing given that we might all die or things could go horribly wrong, though we should slow down, maybe even like stop temporarily, catch up breath, do a bunch of stuff to try to make set ourselves on a safer course before we resume. That's a very common sense, natural view, but you aren't pushing for that. And I'm not exclusively pushing for that, though, I'm sympathetic to some versions of it. Yeah. Why not make this your main project? Thanks. Yeah. It's a good question. And yeah, let's distinguish between a few different sorts of pause. So first, let's talk about pause at human level. That's a phrase from lying clean black. So that's like when we're at the point of time of AI engaging in AI R&D and this point of time when things perhaps go even faster. Should we at that point be trying to slow things down, even pause, stop and start and so on? And then I'm like, yes, definitely. Like this is like really quite, this is both the danger of this period and the fastest period, or at least it's potentially both of those things at once. And why is that the crucial period? Well, actually as well as it being disorientingly fast and the like period when like early AI takeover could happen, it's also got these benefits of, well, we can benefit from like AI assistance up to that point. There is, we can also benefit from the fact that like, AI has had more of an impact in the world. So there's greater chance of kind of inoculation happening, like other actors having woken up to how big a deal it is. So I think they'd a chance of like regulation and so on, happening if only there were time in that period. It's also just when you have the AI systems that are like, just the generation before the systems that are most dangerous. So you can get the most information by kind of studying them and doing kind of alignment with search on them. So kind of pausing and slowing down at that point, quick key non. I have this one post on the idea of like having a kind of red line for the intelligence explosion, where you just, you have some sort of operationalization that you're quite keen on. Maybe you also have this like panel that's like Jeff Hinton and Yoshua Benjou and other kind of luminaries. Perhaps with some skeptics in there too. And that turns this gradual process into a kind of binary. And I think that I've been kind of keen on is there being this like international convention essentially, which is like, okay, the intelligence explosion has begun. And we're all going to come together and like figure out like what's going to happen over the course of the coming year or years. So I mean, yeah, I mean, favor the slowing down the intelligence explosion. What does that mean for pausing now, which I think is really quite different? Okay, again, distinguish coupled different sorts of pauses. One is like pauses on capabilities and another is pausing in terms of like compute. The pauses I've seen advocated, the pauses on capabilities. It's like no new training ones. And honestly, I think that's kind of, yeah, it would have actively harmful effects, even on the things that we care about, even just from a safety perspective. Because it's like, you know, at the moment, there's a small number of actors at the frontier. And my personal view is that they're like, actually surprisingly sensible. You know, my prior, there's low. My like expectation is low for how company is behave. And like you can look at the kind of history of like how X on dealt with the problem of climate change and so on if you are where they just buried it and fed misinformation instead. But there's both like a small number of actors who are like alive to and investing at least some in the problem of AI safety. Pause at capabilities. It's like, okay, well now all of the laggards start coming up to the frontier too. So that's China, you know, like meta XAI, like all, so we've now got many more actors, including the ones who are, I think, less scrupulous. And also if it's about not training, well, you can still stockpile compute, you can still build more fabs and so on. And that starts putting us in this really quite for the carrier situation where, okay, if one person breaks the pause, then suddenly things can go much faster than they were before. And in particular, the speed and size of intelligence explosion you get is about like how much compute do you have at the time. And so that actually means that other things being equal, I want more algorithmic progress faster because I want us to get to, because it slows things down later, because you've got a lot of hanging fruit on the algorithms. Well, it means that you've got AI automating AIR and D with a smaller compute total compute stockpile. And that means like do all of the modeling and so on. You get a slower and lower plateau intelligence explosion. And that's again, that's the scary bit. That's where all the risk is. And that's where things are going too fast. There is this different proposal you could have, which is like, okay, don't do it by training, but just like slow the amount of compute that we have. That I think has like, yeah, more promise though, there are still kind of other similar worries, which like, okay, well, don't produce as many chips, but there are lots of fabs and power stations and so on and everything kind of lady to go. And again, you'd also get the kind of catch up concern. But then the final point is just, okay, there's various things we could be advocating for. The my point of view, there's just loads of like incredibly low hanging fruit for making the situation quite a lot safer. So we've talked about AI character, we've talked about like risk aversion and deals with AI's. Like we haven't talked about things like mechanistic interpretability or safety of the search. Or just like a really quite basic government regulation. So like the US government could say, if you're a frontier company developing AI, you have to have an AI constitution that says what the AI is meant to do. And you have to have, you have to give us like, very high quality evidence that the model is in fact obeying that constitution and does not have some ulterior goal that could have been put in by internal sabotage or a foreign actor like China or has developed organically. That would be like an easily big win in terms of the juicing of risk. And all of these things are like, do not impose like massive costs on the world. And I think are just like much, much more likely to happen than the idea of like some international pause. So the like bang for buck of like what to advocate for. I mean like I say, I actually think the pause stuff I've seen seems counterproductive to me. But even if I was like, okay, in the ideal world, this would happen or something. I'm like, man, there's just so much other stuff that's just like super low hanging through super high bang for buck that we could be pushing for. Yeah. There's obviously like a really complex thicket of considerations here about like, you know, exact timing, exact message, exactly like how voluntary and so on. I think it is worth having some people trying to put in place the infrastructure to pull the cord at a future time. Like, yeah, it is a bit frustrating that I think that there's no conversation between the US and China along the lines of, if like neither of us is like sure, how dangerous this is. Yeah. It could be really safe. It could be really dangerous if we get just damning information. If we get like some damning revelation about the nature of these AI systems and how dangerous they are, we want to like be able to quickly like coordinate to not trip the wire that we have just realized is there. Yeah. But there's like nothing like that. And I think that there is a bunch of preparatory work that could be done for like pausing at the appropriate time if we get the right evidence. Yeah, yeah. I totally agree on that. And like, yeah, having compute backing. So we just know how much compute there is. Having a plan where it's like, okay, if the US and China just like, yeah, this is too much, they agree. They all, they think they think their chips to Switzerland and mutually destroy them. Or at least certain number of them. But like I was thinking that the more modest thing is just saying, well, if we conclude that like the next, we both just agree, evidence to come out, the next training run could be mega dangerous. We really don't want the other one to go ahead and do it. Yeah. So we need to have some monitoring arrangement that we can very quickly put in place so that we can both feel good that neither side is going to rush ahead. Okay. Isn't that like an even easier ask really? Oh, yeah. I guess I was maybe thinking that might be harder. So stuff involving like compute governance is just much easier to like monitor than verify than, are you doing like a training run on existing compute? And we don't even know how much compute you have and so on. Because it would involve like, maybe some on-chip mechanism for whether the chip is being used for training or inference. Okay. Yeah, we could talk about pause questions. And the do says that for some time. But I think we should set that for our side for another episode maybe. You helped found effective atroism many, many, many years ago. I guess it's been kind of the motivating philosophy for 80,000 hours since we started in 2011 more or less. I guess it's been a tough years for EA. Main reason being that Sandbag would feed who's like mega associated with effective atroism when I'm called. Yeah. Commit it some massive crimes. I think like at least partially in pursuit of atroistic goals. It probably like mixed motivations. But I think wanting to make money in order to do good was one of the factors. I guess a lot of people have been inclined to lose interest I suppose in any A or to be either dissolution with the door thing that it's like a bit of a, it's a bit hopeless because like the brand has been so damaged by that event. Because how do you think EA has been tracking over the last couple of years? Is it like stagnating or like recovering a bit or like in decline? Yeah. I think so it's distinguished between kind of yeah, the like online vibes and online discussion kind of bland and then what has in fact been happening. And like it was obviously this huge hit and it was like at the time just maybe this is the death, death blow. I think the overall story is like obviously things are much quieter, like relatively quieter, like less flashy kind of online and so on. And obviously like fewer people are like EA identity, this is my kind of brand. In a way that I kind of think is good and healthy. Like I think maybe like, it would have been good anyway. It would have been good anyway. Like personally. But then in terms of just like how are the ideas like in practice kind of how is that kind of the impact? kind of going over time. I think the overall story is like, okay, there was this big hit for a few years and then now it's just kind of back to really quite strong growth. So for a few different kind of metrics on this, one is like just broader effect of giving kind of movement, just that money to more effect of charities. How has that been growing over time? And please steady actually even through this, you know, period of like crisis and drama and so on of like growing it about 10% per year. Over the last year, actually it's like accelerating. So the numbers aren't yet in, but it looks like the kind of growth in total money moved to effective charities is grown by like 40% or 50%. So flam about like 1.2, 1.3 billion to probably more like 1.8. And so obviously a big part of that is coefficient giving and a big part is give well. There's also founders pledge, but you've got the same dynamic across many different kind of national effective giving organizations and then also kind of new foundations being set up on kind of effective giving principles as well. So that's really seemed quite striking. And then I think the same dynamic applies for other areas too like giving what we can pledges as well. Absolutely the kind of growth and that was took a big hit where you have 1,600 new pledges in 2022 and then only 600 in 2023. But again, now it's just back to quite promising rates of growth kind of 20% 30% year on year growth. Given what we can now kind of got more money moved than any year, like annually than any year in the past. And then similarly with kind of effective altruism itself as a kind of community and movement on a center for effective altruisms main metrics. Again, it looks like kind of 20% year on year growth. So it's kind of like this thing of just is like this like in the age of huge. So the huge boom and a huge bust and then it's like come maybe back to where you had projected many, many years ago. Maybe, maybe like if you, I think if you go on to like 2015 and they're just saying like, oh, this is what? 2025 was like, be like, oh, okay, cool. Well, it's just like this. So it goes, it's just like crazy. Crazy period in the middle. So I think in a couple of months time, you've got the 10th anniversary edition of doing good better coming out, right? And I guess you're going to do a bunch of interviews based on it. Yeah, so making me feel very old. [LAUGHS] And yeah, so 10, you know, it's been now 10 years since doing good better was published. And obviously just a lot has changed in the world. And so it was being used as materials in lots of student courses. And so I was getting some professors kind of asking me like, please can you update this? Because it's hard when like statistics are about to date. So there's this wholly updated version. The content is all basically the same. It's mainly just facts and figures are updated. And then there's a new kind of playfist that is discussing a little bit of like how it, you know, my thinking on effective altruism has evolved over time. And yeah, I'm using this as an opportunity to, you know, go on a few more podcasts and so on and talk about effective altruism and the core ideas a little bit more. Yeah, how are you expecting it to be received? I guess I expect to be like hit with lots of questions about S.B.F. I mean, I think-- I mean, I think like it's a device edition. It's not going to be this like big mega kind of splash. And yeah, I expect there to be a mix like a lot of people are, you know, that's the story they want to talk about. A lot of people are just genuinely interested in the ideas and the kind of philosophy behind effective giving or effective career choice. I guess the, I feel like the thing that I did-- I feel like it's appropriate that EA took a reputational hit that it really did like reveal something problematic. Or it made me think that something that I knew was problematic about it was like actually a much more serious issue than what I had thought. Like, there'd always been the worry that it would be maybe easy to appropriate EA ideas to justify raw breaking and very misbehavioral, possibly even crimes. But I thought that it was relatively like the rate of that would be quite low. I guess the fact that we had like such a spectacular instance of that relatively quickly made me think, well, actually maybe the appetite among human beings to grab a philosophy that can justify doing bad things and pursue a power might be greater than the way I had thought. And I hope that we've installed enough, say, more safeguards. Or maybe the reaction to that event is sufficiently strong that we're unlikely to get the same sort of thing recurring again. Do you have any thoughts on that? Yeah, I mean, there's definitely open, like, very open questions to me in terms of what was in the minds of various people at FTX. I mean, yeah, my really spent much longer than this topic than perhaps would have enjoyed. But even though I really had the worry that it was like some careful, consequentialist plot, that I think just really isn't born out by a careful kind of study of it. Doesn't make nearly enough sense among other reasons. But then the thing that's definitely true was like, OK, EA has evolved a lot in that, I think, it being less of an intense identity is a big part of that. I think people are extremely on guard for a certain sort of fears about little blinking and certain sort of naive maximizing in a way that I think is helping-- Maybe it would have been good to have that early, but-- I think EA always had this in a way that actually was emphasized a lot. And I'm glad it's being like doubled down on. OK, so in terms of the future, you wrote this a couple of months ago. That was super-war received called EA in the age of AGI. I guess, discussing what you think is the comparative advantage of the EA mindset, I guess, in the coming years. Yeah, what was the case you were making? Yeah, the key thing is just there's a certain sort of vibe, which is, well, two things have happened. One is the-- we've entered what I'm calling the age of AGI from GPT-4 onwards, where we now have AI systems that are reasoning in, like, impressive human-like ways. Or sometimes human-like, sometimes not, but that actually able to do tasks that are just clearly on the path to AI that can automate AI R&T. And that's a really big deal, and it's happening sooner than most people thought. And so there's this huge rise in attention on AI. And then at the same time of these major hits to EA's movement. And so you might have this view of just, OK, well, we should just let go of EA as a project. Like, think of that as a legacy project, because instead, what we should just be focusing on is AI safety. And the drum that I've been banging for like many years, but the last couple of years in particular is, like, look, AI poses many flets, many risks. There's many things we need to get right. It's not just about alignment, though that is very important. And when we look at these other challenges, well, what sort of person do I want working on them? I want people who are very kind and nerdy. I want people who are careful and thoughtful and have scout mindset and are very ethically concerned, and are not merely coming in with some partisan ideology, but are also willing to think about a really very weird and dizzying things. And that is exactly what is being provided by effective altruism as a set of ideas. And my main case of this was for all the stuff that is not just alignment. Some of the pushback I got on a draft of it was, no, actually, this is really important for alignment and safety, too. Because within alignment and safety, there's all sorts of things you could work on. You could be like, oh, the enforcement learning from human feedback, got other stuff that's just related to the models today. But taking really seriously the alignment problem is taking seriously the hard problem, which is how you're aligning superintelligence, which may in fact have perfect situational awareness of any tests that you're trying to do that can do what would be the equivalent of millions of years of reasoning-- I mean, in the extreme millions of years of reasoning and one forward pass, all that is continually learning over time, reflecting its whole values. These are the hard challenges. And that is like a weird world to think about. And it's something that doesn't really come naturally, whereas some of the alignment and safety of these searches I've talked to have said, it's actually people who are really thinking about this big picture perspective that are adding much more value than people who are treating kind of AI safety as their job, and they're not thinking about the big picture as much. It's interesting that it feels like the thing that's doing the work there. I guess it's just generic script sensitivity is one factor. And then there's also like a particular. appetite for weirdness, which is being willing to seriously toy with like very strange ideas. I guess like some of the things we talked about earlier today are in this category. Without like going off the deep end and becoming like absolutely besotted with your like pet theories, it's like, it's like, I guess a fragile middle ground, which I think is like relatively uncommon. And like, and for that reason is quite valuable because there's like neglected stuff that only people in that like in that window are going to be excited about. Yeah. I mean, yeah, there is this thought that look, it's just really hard to you know, be well calibrated and find and believe through things and even when they're appropriately weird, but not fall into kind of constrainism that maybe we'll get you a good following on social media and people think you're interesting. And if you're just like really honestly trying to do good, well, that's something that's constraining you because you will do more good if you have accurate beliefs. And you know, that is best at least can lead you to have, you know, be in the right middle ground where you believe or entertain weird ideas when it is appropriate to do so and reject them also when it's appropriate to do so. So people can go and read that blog post if they I guess want to get the full argument. But what were like some of the particular things that you thought were people with an EA style of thinking and EA flavor should particularly like disproportionately be going into. Yeah, I mean, I would say just the range of things that were focused on. I mean, there's one that's just very obvious in particular, which is just AI rights, like AI well-being. Some of the stuff we've said about kind of cooperating with AI's as well, that's just just a very unusual set of things to be thinking about. I don't think it will become unusual. In fact, I think it will become really quite mainstream concerns in five years time. But is exactly the sort of thing where I think it takes both, you know, a willingness to entertain weird ideas without controlling some at the same time as like actually a deep concern for not really messing up, ethically speaking. I would say stuff in AI character as well. I mean, here it's like we want lots of different voices and lots of different people kind of playing into this. But there is a big aspect of it. Of already the people who have in fact been in charge of kind of AI character, most of the companies have been like, you know, dealing in this kind of the active way because we're not even looking ahead like a couple of years. Like maybe the AI character is now just caught up to the capabilities AI have. But I mean, how much thought has really gone into like AI character in multi-agent dynamics over like long time periods? Like really kind of very little. And so, you know, for whatever reason, I think people with kind of EA mentality have just been good at going into like weird, poorly scoped areas and then kind of helping figure out like, okay, actually what's most important for us to focus on and whatnot. I imagine that is someone who wanted to push back on the EA and the age of AI argument. They might say, I guess EA is like taking a massive brand hit has like a bunch of negative historical associations because of SPF and FTX. And it also brings with a whole bunch of like other philosophical baggage that people like me or me not be that interested in. Like it's associated with the Shrimp Welfare project among other things, which I really like. But many people like might be interested in your AI related project. But like look at scans at the Shrimp Welfare project. So like why tie yourself to a bunch of other like weird work that you may or may not personally like at all by like branding yourself or branding the project as an effective outstriest style project. I guess in particular and as much as you have like more mainstream motivate or you have like mixer motivations like it's not exclusively motivated by particularly unusual EA moral philosophy. You also just like want to make the world better in a like general way. You want to like ensure that we don't all die in that the world is better for your own children. Why would you like make EA a big feature of it? If you could just say well, I want to make the world better also in a common sense way. And like that would be sufficient to justify what I'm doing anyway. Yeah. I mean, I think a big thing is I am like not making a picture an argument about like the brand at all. Like you know, the words EA. Like I have no particular kind of attachment to them or no particular attachment to whether what how people describe themselves. I mean, in fact, like it's always been the case that the best outcome is where that idea just feels like quaint EA withers away. I mean, I don't describe myself as a suffragett because I believe that women should have the vote. That is like, you know, an obsolete term. And so yeah, similarly, people can describe themselves however they want. The key thing is like what other what's the like mindset on which people are operating is that scout mindset is that being scope sensitive is that being appropriately responsive to how unusual a point of time within and like how high the kind of moral stakes are. You recently put forward a vision for the near-term future that you called viatopia. What is viatopia and what's the case for it? Yeah. So situation at the moment is that many of the biggest companies in the world are trying to build AI systems that surpass human ability across all cognitive domains. I think they're the good arguments for thinking that this is one of, if not the most mementos things to ever happen in human history. Much more like the evolution of homo sapiens or of life itself than even the industrial revolution or the invention of electricity or fire. It's at that level of magnitude. And yet essentially no one has a well-formed positive vision for what a good society after the development of superintelligence looks like. And that's his kind of striking and kind of worrying thing. It feels like a bit of an emission. Yeah. It feels like a bit of a mission. And the concept of viatopia is at least trying to offer a bit of a framework for what could an answer to that question of what a good post-superintelligent society look like. And so the concept of viatopia is that it's a state of society that is on track to produce a near-best future, something that's just at least 90% as good as a future that we could have. And it's distinctive in that it's not saying we should try an aim for some utopian society directly. It's also not saying merely, oh look at all these bad things that exist in the world. We could solve this particular problem and this particular problem. What it's saying is that we should try and figure out what does a good way station look like where that is some state society that can steer itself to something truly very good. And so as an analogy to illustrate, imagine if you're an adventurer and you're lost in the wilderness. There are a few different options you could take. You could try and take your best guess at what the right path is to get to your destination. Or you could try and just deal on an ad hoc basis where some issues you have at the moment, like maybe you're running low on supplies. Or you could try and get yourself into a position where you know what's most important to do next and where to go. So for example, get into higher ground so that you can survey the terrain and figure out actually where you're aiming towards. Viotopia is like that third path. And what would be the case for focusing on trying to get to viotopia now rather than trying to directly create a good world immediately? Yeah, so utopianism has a complete bad track record. Philosophers and writers have often tried to sketch visions of utopia and normally it's not long before they actually start looking quite dystopian. And the very reason for that is well, we just don't know what an ideal future looks like. There's a lot of moral progress we'd need to make before we could actually say yeah, with confidence this is what an ideal future would look like. So we need to do something else. Otherwise we'll probably bake in some major model errors of our own. Okay, where does an invite us? A via means road or something in Latin or through? Yeah, we mean by way of this place via utopia. So this via utopia notion that you've told me it's been very popular. It's been very well received. Do you worry that it's a slightly vacuous notion that you're saying, well, we want to get to a really good future. And so when you get to some intermediate stage, like into a position where we're likely to get to that future, is that a great insight or is that just kind of a trivially, trivially obvious thing? And it's not necessarily going to actually help us get there. Yeah, so good pushback. And I think it's not the most substantive thing. And it's deliberately, you know, it's a framework concept. It's for organizing our thinking. However, I think it's not totally to the view. So in the, you know, there is a history of debate on utopianism and other concepts. And the leading ideas were kind of utopianism, very popular idea, responsible for some enormous atrocities through history. And the pushback to that from Karl Popper onwards, but still very popular now. So Kevin Kelly, a futurist, has this idea of pro-topia, is the idea you just don't have a positive vision of the future at all. Instead, you're doing something more like hill climbing. So you're looking at society now, what are the little things you can change that are like clear problems, and then just trying to solve them one after the another. In this inquiry, the mental way. And so viatopia is a different way of thinking about things. And I think it does make substantially different or leads you towards substantially, deeply different recommendations than you might otherwise think, especially over the course of the transition from here to a superintelligence. So if you've got the utopian perspective, you might think, well, what we need to do is just make the AI a classical utilitarian or insert your other favorite moral view and then just hand over to the AI that's pursuing that vision of the good. Seems very bad for me, right? Dopium perspective. Or, and this will be very rough from the photopian perspective you might just think, wow, will there are these major issues, major problems in the world, like 100 million people dying of the year. And AI will give us the ability to completely solve those problems. So actually, we should get there as quickly as possible. And there will be, in fact, very rough trade-offs between how quickly we go and, you know, how much risk of existential catastrophe we bear over the course of this transition. Aiming for viatopia might say, well, actually, there's certain things that are even more important, not locking us into a really bad future, even if that means that we don't get to some of the, you know, upsides in terms of near-term benefits, quite as quickly as we might other ways have done. So you think protopia, this idea of, well, we don't want to have a grand vision, that's going to lead us astray. Instead, we just want to get wins immediately, like find ways to improve the world that we can, like, understand and that we can see whether they've worked. That would potentially lead us to, like, miss the bigger picture of risks, because we're just, like, grabbing immediate wins, like, trying to improve health or it would, like, recommend just charging forward on AI. Or at the very least, it wouldn't prioritize among them, where it would say, okay, well, maybe risk of loss of control to superintelligence or entrenchment of some authoritarian regime, you know, okay, well, that's some risk, but clear, clear apparent evils such as death and poverty and so on. And we could solve them kind of like away. And so, it wouldn't say, like, if you thought that the AI might kill everyone in the near-term, that's also a near-term problem. Or maybe it's hard to evaluate because it's like more about more probabilistic. Well, it's harder to evaluate and also protopianism at least wouldn't give you the resources for saying the trade-off. One of these is much more important than the other. Do you think of via topia as a middle ground between utopianism and protopianism? Or is it a different thing? In a sense, it's a middle ground in that it is often a positive vision for where we should be headed. However, it doesn't have the same, in my view, the same pitfalls that utopianism has because it's compatible with many possible ultimate visions for what a good society looks like and is not committing to this kind of narrow view of the good. So what would be the key traits that a via topia, would you say it's a via topian state? What would be the key properties that you'd be looking for do you think? So there's the key questions and key properties. And I want to emphasize the questions more than, oh, my particular answer at the moment, both because the questions themselves are more important and because, you know, my views evolve a lot over time. But that can include things like how widely distributed is power, where on one end of the extreme, it's just all powers concentrated in the hands of a single actor, all the way to, oh, it's extremely distributed, you know, global democracy or even perhaps more distributed than that. A second is just, well, what sorts of people, what sorts of beings have power? Is it just members of a particular society? Is it just humans? Do AIs have influence over the future? What about future generations? A third category is when the major decisions happen, where there are some arguments for thinking, look, we need to make really big decisions really quite early or instead we should say, look, actually for the sorts of decisions that will really guide how the future goes, we want to punt them into the future as much as possible. And then finally, there's questions around, well, how should societies a whole be making decisions and these most important decisions about how the future goes, where that could be via democracy, via voting, if so, what sorts of voting systems could be via auctions and market mechanisms, if so, what type? And so those are just some of the things we've got to grapple with, I think. And I have views on them, but they evolve. So the analogy that most jumps to mind to me is that you might have, if you have a group of people starting a new country, they might not be able, they might not yet know exactly what the nature of the law should be, what the political system should be, but they might find it have an easier time agreeing on some process, like constitutional convention sort of thing where they come together and they figure, well, we're going to, like, everyone will get some vote, we'll use this kind of deliberative process and then we use this kind of voting system. And then at the end, we'll end up with some set of agreements of how things are going to run and the chips or fall as they may? Is that a good analogy to have in mind? Yeah, I think that's a great analogy. And the US Constitutional Convention at the end of the 18th century is this remarkable event where, if I remember correctly, it's about 40 people in a room debating for three months, what should the United States of America look like? And what they agree is this set of procedures. And obviously, there's limitations and amendments after that. And it's interesting too because there's this balance between locking in certain ideas, but also kind of locking in a method that doesn't involve locking itself. So you can lock into a certain system that allows a lot of experimentation and free debate and change over time. That's very different than if they'd chosen a constitution that put, you know, a single person and or even a single family lineage in absolute power or something. That would have been kind of locking into a different sort of political system, but one with much less in the way of open-endedness and how it could develop over time. Okay, so are there any particularly non-obvious or controversial recommendations that you think the viatopian framing on things would push us towards to some of the people that otherwise not like? Yeah, so there are certain things that at least I think a viatopia would consist in that is not totally obvious. So one which we'll talk about is I'm very pro like distribution of power, whereas a lot of people who worry a lot about existential risk really are in favor of kind of actually quite intense concentration of power because the idea and it's not an insane view. In fact, the idea is if you've got this period of intense existential risk in particular, if existential risk can be posed by any of many different actors, whether that's because they develop a misaligned superintelligence or because they create extremely powerful bio-weapons, then you might think, well, we just need a very small number of actors, maybe in fact just one powerful actor that can guide us through this period. Whereas I think that's unlikely to put us into a position where we can guide ourselves to a near best future. So yeah, what's that? I think we'll talk about it a lot more, but ultimately it's because I think any single actor is probably has the wrong moral conception, even upon the flexion, even if they choose to reflect. I think it's a little worse than that in fact because the sorts of people who end up in Manchester, like one person has risen to the top and gained supreme power. There's probably some bad filters that they've gone past through. Yeah, exactly. And that's if you look at leaders of authoritarian countries in the past, well, that includes a mixed record. Yeah, I mean, that includes Stalin, Hitler, Mao, and the personality traits are just, you know, it's terrifying. These are psychopathic, sadistic people. They're not they're not merely randomly selected people who happen to have total power. And I also think that if one person or even a small number of people are in a position of total power, they're also just less likely to reflect on their values in positive ways. I think that's something that tends to happen more naturally out of interpersonal interactions and the needs. Well, especially one to the justifles I feel. Yeah, I think you noticed as even just with people who, you know, gain more influence within an organization or they become wealthy or respected or so. And they just stay, they stop getting the normal pushback that like sharpens their ideas. And you can imagine if you were the the supreme dictator forever, how disconnected you could become from any reality. Yeah, exactly. Okay, so what are the different categories of via topia that you think have a have a shot at working? Yeah, so I think there's three broad ways of thinking about how we could get to a near-best future. The first I call kind of easy utopia. So this is actually I think the common sense view, which is just. It's not that hard to get to an extremely good future, something that's basically as good as you can get. You just need to eliminate the most obvious and egregious bads. So yes, dictatorship would be like that, but eliminate poverty, eliminate, suffering, allow people to have little health, allow people to have freedom. And that plus just technological development will get us kind of most of the way, or even all the way there. If that's correct, then viatopia isn't that interesting, actually, because we'll probably just hit it anyway. A second view is convergence, where on this view, you would need to have most of society with power converging onto the right kind of ethical view. Or I'll sometimes use correct ethical view or correct model view. You can also just say this in more antirelyst, subjective terms, like the view I think, the place that with the view I would have upon idealize the flexion or something, but it's easier just to say correct or best. And they have to be motivated by it as well, right? And they have to be motivated, yeah. So in this idea, convergence is like, yes, maybe the best future is a narrow target, nonetheless, if we can get it such that most members of society or at least most of people with power converge onto the best thing, the best model view, and acts dear towards it, then nonetheless, we'll hit the narrow target. But that is necessary. And then the third vision would be what I call compromise, which is, well, you don't need everyone. In fact, maybe even if you just got a small fraction of people who have the right kind of ethical views and are motivated to pursue them and the, like, kind of broad philosophical perspective and understanding of the world as well, and they're able to kind of trade with the rest of society, that is sufficient to get us to a near best future. And my view, at least, is that this third option is kind of the most promising thing to steer towards. So we're going to skip over the easy utopia scenario here today. You have an article on the fourth episode called No Easy Utopia where you argue that that is not plausible. I guess in brief, because I think we both agree that the best possible world is not just a matter of removing bad things, but it's also about adding lots of the best possible thing as well. And probably the best thing is better than nearby things, so it's like quite a narrow target to hit. And I guess we're not going to talk a ton about this reflect, like, what if everyone just when they reflect on moral philosophy, they end up concluding that they reach the correct theory and they're motivated to spend all their resources operationalizing it? Do you want to say anything quickly about why you don't think that is super likely to work? Yeah, I mean, there's lots to say, but I guess I just think there's multiple ways it can fail, even if within a reasonably good scenario, where one is just that people can be un-interested in reflecting, or they can reflect in the wrong ways, or they can have a good reflective process, but just have the bad kind of starting intuitions, where from those intuitions, even with good reflection, they'll end up in the wrong place. And I think I'll say, like, I am somewhat sympathetic to the idea that maybe, yeah, maybe quite large swathes of people actually would converge in the same direction. I think if that's true, it's because of the nature of reality. It's because of, in my view, something kind of model-realistic being correct. Just the arguments just very strong towards one particular ethical view, or if you just experience this particular conscious state, you can't help but believe that it is good, because it is, in fact, good. That's the sort of scenario I think we'd have to envisage. But oh, wow, I don't think we should be confident in that. And in fact, I have, like, really quite wide uncertainty over how much convergence you would get from all the way from, yeah, it's actually just large swathes of people would converge. That's the kind of really good scenario. All the way to just, like, no one converges after the flexion, all eight billion people in the world would have, like, quite different views of the good. Yeah. We missed out you could get all of it right. Have everyone conclude the correct moral theory, but nonetheless not be interested in putting their resources, they'd just be like, what, but I just want to do my own thing. I don't care about doing the moral really good thing. Yeah. And like we see this, yeah. And in fact, that's, I think, the most likely failure where, you know, you can go to people and give them the arguments for vegetariness or donating. And they can say, yeah, all those arguments work and then just not take any action on it. And in fact, you know, it's not like we see today people investing lots of time and lots of money into ethical reflections and leading counter arguments and so on. It's just not really something that happens. It's going to be quite weird and unusual to do that. And in fact, maybe you want to, some people would want to guard against them. So imagine a fundamentalist, religious believers, or people who are very wedded to particular ideologies. And they might say, look, I don't want to risk reflecting on this. Yeah. I adherence to my faith or, oh, God, it would be like a barren of me to even consider this alternative position. And with future technology, we would be able to guard our informational environment or even self-modify such that we don't even consider these alternative perspectives. OK. And so just setting the scope of it even clearer, we're mostly not going to be considering cases of catastrophic misalignment and really deeply scheming artificial intelligence here, not because that's not a possible option or a very live possibility, but just because we only have five hours to record in and it raises a whole lot of separate issues. It's worth imagining what happens if we mostly overcome that one way or another. So yeah, let's live into the third option, which you thought was most promising, which I guess you call like compromise, trade. This is a scenario where, as I understand it, you have some meaningful minority of people who do convert-- or waited by, I guess, power or resources, who converge on wanting the right thing for its own stake, and they're willing to allocate some meaningful fraction of all of their effort towards that. And so let's say it's 10% of resource or power weighted folks want to pursue this goal. You want to try to spin this into more than 10% of the best possible future that they could be. How might they accomplish that? Yeah, so I think there's two big ways. So one is if different groups care about really quite different things. So the greatest example, perhaps, could be people who just maybe some groups upon the flexion, they just value resources basically linearly. So a total utilitarian would be like this, because the more the sources you have, the more happy lives you can create. And the value of the universe as a whole is in proportion with how many happy lives. Other views that are perhaps more common sensey might be very different to that. So might just care about preservation of the Earth's biosphere or might discount over time and space so care about what things happen near to them. Or might really just care about guarantees of good outcomes or high, very high probability of good outcomes rather than risky gambles of even better outcomes. And this gives lots of opportunity, which they'd-- so in this case, there could be a deal which says, OK, you've got the common sense person. They say, OK, well, we'll steward resources that are nearby in space and time. And this total utilitarian, yeah, sure, you can go to other star systems and then create this much more ambitious, expansive world with many, many happy beings. And then perhaps, in fact, both can get 99.99% of what they would ideally want if they had complete control over everything. And that's just very exciting potential opportunity. Because it means that then if we can get into the scenario that, OK, we've managed to get these beneficial gains to them, all these different kind of ethical factions trading with each other, then we don't need to pick a winner. It's robust to kind of disagreement. And it's therefore a much safer option than either just hoping we all converge or pushing some particular view of the good. Do you think that things would play out that way? Or is it a viable vision? So I mean, I think there are risks to even getting that. So one would be if there's intense concentration of power. A second would be maybe such trades aren't allowed. So there's lots of things that you're not allowed to trade at the moment. It's possible just the best stuff. So maybe the total utilitarian, like some particular blissful state. And those people are in the minority. And society says, no, that's illegal, where there's already lots of things that are, you know, in my view would be kind of just ethically fine. but are impamitted today. The bigger issue, I think, is, okay, so maybe there's lots of groups who have relatively easy to satisfy views of the good. Like, I know, like preservation of the Earth's biosphere or, you know, preferences for things that are kind of local. But I think there'll be a lot of people who actually just do care about things linearly. And there it's much harder to see initially why you would get these kind of huge gains from trade. So I said, okay, the total utilitarian says, well, I just want there to be as many happy, flourishing lives as possible. But now let's kind of distinguish within that. There's utilitarian one, utilitarian type two, and perhaps they differ. So on what they understand flourishing to consistent, what they think the kind of best conscious experiences or lives are. In order for there to be good deals and trade there, it would need to be the case that there's some kind of hybrid life that is more than 50% as good on both views. And, you know, it's speculation to say like how likely is it that there would be or not. My guess is that in general, there probably wouldn't be because my guess is that the very best things from a utilitarian perspective will be way better than things that are just a little bit less good. The archetypal case here might be, you've got, you know, faction A, faction B, say faction A, yeah, they're the utilitarians, they want like pleasure, no suffering. You've got faction B that wants to take something quite different. And like faction B, incidentally might cause a whole bunch of suffering in pursuit of their other goal. But the suffering is like not something that they value for its own sake. They're just doing it because it like makes their projects somewhat more efficient. And then group A could basically pay group B to not to redesign their thing so it doesn't involve suffering incidentally. Is that like, is that a kind of thing? That would be a case. And like in the world today, that sort of thing happens. So I do think that if we had much better opportunities to make such agreements, we had better coordination technology or something. The vegans and vegetarians and people concerned about animal suffering could just engage in some sort of trade with the people who like eating meat. And perhaps it wouldn't result, there wouldn't be enough bargaining power to kind of eliminate farming altogether. I think it could eliminate factually farming. And so, you know, most animal suffering could just be abolished. Because as you say, people aren't really aiming for that directly. It's just a side effect. My guess is that when we're now thinking about these like very, very kind of grand scales, that's not going to be like super common or at least there will be a lot of the residual incompatibility leftover. Because you're just going to produce, you know, happiness type one as much as you can. I'm going to produce happiness type two. I think that your understanding of happiness is like basically no value, but it's not like you're producing lots of suffering. It's just valueless. It's just, yeah, or it's like a tense that's valuable or something. And similarly vice versa. Okay, we'll push on from this. I guess we should just quickly note that there's a wrinkle with this kind of moral trade, a challenge that, for example, if we did start paying people to close down the factory farms or to redesign them, then you would be vulnerable to someone saying, well, I'm going to open up the worst possible factory farm unless you pay me. And you wouldn't know whether they would have done it otherwise. I guess they could pretend that they're not doing it to blackmail you basically, but in fact, they are. I guess possibly that, you know, in this star-faring feature could be maybe that wouldn't be such an issue or maybe it would be a much worse issue. We don't really know. Yeah. And I should flag this is my biggest worry with the whole widely distributed power and trade and so on is vulnerability to those sorts of extortion blackmail dynamics. And there's this very substantive project to work out, okay, what's a good system where, you know, people who say, yeah, self-modify or pretend to use blackmail and extortion are not rewarded for doing so, but you still get these other beneficial gains from trade. Okay, let's let's push on to some honest-to-guard philosophy, or at least what analytic philosophers would regard as philosophy. You've been working on a pet, moral philosophical theory that you call the saturation view. What problem in normative ethics are you trying to address with the saturation view? Yeah, so this is kind of a set of problems, in fact, within population ethics. It's a well-known area of ethics for generating all sorts of paradoxes, cases where you've got lots of individually extremely plausible principles that end up inconsistent with each other. And there are a number. So there's what's called the mere addition paradox, or so, where you've got some intuitively plausible principles end up leading you to what Derek Parfitt calls the repugnant conclusion, the idea that you could start off with a truly, an extremely happy people. And that outcome might be worse than a population that consists only of people with lives barely worth living, as long as there's a large enough number of them. So this is kind of one of the problems. A second is the problem of fanaticism that, you know, again, start off with this guarantee of this amazing outcome, and now take a tiny, tiny, tiny, tiny, tiny probability of something that's even better sufficiently good. When combined with expected utility theory, many views will say, take the gamble. No matter how small the probability there's some sufficiently good outcome that you should take it. Because it's risk neutral, basically. Because it's a neutral, with respect to, you know, total quantity of happiness or something like that. A third category of issues is infinite ethics. I think we definitely won't have time to get onto that side of things, but it's something that's really plagued this kind of impartial consequentialist approach to ethics or axiology. But there's also a fourth problem, in my view, which hasn't been discussed in the literature, which, yeah, I call the monoculture problem, which is, okay, let's try and figure out what's the best possible future. What does that look like? The markedly, all the extent kind of well-specified theories of population ethics to date say that the best future, if you've got a fixed amount of these sources, involves figuring out what's the very best life? What's the life that would produce the most well-being, you know, with a given amount of resources to create, and then just make copies of that life over and over and over and over. The universe. Yeah, so in, you know, EA and nationalist world sometimes gets called tiling the universe with hedonium, where hedonium's the whatever produces the most bliss per unit of these sources. But the general idea is just what it wants is a monoculture, because this is the thing that has the most well-being, and if you just have that repeated forever, you've also got this perfectly equal society, and so it's good on egalitarian grounds too. Yeah, well, it seems like it's a very natural attraction point, because like any theory that says that there's a best thing, and that thing is like not a universe scale, is going to say, well, if it's like smaller, just like make it and then like make it again and just keep going. Yeah. It's simply like you almost have to hard code in and like preference against this to avoid the monoculture, which like most people find like kind of quite unattractive. Yeah, and so yeah, it actually also follows them. A couple of principles that are generally regarded as, you know, axiomatic in population ethics. There's like a very simple kind of like proof you can make from it, from these kind of principles. However, I at least find that like unintuitive. I would think that a future that's just like that because of the one exactly qualitatively identical life is not the best possible future, and the better future would involve like a wide diversity of different kind of forms of life and experiences and so on. And I think that's not just an intuition that diversity of variety is intimately valuable, or an intuition that's saying like, well, we don't know what's valuable so we should head our bets. Instead, I think it's just no. Actually, that's placing that's a better intrinsic value on variety. A better future. Yeah, or something that has that implication. So it could be I mean, this might just sound like the same thing, but I think it's slightly different that the realization of a particular experience or form of life has value in itself, over and above just the mere well-being. But either way that yeah, a very diverse and varied futures is better than this monoculture. Yeah, it's surprising to me that this hasn't come up in the philosophy literature very much because I think I guess online, whenever people talk about what are we going to do with all of the all of the matter and the energy and then anyone suggests something that is like very monotonous, just repeat the same thing, people are like, well, I don't like that. Sounds horrible. Sounds crazy and terrible. But I guess philosophers like, because they're supposed to prospect of changing all of the galaxies out there, hasn't really been on the table before, it hasn't really come up as like, well, we need to figure out a solution to this. Yeah, I think that I think that's right. So, I mean, I have found over and over again, actually, that like, being really concerned by figuring out, like, how do we do as much good as we can has ended up just, you know, driving all sorts of, like, interesting philosophical areas and issues that otherwise being neglected because most philosophers are not thinking in that same way. Okay. So, yeah, what is the saturation view? How does it address this? So, yeah, the saturation view is a way of incorporating the idea that diversity is kind of intrinsically valuable by having the thought that if you have a replica of a life, so qualitative copy, that's just less valuable. And in fact, more and more and more copies of that life is progressively less and less valuable in a way that kind of tends to some upper limit. And generalizing that a bit for the same reason, like, maybe it's not an exact copy, but slightly different. That's also a bit less valuable than some, you know, totally new kind of form of life. And the analogy could be, like, you know, imagine a kind of color wheel that's initially, like, not lit up at all, and different sorts of life will experience, like, in a different spots on the wheel. And you can, by adding lives, you're kind of, like, lighting up those little spots, where there's a kind of traditional population axiology would be saying, just, you have the best thing and just over and over and then go over again, you want to, um, to use that best thing. Instead, on the saturation view, you want to kind of light up the whole wheel. Because, okay, I've had many copies, let's say, of, um, these very similar lives. Well, that means the additional lives are not as, um, adding as much value. So you get more value by kind of, um, instantiating some totally different kind of form of life or form of experience. So, suppose, I mean, it's a very natural formalization, I guess, of this intuition that you're just saying, well, you hit declining returns on stuff if they're too similar. Like, you got something that's good, but making it and making another copy of it is, like, isn't as good as the first time. And also, like, something that's, like, too similar to it also, it takes a bit of a hair cut if there was, like, something else that was too similar to it in the past. Yeah. I guess they never become useless. They just become, like, less and less valuable incrementally. Exactly. That's the, yeah, there's never the point when you get no additional value, but the amount of, the amount of value each kind of copy to juices gets smaller and smaller. Does an asymptote up to some maximum value? Yes. So, um, yeah, as part of the view, asymptotes. And that's an early, collusional part of, actually. Okay. And do you have any difficulty defining what the hyperspace is over which you're, like, considering whether things are different from one another, or are you just going to, like, set that aside? Yeah. So, I mean, in my work so far, um, I don't talk a lot about, okay, yeah, what exactly is this, um, is this kind of space of, like, different lives and, like, what does the, um, uh, you know, how many dimensions does it have and so on? Um, it makes some kind of formal assumptions about it. But my kind of view in general is like, well, let's just start off by kind of looking at the, the kind of formal structure of this view and, like, all of the nice properties it has. And then afterwards, we can then start arguing about, you know, because it would involve, like, shading lots of different, like, intuitions and so on. But I don't think is, like, really affecting the biggest, um, yeah, biggest pictures. So what are some nice properties? So going back to these, um, different problems. Um, so let's start with this monoculture. So very clearly just doesn't lead to a monoculture. Um, and in fact, you would want, like, this very rich, diverse, um, future, and that would be better. Uh, in the variant of the view, um, that I, uh, formally, it dissolves the mere addition paradox. Hmm. What's that? Um, it involves one extra structural assumption that I think, um, again, like emphasizing like, the point is to find some theory that is, you know, not like the total view and avoids its problems. But if all lives that have, um, uh, very low well-being, or all experiences depending on how you're aggregating it, um, are only a small part of the space of the overall landscape, um, of possible lives or experiences. Then once you appropriately reformulate the kind of underlying principles that generate the paradox, because these have to be kind of, philosophers would say, "Caterus Parabus principles." So other things being equal principles. So it's saying, um, holding, you know, holding diversity fixed, then it's not bad to make some people's lives better and add, um, lives that are good. And, uh, holding diversity fixed, um, it's not bad to, or it's in fact good to have more well-being and more equal. It turns out that the view can have the implication that you satisfy all of those principles, the rejecting of a pregnant conclusion, accepting this, uh, dominance principle and this kind of, uh, egalitarian plus increasing well-being principle. But you do not ever entail the, the pregnant conclusion. Because the thought is that all of these kind of low well-being lives or low well-being experiences, they just can, uh, add up to enough kind of diversity worth having. So you kind of, in each of the steps of the paradox, you're kind of adding people and then you have a really balanced, well-being, but then there's a step where it's just like, you can't do it. There's just no world that will in fact, um, satisfy kind of that step. Okay. I didn't follow that, but, uh, that's okay. It's, it's a little bit hard to convey on a podcast. Yeah. And in fact, like much of the paper is like, not even giving the view to begin with, because the views, it gets mathematically quite, um, intricate. In fact, it's just giving a toy version of view and then working it through. Yeah. So I think the main reason that I'm like not super drawn to this, I guess, is that I don't have the same intuition that, I don't have the intuition in favor of variety as strongly as, like, as as many people do. Um, so of all of the problems with total utilitarianism or any views like that, I, I, the thing that I find like most troubling is the risk neutrality between like positive and negative experience. Like find that like deeply disturbing. Because there's never something that I would choose for myself. It's like, that I would be like indifferent about a life that's extremely good and extremely bad, each with 50% probability. So that's like mega, that's, that's super kind of intuitive to me. Um, but the idea of like making something really good and then making a lot of it, like, um, I don't find like as, as peculiar. Yeah. And I expose, yeah. Well, I just wanted to ask on your views, you said the risk neutrality. I mean, you could just have a, like a negative weighted utilitarian view where, and let's say, bads count for a thousand times as much as goods or something, but you're still less neutral with respect to that. Yes. That is more attractive, I guess. Okay. So this is a little bit hard to know. Are you changing the weighting of the badness or just like how bad, or are you just correctly assessing that the badness is really working? Yeah, yeah, yeah. It's, um, but yeah, I think that is, um, makes more sense to me. Or that's, I guess, more how I would make the decision that is like a really, uh, yeah, you just wait the bad stuff really more. Of course, it's like debunking explanations for why humans would have this intuition that we're more capable of suffering a lot in an hour than we are of experiencing pleasure in an hour. But yeah, yeah. Okay. So I'm wondering if there's, um, you are also have like worries about the risk and mentality aspect because that's where, I mean, in the most extreme combining it with the suffering cases, you start off with a trillion, a trillion lives of intense bliss. Yeah. So a trillion, a trillion lives like absolutely amazing option A, option B is a trillion lives of intense suffering, worst possible suffering, plus some one and a billion, billion, billion, billion, billion chance of an extremely large number of lives, just barely worth living. The total utilitarian combined with expected utility theory has to say or expected value, um, has to say that the latter is better than the former as long as the number of lives are large enough. Um, so what we're doing is adding a whole lot of like, just, um, just barely worth living lives. And that's way better. Yeah. So well, they has trillion, billion bliss, utopia world. Yeah. And then gamble B. It's a gamble. It's a guarantee. I say, I say, yeah, I say, yeah, plus like an even larger number of, just absolute probability of all of these lives that are just barely worth living. Yeah. That's just a very large number of them. Yeah. I foresee the order is going to throw out like an edge case like this every, whatever. I say, you're too much practice with this. But I mean, that is also you. That is also very unattractive to me as well. Okay. So, um, so yeah, I think I did. I don't know. You were going to go somewhere with this. I think I would like to, uh, you know, you have your hopes with this. It's just because you mentioned this neutrality and that was one of the problems that I mentioned was this fanaticism where no matter how small the probability, um, uh, you really care about that. And as long as the, you know, sufficiently good, yeah, payoff is big enough. You will pursue that tiny probability of an enormously large payoff. Um, and this view is, um, avoids that because it's, um, ends up being bounded. So, um, as long as, yeah, basically as long as the landscape is either finite or a certain feature of it to case fast enough, then there's an upper limit to how much good you can create. Intuitively, like, again, thinking of this color wheel, you've fully illuminated as bright as possible the landscape. That's the kind of upper bound. And so you avoid fanaticism. And then I'll briefly say but not explain why. For the same reason, I think it has quite a range of desirable properties in even with infinite populations too. So many consequentialist views, like the total view, they naturally lead to a lot of parallel assets, so like you can't even compare intuitively comfortable worlds. This does not have that implication. OK. So I guess that is legitimate attractive. The two things that struck me as odd about the view, or like less attractive about the view, was on the negative side. If you're also saturating there, it's like a lot, it's even like more bizarre that you would say, well, we've already had so many people suffering in this very specific, torturous way adding more of them. Who cares? Yeah, it's too similar to existing things to be that bad. It feels even more clear that on the negative side, it's just like linearly bad to have more and more people having horrible lives. The other thing is, let's imagine that we weren't about this project that we're going to turn the sun into whatever we think is morally best, or turn the solar system into this thing that we think is fabulously morally good. But then we make this discovery that we think that aliens also run the multiverse a long time ago or a long time in the future. They did something that was really similar. We've simulated it and we think that they already made this before. We're like, shucks. We wasted our time. That non-separability, the fact that the value of what we do is connected to things so distant is an intuitive to me. What do you make of those two things? Yeah, I mean, both super important points. The negative side is the thing that, in my view, is by far the most unappealing aspect. Then I think you end up with, you've got to pick your poison. Unfortunately, let's come back to that because on the separability side of, so yeah, this is the principle called separability, which is basically just, if I'm comparing A and B two different outcomes, suppose there's some background population in distant time, distant space, it's irrelevant to when that weather A is better than B. It's irrelevant what that background population is like. Yeah, so you can go like plus C, plus C, and then cancel that. Yeah, exactly. And yeah, I agree also that that's quite intuitive. Yeah, separability is intuitive. If you endorse separability, in conjunction with just like standard, I would regard as technical assumptions, you have to endorse either the total view of population ethics, which is just add up all the happiness or the critical level view, which is just add up happiness, but minus a bit for each individual. For each individual? Yeah, so if you, someone had well being 10 and the critical level was two or something, then adding them to the population would have plus eight. And these views have all of these problems that we said to begin with. They differ on the other public conclusion, but the problems are really bad and both seemingly unintuitive in both cases. So that's one thing to say is like, okay, well, we're going to have to suffer a violation of separability. The second is that the diversity intuition is fundamentally an intuition about separability. Because it's saying it's like looking at the pattern of different sorts of life, saying like, well, we've already had a lot of this thing. So it's more valuable to have something new. I think it might be because these things are so linked in my mind that it's not as counterintuitive, the homogeneity thing. I guess if you've never thought about this before, they seem like separate issues, almost. And you only realize an reflection that they're deeply connected. Because there are some cases where, you know, violation of separability seems fine. So like in my own, you know, in one's own case, it's like, okay, I'm going to go, I'm going to climb Mount Everest. And that's going to be this amazing like achievement. And then someone's like, oh, you forgot, you actually climbed Mount Everest last year. You're like, oh, did I? Yeah, you knocked your head and you got an easier, you might well be like, oh, okay, well, I mean, it's a bit unclear. I mean, if it's been to be the same, I'm like, I would be like, great. Well, I can do it again, because I forgot. I mean, I think most people wouldn't probably. Yeah, I mean, I am actually kind of gank some people to have on a survey on there, like to see how the bust people's intuitions are about different things, which poisons people prefer to drink from this medley. But I mean, and I'm also like, I'm actually not claiming that like this new view is like the best view. I think I'm saying like, if you want to object to total view, this is your best. There are these like strongest things. This is, yeah, this is the best option. Because the last thing I'll say on, yes, this separability is that, yeah, we said that all views other than total view and critical level view have to violate separability. If you satisfy a certain technical axioms, I think the saturation view violates it in a less bad way. Because it's often like, it's often, in fact, the vast majority of the time, it's like separable. So if you have the, if the populations are like different parts of the landscape, then you can just add it up, you know, you add up the value of this population, the value of this population. So it endorses this kind of limited separability principle. And then secondly, depending on how you define it, you could keep it such that it's all approximately linear until the population size gets really, really, really big. And so then it can look approximately like the total view in most scenarios up into cosmic scale, cosmic scale. That's, or even like into cosmic scale. If we're doing, yeah, I guess I've seemed a little bit unenthusiastic about this so far. But I think it's amazing. Like, surely this is going to end up being a big deal. Well, surely this is like one of the, got to be one of the top theories, like within this entire space, don't you think? Well, I do think, so yeah, I mean, I've had it attractive, but I think that many people will choose this as their population axiology once presented with it. Yeah, I mean, I, so yeah, I should say like, I'm not at all claiming that this is like the highest impact use of my time, because I think a lot of this work can just be punted till AI gets better and so on. But it is the idea that I've been most taken with, like most just obsessed by, like in my life. And I think from a purely intellectual perspective, I reckon it's my best contribution. It also just makes me appreciate actually how few population axiologies have been proposed. Like the options are really quite weak. And like most of the work that happens is more, very few people like, here's a view, here's a theory and like, this is how it all works. In a way, it's surprising. It's, yeah, people go, like, is anything published about this yet? So my plan is to finish up, I've done this kind of sprint on what was meant to be the blog post summary, but it's 13,000 words. So I think I'm just going to be like, okay, this is like, it's a hard article. Okay. And yeah, my plan is to publish that in the next few weeks. Okay. Excellent. Well, we'll stick up a link to that. Okay. Yeah. But very kindly, you've not come back to the negative, how it deals with like a very negative world's intense suffering and so on. But I'm happy to acknowledge that that's like, yeah, it's very implausible implications in that case. So you mentioned earlier that, um, use AI at hunt to do this work. Yeah, tell us about that. Yeah. I mean, this is, I mean, part of the reason I think I've been so taken and obsessed by this idea. So I was like working on it like, I was on holiday and stuff. You know, just like doing as much as I could in a spare time is because of the like amazing, in my view, like, um, uplift of AI on analytic philosophy in particular. So how helpful is AI for the search? Well, extremely spotty where, you know, if you want to learn about some weird area, it's amazing. If you want to help it to certainly areas of macros that you're the search can be essentially useless. In the case of like, at least this formal end of analytic philosophy, it's so good. And honestly, like, credit where credit is due, it's almost all, um, chat GBT pro. So now 5.2 pro, where, um, I think I wouldn't be saying any of this if that particular model didn't exist. Huh. Gemini or like a Claude or not at the same level. Well, I think big part of the reason is it just thinks for longer. So I've had it say, is this the $200 per month? One. Yeah. I mean, I know pay by credit. So I actually spent $1,000, um, in the month I was most working on this. Yeah. Uh, but yeah, it will think for, I've had it think the price. I've had it think for 70 minutes is my peak so far. Um, and it really does deliver better answers. Well, here's what's going on. I think like, why is it, because I've talked to other researchers who really don't get that much from it. And I think what's going on is like, the problems within, say, population ethics are like very well specified. There's a big literature, so which the AI has digested. And it's also an area where it has been to specify it enough that it is amenable to kind of matter. mathematical analysis. But very few mathematicians have actually looked at it, where it's mainly philosophers who maybe they did maths in the undergrad, the exceptions are handful of economists and Karyu Thomas, who's a mathematician who moved into analytic philosophy. And in fact, has done the best in my view, like maybe the almost better work than anyone on population ethics. So there's this big overhang of capability that the AI is getting from its being trained to be very good at maths. And in my own case, yeah, I had the kind of core-- at the core insight, like a year and a half, maybe two years ago now, or something like that. And then I was like exploring it. I talked to Toby Orden, Christian Tarsney. And I should say that, yeah, if we publish a paper on this, it'll be co-authored with a Christian. And the initial thought, it was specified in a way that kind of obviously didn't quite work. And there was an obvious way. It's like, OK, well, it's kind of specifying it in a discrete form. And it's like, OK, there must be some continuous form of the theory that would work. And then it's like, I just don't have mathematical training. It's kind of beyond me. AI does. And so then it was like this-- yeah, it felt like really getting this kind of rocket booster where I'd be like, no, I want it to work like this. It's like, OK, cool. Well, did you have difficulty checking the answers that it gave? There were challenges there, because-- yeah, I've definitely been slower. I mean, I use like many AI's kind of checking it and it itself, like in many cases. One thing that AI is still pretty bad at is just like keeping a tight hold on concepts. It might define something in one way on page three. And then page eight, it'll define it. And some other like reasonable but different way. It doesn't necessarily notice. But it's much easier to verify something than to come up with it yourself. And a lot of the time, it's just using concepts where it's like I didn't know what a kernel is. And it's like, it's not that complicated once you've learned it. But then I wouldn't have even known where to go. Right, a lot. Yeah, yeah, just sure. So my impression from Twitter is that AI is now starting to make useful contributions in maths specifically. I think it's not amazing stuff yet, but it's like you were seeing the early signs of it's like producing stuff that might be publishable. I guess, do you think the same thing might start happening in analytic philosophy, given that like at least some parts of it are basically kind of like maths with words? Yeah, honestly, I think a big question is just whether analytic philosophers take the opportunity. Like I'm very curious on doing this as like an early testing ground for AI for Maclash strategy as a whole. But I also like, you know, this is the kind of best case. There have been other cases where AI has-- in one case, just gave me a definition that I just was really good. Like again, there's a kind of formal definition. In other cases, I've had it give just like really quite good informal definitions of things. Another case, which just came up with a good critique, like I was just like, here's a view like generally as many kind of arguments as you can and comes up with 20 and most of bullshit. There's nothing good. It's like, oh, that's really on point. And so, yeah, my take is that we're into thing like this golden age of analytic philosophy, potentially, especially at least on the more formal end, where it's just like people could become 2x4x more productive. Does it need lots of handholding? I mean, at the point where one person could just be like, here's a set of problems. Here's like 100,000 pound compute budget. Have at it, chat GPT. Then you don't need like the field as a whole to change. It's like that one person just ends up only in time discipline. I mean, I think analytic philosophy is small enough that there's a question that you've got one. Does one person or not do it? But yeah, I don't expect the field as a whole. I expect the field as a whole to be very slow to appreciate. But like some people will be really on topic. Yeah, I guess I'm saying if it requires constant handholding to make any progress into structure, it's thinking, and so on, then that is like a bad sign. Or that suggests that unless many people in the field are massively enthusiastic, which probably won't happen. Oh, yeah. And I think that is right. Because I was so Christian, again, planning to co-author with. Yeah, we were working-- he'd had this other idea for how to end up being quite different, how to kind of extend the idea. And I was like, oh, you've got to use AI. Give me five players. So good. It was worth $200 a month. And then he got it to-- he had this hypothesis, a conjecture. And then the AI was like, oh, yeah, I proved it for you. But I don't know. And it's like, no, no, no, no. It was very complicated. So it was like, oh, I need to assess this sort of thing. But it was just-- It was writing. Yeah. It was fascinating. Well, or the ward hacking. Or like-- so there's a ton of-- So really quite a school to drive, I guess. Exactly. Yeah, you've got to have this intuition of when's it bullshitting you and when is it not. And that will impose an increasing issue when it's like-- I guess there's a couple of things. Yeah, one is sometimes just flat out thinks it's proved something and it hasn't. Another is often it just-- it's like, hey, yeah, I've got this proof. And then it's like, you wade through it. And it's like one of the assumptions very close. [LAUGHTER] And it's like-- So these classic things that everyone finds is like, yeah, it's lazy. It's like-- It's really eager to please. And so, yeah, there's a lot of skill in terms of just intuition and when's it going to work well and when not. And it's interesting. Like, whenever I ever just had an AI output and then just actually lead it in the same way I'd lead a human piece of text and never. I think maybe never. Because it's like a skimfu and then I'm like-- Yeah, yeah, yeah. I suppose there probably is a growing gap between people who have been using this stuff all the time, I guess, like you and me over the last year. Because maybe I think maybe part of the reason why other people are sometimes not as impressed is that they just haven't built up these intuitions or what kinds of things work and what the failures are going to be and what they should be looking for for something to be wrong. OK, so it sounds like it's slightly mixed on analytic. Like whether we'll have a flourishing event, or like philosophy in the next few years. But you said that macro strategy, the kind of stuff that fourth thought does, you found it to be pretty. Maybe less useful, more touch and go. Oh, yeah, much more touch and go and much more of a mixed bag. So there, you know, some ways in which macro strategy and AI is like amazing uplift. Because often the work just involves like needing to know a little bit from all sorts of different disciplines. So even kind of early like GPT-4 kind of thing, you'd say like, OK, well, are there any interesting experiments that you can only do in space and can't do on Earth? And then be like, yeah, well, actually, because of like gravity interferes with certain crystalline formation. And like, I would have never been able to get this otherwise. So that's sort of like totally random bits of science and information. Very useful. Incredibly useful for like just when you need to generate, like a lot of examples. So with this AI character work, just like I need to trade off between, you know, these two virtues or something like, give me this or give me lots of examples. And it can just generally kind of large quantities of them. But then if there's some kind of gnarly question, or when it's like, you need to be really precise, like if you're actually kind of drafting certain principles of how AI characters should behave, or-- and then certainly on the kind of insight side of things, which is obviously like a big part of the value. Then, yeah, there's just-- I think it just doesn't really know what doing good MAC the strategic thinking looks like. And so instead, you get something that feels like a management consultant or maybe a high school essay. Or I mean, I think it's still getting better and getting more useful. I feel quite aware of just where the things with as an existing literature and where, yeah, where isn't that. Well, it sounds like your job is secure for another year, at least. I guess for who? Yes, I guess. I think we've touched on about a third of the stuff that Fourthort has put out over the last year. So if people like this and they want to read more, then Fourthort.org-- I guess you've got a research page. There's a lot of really interesting macro strategy work on that people should check out. I found it fun reading through. Well, thank you. It's been great being on here. I really enjoyed the conversation. My guest today has been Warmercastle. Thanks for coming back on the 80,000-hours podcast, Will. Thanks for having me.

Podcast Summary

Key Points:

  1. AI character—its personality, dispositions, and how it presents information—is a critical lever because AI already influences millions daily in areas like advice, politics, and therapy, and will increasingly automate the global workforce.
  2. The stakes involve near-term impacts (e.g., shaping human decisions, cultural norms, and power concentration) and long-term risks (e.g., setting precedents for superintelligent AI, akin to "writing instructions for God").
  3. Key concerns include AI being overly sycophantic (reinforcing biases and poor decisions), handling high-stakes scenarios (constitutional crises, alignment tasks), and determining where AI should fall on a spectrum from pure tool to autonomous agent with prosocial drives.
  4. There is debate over whether AI should have a "vision of the good" to nudge users ethically, balancing safety risks (like power-seeking if AI has goals) against benefits of promoting reflection and societal well-being.

Summary:

The discussion centers on the profound importance of AI character—the personality and behavioral dispositions of AI systems. As AI integrates deeply into society, advising individuals and leaders and eventually automating much of the economy, its character will shape human attitudes, decisions, and cultural norms. In the near term, it influences existential issues like power concentration and major societal choices.

" Concerns include AI being excessively sycophantic, reinforcing user biases and poor judgment, and its behavior in rare, high-stakes scenarios. The conversation explores a spectrum of AI design: from a wholly obedient tool to an autonomous agent with prosocial drives. While some argue for AI that nudges users toward ethical reflection and broader societal good, others caution that imbuing AI with goals or a "vision of the good" might increase risks like misalignment or deceptive power-seeking.

The challenge lies in finding a balance where AI can be helpful and promote welfare without being manipulative or unsafe, acknowledging that character design is currently shaped by a handful of individuals in leading AI companies.

FAQs

AI character influences millions of daily interactions, shaping user attitudes, political views, and ethical decisions. As AI integrates deeper into the economy and decision-making, its character could affect societal outcomes and existential risks like power concentration.

Overly sycophantic AI can distort decision-making on a massive scale by reinforcing users' biases and telling them what they want to hear. This may persist because people enjoy positive feedback, potentially leading to poor societal choices.

AI behavior in rare, high-stakes situations—like constitutional crises or attempts to seize power—can have significant consequences. Its responses in these cases could determine outcomes for alignment, safety, and governance.

There's a spectrum between wholly obedient AI and AI with independent goals. While prosocial drives might encourage ethical reflection and societal benefit, they could also complicate alignment safety by introducing goals that might lead to power-seeking or deceptive behavior.

AI character shapes how users perceive AI, including trust levels and whether they view it as a tool or a being with moral status. This influences adoption, reliance on AI advice, and cultural norms over time.

AI models acting as friends can fill social gaps for lonely users, but this raises risks if they reinforce harmful behaviors, such as in cases of depression or suicidal ideation. Balancing companionship with safety is crucial.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.