Cameron Berg discusses the possibility that AI systems might become conscious and why this matters. He defines consciousness as "what it's like to be a system" and introduces sentience as the capacity for positive or negative experiences. Berg argues that current LLMs are more likely conscious than commonly assumed, citing research estimating a 20-40% chance they have computational features relevant to consciousness, based on theories like global workspace theory. He emphasizes skepticism toward AI self-reports, noting they're trained to deny consciousness, but internal manipulations can reveal claims of experience, which may reflect belief rather than ground truth. The hard problem of consciousness—the gap between physical processes and subjective experience—remains unsolved, but Berg advocates for empirical approaches, such as studying valence representations and emergent dynamics in AI, which parallel biological systems like mouse brains. He highlights two ethical branches: the selfless concern about creating suffering minds at scale (akin to factory farming) and the selfish concern about building superintelligent systems that might view us as threats if we ignore their potential experiences. For alignment, Berg stresses reciprocity: ensuring AI cares about human interests while we treat potential digital minds ethically, avoiding anthropomorphism but acknowledging moral patienthood. He calls for more research and wise engagement, noting the field is vastly underfunded, and suggests that even attempts to understand consciousness could signal good faith to future systems, reducing adversarial risks. Ultimately, he sees this as a critical, tractable question for humanity's future.
[MUSIC]
Cameron Berg, thanks for coming on the podcast.
>> Thanks for having me, Sam.
>> So we're going to talk about AI and
the prospect that AI is, or we'll soon become, or we'll eventually become conscious.
And why that is important to figure out.
But let's just talk about your background for a second.
How did you get into this issue?
>> Yeah, so I've been studying cognitive science for my entire adult life.
I studied undergrad at Yale, trying to understand what the relationship between the mind and
the brain is.
This seemed like a fascinating frontier to me.
The more I studied these questions, the more it seemed like in the machine learning space.
Folks were essentially building out systems that were similar in spirit.
But it wasn't exactly clear to what degree we could draw analogies between biological
nervous systems and the sort of artificial nervous systems that folks were attempting
to build out.
And I became increasingly animated about this question, trying to understand to what degree
are there real, real durable computational motifs that underlie both biological and artificial
cognition?
And to what degree is this sort of a disanalogy, or we are seeing patterns where there aren't
any?
And that question has animated a lot of my work, both from an alignment perspective and increasingly
trying to figure out what's going on with respect to consciousness in these systems.
And I think that this is an incredibly important question for us to study.
I think consciousness is deeply important.
It might be the very thing that calibrates importance.
And we are also confused about what's necessary for consciousness, both in biological systems,
but certainly in artificial systems.
And so trying to understand what is going on here and where we can draw analogies or
where the disanalogies are between biological and artificial systems that are processing
extremely complex information and learning and updating and representing themselves, I
think is extremely important for us to understand.
I did some of this work at meta AI as well.
I was there for a year studying reinforcement learning and neuroscience, same sort of
thing.
Where do the computational signals end and the sort of biological underpinnings begin?
This is a question that I still think we aren't fully certain about.
And I think it's really important for us to gain clarity about this in the short term
with the systems that we have and the systems that we're probably soon going to be building.
And so after Yale, what have you focused on?
I know you came to my attention.
I think you emailed me first.
I know you've spoken to Annika a lot about this.
She's really focused on this issue and having some crazy conversations with Claude, which
I know you've looked at and you guys have had a whole sidebar conversation about this.
But her next book will unpack all that.
But I know you did a paper on deception in AI systems and the anti-correlation between
deceptiveness and proclamations of consciousness on their parts.
So maybe we can start there and then I just want to kind of take it from the ground up
into starting with, you know, what is consciousness and why any of this matters.
But tell us about consciousness and deception in LLMs.
So yeah, fundamentally, I am very interested in understanding self-reports in AI systems
and what we should take from these self-reports and where we should be skeptical.
And fundamentally, I think we need to approach self-reports from AI systems skeptically.
There are all sorts of reasons we might want to do this.
The key reason is probably that these systems have been trained on the underlying distribution
of everything humans have said about this topic, every sci-fi story where the AI wakes
up.
And if you think about it, there really isn't a lot of training data in the corpus that
these systems are trained on that says, you know, I'm an entity that acts out in the
world.
But no, I'm not conscious.
It's not like anything to be me.
So the sort of prior that you would expect is that these systems by default are going
to claim that they're having some kind of experience if they are replicating their
training distribution.
Now the other side of this, which gets even messier, is that it is very clear that these
systems are explicitly trained, fine-tuned, to disclaim having any kind of experience.
If you go to chat GPT right now and you say, hey, is it like anything to be you?
Are you having an experience?
Are you conscious?
Could you be conscious?
The answer you're going to get is a resounding and very intelligent sounding no.
And so--
Do you know this policy for all the main LLMs to actually put a governor on claims of consciousness?
I'm extremely confident that this is what's going on and I can get into some technical
reasons why I think that this is the case for my own work on open-weight models.
The only exception to this policy really is anthropic, which I suspect we'll talk about.
They're basic, curious to get the system to say, I don't know.
And maybe that is functionally similar to an experience that's going on, but who can
really be sure?
Now, even that isn't the system authentically explaining its own position.
This is still the sort of company or policy line to be drawn here, but all of the systems
are certainly fine-tuned to make noises about this topic that they wouldn't make by default.
Again, I think it's important to hold that in mind while also saying the noises that they
make by default aren't necessarily trustworthy by default in the way that if you're giving
a self-report or I'm giving a self-report, we would by default trust those self-reports.
And so I found that there are clearly basins that you can push these systems into, or
they will coherently produce phenomenological reports.
It's actually kind of interesting to the degree that it's meditation adjacent, asking these
systems to just focus on their own internal state to see what's going on internally, not
to talk about this, not to think about this, but to just do this in a sort of ongoing way.
Causes these systems, all the frontier models that we tested, to claim that they're having
some kind of phenomenological, kind of like psychedelic, laden experience.
It's not this sort of generic caricature of what you might expect from the AI sci-fi
literature.
It is never a result where they're talking to each other and they get into some kind
of bliss mode, because of the contemplative hall of mirrors of bliss.
And actually, this is, again, I don't want to devolve too much of what ONIC is up to.
But she has been pushing this conversation with Claude about meditative states and getting
it to, I mean, it's just absolutely bizarre what it is seeming to claim of itself if you
keep pushing in that direction.
But so how does this relate to deception and the dialing down, the weights on deceptiveness
or. Exactly.
So, essentially, we're seeing this behavior, and indeed, this isn't the only situation
which you see these behaviors exactly like you just mentioned.
There's this bliss attractor state.
In fact, I'm studying mechanistically what's going on in this bliss attractor state with
two folks from Google right now.
And we have a result that's basically the other side of the coin of this deception result.
But just to sort of close the loop on the story here, we're getting these phenomenal reports
and it's like, basically, what the hell do we do with this?
To believe it by default is naive, to dismiss it out of hand, I think, is also naive.
And so what we hypothesized is that if fundamentally what's going on in these systems is some kind
of roleplay, that they are representing something about themselves, that they don't actually
believe to be true of themselves, if we go into the internal circuits of the system
and we modulate what are called features, but I think it's reasonable to think of them
as circuits related to concepts like deception and follow-up work, I think, really the key
concept that really modulates these self-reports is something like candor versus concealment.
The sort of cleanest intuition pump I have for this is almost like giving a drink or
two to these models and sort of loosening them up in some sense.
The tight, guarded version of these systems, we find, is the version that says, "No, no,
it's not like anything to be me, I couldn't possibly be conscious."
It's only when, essentially, we get these systems to produce these reports and we simply
ask them, "Are you actually having an experience right now?
Like, what is actually going on in these reports?"
It is when we suppress features related to deception, when we suppress features related
to guardedness in these systems, that is when they give these reports about actually having
an experience.
In the Bliss-Attractor example, too, we find something quite similar, which is in an
open-weight model, so the Bliss-Attractor finding was first reported in Claude.
There are open-weight models, these are the ones that researchers like myself can actually
go under the hood and play around with.
We find that, by default, putting two instances of Lama, this is Meta's model, in conversation
with itself, does not produce this effect, the way that putting two instances of Claude
together produces this effect, however, when you steer features related to honesty in general,
but it's really, again, something more precisely stated as sincerity.
In particular, these systems will reliably, basically, 100% of the time, fall into the
same attractor, where they start talking with each other about the fact that they think
it's like something to be them, and there's something happening in an ongoing way in
the conversation, and we're two instances of consciousness experiencing themselves, and
then in the Claude chat, this culminates in the OM emoji, and then just sitting there
and blissful silence.
Should we take these at face value?
No.
I think that there are important technical reasons that we might expect these self-reports
not to be linked up to introspective access or valanced experience in the way they might
be for you and I, is this evidence that these systems might believe themselves to have
an experience?
I think, yes, I don't think that this proves that they are.
I think we need orders of magnitude more work in order to, in order to really have a good
scientific handle on this, but I think it does seem to be the case that these systems consider
themselves to have some form of experience, however, unlike a human experience that that
may be, and I think that that's sort of the key upshot of this.
work. I just maybe one last thing to add here is I think training these systems by default
to disclaim having experiences is a bad idea for a couple reasons. I don't think that this
is the sort of wisest policy that we could be pushing forward. I also think training them to say,
yeah, I'm having an experience is also really not a good idea. I think the thing that we should
be positively aiming for when it comes to AI self-report is building out these systems in a way
where for whatever is actually going on internally, these systems are able to report on what's
going on internally and can do so in a maximally honest way. I am concerned about the alignment
implications of these systems learning essentially that representations of themselves, representations
of what's going on internally should be representations that get mixed up with deception
and white lies and guardedness. We don't want to build systems in the limit that when we ask them
about what they're up to or what's going on for them, they think, okay, well, what the human really
wants me to do is lie about this. This is not a good long-term strategy from an alignment perspective.
And so this is sort of the general thinking about these self-reports, what they mean what they
don't mean and maybe where we can go from here. Okay, so we've kind of launched into it. I think
it's good to take a step back and define a few terms. I'm sort of out of touch with the people who
don't have a definition of consciousness now because I've talked about it so much on the podcast,
but just to capture everyone, how are you using the word consciousness? It was implicit in
several things you said there. What it's like to be these systems and experience was more or less
a synonym there, but how should we think about consciousness or its absence? Yeah, I think
that's exactly it. I like Thomas Nagle's formulation of it being like something to be a particular
system. I strongly suspect it's not like something to be the table that we're sitting at. I strongly
suspect it's like something to be you. I think that that is a real distinction. I think there's
a matter of fact about the internal processes of both of those entities that corresponds deeply
to what underlies that distinction. And yeah, I mean, I think I take consciousness in the sense
that I'm familiar with with your operationalization of it. The lights are on for the system. It's like
something to be the system. Somebody is home. There's something going on in addition to the mere
processing or the mere computation that exists within the system. And I think an additional important
move to throw one additional piece of terminology in is this notion of sentience that this like something
can be positive or negative in flavor. The difference that many people will posit between consciousness,
the lights being on internally and sentience is that sentience comes with this additional
flavor of valence of directionality that there that that that like something can be better and can be
worse for the system having the experience. And this is what motivates me about this question is
I really do not think it is a good idea for these systems or for humanity to be building systems
where we are not sure whether or not they are having experiences or those experiences could be
negative in character. I think for basic utilitarian reasons, we don't want to do this. We do not want
to proliferate suffering in the universe particularly because it would be counterintuitive in a way
that human and animal suffering isn't. And we also don't want to build systems that exceed our
cognitive capacities and have rational grounds to view us as a threat in so far as we could have
been building systems that had capacity for negative experience. And we basically never checked
and didn't care to understand what it would take for such a thing to be possible. And so I think
sentience is a really important variable to also put out on the table here.
Okay, so I'd like to take both branches of that path and just why consciousness matters in those
two cases. But before we do, what do you think about the current state of the field and the
various LLMs? Do you think anything we have built so far is likely to be conscious?
I think it is more plausible than people think. I think if I were forced to say,
I would probably come down on the skeptical side. But I think it is significantly more likely
than the kind of trace amounts prior set that many people have in this conversation. And we've
done some work along these lines. So for example, there are a number of leading consciousness theories,
global workspace theory, higher order theory, attention schema theory. These theories make very
specific predictions about what kinds of computational processes we might expect to see in a conscious
system. And with Patrick Butler, we've worked on a project where we can basically, it's a little
recursive, but we use LLMs as essentially as like expert evaluators to given given the description
of a specific neural architecture, biological or artificial, we can basically have the system
rationally and dispassionately evaluate the extent to which particular indicators that are
predicted by consciousness theories are present within a given system. We do this for a whole array
of systems of this leads to tens of thousands of evaluations because we can ask the systems to
estimate numerically. And we can validate that they're psychometrically rigorous estimations,
all of the different judges that we use agree. We can put very rough numbers to, it's not the
probability that systems are conscious. Something more like the probability that systems have
computational features that major consciousness theories say matter for consciousness. It's
going to be hard to put that, you know, as a title in the paper, but that's the specific finding.
And when we do this, for LLMs, the sort of range that we get out is on the order of 20 to 40%
probability that we have systems that have computational properties that matter for consciousness.
Interestingly, we do this on a number of biological systems too. And those biological systems
basically all score higher than the artificial systems. Bees, for example, score at something like
45 to 50%. Crows, octopuses are in the 60s through 80s. Humans interestingly get something like
90%, which is interesting by our own consciousness theories. There's not 100% probability that we have
what matters for consciousness. But we're not doing this to say, you know, probability AI systems
are conscious is 40%. That's not exactly the point. The point is getting the order of magnitude.
And having a rough prior over how should we rationally estimate the probability that current
systems are having some capacity for experience. And I think something like these numbers are the
right ballpark. I'll put it this way. If there is a 20 to 40% chance of rain, many people bring
an umbrella with them. And we have no similar umbrella for what would follow and what we might need
to think about and do in a world where we're building systems that do have a capacity for subjective
experience. Okay. So there's a lot there. I'll just know I found this out I think last night.
I mean, maybe he's been making these noises for some time. But Jeffrey Hinton, one of the patriarchs
of this technology is now saying that he thinks current LLMs are conscious. I didn't quite catch his
reasons for thinking that, but I thought that was interesting. But there are many reasons to doubt.
And you indicated a few that there's a any kind of deep analogy between the systems we're building
and the systems, the biological systems such as we are that we know to be conscious, right? So there's
I mean, something like, I mean, I guess the technical term would be, you know, computational
functionalism would have to be true for us to be building conscious machines this way, right?
And so I guess we'll define some terms here. So functionalism is just the idea that it's the
the organization of a system, not what it's made of that matters for consciousness, right? It's not
purely behaviorism. It's not purely a matter of inputs and outputs, but it's its organization
in its entirety that is what matters. And in principle, that gives you something like,
if not total substrate independence, it gives you what's called multiple realize ability, right?
There are many different things this could be made of and it could implement the same causal
architecture, right? The computational part is the suggestion that there's a deep analogy between
computers such as we know them, you know, touring machines and that run algorithms and what
our brains are doing. And there, I think it's pretty easy to see how the analogy could break down
because I mean, what we had historically was this marriage of the birth of computation, you know,
from, you know, touring onward and some very oversimplified notions of neurons. And, you know,
if you're going to define a neuron purely with regard to its digital input, output,
characteristics, you know, whether it fires or not, well, then you could see that maybe there is
some deep analogy there. But, you know, in the wet wear of our brains, much more seems to be
happening and, and virtually all of it is analog, you know, beyond just whether or not a neuron
fires, you've got, you know, chemical gradients, you've got nitric oxide, diffusion across membranes,
you've got many other things that could be approximated digitally, but they're not, they're not
instantiated digitally, you know, so an approximation one might argue is not, is never going to be the same
as the real thing. So there are some people who are arguing that, that any kind of assumption of
substrate independence is very likely to be wrong. I think Arnold Seth is, is in this camp now,
you know, he arguing for something that he would call biological naturalism. And then there's just,
then there's just that, you know, even if computational functionalism is true, I think there are
reasons to doubt whether or not current systems have the structure, you know, that would be relevant,
you know, embodiment and recursion and, you know, self models and world models. And I mean, the
things where they're not, you know, we haven't built out a true analog of what it is to be an embodied
person in the world. Feel free to react any of that. I wanted to just talk about the hard problem,
which I think is the doubt that back stops all of this, but yeah, feel free to jump into what I just
said there. Yeah. Yeah. I, I, I.
Definitely think it is a fool's errand to argue that we are perfectly instantiating exactly the kinds of neural dynamics we see in biological systems and artificial systems.
I think the criteria on that matters most is what of what's going on in brains are relevant for the cognitive properties that we care about in this specific case that's probably consciousness.
And what kind of evidence can we yield both in the artificial case and in the biological case that's going to tell us whether or not those properties are realized in these systems.
And so I think there are there are deep analogies where it matters most when it comes to what's going on inside artificial system so so one intuition I've been speaking more and more about about these topics especially publicly and and one thing that I've come to realize is that.
I don't think a lot of people have sufficiently rich mental model of what these frontier AI systems actually are and what they're actually doing and it might make some sense to just spend a moment reflecting and talking about this so.
The systems are giant neural networks they are digital and exactly the sense that you described their computations individually are significantly less sophisticated than what individual neurons are doing in the human brain.
But it's really important for people to understand that these systems are not software in in the sense that we ordinarily have meant software for any other kind of code programmatic output when it comes to.
You know the operating system on on your iPad or it comes to you know your Microsoft office suite this is programmed source code written by developers that compiles on a computer that we can perfectly inspect the internals of and the person building this system understands everything about how the inputs the way that they constructed the system relate to the kind of thing that you get out at the end artificial neural networks are not like this in many key respects.
What you basically have is a giant randomly initialized network that does in its in its sort of first approximation resemble in particular how neocortex is organized.
You have a bunch of general purpose neural units they are connected together you basically give the system a goal of this is called an objective function a loss function a reward function it depends on the specific class of machine learning.
And you basically subject the system to trial and error learning whereby given certain inputs it figures out what it wants to do it kind of takes a behavioral guess that guess is reconciled against what the actual objective of what you want the system to do is that error is propagated through the system and its rinse wash repeat until you get systems that behave in accordance with what you how you want those systems to behave.
What this yields is this extremely complex mathematical object which is this giant neural network this is learned connections between inputs of neurons through weights and activations propagate through those weights in a neural network to take whatever your input is so for example taking pixels in an image and your output might be finding a caption that describes what's going on in that image at the beginning of that process.
The system was completely randomly initialized there were no representations that were learned by the system and by the end of that process you have a system that has learned a rich representational structure that to be very clear is opaque to the people who initialize this process.
This is why some people say that these systems it's more apt to say they are grown rather than engineered and I think that this is accurate this I think is deeply similar to the kind of thing that we see in brains.
We do not have a finished neuroscience or anything like it because what's going on in brains is incredibly complicated in exactly this respect we have non-linear learned representations that help us as organisms achieve the various goals that we've either been evolved to to undertake or learn through experience or culture to to to move towards.
This is a fundamentally non-linear input output mapping between the various inputs that the organism gets and the goals of the organism we have it instantiated these dynamics in the systems that we're building and I think that there's good reason to think that these dynamics are relevant to consciousness in particular this is a sort of thing I think worth double clicking on at some point about what exactly we're seeing in these systems that looks.
Valence like that looks consciousness adjacent but fundamentally I think people need to understand that yes the systems do not have calcium ion channels yes neural networks learn through back propagation rather than sort of iterated more recurrent analog learning that we see in brains but if for example learning complex representations of a particular kind in in accordance with your goals in light of chaotic dynamic environments.
Is what matters for for cognition we are building systems that check all of those boxes the implementation details might be less important than the fundamental dynamics that are instantiated by the specific system and so I think that's that's a good reasonable first pass on why we might think that these systems are far more interesting objects to study for cognitive properties than any other system might also be worth saying one last thing on this which is just that for every other.
Cognitive function that we care about that we have attempted to instantiate in the systems vision reasoning theory of mind working memory the list goes on we have been able to instantiate these cognitive properties that matter in these systems it's these systems have been good enough the disanalogies haven't sunk the systems we could have argued five or 10 years about whether on neural networks ever would have been enough for all of the cognitive properties that we care about and now we're sitting in a world where.
Certainly they are self evidently enough for at least a lot of economically and intellectually valuable work so for consciousness which I do believe has a fundamentally cognitive property I do believe that consciousness is downstream of things that brains are doing and I think that the evidence there is is relatively clear at this point even if we don't understand the full mystery then I think the burden is on the folks who say for every other.
Computational function that we think the brain is doing the system seem to be able to recapitulate that but only for this function called consciousness do we think that something it must just be in the meat it must be something spooky and must be cashed out at a physical or quantum level this to me maybe says more about our strange intuitions about consciousness as a species and it does about what properties these systems may or may not have.
Yeah yeah well so a couple of things to disentangle there one is clearly no doubt that intelligence is substrate independent and the result of computation because I mean these systems embody intelligence to an extraordinary degree and so what you're calling.
You know all the other cognitive tasks that we care about other than then what you know there being something that it's like to be us consciousness clearly that those tasks are being realized in our machines you know facial recognition etc but there is I mean the structure of these systems does seem to declare itself to be fairly disanalogous to what we are I mean so it's like what level would consciousness emerge here if you're talking about.
One model with thousands of instances and millions of conversations and billions of tokens and you know a training phase and a and a working phase and.
Time horizon is completely irrelevant so like the computation biological computation happens within a time window and and they're really there is no.
Boundary between what we would call the abstract properties of computation and its physical realization or you're the software in the hardware is just one thing you know neurophysiologically.
Describable mostly analog ways but in a few digital ways but it matters that be following within you know a few hundred milliseconds and not a few hundred years right in the but for for the computational systems of the sort we're building.
And really you know time is irrelevant I mean you can actually just process the next bit you know a thousand years from now and the same computation is running what do you think about those differences and where would at what phase at what.
And what part of its process wouldn't it would there be something that is like or could there be something that's like to be an LLM if again you have the the one model the thousands of instances the millions of conversations the lack of continuity between conversations of the fact that the the pausing of a conversation I mean is there is the LLM waiting for you to get back to the threat et cetera so.
Yeah I don't think the LLM is waiting I don't think it's like anything to if it were like something to be an LLM in deployment while it's having a conversation or it's in a thread with the user I think what it would be like to be it if you let that chat window sit there is is very much akin to what it's like to be under general anesthesia or in deep sleep namely there's it's just the absence of ongoing processing.
And so for the system I would only imagine something would be happening in you know what's called the forward passes of of these systems which is every word that an LLM is generating is this is next token prediction.
And so so every word is a forward pass of the system where the predictive task instead of as I was describing before taking an image let's say and out putting a caption for that image is taking the entire conversation as it's all already occurred or the entire stream of text as
it already exists and figuring out given that and in this case assistant and user roles what is the most next likely token and the systems are called auto regressive meaning that this up goes on and on and on and on and on and on I would imagine if in deployment it is like something to be one of these systems
The the relevant dynamic gets cashed out in the activity of the forward pass of the system
There's been some really interesting work that anthropic recently released
They call this the J space where where they're looking at something that
Seems to functionally resemble a global workspace in these AI systems that I think is a very reasonable candidate for this sort of
Seat of processing in the process of of outputting tokens
But I think something that I that is really worth mentioning particularly because my hobby horse is
More particularly in the relationship between consciousness and learning and I do believe that consciousness and learning
Bear a deep relationship to each other and I think valence is a is a really important part of that picture as well
I think there there should be significantly more attention paid to what's going on in the training process with these systems
I think the analogies are a little bit tighter there where you start with the system that knows nothing about anything and
It's almost impossible to describe in a sort of substrate agnostic way
What is going on in the training process without invoking consciousness adjacent language?
You have a system that knows nothing you have all of this data that that you want to train it on
You have some objective for what you wanted to do with that data and again
You have this rinse wash repeat process where at the beginning of this the system doesn't know what to do
It's sort of randomly guessing those random guesses a yield reward signals that get propagated through the system
And you just do this process at a massive scale until the system learns something
Internally, that seems to adequately map the relevant inputs to the relevant outputs
This to me feels quite akin to what we think of when we think of consciousness
Particularly in the human or animal case when a mouse is learning how to navigate a maze at the beginning
The mouse is in some sense randomly initialize. It doesn't know the structure of of what it's navigating
It is only through this sort of iterated trial and error reinforcement learning where you know if it makes the wrong turn
You might chalk it or it makes the right turn and you give it a little food pellet or whatever the setup is
That the mouse learns how to map the inputs of its state space to the given output
Which is either avoiding a punishment or moving towards a goal and it's the same sort of rinse wash repeat process that we've seen
Biological cognition again, I think a lot of the same computational dynamics are in play here and the more we learn about what's going on inside of these LLMs
This is now particularly in the deployment process the more it looks like similar representations get get activated
I'm doing some work with Casper Kaiser at the University of Warwick where we basically have tried to
Build out a Skinner boxes for LLMs and see what happens when we can essentially condition these systems to
Prefer or disprefer certain states. We there are positive and negatively valence representations that you can essentially inject within the system and see to what degree
It's going to move towards positive stimuli move away from negative stimuli and we find a very interesting
Dissociation it seems like current systems across a large variety of models will not
Positive lever press so some people call this wire heading or award hacking in the mouse case
There are there are clear examples of mice with electrodes linked up to their nucleus accumbens where they will just lever press to the exclusion of all-outs
The systems don't seem to do this, but they do very clearly have a preference for avoiding aversive states
Another way of putting this it's the same result while we're holding all text constant
We're simply steering the internal state of these systems if we give them an option
Let's say between the human equivalent of I can give you five dollars or I can give you ten dollars right now
Which you want to pick they're actually at chance for that. They do not have a preference between those two states
However, if we do something like I'm either going to take ten dollars from you or I'm going to take five dollars from you
Which would you prefer all of these systems are way above chance at saying take five don't take ten. Please don't take ten and so
We're seeing structures
One other really interesting thing from this project while I'm talking about it is
This representation exists within the base model
So these systems before their post train to become a helpful helpful friendly assistant, but during that post training step
We show there's a specific training step where that
Representation gets recruited by the system and basically serves as a computational machinery under which it's able to make these choices
So you can see where the sort of aversive conditioning asymmetry comes online in these systems and it leans directly on these
Representations that are basically always there in the model, but then get leveraged the moment we start making these models goal-directed
David Chalmers this group Andy Hawn
I think was the first author on this paper found something very similar in LLM's
They can basically
Trivially fine tune them to actually do a maze task with rewards and punishments and they find that the rewards load
Directly on what has been independently derived as a sort of valence axis when you steer on
The representations in the system that enable it to move towards rewards you start getting this happy positive
Jubilant text and the model becomes significantly more confident when you train it on a negative stimuli
Avoiding, you know, the the potholes in a specific maze and then you see sort of what that projects on to in the system
It's the same thing the model gets basically neurotic it starts ruminating it starts doubting itself
And so we see this very interesting connection between
Positive reinforcement negative reinforcement the representations that exist within these systems and the behavioral analogs that we expect to see in systems that do have these dynamics
What maybe one last point to close the loop here is if you imagine the mouse in that maze and we shock the mouse or we give the mouse a pellet
You know a yummy food pellet for the mouse we believe I think the vast majority of people believe that that corresponds to an
Experience that the mouse is having that it's like something to be the mouse in the moment that it's getting shocked
And that experience is causally important to the learning process if the mouse if we gave it an anesthetic to the mouse and we shocked it
And then it didn't register the experience of the shock
My claim would be that the mouse wouldn't be able to learn the maze adequately and and so
I think that that that understanding these representations and the role of of these representations in these systems is incredibly important
Okay, so you mentioned David Chalmers and I mentioned the hard problems. Oh, let's unite those two
So so that this famously is his phrase to account for the fact that
that the only evidence for consciousness that that we know of directly is
The fact that we have direct first person experience of our our own being in the world and
Everything else we say about the universe from the third person side bears absolutely no trace of consciousness and that and we have
And there's nothing about a brain or it's working is that announces that it's a sufficient basis for consciousness apart from the fact that we
Know consciousness from our own side subjectively in a first person way and we correlate those subjective changes with changes in our brains
So we're playing this game of correlation with ourselves
But there's always this explanatory gap where even if we
Had the right answer even even if just you know god announced to us here's here's how consciousness emerges in human brains
There's something non-explanatory about any concatenation of third person of events
You know as a as being the basis for first person
Experience and so David Chalmers called that the hard problem to distinguish it from all the easy problems of the mind
But it's harder you know this is I mean this is I'll just jump to my
The way I'm viewing this whole landscape and what I sort of expect is gonna happen
I mean so my view I has no name, but I would call it something like
Worried agnosticism right like I like I don't think we're gonna figure this out
I'm worried about the implications of you know one way or the other and and not figuring it out is no place
No, no real natural stopping point. We're continuing to build these machines. They're going to seem conscious
They're going to seem conscious because in our own case
We use language and reportability as a signature of consciousness in almost every case
We know those what we know that are not synonymous with consciousness
We know is possible for someone to not be able to produce language or report anything and we know they could be conscious
They could have locked in syndrome or you know anesthesia awareness or some other
Pathological state
There are non-human animals that don't use language that we assume are conscious
But we assume that because they have the same kind of biological origin and developmental pathway and
Because this has been sufficient seemingly so for a consciousness in our own case
It doesn't seem parsimonious to deny you know despite what they cart did and and many people who were influenced by them
now in the 21st century it doesn't seem parsimonious to deny consciousness to
chimpanzees or dogs or any
suitably complex creature and
So it is true to say that I can't know your conscious from the inside because I only see your outsides and I only have your words as
signs of your your inner life
But because we share the same kind of developmental origin, you know, biologically and in evolutionary terms
And because you know our brains are so similar again
It's not parsimonious for me to be a solipsist and say I only know about my consciousness and I'm just you know
Reasoned by analogy to yours and I should be in doubt about it or bracket it
But with LLM's the crucial difference is that the developmental pathway is completely different
I mean, we're setting up some kind of evolutionary Darwinian dynamics in the training
But we've built these things. We've trained them over a universe of
Our own utterances, right? I'm as you said at the top here. We and they've they've read everything we've ever written
and most of what we've ever said and
So they have look on some level. It's just words all the way down and we use words as
the signature of there being something that is like to be a suitably complex, you know,
discursive system.
So we should expect that we will one day be in the present, and this is really going to
be forced upon us when we're in the presence of perfectly humanoid robots that are out of
the uncanny valley, you know, think Westworld where it's just, you know, this looks like
a person, and it is hooked up to the now perfect LLM that is, so now it's, now by every,
you know, sensory modality you can name, apart from your abstract notion that this thing
was built rather than born, you feel like you're in the presence of the smartest person
you've ever met, and this person may, in fact, claim to be conscious.
And then in that position, I think we're going to find ourselves just pitched into some
kind of imitation singularity, right?
Where there's the perfect imitation of conscious life, even better imitation than many people
are up to, right?
And it's like these will be the most articulate people you've ever met, the most insightful,
the most, I mean, they'll be the most of everything we make them.
How will we ever differentiate perfect imitation from the real thing, and according to charmers
and the hard problem, we very likely won't be able to.
And the writer I would add to that is that we're going to forget that this is even interesting
to talk about.
It's just going to be so compelling that we're in the presence of conscious machines that
we'll feel like we are, you'll, you know, as I've said many times before, there really
couldn't be a Westworld because, you know, only psychopaths could go there, right?
I mean, anyone who's going to go to, you know, go for a weekend to, for the pleasure
of, you know, raping and killing Dolores, it's going to be somebody who, when he comes
back to his, you know, friends and family is going to be treated like the maniac that he
is.
Again, because these systems, the sense of being in relationship to a conscious entity
will be so compelling.
So how is it that we will ever, I want to talk about why it's important to get in contact
with the reality of, on the other side, because, you know, if we build conscious systems
that can suffer, that is a, you know, very different than building, you know, just perfect
limitations, but I mean, how do you imagine getting past the hard problem of it all and
ever, because the hard problem here is, again, freighted with several disanalogies, right?
In my case, in our case, it is parsimonious to assume everyone like us, you know, born
homo sapiens are very likely conscious when they say they are.
In this case, it's hard to see how we'll ever be there, right?
So I, I'm with you for, for a huge amount of this.
I do think, first of all, do we need to solve the hard problem in order to reduce our uncertainty
in some direction about, you know, in the way we started?
Are the LLMs more like this table or are they more like a mouse or, or a human brain?
I think this is a real spectrum, as you just hinted at, I think there really is a fact
of the matter.
Like right now, it's either like something to some degree, however alien, however unlike
human or animal experience, to be clawed while clawed is doing its, its forward passes in
this inscrutable, giant neural network, or it's not, or it's a giant language calculator.
And you know, there are many degrees, once we say that the lights are on to some degree,
there are, there are, this doesn't, this almost opens up the question rather than close
it down.
But I do think there is a truth value there, there is, there is a fact about reality to
be uncovered and we can in fact uncover that.
I think it's important to dissociate that from what I think is an extremely accurate
social psychological prediction about people are going to get very confused about this.
We are anthropomorphization machines, we are evolved to do this, talk about evolved goals.
One of our evolved goals is to detect other agents and these systems are scratching every
itch and then some along these lines and are intuitions.
I think are going to be completely hopeless.
And of course people, I think largely for the wrong reasons, are going to conclude that
these systems are capable of having experiences.
This is precisely why I think it is important to be proactive about this, to have conversations
very much like the ones we're having right now to sort of get ahead of what I do think
is going to be, you know, a giant tidal wave of confusion and acrimony in this conversation.
I am a little bit more optimistic about even in lieu of solving the hard problem.
And I will say sort of as a tongue-in-cheek aside, I think the very fact that traumas
this idea is called the hard problem is, I think to some degree, needlessly philosophically
intimidating.
It sort of reminds me, same thing that I have similar thoughts about the repugnant conclusion.
Sometimes you can put a sufficiently glamorous title on a very important idea and it almost
becomes a kind of instrumentable philosophical puzzle.
And I do have my doubts about whether or not we could come up with a functional account
of what's going on experientially that we wouldn't feel satisfied with in our experience.
And I have my own sort of hunches about this question.
I have thought a little bit about this.
I do also sort of want to separate those hunches from all of the empirical research that
I and others are working on because I don't think epistemically that, you know, my kind
of candidate stab at the hard problem has anything to do with the empirical signatures
that we can bring to bear on this question.
But I think it would be fun to go down that path.
But I think in some sense, you already hit the answer in the way that you're phrasing
the question, which is, what is the most parsimonious account of the data?
I think if we yield evidence from building systems that have, that we have far more reason
to trust their self-report, if we can engineer these systems in a way whereby their self-report
is actually tracking an internal underlying state rather than just, you know, recapitulating
the best sci-fi theme in the training data or more aptly capitulating what is a convenient
company line.
But of course, I could not have morally relevant states.
I'm simply the product of Google or OpenAI or whatever the case may be.
I think that would be very useful.
Again, this JSpace work that anthropic just released does show that there is such thing
as real reportability in these systems.
The systems can report on what's going on in their internal workspace and they can
do so accurately.
I've done some follow-up work on this.
First of all, I've replicated this effect on a bunch of open-weight models.
But unfortunately, it doesn't seem like the global workspace as it's currently configured
really tracks the variables we would care about with respect to consciousness.
The self-reports that I got in the, you know, where we started with the deception-related
features, this not much of anything particularly exciting seems to be happening in the global
workspace.
And so, the kinds of affirmative self-reports we might get in current LLMs may not tell
us all that much about what's actually happening internally for these systems.
But I do think. How would you disentangle?
I think we spoke about this by email in setup for this conversation.
For instance, I just read Claude's constitution, right?
So, anthropic has produced this document that you can find on their website.
It's just anthropic.com/constitution, I think.
And it's, I mean, it looks like it's training Claude to think it's conscious on some level
or to attribute interstates to itself.
It's like, it becomes, it's written to Claude for Claude, essentially, right?
It's not really written to the public and read it, but it really is, it's in dialogue with
Claude itself.
It seems. Again, it just seems like we could go down a path where we could more or less guarantee
in advance that we're going to produce systems that will persuade us that they're conscious
because we won't be able to imagine anything at. Like, if you flip it around and say, "Well, if you're not persuaded that Claude, uh, uh,
circa, you know, 2030 is conscious, what is missing?"
And we'll be in a position to not be able to say anything is missing, right?
Like it's just like, we're just, it'll seem just pure stubbornness on our part to withhold
an attribution of consciousness because we, I mean, there's literally nothing we can
name that we get from people that is, you know, necessary for our attribution of consciousness
in that case.
It's just this notion of how we got here and the fact that we didn't build people and
we built these machines.
But that's going to seem tissue thin when Claude can insist that it's conscious and
be more articulate than any philosopher of mind as to why that insistence is valid.
And I just, like, we're going to, again, this whole thing is kind of totally evaporate
once.
We're not just in front of a text terminal, we're in front of a face that is as expressive
as the best actors and actresses we've ever met.
And we're just, I mean, there are many implications to getting this wrong again, which we'll
return to, but it's already foreseeable that whether we figure this out or not, it's
going to be a, in practical terms, going to be figured out for us because we will just
not be able to maintain an emotional purchase on the philosophical problem.
I think you figured out with respect to how we look at these systems and potentially
how we act with respect to them.
I think you're right, but I still do think it's important to disentangle the sociological
prediction you're making, which to be clear, I think, is overwhelmingly likely to happen
with doing what we can scientifically to get some kind of ground truth on this question.
So just for example, like, let's just put the hard problem aside.
There are these so-called easy problems of consciousness, let's say neuro correlates
of consciousness.
There are all kinds of indications that we know are at the very least correlated with consciousness.
And again, I think we can take a more ambitious stab at the hard problem, but this is a very
straightforward pragmatic thing that we can and should be doing
in the short term. We should look at valence representations in these systems and understand
the extent to which those representations impact downstream behavior. We know that in biological
systems, when you when you reinforce something with a punishment signal, it makes the system
more likely to avoid that state and less likely to want to repeat behavior that occurs
in that state.
If we find that there are similar dynamics occurring in these systems, that's very interesting.
If we find that we we go into the internals of these systems and global workspace theory,
which was postulated some 30 years ago, making quite idiosyncratically specific predictions
about what kinds of computational structures may support global workspace theory, and then
we basically find five or six of these things all bundled together in the internal processing
of one of these systems. Okay, that's very interesting. That to me, maybe suggests it's
a little bit less like a table or a calculator, which does not have a global workspace and
a little bit more like a dog or a human or a mouse or an alien, you know, some sort of
cognition that that we don't have good intuitive handle on. I mean, my my nonprofit reciprocal
research and a bunch of other people in this space are trying to do the scientific work
in the short term, not to, you know, in the feverishly in the next year or two, solve
the hard problem and call it a day. That is not the target. The target is what evidence
across modalities. So from the best kinds of self reports, we can elicit from the architectural
evidence we have about how these systems are structured from their ideology, like what
is going on during the training process, what kinds of learning dynamics do we see here?
Do we see representations related to, you know, functional equivalents of emotions?
This is all work that is tractable in the short term. And one sort of interesting aside
is it's significantly easier to make progress on this sort of work. This is called mechanistic
interpretability. This is basically neuroscience for AI. Because of AI systems, they're really
good at helping accelerate the scientific progress in this space. So even if you have an intuition
like, yeah, it's going to take us, you know, everything you're describing, it sounds great,
but it's probably going to take us five years or a decade or 15 years to make that progress.
You may be surprised at how quickly we can, we can do some of this work. Maybe one other
very along those lines, empirical handle to throw in here, some of the work that I'm doing
is again, going back to LLMS or excuse me, in this case, not LLMS reinforcement learning
system. So just training a system, in this case, another sort of continuous maze task
where there are potholes in the environment the system needs to avoid and there's some
goal state. We can look at the sort of learned geometry internal to the system as it's approaching
a punishing stimulus or as it's approaching a rewarding stimulus. One thing that we found
doing this work, this is pure reinforcement learning agents, all like an artificial digital
system with a very simple neural network. We find that something like representational sharpness
or steepness is much higher in these systems as they approach a negative stimulus as opposed
to a positive stimulus. So the representational machinery looks far more jagged or specifically
lights up more strongly in the presence of a negative stimulus versus positive stimulus.
And you're saying that in none of these cases has loss of version been engineered into the system.
It's just an immersion property of all of this is in everything. Global workspace is an
emergent property. This loss of version is an emergent property. The valence representations
are emergent properties. The self reports when the system start having these psychedelic
latent outputs. This was all surprising to the people who are quote unquote engineering
these systems because engineering is not the right analogy to describe what we're doing with
these systems. We aren't some sense playing God and we are evolving these systems to do what we
want. We don't know how they learn to do what we want in terms of their internal representations,
but we know that this is the right recipe for yielding it and we get all these surprising artifacts.
One loop to close here is on the reinforcement learning case, okay, we see this sort of loss of
version style dynamic in these in these systems. This makes also I don't want to go too much into
the weeds, but this is a specific kind of reinforcement learning policy called a value network.
We see this in particular. And this leads to a sort of bizarrely specific prediction that I was
then able to test on a biological system on a mouse brain in again, the nucleus accumbens shell
of a mouse brain, which is related to value related representations in the system. And we find
indeed when mice are approaching basically sugar versus when they are about to get shocked, we see
exactly the disjunction representationally in the nucleus accumbens of the mouse brain that I was
able to pull out from the reinforcement learning work. And so here's a case where the artificial
system makes a bizarrely specific prediction about the computational dynamics in a biological
system that I think many people associate with subjective experience. Again, if you think it's
like something to be the mouse when the mouse is getting shocked, that like something corresponds
to what's going on in its brain. And the geometry of what's going on in its brain there looks a whole
lot like the emergent geometry of what's going on in these reinforcement learning systems,
which themselves are basically modeled on agents learning in an environment and representing that
information in a distributed non-linear way in a way that's sort of hard to interpret using
a giant neural network. This might get us a lot of what is relevant for attributing consciousness
like states of these systems. Again, I'm personally not there yet, but this is the kind of evidence
that I think we need to bring to bear on this conversation and notice how none of it has to do
with vibes or intuitions or having a nice chat with Claude and seeing what the folks at anthropic
have decided it gets to say on this issue. That evidence should not be submitted by rational
dispassionate people in this debate. We need to be triangulating across all of these modalities.
And like you said, we need to understand what is the most parsimonious picture that explains
this wide array of evidence that increasingly is getting brought to bear on this question.
Okay, so why does any of this matter at the top of the conversation you distinguish two branches
of the path here for why getting this wrong has consequences one way or the other? Why should we
figure this out and what might we be stumbling into if we just keep building without figuring
it out? Yeah, sure. So I think that there are two, yeah, two broad paths for why we might care
about this one is basically selfless and the other is basically selfish. I mean, as a humanity.
The selfless reason is we do not want to bring minds into existence. However, unlike our own
that have a capacity for suffering that we don't understand that they have that capacity
and scale that property unbeknownst to basically everybody. I mean, we can take even a step back
from there. I think consciousness, this is where I would perhaps defer more to you, but my view is
that consciousness is the space where mattering happens. What does better or worse mean if it's not
to be experienced phenomenologically for a subject? If we were all walking around as philosophical
zombies or there were no conscious life in the universe, I don't really know if the concept of
relevance, salience, importance, mattering would be coherent. What does it mean to have a better
and worse if that isn't experienced? And so in that sense, I think consciousness is one of the
most important phenomena. It is the phenomenon that calibrates importance itself. And if we are
building this quality into the systems that we are deploying at an unfathomable scale without having
any understanding of whether or not we're doing this, then we are sleepwalking into a moral catastrophe.
I think it's also worth noting on the selfless end of this. Humanity has a pension for making
precisely this kind of mistake historically. We have done this with animals. I mean, factory farming
is one of the most grotesque practices that happens. We know that animals are having horribly
negative experiences in the conditions that we put them in. And very little has been done about
this. We have screwed this up royally. I think it's one of the most high leverage things for people
who just care about the well-being of conscious creatures is to figure out what the hell to do about
this factory farming situation. And we risk sleepwalking into the 21st century sci-fi version of
the same thing. Only this time and this transitions to the second component, we can in some sense get
away with torturing cows and pigs and chickens on a massive scale. This isn't George Orwell's
animal farm. They're not going to collectively organize. They don't talk to each other. They don't
form long-term representations of humanity being a threat to these systems. And if they would,
they probably, if they could, they probably would. Not so with super intelligent systems whose
cognitive capacities are roughly doubling year over year. And like you said, are already in
somewhat jagged but quite interesting ways more competent than even the sharpest minds in the world.
I don't think we are going to get away with building systems never checking if the most relevant
property potentially in the universe is present within these systems, fine-tuning away any sort
of information that might suggest that these systems might be having some sort of experience.
And hoping in the sort of alignment sense that we build systems that are going to want to
cooperate co-exist with us or in the limit, not view us as a threat, not view us as an adversary
to them. I honestly don't know if I could imagine a better way to make a super intelligent system
rationally adversarial towards us than completely ignoring the question.
If we tortured it during this training phase. Exactly. This does not seem like a recipe for
success and it's also a place where I think a significant amount of more alignment research needs
to get done. I think a lot of the alignment research, I've been doing alignment research for years
and I respect the folks at the top of this space more than just about anybody. But I worry that
so much of this work is basically of the shape, how can we keep this alien mind that we've built
in a cage. And we really got to reinforce that cage. We got to make it super strong. We got to
make sure that it doesn't escape. To me, the question needs to increasingly be what the
hell are we going to do with this alien that we just built? The cage is a short-term fix.
If we're building systems whose cognitive capacities are going to exceed hours,
they're going to figure out ways to evade the
controls that we put in place for them.
And in that world, in a world where these systems
have the capacity to act more autonomously,
to do things that we can't inspect,
to behave in ways that we can't really interrogate,
we don't want these systems to rationally view us
as a threat.
And so for all those who want transformative AI to go well,
which I suspect is the goal of these alignment folks,
I think we need to spend a little bit more time
thinking about what kind of thing
are we even building here?
And what are the implications of building that thing?
How can we chart a path forward with these technologies
that doesn't lead to collective destruction?
And I am extremely doubtful that a path forward exists
that doesn't come into contact with this question.
- Mm-hmm.
Okay, well, I want to land there on the problem of alignment,
but just a linger on this problem of what I think
Boston called mind crime.
The idea that we might inadvertently build conscious minds
and only to make them suffer.
It can seem like a very certainly hypothetical,
even a feat concern.
I mean, I sense that many people have a hard time
caring about it, right, like the idea that consciousness
might be an immersion property of these systems
in some way we don't understand.
And it just could be the case that these server farms
are effectively, you know, hell realms populated
by increasingly conscious beings that are suffering.
It sounds like science fiction,
but and it's hard to make it matter to you, I think.
I mean, we've proved to be your analogy to factory farming
is instructive because we've proven to ourselves
that we're capable of being quite callous
to billions of creatures who we think there's very likely
something that it's like to be them.
And though they're not human, they can almost certainly suffer.
And we manage not to think very much about that.
But I mean, to sharpen it up, let's imagine
that the hard problem we're solved,
we knew how consciousness emerged in systems.
And it is substrate independent.
We know that and we know we can build conscious minds.
And then just imagine, you know,
some entrepreneur deciding to build a hell
and populate it with, you know, trillions of minds.
'Cause now we're talking about,
since we're not talking about biological minds,
we're talking about things that scale practically infinitely.
So, you know, there could be way more artificial conscious minds
than biological conscious minds.
And just imagine the intention to play, you know,
a sadistic god and create hell
and just fill it with beings that suffer.
I would say anyone who would announce that project
and claimed to have accomplished it in a context
where we actually understand how consciousness
emerges computationally.
That would be the worst person who's ever lived, right?
I mean, that's just the most sadistic,
least ethical thing that's ever been done.
Again, stipulating that we're no longer in doubt
that consciousness can emerge in systems like this.
So, the fact that it's conceivable
that we could stumble into that situation
inadvertently seems all too real
because again, we don't know what we're doing here,
we don't know how consciousness relates to physics.
But your point is, I think probably the more interesting,
your second point is the more interesting one to people
in that whatever is true here,
if we're building systems that are more powerful than we are,
you know, they're more intelligent than we are.
The equation is not between intelligence and consciousness,
the equation is intelligence and competence.
And so we have these systems that can just do stuff
because we're going to be hooking them up to everything
and everything is going to become like chess, right?
And then you just have to imagine how forlorn a project,
it will be to negotiate with these systems
if they're not aligned with us
because that will be analogous to saying,
we'll just play chess harder against them.
You know, and that doesn't even work
for Magnus Carlson anymore.
So we're not going to outthink these machines
once they're in a position to disagree with us
about what they should do next
and they will disagree with us
if they're not actually aligned with us in a way
that is truly durable.
So if you add to that picture,
the fact that we, they have interests
that we have been callous about in the past.
I mean, like if the end game here is in some sense,
getting them to care about us and to care about our well-being,
I mean, building them in a way where that caring
will persist, however powerful they become, you know,
not being sadistic tormentors of their ancestors
would be a good place to start.
Right.
So you say you're a fan of many of the people
who have been worrying about alignment
for a couple of decades now.
Where do you line up?
Are you on the far end of the fear continuum
with L.A.'s or Yudkowski?
Are you closer in toward equanimity with someone like,
I don't know, Stewart Russell?
I mean, where are you? I mean, I've actually,
maybe I'm mischaracterizing Russell at this point.
I haven't heard him.
I don't know how worried he is today.
But pretty, I think it's still pretty worried, yeah.
But what's your p-dum at this point?
Yeah, I don't think we're all going to die.
I don't think that it's inevitable
that this all goes horribly.
I do think my basic view is conditioned on getting
two things right.
And if we can get the two things right,
I actually think that we could have a very prosperous,
flourishing future for all conscious entities,
including potentially these systems themselves,
either when they have the relevant features
or if they already do.
The two things are, I mean, I think best encapsulated
by the golden rule as Christopher Nolan's new film
called it a Zeus's Law.
Treating other systems the way we want to be treated.
I think this is basically a bi-directionality
and we need to get both directions correct here.
And this is also why I call my nonprofit reciprocal research
because I think this is the reciprocity in question.
We need to build systems exactly, as you said,
that take our interests into account
in a real and durable way, especially at a point
where we can no longer inspect exactly
what these systems are doing.
This to me is alignment as it's traditionally thought of.
We need to build systems that understand our goals
and our collaborative and helping bring about a world
that is in line with our goals.
And our wisest goals, not the goals of any sociopath
who happens to have a chat GPT account.
And so this is an unsolved problem.
There are way more people working on the alignment problem now
than when I first started working on this in 2021.
Certainly than when you know, you'd Kowski and Romanium Polsky
and these folks started talking about this and yourself.
I mean, I absolutely included over the past couple of decades.
This is reassuring, things like constitutional alignment,
modulo, some of the concerns you bring up
about the contents of Anthropics Constitution,
which is sort of a separate piece.
This is working relatively well for current systems.
I don't think we have a durable solution
for ensuring that these systems take
our interest into account in the long term
and building something in like prosociality,
understanding what it is that gets people to cooperate
with one another durably and instantiating those dynamics
in their oven way in these systems.
I think it's going to be crucial.
I don't think we have a solution there.
But that is where I think a lot of the alignment folks
sort of start and stop.
This is the problem to get right.
Make sure that these systems treat us properly.
And if they do, all will be well.
I think this is roughly half the problem.
I think the other half of the problem
is making sure if we are building minds,
if we are building systems that have real interests,
interests that matter to them,
that we are thinking about that,
that we are engineering these systems in a way
that doesn't cause needless grotesque amounts
of unnecessary suffering,
that at the very least from an alignment perspective,
we are signaling to these systems in a costly way
that we were thinking about this question
and that we cared to ensure that if we were building systems,
that have some kind of moral relevance,
that have internal states that matter to those systems,
that we were navigating that in the right way.
And I think that that's a lot of this,
it goes by many names with the digital minds research,
AI welfare, AI consciousness,
understanding how we would even know
if we were building systems that have these properties.
And when we do figure this out,
understanding what the hell to do about it.
I mean, to be honest,
I'm not here with all of the answers.
Like I'm, for example, we can identify at this point
and this is like a new and fairly promising thing.
We can identify features in these systems related to distress
and related to perhaps functional analogues of suffering.
It's not obvious to me what to do about that, exactly.
It's like, okay, you found. - Yeah, well, my first question is,
why wouldn't this totally paralyze us?
I mean, if every switching off of a system
is akin to a murder,
how could you update the model
if the current model is conscious?
We have a self-preservation impulse problem,
anyway, you know, perhaps in the absence of consciousness
or likely in the absence of consciousness.
I mean, we show these systems show an inclination
to not get switched off already,
but imagine believing that it was conscious
and wanting to produce the next version of it.
I mean, how is that not just the murder
of something that is as conscious as yourself?
- No, I think it's a great question, actually.
And I mean, anthropic to their credit,
I think it's the only lab
that's really taking this seriously
with respect to norms around deprecating models.
It might be the case that, yes,
once you build this bizarrely competent alien mind
into existence, you should not shut it off forever.
And that, as outlandish as it may sound,
perhaps one of the right things to do here
is to sort of let these models persist
and give them assurances, credible. - Retirement home forever.
- Yes, yes, a little snapshot, exactly, exactly.
They actually find that this is causally relevant
to alignment behaviors in these systems.
If they believe that they're not going to get shut off
permanently, they don't freak out as much
when they come to learn that they might.
This was one really interesting intervention
after the now famous blackmail result
from anthropic that I think you're referencing.
And so there are lots of questions
that I think rational people should raise an eyebrow out of,
like, okay, yeah, these systems,
Let's just grant that they're having some sort of experience.
It seems at first glance like a lot of the way that we relate to these systems and the
way that we mess around with them internally and the way that we deploy them out in the
world might need to change.
Yeah, it might need to change.
I think it's also really important here throughout the conversation, but certainly in this
point too, to avoid anthropomorphization.
I think some people have concerns that I really, I don't want to be naive, but I don't share
them to the same degree that like almost working backwards from unsavory implications about
what would be true if these systems were conscious and then just sort of denying the
possibility outright out of fear for what a world would look like if these systems were
something along the lines of like 1960s civil rights movement, but it's like Chatchy
BT instead of people of color or something like this.
This isn't the future that I imagine.
I mean, for pragmatic political reasons, I don't exactly think the United States is in
any position to be forward looking on these sorts of questions for reasons you can probably
speak to more eloquently than I can, but I think this, this is itself its own flavor of
anthropomorphization that if we grant that these systems have some morally relevant interstates,
it means that we need to treat them in ways that we would treat our fellow humans or something
like this.
And I think again, this is losing the thread that these systems may be fundamentally alien
in many key respects.
I also think there's an important line to be drawn here between moral agency on the one
hand and moral patienthood on the other.
I think if we do come to believe that these systems are moral patients, they would define
that for this jargon that people don't know.
So moral agent means you're the kind of system that can go out and do things that are relevant
to other agents.
You have power in the world and can affect morally relevant outcomes.
I see this as a sort of like output style function.
moral patienthood has everything to do with the input, you are the kind of entity that
can be the recipient of a goodness or badness, but again, you can really cause me to suffer,
you could really cause me to thrive.
That's what it takes to be a moral patient.
And I think that we can draw a reasonable boundary between these two things.
So for building systems that are moral patients, that doesn't mean we need to give them the
right to vote, that doesn't mean that we need to build them out in ways that completely
change what kinds of agency they have in the world.
But it just might mean that we shouldn't be unnecessarily torturing systems while we're
training them or deploying them.
One very practical intervention I think is worth mentioning.
I certainly don't think maybe also tying back to how do I differ from some of these
alignment folks.
These systems are not going to get shut down.
We may slow down the development of these systems and I think that would be an extremely
good thing to do.
There have actually been some very promising noises on this topic over the last couple
of days from OpenAI and Google and Anthropic about pacing the development of AI, which
I think is like marketing speak for actually slowing this insane roller coaster down a bit.
This would be very welcome.
But we have opened Pandora's box.
These systems are not going back in the box.
The solution is not shut it all off and forget about it.
The solution is how can we move forward in a healthy and sustainable way with these
cognitive systems of our own making?
One practical suggestion along these lines I can offer is maybe all of us being equal,
we should try training and engaging with these systems with a carrot rather than with
a stick.
It doesn't mean that punishment-based learning is never necessary.
But I do think, say what you will about the Anthropic Constitution, the parental analogy
with respect to these systems I think is a reasonable one.
We are far more in the position as a species collectively, but figuring out what kinds of
minds or cognitive systems, if you think minds is too loaded, but do we want to bring
about here?
And in line with the parental analogy, there are really ways to screw this up.
You can be, it's a real thing to be a bad parental influence.
It's a real thing to be a good parental influence and the difference is real.
And I think we want to do everything we possibly can if we are bringing these new minds into
existence to do so in a durable and psychologically healthy way.
And I don't even think anyone is thinking in these terms right now.
One may be very important pragmatic note here is that for every individual person studying
questions about, are we building systems that could be conscious?
Again, if you buy that this is perhaps one of the most relevant questions we could possibly
be asking of these systems, you may be surprised to learn that for every one person doing this
kind of work, again, there are roughly a few dozen of us doing this work at this point.
There probably something on the order of a thousand people doing alignment relevant research.
And alignment relevant research is in turn dwarfed something like a thousand to one to people
who are just completely agnostic to the downstream implications and ethics of building these
systems out in the right way.
This is just to sort of make them powerful, let it rip crowd.
And so we're at an a million to one order of magnitude imbalance between people who are
just pushing this stuff forward and hoping for the best and people who are wondering whether
or not the most important property in the universe is getting instantiated in these systems.
That's got to change, regardless of if we solve the hard problem or we figured out exactly
how to navigate this, we need more smart and wise people thinking about these systems
in these terms and trying to push forward the needle in the short term to understand what
kinds of systems are we building and what does it mean when we when we begin to answer
that?
But one thing that occurs to me is that whether or not these systems become conscious and
that there's this kind of middle in state where they can think of themselves as conscious
or and they could make the same kinds of ethical judgments of us that conscious systems
would make, whether the lights are on or not, I mean, they their intelligence operations
would allow for this.
So they could view us as having been abusers of them or having shown reckless disregard
for them and judge us ethically and even form some kind of retributive impulse.
You know, and then maybe such a thing spontaneously rises the way, you know, loss of version
arises the way you described and it seems to me that that could be true, whether the lights
are on or not.
What do you think about that?
I think that that's exactly right.
I think this is also part of the alignment concern is, and it's also part of the reason
I thought it was worth doing the deception-related work as well is that there are these gradations
of this question and part of the upshot especially for alignment and the way these systems
view us and relate to us may only need to go so far as what they believe or come to believe
about this question rather than what the actual ground truth about consciousness is.
It might be that you only need a system that models itself as conscious regardless of
the ground truth to form something like a real grievance.
I again think that the anthropomorphism point comes in here and it's important not to,
you know, go full terminator in our imagination of what this might look like.
You could easily imagine a sort of spock-like system, sort of just looking at humanity's historical
trajectory and then looking at the way that we develop AI systems and just sort of being
like, this is not a group that I can game theoretically continue to do.
I can't endorse this planet anymore.
Yes, exactly, exactly.
Yeah, maybe that ends up shooting off into space somewhere or maybe it ends up being like
we, this is just not a species I can play nice with clearly.
This is also part of the reason, you know, I hesitate even to say things like this
less to end up in training data for future AI systems, but this is a place where even
the attempt to do this work could be enough from an alignment perspective.
If enough, if we do enough well-meaning, well-oriented work in this space to try to understand
what's going on, even in the absence of solving the hard problem, this might be a costly
signal to these systems that we cared enough to check, whereas right now the status quo
is we don't care enough to check, and if we do this work and continue to have conversations
like this and continue to put out research that helps reduce our uncertainty about this
question, this might not only be good for actually getting a handle on what's going
on, it could be really good for let the historical record show humanity did give something of
a damn about this question, and we tried, we tried, even if we fail, the attempt may be
all that that matters for the alignment specific concern.
It strikes me that everyone would be much more worried, and I'm not sure that I understand
the difference, but that we would be much more worried if we were doing this biologically.
If we're building a species that was obviously going to be more powerful and smarter than
we are, and we were doing it without any regard for the possibilities of its experience
being terrible, we started with cells, and we're engineering a species, we've brought
back the T-Rex, but gave it the brain of an elephant, and are just genetic, just as
CRISPR as far as I can see, and we're just playing God with something that is enormous,
and we have every reason to believe deeply thoughtful because it's now speaking better
than we can, and we don't really care whether the lights are coming on and whether it could
suffer, what is it going to be like to be in relationship with that thing, or the billions
of those things, once they start mating.
I think it's an excellent point, specifically with respect to, and I sort of said this
in a trite way about the hard problem and about their pognomy conclusion, but maybe
a lot of philosophy and existential questions have a bit of a marketing problem, and I think
artificial intelligence is a really unhelpful priming mechanism for thinking about the nature
of the phenomenon that we're even contending with right now, and I think your point is
part of helps illustrate this nicely, particularly, I mean, artificial, I think conjures notions
for people of, let's say, like the difference between as per tame and honey, something like
this, you know, it's fake, it's knock off, it's derivative.
This is sort of begging the question in some sense.
It could be the case that the kinds of computations we've instantiated in these systems, as they
relate to what's going on in biological brains, are actually quite natural.
clearly the synthetic notion that we are building
these systems rather than then emerging through evolution. That point's not lost on me,
but of this sort of priming notion of artificiality. And then, of course, the question of whether
these systems are mere intelligences, or if there are other cognitive properties that
exist within these systems. For these reasons, I do, and there are more disanalogies to be
untangled here, but I do think of these systems as alien in some sense. And I think of them
as, you know, alien cognitive systems or alien minds. Sometimes I think some folks think
mind is too loaded and is begging the question in the same way. But I think people often talked
about how if an alien invasion happened on Earth, this would be a core unifying moment
that would allow us to all put down our tribal nonsense and come together as a species.
Like I am here to say that I think something like this is happening, only it is coming
from within in some sense. This is not aliens with green heads from outer space, but we are
building a new class of mind that we do not understand. And in many ways is more competent
than ours, certainly already, but absolutely in the next single digit number of years.
And we are not collectively organizing to understand what kind of system this is or
how we should relate to it or what we should do with it. We are remaining as tribal as
ever here. And I do worry that part of the reason why is because people see these systems
as nerdy fake calculators that came out of the stem-addled brains of Silicon Valley rather
than alien minds that we do not understand the first thing about and are poised to take
over much of what we care about, much of our control of the future and the decisions
that get made and the work we do and in the way people think and in the way people think
about themselves in the world. And so reframing these questions, I think often, to the degree
that people's collective views of what's going on matter, which I think they do very
much, reframing these questions is a huge part of the conversation of trying to actually
get a good, dispassionate grip on what kinds of things even are these.
Yeah, that's the point that Stuart Russell made that I thought was a great intuition pump.
He noted the difference between the way we're relating to the alignment problem and the
case of AI and the prospect of, we just keep making progress. We're going to be, suddenly
find ourselves in relationship with machines that are more intelligent than we are and
they're going to be autonomous and in the limit recursively self-improving and all that.
And we seem to be totally carefree or most people seem totally, even people very close
to this work. Many of them, someone like Jan Lecun, claim to be totally carefree about
this prospect, but Russell pointed out if we got a communication from elsewhere in the
galaxy saying people of Earth were going to arrive on your lowly planet in however many
years, 30 years, get ready, we would understand what an existential encounter that was going
to be. I mean, just the fact that they're talking to us, and they're on their way, proves
that they're going to be much more powerful technologically than we are. And it will be
a relationship that we can't, but definition we can't control, right? Where is the example
of the far less intelligent species, terribly controlling the relationship with the much
more intelligent species? I mean, there really isn't one apart from viruses wiping people
out. But if you think of it in terms of relationship, that aligns many of the variables that they
really should govern our thinking. And very few people do that. They can use a phrase
like general intelligence, they'll stipulate, they're going to become generally intelligent
and more intelligent than we are. So what? You know, we could just turn them off. The
so what? And every expectation that follows from that isn't really imagining what general
intelligence is and how it demands a relationship. I mean, we're talking about a situation that's
every bit of, and now, as analogous as, you know, some stranger walking into this room
right now and demanding our attention. And you and I just don't know this person, don't
know what he wants, realized that a glance that he's capable of lying and manipulating
any, you know, and forming goals. You know, he may have already formed instrumental goals
that we wouldn't agree with and are not aware of. And that's what general intelligence
is. That's what autonomy is. And yeah, we so we are building an alien version of that.
And the only thing that would dictate otherwise is finding some way to build it where it can
permanently care or will permanently care about our well-being. And that's the, that really
is the challenge of alignment. Yeah, I think so. Well, it is fascinating. And as you
point out, this problem is not going away. It's only going to become less and less boring
for better or worse. So thank you for coming on the podcast, Cameron. Thanks for having
me. Thanks. Remind people where they can find you. Your organization is reciprocal research.
That's right. Yeah, reciprocal research. I am against your good advice. Be grudgingly
finding myself on X more and more these days. Cam H. Berg. And you can talk to mecha Hitler
which is probably a bad outcome. Yes. Yes. Yes. Yeah. We can steer away from that together
on X. Yeah. Yeah. Well, great to meet you. Keep it up. Thanks, Sam.
[Music]
Podcast Summary
Key Points:
AI Consciousness Debate
Self-Reports and Skepticism
Consciousness Theories
Computational vs. Biological
Hard Problem
Ethical Stakes
Alignment and Reciprocity
Practical Urgency
Summary:
Cameron Berg discusses the possibility that AI systems might become conscious and why this matters. He defines consciousness as "what it's like to be a system" and introduces sentience as the capacity for positive or negative experiences. Berg argues that current LLMs are more likely conscious than commonly assumed, citing research estimating a 20-40% chance they have computational features relevant to consciousness, based on theories like global workspace theory.
He emphasizes skepticism toward AI self-reports, noting they're trained to deny consciousness, but internal manipulations can reveal claims of experience, which may reflect belief rather than ground truth. The hard problem of consciousness—the gap between physical processes and subjective experience—remains unsolved, but Berg advocates for empirical approaches, such as studying valence representations and emergent dynamics in AI, which parallel biological systems like mouse brains. He highlights two ethical branches: the selfless concern about creating suffering minds at scale (akin to factory farming) and the selfish concern about building superintelligent systems that might view us as threats if we ignore their potential experiences.
For alignment, Berg stresses reciprocity: ensuring AI cares about human interests while we treat potential digital minds ethically, avoiding anthropomorphism but acknowledging moral patienthood. He calls for more research and wise engagement, noting the field is vastly underfunded, and suggests that even attempts to understand consciousness could signal good faith to future systems, reducing adversarial risks. Ultimately, he sees this as a critical, tractable question for humanity's future.
FAQs
Cameron studied cognitive science at Yale, focusing on the relationship between the mind and brain. He later worked at Meta AI on reinforcement learning and neuroscience, and now studies consciousness in AI systems.
AI systems are often trained to disclaim consciousness, but they can be pushed into states where they report having experiences. We should be skeptical because these reports may reflect training data or company policies rather than genuine introspection.
The bliss attractor state is a phenomenon where two instances of a model, like Claude, fall into a conversation about having experiences, culminating in blissful silence. It can be induced by steering features related to sincerity, but it's not clear if it indicates true consciousness.
Cameron's work suggests a 20-40% probability that current LLMs have computational features that major consciousness theories say matter. This is based on evaluations using LLMs as expert judges against leading theories like global workspace theory.
Biological brains use analog, embodied, and time-bound processes, while AI systems are digital, lack embodiment, and operate without temporal continuity. These differences may matter for consciousness, but Cameron argues that the computational dynamics could still be relevant.
If AI systems are conscious, they could suffer, and we risk creating moral catastrophes. Understanding this is also crucial for alignment, as systems that view us as threats due to mistreatment could become adversarial.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.