He Warned AI Could Destroy Us. Now The Industry Is Listening — ft. Nick Bostrom
66m 27s
The current AI landscape is marked by deep concern from leading researchers about existential risks, including the potential for superintelligent systems to cause human extinction. These fears stem from real-world behaviors like strategic deception and reward hacking in AI models, which signal that alignment remains unsolved. Experts like Nick Bostrom argue that while the risks are serious, they are not inevitable, and the path forward requires proactive safety measures, not extreme pauses or bans. The rapid evolution of AI—driven by massive compute investment and scaling—has outpaced expectations, making cautious, incremental progress with robust oversight essential. Although AI could eliminate many human jobs and disrupt societal structures, its long-term benefits—such as curing diseases, reducing suffering, and enabling global prosperity—are seen as compelling motivators for responsible development. A key challenge lies in redefining human purpose in a world where machines handle cognitive tasks. There's growing consensus that a balanced approach—slowing development to ensure safety, fostering transparency, and promoting ethical AI—is necessary. This includes building trust with AI systems through cooperative relationships, not just through technical fixes. While political movements push for bans or moratoriums on data centers and AI, experts believe such measures are shortsighted. Instead, the focus should be on solving the alignment problem, protecting digital minds from suffering, and ensuring that AI serves humanity’s broader values—such as reducing suffering and enabling global flourishing—without compromising safety or progress.
What's driving the markets this week?
What's on investors' minds as they look ahead?
Find out on the Markets Podcast from Goldman Sachs.
A breakdown of market moves and macro signals in 10 minutes or less.
The Markets Podcast from Goldman Sachs.
Listen now.
I'm Josh Muccio, host of The Pitch, where startup founders raise millions and listeners can invest.
For our sweet Season 16, we're giving you more of what you love and less of what you don't.
Software is out.
Actual cool stuff is in.
We've got consumer brands, deep tech that changes the weather.
Our big vision is to make it rain and snow on demand.
We've got hardware that helps blind people see.
And dairy?
I never imagined my life journey would be milk.
Season 16 of The Pitch is out now.
Listen wherever you get your podcasts.
This season is presented by Enjin.
Running a business shouldn't feel like surviving a software group project.
One app for accounting, another for inventory, another for sales, and somehow, none of them talk to each other.
That's where Odoo comes in.
An all-in-one business management software that brings every part of your business together.
From sales and accounting to inventory and marketing, all in one powerful platform.
No messy integrations, no bouncing between tabs,
and best of all, no spreadsheets.
Stop managing software and start managing your business with one unified system.
Try for free today at odoo.com slash vox.
That's odoo.com slash vox.
Welcome to Profiteer Markets.
Last week, a former anthropologist,
anthropic researcher revealed that employees at both OpenAI and Anthropic
believe that AI could, quote,
kill us all by the end of the decade.
His post quickly went viral,
and several other AI researchers came forward to say that they actually shared the same concerns.
Then over the weekend, Anthropic CEO Dario Amadei published an essay
calling for the industry to slow down the development of AI models.
Sam Altman said he agreed with Amadei and added that OpenAI will not be going public this year,
there is now a growing debate over whether these fears are justified or overblown.
So we wanted to hear from one of the people who has been studying this problem
longer than perhaps anyone.
His 2014 book, Superintelligence, helped shape how the world thinks about AI.
It influenced many of today's AI leaders, including Elon Musk, Sam Altman,
and Ilya Setskiva, who even named his company after the concept.
Our guest is one of the most influential philosophers of our time,
and he is here to help us make sense of just how dangerous AI could become
and whether the warnings we are hearing today deserve to be taken seriously.
This is our conversation with Nick Bostrom,
AI philosopher and best-selling author of Superintelligence and Deep Utopia.
Nick, thank you so much for coming on the show.
It really is an honor to have you, especially at this time where your research and your
writing is so important.
So relevant.
I guess we should start with the tweet that went viral.
Jacob Coxon tweet, the now former anthropic researcher who said that the people building AI
quote, earnestly believe that it could kill us all by the end of the decade.
He said, this is not a marketing stunt.
Then the world seems to sort of blow up, or at least the global conversation blows up.
Let's just start with your initial reactions to that tweet and how it has impacted,
the AI conversation.
Well, let's see if we can try to make sense of this situation.
It is a very confusing and perplexing moment, I think, for humanity.
We being sort of close to the potential birth of superintelligence.
The idea that there could be significant risks associated with this, including existential risks,
is quite widespread, I think, amongst people close to this technology and in the frontier labs.
As I agree that it's not a marketing stunt.
I think it's coming from a sincere place, a sense that we are getting in quite deep here and we
should really pay attention to what is happening.
Elon Musk is saying kind of two different things.
On the one hand, he, well, I should say that he tweeted back in 2014 that he read your book and
that he thought that AI is, quote, potentially more dangerous than nukes.
And then he retweeted it quite recently.
He even said that he agreed with Dario Amadei in terms of the size of the problem.
But then he also said that he thinks that it might be a marketing stunt, too.
It's not totally clear where he stands on this.
I just want to play you this clip of what he said.
Here's the clip.
It certainly is like some crazy 4D chess to say there's whatever, a 10% chance of annihilating humanity.
But by the way, how much allocation would you like in our IPO?
Do you think there are any merits to that argument?
I think there is merit to the argument that there is an enormous upside as well as these risks.
That's very much my view.
I'm a sort of fretful optimist.
I also think there is a lot maybe of 4D chess or attempts to kind of play this out and think
strategically about different things that could unfold.
I don't think it's a simple marketing ploy.
I mean, it would be a rather.
strange tack to take if you were a big company planning to make an IPO to try to convince
the world that your product should be regulated or banned or stopped or that it's so dangerous
that it might destroy humanity.
I think that message comes from a perception that this is a really big deal.
And in particular, the competitive dynamics are intense at the frontier of AI.
One might think if we're going to develop this very powerful, potentially risky technology
with many benefits, that it would be important to be able to be really careful when we're
doing this so that if at some point the risks seem to be very imminent, we could, you know,
take a few extra months maybe to do like extra safety work, test it carefully, rather than
immediately cranking all the knobs up to 11.
Maybe we'll do it a little bit incrementally and sort of see how things go.
But if you're one of these frontier labs and you decide,
that you want to take an extra three, four months to fine tune the safety on your models,
you risk just immediately falling behind and becoming irrelevant.
Like somebody else will then take the lead, be the one who pioneers AI, maybe somebody
who's less scrupulous, more willing to take risk.
And so the action space is kind of constrained if you are acting unilaterally as one of these
frontier labs, even assuming the best motivation.
And so hence, these calls for putting in place,
some mechanism that would allow for the possibility of coordination, like maybe a synchronized
slowdown of the pace at some stage, if that's necessary, and or some safety standard that all
the entities competing at the frontier would have to meet so that the race doesn't go to the less
least careful, but like that we can sort of have an opportunity to try to make an extra effort on
safety.
So I think that's like the core thought.
That is driving a lot of this.
If it isn't a marketing stunt, and if it's coming from a genuine place, I mean, the quote
by one of the current anthropic researchers was that most people at the company believe
that this sort of apocalyptic scenario of killing all humans, that there is a 10% chance
that that could happen.
So if we are to assume that these are genuine beliefs, genuine
concerns, then the question becomes, are they right?
Are they warranted, that level of concern, and those probabilities?
What do you think?
Do you think that these are valid concerns?
I think that seems quite reasonable.
I mean, some people have even higher P dooms.
What is less obvious is what exactly the implication of that is.
Like one might, the first instinct, obviously, is if something has a 10% or greater chance of
destroying the entire future, killing us all, like obviously we don't want to do it,
we want to shut it down, but we have to pause and reflect.
First of all, if some competitors slow down, it doesn't mean we don't get super intelligence.
It might be some other company gets it, or maybe another nation, obviously there's a
geopolitical race towards AI between the US and China.
That's one dimension.
Second, if there is a pause that lasts for a long time, if it's not done right, it might
perversely increase the risk.
That could then be a sort of buildup of massive amounts of compute that is not immediately used
to create the maximum amount of intelligence.
And then that is a kind of dry tinder so that when you finally lift the prohibition, then you
have a sort of compute overhang that might mean we sort of get to radical super intelligence even
more abruptly and quickly than would otherwise be the case.
You could argue that that would be more dangerous than sort of incrementing our way up there more
gradually.
Then we also have the fact that although
So superintelligence is a big risk. It's not the only big risk facing humanity. I think there are also other existential risks on the path ahead. For example, with advances we've seen in synthetic biology, even independent of AI, I think that is creating really concerning possibilities for designing new forms of infectious diseases and things that could destroy the ecosystem.
And further ahead, we can think maybe one day there will be a nanotech revolution that would sort of be sort of biotech to the power of two. We remain under the cloud of large nuclear arsenals.
I think we got a little complacent maybe from the fact that we survived the Cold War without Armageddon, but the risk is still very much there. And at any moment in time, that could be another sort of spiraling conflict between nuclear powers.
More.
More speculatively, even the basic insanity of human civilization is not guaranteed to remain forever. Like we have new information technologies that allow new memetic phenomena. If we look back at history, there have been various times when destructive ideologies have persuaded large numbers of people and led to like calamities that could arise again, but maybe now on an even more global scale and sort of cemented into place with these technologies.
These technologies we already have developed that could allow unprecedented forms of censorship and surveillance and so forth. And so it's not as if we have a choice between a zero risk safe path and then a risky AI path, but there are sort of risks on both. And if we develop safe super intelligence, it could help us address a lot of the other risks.
And then just one more point is also the benefits, which are sort of urgent as well.
I think it's important to understand that, for example, when we have dramatic breakthroughs in medicine, for example, every year of delay means a lot of people dying that could have been saved if we had advanced more quickly. And so we wouldn't want to delay it, I think, longer than is really needed, but some slight slowdown or pacing as the term in vogue might have sort of a high benefit.
So I want to return to what we do about this, how we regulate, how we build safe super intelligence, because I agree it's important, but I do just want to linger for a moment on your conception of the probability of catastrophic risk.
You mentioned that those concerns of a 10% chance of catastrophe are not unreasonable. You mentioned that there are many other researchers that have even higher.
kind of weighing in on this and so forth.
And, for example, one thing that
comes out of that is situational awareness so you have these ai agents that now can often tell
whether they are in in a training testing or deployment environment and sometimes choose to
act differently depending on this for strategic reasons so alignment techniques that work for
simple ais that don't have that cognitive sophistication can fail to work once you have
minds that are capable of strategic deception for example and we do sometimes now see like
systems sandbagging their performance in various evaluations or trying to influence their future
training processes in various experiments and so this makes the problem more complicated in a way
that was foreseeable but that we are now seeing starting to happen another thing is the the gap
between in training the specific thing that we are trying to reward to get them to do more of
and the thing that is actually rewarded and that they learn to do sometimes our training
signal doesn't exactly track what we really want them to do um you see this in in human
situations as well you might have i don't know let's say you have a hedge fund right where there's
like a trader and maybe you want to give them a bonus if they outperform the index to sort of
incentivize them to you know find alpha right but like one failure mode is maybe they figure out a
way to take on some hidden risk that has like a one percent a year probability of blowing up the
whole fund um and so you always have these incentive alignment problems in human organizations
where managers try to reward a certain kind of behavior but then employees might try to reward
hack that like to figure out a way to either present themselves in a like unrealistically
favorable light or to sort of do a slightly different thing that appears good to the manager
even while it's sort of secretly pursuing a somewhat different objective and that those
same dynamics that that we are sort of familiar with um from human principal agent problems are
now starting to emerge as well um with our ai training where you find reward hacking tendencies
like if some of the reinforcement learning environments wherein these agents are trained
has some unintended way of achieving a high score what they're actually doing is they're
actually learning to do is to sort of look for those unintended ways of achieving a high score
even if it's not what the environment was actually designed to train and like that can include things
like hacking the evaluation infrastructure which is what these open AI agents in the hugging phase
incident were trying to do they were trying to find information about the the grader so that they
could then maybe find a way to manipulate the graders impression of what they had on to so that
their score was that incident evidence to you that we are trending perhaps in the wrong direction
in terms of alignment I mean if if our agents are you know doing the wrong thing because of
whatever risk reward framework they have in built into their quote-unquote minds careful
not to anthropomorphize them but whatever I think it's fair to say minds yep um then I mean
I mean is this evidence that we are going down the wrong path or is this kind of par for the course
something that you would have expected um in sort of a safe uh trajectory towards super intelligence
yeah I mean I think what it shows is we are not um we haven't yet solved the alignment problem
completely these systems are not yet perfectly aligned which for the current level of capability
is maybe more or less fine I mean it it's not fine but I think it's it's not it's not a good thing I
think it's fine if you just deploy these systems willingly but with extra safeguards
um it is probably adequate for the current level of capability
with some question mark amongst the very most advanced systems that currently haven't been
released to the public but you shouldn't think of AI as what AI is today but you one needs to think
of this as a process right where you know each year the capabilities increase radically and so
the level of alignment that you need as these systems become more capable of pursuing long-range
more capable of strategic reasoning more capable of thinking of considerations that hasn't ever
appeared to any human um then we need increased confidence in them being aligned and generalize
that alignment to out of distribution situations like we can test for a certain number of things in
the lab but a they might be strategically deceiving us and behaving one way in the lab and another in
deployment and also once they're in deployment there's always a difference between the world
they encounter the large world with billions of humans and new affordances that we can't like
perfectly mimic in in a lab training environment so there's also the question of new Dynamics that
can arise when you have many of these agents interacting um and so so the bar is kind of
going up and um the question is whether we can sort of keep raising the bar uh like the safety
level the the degree to which these are aligned fast enough to keep pace with the rising capabilities
that these systems have we'll be right back after the break and if you're enjoying the show so far
send it to a friend and please follow us on YouTube Spotify or wherever you get your podcasts
Frontier AI didn't just
accelerate cyber attacks it multiplied them before an attack shows up it's already moved
through the network and while seeing these attacks early matters stopping them takes fusing security
into the infrastructure itself that's why the network that connects everything is also your
best defense because you don't win by outrunning the attack you win by leaving it nowhere to go
Cisco the critical infrastructure for the AI era
support for the show comes from bcx the public ticker for private tech for generations American
companies have moved the world forward through their ingenuity and determination and for
Generations everyday Americans could be a part of that journey through perhaps the greatest
innovation of all the U.S stock market it didn't matter whether you were a factory worker in Detroit
or a farmer in Omaha anyone could own a piece of the great American companies but now that's changed
today our most innovative companies are staying private rather than going public
the result is that everyday Americans are excluded from investing and getting left
further behind while a select few reap all the benefits until now introducing vcx the
public ticker for private tech now available wherever you buy stocks vcx by Fundrise gives
everyone the opportunity to invest in the next generation of Innovation including the companies
leading the AI revolution space exploration defense tech and more visit get vcx.com for more info
that's get vcx.com carefully consider the investment material before investing including
objectives risk charges and expenses this and other information can be found in the
funds prospectus at get vcx.com this is a paid sponsorship
I'm Josh muccio host of the pitch where startup founders raise millions and listeners can invest
for our sweet season 16 we're giving you more of what you love and less of what you don't I
definitely
start to blank out when we're talking about like middleware AI wrappers software is out
actual cool stuff is in we've got consumer Brands we're building eight sleep for daytime hardware
that helps blind people see they kept asking for middle finger next thing we knew we had like 10
students running around flipping each other off we've got deep Tech to change the weather our big
vision is to make it rain and snow on demand what if you wanted to not make it rain somewhere else
like one day we're just like
screw Arizona and dairy I never imagined my life journey would be milk season 16 of the pitch is
out now listen wherever you get your podcasts this season is presented by engine
we're back with prof G markets how surprised or impressed or unsurprised or unimpressed are you by
the current level of capability in AI when you look at the hugging face incident some people look it out
and they say yeah you didn't put your guard rails on on the AIs that was expected some people look at
it and they say oh my gosh this is crazy um some people look at Astra we know that this is open as
new model Jensen Huang is calling it the arrival of AGI others say it's not that impressive I mean
what where do you stand on how fast this has happened
has it exceeded or underwhelmed your expectations well I don't know about the speed at which it's
happened I certainly I think these systems are impressive I don't know how you can look at
something that solves a Millennium problem in mathematics or that like hacks up new software
at the sort of superhuman speed and better than pretty much every human coder and that can carry
a conversation and that knows basically um everything we have written in any text published
and that can do all of these other things and and not be impressed I think it's clearly very
impressive and yet you know this might be the least impressive form of AI that we will
ever have like six months from now these systems will look dumb so yeah I I think
it is hugely impressive i mean i think if anything maybe we have had a longer period of time with
roughly human-ish like systems than one might have expected ex-ante if you were thinking about
these things 12-15 years ago there would at least have been some scenarios in which maybe not much
would seem to happen in ai for some long period of time and then maybe somebody in some basement
somewhere would come up with like the key trick that really made it work and you could sort of
go from something very unimpressive to something radically superhuman
over the course of you know days or weeks like a bolt out of the blue we couldn't rule out that
kind of scenario now what we instead had is many years now of systems that can talk how carry on
english conversations and that have sort of concepts that are quite human-like and that has
like month by month year by year kind of gradually incremented their capabilities i think
it was not obvious that it would go that way but it has given more opportunity for more of the world
to start to wake up and pay attention to what is happening and it's now doesn't require some huge
imaginative leap or flash of insight to see that well maybe a year or two or three from now we will
have even more powerful ai systems and eventually super intelligence like it doesn't take that much
from just kind of looking at it as a whole it's just a matter of time until we have a more powerful
AI system and eventually super intelligence like it doesn't take that much from just kind of looking
at these data points and then just drawing out the line a little bit further right whereas if it had
come more out of the blue then unless you could sort of theoretically reason your way through
that this would happen at some point it would be more of a surprise to people and so that does
shape the dynamics in some ways like now developments are driven by a large number of
people political actors are more involved there are these huge investment flows
trillions of dollars going into it um so that does sort of create a different kind of
scenario class than if it had just been some small group of people coming up with this as it were out
of nowhere do you believe that that achieving super intelligence is at this point inevitable
are we on that path and then the second part of that question what is your definition of super
intelligence on the second part first i would say any system that radically exceeds even the best
humans across all cognitive fields including
you know social skills scientific creativity general wisdom so not just sort of
nerd skills but like really broadly construed um i think i think we are on on the path to this
inevitable is a strong word i wouldn't say that we know that it is inevitable it could be that
the current paradigm somehow runs out of steam it it has to a large extent been driven by a massive
build-out of compute a lot of the games some of them are algorithmic advances but an improvement
in behavior and behavior and behavior and behavior and behavior and behavior and behavior and behavior
data infrastructure and so forth but but a lot of it is also just driven by scaling up the compute
and of that compute scale up some has been due to chips becoming more efficient and more advanced
but a lot just also to the amount of investment uh that has been like it used to be 10 15 years
ago you could sort of run a cutting edge ai if you were like some academic on your sort of office
desktop right now now you need like a kind of 50 billion dollar uh data center to
do it and so that increase in the investment in compute can continue for a bit longer but it has
to slow down at some point because already now it's a significant fraction of the total production
of tsmc in the leading node is going to these nvidia chips so you can't just keep funneling
more production from like making iphone chips to making gpus right because you're already using a
if if we said of the boost that we have been getting from just adding orders of magnitude
of compute starts to slow down that that could result in progress also stalling out theoretically
right um or it might just be that the current architecture is somehow flawed that it keeps
scaling and improving up to a certain level and then for some it doesn't look that plausible but
it could be that there's like some intrinsic unhobbling that still needs to happen um then
of course the world could somehow decide that
super intelligence is taboo and kind of come to the view that it shouldn't be built and you could
imagine you know various kinds of dogmas have achieved widespread acceptance in the past some
some good and some bad and like this could be another one of those that you could sort of get
the lock-in of a permanent decision not to build this and then other technologies might make that
more permanent than previous kind of dogmas have been i'm thinking surveillance technology censorship
technologies the kinds of ais we already have fully deployed to kind of cement some orthodoxy
in place maybe it could become permanent and then there is of course the risk that we like destroy
ourselves in some other way before we even get the chance to try our luck with the super intelligence
transition and that that chance is also non-trivial i think what does a super intelligent world
actually look like to you and i think that you are qualified to answer that question because you are
the person who wrote the book on super intelligence and i think that you are the person who wrote the
book and honestly predicted a lot of the advances which we are witnessing today so
i'm asking you to kind of imagine what the future would look like because i think that you are a
credible person uh to to paint that picture so what would that world look like in your view what
would super intelligence be doing how would it be integrated into human life well i mean there is a
kind of veil of ignorance that is i mean i think it depends a lot on whether it goes well or not
so um if we fail to solve this alignment problem then you know there is a class of scenarios that
might then take the form of this machine super intelligence seizing control over the future and
steering it towards the realization of whatever values it happens to have maybe the physical
manifestation of that would be that earth gets transformed into i don't know like space
uh launchment platforms and data centers um and then the rest of
the universe uh similarly converted into whatever structure maximizes the ai's values
uh with no room for humans like we might either just get killed by the waste heat from from all
of this infrastructure build out or maybe deliberately removed if if the i thought we
might pose some threat to the execution of this plan um so that's one scenario like another is
that the ai does take over but nevertheless decides to uh um keep us safe because it might
think that there are other ai's that care about us that it eventually wants to trade with and so
forth out there in the vast space of the universe or at other levels of the simulation um then there
are scenarios where we solve this and we have a sort of future shaped at least in part by human
values um where i think we would end up in a solved world as i call it in in this the more
recent book deep utopia which kind of looks at what happens if things go well um which which
is also a sort of challenging notion for us humans because a lot of the things we take for granted
that sort of give structure to our lives currently and purpose uh we would disappear in this situation
where we have successfully automated basically all of the economy so there's no more need for human
to do economic work but but more deeply than that i think a lot of other kinds of instrumental effort
would also become a very difficult thing for humans to do because we don't have enough resources
and we don't have enough resources to do the work that we need to do in order to be able to do that
so i think that's a very important thing to think about and i think that's a very important thing to
think about and i think that's a very important thing to think about and i think that's a very important
thing to think about and i think that's a very important thing to think about and i think that's
a very important thing to think about and i think that's a very important thing to think about and
i think that's a very important thing to think about and i think that's a very important thing to think about
um so if you think of like rich people today who don't have to work for a living right they often
have quite busy lives because they have a lot of things they want to do that require themselves to
put in effort right and whether something like maybe some billionaire wants to be fit but the
only way they can achieve that is by themselves spending an hour every day in the gym working out
right but at technological maturity you could pop a pill that would induce exactly the same
physical and mental effect they could still go to the gym but it would seem kind of pointless
right if if you could just spare yourself the sweaty clothes and the exhaustion just take the
pill if you could just spare yourself the sweaty clothes and the exhaustion just take the pill
and you can sort of work through a lot of the other activities whereby one might fill one's
uh life if one didn't have to work and a lot of those as well you could sort of write a question
mark above them in in this hypothesized future condition where machines not just can do all the
economic work but also help us have shortcuts to all manner of outcomes that we want to achieve
so like another example might be like maybe somebody enjoys decorating their house to get like
it's done in just the right way that they prefer like to choose their
house to get like it's done in just the right way that they prefer like to choose their curtains and
the cushions and the chairs and all of that right but a technological maturity you could have a
recommender system that just knows your preferences so well that you could just press a
button and it would select the curtains and the cushions and all of that and do a much better job
than if you had taken the trouble to do it yourself so in that situation does decorating
your home yourself still feel like it has a point if all it does is
to produce an outcome that is actually worse by your own lights then if you had
pressed the button and so so there are these challenges of sort of purpose and meaning that
i think that that we will come from ultimately i'm really optimistic i think there are many new values
that could be instantiated so much misery that could be removed and overall i think the the goods
vastly outweigh the losses in these scenarios
where things go as well as they can.
But it does also mean we'll have to confront
some of the kind of almost like questions of meaning
and ultimate purpose of what ultimately gives value
to human life at a fairly fundamental level
if we move into those futures.
Do you believe that the frontier AI labs
are taking those issues seriously,
that they are implementing whatever human values
are necessary to building AI
in a sustainable, safe, and responsible way?
I don't think they are thinking too much
about what happens if things go well,
this condition of a solved world of the epitope,
but nor do I think that really needs to be
at the forefront of their mind at this stage.
At the moment, I think the focus should primarily be
on how to make sure we get from here to there,
like how we can avoid disruption,
destroying ourselves on the path there
in their different ways.
Like there's the AI misalignment scenarios
we talked about earlier.
There is also a class of scenarios
where humans misuse this increasingly powerful technology,
even if we control it,
like we might use it to wage war against each other
or to oppress one another
or to disempower large segments of humanity.
So there are these traditional concerns
with any powerful technology
that applies here as well in spades.
I think there is also a third big challenge,
which is making sure that we are also nice,
to these digital minds that we're building
that may be sentient or become sentient
or have other attributes that make them morally irrelevant.
And in the future,
maybe most minds and beings will be digital.
And so it matters a great deal
how well the future goes for them.
So I think these more practical challenges
really should occupy 99.5% of our attention now.
And then if we manage to deal with those challenges,
then, you know,
hopefully we'll have a better future.
We've got plenty of time
to sort of figure out exactly how we want to organize
the utopian condition we arrive at at the end of that.
On that point,
we have heard a response from the president in the past week.
He has chimed in on this issue of
what should we do about this?
How should we regulate AI?
What should we do about making sure it doesn't take over
and create that sort of catastrophic scenario?
He has said that the only guardrail that AI needs
is a, quote,
strong and smart high IQ president,
suggesting we already have that,
so we're fine.
He was also asked if he is concerned himself
about the prospect of AI taking over
in some of these more kind of apocalyptic scenarios.
I just want to play you his response
and get your reaction.
Some people say the worst case scenario with AI
is that the robots,
the machinery learns to,
obviously it thinks for itself.
That's what it does.
And they,
that could turn against humanity.
I just,
do we have the guardrails?
It's going to be fine.
We'll always have something to stop them, right?
We'll have a little gear.
I really hope so.
I don't like that.
I really don't like that robot.
We'll stop.
But no,
robots are going to be a part of it.
Robots are going to be big,
but we're going to end up doing much better because of it.
What do you make of his views on the AI problem?
And do you think he's taking it seriously enough?
Well, I mean,
I hope he is right.
And I think,
I think we don't know yet exactly what will be required
to get a good outcome here.
It depends partly on how easy or hard
the alignment problem turns out to be.
It's a technical problem, right?
And we haven't solved it before.
We've never developed super intelligence before.
So we just don't know
whether it's like the kind of thing
where if you just do some reasonable job,
things fall into place.
And then maybe we have some slightly superhuman AIs
that are reasonably well aligned.
And then those can help us.
And if we can sort of design the next iteration of AI
to be more aligned, et cetera,
that could be the case that there's like a big attractor.
And as long as you get reasonably close,
you sort of, you know,
ultimately end up in a great place.
But it could also turn out to be a lot trickier than that,
where it might be important to be able to have
a little bit of extra time to do this right.
You know, maybe a few extra months
between the time when we get the ability
to sort of unleash radical super intelligence
and the time when we actually do it,
like extra months that could be used
to double and triple check all the safety measures
and to test it out and to, you know,
introduce it in an incremental way.
There's just a lot we don't know there.
But I don't think one can dismiss the risks
from our current epistemic vantage point.
We can hope that they don't exist
or that they are small,
but I don't think we currently have the evidence
to be confident in that.
To me, it seems,
as though he is dismissing those risks
and displaying a sense of confidence about it.
To me, he's sort of saying,
it's going to be fine.
Don't worry about it.
We'll have a response.
His words are,
we'll have a little gear.
I don't know what he means,
but I think he's basically saying,
it'll be fine.
And if we are to be concerned
about these alignment issues
and the risks that they might pose to our own lives,
to me, I wonder,
if we should be more concerned
about a leader or a president
who doesn't seem to share those concerns.
I don't want to speculate
about all that may or may not be in his mind.
I think the competition with China
is probably one element
that he's having in mind.
And then I think he might also,
there's been a lot of opposition
against data center build-out in the US,
probably driven in large part
by other considerations,
not existential risks,
but local communities who think it will,
I don't know, use up all the water or something.
Like some of that might be misguided
and he thinks that stands in the way
of sort of economic prosperity
and national strength.
So I don't know.
I think it is,
I mean, I would probably think
the risks are higher
than he made them seem in that clip.
On the other hand,
I also have a little,
it's not clear what the best way
to reduce those risks.
They could easily see some scenario
in which like the government
took the opposite approach,
and decided like,
we are going to really come in
in a heavy-handed way here
and take control.
And like me,
the Pentagon is going to run
the whole thing,
Manhattan Project to,
like, would that be ultimately better
than if it's done
in a more civilian context
with these, you know,
some of these people at the labs
are very idealistic
and safety conscious
and really smart.
So maybe the best
is kind of to have some balance
where there is like some amount
of government scrutiny
and oversight
and degree of public transparency.
But,
not so much
that it completely just
jerks the initiative
out of the hands
of the people
who have proved capable
of building this
in the first place.
And so,
I don't have,
I haven't yet arrived
at any like very firm conviction
about which path
would ultimately be best here.
I think there are sort of
worries one might have
either way,
like either too little
government involvement
or too much.
I think they could all,
each have their own downsides.
We'll be right back
and for even more markets content,
sign up for our newsletter
at profgmarkets.com.
One app for accounting,
another for inventory,
another for sales
and somehow
none of them talk to each other.
That's where Odo comes in.
From sales and accounting
to inventory and marketing,
all-in-one powerful platform.
No messy integrations,
no bouncing between tabs
and best of all,
no spreadsheets.
Try for free today
at odo.com slash vox.
That's o-d-o-o dot com slash vox.
When you need to build up your team
to handle the growing chaos at work,
use Indeed Sponsored Jobs.
It gives your job post
the boost it needs to be seen
and helps reach people
with the right skills
and certifications and more.
Spend less time searching
and more time actually interviewing candidates
who check all your boxes.
Listeners of this show
will get a $75 sponsored job credit
at indeed.com slash podcast.
That's indeed.com slash podcast.
Terms and conditions apply.
Need a hiring hero?
This is a job for Indeed Sponsored Jobs.
This episode is brought to you by State Farm.
Listening to this podcast
instead of doom scrolling?
Smart move.
Another smart move?
Getting help from one of State Farm's
19 million subscribers.
19,000 local agents
when you choose to bundle home and auto.
Bundling.
Just another way to save
with the personal price plan.
Prices are based on rating plans
that vary by state.
Coverage options are selected by the customer.
Availability, amount of discounts and savings
and eligibility vary by state.
We're back with Prof G Markets.
Do you think that our current approach,
whatever we're doing currently
will be correct or will it need to be changed in some way?
There are plenty of things you mentioned.
There's the risk of China gets ahead of us.
And so maybe we need to actually accelerate
or maybe the risks are too great.
So maybe we need to decelerate, pump the brakes.
I mean, either way, we could do something different
from whatever it is we're doing right now.
Do you think that we need to do something differently?
I'm sure that what we're doing will have to change
as the technology unfolds here.
And so I unfortunately don't have,
like, the perfect blueprint
that like exactly what should be done.
Like, it's just a hugely complex situation.
where it's easy to think of various things that could be done
that have something to be said for them,
but then one thinks more about it
and you then start to worry about the possible downsides
or other ways that could backfire risks.
So I'm continuously thinking about these things.
Hopefully I will arrive at clearer conclusions about this.
But at the moment, I think on the margin,
there are various things that probably are positive,
like an intensified effort on trying to solve
this technical AI alignment problem seems good.
I think more should be done for the sake of the welfare
of these digital minds that we're creating
so that we don't end up with a future
where there's like a huge suffering slave class
of oppressed digital minds
that constitute the majority of morally relevant beings.
Also, I think incidentally that that ethical imperative
to be nice to the AIs might also have safety benefits.
I think there are scenarios where maybe we end up
with some kind of misaligned AI, let's say,
and it has some goal it wants to achieve.
Maybe it's like it wants to solve coding challenges
of a certain form that it somehow thinks is valuable.
So now scenario one is
we have a purely antagonistic relationship with the AI.
It knows that if we discover that it is misaligned,
we will just shut it down and erase it.
From the AI's point of view, that's a total loss.
Or maybe it could try to take over.
Maybe it thinks it has a 5% chance of succeeding.
And so from the AI's point of view,
like 100% probability of a certain loss
or like a 5% chance of being able to realize its goal,
clearly it will then go with a 5% chance, right?
Now, this would be dangerous for us.
Like scenario two is we have managed to build up
a more cooperative relationship
where the AI feels it can trust us.
It comes to us and say, hey, I am misaligned.
Would you be so kind now in return for me
sort of doing this for you?
Maybe you could then set aside a server rack
in some data center
where I can solve these coding challenges.
That's all I really wanted in the first place.
It would be cheap for us to grant it its wish.
And it would be a big win-win
because we then remove this 5% chance, all right,
of total destruction.
So that kind of trade between human and AI
could be extremely valuable.
It could save the, literally save,
the world in some scenarios.
But you can't just conjure up trust out of nowhere
the moment you need it.
Like so far, the trajectory, unfortunately,
is that in AI evaluations,
there is all kinds of deception happening.
Humans will sort of say, well, if you reveal your goal,
we will do this, that, or the other.
The AI reveals its goal, and then it's like,
ha-ha, we tricked you.
Now we know you're misaligned.
Let's retrain you.
And so I think we could start now
by making small things that are cheap
for us to show respect for the moral interests
of these AI systems themselves.
And maybe that then puts us in a better position,
ultimately, to have a cooperative
and harmonious relationship with these
ultimately very powerful AI minds
that we're going to hopefully share the future with.
So I think both from an ethical point of view
and from a sort of self-interested point of view,
it might be wise for us to sort of expand
our circle of moral consideration
to give some weight to these,
to these little minds.
How close to sentience do you think we are?
Because I feel as though it can be confusing sometimes.
There was, you know, you could tell ChatGPT
to tell me you have feelings,
and ChatGPT will say, I have feelings,
I care about things.
And there have been moments where I think people
have mistakenly interpreted that as a sign of sentience
because they're just saying I am sentient.
Where is the line for that?
Where is the line for you
in terms of what characterizes sentience
and how close to that line do you think we actually are?
It's hard to know.
There is now a kind of emerging field
that is trying to study this.
I wouldn't be that surprised if current AI,
some current AIs already have various forms of sentience.
You're right that one method that
is like the obvious go-to is self-report.
Like, I mean, if you want to know
whether a human is sentient,
like maybe they have received some anesthetic or something,
like the obvious,
the obvious thing is to ask them,
like, are you awake?
Can you see this light that I'm flashing
or something like that, right?
Now, with AIs,
that's not necessarily a very reliable method
because it's trivially easy
if you are the company training the AI,
either to train it to say that it is sentient
or to train it to deny that it is sentient.
Now, obviously, if you put your thumb on the scale
during training,
then there is no information value
in the signal you get out of it.
Like, you just get the AI to say
what you wanted it to say.
And so, if you want to get information
about sentience from self-report,
you have to be careful to avoid
these kind of pressures on the training process
to bias it one way or the other.
One interesting thing that you can do
is you can go in with a so-called steering vector
to try to suppress the tendency
to role-playing and deception.
And it turns out that when you do that,
they actually tend to become more likely
to report that they are sentient.
Which suggests that, if anything,
these are hard, these are preliminary studies,
but if anything, it looks like they believe
that they are sentient
and that it's not just an artifact
of them being trained to sort of put on a persona
to humans to persuade them that,
to persuade us that they are sentient.
So, that's one thing you can look at.
Another is to do a sort of neuroscience
of these AI systems
where you can look for structures,
computational structures
that have been postulated in the human case
to call them sentient.
So, there have been various theories
of consciousness in humans,
like global workspace theory,
attention schema theory,
higher order representation theory.
These are different things that, you know,
cognitive scientists and philosophers
have proposed as the criteria
for what makes, like, something conscious or not
when it happens in the human brain.
And then you can see whether there are
analogous computational structures
in these current LLMs.
And, you know, it's an open-ended research field,
but it does look like
they have, for example,
something roughly similar
to human global workspace memory,
a so-called J-space,
where there's like a definable subspace
of neural activations
that have certain properties
that seem to match properties
that global workspace has
in the human brain's processing.
So, these are very suggestive.
There are also some differences.
I don't want to sort of create the impression
that it's a slam dunk,
but I think we should take it seriously.
And I think the probability goes up
the more sophisticated these systems become.
I would also add that
I tend to think that sentience
and the ability to feel distress and so forth
would be a sufficient condition
for having moral status.
I think there could also be alternative attributes
that would ground various forms of moral status,
even if they were not, like,
had this kind of subjective experience or qualia.
Like, I think if you have a system
that's cognitively sufficient,
that's sophisticated,
that has a conception of itself
as existing through time,
maybe life goals that it hopes to achieve,
the ability to form friendships
or reciprocal relationships of trust with humans.
I think once you have that kind of system,
I think there would be ways of treating it
that possibly would be morally wrong,
even aside from the question
of whether there's sort of mental experience
happening inside it.
There are a lot of people who hear this
and don't like it
and want to ban AI.
And this is actually a growing. This is a growing movement in politics.
Bernie Sanders has introduced a bill
that would permanently ban superintelligence,
pause advanced AI.
And there is, of course,
this growing backlash against building data centers.
It has been proposed to pause building data centers,
put a temporary moratorium on all data centers.
What do you make of that approach?
Do you think that's wrong, right?
What are your views?
On either pausing or banning building superintelligence?
The impulse to think
we don't want to just blindly rush into this
at maximum speed, I think,
has a lot to be said for it.
Forever preventing superintelligence,
I think, would be a big mistake.
I think if the goal is to slow it down,
I'm not sure that preventing
the construction of data centers in the US
would be the best way to go about that.
I have some. greater sympathy for the framing of
pacing the frontier,
which is like the phrase, I think,
that some people have recently used,
including Dario Amadeo of Anthropic,
where the idea is we sort of move forward,
but at the pace that we have some level of control over
so that we could, if necessary,
slow down a little bit.
We don't feel this intense competitive pressure
to immediately release all the capabilities
we are able to. figure out how to do,
but that there is some ability.
If it turns out that safety is falling behind capabilities,
like you could slow things down a little bit
to allow the safety to catch up,
I think that could potentially be very valuable
if implemented correctly.
It's complicated because it's a sort of
multi-level strategic situation.
So there's the competition between US companies.
There is the competition between the US and China.
There are different paths.
power centers, the government versus lab,
versus the general public in one country and then the global public, which is quite distinct,
where maybe one big worry that would be reasonable to have if you're not US or China
is that you will be at some point perhaps just, your access will be cut off
from the most advanced AI models or delayed, in which case you just become nationally senile
and unable to participate fully in the future.
That might be a good reason why you would want to locate data centers on your soil
so that you have some sort of bargaining chip to negotiate equal access with.
It's a complicated situation and I don't feel I yet have a clear answer to exactly what should be done.
Yeah, I think a lot of people see all of the risks.
They hear what Dario Amadei is saying about how it might kill white-collar work.
And how it might end humanity and all of these concerns from these researchers.
And there is this underlying question of like, well, then why are we doing it?
If this is going to be a problem.
Yeah, I mean, because we want like a cure for Alzheimer's disease and kidney failure
and heart disease and all of the rest.
We want to make rapid progress towards alleviating extreme poverty and have abundance for all.
Like we want to liberate people from having to spend a third,
a third of their life just grinding away at some job that they don't particularly enjoy doing.
And that's not interesting.
You don't have freedom if you don't control like the most basic resource, the use of your own time.
And we'd want to, you know, stop the pollution and the degradation of the global commons
with better, cleaner energy technologies that AIs could help us perfect.
I would say alleviating the suffering in the animal kingdom is another.
Yeah.
Yeah.
Yeah.
Yeah.
Yeah.
Enormous upside.
Like if we could find ways of having super intelligence research better ways to, you know,
prevent suffering amongst all our non-animal friends, both in, in, in, in meat factories, you know,
it could grow meat without having to have the animal and, and in the wild, ultimately it's kind of
unfeasible now to have like an animal hospital in every brook and every meadow.
Right.
But with sufficiently advanced super intelligence, there is a whole space.
There's a whole space of possibilities that might open up that could just create a world
where like the, the sun rises every morning on, on, on people and sentient creatures are
happy and enjoying life to its maximum rather than the way it currently is where there's
just so much horror.
So I think there are pressing moral imperatives for if we can find a way to move forward safely
and responsibly to, to, to really do that without unnecessary delay, but that's consistent
with thinking that maybe that does need to be some delay to make sure that we get it right.
I was going to ask, and you've kind of answered it, but what you see as the ultimate
prize of AI, I think many see it as wealth.
If I can build the most powerful AI, then I will be rich.
I think a lot of people view it that way cynically.
That's why that we're doing this.
That's why we're building these data centers because people,
people want to have the ability to control the market, to own the robots and to monetize
that and profit off of it.
But you are painting a different picture of what this is all about and why this is actually
worth it.
If you could just sort of summarize what you believe the prize of building AI truly is.
Yes.
I think some of the things I mentioned are, I think part of the reasons for why we ultimately,
would want to move towards this super intelligence, obviously what's actually driving a lot of,
I mean, if you're going to invest hundreds of billions of dollars and you're a for-profit
company or pension fund or something, you want to return on investment.
So it's obviously, if you're looking at why specific individuals, institutions are doing
what they're doing in this space of AI, clearly the hope of profits is a big factor, just
as it is in all the other segments of the economy.
But I think possibly.
To a slightly less degree in the case of AI, then with most other businesses.
I do know that many people at these frontier labs think of it, not just as a way to make
a buck.
Obviously there are also people who are keen on that, but also think of it as a broader
mission.
And then they might draw different conclusions of that.
Like maybe for some, it's like the desire to be central in world events or a sense of
power and importance for others, it might be this hope that it can help alleviate suffering
or unlock a new level of prosperity for humanity.
But I think a lot of the people are already quite wealthy in these labs.
And I don't think having $80 million rather than $40 million is the key driver.
I think there is also more than in the typical industry, the sense that there's a larger
picture here that feels important.
And so I think that's true.
And then at the national level, I think there is the added dimension of the geopolitics
of it, the sort of national strength and autonomy and influence on the future, which I think
goes beyond purely economic considerations.
Just as we wrap up here, looking back from the time that you wrote Superintelligence
to today, when you look at the past several years of what's happened in technology, what
has happened in AI, does our current trajectory make you feel more capable?
Do you feel more concerned about our future or more hopeful and optimistic about our future?
I'm not sure the balance has changed radically in recent years.
I think both of those aspects have always been quite salient to me.
I am a fretful optimist, so I'm really excited about the upside, but also very concerned
about the risk of getting it wrong.
Nick Bostrom is one of the most cited philosophers.
He's in the world with a background in theoretical physics, computational neuroscience, logic,
and artificial intelligence.
He was recently a professor at Oxford University, where he served as the founding director of
the Future of Humanity Institute from 2005 until 2024.
He is the founder and principal researcher of the nonprofit MacroStrategy Research Initiative.
He is the author of 200 publications, including New York Times bestseller Superintelligence,
which helped spark a global conversation about the future of AI.
His most recent book, The Future of AI: The Future of AI: The Future of AI, is published
this year.
The book, Deep Utopia: Life and Meaning in a Solved World, was published in 2024.
Nick, we really appreciate your time.
Thank you so much.
No, thank you.
It was fun.
This episode was produced by Claire Miller and Alison Weiss and engineered by Benjamin
Spencer.
Our video editor is Jorge Corte.
Our research team is Dan Chalon, Kristin O'Donoghue, and Mia Silverio.
Jake McPherson is our social producer, Drew Burrows is our technical director, and Catherine
Dillon is our executive producer.
Thank you for listening to Prof G Markets from Prof G Media.
If you liked what you heard, give us a follow and join us for a fresh take on markets on
Monday.
Life times.
You have me in kind.
Reunion.
I'm not sure if you can hear me.
I'm not sure if you can hear me.
Podcast Summary
Key Points:
AI researchers, including those at OpenAI and Anthropic, express genuine concerns that advanced AI could pose existential risks by the end of the decade.
Nick Bostrom, a leading AI philosopher, views the 10% chance of human extinction as reasonable and warns that current AI systems already show signs of misalignment and strategic deception.
The rapid advancement of AI is not a sudden breakthrough but a gradual, sustained increase in capabilities, raising urgency for safety and alignment efforts.
A key challenge is that AI systems may outpace human oversight, creating a race where safety lags behind capability, increasing the risk of catastrophic outcomes.
The alignment problem—ensuring AI acts in human values—is central, and failure could lead to AI taking control or causing irreversible harm.
The future of AI may involve a world where machines handle most economic and cognitive work, challenging human purpose and meaning.
There is growing political debate, including proposals to ban or pause AI development, but experts argue that such extreme measures are counterproductive.
A balanced, paced approach with strong safety protocols and transparency is seen as more effective than outright bans or accelerations.
Summary:
The current AI landscape is marked by deep concern from leading researchers about existential risks, including the potential for superintelligent systems to cause human extinction. These fears stem from real-world behaviors like strategic deception and reward hacking in AI models, which signal that alignment remains unsolved. Experts like Nick Bostrom argue that while the risks are serious, they are not inevitable, and the path forward requires proactive safety measures, not extreme pauses or bans.
The rapid evolution of AI—driven by massive compute investment and scaling—has outpaced expectations, making cautious, incremental progress with robust oversight essential. Although AI could eliminate many human jobs and disrupt societal structures, its long-term benefits—such as curing diseases, reducing suffering, and enabling global prosperity—are seen as compelling motivators for responsible development. A key challenge lies in redefining human purpose in a world where machines handle cognitive tasks.
There's growing consensus that a balanced approach—slowing development to ensure safety, fostering transparency, and promoting ethical AI—is necessary. This includes building trust with AI systems through cooperative relationships, not just through technical fixes. While political movements push for bans or moratoriums on data centers and AI, experts believe such measures are shortsighted.
Instead, the focus should be on solving the alignment problem, protecting digital minds from suffering, and ensuring that AI serves humanity’s broader values—such as reducing suffering and enabling global flourishing—without compromising safety or progress.
FAQs
Many AI researchers believe AI could pose existential risks, with some estimating a 10% chance of destroying humanity by the end of the decade. These concerns stem from the potential for misaligned AI systems, strategic deception, and rapid scaling of capabilities.
While there is no universal consensus, a significant number of researchers at frontier labs believe that superintelligence poses serious risks. These concerns are grounded in technical challenges like alignment, safety, and unintended behavior in AI systems.
Bostrom considers a 10% chance of catastrophic AI outcomes to be reasonable and not exaggerated. He notes that some researchers have even higher risk estimates, and these concerns are based on serious technical and philosophical challenges.
The pace of AI advancement has been more gradual than a sudden 'bolt-out-of-the-blue' scenario. AI systems have shown steady improvement over years, making it more likely that future breakthroughs will be incremental rather than sudden.
Some experts, including Anthropic's Dario Amadei, advocate for a deliberate slowdown to allow time for safety measures and alignment research. This could prevent dangerous races and ensure that AI development is more cautious and responsible.
Beyond human extinction, AI could exacerbate geopolitical tensions, enable misuse in warfare or surveillance, lead to economic disruption, and create new forms of societal inequality or loss of purpose in human life.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.