I remember a conversation I was a grad student and I met with him for the very first time
I think it apps and I told him I was like I sent this thing in this journal
Whatever and it's been a long time. I haven't heard anything. He's like well, you're still alive
So you get that rejection email. You're still alive. That was a good early career advice
You know, I mean yeah, just to get chill a little bit just relax, you know until it's a no
It's it's not a no you wait till it's a no and then you move on to the next thing
Why why assume it's if you gonna worry about it and make it a no in your head
That's that's useless make it a no later when it's actually a no
You know saying that's a good philosophical tidbit for the students you like
Hi everyone welcome back to cheap talk my name is Jeff Kaplow
I'm an associate professor of government here at William & Mary and joining me as always is my esteemed colleague Marcus Homes
Hello Marcus
Jeff how's it going? How's your how's your day going?
My day is going very well Marcus. Did you know that the opinions expressed on this podcast are solely our own?
Wow do not reflect the policies or positions of William & Mary early today. I like that. You are on top of things
That's great. No, normally we wait till the last minute. Can I just say before we start I have a bone to pick
With with AI and the bone and I have to pick is that this technology supposedly is so great and super intelligent that it's gonna wipe out humanity
And yet it can't seem to draft a fantasy football team. That's worth itself. My experience this week in fantasy football was horrendous
I stink like these guys first of all I've never heard of half of them and like like something got like zero points
Zero points and I and as I said last time I let chat GPT draft my entire team
I said computer you have all these vast resources like figure out who to draft it
I followed his instructions precisely to a T and this is the garbage that it gave me so if you're worried about AI taking over
You know, let's let's pump the brakes a little bit it can't even draft a proper fantasy football team Jeff
This might be a socioeconomic thing which model and thinking level where you're using that's a great question
If you if you had more money worth of more tokens to spend you might have had a better team. I was using chat GPT
What is this five point five I have plus I have to I have the $20 a month plan
I don't have the fancy $200 a month plan
I'm on the plus, but I did move the little slider bar over to high
So I wanted to I mean I shudder to think what it would have done if it was on low
I mean this is what it gave me on high. I can't even imagine like what I have a team of all kickers
I mean what would it give me if I didn't have the high version on so high is the halfway
Right, I know I know that it's poorly labeled, but it runs from instant to medium to high to extra high to pro
Okay, in my version. I don't everything after high is there's a little lock icon which implies to me
I can't do it because it wants me to buy it wants me to buy pro
Yeah, sorry, man, so I'm held back by my unwillingness to give open AI $200 a month
Well, you need us some kind of university sponsor to take pity on you
But they don't want to they don't want to spend money on your your fantasy football draft even for science
Well, I could ask them. Well, you're right. They probably don't but but I'm saying like so but that that implies
Flinder goes all the way up to to pro. I'll just I'm using I'm using chat GPT on training wheels is what you're saying
Like I'm using the kind of. Here's the problem right if everyone's gonna be using it in their in their fantasy football drafts
Yeah, then if you want to compete, you got to be willing to go go higher than everybody else
It's an arms race Marcus is what it is
So I'm just surprised I would have thought that the the baby version of chat GPT that I have would have given me something that would have been
I don't even I don't even want to compete
I just want to be like in the in the ballpark like right now. It's not even it's like I'm
I'm playing checkers and everybody else is playing chess like it's not even like I'm we're the same on the same game
You know what I mean like this this team is still horrible
So actually what I did yesterday was I joined another league and I said I'm not gonna use chat GPT at all
I'm just gonna go with my like you know human intuition
Because the season is already started. We're gonna kick off
Pun not intended on week two. So now I have like my team that I drafted purely on human
Resources and then my purely AI team and we'll just see what happens over the course of the of the season
But I am not impressed at the moment. That'll be a fun test to see who ends up
Who ends up ahead? I think my human version is gonna end up ahead sadly. Yeah, I would love for that not to be the case
Speaking of AI
You mentioned it's going to kill us all. What are you referring to there?
Jeff, I haven't seen as much
Sort of hyperbolic discussion about a topic in quite some time and actually I should say like hyperbolic coming from some
Areas of of the sort of industry and then like the exact opposite coming from other areas of the industry, right?
And so what I mean is it seems like over the last week or so
Or the last few days we've had this sort of like drumbeat of
AI
People, you know, whether it's the the CEO of open AI that the penless letter or maybe a clawed
I don't I don't keep track of like through the players are in this. You'll be able to crappy on this
But basically one of the CEOs. Yeah, I think it was anthropic wrote this kind of like long
3000 word, you know, sort of warning basically saying that, you know
We're going towards it. We're moving towards a moment in time where AI is going to become very very powerful and quite
Dangerous and he was calling for basically the industry to kind of slow down kind of calling for some type of you know
regulatory action if if not you know from the government at least you can get together
Maybe as an industry and kind of slow things down and that came on the heels of I think open AI
Employees quitting basically saying like I don't like where this is headed
I don't like the way that the company is being run. They know that this is dangerous
They're not putting the safeguards in they're just kind of continuing down this this road
And eventually it's going to be very very dangerous and as we've talked about on the on the pod many times
This is all sort of on top of the various philosophical arguments that have been made for several years now by you know
Jeff Hinton and other people basically saying like this is AI is going to be very dangerous
And it's something that we need to keep tabs on and regulate so you had all of this sort of like AI is going to lead to
The destruction of humanity. It's going to end the world
And by the way when they say these things like they they do as we're probabilistically
It's some of the estimates are like there's a 2% chance that AI is going to kill everybody
Or there's a 5% or 10% chance it's not like a 100% chance or that like they're saying like we're definitely all going to die
Within the next two years because the the world's not going to exist due to AI, but they're saying like it's a non-zero chance
That humanity won't survive artificial intelligence and so therefore we should take it seriously and try to reduce the risks and and do that sort of sooner rather than later because
You know things are moving along very quickly
On the other hand you have or the other side to this you have you know people like the president who have said that this is a hoax
AI is not a threat to humanity. All you need to stop AI is a smart president
Which we have of course and so that the president has said because I'm smart like we're going to be okay
And I think kind of the the conservative small c kind of way of looking at this is that regulation
You know just generally is kind of a bad thing. We don't want to limit the ability of a private enterprise to do what's thing
This is a competitive advantage that we have vis-a-vis China
Can't let them win the AI race and all that kind of stuff
So you have the people who do AI the sort of like you know people actually programming in and owning the companies that
Are you to create things like Chatchee PT saying this is a sort of existential threat
And the president and some others on the other side of saying this is a big hoax. Don't worry about it
So Jeff I'd say like this is like the the since Chatchee PT came out. I don't remember as much sort of you know sort of
Alarm being raised all at once like it seems like we moved into something different everybody for a while
Was sort of saying yeah, maybe one day is going to be dangerous
Seems like all of a sudden we've we've decided or a lot of people have decided this is super dangerous
And we need to do something like right now like not not tomorrow
But like today, and if we don't do something today
We're gonna bad kind of stuff happen to us. Is that might read to the situation right Jeff?
Yeah, I think it seems that way. Let's back up a little bit and unpack this
Yeah, there's a lot of there's a lot to talk about in in your opening statement their markets
So just just to back it up what
Started I think a lot of this discussion or at least the spillover of a lot of this discussion into the public space
I think these discussions have been happening in the AI research labs for quite some time and as you mentioned in in some
Corners of even our world of academia people have warned about this but
The what really set off the latest round of speculation about the danger of AI
I think was what we've been talking about in this podcast the open AI hugging face hack
This discovery of a series of related cybersecurity incidents coming from all the major AI labs and the kind of
Realization that these labs do not have a great handle on the security around these new models or even a full understanding of what the models are capable of
And then setting off this latest round more proximately was the resignation
Of Jacob Coxon who is a researcher at Anthropic who resigned from his job and then posted on a website known as x.com
That the people building AI he said earnestly believe that he could kill us all by the end of the decade
This is not a marketing stunt
But I hear the these people express these fears privately no other human activity poses this level of danger
and then a current AI research
researcher at Anthropic, Evan Hubbinger, who is the lead for alignment at Anthropic,
which I remember is the idea that I will do what we ask it to instead of something different.
That's alignment.
He posted in response, "Jacob is correct here.
We really do earnestly believe AI could kill all humans."
I personally think it is greater than 10% within the next decade.
I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment
for superintelligence and are not clearly on track, too.
And this is a person still working at Anthropic, so at least the other guy resigned seconds
before posting this thing in public.
But Evan Hubbinger is still there and says, "Yeah, 10% greater than 10% within the next
decade."
Which is high.
That's a high percentage.
I know we all struggle with probabilities, right, but like, if there was a greater than
10% chance that I was going to get in a car accident today, I would not leave the house,
right?
It's too high for comfort.
So certainly the end of all humanity greater than 10%, that's pretty high.
And then the president of the United States followed up on this series of statements that
people are making.
He said, "The industry already has the only guardrail it needs.
A strong and smart, high IQ president."
He dismissed safety concerns as a hoax.
He said that this is in a truth post, I assume.
The Trump administration has stopped AI, quote, people from doing bad, or potentially bad
things, like Dario and Anthropic, who is now pretending to be a perfect little angel.
And we will continue to do so.
We already have tremendous criminal and regulatory power over these companies.
The New York Times goes on to say, "It was not clear what authority Mr. Trump was referring
to or what actions he believes his administration has already brought."
The New York Times writes, "Mr. Trump's own grasp of technology of the technology is superficial,
according to people close to him.
He is an extreme late adopter of tech."
He began text messaging only in 2022.
And his staff friends out social media posts for him to read.
So that doesn't inspire confidence.
And then I think, you know, at the same time, we have an interesting and influential, at
least the subject of much conversation, a post from Dario Amode, the CEO of Anthropic,
the maker of Claude.
We're actually arguing for a slowdown in AI development, a cooperative slowdown across
the industry.
Not a unilateral slowdown by Anthropic.
Right.
Even if they think there's a great room, 10% chance that they'll be a bad idea.
Yeah.
Right.
Why would we want to stop?
But a cooperative slowdown across the industry, and in fact globally, in a post that he put
up called, "We must pace the frontier."
Right.
So taking this step back though, it is striking to have an industry, and not just one person
in it.
It's not like this is one sort of like outlier.
Like there's a bunch of people now in the industry saying, "We need to do something."
Like what we're sort of moving towards is not going to be great, and we need to collectively
do something.
And this is kind of call for collective action.
And as you pointed out in your, you know, a minute ago, it's difficult for, like some
people said, "Okay, Anthropic, slow down."
No one's stopping you.
Be my guest.
Number one, that's not going to solve the problem if it's like other AI companies, you
know, are not slowing down, but also all the incentives are for Anthropic not to slow
down on themselves.
And so they need cooperation.
They need some type of collective action in this industry, whether it's some type of,
you know, regulatory framework or whatever, to slow everybody down at the same time.
So it does seem like we've moved into a new phase of this, Jeff.
This is not just like one kind of outline person being like, the world's, you know, kind
of come to an end.
Like it seems like there are people that are very nervous about this.
Yeah, I think that's right.
And in his post, Amade talks about what us slow down, what pacing the frontier buys
them, right?
That if they do implement this slow down, then they can be able to pay more attention
to operational security.
They can spend more time at every phase, but it's the competitive pressure of trying to
keep up with opening, I guess, meta, I don't know, China, China is pushing them to do
things that he doesn't write this.
Some part of their conscience is saying, you should don't do this.
Don't do this.
And they're pushing ahead.
And my immediate reaction to this is, what the hell, man, just if you believe that what
you're doing is going to destroy humanity, then sure, try to get AI, oh, try to get
open AI to slow down too, but you should just do it, right?
Like, I mean, this is getting ridiculous.
This company is very well capitalized.
There are no, no danger of going out of business that they can slow things down on their own.
And they should.
And so should open AI, and that would be, they could even bill it as a, a cooperative
gesture, right?
And in international relations, we have a whole field of confidence and security building
measures designed to allow different countries that have different competing interests to
take those small steps that are in interest of both together in a way that doesn't disadvantage
either one.
And what they should do is anthropic should say, we are slowing down, effective immediately.
Here is how we're going to structure it.
We invite open AI to join us, and then Sam should join them, right?
And because it's, if somebody takes that first step, it makes it possible, at least, for
the other two.
Now, open AI, knowing open AI will just be like, now's our chance and double their speed,
right?
But so that is a danger.
But I'm not worried about anthropics like long-term business prospects.
And we're not saying stop development, and he's not advocating that they stop development.
He's saying, let's just be more deliberate about how we go about doing it.
So I appreciate, you know, it's good that these people are paying attention to these issues,
for sure.
But it's hard to give them too much credit because it is in their power to just do this
any time, right?
It's not like somebody's forcing them to push ahead in the way that they have been up
to now.
Yeah, and it's possible.
I mean, it could even make some sense from a business perspective, right?
So I know that there are alternative, like, chat GPT type of entities who are pledging
more kind of environmentally sustainable ways of doing it.
I don't know if those are actually like, if that's true, that they're less sort of environmentally
unfriendly or whatever, but there might be like a business opportunity for your ability
to market yourself as like a responsible AI entity in a world of like other responsible
or irresponsible AI companies, I mean, that's also plausible.
But my reaction to you is, it was very similar.
Like, I think, you know, if, if something's going to happen here, I think like kind of
saying out into the, into the void, like, we need somebody to, to step in and like, solve
this problem for us.
Like, just, just take that step, like, show some leadership and say, we firmly believe
that this is a threat to humanity.
So we're going to take the following steps and then others might not follow, but others
might.
Just don't know, I think on the, on the sort of president side, I think what they're
worried about is that it's any kind of slowdown in AI is going to lead to maybe the stock market
not being quite as robust, maybe some of these companies that are, you know, sort of propping
up the stock market in the tech space are going to have, you know, less revenue and that's
going to have an effect.
And so the president and people don't want to see anything like that happen financially.
And so they're worried about, they're also, as you pointed out, worried about China.
So there's lots of, lots of reasons to kind of take the approach to say we can't slow
down, lots of bad things will happen in the short term if we do, but I'm kind of with
you.
Like, if you think this is a threat to humanity, take some leadership and do something
about it instead of just like calling on somebody else to do it.
It actually reminds me of right before the 2008 financial crisis, there were a couple
of banks, not many, who were basically saying like, we need to be regulated.
Like we, the subprime mortgages that we're selling, we're making a ton of money, but
this is quite dangerous and we need like the government to step in and to like stop the,
this chaos and make it so that we can't, you know, do these things and have some control
over what's happening.
Of course, there was no regulation at the time and then when the stock, with the housing
market crash, a lot of people were in big trouble.
But those same banks could have just said, we're getting out of the subprime mortgage
business.
We're not going to, we're not going to do this.
We think it's bad.
We're going to try to convince other banks to do it stop as well.
Instead, like they called for regulation, seems like, you know, a throughopic is kind
of doing something similar here.
Just slow down.
Slow down, take some leadership.
Yeah.
So I just want to give anthropic a little bit of credit here.
You know, when I said they're not taking unilateral action, but they are doing some things
that they feel they can do on their own, on their own.
And so that, uh, amade lays out a three, three part plan, part one of which they're going
to do regardless.
And then part two and three will require some cooperation.
But the thing that they're going to do anyway is, they're calling it embedded evaluators.
And the idea here is, they're going to bring in an outside organization that does AI
evaluation, AI safety evaluation and alignment evaluation for a living, bring them in, have
them sit in house, give them access to all the stuff that they're doing in a way that
they haven't before.
So, so radical transparency when it comes to what they're up to and give them like a check
on their own internal controls, which the implication here is that those are compromised
by the fact that they are trying to be open AI to market for some of these things, right?
So they, the people working on safety at Anthropic, despite feeling like there's a greater
than 10% chance we're all going to die from this stuff, those people.
maybe, I think they probably deeply believe in what they're doing, but they also are subject
to some pressure from the business that they're in, and so you bring this outside organization
to maybe give you a second opinion when things are getting out of control.
Part of what's going on here is that in the past, when these companies have all done
their own safety evaluations, there's some sense that maybe they were selective in what
they were both evaluating and choosing to publish about what they evaluated, and that
there's something to be said for having someone involved in the process who can say, "Why
aren't you talking about this aspect of it, which seems concerning?
Why aren't you publishing about this?
Why aren't we measuring this other thing?"
And so having that outside perspective, I think, is valuable, and it's hard to find fault
with that.
It's also hard to see how that results in, like, a dramatic change, but I guess it's a
good first step.
The other things that Amadeus calling for in this post will require cooperation.
So the first is pacing within democracies.
So I think the idea here is, like, let's coordinate between us and open AI and maybe Google.
And then the third thing is pacing globally.
That is, let's coordinate with China, basically, but also the EU, and trying to reach agreements
about how quickly we should go.
Can we put some kind of a speed limit, he says, on the rate of recursive self-improvement?
So are you familiar with this term, Marcus, recursive self-improvement, or RSI?
I am not, Jeffrey.
Okay.
So this is what these big AI companies are mostly worried about.
While it is human beings who are developing the new versions of these models, we are kind
of naturally rate limited, because human beings have a certain capacity, we can only do so
much, and so we're not going to be able to expand the capabilities of these models.
We're going faster and faster, but the way you go real fast, and the way that is kind
of on the horizon for these models, is called recursive self-improvement, where you have
the model build the next model.
And the concern is, once you hand over that process to AI, we're off to the races, and
things are going to move very, very quickly from there, and also without the element of human
control that we had before.
And so can we agree with China to place the limits on the rate of recursive self-improvement,
so that we're avoiding the situation where AI kind of runs out of control because of
these competitive pressures?
And Amadeh explicitly uses an arms control analogy here.
He talks about the salt treaties, where the number of missiles were capped, particular
classes of missiles in Europe were capped, so as to avoid a situation where things run
out of control, but they still have enough so that we can threaten the other side.
And the idea is, well, you can do enough RSI, enough recursive self-improvement that it
is, keeps everybody moving, and you're not going to give up any competitive advantage
you had, but if we all agree to speed limits, then we can kind of keep things from spiraling
out of control.
Yeah.
I mean, that sounds good.
I mean, it's interesting in that, you know, the timing of this is kind of fascinating,
and that Xi Jinping is coming to Washington, I think next week, in September 24th for a
state dinner and kind of a summit, Trump is obviously, you know, publicly, at least on
the record, that this is, we don't have to worry about any of this, and so we're not
going to put any kind of regulations on AI.
But if the rest of the sort of world is worried about the development of this recursive, you
know, improvement stuff, and we need, we sort of recognize the need for global cooperation,
this is kind of an interesting time to be having a visit, and maybe it's possible that
Xi and Trump will be able to talk about this in some form or fashion.
It's striking to me because I think China, you know, I can't imagine that China doesn't
have similar concerns.
It's like I think, you know, in the United States, a lot of the discourse has been about
like the end of humanity, global catastrophe, et cetera.
I think China's concerns might be slightly different.
I think they're probably more worried about, you know, what will happen in terms of social
stability and the regime and things like that.
But it's not like they don't have concerns about AI as well, and so sometimes I think
we think of the race, you know, it's like, oh, the first one to, as Trump said, the
first one to win AI wins the whole thing or whatever that, I don't even know what that
means.
I think the race to the top, it might be the case that actually, you know, China has
some incentive to say, let's not think of this as much as a race anymore, and like let's
start like putting some controls on this precisely because we don't want to see what's, you
know, we don't want us to have AI take over either for different reasons.
And so if you think about it that way, you know, we're back into this classically cold
war, US Soviet Union nuclear weapons security dilemma where like in either side really
wants to see this conflict happen, but they all, you know, they also have reasons to keep
up, you know, developing missiles or developing technology because they don't want to prevent
or they don't want to see the other side, you know, quote unquote win.
It's interesting too because like the US is approached to China that I've heard, you
know, in terms of like limiting China has been like, we're going to restrict their ability
to get, you know, chips, we're going to restrict their ability to get these technologies,
which you might view as sort of like a defensive maneuver because you're like, well, that's
going to limit, you know, China's ability to do harm, they're going to see it exactly
the opposite, and they're going to look at it in terms of like the United States trying
to win the race and prevent China from developing its capabilities.
And so, you know, it's like this classic kind of prisoner's dilemma type situation where
like, and it's a prisoner's dilemma too, and for all big open AI because they're like,
we want to kind of cooperate, we want to talk, we want to have some type of coordination
here, but our incentives are such that it's going to be very difficult for us to do that.
So I'm thinking that this next week's summit, you know, it might be a good opportunity
to at least start some of these, you know, high level discussions about what this might
look like, and if Trump is, if Trump has beliefs that he's, that are not the ones that he's,
you know, talking about publicly, which might be aimed at the stock market and CEOs and
things like that, and he actually does share some of the concerns that the AI experts share,
then this might, you know, turn out to be kind of not Geneva 1985, but like an opportunity
to at least start a conversation about what this, this should look like.
Monday makes the point that we don't have to successfully reach a formal agreement with
China in order for this discussion to be useful, that changing informal norms around these
issues, sharing information about incidences of misalignment, of security, prep best practices,
that is low hanging fruit in a sense, and again, it reminds me of this CSBM confidence
in security building measures, literature and debate in the 60s and 70s and 80s, when
we were looking for things that are an area where we can agree.
So the US and the Soviet Union are not agreeing on the big picture stuff in this time period,
but there are still things we can do together to reduce the risk of things spiraling out
of control, where it's clearly in both sides best interest.
So like the hotline connecting the White House and the Kremlin, it's hard to argue that
that's bad for anybody, that doesn't give anybody any kind of competitive advantage.
It may not be particularly useful, but it's something you can do.
And from anthropics perspective, there are things along those lines that we could reach
an agreement with China on.
Having said that, Marcus, what do you think are the chances that we come out of this meeting,
the Trump G meeting with an agreement on AI?
I would say the chances are very low, probably lower than AI, the chances of AI, you know,
kills all of humanity, which is to say like less than 10%, probably less than 10%, but
not zero, right?
Not zero.
And the only reason I say that is because I do think that this is probably on the mind
of Xi and Chinese leadership.
I know it's on the mind of Trump because he's been talking about it.
He's been talking about it in a way that suggests he's not interested in having any type
of deal, but you know, like what strikes me is on the nuclear side in the 1980s, like
one of the first steps that was taken between Reagan and Gorbachev was just a joint statement
that essentially said like a nuclear war can't be won, and so therefore it should never
be fought.
And so the very least, both sides were saying publicly like, look, we understand that neither
of us wants to have a nuclear war.
Now looking at this now in 2026, you know, back in the 1980s, it sounds like, you know,
so obvious that it's not even like interesting, but at the time, like it actually was an important
statement.
And so if Trump and Xi can kind of like come to some type of statement where they say,
like we think that we need, you know, sort of ethical dimensions to AI or we need to like
think about cooperation, not competition in AI, or we can agree that it'll be bad, like
the equivalent would be we agree, it would be bad if AI destroyed the world, okay?
But if they just said that, I actually think that that would be helpful.
It's sort of like starting a little bit of a conversation.
There is some evidence I was looking at a couple hours ago, like there are some sort
of track 2.0, maybe you could call them track 1.5 sort of discussions taking place between
United States and China about AI.
Again, similar to what was happening during the Cold War between United States and the
Soviet Union at various points in time, where you had sort of like NGOs and different,
you know, sub, you know, subgroups within both China with the United States and the
Soviet Union kind of having discussions and trying to create frameworks and things like
that.
So there's a little bit of evidence, I guess, that there's, you know, sort of some cooperation
at very low levels between the two.
But I think just like some type of statement like that, you know, it really would really
be helpful in kind of showing some leadership in the international system that this is something
we take seriously and we're going to try to address it in some form.
statements can be really hard to come by. We try
for I think a year and a half to reach an agreement with China culminating in the 2024
statement from the White House after a meeting in November or 2024 in Peru, quote, "the
two leaders affirm the need to mean human control over the decision to use nuclear weapons.
The two leaders also stress the need to consider carefully the potential risks and develop AI
technology in the military field in a prudent and responsible manner.
It took like a year and a half to get China to agree that we shouldn't give the nuclear
weapons to the AI, right?
And I mean, that seems like an exceedingly low bar.
Anything that required them to actually do anything, quite a quite a bit higher bar.
So it'll be interesting.
It'll be interesting to see.
I mean, one argument you can make is that China might actually see it as in their interest
to play into pacing the frontier.
Every time I say that, I'm going to just. I can't keep it.
I could tell the look on your face.
The listeners can't see this, but they'll look on your face.
It's unpleasant for me to use the word pacing in the back.
He struggles every time he orders the food.
I'm going to adopt the lingo of our AI tech overlords and just say, like, if China might
see it as actually beneficial to at least contribute to the idea that the US should be pacing
the frontier, whether that means that they are taking concrete steps themselves.
I don't know, but they might lend some kind of moral support to this such that anthropic
and open AI may be more importantly, the US administration feels like they might be
able to move forward with some efforts to slow down progress knowing that China isn't
like nipping at their heels.
I don't know.
You know, I think it's interesting the extent of the nuclear analogy being used in all
of these discussions.
We talked about this before.
So Amade makes this explicit.
He's using arms control analogies and nuclear missiles to talk about potential progress
on these issues.
I've already brought up the confidence in security building measure analogy multiple times.
And people often will talk about the nuclear non-proliferation regime and maybe we could
use some of the lessons.
This is a thing that international agreements to control the spread of a dangerous technology.
Maybe there's some relevance here that we could apply to the world of AI and Marcus
as the author of a best-selling book about the nuclear non-proliferation regime, I frequently
amass these questions.
I'm sure you will.
How does this apply to AI?
So this is a very like popular analogy to use.
I think a piece of the nuclear story that we don't often talk about in the context of AI
is the discussion of nuclear safety that became like a big deal in the 80s and early 90s.
And there's a great book by Scott Sagan, the political scientist Scott Sagan called
the Limits of Safety.
And this is a book that examines how comfortable we should feel about the lack of any unintentional
nuclear use up to the early 90s when this book was written.
And the conclusion of this book, and there's been a lot of work since this book, that looks
at kind of the safety of our nuclear enterprise, but it basically concludes that it's mostly
luck.
The lack of a nuclear accident is far more due to kind of lucky breaks here and there
than it is to good design of the nuclear apparatus.
And it's kind of interesting, I mean, this work on nuclear safety and nuclear command
to control kind of breaks down Sagan's work at least breaks this down into two broad
theories of how nuclear safety works.
So on the one hand, you have high reliability theory.
And this is the idea that we get kind of from commercial aircraft and actually is interesting
that Amide actually used the analogy to commercial aircraft.
We can have safe systems, and in some sense, a commercial air travel is a miracle, right?
Because we have so many flights all the time, and there's really very, very rarely almost
never any serious accidents in commercial aviation.
And this is a triumph, right?
And the commercial aviation industry is described by this high reliability theory.
And the idea here is that we can prevent accidents through good organizational design and management.
We can design out accidents, we can design in safety.
Safety in these organizations becomes the main priority.
So when you go on an airplane and the flight attendant say we are here first and foremost
for your safety, that is a design feature, right?
I mean, like the whole system is designed to emphasize safety and we create redundancies
and a culture of reliability in order to enhance safety.
So the flight check before the flight, the pilot goes through the checklist twice, make
sure everything's working, and then check it again.
And those redundancies help promote the safety of commercial aviation, right?
So this is kind of one theory, this high reliability theory of how you create safe organizations.
Another theory is called the normal accidents theory.
And in this way of looking at the world, accidents are just inevitable features of complex
systems.
That if you have a complex enough structure, you have these emerging properties of these
crazy software programs that we're designing in the artificial intelligence world, that
accidents, unintended outcomes are kind of inevitable.
Things are just too complicated for us to get a grip on such that we can design out safety.
But in normal accidents theory, we should think of safety as one of a number of competing
objectives of the organization.
So it's not really all about safety, it can't be all about safety in the nuclear enterprise,
it has to be also about, you know, making sure we have a functional nuclear deterrent,
making sure we can get our missiles to where they're going, right?
Like all that stuff is a priority in addition to safety.
And when you think about it this way, redundancy can really cause accidents, right?
It's not just to prevent accidents, but redundancy causes accidents by increasing complexity and
increasing risk-taking because the participants in these organizations know that there are redundancies
back in them up.
You put one of those padded mats in the gymnasium and the kids are willing to take double
flips off the bar because they know they have the padded mat there to back them up if they
fall.
That's a good safety feature, but it also encourages risk-taking.
That's it can also have a great article titled The Problem of redundancy Problem, one of
the best titles in political science, that really lays out this risk.
And I think there are some clear parallels between these ways of thinking about nuclear safety
or organizational cultures and design cultures for safety and what we have going on now in
the AI world where we have these complex systems.
And we are trying to design systems that will be safe, that have these redundancies that
have these double checks to make sure that things don't spiral out of control.
And you know, our initial response when we see an accident, when we see open AI hacking
hugging face or whatever is going on is to look for ways to improve systems so that this
can happen again.
And that's part of what Amade is talking about, like we're going to have better operational
security.
We're going to have outside organizations monitoring us, but this idea of normal accident theory
tells us that we should really consider the possibility that as long as these systems
are tremendously complex, as long as they are in some way human designed, it might not
be possible to fully design out the risk of an accident.
And in that case, just adding more redundancies, more double checks, more safety, outside experts,
it might not help, we may need to look for some kind of resilience to safety risk rather
than trying to avoid accidents altogether.
So I think there is a parallel between this idea of piecing the front-to-face to it.
Oh, man.
There's a parallel of this idea of like limit slowing down our emphasis on getting the best
new thing.
And some of the lessons from this nuclear safety literature, Sagan recommends creating extra
time and slack in the system, in the nuclear system so that if there isn't accident, it
doesn't spiral out of control and lead to global thermonuclear war.
So if we had an accidental detonation, for example, there might be time enough in the systems
to figure out that it was unintended and not respond with everybody's full nuclear arsenal.
So there are doctrines like launch on warning in the nuclear sphere where you see nuclear
weapons in the air, you launch yours, that's super dangerous if those are a false alarm.
Light if those weren't actually launched because now you've just created a global thermonuclear
war for nothing.
And so we need to de-alert our missiles, or use Sagan and many others, so that we're
not doing that kind of thing.
We're leaving enough slack in the system that if there's an accident or an unintended
thing that we can root it out.
Another example is physically separating nuclear components.
So one thing you can do is like disassemble the nuclear weapons, store them disassembled.
That way they can't accidentally be launched and detonated, they'd have to be put together
first.
So South Asia, they've made a use of that as a useful check on the potential for accidents.
And you can imagine similar things in the AI sphere where we separate the agent from
the target that it's trying to get at in one of these cyber security exercises, like
ways to decompose the AI agent so that they don't have the same risk of getting out of
control.
And this is a useful analogy that most literature isn't talking about.
We talk a lot about like global governance and what nuclear can tell us about that.
But this idea of nuclear safety and command and control I think is also potentially useful.
Yeah, no, I agree with a lot of that.
I mean, I, I worry that with AI, it's slightly different in the sense that like trying
to come up with that time, you think about like the time piece, like I could see that in
nuclear kind of instances where you don't have the automated response.
It seems like. the concept of time or temporality if we want to be fancy about it with AI is just so
on a different scale that's hard for me to see like how you could even you know build in a
situation where the AI would stop like you could say like okay AI before you do something really
stupid stop call of a human being let the human being you know that's that's what's going on
like it's hard to it's hard to see an AI who's moving at like the speed of light or electrons
you know bouncing around like understanding what time could even mean to that context but I
agree with a lot of this I mean this is you know this there there is an organizational management
there's like a literature on this type of stuff too where it's kind of like you know we can't
we can't make everything perfectly safe like that's that and that's actually fine like what's
it's just inevitable that we're gonna have accidents um and it's like thinking about what are the
various you know procedures that you can't put in place that makes sense in AI like that's
going to be the next frontier of this like that that seems like a good paper Jeff we should work on
this yeah I mean there's actually another another piece of this command and control story from
nuclear weapons that I think actually does apply to AI as well and that is the idea positive and
negative controls Peter Fever talks about the the always never dilemma in nuclear weapons and the
idea here is that when the present orders a nuclear attack it should always work and systems that
ensure that we call positive controls so this is about reliability so if you say launch the nukes
it they always need to be a functional nuke right we never want to have a situation where
the nuclear weapons don't work those are positive controls but if the president hand or has not
ordered a nuclear attack nuclear weapons should never detonate right that's the never part of
the always never dilemma and systems that ensure that we call those negative controls that's about
safety and security and so you can think of a similar kind of um dichotomy when it comes to AI
agents and models where we want them to be capable of doing the things we need them to do in
their in their day-to-day job of sending emails to random people from from Marcus but we don't want
them to send an email that they weren't authorized to send or accidentally hack into the hugging
face system and so we need the positive controls that make sure that they're reliable and they're
intended purpose and then we need the negative controls that make sure that they don't go off and
do something misaligned and it's interesting to look at the kinds of tools we have in the nuclear
world to address some of these things so in terms of positive controls ensuring reliability
we have things like early warning systems communication procedures that let us know when it's time
to launch we have something called the stockpile stewardship program which makes sure that our
old nuclear weapons still work right if they were if they were to be attempted to be detonated
and you can imagine some parallels to AI in in in this sphere like what is the equivalent of an
early warning system that tells the AI that it needs to be working and we have those kind of
communications processes available and at the same time there are things on the negative side like
permissive action links which is like the security control that makes sure that you must you're
authorized to use a nuclear weapon if you try to use it having clear authorization okay i the
AI I'm about to do a thing that looks like hacking so one of the the issues that anthropic identified
in its hacking incidents was that the AI was told it was in a simulated environment but it wasn't
it really was real they screwed up they gave it access to the internet and the AI is like this
looks like a lot of the internet this doesn't look like a simulated environment pretty cool this
is a great simulation i'm gonna go for it even though like the AI kind of knew that this was the
real internet right but it took some actions and it convinced itself and you can see this in the
in the terrifying transcripts that anthropic released of the like internal monologue of the AI
model that it knew that this was really too good to be true for a simulation but it kind of went
for it because that was its job but if we had like a permissive action link here where there was
a separate authorization every time the model wanted to access what it saw as some kind of internet
system or in network system or codes or some kind of physical security or a two point procedure
which you've seen on in movies about nuclear submarines where you know you need two people to
detonate the nuclear weapon those are those are systems that are negative controls that you could
imagine scenarios where they could put a check on AI agents that wanted to go off and do something
dangerous yeah i mean i guess i don't know how optimistic i am but any of this Jeff terms of
like putting in different controls but i like the idea that we can learn from like we don't have
to reinvent the wheel right we're not starting from scratch like we can we can you know we have the
ability to kind of borrow lessons from organizational management borrow lessons from the nuclear
kind of stuff borrow lessons from other you know like you know the aviation industry and how they
have done things and kind of piece together like maybe a way to tackle this problem um i don't think
trump and gier like up to that challenge right now to do that but maybe like one of the points of
this is that if we can find some intention to have cooperation on this topic you bring together people
in workshops and you know track 1.5 and 2.0 you know kind of diplomacy to sort of start thinking
about what these types of frameworks might look like and borrowing from these other you know industries
and areas maybe we'll actually remove them need a little bit but um it's nice to know that like
this problem has been solved to a certain extent i mean to say his point is he can't we solve it
completely right exactly what say you would say it's like we have not solved this problem right but we
can build in ways for the uh for things to be more resilient and i think that that is an interesting
thing we can solve it to the best of our ability to solve it like you know like we can sort of like
make it as safe as they can be safe um and with the big sort of like you know footnote that
it's never going to be perfectly safe and like we're always going to be potentially in danger
from this technology because it's just too complex and as it if and thus it goes away completely
as long as it exists safety's not going to be you know 100% guaranteed but it's an important mindset
change a shift between these two ways of looking at the world right because i think our natural
inclination is to see a problem and think what went wrong here that we can fix and i've heard
this response to our podcast about the hugging face hack and to all the news reports here which is
like what what should anthropic and open AI do differently so that their models are not kind of
accidentally on the internet and i think there's a school of thought here uh you know and i think
yes they should try to you know make sure that they fill any holes but there's a school of thought
here that that's the wrong question to be asking it's if there's this idea that we're always going to
have safety incidents as you have this very complex technology that's developing very quickly
even if you slow down you're going to have those safety incidents and so we need to think about
how do we design a system that is resilient to those safety incidents that can catch them with
enough time to that they don't do any damage that can eliminate any risk that we get involved in
some international conflict because a open AI prototype model decided it needed to hack into the
Kremlin to find out what some capability was on the Russian nuclear side right so like there are
things we can do that might insulate us from the result of these potential safety incidents
which i think is a useful perspective right it's not the only perspective but i don't think it's
not enough to say let's imagine all the ways the AI could be unsafe and fix those because we're
going to fail to imagine all the ways instead we need to think about how do we respond when there
is the inevitable safety failure that we're going to see great notes ended on Jeff the inevitable
safety failure that we're going to see greater than less we should do a little a little prediction
markets in the next 10 years what are the chances that AI kills us all all humanity next 10 years
yeah that's a long time that's a long time i'm gonna go i i'm trusting the the experts on this one
i'm gonna go between five and six percent five and six percent yeah i think an interesting question
is like how how also do i see it happening you know like just a couple different ways right it's like
AI just takes over and becomes evil and makes us all like they doesn't like us anymore like it's
not aligned with us at all or humans gave it like some some thing that we you know think is a noble
cause and it decides that we humans shouldn't exist but because of that cause like it's going to
figure it out probably more likely is like we you know humans use AI to kill one another and that's
what that's right it builds a bioweapon and yeah yeah i think that that's that's my number one
right now if i'm like if we're drafting ways the AI is going to kill we should do an AI um
killing all humanity draft in the next episode where we pick uh the different ways we could all go
maybe maybe around around Halloween we can do that but i will say if there's a kalshi bet for
the end of all humanity uh that's one that how do they tell well i mean would you bet on it because
like we're never getting your money if you like who's betting on the on the yes side of that well
probably as you as you watch the the sort of like skyfall you'll know that you won you know
i'm gonna have the psychic the psychic benefit it's the same thing with like mutually assured
destruction it's like there's hundreds of missiles coming at us like we're all gonna die anyway
but we want to make sure they're dead too you know what i mean like it's just sort of like
the psychic kind of the emotional payoff the funniest part doesn't matter i'm not seeing a particular
market come up when i search AI kill us all on kalshi maybe that's just a failure of searching but
maybe maybe they'll aren't taking best on that right now i also wonder like what AI looks like
on other uh like planets like out in the in the universe like what is this presumably you know
life exists elsewhere, like the scientists say,
probabilistically it does.
And if life exists, then maybe AI also has been created
in some other form.
I wonder what kind of safeguards other civilizations
that put on AI if any?
Or maybe that's the reason we don't see these entities
anymore, because they created AI,
AI destroyed their civilizations.
And so therefore, there's no evidence
of those civilizations anymore,
thus resolving the paradox of why we have found life out there,
because they've all died because they all created AI.
- You can't get through a podcast without bringing up aliens.
It's really constructed in aliens.
I just said life out there.
I mean, maybe the aliens maybe they're not,
but-- - Well, I didn't mean it in a derogatory way.
I just meant like-- - No, of course not.
- You didn't mean to be the green people with--
- Entity is not from Earth is what I meant.
This has gone off the rails.
I think I'm going to call it.
Marcus, thank you for joining me today.
- Jeff has been a pleasure.
We covered many--
- Maybe an alien sweatshirt in the shop,
the cheap chalk shop?
- I'm going to work on that. - Because I talk about,
I talk about them all the time.
- You do bring up aliens a lot.
- Yeah. - All right.
I'll see what I can do on that.
Folks, you should check out our online stories
at cheaptalk.shop.
Pretty soon there'll be some kind of alien themed
merchandise there, but also send us an email.
[email protected].
Let us know what you think about this whole AI
into humanity, debate, or the alien situation,
either one, interested in either one.
You should subscribe to this podcast.
You can listen to us and subscribe in Apple podcasts,
Spotify, or wherever you get your podcasts.
Unless you get your podcasts on YouTube, in which case,
don't go there or not there.
Maybe it's going to be someday.
Marcus, thanks for joining me again.
Folks, we'll see you next time.
[BEEPING]
Can we do a little sideline on how
we shouldn't let the tech industry use words?
Some linguists somewhere should do a study of the way
the tech industry takes words that exist and then
changes the way they're used so that they mean something different.
So everyone is talking about pacing the frontier.
This is the new term of art.
Do you know what that means?
When they say pacing, they don't mean
like pacing around the room, walking around the frontier.
They mean limit.
Limit the speed.
I know exactly what they mean.
So to use a running metaphor here, Jeff,
if you have a passer, like in a marathon,
that is somebody who's not running the race on their own.
They're running for you.
And their goal is to run a specific pace.
So that as the runner, the one that is in the competition,
you don't have to pay attention to how fast you're
running.
All you have to do is stay with that person who's your pacing.
And that pacing has all kinds of different data
on there, sometimes on a physical piece of paper
that they're looking at.
And they can say, OK, we need to speed up now.
Let's see, you're trying to go for a world record.
And it's a marathon.
And you're trying to break two hours or two or five or something
like that.
You have to hit certain milestones over the course of 26 miles.
And to do that, your pace has to be pretty precise,
because you're on a needle, right?
You're like on a knife's edge.
If you go too fast early in the race, you're going to fall apart.
You go too slow, you're not going to be able to do it.
So the paceer kind of knows the optimum--
or should know the optimum pace at any given moment of time.
And your job is just to follow that lead,
whether they slow up, slow down, speed up, or whatever.
So my interpretation is, this person from Anthropic
is saying, we need a paceer.
We need the equivalent of a paceer in a marathon
so that we don't go too fast and burn out
and never get to the finish line,
because we blew up at mile 13 or whatever.
Yeah, so maybe for non-runners, the analogy
would be like a pace car in a marathon.
There you go.
Where it's like setting the pace out of the gate.
But when I think of the pace car, I think the pace there
is like setting the speed, that's what it's doing, right?
When you say the pace car is there,
the pace we're talking about is the speed of all the cars.
Right.
Pacing the frontier is a different use of that same word
to mean we're going to limit the speed of models
at the frontier of the technology.
It reminds me of like, there are a whole bunch of words
that the tech industry has morphed from their common usage.
Like, or maybe invented.
Like, have you heard the phrase upskilling?
Isn't that, I know, oh, I know upselling,
like when you're checking at the Marriott
and they're like, did you like an upgrade?
Would you like a sweep for $50 more?
That's different, right?
Upskilling is like, you don't have a skill you need.
We're going to upskill you so that you then have the skill.
What would you call teaching?
Oh, you're saying upskilling.
I think you're saying upskilling.
Skill, skilling, I see.
But I don't know what that is either.
Right, so it's like teaching, right?
We're going to learn you some knowledge, my friend.
We're going to upskill you.
Okay, there is uplift is another like AI tech word.
You know that one?
That I know what that means.
Well, uplift, I want to uplift you with my good humor.
And I'm an uplift big presence.
That is a use of the word uplift
that normal Americans would understand.
Yeah, the tech use of the word uplift has to do with,
we're going to increase the capability within a group.
So we need to uplift the workforce, Marcus,
if they're going to be able to compete in a world of AI.
So it's not about your feelings of joy.
It's about your level of capability.
We might have called that training
or improving the capabilities of, right?
But we're calling it uplift here.
So this is like the industry kind of capturing words
that had one meaning, so words that people know
and associate with one particular meaning
and then like add on an additional meaning.
So when that they use them, the people in the know
are kind of like, oh, I know what,
I know what he's talking about with uplifting.
But if you don't know, it just kind of sounds ridiculous.
I feel like the business schools, didn't they?
I mean, like this is, yeah, it is like synergy.
And, you know, yeah, for business schools,
I feel like it's more euphemisms
to avoid the real negative connotations of what they're saying.
So like, we're going to write size our industry instead
of we're going to fire these or come back.
So like, that I kind of, I can't understand.
- You can pass on that.
- But like, I was listening to the Apple's recent
Apple had a big quarter.
They had a blow out quarter, right?
That they say like, but blow out to like,
you know, people in the auto industry
means something very different.
- That's good.
- To parents of children in diapers,
blow out means something very different.
But for Apple, it was like a blow out quarter
and that was like, that's like just a business linko.
- So like, it's a good blow, that's good.
- That's good.
- They like blew the socks off expectations, right?
But it's, I don't know.
I just find this whole thing kind of interesting.
So when amade is like, is the title of his post
is "Pacing the Frontier"?
- Right.
- Is that, or we must pace the Frontier?
Is that language that anyone who's not in that world
knows what that means?
Like, you can kind of get there.
But it's just, I don't know.
I just think it's a strange use of those words.
- To use a different metaphor, when I think of that,
when I think, you ever seem like, I don't think they really
do this anymore, but dog racing.
Like, they have these dogs that would, to go faster,
they put like a fake rabbit or a bunny.
- Right, and it runs around the course.
- Yeah, and that thing was like going like super fast
so that the dogs would run faster.
I guess if the rabbit was slowed down a little bit,
then the dogs would run less fast.
I mean, they would catch up to it, I guess, as the problem.
But he seems to be saying, we need to take the rabbit
or the bunny from the dog racing and slow it down.
Like, we need to like, voluntarily and intentionally say,
this thing should not be going around that fast.
We need to, we need to slow it down.
That's, I mean, as a simple person,
that's the way I took his argument.
Like, just let's slow, pump the brakes a little bit.
Let's slow down.
- Right, well, I think that's exactly what he means.
I just think like the word choice is weird.
- It's weird, it's weird.
Like, there's a header in this, in this post.
Why pace?
Question mark.
- Well, and then he's like the idea of pausing
or slowing a eye has been floated previously.
But that's what he means.
Why slow down?
Why pause?
But instead he says, why pace?
And it's like, he doesn't want to say the word pause or slow.
He prefers the word pace, because it's like,
you're still moving when you're, I don't know.
It's like, it's just not the usage of that word
that I'm familiar with as I say.
- Yeah, I see what you're saying.
I see what you're saying.
- As a lifetime speaker of English.
And someone who I think is a, you know,
reasonable vocabulary, this is just not the usage
of the word pace.
And I think maybe you think some of it might be like,
AI speak too, that it's not quite the right word
for the moment.
And, but AI sometimes does that,
gives you not quite the right word.
- Well, I was thinking like, did, first of all,
did AI write some of that quite possibly it did?
And also if you're around AI speak all the time,
and like that's just what you're dealing with
on a day-to-day basis, and you're not reading
what humans write, maybe you just started
adopting that diction.
You start, you know, it's like all you see
is AI generated text, and so your own text
becomes AI generated.
And actually, I've seen this, this criticism
has also popped up in, you know, other areas
where people are like, well, you know, students writing
a paper, this sounds a lot like AI, maybe AI wrote it.
That's possible.
Or it's possible the student uses AI a lot
and has started speaking and writing like AI does
because it's learned kind of that style.
And maybe it's possible the ideas are actually the students,
but because our text, you know, resembles that of an AI,
you know, chatbot, we assume it's AI written.
So it could be very well-being this guy
is just like around AI so much that he starts to, you know,
talk and speak and write like a chatbot would.
- I will say that I just plugged this post into pangram,
and it comes back as a 100% human text.
- Well, he's also smart enough to--
- He's right, like, Dara's not here.
I'm not gonna, not gonna post something that's gonna come out badly on pan gram, but I don't know I just thought that was interesting.