Building the Cloud for an Agentic World | AWS CEO Matt Garman
56m 18s
AWS has evolved from a startup-focused cloud service to a central player in the AI infrastructure landscape, driven by massive demand for compute, especially GPUs. The company has strategically maintained capacity for early-stage startups, recognizing them as the future of enterprise innovation. As AI agents increasingly write code and manage infrastructure, AWS is building new services and security models—like compute sandboxes, fine-grained permissions, and agent core—to enable safe, efficient agent workflows. Despite intense competition and rising GPU demand from frontier labs, AWS ensures balanced allocation by prioritizing ecosystem growth over short-term gains. The company is also investing heavily in custom silicon, including its Graviton processors and training chips, which deliver superior performance and cost efficiency. Enterprise adoption of agents is cautious, with leaders concerned about autonomy and data safety. AWS is addressing these concerns through robust guardrails, evaluation systems, and secure environments—such as AWS Bedrock—where data remains on-premises. Additionally, AWS is expanding support for open-weight models and fine-tuning, empowering enterprises to build custom AI models. These developments reflect a broader shift in how cloud infrastructure is being reimagined to support autonomous, secure, and scalable AI workloads.
Agents of workflows tend to perform better on AWS than anywhere else.
Compute sandboxes, gateways, agent permissions versus people.
A lot of those are things that we have built and are building and thinking actively about.
The top frontier labs gobble up every available GPUs.
At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow.
How are you thinking about balancing?
From the very beginning, if we launched AWS, startups have been the lifeblood of the core of what we do.
These are the innovators that are at the edge of technology, understanding what's possible.
We're very intentional about keeping capacity available for the startups.
We recently announced we're going to be buying 2 million video GPUs over the next couple of years.
Your graphics, the city has worked 200 to 220 billion dollars for 26.
And we don't anticipate slowing down any time soon because the demand is just massive.
We pull the debate around AI extension risk and the hugging phase attack.
What are CEOs asking you about all these things?
Amazon plans to spend 220 billion dollars in capital this year.
And AWS CEO Matt Garment says they don't anticipate slowing down any time soon.
In this episode, E16Z's Raggle Raggle Rum sits down with Matt to unpack what's driving one of the largest infrastructure buildouts in history.
And how AI is changing the cloud itself.
They discuss how AWS decides who gets scarce GPU capacity.
They're shifting bottlenecks from power and ships to memory and construction.
And the company's bet on custom silicon like training.
They also get into what changes when agents increasingly write code and manage infrastructure.
And that shares what AWS is seeing inside its own teams.
Where agents now write code and engineers increasingly manage teams of agents.
Accelerating how quickly new products can be built.
Welcome to the pod, Matt.
What a time we're living in.
So I have a lot of topics to talk to you about.
Awesome. Thanks for having me. I'm excited.
Yeah, absolutely.
So let's start actually right from the beginning.
So you were the first GM for EC2 and there was 2006, right?
And today you guys are what, 160, 170 billion revenue?
Yeah, about 169 hours every billion, yeah.
Yeah, growing 30, 37%, 37%.
That's 37% at 169.
Yeah, crazy.
There's a ton of updates.
It's interesting to think about it from day one, when the first dollar revenue.
But yeah, the interesting thing is we're still at the early stages of what the business can be.
And what the opportunity is for customers.
Most workloads, there's a huge amount of workloads still that live on prem today.
And the amount of compute that people are doing every single day is more than it was the day before.
And so you see the tailwind from AI, you see the tailwind from migration into the cloud.
And the business has grown really fast and it's been a super fun thing to be a part of.
Yeah, it is, it is, I mean, there'll be writing history books about this and business books about this for a long time to come.
I want to touch on the on prem.
That's sort of boggles in my mind.
I obviously did my best to keep them there for a long time.
You told a lot of stuff on prem back in the day.
Yeah, yeah, I was trying to now get all of that to move in later.
Yeah, so we could talk about that later.
But so if you think about EC2 in the early days and you got your start with obviously selling to startups.
And today the AWS cloud service, I mean, as you reflect on the evolution, what what stands out.
Yeah, part one and then part two is talk about how this nature of how you serve startups has changed.
Sure.
Well, like you said, actually, it's funny.
So I actually interned for AWS in 2005 when we were first kind of as an internal project.
It was my business school internship.
And my project was actually to come up with an analysis of who we thought AWS would be most interesting to.
And the answer was startups, like, surprisingly.
And so from the very beginning of when we launched AWS startups have been the lifeblood of the core of what we do.
For a number of reasons, one is the value proposition is just so attractive.
What AWS provides to startups.
And we spend a lot of time and effort making sure that we're great partners to the startups, helping them not just provide infrastructure,
but also advice and how to get your company up and running and how to think about your architecture.
So it'll scale eventually and a whole bunch of things that we do for the startups.
But we also think that for us, it's just good business.
Because the startups today are the enterprises of tomorrow.
And it's an in-precise number.
But today, we estimate that maybe 30, 40% of AWS revenue comes from companies that was a one-time startup in AWS's lifetime.
So it's fun to see for us, to see the companies grow over time and to get bigger and become enterprises effectively.
And so that's why we invest so much to startups and why we pay so much attention to the brand new two people in a garage type startups.
And not just because of this is outcomes, but also because that's who we learn from.
Pushing our services to say what they'd want more of, what can help them go faster, what can help them achieve their outcomes.
And a lot of times they're pushing more than the banks or the healthcare companies or the governments.
The startups are the ones that are pushing that envelope and it really helps us to be better and make sure that
we're ahead of that wave where the banks are going to want some capability five years down the road that startups want today.
Yeah. Yeah. And compared from then to now, how have the startups changed and what they want from you?
Obviously, they all want a lot of GPUs and we can talk about that, but besides that.
Well, I will say there's a couple of things that have changed.
One is startups started at a much smaller size than when we first started, right?
They might have gotten $10 million of funding and they had a app idea that they were kind of slowly iterating on.
Now it's from day one, they're valued at a billion dollars. They have $200 million of funding. It's a team and an idea and all of a sudden they're worth a billion dollars.
Let's go from our offices to your offices. That's right. That's right.
And the size and scale and ambition of the ideas, I think, requires more capital. They're bigger to start with.
They're more expensive too, right? To go after kind of a training a model or doing something that a lot of the folks are doing today.
So that's number one, which is just the size that they start from is really big.
But I would say number two is there's some things haven't changed, right?
They actually are still thinking about, okay, once I scale, how do I think about an architecture?
How do I think about security? How do I think about performance?
How do I think about having all the capabilities that I need?
How do I think about setting up my IM setups so that when I have more than three employees, this thing is going to work in scale?
Which is a lot of times why they like AWS as opposed to just going to a NeoCloud or something like that.
They actually need all of those other security and capabilities that come around from training.
And so I think that's something that hasn't changed and is exactly the same.
I think the scale of which the ramp up is definitely different today.
I also think increasingly we're seeing teams that want a cloud that is great to work with agents and not just with people.
And I think that's where we've spent a lot of time is how do you think about exactly how do you think about broad scale?
How do you think about performance? How do you think about an interface that is a well-defined API interface that agents can actually easily traverse and work across?
And so that's also something that we both I think we're naturally set up to do well, but that we've actually doubled down on to ensure things like how can you start a database in three seconds and really get to some capabilities that agents are excited about.
Have you introduced any specific new services that are explicitly targeted at agents or people building agents?
I'd say what we've done is we've optimized some existing services so that they can both work for people and agents.
And so the answer is we definitely have some services designed for building agents.
So we have things like agent core and bad rock and pieces like that.
But when we look at just taking an underlying component like S3, which is where the vast majority of companies store their data and have their data lakes.
It turns out like a lot of the use cases for people and agents are similar, but you want to have and one of the things that we have in preview right now.
Or in beta, it's called AWS context that allows you to build a context layer so that agents can actually more easily find all of the data they want across your various data lakes.
And so whether you have your data stored in Aurora or you have it in S3 or you have it in somewhere else in AWS can kind of build this context layer where when people aren't necessarily going to access data in that way, but agents actually are happy to go across lots of those different things.
And so there are some services that we're building like that, but for the most part, you also care about what does latency look like? What does throughput look like?
How do you make sure that the underlying engine is actually fast, scalable?
They actually care a lot about tail latencies, which is interesting, where people are, they don't always care about the P999 S3 latency, like agents do carry block by that.
And so that's the thing that we've cared about for a long time and from a performance perspective, it's one of the things that really pop.
That's why agent core flows tend to perform better on AWS than anywhere else.
Yeah, obviously we see a lot of companies starting out and the very common refrain, of course, is that agents are writing all the code for them, right?
And agents are selecting the databases and the email servers and everything that you can name it, right?
So as there been a lot of thinking that obviously they read the documentation, as there been a lot of thinking on the AWS, like your feels funny to say, quote unquote legacy services that I've been around for a decade on how do you rearchitect them?
I think so, yes, I think there are some things that we thought about actually the core underlying building blocks, I think are in really good shape.
I think there's a usability layer that we're thinking about of how we make easier to use for we'll call out the very simplest use case you call out where somebody is just coding up an app really quickly.
And they ask it to deploy. We find the vast majority of our customers will go until they're coding agent, whether it's Ciro, whether it's Claude, whether it's codex, they'll say, I want to build on AWS.
Here's my credentials, here's my stuff. This is how on deploy it and the agents are great and they go do it. There are use cases, though, where you're just you're already not on a cloud, you're just trying something you say deploy.
And frankly, a lot of times that'll go to some of our partners that have an easier to use layer on top true love, by the way, too.
And we love our partners on those fronts too, but I do think that there are some things that we're doing
Where if you're brand new you don't already have an AWS you already haven't set up here
I am a couple of months ago. It was much harder to start up an account right it's been like this for 20 years
But you have to define your VPCs you have to define your I am roles all these kind of things which are actually super important things
Once you get to be large and what we hear from customers is like it was a hard trade-off because they know that they're going to want those
In months or years down the line, but for now they just kind of want to use the services not worry about that
And so what we've actually launched is if you go create a new AWS account today and we're slowly rolling this out
I don't know if it's actually full yet yet now, but you don't have to do those things you don't have to give it a credit card
You can sign up with your Gmail account
You've got to just all of those things are a handle and default behind the scenes and within less than 30 seconds
You're up and running and can be operating an AWS full AWS account
Which is you know for those type of systems much more what the agents want to be able to do because they don't want to go through all of that
Setting up your VPCs in your yeah, and and pieces like that
So there are some things where we're adding kind of some ease of use to that which we're quite excited about
I've seen some really positive feedback from customers on and we'll keep doing more things like that over time
The great part about that though is it's not like a simplest to get count and then you have to migrate
If you basically say okay, now I actually do want to scale into the
Or I want to develop an organization
I want to go and kind of fine tune some of the things you can easily come just you're already in a real AWS
Accounting you can actually then go and do all of those things later when you need them
And so there's no migration or move later and so that's part of the hard work that we really think about
Which is how do we make it not a choice for the customer but an easy on-ramp into the depth of features that we know customers
And startups are going to walk when they're scaling. Yeah. Yeah. Yeah. Got it. And
What has been
Some of the hardest things to accommodate as agents of techno. I mean you made a name
This the way you guys
Became the phenomenon that you became was serving developers. Yeah
And then infertimes
Now developers are substituted by agents and pretty soon infertimes are being substituted by agents. Yeah
Are there what has been some of the hardest things for you?
You as the largest service provider in the world
To handle that transition. Well, you know, like I think like I said much of our infrastructure was is
Yeah, but it's actually pretty well set up since I handle the scale
Which is good. I think there's been a couple of things which are interesting which are I think interesting paradigm shifts where
um, you could argue that some of our systems are
um
I don't want to say over-engineered, but um
As an example many agents want to create a database
Uh, do a little bit of work and then have the database go away
We really need that database to have five nines of durability exactly. Right. And so there's some things
that we rethink there where
When you're creating an Aurora database like test your production database like you do want five or nines of data
You know you want you you need durability you want availability you want all of those things and for the agent use case
That's it that's arguably maybe not arguably like over-engineered for for what we need and so
You know, it's hard. We don't um, and we kind of have this this uh
Um, like a belief that we don't really want to have like a non-durable option
That's going to cause problems easier because you never really quite know if the database that's created is wants to stay around for a long time or short amount time
And so we're we're um trying to do the hard work to think about how do you accomplish both of those things where you can create it quickly
And throw it away and you don't really waste a lot of resources, but um
But um, but if you do want it and it can be durable and stay for a long time
I actually grow into a big production database
So those are some trade-offs that we think about actively as we think about how does the
The the more traditional kind of this is going to be my production system kind of capabilities
Um
Match with uh, which with some of the more transient nature of the infrastructure that agents
Want to use um, so that that's one. I think but there's there's there's a bunch. I think the
Scale and speed and latency of creation of of durable resources is another one that's interesting
Um, you know, I think the other one that we actively think about too is just
Are there new building blocks that agents are going to want that we just didn't really need both
exactly and so compute sandboxes
um, uh, gateways
Um, agent permissions versus people or or for service role permissions
I'm a lot of those are things that we have built and are building and thinking actively about because they are just brand new building blocks
It's not using the existing building blocks differently, but it's brand new ones where pretty pretty clearly
You want a different permissions for agents. You don't just want to give it
Reggae's permission exactly and let it go do whatever you can do you actually want um, you know
Very time box short-term permissions to just go do a task. I actually may not want to give
Permissions to a whole tool and at all you may want to get very fine-grained commissions
An agent's able to do from a sandbox, right? You actually want a sandbox that um, uh, you want to
What position is having a talk to that's right. That's right exactly. That's and they're they're different
It's similar. It's the same idea, but it's you know, how do you have that lightweight and and think about
Um, unfortunately, we've done a lot of work with um, firecracker or are kind of micro VMs
Okay, a lot of the sandbox companies startups use far
They all use fire which which we invented 10 years ago maybe or something like that and it's it's really
Purpose built for I mean, it's not wasn't purpose built for agents
But it's actually quite good for agents because you can spin them up really rapidly
They have a great security boundary um, and you don't have a lot of virtualization overhead that you have from a traditional large VM
So, you know, I think those are some some new building blocks that we're thinking about and um and more merge every single day
You know, and yeah, so I've rattled enough into agents. I'll come back to it later, but it's a fascinating topic
But uh, let me switch gears a little bit and ask about something that
Every one of our startups face and we get asked most frequently which is
How can we get GPUs, right? Yeah, I mean obviously you guys have running
NASA GPU farms yeah, increasing it every day
But the top frontier labs gobble up every available GPUs more part of that
How do you deal with the internal at the same time you want to promote newer companies that are going to become the enterprises of tomorrow like you said
Um, how are you thinking about balancing? Yeah, it's a question needs and these small companies that don't have a lot of credit and
Not so well. Yeah, it's a great question and and a couple of these things make it more challenging to one is
The the CapEx expense needed to go deploy kind of all of the compute that everyone needs right now is massive, right? And so we are
I think your CapEx is worth 200 or something 220 billion dollars for 26
It's uh, I I one point I saw that it's you know, it's it's a pretty I mean that is a larger expense than we've ever had
Maybe any company has ever had in a in a single year um, and um, and we don't anticipate slowing down anytime soon because the demand is just massive
And so at some point you're limited by how fast can you build data centers how fast can you deploy capital how fast can you get memory and chips and all of those kind of things
And so all of those are at different points constraints that were you know, whether it's
Power data centers capital memory chips. They're all kind of constraints at various times constructs people to build buildings like yeah
Construction people are are our premium today
Um, and so all of those we you know, we we work really hard to make sure happen
Um, and then with the capacity that we're able to deploy which is still a huge amount and and not enough. We know
Um, we really think intentionally about what is that allocation strategy and so
Um, we're great partners with the large frontier labs the anthropics and open a eyes and met in other large customers
And so those folks are are really good customers of ours. Yep
And uh, and we want to make sure that we invest in them. We have large enterprise customers, whether it's
Sales forces or or JPMC's or you know other large companies that that have
Demand for fewer numbers of GPUs, but or or accelerators. Sometimes they want training and sometimes they want
Uh, Nvidia GPUs
But we want to make sure that we can support them as well and we're very intentional about how we make sure that we have capacity for startups
And and so what we do is we actually do allocate and we basically say okay, we're gonna keep
We could you're right. We could sell every single um
Uh, yeah GPU or or AI accelerator we had to probably just little that the big frontier labs and go today
Um, we choose not to do that because we actually want to keep growing that the full ecosystem
They get a large number, but but um, we want to keep supporting
Um, broader set of customers because we actually think that both the whole ecosystem would be more healthy for us
There's some diversification, but it's also just we know these are gonna be big companies over time
Um, and so we try to support them. Um, I saw recently that we say
You know, yes in some way shape or form to something like 60% of the the requests we eventually get
Sometimes it's a little bit later. Sometimes in a different region. Sometimes it's a slightly different configuration than the customers are looking for
Um, but we really try to lean in and and and try to allocate as much as we can and every single startup that you have once more
And so you know, so we're working hard um to try to make sure that we have all of those um
But you know, it's it's hard like that's that's it's you know, like this is the as a lot of people say it's a good problem to have
Yes, it is uh, but it's it's a problem nonetheless and um, and we're
You know, we continue to look at at um, you know, we recently announced we're gonna be buying
Um, you know two million Nvidia GPUs over the next a couple of years like we're planning a massive amount of capacity
And two million two million um, so you know, it's a lot um, and it's over the next couple of years, but um
And uh, and who knows if that's enough, we'll kind of have that that's a at some point again, we're limited by other other components as well
Um, but it is we're we're very intentional about keeping capacity available for for the startups
And I know it's it's painful to not have enough, but we we keep pushing
When did the
your cap X cross, I mean, pure AWS.
I mean, Amazon was a bigger company cross like that.
Let's call it even one billion or 10 billion a year.
- Oh, I don't know.
I'd have to go back and look.
I'm not sure about that.
But it's definitely scaled up over the last couple of years
in a pretty meaningful way.
The AI build out has definitely ramped
our cap X spending.
And so I'd say the last call three or four years
our cap X has definitely accelerated pretty meaningfully.
I don't know when we crossed a billion,
but given we're at 220 now, it's probably the while.
I mean, we've been spending cap X for a while.
And the company has been really good
about funding that, and obviously the AWS business
is a good one that we like to invest in.
- Yeah, yeah, of course.
- And I think we were, for the longest time actually,
we were investing ahead of where the demand was.
And I think one of the most painful things
is that with a real ramp of GPUs,
like a lot of the elasticity has unfortunately,
you know, come gone away.
- Yeah.
- And so hopefully we'll get back to it.
And in our core compute and storage and things
that elasticity is still there.
And that, but, but, you know,
the, but, but so we've been spending for quite a bit of time.
And we feel really good about the spend
that we're making now, which is, you know,
I think I get lots of questions sometimes about
how you feel about that spend and like you're nervous
about bubble and other things like that.
And I will say, you know, we,
because of the position we have like one,
we do take this diversified approach.
And so not all of our capacity is bundled up in one customer.
And I think you go to some of these,
whether they're Neo Clouds or some of the,
the other providers out there.
And, you know, you'll see sometimes concentrations
of 30, 40, 50, 60% with one or two customers.
We're nowhere near that.
Obviously, we're, you know, single digit percentages
at the highest and, and usually it's less than that.
For one, I think we have a lot less risk
on one particular customer.
But also because we have that rich set of services,
like AWS is where people really are coming
to launch their production workloads.
And so the majority of our usage today actually
is either core compute and storage and inference,
which is part of that application.
And so those are the workloads that I think
just aren't going to go away because we see
enterprises getting positive ROI.
You go talk to the customers and you say,
at the capability today and the cost today,
are you seeing positive returns to your business?
And almost to a person, they'll say like, oh yeah.
And so you're like, well, that's not going to go away.
There's no bubble in which they, they stop spending on that.
And so, you know, and you know this, like the VC model,
like is every billion dollar startup going to make it?
No, they won't.
But, you know, that's, that's kind of the game.
And that's been, that's been true for 50 years.
Yeah, they haven't always been billion dollars.
It's changed, but the principles are the same, right?
You bet on 10 and one makes it and pays for the other 10
or whatever the percentage is.
Hopefully it's hard.
It was far, yeah, right?
And so that, that, they're, you know, you saw that
with the internet where there was a bubble
and a bunch of, of internet companies didn't make it.
And the internet still think.
And a lot of the companies that had durable businesses,
the Google, Amazon's, the others like that,
they did pretty well.
And so for us, we think we'll feel really good
about that investment and, and the continued investment going on.
You guys have a view of the demand, that's unparalleled, right?
Because you're seeing across the globe,
you're seeing across every segment enterprise
and the big labs and the A&A of companies and so on and so forth.
So yeah, if anybody should call it,
you should be able to call it first.
I hope so.
Yeah, that's our, that's our plan.
And honestly, we spend a lot of time thinking about it.
We're very intentional about how we spend, you know,
shareholder capital.
And, and we think we're making great investments.
We have a lot of good protections of how we intentionally
invest that money, but, but yeah, we're, we're very bullish on,
and, you know, I think Andy's been public about saying this.
We, the potential for AWS is, is really, really large.
And over the next decade, the potential is there.
And anytime you have an opportunity that's that big, like,
you just want to make, you want to, you want to invest to go after it.
As the scale of these numbers,
go larger and larger, right?
Yeah.
Has your planning and process dramatically changed in terms of,
I mean, you're not writing.
Yeah.
Billion dollar checks, you're writing like 50 billion dollar checks.
So 20 billion dollar checks or whatever it is.
Yeah.
Yeah, yes and no.
I mean, I think, look, a lot of times we're,
the, the, we're still very bullish about the investments
and then lean forward, but they're, they're the process.
So there's a lot of things that have completely changed.
Like funny, even if I think about 20, 28 demand.
Well, and there's things that we have to think about now
that we just never had to think about.
And so, if you go back 15 years, if we needed more power,
we asked the power company to give us another, you know,
couple of megawatts or, well, that's right.
And so megawatts, like, and, and they would just give it to us
'cause 10 megawatts wasn't that much.
Yeah, and that was plenty for us to keep growing.
Now, we have to bring our own, we bring our own power.
And so we, we pay for power projects.
We pay for renewable projects.
We're one of the biggest renewable power purchasers
each year for the last 10 years.
And so it's, you know, we're regularly bringing on
new solar projects, new, nuclear projects, new,
sorry, talking more behind the meter.
Or are you talking about launching with the operator?
We'll do both.
Okay.
And so both of those things.
And so oftentimes with these power projects that we bring on,
we'll pay the capital and pay for the project.
And then it'll go into the grid.
And then we'll get credit for that.
So we'll bring those on.
Sometimes we'll do behind the meter too.
Like, it's a mix at the, at the scale that we're doing,
you have to think about all of those things.
But, you know, that's, that's planning
where you're thinking, you know, 20 years out
of how you're gonna think about power,
how you're gonna think about transmission,
how you're thinking about that capital.
And, and that stuff we'd never had to think about before.
'Cause that's planning we just never had to do.
But, you know, we always had to, it's, you know,
the other things is when we used to think about server demand
that we needed, we, you know, we'd have, you know,
multiple quarter demand things.
And we'd talk with our suppliers and things.
Now we, we work multiple years out.
Just because the size is so much that we have to think about,
kind of what do we need for 26, we need for 27,
what do we need for 28?
And so that's, but, you know,
but it's also one of the value that we bring to customers,
right, that's a thing that legitimately customers
can't do themselves.
Like, they're not gonna do power projects.
They're not gonna plan their memory footprint in 2028.
Like, they're just, they can't do that.
And so that's one of the real values that we bring
to our broad set of customers is just,
that's just a whole set of things
that you don't have to worry about.
And that we spend a huge amount of time thinking about.
- Yeah, yeah.
You guys are generally, I think far that I could,
the largest buyers of practically every component
of a server, correct?
- I don't know that.
(laughing)
- I am sure there were one of the bigger
purchasers of components out there, for sure.
And who knows about everyone?
But, and it kind of depends on how you think about them
and how you measure.
But.
- So the value I was leading to is,
where do you see the constraints being most severe
that's called it 2728?
And where do you see the constraints easing up?
- Yeah, it's funny.
So, I'll answer this around about way,
but, you know, I remember that in undergrad,
we actually, this was long time ago,
and I never thought this would be like a useful book
that I read, but we read the goal.
And then I read the goal, right?
And so it turns out there's never one constraint.
There's always just the latest constraint.
And so you have to think about all of them.
And as soon as you hit one, there's another one, right?
And so, you know, what is the constraint
that's going to happen in 28?
Like, I actually don't think there will be one.
I think it's like every month for us,
it's, do you have enough power?
And then as soon as power is no longer the constraint,
you know, it might be memory.
It might be TSMC capacity.
It might be, you know, HBM.
It might be networking components.
It could be, you know, there's a blip somewhere
in the supply chain and like, you know, connectors, right?
You know, whatever it is, right at some point.
You have to think about all of those pieces.
And it's, you know, it's not also the kind of where
they happen matters to, you know?
So it's, it may be like, Hey, we have a ton of power
in Indonesia, but we don't have enough in Germany, right?
And so you think about kind of where in the world,
you want that capacity too, 'cause it turns out
that not everything is totally fungible.
Some is and some's not.
But maybe disk drives, it might be SSDs.
It might, like we think about every single component
and we have, you know, tens, hundreds of thousands
of components all that we track and think about.
And some we rely on our, you know, our suppliers
to manage some we direct, much of it we directly manage.
And yeah, we have a whole team that does that
and they're fantastic, they're industry leading.
And we kind of, we saw this problem coming
probably a decade ago and really started not just
thinking about, okay, how many servers do we need to track?
But just thinking all the way through the supply chain,
you know, four to years, five to years down.
What is the component that could cause an issue for us
and making sure that we had kind of guaranteed supply
on that?
If you think for your member, I should have remember
when this was, it was over a decade ago
and there was like the floods in Thailand
and no one had disk drives anymore.
- Yep.
- And so that was--
- It was the style of crises,
but then there was memory crisis.
- Exactly.
And so you know, it's like, I think we think through
all of those things and so we also think about
where is there the first occasion?
I mean, the factoring, like all of those kind of pieces
we try to work through and we're never gonna be perfect
on it, but there's always a different supply constraint.
- Yeah.
- So obviously there's a lot of wide-rains in debate
about data centers, right?
And it's clear that folks like us every stand,
but do you think as an industry,
we have not done a good job of explaining
why data centers are good for America generally in the world,
but yeah, and how do you put the internal talk?
- Yeah.
amongst Andy's team, and how do we deal with this?
- Yeah, well, look, I think,
and I think you'll hear more from us over this,
and I agree, I think we need to be more vocal
and be more upfront, 'cause we actually do a ton
that's really beneficial, both for communities we operate in,
for the, we think a ton about how do we bring renewable energy
to these data centers?
How do we think about being water positive?
Actually, our data centers use a really, really small amount
of water we mostly use for air cooling.
How do we think about being great participants
in the communities where we are and how we bring
high-paying jobs in the communities we operate in?
And not all data center operators do that.
I think there are some, there are some all chronicle examples
of others out there that are not great at that,
and they just don't really pay attention to regulations.
They think that the rules don't apply,
or they just launch really quickly without thinking about those.
And I think that it causes problem for the whole industry,
'cause everybody kind of gets lumped into that.
And so, look, I think we'll be, I think you're right,
you know, vocally self-critical, we need to be more vocal
about the benefits that we do bring,
and think about additional ways that we can help communities
understand the benefits that we bring to them.
Both for the services they use, right?
If you usually go to a community and say,
"Well, do you not wanna use Netflix?"
And they'll be like, "No, no, I don't want to want Netflix."
And, you know, it's important for us to think about
and highlight the benefits that we bring
where I recently saw a report where one of the communities
that we operate in, everybody in that county
pays $5,000 a year less in taxes
because of the taxes that we bring to that.
And we don't tell them, they don't even know it, right?
It's just invisible to them.
And so, I think we just need to be more clear
about those benefits that we bring,
'cause I think if you told the communities,
by the way, your tax bill is $5,000 less,
then it would otherwise be if we weren't here,
they might have a little bit of different thought
about the building that's over there.
- Yeah, yeah. - So, but not everyone does that,
not everyone kind of--
- Presumably there's no transparency
that's going to be a lot of fun.
- Yeah, look, I think much of the data center community,
not just us, is actually pretty good actors.
And there's just a few that aren't,
that kind of, I think, have caused some of the angst
recently, and I think we just need to be a better job
of highlighting, you know, who's being good citizens
and who's not.
- Yeah, and before we leave the hardware topic,
I want to touch upon the training.
And your whole history with building your own chips.
Viva Rona, if you're a first partner,
is using Nitro Long-Taragot.
- Mm-hmm.
- Since then, Graviton, and mid tremendous progress.
So, what was the thinking that led to saying,
"Look, we're going to do our own thing?"
And then, how has that progress been and better?
- Yeah, it's actually a fascinating story,
and I think it's a great example of where,
Amazon AWS will innovate, and will iterate over time,
and continue to think bigger about what we can do,
but kind of prove our way there, as opposed to,
you know, and so, as an example,
it's probably now 10, it was probably about 13, 14 years ago,
we were seeing that there was a pretty significant
virtualization tax on the overall number of resources.
And we were kind of thinking about,
how do we, our customers were telling us
I want bare metal performance,
and they're comparing, having all the resources of a server.
And so, the first thing that we did is that,
we took a network offload card,
and virtualized all of our network virtualization,
pulled it off into an offload card,
so that network virtualization got closer
to bare metal performance, and back then,
it was not quite bare metal, but it was closer.
And then we got really excited about that,
and we said, "Okay, what if we could move storage virtualization
off as well?" Right, and none of the network offload cards
could do that, and then we found this one company
who had some R&Pors on an offload card,
and they were doing it for other reasons,
I can't remember their original purpose,
but we're like, "Could you use those
to do storage virtualization in some of these other functions?"
And we're like, "Maybe."
And so, we really iterated with them,
this was the Annapurna team.
- Yeah, this is the Annapurna team.
- And just love that team, like really innovative,
really mission-driven, really wanting to solve problems.
And so, we acquired them, and we said,
"Can you build a slightly bigger card
"that actually could take all the network virtualization off?"
And basically, give us a bare metal server
that has no virtualization on it, no VM virtualization.
Everything is through APIs on the card.
And because we have this view that one,
performance would be much better.
Research, a resource utilization would be better.
We do it again.
- The security isolation.
- Security isolation would be much, much better.
And we tell people, you know,
could then legitimately tell people,
we have no access to any of your VMs that are running there.
And this has been a huge benefit for us for the last decade,
where frankly, like we have been leading,
and others have been kind of slow to do this,
because this is not a generalized thing
that people can do.
But so we got to that, and we basically said,
look, we're making a lot of progress here.
What if we take, you know,
and there's a bunch of ARM cores that were on this offload card,
and we said, "What if we turn that into a server?"
And we did that first with Graviton.
It was a very underpowered, very small server
that we launched, and customers were excited.
They're like, "Oh, I'd love to have an ARM server."
This is super interesting.
And so we went down the path, and Graviton,
and part of what we did is we looked and saw
that there's the slope of, you know,
ARM cores were getting faster.
And where you saw the power utilization and the graph,
and you knew the intercept was going to happen
for where the sark texture was going to be really good
for parts, and they just needed somebody
to drive the ecosystem and get some of the pieces in place.
So we did that with Graviton, and Graviton's been
a runaway hit at this point.
Have you been public about what person
of your fleet is currently?
We land more Graviton ships every year than any other type.
So it's very popular, and it's, you know,
look, we, they're 20% cheaper at a 20% better performance,
and have been like that for the last kind of five, six years.
So that's an easy value proposition
that, and I think the vast majority,
something like 90 plus percent of our top 100 customers
all use Graviton in some way,
shape or form across their fleets.
And so that's been a huge win for us and for customers.
It's been a single big, is single easiest way
that customers lower their bill is to move to Graviton.
They can often, we've have examples
where people have moved their whole fleets
and cut the number of servers they had in half.
Wow, performance is so much better.
It's amazing.
And so half is number of servers each server costs less.
Like it's a, it's a big win.
And so then about five, six years ago,
we said, you know, we saw the rise of, of AI compute happening.
Not nearly expecting what it was today,
but still saw it was going to be a big material mover.
And so we went in, built our first chip in, in training,
and we're now, you know, in market with third generation
in training, three, and seen fantastic results.
So it's, you know, we're sold out for capacity
through probably towards the end of next year or something
like that, and we're trying to, again,
we're trying to figure out who we can save some capacity
and get startups to be able to use some of the capacity
because it's a, we see great results.
So most of, you know, the, the majority of, of traffic
on bedrock all runs on training, and we have great deals
with both anthropic and open AI to build on top of training.
As well as a, a, a set of, of smaller startups.
And, you know, I think we have, half dozen to a dozen startups
that are building on top of training now too.
And so now the name suggests it's a training chip.
Yeah, we're back to it.
Everybody's using it for inference too.
So that is the lean architecturally, and that is it going.
It's a good point.
Look, the vocally self-critical, we're terrible at naming.
And so it's not, it's not our strength.
Originally, we had a chip called inference for inference
and training for, training for training.
And then as the model's got bigger and bigger,
and it turns out you actually want to run the inference
on these really large systems, it turns out the,
training required, turns out that training is actually
maybe the best inference chip on the market right now.
Yeah.
From an absolute performance and cost performance point of view.
And so is it, it's, it's a better memory, memory bandwidth,
where is the, it's just, it has, it's just the architecture
is a little bit different than than others.
And it's much cheaper.
And so from a cost performance perspective,
an absolute performance perspective, training is great.
And so we're, we use it a ton for inference.
And as I said, bedrock, it drives much of the bedrock
inference today.
And, but it's also a good training chip too.
And it's, I think where we, a lot of our broad set
of customers usage is not in training models,
but it's in using it.
And so that's where a lot of people get to use it
under the covers.
And that's where we're excited about it.
But, but a lot of the big customers
are interested in it for training clusters as well.
And particularly as you get to training three and four,
which we announced, we haven't launched training four yet,
but announced it.
Folks have kind of looked at that architecture
and said, yeah, that's the future
of where all my training clusters to be as well.
So we're quite excited about the future
where that goes for these really broad scale training
clusters also.
But yeah, it's both.
- Yeah.
Now, let's get back to talking about ages,
but from a perspective, large enterprises,
or large medium enterprises.
- Yeah, that are they in their adoption?
And what sort of benefits are you seeing them reap already?
And what is the roadmap for them
as far as you can tell from your vantage point?
- Yeah, it's a really good question.
I think it's one that we've spent a lot of time
thinking about.
And when I talk to customers all out there today,
they view, you know, they're getting a lot of value
out of what they've done today.
And I would say the agents that most enterprises
of built are relatively simple and straightforward.
and they're starting to think about,
"And they're mostly non-autonomous," right?
They're still kind of people in the loop, if you will.
And so I think where, and by the way,
their customers are still getting lots of value out of that today.
And so they're really thinking about,
how do I have these be autonomous, but in a safe way?
And I think there's two things
that I think hold customers back today
from just continuing to scale.
And it's already a pretty big business today,
but I think it has a massive opportunity
to really change every single customer out there
and every single workflow and really thinking about it.
And so number one is just how to think about it.
I think what we originally saw was that enterprises
had a workflow and they're saying,
"Great, I would have an agent go do the same workflow."
And what we encourage them to really do is think,
not just replicate, you know,
Bob does step one, two, three, four, five.
So agent is going to do one step one, two, three, four, five.
And then Bob's going to check it at the end.
That's not really the model you want.
You want to actually step back and say,
"If I want to accomplish something,
"how can an agent do it differently?"
It can do it in a massively paralyzed way.
It can try 50 different things and get to that.
And how do you help it get to that right outcome
and rethink how a computer would solve a problem
versus a human solving a problem?
And so one of the things is us just helping customers
understand how to think about that
and really kind of have that blank slate
because that's where you really get value
is not just replicating what you're doing today,
but thinking from a green field approach
about how you go solve a problem completely differently.
And I'm sure that's how many of your startups
are thinking about this too.
How do you help customers green field solve a problem,
not replicate the thing that happens today?
So that is number one.
- People running fleets of agents and swams,
whatever you want to call them.
- Yeah, we're coming these days.
- And you just want to think about it, you know,
and so enterprises are not as, again,
this is one where you learn from the startups
and you try to how do you apply that
to an enterprise world where an insurance company
is not necessarily as forward leaning,
but they would love to figure out
how they can have a better, you know,
approval workflow or something like that.
So that's number one.
But then the second one is this,
how do you turn those into fully autonomous workflows
and how do you actually trust the agents?
And so we're spending a lot of time thinking about
how do we build services to help enterprises
feel like their systems are secure
and that they can trust an agent to make a decision
that can have the right guardrails,
that can have the right permissions on their data
that is not gonna delete production systems,
that it's not gonna make tragic mistakes.
And right now I think that nervousness
is probably holding people back.
Maybe appropriately, by the way,
is holding enterprises back from just saying,
okay, go nuts, like you don't actually want an agent
to just go crazy and I soundly deleted production database.
That's gonna be pretty bad.
And so you, that's how we're kind of actively
working through this with customers on,
how do we both help them architect
and frankly invent new technologies and capabilities
that are gonna help them solve that problem?
And so that is one of the areas
where I think we'll continue to innovate
and we'll get there.
I think we have some really good ideas
and some good technologies brewing
that I think can really help.
- So our enterprises learning how to do eVAL systems
and so on and so forth to keep the agents on the ground.
- They need help straight in there.
- Honestly, like both eVALs,
how do you have a constant loop of testing?
How do you think about goal seeking in a reasonable way?
How do you have your data labeled in such a way
that it actually even makes sense
that the eVAL can actually kind of approximate
what you're gonna be doing in production,
how do you measure in production and back test it
so you're not seeing drift?
All of those things are problems
that enterprises don't know how to solve today.
I don't know if anyone really is great
at solving these today.
It's why you've seen so many FDE teams kind of spin up
and AWS's and our partners are really leaning
into the FDE motion to go and help.
And this is the single biggest area
where customers need help.
And when we think about how you and my view is
we wanna train our customers
to be able to go and do this themselves, right?
This is not the traditional motion
where I wanna have a people driven business
that goes on forever where you just keep paying consultants
over and over and over again.
Our view on how FDE should work and is,
we wanna go into a customer who's ready to accept
the, you know, it's a really kind of accept owning this
and we're done.
And in 45 days do work where we can teach them
how to make an eval, teach them how to get their data
in a labeled way, do it work alongside with them.
And then at the end of 40 to five days
you leave and that customer is good and ready to go
and trained up.
And that's where our customers tell us they want,
they don't wanna be beholden
to an external workforce for the next five years.
Yep, yep.
And, but they need help today.
Now you've made a massive investment in FDE's.
Yeah, so taking that even one step further, right?
Some of your industry peers have said,
look, you can't have all of your data
going into the big frontier model.
What enterprises should really do is to take an open source model
and then post train on your own data
and work clothes and phrases and whatnot.
Where do you stand on that?
Are you seeing customers actually trying to do that?
Or do you guys, how do you think about that?
It's great.
The first point, I wholeheartedly agree on that first point.
Like the customers and enterprise data
is their most valuable asset.
And so from the very beginning,
it's why we built bedrock like we did.
We have a guarantee that your data never leaves your EPC.
And so if you're running inside of bedrock,
your data doesn't go back to the model provider.
They never see your prompts
that stays inside of your own trusted environment.
And so that is why enterprises kind of run,
they prefer to run on top of bedrock
and it's why you see that business growing massively.
Like hundreds and hundreds, I mean, every month,
we see that just business just every week.
We see that business exploding.
And it's why you see opening eye workloads
migrating to bedrock.
It's why you see anthropic really growing really rapidly.
And so whether you're using open models
or close frontier models,
I think bedrock is a great solution
that our customers told us by the way.
Like if you remember three years ago,
I got a lot of seats.
Speaking of bad names, that's a good name though.
Yeah, bedrock is a good name, that's good.
But we got a lot of heat actually
for being slow to the AI world
because we actually built the foundations of this
where we said, look, we're not just going to rush out
a service, we really want to think about
how do we make sure that we protect our customers data
and build a service that we think is going to be durable
for the use cases that we knew about.
And if you remember, we got a lot of heat
and we said, look, we're going to go build the right thing.
And now as people move from proof of concepts to reduction,
vast majority of them are landing in AWS on bedrock.
For much, one of the reasons is because of this.
It's also because of the set of services that we have.
We also offer open models.
We offer proprietary models.
We offer a whole set of capabilities around those agent core.
We build these building blocks.
So it's easier to build agents with any of the models that you want.
Whether they're in bedrock or out of bedrock for that matter,
you can use Gemini or other things for it.
But I think that's a, it's a differentiating piece for us.
And it's a super important thing to think about
because having that data go back into the model provider,
I think is a, is a dangerous thing.
You talk about open weights models though.
I do think that there's a,
it's an area that I'm excited about.
We're really ramping up our support of open weights models
and trying to build a good environment.
And frankly, this is where today I think,
and I think it's true, a lot of customers believe
that they have meaningful proprietary data,
that if they could mix in, you know,
do some post training, do some fine tuning
to an open weights model that they could distill down.
They can actually get a better performing model
at a lower price.
There's a bunch of pieces in here that have to work out,
well, they actually have a good e-vail
to actually prove that that's true.
Most people are doing that in SageMaker today.
I think there's more that we can do to make that easier.
But actually, like, if you go look at where people doing that,
they actually do it in SageMaker on AWS today.
- Oh, really?
- Actually, host the inference via SageMaker.
- SageMaker's getting a new lease of life, hasn't it?
- I mean, it is.
It was always kind of a model building platform, right?
And now, if you think about what enterprises are doing
in this world, that's what they're doing,
is they're basically effectively building their own
custom models.
And so that is SageMaker is a great place of doing that.
And I think there's some things that we need to keep building
on to make that easier and easier to do
and to test across different open weights models,
things like that.
But it's an evolving space.
I think it's a super interesting one.
It's one we want to make sure that we have
all the right things for customers to be able to do
if they have the right data and expertise
to actually go down that path.
- And with all the debate around AI extension risks
and the state of the other,
on the security,
one of the release and so on and so forth
on the hugging phase attack.
How are enterprise, what are CEOs asking you
about all these things?
- Yeah, there's a bunch and they look,
they mostly want to say like,
and it goes back to this like when I launch agents,
how can I trust that they're going to do
what I want them to do?
And so we're increasingly,
we have been for a while.
Like we're basically,
we're heavily investing in building capabilities
that allow people to deploy agents safely
into their environment and think about those controls.
And some of those are,
how do you make sure the agents have the right permissions?
How do they have the right sandboxing?
How do you make sure that you have the right set of guardrails?
How do you really intentionally think
about what you want the agents to do
and not do is they're a human and a loop or not?
And so we spend a lot of time with our customers
thinking about how do you think about safe agent deployment
and get better over time?
And what other capabilities do we need to go build
to help people deploy agents safely into their environment?
So there's a lot there.
The other angle on that,
which a lot of people are worried about,
which is just are some of these really powerful models
going to be attack surfaces and kind of mythosphere.
And so I kind of have a view on their,
Yes, that is a. risk, I think, to customer environments. But it's also a real opportunity. And so we
recently launched a source called Continuum that uses these powerful models to help customers
go secure their environment. And so we'll look across their environment, look for vulnerabilities
with them. We'll use some of these powerful models and help customers find vulnerabilities
they haven't found before. And most importantly, by the way, prioritize which ones, because
we know context about their environment, how it's set up, where their permissions are,
where they may have compensating controls that make it harder or easier to do. And so Continuum
is incredibly popular with customers. Actually, we're really bullish about what's possible
from AI to help with AI-powered security. Because look, at some point, customers are going
to need security and machine speed, not at human speed, not at alarm. Something goes
in, look at it. And so that, you know, we're running fast to go help build that for customers
to help them protect their environments. And I'm very excited about the continuum team
is building on that front too. So with the AWS itself, like how, what's the state of
a usage of agents in the broad? Yeah. Well, Continuum is basically us trying to expose
what we do internally. And so we, we use AI extensively for our own security. We use
AI extensively for our own software development. We use agents, actually, one of the things
that's really cool is we use agents across our entire business. And so we rolled out Amazon
quick to every single Amazon employee. And now I see HR teams building agents to help drive
what what used to take teams of people weeks to do that a single person can now do in a
couple of hours to think about kind of team planning and resource management. I have
finance teams that are building agents to go think about how do they go pull tax rules
from everywhere and ensure that we have compliance on a bunch of different pieces. And super
cool to see that things that used to be blocked by software developers. Yeah. Actually, the
planet is as folks are able to go and unblock themselves and innovate more quickly. And so
quick has been an enormous blow. And that has grown like wildfire. We see customers like
small startups using it all the way to the largest enterprises in the world rolling it out
to their entire customer base to get the benefit of kind of being able to access all of your
enterprise data and easily apply agents and capabilities to help you accelerate your
jobs. And so I guess we use it across everything from software development to security to driving
HR analyses. I mean, you guys are notorious for measuring everything about your operation.
Whatever you've seen the biggest gains. Yeah. I mean, obviously, you know, that's the
real answer to this. The derivatives of that. But yeah, the software, I mean, the speed
of software development has been and really product development as a whole, not just coding,
but just absolutely the case, the case at which we're deploying new products is massively
different than it has been in the past. I think you can see this where AWS has always been
known for rolling out features really quickly. And it's, you know, we've seen a turbo boost
on that in the last year or so. As we've got, we call them frontier teams as they think about
agentic development as opposed to conditional development. And it's, you know, it's not code
completion. It really is agent first, the agents right, all of the code, you're just managing
a team of agents and driving that. And it's been fun to see the pace at which innovations
for customers has been massive. And it has to be because that's the, that's the, that's a
problem. Our customers out there have a almost insatiable appetite for new capabilities.
And that's what we've got to do. Yeah. So organizationally, are you like, do you have any insights
on how organizations should change and how I don't say that agents manage people, people manage
agents? There's like, there's going to be, there's going to be lots of people for a long period
of time. I do think organizations will change. I don't know the magic answer yet. So, but we're
actively thinking about it. Insightly. Insightly. Insightly. We're just trying videos experiments.
Yeah. We're trying experiments. We're thinking about pods. As you think about, you know,
here's one example is in a product organization, you used to have a team that would own a particular
kind of capability for a long time. And you might have 10 people working on that thing.
Well, today you can innovate so rapidly that one doesn't have to have to be 10 people. It can be
three to four people. And they build something so fast. You actually want to move them to different
projects and problems. And so thinking about how do you both operate and maintain the things that
you built while being agile and flexible to move around on an organization that's as big as AWS
is active things that we're experimenting with and playing with. But it's fun. And it's enabling for
our employees. They actually love it because they can build faster and do more, but there's your work
there. Yeah. It's a fascinating time. And so thank you very much for your time. We could be talking
about this for ours together, but thanks for all your insights. Yeah. Thank you for having me and
and thank you. We love having all your your companies as customers and we love learning from them
and appreciate having me here. Yeah, we'll keep sending them your way. Excellent. Thank you. Thanks.
Thanks for listening to this episode of the A16Z podcast. If you like this episode,
be sure to like, comment, subscribe, leave us a rating word of you and share it with your friends
and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X,
a A16Z, and subscribe to our substack at a16z.substack.com. Thanks again for listening,
and I'll see you in the next episode. As a reminder, the content here is for informational
purposes only. Should not be taken as legal business, tax, or investment advice, or be used to
evaluate any investment or security and is not directed at any investors or potential investors
in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the
companies discussed in this podcast. For more details, including a link to our investments,
please see a16z.com/disclosures.
Podcast Summary
Key Points:
AWS has consistently prioritized startups as the core of innovation, with 30-40% of its revenue coming from companies that began as startups.
Startups today are larger, more funded, and demand more GPUs, but still rely on AWS for secure, scalable architecture and cloud best practices.
Agents are increasingly writing code and managing infrastructure, prompting AWS to build new services like agent core and AWS Context for seamless agent operation.
AWS is actively developing new building blocks such as compute sandboxes, fine-grained agent permissions, and time-boxed access to reduce risk and improve security.
GPU capacity is in high demand, especially by frontier labs, but AWS maintains a balanced allocation strategy to support both large customers and emerging startups.
AWS is investing heavily in custom silicon (Graviton and training chips) to improve performance, reduce costs, and offer better control over infrastructure.
Enterprises are adopting agents cautiously, focusing on safety, autonomy, and trust through guardrails, evaluation loops, and secure data handling.
AWS is expanding support for on-premises model training and open-weight models, enabling enterprises to build custom, data-secure AI models while maintaining control.
Summary:
AWS has evolved from a startup-focused cloud service to a central player in the AI infrastructure landscape, driven by massive demand for compute, especially GPUs. The company has strategically maintained capacity for early-stage startups, recognizing them as the future of enterprise innovation. As AI agents increasingly write code and manage infrastructure, AWS is building new services and security models—like compute sandboxes, fine-grained permissions, and agent core—to enable safe, efficient agent workflows.
Despite intense competition and rising GPU demand from frontier labs, AWS ensures balanced allocation by prioritizing ecosystem growth over short-term gains. The company is also investing heavily in custom silicon, including its Graviton processors and training chips, which deliver superior performance and cost efficiency. Enterprise adoption of agents is cautious, with leaders concerned about autonomy and data safety.
AWS is addressing these concerns through robust guardrails, evaluation systems, and secure environments—such as AWS Bedrock—where data remains on-premises. Additionally, AWS is expanding support for open-weight models and fine-tuning, empowering enterprises to build custom AI models. These developments reflect a broader shift in how cloud infrastructure is being reimagined to support autonomous, secure, and scalable AI workloads.
FAQs
AWS intentionally keeps capacity available for startups, even as frontier labs consume significant GPU resources. It allocates a substantial portion of GPUs to startups to support future enterprise growth, ensuring ecosystem diversity and long-term health.
Agents perform better on AWS due to optimized services and low latency, especially in data access and compute performance. AWS has built context layers and fine-tuned underlying infrastructure for agents to operate efficiently across services like S3 and Aurora.
AWS is developing services like AWS Context and Agent Core to help agents easily access and manage data across multiple storage systems. They are also improving sandboxing, fine-grained permissions, and API interfaces for agents to work more efficiently.
AWS is rethinking infrastructure design to support both temporary and durable resources. Agents often need short-lived databases, while enterprise workloads require five-nines durability. AWS is working to balance fast creation and resource efficiency with long-term reliability.
Startups today begin at much larger scale—valued at billions with $200M+ funding—requiring more compute and capital. Despite this, they still prioritize architecture, security, and scalability, similar to enterprise needs.
AWS has simplified onboarding by allowing new users to create accounts with a Gmail email, skipping credit card and complex setup steps. This makes it easier for startups to deploy and begin building on AWS quickly.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.