EP 190: NVIDIA and Marvell Earnings, Hot Chips Hot Takes
53m 19s
Nvidia’s recent earnings showed strong confidence with a 70% revenue growth guidance for the next year, driven by data center demand and supply constraints, particularly at TSMC. While the company raised prices, gross margins are declining, indicating supply chain pressures and reduced pricing power, especially as hyperscalers increasingly source memory directly. The semiconductor landscape is experiencing a period of technological ferment, where diverse architectures for AI inference—ranging from custom ASICs to memory designs—show strong innovation but lack clear standardization. Startups and major players like AMD and Samsung are presenting competing solutions, with many emphasizing custom designs for specific workloads. However, such customization increases complexity and cost, reducing memory fungibility and raising concerns about performance variability across vendors. This fragmentation suggests a future of specialized, vertically integrated solutions rather than universal standards. The dynamic is further illustrated by Waymo’s unique approach to autonomous driving, requiring high-precision inference in a different workload than typical AI models. Overall, while innovation remains vibrant, the path to scalable, efficient, and interoperable AI hardware remains uncertain, with Nvidia positioning itself as a key enabler of the ecosystem despite challenges in margins and pricing.
[MUSIC]
>> Hello, all.
Welcome to another episode of The Circuit.
I am Ben Baharin.
[MUSIC]
>> Greetings, programs.
I'm Jay Goldberg.
[MUSIC]
>> Well, we have a lot of topics to get through this week.
It was a hot week because we had hot chips, which we'll get to.
And we had Nvidia earnings and Marvell earnings and.
[MUSIC]
>> And not our others, but does it.
>> I'm in an on air conditioned hotel room right now.
>> Jay's going to start sweating profusely.
And because we're talking deep silicon today for
anybody who watches this, I have my TPU shirt on.
Here's my glorious TPU shirt since they did talk about TPUs at hot chips.
All right, let's kick it off with Nvidia earnings.
Yeah, the stock did okay, kind of a little bit down on this.
What Friday, August 28th, that we're recording.
But it was up more than it was earlier in the week, but nothing monumental.
I would say the reaction was positive folks remain optimistic that Nvidia continues to grow.
But what was your top line take?
>> So, the quarter and the guide were pretty much as expected.
Probably not good enough.
They were, how do I say this?
They were beat expectations, but everyone sort of expects them to beat expectations.
So that was fine.
What I think really interested investors though, is they gave revenue guidance for
70% revenue growth next year.
And I think that's what people really keyed in on.
This has been a big, big topic of conversation over the last few months.
People have been trying to figure out revenue growth is going to decline.
What's it going to decelerate to next year?
And no, it's actually, it's going to grow 70% they say.
And Jensen kept pointing out they could grow more if they could get the capacity.
So, pretty bullish signal.
>> Yeah, so let's talk on this for a second.
I guess when you do something like that, you kind of set your expectations right in stone.
So, people will just do what they do numbers wise, assume what that growth rate means.
And so now, the upside question becomes like, where can they eke out any additional right surprises?
And I talk to a folks who are like, oh, they're probably being super conservative.
It might be more than that, they're just, they realize it's a supply-contained environment.
So, they're leaving some room for the upside.
But you're right, their key in saying it would be a lot more.
Jensen was asked this on the call and he was like, how much more would it be?
I think it was our curry.
And then he's like, it would have been a lot more.
Okay, so great, but you are supply constrained and he was sort of like,
we got to go figure out how to get more supply, which supply doesn't grow on trees.
So, I don't know where that comes from.
But yeah, I don't know, I guess my point is like, where do they go from here?
Right? They've just guided to that.
So, like, how do you paint a picture going forward other than you're executing on your 70% growth call?
Yeah. So, I mean, the typical investor relations playbook is you under, under promise and over deliver.
And so, people say, oh, he's saying 70%, it's really more like some number that's bigger than 70%.
And I think at this point, they have a fair degree of certainty around their business
because they know pretty well what the orders are, especially because we're talking about data center construction next year.
That's all that's getting lined up now.
And I think to some degree, it's going to come down to TSMC.
And what capacity they can get at TSMC, that's going to be the big swing factor.
And we probably won't know that until early next year, right?
So, yeah.
Yeah, I mean, I think all the things they can do, like, there's other things they can do here, they're, you know,
they're going to start selling LPUs from GROC, we'll see how the market feels about those.
But like, those aren't as capacity constrained.
And so, maybe they get, you know, a little, a little kicker from that.
There's a few other levers that they could pull.
Yeah, they mentioned CPU stuff, which, you know, obviously, they're holding their number.
But, you know, they did announce a deal with Amazon, that includes Vera CPUs and a shipment number of,
I think it was two million more GPUs, but that's a combination of Vera standalone CPUs and attached to head node.
So, do your mix, right, as you will, but they're associating value to that in the CPU side.
And I imagine others, right, but, you know, part of it too is, I guess what I'm sort of curious about is,
like, so all the supply chain work we do kind of seems like 2028, which is not, which is fiscal 2028,
which is not in video fiscal 28, but, but, uh, or year 28, calendar 28, yeah, calendar 28 is when pockets of capacity can start to open
and the other part of this that I think is interesting is, you know, you just look at the lag from order to data center
and kind of those giant orders that people started planning a year ago, like really started ramping this capex wise,
start to show up in energized gigawatts in 2028.
So, it's kind of like supply chain capacity, again, will still be constrained, but it will be better.
That TSMC, which substrates with all the components that they need with wafers, but I, but I kind of feel like the following year could actually be bigger because of an unlock of capacity that I think they're going to remain constrained in for calendar 2027, which is hard to, like, again, that's just nutty, right, to predict my point, though, is they will still be capacity constraint.
It's constrained in calendar 28, but I think less so than calendar 27. That's, that's what I wanted to say.
Yeah, they said they're sold out for this year and they're likely sold out for next year as well.
And like 28 is every, you know, that's, that's everybody's, everybody's guessing right now.
A lot of capacity coming online, you know, it's so far away, who knows?
Yeah, but, but this is, but this is where I think this gets super interesting, which we were texting about, this is where I think Intel comes into the picture.
And, you know, Intel's going to start to have foundry capacity around that timeline. In video is, I mean, they're not, they're as as supply constrained as anybody.
However, I think they have the most ambition to utilize every bit of existential capacity. They're like the most hunger because they know that's going to impact their bottom line immediately, again, true of everybody, but supply chain scale favors and video at this point.
I really think the picture is being clear that Nvidia is most incentivized to help Intel execute here and that they are customer zero, cleaning the pipe to for Intel the same way that Apple did for TSMC, because again, it, it behooves them so much to have Intel ready to go to scale for them to me, because they're, I mean, if it is that high of demand that I think they're the ones that can, can move this needle.
For, in fact, I was joking that like Intel should just full port, full port all external products right now to TSMC and just give all that capacity to Nvidia, because they probably make more money and foundry immediately.
But I think they're too far down the path of where first party products go, but, but you get my point, like they are going to be in Nvidia invested in them.
They're going to be the ones, I think, Nvidia champions to help, help additionally meet their scale beyond what TSMC has.
Yeah, I agree. Nvidia and Intel makes a lot of sense.
But there were a couple, a couple of nits on the Nvidia call as bullish as they were.
I think there's a few, there's a few, a few questions people had.
Like the, the big one that I know a lot of people are talking about is gross margins. So they guided down gross margins, not this current quarter, but the cube, what the January quarter, they're there.
I wish they would ship their queue for their year. They're cute for their to for gross margins go down from 75 to 72% and they stay down at that level to low 70s for the next 34 quarters.
And I think that's, that's weird. Like I had to, because we know in the press, there are reports that they've increased prices, like they sent those letters out last week.
So they've put in place all these pretty significant price increases and still their gross margins are going down, which is.
They said it's, it's partly mixed and it's partly memory as well as mostly memories what they're saying and.
I just don't know to make of that like memory, so it sounds like they're not able to pass on as much memory markup as they used to.
So even they're feeling the pressure there, so the memory guys have the leverage.
not just memory, I think, again, every part of this supply chain, PCBs, analog, everything,
everybody is increasing prices, and I think, again, you're exactly right. And this has
been sort of our question for Apple, too, right? And Apple doesn't have 70% margins on hardware
like Nvidia does, but there's sort of always this kind of challenge of what really is your
pricing discrepancy, and I think they're sensitive to that, to be honest with you. I think
they're sensitive to it, too, because perhaps others, competitors, namely AMD, you know,
might not have to increase prices, right, as much given what they've already secured for
a two-year run ramp. So I agree with you. I think that they're willing to absorb some
of that. I mean, they used to be very clear, right? We passed these costs on. It sounds
like it's a little bit more ambiguous. But I say this to make the point, and we argued
this a long time ago, when I was pointing out, right? The upside for Nvidia and a capacity
constrained environment is not all of a sudden greenfield shipments, it's increasing prices,
but not at the cost of right, gross margin, because it's hardware costs. It's increasing
ASP meaning that you could sell this for $10 million a rack instead of five, but get that
margin, not because your components are that expensive. So I don't know. It's an interesting
thing. I mean, I don't know. I didn't hear a ton of concern over that. I mean, they're
still going to continue to ramp, but I agree with you. I think they just are going to
have to eat some, and that is what it is when everybody in your supply chain is raising
prices on you. I think it's significant here, because Nvidia is able to mark up those
memory costs, those memory prices to their corporate margin, and no one else can do that.
And I have to wonder if the hyperscalers, they're now going directly to the memory companies.
I think that's what's going on. It's rather than Nvidia and they're losing pricing,
negotiation, leverage with its customers, they're just seeing more of the memory business
sort of go around them. And that's really a high margin business for them. So losing that
is, I mean, it's not what I'm saying. It's not a huge deal, but I do think like memories
become so sensitive, people are going to push back on them, paying extra for that wherever
they can. Okay, so let's table the memory bit because something came out of hot chips
that's relative to this that I want to talk about when we get to there. But, you know,
so the other thing I was actually intrigued about, and I was thinking about this relative
to the Amazon deal, and I had forgotten about this, but I shouldn't have forgotten about
it. But remember, remember at GTC, I don't know, two or three years ago when you and I were
in the Jensen QA and somebody asked him something about like, you know, you guys sell racks
or you sell whole compute systems and whatnot. And he was like, we sell, I mean, I'm going
to botch his words, but basically he was like, we sell components to a data center. You
can buy all of them or you can buy just some of them. And we'll continue to sell them to
you. And I think this is important because, you know, we get kind of lost in like what
a neocloud would buy, right, which is a full NVL 72 system. So that's everything, right,
soup to nuts, all in video. Then you have what Microsoft and Google and Amazon buy, which
is parts of that, maybe not the networking, maybe just the CPU and just the GPU. And I
just find it, it's interesting the way that that, and I say that to say that in my brain
is a little bit harder to model because in a rack based environment, let's say they were
selling, you know, Grace Blackville sword ever was between $5 and $6 million whole system.
It's hard to say, well, how much was the CPU content value of that, right? Or what was
the GPU content of it? You could estimate it's whatever, you know, 40,000 GPU, but that's
a bundled price, right? But then you go and you sell a standalone CPU or standalone GPUs
or standalone networking, right? Your markup could be a little bit different there because
you have, it's not a bundle, it's an individual component. And so I'm like, I just say that
to say I'm intrigued by areas where they may not be getting the memory, the hyperscalers
are getting the memory and the PCBs. They're just providing them the components to go build
their racks. And I just wonder if there's ways they can leverage pricing there that they
haven't before versus the whole rack sale, which is where everybody kind of gets, you
know, the whole NBL solution is what Nvidia typically gets benchmarked by, my point is
that this individual component size, there might be variations in pricing too. They can
do there that we're not entirely seeing. So they sort of touched on this because remember
last quarter, they reclassified their revenue segments. And now they say data center and
they break out within data center, they have ACIe, which is enterprise and NEO cloud and
sovereign. And then they also have hyperscalers. And I don't think they actually flat out set
it, but they kind of implied, collette implied that the hyperscaler is a lower margin. And
it clearly, like they want people to model it this way is the long segment lines. And
I think it's safe to say that, you know, when you sell to a major customer like Amazon
or Google is buying in huge volume, you're going to, they're going to pay a little bit
less than a NEO cloud, who is buying in smaller volume is buying the whole bundle, which
is, I don't mean that in the antitrust sense, but it's the whole package with the networking
and the memory. They're smaller customers and they're much more dependent on Nvidia
for, you know, existing. So you're going to pay a higher margin to, they're going to
pay a higher margin. And I've been tweaking my Nvidia model a lot lately, and I'm trying
to hone in on that component of it. I haven't quite nailed it down, but there's clearly
a gross margin component difference here.
Yeah. And I say that again, just to say, this is how difficult the system model, right?
Because I see, you know, friendlies on the buy side or sell side be like, this is how
many are actually going to sell? And I was like, well, yes, that's racks, but there's
also individual components that it's a growing part of their pie. Well, not so the growing
part, but a key part of their pie and what they sell to hyperscalers. And so anyway, I
just thought that was interesting to remember. They'll sell you components as well, not
just the entire rack and you're right. That should impact how you try to model this.
Yeah. I think the other interesting thing on the call was they talked a lot about their
financing programs they're putting in place. Right. They directly addressed the question,
is this sort of the financing? They say it's not. We see it a different way. Was the
words exactly? Yeah. Yeah. And it's complicated. There is news out today. I think the journal
is reporting that it was the Wall Street Journal. Yeah. Yeah. It was a journal was reporting
that there is 500 billion dollar SPV platform they're putting in place is getting reviewed
for anti-trust reasons. And it doesn't sound, at least from that article, it doesn't sound
like the Department of Justice is actually investigating them. It's just that as they're
putting that syndicate together, they're realizing that they have to be sensitive to anti-trust
concerns because of that word bundle. Right. If you know, we're going to provide you
the financing, but you have to buy our product, that gets pretty close to anti-trust law.
That word bundle really is a third rail. And so I don't know. That's all complicated,
but like I added up, they have $579 billion in commitments. About half of that, a little
bit more than half of that goes to the splodging. So buying lasers and substrates in TSMC capacity.
So they're going to add there's another $100 ish billion for a couple of other things.
And then there's just 500 billion dollar platform they're talking about. And then presumably
there's a lot more out there. And they also issued a bunch of debt. So in one quarter,
they've laid it out pretty clearly. It's a big number in terms of outside liabilities
and obligations. Some of which may they may not have to pay, but it's a big number.
So I don't know. That part makes me, I don't talk about this. It makes me nervous.
But I do think, you know, it's interesting because like we talked about last week, right?
That the balance sheet has entered the equation for competitive advantage. And they spated
that specifically. And I just, I think it's a vector. And one of the points about,
right, that I think they're trying to make around why they had to do some of these deals
they did is really the credit rating, right? These companies need to go raise money.
And we pointed this out, I think last year, I think you said, like this is why they're
going to help finance these people. They need them to be able to go get money.
So they basically have to backstop them to help their credit rating. Great. We get it.
That makes sense. The nugget, right, you talk about that is interesting and very hard
to model. Like you could think this could be billions of dollars of revenue sharing
with video clouds, but maybe, you know, like that's not guaranteed. So it's hard to work
that in even if there's upside. So I mean, I guess
Even if Nvidia wants you to like noodle on that, I would say don't, just assume it doesn't
happen.
But I think the meta point is just the way in which they're going to need to use their
balance sheet in order to help the ecosystem grow and again, be able to raise money, build
power, land power, show, et cetera, et cetera, et cetera.
But I guess there's only a handful of companies that can do this, let balance sheet be part
of the competitive advantage.
And I think it was just interesting that they made that point clear.
It is a competitive advantage and we're going to use it.
Yeah, I mean, to their credit, they are disclosing this pretty, pretty well.
But it's just big numbers and it's going to get bigger.
So I always think of this in the context of there in a strategic competition with the
hyperscalers and Nvidia is going to will the neo clouds and do existence and now it has
to use its balance sheet to keep them going.
And to the people who said that circular financing, I think it's unclear.
There's a valid argument that they put out which is this is just upfront, this is an expensive
business, you have to invest heavily upfront.
But the returns are such that like we're just helping them get started and these are viable
businesses long-term and can pay all this back.
And at the moment that works, but at the same time, they're also talking about prices
going up, basically doubling in the next couple of years makes it pretty challenging for
the neo clouds.
Yeah, which means still a lot more.
Yeah, which is true.
I also think that they're like everybody where I continue to double down on this, what's
their content opportunity per gigawatt, which again, I tell everybody, you don't just go,
there's going to be 30 gigawatts activated in 2027 and Nvidia is going to get 40 billion
per gigawatt.
That's not how it works.
It's not getting all of that.
But in an environment where they're the larger, host or beneficiary of that, they're
keep coming back to the content opportunity in 40 billion per gigawatt and growing.
So yeah, I mean, it's interesting.
I think the supply chain stuff in my head is just kind of the most interesting about
where they can move the needle there, not just wholesale, continue to increase prices
and get gross margin as a part of that, but where might they be able to eag out some
supply chain stuff?
So it's going to be an interesting, you know, next, I would say calendar 27, honestly,
because they're a guide for, I think it was 109 billion ish for next quarter, certainly
they'll do that, how much they beat, but I think how this shapes out and can you tease
out bigger growth, outsized growth above, I think where most people landed like, I want
to say it was six, between 630 and 660 billion, which I think most people were adjusting
their numbers to to next year, so like, it's got to be like, oh, we're going to crush
that.
It's going to be like how they figure out to get them to 700 or above is an interesting
exercise, but it's going to have to come from supply chain somewhere.
Yep.
Okay.
Any other nits that we're missing, I want to, I want to cover any, any other nitpicking?
No, you know, it's a $1.00 debt, $1.00 debt and obligations is a pretty, it's a pretty
big nit.
Yeah, it's rough out there on the supply chain folks.
Like we said, LTA Waterfalls, you got to secure everybody to make sure you've got it.
This came up at Hotchips, too, man, like everybody making silicon, we had so much supply
chain conversation.
So we'll talk about that.
All right, let's go to, let's go to Marvel.
I'm intrigued.
Marvel to me like, I'm going to, I'm going to start by saying I assume the sentiment is
because they have an investor day coming up and they want to save a lot of stuff for
that in early October.
And so they didn't show, like, this is how we framed the market opportunity in dollars,
Tam, Sam's, whatever the things, you know, the big things, obviously, this $120 billion
deal with Google, which is potential revenue, not guaranteed and a lot of investors trying
to figure out, you know, give us some idea of the conversion schedule of this and, you
know, what we can think about.
And they just didn't want to do that because of, right, this investor day.
So I kind of feel like it was like great, good posture, sounds all wonderful, but not
enough clarity.
So you're just going to have to wait until investor day and reactions were what they were, like,
because it's still positive, like CXL, all the things, oh, the other thing I was going
to say, because we talked about this a lot.
So I'm glad that they have finally officially chosen their lane of XP you attach.
So they were very, very clear, you know, like, we're not going to not do monolithic compute
tiles if they don't come, but like, that's not going to be, whereas like two years ago,
everybody's assumption was, you got to do the whole thing or you're dead, right?
Or you got to do more in compute or Broadcom eats you.
And then they kind of kept saying XP you attach, XP you attach, XP you attach.
And now it's like firmly our lane is XP you attach.
And I do think that's positive.
It makes you a much more companion to the entire ASIC ecosystem, not having to win the
whole program, but you can do well with your lane.
Everybody's program.
I think that's a good positioning, tougher to model, but I'm just, I'm happy that they've
picked a lane.
But yeah, other than that, it's like, yeah, anyway, what was your take?
I think the, you know, the stock is down 10% today, almost.
And I think the big, the very simple summary of what people didn't like was they have the
unfortunate position of reporting that they after in video, when video is talking about
70%, at least 70% growth next year.
And their, their number was 60%.
Which, you know, in absolute terms is, is fantastic growth for, for Marvel, for any company,
see 60% growth.
But it's still just mathematically that losing share to Nvidia, it's just very, very hard.
And, and to your point, I'm sure like they're going to have their investor day in October
and they're going to, they're going to lay out of a much more positive picture.
But I think for the moment, it's like, just Nvidia's growing 70, they're growing 60.
And then you have to deal with all the complexities of like, what do these guys do and where they
selling and what is this Google deal going to look like and it's just, it's, uh, yeah.
Yeah.
No, I get that.
Like I said, it's, it's hard because you, you have the clear execution timeline position
of, of Nvidia.
And then you have like, it, it actually sounds good on paper, but I need to know exactly
how to like convert and model these timelines for you guys.
So it's like, conviction, like looking more optimistic, but not like a guarantee at this
point versus where like Nvidia tells you, like you said, they're going to grow 70, they're
going to at least grow 70 and, and probably more.
We don't know, right?
The same confidence level for Marvel, but, but I do, like I said, I like that they have
their lane.
I like that CXLs come into the picture even though again, right?
CXL is not going to add a gigantic amount of revenue to them, you know, maybe couples
with billions of dollars, not tiny, but, but decent.
Um, but yeah, the, the Google ones kind of the big one, right?
If that can convert more on of that 120 billion dollars in a, in a quicker timeline, that's
a very different shape of their business, but we don't know and, and they didn't provide
color on that.
Yeah, they were certainly very excited.
I mean, they, they, their tone was very positive.
That was very bullish.
I would say in his comments about that opportunity they face at Google, my, my advice if they're
listening is we know what day Nvidia's holding earnings next quarter.
Don't do it.
Yeah.
And that's it yet, put your date two days before the, you never want to follow, you never
want to follow Nvidia, it's tough, it's tough act to follow.
It's, that's true.
Um, okay, anything, anything else on Marvel, we're missing, I know those kind of, I mean,
like I said, it's, we will, I am sure talk about this after their investor day and kind
of land, try to land, parse out some, the market opportunity in economics, but, um, okay.
Uh, all right, let's, let's talk hot chips.
Um, you know, it was interesting, okay, I'm not going to say, yeah, hot, hot, yeah, hot
chips is, is an academic conference, like professors set up by the electrical engineering
department at Stanford, they have professors, they, you know, it's the kind of thing where
you have, we have academics, PhDs, presenting academic papers, there are posters outside.
But it's an academic conference held in the heart of Silicon Valley.
And so the, the people who attend and speak at it tend to be people from the major companies
doing the most advanced work.
And that is super interesting, but it comes at the, the trade-off of, uh, I mean, we're
of an academic context, you're probably going to present
your best work.
When you have all the big corporates presenting,
even though you have the technical leads,
they're probably holding stuff back a little bit.
But it's a super interesting conference.
The biggest complaint I've heard about it lately
is that it's the amount of investors
who show up as everybody frustrated.
But it's a very interesting conference,
but it's very, very technical.
- Yeah, yeah.
- So, my favorite part about this is always like,
the way the hot chips guys have always sort of framed this,
like whenever I talk about it with them is,
it's engineers presenting to engineers,
showing how they did things, how they solved problems,
what the problems they face,
things that they don't know, things that are unsettled,
that they're trying to solve.
And to be honest with you,
one of the things I really like is
a lot of the questions come from competitors.
You would have like a Google event
and you had somebody from our Val or D-Makers
be like, "Hey, how'd you solve this problem?"
I'm curious how you think about this.
Like, it's amazing.
Like, you're like, "This is great."
- Or the best part is the Huawei guy shows up every year
and asks incredibly sensitive questions
to every single presenter.
And sometimes they answer him.
- I actually didn't hear the Huawei guys much.
This year was owned by a D-matrix employee
who like literally just got up almost every time
and asked questions and people were like,
"Is this marketing stuff for D-matrix?"
- The other fun thing was etched was there
as a sponsor, which was cool 'cause like,
I got to ask them some more detailed questions,
but they had this rack, this black box cabinet,
and it was locked up.
And everybody's like, "What the crap's in this thing?"
Like, I'm assuming it was empty, right?
'Cause obviously they want to sell racks at some point.
They sell individual boxes at this point,
not rack fabrics.
But I'm like, sitting there, I'm like,
"Why is there a, is Houdini in there?"
It's Houdini gonna bust out of there?
Like, give me some things.
So there's a couple of startups.
Samsung was there, which I do wanna talk about
when we talk about memory and custom-based eyes,
but let me give kind of like a high level view.
I do think it was one of the best hot chips in a long time
despite what others thought.
There was an inordinate amount of investors there.
I had two group dinners with both of them.
All of them admittedly, I'm not technical.
I'm kinda just here to check it out,
like, explain to me what this is.
But because there is some presenters
who kind of took a non-technical route,
and I think because they were just trying to tell a story,
knowing maybe it was a more diverse audience,
and then there were others who went super deep
in the weeds, really, really cool stuff across the stack.
Each year it varies.
Last year, there was also memory interconnect ASICs
and CPUs, this year was all three of those.
But like in years past, it was transistor design,
it was next generation foundry.
So it kinda changes, but this year, like last year,
lots of interconnect networking, CPUs, ASICs, memories.
Was essentially the core.
But my take, and I wrote about this,
and this is sort of the high level framing
that I gave is there was a lot of agreement on the problem.
Infrainset scale and serving inference scale
is a computationally complex problem.
How the ASIC works, how the memory works,
how the software moves between those two workloads,
everybody agrees.
But what was really interesting was you listen
to the dematrix guys, or the sambinova guys,
or even Nvidia, and AMD, and everybody who presented,
even the memory guys, everybody kinda had
like a different point of view on how they think
is best to go solve this problem.
And so there was disagreement in how we go
and do these things.
And so I wrote about that this week,
because there's an academic term called the era of ferment,
which I will sort of just briefly explain right
for everybody, but it comes from a 1990 paper
called Technological Discontinues,
and dominant designs, a cyclical model
of technological change by Anderson and Tushman.
And essentially, it is an era of ferment
is when a period after a major technical break,
when an entry tests competing approaches
before one common architecture emerges.
And this is why you're seeing a lot of diversity.
Now interestingly, people who were either technical
in the weeds or other engineers,
or have a bias to a particular one of these problems
was sort of like, hey, you're only doing this in SRAM,
or you know, on an HBM, 'cause you can't get HBM.
So you're hacking around this,
not because it's the best way to do it,
but because you don't have a choice, right?
And that may or may not be true,
but the reality was like you have all of these
kind of competing approaches on inference, on memory,
on ASICs, on the design,
even to some of the greedy interconnects.
And I just thought that was interesting.
This was the first time in a long time
I've seen this dynamic of, we all have strong opinions
and they differ about how we should go solve inference at scale.
And anyway, so that was like high level kind of my take
and like I said, I haven't seen this in a while,
which just to me was kind of interesting.
And the way I frame this too is,
they're all, even the startups, right?
Even the memory guys, you could be clear that, you know,
they're showing next year, they're showing two years away,
they're showing things that they haven't even solved yet.
Like, conceptually, we're gonna go do this,
but we realize heat's a problem.
We're gonna go do this, we're real as yields a problem
or we can't manufacture this yet.
Like, they would acknowledge those changes,
but they're making technical bets right now
in some of these architectures for the future in two years.
And we have no idea what the future of models
or anything's gonna look like in two years.
So it's just interesting to look at bets they're making
and compare and contrast those with others,
again, in an environment that is highly unsettled.
And so we don't know, you know,
we don't know what the standardized inference architecture
will look like, I think it will emerge.
And then again, you know, say that some of the startups
are successful in proving that out, then, you know,
my concern is one of the bigger vendors
just takes those that design and does it if that's true.
But anyway, the diversity of technical approaches,
I found fascinating and thought was super interesting.
- Yeah, it was very interesting to see a number of companies,
four or five companies presented
with CPU architecture. And so in full disclosure,
I couldn't attend this year for family stuff,
but I was reading through all the papers
and reading all the follow-on content.
And it was really striking to me the number of CPUs
that were presented and GPUs that were presented
and just how different their architecture was.
Nvidia has these big cores and ARM has much smaller cores
and just the sort of number of nitty-gritty features
that are including, there is a lot of experimentation
taking place. And, you know, my read on that
is sort of what you're saying, which is,
we don't know what inference architecture
is gonna look like. And you say it's gonna standardize.
I'm not sure that we go down that path.
I think there's, it could remain very fragmented
and very different.
It is really striking how sort of radically different
some of these technical approaches are, so much variety.
It was a fun show. I wish I could have gone.
- Yeah, yeah.
- Well, and not just like, like some of them were like novel,
like you talked to the engineers in your real life,
they're really swinging for the fences with this one.
You know, like, it could work, but it's also really hard.
You know, and so, so, you know, I agree with you.
I don't know, I think at least, and agree,
so my take again, just, and I talk to everybody I could,
I had meetings with six of the leading startups in this,
was like, the challenges, there's a software tax
to all these different approaches, right?
If you're gonna put the KVCache somewhere else in memory
or storage, or you're gonna move it off direct,
the software has to do different work, right?
And so, because it's not just a clean,
it might be efficient from an algorithm standpoint,
but like some of the other agent orchestration and whatnot,
like, there's a little bit more work involved.
And so, there's a clean way to do this,
certain vendors do it, and the software,
and then there's a really hard way to do this,
which again, might be right if X, Y, and Z goes
according to plan, X, Y, and Z,
right, might not go according to plan.
So, but again, like, the idea that everybody's
making technical bets, I think, is super, super interesting.
And I'm assuming we will have some bespoke variety
by the custom hyperscalers, and then a couple of merchants.
But, you know, I think you probably agree with me,
they're 30 or 40 startups don't survive the day.
But maybe with the industry learns from what they tried
and apply it in their architectures,
because maybe it was philosophically or technically sound.
They're just not gonna be the ones that pull it off.
Somebody else will.
- Yeah, I mean, this is dilemma.
It goes back to what we're talking about before,
about NVIDIA sales racks, right?
They don't sell servers, they don't sell chips.
They sell racks. - They sell systems.
- Systems. - Platforms.
They would like us to see platforms, Jay.
I'm not saying systems, I'll go with systems.
I'm not giving it into platforms.
They sell systems, and if you're a startup
that sells, like you can only do so much, right?
And so your biggest ability to differentiate
is gonna be at the chip level, it's right.
And how do you then design a whole rack, whole system?
And I think that's challenging
because then you have to factor in networking, right?
If you're a chip company, so you could have
the world's greatest XPU AI accelerator,
and you lose all your TCO advantage
because you're using the wrong networking stack.
And I think that's right.
And so it's gonna be tough for a startup.
And I think it makes it very hard for them to sell.
And then let alone like everyone's competing
with balance sheet now too,
which you obviously can't do as a startup.
So yeah, I think it's gonna be tough,
but I do think some of them will get acquired
for, I won't put a number on the valuation,
but I think some, there is a likelihood
that these get acquired because they are coming up
with some really innovative things.
And with the market, it's hot as it is,
like I think the big players are gonna recognize
that they need rapid time to market
and pick these things up, but not so many of them,
or however many there are.
Yeah, so no, correct.
Okay, anyway, I thought that was interesting.
Memory I wanna spend some time on,
because the theme of all memory providers,
which again was also the theme of future memory summit.
So it's just carrying on with detail
is custom HBM base dies.
And it's interesting, I listened to this,
Micron didn't give any platform details on custom base die,
with both Samsung and SK Hynix did.
And it was interesting, they kind of point out like,
this is why it logically makes sense
to go to custom base dies.
And, but this is why it's hard, right?
It's a thermal problem, they gotta deal with hot spots.
It's, once you have that level of an ASIC
in terms of wattage, it gets harder.
So they're trying to solve that problem.
But the thing that all three of them kept mentioning,
in fact, and Micron CEO said this at the conference
they were at, I wanna say earlier in the week
or Alaska, I can't remember,
that they think about their work in custom based die,
more like you should think about custom ASICs.
That it is a unique design, customized for that customer.
And because of that, his point was,
and I just raised some questions to me
when talking to the Samsung and the SK guys,
is that it's harder for one vendor
to support three memory players,
maybe even two memory players,
when you've gone down that degree of customization
in based die with that person's memory
and your customization on their based die platform.
And I just gotta got me thinking like, okay,
so what if Micron's based die platform
is more efficient, leads to better efficiencies
with Micron memory than Samsung's does,
and now Nvidia's sitting there going,
I'm not saying this is happening,
like nobody freak about it,
so I'm just saying this is an example,
but then you have performance differences between vendors
for the same accelerator.
Like that's possible.
And I don't know how to reconcile that
with then the point I was gonna make with the Nvidia one,
how does that also not increase the cost
to do this for memory
that you're actually gonna customize a based die platform,
they're gonna charge you more for this,
it makes all the sense in the world
in optimizations and customizations to do this,
but it's gonna be more expensive.
So in a custom based die world,
like memory doesn't go down,
in fact it probably goes up,
and it's gonna behoove one to two vendors,
memory vendors per customer,
because I agree it would be very hard to support
three custom based dies for one GPU.
I think that's an unreasonable thing to expect.
- This is a tricky one, you think about it.
If you're a memory company,
that knock on memory companies is,
oh, it's just a commodity, it's interchangeable.
- Yeah, it's interchangeable.
- And so the memory companies probably see a lot of appeal
to some form of customization,
so that memory isn't totally fungible,
and it's much harder, you get customer lock-in.
But at the same time,
that commoditization of memory is valuable to them,
because it's interchangeable, because it's standard,
right, everything standardized,
and they risk losing that if they start going
everything custom.
But I understand the appeal,
I think it leads to complexity down the line
that it's hard for them to think through,
because the appeal being custom is so big.
- Agree, I also, I'm not saying everybody uses
custom-based type, but I'm saying,
like I guarantee you in video,
and AMD probably will, and maybe somebody else.
But just take your biggest customers, right,
maybe even your XP providers who might do this
for good reason, my point is exactly what you said
is that it becomes more design-in,
it becomes more co-optimized before the full stack,
but it's also less fungible,
and I would argue might have performance variations,
that vary.
Like, again, I just don't like, it would be,
it's a weird world to say my GPU with a micron memory,
or with Samsung memory, behaves differently,
performs better than my SK one, right,
like that's a weird environment to be in,
if that played out even in that fashion, right?
So anyway, everybody walked away like,
custom-based type is gonna be a thing,
we get it, yes it benefits as fenders,
but my point is, in no world does that help memory costs,
it changes it, but it's certainly less fungible
for the memory guys.
So I just thought that was like the talk of this,
everybody was like, this custom-based type stuff's gonna be
incredible, and I was like, yeah,
there's also some interesting trade-offs with it,
if it's gonna be a thing, which it seems it will.
So anyway, that was that.
Crazy.
Yeah, and the other stuff, like Samsung was showing
ZHBM, and then more 3D stacking,
I was like, that's great, if you get there someday,
and you can commercialize this and scale it wonderful,
but that's way far out, that's not.
In fact, I won't get into it too,
but like a guy, I can't remember where he was from,
I don't know if this was from a high band with flash company,
or if he was just giving a presentation on high band
with flash, but like, parts of it sounded positive,
but then it was actually kind of negative,
like had these slides, like, we gotta solve these problems,
and did you think about it this way?
And people were like, I don't think he's showing positivity
on high band with flash.
That would've got some good conversation in the Slack.
But anyway, yeah, it was a great show, like I said,
there was, it was about as busy as I've ever seen it.
I mean, people couldn't sit in the main auditorium,
they'd go upstairs, there was overflow outside,
yeah, it was big, and it's just gonna get bigger next year,
I'm sure, but anyway, it was a good show.
And lastly, I just say like, I like these shows,
and this just doubles down on my point of this era
of ferment concept, my standard view has always been,
if you just understand semiconductor company roadmaps,
you can predict the future because software performance
is directly aligned to where compute goes.
And that's why I like going to the show.
It was just interesting that like,
models all have to go in a very certain way
in order for all these different approaches
to have some degree of success,
and that's very hard to predict.
So it was like my, this confusion doesn't help my prediction
because it's so diverse, I'm not sure what it helps me predict,
but that's why I like going to the show is,
it's a glimpse of what's computationally possible
in the future.
- Did you listen to the Waymo presentation?
- Yeah, I was in Waymo, I didn't know.
- 'Cause I thought that was super interesting
because this sort of talks about,
touches on what we're talking about before
and your age of ferment, which is Waymo
had a big discussion on their approach to silicon
for their autonomous cars.
And it's very clear that their inference workload
is really, really different than what a lot of other people
are sort of with the frontier labs or dealing with.
Like they use, there's this big move to FP4,
it's a sort of less precision because it's more efficient
and you don't lose much, but you can't do that in a car
'cause you actually need super high degrees of precision
to, because you have all these different sensors
that are very, very fine tuned.
And that sort of, that means that their silicon
for their workload is gonna be really, really different.
It's, they're moving against a trend that's been pretty common.
Everyone else is moving down.
They don't wanna do that.
And I think that, you know, that sort of speaks to,
they're gonna be more of those.
We didn't talk about it, but like,
back when they announced earnings, AMD acquired TALUS,
which was a startup, which is a startup that sort of claimed
to be able to do custom silicon
for every individual model.
And they had a really interesting approach
where they sort of had, you know,
lots of different metal layers in their chip design
that were standardized and it's just the top layer
to take a switch around for each
model. So in theory, allowing very, very rapid iteration of custom silicon. And for, you
know, if you, their pitch was give us your model and we'll, we'll optimize the chip for
you. And that's a really, that's a really challenging approach for a lot of reasons. We don't
have to get into, but AMD acquired them. We have to assume for some, for more than just,
maybe just wanted a bunch of engineers, but I think that they're, they're attuned to
this that they're going to have to do more semi custom work for. Yeah, for, so everyone's
going to have a model. You, you bring up a really good point, actually, and the way
Mo guys, like, kind of top through, like, actually, this was the most fascinating part of
this conversation, because I had not thought about any of this, right? But, but he was making
points about, you know, what a, what a, what a, what an autonomous taxi has to do and
know is different than your autonomous car that you're operating and parking in a parking
ladder. And, and what he showed was, you know, this, this car needs to know, for example,
can it park right here? And this is a sign with a pole and it had like eight different
winger you could park their signs, not on Tuesday, not from 8AM on Wednesday. And he was
like, and it needs to know if it can stop here. And I was like, dang, you got, that's
a lot of image processing. What day of the week it is. Like, what time is it? How far is
my next ride? Like, it has to solve different things, right? Then a, you know, your passenger
car. And I just thought, but, but that's again, that's them saying, our problem is unique.
Our problem is also somewhat bespoke. And we need to go solve it a unique way. And
so, exactly like you said, I just thought that was super advantageous. The other thing
they showed was they showed the back. They showed the back of one of their cars from like,
three early on. Like, it was a rack. There was a rack of compute of like a little data center
in the back of this car. And like a wooden slash metal box. It's like, to be an engineer
on that product, be like, yeah, how are you going to as well? You know, we got to figure
out how to get a computer in the back seat. We're going to put this rack. It's connected
with all these crazy fibers. Like, anyway, it was super, super interesting. But it also
makes sense. Why? They feel like, like you said, they need to do their own models and they
need to do their own compute with their own ASIC, which stayed outline because again, they're
trying to solve a problem relatively unique to their vertical. And you've got all these
other things they got to deal with, right? Safety, compliance. You know, anyway, I could
go and other presentation was gone. I could go over, but we'll stop it. You're right.
Not novel approaches for the time being. All right. As a good one. Thanks for listening,
everybody. We'll talk to you next week when hopefully it won't be, I guess I was crazy,
but you never know. Fall is going to be crazy. Thanks for listening.
Thank you for listening, everybody. Leave us a review. Tell your agents. Leave us a review.
Tell your friends.
Podcast Summary
Key Points:
Nvidia guided to 70% revenue growth for the next year, signaling strong confidence in continued expansion despite supply constraints, with much of the growth tied to data center construction and TSMC capacity.
Despite price increases, Nvidia’s gross margins are declining, raising concerns about pricing power and supply chain pressure, especially as hyperscalers shift to direct memory procurement and reduce reliance on Nvidia’s bundled systems.
The semiconductor industry is in an era of ferment, marked by diverse technical approaches to AI inference and memory design, with companies like AMD, startups, and Nvidia offering competing architectures—highlighting uncertainty in future standardization and potential for fragmented innovation.
Summary:
Nvidia’s recent earnings showed strong confidence with a 70% revenue growth guidance for the next year, driven by data center demand and supply constraints, particularly at TSMC. While the company raised prices, gross margins are declining, indicating supply chain pressures and reduced pricing power, especially as hyperscalers increasingly source memory directly. The semiconductor landscape is experiencing a period of technological ferment, where diverse architectures for AI inference—ranging from custom ASICs to memory designs—show strong innovation but lack clear standardization.
Startups and major players like AMD and Samsung are presenting competing solutions, with many emphasizing custom designs for specific workloads. However, such customization increases complexity and cost, reducing memory fungibility and raising concerns about performance variability across vendors. This fragmentation suggests a future of specialized, vertically integrated solutions rather than universal standards.
The dynamic is further illustrated by Waymo’s unique approach to autonomous driving, requiring high-precision inference in a different workload than typical AI models. Overall, while innovation remains vibrant, the path to scalable, efficient, and interoperable AI hardware remains uncertain, with Nvidia positioning itself as a key enabler of the ecosystem despite challenges in margins and pricing.
FAQs
Nvidia guided to 70% revenue growth for the next year, which was a key point of investor interest and signaled strong confidence in continued demand despite supply constraints.
It signals strong market demand and confidence in future growth, especially in the data center sector, despite concerns about supply constraints and capacity limitations.
Nvidia has seen a decline in gross margins despite price increases, suggesting that supply chain pressures—especially in memory and components—are limiting their ability to pass on cost increases.
TSMC's manufacturing capacity is a key bottleneck; Nvidia's growth is heavily dependent on TSMC's ability to scale production, with significant capacity constraints expected through 2027.
Nvidia sells both full compute systems and individual components like CPUs and GPUs, with hyperscalers often buying components separately, leading to different pricing and margin structures.
The split into data center segments (including hyperscalers, enterprise, and sovereign) shows a more granular view of performance, with hyperscalers typically paying lower margins than smaller, more dependent customers.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.