Nvidia’s recent acquisition of Hugging Face and licensing of Poolside’s model factory underscores a clear strategic shift toward controlling the open-source AI ecosystem. By backing foundational tools like Hugging Face’s Transformers and enabling models to run on its CUDA architecture, Nvidia aims to dominate both model development and inference infrastructure. This approach strengthens its hardware sales by aligning model availability with its chips, creating a self-reinforcing ecosystem. However, it raises concerns about vendor lock-in and the suppression of competition among chipmakers, especially in the commodity hardware space. The situation echoes Microsoft’s successful acquisition of GitHub, which maintained open-source integrity despite integration. In contrast, Nvidia’s deeper integration with its hardware may lead to systemic bias in model accessibility and performance. Meanwhile, IBM’s announcement of a dual-processor mainframe chip combining Z-system reliability with ARM-based AI workloads marks a significant step toward integrating AI inference with legacy enterprise systems. This fusion allows enterprises to run critical workloads closer to data sources, reducing latency and easing migration burdens. It also highlights the enduring value of mainframe reliability in mission-critical environments. On a different front, OpenRouter’s mysterious OX Alpha model, revealed to be based on GLM 5.3, generated significant hype through stealth release tactics. While its claimed 100 trillion tokens per day throughput is likely due to scale rather than architectural innovation, the marketing strategy succeeded in generating intense public interest and speculation. The release demonstrates how strategic secrecy can amplify engagement, though its long-term competitive impact remains uncertain. Overall, the week highlights a growing trend of consolidation and vertical integration in AI—driven by hardware dominance, enterprise infrastructure needs, and strategic marketing—reshaping the competitive landscape of open-source and enterprise AI.
This is a clear line in the sand from Nvidia that they are going to depend.
Their future is going to depend on a strong, thriving, many model ecosystem.
All that and more on today's mixture of experts.
I'm Tim Huang and welcome to mixture of experts.
Each week, MOE brings together a panel of brilliant technologists working at the frontiers
of artificial intelligence to lead you through the week's news.
On this week's episode, we've got Gabe Goodhart, Chief Architect, AI Foundations, Ash
Minhas, Technical Content Manager, and AI Engineer, and Skylaspeakman, Senior Research
Scientist at the Software Innovation Lab.
Welcome to you all.
I'm also joined by my co-host today, Ily McCannan, who's a staff writer at IBM Think.
I always say this, like, there's a lot of news this week, but this week it felt like
there was maybe like an overwhelming amount of news.
So, we're going to cover a little bit of a new announcement about IBM's new next generation
dual processor.
We'll talk about OX Alpha and the kind of Mystique around it, which just as of today actually
was revealed to us behind it.
But I actually want to start by talking a little bit about what's happening with Nvidia.
It's a recent line of acquisitions and why it's doing what it's doing.
The headline here is that hugging face announced that it was going to be acquired by Nvidia
for $12.9 billion.
This follows on sort of just another deal not too long ago of a $6 billion kind of acquisition
transaction with poolside.
And I guess Gabe, you flagged both of these stories for us because you were just like,
this is all about open.
And so maybe I'll kick it over to you to give your hot take first.
Yeah.
I think this is a clear line in the sand from Nvidia that they are going to depend.
Their future is going to depend on a strong thriving many model ecosystem.
And right now that is open.
And right now they see that that is a critical lifeline for their company.
So they're willing to put money behind it.
They've clearly started down that track with their work on the Nemotron model line.
And the poolside investment is clearly a further investment in making Nemotron a strong
model line that people can rely on. And then this acquisition of hugging face is kind
of just taking it to the max, right?
Hugging face is the cornerstone of the open AI ecosystem right now.
And that's true of the models, but it's also true of the software.
The Transformers package is owned and operated by hugging face.
It is the reference architecture of literally every LLM on the planet that is released
in open source.
So every downstream architectural implementation follows from the Transformers one.
If I try to get something merged in Lama.cpp before it's merged in hugging fit in Transformers,
they say hold up.
We can't count on this until that other PR is merged.
So this is not, I mean, this is clearly a play for the center of gravity, right?
So Nvidia is making a clear statement here that they want to be the center of open source
AI. And that I think has some pros and some cons, right?
So a big company like Nvidia that is clearly well-moneyed, well-respected and in a very
strong position of power in this industry, it's great to have that support behind the open
ecosystem for those of us that live in that ecosystem.
It means we've got backing for years, which is amazing.
On the other hand, a big piece of the open ecosystem is in a fashion that drives intelligence
per weight downwards.
And in fact, incentivizes making the model smaller and easier to run on a wide variety
of hardware and putting that behind a vendor that is clearly biased in what hardware model
should run on and the size and cost of that hardware creates a bit of a perverse incentive.
So I'll be really curious to see how this plays out in terms of, is it, you know, I think
the obvious analogy that a lot of people are drawing is GitHub and Microsoft where GitHub
is the de facto standard for where code lives in the open.
And thus far, it seems like GitHub and Microsoft have done a pretty decent job of keeping a firewall
between the incentives of the closed source Microsoft ecosystem and the open source code
ecosystem in GitHub.
There's obviously product cross-pollination with co-pilot, but they haven't detracted
from the open source ethos of GitHub.
So the hope here is that Nvidia would do the same with hugging face and that hugging face
would remain a shining light of open source championship and openness and that, you know,
Nvidia would simply be putting their bet behind it.
So we'll see what happens.
I'd love to push more on what you've underlined.
I mean, my question is, who benefits more in this?
Is it the open source community or Nvidia?
You know, what in the best case scenario, how do they each help each other and, you know,
what would be some of the concerns?
And we can, you know, scatter ash if you want to jump in to join the conversation, please.
Yeah, remember when Microsoft bought GitHub?
It's basically that again.
Now, Microsoft has done a very good job of how the acquisition has gone.
I think that one of the interesting things with Nvidia and then requiring hugging face is
going to be that there's a lot of inference providers already providing inference to hugging
face.
And so how will that work now?
Are they reselling the same hardware that Nvidia are providing them that's going to be
some interesting dynamics start to appear from that perspective?
Open weights, open software, but closed hardware.
And Nvidia's play here is to get as many models trained in running on their cuda cores.
That's the larger draw here.
So on the one hand, we do benefit from having easier access to a wide range of models with
an asterisk as long as they run on Nvidia chips.
So I think that's what's really going on here, both with the hugging face purchase and
with the model factory from Poolside.
They want to turn these models out as fast as possible because they're still in the business
of selling hardware.
Well, I think on that note, it really leans into why this is important to Nvidia.
And why Nvidia is well positioned to be the one making this investment.
Pretty much every other significant vendor in this space is in either the model space itself
or the software space above the model.
So if you look at any of the neoclouds that are providing that inference, if you're looking
at any of the big hyperscalers, like their goal is to get you running, it doesn't particularly
matter what the chip in the computer is.
It matters what the API surface you're hitting is and what metering gateway you're going
through.
Nvidia doesn't care.
It cares about the chip under the hood.
So there's no disincentive for hugging face to continue providing a wide range of inference
providers, as long as all those inference providers are buying chips from Nvidia to power
the inference that they're running through hugging face.
So I do think there's an incentive alignment here that's fairly unique to Nvidia.
Again, I think the big losers in this deal will be the companies trying to make commodity
chips that are not Nvidia.
So one of the articles I read about this, clearly positioned this as a defensive play against
all of the labs that are trying to do vertical integration and build their own chips.
Because if the big labs stop needing to buy Nvidia chips, Nvidia needs to sell their chips
somewhere else, and so they're really going to aim at that commodity hardware, go broad
rather than deep type of sales play.
The ones that really lose out are the others attempting to sell broadly to general data
center consumers.
Yeah.
Can I play a bad guy here for a second?
Because obviously there's some discomfort in the room with Nvidia's nefarious plan here
to keep everybody on their closed hardware.
Isn't there another here which is like, get real, right?
Basically, these commodity players were never ever going to catch up.
The video was so far ahead, we're just recognizing reality, which is that like the game has been
lost in terms of certainly open hardware, but maybe even commodity hardware here, just
given the videos like extensive lead.
Do people agree with that?
This is in some ways realism versus them really kind of closing off an avenue for commodity
catch up.
I think to me, yes.
And I also think it's important to note that the smaller hardware vendors that are trying
to do commodity hardware that are not Nvidia, they're trying to soak up excess demand, frankly.
Like right now, we are in a very skewed supply demand ecosystem, and so this isn't going
to change that dynamic.
Now, if the demand really does drop off, that's where you're going to start to see this become
a bit more of a food fight, and you're absolutely right, Nvidia is going to throw its weight
around and maybe it's just calling a spade a spade.
But I think right now, all of the possibility of other vendors losing out is just speculation,
because frankly, if a data center needs chips, they're probably going to look for Nvidia
first.
But if it's a nine month waiting period, they might look elsewhere, and Nvidia's probably
okay with that, because they're still selling as many chips as they can make.
I think it's also interesting whether this is like a new, you know, taking this news,
you mentioned obviously the licensing of the poolside model factory.
It's just like a new, you know, bigger, broader playbook for AI consolidation, because
like in that case, what struck me was, you know, they weren't acquiring it.
They bought the license, they're hiring the team, you know, so it's sort of Nvidia kind
of grabbing land in a variety of different kind of manners, but perhaps that's too simplistic.
I think a key part of that poolside Nvidia conversation was a non-exclusive license.
So Nvidia is not necessarily doing a land grab there.
there are going to be one of perhaps multiple.
companies that are going to be using pool, using pool sides approach to model creation.
So that, that was kind of in the details there of that agreement.
So yeah, we'll be interesting to see who else might be in line to pay $6 billion to pool side
for their license. I don't think there'll be too many customers left in that. But this is not
exclusive pool side and Nvidia in that particular arrangement. Yeah, that's maybe where I want to
kind of like do the final few questions on this segment is, you know, normally when you spend
19, 20 billion dollars, you're usually a little bit short on the money. But we're also talking about
Nvidia here where they could go for a lot more. I guess the question is if Ash I'm curious
if you have any thoughts on like who's who's next up here? Where else would you go if you were
trying to kind of consolidate in the open space? Do you think the video is going to kind of make
another another play or is this kind of like it for a little bit? When we look at sort of like what
does the AI world look like? I mean, the stack is like, you know, chips and infrastructure,
then it's like the cloud providers providing that. And then you've got the application layer.
I think with both of these things, it's like infrastructure locked down. You know, they're in
that middle layer now and like they want to make really good models. They're going to use obviously
that pool side acquisition to make their own Neemotron models better. I'm curious. I mean, I know I'm
responding to your question, Tim, with a question. But the question is it's like, are they going to
go even higher up the stack and start looking at some of that application layer and maybe going
for a slice of that somewhere and making that open or, you know, is that like, you know, going to
go grab open web UI? They've already taken a small step in that direction with Neemotron, right? So
they have put, they're probably one of the largest companies behind an open source agent harness
at this point. So, you know, that there was a big announcement a little while ago about Neemotron,
and it's kind of, I wouldn't say it's it's petered out, but there hasn't been as much hype about
the claws as of late. I think a few other agent harnesses like Hermes have taken up a little bit
more of the oxygen for the generalist agent. But, you know, I guess I would be surprised to see
Nvidia try to go the actual like four sale, you know, by this as a consumer software platform
route. You know, I don't think we'll be going to like ai.nvidia.com and, you know,
having our agent chats there the way folks do with chat GPT or Claude. It just doesn't seem like
that incentivizes the right things because ultimately their business is still built completely
around the chips that are under the hood. So, I think for me the two recent deals we're talking
about here are just a clear sign that they are going for as many flowers blooming as possible,
right? And if they can put their money and their mouth behind specific process that they think
amplify that goal, they will. But I don't see them, you know, trying to then also carve off a
corner of the market and say, but forget all those other flowers come to our specific flower.
I'm going to move us on to our next topic, which I guess is related in some ways, but it's a
it's a different side of the market entirely. We've talked about the mainframe market here on MOE
in the past. And one of the things I love about the market is that it's very different from what we
know of usual, you know, things around AI where the general approach of AI industry has been just like,
all right, let's just launch it and see what happens. With mainframes, you really kind of can't
afford that in some ways. And Ily, I know you flag this one out if you want to introduce it, but
sounds like IBM is basically out with a new dual processor in the mainframe space, which is a kind
of big deal just given, you know, how cautious people are in mainframes. Do you want to talk more about
it? Sure, sure. And yeah, they presented this at this week was hot chips, which is a big kind of
semi conductor annual conference kind of chip makers. Researchers come together and veil new
processors and designs. And, you know, many of the participants were, you know, some of those
gave that you mentioned the frontier companies who've all come up with their own kind of custom
chips that they're they're sharing. But in this context, IBM was announcing kind of the different
type of chip, you know, this dual process architecture. And so this one they've developed with ARM,
the AI software ecosystem and it was different about it is it combines IBM's Z to the mainframe
workload with ARM and all on one chip. And with the ability for the ultimately to be able to
switch between those two workloads, you know, in nanoseconds. So I think the bigger picture, you
know, thinking is that there's a growing challenge in enterprise AI where organizations, you know,
want to run their inference closer to where that data resides, you know, and the reality is that
much of that data sits on mainframes, you know, many of those IBM's Z and Linux One systems.
And then at the same time, you have this ARM ecosystem where a large share of the modern AI
software, you know, is being built. So I think that, you know, the idea is rather than moving
the data between these two places, IBM is sort of, you know, proposing this chip to bring those
workloads closer together. And maybe, Skylar, if you're game to, you know, jump in on this first,
I'm interested, you know, from an infrastructure perspective, why does this matter, you know,
how big of a deal is this? If you get all the way down to the midi-goody details about how these
processors work, there's two different philosophies. One of them is doing multiple simpler instructions.
So if you want to add two numbers, there's one instruction to go read a number from memory,
there's another instruction to add it, and there's another instruction to write it.
The other philosophy is a single more complex instruction that does that all end to end.
And so computer science over decades has evolved from kind of those two different philosophies.
And it's really interesting now to see these worlds somewhat collide in the IBM ARM space.
I do not know the technology behind it about how they're doing both of those instruction sets
on the same chip, kudos to them, very smart people working on that. But it really, I think, is going
to be interesting to, you know, ask questions like, is your bank going to go to the App Store and
update its core banking software? You know, can you go to an ATM and do a transaction and say,
wait a minute, our backend is training a model right now, give us, give us 20 seconds.
And those are questions we had to ask before because these things have been completely separate
ecosystems. And now IBM and ARM are saying still mainframe functionality, but the ability to run
these two different instruction sets. And they have, again, what I want to emphasize here is two
different philosophies of how you go about writing code coming together under one place.
And the thing that it kept bothering me when I was reading this is,
on the one hand, it's cool to see these fences being brought down, but there's also this approach
where you should be asking, why were those fences there to begin with? Do we really want to have
ARM software sitting in the same places as, for example, our government services, banking services,
and the things that have relied on mainframes for decades? So watch this space. I don't know who
all's going to be, can comment on that. But this is, I think this is bigger than just a single chip
sharing two different instruction sets. It really is two different worlds colliding. And it'll be
really cool to see it play out in the coming months. And I think it sort of, yeah, I think to your
point of it being a larger question to it's, is this looking at enterprises trying to integrate,
you know, AI with these mission critical systems, these, you know, established systems that,
you know, they can't just totally rip out and then, you know, kind of bring in, you know,
something totally new. And the analogy, do you do so? The app store, someone, as I was talking
about this news, it explained it like that. Like suddenly, you know, it's like mainframe consumers,
you know, have this whole app store of options, you know, if they're an enterprise, you know,
hoping to bring in the inference kind of closer to where their data is sitting and it's, you know,
harder to get out of. I think both IBM and ARM are winners in this case as of now. I think,
I think right now, I think this really is a, I don't think this is a zero sum game here. I think
there's opportunities for, for both of these kind of established players. It's cool to see an
announcement back in April of the partnership. And then a few months later, the announcement of the
chip will still have to wait for its actual release. Maybe, maybe a bit doubtful for that. But it's
cool to see that sort of quick turnaround on this type of partnership. Yeah. So I want to bring in
the perspective of a software engineer on this, which is, this is going to make a lot of people's
lives a whole lot easier. So, you know, a few months ago, there was a big kerfuffle in the market
that certainly affected us at IBM, where Anthropic announced their migration tool off of Cobal.
And the market worried that that meant the death of the mainframe. And we had a lot of well-articulated
responses out of IBM. But I think the core of all those responses was, look, Cobal is a means
to an end, but the end has not gone away. Right? The reason mainframes exist is bullet proof
for liability. Like I was talking to, I forget who it was earlier this week, who was literally part
of the team that said, you could literally take the mainframe out back and put a bullet through it,
and it would not stop processing transactions. And that's just not true of any other piece of
hardware. Right? Like, you can, you could shoot a bullet through a ramstick on a Z system,
and it wouldn't blink. Now, that reliability is what the foundation of many of our most
pieces, critical pieces of infrastructure is built on. And so, yes, there's a lot of code trapped
in Cobal, which is a pain in the neck, and is probably very difficult for companies to find
programmers that can actually handle that code. But,
the need for that reliability has gone nowhere. And so this is an attempt to say, well,
we can solve that problem in a different way rather than bringing the mainframe code
out and then sacrificing that reliability that you've come to depend on. Why don't we bring
the modern code in, right? And as someone who has tried to cross compile code for arbitrary
random architectures, if you're thinking about import torches, a Python program and you're
like, what does cross compiling mean? Well, let me tell you, Python is a C program. Every single
library that you import in Python is in fact delegating down to a C library if it's got any kind
of performance behind it. Torch itself has backends compiled against every single accelerator architecture
out there. So it is a massive pain in the neck of, you know, ecosystem targeted cross compilation.
And the idea before this announcement of taking something as complicated as an LLM that is
backed by all of this complicated acceleration logic and cross compiling it all to a completely bespoke,
bespoke is the wrong word, completely firewall and separate chip architecture, both for the accelerator
and for the standard CPU processing was a daunting task. And this was the the purview of, you know,
potentially years of work to get, you know, individual LLM architectures ported over to a Z
system. Now with ARM and the fact that a huge number of people are running ARM workloads on
their laptops, on their phones, even starting to be on their desktop processors and server processors
means that a huge amount of software has already been adapted to the ARM ecosystem. So that work is
done for us. So this is hopefully going to open up the floodgates of bringing a ton of interesting
workloads to the mainframe that just could not run there in any reasonable amount of effort before.
Now to your point, Skylar, it'll be interesting to see what that does for the reliability of the
other workloads that are sharing the processor. So the proof will be in the pudding there. But
at least on the surface of it, this looks like this is going to make a lot of people's lives a lot
easier. Yeah, and can we talk a little bit about the future, Ash, like I'm thinking about like in
2050, you know, well, we still have mainframes. It feels like this technology that people like
really gripe about and are always like the mainframes on its way out, the mainframes on its way out.
But to Gabe's point, right, like it is it is rare to find reliability of this kind.
And, you know, I guess the question is I guess in some ways you could almost read the trend line.
It's like actually mainframes will be here and sort of bigger than ever. But do you buy that? I
guess I'm kind of curious about like whether or not this like pretty dusty kind of ecosystem in some
ways is like much, much more robust and interesting than it looks. I think that there are certain workloads
where latency really makes a difference. Okay. And the latency issue just doesn't really exist in
mainframe. You know, you're getting like transactions processed in such a fast time that we will
always need the ability to do that with certain types of transactions. And so I think that we will
trend towards sort of bigger, more powerful computers of this kind, pretty much forever really.
The point that Skylar and both Gabe made it, I also like sort of have some questions I would say
right now and curiosity. You know, this is a great sort of thing that's come out of this collaboration
with ARM. And you know, I guess when I think about this from sort of like a computer architecture
perspective, they could have just put some ARM cores into the chip and they didn't. Okay. It's all
combined in one. And that's really, really cool and really, really interesting. I would love to
like go and talk to the people who made that architectural decision and go, why did you decide
to go the hard way? Because it could have been easier if you did it the other way. And yeah,
also that then raises questions of like, well, you know, we're trying to process these things
on mainframe today, which are like sort of millisecond type transactions. And so now if you've got,
you know, I don't know, let's say pie torch running on there and something else. Okay. What does
that do to the behavior of the overall system? And is that going to mean that it will mean a much,
much bigger computer in 2050? Yeah, that would be a really funny outcome. It's just like we're living
in 2050 and it's like giant mainframes. That'd be really interesting to see.
I'm going to move us on to our kind of last topic of the day. You know, I think it's an
adage in 2026 that it feels increasingly like we're living in a cyberpunk novel of some kinds.
This story I think was very much like it. Basically on OpenRouter, there's this mystery model
called OX Alpha that kind of hit and people are very impressed by its capability. And most
interestingly, you know, at least in the announcement, it was claimed that it was able to serve 100 trillion
tokens a day. And, you know, which immediately raised questions about like, what kind of compute
are you running to be able to offer this kind of thing? I think as of this morning, it has finally
been revealed that it is GLM 5.3 open source Chinese model. And I guess maybe the first thing is we
should just do the usual vibe check test. Gabe, Skyward, Ash, I'm curious if any of you kind of
played with the model. What do you think about it just from a, you know, almost kind of like a
wine review? Like what do you think about it? Is the hype justified? Benchmarks are one of these
things that I think in this like probabilistic world are going to just be contested forever.
So at some point, I don't know about a year ago, I just put together sort of like, you know,
my own level like Benchmarks week, whatever. And I ran that through OpenRouter with it. And like,
thought, you know, well, this is actually pretty good. Pretty good. I mean, it's great that it's
free. It's not free now, but it's still dirt cheap. And I was pretty impressed with the output
that I got. Yeah, I mean, I gave it the smell test as well. It smells like a good model. But like
with wine, I at this point can't actually claim a connoisseur's palette. Because most of the work loads
I want to run can be satisfied by a sub 30 billion parameter model on my DGX sparks. So I think
you know, yeah. So, you know, little plug for small models there. But you know, my feeling is that
it is fantastic to have competition at the top end. Love seeing it. Very curious whether this
throughput number that they are claiming is due to breadth of scaling and just raw compute or
whether they've done something truly novel in the model architecture that enables massive
throughput improvements while maintaining this top end quality. That would be genuinely very
cool if they figured some tricks out about how to, you know, get the bits through the attention
mechanism faster. But at the end of the day, we have a glut of very, very, very good models. And
that's awesome. And I'm still going to try to run as much as I can off my local machine.
My interest in this story is as much about how they decide to go secret. And, you know,
I'm interested in your thoughts on you to presumably intentional. Does it build hype, you know,
is there more to it than that that they, you know, it was going to be revealed, you know, at some
point, you know, why the why the stealth mode? I think the stealth mode worked. I mean, in between
the time where we had these topics chosen for this podcast and when it's actually recorded,
it came out as GLM 5.3. Would we be talking about the release of GLM 5.3 if that was just how it
would came about? Or is it so much more fascinating to have this idea of a hidden model as Tim pointed
out? It's like a cyberpunk novel where you can have these, you know, mysterious people showing up
and competing in a, I don't know, a medieval joust with Helmets on it. We don't know who they are.
And so I think definitely well done to the makers of the model to release it in this way. They got
the hype they wanted to a few days of great speculation. And then they released saying it is
just, you know, an improvement, not just it is an impressive improvement over one of their
previous models. So I think I think they really did a great job with the hidden reveal,
letting me internet talk about it for two or three days. And I will say another thing to that,
which is that timing a model release is very hard. You know, we at IBM just released the
granite 4.2 models, which, you know, we are proud of. They are certainly not going to win the benchmark
race, but we think that they're going to be very useful for the target audience. However,
a day after we did that, we had Quinn drop yet another benchmark busting, amazing model that
also runs on the same local hardware. And, you know, picking when you want to release and how you
want to release is a very difficult game of speculation on who else is going to be releasing
competitive models in the same size and the same space. So, so like you said, Skylar, I mean,
Kudos to the team for choosing a different route to release, because if this had been just
another amazing Chinese model, like what a world we live in, by the way, you know, would we even
bother to talk about it? But choosing a release strategy that almost was competition proof because
nobody else was doing. Now, I don't think anyone else can probably pull this off again for a little
while. I mean, we had it kind of with nano banana for a while. This isn't the first time a lab
has sort of stealth launched a model with a catchy name, but it is a good strategy to mitigate
against. Oh, shoot is one of the other frontier labs also going to release their new model on the same
day or the same week? Who knows? Or sometimes within, you know, I'm recalling one open-eye and
throbic pair of releases within like an hour of each other where it's sort of, it becomes a sort of
playground battle of who gets most attention. I think that they did have like a like an announcement
to say, oh, the
wanted to release it in a sort of stealth mode so they could get unbiased, actual real feedback
from developers to know, you know, how good their model is in all of those use cases.
I do think there's probably an element of that in like why it was stealth, right? You know,
they didn't want to say it's a Chinese model and for it to like introduce any sort of bias
into determination as to the model's performance. I do also think that there's probably 50% of it
is marketing, right? It's a great marketing thing to do. That's great. Well, that's all the time
that we have for today. Gabe Skye, Ash, was great to have you on the show and I always great
co-hosting with you today. Thank you, Tim. And thanks for joining all your listeners. If you enjoyed
what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere.
And we'll see you all next week on mixture of experts.
Podcast Summary
Key Points:
Nvidia's acquisition of Hugging Face for $12.9 billion signals a strategic commitment to dominating the open-source AI ecosystem, positioning it as the central hub for models and software like Transformers.
This move, coupled with its licensing of Poolside’s model factory, reflects a broader strategy to consolidate AI development, ensure models run on Nvidia’s CUDA hardware, and create a closed, hardware-optimized ecosystem that may disadvantage non-Nvidia chipmakers.
While the open-source community benefits from increased backing and access, concerns remain about reduced hardware diversity, potential vendor lock-in, and the risk of diminishing innovation in commodity chips, raising questions about long-term fairness and competition.
Summary:
Nvidia’s recent acquisition of Hugging Face and licensing of Poolside’s model factory underscores a clear strategic shift toward controlling the open-source AI ecosystem. By backing foundational tools like Hugging Face’s Transformers and enabling models to run on its CUDA architecture, Nvidia aims to dominate both model development and inference infrastructure. This approach strengthens its hardware sales by aligning model availability with its chips, creating a self-reinforcing ecosystem.
However, it raises concerns about vendor lock-in and the suppression of competition among chipmakers, especially in the commodity hardware space. The situation echoes Microsoft’s successful acquisition of GitHub, which maintained open-source integrity despite integration. In contrast, Nvidia’s deeper integration with its hardware may lead to systemic bias in model accessibility and performance.
Meanwhile, IBM’s announcement of a dual-processor mainframe chip combining Z-system reliability with ARM-based AI workloads marks a significant step toward integrating AI inference with legacy enterprise systems. This fusion allows enterprises to run critical workloads closer to data sources, reducing latency and easing migration burdens. It also highlights the enduring value of mainframe reliability in mission-critical environments.
3, generated significant hype through stealth release tactics. While its claimed 100 trillion tokens per day throughput is likely due to scale rather than architectural innovation, the marketing strategy succeeded in generating intense public interest and speculation. The release demonstrates how strategic secrecy can amplify engagement, though its long-term competitive impact remains uncertain.
Overall, the week highlights a growing trend of consolidation and vertical integration in AI—driven by hardware dominance, enterprise infrastructure needs, and strategic marketing—reshaping the competitive landscape of open-source and enterprise AI.
FAQs
Nvidia is investing in Hugging Face to strengthen its position in the open-source AI ecosystem, aiming to dominate the foundation of open models and software like Transformers. This move supports a broader ecosystem centered on open models and aligns with Nvidia’s goal of ensuring widespread model deployment on its CUDA hardware.
It provides strong financial and institutional backing to open-source AI, potentially increasing stability and innovation. However, concerns exist that Nvidia’s influence may prioritize hardware compatibility, potentially limiting model diversity and accessibility for non-Nvidia hardware.
Nvidia acquired a non-exclusive license from Poolside to integrate its model factory approach into its own Neemotron line, enabling faster model development and deployment. This supports Nvidia’s goal of expanding a robust, scalable model ecosystem without exclusive ownership.
A major concern is that Nvidia may incentivize smaller, more efficient models optimized for its hardware, potentially disadvantaging open, commodity hardware vendors and reducing innovation in diverse hardware ecosystems.
IBM has launched a dual-processor chip combining mainframe (Z-systems) and ARM workloads on a single chip, allowing seamless switching between workloads in nanoseconds. This enables enterprises to run AI inference closer to data sources without moving data or relying on separate systems.
It simplifies integration of AI models into legacy systems by leveraging existing ARM software ecosystems, reducing cross-compilation challenges and making it easier to deploy modern AI workloads on mission-critical mainframes.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.