Go back

GPU Hoarding is Over. The $401B Reality Check

15m 58s

GPU Hoarding is Over. The $401B Reality Check

The enterprise AI landscape is transitioning from a GPU hoarding panic to a cost-focused audit phase. VentureBeat’s Q1 data shows GPU availability concerns dropped from 20% to 15.4%, while cost per inference and TCO worries jumped from 34% to 41%. LinkedIn’s CTO Iran highlights their strategic advantage of owning the full stack—from data centers to bare metal—enabling optimizations like model pruning, embedding compression, and kernel tuning that are harder to achieve on public clouds. This allows them to balance quality and efficiency while rigorously evaluating ROI. The data also reveals a shift toward private AI: 72% of organizations lack sufficient AI control, driving 17% to adopt full-stack control (up from 11%). Inference now accounts for 60-80% of workloads on hyper-scaler infrastructure, as enterprises focus on token economics rather than hardware acquisition. Iran advises CTOs to treat AI as a long-term strategy, investing in instrumentation and tools to predict future traffic patterns (e.g., from ambient agents) and embedding efficiency into organizational culture. The next 24 months will be about squeezing every ounce of efficiency from existing stacks, not just securing them.

Transcription

2615 Words, 14798 Characters

English
The narrative is shifting in enterprise AI infrastructure. Many large enterprises were quietly over-provisioning their cloud instances, their GPU clusters and data centers, out of pure supply chain panic. They weren't really buying compute to use it. They were buying it as an insurance policy. Now the hoarding hangover has set in. In today's special episode Beyond the Pilot, I'm joined by venture-been analyst Rob Stretchey, an Iran burger, CTO of LinkedIn. Iran was early to talk about the spreadsheet math of AI. Venturebeats Q1 research shows the rest of the industry is finally hitting the same wall. This is Venturebeats Beyond the Pilot podcast. It's about enterprise AI in action. I'm Matt Marshall. Today's episode is presented by Outshift by Cisco, Cisco's emerging tech incubation engine, and driver of a gentick AI, quantum, next gen, infra, and beyond. Gentlemen, welcome to Beyond the Pilot. Thanks for having me. Glad to be on. Rob, let's start with you. We've been tracking this transition from the hoarding phase to the audit phase in our Q1 data. Walk us through what the numbers are telling us about how enterprise priorities have suddenly flipped. Yeah, I mean, I think again that it's a great point that, you know, a lot of the GPU availability concern had dropped over the quarter from 20% to 15.4%. Showing that people are not as concerned about access or availability. But at the same time, the concern about cost per inference and TCO or what I like to call the RO AI has jumped really from 34% up to 41%. And I see that climbing in the next few months as we get more data in. Ron, hearing that data, I'm curious about LinkedIn's reality. Was the GPU scramble ever a real panic for you at your team? At your scale? How soon did you start prioritizing the unit economics of the token? Yeah, you know, it's interesting for the majority of my time here. And I started an individual contributor here building our products compute and thinking about the cost of the feature I was shipping was never a concern. And then we started training, you know, multi-hundred million parameter models. Then we got the, you know, early preview of GPT-3 and sort of and realized, okay, like given our product ambitions now, things we need to operate pretty differently. And so we were able to, you know, starting a few years ago really prioritize. I guess two things, both the foundations, you know, tools infrastructure to do both training and inference efficiently. But also adapt the mindset of our engineering team that hadn't ever really had to think about the ROI of what they were building and make sure that before something was going to ramp to 100% production, we had done the ROI analysis on it. We had the right instrumentation to understand how much compute was using from what piece of infrastructure, especially the high-end infrastructure. And so we today thankfully have all of that in place. And we do have to think about ROI like everybody else, but, you know, we optimize utilization pretty well. It's clear that the game is changing from simply acquiring this, this hardware, as you're referring to, to actually making it work efficiently. And LinkedIn has that playbook, we call that a cookbook. When we talked to you a couple of months ago, just really intricate in terms of that efficiency. Rob, our research indicates a budget shift specifically into these inference frameworks and optimization tools. Can you give us a breakdown of what the data saying? Yeah, I mean, I think again, you know, when they looked at it, the cost per inference. And, you know, like I talked about before, had jumped from 34 to 41% when you started to look at how organizations were saying that they really needed more control and security over this AI. 72% of them admitted that they had a lack of sufficient control over it and that they were really moving towards inference and trying to help understand what was going on. And they're looking at different ways of building it as well with about 11% moving up to 17% where they're actually controlling their full stack from an infrastructure on up perspective. Because that is the best way that they see fit of being able to focus on that TCO and R.O.A.I. of the inference. And inference really, again, in talking to others, you know, the large Neo clouds and hyper-scalers, they're already seeing, you know, 60, 80% of the workloads happening on their infrastructure have really shifted to inference from training. And I think, again, most people have figured out that they're not in the training game and they're trying to weigh that battle between being a token consumer and a token producer and how they solve for that to actually meet their numbers. So let's go to Iran on that point. So Iran, congratulations, but I don't think we've mentioned it so far, but you've just become CTO as I understand it. So you're really looking at the infrastructure from the A.I. side, right? So this is a good question for you. Let's apply it here. We've talked about your cookbook. How do you define and drive productivity of the infrastructure over, you know, kind of simply token activity, right? You know, one thing that's that's pretty unique about us is that we're probably one of the last at scale applied ML shops. So if you take kind of the hyper-scalers out of the picture and we own and we own our full stack, right? From, you know, the experiences our users interact with, obviously, you know, predominantly powered by AI to bare metal. Right. We have our own data centers. We rack our own GPUs and that enable that creates a real strategic advantage for us when it comes to balancing quality and efficiency in our AI, right? We can look at the specific workloads that we want to run. And then we can optimize entire stacks. A really good example is what we talked about last time, you know, Matt, when we're talking about semantic people, semantic job search, we, you know, we can do things like model compression, sorry, model pruning to be able to make sure that, you know, we can reduce model size, which obviously means that we can get more throughput out of every GPU. We can do embedding compression. We can do context summarization. And we can develop tools and infrastructure to do these things to improve efficiency and and have that be something that's really easy for our applied ML engineers to do, right? Because at the end of the day, you know, they want to be focused on shipping products, not thinking about the costs of the products that they're shipping in half. But it goes, it goes a little further, right? You we are, you know, optimizing our own kernels for efficiency, or on GPU kernels for efficiency. We are, you know, adapting networking and storage based on the workload. So, so things that because we run our own, you know, private cloud, we can do. And that kind of optimization is a lot harder to do on a public cloud because you're basically choosing from a menu of, you know, instance and source skews rather than engineering the entire system and to end for the specific thing that you're building. So it truly is a strategic advantage for us. So so Rob, let's let's go to that point, right? So we're here, you know, LinkedIn obviously is conveniently located on the side where I think the markets going or it sounds like from the data we're seeing. So let's pivot to that data protection and ownership. So real quick. So as inference scales, so does the risk, right? And the market is reacting. What is the Q1 data show us about the move toward private AI and software, sovereign clouds? Yeah, I mean, I think again, that's one of the things that was really key to what was going on in here was that a lot of companies are very much like Iran said, they're looking at how do they own that stack and how do they own a lot of their kit so that they can control those costs because they're looking at it as well as a kind of a, I guess you could say, ability to control the narrative around the cost. And I think what they're looking at is they're looking at different technologies underneath the hood, such as RDMA, they're looking at open source in a big way as well. Things like VLLM and LLMD and a number of different distributions and optimizations for inference and moving models around. And they're looking at how do they bring these stacks and understand it. But at the same time, they're challenged by the fact that they just don't have enough people. And I think that was a key to what they were thinking about was that, hey, we got to keep it as, you know, straightforward as possible. We have to have multiple workloads in these environments. It's not just going to be AI, but how do we also balance the traditional apps with those new AI-based apps as well? This series is brought to you by Outchip, Cisco's Incubation Engine. By creating an open interoperable infrastructure, Outchip is enabling agents and humans to share intent, context and reasoning. The cognitive evolution for agents is here. Explore the Internet of cognition at outchip.com. So, Ron, LinkedIn operates on a massive amount of highly sensitive proprietary member data, rate to billions of folks. And as CTO, how do you balance this demand for high-speed, cost-effective inference with that control that's necessary, that data sovereignty required to protect your users? And when it comes to the board, how exactly do you justify this spreadsheet math and the return on investment in AI of these infrastructure choices? I think it starts with is understanding that we are serving our members and our customers first and foremost. And so when you start with that mindset, and that use as a first principle in your product design, and the architecture that supports it that puts, by default, guard rails that prevent you from doing things that are not in their best interest, from what you're referring to, spreadsheet math perspective, at the end of the day, we're running a business, right? And like actually the valuable credit for members and customers is in service of that business. And we need to be able to be really rigorous in evaluating the ROI of the investments we're making. And doing that well requires a bunch of foundations. You need to have the right instrumentation to know for any future I'm shipping, how much is it gonna cost me at scale, not just for the traffic we see this year, but what we will see in the future. You need to know how the business metrics that product AI product is driving is gonna move revenue. And then you need to look at the opportunity, cost of investing that, compute on something else as well. And so all that requires instrumentation, it requires rigor, engineering, rigor, and discipline. - Yeah, just to add to that, I mean, I think from the data, the data is backing that up, right? I mean, people are getting a little bit gun shy of just going and buying and hoarding GPUs. In fact, folks are looking at it, it fell from the 20s down into the teens, from a perspective of that, from 20.8 to 15.4. And it's kind of a retreat from that what we had been seeing. And I think also part of that, and I'm interested in this is too, is just the amount of memory and flash and the cost that's risen, some people were smart enough to hoard that early on. But when you start to look at how that is also impacting their abilities to go and buy these servers with the newest and greatest GPUs, it seems like people are kind of veering away from chasing the brand new fastest GPUs and saying, "Hey, what are the ones, what do I need?" And that goes back to Iran's comment on instrumentation and observability and understanding what does workload look like because AI is definitely different from traditional workloads in the fact that it's, you know, not just a inbound servout, it's inbound thinking out back, there's a lot more going on there than is transpiring with what have been, you know, like a traditional SAP application or somebody else's CRM system or something like that. The key here is doing all of that, but not thinking about the traffic you have right now, but like what you think, the shape of your traffic is gonna look like a few years from now, right, we're gonna have ambient agents, utilizing services like LinkedIn and the internet broadly, all the time, whether your user is awake or not, it doesn't matter whether they're on your site or not may not matter. And so the nature of the traffic is gonna shift, you need to make sure as a company, when you're modeling out your compute strategy, your infrastructure strategy that you account for that, because if you don't, you know, it's really hard to be in a position where actually, you need to stop everything you're doing because all of your resources need to go into efficiency while the rest of your competition keeps moving forward. Like this is really a long-term strategy play. - Well, gentlemen, this has been a great grounding conversation. It's clear that in the last 24 months, for most of us, we were talking about securing that AI stack. It was a lot of, you know, to hell with the costs, we're gonna experiment. And now it looks like the next 24 months are gonna be more about squeezing the stack for every ounce of efficiency. You know, unless, of course, you're an anthropic, you know, or an open AI where they're having arguments about whether they were actually provisioning enough or not, I think for the rest of us, it's really different. Final question, Iran, to close this out, what is your one piece of honest execution oriented advice for the enterprise CTO listening right now, who are staring at that, we get 5% GPU utilization reports, any parting thoughts? - Yeah, I mean, this is a long-term technology strategy exercise. So you need to think about what you need to do now so that you're prepared not just for your physical costs this year, but for the next few years, and that requires putting the right instrumentation in place, putting the right tools and infrastructure around that in place so that you can make it easy to drive efficiencies, ideally independent of the, you know, product use shipping product and it's at a lower level. And then lastly, you have kind of the organizational discipline and mindset to do this well, because for companies like ours, and even companies being started right now, I don't think that that is necessarily, you know, core to their culture and needs to be. - Iran, Rob, thank you for joining us. By creating an open interoperable infrastructure, Outchip is enabling agents and humans to share intent, context, and reasoning. Explore the Internet of Cognition at outchiff.com. For more stories about the AI revolution, like and subscribe to the podcast and check out venturebeat.com to sign up for our newsletters. (upbeat music)

Podcast Summary

Key Points:

  1. Enterprise AI infrastructure is shifting from GPU hoarding to an audit phase, with concern about GPU availability dropping from 20% to 15.4%, while cost per inference and TCO concerns rose from 34% to 41%.
  2. LinkedIn (with new CTO Iran) emphasizes a full-stack approach, owning data centers and GPUs to optimize efficiency through model pruning, embedding compression, and kernel optimization, prioritizing ROI analysis before production.
  3. 72% of organizations lack sufficient control over AI, leading to a move toward private AI and sovereign clouds, with 17% now controlling their full stack (up from 11%).
  4. Inference workloads now dominate, with 60-80% of infrastructure usage shifting from training to inference, and enterprises are focusing on unit economics of tokens rather than just acquiring hardware.
  5. Long-term strategy requires instrumentation, observability, and organizational discipline to plan for future traffic shapes (e.g., ambient agents) and avoid costly efficiency scrambles.

Summary:

The enterprise AI landscape is transitioning from a GPU hoarding panic to a cost-focused audit phase. 4%, while cost per inference and TCO worries jumped from 34% to 41%. LinkedIn’s CTO Iran highlights their strategic advantage of owning the full stack—from data centers to bare metal—enabling optimizations like model pruning, embedding compression, and kernel tuning that are harder to achieve on public clouds.

This allows them to balance quality and efficiency while rigorously evaluating ROI. The data also reveals a shift toward private AI: 72% of organizations lack sufficient AI control, driving 17% to adopt full-stack control (up from 11%). Inference now accounts for 60-80% of workloads on hyper-scaler infrastructure, as enterprises focus on token economics rather than hardware acquisition.

, from ambient agents) and embedding efficiency into organizational culture. The next 24 months will be about squeezing every ounce of efficiency from existing stacks, not just securing them.

FAQs

It refers to the shift from over-provisioning GPUs and compute resources out of supply chain panic to now focusing on efficiency and cost management, as GPU availability concerns dropped from 20% to 15.4%.

LinkedIn focused on foundations for training and inference, adapted engineering mindsets to include ROI analysis before production, and optimized utilization through tools like model pruning and kernel optimization.

Concern about cost per inference and total cost of ownership (RO AI) jumped from 34% to 41%, while GPU availability concerns decreased, indicating a shift from hoarding to auditing and efficiency.

To gain better control over costs and security, with 72% citing insufficient control, and 17% now managing full-stack infrastructure to optimize inference TCO and ROI.

By owning their full stack from bare metal to applications, they customize infrastructure for specific workloads, ensuring member trust while rigorously evaluating ROI with proper instrumentation.

Treat AI as a long-term strategy: invest in instrumentation and tools for efficiency now, and instill organizational discipline to prepare for future traffic shifts, like ambient agents.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.