Go back

Building AI Factories: How Red Hat and NVIDIA Turn Enterprise Data Into Intelligence - Ep. 293

38m 43s

Building AI Factories: How Red Hat and NVIDIA Turn Enterprise Data Into Intelligence - Ep. 293

The discussion centers on the concept of an "AI factory" as essential infrastructure for enterprises to harness AI effectively. An AI factory is described as a five-layer stack encompassing data center resources, specialized hardware (like NVIDIA's GPUs), software orchestration, AI models, and end applications or agents. Its primary purpose is to convert enterprise data into valuable intelligence, enabling companies to innovate, improve efficiency, and grow revenue while transitioning to an AI-native operational model. A key emphasis is on building these systems with trust, incorporating robust security, governance, and compliance to protect sensitive data and integrate with legacy business tools. The conversation highlights the rapid evolution toward agentic AI (e.g., autonomous coding assistants and enterprise search agents), which promises major productivity benefits but requires disciplined deployment to prevent fragmented, insecure "shadow IT." The partnership between NVIDIA and Red Hat is presented as a solution, combining NVIDIA's performance-optimized hardware and software with Red Hat's enterprise-grade platform for orchestration, security, and lifecycle management. Recommendations for enterprises include starting with specific, high-value use cases, using reference blueprints, and adopting a hybrid approach that balances powerful frontier models with cost-effective open models for scalable, production-ready AI inference.

Transcription

6476 Words, 36665 Characters

English
[MUSIC] >> Welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. My guests today are Red Hat's Chris Wright and NVIDIA's Justin Boyzano and we're talking AI factories. Why should enterprises build AI factories? And how can they do so with confidence in building AI factories that they can trust? By a way of introductions, and I'll keep it brief because both of these guys work speaks for itself really. Chris Wright is Chief Technology Officer and Senior Vice President of Global Engineering at Red Hat. And Justin Boyzano is Vice President and General Manager of Enterprise Computing at NVIDIA. Gentlemen, welcome to the NVIDIA AI Podcast. Thank you so much for taking the time to join us. >> Thanks for having us. >> Thanks for having me, Noah. >> Let's get right into it and just and I'll start with you, but always both of you guys feel free to jump in as the spirit moves you so to speak as we go. But Justin, why don't we start with you? Can you talk a little bit about, well, maybe first give kind of a working definition of what we mean, what you mean when we talk about an AI factory. And then get into kind of at a high level, why wouldn't enterprise be interested? Why are enterprises building AI factories and what are some of the tangible benefits that an enterprise can expect to see from an AI factory? >> Sure, Noah. >> Yeah, I think it's important to understand kind of the context of where we are as an industry. And building digital intelligence to power the productivity of organizations is going to be as critical in this decade as energy in running our companies. This is the next industrial revolution and companies are always asking us, how do we build these factories that basically take data in and then produce the intelligence that helps them run their businesses more efficiently. And so as we talk about what is an AI factory, we think of them as really kind of five layers of technology that need to come together. At the base layer, you've got to make sure that you have got the data centers with power to bring into these factories. You've got to have chips, is the easy way to talk about it. But we're at this point in building rack scale infrastructure that's six chips. With the extreme co-designed to build the best token efficiency from the power available to you. The next layer, you typically want to have the software infrastructure to orchestrate everything. And then you want to have models that run that intelligence and then ultimately the apps and the agents on top. And so what every business needs to do though is take this intelligence and build use case specific business outcomes that help them drive innovation, build products faster and ultimately grow revenue, top line through deploying this intelligence at scale. Right. And so these five layers you're referring to, this is the cake, right? The five layer cake. That's right. This is the five layer cake. Excellent. Chris, the world is, I feel like we can say this so often, but things are changing so quickly. Right now, as we record this, there's a lot of talk about open claw and autonomous agents and kind of long running agents. Can you speak kind of a little bit sort of to that and how Nvidia and Red Hat are working together to help enterprise enterprise IT departments kind of step into this new world? Yeah, actually open claw is a great example because there's so much enthusiasm about, I guess, what's possible, what you could do. It's captured the kind of the builders imagination, but also built quite quickly, certainly leveraging AI to help produce code quickly. But not with the enterprise in mind. So when we think about what Justin was describing, that kind of data in to a factory context that produces business value as an output. We're talking about enterprises. That's their data. Those business outcomes are really either driving net new growth or focused on the productivity and efficiency. All of that needs to be done responsibly, safely, respecting access controls, delivering audit trails, things that are maybe not as fun in the builder world, but fundamental to the enterprise world. And so a lot of what we're doing is taking these building blocks, the layers of that five layer cake and making them accessible to the enterprise together. So obviously Nvidia's world class hardware, we're bringing a software layer that enables the higher levels of that cake. And then we're building the right guard rails and security considerations into this combined solution. So there are customers can then feel confident about bringing this into their enterprises. They're all trying to figure out how to do AI transformation, go from a traditional company to really an AI native company. And in that context, not introduce undue risk or essentially undermine the core of their business. Right. So there's research that shows that only 1% of organizations right now have reached the stage of an optimized AI fueled AI native as you were talking about, Chris, enterprise. While over half of organizations still remain in the early stages of transformation. But at the same time, projections have global AI investment exceeding a trillion dollars total by 2029 just a few years out. And of that trillion dollars, these projections are saying, "Agentic systems are going to account for roughly half of that spending." That's a big shift from a year ago, two years ago, you guys know the time for it better than I would. But when agents were kind of this buzzword that nobody necessarily knew there are all these different definitions, etc. And now we're talking about all of this resource and spending going in specifically to agentic systems. Chris, what can we glean from this? And I know you spoke to it a little bit just now, but what are the kinds of things that the AI factory can do for an enterprise, infrastructure-wise, but confidence-wise as you were talking about when it comes specifically to figuring out how to deploy and integrate these agentic systems? Well, if you think about that notion of transforming the enterprise and leveraging internal data and focusing on your core business, about how do you improve it or grow it, there's a whole set of things that are underneath that. Obviously, the data piece that we talked about, but also it is the existing tools that operate your business that are not going to just go away. They're fundamental, they're the baseline, the business's usual components, pretty critical and fundamental. So part of this is how do you carry that forward and really modernize your entire infrastructure to bring these two worlds together, this highly modern AI Native world and the traditional set of applications that literally run the business, because you need to bring AI capabilities not just in the net new, but also in the existing content that runs all the enterprise. And to me, that's exactly what the AI factory does. It helps bridge these two worlds. I mean, the end we've got models, but we also have, as Justin described at the beginning, agentic content or AI-enabled applications and then also the traditional applications. So bringing all of that together and then doing it at a context with consistency across the enterprise, so that you're not asking every team to go figure out their own, choose your own adventure path forward, and that consistency you build best practices across your organization, and then you're ultimately improving your chances for success and reducing the failure race. There's so many studies that suggest a lot of AI projects can fail. There's a number of reasons for that. One of those is having the right tools and having the best practices, access to the data and combining, essentially combining forces as a company to produce and output rather than devolving into sort of the next generation of Shadow IT, and everybody building their own thing and creating this highly fragmented internal environment, which is then difficult to get your arms around if not just produce very little success. Yeah. Justin, are you seeing some other things? Well, I got to say, what's interesting is in the last three months, it feels like the market is really started moving even faster. I'll just say this. You look at coding companies. Every one of you guests who comes on this podcast says that same thing, moving faster, moving faster. Well, you can actually really feel it now. I say that because the first area of agents, product market fit, really was in software development. We see it as a software company ourselves. We can feel these agents doing so much more work for our developers and running longer, more complex software tasks. You give them design goals and they can work towards those goals. At the same time, like you said, this moment of clause came out, and clause basically take it to a new frontier of full autonomy. We're getting to this point where agents are going to have a lot more agency within our enterprise. A lot of those studies that you mentioned where people were having a hard time getting AI to work, I think was at a previous era of the world where people were trying to do chatbots, just very basic chatbots. That was before reasoning and it was before this level of autonomy that I'm talking about. I feel like a lot of what enterprises might have been experimenting with might be a couple of generations. behind where state of the artist right now. And so as we deploy agents internally now that can use a very, I'll call it, deep agent like reasoning framework. They can plan and reason and act across many different business systems to do deep research as an example, to understand kind of the intent of what a user might be asking and help them get to the information across the enterprise in a way that's faster and more efficient than ever, previously thought imaginable. And the nice thing about running this on a factory and AI factory within the context of an enterprise is as Chris mentioned, it delivers data privacy and security by running that all across open models in this on-prem world. And then you can do things where you still potentially use the frontier models but you can use the frontier models in a way where you might only use it for the planning stage of the agent and all the search and summarization is using open models. And so that drives a lot of cost efficiency. In some of our newer blueprints, we see a 30X cost reduction by doing a hybrid model architecture across your private unstructured information. And so that is a use case, the enterprise search, I think is a broadly generalized use case. They get us from, you know, say these early adopters that were seeing the benefits of agents for coding into really how knowledge workers are gonna start to use agents to help them do their jobs in a much more productive and efficient way. - As somebody who sits more on the knowledge worker than software developers, side of the fence myself, getting me more towards that and away from vibe coding is probably a good idea, but that's just my own sort of personal use case there. But that does make me want to double click a little bit on security and governance and things like this, which, you know, I think Chris, you mentioned at the top with the advent of, I mean, joking aside with the advent of, you know, coding tools, vibe coding tools and these more advanced agenda coding tools in the hands of anybody, including folks like me, it's easy to spin something up. I don't know that it has, you know, a hole in it waiting for a prompt injection attack or whatever the case may be, right? And get into that shadow IT world, Chris, you were talking about. So I want to ask you both, and Justin, I'll start with you because you were talking about a little bit just now. When you talk about planning and building an AI factory, what are the non-negotiable capabilities that have to be built in that the enterprise must have to move from, you know, kind of first experiments and prototypes with AI to production, getting into industrial scale production AI use cases with confidence. And, you know, Justin, you mentioned some of these, but there's security, there's governance, reliability, obviously moving to scale. You talk a little bit about some of these factors. - Yeah, and I think, I'll say in this software development, where I'm really good at separating the notion of development versus production. And I think that's obviously the best, best practice as enterprises get going is to separate the two. And on the one hand, you want to help your internal, I'll call it AI development teams, do discovery in the development environment, but separate that access control from, you know, production data until you've basically, you know, proven verification or done the functional verification of the outcome that you're trying to get to, you've QAIDate, you've PEN test it, it's got things like role-based access control so that if a user's using that agent, it inherits their permissions to access business systems. And you're gonna promote, you know, the agent from this development environment into prod in that way. And so I think, you know, I think the worst thing that enterprises can do is overanalyze this though and try and get to, like, how do I get to the, how do I prove the TCO upfront before I start to make the investment? You've got to, you know, believe that, you know, AI is this new frontier and the companies that are able to harness it and put it to work for them are gonna have a massive competitive advantage. And so the sooner you get going, the better. And you can start in this dev environment with, I'll call it narrow use cases that are aligned to your core business goals and then, you know, scale as you start to see success. But to your point, so you want to make sure these agents, you know, have, you know, there's a clear set of governance. There's, you know, clear ability to trace like the data systems that they access and that you can continuously evaluate them against known business outcomes that you're trying to achieve. And then that accuracy against certain use cases is what allows you to promote it then in production. Chris, how can I ask you how things like, well, inference, obviously we did an episode recently about it was energy focus of talking about the coming wave of inference and, you know, the shift of the load moving to some extent from training to inference, you know, maybe in this calendar year or whatever kind of the next wave is. But talking about things like high performance inference and also hybrid cloud agility. How does the AI factory sort of figure in and support these two things in particular? Simply put, inference is your production environment. So training, whether it's pre-training or post-training, those are things that are happened pre-production and inference is where you're bringing this intelligence to life. So scale, efficiency, security, you know, robustness, reliability, compliance with policy, compliance with SLAs or SLAs, these are like the table stakes. And an AI factory is a significant investment for an enterprise. The expectations are produces significant business outcomes. And so we're focused on optimizing that production of outcomes, which you could back up and say, those are business intelligence or you could back up a little bit more and say it's simply tokens, optimize that throughput of tokens in the context of cost and the context of power consumption because we're also power constrained. And so how do we do that? That's through this scaled out inferencing, which is part of the AI factory. It's really the core underlying platform that you're running all of your models and then above that, the agents and AI applications on top of. So to me, it's the critical substrate and the agility that comes with flexibility and choice of where and how you deploy your models or your workloads, that notion of pre-production environments and production environments and where production data versus non-production data is used. You get some choice and where you deploy. And that, to me, is really the hybrid cloud. You have optionality. There's cloud environments, there's enterprise environments, there's even edge environments where you may want to deploy your workloads and take an advantage of all of that with a consistent footprint. Like we're building with this AI factory, it gives you the best of all of your alternatives. And so I think we're bringing the efficiency, we're bringing the flexibility, we're ensuring that we have those confinements whether it's confidential computing or guard rails or any kind of sandbox technology that I think becomes really critical as we're building and delivering these new capabilities. And if you go back in time before the focus on AI, we developed through decades of experience, but Justin highlighted that pre-production, dev test prod kind of best practices, there's a whole set of learnings and rigor and discipline that we built in building and delivering applications into production that we're bringing as part of an AI factory for building and delivering AI applications into production. >> I'm speaking with Chris Wright of Red Hat and Nvidia's Justin Boytano and we're talking about the AI factory and how enterprises can build AI factories that they can go to production with confidence and can scale up to the future and really help transform companies into AI natives as we've been talking about. Want to get into a little bit about specifics, infrastructure and software and platform components and Justin, I'll start with you for customers who are thinking about an initial AI factory footprint and might want to start small but have that ability to scale as they scale. How should those customers think about sizing and selecting Nvidia infrastructure and software? >> Yeah, I think as the customer starts to train build the AI factory, they got a thing through the five layer cake that I mentioned previously. So where do I have data center power? What is the power density of the data center? Do I want to run air cooling or liquid cooling? That seems to be a decision point right now. A lot of enterprises still run air cooled data centers and so platforms like our RTX 6,000s give you very good price performance. That's kind of a general purpose GPU to do experimentation with. So if you don't know where to start, that kind of gives you a great platform for many different use cases. And then from there you start to ask yourself, well what's the orchestration management platform that I want to run my business on? And that's why we worked very closely with Red Hat team. Red Hat AI factory takes care of really the next few layers of the technology stack from software orchestration and management, model delivery, all the, I'll say commercial security, patching, lifecycle management of all of that open source software so that you can run it with confidence and kind of get the factory up and running. And then you get up into the application layers. And the application layers the way we try and make it easy for customers to start as we provide reference blueprints, which are examples of proven use cases that even we run on RAF Actories that in media for things like enterprise search that make it easy to then connect into your enterprise documents and do document ingestion and then start to provide benefits to your users and then from there you can start to expand into your own developed use cases and such but that thinking through that full stack is really the easiest way to get going and then I think taking some of these proven examples is like kind of the quick way to get it in early win with kind of your executive leadership team yeah yeah with them the benefits and then from there usually pivot into you know what's the most important business outcome for the company to be competitive they got ask yourself for Nvidia where chip company where software company and where supply chain company when you really boil it down and so we then go super deep into those use cases to make sure that we're enabling you know tens of thousands of chip designers software engineers or all the people dealing with all the components that allow us to have supply and availability of building this rack scale infrastructure and have a world class at those then we can be world class in market and I think that's generally how companies should think about it Chris on the red hat side what are the key platform components that you see as foundational for this first AI factory deployment thinking about things like OpenShift, Red Hat AI Enterprise AI factory with Nvidia you know when thinking about this first enterprise AI deployment what are the key platform elements to start with and also how should customers think about sequencing them yeah I think for us the stack starts with hardware hardware enablement and then the distributed nature of rack scale architecture how do you make get access to that whole distributed system and then you know going up from there we start getting to specifics of models and and agentic applications and AI enabled applications so the bottom of the stack very clearly that's that's the world of Linux right hardware enablement device drivers low level system software and you know near and dear to our hearts we spend a lot of time in that space and making sure that we work closely together with Nvidia to do that first phase you know right against the metal enablement the next layer above that is the distributed layer bringing that rack scale architecture to life includes a distributed system like Kubernetes Kubernetes is tried and true in the applications space and it's supporting well delivery of agents or models or other content as containers on this distributed system with access to all the accelerators down at the bottom of the stack and then that Red Hat AI enterprise layers on top and this is where we start to integrate directly with some of the key capabilities that Nvidia brings like optimize models like Nemotron or some of the NIMS and that's where we bring that that distributed infrancene stack that is the foundation for intelligence for the business so you know we sometimes call this the the metal to agent stack and starting on you know with that layer right above the hardware building up through infrancene and then supporting the models it is is what we're building together to enable those key reference architectures that that Justin highlighted or or the Val validated Blue Prince or the reproducible plays that you want to bring into the enterprise because I think it's important to have those early wins at Justin highlighted that it's important to have those early wins it's an interesting tension perfection is the enemy of good enough so if you have this like perfect view of your future world where you've normalized all your data and everything is well-defined you'll spend all of your time doing that and you'll never be able to get to showing some business value but if you over rotate to the easiest thing to do the flashiest thing I can show it might not have much business value so picking those right first key use cases and also having in parallel this long-term mindset of it's a pretty fundamental shift in how we operate you know living in that duality that's that's the future on building the right stack to support rapid movement consistent reproducible or replayable plays and you know building from infrastructure that IT operations teams already understand they know Linux they know Kubernetes they're you know they're learning a lot of new things in this context so we'll give them as much stability as we can along the way yeah no it makes sense going along those lines of the getting those first wins right which is a great strategy for for lots of workplace projects to take on but I think talking about such a big shift you know to the AI way of working if you will let's look at those first 90 days that and can you lay out some kind of practical first steps we've got the elements of the joint stack laid out the hardware the software that you know metal to agents as you call to Chris what are some practical things that folks listening to the podcast enterprise leaders can do and kind of structure their first 90 days to get some wins and really start building that AI factory that can grow yeah I'll assume the data center infrastructure is built out let's assume there's built-in infrastructure is built out so you know what we what we publish is what we call validated designs that sort of walk you through a lot of the design decision points of the software and you know you you have to think of like how do I want to you know how do I bring all of my software into this factory how do I make sure I do you know security scanning you know if I'm going to want to rescan everything and and operate it how do I have automation to stand it up and then you know quickly how do I get these first we call them blueprints but think of them as like Kubernetes services that you deploy on the clusters to then get users on the system and then ultimately what we do is we have we call them like user acceptance test teams that we will roll an application out to to have them use the application so they can start to you can start to survey them and understand how are they doing work now versus how do they do it before how much time are they saving versus how they did it before and really that time savings is the productivity gain that you're going after and you can really quickly get to you know from time savings across a user group to productivity you know gains and so if you can get a 2x productivity gain you know across a big population of users then you know you're you're onto something really big absolutely yeah Chris I think that the learning that you'll gather along the way I think is really important and so the notion of starting with a focused you know have a hypothesis and a focused outcome and also iterating as you go so it's about how quickly can you move forward I think that's really important our our experience internally is reinforcing that and we started with some really focused examples of data that we want to bring together within with an red hat the research we wanted to do across that data and having evals I can't understand the important of importance of evals I think that it's up and off often overlooked part of the part of the sack because they help you ensure the the quality of what you're trying to produce and so you know building iteratively towards improving your evals we see this in the public with frontier labs focused on benchmarks and evals but they're just as important within within the enterprise and that iterative process of refining any portion of the stack it could be your prompting it could be how you're managing the data sourcing it could be even the scoping of the problem that you're trying to solve I I think that's that's really important and that notion of picking something that's real so it's not so artificial that you can just show it it's flashy you you get you get high fives all around but it doesn't really change anything internally right I don't think that's particularly useful so focusing on on those things that are real but again not making it too big so right so right sizing and the iterative process of learning as you go is how you start building the thing that ultimately is quite big but yeah I think it's starting small and iterating which we do a lot in open source do a lot software development and having a little bit of diversity we have touch points across every different function in our organization there's different personas but there's also different use cases you know it's more software development oriented it's more finance oriented it's more sort of sales and pipeline oriented each of these brings a little different dimension that is again is helping you flesh out your end-to-end view of what's needed to go through whole-scale AI transformation and be operating with a full-tilt AI factory powering your business yeah and as you get these first projects going and not not to you know sort of skip all the hard work in between but as you mentioned thinking about you know getting something going with an eye towards building out to scale and transforming the whole org looking at it from the other perspective what kinds of guardrails would be not just could people put in place but what kinds of guardrails would you recommend would be appropriate kind of from the get-go to make sure that as things scale as things expand as you know wins or won and people get excited and want to use this stuff more and go faster where do the things you can lay down kind of from the beginning to make sure that you know technical and process and governments and you know the guardrails are in place for these kinds of things hey you know I think so what one thing that that we did upfront was we made sure obviously our security teams were deeply involved as we did this just to make sure I think you learn a lot about your organization as you start to put AI to work and you'll find AI is really good at doing discovery and business systems that it has access to. And you might realize, you've got user permissions over scoped in areas. And so having the security teams understand, are we allowing too broad of access to what we wanna keep confidential with in the organization? You'll discover as you start to connect agents into your business systems. But there's all kinds of techniques for guard railing, data access once you do find systems that have access to, you're gonna realize that you've got to change permissions of many different business systems. You know, ultimately what you're gonna wanna do is scope the agents potentially as users. You think of them as digital employees. So where a lot of people start is they scope them to the user that's using the app, their permissions. So they see access to the information that they've been granted as an employee in the organization. But as we go forward, these agents are gonna start to work more and more autonomously. And we're gonna have to treat them almost like contractors we bring in, you give them least privilege access into your business systems. And then they gotta come back to you and check in with you and ask you for access to more business systems. And you're gonna have to have a process in place where you can slowly get grant them more access to do the job and fully onboard them into the job that we're asking them to do. Chris, I'm gonna turn this one to you first, but Justin, you can be thinking in the background about your answer. We like to end these. The more time passes, the more I feel like it's an unfair thing to ask at the end of the AI podcast, what's the future gonna look like, right? For obvious reasons. But if we look ahead, Chris, a year, two, maybe three years down the line, if you're feeling really bold, what does the AI factory look like? As, you know, agente AI develops and models keep developing and the infrastructure keeps developing. But especially as, you know, more enterprises put these systems, build these factories and use them and put them to use solving real problems and driving new ways of working, what do you think the AI factory looks like a couple of years hence? - I think you take it from a few different points of view that the one angle would be the layer cake picture. And that one, we have a pretty good understanding of the layer cake. So while there might be some subtleties, certainly in terms of specific tools that will come and go over time, that layering of what we're describing from hardware up through AI enabled applications, I don't think it's something that will fundamentally change. So you look a few years ahead, we'll see something that looks quite similar. How it's used by the enterprise, I think is what's going to shift completely in that timeframe. Today, a more sophisticated enterprise has some agents in production, but they're not entirely agentic and it's not translated into the core of their operations essentially. And so that, to me, is the shift that we should anticipate, the autonomous nature of agents and the scoping of tasks will continue to grow. So initially, the simple chatbot, which is just essentially fetching information, then you got a little more sophisticated with stronger and stronger recommendations. You could call that some kind of an assistant, the doing phase of agents and total autonomy were seeing that time horizon just stretch. It feels like almost daily stretch out to be longer and longer so you can, the coding context, you can give coding agents very sophisticated tasks and they will spend hours and hours for just very sophisticated code as a result. That's just the coding example. It's language as well structured. It's a good template for how we should think about the breadth of the enterprise. And so in the end, the AI factory, the layers look similar, the sophistication of the tasks grows and it becomes the core of the business. It becomes the place where we do our operational practices around. And so in the end, it's not that we're going to go through and kind of augment each of today's processes because if you just think of it like that, you take a bunch of questionable in some places, even stupid processes and automate them and then you get an automated stupid process. It's really redefining how we work together completely and to end and where agents take on critical tasks and the business that I think is that future view, which again, you put some time frames on it. We're not talking decades. We're talking quarters away, which is self-confident. But yeah, I think that's to me that's that future outlook. - Well said, Justin, your thoughts? - Yeah, I think the way Chris framed it is right is software development, even I'm saying in the last six months has evolved where you can give AI, I'll say almost like a design document and let it go off and think and produce the code and then do, I'll call it functional verification of that code to make sure that it's accomplished its task, before it comes back to you. And so it's doing very long-running thinking and work that is the work of many, many, many, many software engineers. I'll just say, and I think in the software engineering world, like I said, we've seen this product market fit where we're seeing a two to three X productivity gain with software engineers that can use these long-running agents. And if you extrapolate that out, the productivity gains for the whole software industry is massive. But we're now seeing that move into this knowledge worker world and CAD designers across every industry where they can do the same thing, where they can start to give design document goals to long-running agents, where they can basically explain the exit criteria and give the agent the tools to do the functional verification and say, come back when you're done. And so I think that's what the future of work is going to look like in two to three years. You're going to have different agents working for you that you give these more structured long-running tasks to. They go off and think and do the work and then they come back to check in in a period of time. And that will make us all infinitely more productive than we are today in searching through UIs trained to find information around. And so I think we're going to live through a big change in how we work in the next couple of years. But every company across every industry and every job function will really be transformed with the use of an AI factor. Perfect place to leave it. Chris, for listeners who would like to learn more about your work, the work that Red Hat is doing, places online, they can go obviously a website, social media, technical blog, other places, where would you direct a listener to learn more about what Red Hat is doing with AI factories? The easiest one would be learn more about the Red Hat AI factories in video. So it's sort of an easy thing to search. Chris, you'll find information from Red Hat.com. You'll find more information together with Nvidia on the video website. And that's a really easy place to start digging into the Red Hat view on all this content. Fantastic. Chris, right? Red Hat, Justin Boytino of Nvidia. Again, thank you both so much for taking the time to come on the pod and talk about AI factories. And really the future of work is we landed on Justin. It's an exciting time to be alive. Thank you guys. Thanks, Dawn. Thank you. [MUSIC PLAYING] [MUSIC PLAYING] [MUSIC PLAYING]

Podcast Summary

Key Points:

  1. An AI factory is a strategic infrastructure comprising five layers
  2. Enterprises need AI factories to become AI-native, driving innovation, efficiency, and revenue growth while ensuring data privacy, security, governance, and integration with existing business systems.
  3. Collaboration between NVIDIA and Red Hat provides a trusted, scalable solution combining NVIDIA's hardware and software with Red Hat's enterprise-grade orchestration and security to enable production-ready AI deployment.
  4. The shift toward agentic AI systems (like autonomous coding assistants and enterprise search agents) offers significant productivity gains but requires careful management to avoid shadow IT and ensure reliability.
  5. Successful implementation involves starting with focused use cases in development environments, then scaling to production with best practices for security, cost-efficiency (e.g., hybrid models), and performance optimization for inference workloads.

Summary:

The discussion centers on the concept of an "AI factory" as essential infrastructure for enterprises to harness AI effectively. An AI factory is described as a five-layer stack encompassing data center resources, specialized hardware (like NVIDIA's GPUs), software orchestration, AI models, and end applications or agents. Its primary purpose is to convert enterprise data into valuable intelligence, enabling companies to innovate, improve efficiency, and grow revenue while transitioning to an AI-native operational model.

A key emphasis is on building these systems with trust, incorporating robust security, governance, and compliance to protect sensitive data and integrate with legacy business tools. " The partnership between NVIDIA and Red Hat is presented as a solution, combining NVIDIA's performance-optimized hardware and software with Red Hat's enterprise-grade platform for orchestration, security, and lifecycle management. Recommendations for enterprises include starting with specific, high-value use cases, using reference blueprints, and adopting a hybrid approach that balances powerful frontier models with cost-effective open models for scalable, production-ready AI inference.

FAQs

An AI factory is a comprehensive system that processes data to produce actionable intelligence, driving business efficiency and innovation. Enterprises should build one to harness digital intelligence as a critical resource, similar to energy, enabling them to stay competitive in the next industrial revolution.

An AI factory consists of five layers: data center infrastructure with power and chips, software orchestration, models for intelligence, and applications or agents on top. These layers work together to transform data into business outcomes like faster product development and revenue growth.

NVIDIA provides world-class hardware and infrastructure, while Red Hat delivers software orchestration and management layers. Together, they integrate security, governance, and enterprise-grade features to ensure AI factories are safe, reliable, and compliant for business use.

Agentic systems are autonomous AI agents that can plan, reason, and act across business systems to perform complex tasks like deep research or software development. They are important because they enhance productivity, reduce costs, and enable knowledge workers to operate more efficiently with enterprise data.

An AI factory must include robust security, governance, role-based access control, audit trails, and reliability. It should separate development and production environments to ensure safe testing, verification, and scalable deployment of AI applications.

Inference is the production environment of an AI factory, where trained models generate intelligence in real-time. It requires optimization for scale, efficiency, power consumption, and compliance to deliver business outcomes reliably and cost-effectively.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.