OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti
43m 56s
In this podcast, Sachin Kalti, head of industrial compute at OpenAI, discusses the massive scale of AI infrastructure and the challenges of building data centers to meet insatiable compute demand. He explains that OpenAI views compute as the foundation of intelligence and is taking an active role in building its own infrastructure, moving beyond reliance on partners. Data centers are now giant, liquid-cooled supercomputers that turn electrons into tokens, requiring cooling at every level due to high chip temperatures. Power is a critical constraint; OpenAI adds new generation and transmission to grids, ensuring it does not consume existing capacity, and is exploring nuclear energy for denser, sustainable power. The company’s custom chip project, Halopino, is designed to maximize token output per watt by co-optimizing hardware with specific model workloads. Inference has overtaken training as the primary compute use, and AI now drives its own research, further boosting demand. Kalti emphasizes that the risk is not overbuilding but underbuilding, as any new compute is instantly consumed. Data centers benefit rural communities through tax revenue, jobs, and grid upgrades, and their closed-loop cooling systems recycle water, dispelling myths about high water consumption.
Anytime you're taught you have enough compute, you can slow down. Always negatively surprises like all. We should not have slowed down. Demand for our outstrips compute supply today. So anything we can bring online, we can see in the video. Our obvious worry is that still, at the scale at which we are trying to get compute and build compute. The physical world is on who that fasts. We do believe that the world of recursion is not that far, where AI will design the systems it needs to train and run the next generation of AI. I am Matt Turk, welcome to the Matt podcast. My guest today is Sachin Kalti, who holds what might be the most fascinating and relevant title in tech right now, head of industrial compute at OpenAI. Sachin has an incredible background. He was a professor at Stanford, a multi-time founder, and most recently the CTO at Intel. Now he's leading what many are calling the largest infrastructure build out in human history. In this episode, we step away from the model layer and dive deep into compute and the physical reality of the AI boom. We talk about the staggering scale of the data centers being built. We get into the weeds on liquid-cooled supercomputers, power grid constraints, the potential of nuclear energy, and OpenAI's move into custom silicon with Halopino. We also discuss the broader target strategy and the slightly suffer reality that AI is now beginning to help design its own chips. It is a fantastic look behind the curtain at what it actually takes to power the future of intelligence. Please enter my conversation with Sachin Kalti. I'm sure it's one, if you agree, and two, what it feels like from the inside. How do you view what you're currently building at OpenAI? It definitely feels like one of the largest things a humanity has ever built effectively. Definitely bigger than many of the things that I've put off. I'm not only left to have experience that I've built out, but it feels exactly like what it's on. It's on the belly of the beast. I'm also to speak. Every day is we are making decisions. We are on compute that historically from our previous role, for example, I didn't tell. We're probably taking months to make, given the magnitude of the in those decisions. But the demand is so insatiable that it is growing so rapidly that we have to move very quickly. So it's an intense time, but it's probably the most exciting thing an engineer would want to be part of. Yeah, and I read somewhere that OpenAI was planning on spending about 50 billion computers. Is that still the rough number directly? Actually, that sounds like a cow-garde gap. The whole industry itself was going to be 7 billion in convinced this year as well, so in same numbers it seems. It's probably continuing to grow and a lot of build happening. A lot of that is also going to translate to compute usage from people like us in a year or two. Is the right way to think about this that for OpenAI it's a bit of a new world. Obviously not quite a pivot because obviously the OAI research is going full speed ahead, but building a whole new business within the company is that fair? Is that a happy people think about it? Because OpenAI is building models, one thing building data centers, it's a whole different world. Yeah, I think OpenAI has always had a fundamental belief that computers are the foundation of everything. Computers are foundation for intelligence and the way we keep continuing to scale intelligence and distribute intelligence is by having compute. That has never been different, that has never always been the belief. I think what's becoming clear is to build the kind of compute we need and act this scale. We have to not just rely on getting compute from a partners, we increasingly have to take a much more active role in building and getting the compute that we need. It does absolutely feel like a new muscle that we're building in the company. Maybe to encroach the conversation from the beginning, it would actually be very helpful to talk about what a data seller is in reality. So I think everybody knows that data centers have been built, but you know, going to one's head, I'm not sure that everybody could say, well, what is actually being built? Because we've been building data centers of the industry for cloud, for decades at this point. What is fully different and new about the data centers that we're building for AI today? I think the biggest probably is this scale. We are essentially building large supercomputers. As we think about AI and as we build intelligence and telegram intelligence and models become more capable, we use it for more complex tasks. We need more and more bigger computers, effectively. And so I think the way in the misvideo visualized data centers is giant factories, that are turning electrons into tokens, popular phrase nowadays. It actually has a lot of a ring of truth to it. So how do we take power, how do we take those electrons and actually use it to power chips that effectively are delivering intelligence. And the way I visualized that is large football fields liquid coal because these chips run really hot. The temperatures on these chips are very, very high. And so you cool them with liquids, you can't cool them with air. So a lot of liquid cold basically refrigerators effectively that are sitting in the alongside the building. And you know, that's a big one. We added the calling happens at the day last in 11 or what does it happen at the chip level? Oh, goes. Both. But, right. So you need to cool the data holes, but you also need to cool the chips individually. Because it's not going to be enough to do one or the other. And you also have to cool the things that connect chips, right. And so that's why you need cooling pretty much everywhere. Nowadays, even the cables that are the transformers that distribute the power become too hot. So they also need to be cooled. So everything that processes energy produces heat. And it isn't calling technology that is being used or something that's well understood and it's just getting deployed or is there like fundamental new things happening and cooling right now. I think liquid cooling have been around for some time, but have never been deployed at this scale. And so the innovation is more around how to make it reliable, how to make it cheaper, more scalable. So there's a lot of innovation around that. There's also a lot of new innovation in new kinds of liquids, new kinds of materials that can absorb heat better. Because anything that can improve the efficiency of heat transfer is very important for data centers. So you can then run the chips hotter. Right. And there is a direct correlation between running a chip hotter and how powerful the compute is. So the hotter the chip, the more memory bandwidth you get, the more flux you get. And so there's a strong way of if you can cool well, that also means you can produce more intelligence. All right. So a giant ending factory is lots of cooling. The other part that seems to be very critical to any discussion is power and energy. So how does that work starting at a high level? Do you connect to the grid? Do you have your own power generation? I think the early days we all connected to the grid and we still all would want to connect to the grid. At this point, we are beginning to hit, we are investing in generation infrastructure for the grid transmission infrastructure for the grid. So whenever we build a data center anywhere, we make it a hard commitment that we are not taking power away from the grid. In fact, we are investing in the grid to generate new power so that we can consume it for data centers. What does that mean practically? You have a grid somewhere. It has a certain generation and distribution capability, certain number of megawatts. Obviously, a data center shows out if there was spare capacity, then of course the data center can use it. But if there is spare capacity, then we have to add new gas or solar or high-duration infrastructure. To the grid. So we are investing and funding that build up and then you have to build transmission lines, invest in transformers, substations to distribute that power. So wherever we are building data centers, we are funding the development of all of that infrastructure. And so that's one of the things that we do want to emphasize, which is this is infrastructure that would otherwise or not have been funded if not for these data centers. And one of the side benefits of this big data center build out is the grid infrastructure of America and the whole world for that matter.
is getting upgraded very quickly. And so that's the power piece. So whenever we can do that, we do that. And we consume power from the grid, but it's also good citizens of the grid because we are improving the infrastructure for everyone. Not just for the listeners, but also for households. In some places, we are beginning to hit the limits of how much grid power we can build and consume. And so there are everyone's looking at behind the meter. We are also doing some diameter generation, where we will have onsite power generation and distribution capability that does not come from the grid, but in fact the data center becomes effectively self-sufficient in terms of power. - The gas turbine, what is it? - Today it's gas turbine, especially in the US, because that's the most dense, transportable form of energy. And also the one that is quite widely available in the US. But there you're blocking that by the sub-content. - Do you think the nuclear conversation is interesting? Maybe recording this in France, which has a bunch of nuclear power generation? - New places, so I've come back to the discussion in the US as well. It's something that you think about, anything is interesting. - Absolutely, it can't come soon enough. I think the densest form of energy we can all produce and consume. And it also can. So I think definitely would be a good source of massive sustainable energy or data centers. He obviously outside of France, the rest of the world has a lot of catching up to do and building this infrastructure. But I think it's really a very important role. And data center. - Okay, so that's the great introduction on data centers. The other interesting bits of news that you guys recently had is Halapino. So now for OpenAI in addition to being the application business, I'm sure my enterprise and being in the model, AI research business and then the computing data center business, I think that OpenAI is the chip business. And that's very sort of completely full stack. But I'm sure we'll go into some details of Halapino later. But I'm interested in overall strategy. Where does that fit? - As we begin to, as it is a pretty big faction of the world's population, AI usage is exploding. Infants is obviously becoming a big faction of our work. It's consuming a lot of compute. And one of the other realizations is because we know what is the work towards acting. What is the model we want to run? We can co-design the hardware to be super efficient in delivering those models. And so the strategic thesis we had, Halapino is, how do we take advantage of knowing what the end work is, what the model itself is and design chips that are very efficient in serving those models. So it really allows us to try the efficiency advantage, to drive motor tokens or what. So the key metric that Halapino is optimizing is maximizing the number of tokens you can produce for what. And because the world is constrained by power today. So the motor tokens you can produce for the same number on our order, it's better for everyone. So we look at it at a very critical ingredient and scaling how we deliver intelligence to the world. - Great, so we'll welcome it to help in a second, but you just mentioned in front. And it's such an interesting evolution as well. With that comment, you can see what's going on at OpenAI specifically is inference equally being go much bigger than training these days in terms of like usage of compute. As we shifted from being those very heavy pre-training runs as the major use case for compute to now just inference being the majority. - You know, infants are a big perhaps even the majority on compute. And I think we don't like to make a distinction between training and inference because a lot of training is now inference. So when we train a new model, we are generating synthetic data for itself. That's inference. When we train a new model, we are doing post-training. And that's inference. When you train a model, you're doing test-time compute. That's all inference. - So when we say training, a lot of the compute actually is inference even in that in phase of the work. So infants are the fundamentals building. - Yep, obviously I cannot resist asking the inevitable question around the potential risk of overbuilding given the lag between the demand and usage and how long it takes to build a data center. Then you mentioned somewhere that you were deliberately very paranoid about the problems ahead in the next three years. Very primed out of that. The surprises ahead, which sounds like a very healthy approach. So how do you think about that? Is there any way to mitigate that or is just like, aren't we play a convenient belief that this is the future and which we'll just go live. - We have deep conviction in scale, right? And the history has born as such. So effectively our journey, for example, has tried to compute. We tripled compute and we tripled revenue. And we believe that, I mean, the continues to be to demand far outstrips compute supply today. So anything we can bring online, we can see immediately. So there's no compute that is going with us for our list. So I think that conviction has not changed whatsoever. And if anything, we are saying that scaling laws on research and training continue to hold and potentially the pace at which we are doing research is activated, right, because of AI itself. So AI is doing a lot of AI research now. And so one of the subtle indications of that is is previously our researchers used to run experiments and they needed computer run experiments. But the number of experiments they could run was limited by the number of human researchers they have, which is a scarce resource on the, why there's not a lot of people who can do AI research. Now, AI itself can do AI research. The number of experiments they can run exposed. And therefore, the amount of computing for research also explodes. So we don't see a world where we will have immune to best compute for the foreseeable future. When I was referring to surprises, my worry is more on the downside of we are not able to actually build all the compute we want. - Yeah. - And the source raises the other way. - If we are the big for us. - Right. Because that is consistent in the case. Any tab we have thought we have enough compute it can slow down always negatively surprises like, oh shit, we should not have slowed down. Right. And so our biggest worry is that still. And at the scale at which we are planning to get compute and break compute, the physical world is not move that fast. And physical supply chains, factories don't move that fast. Don't cannot occupy state at first. So for us, the surprise is more on that direction than the other direction. - Look at us, very good. - You alluded to communities a minute ago and obviously that's a key debate. So curious about your perspective on a spectrum where on the one hand, one extreme you'd say, well, the AI industry and computing industry has a PR problem. And there's no problem. It's just the problems that we cannot explain it well enough to the other extreme. Actually those communities have a point. What do you think the reality is? - I think anytime there is new technology which is as, as the evolutionary, as this technology is, there is always disruption that's gonna happen. But in FWV, we have learnt this over history that this all the way leads to better outcomes for society, right? And so how do we draw our line from where we are today to that outcome, right? And explain to the world why this is the trajectory and we all need to be on. I mean, it's our responsibility to do that. On the communities front, there's a little bit of local versus a global issue. - Oh, well. - On the communities, I think data centers are even today are net positive to every community. Because we're building these data centers and rural areas of America, right? Where there's nothing else that is being built and this clear. So we show up in rural Texas, we build a data center that produces new property tax or sinks for the community. That funds schools, that funds hospitals, we show up and we invest in new grid infrastructure, which otherwise would never happen, because there's no demand. So there's a modernized grid that that a can enjoy. We see we produce jobs, right? And so I think one of the things we are investing a lot that is explaining local benefits every time build a data center somewhere and making sure that
it is well understood the kind of upside that decides. And data centers once they are built are essentially very clean citizens. They don't produce any gases or toxic chemicals or anything. Yeah, self-contained, they just produce intelligence. What's your personal question that comes up is water. And I think there's been debolt quite by research that maybe give us a color on my old figure by the water. She is a liquid cold and the liquid is recycled. So we do actually the water content of a data center is shockingly small relative to household water consumption. So I think as you put it, it's been debugged. It's a misperception that data centers consume about water. Is anything they consume so little water for what they do? And all of that water is recycled. So we don't net consume any water. Once we get to a particular point, the water just gets recycled as they is bit cold. Yeah, so all this story is like brown water is just don't make sense because the water at data center happens in like a contained circuit. It's a closed loop. Yes, it's a closed loop. It's a closed loop. By the way, you mentioned Texas and rural areas since we talked about data centers of the big new of this conversation. Why do open AI and other companies pick rural areas? Like how do you select a site for data center? So many factors. So one is of course land, like plentiful land. Number two would be permitting like, can we build these things? And we want to build these things such that they are not affecting a neighborhood. So line that is somewhat remote is there an idea can be take. Of course access to power, right? So strong grid, strong gas availability, all of those are important factors. And then four is neighborhood, right? So how quickly can you build these things? So availability of labor construction, they were qualified electricians, plumbers, all this can be able. So all of those factors go into every single site selection decision. And I know we obviously Texas and pop the wakeers. It fits a lot of these criteria, but it's not the only state. I mean, we have data centers all around the darker in LA. Okay, great. We're going to go into all of this in more detail. But let's talk about you a little bit and you journey. So you're the head of industrial compute at OpenAI, which by the way, to the beginning of this convention, to the title, industrial compute is so I mean, such a perfect title for the moment we're in. But what does that mean? What is the role and how is this all effort organized within opening it to the extent you can talk about it? Yeah, I think think of it as my role and our team's role rather as how do we bring compute online at industry, okay? That's effectively what they tell me. And that's the entire life cycle. So how do we find the ingredients that go into compute? Line, power, shells, chips? How do we finance that? Right. So because these are master dollars. And so how do we make sure that we finance the grid infrastructure? How do you finance the construction on the councils? How do we finance the chips? Then it's about how do you operationalize all this? So how do you actually make sure these things happen on time? They stay up, how do you operationalize all of this infrastructure? So it's that entire life cycle. And then of course, how do you actually use the compute? So a big part of my role is capacity allocation inside OpenAI. So it is always the case results. We show that it's very popular guy. I am not very popular. There is always someone who is unhappy with whatever decisions he makes. But yeah, me actually our team provides the input to make the capacity allocation decisions. So we surface what are the different choice points and what are the what is questions on different allocation choices? That we have some capacity planning and then of course using that to forecast how much capacity we need there. Because it's not just more compute. It's also where what kind, what shape, what chip, what workload you want to run there. So all of those are the steam figures out kind of what should be the forecasting and planning. And that informs that chosen. And so that informs how the grid will go find the next chunk of land and power and chips to put in. And again without going into the confidential although I guess when you guys go public all of this soon all this will be public but like is that thousands of people at this age of like multiple different teams or do you guys like add source a bunch of things and work so as a bunch of contractors? It's a portfolio approach. We are never going to be in a world where we also say everything or build anything up so it's always the mid mix because that's the reason that's the reasonable thing to do. You don't want to put your rights in all in one basket. So we will have hyperscalers, produce probably providing a big chunk of a compute. A majority of our compute. We will have new clouds parts of our portfolio. We will be partnering with design build firms that can build the compute that we need. And of course, namely build some of our results. And so we are always going to have a portfolio approach because at the scale which we need we will need to adapt into all sources of compute. We can just rely on one pretty little mechanism. And your background default of this is your both a professor at Stanford and an entrepreneur or a founder or mostly an academic like just workers through your journey. Bit of all of the above. But yes, I'm at my hard time in academics being a professor at Stanford since 2010. Recently. What are your focus on there? I was a faculty in Condit Science and I've been interested in sharing with a particular interest. My idea of research was neck hooking. I don't know if it's already built in networks, but mobile wireless neck hoots started in the next five days and didn't have both such. But three, four years ago, I did a while out of the Stanford. I did a couple of startups. The last startup got acquired by DM Merri. And that's how got to know that part of the game's Intel CEO. That's how I ended up at Intel. Most recently before coming to opening I was Intel's city. And so I kind of seen all the different things. Acidemia startups, corporate in Intel. And then of course, a mix of all of the above in OPR. Because they have a research lab, a startup and a fast growing company all mixed into one bad open. Yes, isn't what you were you said. What did you say yes to the job when the job came up? It's actually what I just said. That makes this so unique. And it's hard to find anywhere. Because you always have to choose that having a world-class research environment coupled with the hardest technical problems, like the building the largest compute in the world. And so there are a lot of new problems that we need to solve. But also a fast-flowing business. So this biggest I'd like problems that sit at the intersection of business technology and strategies. And so there's a very unique time in history and a unique role, which is very attractive. I'll be sexy. Very cool. All right. So going into a bit more specifics about OpenAI's compute strategy. So maybe let's summarize what you guys currently have. So I think there's some Microsoft. You did this being 20 billion dollar deal with Syribras. There's a bunch of things going on. This is target. Maybe just give us the lay of the land of what you currently have. And then we'll talk about what you're going to name X. We have compute from, I said to many sources. So Microsoft obviously can be a partner in a partner. We also have compute as they have announced from AWS, I and Google. So we have compute from all of that. We also have compute from Colby, for example, so a NeoCloud. And then of course compute that chip partners are supplying now actually in its up directly for an easy compute for us. So I think that's the mixed roughly today. Oh, as we go forward, obviously they'll be building on all of these relationships. But also looking at more options where we design the compute, the data center itself, also potentially even build the data center ourselves. So all of those are ways of scaling the amount of compute that we have. So I think coming back to my earlier answer. It's the answer always probably me try, but it's all up the above. That's helpful. And obviously our preservation makes all of the sense in the world given the scarcity. And then it seems as target as evolved. So from what was going to be a joint venture with Oracle and soft bank,
to what now seems like it's more like an umbrella term for the big strategy these days is that the fair-watching is very bad. Yeah, I mean, we look at StarVet as our compute strategy. And it is varying degrees of us designing or building the compute ourselves. For example, with Vettarical Close Partnership, we help them design, we help them, well, how we to operate a compute, which is a very new kind of compete for us. We work with SoftPantEnergy, which is public. We basically have code design, the WomShell, with them and they're executing on that WomShell. And we will be kind of figuring out how to operate our chips in these data centers, our cells, using that new chips. So StarVet to us is that umbrella strategy for across all of these different things. And think of it as an evolution that we continue to see beyond, because it's never going to be to order to wake up and do only one kind of way of building the compute. Yep. I think StarVet to us is a continuously learning how to scale compute and we're adding on more and more capability. And as part of that, the RD data center has been built by the like, Avalin Tech. Yes. So maybe walk us through that. Well, what is currently being built for people to have some spiritual awareness? We obviously have a big partnership with Ores. That's the Abelian Data Center. That's where, for example, we are training our Nailless Modets. So very excited about that. That's a very big GB Blackbel cluster for our needs. That's up and running. It's up and rising. It's being used for training the last two models more at the end. So it's super-thread. You're seeing how quickly the models are becoming more capable. It's because of these kinds of compute. I think it is something that is currently being built that are good at the brand. Yes. So Oracle is building a number of data centers, all of which are published across Michigan and Texas and other places. So these are coming online in the next couple of years as they get the help and be for whatever chips and that kind training. We are the latest, but these are really meant to be way better clusters that our Mellus Renew Board are training, but also product inference can't compete. The way that deals us structure, the Oracle is the prime building those. Oracle is the cloud. You have the core ten and the core ten and core ten. How does the full list of your building, how does it finance things, strategy work? What was it, the 122 billion? Was it the number recently? The core is of the world, our famous for being very strategic users of debt. Is that part of the financing strategy as well? How do you all think about it? With all of these compute that we have today, we have amazing partners who actually are handling that for us. So Microsoft, Google, Amazon, Oracle, we are off the off take. We have the tenants as we put it, so we commit to consuming that compute to buying the compute where it's online. Across the board, everything, your building stuff as well. So the partners are building partners. And you always the across the board, the tenant, not the owner, the cut. So therefore, financing is our source to apartments. We talked about a jalapeno, let's go into a bit more detail there. In particular, it seems that you guys went incredibly quickly. And Tizaneggit, I think I wrote somewhere in nine months from design to tip out. So maybe walk us through that and what was the reason we went so quick. Yes, it was incredibly quick. Nine months is very, very fast, probably the fastest I've seen in my career. I think several reasons. One is it's a team, it's a strong team. They have many of the team have designed tip-y chips at Google at the past. So very well-expeded team. We have a great partner in Broadcom that have a very strong track record of Delirium and Gexpils A6. So I think that's strong partnership with Broadcom and making this happen. Three, I think perhaps an open AI unique point. In most chip companies or when you design chips, you don't know what you're designing it for. Because you're a vendor, the customer who eventually runs the workload is somewhat different. There's the unique advantage here of us knowing what the future models might look like. And therefore, being able to shock circuit a lot of the decisions you need to make, design decisions you need to make on the chip side so that that's super helpful. And finally, increasingly AI itself, helping design and optimize the chip. That is the only one that takes the longest time. Because you're basically limited by how many human, much human time is there to process all this data and run the experiments. And send me into a lot of those iterations much faster using AI. Yeah. So AI is the linear zone chips now. Yes. I think that world is not very far. I mean, AI like a system in chip design, but we do believe that the world of recursion and not that far, where AI will define the systems it needs to train and run the next generation of AI. And including chips. Including chips. Including chips. You also released a few weeks ago, MRC, which is a networking protocol, workers through that. What is it and why is it a big deal? It's a new networking protocol routing technology, if you will, to scale these really large cluster fabrics. So imagine you have 300,000 GPUs. They need to be connected together. And when you're doing large training runs, they're constantly communicating with each other because these models are so large that the processing of the models is happening over the entire 100,000 GPU cluster, for example. And you can imagine the number of links and switches and nickcards that need to be there to connect all of these chips to them. At this year, CDS are kind of, right? It happens all the time. You can't really read any word all the ways things could fit. So the strategy of MRC is how do you design algorithms and protocols that can gracefully mask all the experience and make sure that the training workload does not get impacted. The network is an abstract system that the training job does not every way. It's always going to be there. It's always going to find a part even if a link fails. So it is all about reliability. It is all about availability. So how do we design protocols that contain the complexity of such a big cluster and make sure that we don't get stopped because of CDS, which are very common in the system. So MRC is like a mock-up path, a spring protocol, where you can spray packets or multiple parts. So between any two chips, there are many, many routes to get there, between any two points in the city, there are many routes. So instead of picking just one route, we will send traffic across all of them and whichever one succeeds, it takes them. And so that even if any one of them fails, it's not a showstopper. So that's kind of the basic intuition behind it. It's obviously a lot more sophisticated than that. But doing this X-Kill on that speed is hard. And so that's why it's quite innovative. What are the bottlenecks that you experience these days? It seems like the. Then show the bottlenecks changing in the computer industry. We usually people talk a lot about memory these days, is that one of them or what else? I think they're bottlenecks, I think, to be honest, to the supply chain. So I don't think there's any one, right? We have bottlenecks and the data center building itself are on permitting and availability of gas turbines, transformers. Those industries have not had much capacity or the last decade or so ago. And they've suddenly experienced the monsoon. And it takes years before you have kept capacity to produce more turbines and transformers. So they're trying to get catch up. And to the jump thing that you mentioned earlier, is that like a shortage of electricians? And my technical people, or trained people that know how to build those things at once are your map? Well, absolutely. Absolutely. So then why we should all do. As AI replaces an only George, like a head of electricians. I think that is indefinitely a shortage of electricians, plumbers.
kinds of tricks, your name, so anything we can do to train more folks to be able to do those, there are way well-paying jobs that a lot of us, all of the hyperscaders, all of the labs, would actively hire you for these new hand-tenic conversations. So that is definitely about tonight. I think they become an increasing bottleneck because we are all trying to build more and more and we have unlimited number of things, these kitabinities. Let's talk for a minute about the business side of things. Another thing you want to visit is a guaranteed capacity for customers to lucky and compute. Which is interesting right now, it's almost feels like open-air, also becoming a utility company providing convenient to others. What's the story maybe in that in the strategy? Yeah, so guaranteed capacity is guaranteed tokens. Right, so we are effectively saying we will guarantee you a certain dollar's worth of tokens of intelligence. I mean it makes sense right, so in a world where computers are shortage, like what tokens are always going to be at a premium and there's a shortage of tokens that we can produce even the limited compute that we have. And as this becomes a fundamental input to the enterprise, right, so enterprises are going to need intelligence more and more of it to run. Right, and so this is a way of enterprises gaining their assurance that the tokens of intelligence that they need will be there for them. And so they don't have to take business risk, right. And so I think it's good business hygiene. If you have to credit those supply resources and every enterprise intelligence is kind of the most important supply item issue, it makes little sets to make sure to secure that supply. And so the fact that in mind we're trying to socialize, I know it's a new concept, like what does it mean to have guaranteed capacity for the intelligence thing? But I think that's what intelligence is becoming. It is becoming a supply unit, not every digital enterprise. So maybe to end on the fun one, and you can answer either with your opening hat or not. The other center's in space is that is an exciting, is that science fiction? Is that needed? Is that something that people like to talk about just because it's cool or where when you land? For the geek in me and the G&M, it's definitely exciting. It's what I'm just thinking, it's super cool to see a constellation of satellites producing compute. I actually think it is, it will become feasible, right. There's engineering problems we solve, but they can be solved with enough time and investment. Whether it is needed, I think there is room for our little compute. I don't think it could serve all the compute beats, but it's definitely going to be a component in the arsenal. I think what we are waiting for is when does the economics of the launching exactly like change? And when does the economics of the hardware change? Because we need to get to a point where it's cheap to launch the hardware, and it's something fails, it's cheap to throw it away. I think it's the kind go up and fix it, unlike on the ground. And so I think that inflection point will see it happen soon. I know that time it bit them slightly. Okay, so cheer has been the wonderful, thank you so much for spending time with us. Appreciate it. Thank you. It's been great to have this chat. Cut.
Podcast Summary
Key Points:
Demand for compute at OpenAI far outstrips supply; any new compute capacity is immediately utilized, making overbuilding unlikely.
Data centers for AI are massive, liquid-cooled supercomputers that require cooling at the chip, cable, and facility levels due to extreme heat.
OpenAI invests in grid infrastructure (power generation, transmission) to avoid taking power away from local communities, and explores onsite gas turbines and nuclear energy.
The shift from training to inference is significant; inference now dominates compute usage, with training also heavily relying on inference for synthetic data and post-training.
OpenAI’s custom chip project, Halopino, aims to maximize token output per watt by co-designing hardware with specific model workloads for efficiency.
AI is accelerating AI research, increasing compute demand for experiments beyond human researcher limits.
Data centers are net positive for rural communities, providing tax revenue, jobs, grid modernization, and using recycled water in closed-loop cooling systems.
Summary:
In this podcast, Sachin Kalti, head of industrial compute at OpenAI, discusses the massive scale of AI infrastructure and the challenges of building data centers to meet insatiable compute demand. He explains that OpenAI views compute as the foundation of intelligence and is taking an active role in building its own infrastructure, moving beyond reliance on partners. Data centers are now giant, liquid-cooled supercomputers that turn electrons into tokens, requiring cooling at every level due to high chip temperatures.
Power is a critical constraint; OpenAI adds new generation and transmission to grids, ensuring it does not consume existing capacity, and is exploring nuclear energy for denser, sustainable power. The company’s custom chip project, Halopino, is designed to maximize token output per watt by co-optimizing hardware with specific model workloads. Inference has overtaken training as the primary compute use, and AI now drives its own research, further boosting demand.
Kalti emphasizes that the risk is not overbuilding but underbuilding, as any new compute is instantly consumed. Data centers benefit rural communities through tax revenue, jobs, and grid upgrades, and their closed-loop cooling systems recycle water, dispelling myths about high water consumption.
FAQs
Demand for compute far outstrips supply today, so any compute brought online is immediately used. The main worry is not overbuilding but being unable to build enough compute fast enough due to physical world constraints.
OpenAI invests in grid infrastructure, adding new generation and transmission capacity so data centers do not take power away from the grid. In some cases, they also use onsite gas turbine generation for self-sufficiency.
Halopino is OpenAI's custom chip initiative designed to maximize tokens per watt by co-designing hardware with models. This efficiency is critical for scaling intelligence delivery given power constraints.
Key factors include plentiful land, favorable permitting, remote locations to avoid neighborhoods, access to strong grid and gas, and availability of skilled labor like electricians and plumbers.
Liquid cooling is used at both the data hall and chip level, including cables and transformers. Innovation focuses on reliability, cost, scalability, and new liquids and materials to improve heat transfer efficiency.
The role involves bringing compute online at industrial scale, covering the entire lifecycle from ingredients like land, power, and chips to financing and construction.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.