Go back

Andy Hock - How Cerebras Plans to Kill Nvidia

77m 22s

Andy Hock - How Cerebras Plans to Kill Nvidia

Cerebrus is a leading AI computing company that has developed a revolutionary wafer-scale chip with nearly one million cores, designed to deliver unprecedented speed in AI training and inference. Unlike traditional GPUs, its architecture integrates compute and memory on a single device, enabling faster, more efficient processing by eliminating bottlenecks in communication and memory access. The chip’s design is flexible and scalable, evolving from early training-focused systems to optimized solutions for high-speed inference—particularly in coding, voice, and reasoning applications. Cerebrus argues the current surge in AI investment is not a bubble but a reflection of genuine, growing demand and untapped value in AI technology. The company’s systems are especially resilient to memory supply constraints, as they use on-chip SRAM instead of HBM memory. A key market shift is toward inference, where speed directly impacts user engagement and product value—making fast, responsive AI essential. Cerebrus has secured major partnerships with OpenAI, powering models like Codex Spark, and with G42 in the UAE to build AI models in Arabic and other local domains. It is also working with the U.S. government on national AI initiatives, including the Genesis program, to advance scientific discovery and enable technology exports to allies. While data center construction remains a challenge due to physical and logistical constraints, Cerebrus is pioneering solutions through efficient design, energy optimization, and strategic infrastructure partnerships. The company believes that faster, more accessible AI computing will unlock transformative applications across industries, from health and science to global digital economies.

Transcription

12310 Words, 69004 Characters

English
I don't think we're in a bubble because I don't think we've really discovered the full value potential of AI. What if we built one big chip and then left it whole? This is the chip, this is the Wafer scale engine, nearly 1 million cores. They're going to buy 700 Dmw of compute by cerebrus to power their applications. The coding model that uses cerebrus now is Codex Spark. Inferences is where you ring the cash register. Medium speed tokens at really low cost. That's not a market I'm interested in percent. The market for fast tokens for the future of AI applications that are at a competitive cost. That's what I want to go hunting. I know Dario and the team really well, extraordinary team, extraordinary technology. I hope we get the opportunity to do more. Most computer chips are smaller than a posted stamp, smaller than a bottle cap. cerebrus builds one the size of a dinner plate. In this episode, I chat with Andy Hawke, who's currently the Chief Strategy Officer at cerebrus, who is just IPOed and secured a 750 megawatt compute contract from up in AI. Andy makes his case that the trillions of dollars pouring into AI infrastructure isn't a bubble. He explains why memory stocks are ripping, why DTU prices out to follow, and why their chip actually benefits from rising memory prices. He tells me why the stat that 95% of enterprise AI projects fail is measuring the wrong thing. He talks about the US selling chips to China and the possible job of the people who make that decision. Let's get into it. Andy, you are the Chief Strategy Officer at cerebrus. Can you tell us about what that means and what cerebrus is? Absolutely. It's a pleasure to be here, Vazel. Thank you. So cerebrus for those members of your audience that aren't familiar is an AI computer systems company. We observed nearly 10 years ago that AI had transformative potential, not just for things like cat videos and consumer applications, but also for fundamental science, industry, society. So we saw this big potential, and we can go into this more a little bit later. What we saw limiting that potential was the speed of computation. Took too long to build models. Once you had a capable model, took too long to deliver answers. So it's cerebrus. We developed a new chip and a new system that was specifically designed for artificial intelligence compute. Fast forward to today. Those systems are delivering AI inference and training more than 10X faster than like a C general purpose processors like GPUs. And enterprises are starting to come to us to work with us because AI has moved from being a curiosity to being existential, and once a capability like this is existential, everybody wants it to be fast. And that's where we live in the market. That's what we do. In my role, I'm our chief strategy officer. I started with the company nearly 10 years ago and actually started as our first head of product. So working directly with our engineering teams, our founders, our customers to help define what we should build in terms of the hardware where it sits in the data center, the software stack that meets users, where they are. And work with the team to build towards those market and customer requirements, a computing solution, hardware and software stack that would deliver on that promise of accelerated compute for AI. These days, under the strategy banner, my charter includes product strategies, so thinking about where we want to build into the future. But also many of our largest strategic customer engagements. So think US government, think large commercial and international like G42 and the UAE are recently announced partnerships with OpenAI and AWS, the kinds of big partnerships and big technology moves that are going to carry us from our first decade into our second NBI. >> So when Cerebra started, I'm assuming that was before LLM's, pre-LLM. >> Pre-LLM. >> So what were you building towards back then? What were you thinking was going to be your customer or who? >> I love that question. I mean, I think for many of your viewers, your listeners, this is going to take us back in time. So let's put ourselves in 2017, 2018. The canonical AI problem of that era was training convolutional neural nets, like ResNet for image classification and object detection problems. So image net, right? Find a cat. >> Hot dog, no hot dog. >> Hot dog, 100% hot dog, no hot dog. So that was the problem. So it was training, cognets for image problems. And at that time, that was the canonical problem of the day. And so we thought about that as an archetypal problem to inform our hardware and software development. But we also knew that we didn't want to be limited about that because by that rather. Because in some sense, we had a hunch, a hypothesis that AI was going to be transformative, as I mentioned before. But we didn't know exactly how. And so we'll get into the chip and what makes it special, I think a little bit later. But long story short, we didn't design the chip at the hardware systems specifically for training or for inference. And we didn't design it specifically for convolutional neural networks or transformer networks. We designed our chip at the hardware system to be a general purpose accelerator for all of AI compute. But I don't think we really articulated it this way at the time. But I think what that bought us as we then moved from 2017 to 2026 in present day, what that bought us is durability and flexibility to accommodate large language models, inference as well as training. And be able to accelerate all those workloads by building a general purpose AI accelerator built around the first principles of all AI compute, rather than building a truly sort of problem specific chip that would just be good at one thing. Yeah. So, when, so I guess was GPT-2 on your radar back when it came out, I'm assuming so. And if so, did you have to change anything about the chips you guys were building for? Hey, like, LLM's are going to be huge over the next few years. So yeah. Great question. I mean, this is the fun part about this history, I think, when we're building a fundamental technology that is a new technology that's sort of at the base of the stack, I think you have to acknowledge, particularly in the hardware world, that it's going to take some time. And GPT-2, yes, when it came out, it was definitely on our radar, but I think is you just articulated the question, I'm not sure if we really knew at that time that LLM's were going to be what they are today, right? We knew that it was a really compelling answer to, I think, the fundamental value question of AI. Because that was the moment in AI where AI went from being sort of an interesting curiosity to being demonstrably valuable, right, was that moment, that sort of chat GPT moment. But we still didn't know that it was going to, that that architecture was going to endure, and that it was going to be as large in the market as it is today. So I'd love to say we did, but I don't think our crystal ball is particularly better than anybody else's. I think we just built a technology intentionally that could be durable and flexible to changes like we've seen in the field over the past decade. Okay, so then to your question, did we have to change anything? The short answer is yes, but not in the hardware itself, so the chip, the heart of our system has remained the same throughout, throughout all the way, from, you know, training convenettes for image problems to inferencing language models for, you know, coding an agentic, a generative AI. The chip at the heart of the machine has stayed the same, generation over generation has been improvements, but the fundamental architecture has stayed the same. At the software level though, that's where we were able to pivot and respond to the market. So our original systems were, were delivered with a software stack that was optimized for training, and that had a lower level library of compute primitives oriented towards convenettes, right? That makes sense. That was the market of the day. That's what we were delivering and building for. At that time, we were also integrated with the ML framework TensorFlow to, so go back in time, shout out to Google and the TensorFlow team. Before PyTorch really came into the market, and we came to DeFacto ML framework for, for ML model programming. As the market shifted then from training and convenettes to, to language models, still training, language models then got big, right? We might remember the transition from, from, from Bert to larger language models, right? Language models started to get beyond, you know, one or 10 billion parameters in the hundred billion parameter territory. And Bert was like 200 million or something? If I recall correctly, yeah, something like that. So it's still modest, by, by today's standards, right, right, now, yeah, with a lot bigger. There's still a lot that can be done with those, those, I'll call them medium-sized models, and smaller. But yeah, once training became, went from images, images, sorry, convolutional models to transformers and language models, and then from, I'll call them small to medium language models to large language models. We had to update our software library, and then we actually had to change the way software used our system to train those much larger. models. So we basically did a significant revision on our compiler and software library to be able to train those very, very large models, not because they were language models particularly, but because they were large. And then when, you know, in that post-GPT moment, when the market became interested in inference, the same thing happened again. So we recognized that there was an opportunity to accelerate inference, but our software needed to use this engine a little bit differently than it had for large training problems. I see. And so part of the machine stays the same. The software around it and how we present the problem or how we compile the problem to that engine has changed. Yeah. Could you maybe now give us a one-on-one lesson on the chip? And just like, how does it work? How did, you know, how did GPUs work? How do they differ? That's what I think. Sure. That's a long question. Okay. You reign me in. And I'll start. So at the beginning recall that I mentioned, we weren't building this device specifically for training or inference or calmness or transformers, but that we took a step back and looked at the first principles of AI compute. So if you ask yourself them, what are those first principles of AI compute? Obviously, you need a lot of computing operations. You need a lot of flops. But not just any flops, AI compute tends to want lower precision, like 32, 16, maybe eight, four-bit. Also, it needs the ability to handle sparse operations very efficiently, because there's a lot of zero-value data in the input data and the model weights of AI problems. So you need a lot of compute, but you want it to be optimized for sparse linear algebra and relatively low precision. Cool. Lots of compute check. The second element that you need is fast communication between computing elements, because AI is fundamentally a problem of handling data in motion. So in the analogy of a multi-layer neural network, the computing element that's computing on one layer needs to be able to quickly share the output to the next layer in order for the whole thing to go fast. So computing elements need to talk to each other at high bandwidth, so you need high bandwidth communication. Check. Last thing is you also need high bandwidth memory. And I don't mean that in the capital HBM sense, but you need memory that can be accessed very, very quickly by those compute elements, because you want your compute to be able to quickly pull the stored weights of the model and then combine that with inbound activations, push it out to the next computing element in the chain, whether it's training, or inference, condonets, or transformers, all those things hold true. Okay. So then you say I want a computer that has a lot of sparse linear algebra flops, high communication bandwidth and high memory bandwidth, and you look out at the computing landscape of 2015, 2016, 2017, and you have x86 machines, CPUs, and GPUs, the predominant compute engines. And they have a lot of flops, so they can bring the compute to bear. Maybe they have more high precision or whatever, but like they can bring the flops, but their challenge, those architectures are challenged by the latter two attributes that we talked about, high communication bandwidth. They both are what are known as traditional von Neumann architectures, so they have a compute substrate, a logic die, and then some memory. So GPUs, for example, use capital HBM, HBM memory, DRAM, off the chip. So those are the memory is separated. And then if you want your processors to talk to each other, then you're wiring your processors back together in a cluster, right, over copper, through a network interconnect, over fiber, across the aisle, and a data center to like some other machine, through another router rate. So the communication memory bandwidth is really a challenge. So let me say, well, look, what if we just put as much compute as we could together on a chip, on a big chip, and then allowed the compute elements to talk to each other over silicon, so we wouldn't have to wire a bunch of small chips together. And we can, if we have a big enough chip, we can also put a bunch of memory right on the chip, and then we don't have to reach out to some external or separate source of memory for memory bandwidth intensive operations. We said, okay, so that would, how big of a chip can we build? And then you learn that the way every chip starts is a 300 millimeter circular disk of silicon, and use photo lithography to basically snapshot out or stamp out circuitry on that chip. And then traditionally what happens after you have that full disk printed out is you then slice it up into lots of small chips, and then everything we talked about before happens that is people spend their entire careers and industries figuring out how to stitch those chips back together and get them to all work on big tasks. So we said, well, what if we just didn't do that last part? That would give us all the ingredients that we just talked about, and it would be a fundamentally better design for the first principles of AI work and allow us to achieve significant acceleration on this class of work that we think is going to be so important for the future. So that chip, it's just a bunch of transistors, but where is the memory? Yeah. Okay, so let's talk about the chip. So for your viewers, this is the chip, this is the way for scale engine. I think you've seen it, but you can take a look. This particular one happens to be our third generation design, but the first and the second were also the same size. The fourth is also going to be the same size. This is the biggest square you can cut out of those primary 300 millimeter circular wafers of silicon. Also random, but why a square? Yeah, so okay, so why a square? It turns out A, it's a lot easier to physically package. That is once you build this thing, then you have to figure out how to deliver power to it, how to deliver data to it, and how to keep it cool. And it turns out that a square is a lot easier to package. The second reason actually is software reason, because it's a lot easier to tell a compiler how to compile a workload like a neural network to a square array, rather than an array that has all these sort of tangential parts to it, and that's non-U. I see, I see. Yeah, so we conceived of this design, our original CTO, Gary Lauderbach, our other co-founders Michael James, Sean Lee, JP Fricker, and Andrew Feldman conceived of this design. They looked at other solutions like a bunch of small chips on a common interposer. We looked at alternatives because we knew this was going to be really hard, but kept coming back to this because of its performance attributes on the work. And then we worked very, very closely with TSMC, who is our fabrication partner, Taiwan, to basically develop the fab technologies and the methodology needed to, instead of dicing this up into a bunch of small chips, keep it as one. And so this is then the wafer scale engine three, and in today's systems, it has nearly one million cores on one device of 900,000 in our production machines. They're in a 2D grid across the entire surface of the wafer, and then all those cores are directly connected to each other, and across the entire wafer over silicon. And each core has its own local bank of SRAM memory. So the memory is on the chip, too, back to your question. So we're coming back to those first principles, lots of sparsely in your algebra compute, high communication bandwidth, high memory bandwidth. This design checks off all those boxes, and gives us what you might think of as a cluster worth of AI compute on a single device, and literally three orders of magnitude, more communication bandwidth and memory bandwidth than is possible with traditional chip designs. And I'm sure we'll come back to it later, but it's that combination of resources, particularly the memory bandwidth that allows us to just scream on today's agentic inference workloads. Yeah, so it seems like this chip is better and basically like every dimension. So why do we still use GPUs? You know, it's a great question. I think there's two factors here. One is that different parts of the AI ecosystem and problem landscape need different engines under the hood. Like what people have done with GPUs and industry to date has been incredible. And there's still lots of problems that might not require, you know, blistering feed speed instant answers. I happen to think that the market is going into a direction where blistering speed and instant answers are going to be the default, right? And we can come back to that as well. But I think the future of AI compute is heterogeneous, right? There's going to be different kinds of compute engines that answer different problems, just like the compute industry today. But I do think the market needs a different and faster solution for AI training and inference. And that's what we've built here. I think that the second thing, honestly, that was a challenge and opportunity for us in the beginning was building a software stack to meet users where there were, right? This is a very different kind of device. We couldn't use existing compilers, existing software libraries. We had to write our team how to write the software to make this device easy to program and use. And as we fast forward today, it is, right? We meet users exactly where they are. For training problems, you can program the machine in standard PyTorch. And for inference problems, actually, the programming challenge gets even easier because for most inference problems, what application developers want is just an API to some inference model under the hood. And so for inference, we just surface an open AI standard API for inference on our machines. So what was once a sort of a challenge for us to build and something that made adoption trickier has all but dissolved. And today, it's easy for customers to do either training or inference on the machine with standard framework. Yeah. You mentioned earlier that maybe a few years ago, you were more focused on training. When did you have to make that transition to, oh, like, inference is the more important problem now? Is that the case, would you say, first? I think, yeah, yeah, yeah. More important is interesting. What I would say is inferences are focused and inference is the primary area of growth in the market right now. But I think actually one of the under told stories is the remaining importance of training. The training market is still really growing, but inference is taking off extremely rapidly because that's where AI delivers value, right? So as soon as AI is valuable for a problem, then inference just takes off. Just because the demand is insane. Yeah. I mean, the way I think about it in normal terms is training is a cost center. That's where you build your model. That's where you teach it how to do what you want it to do. This is a value center, right? Once the model does the thing you need it to do or the thing that's valuable for you or your end users, inference is where you ring the cash register, right? And so that actually goes back to your question. When did we make that shift? We made that shift to start building the optimized software stack and our cloud services for inference when inference started to become the predominant problem in the field. So that was probably about 18 to 24 months ago. Yeah. And would you say most of that usage right now is coding? Yeah. Most of the usage right now is coding. So I think the other areas where we get a lot of usage are voice models and reasoning models for various chat and assistant applications. And timely that you're here today, we actually just launched today a private preview of Google's Gemma 4 deep-mind model, which is a multi-modal model. No, consumers. Yeah. Yeah. And it's, so private preview today and we'll come to this later too, but like all the models that we bring up, it is now the fastest frontier multi-modal model on the market. We're serving Gemma 4 at 1,500 tokens per second. It's 10 times faster than any GPU implementation and 15 times faster than something like a cloud's high-coup model. So yes, I think where there's a massive value center for a lot of our customers today is fast coding agents where speed equals software engineer productivity. But I think that will remain very likely into the future. But I also think in the future, things like agent applications for other workloads, multi-modal workloads, AI for science and security applications, right? These are all sort of just over the horizon and I think they're going to be really exciting in areas for fast inference in the future. If you had to make a prediction on the next couple of years where the share of usage is going to shift, do you think that coding is still going to be the biggest share in, I don't know, three years, or do you think life sciences is going to be that bigger share, voice, whatever? So you're saying I do have to make it per day. No, look, I'm a recovering physicist. I think it's a reasonable hypothesis. So much of the large enterprise and consumer industry that we touch every day is digital. I think it is very likely that that will remain true. And if that remains true, then how will we continue to update that, make it better, develop new applications on top of that is going to be coding under the hood. And so I think that's a durable application. Whether it remains the predominant application in three or five years, I just can't wait to see what people will build. So I don't want to rule anything out. Yeah, yeah. You know? Okay, cool. I want to shift the conversation a little bit to everyone saying, oh, there's like an AI in for bubble going on right now. Do you agree with that? Disagree with that? You know, all the big hyperscalers at Google or whatever, like in, I guess combined, they want to spend trillions of dollars over the next few years on AI infrastructure. So how do you reconcile those two things that on the one hand, everyone saying that, oh, we're in a bubble. On the other hand, these huge companies who are supposed to be the most sophisticated buyers in the world are spending a lot of money on this stuff. So yeah, like, how do you think about that? Look, I think people are investing behind a massive, demonstrable value. That's what I really think. Is it big and how does those investments sort of come online quickly and grow and really rapidly? Yeah. Will there be a settling in the market? And I don't necessarily mean like a correction, but will there be a settling out or a reconciliation where some investments pay off and others have to be redirected? Yes. That's how markets work. But I actually, I don't think we're in a bubble because when I think about a bubble, not an economist, when I think about a bubble, what I really think about is a value that's inflated beyond the inherent value of the underlying technology product services assets. And I don't think we're there with AI. I don't think we've really discovered the full value potential of AI. And so I think these really sophisticated organizations that you mentioned, some of whom are our partners, I think they're investing into and ahead of a legitimate value hypothesis. And a lot of what I heard is we're basically you're not able to build data centers fast enough to service all the demand. So is that a problem? Yeah, it's a problem. It's a great and interesting problem. So I think coming back to cerebris and what we're doing here and my own view on the market, there is a fundamental mismatch between the timescale of development for hardware solutions and physical infrastructure and the software applications that people build on top of it. So what we're seeing now is this massive acceleration and uptake of AI software applications services that are fundamentally built on top of this physical infrastructure of metal racks, power lines, fiber cables, big boxes with wafer scale chips inside. It takes longer in real clock time to build these physical systems and poor concrete and raise up data centers. There's almost no way that you can build that physical infrastructure as quickly as it's being adopted when you're in a moment like this. And so yeah, I think it's a problem, but I think it's a really, really interesting one. At cerebris, I think we did get ahead of it a little bit, but recall that we started this project almost 10 years ago, seeing that we would need to invest for years in order to build the right instrument for the future of this work. We knew we had to be ahead of it because we thought it was going to be big and important. And it was different enough to be valuable. We also knew it was going to take time. And to be candid, I think we as an industry probably didn't have that same lead time opportunity with inference in the post chat GPT moment to get a head start, pouring concrete and building data centers and developing like thoughtful energy strategies for global infrastructure. So I think as an industry, I think we're playing a little bit of catch up right now. But I'm optimistic on what that means that our sort of community of innovators around us are going to create, right? What that challenge puts pressure on creative people and investors and we're starting to see people develop, for example, novel energy solutions, storage and generation. We're starting to see people come up with clever approaches to reuse physical facilities and finance data center expansions, right? So I think it's a problem, yes. But I think it's a challenge that is going to be really exciting for us as a community to solve and smiling because I think if we solve it, then we get together to see this potential of this technology have some really, I think, positive impacts on the world. Yeah. I guess one of those bottlenecks is data. center development. The other one is memory. I think a lot of people are, you know, seeing all these memory stocks just blow up. Why is that happening? Like, why is there a shortage of memory? Yeah, could you explain that? Sure. I mean, so back to the fundamental architecture of a traditional chip is that logic die and then the memory die. And what we're really talking about here is that type of memory, which is used by GPUs called HBM memory. That memory is, you know, subject to fabrication limits, manufacturing limits, similar to what we talked about before, right? There's physical facilities, machines that need to be built in order to turn up capacity. You can't just turn a knob and produce 10x more. And so that memory is a bottleneck right now for those types of chips. For us, it's not, right? We have all of our memory on the chip, as right here. It's not part of that supply chain. So we don't have the same challenges being subject to HBM supply chain limitations as other chips do. So why, so this is called SRAM? Yeah. Yeah. So why is there not a shortage of SRAM, but there's a shortage of DRAM? It's really about the fabs and manufacturing facilities, right? So for a traditional chip, it goes through, say, the same fab that we do at TSMC or Samsung. And it's one process, right? Building memory chips uses a different fab, a different manufacturing facility, different supply chain, because we can leverage the traditional chip process. We don't, we aren't subject to those. We don't need the other manufacturing facilities for our systems. So I guess that's a, that's really good for you guys, because if the memory, memory costs are going up 5, 10x, then all the GPUs are also going to have to go up in price. I'm assuming, but you will not have to. I think that's, yeah, I think that's, that's about right, right? When we, so when we build big clusters, we have other systems in the clusters that do use that memory. So it's a, it's a part of our solution, but we, we're not nearly as reliant upon it. So I would say that that's, it's, it's an opportunity for us to be able to deliver sooner for customers and to be able to deliver ultra high performance, leading performance in the market at an even more competitive price point. So the other bottleneck in data-centered development, everyone talks about is energy. Do you agree that that's a bottleneck or why, why is that a bottleneck, I guess? bottleneck is an interesting word. I think, yeah, it is a constraint on the rate of development. And I think it's also an important constraint to be mindful of when we think about environmental impact, impact on the communities. Is it a, is it a fundamental bottleneck that is, you know, is there enough energy on the planet? I don't think we're quite there yet. But I do think as a community, as an industry, we need to think about not just how to make AI bigger, faster, and more capable and more valuable. We also have to think about how to make AI more efficient. And that, by, by efficient, I mean more efficient at learning, right? Using maybe fewer weights to learn the same thing, or using fewer compute flops or watts to, to learn the same thing, or less data. So more efficient learners. But I also think we need to make AI that is more efficient in terms of execution and output. So we need more efficient chips for that, and we need more efficient ways to deliver energy at scale. Whether that's through updates to the grid, or improvements to energy storage in communities and at large data center facilities, or if it means changes in how we generate, whether that's things like small modular nuclear reactors, or dare I say at data centers in space, where you have all the energy of the Sun pouring down on solar panels. Right? So I think that in that world, right, we're very much a part of the solution, because, but by building this big chip, not only are we faster, but we can also save energy on things like communication of data and memory access, because we're not accessing memory from a different chip. We're not pushing data between two different chips in a cluster meters apart. We're accessing data from our own SRAM on the same chip and pushing it microns and nanoseconds from core to core. So we're a part of this solution, I think, but as a community, and I think this is what a lot of folks that aren't in this industry are probably thinking about, we also need to think about how to deliver energy more efficiently, make this whole process more energy efficient, and really think of it holistically as a problem. So I do think it's a bottleneck, but not in a way that it's sort of a fundamental limitation, but in the way where we need to figure out how to build more efficiently and deliver this capability, more efficiently, particularly per kilowatt-kilowatt hour. Yeah, so I want to shift the conversation a little bit to, like, you building your own data centers. Are you guys building your own data centers? Or, yeah, why is that? Why is the chip company making their own data centers? Yeah, so to be clear, I often joke that working on this at Cerebrus on the wafer scale engine in our systems, I never thought I'd learn so much about plumbing. Our systems are water-cooled, so when we deploy to a data center, we take cool water, we push it across the back of our waifers inside the machines, and that allows us to keep the chip cool. As a result, we have a bunch of pipes running through floors and ceilings, and we have to think about how to keep the water cool. Sometimes things leak, so there's questions about plumbing, and that turns into conversations about things like joint, you know, connecting pipes together in wrenches. It's a real physical infrastructure problem. I think that, in that sense, we're not building our own data centers to be clear, right? We're not actually pouring concrete from, like, trucks with Cerebrus logos, and, you know, setting up facilities that way. What we're doing is we're contracting with folks that either own facilities already or are developing facilities, and we specify to them a particular layout of the floor, power requirements, inbound data requirements, cooling requirements, and then we work with them to basically contract and unleash that facility. I think that's your question, but I wanted to clarify that for the audience, because I think when you think, like, building your own data centers, like, man, one of these guys getting into, right? I mean, like, at the end of it, do you own it, or does a different company own it? Are you kind of overseeing the development? It really depends on the customer, but typically we don't own it. We're leasing the facility, right? And renting the power. There are cases where we have our own ownership, or we're operating a facility on behalf of a customer who might own the facility. But often what we're doing is we're leasing the facility. In that context, right, we chose to effectively roll our own data centers and build our own cloud, fundamentally, because we had so much demand for what we were building that is serving fast inference that we needed to sort of own our own destiny in rolling out data centers. And we also needed a particular, you know, for each compute provider, this is true, but we wanted a particular layout of power, network, cooling, and racks. And that allowed us to go quickly in response to the market. Okay. I want to shift the conversation back to the chips, actually. But more, I guess, on the economic side. So how do you think the cost? Well, first of all, what are the costs involved in building a chip? And then how do you think those costs are going to evolve over the next five years? This is another good one. So I'll try to keep it at a high level, because each one of these questions are great. Each question could be sort of a one-on-one course. But I think the costs of developing a system, a chip in system, are fundamentally the design, the engineering development, the prototyping, and the manufacturing at scale. And particularly the last one, the manufacturing at scale also involves the inbound supply of all of your sub components, right? So just to give you a sense of the sort of the timelines involved in this, company started in 2016. I started in 2017. We at that time had designs for the first wafer scale engine. And if I recall correctly, we launched the first version of our cerebra system CS1 based on that first wafer scale engine in 2019. So you can imagine yourself, it took about two or three years to get from a design into a workable device. And then it took another few years, probably, between our first generation system CS1 and CS2, really to iron out most of the system and software challenges to make those units really reliable at data centers. standard and production scale. So it takes time and it has all those different cost centers. It's not a cheap project by any stretch of the imagination. Now back to your question, where do those costs go in the future? One of the areas I'm most excited about is AI-accelerated, AI-augmented chip design. There's some really interesting work building foundation AI models to enable faster chip design, or maybe even use AI to design chips that compute AI that can then develop new chips. So that area of chip design, I think, is one that is, in some sense, ripe for acceleration and potentially ripe for cost or barrier to entry reduction in the future of computing. I think once you have a chip design, even if it's faster and cheaper than it was before, you still have to figure out how to package that device, that is bring power, data, cooling, put it on a board that integrates into a system that can go into a server or a data center rack. So you still have to figure out the physical packaging and the prototyping. I think those things will get faster, but probably not at a sort of order of magnitude kind of X-factor. Maybe we get 2X-factor. Maybe there are tools or equipment costs or learnings from industry that drive a cost down by a factor of one and a half or 2X, right? Just hit pocket numbers. Yeah. Then the manufacturing at scale, the last piece. I think it's going to get really interesting. I think some of the fundamental supplies, like you were talking about, memory. There's a lot of conversations about rare earths and materials that go into these systems. The supply chain that goes into manufacturing at scale. I don't think any of that gets cheaper, right? Or I would be surprised if it does. I think the entire history of human species in industry would suggest that as interests, as demand grows, right? And supply sort of stays the same or gets smaller unless we can find new sources, right? That's going to get more expensive. Then the manufacturing that is the actual chip fabrication, system manufacturing. I think that's an area that's also right for more efficiency and cost reduction, right? We're seeing teams here in the Bay Area build up prototype manufacturing lines and build scaled production manufacturing lines far faster than they could five years ago even and use more automation and more data to drive not just manufacturing velocity, but drive costs down. So in summary, I think cost of design, opportunity to decrease, cost of packaging and prototyping, met and cost of manufacturing at scale might also be a wash if we see supply prices go up and manufacturing prices at scale perhaps go down. Are there certain rare earths that you see as like these are the most important in terms of or maybe these are the ones you're focused more on because there's some, you know, some country has 90% of the reserves and things like that. Yeah, I mean, just to be candid, not at this moment, no. But I think we're as an industry, we're keeping an eye on that. And we're also excited to find other sources for those materials. Yeah, so Google, Google has like a full supply chain of, you know, they basically own everything like they own the software, they own the, they're building their own TPUs, they're offering it as a service through GCP. And that makes it probably like very cost efficient for them to offer it to other companies if they want to. Do you think they will be the lowest cost producer of tokens, I guess going forward? It's an interesting question. I'm not sure I want to speculate on who will be lowest cost. But I do think you're fundamental there on, you know, they have a vertically integrated product that they can offer at scale. That is true. I came from Google, I know that team, outstanding team. That being said, I think rather than focus on cost, what we're focused on is performance. So delivering the fastest tokens per second at a competitive price and with greater power efficiency. And I think that back to our conversation about the future of AI computing, right? Well, I think TPUs are an extraordinary machine. I think in the context of the future of AI, there's a place for a machine like that and there's a place for an ultra-fast race car machine like ours. And to the premise of your question, one of the reasons why we vertically integrated is so that we can make the fastest AI computers on the planet, but do it at scale and be competitive to those costs. So I think there's going to be future systems that deliver, I'll call them medium speed tokens at a really, really competitive price. But for applications, for the user applications that we were talking about before, right? Voice, reasoning, agents that I think are really the future of AI computing. Their speed is paramount, right? Those applications just don't get built or aren't feasible if they don't answer quickly enough. And so yes, there's going to be a future of compute that delivers medium speed tokens at really low cost. That's not a market I'm interested in pursuing. But the market for fast tokens for the future of AI applications that are really competitive costs, that's where I want to go hunt in a minute. Yeah, so why is speed so important? Like we know this from the entire history of the internet and I'll give two examples. One, we just talked about the company. Let's talk about Google search, right? Would you wait for a Google search answer? One was the last time you saw the little sort of pinwheel of death on a Mac pop up from a Google search query. I think that, and before you say it, AI summaries, right? I think early AI summaries made us wait. And even just maybe six or eight months later or a year later, we don't want to wait for AI summaries anymore. So I think speed matters and actually Google search and the team there did a lot of research to show the effect of speed on user satisfaction and user retention. And what we know from that work and by proxy the entire history of the internet is that every millisecond of latency matters for these valuable or day-to-day important applications. Every millisecond that you delay, your users get upset, you use to walk away, they go to your competitor. So for any applications that is like that, speed matters and that's where we're seeing voice reasoning agents going. The second thing that we know about the history of the internet is when things get faster. It's not just that the same group of applications exist but go faster, we see entirely new industries. So the example that I have here that is fun is Netflix, right? When internet speeds went from broadband or from dial up to broadband, it's not just that Netflix delivered more DVDs by mail faster. When speeds got faster, Netflix went from delivering DVDs to being a digital production studio and streaming entertainment company. We created an entire new, effectively industry and hundreds of billions or trillions of dollars of revenue and lots of user delight from people like me that like to sit on the couch and stream Netflix. So I think speed matters full stop. And when we're working with customers today, like OpenAI, AWS, cognition, notion, in those areas of coding agents, voice agents, reasoning models, they know our partners know that if they can deliver output and answers to their users faster, their users will be happier, more creative, more productive, deliver more value and use the product more. And honestly, that's why I'm so excited about where we are in the market and to see what our customers are going to build on top of our systems. Yeah, so on the OpenAI deal, could you talk a little bit about that? So is it when I use codex? And I see that they're offering me 2x faster speeds for twice the, I forget the cost for it. But is that 2x faster speed option using cerebrus on the back end? Yeah, so we're building. I'm happy to talk a little bit about that. We have a great partnership with OpenAI. We got to know their team actually from the early days of cerebrus. We saw their ambitious vision to build state-of-the-art frontier AI models and AI. And they saw our crazy ambitious vision to build a computer system that would be the right system for the future of AI. And it's great and serendipitous that we kept track of each other through those almost 10 years. And then we're able to reconnect and partner to deliver fast AI for coding and agentic workflows now. So we have a deep code design collaboration with OpenAI engineering team. So the public terms of the deal. they're going to buy 750 megawatts of compute by cerebrus to power their applications. The coding model that uses cerebrus now is Codex Spark, and we're working with the OpenAI team to bring up additional models in the future from their GPT, frontier models to next-generation coding models, all of which should be the fastest examples of those models in industry. Yeah, and I'm assuming the thing stopping you from working on the bigger models right now is just the data center issue. You know, it's really a matter of product strategy and prioritization. Oh, really? Yeah. I mean, certainly more computers will be better. OpenAI has incredible reach into the market, incredible user base. If we could serve all of them today with the fastest tokens, I'd like to say that we would. So, yeah, you're not wrong, right? Like, if I had more computers, it would be better. But we're also working really closely with their product teams to say, look, here's the library of models and applications that are really important new users today, which are those where speed is the most valuable or most enabling. And then let's look at the compute resources that we have together, and let's rack and stack and then develop a staged approach to bring the right fast models to your users first based on what we have. It's a fun exercise, actually. It's really cool to work with their team. That's really cool. Are you planning on working with Anthropic in the future? Yeah. So, right now, I think you just mentioned you have a deal with Anthropic for, was it 750 megawatts of compute? Open AI. Oh, sorry. Open AI. Yeah, yeah. If we're so like a few years ago, I think talking about 750 megawatts of compute, like that would be insane to talk about. And now, you know, people talk about gigawatts. And where do you think that goes in five years? Yeah, I mean, I'm showing my age here, but maybe you and I'm sure some of your viewers and listeners remember back to the future. There was a joke about 1.21, they call it gigawatts, but 1.21 gigawatts to get the car to go back to the future, right, from a bolt of lighting. And you know, in that, you know, in that sci-fi movie of my youth, it sounded like a made-up number, you know, it sounded crazy and funny, but unbelievable in just the right kind of way. And now we're talking about those kinds of power envelopes in a very realistic way in the next, you know, 12, 18 months for some of these systems. I actually think it's sort of an awe-inspiring trajectory for the industry that, and a signal of fundamental demand and value that we're talking about building infrastructure at this scale. The only analogy that I can think of is through the building of the railways or laying of undersea fiber to enable the internet, right, this is a generational infrastructure project that's being taken on by a coalition really of individual members of industry and governments to build this, in some sense, the rail lines of the AI industry future. What do I think we're going to be talking about two or three years from now in terms of power? I'd venture, maybe, I'll get some blowback from this, but I'd venture we're still talking about gigawatts, maybe 10 gigawatts, but I think we start to run into some limitations of physics. Yeah. And without new generation or significant grid updates that do that does take some time. I think in a few years we'll probably still talk. So you're also leading U.S. federal programs. Could you talk about, like, what does that mean? Sure. So at Syribris, the U.S. government was one of our very first customers. We worked with the Department of Energy, Argonne National Labs, Lawrence Livermore National Labs, we're now doing a lot of work with Sandia National Laboratory. We also did some early work through micro-process or R&D with the Department of Defense and DARPA. So the U.S. government has been a long time customer for us, but also a technology development partner. And I think for folks that are building new companies today, I'm happy to share more about this in another form of people want to reach out directly, but I think that the U.S. government can be a great partner. Many organizations within the government actually have the charter to be early adopters of new technologies. In our case, the Department of Energy, who builds world-class supercomputers for scientific simulation, also has the charter from the U.S. government to be an early adopter of new computing technologies, so that the government and its research scientists can sort of get early peak at what industry is building and participate in it, and also inform its development so that it can contribute back to science and society and public interest. Maybe I'm dorking out a little bit too much over it, but I think that that's a really, really incredible mission and function that the government plays. And for people that are building new deep tech or other technology startups, I would say, look, don't be afraid of the government. Lean into the government, and there may be ways that they can help you get revenue or get access to technologies or partners that your industry partners may not be able to. So at Suribis, government has been an early customer and a great technology development partner. We're currently working with different parts of the U.S. government on future I/O or interconnect technologies, as well as future memory technologies coming back to what we talked about before. So we're continuing to do work with the government on those kinds of projects. We are also a part of the U.S. National AI Initiative, which is called Genesis, which in summary aims to build AI computing infrastructure to fundamentally accelerate how we do scientific discovery, everything from biology and life sciences to physics and chemistry and drug design to space in the environment. It's an incredible mission and actually I think it's something that the U.S. needs to do in order to keep pulling industry forward and help us as a nation and as an American industry lead in this technology that in some sense we had a big hand in inventing. In other words, if we want to continue leading and extend our leadership of AI technology industry in the world, it's going to take a big initiative from leadership in the country to do it. So we're a part of the Genesis program and we're also working with the Department of Commerce to figure out how to take what we're building here in the U.S. with fast AI wafer scale systems and cerebral super computers and figure out how to package that with other U.S. AI technologies and quickly and easily export those technologies to friend and ally nations. So that's where our work with the federal government also intersects with international sovereign initiatives. That is how can we enable our friends and allies with world-class AI technologies that are built in the U.S. At least for me, I'm a believer that we get closer to realizing the full and positive potential of AI if the right AI computers and AI technologies are in the hands of more, not fewer, organizations and countries. So one of the reasons that I'm at cerebris is to build that right AI instrument, build that right AI computer, and then put it into as many developers' hands as many organizations, hands, companies, and countries' hands as I possibly can because if everybody can build faster, then we're much more likely to discover that next great thing that fundamentally changes our health, happiness, ability to do work, productivity, then we would be if all these computers just lived in one place. So the work with the federal government helps us do that and helps us get access to foreign markets. Cool. What are some of those projects that you're working on to export to foreign markets if you're able to talk about some of that? Yeah, absolutely. So there's actually a new, relatively new program led at the U.S. federal level by the Department of Commerce called the AI Export Text Act. And so that project is basically the U.S. government providing a framework for industry members to come together, define a common stack or a menu, if you will, of AI technology solutions, and then the U.S. government will work with industry to streamline the export policy for those so that industry members can quickly package and sell and export those technologies to foreign and ally nations. That's the general overarching goal of that program and where we fit in as a technology provider and partner. And we're also helping the government define what that stack looks like and what the menu of solutions might look like. For us, though, that's actually strongly informed by the work that we started four or five years ago now with our strategic partner in the UAE called G42. There's lots of information online about our work with G42, but long story short, they came to us about five years ago and we wrote together an MOU, a memo of understanding, that articulated a vision, that they wanted to build world-class supercomputers and world-class AI models in Arabic language, in health, in science, in finance. And the MOU was that we would do this together. Fast forward to just a few weeks ago when we had our IPO, I was sitting down with the CEO of G42 and we were reminiscing a little bit that we wrote that MOU and all of it came true. We've built with G42 as a commercial national champion for the UAE. We have built for G42 world-class supercomputers. We've collaborated on state-of-the-art model development. We worked with UAE researchers to build the first and still best open Arabic GPT model so that state-of-the-art AI can speak the language of local people and reflect the cultural interests and social and commercial interests of that market. We've worked with their healthcare teams to build state-of-the-art AI healthcare models. That work, I think, was reflected this ambitious vision of the UAE to be a leader in AI technology and I'm humbled to say that I think we played a helpful hand role in bringing them on to the world stage and where they are today. And I think if we can do that for more of our international partners through, for example, the current US Department of Commerce program, then we're all winning. Yeah. Is it easier to build data centers in the UAE than the US? Is it faster? You know, that's a good question. I actually don't know the answer to that. Let me put it this way. I'm looking forward to seeing how fast they can build data centers. What's the most surprising thing about working with governments? Is this your first time working with governments? No, actually, for the past 20 years. I used to work at Google before that. I was building satellites and before that I was building image processing algorithms and software for environmental surveys. In all cases, actually, I was working with governments. But I think what surprises me most here in AI particularly is I think some governments are really moving fast. And by design, moving fast is not a thing that most governments are good at. In fact, I think as a citizen, you generally want your government to be deliberative and fair and thoughtful and take its time. But I think some cases velocity really matters. And I think emerging as a leader or developing state of the AI, velocity matters right now. And so I think the surprising thing is seeing some governments really, really lean in. I think UAE is a great example of that. I've been thrilled to see over the past 12, 24 months how the US government has put the proverbial pedal of the metal on AI. And I think other nations around the world are starting the same thing. Yeah. So on that note, should we be selling chips to China? That's a, as you know, that's a complicated question. We've had the opportunity to work with China in the past, cell chips to China in the past. And we decided not to. I think there's an incredible developer market, there's an incredible research market in China that I think as a society, we would love to see go faster. I think if we can answer some fundamental questions and come to agreement collectively on things like security, then we absolutely want to be able to work together. Yeah. I've heard some people say, hey, we should sell to China. We maybe shouldn't sell our best chips to China. What, like, what do you think about that? This is, so I also work pretty closely with the part of the Department of Commerce that administers export policy. It's called the Bureau of Industry Security BIS. Huge shout out to the people that work in that part of the Department of Commerce, they have what I think of is probably one of the hardest jobs on the planet because they, at the end of the day, they are charged with balancing US national security interests with the interests, the interests of US industry and our economy, right? So on one hand, you might say, open up the floodgates, sell everywhere, it's good for industry, it's going to be a boom for our economy. On the other end of the spectrum, you might say, we have fundamental questions about how these chips are going to be used. That could threaten national security, or we just don't know. So let's just stop for now. Both of those endpoints have challenging consequences. So finding the right middle ground for any given global geopolitical state of play on any given day is incredibly difficult. If you hold back too much, you over and incentivize the development of competitors. If you deliver too much, what security might you put at risk? It's almost an impossible problem to solve. So I think the nuance of getting that just right is incredibly difficult. I will not claim to have the answer of what is the right thing to do. I think, as I said before, if we can answer this, fundamental questions about security, if we can have the right diplomatic engagement, if we can agree on the markets, we would love to be able to help the folks that are working on fundamental science, AI, new commercial AI, new open source AI in those markets go faster. Yeah. So what are your thoughts on Europe? Everyone says, hey, Europe is so far behind. You know, their policy is, or there's way too much regulation, way too much taxation. Have you guys worked with Europe? Yeah. And where are you seeing that headed over the next couple of years? Yeah. In fact, I'm going next week to Germany to a conference called International Supercomputing, where we're going to be sitting down with European owners of state-of-the-art government public sector supercomputers, but also owners and developers of new commercial AI infrastructure and data centers. We have a long relationship, an extraordinary system and partnership with EPCC, which is the supercomputing center in Edinburgh, Scotland. So we actually have systems in Europe. We also have systems in Germany. And we are working to deploy data center for our commercial inference services in France. That's all just context. Back to your question. I'm trying to encourage European AI infrastructure developers to think differently and move faster. I do think that there is an awesome opportunity. I think back to our conversations about our conversation about government. I think European governments are trying to be really deliberative and make the right choice, just typically in the interest of the population. But I think right now industry is moving faster. And if they also want to lead as friends and allies, I want to encourage them to move faster. And so that's one of the soap boxes that I'm going to be on next week. As a community, let's think a little differently about AI infrastructure, not just GPUs, speed matters, where speed matters, their solutions like ours. But also, how can we help you move faster, whether it's through commercial data centers or public supercomputing project? Cool. I want to shift the conversation. This is going to be like the last couple of questions. What predictions do you have about where enterprises, or how enterprises are going to be using LLMs over the next couple of years? Like right now, maybe coding is a. I mean, I guess we talked about this a little bit. We talked about it a little bit before the podcast, too, that right now AI is showing this massive potential. And clearly, there's a value proposition to be harvested in coding. But how we as a community build those tools to put them in the hands of enterprises so they can use them at scale, that's a really tricky nut to crack. And so I think even in coding today, we see massive adoption, massive value, but we still see large enterprises figuring out how to adopt at that scale. How do we integrate these tools with our workflows? How do we integrate these capabilities with our current teams? Right. So I think step one is we will see enterprises working with industry to adopt things like applications like coding and fast coding agents like those on our systems at scale. So step one is like increase adoption, right? Or absorption of that capability. I think then once that gets in, what we're starting to see from some of our partners in enterprise is development of new, I'll call them domain or industry specific models, right? Maybe it's not a language model for coding or chat, but maybe it's an LLM or a sequence style model for drug development or for material design or an AI model for physical industry for say manufacturing to say control robots or analyze data offensive manufacturing. line. That realm of broadly speaking physical AI and domain-specific models for industry, I think is a hugely exciting and active area of development that I think is over the next couple of years we're going to see some big changes. So we have long-time partners like GlaxoSmithCline that for the past five plus years have actually been building AI models to accelerate and improve the development of new therapeutics. We have customers like the National Labs that are working on fundamental AI models for the physical and life sciences. I think that area beyond sort of beyond coding and chat bots, assistant style capabilities and consumer agent models, I think that area of physical AI is going to be big over the next couple of years. Another tangential question off that. Why do you think you've probably read those stories that, hey, like 95% of enterprise adoption projects for AI are just failing? Why do you think that is? In some sense, I think they measure the wrong thing at the wrong time, right? I think those studies, it's important that we measure sort of value return, but I think it's far too early to ask that question really at scale, or not really to ask the question at scale, we can ask any time. It's too early to come to that conclusion, right? The way that enterprises typically trial and then adopt and scale technologies takes 12, 18, 24 months or longer for some big enterprises, especially if that enterprises the government, right? You start a pilot. You run maybe a series of pilots. It's intended and known that maybe eight out of those 10 pilots will not show results, but then you have some subset of those pilots that show positive results. You expand the trial audience, you start to scale it up only after you scale it up, and then maybe after a year of it operating at scale, do you actually start to see returns, right? That's how the life cycle and time scale of enterprise adoption and new technology works. And like I said, I do think it's important to ask the question of, are we getting value out? But I think it's a little too early to come to that conclusion. Yeah, I mean, to me, I definitely see a lot of people getting value, especially startups, are definitely seeing that they're creating value. I think the problem with enterprises is maybe they're a little slower to adopt than like startups. So yeah, and exactly what you're saying, right? You have to run 10 experiments. Two of them are going to work out and then you just scale it. And you know, it's funny, right? I think I like where you're going with that and I was chuckling about the problem with enterprise. I happen to like working with big enterprises because they have incredible reach, right? But to your point, they also tend to be more deliberate and move a little slower. There's more people involved, right? And there's intentional controls in place. So good. It does take time. So good. I think you're right, though, to bring up startups, right? For a green field problem, for a brand new team, we're seeing people build from scratch with AI. And I think those are going to be some really, really compelling disruptors is those teams that are AI native from zero. Yeah, I've been talking a lot of engineers at some startups and it's kind of crazy to me that, you know, all of them are telling me, I mean, these are like really, really amazing engineers. And they're like, I haven't looked at a line of code in seven months, eight months because they just use agents for everything. Yeah. It's pretty crazy. No, I mean, look, we do internally at cerebrus as well over the past six months where we had probably a half a dozen different programs across our entire engineering teams from design in hardware to program management and product management to software engineering and AI. All of them now are effectively being AI native for all the code design. And I got to say that's really reinforcing to us here too, right? Because back to the OG hypothesis, if we can build the right instrument for this work, we think this work could be really important and transformative. And now we see this, that same workload literally transforming how we're doing what we're doing. And so it's awesome and encouraging to see that sort of substantiation even in our day-to-day work. Yeah. What are some projects that you are currently working on that you're really excited about that you want to share? I think, I mean, everything is a cheater answer, but I come in to this company every day and I am just at awe and humbled by the ambition and curiosity, hustle and resilience of everyone of our staff, right? Not just our engineers, our program managers, people that are working at the office. It's awesome and everybody is incentivized in some sense to be a part of building this and changing industry. And there's no task too small and there's no vision that's sort of too big. And so across the board, I'm excited from model research to software development and systems, but I think what came to mind immediately is it's also really cool to work at a hardware company because for those of us that spend most of our time in front of a screen or working with documents or PowerPoint or God forbid just spreadsheets, being able to take a break and like walk into the back and see people turning bolts on things and see the things that you're creating in a digital realm actually take shape in the physical world. That's, it's an incredible opportunity. So I think to your question, I'm really excited about our future hardware actually. At the end of the day, we are a hardware company. And so we are on CS3 now. You can imagine that there is a 4 and a 5 and a more beyond that. I think what we've seen from the market is that there's extraordinary value and fast inference, fast training in some sense. I think we've proven out the value and burned down the risk of the the wafer scale processor. I'm really excited to see what we build in our next generations to go faster. As I mentioned, we've got some fundamental investments going right now in IL and memory technologies and future systems. So in some sense, I think the wafer scale engine of generations one through three where we are now is been incredible, but it's just the beginning. And so for for the audience out there, I would say stay tuned because we've got some really, really exciting stuff in the hardware pipeline. And I think that could transform computing very much in the in the same way that the wafer scale engine did in the first place. Awesome. That's all I had. Thanks for coming. It was a lot of fun, man. Thank you so much. I hope we can continue a conversation. We should check in on those future predictions sometime and see what came true and see what was a puff of smoke. Yeah, awesome. Thank you. Is there any topic that you wanted to cover that you know, I didn't cover? I thought it was great. Okay. No, I mean, I'm sure I'll think of something later, but I just enjoyed the conversation. I hope it I hope it it sticks with your audience. Yeah. And I really appreciate it. Yeah, I learned a lot. Awesome.

Podcast Summary

Key Points:

  1. Cerebrus has built a wafer-scale AI chip with nearly one million cores, designed for high-speed, sparse linear algebra and efficient memory bandwidth.
  2. The chip’s architecture enables up to 10X faster AI training and inference than traditional GPUs by integrating compute and memory on a single device.
  3. The company’s software stack is adaptable, evolving from training-focused to inference-optimized, especially for fast coding and multi-modal models like OpenAI’s Codex Spark.
  4. Cerebrus is not in an AI bubble, as massive investments reflect real demand and the full value potential of AI remains undiscovered.
  5. Rising memory costs benefit Cerebrus because its systems use on-chip SRAM, avoiding reliance on HBM memory supply chains.
  6. Inference is now the primary market driver, where speed directly impacts user satisfaction and application viability.
  7. The company has secured a 750 MW compute contract with OpenAI and is expanding partnerships with the UAE (G42) and U.S. government on AI infrastructure and export initiatives.
  8. Future growth will depend on solving physical infrastructure bottlenecks, energy efficiency, and global scalability through vertical integration and innovative design.

Summary:

Cerebrus is a leading AI computing company that has developed a revolutionary wafer-scale chip with nearly one million cores, designed to deliver unprecedented speed in AI training and inference. Unlike traditional GPUs, its architecture integrates compute and memory on a single device, enabling faster, more efficient processing by eliminating bottlenecks in communication and memory access. The chip’s design is flexible and scalable, evolving from early training-focused systems to optimized solutions for high-speed inference—particularly in coding, voice, and reasoning applications.

Cerebrus argues the current surge in AI investment is not a bubble but a reflection of genuine, growing demand and untapped value in AI technology. The company’s systems are especially resilient to memory supply constraints, as they use on-chip SRAM instead of HBM memory. A key market shift is toward inference, where speed directly impacts user engagement and product value—making fast, responsive AI essential.

Cerebrus has secured major partnerships with OpenAI, powering models like Codex Spark, and with G42 in the UAE to build AI models in Arabic and other local domains. S. government on national AI initiatives, including the Genesis program, to advance scientific discovery and enable technology exports to allies.

While data center construction remains a challenge due to physical and logistical constraints, Cerebrus is pioneering solutions through efficient design, energy optimization, and strategic infrastructure partnerships. The company believes that faster, more accessible AI computing will unlock transformative applications across industries, from health and science to global digital economies.

FAQs

No, the AI industry is not in a bubble. The massive investments from major hyperscalers reflect a strong belief in AI's demonstrable value. The full potential of AI has not yet been realized, and these investments are aligned with a legitimate value hypothesis.

Cerebrus's chip is a wafer-scale engine with nearly one million cores, designed for high sparse linear algebra compute, fast communication, and high memory bandwidth. Unlike traditional chips, it integrates memory directly on the chip, eliminating bottlenecks from external memory access.

The chip delivers up to 10 times faster AI inference and training compared to general-purpose processors like GPUs. This is due to its optimized architecture for sparse operations, high communication and memory bandwidth, and reduced latency across cores.

Traditional memory (like HBM) faces supply chain constraints and manufacturing bottlenecks. Cerebrus avoids these by integrating memory directly on the chip, making its systems less dependent on volatile memory prices and enabling cost-competitive performance.

The primary use cases are fast coding agents, voice models, and reasoning models for AI assistants. Cerebrus’s systems are now the fastest for multi-modal models like Google’s Gemma 4, delivering 1,500 tokens per second — 10x faster than GPU implementations.

Speed directly impacts user satisfaction and retention. For applications like search, coding, or voice assistants, every millisecond of delay reduces user engagement. Faster AI enables new applications, such as real-time agent workflows, that were previously infeasible.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.