The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
72m 41s
Andrew Feldman, CEO of Cerebras, discusses how his company built the largest chip in computing history—58 times larger than a GPU—to address the critical need for fast AI inference. He explains that speed, measured in tokens per second per user, has become the dominant conversation as AI transitions from a novelty to a productive tool. Unlike GPUs, Cerebras's chip avoids three major industry bottlenecks: HBM memory, TSMC's CoWoS packaging, and 3nm fabrication, using 5nm technology and no scarce DRAM. This design allows for faster processing, making inference feel real-time. Feldman notes that perfect timing came after a decade of struggle, with initial indifference in 2020, but Nvidia's $20 billion acquisition of Groq validated their approach. He argues the market is not a bubble—demand outstrips supply, unlike past overbuilt infrastructure. The rise of AI agents also drives CPU demand, as they require action-taking. China is a real competitor with power advantages but lags in chips; Cerebras avoids selling there due to geopolitical and regulatory reasons. Ultimately, Cerebras focuses on data center compute, leaving edge devices for later, and believes a multicellular chip ecosystem is healthy for the industry.
This is the largest chip built in the history of the computer industry. It's 58 times larger than the GPU. And for AI, paper chips process information more quickly. And therefore you get answers at last time. For AI work, big chips are undoubtedly the best way to go. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the class. You solve the problem that nobody in the history of computer solved. And we delivered it in 2020 and nobody cared. No big care. No re-bought it and nobody cared. Everybody said we were crazy. It would never were. So then we built the next one. Hi, I'm Matt Turb. Welcome to the Matt podcast. My guest today is Andrew Feldman, co-founder and CEO of Syrobras, the company that built the largest chip in the history of computing and just pulled off the biggest semiconductor IPO of all time. Andrew has been everywhere talking about the headlines. The $20 billion plus opening ideal, the IPO, but this conversation is a bit different. We started from what is a wafer and built up step by step. Why GPUs struggle with fast inference. The three shortages nobody talks about. The decade in the desert when nobody wanted this chip. And why Andrew believes that CUDA is no longer a mode for Nvidia. If you want to actually understand the current chip landscape and how AI inference works at the silicon level, this episode is for you. Please enjoy this fantastic conversation with Andrew Feldman. I've had a front-place G-Star. It would be to talk about speed. So it has speed become the dominant conversation for AI today. What happened I think was for a long time AI was sort of a novelty. What I do is like a parlor trick. It was cool, but not useful. And what happened somewhere around the middle of 2025 was the AI got smart enough such that people began to use it. And remember we make AI with training, but we use it with inference. And suddenly people wanted to use it. In the minute you want to use it, the minute it's productive, right? Speed matters. Fast tokens are more productive. And so the conversation moved from everything else to how do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens. We get more done in less time. And they're forced more by the-- Right. And what does speed mean? Is that the equation of token speed? Is that the completion of the task? What's the right metric? Should the right metric is tokens per second per user? That's how fast you get the first token all the way through the last token in your response. And it's true for queries from chat, but it's also true for agentics loads, right? If there are sort of multi-cycle turns waiting is amplified. And so what you want is blisterly fast responses, so that the AI feels like it's in real time, you can engage with it. Hey, Bing? So it's the braunbad moment for us, Ray. I say, I think that's right. And I think that's a very good analogy. I think if you think of something like Netflix, right? When the internet was slow, right? Netflix delivered DVDs and envelopes. You would get a DVD at an envelope. And when the internet became fast, they didn't get more efficient at delivering DVDs and envelopes. They became a movie studio. Right. The speed enabled them to become something completely different. And that's what speed does in general and in particular for AI. It opens up a whole new domain that allows you to use the AI differently. You will stay longer. You will come more often and you work on harder problems. Yeah. So it's literally a question of the UX, right? That's just, nobody wants to wait a few seconds. That's right. How big is the market for slow search? How big is the market for dialogue is zero? How big is, how long will you wait for a website to resolve? Will you wait eight seconds? Or nobody waits? See? And so is the exact same with AI. Yep. So no more people waiting. Was there laptops open while the agent? That's right. Well, it's running and running and running. I think that is not what people want. Okay. Great. Wonderful. I would love to talk about the landscape of the chicken just for right now to help people visualizing what you guys are. So there used to be basically, this comes up about one chick to do it all. And obviously this is evolving dramatically. This you guys, this, this grog, this, there are GPUs. There are then people have a lot of training. People have a lot of GPUs. So help us compare and contrast. Who does what for what? Well, I think there used to be one chip to do it all. It's called the CPU. Yeah. Full through right. And there emerged a co-processor to do discrete graphics. And as the AI workload became interesting, we focused chips more and more on that particular workload. And today several companies make traditional GPUs. So in India, AMD, it made very standard GPUs. There are a group of companies who, the hyperskillers make, make some of their own parts. So the TPU. So today is Google. The TPUs Google and training them is going to go US. And then there were group light saber. We were among the pioneers to build a part from scratch, optimized for AI and nothing else. And we were inside of a hyperskillers for, we were optimized for a hyperskillers problem or for one lapse problem, we were building a chip for a collection of AI problems. And all our thinking was around how to accelerate AI. That's sort of the landscape to this day. Yeah. This even more specialized ones, right? Like edge developed, specifically, version formers. And those are called asics. That's why I couldn't, can you maybe define what that term means? And basically it's an application specific integrated circuit. And it's a word that now has a wide range of meaning. It means that you have made a series of choices away from general towards a narrower class of problem solving. And that you've made some decisions in your architecture that make it much better at some things and much, much worse than others, right? And that's choices that are made across the spectrum. So the TPU has made some choices like that. It can't do graphics. It's very good. It made it's multiply. It's not very good at a collection of other things that we do in mathematics. Saying for the GPU, we've all made different choices. Right now in production, there are, it really at AMD, there are a collection. There's the TPU for Google, Traidiums. Just coming up as a Maya part, as a part from Microsoft, Serebris. And it was one other, a growth that got acquired by Nvidia. Yeah. And to the general chip versus business-wise chip, it was actually increasing because I you guys in your CTL celebrated one GROC was acquired. Was that, what was that? Was it a recognition by Nvidia that your vision was right all along? Yeah. I think one of those same ideas, sort of most durable modes was the perception that the GPU could do every thing. And it was all you needed for AI. And the acquisition of GROC for $20 billion and the structure and the speed with which they chose to do it, made clear to everyone that that wasn't true. That the GPU architecture couldn't do, could not do fast at trends. And that this market was large and growing quickly. And we were the fastest at it and the largest. And our sales were more than 10 times the GROC's and they paid $20 billion for the number to collect. So that was a good day. That was a good day. And the recent announcement of Hallotino, we've been opening up which is a very large customer of yours and Broadcom. Why does it fit in that picture? Is that one of those highly specialized AZX-type? Right. GROC is for inference. That's right. Hallotino is a part that has been a long time coming. It was announced with Broadcom before OpenAI did the deal with us. Remember, we did a huge deal. This is quite the largest deals in Silicon Valley has to be worth $20 billion. I think they
We have a yawning need for silicon. And one of the things that OpenAI has been, I think, the best Hada in the industry at is looking at an exponential curve, adoption, rate of growth, not being afraid of what it says sweet, right? Others have had to go out and strike really bad deals to get capacity. After OpenAI saw this coming, they struck big deals for memory, for compute with us, with others. They've really been sort of visionary in understanding what it means to extrapolate from an exponential curve. And that curve is AI usage, right? It is growing so unbelievably fast. The whole industry is chasing demand, right? Usually, well, often it's the other way, often people are building it, hoping it will come. And in our case, all of us are chasing what people already want to do, let alone what they might do in the future. That would be the metamilia, because of this such a interesting topic. To finish on GPUs versus inference versus ASICs, is that the permanent situation for your perspective that we're going to be in this forever multicellular kind of environment? Yeah, multicellular kind of environment is a healthy ecosystem, right? I don't think anybody would say that the X86 environment was a healthy place. There were 20 years, there were two players. And if you notice, when a new workload came around, a self-alwork load, it was very similar, but required low power, and battery, and both lost, right, until that zero share, and that zero share, and a company nobody pretty easily heard of became the largest seller of computer in the world, and that's our A and that healthy ecosystem have lots of different ways to solve problems. What is China, so this Huawei, Sandchips for deep sea, can this emergence of a full stack, Chinese AI factory for like our better tone? Is anybody happening or is that overstated? No, no, that Dubai is real. No, I think they are an industrial adversary. I think they have really interesting investments that made investments in power in the grid, which is real weakness in the US. You share in France, you have nuclear, which sort of turns out to be pretty cheap power and pretty clean. But in China, they made huge investments in their grid. They are behind in chips. Their approach was at the Netslangle as open source models, where they produce some extraordinary models, not as good as G-PT or Anthropic or Google's Gemini, but very good. I think they don't have the Sandchips strikes at the US, but they have that leadership and other domains, and particularly in power, which is what we need for Bay of the Suners. Well, what is China following their race strategy? We don't sell. Corrections. For geopolitical reasons, for business reasons, for regulatory reasons, and for some geopolitical reasons. Okay. Is there a role for local chips and video ads, some announcements around just building chips for Windows computers? Is it foreign currencies? Is that something that you guys think should be part of the multi-sidicone ecosystem? Yeah, I think if you look at the way the ecosystem for apps and cloud emerged, everything you can do, you should do it in your phone or your laptop. But the ability to get real process in power to a phone or to a laptop is constrained because they're generally working off a battery. And so you want to do as much as you can, as close to the data as you can, but the truth is, in most situations, for real compute, you have to go to the cloud to the data center. And that's exactly the way it's going to be with AI. We're going to do a lot of work on the cell phones on the laptop. But for the big work, you're going to go to the data center. And that's where our focus is, our focus is data center compute free. Great. So as you mentioned, the how powerful the demand was and demand, you know, right of production. So just to ask the inevitable question around market equilibrium, bubble or not. And so like the most recent thing we also said, there was a little bit of a flash crash earlier in June where 1.5 trillion in market value was lost and in the end of the last 300 billion that day. Now that was corrected within, you know, 48 hours. So the market seems very jittery around that general concept made up just, uh, definitely worth senior. That's what just based on what you said, that's like you'll, you'll resivably on the other side of that. Okay. Concern. I think that if you spent a lot of time staring at the public markets and watching their ups and downs every day or you're watching the wrong thing, I think that, right, the wisdom I received from the great thinkers of public market investing is that in the short terms, the public markets are voting mechanisms who is most popular. But in a long term, they're voting mechanisms who's created the most value, the most weight. And so what we can do is a, it only public company is as we cannot look at the day-to-day fluctuations and we can focus on building extraordinary technology, winning the customers, making our existing customers happy and sharing it by more. And being better every week. And so I think the day-to-day fluctuations are out of our hands. And that's about the voting. Yeah. And we mean just to play devil's advocate on the demand. It seems like a lot of the demand comes from the big labs, which, you know, themselves are financed by a venture capital or private equity edge funds where your quality tends to be silver and investors. Is there any concern that a demand for chips comes from the labs, which are maybe artificially financed? And that the, you know, if you add a few circle of deals here and there, do you sense any differentiating? That's not exactly our experience. I mean, obviously we have enormous demand from impolokidan. But we have huge, you know, dozens of other customers who would find a place in very big orders or hubs. And historically, bubbles were when supply got out ahead of demand, right? When in the '90s, we built out a telco infrastructure. Right. We built out fiber years before it was going to be used. And it took six or eight years and it all got used. But it was a sort of, if you build it, they will come mentality. Whereas what's different about AI right now is we're all trying to catch up. We're trying to build data centers faster. We're trying to increase our demand, our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. And we're ready. Those sort of overwhelmed with compute were overwhelmed with demand for memory, which was a real weakness in the GPUs. It's not a problem we face. In the ecosystem, there are constraints left and right. And that doesn't feel like a bubble. Did you know actually the bulk take on that? Because that's sort of dressy right. It seems like the bottleneck keeps moving because what's the current problem of memory shortage that people may have heard of? So there are three major bottlenecks right now. The first is all GPUs and all ACES except us use a type of memory called DRAM and a particular flavor of that memory called HBMS. And that memory is made by three companies in the world. The clinics, Samsung and Microsoft. Three companies. And they're sold out. That's the number one problem. We don't use it. And it's sold out as in. It's like the only way to. As old things are to work, the only way to increase supply would be to build new factories, build more factories. And so you have this huge problem for
GPUs and even now for CPUs, you can't get this memory. Because of our architectural choices, we don't use it. The second constraint was a process at TSMC. Process was called COOS. And this used silicon as a motherboard. And this is what in ZEDIA, AMD, and others use as a motherboard so they put their chips and the memory chips on it. And it's more efficient than a traditional motherboard. I had. That process is sold out. It's highly constrained. We don't use it. Third limitation in the space is 3-dynameter factory space at TSMC. So this is the most aggressive factory. So with the most advanced technology, our chips are at 5-dynameter. So we don't use the 3-dynameter technology. And so. And then that's so fascinating. So to ask the layman's question, so what does that mean, StreamLiger factory? So this is a solid-edicated factory. That's right. They build factories and the factory etches, transistors that are a certain distance apart. And the smaller the distance apart, the more transistors per square millimeter. And the more you could do per square millimeter. And so the history of the computer industry making chips has been we've gotten better and better that put in transistors closer and closer. Yeah. And so we used to be at 16-meters apart. And then we went to 10-7-5. It's three and they're working on two, one point eight. And so right now the bulk of the GPU world is off three. And there's tremendous congestion at factory. And that's literally different tooling and the factory. And we use the 5-nm factory. And this is our shit here. Yes. Wow. You can't market this is the largest chip built in the history of the computer industry. And it's 58 times larger than the GPU. It has, you know, 2.5 or 3,000 times more memory bandwidth. Of a, and 4 AI, bavre chips process information more quickly. And therefore you get answers at less time. And obviously you don't want a bigger chip for a laptop or phone or for lots of other things. But for AI work, big chips are undoubtedly the best way to go. Okay. That's really great. So let's, so I want to do a bit of a deep dive on that on the tube itself, in a second. But to close on this, so three bottlenecks. You also mentioned CPUs, a couple of tumultuous conversations. And this seems to be a theme around the emergence of like a CPU shortage as well. Is that so why is that true or what causes it? So, a gen to AI is a world in which AI doesn't just provide the answers. It initiates action. So that action might be good on a website. It might be learned, gather some data from a website. Bring it back, take another action. Those actions are done by CPUs. And so as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more more CPUs. And that is driving up the consumption of CPUs, and therefore the demand for CPUs. And so this sort of huge push for more CPUs is being driven by AI on machines like ours and GPUs, doing a gentick work. And asking the CPUs to take an action. To go to a website, to order a burrito, to find a piece of information, to pull it from storage to all that work is being done by the CPUs. Including in a system like yours. It is just to like it. Oh, it's interesting. What does it mean? So you have your chip and CPUs on the side, and those will be a system? It means that AI, when we do AI in a gentick, the AI processor is like the brains. And the CPUs are like the body. They're taking action. They're doing things in the digital world. Under the direction of the AI, which is running on the accelerator, on the 3 bar system. And so as we do more and more AI work, and more and more agentic work, we're making more and more calls to CPUs. And therefore the demand for CPUs is through the roof. And CPUs experience the same memory as yours. The same memory as yours. You see exactly the same memory. That's right. Yeah, it's going to make some products in a second. But let's talk about you and your journey a little bit for people to contact. So obviously you've been pulled in. Well, that's an incredible idea, always perfect timing. And I know I know you've lived a little bit. The way you have perfect timing is to have a horrible timing for 10 years. That's so you have perfect data. And very much to this point. So you guys started building this bigger chip, focused on inference, or in decade ago. To 2016, which makes perfect sense in retrospect. Given how long it takes to build a technology like this. What was the vision? 2016, I guess it was four years after ImageNet. It was all Vian. It was all Vian. It was all Vian. Yeah, Abolution. And that's why we don't believe the right thing to do is to embed the latest and coolest model into your circuitry. That's a mistake. We're the fastest in the world to transformers. And our architecture was set before transformers existed. Right. We're the fastest in the world to diffusion before an architect was set before diffused. What you want to do is take the underpinnings of those so that when the market moves, you can be good at that as well. Otherwise, because of the short life. To the ASX question, or earlier. So that's what others do. They build the architecture of the model into the-- Some's half, of some half. Okay. And historically, that's been a structural mistake. It has turned from the sixth. So you were doing the broni or is in toll platform on the way. That's true. So you're starting with the model. You want to think about what is the underlying calculation? The underlying calculation at all this work is Sparks linear algebra. And if you can accelerate that, whatever the bottleneckers invent, you can make faster. And that was our approach. Yep. So you started in 2016. You had a prior company that just hold to your AMD. Or some of the lessons you're on there that you took into your service. I think the lessons are large and medish. I think what are the things that I guess experience is another name for having made mistakes and learning from them, right? I think around-- We have, as a team, built dozens of chips of the past 20-clock years. And there were terms to experience and chipmaking as an artist. And we built a different type of computer at C-Micro, a type of computer optimized for low power and optimized for a workload that was very different than AI for something like WebRozon. But the fundamental underpinnings, the questions you ask as a computer architect are always to say, what can I do to make this work faster? And is there enough of it to make it worthwhile? So I think the two questions we ask, should we build a part for it? Should we-- what could we do to build a chip optimized for AI? And will there be enough AI so that you can build business around it? Those are the questions we asked in 2016. You know, the flip side of that was-- I don't know. Wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years, and been pushing pixels to a monitor, was suddenly good at a new world? It'll-- what would that be, a certain nipitous? And we can't believe it wasn't the right architecture for it. It was just better than the CPU. And that we could build an architecture that would be faster, that would use less power, and could drive down the cost of the app. And that will was the journey. So, Sir, we're already-- Then we spent time in the desert. We can visit. And we wanted-- We wanted-- in the desert. Yeah, but-- Yeah, maybe walk us--
a little bit through those years, for any, you know, founders, especially NDPEC founders, so listening to this. So, what was the, so first of all, what was the issue with the market timing? Was it the technology was not working? And then how did you go about it? As a team, you know, I get you'll be on board and you're investors and raising more rounds. So as you presume of any, then I have the full points that you're one beneath. So like for us, we, at the beginning, we were honest with our VCs and told them we're, you know, attack a really hard problem. We worked in a build something that was a little bit better than a GP. And that our idea, our strategy was that you will never be a great company like NVIDIA by doing something a little bit back. See that they're going to buy everything for less. They're going to have price and pressure. They're going to be able to bundle. And the right strategy would be to do something incredibly hard and engineering that was way better. 10, 16, 20, 30, 60 kinds better. But to do that, the ordinary and obvious paths are all closed. Everybody else has taken the bar. And so what we observed was that speed in inference was going to be a function of members. And that there were two types of memory. There's this g-rab of hbm and they can store a lot. But there's four. There's another type of memory called Esther. It is unbelievably sast. But per unit area can't store very much. And for graphics, everybody had always used DRAM that used hbm. And that the AI workload was fundamentally different. In graphics, you move data to the GPU and then you work on it for a long time. Then you send the result. So the time spent in total of movement plus work is dominated by work. In inference, the AI gets the exact opposite. You move a huge amount of data all the weights from memory to compute. And you may want calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data. So that's why g-teams have so much trouble being fast. So we observed that if we chose a strategy, using sram, we could be fast. But then we have to overcome the trade-off. The sram can't store very much. Now, let us to the solution if we could build a part vastly larger than any parting history. Right, the size of a dinner plate, we could stuff it to the gills with sram. And thereby get the benefit of sram that it was fast and overcome the weakness that it can't store very much. And that led us to a strategy called wafer scale. This chip is made from a single wafer. There it comes out of TSMC. Don't define what a wafer is. A wafer, all chips are cut from a wafer. A wafer is a circular piece of silicon that's 300 millimeters across. And the process of chip making stamps out like your mother does with a cookie cutter. Stamps outchips. The biggest chip that had ever been built before us was 800 square millimeters, 840 to be exact. And this is 46,000. So we have to invent all sorts of new technology to build a chip this big. And once you build it, you have to invent ways to power it, to cool it. There are no vendors waiting for you. Because you look like nothing else ever made. And that took years. And it was a deep tip problem. It had never been solved before. And we had an 18-month period where we were spending 8 million a month and we couldn't build them. So if you're deep tech founders, you have 8 million a month. Yeah. Why so much? What was the cost of just understanding how much business has been? Because what everybody thought was hard, we solved quickly. And what nobody else knew about because they'd never actually done it turned out to be really hard. You know, imagine that I tell people, imagine the first group that was going to climb Everest. And they're at base camp. And they're having tea with a group that just failed. And the group that just failed says halfway up. There's this part. It's unbelievably hard. We couldn't do it. Okay. Your team climbs up, makes it all the way that comes back. They're having tea again. And the team that made it leans over the team that hadn't made it said that part in the middle. That wasn't the hard part. Because nobody got in past it. Nobody got in past certain things. So they didn't even know what to be afraid of. We now know. And it was something a step called packaging. And that's how you fix a way for to motherboard how you deliver power to it and how you call it. And nobody did that before. And over that 18 month period, we became the best in the world that from approximately zero. Hey, and we did that by sailing again and again, and using good engineering methodology and do a lot of failure analysis every single failure. So we failed differently. And hand it again and again and again. And we told our board, you know, we met with our board over six weeks. And yeah, this is the strategy. No reason we're done this before. This is what we're going to do. And we had to invent new materials. We had to invent new techniques. We need uh, ended up building things that that everybody else that cardierers who could do. But at all this to 2019, when else would solve it. And it'll, my co-founders and I the first time it worked, it was sitting in a tiny little lab. And we couldn't believe it. We were the first few things this tree can be to bake one more. And we just stared at it and watching a server run is about as exciting as watching paint their eye. And here we were. The five of us just staring at this machine not believe it might work. It didn't work. Was it a bigger moment that I'd actually ring into the bell or than just completely different events emotionally? Emotionally, it was a completely different thing. It was that we had made our idea work. And I think for deep tech founders, it's a particular home. There's always this little thing in the back. You're buying this. And it maybe is shit. Maybe we actually crazy. That's right. Maybe it's no fail. And maybe it's not going to work. It has maybe I don't have time. Maybe we're going to run out money. Maybe maybe maybe. Maybe. And the flip side of that is the joy that this was my co-founders ideas. I mean, these are their ideas manifest in the world. And that is an extraordinary thing. And it is when someone's ideas take physical shape. And then the next step is when you watch other people's works and on top of your idea. Then you know you love making infrastructure. And so when we ran the bell and we went public on May 14th this year in the largest semiconductor IPO artistry. And we did something unusual. We invited all our engineers who'd been with us to start. Everybody who'd been with us for the nine years and their families. And we all rang the bell together. We all got up on stage. That was sort of a moment of a different pride. We had done this together. And that we should manage not to die. It's untagridored startup. You know, people are honest. They don't tell you that. But we'd avoided some, we'd made plenty of mistakes. But we've been avoided the saddle ones. And we made it through to a level of success that gave us the opportunity to pursue a new level of success. We asked what an IPO is. It's not the end. So you've gotten to a plateau that you can chase a new level of success in the public market. And that felt pretty good all of my amazing, amazing, thanks for sharing. So just to go a little deeper on some of the product and and technical stuff. So one of the trade-offs of building a bigger chip, one that may come to mind is failure mode. Sure. If you have a lot of little chips on a big way for presumably you can isolate the problems. If you have a big one, then everything won't fail in the same time. So you often think very cherishing about failure mode. And we invented a technique that had about a million identical tiles. And if one fails, we'd shut it down and we can use or redundant ones, then she'd go A and so E it should go to
go big, you have to think about in the very architecture of the computer, how you're going to manage daily. I mean, the JPs have a huge failure rate, so I'm sure you guys have spoken about this. Infit mortality is enormous, and they fail all the time. Because big data, they say, spoke put out a paper on a number of failures that get in a big cluster. Now, we can do some other things because we have all this compute in one spot, we can invest more decoult. So we pioneered water quality in AI systems, and we run these much colder in JPs, and the failure mode in electronics is temperature. And so by running the standard, we are more reliable. And so that was an advantage. But again, we have to adapt the technique to cool off. You know, chip this big. And so I think we have a system mentality, right? If you're going to do something big, you've made trade-offs. And you have to think about an architecture that allows for redundancy and repair. You have to think about the pros and cons of every aspect of the architecture. And that's how you do something different and new. And again, just drive the point of home to make sure this, I guess, clear take away for people from this conversation. So explain it to me, like I'm maybe not five, but 15. Why is this faster than a GPU? What fundamentally makes it faster? And what country is super fast with a GPA? How tokens are generated in inference is why it's faster. In an inference to generate a word, and our answers are a whole stream of words. They can be code. They can be pixels, but call them words. Talk. The weights of the model are moved from memory to compute. Calculation occurs, and that generates the word. Is a pre-fuel versus NECode? That is both steps. Okay. And you want to maybe explain what pre-fuel and NECode are. Okay. There are two steps in the computation necessary to do inference. When you type into chat, GPT, it's like maybe the history of this village, prior to World War II. And it can't see you. Two things have happened. The first thing is your prompt has been processed as step one. And step two, your answer has been generated. And the way your answer is generated is called decode. And decode is sequential. Processing your prompt, which would call pre-sale, can be paralyzed. So you can process many of them simultaneously. But the speed with which you get an answer is a function of the decode, and it is step by step by step in sequence, and that can't be changed. And how you do that step is you move weights, which are the intelligence from the model, to compute, to generate a word. You do a calculation, and you get a word. And then that word is used to generate the next word, where weights are moved from compute. So the process is what is moving weights from memory to compute. So how big are the weights? In a little model, like a 70 billion parameter model. The weights are about the size of 100 HD mookings. So to generate a single word, you move 100 HD mookies from memory to compute. And then you have to move them again for the next word. And you want to do this a thousand times to get a good answer, a thousand words. This is where HBM is slow. This exact step is where HBM is slow. And that exact step is where by having all the S RAM here, we're blisteringly fast. And so the speed of moving weights to compute, if about two and a half thousand times to ask their here, then on a will then GP. And so that's the essence of what we've been able to do here and why it's so much faster. Yeah. Fascinating. Do you think that models need to really evolve as well in that fast AI world or is that just to. No, I don't. Exactly. Because remember two things are happening. One, we want the models to be scarters. Or two, one of the ways models who get in smarter is with RL. And RL uses inference inside a training. And so the faster you can do the inference inside a training, the faster you're creating. One of the things that is at a market for you guys in the AI world, that's a miracle for us as well. So you're not just. And perhaps you're still so, we do. For RL and we do traditional training too, not for the largest models, for the largest lab, but for the next tier. For pre-training, for pure pre-training. For training, fine-tuning, full set. So you should do pre-training. Plus training was RL, some pre-training for the other labs. I mean, GVU is just a little better than you for one job there. So GPUs in trainings, have some challenges that have been solved by very narrow selections of the community. GPUs are a very small chip. And the calculations that we need to do in training are very large. And one of the most complicated parts of training is the breaking up the calculations and spreading them apart on multiple GPUs. And that's called distributed compute that has historically been the domain of the super cookie world. And is very difficult, not just because cracking a problem and having lots of others worth thought it is hard, but they have to constantly share information. And that sharing information is why they need it to buy mailbox. Is they needed to control a sabreck over which all is sharing in order to get an answer would happen. That breaking up a big matrix multiply, a big calculation, is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And the best labs in the world are going to be good at that, but nobody else is made. When we run training, we don't have to run that way. We run what's called model data per watt. And data parallel is very simple. And so it allows teams who were good and very good to quickly test ideas in training. And so we are easier to use and faster because we allow them to use a technique which is much simple. Let's talk about the energetic world. Maybe starting with reasoning. I think I saw a blog post where you guys have a given state of that reasoning was not always the right solution for all problems. How do you think about this and what that fits in the answer to the world? Reason is a technique. A little bit like what you were in eighth grade and you wrote different graphs of the paper. The single shot in front of the city. You get an answer. You write a query. You get an answer. The easiest way to think about reasoning is you are going to deced several graphs. It's going to break the problem up. It's going to solve the parts. It's going to bring them together. It's going to review the results. It's going to improve the results and then it's going to give you an answer. And so that is going to take more compute. And if your computer is slow, that's going to be more than an error chip. That could be crippling. And so as the best models, all of them, whether they're domestic or Chinese, whether you are open AI or entropic or Gemini, whether you're any of the Chinese mouth, they all move to a reasoning approach. But that meant more compute was being used during inference. That was a huge advantage for us because we were fast. And it made the GP's slowness stack up and it made our speed have an even bigger advantage. And so this was a huge blow in for us. We think
this is a kind of stay as a fundamental essence of the way these models are run right now. You also wrote about verification and whether what was the bottleneck in agents was whether the reasoning was good enough whether we're smart enough or whether there was a verification problem. Sure, I think the verification problem is a little bit like the guardrail problems. What you'd like to do is after you've written two or three graphs of your answer, you would like to be sure that it was a wrong and that's your verification step and you can do that maybe with a different model. You can do that by asking your existing model a similar question in a different way. Right? All of these are ways you can pressure test your result. Guard rails could work the same way. You want to review either with a model or with another technique through a scoring mechanism that this question isn't out of line that this answer isn't about how to make biological weapons or calling upon information that that you are directed to FBI. Right? And all of that takes compute time. And so whether you're trying to improve to a reasoning or whether you're running guardrails by being faster, all right? You can get results in less time. Hey. So what do you think the world is evolving? Quagions are the bunch of smaller models running faster during more verification versus a large model? But I think those work together. I think your big model produces an answer then you want to double just want to check your data. I mean in the journalism industry, people would write papers that have data checkers, right? Somebody would go and make sure back in the world journals check data. Or fake the data checker, right? That was a job. Each claim was checked independently. That's a different model. Right? The made model wrote the piece. And then a little model checks some of the answers. And I think that's a very, a very good way to go by. Where does a multi-windowality or the larger models falling in your world? We just announced that we were fastest and the world on one of Google's multi-modal models. I think the truth is that there's very little text that doesn't have charts and graphs, right? You must understand text, be able to understand illustrations, graphs, charts. And so that's sort of the first and easiest part. And then you ought to be able to create, and then you want to understand energies. And I think the new models are very, very good at that. Obviously what follows that is video because a video is just a collection of images. But that takes an enormous amount of compute right now. And that's what are the resources in the sort of set aside by deleting labs. So unbelievably, so reputation in tends to be great. This will be a bit of a bit of a business. So you guys obviously make our provider as we discussed. You're also a cloud provider, metacenter provider, the world which, so one of the different parts. We may compute. And that computers optimize for AI. There's the fastest today I go. If you have a data center, we will sell you hardware for deployment in your data center. If you don't have a data center, and you'd like to rent it by the month for the year, we have data centers. They have so you can rent our, our equipment through our data centers and through our cloud. And so that allows us to get AI natives, as well as large enterprises and governments. And what's the, and they're sharting the end of the 50 50 last year. And I think this last quarter, it was maybe maybe 75 25 and save a row of hardware sales. I think this year it might be 50 50. And as our open AI deal continues to unfold, it will probably be 30 70 with 30 on premise deployments of hardware, the 70 cloud. Great. Let's talk about that open in the ideals and says such like a major historical milestone record making. So it's providing up to was 750 megawatts, which is interesting. By the way, as a metric because of a chip provider, but this is power or so, he's a shorthand for us. It's a shorthand. I mean, it turns out right now. And we didn't talk about this because there's sort of in the adjacent supply chain. We went through the shortage of of memory, or we had to do the shortage of a process of co-awesome, three nanomere capacity. The other limitation in our industry right now is data center availability. And I mean, that is the only way for everybody. And that's why anthropic did issue chose sort of very expensive deal with the law for the data side capacity. Our deal with them with with open AI was because data center capacity is a limiting constraint measured the way data centers are measured in in megawatts. The deal is 750 megawatts 250 megawatts and 26. On a multi year lease. Additional 250 megawatts and 27 on multi year leaks and additional in 28 multi year leaks. Any during the data center for them, we're providing the chips that we're going to the data center for them. We're delivering the full-cloud solution so they connect to us via an API. Basically, I mentioned 20, 26 devices in muniply that's channel one. I am looking for data center. My next being is in fact with a data center provider. We're doing a lot in Europe right now. A lot in the in the Nordics. Is that because it's a closer to power sources? Yes, because there is low cost power. I'm not clean low cost power and low cost pool. Doesn't matter just like in the company's whether data centers is is located in terms of spin. Yeah, there is an additional latency called transport latency and that's. The speed of light through fiber heard to get from. Helsinki to New York and you have to account for that if your customers in New York and your data centers and Helsinki. It's usually about two thirds of speed of light. Be it case, how long it takes. You would like data centers on the same kind of. That's the opening ideal. There was an exciting deal with AWS as well where you. It's a it's a coach chip solutions right. It's the disaggregated solution you mentioned before where their premium part. Is doing the pre fill and is doing the paralyze up all stuff so it's training as trading is doing that the pre fill staff and our chip is doing the decode. And so you get a lot more. More smoothly fast tokens. It's a good deal for us. It uses their data centers. So these are deployments in the AWS data center. It's fascinating. I like talking about this industry is how it seems that flexibility is so important. There's just something very mis-willuxion like unit chips and get chips in your data centers. Everybody is buying from different suppliers to reduce dependency. I think that's one of the reasons why we. We. What from being a traditional chip consistent provider to also offering data centers is that what what our customers want or fast tokens. And anything we can do to make the delivery of fast tokens easier for some of them that's in their data center for some of them that's with an API just played your traffic to us. And we'll point the fire hose of fast token back. Yeah. And just to get a sense for where you start and where you stop is you do not. So that at least currently provide the the the convergence. If I want to run to me like I know you have like incredible stats for a Q and Gemma into those being different to mention them. But you don't run those as a service or do you do you do you do you have a service was competing with the base 10 some fireworks. Yeah.
- Yeah, I think we have an on-demand service where you can come to our site and book a month. I think you can even buy buckets of tokens for Kimmy, or GLM, or some of these models. Many of our customers come there, get excited about it and then move to a dedicated offering where they take hundreds of machines for a year or two or three or four. Each one save proof it out. The better set for them in their work. Often they do a B tests, not surprising people like Faster. - So it's more like a testing. - It's a fully environment. Okay, you can go and use it. It's at 3.5. You can go and play around. - Fascinating, but that can become like a yet another big part of this. - Yeah, that's great. So you have chips, you have data centers, and you have a client business running in front of Salt and got 19 cents to go and first thing. - Thinking about modes, so, you know, famously, NVIDIA as CUDA as well, just because mode, what's your equivalent of? - Well, I don't think CUDA is a mode. - Correct. - Korean, we should talk about that. - I think two years ago, every state of the art model was trained in a CUDA flow. And right now, Gemini is trained without CUDA. Anthropoclaw is trained without CUDA. Open AI is trained with CUDA. So in a one or two year period, they lost 70% share of training models or is the state of the art. All because Gemini is trained on TPs, that's very good. Out of Anthropic is trained on training of. Not age, you can go. - Take. - And so, I think the story of the mode is still present where the data, I'm sure the mode is clearly shrinking. And there's no mode in inference. - Yeah, it takes you eight keystrokes to move from a GPU to us in the cloud. And it's for a suite. - Oh, a suite. - A suite, a suite, a suite, a suite, a suite. - A suite, a suite, a suite, a suite. - That's it. - To mode your traffic from a GPU API to us. And so obviously CUDA was sort of enormously important in the creation of our industry and didn't allow the graphics process in here to be more general than graphics process. But since 23, 24, quick, five, I think it's a ability to solve the durable mode, and she's short of substation. - Hey, hey. - Any of the whole ecosystems, strategies, you think about your mode? Do you think some of that any mode can be present in this industry? You're building a whole ecosystem? Is that, is that, is that, yeah? - Yeah, we built an ecosystem. I think our mode comes from the fact that by virtue of our architecture, we are doing things no one else can do. - So it's, and it's not, that they can spend more money or where they can't, it's, hey, more for this GPUs, if you want to fast, you can't fail it. - Well, I mean, Jake, it, it doesn't work. And so that's where we're building this sort of our strength, and that's how we're, where to deliver value to customers. - I do think about supply chain, we mentioned supply chain constraints for others, whether I want to use supply chain constraints, are you okay with your old CSMC? - New RTSMC. We have very close collaboration. They were investors, they've been exceptional partners. I'm not gonna, I'll tell you an unusual story in 2017, we showed up, and we were about 30 guys total. And we showed up in August, I'm horrible trying to Taiwan. I don't go to Taiwan in August, the way is we grew roots. - That that is so nice here. It's only 90 years. - You're talking about bears today. - Yeah, but all the humidity, again, yeah, you gotta put on a sands. And we met with the leadership of CSMC, we said we would like a little Pipsqueak company. We believe we can solve a problem that nobody solved in history. And here's how we would modify the way you make chips to make this possible. They thought about it. And they said, we agree, let's do it. - In the meeting. - In the meeting. - In the meeting. It wasn't going for a month in the meeting. - Was that a prepared mind or that weren't just exceptionally fast on the issue? - First, the salesperson had gathered the decision-breakers. Second, we, our proposal was sort of really good at allowing them to use what they were good at. And it didn't require them to change a huge amount. But it did require them to make real changes. And I think they saw this as sufficiently bold that they would learn as they did it. And they also knew that AI was better on big chips. - Hey. - And so the combination of fair mind, willingness to take some risk, bold thinking from a very large company. - It is fascinating. I mean, that's how big companies went. That right. And how rare is that was explored there. - And what happened next? Like how long does it take between the decision in the meeting like about the expenses exceptionally fast to your - Two years. - Two years. - Chipmaking is a lot of hard process. And most of the time your first chip isn't a winner. And there are lots of startups now. So with really smart guys, like their first chip will not be ordered. - Well, the TPU, Google had some of the best guys industry first chip wasn't a winner nor the second nor the third. Fourth was really good. Now they're on their eighth and it's really good chip. - Hey. - The Anapolina team at AWS first chip wasn't great. - And the second third chip really good. - It takes time. And so we built a chip, we delivered it. And this gets to an earlier question. You know, we solved the problem that nobody in the history of computer solved and we delivered it in 2020. And nobody cared. - No big hair. Nobody bought it and nobody cared. - Yeah. - We were like, oh my god. Nobody, everybody said we were crazy. We'd never worked. It not works. Like nobody wants to solve it. - Yeah. - So then we built the next one. I mean, the first one we probably solved quite a year. - And nobody wanted it because the marker was not ready or because the product was not good enough? - Nobody wanted it because AI was a hobby. It was right. And who cares if your hobby is really fast. You care about fast, what it's in production. You can care about fast when you use it every day. And so we built another one. And you know, that one we solved three or five clinders. - Yeah. - Hey, we built the third one. And we solved tens of thousands. - Right. - Yeah. - It's amazing. - Yeah, isn't that interesting, Amaz? - And so we're going back to supply chain. Do you need to think about on-choring, diversification. So it's very hard to diversify away from TSA. Chips are so hard. And you actually, when you design a chip, part of the design is for the rules of that factory. Right. So you can't take your design from TSMC and go to somebody else because a huge amount of the work was meant to be sure your design is within their rules. And so Eamland and I think, only with one or two exceptions in history, each chip generation goes to one fat. So that's where we're going to be with TSMC for our next generation as well. - Hey, I think some out of them, we have a supply chain that is built in many parts, but we bring the chips back from TSMC to the US, repackage in the US and we assemble in the US. We do them that if that's churing in the US and then we ship from the US. I think when you're growing this fast as we are, right? There are a whole range of garden variety supply chain challenges. No vendor screws up a batch. I thought that one gets stuck in customs. The number of ways that things can go wrong with supply chain is unbelievable. But we manage these every day and we're increasing our manufacturing through code exponentially. So it's really that part of the business in Colbert. - Incredible. So maybe to zoom out as a last question, what's your best guess?
about where all of this is going, you know, obviously, who knows any eye, the next like fears, but like in the next year or two. Well, we know some things. We know that the model we use that you're using today, if they bowl or or any petyflage six, we'll be the worst model you ever use. And whatever you think is cool about it right now is going to be boring and backward to the six plus. And that is still exciting. And I think I watched the way our young engineers use it and it's very different. It didn't the way I'm using it. And it's sort of a fun time where you can learn from your young team members that they're using the AI very different way. I think the business of dashboarding and the business, it was just the eye doesn't. So the damage is doing the sass is that thing unrepairable. If our moneys, you know, you can ask your AI, they'll me a tool like Salesforce 30 seconds later. You have a working tools that I just unbelievable. And all the things that were difficult because they cut across your inside organizational silos, right? One of the things that's really hard. I actually want to know for your top performing people, when was the last time they got a stock option refresh and how much holding powers left. So how much invested stock they have left at today's price rate. It was like five systems. You're in work days. You're in your stock. You're in your car. Car, car, none of the, you know, when that's when to see you at once. How much holding powers for my top guys? And I used to have little tools I wrote for this and boom, I've got our little app that I had. And Spark, right? Where? You go. Well, what a story. It's just incredible to hear all of this from you, what a journey and what an exciting future. So thank you very much. I learned a lot and this was terrific. Thank you, Andrew. Thank you for having me on your show. I really appreciate it.
Podcast Summary
Key Points:
Cerebras built the largest chip in computing history, 58 times larger than a GPU, optimized for AI inference.
Speed (tokens per second per user) is now critical for AI usability, shifting focus from training to inference.
The chip avoids three major industry bottlenecks
Cerebras's chip uses 5nm technology, not the constrained 3nm, and doesn't rely on scarce DRAM or HBM memory.
AI agents require more CPUs for actions, creating a new CPU demand surge.
Andrew Feldman emphasizes that perfect timing came after a decade of struggle; Cerebras delivered its chip in 2020 with initial indifference.
Nvidia's $20 billion acquisition of Groq validated that GPUs cannot handle fast inference, boosting Cerebras's market position.
The market is not a bubble; demand exceeds supply as AI usage grows exponentially, unlike past overbuilt infrastructure.
China is a real industrial competitor with power grid investments but lags in chips; Cerebras does not sell there for geopolitical and regulatory reasons.
1
Cerebras focuses on data center compute, not edge devices, which are limited by battery power.
Summary:
Andrew Feldman, CEO of Cerebras, discusses how his company built the largest chip in computing history—58 times larger than a GPU—to address the critical need for fast AI inference. He explains that speed, measured in tokens per second per user, has become the dominant conversation as AI transitions from a novelty to a productive tool. Unlike GPUs, Cerebras's chip avoids three major industry bottlenecks: HBM memory, TSMC's CoWoS packaging, and 3nm fabrication, using 5nm technology and no scarce DRAM.
This design allows for faster processing, making inference feel real-time. Feldman notes that perfect timing came after a decade of struggle, with initial indifference in 2020, but Nvidia's $20 billion acquisition of Groq validated their approach. He argues the market is not a bubble—demand outstrips supply, unlike past overbuilt infrastructure.
The rise of AI agents also drives CPU demand, as they require action-taking. China is a real competitor with power advantages but lags in chips; Cerebras avoids selling there due to geopolitical and regulatory reasons. Ultimately, Cerebras focuses on data center compute, leaving edge devices for later, and believes a multicellular chip ecosystem is healthy for the industry.
FAQs
The largest chip is built by Cerebras, and it is 58 times larger than a GPU. For AI, bigger chips process information more quickly, leading to faster answers, making them ideal for AI inference.
Speed matters because around mid-2025, AI became smart enough for practical use. Faster tokens mean more productive AI, enabling real-time engagement and new use cases like agentic workloads.
Cerebras builds application-specific integrated circuits (ASICs) optimized solely for AI, unlike GPUs which are general-purpose. Their architecture avoids bottlenecks like DRAM memory shortages, using a large, single-chip design for faster inference.
The three bottlenecks are: 1) HBM memory shortage, made by only three companies and sold out; 2) TSMC's CoWoS process, used for chip interconnects, is constrained; 3) 3-nanometer factory space at TSMC is congested. Cerebras avoids all three by using different technologies.
AI agentic workloads require CPUs to take actions like fetching data or ordering items. As AI usage grows, more CPUs are needed to execute these tasks, driving up demand and creating a shortage, especially since CPUs face the same memory constraints as GPUs.
Cerebras started building its large inference chip in 2016, a time when AI was a novelty. They faced skepticism until AI became widely used around 2025, making their focus on speed and inference suddenly critical and perfectly timed.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.