Cerebrus is a leading AI computing company that has developed a revolutionary wafer-scale chip with nearly one million cores, designed to deliver unprecedented speed in AI training and inference. Unlike traditional GPUs, its architecture integrates compute and memory on a single device, enabling faster, more efficient processing by eliminating bottlenecks in communication and memory access. The chip’s design is flexible and scalable, evolving from early training-focused systems to optimized solutions for high-speed inference—particularly in coding, voice, and reasoning applications. Cerebrus argues the current surge in AI investment is not a bubble but a reflection of genuine, growing demand and untapped value in AI technology. The company’s systems are especially resilient to memory supply constraints, as they use on-chip SRAM instead of HBM memory. A key market shift is toward inference, where speed directly impacts user engagement and product value—making fast, responsive AI essential. Cerebrus has secured major partnerships with OpenAI, powering models like Codex Spark, and with G42 in the UAE to build AI models in Arabic and other local domains. It is also working with the U.S. government on national AI initiatives, including the Genesis program, to advance scientific discovery and enable technology exports to allies. While data center construction remains a challenge due to physical and logistical constraints, Cerebrus is pioneering solutions through efficient design, energy optimization, and strategic infrastructure partnerships. The company believes that faster, more accessible AI computing will unlock transformative applications across industries, from health and science to global digital economies.
I don't think we're in a bubble because I don't think we've really discovered the full value potential of AI.
What if we built one big chip and then left it whole?
This is the chip, this is the Wafer scale engine, nearly 1 million cores.
They're going to buy 700 Dmw of compute by cerebrus to power their applications.
The coding model that uses cerebrus now is Codex Spark.
Inferences is where you ring the cash register.
Medium speed tokens at really low cost.
That's not a market I'm interested in percent.
The market for fast tokens for the future of AI applications that are at a competitive cost.
That's what I want to go hunting.
I know Dario and the team really well, extraordinary team, extraordinary technology.
I hope we get the opportunity to do more.
Most computer chips are smaller than a posted stamp, smaller than a bottle cap.
cerebrus builds one the size of a dinner plate.
In this episode, I chat with Andy Hawke, who's currently the Chief Strategy Officer at cerebrus,
who is just IPOed and secured a 750 megawatt compute contract from up in AI.
Andy makes his case that the trillions of dollars pouring into AI infrastructure isn't a bubble.
He explains why memory stocks are ripping, why DTU prices out to follow,
and why their chip actually benefits from rising memory prices.
He tells me why the stat that 95% of enterprise AI projects fail is measuring the wrong thing.
He talks about the US selling chips to China and the possible job of the people who make that decision.
Let's get into it.
Andy, you are the Chief Strategy Officer at cerebrus.
Can you tell us about what that means and what cerebrus is?
Absolutely. It's a pleasure to be here, Vazel. Thank you.
So cerebrus for those members of your audience that aren't familiar is an AI computer systems company.
We observed nearly 10 years ago that AI had transformative potential,
not just for things like cat videos and consumer applications,
but also for fundamental science, industry, society.
So we saw this big potential, and we can go into this more a little bit later.
What we saw limiting that potential was the speed of computation.
Took too long to build models.
Once you had a capable model, took too long to deliver answers.
So it's cerebrus. We developed a new chip and a new system that was specifically designed for artificial intelligence compute.
Fast forward to today.
Those systems are delivering AI inference and training more than 10X faster than like a C general purpose processors like GPUs.
And enterprises are starting to come to us to work with us because AI has moved from being a curiosity
to being existential, and once a capability like this is existential, everybody wants it to be fast.
And that's where we live in the market. That's what we do.
In my role, I'm our chief strategy officer.
I started with the company nearly 10 years ago and actually started as our first head of product.
So working directly with our engineering teams, our founders, our customers to help define what we should build
in terms of the hardware where it sits in the data center, the software stack that meets users, where they are.
And work with the team to build towards those market and customer requirements, a computing solution, hardware and software stack
that would deliver on that promise of accelerated compute for AI.
These days, under the strategy banner, my charter includes product strategies, so thinking about where we want to build into the future.
But also many of our largest strategic customer engagements.
So think US government, think large commercial and international like G42 and the UAE
are recently announced partnerships with OpenAI and AWS, the kinds of big partnerships
and big technology moves that are going to carry us from our first decade into our second NBI.
>> So when Cerebra started, I'm assuming that was before LLM's, pre-LLM.
>> Pre-LLM.
>> So what were you building towards back then?
What were you thinking was going to be your customer or who?
>> I love that question.
I mean, I think for many of your viewers, your listeners, this is going to take us back in time.
So let's put ourselves in 2017, 2018.
The canonical AI problem of that era was training convolutional neural nets, like ResNet
for image classification and object detection problems.
So image net, right?
Find a cat.
>> Hot dog, no hot dog.
>> Hot dog, 100% hot dog, no hot dog.
So that was the problem.
So it was training, cognets for image problems.
And at that time, that was the canonical problem of the day.
And so we thought about that as an archetypal problem to inform our hardware and software
development.
But we also knew that we didn't want to be limited about that because by that rather.
Because in some sense, we had a hunch, a hypothesis that AI was going to be transformative,
as I mentioned before.
But we didn't know exactly how.
And so we'll get into the chip and what makes it special, I think a little bit later.
But long story short, we didn't design the chip at the hardware systems specifically
for training or for inference.
And we didn't design it specifically for convolutional neural networks or transformer
networks.
We designed our chip at the hardware system to be a general purpose accelerator for all
of AI compute.
But I don't think we really articulated it this way at the time.
But I think what that bought us as we then moved from 2017 to 2026 in present day, what
that bought us is durability and flexibility to accommodate large language models, inference
as well as training.
And be able to accelerate all those workloads by building a general purpose AI accelerator
built around the first principles of all AI compute, rather than building a truly sort
of problem specific chip that would just be good at one thing.
Yeah.
So, when, so I guess was GPT-2 on your radar back when it came out, I'm assuming so.
And if so, did you have to change anything about the chips you guys were building for?
Hey, like, LLM's are going to be huge over the next few years.
So yeah.
Great question.
I mean, this is the fun part about this history, I think, when we're building a fundamental
technology that is a new technology that's sort of at the base of the stack, I think you
have to acknowledge, particularly in the hardware world, that it's going to take some time.
And GPT-2, yes, when it came out, it was definitely on our radar, but I think is you just
articulated the question, I'm not sure if we really knew at that time that LLM's were
going to be what they are today, right?
We knew that it was a really compelling answer to, I think, the fundamental value question
of AI.
Because that was the moment in AI where AI went from being sort of an interesting curiosity
to being demonstrably valuable, right, was that moment, that sort of chat GPT moment.
But we still didn't know that it was going to, that that architecture was going to endure,
and that it was going to be as large in the market as it is today.
So I'd love to say we did, but I don't think our crystal ball is particularly better
than anybody else's.
I think we just built a technology intentionally that could be durable and flexible to changes
like we've seen in the field over the past decade.
Okay, so then to your question, did we have to change anything?
The short answer is yes, but not in the hardware itself, so the chip, the heart of our system
has remained the same throughout, throughout all the way, from, you know, training convenettes
for image problems to inferencing language models for, you know, coding an agentic,
a generative AI.
The chip at the heart of the machine has stayed the same, generation over generation has
been improvements, but the fundamental architecture has stayed the same.
At the software level though, that's where we were able to pivot and respond to the market.
So our original systems were, were delivered with a software stack that was optimized
for training, and that had a lower level library of compute primitives oriented towards
convenettes, right?
That makes sense.
That was the market of the day.
That's what we were delivering and building for.
At that time, we were also integrated with the ML framework TensorFlow to, so go back
in time, shout out to Google and the TensorFlow team.
Before PyTorch really came into the market, and we came to DeFacto ML framework for,
for ML model programming.
As the market shifted then from training and convenettes to, to language models, still
training, language models then got big, right?
We might remember the transition from, from, from Bert to larger language models, right?
Language models started to get beyond, you know, one or 10 billion parameters in the hundred
billion parameter territory.
And Bert was like 200 million or something?
If I recall correctly, yeah, something like that.
So it's still modest, by, by today's standards, right, right, now, yeah, with a lot bigger.
There's still a lot that can be done with those, those, I'll call them medium-sized models,
and smaller.
But yeah, once training became, went from images, images, sorry, convolutional models to
transformers and language models, and then from, I'll call them small to medium language
models to large language models.
We had to update our software library, and then we actually had to change the way software
used our system to train those much larger.
models. So we basically did a significant revision on our compiler and software library to be
able to train those very, very large models, not because they were language models particularly,
but because they were large. And then when, you know, in that post-GPT moment, when the market
became interested in inference, the same thing happened again. So we recognized that there was
an opportunity to accelerate inference, but our software needed to use this engine a little bit
differently than it had for large training problems. I see. And so part of the machine stays the same.
The software around it and how we present the problem or how we compile the problem
to that engine has changed. Yeah. Could you maybe now give us a one-on-one lesson on the chip?
And just like, how does it work? How did, you know, how did GPUs work? How do they differ?
That's what I think. Sure. That's a long question. Okay. You reign me in.
And I'll start. So at the beginning recall that I mentioned, we weren't building this device
specifically for training or inference or calmness or transformers, but that we took a step back
and looked at the first principles of AI compute. So if you ask yourself them, what are
those first principles of AI compute? Obviously, you need a lot of computing operations. You need a lot
of flops. But not just any flops, AI compute tends to want lower precision, like 32, 16, maybe
eight, four-bit. Also, it needs the ability to handle sparse operations very efficiently,
because there's a lot of zero-value data in the input data and the model weights of AI problems.
So you need a lot of compute, but you want it to be optimized for sparse linear algebra
and relatively low precision. Cool. Lots of compute check. The second element that you need is
fast communication between computing elements, because AI is fundamentally a problem of handling
data in motion. So in the analogy of a multi-layer neural network, the computing element that's
computing on one layer needs to be able to quickly share the output to the next layer in order
for the whole thing to go fast. So computing elements need to talk to each other at high bandwidth,
so you need high bandwidth communication. Check. Last thing is you also need high bandwidth memory.
And I don't mean that in the capital HBM sense, but you need memory that can be accessed very,
very quickly by those compute elements, because you want your compute to be able to quickly pull
the stored weights of the model and then combine that with inbound activations,
push it out to the next computing element in the chain, whether it's training, or inference,
condonets, or transformers, all those things hold true. Okay. So then you say I want a computer
that has a lot of sparse linear algebra flops, high communication bandwidth and high memory bandwidth,
and you look out at the computing landscape of 2015, 2016, 2017, and you have x86 machines,
CPUs, and GPUs, the predominant compute engines. And they have a lot of flops,
so they can bring the compute to bear. Maybe they have more high precision or whatever,
but like they can bring the flops, but their challenge, those architectures are challenged by
the latter two attributes that we talked about, high communication bandwidth. They both are
what are known as traditional von Neumann architectures, so they have a compute substrate,
a logic die, and then some memory. So GPUs, for example, use capital HBM, HBM memory, DRAM,
off the chip. So those are the memory is separated. And then if you want your processors to talk to
each other, then you're wiring your processors back together in a cluster, right, over copper,
through a network interconnect, over fiber, across the aisle, and a data center to like some other
machine, through another router rate. So the communication memory bandwidth is really a challenge.
So let me say, well, look, what if we just put as much compute as we could together on a chip,
on a big chip, and then allowed the compute elements to talk to each other over silicon,
so we wouldn't have to wire a bunch of small chips together. And we can, if we have a big enough
chip, we can also put a bunch of memory right on the chip, and then we don't have to reach
out to some external or separate source of memory for memory bandwidth intensive operations.
We said, okay, so that would, how big of a chip can we build? And then you learn that the way
every chip starts is a 300 millimeter circular disk of silicon, and use photo lithography
to basically snapshot out or stamp out circuitry on that chip. And then traditionally what happens
after you have that full disk printed out is you then slice it up into lots of small chips,
and then everything we talked about before happens that is people spend their entire careers
and industries figuring out how to stitch those chips back together and get them to all work on
big tasks. So we said, well, what if we just didn't do that last part? That would give us all the ingredients that we just talked about,
and it would be a fundamentally better design for the first principles of AI work and allow us to
achieve significant acceleration on this class of work that we think is going to be so important
for the future. So that chip, it's just a bunch of transistors, but where is the memory?
Yeah. Okay, so let's talk about the chip. So for your viewers, this is the chip,
this is the way for scale engine. I think you've seen it, but you can take a look.
This particular one happens to be our third generation design, but the first and the second
were also the same size. The fourth is also going to be the same size. This is the biggest square
you can cut out of those primary 300 millimeter circular wafers of silicon.
Also random, but why a square? Yeah, so okay, so why a square? It turns out
A, it's a lot easier to physically package. That is once you build this thing, then you have to figure
out how to deliver power to it, how to deliver data to it, and how to keep it cool. And it turns
out that a square is a lot easier to package. The second reason actually is software reason,
because it's a lot easier to tell a compiler how to compile a workload like a neural network
to a square array, rather than an array that has all these sort of tangential parts to it,
and that's non-U. I see, I see. Yeah, so we conceived of this design, our original CTO,
Gary Lauderbach, our other co-founders Michael James, Sean Lee, JP Fricker, and Andrew Feldman
conceived of this design. They looked at other solutions like a bunch of small chips
on a common interposer. We looked at alternatives because we knew this was going to be really hard,
but kept coming back to this because of its performance attributes on the work.
And then we worked very, very closely with TSMC, who is our fabrication partner, Taiwan,
to basically develop the fab technologies and the methodology needed to, instead of dicing this
up into a bunch of small chips, keep it as one. And so this is then the wafer scale engine three,
and in today's systems, it has nearly one million cores on one device of 900,000 in our production
machines. They're in a 2D grid across the entire surface of the wafer, and then all those cores
are directly connected to each other, and across the entire wafer over silicon. And each core has
its own local bank of SRAM memory. So the memory is on the chip, too, back to your question.
So we're coming back to those first principles, lots of sparsely in your algebra compute,
high communication bandwidth, high memory bandwidth. This design checks off all those boxes,
and gives us what you might think of as a cluster worth of AI compute on a single device,
and literally three orders of magnitude, more communication bandwidth and memory bandwidth
than is possible with traditional chip designs. And I'm sure we'll come back to it later,
but it's that combination of resources, particularly the memory bandwidth that allows us to just scream
on today's agentic inference workloads. Yeah, so it seems like this chip is better and basically
like every dimension. So why do we still use GPUs? You know, it's a great question. I think
there's two factors here. One is that different parts of the AI ecosystem and problem landscape
need different engines under the hood. Like what people have done with GPUs and industry to date
has been incredible. And there's still lots of problems that might not require, you know,
blistering feed speed instant answers. I happen to think that the market is going into a direction
where blistering speed and instant answers are going to be the default, right?
And we can come back to that as well. But I think the future of AI compute is heterogeneous,
right? There's going to be different kinds of compute engines that answer different problems,
just like the compute industry today. But I do think the market needs a different and faster solution
for AI training and inference. And that's what we've built here. I think that the second thing,
honestly, that was a challenge and opportunity for us in the beginning was building a software stack
to meet users where there were, right? This is a very different kind of device. We couldn't use
existing compilers, existing software libraries. We had to write our team
how to write the software to make this device easy to program and use.
And as we fast forward today, it is, right?
We meet users exactly where they are.
For training problems, you can program the machine in standard PyTorch.
And for inference problems, actually, the programming challenge gets even easier
because for most inference problems, what application developers want is just an API
to some inference model under the hood.
And so for inference, we just surface an open AI standard API for inference on our machines.
So what was once a sort of a challenge for us to build and something that made adoption
trickier has all but dissolved.
And today, it's easy for customers to do either training or inference on the machine
with standard framework.
Yeah.
You mentioned earlier that maybe a few years ago, you were more focused on training.
When did you have to make that transition to, oh, like, inference is the more important
problem now?
Is that the case, would you say, first?
I think, yeah, yeah, yeah.
More important is interesting.
What I would say is inferences are focused and inference is the primary area of growth
in the market right now.
But I think actually one of the under told stories is the remaining importance of training.
The training market is still really growing, but inference is taking off extremely rapidly
because that's where AI delivers value, right?
So as soon as AI is valuable for a problem, then inference just takes off.
Just because the demand is insane.
Yeah.
I mean, the way I think about it in normal terms is training is a cost center.
That's where you build your model.
That's where you teach it how to do what you want it to do.
This is a value center, right?
Once the model does the thing you need it to do or the thing that's valuable for you
or your end users, inference is where you ring the cash register, right?
And so that actually goes back to your question.
When did we make that shift?
We made that shift to start building the optimized software stack and our cloud services
for inference when inference started to become the predominant problem in the field.
So that was probably about 18 to 24 months ago.
Yeah.
And would you say most of that usage right now is coding?
Yeah.
Most of the usage right now is coding.
So I think the other areas where we get a lot of usage are voice models and reasoning
models for various chat and assistant applications.
And timely that you're here today, we actually just launched today a private preview of Google's
Gemma 4 deep-mind model, which is a multi-modal model.
No, consumers.
Yeah.
Yeah.
And it's, so private preview today and we'll come to this later too, but like all the
models that we bring up, it is now the fastest frontier multi-modal model on the market.
We're serving Gemma 4 at 1,500 tokens per second.
It's 10 times faster than any GPU implementation and 15 times faster than something like a
cloud's high-coup model.
So yes, I think where there's a massive value center for a lot of our customers today
is fast coding agents where speed equals software engineer productivity.
But I think that will remain very likely into the future.
But I also think in the future, things like agent applications for other workloads, multi-modal
workloads, AI for science and security applications, right?
These are all sort of just over the horizon and I think they're going to be really exciting
in areas for fast inference in the future.
If you had to make a prediction on the next couple of years where the share of usage
is going to shift, do you think that coding is still going to be the biggest share in,
I don't know, three years, or do you think life sciences is going to be that bigger share,
voice, whatever?
So you're saying I do have to make it per day.
No, look, I'm a recovering physicist.
I think it's a reasonable hypothesis.
So much of the large enterprise and consumer industry that we touch every day is digital.
I think it is very likely that that will remain true.
And if that remains true, then how will we continue to update that, make it better,
develop new applications on top of that is going to be coding under the hood.
And so I think that's a durable application.
Whether it remains the predominant application in three or five years, I just can't wait
to see what people will build.
So I don't want to rule anything out.
Yeah, yeah.
You know?
Okay, cool.
I want to shift the conversation a little bit to everyone saying, oh, there's like an
AI in for bubble going on right now.
Do you agree with that?
Disagree with that?
You know, all the big hyperscalers at Google or whatever, like in, I guess combined, they
want to spend trillions of dollars over the next few years on AI infrastructure.
So how do you reconcile those two things that on the one hand, everyone saying that, oh,
we're in a bubble.
On the other hand, these huge companies who are supposed to be the most sophisticated buyers
in the world are spending a lot of money on this stuff.
So yeah, like, how do you think about that?
Look, I think people are investing behind a massive, demonstrable value.
That's what I really think.
Is it big and how does those investments sort of come online quickly and grow and really
rapidly?
Yeah.
Will there be a settling in the market?
And I don't necessarily mean like a correction, but will there be a settling out or a reconciliation
where some investments pay off and others have to be redirected?
Yes.
That's how markets work.
But I actually, I don't think we're in a bubble because when I think about a bubble,
not an economist, when I think about a bubble, what I really think about is a value that's
inflated beyond the inherent value of the underlying technology product services assets.
And I don't think we're there with AI.
I don't think we've really discovered the full value potential of AI.
And so I think these really sophisticated organizations that you mentioned, some
of whom are our partners, I think they're investing into and ahead of a legitimate value
hypothesis.
And a lot of what I heard is we're basically you're not able to build data centers fast enough
to service all the demand.
So is that a problem?
Yeah, it's a problem.
It's a great and interesting problem.
So I think coming back to cerebris and what we're doing here and my own view on the market,
there is a fundamental mismatch between the timescale of development for hardware solutions
and physical infrastructure and the software applications that people build on top of it.
So what we're seeing now is this massive acceleration and uptake of AI software applications
services that are fundamentally built on top of this physical infrastructure of metal racks,
power lines, fiber cables, big boxes with wafer scale chips inside.
It takes longer in real clock time to build these physical systems and poor concrete and
raise up data centers.
There's almost no way that you can build that physical infrastructure as quickly as it's
being adopted when you're in a moment like this.
And so yeah, I think it's a problem, but I think it's a really, really interesting one.
At cerebris, I think we did get ahead of it a little bit, but recall that we started
this project almost 10 years ago, seeing that we would need to invest for years in order
to build the right instrument for the future of this work.
We knew we had to be ahead of it because we thought it was going to be big and important.
And it was different enough to be valuable.
We also knew it was going to take time.
And to be candid, I think we as an industry probably didn't have that same lead time opportunity
with inference in the post chat GPT moment to get a head start, pouring concrete and building
data centers and developing like thoughtful energy strategies for global infrastructure.
So I think as an industry, I think we're playing a little bit of catch up right now.
But I'm optimistic on what that means that our sort of community of innovators around
us are going to create, right?
What that challenge puts pressure on creative people and investors and we're starting to
see people develop, for example, novel energy solutions, storage and generation.
We're starting to see people come up with clever approaches to reuse physical facilities
and finance data center expansions, right?
So I think it's a problem, yes.
But I think it's a challenge that is going to be really exciting for us as a community
to solve and smiling because I think if we solve it, then we get together to see this
potential of this technology have some really, I think, positive impacts on the world.
Yeah.
I guess one of those bottlenecks is data.
center development. The other one is memory. I think a lot of people are, you know,
seeing all these memory stocks just blow up. Why is that happening? Like, why is there a shortage of
memory? Yeah, could you explain that? Sure. I mean, so back to the fundamental architecture
of a traditional chip is that logic die and then the memory die. And what we're really talking about
here is that type of memory, which is used by GPUs called HBM memory. That memory is, you know,
subject to fabrication limits, manufacturing limits, similar to what we talked about before,
right? There's physical facilities, machines that need to be built in order to turn up
capacity. You can't just turn a knob and produce 10x more. And so that memory is a bottleneck
right now for those types of chips. For us, it's not, right? We have all of our memory on the chip,
as right here. It's not part of that supply chain. So we don't have the same challenges
being subject to HBM supply chain limitations as other chips do. So why, so this is called SRAM?
Yeah. Yeah. So why is there not a shortage of SRAM, but there's a shortage of DRAM?
It's really about the fabs and manufacturing facilities, right? So for a traditional chip,
it goes through, say, the same fab that we do at TSMC or Samsung. And it's one process,
right? Building memory chips uses a different fab, a different manufacturing facility,
different supply chain, because we can leverage the traditional chip process. We don't,
we aren't subject to those. We don't need the other manufacturing facilities for our systems.
So I guess that's a, that's really good for you guys, because if the memory,
memory costs are going up 5, 10x, then all the GPUs are also going to have to go up in price.
I'm assuming, but you will not have to. I think that's, yeah, I think that's, that's about right,
right? When we, so when we build big clusters, we have other systems in the clusters that do use
that memory. So it's a, it's a part of our solution, but we, we're not nearly as reliant upon it.
So I would say that that's, it's, it's an opportunity for us to be able to deliver sooner for
customers and to be able to deliver ultra high performance, leading performance in the market at
an even more competitive price point. So the other bottleneck in data-centered development,
everyone talks about is energy. Do you agree that that's a bottleneck or why, why is that a bottleneck,
I guess? bottleneck is an interesting word. I think, yeah, it is a constraint on the rate of
development. And I think it's also an important constraint to be mindful of when we think about
environmental impact, impact on the communities. Is it a, is it a fundamental bottleneck that is,
you know, is there enough energy on the planet? I don't think we're quite there yet.
But I do think as a community, as an industry, we need to think about not just how to make AI
bigger, faster, and more capable and more valuable. We also have to think about how to make AI
more efficient. And that, by, by efficient, I mean more efficient at learning, right? Using
maybe fewer weights to learn the same thing, or using fewer compute flops or watts to, to learn
the same thing, or less data. So more efficient learners. But I also think we need to make AI that is
more efficient in terms of execution and output. So we need more efficient chips for that,
and we need more efficient ways to deliver energy at scale. Whether that's through updates to the
grid, or improvements to energy storage in communities and at large data center facilities,
or if it means changes in how we generate, whether that's things like small modular
nuclear reactors, or dare I say at data centers in space, where you have all the energy of the
Sun pouring down on solar panels. Right? So I think that in that world, right, we're very much
a part of the solution, because, but by building this big chip, not only are we faster,
but we can also save energy on things like communication of data and memory access,
because we're not accessing memory from a different chip. We're not pushing data between two
different chips in a cluster meters apart. We're accessing data from our own SRAM on the same chip
and pushing it microns and nanoseconds from core to core. So we're a part of this solution,
I think, but as a community, and I think this is what a lot of folks that aren't in this industry
are probably thinking about, we also need to think about how to deliver energy more efficiently,
make this whole process more energy efficient, and really think of it holistically as a problem.
So I do think it's a bottleneck, but not in a way that it's sort of a fundamental limitation,
but in the way where we need to figure out how to build more efficiently and deliver this capability,
more efficiently, particularly per kilowatt-kilowatt hour.
Yeah, so I want to shift the conversation a little bit to, like, you building your own data centers.
Are you guys building your own data centers? Or, yeah, why is that? Why is the chip company making
their own data centers? Yeah, so to be clear, I often joke that working on this at Cerebrus
on the wafer scale engine in our systems, I never thought I'd learn so much about plumbing.
Our systems are water-cooled, so when we deploy to a data center, we take cool water,
we push it across the back of our waifers inside the machines, and that allows us to keep the chip
cool. As a result, we have a bunch of pipes running through floors and ceilings,
and we have to think about how to keep the water cool. Sometimes things leak, so there's questions
about plumbing, and that turns into conversations about things like joint, you know, connecting pipes
together in wrenches. It's a real physical infrastructure problem. I think that, in that sense,
we're not building our own data centers to be clear, right? We're not actually pouring concrete
from, like, trucks with Cerebrus logos, and, you know, setting up facilities that way. What we're
doing is we're contracting with folks that either own facilities already or are developing facilities,
and we specify to them a particular layout of the floor, power requirements, inbound data
requirements, cooling requirements, and then we work with them to basically contract and
unleash that facility. I think that's your question, but I wanted to clarify that for the audience,
because I think when you think, like, building your own data centers, like, man, one of these guys
getting into, right? I mean, like, at the end of it, do you own it, or does a different company own it?
Are you kind of overseeing the development? It really depends on the customer,
but typically we don't own it. We're leasing the facility, right? And renting the power.
There are cases where we have our own ownership, or we're operating a facility on behalf of a
customer who might own the facility. But often what we're doing is we're leasing the facility.
In that context, right, we chose to effectively roll our own data centers and build our own cloud,
fundamentally, because we had so much demand for what we were building that is serving fast
inference that we needed to sort of own our own destiny in rolling out data centers.
And we also needed a particular, you know, for each compute provider, this is true, but we
wanted a particular layout of power, network, cooling, and racks. And that allowed us to go quickly
in response to the market. Okay. I want to shift the conversation back to the chips, actually.
But more, I guess, on the economic side. So how do you think the cost? Well, first of all,
what are the costs involved in building a chip? And then how do you think those costs are going
to evolve over the next five years? This is another good one. So I'll try to keep it at a high level,
because each one of these questions are great. Each question could be sort of a one-on-one course.
But I think the costs of developing a system, a chip in system, are fundamentally
the design, the engineering development, the prototyping, and the manufacturing at scale.
And particularly the last one, the manufacturing at scale also involves the inbound supply
of all of your sub components, right? So just to give you a sense of the sort of the timelines
involved in this, company started in 2016. I started in 2017. We at that time had designs for
the first wafer scale engine. And if I recall correctly, we launched the first version of our
cerebra system CS1 based on that first wafer scale engine in 2019. So you can imagine yourself,
it took about two or three years to get from a design into a workable device.
And then it took another few years, probably, between our first generation system CS1 and CS2,
really to iron out most of the system and software challenges to make those units really
reliable at data centers.
standard and production scale.
So it takes time and it has all those different cost centers.
It's not a cheap project by any stretch of the imagination.
Now back to your question, where do those costs go in the future?
One of the areas I'm most excited about is AI-accelerated,
AI-augmented chip design.
There's some really interesting work building foundation AI models
to enable faster chip design, or maybe even use AI to design chips
that compute AI that can then develop new chips.
So that area of chip design, I think,
is one that is, in some sense, ripe for acceleration
and potentially ripe for cost or barrier to entry reduction
in the future of computing.
I think once you have a chip design, even if it's faster
and cheaper than it was before, you still have to figure out
how to package that device, that is bring power, data,
cooling, put it on a board that integrates into a system
that can go into a server or a data center rack.
So you still have to figure out the physical packaging
and the prototyping.
I think those things will get faster,
but probably not at a sort of order of magnitude kind of X-factor.
Maybe we get 2X-factor.
Maybe there are tools or equipment costs
or learnings from industry that drive a cost down
by a factor of one and a half or 2X, right?
Just hit pocket numbers.
Yeah.
Then the manufacturing at scale, the last piece.
I think it's going to get really interesting.
I think some of the fundamental supplies,
like you were talking about, memory.
There's a lot of conversations about rare earths
and materials that go into these systems.
The supply chain that goes into manufacturing at scale.
I don't think any of that gets cheaper, right?
Or I would be surprised if it does.
I think the entire history of human species
in industry would suggest that as interests,
as demand grows, right?
And supply sort of stays the same or gets smaller
unless we can find new sources, right?
That's going to get more expensive.
Then the manufacturing that is the actual chip fabrication,
system manufacturing.
I think that's an area that's also right
for more efficiency and cost reduction, right?
We're seeing teams here in the Bay Area
build up prototype manufacturing lines
and build scaled production manufacturing lines
far faster than they could five years ago even
and use more automation and more data
to drive not just manufacturing velocity,
but drive costs down.
So in summary, I think cost of design,
opportunity to decrease,
cost of packaging and prototyping,
met and cost of manufacturing at scale
might also be a wash if we see supply prices go up
and manufacturing prices at scale perhaps go down.
Are there certain rare earths that you see
as like these are the most important in terms of
or maybe these are the ones you're focused more on
because there's some, you know,
some country has 90% of the reserves and things like that.
Yeah, I mean, just to be candid,
not at this moment, no.
But I think we're as an industry,
we're keeping an eye on that.
And we're also excited to find other sources
for those materials.
Yeah, so Google, Google has like a full supply chain
of, you know, they basically own everything
like they own the software,
they own the, they're building their own TPUs,
they're offering it as a service through GCP.
And that makes it probably like very cost efficient
for them to offer it to other companies if they want to.
Do you think they will be the lowest cost producer
of tokens, I guess going forward?
It's an interesting question.
I'm not sure I want to speculate on who will be lowest cost.
But I do think you're fundamental there on,
you know, they have a vertically integrated product
that they can offer at scale.
That is true.
I came from Google, I know that team, outstanding team.
That being said, I think rather than focus on cost,
what we're focused on is performance.
So delivering the fastest tokens per second
at a competitive price and with greater power efficiency.
And I think that back to our conversation
about the future of AI computing, right?
Well, I think TPUs are an extraordinary machine.
I think in the context of the future of AI,
there's a place for a machine like that
and there's a place for an ultra-fast race car machine
like ours.
And to the premise of your question,
one of the reasons why we vertically integrated
is so that we can make the fastest AI computers
on the planet, but do it at scale
and be competitive to those costs.
So I think there's going to be future systems
that deliver, I'll call them medium speed tokens
at a really, really competitive price.
But for applications, for the user applications
that we were talking about before, right?
Voice, reasoning, agents that I think
are really the future of AI computing.
Their speed is paramount, right?
Those applications just don't get built
or aren't feasible if they don't answer quickly enough.
And so yes, there's going to be a future of compute
that delivers medium speed tokens at really low cost.
That's not a market I'm interested in pursuing.
But the market for fast tokens
for the future of AI applications
that are really competitive costs,
that's where I want to go hunt in a minute.
Yeah, so why is speed so important?
Like we know this from the entire history of the internet
and I'll give two examples.
One, we just talked about the company.
Let's talk about Google search, right?
Would you wait for a Google search answer?
One was the last time you saw the little
sort of pinwheel of death on a Mac pop up
from a Google search query.
I think that, and before you say it,
AI summaries, right?
I think early AI summaries made us wait.
And even just maybe six or eight months later
or a year later, we don't want to wait
for AI summaries anymore.
So I think speed matters and actually Google search
and the team there did a lot of research
to show the effect of speed on user satisfaction
and user retention.
And what we know from that work
and by proxy the entire history of the internet
is that every millisecond of latency matters
for these valuable or day-to-day important applications.
Every millisecond that you delay, your users get upset,
you use to walk away, they go to your competitor.
So for any applications that is like that, speed matters
and that's where we're seeing voice reasoning agents going.
The second thing that we know about the history
of the internet is when things get faster.
It's not just that the same group of applications
exist but go faster, we see entirely new industries.
So the example that I have here that is fun is Netflix, right?
When internet speeds went from broadband
or from dial up to broadband, it's not just that Netflix
delivered more DVDs by mail faster.
When speeds got faster, Netflix went from delivering DVDs
to being a digital production studio
and streaming entertainment company.
We created an entire new, effectively industry
and hundreds of billions or trillions of dollars
of revenue and lots of user delight
from people like me that like to sit on the couch
and stream Netflix.
So I think speed matters full stop.
And when we're working with customers today,
like OpenAI, AWS, cognition, notion,
in those areas of coding agents, voice agents,
reasoning models, they know our partners know
that if they can deliver output and answers
to their users faster, their users will be happier,
more creative, more productive, deliver more value
and use the product more.
And honestly, that's why I'm so excited
about where we are in the market
and to see what our customers are going to build
on top of our systems.
Yeah, so on the OpenAI deal, could you talk a little bit
about that?
So is it when I use codex?
And I see that they're offering me 2x faster speeds
for twice the, I forget the cost for it.
But is that 2x faster speed option using cerebrus on the back end?
Yeah, so we're building.
I'm happy to talk a little bit about that.
We have a great partnership with OpenAI.
We got to know their team actually
from the early days of cerebrus.
We saw their ambitious vision
to build state-of-the-art frontier AI models and AI.
And they saw our crazy ambitious vision
to build a computer system that would be the right system
for the future of AI.
And it's great and serendipitous
that we kept track of each other
through those almost 10 years.
And then we're able to reconnect and partner
to deliver fast AI for coding
and agentic workflows now.
So we have a deep code design collaboration
with OpenAI engineering team.
So the public terms of the deal.
they're going to buy 750 megawatts of compute by cerebrus to power their applications.
The coding model that uses cerebrus now is Codex Spark, and we're working with the OpenAI team
to bring up additional models in the future from their GPT, frontier models to next-generation
coding models, all of which should be the fastest examples of those models in industry.
Yeah, and I'm assuming the thing stopping you from working on the bigger models right now
is just the data center issue. You know, it's really a matter of product strategy and
prioritization. Oh, really? Yeah. I mean, certainly more computers will be better.
OpenAI has incredible reach into the market, incredible user base. If we could serve all of them
today with the fastest tokens, I'd like to say that we would. So, yeah, you're not wrong, right? Like,
if I had more computers, it would be better. But we're also working really closely with their
product teams to say, look, here's the library of models and applications that are really
important new users today, which are those where speed is the most valuable or most enabling.
And then let's look at the compute resources that we have together, and let's rack and stack and
then develop a staged approach to bring the right fast models to your users first based on what
we have. It's a fun exercise, actually. It's really cool to work with their team. That's really
cool. Are you planning on working with Anthropic in the future? Yeah. So, right now, I think you just mentioned you have a deal with Anthropic for, was it 750
megawatts of compute? Open AI. Oh, sorry. Open AI. Yeah, yeah. If we're so like a few years ago,
I think talking about 750 megawatts of compute, like that would be insane to talk about. And now,
you know, people talk about gigawatts. And where do you think that goes in five years?
Yeah, I mean, I'm showing my age here, but maybe you and I'm sure some of your viewers and
listeners remember back to the future. There was a joke about 1.21, they call it gigawatts, but
1.21 gigawatts to get the car to go back to the future, right, from a bolt of lighting. And
you know, in that, you know, in that sci-fi movie of my youth, it sounded like a made-up number,
you know, it sounded crazy and funny, but unbelievable in just the right kind of way. And now we're
talking about those kinds of power envelopes in a very realistic way in the next, you know, 12,
18 months for some of these systems. I actually think it's sort of an awe-inspiring
trajectory for the industry that, and a signal of fundamental demand and value that we're talking
about building infrastructure at this scale. The only analogy that I can think of is
through the building of the railways or laying of undersea fiber to enable the internet,
right, this is a generational infrastructure project that's being taken on by a coalition really
of individual members of industry and governments to build this, in some sense, the rail lines
of the AI industry future. What do I think we're going to be talking about two or three years from
now in terms of power? I'd venture, maybe, I'll get some blowback from this, but I'd venture we're
still talking about gigawatts, maybe 10 gigawatts, but I think we start to run into some limitations
of physics. Yeah. And without new generation or significant grid updates that do that does take
some time. I think in a few years we'll probably still talk. So you're also leading U.S. federal
programs. Could you talk about, like, what does that mean? Sure. So at Syribris, the U.S. government
was one of our very first customers. We worked with the Department of Energy,
Argonne National Labs, Lawrence Livermore National Labs, we're now doing a lot of work with
Sandia National Laboratory. We also did some early work through micro-process or R&D with the
Department of Defense and DARPA. So the U.S. government has been a long time customer for us,
but also a technology development partner. And I think for folks that are building new companies
today, I'm happy to share more about this in another form of people want to reach out directly,
but I think that the U.S. government can be a great partner. Many organizations within the
government actually have the charter to be early adopters of new technologies. In our case,
the Department of Energy, who builds world-class supercomputers for scientific simulation,
also has the charter from the U.S. government to be an early adopter of new computing technologies,
so that the government and its research scientists can sort of get early peak at what
industry is building and participate in it, and also inform its development so that it can contribute
back to science and society and public interest. Maybe I'm dorking out a little bit too much over it,
but I think that that's a really, really incredible mission and function that the government plays.
And for people that are building new deep tech or other technology startups,
I would say, look, don't be afraid of the government. Lean into the government, and there may be ways
that they can help you get revenue or get access to technologies or partners that your industry
partners may not be able to. So at Suribis, government has been an early customer and a great
technology development partner. We're currently working with different parts of the U.S. government
on future I/O or interconnect technologies, as well as future memory technologies coming back to
what we talked about before. So we're continuing to do work with the government on those kinds of
projects. We are also a part of the U.S. National AI Initiative, which is called Genesis, which in
summary aims to build AI computing infrastructure to fundamentally accelerate how we do scientific
discovery, everything from biology and life sciences to physics and chemistry and drug design to
space in the environment. It's an incredible mission and actually I think it's something that the U.S.
needs to do in order to keep pulling industry forward and help us as a nation and as an American
industry lead in this technology that in some sense we had a big hand in inventing. In other words,
if we want to continue leading and extend our leadership of AI technology industry in the world,
it's going to take a big initiative from leadership in the country to do it. So we're a part of
the Genesis program and we're also working with the Department of Commerce to figure out how to
take what we're building here in the U.S. with fast AI wafer scale systems and cerebral super
computers and figure out how to package that with other U.S. AI technologies and quickly and
easily export those technologies to friend and ally nations. So that's where our work with the
federal government also intersects with international sovereign initiatives. That is how can we
enable our friends and allies with world-class AI technologies that are built in the U.S.
At least for me, I'm a believer that we get closer to realizing the full and positive potential
of AI if the right AI computers and AI technologies are in the hands of more, not fewer,
organizations and countries. So one of the reasons that I'm at cerebris is to build that right AI
instrument, build that right AI computer, and then put it into as many developers' hands as many
organizations, hands, companies, and countries' hands as I possibly can because if everybody can
build faster, then we're much more likely to discover that next great thing that fundamentally
changes our health, happiness, ability to do work, productivity, then we would be if all these
computers just lived in one place. So the work with the federal government helps us do that and
helps us get access to foreign markets. Cool. What are some of those projects that you're working
on to export to foreign markets if you're able to talk about some of that? Yeah, absolutely. So
there's actually a new, relatively new program led at the U.S. federal level by the Department
of Commerce called the AI Export Text Act. And so that project is basically the U.S. government
providing a framework for industry members to come together, define a common stack or a menu,
if you will, of AI technology solutions, and then the U.S. government will work with industry
to streamline the export policy for those so that industry members can quickly package and
sell and export those technologies to foreign and ally nations. That's the general overarching
goal of that program and where we fit in as a technology provider and partner. And we're also
helping the government define what that stack looks like and what the menu of solutions might look
like. For us, though, that's actually strongly informed by the work that we started four or five
years ago now with our strategic partner in the UAE called G42. There's lots of information
online about our work with G42, but long story short, they came to us about five years ago and we
wrote together an MOU, a memo of understanding, that articulated a vision, that they wanted to build
world-class supercomputers and world-class AI models in Arabic language, in health, in science,
in finance. And the MOU was that we would do this together. Fast forward to just a few weeks ago
when we had our IPO, I was sitting down with the CEO of G42 and we were reminiscing a little bit
that we wrote that MOU and all of it came true. We've built with G42 as a commercial national champion
for the UAE. We have built for G42 world-class supercomputers. We've collaborated on state-of-the-art
model development. We worked with UAE researchers to build the first and still best open Arabic GPT
model so that state-of-the-art AI can speak the language of local people and reflect the cultural
interests and social and commercial interests of that market. We've worked with their healthcare
teams to build state-of-the-art AI healthcare models. That work, I think, was reflected this
ambitious vision of the UAE to be a leader in AI technology and I'm humbled to say that I
think we played a helpful hand role in bringing them on to the world stage and where they are today.
And I think if we can do that for more of our international partners through, for example,
the current US Department of Commerce program, then we're all winning.
Yeah. Is it easier to build data centers in the UAE than the US? Is it faster?
You know, that's a good question. I actually don't know the answer to that.
Let me put it this way. I'm looking forward to seeing how fast they can build data centers.
What's the most surprising thing about working with governments? Is this your first time working with
governments? No, actually, for the past 20 years. I used to work at Google before that. I was building
satellites and before that I was building image processing algorithms and software for
environmental surveys. In all cases, actually, I was working with governments.
But I think what surprises me most here in AI particularly is I think some governments are
really moving fast. And by design, moving fast is not a thing that most governments are good at.
In fact, I think as a citizen, you generally want your government to be deliberative and fair
and thoughtful and take its time. But I think some cases velocity really matters.
And I think emerging as a leader or developing state of the AI, velocity matters right now.
And so I think the surprising thing is seeing some governments really, really lean in.
I think UAE is a great example of that. I've been thrilled to see over the past 12, 24 months
how the US government has put the proverbial pedal of the metal on AI. And I think other nations
around the world are starting the same thing. Yeah. So on that note, should we be selling chips to
China? That's a, as you know, that's a complicated question. We've had the opportunity to work with
China in the past, cell chips to China in the past. And we decided not to. I think there's an
incredible developer market, there's an incredible research market in China that I think as a society,
we would love to see go faster. I think if we can answer some fundamental questions and come to
agreement collectively on things like security, then we absolutely want to be able to work together.
Yeah. I've heard some people say, hey, we should sell to China. We maybe shouldn't sell our best
chips to China. What, like, what do you think about that? This is, so I also work pretty closely
with the part of the Department of Commerce that administers export policy. It's called
the Bureau of Industry Security BIS. Huge shout out to the people that work in that part of
the Department of Commerce, they have what I think of is probably one of the hardest jobs on the
planet because they, at the end of the day, they are charged with balancing US national security
interests with the interests, the interests of US industry and our economy, right? So on one hand,
you might say, open up the floodgates, sell everywhere, it's good for industry, it's going to be a
boom for our economy. On the other end of the spectrum, you might say, we have fundamental
questions about how these chips are going to be used. That could threaten national security,
or we just don't know. So let's just stop for now. Both of those endpoints have challenging
consequences. So finding the right middle ground for any given global geopolitical state of play
on any given day is incredibly difficult. If you hold back too much, you over and incentivize
the development of competitors. If you deliver too much, what security might you put at risk?
It's almost an impossible problem to solve. So I think the nuance of getting that just right
is incredibly difficult. I will not claim to have the answer of what is the right thing to do.
I think, as I said before, if we can answer this, fundamental questions about security,
if we can have the right diplomatic engagement, if we can agree on the markets, we would love
to be able to help the folks that are working on fundamental science, AI, new commercial AI,
new open source AI in those markets go faster. Yeah. So what are your thoughts on Europe? Everyone
says, hey, Europe is so far behind. You know, their policy is, or there's way too much regulation,
way too much taxation. Have you guys worked with Europe? Yeah. And where are you seeing that
headed over the next couple of years? Yeah. In fact, I'm going next week to Germany to a conference
called International Supercomputing, where we're going to be sitting down with European owners
of state-of-the-art government public sector supercomputers, but also owners and developers
of new commercial AI infrastructure and data centers. We have a long relationship, an extraordinary
system and partnership with EPCC, which is the supercomputing center in Edinburgh, Scotland.
So we actually have systems in Europe. We also have systems in Germany. And we are working
to deploy data center for our commercial inference services in France. That's all just context.
Back to your question. I'm trying to encourage European AI infrastructure developers
to think differently and move faster. I do think that there is an awesome opportunity.
I think back to our conversations about our conversation about government. I think European
governments are trying to be really deliberative and make the right choice, just typically in the
interest of the population. But I think right now industry is moving faster. And if they also want
to lead as friends and allies, I want to encourage them to move faster. And so that's one of the
soap boxes that I'm going to be on next week. As a community, let's think a little differently
about AI infrastructure, not just GPUs, speed matters, where speed matters,
their solutions like ours. But also, how can we help you move faster, whether it's through commercial
data centers or public supercomputing project? Cool. I want to shift the conversation. This is
going to be like the last couple of questions. What predictions do you have about where enterprises,
or how enterprises are going to be using LLMs over the next couple of years? Like right now,
maybe coding is a. I mean, I guess we talked about this a little bit. We talked about it a little
bit before the podcast, too, that right now AI is showing this massive potential. And clearly,
there's a value proposition to be harvested in coding. But how we as a community build those
tools to put them in the hands of enterprises so they can use them at scale, that's a really tricky
nut to crack. And so I think even in coding today, we see massive adoption, massive value,
but we still see large enterprises figuring out how to adopt at that scale. How do we integrate
these tools with our workflows? How do we integrate these capabilities with our current teams?
Right. So I think step one is we will see enterprises working with industry to adopt things like
applications like coding and fast coding agents like those on our systems at scale. So step one
is like increase adoption, right? Or absorption of that capability. I think then once that gets in,
what we're starting to see from some of our partners in enterprise is development of new,
I'll call them domain or industry specific models, right? Maybe it's not a language model for
coding or chat, but maybe it's an LLM or a sequence style model for drug development or for
material design or an AI model for physical industry for say manufacturing to say control robots
or analyze data offensive manufacturing.
line. That realm of broadly speaking physical AI and domain-specific models for industry,
I think is a hugely exciting and active area of development that I think is over the next
couple of years we're going to see some big changes. So we have long-time partners like
GlaxoSmithCline that for the past five plus years have actually been building AI models
to accelerate and improve the development of new therapeutics. We have customers like the National
Labs that are working on fundamental AI models for the physical and life sciences.
I think that area beyond sort of beyond coding and chat bots, assistant style capabilities
and consumer agent models, I think that area of physical AI is going to be big over the next
couple of years. Another tangential question off that. Why do you think you've probably read those
stories that, hey, like 95% of enterprise adoption projects for AI are just failing? Why do you think
that is? In some sense, I think they measure the wrong thing at the wrong time, right? I think
those studies, it's important that we measure sort of value return, but I think it's far too early
to ask that question really at scale, or not really to ask the question at scale, we can ask
any time. It's too early to come to that conclusion, right? The way that enterprises typically trial
and then adopt and scale technologies takes 12, 18, 24 months or longer for some big enterprises,
especially if that enterprises the government, right? You start a pilot. You run maybe a series
of pilots. It's intended and known that maybe eight out of those 10 pilots will not show results,
but then you have some subset of those pilots that show positive results. You expand the trial
audience, you start to scale it up only after you scale it up, and then maybe after a year of
it operating at scale, do you actually start to see returns, right? That's how the life cycle and
time scale of enterprise adoption and new technology works. And like I said, I do think it's important
to ask the question of, are we getting value out? But I think it's a little too early to come to
that conclusion. Yeah, I mean, to me, I definitely see a lot of people getting value, especially startups,
are definitely seeing that they're creating value. I think the problem with enterprises is maybe
they're a little slower to adopt than like startups. So yeah, and exactly what you're saying, right?
You have to run 10 experiments. Two of them are going to work out and then you just scale it.
And you know, it's funny, right? I think I like where you're going with that and I was chuckling
about the problem with enterprise. I happen to like working with big enterprises because
they have incredible reach, right? But to your point, they also tend to be more deliberate and move
a little slower. There's more people involved, right? And there's intentional controls in place.
So good. It does take time. So good. I think you're right, though, to bring up startups, right?
For a green field problem, for a brand new team, we're seeing people build from scratch with AI.
And I think those are going to be some really, really compelling disruptors is those teams that are
AI native from zero. Yeah, I've been talking a lot of engineers at some startups and it's kind of
crazy to me that, you know, all of them are telling me, I mean, these are like really, really
amazing engineers. And they're like, I haven't looked at a line of code in seven months, eight months
because they just use agents for everything. Yeah. It's pretty crazy. No, I mean, look, we do
internally at cerebrus as well over the past six months where we had probably a half a dozen
different programs across our entire engineering teams from design in hardware to program management
and product management to software engineering and AI. All of them now are effectively being AI
native for all the code design. And I got to say that's really reinforcing to us here too, right?
Because back to the OG hypothesis, if we can build the right instrument for this work, we think
this work could be really important and transformative. And now we see this, that same workload
literally transforming how we're doing what we're doing. And so it's awesome and encouraging to see
that sort of substantiation even in our day-to-day work. Yeah. What are some projects that you are
currently working on that you're really excited about that you want to share? I think, I mean,
everything is a cheater answer, but I come in to this company every day and I am just at awe
and humbled by the ambition and curiosity, hustle and resilience of everyone of our staff, right?
Not just our engineers, our program managers, people that are working at the office. It's awesome
and everybody is incentivized in some sense to be a part of building this and changing industry.
And there's no task too small and there's no vision that's sort of too big. And so across the
board, I'm excited from model research to software development and systems, but I think what came
to mind immediately is it's also really cool to work at a hardware company because for those of us
that spend most of our time in front of a screen or working with documents or PowerPoint or
God forbid just spreadsheets, being able to take a break and like walk into the back and see people
turning bolts on things and see the things that you're creating in a digital realm actually take
shape in the physical world. That's, it's an incredible opportunity. So I think to your question,
I'm really excited about our future hardware actually. At the end of the day, we are a hardware
company. And so we are on CS3 now. You can imagine that there is a 4 and a 5 and a more beyond that.
I think what we've seen from the market is that there's extraordinary value and fast inference,
fast training in some sense. I think we've proven out the value and burned down the risk of the
the wafer scale processor. I'm really excited to see what we build in our next generations
to go faster. As I mentioned, we've got some fundamental investments going right now in IL
and memory technologies and future systems. So in some sense, I think the wafer scale engine
of generations one through three where we are now is been incredible, but it's just the beginning.
And so for for the audience out there, I would say stay tuned because we've got some really,
really exciting stuff in the hardware pipeline. And I think that could transform computing very much
in the in the same way that the wafer scale engine did in the first place. Awesome. That's all I
had. Thanks for coming. It was a lot of fun, man. Thank you so much. I hope we can continue
a conversation. We should check in on those future predictions sometime and see what came true
and see what was a puff of smoke. Yeah, awesome. Thank you. Is there any topic that you wanted to
cover that you know, I didn't cover? I thought it was great. Okay. No,
I mean, I'm sure I'll think of something later, but I just enjoyed the conversation. I hope
it I hope it it sticks with your audience. Yeah. And I really appreciate it. Yeah, I learned a lot. Awesome.
Podcast Summary
Key Points:
Cerebrus has built a wafer-scale AI chip with nearly one million cores, designed for high-speed, sparse linear algebra and efficient memory bandwidth.
The chip’s architecture enables up to 10X faster AI training and inference than traditional GPUs by integrating compute and memory on a single device.
The company’s software stack is adaptable, evolving from training-focused to inference-optimized, especially for fast coding and multi-modal models like OpenAI’s Codex Spark.
Cerebrus is not in an AI bubble, as massive investments reflect real demand and the full value potential of AI remains undiscovered.
Rising memory costs benefit Cerebrus because its systems use on-chip SRAM, avoiding reliance on HBM memory supply chains.
Inference is now the primary market driver, where speed directly impacts user satisfaction and application viability.
The company has secured a 750 MW compute contract with OpenAI and is expanding partnerships with the UAE (G42) and U.S. government on AI infrastructure and export initiatives.
Future growth will depend on solving physical infrastructure bottlenecks, energy efficiency, and global scalability through vertical integration and innovative design.
Summary:
Cerebrus is a leading AI computing company that has developed a revolutionary wafer-scale chip with nearly one million cores, designed to deliver unprecedented speed in AI training and inference. Unlike traditional GPUs, its architecture integrates compute and memory on a single device, enabling faster, more efficient processing by eliminating bottlenecks in communication and memory access. The chip’s design is flexible and scalable, evolving from early training-focused systems to optimized solutions for high-speed inference—particularly in coding, voice, and reasoning applications.
Cerebrus argues the current surge in AI investment is not a bubble but a reflection of genuine, growing demand and untapped value in AI technology. The company’s systems are especially resilient to memory supply constraints, as they use on-chip SRAM instead of HBM memory. A key market shift is toward inference, where speed directly impacts user engagement and product value—making fast, responsive AI essential.
Cerebrus has secured major partnerships with OpenAI, powering models like Codex Spark, and with G42 in the UAE to build AI models in Arabic and other local domains. S. government on national AI initiatives, including the Genesis program, to advance scientific discovery and enable technology exports to allies.
While data center construction remains a challenge due to physical and logistical constraints, Cerebrus is pioneering solutions through efficient design, energy optimization, and strategic infrastructure partnerships. The company believes that faster, more accessible AI computing will unlock transformative applications across industries, from health and science to global digital economies.
FAQs
No, the AI industry is not in a bubble. The massive investments from major hyperscalers reflect a strong belief in AI's demonstrable value. The full potential of AI has not yet been realized, and these investments are aligned with a legitimate value hypothesis.
Cerebrus's chip is a wafer-scale engine with nearly one million cores, designed for high sparse linear algebra compute, fast communication, and high memory bandwidth. Unlike traditional chips, it integrates memory directly on the chip, eliminating bottlenecks from external memory access.
The chip delivers up to 10 times faster AI inference and training compared to general-purpose processors like GPUs. This is due to its optimized architecture for sparse operations, high communication and memory bandwidth, and reduced latency across cores.
Traditional memory (like HBM) faces supply chain constraints and manufacturing bottlenecks. Cerebrus avoids these by integrating memory directly on the chip, making its systems less dependent on volatile memory prices and enabling cost-competitive performance.
The primary use cases are fast coding agents, voice models, and reasoning models for AI assistants. Cerebrus’s systems are now the fastest for multi-modal models like Google’s Gemma 4, delivering 1,500 tokens per second — 10x faster than GPU implementations.
Speed directly impacts user satisfaction and retention. For applications like search, coding, or voice assistants, every millisecond of delay reduces user engagement. Faster AI enables new applications, such as real-time agent workflows, that were previously infeasible.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.