Neil Mova, founder of Sail Research, is building a "token factory" that provides the cheapest possible inference for open-source language models, specifically targeting background agents that run for extended periods rather than real-time chatbots. His core thesis is that the future of AI lies in long-horizon tasks—like deep research, cybersecurity, and proactive personal assistants—where latency matters less than cost, and agents operate on human timescales without needing constant human input. This shift from latency-optimized to throughput-optimized systems exploits a fundamental GPU trade-off: batching work increases efficiency and lowers costs, but sacrifices response speed. Drawing on his Nvidia experience, where he learned to chase "speed of light" performance, Mova's strategy involves building software that maximizes efficiency on any chip, including undervalued options like AMD, and "scavenging" resources others overlook—from less-desired hardware to intermittent renewable power sources. He embraces distributed, small-scale data centers with lower uptime guarantees (e.g., 95%) because background agents can tolerate failures and reroute work, enabling access to cheaper power and land that larger players ignore. Mova predicts a future where 90% of inference is background-oriented, unlocking abundant, low-cost intelligence for verifiable problems, while leaving creative tasks to humans. He also questions the durability of frontier labs' premium for being months ahead, seeing distillation as inevitable and open-source models as a resilient counterweight. Ultimately, his vision is one of intelligence abundance, where tokens are so cheap that proactive agents become ubiquitous, limited only by the questions we can ask.
Ramp is the only platform built to make your finance team leaner, faster and better,
saving businesses 5% annually on average so you can stay focused on growth.
Ramp customers grow revenue 3.2 times faster than the average American business.
Visa, Versel, Cursor, Stripe, Notion, 11Lab, Shopify, and 70,000 other businesses all now run on Ramp.
Mind us too and so should yours.
Learn more at ramp.com/invest.
OpenAI, Cursor, Anthropic, Perplexity, and Versel all have something in common.
They all use WorkOS.
To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO,
SKIM, RBAC, and Autologues.
Instead of spending months building these mission-critical capabilities yourself,
you can just use WorkOS APIs to gain all of them on day zero.
That's why so many of the top AI teams you hear about already run on WorkOS.
WorkOS is the fastest way to become enterprise ready and stay focused on what matters most your product.
Visit workos.com to get started.
Felix Byrogo is a personal finance agent that turns a single prompt into finished,
client-ready work using your firm's own templates, context, and standards.
Send Felix an email like, "Take these comments and turn them for me."
Or, update my tracker with the context of these emails,
and Felix sends back Finnish PowerPoint decks, Excel bottles, and source research.
Felix works the way your team already does, delivering work quickly and accurately around the clock.
Learn more at rogo.ai/feelix.
Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like The Best.
This show is an open-ended exploration of markets, ideas, stories, and strategies
that will help you better invest both your time and your money.
If you enjoy these conversations and want to go deeper,
check out Colossus, our quarterly publication with in-depth profiles
of the people shaping business and investing.
You can find Colossus along with all of our podcasts at Colossus.com.
Patrick O'Shaughnessy is the CEO of Positive Sum.
All opinions expressed by Patrick and podcast guests are solely their own opinions
and do not reflect the opinion of Positive Sum.
This podcast is for informational purposes only
and should not be relied upon as a basis for investment decisions.
Clients of Positive Sum may maintain positions in the securities discussed in this podcast.
To learn more, visit psum.fc.
My guest today is Neil Mova, the founder of Sail Research.
Sail is building what Neil calls a token factory, an inference company designed
for a specific kind of future.
One where AI agents run in the background for hours or days at a time
rather than answering a human in real time.
In that world, latency matters last and cost matters much more,
and Neil has built the entire company around driving the cost of a token
as low as it can possibly go.
What makes this conversation special is as one of the most detailed tours
I've ever done through the full stack of intelligence, the software,
the chips, the power, and how the three connect.
Along the way, we cover the trade-off between speed and cost
that lives inside of every GPU, his scavenger strategy for buying the chips
and power known else wants, his contrarian view on Nvidia,
and why the premium, the Frontier Labs charge for being three to six months
ahead may not last.
Please enjoy my conversation with Neil Mova.
I think it's important early in these conversations to just say the thing,
literally what you're building and what it does today.
So maybe just orient us there with a brief description,
like literally what the system is that you're building and why it should exist.
Cell research is a token factory.
We have an API where anyone can send us requests where they can use
large language models, open-source language models for any test they want.
We will serve those tokens to them at a price that is unbeatable in the market.
We also support their ability to build agents on top of this.
We host what we call sale boxes,
which are long-running agent virtual machines hosted in the cloud
that are designed for agents that run for hours, days or weeks.
So I should think about you as a peer company to others that serve different kinds of inference.
You're serving one specific kind of inference,
and your goal is to be the absolute cheapest provider
and a neighbor of a certain kind of use of intelligence.
Exactly. The theme of our company is abundance.
We want to deliver this new commodity of intelligence to as many people as possible
at a cost that is sustainable for almost every industry.
We think that whenever you make something test cheaper, it's a new product category.
We aspire to do that for tokens.
We think it's so profound that the machine can think
and now our jobs to make as many machines as possible in the world work towards thinking.
So if you think about the theme of the day being token costs,
is token costs the right way to think about this?
Is there some other way you'd put it?
To start with, absolutely, token costs.
Today, my north stars, I want to have the lowest cost for token in the industry
and do that by a mile.
I don't think tokens are the final unit of work or intelligence,
but they are what we use today.
After tokens, you start to move more towards more outcomes,
which is like a vague direction.
You can imagine, for example, today when you can zoom tokens through an agent,
you don't actually control how many tokens the agent reasons for.
It can reason for a certain amount of time or it can call a certain number of tools.
And increasingly, I think we will have agents do some unit of work.
Take as many shots on goal as they can.
And however many tokens they used to get there is going to be dependent variable depending on the task.
So you think about agents that self-administer a token budget
as opposed to a company sending a budget for how many tokens engineers can spend for a month.
Why is there an opportunity that you can tackle?
It seems like the entire world is oriented around more, better, faster, cheaper tokens right now.
It seems like the world is trying to solve this problem very aggressively.
What was the unique opening that you saw that's maybe the market's not being efficient
and it's attempt to tackle this?
So I think there's two things that are tailwinds for our company.
One is got to be the rise of open source had to tackle that first.
I think we are starting to see increasing number of our customers
and the broader market care about owning intelligence.
They want to have control sovereignty over the thing that they depend on.
That created a much more robust market for a customized models.
Or even just like these vanilla open source models that no one can ever take away from you.
You always have the weights, you always have the right to deploy them however you like.
In that world, there's been a reasonably robust market for the past couple of years
serving these models at large scale.
The challenge is all those companies, you could take your pick based on fireworks together.
They all focus on low latency inference and they were pulled in that direction by one very
important customer.
Cursor.
I think that that was the right choice about a year ago.
And as a six months ago, it started to look like maybe low latency wasn't the only thing
you wanted from an agent.
You wanted more persistence, more long horizon tasks.
And now it's to me very obvious that the future of ancient inference is long horizon tasks.
You're going to run the machine for hours or days at a time.
It doesn't matter if it's been set tokens at a hundred tokens per second.
Maybe ten is just fine.
If that comes with corresponding advantages and efficiency.
Why are you so confident in that?
To me, it seems like I want everything as fast as possible.
When you're waiting on it, you absolutely deserve the fastest sensor possible.
My trick is I don't want you to be waiting on it.
I want it to be proactive.
I want it to be in the background.
One way to say it is like the best latency is no latency at all.
When you wake up in the morning, the works are already been done overnight.
You didn't even have to ask for it.
That's the dream.
We're not quite there yet.
More importantly, I think the more you're in the loop as you prompt agents and wait for
a response, in fact, you're the bottleneck in having the agent do more or less work.
What we'd like is the agent to operate on more human time skills.
You don't manage your colleagues every five minutes.
You ask them to do a high-level task and you come back and check in maybe every day,
but more likely once a week.
And that, to me, is the future of human agent collaboration, more like human time skills.
Say more about the early indications that this is happening, and therefore you should be
building this company.
The first and most important thing is the idea of test time compute scaling.
The idea that you can give an agent more time, and it will give you a better answer.
That was theorized about two years ago now, but it wasn't really something that we could
actually bet on until, I would say late last year, with Opus 4.5.
Opus 4.5 was the first agent that was at all suitable for longer horizon tasks.
It was pretty mediocre when it first came out, but you look at the more recent models
and what we've done on open source as well.
You see that agents are capable of running for an hour at a time.
I wouldn't say it's days, but definitely an hour is quite suitable today.
Seeing that like average task length get longer and longer, doesn't take many points
to how do you draw out the exponential and see that agents are worth running for a longer
period.
What do you think will be the market share of long-running agents in three years or something
like this?
I love this market because it's unbounded.
There's no human in the loop so you can consume as many tokens as you like in the background.
Versus human attention span, if you tell me to consume 10x as many tokens at Codex or
at Cloud Code, I'm actually not sure if I can anymore.
I'm already in the loop and locked in coding for most of the day that I'm at the laptop.
What is unbounded is how many tokens can be consumed in the background or proactively.
Long-term, I think we're going to end this year at maybe 50/50 background and real-time
workloads, but I see this going to 90/10 in favor of background.
What are your favorite examples of something that gets accomplished much better as a background
task than as a human in the loop task?
Most deep research.
Questions where you want to have a definitive answer over not 100 sources, not a thousand
sources, but 10,000 sources or more.
If you want to build an authoritative index of information, like, for example, one of
our customers parallel web systems seeks to do.
They want to build an index over the whole internet and they want to monitor the internet
in real time for changes.
That is the kind of crazy, exabyte-scale task that you need a very different kind of intelligence
or scale of intelligence to achieve.
Deep research is a top category for us.
And increasingly, we see cybersecurity following this direction.
If you think about, yes, there's so much code you can generate, but there's exponentially
more ways to break that same code than it's to generate that code.
There are some great customers out there who are working very hard to find agents that
can break any pieces of software and proactively patch them.
When Fable first came out, for example, or Mythos first came out, basically there was this
push in the cybersecurity community to run Fable against every line of code we've ever
written and look for bugs in 20 different ways, meaning you're looking for both memory
errors, you're looking for business logic errors, and looking for network vulnerabilities,
all these things.
And these are all actually things that you would write specialized agents for.
you wouldn't just have Fable look at the source code once,
you'd have it actually set up environments,
where you can pen test these applications.
At some point, people started to make this joke
that security has become proof of work.
When you want secure software, it's really a question
of how many dollars did you spend on Anthropics APIs,
trying to break into your software.
That is the best indication for how secure it is,
because that's the best tool in the world.
And increasingly we found that open source models,
well, the frontier of intelligence here is quite jagged.
It's not the case that Fable finds a superset
of all bugs in software.
You would find some bugs with a very small model
that you don't find with a large model.
You'd find some bugs with Hiku,
that you would find with Hable and vice versa.
So it encouraged this very diverse approach
to sampling and trying to build cybersecurity agents
that break software autonomously,
such that you can patch them.
If you were to get speculative and imaginative
about the sorts of things that very cheap,
very long-running agents can enable,
we talked about some very practical examples,
deep research, cybersecurity, et cetera.
If you get a little bit dreamier about the use cases,
new product category, that this sort of inference will unlock,
I guess the question is just like, so what?
If you're maximally successful,
dream a little bit about what that might enable.
I think for individual users,
what I'm excited about most is this idea
of proactive, intelligent agents.
You can imagine a series that is running in the background
all the time to understand all the emails you received
in a day, all the text messages you receive in a day,
and has a much more encyclopedic view of your life
and how to be helpful in that life.
Right now, there's still point solutions
you're doing a lot of prompting.
Series number you're proactive.
So we can fix with abundant inference.
If you trust the machine enough that it's reliable
and also trustworthy as in private,
you might even imagine the machine can understand
how you interact with it and practically surface
your next action whenever you open your phone,
can we build a good model of what you're gonna do next?
My estimation is yes, we totally can't.
And the key to that is incredible cheap intelligence.
You have to be willing to spend tokens
without any promise of return.
That is the unlock.
The long lens view to take on this is that
we have a form of intelligence that can tackle
any verifiable problem.
Any verifiable problem means most software.
It means a lot of formal math proofs and similar.
And it could also mean scientific discovery.
These are all relatively verifiable problems.
And all those things currently have a dollar cost attached
to them, essentially, that's a hidden one.
It's like how many tokens could you possibly harness
to make this work?
We have started to bring it within view a dollar cost
for these long horizon tasks that is reasonable.
It's not millions, it's thousands.
And maybe it could be hundreds or even tens of dollars
in the near future to have a definitive answer
to any scientific question, to any research problem.
So if we dream about that future,
we then become limited just by the questions
that people can ask basically.
Pretty much.
The questions we can ask, the models are on the cusp
of basically taking even a high level question
and chasing it down every possible follow-up.
You can have the model, essentially, take that on its own.
And the question is, what is your token budget?
And we will solve the token budget problem.
What about non-verifiable tasks?
Those are basically the entire category of human taste
into that category.
We have not solved human taste yet.
And I don't know that it fundamentally can be.
I'm excited to be surprised here,
but we are focused on very quantitative problems.
We leave the quality of writing.
We leave the beauty of art to people.
Vanta automate security and compliance
for over 16,000 fast-moving companies
like Ramp, Kersher, and Harvey, keeping
an audit ready around the clock.
It's the number one agentic trust platform.
And it now helps companies like yours
watch for the risks that show up between audits,
across your vendors, your AI tools, and your whole environment.
Every new tool your team signs up for,
every vendor that turns on AI features
is an opportunity for something to go wrong.
And most security programs weren't built for AI's pace of growth.
The Vanta agent works like a 24/7 GRC engineer
in the background finding issues,
drafting fixes for you, and cutting vendor assessment time
by up to 50%.
Whether you're a fast-growing startup
or a global enterprise, Vanta helps you earn and prove trust.
Invest like the best listeners get a special offer
for $1,000 off at Vanta.com/invest.
RidgeLine is the first end-to-end system of record
with embedded AI for investment management firms,
running portfolio accounting, reconciliation, reporting,
trading, and compliance on one unified platform.
Firms are moving off legacy technology
and onto RidgeLine because of how far ahead
RidgeLine's AI features are compared to anything else
in investment management software.
Which is why I believe that firms that come out ahead
in the AI era will be the ones running on RidgeLine's
unified platform.
If you're serious about your firm's AI strategy,
RidgeLine should be part of that conversation.
You can request a demo at RidgeLine.ai.
(upbeat music)
All right, now let's talk about the very clever stack
of solutions that you hope to build.
Ultimately, they have this giant token factory supplier
of extremely low-cost intelligence.
I think you think about this in terms of software hard work
and power.
Talk through what your master plan is to approach
this challenge that's so different
from what others are thinking about doing.
We always have to start with software.
Where is the opportunity on today's data centers
to improve efficiency?
And the first thing we did was we tried to build
the entire LLM software stack around peak GPU efficiency.
Meaning we're using embedded GPUs.
We wanted to squeeze out more tokens from the same chip
than anyone else in the world.
And that starts with the lowest level of programming kernels.
It's actually my background.
I spent my whole professional life working on GPUs
and kernels in videos my first job while I was in college.
I got to see how the tensor cores got to earn
their right to be on the chip.
What does that mean?
What is the tensor core?
Tensor core is a specialized unit on the GPU
that accelerates matrix multiplication.
Simple as that, there's been a long history of how we evolved
that tensor core over time that we'll get into.
And why is matrix multiplication so important?
I cannot say that there is a divine truth of the inverse
that explains why matrix wall applies
seem to be the atomic unit of computation.
But one way effort to describe to me is,
well, it's a really succinct way to mix
two blocks of numbers together
and have them interact in some interesting way.
That's as much as I can say about it.
It is really convenient that linear algebra turns out
to be a very compact representation
of arbitrary relationships and data.
So Nvidia, great graphics company,
had market share dominance in GPUs
and gaming graphics for quite some time.
And then starting in like the mid 2010s,
they started to actually start these like Skunkworks projects
to make the graphics processor more suitable
for machine learning tasks that they were tracking.
I remember actually reading some of the lab notebooks
of some of my managers when I was in Nvidia,
they would visit these small ML conferences
like ICML or NERIPS at the time.
They would just take note of these papers,
like, oh, this deep learning thing seems to be catching on.
And what's really interesting is that these grad students
are using gaming and Nvidia GPUs in order
to train their large models.
We should double click on this and figure out what's going on here.
By 2015, 2016, at least Jensen had the conviction
to kind of double down on, hey, this usage of our chips
is only going to grow.
Let's start allocating more and more precious silicon
die area to this capability that seems to be emerging.
Let's put the first version of TensorFlow's on the chip.
So we're talking about taking this gaming chip,
which is designed for painting pixel in a screen
and adapting it to do magical supplies.
It was early, you would be competing against the graphics
teams, essentially.
When you ask for more silicon area on any chip company,
there's always competition for that.
It is something that the designers guard so carefully.
You don't ever want to invest in the wrong technology,
because that's opportunity cost that you could have allocated
to some other functionality.
We fought tooth and nail and got just a tiny bit of dire.
It may be like 5%, 10% something like that
for the first generation of these chips to get
some amount of acceleration for basic convolutions, which
were the fundamental operation for computer vision models
of the day.
And then we had a software team that
was trying to squeeze all the performance we could out
of the chip.
And I think on that software team, which is where I work,
that's what actually taught me the most
about the ethos that Nvidia has around this term called
speed of light.
They always chase the speed of light for any piece of hardware
that they make.
It is so ingrained in every engineer's mind
that if the machine can do it, we're
going to push the machine to the frontier until it does.
What we think is the speed of light is the edge of what's
possible.
The speed of light is the edge of what's possible, exactly.
If we think the chip can run at this frequency
and produce this many multiplies per cycle,
we're going to get there.
We're going to break every bottleneck
and get to that peak level of performance.
To this day, I tell all my engineers,
we're chasing 100% speed of light.
I don't care about relative numbers, first of the competition.
I only care about absolute numbers.
What are we able to do on the chip?
How do we achieve that?
Before we leave that chapter of your time in Nvidia,
anything else beyond that cultural touch
point that changed the way you think about things
or that stood out the most about how the business ran back
then or its culture?
I have a ton of stories about Nvidia.
I could tell you a few of them.
One of my favorites is that on the 10-year side,
a lot of people I worked with in Nvidia in 2015, 2016,
are still there today.
That company has incredible retention,
and these are the best engineers.
On the Silicon side, at least I've worked
with my whole career.
They're extremely, extremely motivated and passionate.
They believed in parallel computing as a concept
through its various incarnations and have loved
seeing the chip evolve.
This is their life's work, and they're extremely competent
in that direction.
They're also a very frugal company.
Nvidia and all of it, Silicon Valley companies,
after 2008, they had some cutbacks and perks.
So no free lunch, for example.
Nvidia took it one step further.
There was no free milk in the fridge.
So if you wanted to drink coffee at Nvidia
and you wanted some milk, you actually had to chip in
a dollar every month to the milk club,
and the milk club would stock Costco milk in the fridge.
And I remember that distinctly.
We don't do that at sale.
It's a frugality that permeates the company.
So coming out of this time there,
you get this experience of what it's like
to develop more efficient usage
of the underlying hardware through software.
So link that to today's environment.
- The GPU is fundamentally a throughput machine.
The GPU's happiest when you give it a lot of work to do
and let it chew through that work
at peak utilization of its compute units.
But that's not actually not the way that we've taken AI
in the last couple of years.
We've pushed AI to be an interactive chatbot tool
is the most common form of AI usage today.
In that world, you care a lot about actually
spinning answers out to the person of the keeper.
as quickly as possible, to your point about don't make the user wait, I want things as fast as possible.
That's actually quite interesting for the GPU.
It's very difficult to put the GPU in its happy path of being fully compute utilized
when you're trying to spit out tokens quickly.
There's a fundamental trade-off on the GPU between being throughput oriented or latency optimized.
And everyone has chosen latency optimization because the shape of usage was jackpot oriented.
I believe that's the most profound change we're going to see in the next year.
We're going to move away from chatbots to more proactive or background agents.
In that world, it makes a lot more sense to build a stack around throughput.
Can you explain technically why the trade-off between throughput and latency is unbreakable?
Why can't we have both from the same hardware?
It's quite foundational in almost every system that you could ever possibly look at.
There's always a trade-off between getting a small amount of data through the system as quickly as possible
and leaving a lot of buffer room for that or trying to run wide and slow.
Narrow and fast are wide and slow is like a classic trade-off in all computer science.
But for GPU specifically, I think there's one thing to focus on, which is there's this concept of like batching on the GPU.
We want to group many users' work together into a batch that we can run all at once on the GPU.
That's the parallel processing of the GPU we'd like to have a lot of parallel work to do.
The thing is though, you're doing net more work when you run a large batch of compute together.
So you might be filling all the units, but every step along the way as you carry a batch of work through the GPU,
there's more work to be done. Any individual token or any individual user's request in that batch,
it's going to spend a longer time on the GPU being carried with other people's traffic.
Maybe the way to say it is, if you want to get downtown NSF, you can take the bus or you can take a private transit.
And the private transit is going to have its own direct path as the crow flies or using exactly the roads that you want from pointed to point B.
A bus, it's going to have to serve many more people and it has to fundamentally do something that works for everyone.
So it takes a slower path and it stops and waits for other people to get on and off.
I think the bus versus car analogy is pretty accurate.
It's a great analogy and step one for what you're trying to do is create the best possible bus on top of Nvidia GPUs.
Like that's step one of your optimization.
That's exactly right. It means we explore things like different parallels and schemes.
And maybe that's another example I can give you is with Nvidia GPUs,
one of the things that they've really innovated on into a great job with is the NV link interconnected between GPUs.
And in fact that NV link system is so good that if you have a large matrix multiply that you want to perform faster,
you can actually cut that matrix multiply in half and shard it across two or more update.
Let's say, Nvidia GPUs and have them all work on pieces of that larger matrix of the play and have them connect to the results together at the end.
Reduce the results back together at the end.
This is a great, great way to cut the minimum latency of an operation.
Each GPU is ending one eighth as much work, let's say, and therefore can finish faster but not eight times faster.
It's sub linear scaling. You'll use eight times more hardware, but you won't get eight times the speed.
You might get like four to five x the speed. You're not going to get strong scaling.
This is because of communication overhead because every GPU is going to be a little less efficient working on a smaller tile of work than a larger tile of work.
It's the only way to speed up if you want the minimum latency possible.
You can do that, but it is not the choice I would make.
For example, I would prefer to use a different parallelism scheme like expert parallelism or pipeline parallelism.
We may do interesting things to overlap and hide the communication latency in a way that you would have less ability to do that for a low latency service.
Which is the right way to think about NVLink as a technology which improves latency performance and only latency performance.
Which will segue into the next segment of what we do differently as a company.
But yes, NVLink is mandatory for low latency inference.
So NVIDIA is excellent at low latency inference, and I'm telling you that we don't really care that much about low latency inference.
So where does that leave us?
I'm not holding my breath for other companies broadly to figure out NVLink quickly.
It's a challenge of technology to figure out. It's hard to scale. It's hard to productionize.
If I do have some other vendors chip and it is good at the foundational compute components,
it can still be made to multiply as really well. You just can't communicate those results across its peers quickly.
Well, maybe there's room for that other chip in my stack as a really, really good compute per dollar option.
That's what I actually optimized for. In most cases, is how many flops does this chip have?
And how much is it going to cost me per hour to own an upgrade?
There are other chips that definitely rank higher than NVLink on flops per dollar, but they may not have as much interconnect.
So it's my job to figure out what parallelism scheme I'm going to use that's going to make this chip suitable for inference.
It's not going to be tensor parallelism in video. It's basically mandatory for that.
But other techniques may work well for me.
So before we leave the latency part of the story, can you comment on companies like Serebris or others
that can perform incredibly fast operations? I'm curious like what you think about those approach to those companies,
what might happen in the future? What is your prediction for the future of very low latency focused hardware?
Serebris, GROC, and a couple others that are coming out of stealth now.
I think have made a very interesting bet on not just building another GPU,
but actually building a different kind of accelerator that focuses on a different memory hierarchy.
They want to maximize the amount of SREM on the chip and use that as very, very fast memory for weights and KV cache.
So SREM versus DRAM, there's two ways to make memory for a chip.
One is to integrate the memory on the logic diet self.
Meaning you tell TSMC, I want this many megabytes of storage on my chip, and there's a way to build that.
TSMC has a standard cell library you can use, and you can just print out a bunch of cells of SREM.
The problem with SREM is it takes a lot of area on the silicon die.
So if you want to build a large die like let's say an Nvidia Blackwell at 800 millimeter square,
if you made that whole die SREM, it would be in the maybe like single-digit gigabytes.
It feels like it's not a crazy amount of data storage.
Compare that to if you're willing to take a different process technology entirely.
So not TSMC anymore, but now micron, SKH next Samsung.
Today you build DRAM, which is a whole different way to build memory, and that's more focused on capacitors than transistor cells.
The standard way to build SREM is what's called the 6T transistor cell.
It's a stable transistor arrangement that allows you to write a bit to it,
and then it holds that state in that bit regardless of whether you keep it playing.
Well, you have to play some power, but it's holding that bit without any sort of like active management.
It's static.
Now dynamic RAM, DRAM, it's dynamic because what you do to write some data is you write a charge onto a capacitor.
And as soon as you write that charge with that capacitor, the charge is dissipating.
The dynamic part of DRAM is that you must every 50 milliseconds or so refresh every bit to your friend.
So you're constantly juggling billions of balls in the air, essentially billions of bits,
had to be managed by a memory controller, which is reading and refreshing every bit on the DRAM.
Now the benefit of that is you can get much, much higher density.
And it's a whole different process technology. There's a ton of different trade-offs.
Hence, where we split the DRAM manufacturing into an entirely different company,
like Micron, SK, Linux, and Samsung. These are the best companies in the world to do this.
They build DRAM.
And if you take DRAM from those companies, and you stack it into many layers,
and you kind of print them or solder them around the main logic die that you get from Nvidia,
you can now get hundreds of gigabytes.
Like Blackwell has 298 gigabytes of HBM capacity around the logic die.
And the logic die itself, maybe only has like 500 megabytes of SRAM.
So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus SRAM.
So let's go back to the reverse. What are they doing?
They see this problem. There's not really an obvious way to increase SRAM density on the chip.
But the thing with SRAM is because it's so physically close to the logic gates that actually do the computation.
The arithmetic logic units are right next to the SRAM that they're going to pull from the compute units
that are doing the magical supplies can pull data from SRAM at mind-boggling speeds.
This reverse quits petabytes per second, 21 petabytes per second for their wafer scale engine three.
Compare that to HBM on an Nvidia Blackwell, is that 10 terabytes per second or so in that range.
So once again, many orders of magnitude difference, more capacity,
but proportionally less bandwidth, essentially.
What the reverse does is they say that we're going to take as many of these dyes as we can.
We're not going to limit ourselves to the 800 millimeter reticle limit, the TSMC,
8th and square millimeter limit, the TSMC imposes on us.
We're going to take the entire wafer and have every die connect to every other die over scribe lines.
And we're just going to try to get as much SRAM as we can on the whole wafer.
And we can get to like, let's say, 50 gigabytes of SRAM per wafer.
And then we're going to stack many waifers together in a pipeline or similar.
And now we can have up to a terabyte of memory, a very, very fast memory.
You do all that work just to get to the ability to read data from SRAM at 21 petabytes per second per wafer.
Therefore, you can now serve these language models at extremely high tokens per second
because you can move the entire parameter count of a large model like Kimi.
You can move all that data in and off the logic cores in about a millisecond or something like that.
So there you go.
You have a path to a thousand tokens per second.
So what is your prediction for like that segment of the market?
What happens to them is some hybrid.
We had to pair this rubric chip where it's very strong.
It's very, very good at fast access to memory with something that has more capacity for memory.
It's true that you can take a one trillion parameter model like Kimi and fit it on a large number of SRAM's die waifers.
But you can't do something about the KB cache very easily.
The KB cache is something that grows as people use the model more.
And that is always dynamic.
You don't even know how much KB cache you're going to need.
Depends on how many users you have and how many users you want to serve.
Can you explain KB cache?
KB cache, when you ever use a language model, every token you send through the language model stays in the context window of the language model.
For as long as you're having a conversation, we talk for a hundred thousand tokens.
The hundred thousand than one token is still in the conversation behind us.
And the model is referencing all the past conversation history in order to make better predictions about what the next thing we're going to say is that KB cache, it's a bunch of memory.
You have to store a representation.
for every token that we send through the language model.
And it frequently gets to be larger
than the weights of the model themselves.
You have this crystallized knowledge in the model weights,
and you have the dynamic knowledge
of the exact conversation we're having
in the KB cache is the way I like to think about.
- And this is why sometimes people would observe
deep in a conversation things start to degrade
'cause there's some sort of technical problem.
- So the KB cache is quite interesting in that regard.
The KB cache is an exact representation
of everything that came before.
We store all the information that we've seen
in the conversation.
However, during training, the model did not get trained
primarily on very long context conversations.
It got trained primarily on, let's say,
8,000 token conversations or 16,000 token conversations.
So if you take the model to 200,000 tokens,
there was some training that happened at that context length,
but it's not the models like core strength.
And so there's always been a challenge
for the frontier labs to figure out
how do we make the model exactly as intelligent
at 10,000 tokens as we expected to be at 200,000 tokens.
And it's going to be a perennial battle for us.
We've had one million context windows
as a concept for years now.
And for up it goes, I think the first
to hit that one million context window length.
Len, I still use slash compact in my cloud code.
Well, before one million context length,
I don't think it's actually great to hit the full length.
- These extremely fast, extremely low latency approaches,
ultimately are limited by this factor.
- Yes.
- You can do whatever you want for the weights.
- It's very possible to have
which unbeatable performance on weight storage.
However, the KB cache is going to be
a big thorn in your site.
- So three years from now, five years from now,
what role do you think these kinds of chips play,
like what sort of market share do they have
in the heterogeneous chip market?
Like Srebris and Grock and maybe a couple others.
You should think of them as accelerators.
What they are really good at is being used in conjunction
with a more traditional GPU like device
that critically has this off-chip memory built in.
You want off-chip memory for capacity
and on-chip memory for speed.
We want to hybridize these two things.
So if you take transformers in the limit,
you take a transformer to a million context length.
What ends up happening is you have this compute down stage,
which is the actual matrix well applies
for the what we call the MLP,
which is where most of the models
acknowledge world knowledge is encoded,
and then you have the attention layer,
which is where we're dynamically adapting
to the current conversation.
Attention in the limit is usually memory bound.
The MLP in the limit is compute bound
at large enough batch size.
I would say the original sin of transformers
is that you've taken this extremely
fundamentally memory bound layer
and juxtaposed it right next to a compute bound layer.
It is very difficult to have a single chip
that is going to both compute operations
and memory operations.
The GPU is quite balanced in this regard,
but you had to choose one or the other.
Serubis has a very fast memory access
for something like a Mitch's will apply,
and it's really good to host the MLP
the weights essentially on this rubric chip.
But the GPU has the capacity to scale
to really long context lengths.
You would like to put the attention
possibly on the GPU and the MLP on the Serubis chip.
Can I believe this is what's happening with Nvidia and GROC?
Can you refer me to just on transformers
for people that again aren't deeply familiar
with what this innovation was in 2017?
What its strengths and weaknesses are
and whether or not you think it will remain
the dominant architecture or a dominant architecture
for the future of AI?
What it did, it allowed us to learn
on supervised data really effectively.
Because transformers, what they're all about
at the end of the day is taking any sequence,
any arbitrary sequence of data
and trying to find patterns in that data.
And they critically, the attention operation,
which is the headline component of transformers,
it allows the model to dynamically adapt
to what it thinks is the most relevant component
of the sequence.
Every step you take through a transformer,
you are essentially like re-weighting the input
that you looked at before and figuring out
which is most relevant for your next prediction.
It's extremely amenable to learning arbitrary sequence
data and the most interesting sequences of data
that we produce on a regular basis is language.
And that's how we got two dominance
in the language regime.
To zoom out even further, I think what
transformers really did well is that they scaled.
Transformers make no such human prior.
Transformers just say, well, there's
going to be a pattern in the sequence of data.
And if there is a pattern, I'm going to find it.
I'm going to throw more and more parameters
at this problem until it works.
Transformers benefit from a lot of the computer vision work too.
For example, one of the challenges in computer vision
was we had a hard time going from hundreds of thousands
of parameters, which you get for linear models,
like support vector machines or other legacy machine
learning models.
Those had thousands of parameters.
Then we got to deep learning and got to tens of millions
of parameters with computer vision.
The biggest models were 150 million parameters
was a huge model for computer vision.
And now we routinely talk about trillions of parameters
and transformers are the link to go
from millions to trillions of parameters.
So if I think about the important units of scaling
being data and compute, does it stand a reason
that you think transformers will just stick around?
Because that's the thing that we're
good at getting more of those two things.
Transformers are such great sponges.
You increase the compute available to a transformer by 10x
and you'll get some log improvement somewhere.
And so far, the scaling was really work.
They're really quite beautiful.
And to the point about what do transformers do really well,
they extend to almost any data set you can throw at them.
They're extremely powerful general learners.
And I think what's especially useful about transformers
over other techniques that we've tried to replace
attention is transformers represent
any paralyzed relationship that you want.
Any token in the sequence can attend to any other token
in the sequence.
So if there's any relationship that's in the sequence at all,
you're going to find it with transformer.
Now, maybe the case that you don't need all to all modeling.
You don't need every token to look at every other token.
But if you need to, transformers give you that option.
And until we know a better way to prune that space down,
a better way to kind of have information modeling
be more selective.
Attention is a very, very good operation.
This is another trick that we learned in the Confederation days.
One of the old Carpati sayings is that,
if you have any data set that you want to train a model for,
your first goal should be to over parameterize the model
and try to overfit the data that you have
to prove that there is a relationship that you can model
or memorize that you're learning algorithm works,
that you can instill knowledge into the model.
Once you can overfit, then you can compress.
And the compression is how you get generalization.
You don't want to actually memorize the data
that you have in front of you.
You want to generalize.
And therefore, once you overfit the data set,
then you can kind of work backwards
and try to find the general patterns
that fit into the smallest parameter count possible.
What's your prediction for the future of data
and riff on the importance of data in this whole story?
I like the phrase that intranet was a one-time subsidy on data.
We got it for free.
It's extremely high-quality.
There are about 30 trillion tokens of high-quality text,
300 trillion tokens if you take a wider view
on what qualifies as a good text.
And we've basically looked at it all already.
The models have seen the entire internet
many times over at this point.
And there is not a whole lot more to be done
on human data from the internet.
The next phase of data, in my mind,
is model self-improvement through RL environment gyms, basically.
Now, in fact, we don't even benefit
from getting more random user reactions with AI.
It used to be that the new type of data
that we cared about a lot was the interaction data
from people using HTTPT and giving HTTPT signals
on what they liked and didn't like.
I like the argument now that the median model
that we serve is so much more advanced
than like a random human giving feedback
that the signal you get from random human preference,
unconditioned human preference,
is not actually worth anything anymore.
You want expert human preference
at this point.
The model has outgrown.
- Every day, generally.
- Yeah, every day too.
- Exactly.
So the feature of data, to me, is giving the model
a hard, verifiable task and letting it run in this gym
where it's kind of isolated
and it just has a problem and it can make progress on it
and get measurement of whether it made progress
on that problem or not.
You can imagine coding problems are in this category.
Math problems are also in this category.
More and more,
we can just give the agent a computer essentially
and have it act like it's a human worker
and give it feedback on whether it's making progress
towards the target outcome.
That environment becomes the data.
I think this is not a super differentiated take
but it's been really, really productive
for what I've seen so far.
- And you think that just goes on
for a really long period of time
or is that another like,
if I think about the internet as this one big block,
like this is another big block that will have its day
and the sun and will kind of get it all
and then we'll have to move on to something else.
- I think it's actually more profound
than that.
Basically, the idea is that if you want artificial
general intelligence,
the best way to get there is to just keep stacking
specialized intelligences until you have no more gaps to fill.
The only thing you need to make sure you do
to make this work is you must make sure
that your task is verifiable.
You need to get the model a self-grading system.
If you have that, you have the recipe for self-improvement
on any task you like.
I think you've seen this held up by the way
Frontier Labs spend,
they just spend that much on data.
Now they spend a lot more on our environments.
These environments absolutely capture that relationship
of recursive self-improvement on a verifiable task.
- Coming back to your initial task
of making existing hardware more efficient
by being more in control of what's going on
at the hardware level through software.
Keep going on what you've done so far
and what you want to do.
And then we're going to jump to hardware
and then jump to energy final light.
- I mentioned kernels.
It's surprising.
People think kernels are done.
They're great people like treat out who write excellent kernels
and they form the bedrock of all of our modern deep learning
is built on flash tension.
Modern transformers are built on flash attention.
But if you deviate from the happy path at all,
if there's a new model that comes out
that has a slightly different way
to embed positional information,
the change of the rope system.
Suddenly, the kernel that we had is not suitable
for this new model.
And we may have to make a patch to this kernel.
I wouldn't say we're in the phase
where we had to invent new kernels from scratch,
but having the ability to quickly modify existing GPU kernel.
Sorry, a kernel, by the way, is a general term
for any programming run on the GPU.
Historically, kernels tend to be put into a library
where every kernel has a very, very scoped purpose.
You typically have a kernel for a major supposed apply.
You have another kernel for you and something
as simple as addition.
You want to add two tensors together.
That's another kernel.
And then increasingly, we've started to fuse
those kernels together.
if I do.
a matrix will apply and then I want to add it to another matrix that I've also multiplied,
maybe those two become one kernel and I just fuse the operations where sort of writing the
data out to DRAM and then reading it back in just to the addition, they make us do this easily.
Why are humans still doing this? It seems like the sort of thing that AI's would be exceptionally
good engineering, more efficient kernels. Maybe that's where we're going and we're just not
quite there yet. If we aren't there yet, is that where we're going? If we're not there yet,
why humans still doing this? Why is tree now so well known? It's a name I know.
I don't want to speak for tree, but what he taught me was you shouldn't write kernels by hand
anymore, necessarily. I like to say we write kernels in the whiteboard. We go to the whiteboard,
we describe what we think the machine should be doing, then we succinctly describe that in
natural language to a model and then the model is able to do the execution of here is my input and
output, here is the strategy of how we want to dispatch this work onto the GPU. I'm going to go
implement it. We're doing the conceptual design. Exactly. I'm not sure exactly why models are not
superb at doing this. I don't think this is like arm out or anything like that. I'm sure in six
months time, we'll have much better models on kernel engineering and I'm sure the labs would tell
you that they already do a lot of their kernel engineering in a fully automated way. So software
as an edge, if I think what software is maximally near speed of light, efficient usage of an
underlying piece of hardware, is going to trend towards not being an advantage for a company like
your software time. That's right. The rising tide of something like mythos or GPT 5.6 sole,
that lifts all boats. It really does. I actually don't think there's a point in specializing
to say we work on making the model better for just kernel engineering. I think that's actually not
the most meaningful subset of coding in general. Kernel is engineering in particular. Maybe there's
some privileged information you inject into the prompt that's like a useful way to steer the model
to be better at writing kernels, but probably speaking, yes, we're all downstream of the frontier
in terms of this capability. I always love this from the history of energy. There's always this
pendulum between the raw source, let's say coal. And then if there's a certain amount of energy
available inside of a hunk of coal, what percent of it we can harness and use? A big part of the
history of energy was getting that number from 10% to 95% or whatever. If I just think about a
blackwell or something and blackwell is the piece of coal, what percent do you think we're at?
How efficiently can we use any distinct piece today? There's a lot of different ways to analyze
that. In some level, we are really efficient at optimizing the performance when the GPU is doing
the thing that it's most happy doing, which is a large dimension matrix will apply. That operation
runs at 70-80% of p-cutalization and it's limited not by software, but by power. The way NVIDIA
quotes p-flops is optimistic. You never hit that because of power throttling. Exactly, thermals.
In practice, you don't spend the majority of your time in a transformer in that happy path where
you're doing a large batch matrix will apply. Our job is to basically build the engine around the
chip such that we are feeding the GPU these large batches of work at all times. One of the most
profound transitions we've had in the GPU world in the last year has been this moving of you don't
program one GPU at a time anymore. You should think about the whole rack and maybe you should think
about the whole cluster the entire data center at a time. With NVIDIA, they've started shipping not
just a single GPU or a single motherboard, but actually the whole rack system is something that
they prescribe. They call it NVIDIA 72. Their latest chip, the Grace Blackwell 300, that ships
as a rack of 72 units and it's an open race to figure out who can program the whole rack scale
computer as efficiently as possible. And my belief is that that shape of compute is the future of
efficiency and speed. In fact, the video is a great job of if you want the lowest possible latency,
you should be using that chip and if you want the highest possible throughput, you should probably
also be using that chip as a right now. And it's all comes down to like this is a very new paradigm
of programming. One of the things you hear is that the market for the best chips Blackwells, let's say,
is like a drug market or something right now. There's all sorts of fascinating things happening
to get as many of them as possible because everyone's so short. It would be to react to that analogy
is that what it feels like. Then also to talk about what the market is like for like not the
bleeding edge chips. If I am willing to accept a slightly or moderately inferior chip, what's that
market like? Let us into that world. Media has a long-term view on all their chips. They see
this immense demand for the Blackwell chips. They can do what other suppliers have done in the past,
which is crank prices, meet the market, supply and demand curves will correct. They'll intersect
at some point and everyone will be technically happier. But in video sees, if they just let the most
deep pockets by all the chips that maybe hurts them in the long term if that customer ends up
accruing a lot of more power, they understand the compute is power today. They're fried strategic
about how they allocate compute. That's the first thought. The second thought is that relationships
matter a lot. Nobody wants to have a huge order of a chip rental come in from this new startup
that says, "Oh, yeah, I'm going to rent 10,000 Blackwells for three years or five years."
This startup has only been operating for months, typically. Who knows whether they're good for the money?
The way you convince someone to give you access to compute is quite challenging these days and
requires some pretty great relationships or just incredible financial backing to make this happen.
On the Nvidia side, and it's all because this scarcity is so high and demand is just off the charts.
Now for other chips, I wouldn't even call them inferior. I like to say there's no bad chips,
there's only bad pricing. I will make any chip work at the right price. That's one of the
ethos of the company. Let's talk about AMD. AMD, I think great chips overall, the challenges that
people don't understand how to program them. Very well. I've been talking to you about how we
have a great kernel team or so serious about squeezing the performs out of the hardware. Nvidia is
pretty good at doing that for their own chips, frankly. There's some alpha that we can squeeze out,
but there's a lot more to be done on their chips because the vendor does a little bit less work than
the Nvidia does to make the best kernels out of the box. Or even better, there is alpha and just
other people have this perception that AMD is not as good as Nvidia. That's music to my ears. I'm
very happy for them to sleep on this chip and for me to buy as much as I can. Now I think that's not
actually super true anymore. I think AMD is actually somewhat popular amongst some large buyers.
I think publicly meta and opening, I have bought a ton of AMD chips. We're increasingly seeing that
all the AMD supplies also being allocated, but there's a long tail of other companies that are
popping up and that new companies are great, such as Edge, or Sombanova, or Dmatrix. All of these
companies are popping up. I think the main challenge for them is scale. Can they actually get enough
wafer allocation from TSMC to pump out chips to making it into the market? If there's a new chip
on the market, I'd like to know about it as quickly as possible and evaluate whether we can buy
a good fraction that's apply. And so it's fundamentally an arbitrage for you. If you can be much
better at eking out performance from chips that have received less attention, you can then resell
that at a margin and it can be a great business. Exactly. And I think that it's not the case that
everyone else has a skill issue that they can't make these chips work as well. I think we're quite
competent at this. I think we're probably one of the best teams in the world to use multiple
silicon architectures and be quite aggressive in chasing down performance in unlikely places.
But yeah, I think it's the speed at which we're willing to kind of build our stack around a new
chip. We don't have a huge amount of incomeancy around while our data center providers are only
stuck with this class of chip and it's going to be a huge pain for us to deploy these net new chips.
We have some very creative data center partners who are willing to move very quickly. And there's
a new class of this that we can talk about. Most importantly, we don't try away from the challenge.
That's frankly a big part of this is just saying, yes, we love TPs. We're going to make TPs work.
Yes, we love training. We're going to make training work. And if it doesn't work easily,
we're going to find a way to fit it in with the heterogeneous serving system. It will have a place.
Every chip has a comparative advantage. We had to find that advantage and then squeeze it in that
direction. Just as an interlude before we get to hardware data centers, energy, etc. Which will
be really fun part of the conversation. I'd love you to talk about your perception of
the investor classes worry. Like you look at memory stocks or my current favorite is you look at
the chart that plots the percent of the S&P 500 that semiconductors. Historically it was like
two, three, four percent. Now it's 19, 20, 21 percent. And it just sort of looks like if you're a
student of market history, you get all these things through time that are sort of reach some crazy
near-term peak and then collapse back to long-term norms. That has all investors worried. A lot of
people have made a lot of money in micron and SK Heinix and companies like this. But everyone feels
like on-term like computes a commodity. They will not represent a quarter or fifth of the entire market
capitalization of the world. They're scared and that's the setup. Everyone acknowledges that
there's a huge shortage. Everyone sort of feels like we'll figure it out and these things will
revert back down to their normal place in capital markets. I'm curious what you think about that
narrative. I'm less of a student of history as more of a member of history. I was born in 1997
and my mom worked in Intel. They run up to the year 2000 and the dot com boom and crash. I remember
the time where Cisco was the most valuable company in the world. And in Intel was close behind
and mostly draw parallels to that period of history from 25 years ago to today. And I think the
main difference is that a lot of the investment in networking equipment historically was speculative.
We anticipated this future demand for users that never came. And I think what's interesting about
token consumption or AI consumption broadly is that it's no longer speculative. People buy tokens
because they're immediately valuable to them. You don't hoard tokens. You use them immediately.
This is also even different from what we had two years ago where there was a supply crunch for
hopper generationships in 2023, 2024. In that period, it was all training oriented spend.
And training is inherently speculative. Now it's ever instituting caps on how much you can spend on
a cloud code. It's a very, very different world to be talking about inference spend and predicting
inference spend to go up. I do think inference spent monotonically increases. There's no speculation
on inference. Coming back now to your take on hardware, the unit level is interesting to me.
Talk about chips, talk about racks, talk about clusters. I'd love to talk about data centers.
You said you've had some interesting partners doing some cool things. Talk is about the present
and future of data centers as you see it because of this seems like obviously a critical thing
for being able to serve all day.
This inference has lots of innovation in this part of the world and obviously you're focused on it.
I think one of the themes in our conversation has come back to what is training versus inference.
So what is the difference between these two categories?
And you know what was different about two years ago being training oriented and today being inference oriented?
And I think the most conservative players in the entire ASDAQ have got to be the infer players,
whether the data centers or even more conservative SMC that ship in for people.
Data centers, they ever built for training.
Training is a superset workload over inference.
You can make any training cluster work for inference, but maybe not vice versa.
The difference there is networking, how much do you invest in bandwidth between ships and how large of a cluster do you need?
There's actually a dis-economy of scale to data centers in some way.
It's way more expensive and difficult to build a 100,000 GPUs in one data center than it is to build 10,000 than it is to build 1,000.
Now we just talk about how many megawatts or gigawatts do you have?
And basically there's no way to build a gigawatt data center in the United States easily anymore.
Even 100 megawatts is increasingly hard.
It's basically impossible unless you're a very special set of customers.
10 megawatts is probably on the edge of what's possible today.
And one megawatt I would use plentiful.
So there's this incredible floor on the market where you can find lots of aggregate power, but it will not be concentrated.
And that was not interesting to anyone who's building training data centers because you just assume the stallion on spot for no one wants to deal with cross data center training.
And the market has lag in it. I think that the market still assumes that we have to go shake down those 100 megawatt and 10 megawatt data centers wherever we can find them.
It's still the attitude I hear from a lot of data center developers, but increasingly we're seeing a few new thinkers realize that inference is going to be suitable for these distributed 1 megawatt data centers.
We're quite in agreement with that and we are very happy to buy small pools of compute across the United States and use that as our inference fleet.
Give us a sense of a literal physical size of one megawatt versus 10.
Yeah, well, so this got really wonky with the advent of liquid cooling.
Now you can pack insane levels of power density into a single physical rack, a megawatt of compute.
You'd imagine this massive data hall, like a huge warehouse basically.
And now you can actually pack that into around like eight racks for the compute.
Each rack is about the size of a refrigerator.
You can just imagine eight of them lined up.
Your view would be that the future that you want to help build is a whole bunch of different chips that can be used together that you can buy.
You're a buyer to eke out the most per chip and that those chips can then be coupled in very small data centers to just do inference.
Those two steps of a whole bunch of random compute, some of which is cheaper than it should be your ability to eke more out of it.
And then small units of expression in the data center equals way cheaper intelligence.
I certainly think so.
Yes, there's a lot of ways to access cheaper flops if you're able to be creative with what you take.
And so one of the ways that I describe what we do is we will buy any chip anywhere in the world for any duration of time.
That is a level of flexibility and liquidity that no one else has right now.
We're very aggressive about putting our money where our mouth is and we will take any capacity and find a way to make it work in our fleet.
And that is a big part of our advantage today.
Long term, we had to create more of that advantage by investing in these data centers that other people are going to be skeptical of because what's going to happen when you set this army of a thousand small data centers versus the one big gigabyte data center.
Few things. You're not going to have power redundancy quite often.
You're not going to have backup diesel generators on site.
Those are all very expensive. We cut all that overhead.
We're not even going to have we're done a networking in a lot of cases.
We're going to put these in facilities where we have good access to power, a single source of power.
And we're going to trench one line of fiber to these data centers.
But we're not going to have three lines of fiber with redundancy and failover in SLA's.
She's going to go down sometimes.
I won't be surprised if some of them get down to like 95% up time, which is bad.
Very bad. That's fatal, atrocious for anyone else.
Couldn't survive in a big gig last night.
You have basically zero buyers for a data center that is 95% up time.
I'm that first buyer.
I will buy that 100% up time.
And the reason for that is because of this background engine thing.
There's things running in the background, you don't care. Partially, it's actually two things.
One is that we have a really robust control plane that is going to be fine handling any
single failure in any single data center as long as it's not correlated with other data centers.
And I can just move the workload somewhere else.
I'm cool with that. The failures happen at some rate.
And I am basically linearly happy with a data center that is 95% up time, which is 98%
versus 99%.
It's just linearly good or bad for me.
You do need that async piece that I mentioned.
We serve these long cryon agents because what happens when a request fails is that I'm
going to have to go find a new GPU to put that request on.
And that means that for that single turn of the agent's work, it's working for an hour,
but then hits a roadblock because it's GPU got pulled away.
In that moment in time, that agent is going to experience maybe like an extra minute or
two or three, maybe even 10 of latency.
My argument is that my customers don't care because their agent was running for hours.
They're sleeping.
It doesn't matter if it's a matter of like a single turn occasionally becomes a little
bit longer.
We tell our customers like our average throughput is going to be very competitive, but our
P99, our 99% tall latency is not going to be controlled.
It cannot be.
And in return, I'll give you unbeatable economics and I think that's the right fit for
background agents.
Talk about power as a category.
What the innovation that you're seeing is where you think it goes from here.
What are you seeing that's interesting, innovative, what do you think this goes?
Okay.
So I said, I want 95% up time on my data centers.
But I even take 80% up time at the right price, probably.
And what does that mean?
Well, I'm a Senate California.
I love solar and wind.
I think solar and wind power is way undercapped in the United States.
And the challenge has always been this intermittency.
You would even consider solar and wind unsuitable for data centers because you have a persistent
base load and an intermitting power source.
What are you going to do?
Well, I think we're actually not that far from solving that problem.
I am totally capable of tolerating a outage from a data center that's measured in even
days or weeks, which is like the worst case nightmare scenario for a data center is that
we're going to have a long term outage because the wind is in blow and the clouds are in
the sky.
Fog is hanging over the valley for some time.
That's the worst case scenario.
It's in fact highly predictable.
And I can just call in capacity in some other place of the world whenever that happens.
I'll just model the weather and figure out when my data center is going to be offline,
move my work when somewhere else.
And it's fine.
The trick is that it's going to give me better access to power that no one else is going
to touch because it is so annoying to deal with that kind of outage.
One of my chips are cheap enough.
They're probably not going to be in video racks.
I don't mind the capital cost of having idle chips.
I've heard you describe this entire system as like a scavenger strategy.
Yeah, that's right.
Impact that analogy a little bit.
Well, first we scavenge chips and then we scavenge power for those chips.
The idea is in both cases, I do not want to be bidding against anthropic or up an AI.
For it could be capacity, I'm not going to win against them and I don't want to.
I want to be more creative and use the supply that they don't find legible today.
And over time, I am asked enough aggregate supply.
I'm never going to get concentrated supply.
I will always get aggregate supply.
And over time, I build my aggregate factory that is unbeatable in economics.
We are building a factory.
We're trying to build the best steel factory in the world.
But it will come through mini-mills, not through large, monolithic steel plants.
And if I imagine the different versions of this, how vertically integrated you can be,
the extreme would be you own everything so that it's a very capital intensive business.
You own the power source, you build the data centers, you design your own chips, you control
the software that eaks the most out of those chips, and you sell the end finished token
to your user.
Your user is me and you just own the whole stack.
But you can imagine many other permutations of the business where draw the line anywhere.
You can be incredibly capital-like, own nothing and just be like the coordination plane
across all this stuff, the virtual scavenger.
What do you think about that question of which type of these businesses to be?
You know, there's actually two parts to me to receive that question.
One is the CEO of a company that needs to work every day and grow sustainable and as
quickly as it possibly can.
The other is a founder.
And the founder is much more imaginative and loves this stuff.
The founder in me wants to do everything.
This is my entire life.
I said my entire life thinking about chips, power, energy, all I care about is this stuff.
So of course, I want to be maximally ambitious.
I don't want to stop ever.
I will never stop until I have built the most efficient system from soup to nuts.
You're doing real life factor, basically.
Very much so.
Very much so.
That's like the emotional from the hard answer on the CEO side.
I think we have to be more pragmatic.
The capital we're looking at for owning everything is insane.
Software has high leverage, so we have to start with software.
But ultimately, do we own power generation or can we get great power purchase agreements
with utilities?
I'm more inclined to pursue planning other people specialize in the things that there is
to work we get at and then see if we can get to the scale.
I think of it as like, I want to get to the scale where I earn the right to take this
under our wing.
I absolutely think that there is efficiencies to be gained everywhere in the stack.
If you can break the assumption that people I would be buying from, they made assumptions
about who their customers would be.
And I maybe break those assumptions.
It's a pretty optimistic view.
I think it's only possible because we're actually trying to underwrite the largest market
for compute in the history of computing.
We're actually going to build so many millions, trillions of dollars of investment into inference.
Because of that focus, it makes sense to build a lot of things that are custom for inference.
And it's my job to seek all the places where that's possible.
And then as it become obvious to me and my partners, they'll look at my partners to build
custom things for me.
And if they can't do it for me, I will do it myself.
If you had to just zoom out of this entire system, software hardware, energy, et cetera,
and stack rank the places that you think that we are the most inefficient today at producing
useful, intelligent tokens, what does that list look like?
I think compute scaling is actually very efficient.
You give me more flops and I will use more flops.
I would say we're actually fairly judicious already with our use of flops.
If you look at a modern MOE model, there are very few models that are more than 10%,
dense, meaning 10% of the possible number of experts you can activate are activated.
And I think the frontier models are close to like 1%.
Fairly sparse already, I don't think that we're wasting too much on the MOE side.
People have been working with MOE's for quite some time.
They're pretty good at squeezing MOE's.
Where we are not good is attention.
And it's use of memory, specifically.
The KV cache is quite uncompressed right now.
I think if you look at the entropy in the KV cache, it's not earning its keep.
Like we're storing many kilobytes of data in the KV cache per token.
And that's probably off by an order of magnitude or two.
I don't know the frontier labs do, but DeepSeek certainly publishes really interesting work
to compress that further and further.
And they're making it progress.
And I think the fact that they're able to make order magnitude and progress here every year or so
signals that there's a lot more room to go.
If you zoom out further, I think that we actually don't marshal our compute effectively at all.
We have all this compute in the world.
In reality, it's pumping out 5 million black wall chips this year.
Where are they all going?
Are they all being used at all all the time?
I certainly doubt it.
I think that at some level, we just need better orchestration of compute across the world.
This is very difficult to do because a lot of the compute disappears into private pools
or compute that we'll never see the line of day.
And those GPUs sit very sadly idle.
It pains me physically to see that those GPUs are just silicon and power going into that.
And it's just sitting idle.
And I want to fix that.
How we organize and orchestrate the world's compute as a shared resource and pack it more
efficiently.
I would estimate that we all make fun of XAI for having some challenges with total flop
utilization on its clusters.
The reality for the rest of the world is as far worse.
A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific
customer.
Just don't get utilized.
You're attacking the efficiency of that very directly.
That's way more effective, yeah.
What about FABs?
What do you think is the future of FABs themselves?
I think everyone is wondering, will the memory companies, will TSMC, will Intel and others,
how will they expand capacity, basically?
Well, we do it here in the US.
We have fun fabrication of chips themselves, like if we could just snap our fingers and have
100 times the chips in the stock today, we'd probably have way cheaper tokens.
That seems like an important part of the universe to hear your view on.
Everything grows in balance with each other.
If we snap our fingers and double all those things, you might fix a TSMC bottleneck.
You're just going to run into another bottleneck.
You make 20% more chips than you have another bottleneck immediately.
I will say though, it is interesting what they consider to be a must deliver, like an
invariant that their customers, me, are always going to want versus what I think as a more
fluid relationship.
I think that if the FAB exposes more of their trade-offs, to me, I'm able to make more
intelligent decisions about what I can do.
One of the most interesting examples here is that any FAB has a lot of spread in their
worst chip that comes out of the production line and the best chip that comes out of the
production line.
There's a lot of variants in how chips are made.
The question is, if you have a company like TSMC, they work very, very hard to tighten
what we call these process corners.
We want to keep the worst chip as close in characterization to the best chip, and then
they go to great lengths to make that possible.
That means that they are adding a lot of controls in the process that maybe I don't need.
Maybe I'm actually willing to find a place for that worst chip.
You don't need to tighten the process control as much, which takes more time and cost.
Maybe I'm willing to take a lot more rejects.
I think for us, it's like a more holistic optimization around cost of the dies, supply
of the dies, and then the cost of power and places we can put them.
My whole goal is to so dramatically expand the supply of power across the United States
that I have a home for a lot of ships that otherwise would not have earned their place
in a data center.
Can we talk about how you design the system of your own business?
What lessons have you learned?
You talked about some interesting and video lessons.
Bring me into the culture and how you structure a team and a business where this is the
North Star.
There's a lot of in the limit thinking we don't worry about the immediate nature of when
we start working on a model, the efficiency is not going to be very good.
We think about where we could end up in like a month, six months or a year's time.
We don't accept the state of the machines we work on as fixed.
Even something like the Blackwell chip, if we think that there's some bottleneck that
is holding us back from achieving this performance, it's very important to me that we understand
and characterize that very well and write it down so we can both a tell and video about
it for friends and also to basically keep this in mind for future ships that we buy.
We want to learn things that are invariant for us or the company long-term and kind of
fold that into future decisions that we make.
We're very collaborative.
I think one of the most important traits that we look for are people who either who are
both good students and great teachers.
A lot of our people on the team were TAs in college and loved the experience of sharing
knowledge in this way.
We do whiteboard sessions all the time.
The collegial environment where everyone has something to teach and something to learn
is extremely important for us.
What are the attributes of people that you would want to hire that you think will be resilient
to the work environment three years from now when more stuff is handled by machines?
Curiosity.
100% curiosity.
The one thing I cannot teach is love for performance.
Love for digging into every microsecond.
The machine is working and understanding what's happening on the machine at that time.
That to me is the most important trait for a performance engineer is what I look for.
I don't look for lots of AI experience.
I don't look for CUDA experience at all.
That's actually a huge red herring.
CUDA has a concept or GPS is a concept of evolved so much in the last five years.
It's no point asking for 10 years of experience.
I want to teach that, but I cannot teach the love for performance engineering.
That is what I seek.
Can you give your assessment of the major labs one by one, but also then the relationship
of close source as a category to open source what you think is happening and will happen?
In a line, I would say the labs pay an immense premium to be three to six months ahead of
everything else.
I think that's probably still worth it.
I think it makes perfect sense for it.
And I'm dropping to do what they do.
There's a sensitive topic around distillation, which I think is a very core piece of the
relationship between closed and open frontier.
I'd like to offer an alternative view on that, which is there is the sense that distillation
is theft that you are taking something from the frontier models when you distill on their
outputs.
And in fact, even if that's not your intent, even if you don't ever try to scrape data
from anthropic, one thing I'll offer is that an increasingly large percentage of the artifacts
we put out on the internet are agenerated.
You just look at GitHub alone.
What percentage of repos created in the last year?
Do we think we're created by cloud code?
Do we consider that to be distillation?
Because it's probably all we need.
I would not be surprised if you could train a fable class model only on the outputs of
code you consider good on GitHub that's open source.
And certainly if we take the position that users own the outputs of their interaction
with AI, and they choose to put that up on GitHub, which a lot of them do, we're going
to have latent distillation for a long time.
It seems fundamentally possible for me.
I don't think it's fundamentally possible to prevent the diffusion of information or
model capabilities.
It will happen.
The question is just how fast?
And so then the question becomes do scaling and improvement laws hold forever for a really
long period of time?
And if they do, then there's value of being three and six months ahead.
And that was just last as long as it lasts.
And they can charge a huge premium for those tokens relative to a very cheap open source
token.
But the right way to think about it.
I think it's possible.
I don't know the premium for being three to six months ahead is going to last that long.
I mean, if you look at like enterprise deployments, they don't move at three to six months speed.
A lot of enterprises are probably still handling four six.
Opus four six or opus four seven, you know, adopt the bleeding edge rapidly.
There's a lot of questions that people have around rolling out any change at all.
We're just so early in scratching the surface that I don't think there's any way to call
a winner in this race.
Certainly, I don't even think this is a race that can be decided ever.
It's a continual process.
Fundamentally, I don't think open source ever goes away.
If there's a vacuum because you know, one leader steps out, a new leader will step in.
There's too much incentive and too much tailwinds, too.
It gets easier every day.
It should train a frontier class model.
So your hope of what the future looks like is what balance between closed and open?
What balance between model companies doing everything, because they have the advantage
of owning the stack or whatever, I'm entropic, can do that.
It's like the new Google will just do that or something.
What do you hope the future looks like?
I want abundant tokens and diverse harnesses.
I want everyone to build their own harness.
Every company, every user, even, make the engineers your own, or not that far away from
that level of customization capability.
I want people to own their intelligence, and I want that intelligence to be customized
probably not through weight fine tuning, but probably through more in-context learning.
That's a more technical detail, but the underlying input to this abundance future is about cheap
tokens.
My job is to make the tokens as cheap as humanly possible.
I will achieve that, and I will do it through every layer in the stack available to me.
I love the supply side levers.
I will use every chip.
I'll use every source of power, and I'll use every piece of land in the United States
that's suitable for this.
In return, people will have the incentive to explore what it's like to have abundant intelligence.
We still treat the agent as a person that is expensive to consult, and you should ask
them when you have a hard question.
That's not the way to think about intelligence.
It's incredible that the machine can think, and we should try to get that into as many
hands as as many people as possible.
You sit in such a unique seat, and you have such a unique perspective on what you're
trying to do to make this future a reality.
What do you think are your most divergent views of the world versus your friends who are
really well informed and interested in this stuff?
What ideas of yours make your friends look at you like you have through heads?
Most of the ideas on chips, I would say, you know, when I talk about building custom
chips, and they ask me, "Oh, so what's different?"
It's about side-stepping the HBM shortage and focusing on more extreme offload to other
forms of memory, such as flash.
I'm quite passionate about that idea.
Everyone on my team knows that I keep banging the drum over and like, "What do we have
to change about the model architecture to make offloading KB Cash to Flash work at a
much greater level?"
That's in the community of inference people.
We have some divergent views on what you can do if you design a system around serving
at one to ten tokens per second, which is our whole North Star.
More broadly, I think there is this larger sense around how do people consume a trillion
tokens per day?
That's the world we want to create, the capability for them to do.
do that. What's the trillion tokens like ground to see how much that is? A trillion tokens,
well, okay, an opening eye pricing, that's at least $5 million at the very last for
$5.5 or $5.6. Yeah, I think the dollars are probably the most good metric. Yeah, it's
million dollars. Yeah, so what's the world in which we consume what currently cost $5 million
per person per day? We were asking for at least three to six orders of magnitude improvement
in cost per token. Get that into 5,000, you probably have some customers. And in fact,
I would argue that for some size of model we are approaching a trillion tokens being measured
in $10,000 dollars. And that's something that you could imagine running for a single job.
Are you at all worried that just like the average person just can't and won't do that,
doesn't do that now with their own brain? There actually isn't that much demand for intelligence
in the world? I never will believe in that. There is always demand for intelligence in the world.
The on-ramps to that intelligence are our challenge as a product community. I'm not a product
of my person, so I cannot say I had the best vision. You want to enable those people. I want
to enable those people. I want them to never be held back by the sense that my free tier users,
I can't afford to give them as many tokens. And I hear that from my customers all the time.
We want to fix that. What about the inverse question, not what you think is craziest, but like what
consensus thing you think is wrong? One of the things that he coming back to is this question of
Nvidia. I am bullish on Nvidia in the short term and Nvidia, you should never bet against them,
they're always going to reinvent themselves fundamentally. I think one thing that surprises people
is when I tell them that, hey, if you look at Hopper and a Blackwell to Rubin and you compare
like for like, what is the performance per watt of B-float 16 multiply? It hasn't improved all
that much. Or you take that one step further and go to TSMC. If you look at TSMC 5 nanometer versus
4 versus 3 versus 2, the performance per watt on these chips doesn't change like a dramatic amount.
The consequence of this is people lose their minds of regime politics. What would happen if we lost
access to TSMC for any reason? My contrarian take is that it wouldn't be that bad. The supply would
take a shock for sure, but the best processes that we have in the West, like Intel, not that far behind,
at worst, like maybe 2x, worst performance per watt. The gap is just far smaller than you would
make it out to be if you follow like the chipboard dialogue. What else is happening in the AI world
that is not in your path, meaning it's not like a component of this whole system that you would
end up doing something in that interests you most? Well, we're fully downstream of models.
So the model people get to decide how to design their architectures. I don't have any input to
up an AI or Anthropic, but I can only pray that they go in the direction that is amenable to me,
or I have to like do my best to predict where I think they're going to go and build my serving
architecture accordingly, both software and hardware choices. They have I think the most interesting
game in some ways to play. Once again, this game actually like the profundity of the machine thinking
and how consequential it is to decide to use something like sparse attention versus dense attention
or how consequential it is to like use a different data type. We were training in B4-16, but now we
can train in FP8 or FP4 or lower precision data types. That is just an arbitrary choice it feels
like, but it has profound implications for what chips I can use and how I should build my hardware
and think about the future of compute. If you had a hundred entrepreneurs in a room, all of whom
wanted to create some new compute startup, and let's say they were specifically wanted to make
hardware chips or systems or racks or whatever, what advice would you give them on how to
orient their companies or like the type of company, not the specific choice they're making on a
tech bet or something like this, because it seems like we're going to try everything and that
will be great for the world. Some stuff will work, but if you had to give them advice on how to
orient their business to be successful in this coming world, what advice would you give them?
It's all about the bottlenecks on supply chain, so you need to first convince me or convince
an investor that you understand the three to five bottlenecks that dictate modern chip supply. There's
TSMC wafer capacity, there's HPM capacity, and there's advanced packaging, maybe a fourth one would
be power. Where will you get the power? How will you build these racks? You should have a great answer
to each of those four bottlenecks and how you're going to work around them because it's all
arbitrary to the end of the day. You're building a chip because you think that NVIDIA has made some
choices that are difficult for the change, which is true. NVIDIA makes a lot of choices that are
difficult for them to change. They're not perfect. They're just really well balanced. You want to be
spiky. You want to pick something and say, I think they've underpriced the impact of how short
and we're going to be on HPM. We're going to push really hard in this other direction instead,
which I do think is probably the thing to attack most. Why? There's no easy way to bring on a lot
more fabs of memory. So it's going to be a while until we have. Yeah. Then the boys and boys
don't love huge catbacks for cyclical. They've been burned on that many times. Conceivably like
because of that shortage, the world is just going to route around it by making everything else in
the system more efficient. They're going to make everything else more expensive. I think that
iPhones will cut their memory. iPhones are going to go up in price. We're just going to deal with it.
Why doesn't NVIDIA go all the way to the end and sell tokens, do you think? NVIDIA is really
smart about those. They don't compete with their customers. NVIDIA takes a long view on everything.
Why don't they even start with the NeoCloud? Why don't they just sell, compete at the back door?
Well, NVIDIA is really good. Jensen is really good at making his friends billionaires.
He's made a core weave, a billion dollar company, a many billion dollar company, and there's no need
for him to destroy that goodwill. He wants to create a diverse community of NeoClouds and
inference providers who are all jockeying to create demand for NVIDIA such that if any one of them
decides to, I don't know, vertically integrate or go with AMD or any other option, he's got three more
people hungry to fill that position. It's great to have competition amongst his buyers.
My favorite closing question for everyone is what is the kindest thing that anyone's ever done for you?
My immediate first thought is like all the mentors that I've had over the years. It's a rare person
who takes a lot of time out of their schedule and makes it like their personal interest essentially
to make sure that you understand something or teach you something or ingrain some value in you
that they think that you're on the cusp of understanding but just push you over the line for
understanding. A lot of the people in NVIDIA that I mentioned earlier who instilled that love
of performance engineering in me, but also my professors in college who I remember like my advisor
in like sophomore year, I was very impatient student as I showed up at his office hours and say,
I want to build AI trips. I know what I want to do. Why am I wasting time taking all these like
other basic classes in networking and operating systems? And he just laid out basically like a whole
stack and showed me the beauty of understanding every piece in the puzzle. He took my entire path
of like trying to focus on one piece of the system and said that it's so rare that someone can
actually understand the entire stack from the gate level silicon all the way to building a great
internet skill service. You should aspire to be someone who over the course of your lifetime
achieves that level understanding. It is such a rare trait that level expertise is so noble
to chase and that stays with me quite a bit. Amazing conversation. Thanks a lot for your time.
Thank you so much for having me. If you enjoyed this episode visit Colossus.com you'll find every
episode of this podcast complete with hand out of the transcripts. You can also subscribe to Colossus
our quarterly print digital and private audio publication featuring in-depth profiles of the
founders investors and companies that we admire most. Learn more at Colossus.com/subscribe.
You know how small advantages compound over time that's true and investing and just as true in
how you run your company. Your spending system is your capital allocation strategy.
Ramp makes it smarter by default, better data, better decisions, better economics over time.
See how at ramp.com/invest. As your business grows,
vanta scales with you. Automating compliance and giving you a single source of truth for security
and risk. Learn more at vanta.com/invest. The best AI and software companies from Open AI to
cursor to perplexity use work OS to become enterprise-ready overnight, not in months. Visit workos.com
to skip the unglamorous infrastructure work and focus on your product. Ridgeline is redefining
asset management technology as a true partner, not just a software vendor. They've helped
firms 5x in scale, enabling faster growth, smarter operations and a competitive edge.
Visit ridgelineapps.com to see what they can unlock for your firm.
Every investment firm is unique and generic AI doesn't understand your process.
Rogo does. It's an AI platform built specifically for Wall Street, connected to your data,
understanding your process and producing real outputs. Check them out at rogo.ai/invest.
Podcast Summary
Key Points:
Neil Mova, founder of Sail Research, is building a "token factory" focused on serving open-source language models at the lowest possible cost, prioritizing background agents that run for hours or days over real-time chatbots.
The company emphasizes throughput over latency, leveraging a trade-off inherent in GPUs: optimizing for speed reduces efficiency, while maximizing throughput (via batching) lowers costs—a strategy suited for long-horizon tasks like deep research and cybersecurity.
Mova's background at Nvidia, where he worked on tensor cores and pursued "speed of light" performance, shapes his approach: building software stacks that squeeze peak efficiency from chips, including less-popular ones like AMD, to gain a cost advantage.
He predicts a shift toward 90% background inference workloads, enabled by cheap tokens, enabling proactive agents that operate on human timescales and tackle verifiable problems, while leaving non-verifiable tasks like taste to humans.
The strategy involves "scavenging" undervalued chips and power sources, including renewable energy with intermittent availability, and accepting lower uptime (e.g., 95%) in small, distributed data centers to cut costs, which is feasible due to the asynchronous nature of background agents.
Mova views the future as a hybrid of closed and open-source models, arguing that distillation is inevitable and that the premium for frontier labs' 3-6 month lead may not persist, while emphasizing the importance of curiosity and performance engineering in his team.
Summary:
Neil Mova, founder of Sail Research, is building a "token factory" that provides the cheapest possible inference for open-source language models, specifically targeting background agents that run for extended periods rather than real-time chatbots. His core thesis is that the future of AI lies in long-horizon tasks—like deep research, cybersecurity, and proactive personal assistants—where latency matters less than cost, and agents operate on human timescales without needing constant human input. This shift from latency-optimized to throughput-optimized systems exploits a fundamental GPU trade-off: batching work increases efficiency and lowers costs, but sacrifices response speed.
Drawing on his Nvidia experience, where he learned to chase "speed of light" performance, Mova's strategy involves building software that maximizes efficiency on any chip, including undervalued options like AMD, and "scavenging" resources others overlook—from less-desired hardware to intermittent renewable power sources. , 95%) because background agents can tolerate failures and reroute work, enabling access to cheaper power and land that larger players ignore. Mova predicts a future where 90% of inference is background-oriented, unlocking abundant, low-cost intelligence for verifiable problems, while leaving creative tasks to humans.
He also questions the durability of frontier labs' premium for being months ahead, seeing distillation as inevitable and open-source models as a resilient counterweight. Ultimately, his vision is one of intelligence abundance, where tokens are so cheap that proactive agents become ubiquitous, limited only by the questions we can ask.
FAQs
Ramp is a platform designed to make finance teams leaner, faster, and better, saving businesses 5% annually on average. It is used by companies like Visa, Cursor, Stripe, and Shopify to stay focused on growth.
WorkOS provides APIs for enterprise features like SSO, SCIM, RBAC, and audit logs, enabling companies to become enterprise-ready on day zero. It is used by AI teams such as OpenAI, Cursor, and Anthropic.
Felix Byrogo is a personal finance agent that turns a single prompt into client-ready work using your firm's templates and standards. It can create PowerPoint decks, Excel files, and source research by emailing instructions.
Sail Research is a token factory that provides an API for serving open-source language models at unbeatable prices. Its goal is to be the absolute cheapest provider of inference, focusing on long-running background agents.
Sail Research believes the future of AI is long-horizon background agents that run for hours or days, where latency matters less than cost. This allows them to optimize for throughput, making tokens cheaper.
The KV cache stores a representation of every token in a conversation, allowing the model to reference past context. It can grow larger than the model weights, impacting memory usage and performance.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.