OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha
from Latent Space: The AI Engineer Podcast
80m 43s
Open Router was born from a deep understanding of the critical gaps in early AI development: developers needed accessible, reliable, and diverse open models without being locked into proprietary ecosystems. The founders identified that traditional model deployment was fragmented and inefficient, with researchers failing to translate model training into usable developer tools. Drawing from early experiences with crypto communities like mid-Journey and Discord, they realized that real-world adoption required not just models, but a marketplace where developers could easily discover, test, and deploy them. Open Router’s core innovation lies in its PubSub-inspired architecture—enabling continuous data publishing and subscription—alongside a neutral, developer-first platform that provides seamless API access, versioning, and governance. Unlike closed systems, it empowers users to switch models dynamically, creating a feedback loop that drives innovation. Early skepticism from investors, who saw AI as a monopoly risk, was countered by the emergence of open models like Alpaca and Mistral, proving that decentralized competition was not only possible but necessary. The platform’s success was amplified by its ability to serve both end users and developers, with communities like mid-Journey proving the demand for visual, interactive AI tools. Crucially, Open Router’s value wasn’t just technical—it was strategic, offering a scalable, transparent layer that democratized access to AI. This combination of community insight, technical design, and developer focus helped it grow rapidly, eventually becoming a central hub in the open AI ecosystem. The journey also highlighted how focus—on specific user needs like coding or moderation—was essential in navigating early competition, proving that clarity of mission drives sustainable growth in complex, fast-evolving fields.
(upbeat music)
- Okay, we are here in Anjus House.
(laughing)
It was just where all the start-ups in San Francisco start.
- Howdy.
(laughing)
- And congrats on cursor, Mr. Al.
I don't know what else you got so much stuff going on.
(laughing)
There's a lot going on.
- Open router is probably the most I would say,
one I'm excited about, Chris.
- Yeah, yeah.
And we have Alex, first time on the pod,
but you've been there.
- Thanks for having me.
I appreciate every time you've shown up for the community.
Congrats, I just, like, what a journey.
When I was looking back at your past post,
one of the earliest principles that I saw you write
as a sort of product person,
is PubSub as a product principle.
And I wanted to, you to maybe explain
how you think about what should exist in the world.
- Yeah, the PubSub piece, which was early 2023,
I didn't think about it until we talked like 10 minutes ago,
is about how there is like a way of thinking about products
as an intersection between subscribing to data
and publishing data.
And marketplaces are an easy, easy, easy example of this.
You have suppliers that are publishing some kind of product
to a skew.
And the skew is kind of like a PubSub topic
that a consumer is subscribing to
and just going to like consume whenever they want.
And humans consume in a very like a discrete ad hoc way.
It's not very scalable.
All their attention is on the topic
when they're buying the thing
and their attention is nowhere else when that happens.
Agents and consumers of inference don't act like that.
They're consuming continuously
and they're changing the skews
that they consume from all the time.
So open router is sort of like a blend between a normal API
experience and a marketplace
where we create model slug.
We have the auto router.
We have all kinds of like product skews
that you can subscribe to.
And then you can like continuously at like a derived value
and make decisions based on those consumers.
Yeah.
This is something that was more consensus now
but not consensus when you guys started
which is that there is such a demand for swapping models
and changing things out and that people would not use
the native SDKs.
I guess for each of you,
what was your sort of realization moment that this would be it?
You've given a talk at the IE about Alpaca
as like one of your inspiring moments.
Alpaca, I can like rehash the Alpaca moment for a sec.
Like the very beginning at the end of 2022,
opening I was the only game in town.
There was like opening I cohere
and then a smattering of like early attempts
at open weight models.
Yeah.
When Wama came out in January of 2023,
it was like well, really exciting.
This was really big.
It outperforms GPT-3 on one or two benchmarks.
But you can't chat with it.
It wasn't like an actually an engaging model.
But it seemed like someone just needed to fix
a couple of things and do some RLHF on it
to get it all the way there.
And Alpaca was the first model that I saw that did that.
It only took $600 to do.
A team at Stanford generated a bunch of synthetic data,
fine-tuned Wama, and made Alpaca,
seven billion parameter model.
It was maybe it was 13 billion parameters.
And it was so good.
Like I was just like on an airplane using it.
I, you know, in many cases,
like you could not discern a chat GPT versus an Alpaca result.
And I figured it was this easy to make a model.
One, we have a whole new way of monetizing data
for the first time.
You can just like take really valuable data
and turn it into a service in $600.
And that cost will probably go down over time.
- Two.
- When you say monetizing your data
as what eventually will become an MCPM point
or as a training data for a model.
- Yeah, training data for a model.
By an abstract way of saying like, hey, I have this data.
- I press it into a model.
- Like it makes sense for me in my product.
But like I could repackage it in the form of a model
and sell it.
And so it's just a whole new business model for the economy.
It also of course provides like, you know,
a way of following what frontier labs are doing.
But in a way that like a single developer
or a small team of developers can rule on their own.
And so that whenever you have an example of that,
like a breakout app that's doing really well
and then some kind of framework for imitating it
with your in your own flavor,
you have an immediate ecosystem of like an immediate ecosystem
like should arise because there's just a huge gap
between the like decisions that the single company is making
and all of the variations in those decisions
that like a wider ecosystem can create themselves.
And so then, you know, you need a marketplace
to like discover all of those services
and all of those products.
There wasn't any place on the internet
that like was like a home base for LLMs
in terms of seeing how much they were being used
and seeing who was using them and why.
- The closest would be hugging face.
- Hugging face was the closest of the fact.
- They just started hug like a few years ago before that.
- Yeah, and hugging face also didn't have
the close source models.
- Yeah, and they didn't, you couldn't use the models
at the time and there wasn't data about who was using that.
There are like a bunch of differences
between open router and hugging face
and those differences felt really critical to me,
especially when I was just trying to learn about LLMs
and like why people are choosing like these different
little ones that are emerging over time.
Got it.
- On no stranger to wanting more model diversity
at the time, you know, you're a couple of years
into your anthropogenic journey
which you've covered in the previous podcast as well.
What was your introduction to Alex?
- Well, the introduction was, I think, 13 years before that.
Oh, but the open router handshake actually happened
right over there if you remember.
- Yeah.
- Which is, Alex and I met, I believe it was sophomore's
now I have a Stanford review meeting for the first time.
So yeah, yeah.
So Stanford review was the Libertarian newspaper
on campus at Stanford that Peter Tiel started back in the day
and for whatever reason in Alex and I both showed up
to one of the meetings.
And I remember the editor-in-chief
was a mutual friend of ours, Lisa was really a really great
editor-in-chief.
Part of an editor-in-chief's job is to assign responsibilities
to people and make sure the work gets done.
And I maybe misremembering the details, but I remember
wanting to, it was kind of surprising to me that at the time
there was no dedicated technology section in the newspaper.
- Because it's political, right?
- It's primarily because of the thought.
- It started as like a political,
and yeah, states and things, yeah.
- But to take us back in time, you may remember this,
but there was this technology kind of legislation
that was being debated called the Net Neutrality Act.
And Net Neutrality is inherently this political concept.
It's about the regulation of internet broadband access.
And so there was a community of us who were kind of technologists,
but also debating the politics of the technology.
And I thought that a review would be a great place
to write about that.
And I was working on Net Neutrality article.
And I remember proposing, well,
maybe you should start a technology kind of section.
And Alex was only the only people who said,
yes, that would be cool.
And I forget whether we end up writing stuff together,
but that's when we first met, was 2011 or 12,
I forget which year it was, was one of those.
- Yeah.
- Is that old union, if I remember correctly,
'cause that's where we used to meet.
But along the way, Alex and I've had a chance to hang out often
and probably the time when we had the most professional overlap
was when I was running the platform at Discord,
and it had become this explosive kind of platform for crypto.
- Yeah.
- And NFTs in the middle of the pandemic.
- Which also, by the way,
you were in charge of safety and security as well, right?
- I was the head of platform,
which meant all of the crypto-dow and NFT launch
security debugging fell on me.
- And the messaging.
- The phishing, the social engineering attacks,
that Katana DDoS, that we were getting hit by,
but so on the time it started teaching security,
it's Kale, it's Stanford, CS153.
And Alex was at OpenC at the time and I was trying to figure out
how we could defend against all these attacks that we were,
and at Pika, I forget if you remember how much NFT volume
was running through Discord.
- Discord.
- But it was a meaningful amount of like,
it was like several billion dollars in NFT volume
of GMV, so to speak we were running through the platform
and it was all coming from OpenC.
It was these like, buy, sell, trade.
- I mean, the servers, the D in Dow is Discord.
(laughing)
- So that's when I think we had hung out professionally,
but a year after that, OpenAI gave Discord
early access to GPT, sorry, GPT-3.
- No, it was GPT-3.5, actually.
GPT-2.5, which is the RL version of GPT-3.
And that's around the time we made a Discord bot
with OpenAI for internal deployment,
and that's when I realized we would need,
like since I was part of the deployment team,
what was the use case?
- There were two that were, and there's actually a post
now called Discord is your place for AI with friends
that somebody sent me recently that I wrote.
And that's it.
published in 2023, but there were two use cases. One was Clyde, which was like a first-party
friend inside of Discord that could help you set up your Discord server and talk to you
about onboarding and get your friends standing out more. And then there was content moderation.
And one of the realizations we had with content moderation was it would refuse to moderate.
Like we just refused our prompts because the RL, the post-training was, we were very early in the
post-training era. And it would just, our prompts would trigger it. It's like guardrails.
And we told OpenAI, "Hey guys, we need access to the way it's because if we're going to be doing
content moderation at scale, we have 250 million monthly active users. We need more reliability that
the model will do what we needed to." And they said, "Well, sorry guys, that's not how this works.
We're a closed source company." And so that was my first realization that we needed open models.
And the enterprises would need more control or over capabilities. And then ultimately
would need some kind of control play and a management system to orchestrate these open models.
But I wasn't no good up. They were no good open alternatives until maybe six months later when
Lama came out. And six months after that, I led the Susie into Mistral, which was started by Guillaume
and the Lama team. And that around that time is when I remember hearing what Alex
launching OpenRouter and going, "These worlds are going to collide. And I don't know when it'll
make sense to team up." But Alex was so early and could see, I think he was totally right about
this ecosystem starting with Lama that then needed like an easy layer to manage for it.
Especially for, I was approaching it from the enterprise perspective because I did that,
like that's the VP platform at Discord. It was my job to ensure that when we deployed
models to like 250 million users, they did what we wanted them to. And that was very hard.
Because if you outsourced it to the labs and they controlled the guardrails and their guardrails
or their safety policies forbid the model from responding to your prompts, that was quite catastrophic.
Yeah, well, you know, a moderation is the thing that they want to support. And obviously,
beyond that, they would work, opening up woodwork with you presumably to give you a moderation
endpoint, which they offer for free. It was an interesting use case that they,
so they did give us a moderation endpoint. However, as you guys know, every Discord server
is like a mini deployment of itself. And so the use case was instead of having human moderators
that have to interpret the norms of the community, you just give the, they often like every,
you know, subreddit, Discord server's public ones have their own rules that the user,
the users create randomly and space the squad in, yeah. And then humans used to read those norms
and then enforce it every day manually, like observing each message in these communities.
And these communities are like million of users. So we had a 5,000 plus person team
globally on the Discord content moderation team. These outsourced contractors were a really
tough job. And so the idea was instead, if you could give the norms of that server to the LLM,
then the LLM would do custom moderation for that server. It's almost like a, like in context,
moderation for that server. And many of those servers norms just violated OpenAI's rules.
And so that, it was like, we had our own custom e-vails, so they each server had its own custom
e-vail, but at the time OpenAI's e-vails, we were also primitive in our thinking about how to
deploy these LLMs. That often the post-training prompts were super heavy handed. It said,
oh, anything about Harry Potter, anything that has trademarked content, you know, don't refuse.
And it was a fan, Harry Potter fan community, this is a real use case that had a content moderation.
The LLM would just refuse. And that was just not precise enough.
Another one that we heard was like, if someone was trying to write like a detective story,
and there's one chapter with a lot of violence, like maybe someone, like kill someone,
the LLMs would just refuse to help with that part of the story. And then these would be like,
okay, this is not structurally inherent to LLMs. There must be some choice out there
so that I can switch to another model when I'm getting a refusal or a bad result from the main
one that I have. And that tension also drove me for a marketplace.
Yeah. I think that is well accepted now. What was it like back then when you were raising or,
you know, starting this, did people get it? You know, what was some of the struggles?
I basically, I like getting stories out of him about how other VCs don't get it.
So, like anything you want to talk about, you know, now that that's, let's call it that,
you know, that early journey of LLM is done, right? You can obviously talk about some of the early
day stuff. Well, I was going to say that, like the biggest objection we got is big model win,
which is all you scaling loss. Yeah, scaling loss. And natural network effects are just going
to kind of a crew to one company, which will be like, it'll be a Google style monopoly,
just like how Google won the search market by a large, large margin. And you'll just be fighting
for its graphs at the end, basically. That was probably the biggest objection we got.
You know, it is interesting. Like Google won the search engine race with such a huge margin.
You know, I think like, had there been more interesting benchmarks or I had like,
search engines been, you know, have people like seen them a little bit more like LLM's where
there are services that you can build companies on top of that might not have been the case.
But LLM's don't merely have a user interface. They're also like ways of building entirely new
businesses. And, you know, a Google level monopoly would be like the Dutch East India company times,
you know, quadrillion in magnitude because the whole economy ends up like depending on the one
monopoly as well. So it didn't seem like, you know, like would be a really crazy outcome if that
happened. And it's also less likely because the economics of like creating good competitors are
much like much more decentralized. Everything Alex said is true. And I came at it from a completely
different perspective, which is yes, this is why we're here. The skating laws were never like,
in my mind, we're always a feature, not a bug for why open router would be very valuable because
I was one of the first investors in entropic. And it was obvious to me that other researchers in
our friends group, I went to a grad school for machine learning. And I just had a lot of friends
in the ML community who it was a way obvious to us that the bitter lessen holes. And so I was like,
fantastic. Now we have at least two proof points that compute scaling works. It was the opening
eye and entropic. And by the time I think we decided to team up an open router, I had already invested in
Mistral and Black Forest Labs and Luma. So there was multiple model companies and teams that I was
working with. But you did other modalities, whereas this is a lot of different modalities.
Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being
created. And that this whole narrative of like only one company will dominate like Google
was like, maybe true. But one, I don't believe that. But two, there was so much extraordinary
innovation happening across several different research teams. But the shared problem I was noticing
across all of them was often, you know, the research teams were fantastic at figuring out how to
reason about new capabilities. They think in terms of capabilities, but never like are not
developer mindset oriented. Like what happens after the training is done and the checkpoint comes
out, like you'd be shocked how similar they're early pre training teams at OpenAI and entropic,
BFL, Mistral were in their like default approach to taking their research out of the lab and
kind of scaling their impact, which is often, oh, the checkpoint is done. Put it out as an API
done. And then there'd be crickets. In the case of Claude, the first Claude checkpoint was
actually done a year before they released it internally. And then Chachi Pt came out and we decided,
okay, yes, it's a good idea to release a Claude version externally. And they had no plan.
Like no, no plan for how to get developers to actually try it out. And so if you go to the Claude
one blog post, you'll notice they like three kind of developer examples for users of the API.
And one is a discord bot. And the second is Vivian. My wife's startup called Juni Learning.
Because and then there was like notion, because these were all friends of like the entropic team,
because that's how like last minute the planning was around, hey, once the model's done training,
how do you get it out to the world? There was no distribution platform that understood what
developers needed. All the key management provisioning, like simple, like an endpoint management
versioning control, like all these things that the scientists and researchers go, I mean,
that's plumbing out under implementation, right? And instead, Alex came at it from that perspective.
And so, you know, it was so obvious to me that like every single lab I was funding would spend
literally sometimes billions of dollars into training. And then a checkpoint would be done.
And it'd be crickets like during early access, because they're like, oh, that's right.
Like, it's hard to use a checkpoint to make anything. You actually need a whole bunch of plumbing
around it to make it usable by developer. And so by the time I think we think it was so
obvious to me that a distribution platform like OpenRouter was critical to have in the ecosystem.
If we wanted there to be competition to Google, like on less, you know, with Google,
Deepind is done training a new checkpoint. And then they push a button and it gets blasted
out across all their surfaces from Google.com.
talks to, you know, everywhere, even if I don't want to.
- Everywhere, you want to know, like on Android,
like overnight they can deploy a new checkpoint
to like a billion devices, right?
And that invisible infrared vantage,
distribution advantage, most people don't realize,
but until OpenRouter showed up,
you had to think about all of that yourself as a model lab.
And it was very daunting, you know, an anthropic,
I think it took more than 12 months
to get to our first 10 million in revenue.
And contrast with Black Forest Labs,
I remember the early days, you guys had a conversation
with the BFL team.
And it was so simple for OpenRouter to say,
oh, no problem, like the day you launch,
we can send a million developers to you.
You know, that was crazy.
That was like a step function change in like, power.
- Is that a real number, a million?
- I think today it's like a million.
How many developers are on OpenRouter today?
- Over 10, but yeah.
Over 10 million, but like, it's hard to do.
- Yeah, we do a lot of like, you know,
account de-duping work, but yeah.
- If you could get a thousand developers,
just to put in context.
If you get a thousand developers who actually try
the model on day one, after you release it,
and just like, do inference and give you feedback,
that's a thousand more developers than they knew
how to get to on their own.
- Well, you know, BSO had a reputation.
- But yes.
- They had where I would stable diffusion.
And with Mistral, I don't know if you guys remember,
but the first checkpoint they released was like torrents.
There's like torrent weights.
- Yeah, they just put up a magnet link.
- Yeah, there was no API.
'Cause they weren't in for people.
(laughs)
- You're not like, say, okay, download these weights
and you guys go ahead and have a story on his side, yeah.
- Yeah, I mean, in addition to the,
like, building a really good developer experience around it,
the marketing that we do, like, for different models,
is totally different and perceived totally differently
from the marketing that a Model Lab does for itself.
- Yes, we are like a neutral layer looking at this market.
Like, it's a big dark room with all the corners,
completely obscure to users and users are walking into the room
and like, feeling around and trying to figure out
what objects to grab off the tables
and like, build into their companies.
And it's just an insane way of working.
Like, models are not products where you can just
enumerate all their features onto a web page.
They're all black boxes, including the open weight ones.
So you need to like shine lights on all corners of this room
so that people can see what makes this model good.
And you need the company shining that light
to be a neutral third party, which is what we specialized in.
So the, like, in addition to developer experience,
there's also like a very important, like, marketing
and product packaging components
and a way of, like, routing and discovering models
becomes, like, critical to your go-to market
as a provider or a Model Lab or a server tool
and more in the future.
- And this value to your earlier point about how many VCs,
like, you know, just don't, one of my biggest frustrations
is that venture capitalists, many of them,
like, just don't have any operating experience in the field.
You know, so unlike a traditional investor
who's just maybe come up through the ranks
as, like, of associate working on financial modeling
or maybe hasn't been a real operator in the field
for, like, more than 10 years,
which is a big part of the industry now.
I had just arrived at A16Z, like, a year after running
the platform.
And so I knew what the challenges were of, like,
building a real great developer experience
and actually, like, being able to create a working piece
of software with a model.
And there were a few, I won't name names,
but there were investors who were looking at OpenRouter
and, you know, felt at the time,
like, when I would compare notes with people,
that it was just, I quote, just a marketplace.
- Yeah, it's just a thin layer, just a proxy,
just a wrapper or whatever on other people's APIs.
And I was like, you have no idea how strategic,
the value that OpenRouter is created by being able
to orchestrate even three APIs in production.
The amount of both engineering work and community design
that goes into getting that actually live
and running in production at the scale,
the OpenRouter team had started.
This doesn't happen by default.
And that was one of the things that stood out to me
about Alex from the earlier days.
Like, he just understood, like, these,
from a systems perspective, like,
how do you get these fly wheels going?
Like, that stood out to me with OpenC
when we were working together on the NFT integration
that discord, like, Alex had a level of community,
like, systems thinking on how you get these fly wheels going.
The most scientists and machine learning people
just don't think of, like, we often think
in terms of pre-training, maturing, post-training.
- It's a linear stage.
- It's a linear pipeline.
- It's an old loop, yeah.
- Yeah, it wasn't much too much later
that the modern context feedback loop cycle
really got standardized in the industry, but at the time,
if you remember, machine learning was like,
like, mostly we did a lot of ML,
like when I was in grad school on a laptop.
So you just, like, downloaded a dataset
around some populations and you looked at the lost curves
and you're like, great, I made AI.
And the idea that you have to, like,
deploy those capabilities, collect feedback, trajectories,
then, like, put those into a continuous loop,
might came much, much, much later.
And it's very counterintuitive to the science,
like, the traditionally, I mindset.
I do remember doing the investment phase for OpenRouter.
I just didn't try and re-educate a bunch
of other VCs on White was not just a marketplace.
I was like, you know what, I'm just gonna invest.
And I'm going to, like, take the opportunity
to partner with Alex.
And if not, no other VCs get it, that's totally fine.
'Cause at the time, it was not obvious,
I think, to several of the investors that, like, OpenRouter
was not more than just a wrapper around APS.
That infuriated, you know, I was like,
you know, I don't have time to debate you.
I'm just, we're gonna invest.
And then I think, like, a month later,
Matt Murphy marked it up by 10X.
Like, I think, I forget what the exact post money was
and so on, but, you know, to his credit,
Manlo Ventures realized, okay,
there's actually much more strategic value here as well.
Maybe you didn't hear all these conversations
behind the scenes, but that frustrated me a lot.
You know, there's a lot of this, like, opining about wrappers.
And if you're like, oh, an app is just a wrapper on no model.
Then, like, and OpenRouter is, like, this wrapper
on top of other APS, and this is the most stupid,
reductive framework.
And so it's clearly somebody who has no experience
deploying products.
- The thing you dismiss other things with,
like, everyone's a wrapper on everything, right?
Like, if there's some point, some wrappers are value.
- I mean, investors are wrappers in L.D.'s, right?
Like, magic wrappers.
So, I mean, yeah, it's all wrappers done,
all done to bare metal, I guess.
- Yeah, it does.
- Okay.
- When I started the whole AI engineer,
I guess, the coining in 2013,
like, that was the number one pushback,
is that this is no value.
You should actually just train models.
- Right.
- And yeah, I mean, obviously, this is like,
you guys are one of the testaments to the fact
that you can actually build very valuable wrappers,
but also very valuable model companies.
- It's so hard to be, like, the day a model launches,
the fact that you have an open router end point
for that model, frequently at the top of hackers news
on day one, people don't realize the amount of work
that goes into accomplishing that.
An open router used, like, that would happen over and over again.
And I remember going, people have no idea how hard that is.
You know, that's not, yeah, we've covered some
of the inference engineering that goes behind some of the,
yes, with base 10 and all those.
- Well, today you have, you know,
all those like cool code name things,
there are people guess what oxy alpha is and all those things.
But I guess one of the things that you're teasing is,
how do you get that initial flywheel going, right?
Because today you have your scale
and your reputation and all these things.
So obviously you can't, you drive an immense distribution,
but when you're early on, when it's mostly--
- The boot strap.
- Yeah, well, how is the boot strap like?
- I mean, to bring it back to early Discord days,
I think we like initially connected with,
this is an open C story, technically.
But we initially connected, when you were at Discord,
and we talked about, like, XB, XC Infinity Sturver.
- Yes, yes.
- This server was like the biggest server at the time,
at Discord. - Yeah, that's right.
- And you were kind of, like, constantly bumping up the window.
- The limits on the server, oh my God,
would be in-- - For no server, like 10% of Philippines
was XC, was on that server.
- That's, it was like a meaningful controller
to the DDP of the country. - NFC, like, crypto game,
but-- - It was like a Pokemon breeding thing.
- Yeah, similar, yeah.
- There was battling, there was breeding,
and then there was, like, a marketplace for trading.
- Free to earn as well. - Free to earn, yeah.
- Free to earn, and, like, the graphics were really cute
and fun, and you kind of, like, you, you know,
you get kind of emotional about your action that you make.
So, to, like, start a community like that,
which we had to do many times in OpenSea,
with basically every early project
for us to create a marketplace for it,
we need to make sure that, like, the community actually wants it.
And it's kind of, like, building something that people want
and going and telling them about it,
like, you can do that on a one-on-one basis,
but as a way higher leverage to do that in a community,
where everyone can talk to you at the same time.
So, we spent a lot of time, like, building things
that the community really wanted.
We did the same thing for OpenRouter,
and, you know, like, the Axi community
was one of, like, a zillion communities we did that with,
and on just, like, saw us doing it,
and, 'cause you could just see people sharing OpenSea links
constantly in that discord.
Like, user sharing links is a really clear indicator
that, like, something important is going on.
So, we spent, you know, a lot of time, like,
first, figuring out what the gap is in the technology
that people care about, like what was the actual problem
that needs to happen?
to be solved in early LLM days, it was, you know, open AI refusing to finish the prompt
or like complete the task. It was also inability to customize models. And so there are communities
that are just completely blocked on that issue. And those are the communities that are most useful
to sort of learn about and and diving to and explore. Something that really struck me at that time,
as I was just hearing your talk, I remember noting how you may not remember this, but we were,
we had these like working Zoom calls that we were doing a sprint around for like this open
CD integration with Discord. And you know, we'd get, it was myself, my engineering team, I think
you were there. And I remember, you know, Alex in the middle of one of those calls, just like,
there was like silence. You know, we were, we were like, oh, yeah, this totally makes sense,
let's do this. And then there's some everybody aligned. And Alex was like, no, this makes no sense
to me. And everyone's like, I remember going, what? Like it works. Like you click on a link,
and then it bounces you out to like open C. And he was like, it's not a good user experience.
Yeah, we should not do this. And I remember going, you know, he was the only person out of all of us
to actually raise his hand and go, yes, it made sense from a technical implementation perspective,
like we're bouncing the user out into the into open C. And so it kind of checked the box of the
product managers requirements on both sides. But Alex went one step further and I was like,
you know what would be better guys, if we just embedded the experience right here inside of
Discord. So the link opened up as an embedded iframe. And you can just check out right there.
And not one person on the corner, like seven of us who had met like, you know, we got to read
the guy who doesn't work for Discord. And as a guy who doesn't work for Discord.
Like technically you benefit if they, exactly. And that was like adversarial to keep the user
inside of Discord would be adversarial to open C. And yet Alex put that user experience first.
And I was like, that's special. Wow. Because it's very hard to have somebody who's technical like
Alex and understands the developer flow, but also understands the best user experience. I want
to prioritize that. And that's two sides of the fly we let you can get spinning. Like it's
often hard to stop. And you just reminded me like that one was one of those moments where I go,
I really, I got to be better at user experience because I should have been the one who came up with
that. And I didn't. And I learned from you. And I think that went into one of our case studies
for the PM training program at Discord. I don't know if it's there.
Because you need an Alex, this is a conclusion. Yeah. Yeah. Yeah. And this is why I'm not,
you know, nobody should be surprised why Stripe decided like they had to buy open router because
it's a really rare combination of people who understand the machine learning community,
the developer experience, and the end user experience. And putting all that together
has resulted in this extraordinary scale that very few other marketplaces have been able to achieve
over the last, you know, five years. Yeah. Well, we should talk about the other reasons for
acquisitions, which you've written about. I wanted to sort of proceed somewhat chronologically as
well. So there is a point that, you know, one of the questions that Dave from H of 0 send in was,
when did you know it started, really started to work? And you brought up a mixture. I don't know if
you want to bring up that song. Oh, yeah. Which wasn't your overlap with. So yeah, the MOE was,
I don't know when, I mean, there was no like one moment where I was like, oh, this is,
you know, officially starting to work. It was like, well, we have increased extension. They really
super early on. Oh, yeah, but well, yeah. So before open router, I wanted to like explore a
bring your own model experiment. And anyone familiar with crypto is like, yeah, phantom in all these
things. Yeah. Yeah. So it felt like doing a meta mask analogy for AI would be kind of a fun
way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps
that were like hitting AI, like hitting an LLM via an API call. As there were like games,
just doing it in JavaScript. You know, basically like there was a there was a moment in time,
where it could have been the case that web apps call LLM through the browser, like through some
kind of desktop managed app that is controlled by the user. And of course, there are like,
I think many reasons that that did not happen. But back when the when the days were that primordial,
I built a Chrome extension called window AI. And with plasma with plasma, I had come across
early on and I was like, who's going to actually use this? You did. Plasma had a couple,
like, I think phantom was using it. There were some other like like like real companies. There was
like a shim basically reacts for code extension. It compiles to all these. Yeah, kind of like next
JS for cool things. Yes. Okay. And yeah, built window AI on top of it, the creator of plasma,
like started contributing code to window AI in GitHub. And that turned out to be Lewis Vichy.
Oh, you're the co-founder of OpenRouter. That's you have told me to tell you about Lewis.
Yeah. Okay. So that allowed users to kind of like configure which model they wanted to use
for a web page in their browser. And then like the app would just call out to that model
wouldn't need to do things. You know, not the right form factor for LLMs. But you know, it's like
fun experiment. You learn a lot. And like I open sourced it. And the main learning is like,
okay, this has to be an API. And it has to look a little bit like there has to be more of a
developer experience here and more of a discovery experience as well. Like I don't know where to use
these models. And a little Chrome extension is not going to help me discover. It's not enough
real estate. I need more space. I need visuals. I need graphs. I need, you know, examples. I need
images. I need to like, I need to be able to like explore both as a human and as an agent. So
that's kind of how how OpenRouter came to be. You know, a meta point that I think is under
appreciated. But Alex is reminding me is that we were quite lucky that we were so we were like
adjacent to the crypto community in those days. Because in hindsight, crypto ended up being kind
of like a dress rehearsal for generative models. Right. If you think about the the Axi experience,
you know, Alex is totally right. They were not that many AI apps at the time. And while I was
dealing, you know, my job was to be the head of platform at Discord, which meant to be a general
purpose place for communities and friends to create for developers to create apps and bots and
you know, the services that could be deployed across Discord. And while 80% of the attention
of the time was being spent on crypto because that's where all the NFT volume was, there was like
20% of my time I was spending with a friend who would get hotpot with me and ask me for we
played Magic the Gathering on weekends. And he was working on a little Discord bot that could
take a text input and turn it into an image. And it was called Mid-Journey. You know,
it was David Holt. That was a good friend. And David and I both then sort of failed ARVR founders,
you know, in the last before that. And I remember this mid-Journey was one of the fastest growing
communities we had after Axi infinities started to beat her off. And many of the the like the
abstractions and the infrastructure decisions we made to scale Axi happened just in time because
they inaxi did this and then fell off a cliff. And then as mid-Journey was taking over, we like
explicitly decided to help David make the server, the mid-Journey server, the primary place for
interaction with the with the model because it was very hard for people to understand how to use
the model if they couldn't see other people using it and copy them. And so the single player mid-Journey
web app on its own like midjourney.com had like terrible retention because people would show up,
they'd see this empty field. It was kind of like Dolly too. And they would type in like cat
or dog and it was like paralyzing for them to have this blank canvas that they had to fill because
they never used an AI model before. But instead in a discord server, you could see other people
using it and riff off of their prompts and the engagement was off the charts. And so scaling,
you know mid-Journey from zero to like 10 million monthly actors was a much smoother approach
both Axi infinities. And so don't forget the best of four pictures in the end. The best you
know. Yeah, which is the feedback. The Arlate, Arlate our feedback loop. Which by the way,
separately, Ekton Brown, David and I used to play Magic Gathering on weekends. And so like it
was one group of friends would hang out with like these concepts were all being discussed all
the time. But you know, there was, I think there were few of us who bridged both the crypto worlds
and the AI worlds. And compared to crypto where it was all, the question was always what's the
use case, you know, for this technology, there was never any need to ask that for you,
because the use case was so visceral. I can create now anything I can imagine. I can write novels,
I can code. And the infrastructure that those of us who believed in the distributed systems,
like value of crypto, like the censorship resistance part found this use case that was explosive.
And I think between mid-Journey, you know, Claude was a discord bot pre launch, you know,
that we were using internally as an LEM. 11 labs had a TTS model that we had on discord as well.
Like discord became this ptry dish for like early apps to innovate. And I don't think it's a
coincidence that they found a home there before open router gave the world like a public home store
or like a, you know, storefront. Discord was this like almost kind of ptry dish storefront
that was kind of like piggybacked on the infrared we built for crypto communities. And then I think
Alex was one of the first people to realize, wait a minute, like these apps need their own home
On the internet, an open router to me was a continuity.
of that, of that community's needs.
And of course, there was the crazy distribution
that you enabled for a lot of these developers.
- So then my question is how come you were,
my recession is open rudder is not that discord-centric, right?
You have a discord and you use it to engage your community,
but it's not like mid-journey where like,
no, that is like the primary way
of people experience open rudder.
- Yeah, mid-journey, like it really helps us
see visually, really quickly how people
are using the model and how to prompt it.
- And I think that is partly why the server was so critical.
It's like, it is the user experience.
It actually adds a ton, and you can go the whole mile
with just like prompting via mid-journey,
like via the mid-journey discord server,
getting your images and then sharing them and having fun.
For OpenRouter, for LLMs, like,
you need a lot of user experience around LLMs,
make them like really use the tools.
- Charpoint.
- And yeah, like seeing the examples of other people
is also not as useful, 'cause it's a lot of stuff to read.
It takes a long, long time.
- Yeah.
- You need like code-based integration,
not possible to do in a discord server.
You need, or technically, it's possible,
I shouldn't say that.
It's just not great, great developer experience.
You need like, you need governance
for, at the point where you got code-based integration,
now you need governance for managing the LLMs
that have access to it, the data policies,
which teams, all that stuff needs a lot more
than a discord server can provide.
So it's just like it's not the right.
- Well, in addition, you're not wrong,
but also there's the very important distinction that,
you know, mid-journey was an end user application.
- Right.
- And, you know, that's why discord
has 250-minute monthly end consumers, you know,
made sense for discord to, to kind of,
to be a host for that application experience.
What I knew was gonna happen soon after
mid-journey found explosive product market fit,
'cause we, I think when mid-journey launched,
from launch to 100 million revenue run rate,
it was less than eight months.
And shortly thereafter, stable diffusion launched.
And, you know, all of us used to hang out
in the discord server.
I think it was the--
- Distributed Discord?
- It was the Lion.
- Yeah, Lion.
- The Lion.
- It's the image community that's falling
to stable diffusion. - Yeah.
- And so when stable diffusion came out,
I realized, oh, now,
other people can build their own mid-journey.
Because until then, mid-journey did not have an API.
So they were a full stack company, right?
They were training their own models
and they were deploying them as an application.
But if you want to build your own mid-journey,
there was no API of that quality.
And I think Dolly too was still quite primitive.
Like, mid-journey actually had great quality.
And then when stable diffusion came out,
suddenly there was this new person
who could, there was new, this new capability
in the world, which is a developer
could create their own mid-journey.
And that, I think, created the need
for something like open router.
'Cause then you need an API to,
if you had the kind of creativity of David Holtz,
and you had stable diffusion as the model,
and you wanted to put these things together,
how could you do that without having
to figure out how to host the weights?
And what open router,
the shape of open router enabled is that, right?
When you have open model alternatives
to closed sort of applications,
open router's value in the world becomes extraordinary.
'Cause now I need developer can just show up
and use the old model diversity.
- Did you just say the shape of open router?
- No, no.
(laughing)
- It's the real one.
- I've been, I've been, I'm misaligned now.
I've been over-trained.
I've been using CloudWriffy much, haven't I?
- Slotish is what people have showed us.
- Slotish, oh God, I got to train myself.
- Okay, and I just want to cap off the Mistral side.
My, my, my TLDR is, there was a Mistral price war,
is what I called it, right?
Like, run about in Europe since 2023.
- Yes.
- For the December.
- They launched the Mistral by eight by seven B.
And the price went down at 80% to me, that's very positive
because it's like the first, like real competition
to post Mistral, is there more?
- Yeah, that was, I'm like trying to remember it,
all the things that happened.
Like, we saw that model come out
and immediately saw people say
that it was the best model in the world.
Like, this was, to my knowledge,
the first time an open weights model was called that
in real seriousness.
- It's hype, right?
- Is it, you know?
- It was hype, it was hype.
It was also like hype from AI influencers at the time.
And there were many examples
where it was like outperforming GPT4.
So, people really wanted to try it out
and see how it, this can be true for me too.
And if so, at what price?
And the like inference landscape was really messy.
We cleaned it up, it allowed providers to compete on price.
So, we could give users the best price in one spot.
And so, I think the first clear example
of like a provider marketplace working
in a way that adds value to developers.
- Sean, you may not remember this,
but I think we met for the first time,
a few days after Mistral came out at Nureps,
at a luncheon.
- Yeah, that's where the BFL as well.
- And Guillaume was there, Nureps at that time.
- You were there too.
And we had just announced the Mistral investment.
And Guillaume was over there.
And I remember turning to Guillaume and asking him,
like how are you feeling
after the launch of Mistral in 7B?
And his typical French fashion was like,
I mean, it's an okay model, it's not that great.
It was so, you know, in contrast.
But I remember him also saying that part of the reason
he felt a lot of people thought
that it was better than GPT-4 was because of the speed.
You know, it was an MOE model
that they had like absolutely kind of figured out
how to make super efficient.
It was on the period of frontier.
And this is an important thing about LLM's, right?
Sometimes when they're faster, you think they're smarter.
Even though like if you did end of, you know,
these common like e-veils that are seven,
you do seven tries.
And I don't actually remember,
I think we should go back and figure out
what the data says, but I wouldn't be surprised
if it turns out on an end of seven attempts,
GPT-4 was smarter on e-veils,
but the perception of on like,
or correctness would be smarter or more accurate.
But, you know, people like from a human preference perspective
felt that it was faster because it was,
or smarter because it was so fast.
- Yeah, and actually most queries do not take that level.
- Don't take that.
- That's true.
- So this is the start of humans as router,
which then eventually becomes open router as router
of like the auto level.
- Oh, that's interesting.
- Because humans are the routing mechanism,
like I will ask the fast model first.
And then if like, not good enough,
I'm gonna upgrade manually.
- Yes, yes, yes.
- But then he's gonna auto it.
- I didn't, I didn't auto it that way,
but that, I mean, that makes sense.
- Which then there's, there's a lot more techniques,
like fusion, fusion is the thing that we should talk about.
Before I move on to those things,
I just wanna close off the sort of early years.
One thing that I observe,
which you are also an investor in arena.
- Right.
- And we talked about mid-journey having that feedback group
of ABCD and choosing that being very important.
And you understand the flywheel.
So how come you didn't build arena
and how come arena didn't build open router?
- Well, arena started before open router, right?
- They had, they had the school project
and then they became a computer arena.
- Yeah.
- And I know you had some arena experiences,
like the heads up comparison type things.
- Yeah.
- But you never really went as hard as arena did.
- And doing heads up experience.
- And Ellens has actually did have a router project
based on Elmerina Elos,
they which they never commercialized.
- It's hard to do a company that does both
because one company is taking data and selling it
and the other company really can't (laughs)
by default.
So, you know, I think there is like a branding reason
that there are two companies here.
Like, when you set up open router,
there's no training, there are no prompts,
like aside from what your provider policy set,
like open router can't see your prompts or completions.
If you want to see that as an org,
you have to opt into it and enable it.
And so we're like pretty conservative
and careful about data policy and security and privacy
and Elmerina is like, their business model
is like oriented around the labs and-
- You don't have a fee, right?
You don't have a fee to give a fee.
- Yeah, I mean, we do give some,
we like have free endpoints too,
but like those free endpoints,
we think we're not collecting any prompts,
we're not like monetizing the data
unless you opt into it for some reason.
- This compares, I mean, you're not the first person
to ask me this and Alex knows this,
but I, you know, I was the first CEO of arena
for the first five months when we were helping
on the status in Whelan, kind of spin out a Berkeley.
And I did invest in that before open router,
but it was very strange to me,
the comparisons that outset, you know,
folks would make between the two projects
because the missions were completely different.
The founding entity for arena,
we called it the AI Reliability Institute
because it was actually there as an EVAL service.
Like the data, so to speak,
that they were originally kind of offering the labs
was how do you make the evaluation
of models more reliable,
then kind of like the state of the art at the time,
which is like really just finger in the wind.
- Yeah, that's kind of what Anastasia's
in Whelan's PhD work was, as scientists at Berkeley,
was on statistical methodologies for sort of correcting,
you know, EVAL estimates based on like intrinsic biases
and how you collected the data.
And stuff control and stuff like that.
And which is very much like a, hey,
how if you're a scientist,
And you're trying to kind of, the highest expectation customer for Arina was always like a post-training and like a research at a lab.
Whereas the highest expectation customer from my perspective that Alex really understood and the mission was to serve was like a developer who then takes the result of the research and then produces an application that's deployed to the world.
It was actually a completely different problem and person that these two teams were focused on.
And so from the outside in, actually I don't know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet together for OpenRouter at given you a call
because we were trying to get a pooled data set together from OpenRouter and from Arina to create like an open source repository of prompts.
I mean, these projects were so kind of different in their goals that it was totally normal to me to be like, let's call Alex and see if he'd want to team up on pooling data because they're so different.
We need, we actually don't have that kind of data at all. We didn't have API prompts.
We didn't have like what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model.
Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000 for level you could, I guess you could kind of conclude that Arina and OpenRouter are adjacent.
But the road maps, the missions and so on at the time at least were like in very different sort of directions.
That ideal customer I get, I totally get that. As a founder, I want to own everything. Like this is clearly the adjacency and I'm like, I'm going to explore that.
Oh, and everything meaning like you don't know what to do yet. So you want to like make sure you catch PM.
No, I think what he says is you want to own the entire infrastructure space. And so you kind of expand to whatever the way that you can capture.
Yeah, I think that's that's hard at, you know, in reality, because serving multiple customers is fairly tell you, you know, this is when it looks as a focus.
Yeah, I still think even in the age of AI, like focuses is underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.
The world can map like, Oh, I have this issue, which brand out there is going to help me with that issue. This is the brand that's known for that focus.
So like, if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it.
To underscore Alex's point, but how important focus is in the early days of entropic, it was not easy to like people think that the early days of entropic were like super easy because they were in the GP2 guys who left.
But it was actually very competitive. The company was starting $10 billion behind open AI. Right. And so to get to the frontier, like the big question was, what do we want to be known for?
What's the mission? And the mission was at age. I've got bare programming. And so to the exclusion of all kinds of other things that were really shiny at the time, like image models and video models that were getting lots of, you know, momentum.
The entropic team was like, we just got a focus on coding like that is the core capability that we're focused on today. You can see the results. Right. It's a trillion dollar company within five years.
And that focus, I think, like the high, the focus on who your highest expectation customer is and how you exceed the expectations, because exceed anyone's expectations is hard.
And doing it for multiple like customers is so even more difficult. It's part of the reason why open rather succeeded in entropic as well.
Was the focus on coding that early, though, or did it come later?
Literally from day one, it was AI, bear programming is responsibly commercialized in AI, bear programmer was the seed of memo. That was when I invested, right. We actually like refine that memo a lot.
Well, you got to ask Dario and Tom for permission of that. But it's an extraordinary piece of writing that they had put together and AI, you know, commercializing it responsibly commercializing AI, bear program was the mission, you know, from day one.
And I would say there was maybe like a couple of moments in the company's history where like they did experiments to kind of see if like little detours made sense, like a general chatbot like Claudia when chat GPT was really taking off.
But you know, at the end of the day, especially once they got their really significant pre training computer online, I think, like all the main evals of the company, for example, have always been coding evals long horizon,
programming from day one, that was always this one like what instant came out and clawed to came out. Yes, I remember the marketing mostly being focused on pros, like this model.
Yeah, right better. Yeah, and long context.
The first of all context.
This directly affected me because there's something on that.
What did you think?
A small developer, which is my Devon before Devon. Yes, it's small.
And you know, so I think there's there's all that really like good like focus is there anything that is a question that people do want to ask.
You know, you could have built any any other things like and obviously open road was working, working, working with their other ideas that you wanted to pursue that you turned down.
You know, just the past road's not taken.
We made a couple prototypes for things that we didn't launch one was a fine tuning model as a service.
Yeah, lots of that where open pipe in all those things, but it was it kind of was at a very consummary form factor where you would give us a YouTube video or two or three.
We would then extract all the transcripts from it and try to fine tune a model to talk like the person in the YouTube video or the people in the videos that you sent.
So like a really, really easy way of creating a fine tune model based on like some kind of videos that you like.
That would be so useful. We made it too. It was like, it was, it was, we didn't actually like test it with that many people because the model marketplace was our main our main focus and it and it was like growing and we're building more conviction and over time.
Just just. Yeah, why was as a creator? Yes. Yes. I have 500 hours of recorded voice of myself. Make a thing of you, your charge access to it works for only fans doesn't work for us as regular people.
I think all this mostly it's just a glorified ragpot. Whether it's in the weights or it's outside the weights doesn't really matter.
You're just doing rag on the videos. And people ultimately always just want to find the source video directly answers it. My use case was mostly to practice with myself because I often like to see what like the way I practice for a job interview.
I'm hiring a candidate or public speaking or whatever. I wish there was like a good mini me that I could like critique because it's kind of hard to pull yourself out. I would never get.
I would never offer it to other people service like pick your top five mentors that didn't talk to them instead of. That'd be cool too. Yeah. That was that's the character.
And that was the use case we are aiming at. I see. It's like you want to create an experience. Yeah, I see jobs and A I see jobs was the initial use case.
That's like that's not allowed. That's like that's a lot more. It's about a GCSEs fine tuning as a service as part of the router service is something that I would typically think about as well, right?
Like why don't you do that? Because if people are running already their inference through you store everything log everything fine tune to a smaller model that is cheaper faster.
All these things that that's within your control. Right. You didn't do that. But like other people would have pitched that in the general state of of the infrastructure start up.
Yeah. I think you were just maybe a little bit early because today that's an extraordinarily fast growing segment like, you know, from a straw where they do a lot of enterprise deployment and stuff and fine tuning is a custom model for ASMR or whatever.
Not often, but not as a router. They're just like, I come to you because I like your Mischal models. I want custom Mischal model. Right. It is not.
I want to run all my open I promise get the store all my results and then just move off of it. Right. They're not doing that as as like a way to export off of dependency on a on a frontier lab. I've not seen that yet. Yeah, which which was your kind of your.
I mean, we decided really we like leaned into our focus and figured that like there aren't like we just saw the ecosystem develop over time. All these inference providers that do want to help companies do that.
Like it makes sense for us to partner with them and to like give users lots of choice and to like figure out what makes them what gives them competitive advantages.
It's a whole new business basically and there's there's value in being a neutral marketplace that just kind of like works with those companies.
Could you share a little bit to Sean's point like how you prioritized what are some ways you prioritize features because you've always done it so elegantly that never.
It just happens and you make all the right decisions that always have product market fit from the outside looking in but consistently you seem to have prioritized.
You know, a lot of hit features that worked and maybe I have samples that bias or whatever, but Sean.
What is what you think hit features worked like leaderboards like leaderboard. Yeah, you know, like from day one, they're like the feedback charting B.Y.
Okay, but like he had like plugins, you know, he had like I think there was a whole thing I want to get into about like completions versus.
Yes, check check which is with the completions and then also let's call it like the rise of the reasoning models and how you deal with those multi modality all those things B.Y.
Okay, yeah, there's one I think it was in early 2024 very early 2024 we thought it might be interesting to fuse the results of multiple models together.
And we launched a prototype called mom mixture of models that let you like pick a couple models we pick them for you and then it would fuse the results together at the end.
and would show you all the intermediate results
in this big, con-bon-board-looking product.
What does diffusion at the end in other models?
- Another model.
- The smartest of the set, of the three of the set.
- So this is a good console idea.
- It was a model, it was like a very early LOM council.
- This is a multi-agent swarm as like,
they would call it at one of the frontier labs.
(laughing)
- There it is, you know?
- Yeah, like some of those ideas are like,
going the right direction, but the devil's in the details,
there's a lot of like product refinement needed
to make them really work.
They take your focus away from like, you know,
whatever else you have going on.
And there's a lot of like community building
and learning that you need to do.
And the technology might be too early.
So like all kinds of reasons they might go wrong.
And in our case, the technology was a little too early.
In other words, the fused result was a little bit worse.
- Sometimes the same.
- As the best model that was being used to fuse
because the best model was so far ahead
of options two and three at the time.
You know, over time, the top three or four LOMs
have gotten closer together, still neurodivergent,
but like all capable of inserting like pretty interesting ideas.
Like RL has basically like expanded the surface area
of creativity for machine learning researchers within each lab.
And so they can, you know, diversify the reasoning power
of different models more effectively.
At least that's my theory for why fusion is,
it like works better than it used to, but early 2024.
And so the technology was a little bit too primitive.
The form factor was, was not right.
And so we would have had to go through
like a couple more iterations.
And so we decided to just delete all the count.
(laughing)
And then years later, middle of 2026,
or early 2026, we're like, let's bring it back.
Like the research is looking kind of promising for fusion.
The models now have like two, three, four
top frontier models that are all really good.
And like, like I'm frequently trying
to like consult multiple models to get the best results.
Like, and then I, you know,
I ran a little personal experiment where I was like,
I'm gonna like do a, an architecture plan for a code change.
I'm gonna give it to all the models.
I'm gonna fuse the result.
And then I'm gonna ask all the models if the fused result
is better than the individual result
each model came up with.
And they all said yes, that the fused result was better.
And this happened a couple of times.
And I was like, okay, spot check, pretty good.
We should like benchmark this.
And that's how we built fusion.
- Yeah, and it came on your fable.
So you were like, this is fable level.
- Yeah, yeah.
- Let's start leading up to this year.
(laughing)
Should we have been gone to this year?
You know, can you mark out the main milestones
in the journey?
I think it seems like your promise was, you know,
routing, you decided the business model
very early, you take a cut.
And like, you know, what are the major milestones
that inflect the growth, right?
Like, you're going like 9% week and week now.
Is this the official number?
- In terms of token volume, I think that sounds about right.
- Yeah.
So just like, can you mark out the sort of brief history
of open router up to, you know, the acquisition?
- Let's call it, we're just talking about, you know,
people are, you have sort of your birth movement
with the missile stuff where people are really competing.
You have your state of AI thing where it's very cute.
You have 100 trillion tokens, ha ha ha.
Because now you're doing 10 a week, you know?
- Yeah.
- We're doing 10 a day.
- 10 a day now?
- Yeah.
- So yeah, you do this in 10 days.
Like, what are the major points there?
You know, I just wanna, like, there's this move curve,
but like, you feel the infections.
- A lot of this is kind of oriented around model launches.
We had, you know, a huge focus on pros,
all the way up through May of 2024,
because coding was just not there.
And no apps were able to build much on top of that.
So, you know, a diversity in models,
but not a wide diversity,
not a wide diversity in use cases.
Dream Tavern was one of our top apps at the time,
the creator of Dream Tavern now runs product at Cognition Devon.
Then we, in the middle of 2024,
we saw Claude Sonnet 3.5.
That came out incredible leap forward in coding.
And we saw the dynamics of like apps
building on top of us change.
We saw a huge surge in volume,
in like users using OpenRouter.
And this is when I think people started to look
at the like, money that they were spending
and get a little bit like, whoa, what's going on?
I might need to like, think about like,
more cost efficient, but equivalent models.
And shortly after that, I think it was after Sonnet 3.5,
mixed role, 8X7B came out.
And everyone was like, what, this is the model?
Like, the open weights community delivered.
And so it was really good timing from the strong.
- Basically, all of Anjus Portklausage is helping you.
(laughing)
- It takes an ecosystem to go in OpenRouter.
- Yeah, that was the, yeah, it was like,
it was like this early ecosystem,
it was like a swing action,
where like model labs would come up
with some sort of frontier innovation.
Like usage would surge,
then users, you know, look at their invoices,
30 days later, and like, whoa, what's going on here?
And then open weight models would deliver
like a cost effective options to three months later.
We saw that happen several times.
- One thing, one thing for the coding agents
was that you broke out, which are the top coding agents?
And they love that, they love that leaderboard.
The client versus the root code versus the what have you.
- Yeah, yeah, like client was like the top
of our leaderboard at the time.
We then, at the end of, and I'll skip forward a little bit,
at the end of 2025, there were quite a few coding apps
on the leaderboard, but they were all IDs
or, you know, terminal based agents.
And at the end of 2025, we saw OpenClaw up here.
And OpenClaw was like particularly interesting
because one, it was like a new form factor
that like brought in a new type of user,
not just a developer, but like a productivity
or sort of like an internet creator came to AI
for the first time.
And it also had an interesting architecture
where it was like calling your chosen model
for these heartbeats to see if it was still alive
in addition to actually using the model for real task.
And the heartbeats are like, they're kind of,
you don't wanna pay a lot of your new heartbeats.
So the auto router that we provided
was really, really useful to this like wide range of users
all of a sudden.
And so we just saw it rocket exponentially.
And then we saw, you know, like OpenClaw just blow up
and a couple other apps lean into that new paradigm
and do something similar.
Hermes came out and really leaned into things
like the auto router and built like a really good community
and leaned into like basically skill management
and making a really easy and effective
for people like, like set their memory in the agent
and build really good skills.
- Which another thing you never did memory skills,
sandboxes, always like adjacent things, you could have done.
- Could have, but like, it's, I think like--
- It's hard to bet.
- They're also very, there are things that developer
that really matter for like the developer use cases
that were coming out at the time.
Like developers wanted to architect those things.
Those are kind of critical to building a good user experience.
It was like hard, it's been hard for companies
to find abstractions that work for all developers
on the memory layer.
It is, you know, there are some, like Mastera
has done a pretty good job, for example,
but like developers have like lots of very preferences
for them.
And then we saw, you know, the way our leaderboard has changed
over time is kind of like a movie of how the AI space
has changed over time.
If you just sort of like go to the way back machine
and look at the rankings leaderboard
and the apps leaderboard over time,
it sort of shows you like what's happened in AI
over the last couple of years.
- To me, the coming of each moment was,
Andre Kapati was like, I no longer read Logolama
'cause like I just go to open routers leaderboard.
(laughing)
Which I remember that.
I think you probably like said, like,
sorry guys, I'm gonna send a bunch of traffic to you.
(laughing)
- So I was gonna bring it into the strike thing.
How does that kind of conversation start?
- We had this long standing relationship with Stripe, though,
from, you know, like many different projects
that we had worked on with them.
We invest, you know, a lot of effort in countering abuse.
- So good for odd.
- And token for odd.
- Can you give some numbers just so people understand?
- I think I like, I posted about this.
We blocked 10X as much dollar volume last month
as the month before.
And the types of token fraud are diversifying quite a bit.
You know, there are like fraudsters
going after typical stolen credit cards,
but they're also, you know, people trying to resell traffic
against the terms of service.
There's like hacked accounts.
There's people who just lose act.
You know, like, their whole company is compromised.
don't even realize it. We help them regain control and detect it. There are accounts that
are like reselling inference on the side. There are accounts that are dealing with an
accidental runaway agent and they don't realize it. Not a hack, but it's something that
blows up and the company doesn't want it. So our trust and safety team works a lot on
all of these categories of problems and helps block it and detect it. We have models around
them. We worked closely with Stripe for a while on this. I think it's going to become a huge
problem in the ecosystem. We're already seeing a lot of companies start to see these fraudsters
like spread and look for other ways, other than open router to other fraud vectors. And if you're
making a gateway or selling generalized inference, you are a target for fraud. If you're selling very
discrete intelligence products, intelligence products that are doing something pretty specific,
but not just reselling inference with some added capability, then you're way less likely to get
these fraudsters. I think we'll see companies also move away from just reselling inference
with some added capability and move towards discrete tasks and charging for those tasks and
charging for those enhancements and letting people bring their own inference in a third party way.
Whoa. Okay. Obviously, you would power that. But people will pay for outcomes or per task.
I think people will pay, you know, I think like the data dog pricing page is a good look at
like the future to come. It's like companies, like infrastructure companies will like charge for
different types of events that they're providing. And there'll be lots of like continuous pricing
models that look like that. And of course, there will be like, if you go down towards consumer apps,
you know, simpler pricing, more subscriptions, you know, fewer events to worry about,
and ones that like are not focused on just adding a markup on top of inference.
Not just because fraud is hard, but also because the pressure from the labs and from
like good inference providers to like do a commit and then bring your inference elsewhere is
going to be very happy. Any comments? Two, one, you know, I think Alex has done a very eloquent
job of describing something, you know, counterintuitively, I knew would be a thing at scale like four
years ago because of discord. And the particular experience that taught me this was, you know,
as we started scaling my journey, you know, the one of the primary ways that we used to give away,
or like get people to try my journey early on to get to the first 10 generations. Because,
you know, 10 generations of 10 images generated was roughly the magic moment activation point we
found. Like once you've done 10, you were like, this is extraordinary. But for that, we, so we had a
free trial with mid-Journey. And one day I woke up because I was the head of platform and had to
monitor all these dashboards. You know, I had like three missed calls from David and it turns out
like there had been this flood of new users overnight. And we were like, this is great. And he was
like, no, actually we shot down the free trial. And I was like, why is that? And he said, I don't
look at the geolocation IP addresses. And basically somebody in China had started to resell mid-Journey
free, you know, subscriptions with the free trial as a way to like basically, you know, it was
fraud abuse, right? And was it very specialized model like mid-Journey? Yeah. And that was an
action application. So this idea that I think the big picture of realization I had back then was,
hey, there's a new type of unit of value that's being streamed across the internet called a token.
And over the next 10 years, the entire internet value chain was going to have to deal with the fact
that like the more valuable tokens got, the more bad actors were going to go to try to get their
hands on those tokens. And anytime you scale something and the payload gets more and more valuable,
more bad things people try to get access to that value. And so it was very obvious to me back then.
And so look, to this day, I don't think there's a free turn, like I don't think mid-Journey's ever
actually turned on the free trial since then because it was really not an easy problem to solve
in terms of trust and safety. And so I try to start teaching the class security at scale. It's
Stanford. Like it was like one of the that and the anthropic learnings to me was clear that the
need for security at scale is going to be enormous a few years from then because if you just do the
math, right, think about if where, you know, online payments, you know, it has started roughly in
the 80s and 90s, right, and grew to over a trillion dollars over the next 10 years. And we needed to
build entirely new payment solutions to deal with online fraud. Where we are today is roughly there
on tokens. But over the next even five years, we're expecting the token economy to get to like
roughly five trillion dollars. And over in the next 10 years, I'd be shocked if you want to 10
trillion dollars of token flow. And so if we were starting to see such aggressive abuse and fraud
at sub-scale mid-Journey, remember mid-Journey at this point was like less than 300 million revenue
run rate a year. I just realized we were going to need like entirely new like systems to deal with
the fraud that was going to happen for trying to get into the token flow. And so my, I, you know,
when I forget the board meeting was when you brought up that you know, Stripe wanted to partner
up and it made so much sense to me because Stripe rate are when I was a client or 10 years ago,
we invested in Stripe and the whole pitch that, you know, Patrick and John communicated. So
eloquently was like, hey, unlike traditional payment tools like brain tree that do a seven day
verification like KYC and AML to get get the fraud out of the way, we actually just bite the
fraud cost up front as a customer acquisition cost and give tell the developer like just use five
lines of code and we start accepting your payments in five minutes and what'll happen is over time
we'll collect all this data on the developer's. Clothel model. Is the cloudflam model, right? And
they did five years later they launched Stripe Rader and Stripe really today is a security company.
That's the real people think it's a payments company. No, the reason there's lots of other
payments providers today that give you like cheaper payments transmission. But the reason Stripe
keeps, you know, being the dominant one here in Audi N and Europe is because they have extraordinary
fraud detection that they've built, you know, over the years. This is the same story with Elon and
Max Levchin and. And a firm. Yeah. So, you know, I think the story shows up over and over again
where every time you have value streamed across the world in large amounts, you need new protection
and security infrastructure to fight to keep the bad guys out and allow the good people to like have
their transactions happen really fast. And so I think, you know, the, this is why, from my perspective,
like the Stripe and Open Router story is a security story for the internet ecosystem for the
frontier ecosystem without a partnership like that. It becomes very hard to defend the quality
of experience and the speed and all the good stuff without letting the bad guys get in the way.
The second is that, you know, there's this under appreciated thing about like the fact that you
need to, like these, all the bad things that Alex described as being perpetuated by humans right
now is going to be perpetuated by AI agents over the next 10 years. So think about the like
recursive scale we're about to see of bad actors. It's not just bad human beings. It's all the bad
agents that are going to be attacking the token flow. And it's very hard if you're a researcher
and an AI lab to reason about that problem because the only data you have is how agents your training
are going rogue. But that's just a fraction of all the bad behavior on the internet that's we're
going to see. And so what you need is defenders, new sheriffs and down, which in cowboy hats,
that can can see all the bad behavior from AI agents across the ecosystem from different model labs
and different post trained deployments and different developers and take all of that data and say
we're going to build a shield for the entire token economy because without that, you know, the
amount of fraud we're going to see of the $10 trillion in GMV and global GDP growth is like a huge
percentage of that, I think is going to be fraud abuse. And we might never get there if people just
don't trust tokens. And I don't think this infrastructure exists. So you have your work
out for you with Stripe, but I don't think people have realized the scale at which agents,
agent, agentic fraud, like bad behavior perpetuated by AI agents is about to hit us like a tsunami.
Yeah. I mean, there's a lot to dig into there. I want to give you the last word we do have to wrap.
What can people expect from OpenRouter and Stripe? I mean, I think this is a really good way
for us to accelerate go-to-market and to go out market more quickly. It's also, you know,
as I'm eloquently described, there's a really clear better together story here when it comes to
improving trust and safety and making it really easy to accept tokens and let people bring their
own inference to your app and to help developers just like build on top of inference going forward.
We have a really strong brand with OpenRouter and we're keeping the brand. So like OpenRouter,
like as a product and the roadmap and the name and the brand, like, you know, is saying the same.
And so what you should expect, you know, in the next six months is that most things will be like,
like what we would have done had we been independent except everything.
will be moving faster. And that's kind of like our, you know, near-term goal, a longer
term, hopefully I can comment on it. So yeah, but I can't. Now.
Okay. Well, we'll hopefully do a follow-up at some point. But thank you for being so generous
here your time and congrats on the partnership. I mean, it's just one of the most beautiful
bromances I've seen in the idea. Starting from Stanford to here.
Yeah. Lots more to do. Lots of sheriff's, policing to do of the, of the token economy.
We need new, we need new sheriffs for sure. Yeah. Awesome. Thank you. Thank you.
Podcast Summary
Key Points:
Open Router emerged from a need to address the lack of developer-friendly, scalable infrastructure for open models, enabling continuous data publishing and subscription.
The concept of PubSub as a product principle reflects how marketplaces and agents consume data dynamically, unlike humans who act in discrete, attention-limited bursts.
Early breakthroughs like Alpaca demonstrated that accessible, low-cost fine-tuning could create high-performing models, paving the way for a new model economy based on data monetization.
Open Router was built to solve a critical gap
The platform’s success stems from its neutral, developer-centric design that prioritizes user experience, API accessibility, and governance—contrasting with closed model ecosystems.
Initial resistance from investors came from the belief that large models would dominate via network effects, but the ecosystem’s decentralized innovation and open access challenged this view.
The crypto and Discord communities, especially mid-Journey and NFTs, served as early testbeds for AI applications, demonstrating the urgency for a scalable, open marketplace.
Open Router’s value lies in enabling real-time model switching, continuous feedback loops, and developer autonomy—key for building diverse, innovative AI applications.
Summary:
Open Router was born from a deep understanding of the critical gaps in early AI development: developers needed accessible, reliable, and diverse open models without being locked into proprietary ecosystems. The founders identified that traditional model deployment was fragmented and inefficient, with researchers failing to translate model training into usable developer tools. Drawing from early experiences with crypto communities like mid-Journey and Discord, they realized that real-world adoption required not just models, but a marketplace where developers could easily discover, test, and deploy them.
Open Router’s core innovation lies in its PubSub-inspired architecture—enabling continuous data publishing and subscription—alongside a neutral, developer-first platform that provides seamless API access, versioning, and governance. Unlike closed systems, it empowers users to switch models dynamically, creating a feedback loop that drives innovation. Early skepticism from investors, who saw AI as a monopoly risk, was countered by the emergence of open models like Alpaca and Mistral, proving that decentralized competition was not only possible but necessary.
The platform’s success was amplified by its ability to serve both end users and developers, with communities like mid-Journey proving the demand for visual, interactive AI tools. Crucially, Open Router’s value wasn’t just technical—it was strategic, offering a scalable, transparent layer that democratized access to AI. This combination of community insight, technical design, and developer focus helped it grow rapidly, eventually becoming a central hub in the open AI ecosystem.
The journey also highlighted how focus—on specific user needs like coding or moderation—was essential in navigating early competition, proving that clarity of mission drives sustainable growth in complex, fast-evolving fields.
FAQs
OpenRouter is built on the PubSub principle, where products are seen as systems involving both publishing and subscribing to data. This model enables continuous, dynamic consumption by agents and consumers, unlike traditional human-driven, ad-hoc interactions. It creates a marketplace-like experience where users can subscribe to various models and make real-time decisions based on data flow.
Unlike Hugging Face, OpenRouter offers a more comprehensive marketplace experience with open models, better developer tools, and a neutral platform that helps users discover and manage diverse models. Hugging Face lacked closed-source models and had limited visibility into usage patterns and user needs, making OpenRouter a more complete ecosystem for developers.
The realization was that developers needed a marketplace to easily access, switch between, and manage multiple open models. Early models like Alpaca demonstrated that high-quality models could be created affordably, leading to a demand for a platform that would allow developers to try, compare, and deploy different models efficiently without relying on closed APIs.
Before OpenRouter, there was no centralized place to discover or use open LLMs. Developers faced fragmentation, lack of developer tools, and poor model discovery. OpenRouter filled this gap by providing a unified, easy-to-use interface where developers could find, test, and deploy various models, accelerating innovation and reducing the friction of model adoption.
MidJourney’s explosive growth showed that developers and users wanted easy, accessible, and visually engaging AI tools. This highlighted the need for a marketplace where users could try different models without complex setup. OpenRouter emerged as a solution to this, enabling developers to build on open models with intuitive APIs and developer experiences.
Community feedback was crucial in shaping OpenRouter’s product direction. Early engagement with communities like Discord and Axi revealed real user pain points—such as model refusal, lack of customization, and poor developer experiences—leading to features focused on usability, discoverability, and real-time model switching.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.