Building a successful infra product between all the AI apps and model providers (chat with Louis from OpenRouter)
33m 56s
Louis Vichy, co-founder of OpenRouter, explains the company's evolution from a risk-reduction tool for AI developers to a leading inference platform. Initially inspired by hackathon stories of developers facing huge bills from free tiers, OpenRouter launched a feature where end users paid for tokens directly, allowing developers to build without financial risk. This attracted a developer base, leading to growth from 4 billion to 5-6 trillion tokens processed weekly, proving product-market fit. Today, OpenRouter serves as a reliable, neutral layer for sourcing inference across multiple providers, handling complexities like model fallbacks, pricing, and API differences. The platform saves developers time by managing model updates and offers features like consolidated billing, compliance controls, and presets for enterprises. Louis emphasizes a focus on practical needs—cost, speed, and reliability—over "magic routing," and highlights engineering challenges like routing across endpoints with varying capabilities. Looking ahead, OpenRouter aims to be a utility layer for AI agents, potentially adding memory support and other developer tools. Louis's hot take is that model monopolies are transient, easily disrupted by new entrants, so the value lies in harnesses and reliable infrastructure. He invites engineers with good taste and TypeScript skills to join, and points users to OpenRouter.ai for rankings and model access.
(upbeat music)
Welcome to the Inn for a Pod.
This is a tip from SNVC, and you let's go.
- Hey, this is Ian, lover of all agents,
couldn't be more excited to talk to Louis Vichy today,
co-founder of Open Router.
Louis, what got you started on Open Router,
and why'd you dive in to start in the company?
What was the crazy idea in sight you had
that got you all started on this journey?
- Ooh, starting from the bottom.
- Essentially, one of the thoughts that we had
was building with any new technology
would incur a lot of risk.
And this actually was inspired by a lot of the
hackathon journey I went through,
so I went to a lot of hackathon early on in my day back in college.
And every time I heard a story about a kid
who used the MongoDB free plan,
and then they used it so much,
and then they're publishing an app, right,
with the production of pipeline,
and then the free tier, you know what,
it kind of expired, and now they're like,
on the hook for like $70,000 from MongoDB,
and you'll be like, "Oh my God, how are they gonna handle that?"
And I would imagine AI, this is early turn and turn,
I imagine AI would have the exact same issues,
meaning people were gonna build crazy subordance,
and people were gonna consume a ton of token,
and it would be so risky to build
with a proprietary AI model,
unless you have to do some local models, right?
That would be much more sustainable,
meaning the model's running globally,
and you wouldn't get charged much for it.
But then the problem is local model sucks, right?
At the time, none of them were good.
And so the first thing the OpenRouter ship
was really a way to did risking developer,
and entirely when they built AI apps.
We built basically this feature
that allows you to do sign in with OpenRouter,
where the user of the app paid the LM directly, right?
So essentially, it's a flow
where you would create an OpenRouter account,
and then we will mint a key for you,
and that key will allow developer
to charge the end user directly.
So a lot of app that started with us,
that built with us early on,
they would consume billion of tokens,
and the dev paid nothing,
because the end user paid directly.
So that is really is the thought process
of the first feature.
And then once we got that feature out,
why when the gate is open,
where a lot of developers built, right?
And they can build crazy, crazy agent and loops with us,
with practically zero risk.
I mean, there are some risks, right?
But it's not like financial run.
We start thinking about how can we scale the model,
pipeline, the model evaluation, and the model,
how can we make the model more competitive, right?
How can we make all the model apps more competitive?
And that is the origin story for the model ranking page,
where we are now, if you look at OpenRouter AI slash rankings,
you can see we're doing more than north of Sichuan token
a week, and that hasn't been a crazy journey.
The first year when I showed to my dad,
we were processing about roughly four billion tokens
in a week.
The second year, I showed it to HF zero,
one of our pre-seed investor accelerator.
We were processing about 150 billion token in a week.
And last year, and essentially growing now,
five trillion to six trillion token in a week.
So there's clearly a growth curve of the thing
that we are processing out of the browsers.
And it's kind of proof like there's a PMF there
for our other fit.
And also proof that more and more models
are being more useful and being more competitive, right?
So essentially, that is like the--
from the initial journey of OpenRouter so far.
And where were you today, like, if you were to sit down
and say, OK, obviously OpenRouter's
tried to hit like many things very humble,
but insightful beginnings, many people who are in building
agents have heard of OpenRouter one way or the other.
Could you help people understand what OpenRouter is today
and how that vision has grown over time?
I'll say today, we aspire and we hope we are the most
reliable way to source the intelligent,
to source essentially inference across
six to different providers.
When you need inference, you go to OpenRouter's, right?
And we hope that what we're providing
is the add-on on top of the models, labs,
and the provider themselves, right?
To help the developer dealt with, essentially,
the risk in them, right?
The risk in them from beside financial room,
also vendor lock-in, right?
A lot of time, we also help the developer
like dealing with the relationship directly
with provider too, as needed, right?
Because a lot of time, they would just then bring the key
from the provider over to us so that they can manage
all the AI spend on OpenRouter's, right?
So essentially, we are now slowly evolving
from the best place to get inference
into the control and plan for all of your inference need.
In inference, essentially, you call an LM,
you call ChatGbT and so on to get some token back.
- And so I think like I'm a developer, Ian is a developer,
we're all very much all using OpenRouter somewhere,
which is like what a fun part of it, right?
And it's out there so widely.
But if you ask me even maybe two years ago,
there will be an infra company just be routing,
not just be, but there will be a centered around routing
to already LMS, it will be a little bit unthinkable,
but like, okay, how big of a problem that will be,
why we need to layer like this?
Can you talk maybe a bit more from an infra layer,
why is everybody adopting you?
Like what we're thinking about is this too much complexity
when they're dealing with a variety of model
plus in mind, providers, or there's some,
even like some lot more nuances,
make unpack like a complexity here.
Why is it such a hard thing
that people don't want to build themselves
or rather use OpenRouter today?
- I think, well, to your point, right?
So many people early on for the first,
so the company been around for three years now, I think.
And the early journey, people would load us
and the consumer, the enthusiasts would be the most stickiest
and we stuck to them, right?
So they would be the highest advocate for us.
And we understand the pattern.
The pattern is that model apps, I mean, the bet is,
they're going to be better model in the future, right?
That was the cold bet for the first year and a half.
And with that bet, every time there's new model
coming up on Tuesday, I don't know why,
but they always ship it on Tuesday,
your company is gonna spend an hour
with all the engineers scrambling to support that new model.
If you don't, your customer will be like,
"Hey, I want Sonnet 4.5."
Oh, I want the latest GPT model, right?
And then you have to add all of those code.
And now you also have the cat tracking for pricing,
billing, and all of that stuff, right?
And it's even worse when you have to move to
a whole different API shapes,
like moving from Open AI to Anthropic,
totally different shapes, right?
And I think after people like onboarding about 10 model,
they were like, yeah, this is enough, you know, like,
I'm done.
It's very, very, like labor, it's basic labor.
And so we have built all the ARM and system
to track provider.
We work on the spec to help provider,
like, you know, communicate with us,
you know, the changes of all of these endpoint
over the past years or so.
And so we have infrastructure to just keep track of all that
together with pricing as well.
And so I think that is the main reason why, you know,
people was like, yeah,
let's just stick to Open Router for now.
And then, you know,
probably easier for the entire team.
And we basically save you like, you know,
an hour or almost every week, right?
On engineering time, doing some minion work.
And then on top of that, over time,
we also add, you know, all the bell and whistle feature, right?
Like we add, you know, plugins,
I mean, over time, we actually built feature
for both developers and also their manager, right?
'Cause like, if you have AI spend across a bunch of vendors,
it's just like a nightmare, especially for a small team.
And your finance, your CFO is not gonna love it.
But the nice thing about Open Router is you match one bill
and you have, you know, essentially consolidating usage
across two different providers, right?
And your CFO would be decently happy.
I would say, too unboring us.
- Yeah, that's actually, that's amazing.
Because as you mentioned, like model,
number model is just increasing, you know?
And I think we are also in this huge debate
between like Uber models versus even smaller models
and all this kind of stuff.
Actually, maybe it's a fundamental question.
Like, I remember when I was looking at a variety of people
building LM apps, there actually has been like,
a variety of libraries out there
to try to abstract away, you know,
calling different models as well, right?
And so I think the probably the most common pattern I've seen
is like some GitHub library project
that lets you able to like,
abstract away different models and stuff.
Where you guys are totally not just that way.
There's a hosting environment and stuff like that.
I'm sure like, maybe the way my question here is like,
why do you guys want, what got you guys
to be able to be almost like the number one adopted product
when there's actually a lot of spring
of lots of things happening.
I remember at the time.
- I think this applied to almost any company out there
as well, by the way.
So when we were building out open routers,
there were also a lot of competitors, right?
But a lot of them were focused on the wrong problem.
Meaning they're trying to solve for the magic routers.
A lot of early model did not make the bet
that there we always better model in the future.
And they make the bet that the current model
are good enough.
So let's try to find a way to find the best model
of the current model.
But we focus heavily on,
there's going to be better model in futures.
Let's focusing on measuring what we have on the current model,
but only the pipeline to get the next model.
And.
build the relationship right to get like as much capacity can get on the next top tier model right
and build a very strong relationship with all the AI apps because I think that's like the awesome
part of open routers we stood at the center of all the AI models lab and the app developer and
bridging a gap right across apps and all the AI models and going back to the point of like doing
the thing that is like we didn't try to build like this magic router the router is very um
it is a very uh the here stick is decently simple it's essentially falling back
and we're routing based on either pricing throughput or latency right these are stuff that you can track
these are stuff that you can track you can really see track you don't have to leave out the model to
see if it's smart or not and these I think are the core inside that we have done over the past year
and a half where let's focus on that first and build the smart model later right because the core
need of a lot of user though is still can we optimize the cost I actually just need models extremely
fast the fast deployment right and I think just listen to your customer to see what they actually
need the moment I'm probably not not a lot of people need this magic router to wrap between model
and I think this applies over to every other technology company out there too is focus on what can be
realistically do and focus on the customer first and naturally customer work you know stick to you
and ease you more and naturally spread you like Wi-Fiers and I think that's a pit of success for
to go from now on so far yeah kind of curious as models evolve as new model types you know we
all different multiple models we models are great for ph browser to text different reasoning coding
whatever I'm curious do you think we enter a world of like there's like a long tail of specialized
models and the use of the model is like hyper task specific so like today let's be honest like today
you have an agent let's say cloud code right and that agent is very specific to coding and Opus 4.5
it's on it and like the coding specific like models are like it's that's its jam and what we've learned
also is that coding is increasingly becoming a soft task and I'm curious do you see a world where
we need like a generalized task model routing layer that's intelligent to be like hey I the user
wanted me to do this like what do I do or or do you think that there's simple heuristics and simple
agents where we end up like what's the complexity of like what tasks agent at high level when I say
agent like the top level agent for like you know a dummy like a chat GPT or whatever like what's the
level of complexity on the routing the task and how that's interrelated with sort of the prompt
that the user's asking my hot tech this might be a hot tech we love hot techs yeah I mean
they were here for them my hot tech is it's possible you can you just need a very small model to make
two code decision it could be a bird you might just need a very small model to do like some kind of
edge routing on like the kind of classification decision decision workflow and then douching sure
can be a mixture of huge on model because even look at cloud code cloud code except actually just
close on it right and cloud opus the model is opus it's very generic the thing that make
cloud code is the harness the agent harness which is like it's a library I'm sorry it's a binary
rank on bun probably written in tiscript and that can certainly be replicated by a lot of people
and there's a lot of people trying to replicate what cloud is doing already right in the open
source space like open hand there's also like other other a tool of close source solution like
Devon as well right like in fund condition so I would say my hot tech is the base decision layer
can be a very small model but it will fan out it will fan out to like more space of model
that's it and I'm curious as open router evolves and if you think about the space and I think what's
happening is that sort of the open router vision so we're going to be right there along the way
and get smarter over time or how do you think open router evolves as the space evolves and I'm
all secure is like how do you think the the complexity of these agents evolve as well like that's
another way to ask a previous question but a different one yeah I think the agent is not run for
longer for sure people want to try to make them run longer running overnight even one of my dream
right now I'm trying to build this thing as a as like my personal goal is that when it goes to
sleep I have like a farm of agent just you know building stuff for me right we already have that
but like that be fun to have them just keep on running I mean I don't know but basically my
current habit every day is before I go to sleep I just spin up like 10 different taps of
clockhood and just say hey refactor my code base for me and then in the morning you know hopefully
will come up with something I'll review them their PRs which is but I would like to automate that
even more I think the current baseline for router is still like let's make sure we are the most
reliable way to source inference so let's make sure that the base layer is extremely reliable
scaling or infar making sure that we have enough like all the qr you know acting not netting
all the description are like flowing so that we can serve the baseline well open router essentially
become like the utility layer to get the monitor to get like the inference which is like how we can
call tooling and so on and then the and then let the developer build it out the developer will
build out the agent harness the infinite loop you know the temporal workflow that's running on top
open routers right that's like how I right now we are thinking about like the positioning of open
routers right to become a very reliable infrastructure for like all the AI agents and talking about
reliability and stuff I think looking at your website obviously performance reliability is almost
like the forefront and then there's just billing and stuff like that um you know when I talk to
friends even two years ago uh or even now like opening it being down and talking being down it's
almost like a given every day almost like uh like stops now I can't cloud I can't five code because
cloud is not working it seems like it's so common and I guess since you're in the middle right
you know if that's off it goes on you can't really do anything I assume I guess oh yeah so I
guess the question is like what is what what are some of the things you're trying to do to really
uh maybe even technical challenges that are surprisingly didn't even uh that people don't know of
trying to solve this problem is uh give us a little bit of like a a taste of what are actually some
interesting problems you solving for customers because we don't even know that well essentially
let's say anthropic is down right but it's only anthropic first party API there are still vertex
that's serving that anthropic model and also there's bedrock that's serving that anthropic model right
so essentially the router would fall you back on the same model across those endpoints and essentially
the back the back on in the back of the room we would talk with you know like google and bedrock
to you know execute you know like a smart capacity we we can right to serve when there's downtime
with your anthropic first party so if you went to any of the like some of the anthropic model even
opening a model right in the uptime chart there are time when them the upstream endpoint
went down but our uptime is still you know like in the 99 and it's all because you know we have
this mechanism that we have engineered and designed so that you can fall back and these fallback
mechanisms is very model specific right each model has a bunch of providers and that's like one of
the challenge and then there's all the challenge of like across these endpoints so for a given model
there are let's say five different provider a model and a provider within open data link we call
them an endpoint and so we have fine endpoint each of these fine endpoint though might actually
have different way of serving the model some might have two calling some might have structured output
some might have none so these are the thing that we would then have to look at your request
and figure out okay which one is the best endpoint to serve your request even given for a single
model I don't know if it's well known but we have like a parameter in the API that allow you to
specify the max pricing that you're going to pay for this request and so we would have to then
you know like measuring how how expensive your API code would be to essentially estimate that
and route you to like the right endpoint for example so and so forth right so there are so many
small little thing that really our developers have worked with us and kind of expressed you know
to us and we just build them out yeah yeah that's that's actually really super interesting because
that there's so many little problems like that just using your normal models you think or just I'm
just calling a model but behind I think there's a lot of thing that need to happen um actually want to
maybe double down some other other aspects of things because I think today when we look at models
used to just be chat GBT sending a prompt but today is a lot more about the rags and they even
beyond now as a memory right like actually yeah how do people consume and how to remember the prior
you know either as a multi-turned-state kind of thing or there's going to be a waste remember
prior states and some level like this open router has a role of memory in your spreadsheet like
if people want to have persistence or even memory across agents calling a model is it's just like okay
that's the last one you call it go there or do you have even some aspects of like I want to
able to help sell save memory and go across to the end points we are definitely very interested
and we are talking actively with customer right and I think this is the part where we always
we want to have deep customer empathy to really understand like how can we serve memory without
and attention, right? Because a lot of time we've built memory.
Memory is also not a thing where it's very,
if we do it bad,
we will just trust with the end users, right?
So I think it's crucial that we design it with,
but to your question about,
are we interested in like kind of expanding the API
to be slowly improving the developer experience
when they're building up Asian?
I would say to answer definitely yes, right?
'Cause I think more and more people are building out Asian.
And so any feature that the risking developers
and allowing them to build like better workflow,
better agents, like we are all for it.
Yeah, a lot of model providers, right?
Doesn't really parse PDF, right?
A lot of times they don't.
We can add them to parse the PDF,
but a lot of times they are serving the basic LM models.
There's no PDF parser.
And so we would then add a layer on top of our API
to always parse for PDF.
And so when you send a request with a PDF,
it is automatically like parse for you.
Yeah.
So those are like the belt whistle and stuff
that we added to the API to help the developers.
I mean, there are also other stuff that we build as well.
I mean, as we grow and as we get adopted by bigger customer,
we now also look into compliance,
look into like security posture, right?
There's like, we have a new feature coming out soon
called Godbrow, which allows you actually controlling, right?
How your model for your entire organization, right?
Because when you see so, you kind of want
to ensure that your organization is using the right models
or using a model that have data center in the US
or have like software compliant,
people compliant, or even have like, you know,
like you want to use endpoint, right?
Endpoint meaning a model served by a certain provider
that meet your, you know, like maybe data retention criteria,
like ZDR, right?
So we also have all this feature built out for, you know,
enterprise to serve their need.
And even for enterprise customizability as well,
we have this feature called preset.
So with a preset, you basically make an alias
for a model with a custom display name,
with a system prompt that you can add to it, right?
So now you have this nice little custom model
that your company can use across your organization, right?
These are the kind of thing that we are,
you know, incrementing building out, right?
As a customer grow, because interestingly,
the jury open router is a lot of a customer
become more mature, right?
They come in, they build, they build a small little company,
and then either they grow to be bigger,
or they bring us to their actual work, like to their boss.
Yeah, so it's a crazy journey and it's very fast.
Like the iteration is so fast that a lot of them have graduate
from like, you know, like from literally a one person team
to now look like a 20, 25 person team something.
That's like, yeah, that's like into like a CRA company.
So it's fascinating to see like your customer growing
just as fast as you, I would say.
- Yeah, AI is nuts.
Well, it is nuts.
- Let's jump to our favorite section,
what's called a spicy future.
This is where you give us,
what's your spicy hot take around AI or infra
at the most people don't believe in it yet?
- Yeah, I was thinking about that,
but then you know, like as we're speaking,
I'm like, is that like a good idea?
Is that a fun idea, right?
Okay, the question is, what is something that your thing
is true, but everyone thing is not true right now,
but focus on like, yeah, infrastructure.
- First, AI in general.
- Yeah, it doesn't happen.
- AI in general, yeah, yeah.
- I know, the problem is, you know what?
- So one thing with me though, right, is,
but this is like a separate topic.
But I'm just like rents on me or whatever it's your thing.
I had a lot of these thought about half a year ago,
but as it come to grow,
I have to then condense or thought into a specs,
you know, RPD and thing, right?
And then present it to my team
so that we can actually internalizing it.
And then God damn it.
And probably just release a bunch of freaking paper
and article about certain things like long running agent
and skills and tool calling.
And everyone's, oh my God, yes, we gotta do it.
(laughs)
I'm like, okay.
- I have my head in there, you know, somewhere.
- No, there's a lot of that.
- We would be like R-T-P-R-D or something, right?
- Yeah, so it's very hard to get a take now,
at least from my end right now,
where people don't agree because it's like,
it's not all news, unfortunately.
- It might be a hot take on its own, right?
Which is like, people, anyone's believing that they have,
you know, revolutionary thing, a spec or something new,
it's not new in AI quickly, right?
Just the level of like changes.
And the state of art model that we have
can only stay state art like a week now.
(laughs)
- Yeah, very much a week or so.
- It's crazy, I'll stupid, like low,
in femoral these models can get,
which is what plays in your favor.
So maybe you can talk anything in this line of wording
at all, actually, right?
It doesn't have to be a very specific product
or project or technique.
It could be just like, everybody's still police,
you know, open AI would just win over, right?
Or anthropic and everybody.
- So that was the bed though, yeah,
but that was the bed that we made originally
with open routers.
- But I still feel like today,
may I be still universally bleeped, right?
So it doesn't have to be like, no one believes,
it can be like, majority don't believe.
- I also don't want to be too forward leaning
to be favoriteism or like criticizing one partner
'cause they are our partners.
- Yeah, they are our friends, I know, I know.
- Yeah, they are.
- They are our partner.
- We don't want to do that, we don't want to do that.
- You don't want to do that.
You're a very neutral party, I understand.
- So it's also tough for me to get an actual hot take, right?
'Cause like, for example, some guy from anthropic
technology came out and was like,
nah, too calling, open AI is like a mid implementation.
And then I'd be like, yeah, I guess, you got the right
to say that because they compare it, right?
They can short it each other.
- Yeah, yeah.
- Maybe, so yeah, maybe another way to do this
is could be like, what are something that you see
or start to happen but hasn't happened everywhere yet?
- Like, so one, I'm starting to see a lot of company
do a small company experimenting with this.
So linear, just release their linear agent, right?
So now you can go on Slack and tell linear,
hey, let's make a ticket for me.
That is not open in public.
Fern also have a chat box now where you can tag Fern
and it will improve your doc for you.
And I think there'll be more and more,
are you familiar with Devon from Commission,
the AI solve engineer?
There'll be more Devon for X.
Devon for Excel spreadsheet, Devon for update my docs,
Devon for HR that will live either in your Slack,
in your team message, in your bunch of your, you know,
your chatting the face, essentially,
wherever you consume or chat with other people, I think,
that would be more of that.
- So I think seeing more like the Devon
or people, someone even used the word cursor for X
or whatever it is, that usually implies
there's gonna be one dominant player that does that
because encoding, even though there's many, many other players,
cursor has been like almost like the number one player
everybody uses.
- Cursor is awesome.
But I think clock code is like getting huge traction
because even look at the way that acrobic
has positioned in clock code, right?
The clock code agent can be decouple.
And as mentioned before, the clock code is just a harness.
You can take clock code and run it on a clock container
and then use it to run like an Excel spreadsheet
or to running your entire finance team
or like to run a bunch of interesting stuff
agentically in a background, right?
You can now implement it with clock code.
My prediction is clock code will be rebranded as clock.
So it will be just a clock agent, it's a base agent.
And then you can deploy clock agent to almost every word.
- And I mean.
- Co-worker thing they launched, right?
- Oh yes, they did, and they launched that,
which I'm saying, right?
Any heart attack?
- They have a really retired code.
I think they just added a coworker
that doesn't hit your computer only.
But I think you're being doing further,
like it's not just your CLI for coding,
it's not just doing stuff in your computer.
There will be more, I guess, environments
that runs in other data and other apps running.
- Great, so yeah.
And everyone can actually build that.
Like not just anthropic.
Everyone would actually leverage what anthropic is built
to build their own version of it, I would say.
Because the main reason why I've had a lot of friends
try out code work and the main response
has been is, it doesn't seem like it will help me
in my workflow.
But imagine if they can build core
with their own workflow.
And if that building process is very simple,
the thing more and more people will do it, right?
- Maybe even spicier.
Because like CLI, obviously come out anthropic,
will just support CLI, you know?
Cursor of the world are independent toolings.
They support all models, right?
And so a lot of these products are using yours.
But like, you had the advantage.
I'm got some curious for you.
Do you see or believe in the future,
everybody doing finance will all use one or two models only.
And that's it, you know?
Or do you think like the tools will be the most valuable
in some sense and because the models are all different,
basically everybody will use different models.
And there will be like a winner take all
in all the tools models, you know?
Because I think that open route will be the most interesting
layer because no one, no one single model went out.
Everybody, there's always a need for a variety
and ongoing,
improvement models. Do you think there's going to be a monopoly model every single sector here or or not?
My take is that's even if there's a monopoly, it's very easy to disrupt it. If there's a monopoly,
a new model lab will just come out of China and just disrupt the whole thing.
It's just an monopoly. I mean, if you're a model, it's like a day, that's it doesn't count.
Yeah, a lot of Joe. Yeah, essentially. That's just like a blip and a blip. Yeah. So in my mind,
it's really this really strengthened, like, kind of positioning as like where the liar is a reliable
liar to like get access to all these models, right? And so the harness, except from the model,
essentially. And I think the best harness should be able to reliably extract a workflow from any
model to be honest for the most part. Yeah. And then sure, if the model is dumb, switch model,
open route is there. Yeah. Yeah. That's amazing. Okay. I think that is a really interesting
take here. So for folks that, I mean, I'm sure everybody has heard open route or for the random
people that hasn't, we don't want to check out more or even like want to try out some of the models.
Where would they go to find open router or find you? Just go to open route AI. We also on on
X and open route AI slash chat. If you want to try out some of the model, you'll have to sign up
for an account though, but there's a plenty of models of you can read offers for free together with
some of our partners. You can also go to rankings open route AI slash ranking is a most interesting
page. A lot of people have, you know, what retweet this page, it shows the actual token being processed
by open routers and you see it's grouped by models, but as you scroll down that page, you see
it's grouped by some use cases like programming, like marketing, like, you know, copywriting,
and so on and so forth, right? We also have like model group by languages and eventually we'll have
some, I think we'll have a ranking based on throughput as well. Yeah. So that's what,
AI, you know, precisely. Yeah. It's a Google trend for AI. You take you like a lot of company would
have literally a screen that's showing that renting page just so they can track, you know,
themselves versus other people, which is, it's a, it's a very flattering thing. So also,
an open router we are looking for like extremely high, high HSC engineer. So if you're looking for
a place where you can basically use any AI tooling, you can use any agent out there to build
amazing stuff and almond workflow. Let us know ping us, you can ping me directly too,
Lewis open on our AI. Amazing. Yeah. So even a stack engineer with a ton of agents C1 try every
any AI dev tool out there to be in the hotness of everything open routers in some place to consider,
right? And make sure that we love engineer with a good taste. And especially if we do type
strip because we use type strip for most everything. Okay. Yeah. Well, taste and type screen,
not everybody will put them those two words together sometimes. There's a reason why anthropic
bot bun, right? And a bunch of the type strip ecosystem tooling has, you know, like even though it's
too very new, there are certainly strong tastes in a sense of like, when you use type strip,
why do you need type strip? And like, I mean, the whole wide type group exists, right? Is I think
a fascinating story by itself, I would say, right? It helped with the type safety of JavaScript,
which is already a crazy language, but it's made it really the web, right? So I think it's like,
I think it's a good heristic to five people who are, you know, like, like truly lean into,
you know, a certain ecosystem, I would say. Super cool. Well, hey, thanks for your time. Thanks
so much. I think we could ask some more questions, but just based on time, I want to stop here.
Yes. Thanks for being on our pod, and I'm sure I'll have a ton from Operator.
Podcast Summary
Key Points:
OpenRouter was founded to reduce financial risk for developers building AI apps, starting with a feature allowing end users to pay for model usage directly.
The company expanded from a routing service to a comprehensive inference platform, processing 5-6 trillion tokens weekly, and now offers billing consolidation, model rankings, and enterprise features.
OpenRouter differentiates itself by focusing on reliability and future models, not "magic routing," using fallback mechanisms across providers to maintain uptime even when upstream services fail.
The vision includes becoming a utility layer for AI agents, with potential features like memory support and PDF parsing, while enabling developers to build custom agent harnesses.
Louis predicts more specialized "Devon for X" agents (e.g., for spreadsheets, HR) and sees model monopolies as fragile due to rapid disruption, positioning OpenRouter as a neutral, reliable access layer.
Summary:
Louis Vichy, co-founder of OpenRouter, explains the company's evolution from a risk-reduction tool for AI developers to a leading inference platform. Initially inspired by hackathon stories of developers facing huge bills from free tiers, OpenRouter launched a feature where end users paid for tokens directly, allowing developers to build without financial risk. This attracted a developer base, leading to growth from 4 billion to 5-6 trillion tokens processed weekly, proving product-market fit.
Today, OpenRouter serves as a reliable, neutral layer for sourcing inference across multiple providers, handling complexities like model fallbacks, pricing, and API differences. The platform saves developers time by managing model updates and offers features like consolidated billing, compliance controls, and presets for enterprises. Louis emphasizes a focus on practical needs—cost, speed, and reliability—over "magic routing," and highlights engineering challenges like routing across endpoints with varying capabilities.
Looking ahead, OpenRouter aims to be a utility layer for AI agents, potentially adding memory support and other developer tools. Louis's hot take is that model monopolies are transient, easily disrupted by new entrants, so the value lies in harnesses and reliable infrastructure. ai for rankings and model access.
FAQs
OpenRouter was inspired by hackathon stories where developers faced huge bills from free tiers. The founders wanted to reduce financial risk for developers building AI apps by letting end users pay for model usage directly.
OpenRouter is a reliable way to source AI inference across various providers. It helps developers manage risk, avoid vendor lock-in, and consolidate AI spending, evolving into a control plane for all inference needs.
Developers choose OpenRouter because it saves them time and effort. It handles the complexity of integrating new models, tracking pricing, billing, and API shape changes, so they don't have to scramble to support each new model release.
OpenRouter uses fallback mechanisms to route requests to alternative endpoints serving the same model, such as through Google or Bedrock. This ensures uptime remains high even when a provider's first-party API goes down.
Yes, OpenRouter adds layers on top of its API to automatically parse PDFs and other file types that model providers may not support natively, improving the developer experience.
OpenRouter offers features like Godbrow for controlling model usage across an organization, ensuring compliance and data retention criteria, and presets for creating custom model aliases with system prompts.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.