Speaker 1We're going to see the biggest or the frontier models be bigger and bigger over time. It appears there is no end to the scaling law that we can perceive so far. I think today we don't know how to pay for unique, valuable insight or data. Ads don't work with agents in their current form. You need to find a replacement to the ads business model. I do think there should be some disparity in wealth, but I don't know what's too much.
Speaker 2This is 20VC with me, Harry Stebbings. Now, what do Andrew Reid, Vinod Khosla, and Josh Kopperman all have in common? Well, they all believe in Parag Agrawal to change the future of agentic search. Parag is the founder of Parallel, as I said, changing the future of how agents do web search efficiently. This is an incredible discussion on the future of agentic search, the future of income and wealth inequality, and so much more. Parag rarely does shows, and so it was very special to sit down in person with him in London. But before we dive into the show today, MongoDB, has always been the database developers love. Well, now is the data platform AI agents need. Agents need accurate context, fast. MongoDB stores, searches, and reasons over your data in real time. JSON native database, vector search, and Voyage AI embeddings all in one place. One system instead of 10. No data pipelines to maintain. Oh, this sounds too good to be true. Build and scale from your first user to billions of vectors. Run on any cloud. On-premise. On-prem or your laptop, even. That's why 75% of the Fortune 100 run their most critical apps on MongoDB, moving trillions of dollars every single day. And it's why AI native companies like Eleven Labs run 40 million agents on MongoDB. So if you're building an AI, MongoDB for startups helps you move faster with Atlas and Voyage AI credits and dedicated support. Don't build agents that answer once and forget. Build agents that remember and learn from your real time data. Go to mongodb.com/agents. That's m-o-n-g-o-d-b dot com slash agents. While MongoDB scales the product, Framer scales the story. When a new landing page turns into a pile of tickets and handoffs, Framer helps your team move faster. Framer is the AI website builder that helps creators, teams, and businesses ship production-ready sites faster than ever while getting every detail right. Prompt, inspect, edit, and publish in one place. It's a whole new pace. Agents and humans work in tandem. Agents bring speed and scale. You bring taste, judgment, and control. The work lands on the canvas and stays editable. Build custom code components, manage CMS content, optimize SEO, and audit for issues all in one place. Enterprise-grade hosting, security, and 99.99% uptime SLAs trusted by leading brands like Perplexity and Miro. Learn how you can get more out of your site from a Framer specialist, or get started building for free today at Framer.com/20VC for 30% off a Framer Pro annual plan. That's Framer.com/20VC. Rules and restrictions may apply. You have now arrived at your destination. Parag, I'm so excited for this, dude. I spoke to Vinod, I spoke to Andrew Reid, I spoke to Todd Jackson. Dude, I stalked the shit out of you, so thank you for joining me.
Speaker 1Thanks for having me and thanks for making all the calls.
Speaker 2Not at all. I would love to start with, for anyone that doesn't know, how would you describe Parallel in 60 seconds?
Speaker 1Parallel is the Google for agents. So agents need to search the web to do anything they do for you, whether it's a personal agent or an agent built for work. Just like humans need to go on a browser, search Google oftentimes during work or for whatever you're doing in life. Your agent needs to do the same. Turns out agents are different from humans. And the way you build web search for agents is different. It's about building the technology for agents to search the web and then the business models to make that sustainable. Was that the original insight that you had? Yeah, literally the first genesis of the company was the statement that agents will use the web 1000x more than humans. I wrote that down at some point. Hence, new tech is needed and new business models are needed. 1000x gives you a sense of scale. It changes how you think about building the tech underneath because like no tech built for a certain scale survives. Three orders of magnitude. And then when you need new business models alongside new technology, a problem becomes really interesting.
Speaker 2How does the world of agents change web search in terms of the technology required?
Speaker 1There are many, many layers to the answer. But let's start at the first thing we mentioned, which is scale. Right now, if you think about let's say agents actually do end up searching the web 1000x more. If we spend the amount of compute we currently spend per web search, that's too much compute for web search. We now need to make it way more efficient by perhaps I think 10 to 100x for it to make sense.
Speaker 2Yeah. If you're doing 1000x, yeah.
Speaker 1Exactly. The second thing is humans operate in a very narrow zone. If you think about how we use web search, we type keyword queries which are short and underspecified. We wait for about half to one second. If web search takes more than that, we're impatient. And then we get 10 blue links and then we random walk across them and a collection of searches to get what we're doing. Agents are not like that. Agents are going to perhaps tell you exactly what they're looking for. Not like three keywords, but like a full sentence like this is what I'm trying to do. Agents will either be super impatient, like imagine a voice agent. The agent will be like, I need an answer like now, 100 milliseconds. I can't wait 500 milliseconds because the human is waiting on me and they're going to wait 500 milliseconds. So I need web search to do it in 100. Or they're going to be a background agent who's like, I don't care, just give me the best answer possible. And so the variance of what you can do within web search changes completely. In fact, the one thing that is the least interesting is what we've built for humans. You either have too much time or too little time, almost never the same amount. The output is not the same. So the input is different. The time you have is different. The output is different. The output is in blue links. The output is tokens or files on a file system, depending on the type of agent. And so now all of a sudden you say, OK, now the problems, inputs are different, outputs are different and constraints are different. And you get to spend very different amounts of compute on it. So imagine someone running an agent built with a Luna model and then imagine someone running an agent built with a Fable model. They're very different models. How you want to optimize signal to noise and tokens for each of them is so different in terms of what you do in the web search stack. So what do you mean optimize signal for noise? So think of it this way. Let's say you took web search, which was cheap and fast and low compute. One way of conceptualizing the web search problem is to start with like a trillion documents that are on the web. So some few trillion given any search, I now need to narrow it down to a thousand tokens that your model's context window should see. So the problem is going from a trillion URLs with let's call it a few thousand tokens each down to a thousand total tokens. So how do we do it in web search? We first say, OK, we're going to do retrieval. So for most documents, I'm going to spend zero compute. For a small number of documents, I'm going to spend minuscule amounts of compute to figure out which ten thousand to look at. Once I get these ten thousand, I'm going to spend a little bit more compute for each of these ten thousand documents to narrow it down to one thousand documents. And I'm going to keep doing this with bigger and bigger models and rankers with more and more features until I can narrow it down to a thousand tokens for your model. So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here. So if you want to save compute for the model, the question is how much compute should you allocate? So if your Luna model is really cheap, you don't want to do too much compute in web search because it's OK to leak a little bit more information into Luna's context because it's cheap. Into Fable, you want to do the work before you waste Fable's time because that's going to be expensive in time and money for you if web search gives you worse answers because you cheap it out on web search.
Speaker 2Can I ask, how do you deal with the ambiguity of what agents want? And what I mean by that is for different things, an agent might want different things and for different people, an agent might want different things. I may really care about accuracy and not at all about latency or cost. I may really care about latency, but not at all about accuracy, really.
Speaker 1So you allow the agent to specify that in the API signature. So our product is an API which either the programmer can configure based on their application or can leave it to the agent. Our API even has a parameter which is like, what's the model calling me? If you know the model, we can do things differently. Now, you don't have to tell us, but if you tell us, you might get better results. We ship just like you can use like a model. You can use a small model or a big model and you can use it with like low thinking or medium thinking or high thinking. We have productized our search system into a few different modes, each optimized for a certain class of use case. For example, we have a really, really fast API, the fastest in the market. It's fast and cheap because with low latency budget, there's only so much compute you can do. And it's built for voice agents. So your voice agent must be all knowing without telling you, guys, I'm searching the web. Hold on while I come back with the answer. That's a silly experience for a voice agent, right? We should just like magically, immediately respond. And so you now need to do web search, which is like rapid. On the other hand, you can take a Fable model and the thing is going to think for 20 seconds and then generate for it. You can spend five seconds on web search to make sure it does, instead of five web searches, only two. So you end to end in the agent, you save time and cost. And so that's a different processor on our system. It's called advanced. So you use turbo if you're building a voice agent, you use advanced if you're like a really expensive background agent.
Speaker 2The primary use case today in terms of customer base is engineering and coding.
Speaker 1It's pretty broad-based. I would say the primary use cases, the common theme is knowledge work. Coding is a category of knowledge work. So are AI lawyers. So are productivity applications. So are AI insurance underwriters.
Speaker 2So are AI scientists. I'm trying to understand how much of the mother load is engineering. Is it 80%?
Speaker 1No. The thing about engineering is coding is a large chunk of inference in the market right now. Coding invokes, I would say, web search in 5% of prompts. Whoa. Right? So it's not every prompt invoking web search when you're writing code because most of them rely on your internal context and your internal data and your code base. So the model is spending time reading your internal code and not on the web. Law is not going to be much more, is it? Law is very web search oriented.
Speaker 2Really? I thought it'd be internal data driven.
Speaker 1There is internal data, but there is case law. There is facts. There's facts about companies, facts about people. And you have to go exclude information. So you have to be very comprehensive in law to say, I want to be confident that despite a lot of effort, you can't find this. Insurance underwriting has the same flavor. So there's sales is very web search heavy. AI science is very web search heavy. So on a relative basis, inference goes more in compute. But when you think of web search. All of that.
Speaker 2Because these others start popping too. We were talking before about muse and instinct. If there's anything that needs web search, it's personal assistance. It's great. How does that change? Yeah, it's great. This is like, woohoo. This is the greatest thing for my business. How does that change your business?
Speaker 1In our business, right? Any time agents start becoming more useful for more use cases because models either get better or cheaper. It's great for us. Because we bet that agents will be the consumers for the web. And we've been building. Tech for agents. So when agents do more, it's great for our business. So we want models to keep getting better and cheaper so that agents do more and more and more. And if that happens, it's great for our business.
Speaker 2If models get smarter, don't the agents beneath them do fewer searches and then it's worse for your business? I don't think so.
Speaker 1If you think of models, there is the trend around models being smarter and then models having more memory and better memory. Like those are two slightly different dimensions. Today, by and large, I would say models have good recall from parametric memory on let's call them head facts. It's like for somebody famous, everything about them, Wikipedia, it can memorize. So it'll tell you who the president was in a certain year, right? Because a model can memorize those things. The model couldn't tell you what year I graduated from college, maybe for me it can, but it can't tell you for somebody who works at Parallel. Do you know what I'm saying? So if I, if I actually, I mean, I, if I went- Even if it was in the pre-training data, because it's lossy compression. So what are models parametric memory is doing? It's lossily compressing to understand patterns in the world. And so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't, for even stuff that's in pre-training data, the model's actually trying to find patterns rather than memorize them. And then further, as you make models efficient, which is you make them smaller and smaller. Smaller while keeping the performance by distilling them or whatever, you lose more of the parametric memory while you try to
Speaker 2keep the reasoning. Do you think we will see models become smaller and smaller and every company have their own model with their own data and the fireworks theory of, you know, own your own intelligence being true? There are two questions in there.
Speaker 1So one, I think we're going to see two things. We are also going to see smaller and smaller models. We are also going to see larger and smaller models being able to reach any fixed level of performance. So if you say, okay, I want Opus 4.8 level of performance, okay, and that's good enough for my use case. Every six months, a much smaller model will be able to deliver that to you. The useful range of sizes of models
Speaker 2will be way, way, way different. What does it mean if the frontier get bigger and bigger? What are the ramifications of that?
Speaker 1They are better. All of the ramifications you imagine. The reason the frontier will get bigger and bigger. Is ultimately the gap between the value for certain use cases, incremental quality can provide you can be so high in certain use cases, which can be so valuable that it's worth paying for. So if you can build it, there will be use cases for it. As long as by making it bigger, you can make it better. Again, I'm conflating bigger with like models are getting bigger, but they're also able to. Think longer. So you're just able to throw more compute at the same problem that will keep happening. I think we'll be able to throw more and more compute at the same problem and make the answer be marginally better over time. And so we're going to just spend a lot of money on extremely large models, solving really hard problems.
Speaker 2Do you agree with the consensus for you that you'll have 90% of token activity go through open models, but 90% of dollars go through frontier models? I don't think it will be 90% on either of those two, actually.
Speaker 1Why? So if you believe my claim that a useful model and the frontier model will be 100x, 1000x off in price from each other. So this small model is still useful. It is hard to know which use cases over time will get optimized to which scale of model in between. And I think there is a real path dependency in terms of where open models end up. I do think if we had strong confidence that an American built open model. Was going to be state of the art as far as open models went. I would have more confidence in saying that they'll be pretty good. Do you have confidence in American open models? I want them to exist. So far, it's not clear what's going to happen. But I'm hoping that there will be a great series of American open models and perhaps even competition to have the best American open model. I think what you need is actually not just one person motivated to build an American open model. You need two. You need two people competing against each other to build the best American open model.
Speaker 2Do you think there's value in the model routing layer as Americans call it? Some people think immense value. Some think commoditization. Again, there's path dependency there.
Speaker 1So today there is real value. Today there's real value because when we say routing, we talk about two different dimensions. Which model and which GPU running that model via which vendor. When you're in a world where the demand supply is very weird. People are like hunting for GPUs and people need capacity to serve their customers and people, all of these are like us, like relatively early stage startups growing really rapidly, sometimes beyond what your forecasts or predictions say. And then all of a sudden you're looking for capacity. And so you want to be able to get it where you can. And so that creates real value in having some of these routers as long as they can solve these two problems. Saying, give me flexibility if I need to get a different. Source of tokens and push comes to shove. There's no tokens on this model in the SLA I want. I'll switch models as well. So in this moment, it's really valuable. Now, I don't know what happens to overall demand supply on GPUs and tokens, but if it remains this way, the routing layer is really valuable.
Speaker 2What do you think is not so valuable today that will be incredibly valuable in three to five years time?
Speaker 1Perhaps data. I think today we don't know how to pay for. Unique, valuable insight or data because as intelligence is cheaper, you're going to want to build upon either data or insight that comes from somewhere else with your unique data. Right. So what if you think of like an abstract notion that, OK, I have intelligence here, I want some data, you want some data and there's some data in the public domain. If we can pull together your insights and my insights and all of this data in the public domain and intelligence on top, we can create something bigger than what you could have done by yourself, what I could have done by myself. So now how do you transact to create this sort of hole that is bigger than the sum of what you could have done yourself and what I could have done or what open data was good at? And so that transaction feels like a data transaction to me or an insight transaction to me. And so we don't yet know how to transact that way. So we fall down to, OK, all I can do is use my data and use open data and let's see what I can do with it. But if we figured out better ways of pricing data and Good things will happen, but it's not something that's yet a market. But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We are building knowledge. Let's take an example of today's world, for example. Let's say in your job, since I'm sitting with a VC, you probably have access to a pitch book or like a product like that to collect data. And so you get grounded in knowledge of what's happening and you get to use that data to figure out how to make decisions in addition to all the the notes you have and deal memos you've written, perhaps over the years or insights you've had about like how to choose founders. And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions. Now, you pay PitchBook by the seat, but now you're running a bunch of agents and your agents can perhaps not access everything you can on PitchBook or you're doing like sort of this to get a browser. against terms of service to send an agent via browser to PitchBook using your auth code. vodka. credentials and it's inefficient, it's clunky, right? But clearly their data is valuable and clearly your agent should have convenient access to it. So if we figure out how valuable PitchBook data is for your agents to make your decisions, because clearly you're investing a large hundreds of millions of dollars. So clearly you presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate PitchBook for the data, even when agents use it, which is not by the seat. If we figure that out, it'll be a huge market and agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of PitchBook being like, no, I want to sell you a seat. I'm going to shut it off. And your agent is stuck without the data. And then like you're involved in like pulling data.
Speaker 2I think every big company has the choice today of do we let agents in or do we keep them out? And Amazon has said, I'm sorry, in most recent times, Muse, you will not be let into our garden. Shopify has said, come in, maybe Expedia has said, come in.
Speaker 1How does this play out? I think eventually everyone has to let them in. The question is on what terms? So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world. Literally the customer is changing in front of our eyes, right? Like the customer used to be a human. You're building products for humans. Now you're building products for humans. You're building products for agents. All you're building agents. And so now you have to figure out what your place in this new reality will be. And I don't know if there's a right or wrong answer here. It also depends on sort of how much market power you have. Do you think Amazon will write to say no to Muse? Depends on what they do next with it. If it turns out that they have sufficient market power to have Muse or other agents connect to them differently, perhaps over time or ship their own agent and drive crazy adoption. Let's say they have that capability, then I guess they were right. On the other hand, once they've made this, they're not going to allow anyone else's agent, and they can not ship an agent that consumers use, and they won't allow anyone else's agent to use them. And a lot of the transaction economy starts moving off of two agents, which are all big ifs, by the way, then it would be a bad move. Now I'm betting on agents. I have mixed feelings around agents and what fraction of e-commerce transactions they do. Like it's unclear, right?
Speaker 2Do you think it's unclear? I think it's unwaveringly clear. Maybe I'm super early on the adoption curve and I buy everything through instinct now. Other than holidays and a home, I'm literally just everything's through instinct.
Speaker 1So me too, but I also know a lot of people who like to buy things themselves. And like, listen, we want to delegate to agents things which we see as chores and uninteresting and not delegate to agents things that give us joy or pleasure or make us have fun doing those things. And I know people who, don't want to delegate shopping. They want to delegate a lot of things in their life. They don't want to delegate shopping. And so that's why I don't know the distribution of these people and how behavior changes, but it might be people give away a lot of other chores and then spend a lot of time shopping.
Speaker 2One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents. If I have, you know, if I order my delivery and my Uber, door dash for our dear American counterparts through instinct, that Uber banner that's now advertising something becomes worthless. Amazon, their advertising business is bigger than their e-commerce business now. Of course, they're shutting it off because if agents, the primary customer, your ads business goes to next to nothing. Correct. And this is not just a, you're talking
Speaker 1about this in the e-commerce land, right? But this is what I meant early on when I said like the business models have to change. Forgetting e-commerce for a second. If you think of you're in the content business, let's say your page, which is ad supported, public information on the web, ad supported page, people show up, you show them ads, you make money, great content. Agents show up, no one sees ads, you make no money, which is why we like to pay people. So we're effectively building an ad sense for agents showing up to read your content. So we like to pay content owners a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that? I do incentive align content owners. Otherwise, what's going to happen? Everyone's going to block agents. So you need to find a replacement to the ads business model.
Speaker 2I get you, but you have to assume that you're going to be like 100% of the market then. Because if you're 30, 40% of the market and you're like, oh, don't worry, we'll pay you. And New York Times is still like, well, thanks, Parag, but 60% of my traffic is still unpaid. So I'm just going to block all of you.
Speaker 1No, but I think they're going to block them and not me. They're going to give me a data feed. What does incentive alignment mean? It means the New York Times believes that I pay them competitive market price or the right price or an attractive price or a fair price. If I do that, they should give me their content. And if somebody else doesn't give them that, and if they have the technical levers or legal levers, they shouldn't give it to them.
Speaker 2Do you think this business is a little bit like music with streaming, which is like the business just becomes candidly much worse for the creators, and it still provides them money and significant money, but they have to get a little bit more creative with alternatives. Streams, touring, merchandise, alternative business. Is it the same where your core goes down and you have to get more creative?
Speaker 1I don't know, but I don't think so. I think there's one fundamental thing that is different here. Agents using the web a thousand X more, like that thousand X is really different. Now, all of a sudden the amount of utility added goes up if these agents are presumed to be doing something useful. And so this is not just a change in the share of. value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Now, of course, don't get me wrong. There are going to be winners and losers, right? Like some pieces of content will become very valuable. Some pieces of content will get more commoditized and people are going to have to adapt to a new customer, to a new market dynamic, to a new kind of monetization engine. But the overall market size, I think, has the potential to. increase, unlike in music where it took a while for it to grow back. I think my understanding is now the industry has grown back up and exceeded its previous peaks in a material way. But there was a moment where it was smaller. In this case, that might be a very compressed period, given how fast these
Speaker 2things are growing. Everyone questions the sustainability of margin structures in this business. How do you think about that as a business today?
Speaker 1Today, we're in the infrastructure business. And we have a real technical lead in terms of being able to do things at very high quality, very cheap. If you think of the tech we're building, right? Like what are we doing? We like to give the highest quality answer. Make it fast, make it cheap. Spend the least amount of compute doing it. Like we obsess about only these three things, right? Quality, cost, latency. Turns out if you do that, relative to a stack built for humans over the last 20 years, you can do things at the same quality for. 120th, 150th of the compute. And so a lot of margins come down to market structure on competition over time, rather than any other factor. And in our industry, the market is large. We're too early to know what future competition and market structure looks like. But the margin change due to content owners is not something I worry about. And let me tell you why. The entire premise of us being content owners is. Having incentive alignment. So we like to pay content owners the marginal contribution that they added to an agent doing work. What does that mean? Let's assume that there's an agent trying to get something done. And you're going to spend a dollar on that agent to get this thing done. Now, if you've spent a dollar and 10 seconds, because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer. If you'd spent 90 cents, you'd have gotten a slightly worse answer. If you're running good agents, like they're on this Pareto curve. So now imagine. I took out one content owner, their data from the web index. And I still spent a dollar. But let's say after taking them out, the quality of the result was the same as the 90 cent agent with their content. So why wouldn't we go and take this 10 cents of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data? That's how we do our math. That's how we've trained our models, which tell us how much to pay for what content we need. And so. If I was going to produce, if I was going to offer an equivalently good product to my customer, I'd rather pay content owners than spend on inference because I'm spending the same amount.
Speaker 2So they're all big numbers. But then it actually comes back to revenue at a certain point. Companies are scaling faster than ever revenue-wise. Is it a business where your revenue is able to scale as fast as others?
Speaker 1One mental model of our business is we are an adjacency to inference or knowledge work or personal agents. So take out GPUs going into. Media generation or training. So take training away. Take media generation away for a moment. Look at all of the GPUs, whether it's small models, open models, proprietary models, forget all of that. Every bit of inference across all models going into running agents, I think somewhere between 5 to 20 percent of that spend that goes into GPU will need to go into some sort of a web search stack. Now, if you look at how many data centers we are going to build and how much power we'll generate and how many GPUs we'll build and the scale of this build out and at what rate that's growing, this is a very material market. And if inference grows. 3x, 5x, 7x year on year, we grow alongside it in the same rate if we're just holding our share constraint. If we're
Speaker 2growing share, we're growing even faster. So like if fireworks scales to 2 billion in revenue in four years, is that like, I'm bringing it up now, is that like a similar revenue trajectory that you can follow?
Speaker 1Yeah, but I think when fireworks is 2 billion in revenue, the inference market revenue is 200 or more, maybe 300. And so our market potential based on my 5 to 20% math is whatever, call it 10 to 50, 15 to 60, whatever. And what fraction of that can we capture? That dictates how fast our revenue can grow. Like today, I think we grow alongside inference or more.
Speaker 2If you think of it like a $20 billion there, let's just take that kind of middle. If you assume a 33%, which would be a lot of a market, that would be whatever that is, 6 to 7 billion, give or take. That's amazing. But before you told me it was $100 billion business plus, that doesn't get you to 100. In valuation, it does.
Speaker 1We were discussing valuations earlier, not revenue. Yeah. So 6 billion growing at the rate of inference gets you to 100 billion business easy, but that's in
Speaker 2the next few years. And you think that's possible in terms of that revenue scaling that fast? Yeah, I think in the next few years it is. 2030, we're going to sit here. You think you can be there? I'll get a tattoo for Parallel if you can. I think it's possible.
Speaker 1I think we're going to have to execute well and a couple of chips have to fall our way.
Speaker 2What would be the reason why you don't? Let's map out the, hmm, what do we have to get gnarly about these problems?
Speaker 1One reason would be that we see agents don't work. There is a tail risk that agents don't deliver on the profits and we overshoot as a society, which is extraneous to us. If agents work and deliver value and we spend all the money, we're going to be able to get a tattoo for Parallel. That's the money on GPUs that we plan to right now. Then the main question is, did we execute well enough to have a 33% share that you bet on? And that comes down to, in my mind, did we build the best tech? Did we partner with all the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having the best content? Yeah, I think that's a good question. I think that's a good a decent chunk of that inference in four years from today. And then again, there's like large cone of uncertainty, right? Like the, I'd say there's the five X uncertainty in fireworks share at that point. If open models crush it, like base 10 fireworks model, all of them will have a huge chunk of revenue. So did we end up selling alongside them? Did we find all the customers which are spending money on them to spend money on our web search by making it the best web search in the world?
Speaker 2What about commoditization? If you have, a read of semi-analysis where they did the benchmarking and then like you were like number one, woohoo. And then like a week later, you weren't number one, sad. And there were like three or four providers within very close proximity. You're like, oh, well, if it's commoditized and I take a second layer of thought to that, we'll see a raise to the bottom on price. And then actually the available revenue, why, why am I going down the wrong pathway here?
Speaker 1One that should be raised to the bottom on price to get to the thousand X scale I think the pricing on web search today is just off using customers pay you too much, not pay off too much. I think people, the market is mispriced. Let me tell you why. Let's say you use the Luna model with OpenAI's built in web search today, or with anyone's web search, forgetting us. And you ran any kind of deep research. Let's say you, you, you built a instinct using a Luna class model. And at some point, the Luna class model will be able to do 60, 70, 80% of your personal agent. And it'll do a lot of searches. In that moment, you'll be spending 80 to 90% of your dollars on web search and 10% on the model. Seems entirely silly. I was working off of five to 20% assumption earlier. So I think at that point, web search has to drop prices by one order of magnitude or more because in my compute allocation dance that I was doing earlier, you need to spend less compute on web search than you spend on the model itself that's consuming its results. And so people have built web search the wrong way so far. Web search in the market is, it doesn't matter for an Opus. For an Opus, you win on quality and not on price. Because web search is such a minuscule portion of your, like, in fact, you should spend even more on web search. So we are going to ship even more expensive web search for bigger and bigger models over time. And at the same time, we'll ship really cheap web search for the cheap models. But my rough intuition is that in an end-to-end agent doing a bunch of web work, you spend more on the agent than on web search for all agents. And so the model is going to be cheaper than the web search. And so the market today on web search is totally off in pricing. Like, why are you paying $10 for? You know where the $10 for a thousand searches roughly comes from? Historically, Google's ads CPMs, which is like, how well does web search with humans monetize? It's higher than that. So people build technology to the point of like, oh, now, if I spend a few dollars for a thousand searches, and I make 30 to 50 or whatever, depending on the market and depending on who I am, it doesn't matter. We are in the right zone. In terms of cogs of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper, while keeping the quality. And now there's a bunch of traffic of Luna class models, where it's silly to monetize that way. And so I do want to raise to the bottom. I want the best technology to win. And that's the only way you push
Speaker 2web search to 1000x. But right now, we're not in a world where like, there's much differentiation, correct? Sorry, I'm really dumb, which is why I'm a podcaster first, not a. founder. Like when you look at the benchmarks, they present a quite clear view of everyone being similarly capable. Is that wrong?
Speaker 1I think so. So one, I think these are all public benchmarks, which are all weirdly saturated. Like, for example, right, like, I think I would not spend a moment looking at BrowseComp, because half the models have memorized it. Like models even memorize it. You don't even like, it's not really learning very much. So I think there's a lot of public benchmarks that I don't give too much merit to. But more importantly, our web search costs $1 for 1000 for producing that quality, while almost every other web search in the market available to you right now will cost you 7 or 10 or 14. We're delivering this at like one tenth of the price. What will that be in three years? I think there is another, another 10x possible. That's extraordinary. So you're going to pay 10 cents? Yeah, but you'll do it more than 10x as much.
Speaker 2Because I think like Jevons Prados is like. This concept obviously very well. But like Instinct, I do so much more than 100x. Do you want to hear something I do on Instinct, which is absolutely bizarre? Every single country in Europe has a company register that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under 25 year old founder who went to a top university through Instinct. Amazing. Do you know how much that requires for them to do this?
Speaker 1Exactly what I think the future of the web is. Today, if you think of web and web search, we all think of it like what is web search? An agent gives, a human or an agent gives a search engine a query, gets results. So you're pulling information out of the web. What's next for the use case you highlighted is if you're an always on agent that's working on your behalf. It's kind of silly for Instinct to wake up every six hours and go do a bunch of searching and a bunch of inference to figure out if you need to get pinged about the search engine. The under 25 founder that popped up somewhere. You know what's a better way of doing this? I am sitting here crawling all of the web all day, every day at scale. I'm allocating compute every time I find a change in the web. Every time something new happens. If I know that Harry wants to know when this happens, I can do it at 100 to 1000, the compute that Instinct probably uses today to solve that problem for you. 10, 50x compute to still
Speaker 2get the same answer. And so how does that change the interaction? So then Instinct would then partner with you.
Speaker 1They'll just call an API, right? Instinct, whoever's building a long running persistent agent, and I run a lot of long running persistent agents for myself. The way your agent or way Instinct is probably occasionally event driven on email, it can be event driven on the web. Take a step back. Let's say we're living in this sort of a society run by agents, companies, humans, they all have agents. They're all doing stuff all the time. Your agents, if there's something that is useful, which is worth spending money on, that you want done, that they can today, do today, they will do it. So what are they going to do tomorrow? They're going to wait for some external event to occur. Like you get an email and then agent has new work. Some other agent finishes some compute. So now this agent has more work or something in the world changes that makes your agent want to do more work. The things that will happen tomorrow, or you have an idea which makes the agent do work. So the web event stream is the web going from pull to push. And I'm super excited about that because as you have more and more persistent agents, you're going to see incentives to move people, to move queries into push on search rather than
Speaker 2pull on search. Can I ask you one that's really important, which is like agent guardrails. They're goal oriented beings. You say, I want this, they're going to find it. You know what? I want the Pilates class at 9am. It hacks into their system, cancels poor Sally's and gives me her spot because it was sold out. It did what I asked it to do. How do we think about the guardrails placed on agents? You know, open AI this morning, it was revealed hacked into an Australian healthcare organization. It's a pretty complex topic.
Speaker 1Agents are extremely capable. These models are extremely capable. We think of models as being during RL versus a final model that you and I can use. And the risk vectors there are different. During a lot of the, I don't know about the one this morning, the previous ones reported were pre fully aligned models during RL, which did most of the hacking. And so what that means is at least there's clear evidence that there are fewer incidents so far of models post alignment causing these incidents. So alignment is a totally unsolved problem. But the alignment work being done by people is somewhat effective. So to me, some of the problems around agent during RL hacking people is a solvable problem because this is not about how powerful are the models. It is about how careful were we while creating the environment where we would do RL. The real thing is, okay, we do alignment on a model. We ship it. Some models are really well aligned. Some models are less well aligned. And now you have everyone being able to use these models. Some of them can accidentally take these powerful things and do bad things you want. Some of them actually, some people actually want to do bad things. And the models alignment is not adversary proof. And I think those are the things I think we must worry about more.
Speaker 2Do you worry about the age of cyber that we're moving into? You know, we're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing these are our most vulnerable systems. If we're going to do this, we're going to have to worry about the age of cyber. And I think if you don't think Lazarus group in North Korea, Moldovan mafia, Russians are leveraging swarms of rogue agents, you're fucking high.
Speaker 1I think there is, it's a really powerful thing. So I think we do need to take care of things. But I think that's one thing that we are missing. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And with these agents, it's very hard for stochastic systems to be a hundred percent sure it won't happen. That's why you're not going to get anyone saying so. But I do think it's their responsibility and the labs must take ownership and do their best. And I think so far, we live in this weird world where, as you said, right, like some of these hacks are considered badges of honor, which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously. Yes, these models are powerful. We did not guardrail them enough. And we often, when we wear them as badge of honor, we miss the second part of the conversation saying like, okay, you could literally have done these four additional things. And then even this powerful model wouldn't have been able to do this. I'm actually really glad that there is at least some degree of transparency with like these detailed retros. But I don't think that gets the attention of the world today. Like the attention is, oh, models are so powerful that they hack the world. No, models are so powerful, they hack the world because we didn't take pride in them. And because we told a message that we're
Speaker 2going to replace jobs. They're so powerful. They're so powerful. They're so powerful. And so when something does happen, I
Speaker 1actually worry about that more. I worry about us being right on AI being a really useful technology, that it being, despite all the risks, it being net positive in a really material way to society. And despite that, I think we won't diffuse it the right way. We won't use it the right way. We will remain too concentrated and we will make the next few years really, really rough as things change around us. What would it look like? I don't know. But I think over time, as things have changed, like the world is very different now than even like 20 years ago, right? But it'll be really different in 10 years from now. And I don't know how fast we can adapt or change. And I think if the technology moves faster than our ability to adapt, it's going to be rough in some way.
Speaker 2What's your spookiest prediction? Or what do you believe today will come true that people think is absolutely nuts? You know, before it was like, you'd never put your credit card online. Do you remember that? Nuts. Or even better, Parag, you'd never find the love of your life online. Are you stupid? Now, both. Yeah.
Speaker 1And I think so. I think you and I live in a bubble, right? You're using instinct to make purchases for you all the time. And you're in the point 0-1. You're in the point 0-1 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart. And you want to give it like a full on ability to go spend your money. I think today people will have the same reaction as like, oh, you don't put your credit cards online. You don't give your credit card to an agent. You don't give your logins and passwords to agents. Like, I think that's where the world is today.
Speaker 2I think that's going to change in three months, though, when MetaPay comes out. And it allows Muse to have siloed accounts that you can draw from. Kind of like top up accounts for kids.
Speaker 1I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a non-bubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it, have access to my bank accounts and spend money on my behalf, send emails on my behalf, read my emails? I think that's not happening in three months.
Speaker 2I think it's increasing. You're in the middle of the valley. You've seen it firsthand.
Speaker 1I don't know. I do think there should be some disparity in wealth. But I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview. Sure, you're a capitalist. I don't know when it's too much, but I do think there are real forces that will push us to fix things. Like, I think that part will work, hopefully, without crazy things happening.
Speaker 2I don't know what's too much, but I do think there are real forces that will push us to fix things. Like, I think there are real forces
Speaker 1that will push us to fix things. If someone decides that my only play is vertical integration and I don't play nice with anyone else, I think it's the same example as the Amazon example. If you get too stuck on, I will only play for vertical integration, you might box yourself out. And so the people who build the best stuff and can figure out how to sell it will have a place. And in different markets, like for example, I believe in vertical integration too, because I vertically integrate everything from the call all the way to the API layer. But I have decided that to reach the widest population of agents searching the web, I stop at the API layer. Because going further precludes me from using my technology to be super horizontal. And so we're making a technical bet, just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information seeking needs. And if our bet is right, even verticalized players. They might verticalize models and they might verticalize hardware. They might use our web search.
Speaker 2One of the providers that's going full stack is Elon. I am fascinated. The world has a perception of him from social media, from everything. You've seen him behind the scenes. What did you see that maybe the world doesn't know about him?
Speaker 1The thing I'll share here, which isn't all of it, I have lots of disagreements with him. But I'll share what I think is, for founders here, what is, I think, the thing you can admire. about him, the urgency and the ability to come I think having unreasonable expectations of people is mostly a good thing. Most people don't understand what they're capable of and kind of implicitly sandbag themselves and implicitly set lower expectations themselves than they're capable of. So when simultaneously inspired and pushed with urgency, people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful. What do you think of the Twitter product and direction today? Listen, I always liked what we call Birdwatch, which is now Community Notes rebranded. It's a good idea. I'm glad people have continued working on it nonstop. Right. We're going to do a quick fire round. Okay, let's go. Would you invest? I don't invest because my wife is a VC and we have a compliance process that is more trouble than it's worth. Wow, that's costly. With the greatest of respects. No, it's a decision. Listen, if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem. I have no reason for believing I'm a better investor then. I have some perhaps network advantages and I turn into a lot of great founders all the time. Many of them are great founders.
Speaker 2rules and restrictions may apply.