20VC: Are OpenAI and Anthropic Overvalued? The Open-Source AI Reality | How Token Costs Will Fall 10x And Usage Will Explode 100x | The Future Is Not One AGI; It's Millions of Specialised Models with Lin Qiao, Founder and CEO @ Fireworks
77m 6s
In this interview, Fireworks founder Lin Qiao argues that the future of AI lies not in a single, general intelligence owned by one company, but in specialized, private intelligence derived from enterprise data. He contrasts this with the AGI-focused approach of frontier labs like Anthropic and OpenAI, which he views as essential infrastructure—like power lines—but not a replacement for customized solutions. Lin predicts a 10x reduction in token costs over three years, which will drive 100x usage, and believes open-source models are key because they give enterprises full control over customization and cost. The investor, Harry Stebbins, justifies his $10 million investment after a 15-minute meeting by citing the world-class team, the rapidly growing inference market, and Fireworks’ potential to be a $500 billion company in a world of many specialized models. Lin addresses concerns about Chinese open-source models by stressing that enterprises can implement their own guardrails to align with their unique values and needs. He emphasizes that product-market fit and durable business are now separate concepts, as scaling with expensive general-purpose APIs can lead to bankruptcy, whereas open models offer affordable, customizable alternatives for production-scale use.
What I don't want to see is there's only one company owns intelligence. I think that doesn't make sense to me. I think last year is the year of coding. And this year is the year of co-work. I do think the cost of token will go down drastically. 10X cost reduction in the next three years. And this 10X cost reduction will drive 100X usage. We absolutely are not going to move into application layer. Very clear to us. Whether we will move down into view data centers, and so on, that could be always be on the table. This is 20 VC with me, Harry Stubbins. And in the hot seat today, a founder who I wrote a $10 million check for after just a 15-minute meeting. Lin Kwaow, founder at Fireworks. This was one of the easiest investment decisions that I've made in a 10-year investing career. Number one, the team is one of the best in the world. They've worked together for years. Number two, the inference market. It's growing insanely fast, and it will be one of the biggest markets in the world. Number three, the traction. Honestly, I get in trouble for this. Everyone's like, oh, don't say triple, triple, double, double, double, dead, Harry. It's because that's what they are. Well, that's the sad truth. When you can invest in Fireworks, which is scaled to 1 billion in error in four years, that's what you should do. Number four, hiring ability. Honestly, the single best founders, they're able to hire the best in the world. Lin just hired George Hugh. He was the former president of Salesforce. And then number five, upside. How much money can this make? Well, do you believe AI will be transformative to the way that we live? Do you believe we will live in a world of many models, which are horizontal and specialized? If so, Fireworks, honestly, it can be a $500 billion company. I don't know about you, but this train is leaving the station. Juju! But before we dive into the show today, founders face a different set of challenges at every stage of growth. For Sid's shade, co-founder and CEO of Dmatrix, JP Morgan delivered the guidance and expertise to help navigate what came next. He credits JP Morgan's high touch approach with supporting Dmatrix as it grew and expanded internationally. Whether you're in the early days or expanding into new markets, JP Morgan helps startups navigate complexity with real confidence, offering personalized guidance and deep sector expertise. Find out how JP Morgan helps founders at jpmorgan.com/growwithoutlimits. JP Morgan is the bank of the innovation economy. While JP Morgan powers your finances, Navan keeps your team moving. Did you know the industry average for booking a business trip is 45 minutes as a massive waste of your team's time? Or would Navan, your employees can book a trip in just seven on average? Navan is the AI-powered travel and expense platform designed for companies that value efficiency. It drives real business impact through high employee adoption and automated policy control. Now, the built-in AI approves in policy bookings and blocks the rest automatically. This allows finance teams to stop chasing receipts and skip the month than chaos. And you get this real time visibility that can save your company up to 15% on your travel budget. And that's why leaders like Visa, Stripe, Figma and even anthropic rely on Navan these days. Go to navan.com/20VC today to see for yourself and you'll get a chance to win two business class flights anywhere in continental US. No purchase necessary rules apply. Head over to navan.com/20VC now. While Navan keeps your team moving, Base 44 helps you build faster. You have the idea, but with most AI tools, you hit a wall. The setup, the config, the gap between what you pitched and what you actually ship. Well, Base 44 is where that wall disappears. You just grab it. Yeah, Base 44 builds it. Apps, websites, AI agents, real working products, built-in minutes using nothing but plain language. And it's all batteries included. The backend, the database, the authentication, the hosting, the heavy lifting is handled. So you just really stay in the flow. This doesn't just take the busy work off your plate, but it gives you an advantage and pushes you past what you thought you could build alone. So in this market, fast is the baseline. To win, you just have to be first. Base 44 is that edge. The move that skips the troubleshooting and gets you straight to the breakthrough. Build your next thing at base44.com. That's base44.com. You have now arrived at your destination. Lynn, I am so excited for this. I had so many great things. I just got off the phone with your co-founder, Deema. I speak to Alfred Lynn, Sonia, Matt Miller, many more. So thank you for joining me. Oh, thanks for having me. Now I heard Eric Vissier has a rule, don't invest in big tech directors, but he broke that rule with you, which is very special. I think so too. So funny story. After we decided to handshake, he did call me and said he talked with one of his advisors. And his advisor questioned him, hey, how many big tech exactly have seen being successful in starting a company? Very few. And he told me that I was surprised. Like, are we breaking our handshake now? No, but since then, we work very closely with each other. Eric is one of the best. You also started the company when you were 48? Oh, yeah. That's quite late. How do you reflect on being a 48-year-old founder when we glorify starting a company when you're pretty much 15 these days? I didn't think deeply about that. I always want to have a tech business myself. I actually want to start a business in 2015, because I'm a first-generation immigrant and came to US in 2000. I did my PhD in the distributed system, convert science, especially focus on databases. And databases are a very concept system to build, a lot of different objectives optimized for, and pretty much touched after I joined Research Lab, and pretty much touched every single aspect of processing data. And then I moved to LinkedIn to further it down to build systems and products to be used, drive real impact. And at that time, I feel I'm ready to start a company. I know all the tech, I know what product to build. I have a business proposal. I have a list of people I want to start a company with. And I spend time thinking about it in that past, because I don't think I have the skills that I'm people to build a company. It's not just about product, it's not just about tech, it's actually about people. And I decided I want to go to a place I can learn the most of people. And the best company at the time is Facebook. It's a rising star in Silicon Valley. And secretly, I was planning to learn for one year and live and go back to do my own business. I stayed there for seven years. - So with fireworks, you saw something in inference that the world was not focused on. The world was focused on training. I think it's helpful for people to understand kind of the stack, because beneath you, there's obviously kind of chip providers and you're in videos of the world. And then you've got above either model providers and you sit in between. Why is that a valuable part of the stack and not a commodity? - That's a really good question. But why bother specializing intelligence? Why not just use generalizing intelligence and you worry less things, right? You just kind of build on top of a API that provided by Frontier Labs. Wouldn't that be much easier? So argument is following. If you're thinking intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model. The training data is coming from public internet and the label would data. Public internet is very small. Corporate of data compared with world data, majority of world data actually private a locked inside application locked inside enterprise. It will never get shared with anyone else because this company's proprietary IP. If you look at the space, then becomes very interesting because majority of data is not being activated to derive any intelligence and that's where we believe in is to activate that data and we believe the future of the frontier of the intelligence actually private intelligence are specializing intelligence. So that's kind of where far was from the beginning we had been focusing on driving the value. - I have so many questions to ask you. I totally understand you in terms of the values in private data within some of these largest companies. Is that not the premise of what anthropics enterprise businesses though with Clorco work and with a lot of their adjacencies that they're building would dare or not say that that's exactly what we're going after. - That's interesting because I view and topic as a company fully believing AGI. The definition of AGI is there's this one model that can solve all the problem in the best way. That to me that's the definition of AGI. To me that means you do not need to specialize and that one model should be able to solve all the problems is so intelligent so much knowledge of every parts of the businesses, every part of the jobs it can fulfill then why do you do bother specialize. So that itself is a validation that we're living a world that's not ruled by one principle. We are living a fully diversified world give you one example, right? Different region will have different value systems will have different policies will have different way of conducting business will have different lifestyle. It's all taste choices judgment combined. I think that's what define us as human. We are not robots. If our future world is going to be ruled by one standard a taste dictate by one company we turn ourselves into an army of robots. And that's very depressing to me. I think that's the way
What's separate out for Persepian for other species is the creativity. Is the deep desire of pursuing new things, of discovering new ways of living, that defines us as human beings. That part cannot be copied. That's my fundamental belief. That's why, in Silicon Valley, there's so much creativity across the world, there's so much creativity of building new businesses. I had this interesting conversation with Jensen after he's GTC keynote. We actually recorded it. I watched it. It was great. Yeah. Recording with Jensen is not really recording. He just started having a conversation with me. I didn't know his crew already started recording. We just keep talking. It's so easy. We talk about this specializing intelligence. He said one thing to me. "Lin, you're right. There's no specialized general company." As in, every company is built on a special belief of doing things. Otherwise, there's no reason they should exist. It feels yes logical, but then I start to think back about what he said is profound. Because every single company is doing something unique. That justify their existence. And this something unique is deeply baked into their product design. It's deeply baked into their software design and system building. That's deeply baked into the data and the interaction with their user and their deep understanding of their user intent, interacting with their product and engagement and so on. All of that is a fundamental base of why a company should exist. That is not learnable or shared by another company sitting outside. When you help me understand it, as a podcast, I specialize in asking basic questions. So forgive me. But why then do people like Dario, like Sam, like Larry and Sergey talk about AGI and the way that they do as an inevitable? I think what they build is fantastic because they are basically building power line to distribute a really great source of intelligence that everyone else can build on top of. That's how I view their contribution. If we don't have this fundamental infrastructure, then we will not have all kinds of appliances living in our home. I love my coffee machine. And it's a special rendered, right? So, but without that power, then we don't get to do the things that are fun. That's unique. That's a special that ingrained my, I will encode our taste. So I do think that's very, very important. But the question is, is this power line going to replace everything we do? I don't think so. The question for me as an investor is, are power lines good businesses? You said about pie torch and open and the open ecosystem, open source in the last, I would say three months. We've all realized is actually accelerating so fast and the capabilities have increased to such an extent that it's not comparable quite, but it's getting 90% as efficient with 15 times to trimass statement, more cost effective. Our power lines good businesses in a world of open source. So here's how I view open source. So early on when we found the company, we have pretty deep debate among the co-founders. What do we do? Do we build our own models or we build on top of our open models? At that time, open model was not almost like at its infancy. It's a big bat. If we're going to take that direction, it's a huge bat that it's going to do well. But with our pie torch experience, we believe in the open community. We believe in openness. That's a fundamental principle we operate with because openness gave control to the user. Think about open models, right? Once the model is released, you have the full control of the weights. You can change your however you want. It's yours. And then you can build on top of it. So that is a fundamental, different operating principle that we believe in because of a roots in open source before. So we took that bat and it did pay off in the sense that both open model and closed model, the quality significantly increased, improved over the past two years. To the point, both of it, both of these two streams cross a quality threshold, it can solve so many problems. So within, within far away, obviously, we do our own product. We use open model to drive our recruiting process, candidate sourcing and feedback collection. We use open model to even drive some internal finance processes. Obviously, for coding, we use open models to help us debug. We have a ton of agents within Firefox ourselves and we are cost conscious. So that is important. Both model categories across the thresholds solve so many variety of problems. Second is open model across the threshold is so much easy to tune. So being able to steer a model is intelligence. It's part of the model intelligence. And the model intelligence has passed through to is much easier to steer, especially with small amount of data. A small amount of unique data a particular company has and then we can heal climate towards your evil. And often time, the end result of heal climbing is to solve your unique problem with your data, you are better than a general purpose model. 90% of enterprise workflows can be done as you said that the incredible array of functions that you now use open source for with open models. So the usage for frontier models will not be as large as it was if it was needed for everything. So are these companies actually dramatically overvalued and overestimated if the majority can just go through open? I think people start to realize it. I remember following two years ago, I went to different places and talked about an interesting phenomenon. It doesn't exist in the past in the SaaS era. During SaaS time, product market fit and the durable business almost are equivalent to each other. The hardest thing is find power market fit and then you will find it, it just scale as fast as you can. Because CPU is a commodity, the infrastructure built on top is almost like a commodity. You don't even worry about that as your cocks. Now, product market fit and durable business are two separate concepts. For startups, you know, great company that have power market fit, customer want to pay them and they really value their product, but they cannot scale because once they scale, they could scale into bankruptcy. Have you heard about scaling to bankruptcy? So that's real problem. It's even a bigger problem for incumbents. So the big companies for digital native, because they have the traffic. They have huge amount of traffic. They're the winner from a decade ago when there were startups. And they have so much traffic. Once they draw all those AI features, they're going to reach to all their customer base and they cannot afford to do it. Because they're still for a look at their cost proposal, cost forecasting is kind of just no way. You can justify this, right? So then it becomes a real problem to all those innovators. Hey, we really want to plug in to this new technology, new disruptive technology, but we cannot afford it. And we need to find an alternative to be able to afford it. And the alternative is to have the control over your open weights model and roll out your own model. It's all you see what Sam Orton was releasing last few days, which is just dramatically lower cost models. I can't remember the amount today, but I think it's like half as expensive or maybe three times cheaper. Is the next step actually we just see a massive reduction in price from the frontier models. It could be, but I think at the same time, it's just a very different operating principle because for open weights model, because it's just there, basically a model acquisition has no cost. No, obviously some company trained those models and willing to open it up. I know within US, there are multiple companies doing that, including Nvidia, streaming, Nimo Chow, we're working obviously very closely with them. So once the model is there, whoever using those model, there's literally no cost, but there's fundamental cost for the frontier labs to investing those models and recouped R&D cost back. So in second is you just cannot customize those general purpose models and you use it as is on top of API. You have no control over versus with open model, you have full control. You can tune however you want, you can use it however you want, especially fireworks, we are special specializing intelligence platform. We offer all sorts of tools for you to easily customize the model for one specific use case and after that model is tuned with high quality and then we further optimize for inference deployment. Think about fireworks, we think about every single model deployment as one size fits one. It's unique for your workload only. It's optimized for your workload only from quality speed cost point view. So we believe that's that's absolutely needed because once you think about a production scale of reaching to millions of users, tens of millions, billions of users, then even 5% of cost reduction means a lot. It's a massive amount. Let alone what we have seen in the past is five times to ten times cost reduction. The one question that I do have to ask is the concern that enterprises have is national security concerns. When you look at open Russia, I think the top six models today are Chinese models and their incredible quality, the speed of development is incredible, but they are Chinese models. Do we have serious national security concerns when analyzing the power of Chinese open source? I think it's a huge debate happening right now across the industry. Once the model is open, you'll see that.
You can put all kind of guardrails specialized to your business around. I would say to all models, it doesn't matter if you open or close, you should put your own guardrails around it. The fundamental reason is the following, a model provider will infuse their own judgment, their own taste into the model training process. You cannot guarantee it matches yours. Remember, it goes back to Jensen's comment, "There's no specialized general company, every company is special. Every company will have a special design principle. Every company will have a special taste. Every company will have a special target audience to serve." Because of that specialty, it's guaranteed that the judgment, the taste, the design principle from one company will mismatch, will misalign with your company, which is a special problem. That is a reason you need to tune those models to match yours. I really believe the future will not be a few small numbers of AGM models. It may be scary, but I think that's true. It will be millions of specialized models, one propagation per use case. We saw in the last week actually reports that China were looking at actually restricting access to their open models because they were seeing the development being so fast and so good. What would happen in a world where China actually started restricting access to their open models given the lack of open models we have in the U.S.? I think it will be a big impact in the short term. But the beauty of open ecosystem is it's not one provider. That's why it's open. Usually I track many, many, many interested parties to participate. I do believe in terms of talent density and the resources. I do believe U.S. will be able to build that open system by ourselves and we should. I have seen this happen again in many open systems. There are thousands of lower bars and that's the beauty of that. When we talk about the specialization of intelligence within enterprise as you have done just there, if we take a very prime example, which I don't particularly want to take because I'm an investor in LaGoura and I think I know which side you're going to fall on here, but you have two companies that compete in the legal space. Harvey and LaGoura and Harvey have committed to building their own model and then LaGoura have not. A year ago, it looked like companies that didn't commit to their own model were right because front-end models were increasing so fast in terms of capability. Now it looks like they're wrong. Should companies like Harvey and LaGoura be building their own model and actually if you don't, what happens? So here's one observation I had is software development, especially SaaS. Space has been significantly disrupted because of the general intelligence of coding. The application development lifecycle has significant claps in terms of the timeline and the resource needed. In the past, it requires tens or very strong product engineers and PMs to convert for ideal to implementation to production scale, multiple quarters of years of investment. That's a deep mode. And today, one person, a few weeks, can possibly launch their ideas into product and scale quickly. This is unprecedented and that's also create interesting dynamics in redefine where the competition is because it's really hard just to compete on the idea of the application by itself because many people have similar ideas now in competition is no longer such a big barrier. Is that actually true though? When you're looking at enterprise deployment and enterprise rollout, if you're working with some of the biggest law firms in the world, I mean the enterprise sales cycle is at least multi-year with relationship build that's very tough and then you have deployment that's very customized. It's not like a laven labs where you pick it up and go. It's different. And also, I think legal space is particularly challenging because lawyers are usually more conservative. Legal is also not tolerant at all on arrows, right? Because that's why lawyer got paid. It's going to be the very like rock solid case. If something hallucinates and generate wrong judgment, then you're in trouble. So I do think the legal space is a very interesting space to penetrate and these companies are both doing great job. But on the flip side, I do think both companies are owning proprietary knowledge and information, how to build those assistants to the case studies, to go deep in driving legal research and all this. I might understand your legal so shallow, but there's so many different versions of flavors of cases. So I do think they are in unique position to convert that deep understanding and they all have data. It's not just about how defensive their business is. It's about, hey, oftentimes when they build those assistants, there's a harness integrating and deciding orchestrating which AI tool to use, which tools calling to, and this is Facebook, this is customized and accuracy of calling those tools and calling to what kind of tools is important. Even that harness need to be co-trained with a model powering it. There's just kind of ample examples of driving that business to excellence by owning their own intelligence of how to do that in the workflow there. So maybe it's a timing, coding for example, I think in coding space, cursor is probably one of the pioneer starting to tune their model and now almost all coding companies tune their own models. The standard pace of model development slow down because every single day it seems like we have a new model with a new capability and it's like, oh my gosh, class is newest model is amazing. Next, we have, Mr. Arznews model is amazing, a Gemini's newest model is amazing. In three years time, will the pace of model development still be so fast and model superiority be so transient by one day it's one and the next day it's another. So there are a few layers of model advancement. There's base, general IQ advancement. So those will take step functions. So that's why when they release there, there's always major release or minor releases, right? The major release of step functions. As you remember beginning of last year, there's whole this thinking, the thinking process is new. But the model just don't spit out in serimidially, the model will think by self in speed all in serim is much better that way. So that's one step function. There are many step function we have seen through, but I see those as every year or every three quarters as a major leap. But at the same time, build on top of those, the best base models and I can see the specialization start to accelerate because as I said, it's really like a tree, right? There's so many branches and leaves that can possibly hang out on the trunk. And as the base model quality start to have step function, leaps, and there's so much more we can do to specialize. So I do see in the world specialization is going to accelerate much faster than the general intelligence part. When we think about the general intelligence part just before we move kind of further into the stack of like multi model, some profit, the 5% gifting of open AI and others to the administration. Do you think we've reached this age where model development is so advanced and so important to society that they will impart the government or administration owned? That's very interesting question. I think there were presidents of that. If we think about the foundation tier of those general intelligence model as fundamentally a basic infrastructure for the big, big economy to operate around, there has been presidents of like PG and E, owns electricity and gas and so on. So what I don't want to see is there's only one company owns intelligence. I think that doesn't make sense to me because they're different. I said that they're different flavors of intelligence. There's this general common intelligence that benefits everyone. And then there's a specialized intelligence that actually help us advance in history to think differently, to create new paradigm of living or new paradigm of doing business and shaping the industry. I don't want that to die because there's only one company can do that. I don't think that makes sense. With the many models blooming theory, does the idea that you will route tasks to different models dependent on what they specialize in? I think so. If that in mind, will you not build your own open rupture of the world to cater to that? Yes. You can argue they're the best people there because they deeply understand their use case and they have the evils. So again, my thinking of what is a frontier is not just this one model. The frontier could be your special routing mechanism for your business and you decompose that based on hitting orders to fulfill this task and you need a highly intelligent layer, maybe the most expensive close models to judge at the highest complexity. And usually people will also be sub agents solve smaller problems than those can go to smaller open models and those can also further be customized to fitting into your special design. So I've seen a lot of people already doing that today. And we also think there's a space to build an automatic routing system that can learn by itself and that compound with automatic tuning system eventually, with things you should all be automated. And then you can see a self-evolving system based on what flows through your product and your product keeps evolving. Your product is live, right? So you keep deploying and launching new features and to interact with your users. just kind of it will be um. totally self-evolving automated system. - Do you think then that rooting layer of the stack is valuable? If it can be automated or it can be built on its own, is that a valuable layer to have? - I definitely think so. - You do think so? - I do think so. - If it can be automated or companies can build it themselves, why would you need a request or an open router? - You probably don't. - We're not there yet, but I do think this could be area of innovation. - You said cursor being the front runners in terms of how innovative they've been. I completely agree with you. But I heard, and I really stalk you before shows, but I heard that CTO DEMA was embedded at cursor for months building the RL infrastructure. Is that how it has to be done and is that scalable? - So what's happening is usually in the early adoption curve of new technology, the early adopters are hackers. The hacker is not in a bad way, it doesn't have negative connotation. They have deep expertise in certain area, and they want to control all the things. Versus in the late stage of a new tech adoption curve it starts to get more accessible to a much bigger cohort user, doesn't have deep expertise, and they need less control. So it always goes into deep control first, usually, and the little control later. So we definitely are aiming towards the later stage as the ultimate 10 want to target, but it's also extremely valuable to understand what is required to get there. So that's why we partner deeply with cursor, they are the pioneer in trying those ideas. They do have researchers from Frontier Labs, and they want to control every single thing. And at the same time, we're also pushing to the boundary, we're doing things that never existed before, we've been system never existed before, because we push the boundary that is unique to this particular setting, okay, what is unique is here. Typically, if you think about training, training happens, training is very kept to intense, and usually happens in big companies. They have a lot of money, they put those money to buy very expensive training cluster interconnected with each other, super expensive. And then once you have those expensive large fleet, usually you don't need to think too deeply, how to be efficient, you just focus on doing your work. Curcer is like us, their startup, both of us are very kept to conscious. We want to be efficient while we don't want to slow down the research innovation. So together we figure out a very smart way to drive their training process is they do massive post training, which is reinforcement-learning-based. And the reinforcement learning we break that into P2 pieces, one is the trainer that is taking the weights of the model, and the basic general new model version constantly. And that new model will deploy to, we call the RL rollout. It basically is a, deploy that new version, interact with a synthetic environment, a synthetic like coding environment or real coding environment, and then get the reward back to judge if that model is, version is good or bad, right? So that's a good rough process. And we decouple these two. In the past, in light hyper-scaler, they run that altogether. If you think about you get 10,000, 100,000 chips, or interconnected together to infinite band, it's extremely expensive, it's really hard to find. But then to go really quickly, and we designed fully distributed system, we run across five, six data center regions globally and tapping to scatter GPUs, and they are able to run massive jobs, our jobs. But the challenge there is we need to sync model weights across all these different regions. And then we can, how hard can that be? It matters because the latency of delay of sending these weights over is gonna dictate how fresh the rewards are. And then if it's too stale, then you are too off. So it's a balance. But we innovate a way we can distribute fresh model weights quickly. It's not too off. So numerically, it's still sound. While we are not limited by our very expensive deployment of GPU fleet. So those are the innovation we work together with Cursor to push the boundary and leading to their recent model. Longchitz were very proud of them. Incredible customer to have amazing progress they've had with you. It's a very large customer for you. How do you think about the concern of a Cursor churn in the wake of a space acquisition? Yeah, everyone's concerned. The whole industry in terms of application innovation is by model. In the sense, there are few companies are very successful. They escape velocity, but few of them. So that's the shape of the whole industry. And last year Cursor is one of the few. I would say all model companies are concentrated on Cursor. We concentrated on same group of app companies. And since then, it does change. We do have a very healthy diversified customer base. Especially, I think last year is the year of coding. I think all major coding companies are ours. And co-work is much more diversified by itself than coding. Because their general purpose co-work, for example, general purpose co-work to help you do all kind of research. You want to ask, hey, what will be the MVP that GP will prize two years later? So those are deep research, general purpose deep research. Or there are so many different categories of special purpose co-work. Ligo, we just talk about two legal companies, finance, customer support, recruiting, sales, marketing, healthcare. So there's very broad set of co-work space of innovation app company. They are doing really well. And we have them as our customer base. And then more interestingly, we start to see an uptake of consumer facing company are all start to looking to genie technology. They are changing how they are thinking about their traditional business of doing recommendation, for example. And that's very interesting to me because we have obviously well worked that huge recommendation system in the world matter. And we are very eager to see how that transform into a new economy for us. I'm sorry for being naive here. Do people work with just one provider in the inference space like you? Or do they work with you and with together or anyone else in the space? I think people are more intuned to multi vendor strategy in this space because they don't know what's happening if you're safe to have multiple providers to balance things out. But we don't view ourselves as an inference provider. Again, we view ourselves as delivering these specializing intelligence where we help companies to their model, give you some numbers. Today we process more than 40 trillion tokens a day. So majority of those tokens are coming from a customized model, not from off the shelf models, a coming from customized model. What will that token count be end of next year? Anywhere ranging from 20 to 100 X could be possible. 20 to 100 X. Yeah, we're at a very early stage of S curve of explosion right now. 20 to 100 X. If it's 20 to 100 X, the idea that we are in a catapag's bubble is ridiculous and we are desperately needing far more catapags than we are ever suggesting for compute. Is that right? That is right. At the same time, I think Jensen has a five-layered AI cake from top-down application model infrastructure chips energy. We are bottlenecked by the lower part of the AI cake in terms of supply chain. Being energy. Being energy, being chips in the physical world, how fast we can manufacture. Because in the history, all these industries are not designed for massive scaling. Speaking about the 100 X scaling, no one was designed for that. I talked with many manufacturers. It's kind of we're bottlenecked by small parts. Transistor. The smallest tiny parts that hold off the whole manufacturer line of servers that can deploy to the center and be used to generate tokens. Do you have to be full again to Jensen's five-layered AI cake? Do you have to then be full stack to win or to reduce dependencies? We've seen open AI come up with jalapeno, terrible name, anthropic talking Samsung about building their own chips, deep-seeker building their own chips, Zuck came out with meta building their own chips. You have to be all of it. It really depends on the company philosophy. To us, agility is everything. And we need to earn the rise of building anything. So focus is everything for us. And we want to focus on where we add the biggest amount of value based on our strength. We would like to leverage other people's strength to build on top of. So in particular, we want to run everywhere, all possible AI chips in the world. We don't want to limited by how much chips we can bring into our data center where they were constructed or rented. But over time, when the business grows very big, right? So I still remember when matter was young, they don't build everything. And when they're big, they make sense to build. You earn the rise to build for your own giant traffic. And if it's saved like five times more cost, then you should go do it. But I think at the early stage, that's why I give tell you an interesting story in the coding space. I would say cursor is the first company they have decided to work with us early on. I remember when they worked with us, they were single digit million dollars. Very small. It's only two years ago. They go by 100,000 x.
I'll be right over to you. But they decided to work with us early on because they recognize they only want to focus on product innovation and later on research. They do not want to focus on this platform innovation. They know we are putting all our earned in there and they want to find the best partner to win big. I do think that's the right mentality to specialize and we want to specialize. We do not want to own the whole intact stack. That's not our goal as a company. I'm sorry to be hopping on that. Why does Jensen skip your layer of the cake? Because he's doing neemotron with models. Why does he not want to cannibalize your business too? Jensen is not building a cloud either. You can say, hey, Jensen probably have all the rights to build an immediate cloud. So he's not building a cloud infrastructure. I think he mentioned that as well. I mean, who asked him that question? And he also mentioned he wants to specialize in what they have the rights to do. Why models? I think it's pure supply chain question. Is if US doesn't have a US native open model, it's a problem. It's a supply chain problem. So he is solely there to solve the supply chain problem. But if there's no supply chain problem because the company, like us, are providing this specialized intelligence platform layer, then he doesn't to worry about it. So he just want to make sure the whole entire five layers of a cake is flowing. There's no blockage. And if there's a blockage, he's interested in solving those problems. Mark Beningoff, one of your investors, I think, in the new round, which obviously this will come out after round is announced, said that he spends about 3.8% of developer salaries it sells for on anthropic and core code. And I think it's a useful analogy, because if you assume that that is what spent on core code encoding tools, that says one side of the market. But if it's 20%, wow, we're underestimating how big these companies can be. When you think forward a year or two, what percent of developer salaries do you think will spend? Is it less because these tools will get cheaper? Or is it more because they'll get better and better? Because it hasn't so far. It hasn't so far because of supply chain constraint. But we are living in a free economy. So think about whenever this shortage prices high, prices high, high prices will invite a lot of people coming to solve the problem. And in we invite competition, competition will bring down the cost. And then eventually we're leading to a very economical solution. But actually, that's good for everyone, because much more affordable infrastructure will invite more usage. So my prediction is with the decrease of the infrastructure, that's where it comes to my prediction of how, how far next year will look like. Because the infrastructure costs will go down, and usage will explode because of that. So the moment you don't think about that as a problem for you, and if it's a utility, you just use it. How much will token costs come down? Is this just how it means that is it like a halving? Is it like, oh, it'll be 100th of the cost? So there's different ways to think about it. Not all tokens are equal. I think we should establish a best practice to evaluate the token economy per task. Because different models have different way of speed out token. Some are much more verbose than the other. So you can imagine one model is too extra than the other, but it's too much verbose to solve the same task. And then they're the same cost. But overall, I think as a model quality improve, I think being precise is going to be part of the optimization. And so that's one level of optimization is to solve one task we should need less tokens. And the second is for one token, and how to do that is you need to customize the model to solve your problem, especially better and more precise that goes into model tuning. And second is for each token speed out from those models and process by those models, we also specialize in making the unit of the economy much better through our platform. And third is underlying infrastructure like the GPUs, the surrounding like memories, and all this today is under stark supply chain constraint is going to get much better, situation will get much better. I don't think probably in the next one year or in year half, the situation will not change. But in the long term, 23 years, you should change. And that cost will compress. So overall, I can imagine 10X cost reduction in the next three years. You said there about token efficiency and how you enable your customers to be much more efficient. With that efficiency, you do charge more. You know, when I did the research, when it compared to competitors, I got like, together's price king. And I didn't mean this disparagingly, but like, that cheaper. If you want cheap, you go there, and respectfully if you want better quality product, you go to you, but it is more expensive. Do you think that's a fair assessment and a fair analogy? I think we're probably not comparing apples to apples in the sense that again, goes back to our business. Majority of our traffic is customized model. We optimize for quality. Number one, always quality. Quality as in model quality towards your applications, your specific business, your use case, and so on. The second is when we deliver those model in inference, it's also quality. And we care quality so much we do extreme things. For example, during training time, this is a very hard thing to achieve. It's called zero KOD. It's a little bit technical. The idea here is-- Zero KOD. KOD. KOD is a measure of quality. And what it means is between the training system and the inference system, when model move over, we have bit equivalents. So as in the numerics are fully the same. We do not lose a bit of accuracy. That's really hard to achieve. But the reason we push that-- well, we deliver that. And the reason we push that is because we know primary business is in model customization and inference of customized model. And we want our customers every single dollar investing training, maximize it. And then if cross-training inference boundary is not bitwise equivalent, they just drop the quality down. And then it's like you pay your training investment by discounted quality while you do that. So quality first and quality does bring additional value. And that's why we are not interested in commoditized, one size fits all. This is off-the-shelf model deployed in the same way for everyone that kind of business. We're always customized model deploying a unique way for your particular workload. Two questions. Do you have to have an FDE model to make the customized model efficient? As a matter of fact, we do have a FDE team. It's called applied machine learning engineering team. So their primary job is to accelerate this customized deployment. As a matter of fact, it will also be with the agent to automate a lot of deployments. So-- Given where we are in the stack, a lot of the complexity that we have, we have a margin structure that's a little bit different, like traditional size, being 80%, I don't know, the margins are precisely here, but the addition is in the 30% to 40% range for where we are. Is that the new normal for where we are? I don't think that's new normal. That is a reflection-- at least for us, I don't know other companies. For us, it is a reflection of we are in hyper growth space. During hyper growth space, you have the choice, right? You either optimize-- to me, margin optimization is a constraint problem. As in, hey, we want to go to 70% margin. We want to go to 80% margin. And then we are going to go backwards and impose those constraints to guarantee those margin. And usually, constraints slow down innovation. So for-- give me an example, during system development. And in a high velocity system expanding phase, we don't want to over-build, because we're in kind of high experimentation, we're testing what will stay, will not stay. Optimization doesn't make any sense. Once we know this system that we want to build 100%, and we are going to scale this 1,000 times bigger, then we go optimize the heck out of it. I think-- mind now, you will think about businesses the same way. Well, in the hyper growth, if our focus is only optimized growth margin, we absolutely can do that. But we are sacrificing the speed of growth, because we want to go everywhere. We want to go into different geo regions. We want to go into tackle different use cases. We want to constantly create different product lines. And those are not the time for optimization. That's my opinion. So we will be able to increase margin without moving into different layers of the stack. Not into-- we absolutely are not going to move into application layer. Very clear to us. Whether we will move down into view data centers, and so on, that could be always beyond the table, but the question is timing. Isn't the statement you either die, or you live long enough to build your own data centers? As Elon or Zach now have spent, I think, 10 billion on the latest data center in Canada. Would you like to build data centers? So I have built the center set matter. Also, lots of innovation possible there. There's no one side space all as well. And building a GPU native data center is also interesting. So it's a trade-off, right? From Operation Point View, it's much better to build a heterogeneous deployment. It's all the same chips, all the same skew, as big as possible.
and run multiple workloads, so it's fungible. It's very demanding. You build one principle, one process to maintain this operation. Again, it goes to optimization, but once it's so big, then any optimization is going to drive a lot of economic return. For example, we're talking about a video recently acquired company, also called Gwok, with Q. It's a large SRAM-based ASIC accelerator. I spoke to Jonathan before this show. Jonathan is excellent. He said what a fan is, he's a views. Oh, also fan of his. So, but it's a great combination between a Flops, Intense, GPU, and SRAM Intense ASICs, because Flops Intense is really good for first half of LM processing. It's pre-fuel. It's called pre-fuel processing and prompt and so on. And SRAM Intense is really good for generation. That's just the nature of the model architecture. It's great to combine these two instead of running heterogeneously on the same chip. But that requires a very unique system design and deployment into data center. And it is heterogeneous actually before I really mean homogenous design is much better for operation. And this is heterogeneous. And then how to operate this heterogeneous design requires unique innovation in data center deployment. And so on. So data centers aren't commoditized. You can specialize in data center deployment and one data center is better than another. And data center deployment can be done well and badly. Data center is so complicated, right? If you think about the beginning, all the way from construction to power deployment, you have the right power to coming, right? Fiber channel, the right cooling, especially, new chips requires liquid cooling. To get all this right and the parts can fall apart and how to replace them, it is all very deep expertise. It's no joke. It's not tomorrow I can be a data center operator I cannot. Is that not where you would bet long on China with the greatest of respects? Especially in the US, one of the biggest barriers to data center deployment is policy and is local legal infrastructure that prevents it. In China, you don't have any of that. Data center deployment is much, much faster. I think in general, infrastructure, the base physical infrastructure in China is going really fast. I literally see some crossover bridge being built within a week. The velocity is very high there. And there's a highway close to my home. After one year, it's not done yet. So it's also a crossover. So I do think there's a unique strength, probably, because of the population density and they are specializing those kinds of construction. But I do think we also have those specialty. People is just, even, I heard, even electrician is under severe shortage. We are under global supply chain constraint here. What change would moving into the data center that I caused to margins? Would that take it from 30 to 50? Would it be not that meaningful? Like, what would that change due to margins? How we calculate ghost margin is interesting these days, because how long does hardware depreciate has significantly changed? Yeah. In the past, it's six years, solid six years. And hardware releases usually three years, that's fast. And now, within a year, from one vendor along, we have three skews. The newer model usually runs the best on the newest hardware. Model depreciation is also very fast. Every week, we launch a new model. And then the model is kind of peak in its value before the next model comes out. And the new model likes the newest hardware. And imagine these cadence after two years, which model runs on the two years old hardware? It will be two year old model. Those models do available. So I think that's kind of the real dynamics we are facing right now. The hardware will last for six years, too. But what you're saying, the speed of model development, fire outstrips the speed of chip and hardware? The speed of the model definitely is the fastest. But even the hardware innovation itself is the fastest. So after three years, if every year there's three hardware skews, after three years, there are nine hardware skews in between. Do you still want to go back to nine generation older hardware running three years old model on that? That's questionable. Maybe there's a war. We still, it's still valuable. But with these pays of innovation, it's questionable. Now, with a different depreciation cycle, it changes the dynamics of beaut versus own, beaut versus buy. And again, it goes back to my origin thesis of do you optimize for growth or do you optimize for growth margin? It's all about timing. How do you think about that question? When you're sitting there in an armchair on a Sunday afternoon thinking, "Hmm, we're optimizing for growth now." When is that time to optimize for growth margin? Why would we optimize? We want to optimize for both. So here's our high think about it. Optimize for growth requires a lot of business planning, assuming there's a part of market fit. Optimize for growth margin is optimized for differentiation. I think I want to avoid over optimizing for growth margin, but we should optimize for growth margin continuously. As in, we should optimize for product differentiation continuously. There's no question about it. I think we want to continue to optimize towards a healthy growth margin, which allow us to grow really fast. And it's a trade-off, and we don't take compromises. The compromise as in, we over optimize growth margin to resulting in very slow growth. One possible way to optimize growth margin, we do not grow at all. We just often have a heck of it. I know we can give a climb to a high number, but that's absolute disaster outcome. Okay, interesting. If we just said, hey, so growth margin, we're going to take it from 30% to 10%. Is it a winner-take-who-market where we could eat up everyone else's lunch and then optimize growth margin later? Winner is all probably not a snapshot in time. It's going to be a long-term situation. We do see a particular industry will oscillate and start to settle with a few good ones. Take a look at, for example, I was on a dinner table, and interesting seems like there were a lot of those companies around two years ago, but now it's pretty much two. I think it's a long game. How do you see the more mature state of your market? Is it like a cloud market where you have, obviously, as you're AWS GCP, or is it an Uber and a Lyft, where one takes 90% and the others fight for scraps? We're in the adoption curve, where a lot more companies, they are in the airspace, start to seriously think about moving to specializing intelligence. To start to seriously think about owning their intelligence is better than renting. Because going back to this optimization, when is the good timing, right? So it's the same question we're answering for ourselves when people would versus buy. And our customers also think about being versus buy, or being versus rent, or owning versus rent, right? I think AI, or AI adoption journey has gone further along into a lot of companies as meaningful traffic. A lot of companies is deploying AI into production. A lot of companies is at the face of scaling, and that's where optimization kicks in. When optimization kicks in, you need to have control to optimize. If you don't have control, you just don't have the range to optimize. And for you to have the control, then you have to build on top of some open model. You have to kind of turn your data into your intelligence. That's pretty much the part that we've seen so many companies crossing the street, the Richardson conclusion that are moving towards distraction. Being of owning your intelligence versus renting it, that does apply to like a national layer. And when we've seen, you know, like Fable be banned, in some case, by the administration briefly for 19 days, especially in Europe, we suddenly went, "Oh my gosh, we cannot be at the hands of OpenAI and Anthropic, where we can just be banned in our health services," said on the infrastructure of something that an administration can turn off. Do we see a future of sovereign models where large nations or nation blocks own sovereign models? I definitely see that possibility. I also see, if we think about the general intelligence model as the electricity there, as a power line, every country should own their own power line. Right? So, I think that is a very scary moment. Is my power line is going to be cut off and all my fundamental day to day is going to not working because I feel so frustrated. Whenever there's power outage in my home alone, I feel so frustrated when I cannot access my wifi. I feel so anxious. So, I mean, obviously, operating the country is extremely important beyond top of this fundamental baseline. And for every single company, the same thing, is not just whether a country should have their unique sovereign independence, but every single company should have their independence. You don't want any single person to cut off. That's extremely scary moment. Why would you move into the data center space but you wouldn't move into the chip space? Because I know the beauty chip is extremely hard. I thought so too. But then, how come everyone is seemingly doing it as if it's like just another product? As I said, open AI and thropic deep-seek matter, we're building our own chips now. I think Matt has been building their chips for more than five years, more way more than five years. And MTIA has been project, sings, when 18, maybe earlier. Because Matt has been investing AI for a long time, pre-GNI. And they have a huge AI workload to focus on ranking recommendation. And Matt has been building other hardware as well in the past. So whenever the--
The usage has passed a certain threshold. It makes economic sense for you to build an underlying supply. And then you can specialize towards your workload. And that's another form of specialization. It's specialized to bake your logic into hardware and it's hardware's purpose-beautiful, your particular workload. And you better make sure this workload doesn't change. Because it's really hard, once the hardware is taped out, it's really hard to go back and change it. It's possible, it's very costly. So once your workload stabilized, once your business stabilized, it doesn't change too often, then that's the time to consider building a chip. I still see the whole AI world, especially models, customization is very dynamic. Workable patterns, very dynamic. So think about how much energy in the application space people are experimenting, all kinds of things. You don't know which one is gonna take off. And they will just take off quickly. And once they take off, I wish I was gonna sustain. And a few ones will sustain, then that's the time. Oh, now we know this is a pattern. And now we should probably encode this pattern into hardware and bring this hardware into a data center. And so it's all cascading. And then it's gonna cascading down to me. It's a funnel question, where are we in the stage of funnel? A maturity funnel, I mean. We're still in the early stage of workload maturity funnel to warrant a chip that will be durable. So now you go back to all the, we have so many accelerators, they are successful, some are really successful. But remember those AC company, they started before GenieI. They started on some seasons to optimize some workload. And they appear to AI and trying to kind of fitting AI work loads almost like you bat before this AI workload emerges and now it becomes a serendipity question. Are you lucky enough this just worked, right? And some really worked. Some fundamental design of putting a lot of SRAM on the chip is great for AI model because they are memory hungry. And this really accelerate the execution of inference and so on. So those works and some don't work. What do you see as the greatest bottleneck today? You know what I think it was when I had Jonathan from Groc on the show who said like HBM was the greatest bottleneck and that's why you've seen the 5x increase in price. What do you see as the greatest bottleneck that people don't talk about enough? I still think we don't have a great system for a very large model. I really believe the fundamental low-level infrastructure cost will go down. For solving tasks we should need less token. That will increase. So collectively the cost will significantly reduce. Therefore we can run the highest intelligence model much more ubiquitously in the future. But we don't have a system designing for that. For example, we don't have a great system designed for 10 trillion prime models today. That will require very smart engineering code design from the model to the customization serving platform layer all the way to chip layer. The chip is not the individual chip of the system. Collection of chips in system and all as a total package. I think there's still a lot of innovation we can do. I think recently you announced you were at 800 million in there are incredible feet and scaled so fast. What is that at the end of this year? We think we can at least double. By the end of the year. Wow. It's so interesting for me as a venture investor. I've been investing for 10 years. We used to be in the day where Slack was the golden child while at one to 10 million in revenue in 18 months was like amazing. And now we have companies like fireworks. We used to get to 800 million in revenue in a matter of years. And you mentioned cursor scaling to billions in revenue in a matter of years. The speed of company revenue growth is just unparalleled. I think it's because there's a fundamental disruption in this technology that is all empowering. All empowering the sense it reach out to every individual one of us to be creative. And it unleashes a lot of creativity that we just don't have access to. And that's why I was seeing this phenomenon of extreme fast growth because of a demand. Final one before we do a quick fire. You hired George here who was president of Salesforce. He's exceptional. He's one of the most direct, no BS operators I've ever met. But you met him a couple of years for a year before and you were like, oh, we're not ready for you yet. Why did you say that? And why did you decide now was the time? Sorry, a year ago, I think we're probably just 50 people. So today we're 200 people. We're still not that big. Wow, you're four million ahead. At 50 people, I'm more thinking about skating the product first than skating massive scale the business. And we talked and I have huge respect to him. I know his legendary, his legendary operator in Silicon Valley. I just feel like we're too small for him. And I told him that we're probably too small for you. But I would like to work with you at some capacity. So he helped me actually build out the team, interview a lot of executives. His feedback is always well balanced, worth thoughtful. And we start to work together in that capacity until I think in the last year, we're growing really fast. And he knows. And we start talking seriously. And that only relationship paid off. He's really cool in the sense that he-- He's so cool. He did a lot of things. Great accomplishment in the past. I find a unique character about him is his extreme experience has high attitude of business vision. But he is also very curious. It doesn't make assumptions. I know it all. I've seen all the movies. It's the same movie. And let me just kind of direct this movie as I did in the past. So he didn't come with that attitude. He knows AI goes in insanely fast pace, learning a lot of way, but also fully embraced AI. Actually, his team, our GTM team, is using all kind of AI agents. They're sharing skills. So they maximize their productivity. And he knows we have a super linear demand curve. There is just certain pace we can build our GTM team. In order for us to catch this curve, we need to build our team, but the team need to also have increasing productivity to match. So that's the problem he's solving. And I feel very fortunate to work with him. In general, I feel in the AI space, the unique part is people need to have very special traits. Almost like contradictory characteristics. For example, very experienced, but super curious and the fast learning curve. Or Dima, we talked a little bit earlier. He is brilliant, high intellectual horsepower, but extremely humble. It's weird combination. And he's almost like cynical in Eastern European way, but also at the same time, very humble. - Can I do a quick fire round? - Okay, let's do it. Okay, what have you changed your mind on most in the last 12 months? I think how fast we grow, I change my mind, because I have been quite worried about too big a team to early. So that's why when I met Georgia, I told him we're too small for you because I don't intend to grow very fast in terms of people. I worry about getting slow down and lose the agility and the velocity very deeply. But since then, we have been very aggressively using air tools. We have developed our own unique way of hiring certain type of people that we know they will be charging forward with high velocity as extreme sense of ownership, very communicative, and never take no as an answer. So we also learn how to get those people and now I feel much more comfortable getting really fast. - What's your type of people? And then that sounds weird, but our type of people is actually really specific, pretty much only high immigrants, British people don't work very hard. So I'm very scientific and rigorous, use data for most things, actually think creativity often comes from data and is informed by data, and unwavoringly accountable on ownership. Like nothing is anyone else's fault, it's all my fault, even if it's someone else's fault. That's a 20 VC person. What would you say, you also say? - In weird way, it's not competence. We want people with the high competence, but more importantly, the strong indicator whether they will do well in this way, especially in fireworks is whether they are really built for taking extreme ownership. Extreme ownership, as in, we are not putting people into any, any boxes, and we're just stacking the box together into a tower. We, people just automatically claim, hey, this is an end-to-end problem, I'm gonna see through the whole thing, and work with a bunch of people to make it happen, and I'm gonna deliver it no matter what. So those kind of people has the highest, longest mileage, and their growth curve is amazing also. - What's your biggest lesson from working with Jensen Huang on what makes him so special? - He's everywhere. - I seriously think he has a clone of hundreds of Jensen. Somehow, for some by sending him an email, you reply in one minute. I just don't understand how he's like constantly in details, but now I operate a company for four years, I understand why he's doing that, is not defines velocity, because what is leadership? Leadership is just judgment. It's not privilege, it's judgment. You basically have the context, and you have the right context to make the right judgment for the team, and especially in a high velocity space, if you do not know what's happening, what works, what doesn't work, what are the gaps, you make the wrong call. In a slow-moving space, you can wait for the cascading information up and down, and make those calls, but in a fast, adoration space, you just cannot wait. Because it's guaranteed there is information loss.
transition, layer after layers, people of people, it always happens. And not knowing what exactly is happening and having the position of make judgment makes bad leadership. He is demonstrate through his own example, even before this crazy AI thing is his operating that way. And before I was admiring him, his sheer amount of vulnerability of doing that, now I understand wisdom behind that, because I also operate that way. I need to know what's happening on the ground to make the judgment for the company. What did you wait on in the fireworks journey that you wish you hadn't waited on? Marketing. We talk about it. So we are a little bit nerdy in this way. At very beginning of our journey, we kind of didn't even discuss it, but we feel proud I was speaking for self at the end, proudest sense. And we want to devote all our effort and focus on building product, working with customer, validate product market fit and go find there. And we didn't spend much time marketing at all. We didn't prioritize educating our customer what's the right direction to think about the trend and the value. But we do think now I do think it's important. Marketing is not about flows. It's more about education. It's more about clarity. And we're working on that. What area of AI is underinvested in stay in your mind? You mentioned like cooling or servers. What areas are underinvested in? It has the sexy part of this is such a innovative creative technology. And the view something on top of it is the focus. But monitoring the R.I. I think the industry started to kind of pay attention to it. But eventually that's what matters. Not how much spend is, how much what is return and what is the cost and what is attribution. In the next couple of years, as AI is getting more and more into production, there will be a lot of focus in getting that clarity and getting that discipline out. The token maxing is just a thing in time, but we're quickly moving to our i-maxing, which is about all about running a business. What large customer do you not have that you would most like to have? So we haven't spent too much time in traditional enterprise segment. I think that's just because we was very small. And now as we build out our company, I do think even without our investing, we have customers like Geico, like Capital One, all these companies. So even without us pursuing enterprise, traditional enterprise, they come to us. But I do think that's a very big market. What has to happen before the end of the year that hasn't happened for you to consider it a good year? I'm confident in our capability of driving the business. And to me, this is a year I want to prove we can scale quickly by keeping the same velocity. And that's very important to me. If we reach that point, we should prove point, and next year I have a lot more confidence, continuous scale, extreme aggressively. I want to make sure we do it right this year. Final one for you. What does no one see about the next three years? The UC very clearly happening or not happening? I really see people who own their everything company own their own intelligence as a must have. It's not optional. That's a trend I'm seeing. Because there's an analogy to software is there's a reason why every company build their own software stack. There's no standardized software you just use of the shelf to solve a problem because every single company is solving a unique problem. And they want to build software because they want to have full control. And obviously they were picking choose. Which part of the stack they want to build themselves? Which part of the stack is common knowledge is no point of building. But every single company own their own software stack. Obviously we're talking about this in the SaaS time, right? So same. I think every time every single company should own their own intelligence. Then it was mad that introduced us first. I've had the joy of getting to know you and obviously George, I can't thank you enough for joining me for coming in person. It is so wonderful to do it in person and you've been fantastic. That's an amazing studio. You ask a lot of interesting questions that I have a lot of fun talking with you. We do a lot of research before, huh? Yes, you did. But before we leave you today, founders face a different set of challenges at every stage of growth. For Sid's Shate, co-founder and CEO of Dematrix, JP Morgan delivered the guidance and expertise to help navigate what came next. He credits JP Morgan's high touch approach with supporting Dematrix as it grew and expanded internationally. Did you know the industry average for booking a business trip is 45 minutes as a massive waste of your team's time, or would Navan your employees can book a trip in just 7 on average. Now the built-in AI approves in policy bookings and blocks the rest automatically. This allows finance teams to stop chasing receipts and skip the month than chaos, and you get this real-time visibility that can save your company up to 15% on your travel budget. Go to navan.com/20Vseatsday to see for yourself, and you'll get a chance to win two business class flights anywhere in continental US. Head over to navan.com/20Vse now. You have the idea, but with most AI tools you hit a wall, the setup, the config, the gap between what you pictured and what you actually ship, or Base 44 is where that wall disappears. You describe it. Apps, websites, AI agents, real working products, built-in minutes using nothing but plain language, and it's all batteries included. The backend, the database, the authentication, the hosting, the heavy lifting is handled, so you just really stay in the flow. That's base44.com.
Podcast Summary
Key Points:
The founder believes intelligence should not be owned by a single company; this year is about co-work, and token costs will drop 10x in three years, driving 100x usage.
Fireworks focuses on the inference layer, specializing in activating private enterprise data through open-source models, which are easier to customize and more cost-effective than general frontier models.
The investor highlights Fireworks’ exceptional team, rapid growth (scaled to 1 billion in error in four years), and potential to become a $500 billion company in a multi-model world.
Open-source models allow enterprises to retain control, customize for specific use cases, and avoid scaling costs that could lead to bankruptcy, unlike reliance on costly general-purpose APIs.
National security concerns about Chinese open-source models are addressed by emphasizing that enterprises can add their own guardrails to any model, ensuring alignment with their unique tastes and principles.
Summary:
In this interview, Fireworks founder Lin Qiao argues that the future of AI lies not in a single, general intelligence owned by one company, but in specialized, private intelligence derived from enterprise data. He contrasts this with the AGI-focused approach of frontier labs like Anthropic and OpenAI, which he views as essential infrastructure—like power lines—but not a replacement for customized solutions. Lin predicts a 10x reduction in token costs over three years, which will drive 100x usage, and believes open-source models are key because they give enterprises full control over customization and cost.
The investor, Harry Stebbins, justifies his $10 million investment after a 15-minute meeting by citing the world-class team, the rapidly growing inference market, and Fireworks’ potential to be a $500 billion company in a world of many specialized models. Lin addresses concerns about Chinese open-source models by stressing that enterprises can implement their own guardrails to align with their unique values and needs. He emphasizes that product-market fit and durable business are now separate concepts, as scaling with expensive general-purpose APIs can lead to bankruptcy, whereas open models offer affordable, customizable alternatives for production-scale use.
FAQs
Fireworks AI focuses on specializing intelligence, particularly by activating private data and enabling customized model deployment for unique enterprise workloads.
Open-source models give users full control over the weights, allowing customization and cost-effective solutions for specific business needs.
He believes in a world of many specialized models rather than one general intelligence, as companies have unique data and design principles.
It's when startups have product-market fit but cannot scale due to high inference costs, leading to financial unsustainability.
He predicts a 10X cost reduction in tokens over three years, driven by competition and the need for affordable AI, which will boost usage 100X.
Fireworks sits between chip providers and model providers, offering a platform to customize and optimize open models for specific workloads.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.