Go back

20VC: How to Build Your Own Data Center & Why Every Startup Should Do It | How ElevenLabs Leapfrogged Us: What I Learned | The AI Talent War: How Your Hiring Process Needs to Change with Cliff Weitzman, Speechify

65m 13s

20VC: How to Build Your Own Data Center & Why Every Startup Should Do It | How ElevenLabs Leapfrogged Us: What I Learned | The AI Talent War: How Your Hiring Process Needs to Change with Cliff Weitzman, Speechify

Cliff Weitzman, founder and CEO of Speechify, joins Harry Stebbings for a wide-ranging discussion on AI infrastructure, competitive strategy, and the future of voice technology. Weitzman details why Speechify invests heavily in purchasing NVIDIA GPUs rather than renting, explaining that the economics favor ownership: renting an H100 for a year costs 1.5x the purchase price, and owning hardware enables co-located memory for large-scale training runs. He also notes that older GPUs remain useful for inference, and NVIDIA's recent underwriting deals with major banks are creating a liquid secondary market. Weitzman candidly reflects on his biggest strategic mistake: dismissing the B2B API market as commoditizable, which allowed Eleven Labs to leapfrog Speechify. He now advocates for a "compound startup" approach where companies continuously innovate beyond their initial wedge, offering products for free or cheap to gain adoption before expanding into adjacent offerings. He defends Speechify's move into B2B despite competition from Sierra and Eleven Labs, arguing that the market is oligopolistic and that being in the race is essential. On hiring and engineering, Weitzman emphasizes technical aptitude and learning speed over existing expertise, noting that AI agents amplify individual productivity. He describes how Speechify engineers orchestrate agents for long-horizon tasks, with credit only given when features ship to production. He also discusses the importance of owning compute to unlock team velocity, comparing it to giving Michael Jordan a basketball hoop in his house. Finally, Weitzman shares his excitement about applying AI to biology and rare diseases, describing his personal efforts to sequence genomes and run analyses on GPU clusters to solve his brother's autoimmune condition. He believes voice will become the dominant human-computer interface, and he remains optimistic about founder-led companies like Meta and SpaceX despite their challenges.

Transcription

14291 Words, 76697 Characters

English
Speaker 1100% is on. It's the biggest strategic mistake I made in the history of Speechify. The best way to lose is not to be in the race. Be in the race. You don't want to be a fat manager who is like a general sitting in the back saying, take that hill. You want to be the warrior who runs up with their sword and engages the enemy first.
Speaker 2This is 20VC with me, Harry Stebbings. Now, I am fed up of the simple question-answer, back-and-forth podcast. Today is a real frickin' discussion. Cliff Weitzman, founder and CEO at Speechify, one of the fastest-growing text-to-speech startups in the world, on the show where we have a real debate about whether it's right to scale into enterprise from a phenomenal consumer business, what it takes to build an amazing go-to-market motion when you've already built this amazing consumer business. And then he also tells us some wild frickin' stories about spending tens of millions of dollars on NVIDIA GPUs and why so many more companies should be doing that over relying on other providers. This and so much more in the episode today. But before we dive into the show today, today I want to tell you about how the first AI law firm, Crosby, helped us close a big sponsor. As you know, some of the biggest companies in the world advertise on 20VC. My British dulcet tones clearly convert well. I was working to close this big sponsor, and they wanted to get through legal review quite quickly to close the deal. Crosby turned red lines around in three hours and caught major issues that would have caused us serious problems in the future. Crosby combines AI, some of the best engineers in the world, from companies like Ramp, and Stripe, and some of the best attorneys in the world from top 10 law firms. Customers get the best of both worlds. An elite human attorney reviews every contract, but they move incredibly quickly, returning red lines in under four hours. They help the fastest growing companies like Cognition, Ramp, and Clay close deals in hours, not weeks. Learn more at crosby.ai/20VC If you want to red line NDAs, MSAs, DPAs, and any other procurement contracts faster, go to crosby.ai/20VC. It's speed that you can really trust. While Crosby keeps your numbers sharp, OneMind keeps your customer conversation sharper. Our friends over at OneMind have a hot take. The B2B GTM model we've been using for, well, the last 20 years, it's collapsing. Predictable revenue isn't so predictable, and buyers are just tired of explaining themselves at every handoff between SDRs, AEs, CSMs, and support. You feel it. You feel it in your board reporting. Your sellers feel it in their coverage. Your buyers feel it as they wait for answers. Well, enter OneMind and their GTM superhumans, HubSpot, AltaSense, on an investment that helped us close an $8 million deal. $8 million, baby, that's a lot of money. That's why I'm genuinely excited to have them as a partner on 20VC. AlphaSense combines AI with one of the world's deepest libraries of market intelligence, including expert interviews, broker research, earnings calls, company filings, and real-time news. Every answer is grounded in this incredibly trusted evidence and fully traceable to the original source, which is so important. So you can make really high conviction decisions with confidence. But the best part? They're building super analysts and always on AI analysts. So instead of starting your day with another search, you'll start with work you've already done, your coverage monitored, the important developments surfaced, and your investment brief already waiting for you. See for yourself. Head to alphasense.com/20VC, that's alpha-sense.com. You have now arrived at your destination. Cliff, it is so good to have you back in the studio, dude. I was looking forward to this one, because when I was writing it up, it's a very different thread of conversation to how I'd normally go. And so thank you so much for joining me again today, dude. My pleasure. Glad to be here as always. Now, I wanted to start with you're spending tens of millions of dollars on NVIDIA GPUs, and you're paying an additional $100,000 per GPU to receive them four months early. Why? Like, what do you know that the market doesn't know?
Speaker 1So in 2022, we bought a huge rack of GPUs from NVIDIA. And the reason we bought them is for training, right? We have a bunch of models. The newest Speechify Simba 3.2 model is ranked number one in the world for quality. Above all the Frontier Labs, 10x more affordable and stuff like 11 Labs. And we used to rent GPUs. And we found that engineers at Speechify would be parsimonious with how they use the GPUs, because they were like, "Oh my God, I'm costing the company tens of thousands of dollars. Like, I don't want to do that." And the analogy my brother and I came up with is imagine you're Michael Jordan, and you want to be in the NBA. It's the only thing you care about. And you need to pay $20 an hour just to train in a basketball center. Well, that sucks. You want one that you can go to whenever you want to. In fact, you want a hoop in your house. And so our initial idea was we want a hoop in our house. And so we bought a bunch of our own GPUs. And that deal ended up being really good for us. And we ended up training really good models. So with time, we invested more and more and more and more. So that's the first part. The second part is actually how the economics work out. So if you look at it, the Transformer was invented inside of Google in 2017. NVIDIA came out with A100 GPUs in 2019. Shortly after, they came out with H100 GPUs. The original Chet GPT was trained on A100s. And then they came out with Blackwells, so then B200s, B300s. And now they came out with Rubens, which is the GPUs that Elon is sending to space. And they're liquid-cooled. They're very, very cool. And we're like, "Okay, huh." One, every class of GPU is more affordable per one trillion flops. So a flop is additions. Additions, subtraction, multiplication, any mathematical operation. And you measure them in how many trillion of operations happen per second in a GPU. And so they're more affordable as it relates to this. If I was to buy an H100 for, let's say, $30,000, that's how much a single card would cost. If I wanted to rent an H100 for one-hour spot instance from GCP, it could cost me $5. If I rented it from Azure or AWS, maybe it'll cost me $3.5 per hour. So if I multiply that times 24 hours, and then times 365 days in a year, I'm actually going to end up paying $35,000 to $50,000 to rent that GPU for one year, but I could buy it for $30,000. So it's 1.5x the cost of owning the hardware to rent the hardware for a year. Now, the hardware is typically warrantied for three years to work properly. But it'll keep working up to the warranty for, I imagine, I don't know, 10 years. So the math just maths where it makes way more sense to buy them. The other big part is if you want to do large-scale training like we do, you need the memory to be co-located with a large cluster of GPUs. I can't just rent from Google or Microsoft or even base 10 and run the size of training that I want because I need a gigantic memory card next to it with all of my data that all the GPUs are accessing. So that's why we first started buying them. The next thing that we found is actually if you run open source models for coding, you could pay Anthropic. You're paying for all the tokens. And the fact that you're doing that, the branded, right, Fable 1. Or you can run an open source model. And instead of running it on a spot instance from Azure or anyone else, you run it on your own hardware. And then you're paying a fraction of a fraction of a cent per token. And so for all those reasons, it made a ton of sense. But we can go into all the depth that you want.
Speaker 2I just want to dig in. The first thought that I have is I completely understand the rationale there. But chips depreciate. You have chip cycles, and they are accelerating. We are seeing newer and newer chips being created. We're seeing specialization within chips. By buying, you're locking yourself in, so to speak, to one chip architecture. How do you think about that?
Speaker 1At Speechify, we still use K80s for a lot of specific operations for inference. And we use older models of GPUs constantly. And there's essentially a difference between when you do inference and when you do training. For training, I'm like, okay, I have this hypothesis. I want to know the answer to this hypothesis as soon as possible. Like every minute that it doesn't come out, I'm in competition with everybody else. And so having a GPU architecture that is much faster by orders of magnitude is a huge advantage. But if you go speech-to-text or text-to-speech with Speechify, I can afford to give you a lower quality GPU, and it'll give you what you need still in 100 milliseconds. So it's totally good. And so I can always use these older GPU models for inference. That's number one. Number two, we have so many experiments that we're running at every single point in time. Not all of them need to run on the newest hardware. So the analogy I always give, let's say you bought an iPhone back in 2011, and it's an iPhone 3G. And then you bought another iPhone and another iPhone and another iPhone. You could have a drawer in your house with like five iPhones that are collecting dust because you can only use one iPhone at a time. But if I own 100,000 GPUs, I'm still going to use all of them at the same time. And so I'm not losing anything by having more GPUs because not only do I own a bunch, I still rent from the hyperscalers all the time. And I rent both dedicated instances that I prepaid for, and I rent spot instances. For example, more people use Speechify in September because everybody goes back to school. So I need to like level out the load. And so the parts of that load that I know for sure I'm always going to use, whether it be training or it be inference, I might as well just own it. And then on top of that is also the case that I have so many other friends who are running training and running inference, I can always rent it out to other people if I have excess capacity, which I don't expect to have. But like every once in a while, you have an interesting situation. So for all those reasons, it just makes mathematical financial sense. Lastly, if you have excess capital, really you either stick it on the bank or you buy a bond. The best bond you can buy long-tail will yield you like 5% or you can buy a GPU. And because renting it would cost me 1.5x buying it for the year, the return is like way higher. So how many GPUs do you buy then? Let's talk about Rubens, for example. So Rubens come in the form of 72 cards in one rack. So we'll buy multiple racks of Rubens. And on top of that, we'll buy B300s, which are like the newest form of Blackwells, because we can get them earlier. And then the same thing, like when we bought our first instances of DGX H100 GPUs, we bought just like a bunch of racks of those. And then those get delivered in a truck to the data center. We rent the data center, and so the data center provides the networking capability. It provides the energy, which is actually the largest constraint now. And it provides physical engineers that take it off the truck, they install it. If it has an issue, they fix it. And then it just runs. Does 11 Labs do this? Yeah. 11 Labs is amazing at this. 11 Labs, I think Piotrek at 11 Labs literally bought a bunch of GPUs early, early on and set them up in his house. And then they just kept building bigger and bigger and bigger clusters. They do the same thing that we do. How do you think about forecasting chip
Speaker 2buying? It's incredibly difficult to know A, demand, but also B, supply of chips. How do you think about forecasting chip purchasing?
Speaker 1Yeah. So number one, I want to explain again, it's very different than buying an iPhone or buying a MacBook. I can only use one MacBook at a time, one iPhone at a time, but I can use all the chips I have at any given point in time, and I still will have more demand, especially when I have multiple teammates and 60 million users who are using inference on my Speechify software that's providing text-to-speech and helping them read their work and dictate. And, you know, use Speechify work, which is our newest product that's a GenTech, kind of like Jarvis from Iron Man. And so I go, okay, let's imagine I have a hundred percent capacity that is the average usage per month that I need for GPUs for training of my AI models and for inference on my AI models. Inference is when you actually make a call to Speechify and you give me text and I give you back audio, like there's math that happens in the background, that's inference. Training is I take a gigantic amount of data. I take all the architecture and software engineering that we're doing. And I take all the data that we're doing and I take all the data that we're doing and I go, I think that this will give me a better model. I'm kind of baking that model in the oven and I'm going to come out with a new black box. And then when I give you text, that black box is what calculates it and gives you back the audio. So those are the two usages. Let's say I have a hundred percent, which is what I would have in, let's say a month like November. In October, I'll have 140% because it's like a big month for us. In December, you know, everyone's at home, you know, they're not necessarily studying or working. So I might have 80% utilization. Okay. So I go, cool. Well, I can take 20% of the usage that is normal and let me buy it because it's the best deal. I'll take another 25% of the usage and I do long-term contracts with hyperscalers. The rest, I'll rent what's called spot instance from the hyperscalers. And then I'm still not even close to overcommitting myself. And so that's kind of how we think about the math. And then we go, okay, well also we have 45 engineers, but we want the team to be 150 engineers. And even inside of my 45 engineering person team, like there's a couple of people who are rock stars. They have dedicated like DGS, racks just for that one person. And 25% of my team are almost like waiting. And I want to double the size of the team. It's like you have a football team and you just need another field because they don't have enough field to practice on. And so that's how I think about how to allocate. And then in terms of depreciation of the asset over time, I go, okay, well, these are still amazing GPUs. Like even A100s, you can run amazing experiments on. It's completely valid to use that as long as it's hooked up and as long as it's not stopping to work. And so think about the mileage of a car. If a car gets like 250 miles an hour, it's going to be like, oh, my God, I'm going to break at this point. That's not necessarily true for a GPU because it doesn't have as much wear and tear. Yes, it's moving. And yes, all these things, but like it's in a very clean environment. It's very much cooled. It has constant maintenance because it's not moving around. It's very expensive. And NVIDIA just does a really good job. And so that asset is going to stay for a very long time. And let's say it got so not good, so outdated that I no longer can run training on it. Cool. Now I'll use it for inference. There's one more thing that's very interesting that just happened. I believe earlier this month, NVIDIA did a huge deal with Blackstone, BlackRock, Apollo and Goldman Sachs. And they said, listen, we want more people to buy more GPUs. We're going to underwrite for you up to 25% the value of a GPU. That if you lend money to someone who buys a GPU, let's say Google or a startup, CoreWeave, and that startup goes out of business and you have that GPU as collateral against that investment, we'll buy back the GPU for up to 25% of the value of the GPU. And so they're succeeding in creating a liquid second. They're a market for GPUs that they're underwriting. So now the large banks have an incentive to loan money at much better interest rates. This is actually exactly what Elon did in the beginning of SolarCity. He went to Morgan Stanley and Merrill Lynch and got them to amortize the price of solar panel over 30 years. So the whole invention behind SolarCity was the fact that you could take a loan against the collateral of your solar panel. So NVIDIA has done an amazing job now in creating a clear floor for the value of the GPU over time.
Speaker 2Do you think the circular economy fears that people often, past against NVIDIA, are justified or not? We saw their CFO push back on them and say, "Enough. Enough of this bullshit." Do you think that justified or not?
Speaker 1I think that a lot of the things that about a year ago, we're going on between Oracle and OpenAI, that was way too much. That was ridiculous. I think the NVIDIA stuff is not, because you're talking about a real asset. So if you think, for example, about the logic behind the value of Bitcoin, Bitcoin, what is the intrinsic value of Bitcoin? I can't really tell you, right? What's the intrinsic value of gold? Well, gold, you can use it for some medical stuff, because it's a really amazing metal, and it's jewelry, whatever. But a GPU, it has intrinsic value. You can actually use that asset for something that's really, really valuable. And it doesn't matter where that GPU is. It could be in Iceland. It's still useful to anybody all over the world, as long as it's networked. And so actually, it has a pretty good store of value. Even if new GPUs come online, really my one question, and this is the math for everybody to come back to, is how many teraflops per second can this device do? And that is essentially a token. That's the value. And so there is intrinsic value. So yes, you can have all these circular things, but at the end of the day, NVIDIA is making a product that's real. It's not complete tulip mania. There's a real, real intrinsic value here.
Speaker 2What does no one know about buying chips that they should know? What's the like, "Oh, my God, people are so naive about this."
Speaker 1I mean, it's not that people are naive. It's just they haven't been in the space. So I'll give you an example. Imagine you're buying a GPU. Well, you're going to buy it. It's an NVIDIA-produced product. But NVIDIA is not going to buy it. It's not going to buy it. But NVIDIA is not going to waste the time talking to Cliff Weitzman. So who do I buy it from? Well, one of the best-rated vendors is Dell. So everybody thinks Dell is a personal computer company. No. Dell is a GPU rack supplier at this point. And then, okay, I want to buy it from Dell. Well, Dell has a constraint because there's not a lot of Blackwells out there. Well, it happens to be that they have some in France. All right, well, I'm going to order mine from France. Okay, shoot. It was supposed to come a month ago and it's still not here. So then you need to negotiate to make sure that you get it, which is why we're very willing to pay $100K per month extra to get them earlier. So you'll call up Pierre in France and say, "Hey, we'll give you an extra $100K kicker if you get them here in a month?" Even more than that. So in that France situation, which is something that happened to me, I was like, "Pierre, what the heck? We have a contract. You're not delivering on time." And so it is the case that we had a contract with another company beforehand, and they were a few weeks late. And I called them and I was like, "Listen, I've got a better deal. I'm canceling our contract because you didn't deliver. So I'm going to go with this other contract." But if you have a better price, we'll go with you, but I just need the GPU now. And remember, I'm paying for the renting space of my data center. So the most expensive part of a delivery of a GPU is if it's late, I'm still paying rent for that data center space. Now that GPU... And so you put pressure on Pierre to send you the thing when he said he was going to send it to you, and then you go to Nvidia or Dell or whatever, and you're like, "Oh, it's a market. Hey, can I pay more to get it earlier?" Skip the queue. Yeah, you can. Cool. Now there's a truck somewhere in the United States with a GPU whose value is the value of a house that's coming to my data center. Well, I should have insurance on that, right? Because if that truck gets hit or there's too much humidity or the GPU gets flipped, I lost multiple houses worth of GPUs. So okay, the insurance is really, really important. And then also the value of the amortization is really, really important. And it's like, there's all these nuances of how to do the math through. Then there's the cooling, right? So you're not only paying for the physical space and the networking and the energy, and the energy is the biggest constraint. We'll talk about it in a second. Well, how do you cool that thing? Because you have a thing that's just like... It's just like moving, and moving, and moving, and moving, and moving. Well, the thing that's most new now is liquid cooling because air is just not enough. And the thermal load of water is much better. And there's other liquids that are even better than water. And so Rubens are liquid cooled. But most of these data centers don't have liquid cooling installations already approved. So we had to do a bunch of research and we find, okay, we could buy what's called a side cart of liquid cooling that you enter into the data center. Then you pay someone at the data center to install it for you. Cool. Now you can have the rack that you want. And so there's a big difference between running a purely software company and running a company that includes hardware.
Speaker 2But when I listened to all of this, I'm now more sure than ever that it is a mistake to price optimize and to spend the money to buy it versus to rent it. Because I get you on the optimization, but you're not saving 10 times more. It's 0.5X more per year. No, no, per year. Exactly. Yeah, per year. But you have the flexibility to tailor it up and down. You don't have any of the logistical nightmares of insurance, transportation, security, all that kind of security, water cooling, logistics. And then you can build your product that actually what matters most against Eleven Labs who are fucking running fast. I don't want to worry about water cooling and insurance for a freight truck.
Speaker 1Eleven Labs worries about the same thing because for them to train excellent models, they need to have co-located GPUs with a lot of memory available. You can't do it if you rent it. You can. It just becomes one, ridiculously expensive. Two, you need to commit for many, many years ahead of time because you need to build a co-located cluster. And then you don't have as much control because you don't own it. So it's difficult to... You suddenly need an InfiniBand cable, which allows for the memory to flow from one DGX to the other one. And the answer then is, well, now my ability to train is so much bigger. I can have a larger AI team. Every person in the AI team is leveraged. I can shoot ahead of everybody so much faster. And let me just make one thing clear. If I want a Reuben, which is like these much faster GPUs, I'll get it faster if I buy it than if I wait for Google to buy it and then there's other people in front of me in line. So I'm going to skip to buy like a lot. And then I'm going to have like a year of access to Reubens before everybody else does. The way I think about it is the following. How do you build an amazing company in a world where there's so much competition today. The team is the most important part, but the team is the most important part because the team gets you the other resources. And so what are the missing pieces? The missing pieces are data, compute, and architecture. In a world where intelligence is commodified and no one needs to handwrite code at all anymore, our engineers, really what I'm looking for is 10 really good decisions per day, which is very tiring, not like optimizing random parts of the code. And each one has five to 18 agents running at any point in time, doing long horizon tasks on these GPUs, coming up with theses, testing them, going back and forth, back and forth, back and forth. If they don't have the capacity to train, the team is limited. If they don't have the data to train, the team is limited. And by the way, a lot of data, cleaning the data, right? You get this raw data in the beginning. Well, you need to organize it into data sets. And so you need the GPUs to also organize data sets too. I have one of my best engineers right now is not even writing models. He's making synthetic data sets to train models. And so it really becomes an indispensable asset. Would you ever buy data? We have, but like small data sets.
Speaker 2So I suggest you use Fireworks. But I mean, Fireworks is amazing. Lynn and the founders, she's one of the co-founders of PyTorch. But I had the very obvious realization that you'd have every company having their own specialized models of a certain size, trained on their own data. But you would need supplemental data, like this synthetic data, or real world data that you just don't have. And- And that you would buy that from data providers like Mercore, which is why I was saying-
Speaker 1Yeah. Micro One, Surge, all these companies are amazing. And they shorten the cycle, by the way, to getting to revenue. A hundred percent. Because if you are a Mercore, shout out to Brendan Foody, Eleven Labs or OpenAI or Anthropic is going to make money for the next decade or two on the data that they bought from you. And so they're willing to pay a fraction of that 10 years of revenue to you today to supply them the data. And again, it's all a speed thing. Yes, Eleven Labs or, you know, OpenAI can go and make a team that will get the data, but they don't want to manage it. And so the data is key. All this, like you need all three things. You need compute, you need data, and you need a team that writes great products. And ideally, you need users that use you a lot, and a lot of them, to have a feedback loop of whether this stuff is good or not. So benchmarking.
Speaker 2And so for me, when I was doing the, as a venture investor, we do outcome scenario planning, which is the most bullshit exercise to pretend like you're smart, predicting the future. We do it because, you know, it makes us feel important. But, you know, they've predominantly settled to frontier labs today, and that's where 90% of their revenue comes from. And so, you know, they're getting their revenues from, with the rise of specialized models on a per company basis with their own data. I believe that you move that customer base from purely frontier labs to every large scale enterprise who needs supplemental data. If that is the case, how big a outcome
Speaker 1is the data marketplace? So the first problem to understand about the data marketplace is it's not ARR, right? It's not annual recurring revenue. It's a one-time deals every single time. So the buyer of the data is not required to buy it from you again. So it's a very risky business. And if you look at early days of companies like Mercore, they didn't raise significant funding off the bat because investors were very skittish about that fact. Let's put that aside. Well, very important, the T's are crossed and I's are dotted about how you got that data. You need to indemnify the companies who are using you. And that's part of why they buy it from you as opposed to sourcing it themselves, right? We've seen the lawsuits, but it's a great business. And if you could do it well, but like you need to be an ops monster, like you need to be really, really good at operations. You need to be very fast. And really the key is, is the company training on your data needs to actually see improvements in their model at the end of the day. The thing that has always been challenging for Speechify compared to other companies is B2C customers pay a lot less than B2B customers. So Eleven Labs, huge credit to them, leapfrogged us because they sell to B2B. Well, historically we've only sold to B2C. And so our big constraint is we need to do this on a cost basis of it needed to cost us less than $10 per million characters. Eleven charges $100 per million characters. The open AI model on the benchmarks cost $96 per million characters. So ours, when we sell it to other B2B companies now, we just launched our API, Simba 3.2, it costs $10 per million characters. Dude, I'm too old to not ask the painful
Speaker 2questions. And I think the joy is when you ask something kind of less, you know, worried you're about asking, you said Eleven Labs kind of leapfrogged you. Is that on you for not doing B2B?
Speaker 1Yeah, a hundred percent it's on me. A hundred percent it's on me. It's the biggest strategic mistake I made in the history of Speechify. How do you reflect on that? So I met Piotrek and Mati. I was living in London, at the time, in my house in London. I think it was 2022. And we were very impressed by them. And we wanted to use their model, by the way. It was just too expensive for us to use. And I looked at it and my thought to myself was, they're very smart. They're going to do well, but I don't like their strategy because I think that an API for Texas Feature will become commoditized with time. You're going to get to the point that you can run that API on your computer and then on your phone, and then like, what are they selling anymore? So I don't want to go into that business. And I made a critical error. What I didn't understand is that the point of an AI lab like Speechify or like Eleven Labs is to continuously innovate. And the first product that you release is your wedge that gets other people to then later use your other technology. So for example, if you're in text to speech, you build the best text-to-speech model in the world for one specific voice. Cool. Well, now you can do other voices. Now you can add emotional prosody. Now you can add voice cloning. Now you can add speech-to-text. Now you build duplex models where it makes the um, ah, laughter, interruption handling, turn-taking. You add a harness for voice conversations. Then you optimize it for sales. Then you optimize it for sales. Then you optimize it for sales. Then you optimize it for customer support. You optimize it for all these things. And so what they did is they first built an amazing API. They were great at launches. They built a really great product for creators. Then they built their best product ever, which was Agents. Agents is amazing because the buyer is no longer a software engineer. The buyer is a CTO, CIO, CEO, executive in the company. Sierra has this concept called outcome-based pricing. Brett Taylor is amazing. And so you can start finding on the outcome and having a AI agent is like having an AI coworker. But it was my mistake to think that an API product was a bad strategy because I thought it was something that would become commoditizable. And I forgot the central thesis about Silicon Valley, which is constantly innovate, get the user to start using your product. I don't care if it's free. Then you sell them other things. And so that was my big,
Speaker 2big, big, big mistake. How possible do you think it is? I think people underestimate the complexity of building out a B2B GTM. Super hard. I think it's a strategic mistake for Speechify to go to B2B. Tell me your position. You are now competing against Eleven Labs and Sierra. And those two are competing. Whether they like to admit it or not, they absolutely are competing. And they will, I'm sure, if you ask them off camera. That's Brett Taylor. Yeah. You don't want to compete against Brett Taylor. Motherfucker, I don't want to compete against Brett Taylor. And that is the tidal wave of Eleven Labs. Now, Eleven Labs is an unstoppable machine at this point, to the point where it has government buy-in across all of the large major Western democracies. Actually, it's insane the government buy-in they have. And they started three months ago. You just said the key thing. They started three months ago. Yeah. And so, you know the graph of opening... But I think they've reached a tipping point where actually, they've just taken the market. I think Sierra are running behind them chasing. And they're doing a decent job of it, but they've got Brett, and they've got Sequoia and Green Oaks and every royalty of Silicon Valley behind them. And they're still running behind chasing Eleven Labs with Sequoia kind of pretending to be neutral because they're in both of them, which is incredibly challenging. And I just think being third, the Postmates effect is never a good market to be in when I could be the dominant consumer brand that leads with a really different and compelling story.
Speaker 1So, here's the two things to consider. The first one is if you go to the app store and you search Texas Speech, Speechify has 98% of the installs in Texas Speech for B2C. Speechify has served more than 770 billion words to users over the last few years, which in terms of times of listening, it's like 6,000 years of listening. And that's a lot of time. And that's a lot of time. And that's a lot of years of listening. If you go from today to zero BC and back, you still have thousands of years left. So, we've completely dominated that market and it's still a business that's growing really, really fast. And we're constantly adding more features into that product. The thing is, we have a pretty big engineering team and now everybody is capable of doing 10X what they did before. So, I have extra staff. I have a huge AI engineering team with ability to make amazing models. So, where is the highest ROI for that to go? Well, it needs to go both B2C, but it should also go B2B. And one thing that I will never be is a person who doesn't learn. So, I might as well freaking learn B2B. Now, to your point about competing against giants like Sierra or Eleven Labs, hey, Anthropic came into the market as a second to OpenAI and they were second for a very long time. And now they're not second. Facebook came as a second to Frontster and MySpace, and now they're not second. And so, the nice part is this space is not a monopolistic space, it's an oligarchical space. And if you look at what happened with Eleven Labs, I'm going to exclude Sierra because Brett Taylor effect is huge. It's just amazing to see how good of a business that is. And so, it might very well be, but it's not going to be. It's not going to be. It might very well be that for the core offering that they're currently winning on, I will not win. But what did I learn last time? It's fine if I offer my product essentially for free because I'm an AI research lab. And as long as people start to use me with time, I'll be embedded in the system and I'll keep coming out with more and more and more innovations that are useful to them. And so, there's unbelievable demand from all these companies and governments and everybody else for great tools, whether they be AI agents or APIs or products. I just want to be on your phone if you're a user or in your stack, if you're a company, and supply you with the best front deploy engineer experience, an AI orchestration experience, an API experience to give you an amazing experience,
Speaker 2and there's room for everybody. I agree there's room for everybody. Value accrues to top one player. I agree. I think it's kind of like the inference market where fireworks will be a multi-hundred billion dollar company. And then I think a genuine base 10 will be a hundred billion dollar company and then together and a load of the others will be 15, which is amazing. It is completely true. Power loss, yes. Hugely amazing, valuable companies.
Speaker 1But you would then think that open AI would be the place where value accrues for voice AI, right? That's what you would have thought three years ago. And that's not what ended up happening. So you can't not go into the race because there's a big incumbent. Well, I think with all candor, that's because of incredibly poor management.
Speaker 2I agree. And hiring. That was theirs to take, and they fumbled the bag across every spectrum.
Speaker 1And for every company in the world, no matter how exceptional the leadership team is, niches get fumbled, right? So voice AI was a niche for open AI, right? LLMs were the core. And by the way, they also fumbled AI coding. Now they're trying to cash because it's such a big space. All respect to Piotrek and Mati. I think they're absolutely amazing, and I love working adjacently to them. I just don't think they're going to fumble the bag. That's my trouble. But they have so much in their net right now, and so much is getting added to the net constantly. That's true. And so you have to go where the football is going. And so I think that it's too expensive for speechifying not to be playing in B2B as well as playing in B2C. The best way to lose is not to be in the race. Be in the race.
Speaker 2In terms of the products that we built, we were chatting earlier. You said that every startup has to be a compound startup. Can you talk to me about that and how you think about that?
Speaker 1It's not that every startup has to be a compound startup. It's at a certain point, you can't afford not to be that.
Speaker 2Do you not think there are a few companies that are just absolutely fucking running rings around everyone else?
Speaker 1Yeah, absolutely. Those are the winners, right? Eleven Labs is an example. Anthropic is an example. Ramp is an example. Speechify is an example. All the companies that have absolutely maniacal leadership teams and engineering teams, that's why people care about team more than almost anything else. Because the right team will iterate fast, get there, and then figure it out. And now… You know, when everything can be turned into a reinforcement learning problem, where you can have long horizon agents and orchestrating agents thinking about the problem for like two weeks at a time, if you set that up, of course, you're going to win.
Speaker 2I got into a lot of trouble, as I always do with most of my social posts. I used to be quite a sweet little boy, actually. No, really, I used to be like the Harry Potter of venture capital, and now I'm more like…
Speaker 1Yeah, you lost the glasses.
Speaker 2Lost the glasses, and I kind of became more like Piers Morgan, if you know Piers Morgan in the UK. Highly, very opinionated. But a question that I have is like, I said, if you're a startup, it's never been harder to hire great talent. OpenAI and Anthropic have such a carrot reward mechanism in front of you, especially with impending IPOs, that the best talent just wants to go there. And talent follows talent. You've seen the fucking founder of Monzo, a multi-billion dollar bank in the UK, go there from YC as a partner. Matt Clifford, the founder of EF, which is a multi-billion dollar company. I mean, he should be fucking prime minister, and he's going to join Anthropic. Am I wrong that this is the hardest time ever for startups to hire? Because the price… The prices of Anthropic and OpenAI are so great.
Speaker 1My favorite type of person to hire is a CTO of another company. We have, when we were 21 people at Speechify, 18 of the folks at the company were previously either CEO, CTO, or VP of engineering of the last company. Anthropic, I have never seen a company like this, hires so many CTOs of publicly traded companies and other successful startups. Workday, one of them. The reason is they build the best, most beloved product for engineers in the history of the world. So it's easy to hire CTOs. By the way, they hire much more CTOs than CEOs, because CTOs are the ones who get the most excited about this product. And you're right, they're the fastest growing company ever, especially at the scale that they are. So they're going to keep growing. You had this like very condensed period, like fireworks of growth in both of those companies. Yeah, it's very hard to hire. But remember, they're hiring people that their annual compensation needs to be $15 million a year, minimum. What startup is hiring someone and paying them $15 million a year? You're not, like your seed founder. That was not something that you were going to hire. And so I will push back against it. The competition for growth stage companies hiring exceptional leadership talent is more difficult. For seed companies, I would say it's the easiest time ever, because the impact of even just the founder on their own is bigger because they can orchestrate agents. But the same thing for hiring. So one thing that we have changed about our hiring in the last even six months, we really cared that you read a ton of textbooks about software engineering and that your handcrafted code was amazing. I still care that you read a lot of textbooks about software engineering and you understand it. But the thing I care about the most today is technical aptitude and just like raw technical intelligence, because I know that we could teach you everything else in six months. You could be a machine. And so we hire a lot of math Olympiads and elite coders, Kaggle award winners and people who studied physics and math. They might have even not coded before because I just need the hunger and the work ethic and the intelligence. Anyone can become so good so fast. Anyone can become so good so fast. And so the pool for hiring exceptional talent is bigger than ever before. And Duolingo did this really well. They love hiring college grads and coaching them. And so I wouldn't say that it's harder to hire than ever before for seed companies. Seed companies now, almost anyone can be someone that you hire if they're smart and hardworking because you can teach them very fast. What is more challenging to hire is for growth companies because you're fighting with just absolute juggernauts.
Speaker 2And you're not a growth company.
Speaker 1So it's challenging for us. Why do you think it's hard to hire a really good salesperson? I totally get that.
Speaker 2And I completely agree. I will see CRO packages in the $50 million plus range, by the way. Yeah, exactly. $15 million is like kids play. By the way, with the greatest of respects, I will even see $15 million on the table for comp packages for seed companies today.
Speaker 1That is the dislocation that I think, with the greatest of respects. Wait, wait, wait. Sorry, sorry. But this is a seed company that has raised how much money at what valuation? I mean, you've got to understand a seed round today will be $150, $200 million.
Speaker 2And... And there are several of them. I mean, there's 30, 40 companies. That at seed have raised $100 to $300 million.
Speaker 1And this is a company of like a guy who's like one year out of university?
Speaker 2No, no, no.
Speaker 1This is a guy who's probably spent four years at OpenAI or spent four years at SeedMind. So then what about the company that's like, you know, the guy who's been in university for like two, three, four years and now they're starting a company? Or do you think that those people are out of the water now?
Speaker 2No, I think that that's just a very different world.
Speaker 1And so, yeah, they'll raise $10 million seed rounds. Yeah. So for the company that you just described, they raised a seed round at $150 valuation. And they raised... I don't know. $20 million? No, I said it was a $150 million raise. Oh, I wouldn't call that a seed round. Maybe that's the name.
Speaker 2But my point is, and that's my point though, which is like the talent is concentrated. The people who really fucking get AI and systems and have seen the magic inside OpenAI, Anthropic, SeedMind.
Speaker 1Yeah, I agree with you that if you have a company that's raised $150 million at a $500 to $2 billion valuation, definitely that company should give $15 million comp package. And there's a lot of them. Yeah. That makes perfect sense.
Speaker 2But there's a lot of them. There's 30. And those 30 take 30 people. And there is 1,000 people now.
Speaker 1And that is fucking hard. But what you just described is exactly what used to happen with Google and Meta, let's call it six years ago, which is if you were really cracked, there was essentially a maximum amount that you can get paid at a company like Google or Meta. And the best way for you to make a life-changing amount of money is to go to a company that is small and ride from the beginning all the way through and be a really solid founding engineer. I think people want more certainty of cash today than upside, which sounds... No, I think that the equation is the same as always, which is each person has their own equation in their head of how much certainty and how much risk they're willing to take. It hasn't changed. It's the same. Humans are still humans.
Speaker 2But I think people would rather know that the certainty of a $10 million from Anthropic versus a 60 from that quirky startup they could make.
Speaker 1This is the reason why companies IPO. There's two reasons. Either you want a ton of money or you want a lot of credibility and B2B like Zoom. Or you're hiring and the value of the package that you offer is so much better when your stock is liquid.
Speaker 2When we look at that dev team for you today, you said, hey, I wanted to go in. I want to see how we're orchestrating agents. What did you find? What did you learn in that discovery process around agent orchestration internally?
Speaker 1So inside of our AI research team, everybody's orchestrating agents. It's when you go lower. If you go then into the product facing things that we build, for example, the platform team or the iOS team or the Mac team or the Chrome team or the web team or the Android team. These are super smart folks who have been working in those domains for like 10 years and they know iOS like the back of their hand. They know Kotlin, JetBrains like the back of their hand. And so it's very easy for them to hand code things because you're not dealing with something that's like super, super new. So why change? It's hard to change. Right. And so you just need to force them to change. So one, the best thing is to inspire. So you do a Zoom screen share and you show them how the best engineer in the team is orchestrating agent. And they're like, oh, wow, I didn't know you could even do that. And then you go, yeah, like, please do it. You recommend. A blog post for them to read, books for them to read, Twitter threads for them to read. What's the team using? Cloud Code, Cursor, Codex? Cursor and Cloud Code. Those are the two most popular. Yeah. It's a little bit of Codex usage. It's not mapping. I would say Cloud Code is number one, then Cursor, then Codex. We want you to use as many tokens as possible in whatever harness way is the best for you. You mentioned Linear. Linear is amazing. Like automatically cutting tickets from Linear is fantastic. And just like being able to go into your agents and be like, OK, I have these like six Linear tickets. Start on them. And then. Really, a good engineer today is just an exceptional QA. The AI will make them feature. You will test the feature, see if it's good. You'll figure out where the edge cases are. You'll prompt it to fix it. And then you try to make it as efficient as possible, which is hard to do. And then you need to make essentially like, you know, roughly 10 really good product and engineering architecture decisions a day.
Speaker 2How do you think about token allocation internally? You know, we've seen leaderboards be used, which is, I think, the most fucked up form of incentive kind of playing. You don't want to.
Speaker 1Prevent. There's a lot of people who are a lot of talk and I'll ask for examples and I'll read the examples and like, you know, I'm doing this, I'm doing this, I'm doing this, I'm doing this. And then you look and I'm like, and so I think about it in terms of demos. Can we hop on a Zoom call and you'll show me what you built and then I use it myself and I'm like, wow, that's amazing. Or you send me a screen recording of a feature or technology that you built and I'm like, wow, that's so good. And so we give credit when things get shipped to production to users. So even inside of the AI team, if you build a really amazing. This is part of why Speedrify ended up waiting. You asked, "How did you build bigger labs?" shipped to production all the time. That's how we won. We are not in the theory space. We are an applied AI company. That's why we win. And so if you're an engineer at Speechify, the analogy I always give people is imagine that you are in the milk delivery business and you make me a beautiful bottle of milk and you leave it down the road. The milk will spoil. You have to get it to my door. Knock. If you didn't do that, you get no credit. If you carry the football all the way to the line, but you don't cross over to the end zone, if you don't kick it into the goal, you get no credit. If you bring the ball just to the rim, you don't put it in the rim, you get no credit. And in the rim means push to production with no bugs and users are actually using it. And then we get feedback. How many companies do that iteration cycle fast? Almost no one. Definitely not with a user-based number that Speechify has. And so in the AI team at Speechify, you make some amazing discovery. We're like, great, push it to production. And then you go, oh wait, there's this QA problem and this QA problem. And if you have this many people use it on the AI serving layer, then you have this other issue. Cool. You get no credit from me. It's not in production. I can't use it on my phone. When I can use it on my phone, I will give you credit. And so this morning, actually not yesterday, yesterday I had a call with our AI engineering team. And I said, listen, the projects that we have running for duplex models and for AI conversational harnesses is something I'm really excited about. And it's been moving fast. I want it to move faster. Here's like 14 notes that I want. And then what I do always is I'm on a Zoom call. I flip my computer around to face my phone and I use the product in front of them and we record it. And so then they see all the bugs. And then I send the recording in the chat. Someone on our team, he's 19 years old, sent me a demo this morning off of that conversation that solved all of my problems. And he was like, hey, I was waiting for like three training runs to finish. So I had a little bit of time while I was waiting. So I implemented everything that you asked. And it blew my mind. It was so good. That's using AI correctly. So it's not a token leaderboard. It's what did you show in production that was good?
Speaker 2How many companies do you think are actually as token pilled? AI-centric as we think in terms of devs?
Speaker 1I think there's a guy, Jason Yeager, who used to work at Speechify. And now he has MyTech CEO on Instagram. He's super funny. And so he makes a lot of videos about like, you know, crazy CEOs who all use tokens, use tokens. I think all founders in some way have that animal inside of them because you know that it's the right path. But there is a difference between reality and theory. And you need to make sure that you don't overdo it. Do you have any price sensitivity on tokens? Yeah, of course. Absolutely. I mean, I'll lose my mind if to implement a tiny feature, you use 15,000 tokens. Like, why did you do that? And like, we will let people go if they just go bananas with something for no reason. Are you able to accurately budget tokens? Not accurately, but within bounds. The other thing is like a lot of engineers are, look, you go into engineering because you like optimization. Most engineers are not blind, and it physically hurts them to overspend tokens. And again, I always think that the best way to interact with AI is you are chatting in the chat or actually doing it verbally. And you're essentially pseudocoding with your words constantly, and you're explaining architecture. And a great example would be, I know someone who has no engineering background, and they wanted to build an app. And they built exactly what they want. It took them two hours. But they needed an API call, and they need to scrape this website. And they basically scraped every single page of the website, every single part of the website. And so the bill that they got for the scraping was gigantic. And then I was like, why are you doing like that? Why aren't you going into the database to this exact URL and then scraping that from the URL? So the amount of nodes they needed to hit became like 20 instead of 25,000. And so an engineer will spend their time doing that. And they're like, oh, I don't know. I don't know. I don't know. I don't know. I don't know. I don't know. I don't know. I don't know. I don't know. And so it takes a lot of time making sure that the thing is optimized like that. So that's how you build like a good database or a good architecture system, whatever. You do the same thing when you're interfacing with the agent. You want the agent to take the path of least resistance, not the path of most resistance. Totally get that and agree with you. I think one of the biggest problems is that agents are goal-seeking. And so they are like... It's all about the target. You need to be good at picking the right target. And I think Anthropic published this paper when Fable 1 came out about long horizon tasks with Fable. So the first thing is it was much better at like running... Running a two-week task. And it could burn $12,500 worth of tokens in two weeks and basically make a better model with that. That's a perfect, amazing way of using tokens. That's exactly what you want. And what you don't want is burning 12,000 tokens in the span of five hours doing something that's like totally unnecessary and doesn't make any sense. You need the loops to happen and then you need to check the result. So what you want to build, and Boris, who's the inventor of Cloud Code, talks about this all the time. It's all about the loop. You say, here is the target. Here's how you measure the target. Now iterate against the target over and over and over again until you get it.
Speaker 2What did you not know about building an AI-centric dev team that you wish you had known? How useful is it to own your own GPUs? What was that realization moment? Just did you see a build one there?
Speaker 1The realization moment was when we realized that we had really talented engineers who were essentially moving at one-seventh of the speed they could have if they had the computer. And we were able to get them to execute one-to-one with their creativity and ideas.
Speaker 2If you're a founder listening to this, how should I change my hiring process in a new AI world?
Speaker 1Number one, functional interviews. Build this and then you see if they can build the thing. And then you run it through unit tests. The second one is give them a large code base, even an open source repository, and have them understand the code base, make changes, and then check what they broke. And then, yeah, like they have to be able to orchestrate agents well. And if they're not doing that, it's kind of not worth to have the person. And then the next thing I'll say is, it is more fun to have a smaller team. Having a big team is great as long as everyone's carrying their weight. But the way I kind of think about it is, yes, I can have multiple agents running on my computer, or I can have several Slack chats with really smart people who are bigger domain experts than I am. And basically that human being is the outcome owner for that task. And they have the agents. And so I can run, as a founder, multiple projects at the same time to a really amazing level of granularity. And so I think about moments earlier in the year when my brother Tyler would literally have an alarm to wake up at three in the morning because he needed to check what the agent was doing at three in the morning. And then you wake up, make sure it's good, go back to sleep. You want to babysit your agent basically every three hours. And the beautiful thing now is you can go work out, and the agent will tell you the answer. And then you voice note back with Speechify what you want to do next, and that'll happen. And so you want people who are essentially that level of addicted. Obviously, that creates massive AI fatigue, so make sure your teams don't burn out. But you want someone who is that level of excited. And so I think hiring for Slope more than Intercept is more important today than ever before. Said another way, I look for the potential the person has more than I love it. So I think hiring for Slope more than Intercept is more important today than ever before. So I think hiring for Slope more than I love it. So I think hiring for Slope more than I love it. So I think hiring for Slope more than I love it. So I think hiring for Slope more than I love it. So I think hiring for Slope more than I love it. When I look at
Speaker 2Whisperflow and Willow, and I did this tweet, and I deleted it because I don't ever want to be socky and miserable. And it's an amazing thing to build a company. You should be incredibly credited for doing so as an entrepreneur. But I found Whisperflow's product was just getting worse. And I said it on Twitter just because I honestly just wanted alternatives. I really need this product. And I wanted alternatives. I got 500 different alternatives. And I was like, we'll talk about the commoditization of a market. That is not one that I want to be in. Can you help me understand? And we've seen the complete commoditization of that Whisperflow, Willow, speech to text for productivity.
Speaker 1What they came out to the market with first was not necessarily their own model. Part of the reason they got worse is they switched their own model because it's a lot more affordable. And so they had a harness that ties together a bunch of other things. Probably it was DeepL or DeepGram under the hood with a bunch of optimizations. Now they're trying to do notes and they're
Speaker 2trying to move more into productivity. And I think they're being successful with it. Yeah.
Speaker 1A hundred percent. Yeah. So that's to your point of the compound startup.
Speaker 2When you look at them, do you not reflect on your, we said before, not announcing fundraisers, not announcing anything. They've announced everything. They announced going to the bathroom. And hence they have, I would say, a bigger brand.
Speaker 1Not in terms of users. If you walk down the street in New York City, way more people will know Speechify than know Whisperflow just by virtue of the fact we have way more users. But in the tech world, way bigger brand, right? Investors know who Whisperflow is because they announce. We intentionally don't announce. But we do. We don't have any competitors. Who are you going to use instead of Speechify to detect speech for your models? The closest thing is 11 Labs and we're so much bigger than 11 Labs. We are unique in our market. So because Whisperflow was so public about it, they now have a lot of competition. And so Peter Thiel, only losers compete. Try to not compete. What happens to that market? Whisperflow take majority and then there's thousands of ankle biters? I don't know. I mean, I want them in that market, right? I think that market becomes oligopical as well. And this, by the way, I think it's another mistake that I made. I built my own speech to text, experience that I've been using on my computer for the last seven years. Side loaded on my iPhone and on my computer. But I figured it's a commoditized product, right? Apple's going to release it instead of the button. It'll be great. And there you go. But Apple keeps not doing it. If you remember two years ago, Apple announced a partnership with ChatGPT that will improve Siri. Nothing happened. And so that's also the reason why I never went after Siri. And so now we've launched a product to compete with Siri and we've launched a product to compete with Whisperflow and we launched a product to compete with 11 Labs because I learned a lesson that I should have learned before. Which is the same lesson from 11 Labs. The way that you win is you offer an excellent product for free and then you have a wedge and then you add more and more and more things. And so I don't know what happens with the Whisperflow space. I just know that if you're a founder, you should
Speaker 2also always try. Final one. Another one that I get in trouble for, but I stand by strongly is I just think the customer support market's a challenging market to really get behind. You have Sierra and Dacagon out in front with the majority of funding and attention. But to say that there are 18 companies that have now raised over a hundred million in the last 18 months, there is kind of what I call the mid-tier, which is your intercoms risks, and your crescendos, all these ones. Whereas they're not old, but they're old enough, eight to 10 years old, and they're pretty good. And then you've got Salesforce, Atlassian, and the much older ones. And then the worst thing about this market is that for any sophisticated buyer, an Airwallex, a Klarna, a Navan, a technology-facing company, everyone has built their own because they need a sophisticated- Of course, and why would you pay a tax for it?
Speaker 1What am I missing? Yeah. So the first thing you're missing is the core product that we're offering B2B is the API, not the agents. So Sierra doesn't have their own model team. They use other people's models because the value of Sierra is the go-to market. It's Brett Taylor. And so that's why if you talk to Monty and Piotrek, they'll tell you we're not competitive with Sierra because their main business historically has been the API. So that's the first thing. In the API business, you have Speechify, 11 Labs, Gemini, SpaceX is now in the race, and that's kind of it. And so that's not that competitive a space compared to the B2B customer support thing. Everybody's in that space, Finn, everybody. I'm not building that product. What can I offer you that's 10X better than the next person? Not much. And so in the core API side, I can offer you better quality, faster speed, and 10X cheaper. Good offering. But then I have to also offer agents because there are so many pockets of value that have not been unlocked. And unless I am, again, I have this model for leadership. You don't want to be a fat manager who's like a general sitting in the back saying, take that hill. You want to be the warrior who runs up with their sword and engages the entire team. You want to be the warrior who runs up with their sword and engages the entire team. You want to be the warrior who runs up with their sword and engages the entire team. You want to be the warrior who runs up with their sword and engages the entire team. You need to be the same thing with your product. You need to be the number one user of your B2C product, and you need to help your customers use your product better. And if you do that, you will learn their problems, and then you will figure out what the next product is that you need to offer them. So unless I have front deployed engineers working with my B2B customers, building agents for them using our technology, I will not figure out what the really amazing next innovation across the hill is. And so you mentioned the right thing, which is 11 Labs now has all these partnerships with governments. Governments is not exactly customer support. They would have never gotten to governments had they not done a great job on the private sector first. I agree. 11 Labs is, in addition to OpenAI, is the most integrated company right now, AI company with governments. That means they figured something out, but you got to start in something like customer support. Again, if you're a founder, you need to try. You cannot not try. You cannot give up before you're even in the race. What will be a bigger company in five years, Sierra or 11 Labs? I think they're both going to be massive. Give me one name. Brett Taylor has the best resume, I think, of anyone in the world. I think he started Google Maps. Then he was CTO of Meta. Then he was co-CEO of Salesforce. He's on the board of OpenAI. And now he founded Sierra. I would never try to fight Brett Taylor. And I think the field is so large. We don't understand how big the space for AI agents is, like AI voice agents, not even close, in the same way that people didn't understand how big the field was for LLMs in 2019, and the same way people didn't understand how big the space was for AI coding, agents in 2021. This is the next huge space. And so both those companies are going to be massive.
Speaker 2They're playing very different games. I think Brett Taylor's actually trying to recreate a next generation of Salesforce. He is absolutely not playing the customer support game. He's moving into pre-sales. He's moving post-sales.
Speaker 1Neither is 11 Labs. 11 Labs has a product that also does code, too. That's why I call it AI
Speaker 2agents, not customer support. 11 Labs is not a fan. A very opinionated, voice-centric company. Correct. It's voice-centric. Oh, yeah. Yeah. Yeah. I think Brett Taylor is doing all of it.
Speaker 1To put it another way, if you use a tool like Sierra, the wedge right now is voice, but the important part is tool calling. 11 Labs lets you do some tool calling, but that's not the bread and butter. There was a really good presentation that Brett Taylor did, a screen share of him building a guitar store on Shopify, and how he uses Sierra to do customer support and sales and everything else. It was extremely impressive. If you haven't searched this, you should search this. Brett Taylor is a big guitar guy. That is a very different product than what 11 Labs is doing. And so they're both going to crush. I agree with you on the Sierra conclusion. What today is a no, and in five years' time, we'll be like, yeah, of course. Human-computer interface is going to become primarily voice, as opposed to a screen. So part of the reason why Google succeeded, it is a very simple interface. There's a text box and a button. That's it. Anyone can learn how to use it. The reason why ChatGPT worked as opposed to GPT-3 is because it was also a very simple interface, just chat. There's a text box and a button, you get a response. That's it. The simpler version of that is just having a conversation. I say something, I hear something in response. If you use voice AI from ChatGPT right now, it sucks. It's too slow. The LLM is much dumber than the core LLM. The escalation to the higher quality LLM is pretty weak. I think what will happen, and Meta has the right idea, by the way, so go Chris Cox, is people are going to be talking to their computer and phone and some wearable constantly throughout the day and using screens a lot less. You can buy one SpaceX or Meta. Where did you buy? Why? Elon's distracted.
Speaker 2Is he distracted or is he building full stack? Because actually, I think he's never been more strategically positioned and he has an outlet for each of the different products that he's built and each one feeds the next. When you look at Zuck and Meta, bluntly, the compute spend that he's producing, the outlet is increased conversion on an ads business, which is the biggest ads business in the world. So 7% on $240 billion is a lot of fucking money. But it's actually not in the same quantum league as doing...
Speaker 1Space data centers. Yeah. So let's take the space data centers out for a second. I think the space data centers is a very interesting idea. And what it does really well is it lets me underwrite a gigantic TAM. It ruins all estimates. Right? And so that was a great rabbit out of the hat by Elon in order to pitch investors really well. Let's take that out for a second. And I'm going to talk to you about SpaceX and Tesla like they're one company, because really I'm assessing Elon. I'm not assessing everything. SpaceX is an individual stock. For data centers, the biggest constraint... The biggest constraint right now is memory cards, and then very soon it's going to be energy, and it's energy a lot of the times. So what do you need for energy? You need energy supply and you need energy storage. And so the best energy storage right now actually comes from Tesla. Tesla also has a chip manufacturer that they're doing, basically competing with everyone else. That's going to do really well. And if you saw that Joe Rogan interview with Elon maybe two years ago, he was explaining that the hard part is not building the product. The hard part is building the manufacturing for the physical product. So Elon is number one in the world for manufacturing complex items like that. That's very exciting. And so the TAM for Elon's companies are bigger. However, I think the MetaTrade's... What is MetaTrade at right now? Less than SpaceX. It's been less than SpaceX. It's fucking dumps. And so I think Meta has more data than anybody else in the world. I think Meta is actually super hampered by laws like GDPR. If GDPR didn't exist and the other laws in the US didn't exist, Meta would be ripping. They just can't train on their data properly. And so they'll figure that out at some point in some way. I don't know how, but I believe in Zuck. And at the end of the day, I'm a huge believer in founder-led companies. And so we're talking about two of the best founders in the world. And the last thing I'll say, look at Zuck's age and look at Elon's age. And Zuck's not going to stop and Elon's not going to stop. But at a certain point, one of them will expire. And so Zuck has like 20 extra years. And so depending on how long you're investing, I'm younger than Zuck. Let's see what happens. I think if Zuck expired... Meta's dead.
Speaker 2I think if Zuck expired, Meta's stock price would increase. What? I disagree completely. No, because you'd have a CEO who comes in and understands. And this may be a short term. Yeah. But they're saying we're going to invest more and more and more and more and more. More and more and more and more in CapEx. Yeah. When we don't have an outlet for it, you'd actually see a stock price appreciation in the short term. Every time Zuck steps out on the podium and says CapEx, CapEx, CapEx,
Speaker 1he's like fucking hammered for it. But that's why Meta is a good investment right now. Because what Meta doesn't have is what Palantir has, which Palantir has the Alex Carp effect. Alex is really good at pumping up the PE ratio of the stock. And Zuck, I agree, is the opposite. Same as Elon. It's the Elon premium. Same as Elon. Exactly.
Speaker 2If Elon were to be removed, he loses 70% of that value. Exactly. If Zuck is removed, you definitely don't lose 70%. You maybe lose, I don't think you lose anything. I think you get to experience Zuck in who says we're an outs business.
Speaker 1Charlie Munger and Warren Buffett, actually no, it's Benjamin Graham has this concept of the cigar bot, right? Like what's the intrinsic value of a company? And they approach it from an accounting perspective. I think about it as from an underlying technology and business perspective. So the underlying asset, the intrinsic value of Meta is so large in relation to how it's valued in the market today. And you're correct. What's the PE ratio of Meta? 32, something like that. And SpaceX is insane. Tesla is also in multiple hundreds. I think that there has to be a correction that happens unless Elon succeeds with a big, big, big, big, big vision. In which case, then he wins. What are you most excited by? I'm most excited by applications of AI to pharmacology and biology. I have a family member who has very severe autoimmune neural inflammation. He's had it for six years. I took a blood sample from him every week for 15 weeks, sent it to a lab. Sequenced his genome. Did proteomics on it to figure out how the proteins are expressing in his body. And run an RNA analysis in each one of those weeks. And then I compared that to self-reporting data on what the quality of life is and what his mood effect was. Every day I have like six years worth of data on him. And ran it on a GPU cluster. And I found so many things that no doctor could ever tell me. And he has a very rare disease. It was called like an orphan disease because there's not that many people. There's a Facebook group for this disease. I'm buying now basically like, you know, it's like a $5,000 device you can fit in your pocket. But if you put a piece of hair or saliva or blood into it, it can sequence your entire genome. And so I'm organizing meetups with all the people who have this disease to sequence all of their genomes and then compare them all on a gigantic GPU cluster to figure out what epigenetic common thread there is between them. And I know I'm going to solve this disease. I would have never had an edge to do that in the past. And like, it gets even more beautiful because I can then take all the conclusions that I have about it and put it into alpha fold from isomorphic. And I can design design, not just the protein that is creating these issues, but I can design molecule that needs to bind to that protein to either turn it on or off. I can use CRISPR to do the same thing. I can use a lab like Twist where I can tell it, I want you to make me this RNA sequence or this DNA sequence, and it can make it for me and ship it to my lab or my house. And I can create amazing outcomes with it. And I can simulate all of it on my computer that's SSH into my GPU cluster in Scottsdale, Arizona, and I can cure my brother. And so my experience is I'm a kid who, when I was eight years old, I couldn't learn how to read. And my dad had to open a book and read Harry Potter to me. And that's how I learned how to read. And then when I was 13, I moved to the United States of America and I didn't speak English. And I listened to Harry Potter audio books 22 times in a row and I still had the first chapter memorized. And then I couldn't get into the private high school that my brother went to. And then my sister went to, and I was really bummed. I went to a lower quality high school, whatever. And I didn't get into AP US history because I made a bunch of spelling mistakes in my essay and I couldn't read the passage in time. And I needed to train myself to read the SAT English portion. I wouldn't read the answers. And then I'd go and hunt for the answer. And then when I got to college, somehow, by the grace of God, I ended up going to Brown and starting a major for renewable energy engineering because I couldn't do literature. And I built a text-to-speech tool that would read out all my books to me. And that's why I graduated. Technology solved my dyslexia and it solved my ADHD and it's going to solve my brother's disease. And it's already solved my dad's prostate cancer because I figured out with a bunch of help from other people how to use GPUs to identify where in his body the lesion was. That's what I'm excited for, is there's better quality of life for literally everybody because you have this magical machine that can run a trillion operations per second on as many GPUs as you want and it can solve problems that we can't.
Speaker 2I find it staggering that still today, we have orphan diseases, which is like, oh, there's too few people to make it economically viable for us to try and solve. And there are hundreds, thousands, low thousands, but low thousands.
Speaker 1And again, it's the same thing. You just need data, you need compute, and you need to ask good questions. Like I said, 10 good decisions per day, either hypotheses or actual product decisions, and you can solve these problems. Freaking amazing.
Speaker 2Cliff, it's been so great to have you on the show. I much prefer it when it's a discussion. Talk to you soon. But before we leave you today, today I want to tell you about how the first AI law firm, Crosby, helped us close a big sponsor. As you know, some of the biggest companies in the world advertise on 20VC. My British dulcet tones clearly convert well. I was working to close this big sponsor and they wanted to get through legal review quite quickly to close the deal. Crosby turned redlines around in three hours and caught major issues that would have caused us serious problems in the future. Crosby combines AI, some of the best engineers in the world, from companies like Ramp and Stripe and some of the best attorneys in the world from top 10 law firms. Customers get the best of both worlds. An elite human attorney reviews every contract, but they move incredibly quickly, returning redlines in under four hours. They help the fastest growing companies like Cognition, Ramp, and Clay close deals in less than a year. Crosby is the first company to have a redline, and it's not just a matter of weeks. Learn more at Crosby.ai/20VC. If you want to redline NDAs, MSAs, DPAs, and any other procurement contracts faster, go to Crosby.ai/20VC. It's speed that you can really trust. While Crosby keeps your numbers sharp, OneMind keeps your customer conversation sharper. Our friends over at OneMind have a hot take. The B2B GTM model we've been using for, well, the last 20 years. It's called the B2B GTM, and we've been using it for, well, collapsing. Predictable revenue isn't so predictable, and buyers are just tired of explaining themselves at every handoff between SDRs, AEs, CSMs, and support. You feel it in your board reporting. Your sellers feel it in their coverage. Your buyers feel it as they wait for answers. Well, enter OneMind and their GTM superhumans. HubSpot, AlterX, and ZoomInfo are just three of the tech companies using OneMind's superhumans to qualify buyers, ride along with sales reps, and coach customers. Whether you're ahead of target, or under pressure to increase bookings ahead of hiring, or behind on target even, and need every rep to hit quota, OneMind is the solution. See for yourself at OneMind.com. That's OneMind, M-I-N-D, dot com. While OneMind helps convert more buyers, AlphaSense helps you spot what matters next. We used AlphaSense on an investment that helped us close an $8 million deal. That's a lot of money. That's why I'm genuinely excited to have them as a partner on 20VC. AlphaSense combines AI with one of the world's deepest libraries of market intelligence, including expert interviews, broker research, earnings calls, company filings, and real-time news. Every answer is grounded in this incredibly trusted evidence and fully traceable to the original source, which is so important. So you can make really high conviction decisions with confidence. But the best part? They're building super analysts, and always on AI analyst. So instead of starting your day with another search, you'll start with work you've already done, your coverage monitored, the important developments surfaced, and your investment brief already waiting for you. See for yourself. Head to AlphaSense.com slash 20VC. That's Alpha hyphen sense.

Podcast Summary

Key Points:

  1. Cliff Weitzman explains why Speechify spends tens of millions on NVIDIA GPUs, arguing that owning hardware is mathematically superior to renting due to the 1.5x annual rental-to-purchase cost ratio and the need for co-located memory for large-scale training.
  2. Weitzman admits his biggest strategic mistake was dismissing the B2B API market as commoditizable, which allowed Eleven Labs to leapfrog Speechify by building an API wedge and then expanding into agents and enterprise.
  3. Weitzman advocates for a "compound startup" approach where companies must continuously innovate beyond their initial wedge, and he emphasizes that hiring should prioritize technical aptitude and learning speed over existing domain expertise.
  4. Weitzman describes how Speechify engineers orchestrate AI agents for long-horizon tasks, with credit only given when features ship to production, and he stresses the importance of owning compute to unlock team velocity.
  5. Weitzman believes voice will become the primary human-computer interface, and he is personally excited about applying AI, genomics, and GPU compute to solve rare diseases, inspired by his own dyslexia and his brother's autoimmune condition.

Summary:

Cliff Weitzman, founder and CEO of Speechify, joins Harry Stebbings for a wide-ranging discussion on AI infrastructure, competitive strategy, and the future of voice technology. Weitzman details why Speechify invests heavily in purchasing NVIDIA GPUs rather than renting, explaining that the economics favor ownership: renting an H100 for a year costs 1.5x the purchase price, and owning hardware enables co-located memory for large-scale training runs. He also notes that older GPUs remain useful for inference, and NVIDIA's recent underwriting deals with major banks are creating a liquid secondary market.

Weitzman candidly reflects on his biggest strategic mistake: dismissing the B2B API market as commoditizable, which allowed Eleven Labs to leapfrog Speechify. He now advocates for a "compound startup" approach where companies continuously innovate beyond their initial wedge, offering products for free or cheap to gain adoption before expanding into adjacent offerings. He defends Speechify's move into B2B despite competition from Sierra and Eleven Labs, arguing that the market is oligopolistic and that being in the race is essential.

On hiring and engineering, Weitzman emphasizes technical aptitude and learning speed over existing expertise, noting that AI agents amplify individual productivity. He describes how Speechify engineers orchestrate agents for long-horizon tasks, with credit only given when features ship to production. He also discusses the importance of owning compute to unlock team velocity, comparing it to giving Michael Jordan a basketball hoop in his house.

Finally, Weitzman shares his excitement about applying AI to biology and rare diseases, describing his personal efforts to sequence genomes and run analyses on GPU clusters to solve his brother's autoimmune condition. He believes voice will become the dominant human-computer interface, and he remains optimistic about founder-led companies like Meta and SpaceX despite their challenges.

FAQs

Buying GPUs is significantly cheaper long-term than renting, and owning hardware gives Speechify's engineers unrestricted access to train models without worrying about per-hour costs. Additionally, large-scale training requires co-located GPU clusters with shared memory, which is difficult to achieve through typical cloud rentals.

Receiving GPUs early allows Speechify to skip the queue and gain months of access to newer hardware before competitors, which is a major competitive advantage. The cost is offset by the value of having leading-edge compute for training and inference sooner.

Cliff initially dismissed the B2B API market as commoditized and chose not to pursue it, which allowed Eleven Labs to leapfrog Speechify in the enterprise space. He now believes the key lesson is to continuously innovate and get users on your product, even if it starts free, then expand into adjacent offerings.

Speechify offers its Simba 3.2 model at $10 per million characters, significantly undercutting competitors like Eleven Labs and OpenAI on price while claiming top quality. Cliff argues the market is oligarchical rather than winner-take-all, and Speechify leverages its massive consumer user base and AI research capabilities to remain competitive.

Speechify now prioritizes raw technical intelligence, hunger, and work ethic over existing coding expertise, because AI tools allow smart people to become highly productive in about six months. The company values candidates who can orchestrate agents effectively and ship features to production.

At Speechify, engineers only get credit when their work is live in production, used by real users, and free of major bugs. Cliff compares it to delivering milk to the door rather than leaving it on the road, emphasizing that applied AI wins over theoretical work.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.