Go back

Wispr Flow CEO Tanay Kothari - voice AI deep dive

35m 41s

Wispr Flow CEO Tanay Kothari - voice AI deep dive

Whisper Flow, a voice-first AI assistant, aims to make keyboards optional by converting speech into contextually appropriate written text, unlike traditional dictation tools that transcribe verbatim. Co-founder and CEO Teneca Thari explains that while platforms like Apple or Google focus on word-for-word transcription, Whisper Flow optimizes for a "zero-edit rate"—producing ready-to-send content tailored to the medium, whether casual messages or professional emails. This approach yields an 85% perfection rate, far exceeding the ~10% of major tech companies. Users often shift from 20% to 80% voice usage within months, applying it to tasks ranging from emails to coding. Challenges include iOS limitations, with Android offering better integration prospects, and accent diversity, as many users are non-native English speakers. While voice isn't suited for everything (e.g., some prefer typing for journaling), it's gaining traction in workplaces, with discreet microphones enabling productivity in shared offices. The discussion highlights voice AI's potential to transform communication, especially for ESL speakers and technical professionals.

Transcription

6545 Words, 35015 Characters

English
Hello fellow data nerds, welcome to World of Dazs. I'm your host, Orhan Hoffman, C of Incubate and GP of Flex Capital. Discover more episodes, get weekly data as the service news, original content, articles on data, and more at worldofdazs.com. That's worldofdazdaas.com. Hello fellow data nerds, my guest today is Teneca Thari. Today is the co-founder and CEO of Whisper Flow, a voice-first AI assistant aiming to make the keyboard optional. By the way, I use it. I love the product. Whisper Flow has been growing at user base by 50% of a month over a month, and they've raised $30 million series A earlier this year. Today, welcome to World of Dazs. Orhan, thanks for having me. I've been looking forward to this. Yeah, okay. I'm a big user, by the way, of the product. I have a new user actually, so I'm excited to talk to you. I'm just interested in dictation in general. Apple's got native dictation, Google's got native dictation, Microsoft's native dictation. They don't really work that well. Why haven't they nailed it? I think at the end of the day, what we saw happen was I honestly wish they did, but the problem that they were trying to solve is very different than the problem that people wanted solved. So every single dictation to until today, what it tries to do is it takes what you say, and it tries to write it down word-for-word. And while you're saying it too, it doesn't like wait for the context and stuff, right? While you're saying it. So one, you're distracted reading the same thing that you said half a second ago. Two, the way you speak is very different than the way you write. You don't actually want your rambles on a piece of paper. What you want is something that is ready to send that is in your written tone for the purpose you're in. Yeah, if you're like, you don't want those in there either too, right? No, not at all. If you're sending an i-message, you want it to sound casual. If you're sending an email, you want a nice, structured, professional sounding. And the way you speak is the same in all of these cases. It's just the writing that is different. So that is the first insight that we had. The second insight that we had is what people want from these systems is perfection. If you dictate something with city, and in every sentence it's going to make one mistake, the time you spend fixing that mistake just takes away all the productivity boost that voice gave you. What we aim to do with whisper flow is, hey, the metric we want to optimize for is what we call zero-editrate. What percentage of your messages are perfect? Apple, Google, Microsoft, all of these large players. I read about 10% zero-editrate, which means 10% of their dictations are perfect and ready to go. Whisper float today is 85%. Yeah, and you probably even need to get into the 90s to make sure it gets used all the time. There's probably some sort of number where you need to get to where people will use it all the time. Even the other night I came from an event where they had really good sake, the Japanese wine. And so I was sending a text to someone saying, hey, thanks for bringing this amazing sake. And it kept messing up. This is in whisper flow. This was Apple kept messing up the sake. And there was just, I did like four takes. And finally, I'm like, okay, I just went to whisper flow and did it and worked great. But it's just like those types of things. Maybe it's not like the most obvious thing that someone's going to talk about. Maybe sake isn't a common thing people talk about. Most people aren't always having obvious conversations. No, almost no one is. That's why you haven't ever heard anybody say like, hey, I love city or I love Google's voice because it never gave you that aha moment. It's good if you have a very specific instruction. If you're like, hey, Alexa, play this song by Taylor Swift, it will generally find it because it knows, okay, you're saying play, which is good code. And then, okay, you usually talk to me about music. And therefore, if the surface area is really big, like it is with text messages and emails and those types of things, it really starts to break down really fast. Oh, 100%. Even when they're doing translation on a voicemail. So when you get a translation, a voice to text on Apple or Google, those are also laughably wrong. And those they have plenty of time. They're not doing it. We'd real time. They can take it. They have plenty of time. They can then do the thing. And then they could send it to you. Is it just they're not putting compute behind it or why is that so bad? That is the second biggest insight that we have. That is a mistake people are making today, which is when they're using voice, they're just like, hey, here's the voicemail, let me transcribe it. So they just use the audio. At the end of the day, even when you and I are speaking, just like humans understanding audio, right? There are so many different ways you can kind of hear the same thing. But what matters is the context. And there's been so many studies on this that if you have context and humans understand perfectly what you're trying to do. And the same thing applies to models. So if in the voicemail, it understood a bit about who sending it, who it's sending to, the voicemails always get my name wrong. And Apple should just be like, no, this dude's name is Tane. That is likely what they're saying and not some other random word. Right, because they're calling you. Of course, the person calling you might get your name wrong because it's not a common name. People get my name wrong all the time. Yep. And so that is the second thing that we added in. And I think going forward in the future, that is just going to be common across all voice platforms that are super excited to be the company that is driving that change. No, is there a way to get one of the problems is it's just hard to get natively in to iOS or into your pixel or something like that. So on my laptop, if I want to use a voice service like whisper, it's pretty easy. You can just hit a key stroke and you can do it where on your phone, you've got to go to a new keyboard or something. So it's like three buttons to get there often. And obviously, it's not an Apple or Google's best interest to get another voice system in there. So how do you circumvent that in a way to make it easier for the user long term? So most other voice keyboards are five, six taps. We made whisper down to okay, on frequent usage. It's just one tap to get it to work. If we want it to just be that consistently, the best experience, then all right, you got to message them and do something about it. Android is going to be a phenomenal experience. I don't think Kim Cook listen to this podcast, but Sundar might be. So we can at least talk to him through there. Android is going to be fantastic because iOS is just heavily limited. They don't let you do a lot. It's just every limited. Okay. And it's going to work like a charm. So we're going to be shipping that end of this year. Oh, awesome. Okay. Cool. Hopefully that gives Apple the impetus. Yeah, but there's all these reasons for me to switch to an Android phone. So it's very possible by next year. I might be switching for a lot of reasons because the Apple phone is just getting stupider and stupider compared to the Android phone just because it's not AI enabled. Oh, and we're going to have so many people in the comments. Totally. Totally. I mean, look, I love Apple. I use it today, but cheese they're falling behind fast. I find myself using voice way more today than even a few months ago. How are you using it? What do you do is doing in voice? You're probably in the top one percent of voice users. I assume. So how are you using voice that maybe the average person isn't? I'll tell you something shocking. I was looking through our user base and I personally am not even in the top one percent. Okay. Because of how deep voice penetration has become because of whisper flow. So I'll give you one stat. When people start using whisper, the first month, about 20% of the work on their laptop is done with whisper flow. 80% is their keyboard. How do you know? You just estimate it basically. Do we just see the number of keystrokes? Number of key presses that you do, we see that number of times they use flow. We see that and so we can put down. Hey, what is the flow versus keyboard look like? Four months in, it is 80% flow, 20% keyboard. And that was shocking to me. And this is across every single flow user ever. Yeah, I'm not anywhere close to that today. You'll get that in a couple of months, but what we see is people slowly build that habit of voice. And this is the first time in human history people have stopped using the keyboard all together. And that leads to the question that you were asking is what do you use it for? Well, when people first start, they use it for one thing that they do often. They might use it for all their AI prompts or all their emails, all Slack messages, all documents and journals and notes that they're writing. And by the way, the thing is like when I do an AI prompt in chat GPT or whatever, it actually does do the voice transcription pretty well. If I talk to chat GPT, that doesn't have a problem actually. It usually does it quite well. It's like all the other things don't work. Yep, you would be surprised. Chat GPT is actually the most popular tool for people use whisper flow. Okay, just because they're already in it. Even though chat GPT literally has a mic button embedded in the product, but people love whisper more than they do open AI's own tool. And so they just use whisper inside of it. But coming back to the thing, what we see happen a week or two later is people hit this realization of, wait, why don't I use this for everything? And then what you see is the number of applications they use whisper and blows up. The media number of applications per person is 69 where they use whisper. So they're doing it everything from sending a text to their mom to the two word Google search to writing long messages. Couple of people have written entire books with whisper. And so it is genuinely replacing everything that she would do with her keyboard otherwise. And it is bitter when I'm by myself. Like if in my office by myself or whatever, when I'm in office with other people or when I'm on a airplane or something like that or on a train or whatever, I don't feel as comfortable. Maybe that's just me. I don't feel as comfortable talking even whispering in the thing. It just seems weird. So I'm more typing words. I'm by myself. I'm like talking all the time or people having similar experiences. A lot of people start off like that. And so what we've seen is a lot of our users are in open offices. And so about 50% of them, what they do is they have a room where they go to take zoom meetings. We all do this all day long. So they just have a tour block in their day where they go to this room and they get all their work done with whisper flow. Yeah, yeah, yeah. What I've started to see and this is absolutely incredible. We're seeing actual workplace transformation is people are in their offices that having these mics that come up with their mouth. Now, this is a fancy multi hundred dollar mic. People have these 10 dollar mics, these podium mics that come up with their mouth. They're silently whispering to it. And there are some offices where the entire office everybody has this mic that all just was put into their computers. It looks a little surreal. Yeah, but you see the productivity levels shoot up like nothing else. And it starts with one team using it. And initially, people are like, oh, what are these people doing? And then the boost is so visible that the rest of the office is like, okay, yeah, we want to do whatever they're doing. We want to try. Yeah. And we've seen this happen across Fortune 500 companies and the startup and beyond what I was thinking, I'll write a blog about how we are seeing this change happen. It's interesting because for me, and I think for many people, I can input things very quickly with voice quicker than I can with typing often, but I still prefer to read. I can read much faster than listening. And so the text is very helpful. It's not like I want the two-way conversation. I don't need it to read it back to me. I can rock it. I can read much, much, much faster than maybe even three X faster than I can listen. So it's very interesting how it's like a one-way input, not a two-way input. Yes. That was very intentional. I think a lot of people ask us if we're ever going to do text-to-speech. And the usual answer is people actually don't want that. It's really annoying to listen to something very slowly when you could just read it like that. It also is like sometimes you realize just why you're reading it. You don't need to read all of it. You read the first sentence of the paragraph. I can move to the next paragraph. Whereas if you're listening, you don't have the ability to scan the document and move ahead. It's just not always that relevant that comes there. I find when I'm in the car often, I'm talking, like let's say chat, GPT or GROC or some of these things in the car. And it's just taking so long to have this learning conversation. Whereas if I'm not, I have no choice but to listen when I'm in the car. Whereas when I'm engaged, I can learn so much faster because have to stuff their time already now and I can move around or I don't need to know that or it's just flopped for whatever it is. Yeah. No, that's bang on point. We're low-dass isn't just a podcast. We're low-dass membership is a private invite-only community for founders and executives of important data and AI businesses. It's a place for off-the-record, high-trust conversations focused on the biggest questions data leaders are facing. The community is about peer learning, curated intros, and access to private events. Everyone in the group is building or leading something meaningful in data. If that sounds like you, apply to [email protected]. That's worldofdaas.com. Now, what about accents? What I found, okay, and this is my very small thing, is these things work super well with certain accents and terrible with other accents. I found even with strong accents from India, it works pretty well, but strong French accents, it can't handle at all. Are you talking about Whisper Flow? I haven't tried Whisper Flow with a strong French accent because I don't speak with a strong French accent. Yeah, I can probably believe you. But I don't really know how to do it, but I just found certain accents, these things generally tend to work better, like if I use granola or some of these other things, they tend to work better depending on the accents that are. I don't know if you've come up with some of those things yourself. That was actually a really hard problem to solve. So what you'll see, and I guess if you don't have a strong accent, you wouldn't, but what a lot of other people see is if they speak English with a strong Russian accent, it writes it in Russian. And the models just do this for you. And it translates the whole thing into Russian, and then you're just like, what? It knows it somehow figures out that you're talking in Russian, and it literally puts it out in Cyrillic or something. It does that, but you're actually speaking English in a Russian accent. Yeah. And it writes all of that in Cyrillic, and just transays the whole thing in Russian. Oh my gosh, that's crazy. I didn't know that. Okay. Yeah, this happens all the time with voice users that have thick accents. And so that was a key problem that we had to solve. Oh my gosh, I had no idea. That's so cool. Yeah. It's cool to first and then review and knowing after that. Yeah, I'm really knowing. 60% of a user base is actually outside of the US. 70% of our users speak a language other than English. That was one of the things that we had to solve. And I think actually whisper flow is the first voice AI tool that solves that. You can have a thick accent in any language, and it just gets you still. One of my friends is mom lives in the US, but she's originally from India, her first language is Hindi. She is pretty thick. She moved here when she was 35 or something. She's pretty thick accent. And when she now is communicating internally in her work, she now talks in Hindi, and then it automatically comes in. And then it also translates it to English for her, because her writing is good in English, but not amazing. And then she then can like send it on to her colleagues and stuff like that. And she says it's like a game changer for her in the office. Yeah. No, definitely. I think that is there are so many problems that ESL or English has a second language group face when they're working. And this is not just in the US. People are all over the world. English is the international language. English is the language of work. So if you're doing international business, you're using English. Yeah. And there's a lot of support that they need to be able to feel confident in their writing, their communication, because it's all important. So there's a lot that we're also going to be building there to make lives better for ESL speakers, because that's actually majority of the world population. Yeah, there's more English speakers. She's speaking this as a second language to some people who speak as their first language. It's really the only language like that that exists really. What do you think the use cases are that will never move to voice or maybe not in the next five years? The only one that I have seen there to be still people prefering other modes is journaling. Journaling. Oh, there's something about journaling that people find therapeutic when they're writing on a piece of paper. Are they actually writing or, for example, physically writing? I would have thought journaling would definitely go to voice pretty quickly. The depends person by person. Some people, if they're just speaking out loud, they just ramble. And so they need something to ground them, which either they're writing physically in a notebook or they're like opening a document and then writing on their keyboard. That I've seen for some people is just they can use other tools, but it's never going to work as well, because that's just how their brain works. What about things that are super technical when I'm coding? I don't feel like I want to use voice there. I like using the AI assist to help me or I can imagine some sort of AI assist for lawyers. I think that might be great, but I don't know if I want to do that by voice. I probably want to use it text and then it pops up and gives me some options. So I want to do that. It seems like those maybe move a little slower to voice or do you disagree? He would be shocked. Okay. I was shocked. I'm just developers using it now voice. Developers are starting to become one of our biggest segment. And they're using it in closer. They're using it for every single thing. People aren't writing code by hand that much at all at this point. And same for lawyers. Lawyers is also one of our big, big segments. Okay. Interesting. All right. They're dictating contracts with Whisper because they're like, this is just fast-straining exactly what I want to say. And Whisper just writes in their self. But those two things I was not expecting at all. Okay. I know when you guys originally started out, you're building a wearable device, inhibited, but think of the wearable world. Where do you think the AI wearable world is going? There have been so many different instantiations offered that we've seen in the last couple of years. There was the humane pin, which you could just attach on your jacket or your shirt. That was the rabbit R1. This is little orange neon thing that you could have in your hand. There are companies like EO. That was just like an in-ear, voice-in, voice-out system. And then you have the smart glasses, which are a lot better, which I love. I use the medical glasses and I love it. I think it's great. Doesn't do that much, but what it does, it does really well. I'm excited for when it has displays. And so the point we were getting to earlier, which is voice as an output is I think one of the most seamless ways for humans to interact with technology. However, displays as an input is the fastest, highest bandwidth way in which we absorb information. And so to me, just like that, first principle is thinking, if you're building a core piece of technology that is going to be with people that is their primary personal computing device, it needs to have a display and it needs to have really solid voice input. Those two are non-negotiables. Correct. Yeah. Now, when you look at that stuff that comes top of mind is number one, AR glasses with a display in them, similar to the new MetaRay bands, that I'm extremely excited about. It's also really hard to be able to get a tiny display. Super hard to replace what your phone does. It's possible, but it's hard. Are you buying that one? The one with the display? I buy every single piece of hardware to test out to play with. Yeah, to test out, of course, yeah. And the second thing is again, smart watches. If you remember the movie, her. Yeah. I think that showed it perfectly where you had an in-ear device. So it was vice-first throughout. And then you had a little display that you could pull out time and again. Yeah. The biggest change that happens that I think most people are not talking about is displays become secondary. And we have been in a GUI first world for the last 25 years ever since MSDOS. This is something that is going to be one of the biggest technical revolutions that we'll see is the death of the GUI. And that is going to lead to software being built differently, experience is being built differently. And a new generation of hardware that just adopts that. Because even in a car you can imagine, well, just tell me where you want to go. Okay, I want to go to Mary's house. Great. Okay, we're going to Aunt Mary's house. Just sit back. We'll take you there. Give us a sense of the type of music you want to listen to or whatever. And then you don't have to know what's going on. You don't have to know how fast it's going. All the dials on the car you normally need to look at. You just could imagine. All those get abstracted away. If you want to look at them, sure, but like you don't have to look at those. Just like when I ride in the back of a car of Uber, I don't look at anything today. So the driver might be, but I'm not looking at it as types of things. I just tell the driver where I want to go and maybe what kind of music I want to listen to. And then we're off. Yeah. That is a really good analogy actually. I'm going to use that. Oh, perfect. All right. Great. Great. It's free. Don't need to give me any stock or anything for it. When you guys made the pivot originally from hardware, I think you had 40 employees. You went down to like five overnight. I imagine this was a very trying time. Walk me through that time a little bit. When we started Whisper, we had the vision that, okay, in the next few years, everybody's going to be using voice. When that happens, you want to use it when you're around people. This is going back to the conversation we were having 10, 15 minutes ago. And you want to not disturb them, but you also want privacy. You don't want to feel conscious. And so we wanted to build this wearable device that could understand what you're speaking silently. So it went from your thoughts to text and let you do everything you want it. And so we put together this incredible team of some of the best engineers and scientists in the world. PhDs across machine learning, neuroscience, signal processing, electrical mechanical. And after working on it for three years from 2021 to 24, we actually had this thing work. It is, I think, still until today, the most magical piece of technology I've ever tried. And once it worked, we tried connecting it with Siri, chat GPT, Alexa, and they all suck. We wanted something that went from your mental rambles to something that was structured and usable. And we called this little project flow, the flow OS. And we ran it on this hardware for good. Then I packaged it up because I wanted to give it to some friends who didn't have the hardware device. So I packaged it into a desktop app, sent it to them. That two friends became five, became 10, became 100, became a thousand very quickly. And what hit us like a brick in the face was the realization that people still today don't have the behavior of voice. We're giving people a device that lets them use voice better. If they're not using it already, you ought to go, no one was using voice. This company is going to tank. The first thing we need to do is build a behavior of voice. Now again, I wish somebody else had built that so you could keep working on the hardware. But the way I think about the world is like, hey, this is the state of the world today. This is what I wanted to be. I wanted to be in a place where interacting with devices feel just as effortless as talking to a close friend. Technology just gets you. So with you 24/7, it lets you be present, doesn't pull you into screens. And when I think about a world like that, the way I think about what we do with the company is whatever is needed in the world to get people one step closer to that. In mid 2024, that answer was we need to build voices and interface that people can use. That was the hard decision that made us switch over from building a hardware device to the software device. Now, majority of our team were neuroscientists. We're collecting data for our brain computer interface. We had a massive hardware team that was building the actual hardware device. And with this new product, there was none of that. And so it was insanely hard at that time. Think for two months, it was probably the lowest low for me. And it was really tough for a lot of people. And that was what I was feeling secondarily as well. Because I loved all these guys. I'm still very close to a lot of them. But we were doing this decision not because we were out of money. We had more than half the capital we'd raised. Now because they wasn't doing well, we were ahead of schedule, which is crazy for a deep tech company. Team was incredible. It was starting to work. But it was not the right thing to build. You can't do both. You have to focus. You have to put all your eggs in one basket. Getting one thing to product market fit is hard enough. And so decided to make the hard decision to move over here. And looking back a year later, because this is all happening August of last year. This is still fresh. Yeah. Sorry, we're ripping off the band-aid area. Trust me, I've ripped it off myself hundreds of times before. But looking back, I think it was one of the best decisions we made. Because we are closer now to I think where we want the world to be. That decision was the actual decision that was ripping off the band-aid. I'm obviously these are talented people. Can I fund some of them? I'm sure that I'm going to start interesting things and things like that. Are they out there? Like any of them starting companies? A couple of them actually have. A few of them are leading the AI teams, their chief scientists that a number of other companies were talking like it's sleep and aura and a number of other new neurotech companies. People are building metaria band glasses. Other people are at neural link. And so it's crazy to see where people have gone. The thing that makes me happy is I think every single one of them found a new home. And we helped a lot of people through that. Because again, this is not their fault. I tell people today shouldn't be a problem. That shouldn't be a problem in today's market if you're talented. Now I know that you're big on overall optimization, health, productivity, etc. What are some of the counter intuitive things you do to optimize productivity? The best advice that I have gotten is you're going to let some fires burn. When you're living life, when you're doing anything hard that has gone like slots and moving pieces. There's a lot of fires that come up all the time. And if you try to go and try to fight all those fires and solve all of those problems yourself or even as a team, you are all going to get burnt out. And also your context switching a lot. There's so much that happens because of that. I realize, hey, there's some fires that are important that are existential to the company. There's some others that if we just let them burn, it'll be fine. And just actively saying that here, let some fires burn. The moment, whenever I see the people around the table or anxious, something's bothering them, just saying that one sentence puts people at ease. Okay, they're like, all right. Yeah, we understand we can't do everything. Everything. It's okay. And you just see them relax and sit back in their seats. And having the mentality of when you're not anxious, when you can think clearly and calmly and you're composed, you're just able to do so much more clarity of thought as I think one of the most underrated productivity hacks. And anything that gets you there is in my opinion way better than literally everything else that I've tried heard or any amount of lion's mane or gaffey in that you can take. All right, a few personal questions. So on your bio, it says your Forbes 30 under 30, you might be one of the only ones that haven't yet been indicted in that. How do you feel about that given all the other kind of like famous people that are there? It's like a badge of honor. And also a worry badge too. Yeah, because I've gotten so many wrong. It is funny. So in session, I was out in the company. The first thing that we did before we even decided to do all the trigger is we wanted to align on values, things that are the highest priority for us as individuals. And the number one thing is integrity. I think as a person integrity is all you have. And if you lose it once, it's gone forever. And that has become just core to how me and him operate, core to the kind of people we bring on in the company, both as employees and investors, doing things like the people who've gotten indicted. It doesn't worry me because that is so far away from the things that I actually care about. Last question we ask all for gas. What conventional wisdom or advice do you think is generally bad advice? When you're building a new product, give it to your users for free. A lot of people that are like, hey, first see people are actually using it. And you can charge them later. You can get a lot of people scale the product. And then you can always charge people later. What it does is you never genuinely learn if you're actually adding value. Because if you're building something that adds value to people's lives and it's actually solving a problem, they will transact back. And that is the strongest signal that you as a founder can use for deciding if you're building the right thing or not. Why would that give me the signal? If someone bought Whisperflow, but only used it once a quarter and you gave it away for free and someone was using it seven times a day or twenty times a day or something like that, wouldn't you get more signal from the latter than the former? Do you know by usage it's hard enough just to get some product market value just so you learn like, okay, whether they pay for it or not. It's like sometimes easier to get something to pay than get something to use it. In some cases, maybe if these two examples are the only things we're comparing, of course the person who uses it gives you more signal. But from what I've seen, the person who pays and then doesn't use it, it's very small minority of the people who would eventually. Yeah, I mean, sometimes when you pay, you almost like force yourself to use it. You force yourself to use it or you have used it enough where you want to pay. And so even in the product you see it right now, when you sign up, no credit card. We actually give you two weeks of the full version for free and then you get dropped down to the premium one and even then, we don't bug you with KPF for this product pay for this product. Nothing. Just keep using it. We're happy, you're happy. But once you exceed the limits, once you're using it a lot. Enough. Yeah, Dan, you jump in. Then the product's like, hey, now you got to actually pay. Yeah, yeah, okay. So you are doing a free product. You're completely free. And then once people use it enough, you ask them to pay, you're almost going back to that thing. You are, but then you also block people's usage when they are hitting that. But wouldn't all premium products be that way? Exactly what you said. Is that the whole goal of a free product is to get people to use it enough so that they pay? There's two nuances here that I would say are important. If I was now now going down into the actual tactical details of it, what percentage of your engaged users are actually paying you? That is important. If, for example, that's 1%, you got to really go back to the drawing board. And see, like, why do I have millions of users? And only a thousand or a couple thousand people pay me money. If you added about three to four percent, it's good. It's decent. That's what you see for a lot of premium products. With Whisper, we actually saw that number be 20%. 20% of our monthly active users pay us. And that, again, we're talking about strength of signal. And so I generally think there's very standard to do this thing. This is perfect, but it all gets into the nuances. So this is the level of signal that you're getting. And that to us was like, okay, wow, this is extremely rare to see this kind of traction. And that is something really special here, enough to convince us to double down, to leave everything else, to just focus on this and drive a lot more conviction. Worst is if it was 1%, maybe you wouldn't be here talking today. I was been amazing. Thank you to Nekathari for joining us in World of Das. I follow you at TANCOT's T-A-N-K-O-T-S on X. I definitely encourage our listeners to engage you there. This has been a ton of fun. Super interesting. I love your product. Keep building it. This has been great. I'm not even an investor or anything. I'm just showing your product just purely out of love. I wish I was an investor, purely out of love of the product. So this has been awesome. Really great chatting with you. All right, and thanks a lot for having me. This is fun. Thanks for listening. If you haven't already done so, please subscribe to World of Das wherever you consume your podcast, YouTube, Apple, Spotify, and more. And please help us get discovered by leaving a review. And check out worldofdas.com. That's worldofdas.das.com. And of course, connect with me on Twitter @orin. That's A-U-R-E-N. Would absolutely love to hear from you.

Podcast Summary

Key Points:

  1. Whisper Flow is a voice-first AI assistant that focuses on converting speech into polished, context-aware text tailored to the communication medium (e.g., casual for messages, professional for emails), rather than verbatim transcription.
  2. It achieves an 85% zero-edit rate (percentage of messages perfect upon dictation), significantly higher than the ~10% rate of major tech companies, by leveraging context and user-specific data to improve accuracy.
  3. The platform sees high user adoption, with many users transitioning from 20% voice usage initially to 80% within months, applying it across diverse tasks like emails, AI prompts, and even coding or legal document drafting.
  4. Challenges include integration barriers on iOS, which Android may better support, and handling diverse accents and non-native English speakers, a key focus given most users are outside the U.S.
  5. Voice is not ideal for all scenarios (e.g., some prefer typing for journaling), but adoption is growing in workplaces, with users employing discreet microphones for productivity gains in shared spaces.

Summary:

Whisper Flow, a voice-first AI assistant, aims to make keyboards optional by converting speech into contextually appropriate written text, unlike traditional dictation tools that transcribe verbatim. Co-founder and CEO Teneca Thari explains that while platforms like Apple or Google focus on word-for-word transcription, Whisper Flow optimizes for a "zero-edit rate"—producing ready-to-send content tailored to the medium, whether casual messages or professional emails. This approach yields an 85% perfection rate, far exceeding the ~10% of major tech companies.

Users often shift from 20% to 80% voice usage within months, applying it to tasks ranging from emails to coding. Challenges include iOS limitations, with Android offering better integration prospects, and accent diversity, as many users are non-native English speakers. , some prefer typing for journaling), it's gaining traction in workplaces, with discreet microphones enabling productivity in shared offices.

The discussion highlights voice AI's potential to transform communication, especially for ESL speakers and technical professionals.

FAQs

Whisper Flow is a voice-first AI assistant that aims to make the keyboard optional by providing accurate dictation that adapts to the user's writing tone and context, unlike traditional dictation tools that transcribe word-for-word without considering how people actually write.

Whisper Flow focuses on achieving a high 'zero-edit rate' by understanding context and adapting tone for different purposes (e.g., casual messages vs. professional emails), whereas native tools often transcribe verbatim and have lower accuracy, leading to frequent corrections.

The 'zero-edit rate' is the percentage of dictations that are perfect and ready to send without edits. Whisper Flow achieves about 85%, compared to around 10% for major native dictation tools like Apple, Google, and Microsoft.

Whisper Flow is designed to accurately transcribe English spoken with thick accents and supports multiple languages, addressing common issues like misinterpreting accented English as another language, which benefits its diverse user base, including many ESL speakers.

Users employ Whisper Flow for a wide range of tasks, including sending messages, writing emails, generating AI prompts, creating documents, and even coding or legal work, with many eventually replacing most keyboard use with voice input.

On mobile, Whisper Flow aims for one-tap access, though iOS limitations make integration harder compared to Android, where a smoother experience is expected. The goal is to minimize steps for users despite platform restrictions.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.