Go back

Voice AI’s Big Moment: Why Everything Is Changing Now (ft. Neil Zeghidour, Gradium AI)

82m 49s

Voice AI’s Big Moment: Why Everything Is Changing Now (ft. Neil Zeghidour, Gradium AI)

The discussion highlights that voice AI is having a breakthrough moment after years of lagging behind other AI modalities. Recent progress in reducing latency and improving naturalness has made AI voice interactions more convenient than ever, even rivaling human conversation in some cases. Historically, the field attracted fewer visionaries and less prestige, despite early deep learning successes in speech recognition. This has resulted in a very small global pool of experts, as the discipline requires a rare combination of skills across machine learning, signal processing, and psychoacoustics. The future direction involves developing full-duplex systems that allow natural, overlapping conversations and incorporating emotional intelligence. Voice is also poised to become the main interface for new, screen-less hardware. However, challenges remain, such as enabling AI to function in noisy real-world environments and navigating social norms around speaking to devices in shared spaces like offices.

Transcription

13894 Words, 75092 Characters

English
For the first time, it's actually can be enjoyable and even more convenient to talk to an AI on the phone than talking to a human. I don't want to be mean to my people, the speech scientists, but historically, for some reason, voice did not attract the visionaries in shillon. All the new hardware companies have voice at the heart of the product. All of these devices, they got rid of keyboards. They don't really have a screen or an interface, and voice is going to be the main one. Hi, I'm Mette from FirstMark. Welcome to the Mad Podcast. Voice AI is having a big moment. For years, the field was stuck in the uncanny valley, lagging well behind other AI modalities, robotic, slow, and frustrating. But in the last 18 months, everything has started to change. My guest today is Neil Zegidor, CEO of Gradia, and formerly of DeepMine and Meta. Neil is one of the very top AI researchers in the field, and a key architect of the rapid evolution of voice AI towards real-time native audio intelligence. This conversation is a deep dive into everything you need to know about voice AI, where we explore many key concepts in a very accessible way, and discuss plenty of fun stuff, including why voice AI has so few experts, the massive challenge of building native audio models, and the rise of autonomous voice agents. Please enjoy this terrific and very educational conversation with Neil Zegidor. Hey Neil, welcome. Hey, thanks for having me. So, a lot of people in the industry are saying that voice AI is having its big moment. There's certainly a lot of activity, there's a lot of funding rounds. From your perspective, so you've been in this field for many years now, DeepMine, Meta, NAC Gradium, is voice AI, indeed having its big moment, or are we still early? I think it's both having a big moment and we're still early. It's having a big moment because there is progress all around AI modalities, and in voice, for example, the progress in latency, naturalness, accuracy have been really, really huge in the past years, in particular in the two last years. And at the same time, text models have evolved into what we now call agents, which are not only text models, but they can actually make actions and manipulate data, access, information, and so on and so forth. And now when you bring both together, you can have voice interfaces that at the same time are going to solve complex problems. And so I think there is a moment now because for the first time, it actually can be enjoyable to and even more convenient to talk to an AI on the phone and talking to a human because you can call any time of the day or night. And the interaction is working pretty well and it sounds really nice and the latency is low and so on and so forth. So it's definitely having a moment because I think in a way, it's now it can be used in much more use cases than it used to, but it's still early because it's still quite experimental. So anybody who is using even the most advanced voice agents and compares that to the horror movie from 12 years ago, you know, it's obvious the gaps that is still remaining. And there are so many topics that are completely unaddressed at the moment, in particular, you know, every time you watch the voice agent demo, just realize that it's someone talking to a phone in a quiet room. So the day where you will have someone shouting to a robot in the middle of a factory and having the robot understanding what's happening and who's talking to them, that will be, you know, like, will be there and we don't know that they are too. So we'll get into some of the technical details in a minute, but at a high level, why has voice AI being, I guess, the most underdeveloped modality? There's been obviously extraordinary progress on text AI and then image AI and then video AI, but it seems that voice has been a little bit the the the pro parents in terms of progress. Why is that? I don't want to be mean to my people, the speech scientist, but historically, for some reason, voice did not attract the like the visionaries in machine learning. So even if you looked at the dynamics in conferences, if you propose a new method, like fundamental algorithm and you wanted it to be accepted in prestige use venue, you had to have an application as a computer vision like image classification or an LP. If you did it in speech, you will get rejected because it was like a two speech and at the same time the prestige of speech conferences used to be much lower than that of a computer vision or an LP. So honestly, I don't really know why because when you look at the details, the first big success of deep learning everybody knows the Alex Net model in 2012 where for the first time, you know, you had a deep learning model outperforming every single alternative on image classification. But actually, the really first big success in deep learning was speech recognition. We've worked from Geoffrey Hinton himself. And when was that? It was in I think 2007 or 2008. So way before? Yes. And so I think it's, you know, it was just not as prestige use. And so it wouldn't attract a lot of people that could have made significant contributions, I would say. And then what happened and was very nice is there was kind of a convergence of algorithms around transformers and LLMs, so that now pretty much regardless of the task or modality, you are looking at, you're always looking at the same technology. And they're started to be much more progress also thanks to now the similarity between as different modalities because in particular what I contributed to with my team was to take a lot of inspiration from successes in vision and NLP and apply them directly to speech. But it's still interesting because in a way, there are way fewer people who can train a competitive speech model than in text or in vision. So I mean it's a good position to be in because it's very few people have really going to like the depth of this topic. And it's one that is very challenging because it's bringing, ideally, if you want to solve the problems, you need to understand machine learning, signal processing that is much more, you know, completely different literature around telecommunication, audio compression and so on, along with psychology and not like cognitive psychology, but psychoacoustics. How does the human earring works? How does speech production works in humans? You know, to all of this, when you bring them together, you can make competitive models. But it's it requires kind of a very wide scope of expertise in very different domains. Fascinating. To put it in numbers, how many would you say people there are in the world with that expertise? Are we talking about 100, 500, 10 between 10 and 100? No, I would say 50. I don't know. So it's hard to say. Yeah, I think it's very few and really meaningful contributions that I've pushed the field forward have been made by very small groups of people. And I think that's also what's nice. So AI, I think in general, is one field where individuals can have a disproportionate impact because, you know, the amount of things you can do by yourself, you have access to to compute and and that assets is huge. And in voice in particular, since the required compute is much lower and is at the same far data, really a few individuals can make a stuff that is completely, you know, just changing applications at very large scales. Great. So you mentioned a minute ago, which is the inevitable reference for any conversation about voice. What is the ultimate success in voice that the field is working towards? Is that super low latency, expressiveness? What is what is great? So latency, latency is already something that only makes sense if you are in a turn-based conversation because latency then the definition of latency is how much time there is between two turns. One of the things we contributed doing is getting rid of speaker turns completely with what we call full duplex conversation. So I think that it's always listening, always speaking. And when it's not speaking, it's just that it's producing silence, but you know, it's always on. And in that context, there is no real latency anymore. So because the model is just basically can talk at any time and it can talk over you and you can talk over it. And that makes the conversation really natural. Then naturalness, it's not only this dynamics in terms of tempo, you know, like when the model can jump into the conversation, when they should remain quiet, there is this dynamics question. And then there is emotion. And so there is emotion in what the AI expresses. Its emotion is natural, but also appropriate that if you start feeling confident enough to start sharing about stuff that makes you unhappy or sad or feeling miserable, it's not saying, "Oh, I'm so sorry for you. Let's talk about it." And it should also understand when you're getting upset, when you're getting confused and so on and so forth. This will make already voice AI in terms of interaction, extreme natural and the as close as possible to you, man, which is basically what is one of the things we see in the harm movie, which I hate mentioning as a reference because it's so overused, it's annoying. But at the same time, everybody understands, you know, the gap between where we are right now and the movie. So that's I think it's still a relevant one. And it's even interesting how relevant it still is, despite the fact that there is so much work around voice. But then there will be other questions about how voice is integrated into our life. So, you know, there are paradigms in voice AI such as Wake World. detection. So you know, when you use Google Home or Alexa or whatever, you have a wake world that is going to turn the speech to text on. So now let's say you want to work with your assistant that is always listening to you. So in a way, you will have something that is just running constantly without having even, necessarily a wake world. So all of this, I think, is going to be both technical challenges and product challenges around where do they sit, how are we interacting with them, the link to the hardware as well. So I think what is a good sign for voice as a field as well is that in my perception, all the new hardware companies have voice at the heart of the product, all the prototypes that we see whether it's glasses or pendants or you know, like the new stuff that are Geneva and some of them are working on voice is at the heart of the product and will be the main way of interacting. So all of these devices, they got three of keyboards, they don't really have a screen or an interface rather on, you know, and voice is going to be the ones that is the main one. What's your vision of the future, where does voice fit in, is that a voice and text, is that primarily voice for certain use cases, you know, that's certainly an argument that you're hearing a lot of people saying voice is great, but like most of the time, like I'm at the office, like the last thing I want is like for people to hear my conversation and therefore I don't want to talk to your machine. So where does voice fit in that vision of the future? So for example, I used to think that one of you's application where voice was kind of irrelevant to us coding because it's fundamentally you're not going to read code out loud, right? Yet now since coding is going more and more towards vibe coding, which is natural language, it makes a lot of sense to do it by speaking. And now people are developing products that allow you to, you know, dispatch orders to code agents in a way that is much more efficient that if you had to type in each different window to each of them. Even prompting LLMs now is doing by voice is much more convenient rather than typing. I still agree that there is one part which is more social about what the office's environment will look like. Oh no, maybe we will just also rethink the way we just structure office environments. What is sure is now people have AI assistants that are almost colleagues, right? I mean, you talk to any software engineer, the anthropomorphization of cloud code is I find it extremely funny. Even the verb coding is going to be coding pretty soon. And so these people, you know, they will, if it's more convenient to interact with their main tool, move voice that will justify also rethinking office spaces, I guess. So, yeah, I think there will be work around that and we will naturally find them if voice becomes the main way. I mean, if it's more practical to interact with AI through voice. It's super interesting. Before we go further, let's talk about you a little bit and your journey and the company. So I mentioned DeepMine and Meta and now Gradium and QTN in the middle. Just like what worked through your life story and your work. So I stood in mathematics and I started my career in a short internship in quantitative finance. And that was in Paris, right? Yes, in Paris. And I was born and raised there. And what was interesting doing my internship, I had access to Bloomberg, terminal. And so I will see the news, like the constant news in the below the screen. You know, I was thinking what if I could have an algorithm that just reads these news and take positions on the market faster than anyone because it was just able to analyze the news live. And I was looking, but I had no keywords about that, right? So I literally googled how can I analyze text automatically or whatever. And I found machine learning. I, you know, it was epiphany decided to completely stop. Started studying again. My goal was to go back to finance with AI and machine learning. But I got passionate about all the possibilities there was around the, so by the end there were no LLMs and it was not really about generative AI. It was about medical imaging, text understanding, a lot of things around audio speech recognition obviously. And I looked for an internship which was about unsupervised learning, which was already pretty cool. And I just wanted to do it. And so I pretended that was passionate about language. And I got the internship. And then Jan LeCamp and Facebook Paris, I was able to interview. So for the anecdote, I did my coding interview with Sumich Intala, who then invented PyTorch. And now I think it's a city of thinking machines. And so I didn't know how to code because I had only studied mathematics. And so he asked me to implement k-means, basic algorithm. And I asked if I could do it in MATLAB because I didn't know. I didn't even know PyTorch. I knew a very basic Python. And it was kind enough to let me do my coding interview in MATLAB. And I got the job which when I think about it, it's so cringe because, oh my god, thank you, Sumich. And yeah, I did my PhD there. It was very interesting to already on Ron's pitch. And I was spending half my time at Facebook and half my time at Ecolon normal superior in Paris in a lab that was studying language acquisition in babies. And in particular, the main observation on which the lab was built was that humans learn language from mostly two speakers, their parents. With few hundreds to one thousand hours overall in the first four years, with huge variance between social backgrounds. And without annotation, right? Because you learn to speak before you learn how to read. And that still makes us already pretty okay for conversation when we are kids. And you know, speech recognition, vaccine was trained with already hundreds of thousands of hours of annotated data. Now it's millions of hours of annotated data. So the topic was more around efficient learning, which is interesting because it was ten years ago, but now it's still as relevant as I think there was a new company that raised a last year recently to make learning more efficient. So it's still as relevant as it used to. And then I joined Google at that time. It was interesting because so I joined working on speech in Google Brain and there were almost nobody working on speech in Google Brain. It was not considered vibrant resource topic. It was like a product topic. A lot of people were saying, "God, but it's solved. Oh no, it's solved. It just works." So I already back then. And what were you always at? 2019. And so I found someone to work with me and we did a lot of work around speech. And then I got excited about the specific topic around compression with your compression. So it was just out of patient. I wanted to do a new compression format that would not be MP3. It's like a second value at the HBO. Exactly. Absolutely. And I wanted to do it with neural network. So the idea was that it will be computationally more expensive to compress and decompress the audio. But then you could compress it much more efficiently. And that's something we worked on for Google Meet. And it was called Sound Stream. That was the first what we call neural audio codec. And I had no plan of doing generative modeling, like Zen. I really didn't care about that. But I was very lonely in a way. And I just wanted, I was trying to lure some people around me who are working in reinforcement learning. I wanted to get them to work on speech with me. I started a project around diarization, which is a task of your listening to a conversation. And you have to tell who said what, which is probably the less sexy research topic out there. I'm sorry. I think it's special, but you cannot get people excited. I mean, it's very hard to get people excited about that. So within speech, which is not very sexy, this is the less sexy part. Like the monk project, you know, like very lonely and very, yeah. I was not very successful with that to get people to work with me. And I thought, okay, generative models, the nice thing is that if we generate speech, people will listen to it and say, oh, that's cool. My thing, you know, my method has generated speech. So I think it was very opportunistic. For me, I thought that it would be a good way to get people to work with me. And so we started a project. And the idea was that we started to see success around language models. So it was 2021 way before LGBT, but you know, internally at Google, there were already quite a few projects that were successful around language models. And so I do was that just after the work we had done on the neural codec. So now if instead of using your codec for a real time communication, but you just use it to compress audio, now you have repeal on a different quickly what codec is. Yeah, a codec is just a compression, a compression. So you have a node, you're right. And you want to send it over the network when you're having a zoom meeting. And you're not going to send the uncompressed wave file because it's too heavy. So you're going to compress it in a much lighter file that you will send over the network. And then the receiver can decompress it into back into audio. And the secret is based on a lot of science and knowledge around human hearing. We know what kind of information we can remove from audio so that it won't create a perceived degradation basically. So it's there is a lot of science around what specific information you can remove from from an audio that will make it almost as good for a human as the original one, despite the fact that we removed a lot of information which allows you to compress. And the main idea was that instead of using art coded rules to do that, we would learn from data what are the transformations that allow to compress audio while making it as transparent as possible for the human ears basically. And so now we add this way of compressing extremely efficiently. plus de plus de 3 ans en plus en plus. Et dans le cas où vous pouvez considérer que c'était trop compressé que c'était plus comme texte. Et donc avec un simple truc que nous avons, nous avons juste trainé l'al-lm à produire cette compresse audio par le texte de l'al-lm. Et puis vous pouvez faire exactement ce que vous pouvez produire. Vous pouvez prendre, pour le texte du seconde audio, compresser, passer à l'al-lm et laisser le texte du seconde audio. Et nous réaliserons que dans un moment qu'on avait inventé l'instant de voice cloning, donc nous pouvons repliquer un voice avec un petit seconde audio. Et oui, ce qui était très successe. Parce que c'était tous les advantages que nous avons avec l'al-lm, nous pouvons bénéficier. Donc l'al-lm est très bien aidé à modéling long contexte. Ce qui est très bien aidé. Donc si vous avez un modèle large, vous vous pouvez juste accueillir le modèle. Ce n'est pas un abuse comme ça, mais pour beaucoup de architectures, c'est très difficile de aller à 100 millions de paramètres à un parlement de paramètres. Avec transformé l'al-lm, c'est un abuse. Et je peux aller dans les détails, mais en. Qu'est-ce que c'est un project called? Audio-lm. Et puis il a dit "Musique-lm". Et puis il a dit "Note bouquet-lm". C'est un podcast automatique. Et ça a été un standard framework pour audio-generation. Il y a 2 familles qui ont été combattés à faire des fois des défusions. Les modèles, ce qui est ce que je pense que 11 labes ont été faits avant. Et nous avons été le modèle audio-longuage. Je pense que aujourd'hui, tout ce qui est audio-longuage. Parce que, depuis des autres autres, ils ont été en mode de streamer. Ils sont naturellement compatibles avec des inférences de temps. C'est un peu le plus grand de la voie. Et donc, tout le monde est utilisé cette technologie aujourd'hui. Et oui, c'est un très très succesful. Et c'est très facile d'appliquer une nouvelle table. Donc, vous savez, on a dit en speech, la première. Et puis nous avons collecté le modèle de performance piano. Et puis nous avons le modèle de piano. Et puis nous avons fait plus de musice générale. Et puis nous pouvons faire très bien. anything. C'est aussi une chose par un profite non-profit. L'ab, qui est un humain. Sorry, animal vocalisation. Donc, il tries de décoder le longuage de l'animalisme de l'espoir de l'Earth-Species Project. Donc, vous pouvez vous dire, comme l'Element de l'Earth-Specialité, c'est facile. C'est facile. Vous avez joué un très important piano-ningeré, le rôle, dans le cas de la voie de la voie. Et puis, c'est le reste de l'element. Et puis, au moment où, Geminai a commencé à Google, c'est quand je suis là. Donc, je voulais vous dire, pour créer un petit ressort de l'environnement. C'est pour me reminderer des derniers jours de faire au Google Brain. Donc, un très petit team, élite, n'est pas une distraction. Sorry, pour dire que n'est pas un product manager, comme un ressort scientiste. Non, il est en train de réciter, juste à la maison, avec les machines et les fonds sur les sciences. Et, en particulier, ce que le rôle de moi a été vraiment pour travailler sur le research et pour placer le monde et les entraîner les students et tout. Parce que je suis très très grateful pour avoir été able à faire le research dans un certain nombre d'environments. C'est aussi aussi l'obstacle de moi et pour ça, je pense que je me récoute 100% avec le piano de le piano, pour le faire de ce qui est ce qui est récréable et dénémique et dénémique en 2012, pour où nous sommes aujourd'hui, c'est un ressort open. C'est parce que c'est un monde qui est un collaboration et tout le monde est dénémé et c'est pour moi, c'est important pour le monde et pour le monde. Et donc, on a décidé de créer un profit avec l'invasion de l'exéminorité et de l'exéminorité de l'Ordor Sade. Donc, pour le code "N°1" est faible parce que c'est le nombre de les restaurants où on a discuté le projet. Et donc, on a demandé de nous ne pas le dire "n'exéminorité" évidemment. Donc, on a juste demandé de la chat Gpt pour le nom de la chute d'Espher et le japonais en couteau "Espher" et il était "AI" donc c'est comme "OK, c'est le nom du lab" et donc c'est ça qu'on a créé le couteau "Qui". La première personne que je récite, c'est l'exéminorité de l'exéminorité qui est aussi maintenant un fondeur de l'gradier et c'est notre type de science-officeur. Parce que nous avons fait notre PhD ensemble avec Facebook et nous avons fait des réveils parce que nous avons fait le même et à chaque fois, nous avons fait un autre qui ne nous aiment pas parler de quelque chose. Parce que c'est comme "What are you working on" ou quelque chose. Et donc, on a un petit team mais avec des experts qui ont pu faire des choses et nous avons décidé de faire des choses avec une autre décision opportuniste que nous avons fait. Donc, je l'ai regardé à un autre chose et pour un lab, on a 1000 GPUs qui sont aussi beaucoup en 2023. Il a déjà déjà fait un autre qui est un autre qui est plus large et plus large. Donc nous avons fait le projet où nous pouvons faire des choses par le fait que nous avons 4 personnes. Il ne faut pas trop de compute et il va être très innovatif par le fait que, par le fait, nous pouvons faire une différence. Nous avons fait des conversations avec les réveils et des conversations de rétimes. Et il a fait des conversations avec les rétimes et des conversations de rétimes. Il a fait des conversations de rétimes. Et il a fait des conversations de rétimes. et aussi réalise que l'académie de la prestige est une chose mais de faire une impact is having your models being used in the real world. And so for me, it's the ultimate impact we can have. We still do science in particular, QTI keeps doing open source and open science and so on. For me, the upside about it is mostly to be able to train the next generation of AI researchers and keep the field alive, as I said, because I think it's to have a healthy and vibrant AI field, you need to have scientific dissemination, so scientific exchange between institutions. And they also, you know, like the Chinese lab, I'm making a remarkable work. And it's the kind of, are forcing everyone to stay open to some extent, because otherwise, it also hurts the ego, I think, of the people who are in the lab that don't publish. That was also something I was very opportunistic about. So my strategy was like, if we publish in a world where the authors don't publish, they will get so pissed, you know, of us claiming all the inventions, that it will make them join us eventually. And it's true that, you know, it's kind of hard, because some people, they want to be in the place, that is the cool place, where the cool stuff happens. And it's not only about composition and so on, it's now, it's, I would say everybody working in AI and doing a good job in AI, is going to get good economic outcomes. So then what can you, you know, glory is also very important. And it can be scientific or it can be just being proud that you are making the best products. But I think it's also an important part. Important part, yeah. Great now, Gradium is an actual company that was launched a few months ago. And obviously you are a new entrant in the field of AI, where there's tremendous amounts of competition. So I think for voice in particular, like the obvious question is, why has OpenAI or Google, or made it not already one voice AI? And I think you probably alluded to some of the reasons up front. But what, why is that? Why can a small company hope to become the leader? So one thing I mentioned was, if you have the right team, it can be extremely small and still make a significant impact. Other arguments, I think, is one is focused. So for example, if you look at large multimodal models, right? Like these generic models that understand images and can generate text and can produce code and so on. You have like a limited budget, which is a number of parameters and data you're going to fit your model. When you want to add speech to them, you're fighting with coding and image understanding and so on. So you are playing with a lot of trade-offs that are irrelevant to the tasks that you want to solve. And at the same time, these models are so large, they cannot run at scale because they will just make everyone lose money in the process. The only format that makes sense for speech models to run at scale is to be extremely compact, which also means that the training resources you need to train them are much smaller than what you need in a to train other kind of models. So the resources are not as challenging as for text models. I think also another aspect is in a way not trying to just make a conversational product. So really making building blocks so that people can build the product. So we could make the gradient conversational assistant and think a lot about its capabilities and what it can do and what it cannot do and what will be the use cases for people and so on. And hope that like the voice mode of OpenAI, it's used for this task and this task and so on. But it's impossible to cover everything with that. And so now if we want to cover NPCs, fake sport commentors in video game and a language learning app and an annoying character in a cartoon and the customer center agent and so on, then we just make the building blocks. And this again is not really I think in the DNA of big companies to do this kind of very specific models that are targeted towards developers, rather than trying to solve a lot of things at the same time. And then there's an additional aspect to this, which is that voice can also and should be pre-offent on device versus an API called to the cloud. Is that fair? I think what is very challenging right now is if you want to have the full intelligence on device, like the full conversational AI on device. And I say I would say at this point, if you want such a model to be useful, we are not there yet, right? You can have something that can cheat chat a bit and it will be decent or we also have shipped models on device but there are much more constrained in terms of applications. So for example, we started a year ago with a on device to speak translation, which is something that makes a lot of sense because when you're traveling, maybe you don't have a data plan that is going in every country, so it makes sense to have something that works on your phone if you want to order at a restaurant, so feel like that. I mean, so particularly adapted use case. But now we also, we released two weeks ago a model called Pocket TTS that not only is on device but CPU only. So there are already modes for AAA video games where the NPCs can be powered through these voice models. And now you unlock a completely new kind of applications because on device models allow to do very large scale a personalized content that will be economically not realistic with an API. So again, these kind of things is, you know, if you want to make meaningful progress in that direction, making small models in voice is much more difficult than making large models in voice. So keeping the quality while reducing the size of the model, that's where the big challenge is. And our CPU model in terms of algorithm it's like the cutting edge of what we know. It's really the later generation of everything we've been doing so far. So just a few days ago there was this big announcement by Alibaba/Quenn that they were open sourcing the Quen 3 TTS family for voice design, clone, and generation. How do you think about open source in your world as gradient is that a friend, is that a foe? So if you read the paper, you will find our names. In several pages, it's mostly inspired from the motion architecture, like pretty much every model right now, even the virtual model that was released by Mral Twixgo is also based on our framework. I think that's really interesting because this proves that there are things that we do right because everybody is building on them. At the same time, I would say it's quite an advantage because I would say not there are two kinds of research papers. There are research papers that are meant to be as explicit and reproducible as possible, which is what we try to do when we do one. And there are some that are more about, I would say, marketing in the sense that they are mostly focused on the results and the performance rather than explaining the underlying mechanism and the data and so on and so forth. So nice thing of people building around the frameworks we introduce is that even when they don't give details, we can infer the details. So in a way, in a competitive landscape, I think it's quite an advantage because in a way, it would be more challenging if people were transitioning to something completely different from what we've been doing because then we would not infer anything from one reading their papers. Now all of this is very familiar when we read those. And at the same time, we have all the issues that people are facing. We have been facing them for a while and so we have already worked around and new versions and so on. Open Source ad cuties as like the end goal of the lab. Ad Gradiums, that's not the end goal of the lab. The end goal of the lab is to make the companies to make competitive products that all perform every alternative. But Open Source is, I think, a good way of following developers to prototype stuff, understanding what people expect, what they want. You know, it's also a way to train talent. A few days ago, we released the Hibiki Zero so that Hibiki was our speech-to-speech translation system. Now it's a new algorithm that makes it even lower latency, better voice-cloning, multi-lingual, and so on and so forth. And the PhD student who was working on that project at Qtai is joining Gradium for a few months to do visiting PhD and then we'll start this PhD again. So I think there are very healthy relations between Open Source and Close Source that can be done. At the same time, I don't think it hurts any defensibility-wise opposite honestly. And so I think it's also for talent. It's very attractive because people, when they join us, they know that they can work on really competitive products. But at the same time, they can keep sharing models and more exploratory research and do really frontier stuff. And I think if we want to stay at the cutting edge, we should be a product company, but also be a frontier lab. And being a frontier lab, you need to do fundamental research. But what's the gap between Open Source and Close Source in voice? Obviously in text, there is this cat and mouse game and it seems that the commercial labs are constantly like pretty far ahead of the Open Source. Is that the same thing in voice AI? So what is interesting in voice AI is now I think a high school student could make some things that is decent. But then what people want is the last mile. And the last mile is extremely difficult. And the last mile is pronouncing, well, all the difficult cases. Is having a latency that not only is low, but is robust and almost zero downtime. And being able to clone voices, regardless of facts and so on. So all these hard cases, for me, the only way to find the energy to solve them is because that's your business. Because otherwise if you look at benchmarks. like the benchmarks for TTS, it's a Libris pitch. So it's books. A lot of them from Dusty Fski. So it's more about whether you are going to miss the I or Y in a Russian name. Or it doesn't evaluate, can you pronounce for numbers and email addresses and new arrays and all these stuff. So when you optimize for these benchmarks, you lose a lot of the actual real world cases. Yeah, so I think the main difference is the incentive to do things really with the finance details only makes sense for when you want to be competitive in product. It's not only about the models, right? The infrastructure to run these models at scale is extremely important. I think particularly in API business model, your margins mostly depend from the efficiency of your inference. And so in particular, I'm very lucky to have in my team one of my co-founders, Laurent, who did several years in a-- he did all his career in quantitative finance, mostly at Genstreet. And now we have more people coming from quantitative finance. And these people, they are really passionate about efficient inference. That's our bread and butter, you know? Because in financial-- like in high-frequency trading, that's kind of-- that's the only way to exist. And the engineering challenges are really significant. And I think A team in engineering as well, and not just like training models, makes a huge difference. And where now, if we are talking then about on-device inference, that's even more complex, right? If you want a model to run on all Android and iPhone and so on, that's big engineering challenge. You mentioned a benchmark a minute ago. And that seems to be a really interesting question for voice and video as well. And images is how do you make a case that your technology is better than the next provider? Because some of it seems to be a little bit around vibes, like how you feel when you're on the receiving end of a voice AI. One thing that is clear is that you can only trust human judgment. People have tried to make objective proxies of human judgment. Like, that would be a neural network that listens to an audio and gives it to grade. It sucks. Like so many people try, and it works on their constrained setting. And on real audio, it doesn't work at all in completely breaks. So we don't trust anything about our ears. So we do a lot of blind tests internally. We do a lot of blind tests externally. So we're working with human judgment constantly. So every single decision we make is based on human listening. We don't trust metrics at all. So it's fundamentally subjective experience, the quality of audio. But there are some things that are going to be widely shared. Reserves like the prosody, so the tone and the rhythm are natural or not. A lot of people would agree on that. Is the voice nice or not? Nobody agrees on that. And so then the only way for me to claim to have the best solution is to have the largest catalog and most diverse set of voices that people can pick. Because then it's the kind of what voice are people going to like this three depends on between people. I had faced that in the past. When we did music LM at Google. So for the first time, we made text to music. So you could type like a death metal with a marine bar or whatever, and you get death metal with a marine bar. And so we put like a website online where people could do this stuff. And at the same time, we were already planning to do for the first time a RLHF. So reinforcement learning from you might fit back on music. And so what we had designed is people could give-- so there would be two generations every time. And people could give a trophy. I had planned that because I wanted for the first time to have integrates humans in the loop of music generation. Because I think scientifically that would be really huge. And what was interesting is that-- so we made a paper about that called music RL. And it was quite nice. But the results were not extremely convincing. And then what we did is that we just did judgments our judgment ourselves. So with people with my colleagues, we took like, I don't know, 10 or 20 pairs of audio. And each time we choose our preferred one. And there was zero agreement between people. So there is no way, like, an algorithm is going to learn human preferences when there is so much subjectivity. So the only thing you can make is make it more steerable for each user to be able to customize it. Because there is nothing that is going to please every user consistently. The obvious next question is if nobody can agree on whether this model is better than that model, then aren't all the models more or less the same. And therefore the entire voice AI model industry is sort of like commoditized. All the models being the same. I think then people talk about TTS already, right? Which is much more constraints than voice AI. So in text to speech, I mean, there are factual metrics about accuracy and latency. And then there are more subjective things around the expressivity and so on. But you can, again, you can really make a difference by making it more controllable, more customizable, and so on and so forth. So I don't think that it's clear that the best TTS, the most controllable, the most robust, the most smart in terms of expression is in front of us. Nothing is close to it yet. And then there is everything that is not TTS. You look at transcription. Now again, what if a lot of people are speaking? So I was talking about the realization as the least sexy problem ever. At the same time, it's a semi useful problem and you look at the error rates and they are very bad. I mean, it's just not working in difficult cases where you have a podcast with a lot of people talking at the same time, it just completely breaks. Full duplex, we did machinery a year and a half ago. Still nobody has made it into a product. There are so many things that are just not existing today that I think the communication maybe it will happen someday, but we are very, very far from it, honestly. And the gap in, there is already bridging the gap in quality and abilities to make that as powerful as human production, as speech and understanding, as speech in complex environments and so on. And then even if you reach this point, there will always be challenging about getting the same performance with smaller models. Because then again, like a full duplex smart voice AI that can run on a gamer GPU or on a iPhone. Good luck, you know, like that. (laughs) Okay, awesome. I'd love to spend a little bit of time on the technical aspects of voice AI. So we talked about some of them alluded to others, but I think it'd be nice to bring everything together. So in terms of the fundamental innovation that you described around speech to speech models, can you, for us, compare and contrast like the old way, which I believe is the cascade versus this new generation of models, how does that work? The Gatscaidid system, which is the old way, but it's actually pretty new. So like the old way is without elements, you know. (laughs) So we talked, you know, the old elements, so the old way is like Alexandria and so on. It's back two years ago, back in the next years ago. The old way is natural language understanding with parsing and with very limited vocabularies and so on. But the Gatscaidid idea is you basically just take a LLM and you wrap it into a speech input and speech outputs. So you have like streaming transcription and streaming text to speech. The nice thing about that is you can plug in the LLM and you can customize its behavior through its prompt, like you will do with your text model. A main imitation is around first latency of the whole thing because you're running through model in the cascade and each of them is going to add to the overall latency of the interaction. And by going through the bottom neck of text, you lose what we call parallel linguistic information, which is all the information we convey when we speak on top of what we say. So that's emotional state, irony, lying is one. You know, like a lot of, you know, it's a, you know, a lot of information is conveyed that is not, that is not in what we say. Particular, I encourage people to look at the, every few years, interspeech, one of the main speech conferences organizes a parallel linguistic challenge with challenges that are more on them every time. Like, so lying detection is one, but then you have trying to recognize the origin of the parent of someone from them speaking. Or there was one about people are speaking while eating and you had to recognize what they are eating. So you know, anyway, when we speak, there are a lot of information that come about us. And this is lost through, so what you are eating, obviously, is maybe not the most, like a relevant one, but understanding in the customer care context and the things that are getting lost or annoyed and so on, it's extremely important to be able to recover the conversation into good state. Sometimes it's so used from the text, like if the AI says good to hell, yeah, probably they're upset. But sometimes it's not that that abuse in that kind of cues can be, can be helpful. And so that's the only way, I mean, one way to address all of that altogether is speech to speech. And full duplex on top of that addresses what I think today is the worst part of voice AI. I hate it from the bottom of my heart. It's turn taking. In any introductory course to a conventional network for image classification, there is always this analogy that is made that you can make rule to tell whether an image is shot in the morning or on night, just based on colors, right? You can just look at colors, the values, and make an handmade rule. But now if you want to recognize whether it's a cat or a cat. d'un dog. Vous ne pouvez pas faire rouls que vous reconnaissez tous les shapes de cats et dogs et tous les ingots et tout et tout. Donc c'est pourquoi nous faisons le machin de l'anning, juste là, en fait, de la data, quand vous ne pouvez pas faire rouls. Nous avons fait un "king", nous sommes à l'arcaïque era de "and made rouls", qui est ridicule. C'est comme, vous avez un algorithme, c'est que la voie d'activité d'activité d'activité d'activité, que vous êtes silent ou non, et si silent plus que 100 de minutes, et puis cette roul, et cette roul, et cette roul, il est une interruption. Mais si ça ne fait pas un interruption, et puis nous sommes rouls en top rouls, en top rouls, pour décliner le modèle de la taille, et ça fait ça, je sais que vous savez que tout le monde est avec les systèmes de caster, maintenant, vous vous needz de décevoir quand vous parlez à AI, vous vous needz de l'adapter à la flotte, otherwise, vous êtes bien faite et bien confus, et interrupte, et ça c'est extrêmement annoying. Plus, encore, les gens peuvent parler aujourd'hui, c'est pas vraiment pouvoirs, mais la qualité de l'interaction est enceinte. En ce point, vous pouvez vraiment prendre 5 secondes pour penser à ce que vous allez dire, et puis le modèle ne va pas être faible, il va juste être extrêmement flexible. Et ce que nous avons dit, c'est que c'est un très simple, donc il a fait comme un modèle audio, instead de avoir une strima de tokens, nous appelons le multistreme, juste deux strima de tokens, une fois le user, une fois le AI, et nous aussi peut être actifait de la même manière, et il n'y a pas de temps de prendre enceinte, et vous ne faites pas de terrain de la table, les gens parlent de la telefon et vous avez un personnel enceinte le même et un enceinte le même et votre modèle est le même. C'était un point integrated de tyres très satisfaiblique, 아니라 un spicy tragedy. Je commence à dire que si on avait le懷 le chien, on a les pourb 않고 , et nous diğerons legumes. - Je me devrais trans silhouper una fois. Comme que il tikais un autre déplet, mais apportait longtemps. Pour qui avoiding plus de地方, est-ce que tout n'a pas remarqué récemment, comme un petit peu de politique. C'était très difficile, car on est en train de dire qu'on a dit, "Hey, je suis en train de l'Amérique et tout le monde est un nouveau." Et on peut dire que vous avez des questions sur le point, et tout ça. C'est un genre de paranormal expérience très fun. Je pense qu'il est évidentement possible de choses que nous devons faire à la suite. Donc, évidemment, pour nous, comme le «N-Go» est un coup de complexe, c'est toujours notre motivation. A la fois, ce qu'on fait est l'assurance d'un système cascade, car c'est où le marché est là. Il y a aussi des gens qui sont éterrétés sur les textes de l'irlage, que ils veulent utiliser, sur les textes de l'irlage et tout. Il y a beaucoup de progress sur le texte, un gros espèce de model speech. C'est que, depuis tout ce que c'est intégré, quand vous allez sur un model texte de l'irlage, vous avez besoin de fin de la fin de la fin de la fin. Donc, maintenant, c'est le cost de la switch de l'irlage. Le model texte est extrêmement high parce que vous allez avoir besoin de fin de la fin de la fin de la fin. Vous voulez quelque chose qui est modulant, de la fin de la fin de la fin. Ce qui nous permet de faire le même flexibilité à la cascade de l'irlage, pour que vous changez le backend de la fin de la fin, et vous obtenez le même customité et de la customité que vous avez avec le model texte, mais avec le model de l'irlage. Et ce que je pense que c'est de être un évident pour la solution pour tous les limites. - C'est le frontier, un des fois. - Une des les frontiers, je dirais. Mais, le frontier, je dirais, je dirais, je dirais, pour tous les limites de la switch de l'irlage, qu'elles ne se sont pas en train de la fin de la fin, c'est un robot dans le model et le factory. Il y a beaucoup de machines de l'irlage, et vous avez beaucoup de gens qui ont des robot et le robot, pour les choses qui ont des choses qui ont des choses qui ont des choses. Je peux vous dire que ce n'est que ce qui est électionnant. J'ai passé le碼 qui est dans le bras qui est la plus ace qui va être un très Chaneling, je pense que ça doit être un voilaicht. Même plus la plus faible, même plus du plats. En effet, fortement plus oyster. J'ai à ce que cela permet d'une pièce với les trous que vousładons. Si vous savez autrement, qui le thiome votre gênage à ce vol et vous le acostez à manger au P Toulou, vous voyez si votre gênage est inv 아니ide et si et il vient sur la pliga, si vousakatons une highly democrieraie, avec l'inspiration de son p Progressive, je suisGenerale la panneux sur argued avec une Chaleneurстанов. Il y a des étudiants qui ont eu des médecins qui sont là. Ils sont là, ils ont eu des étudiants. Ils ont eu des étudiants qui ont eu des étudiants. j Well, je vois, j'ai assez d' proyects durant longtemps. C'estṭigu, c'est zeroprogaix. Vous avez déjà mentionnéod Xuecon-game et R2 J'ai fait Спасибо pour vous de faire! J'ai les tests qui ai scheposeux liters\ Plus ils sont vraiment lanié chez l'intro Mais sentir leurs étudiantes estaquelle Léa aussi Des pseudon dokładaire OU recommendation Nous était quelaire est ce que leur va faire Je sais pas, un modèle basic qui va faire un 100 millions de millions de dollars de speed C'est quelque chose qui est un peu un moment de vue qui est très hard à faire. Je pense que c'est une très très très question, qui est une discussion de beaucoup. Et tous les gens ont des theories. En particulier, un impact, un, je dirais, un attribut de speed data C'est que si vous traîne un modèle conversational sur speed data Il va être plus moins intelligent que le test de la mode de speed. Et je pense que quand vous vous écoutez de speed data, la capacité de formation est très très importante et vous avez des textes. Donc, vous ne avez pas de speed data, vous ne avez pas de speed de speed de speed de speed, etc. Vous avez des modèles de learned de la mode de speed de la mode de speed de speed de speed. Je pense que c'est une très très bonne idée. Je pense que vous vous commencez. Nous avons des textes de mode de speed. Donc ce que nous avons fait, nous avons commencé à faire des textes de mode de speed. Et puis, nous avons fait ce modèle de test et de la speed de trainement. Et nous avons fait de la speed de trainement pour la façon dont il est possible de la dose de l'intelligence. Donc tout le temps, nous raccourcons le texte de mode de speed. Et nous avons dit que nous avons fait de la façon dont nous ne sommes pas contentement. Mais, en fait, la qualité de la speed de date est très valuable. Je pense que nous avons fait de la pace de la speed de trainement. Donc, nous avons fait de la speed de trainement de speed de trainement de speed de 7 millions d'héroses. Je pense que nous ne pouvons probablement que nous avons 10 000 hés de la speed de la speed de trainement de speed de mode de speed. Nous ne savons pas sekés parce qu'on s'apprenaît desioxins 19-20%, parce qu'on déteste le sibling pour les équipements de promoted 60% JS Hospital Syms Parfois, ik really, accommodate deux- fuime. C'est une source de data, une grande volume, une petite conversation, comme publics, l'available audio-boots, etc. Et puis, vous êtes pour une spécifique de data pour vos applications. Par exemple, pour TTS, vous voulez avoir une expérision d'héroso-high-quality. Vous ne voulez pas avoir quelque chose qui est récordé avec des conditions arbitrariques. Vous voulez avoir étudié, récordé, très low level de noise, professionnel ou subi-professionnel acteur. Et puis, vous avez des textes de la voie de vous donner pour vous dire. Et donc, c'est aussi une question très intéressante. Parce que si vous voulez générer des 10000 ou 10000 horaires de la scripte, pour les secteurs, c'est pas facile. Vous ne pouvez pas avoir une clode de chat gpt, des 10000 horaires de la scripte, et vous faites des horaires comme des choses. Donc, ça ne fonctionne pas. Vous pouvez avoir une clode de chat gpt, et vous faites des horaires de la scripte, même si vous avez des horaires de la scripte, vous avez des horaires de la scripte, et vous faites des horaires de la scripte, et vous faites des horaires de la scripte, et vous faites des horaires de la scripte, vous faites des horaires de la scripte, et vous faites des horaires de la scripte, et vous faites des horaires de la scripte, et vous faites des horaires de la scripte, Nous avons toujours eu un jeu de la santé, nous pensons de la peinture et ce pays est un peu faible. Parce que, comme je dis, dans les ressources compétationales, nous ne avons plus besoin de la trainée large de modèles ou vidéo. Mais pour le speech, ce n'est pas vraiment du volume de la table, mais de la qualité de la table est très useful et très important. Et le accuracy, par exemple, de la notation, c'est par un amount et c'est beaucoup de la nature humaine pour être able à ne pas être plus facile. - Il y a des aspects de la table? - Oui, des aspects de la table est intéressant parce que la notation automatique fonctionne très bien. Et où les humains sont utilisés, c'est vraiment pour les quelques manches que les automatiques automatiques vont maintenant. Et donc, la seule façon de faire le temps de la table, c'est qu'ils ont une belle notation. Ils ont une belle notation, pas très belle. - Et comment ça fonctionne si vous voulez avoir une nouvelle language, spécialement, pas de la peinture, mais comme, je ne sais pas, ou d'Africans, comme ça, quelque chose, quelque chose, des moins de données sur le internet, presumably. - Est-ce que vous avez des acteurs d'acteurs d'acteurs pour les gens? - Comment ça fonctionne? - Ce qui est intéressant, c'est que nous nous verrons les étrangers de la culture, en particulier, des étrangers de la même famille. - Par exemple, avec les hibiquies d'héroes qui ont été lastues, donc il était en train de faire 50 000 h, ou peut-être 100 000 h, pour 50 000 portes qui sont en Spanish, et puis pour faire l'Italian translation pour l'English, 1000 h, c'était bien. - Je vais vous dire que pour l'angouage qui est complètement différente de la famille, c'est un peu plus de challenge. Ce qui est bon, c'est que, fondamentairement, tout le monde qui est en train de faire le même, tout le monde. - Maintenant, nous avons été facé de quelques languages pour trouver les récites qui se croient, et puis, ça peut être un apply pour plus de languages. Ce qui est très difficile, est de ne pas rétablir les languages, les languages qui ne sont pas un système officiel, des systèmes de la santé, c'est un lot de dialectes qui sont des gens qui ont mis les languages et ont un peu de frein et un peu de créole et un peu de quelque chose, etc. C'est très difficile, car le plus difficile est de faire des détails, car vous avez beaucoup plus de détails que vous pouvez faire, mais ce sont les plus plus de challenge que vous dites. Mais ce qui a été un challenge pour un long temps et il y a beaucoup de projets sur les détails de collections et des sessions. C'est intéressant, c'est un, pour exemple, le contenu que vous pouvez trouver est le plus de languages qui est un bain de la bain. et le désormais c'est beaucoup d'exèges sur le monde le plus left seul et volé de Fucking. Le bacon à 3Ч, て cuivre Sanghous��, La facile de tout ce que很有ис graduatingé maladie, c'est pas�� immigrant. Alors, ils font partie différentes à son cher valued. ochir disfrut de lesalies politiques, allez-y�� smoking en voici les combinaires. Ça déivide. Ce sont des plus grandes vari華 dans buscar hein! L'ambultataste suite actuellement. Vous utiliserez une nouvelle wasting. Vous voulez avoir des modèles voici à l'esprit et de faire une sorte de. dans des économiques, vous aussi vous voulez avoir cette modèle compact. Si vous pensez de faire des NPCs dans un jeu, vous pouvez avoir une paix sur 70 ou 90 de la paix sur un jeu. Et puis ils vont parler de 3 h/h au jour, car c'est un très bon jeu, donc vous spendez beaucoup de temps sur ça. Vous pouvez avoir une modèle large, impossible. Ce n'est pas seulement dans le GPU, mais même pour API, ce n'est pas une sorte économique. Donc, à fond de l'entrée, je pense que ces modèles ne sont pas très bonnes. Mais parfois, ils ont pu accéder à l'arbre large, car ils sont en train de se mettre en place. Je suis très. En fait, c'est pas très obvious, mais. C'est légitime et adaptif, compuieuse, en contexte sur le contexte et le difficulté de la tasse à la tête. La tasse est aussi précise comme le vol de l'autre. Je peux juste faire ça, comme ça. Parce que quelqu'un dit "je vous parle et je vous parle" ils appellent pour le vol de l'esprit et je dois faire une internet de l'émergage. Donc, c'est très sélectif pour qu'on puisse compter, je pense que c'est la seule chose pour tout ce qui fait ça. Écondémiquement. On va parler un peu de la producte et de la business aspect de l'VICI. Vous vous imaginez que vous êtes une compagnie de modèle, aussi une compagnie de modèle. Qu'est-ce que c'est un produit dans l'VICI? Et les deux choses que je pense sont comme une cloning, et puis les agents qui sont les deux qui sont les gens qui sont en train de parler de beaucoup, de prendre une autre. Oui, donc, je pense que cette chose, pour tout ce que le produit est en train de parler de modèle. Parce que nous pensions que les gens sont les agents de la voie. Donc les agents de voie, je pense que c'est un peu comme un agent de business, comme un agent de customer carré, ou quelque chose comme ça. Mais un NPC, dans un agent, parce qu'ils ont une interface de voie, ils ont une modèle de modèle de modèle de décembre. Éventuellement, dans le vidéo, ils vont être able de faire des fonction de la cloning, de la cloning de la cloning, pour les décisions que vous avez sorti, ou pour les contrôler des actions, et tout ça. Ce que nous faisons en faire, c'est de faire le plus de technologie pour les gens qui veulent être les agents. Donc nous ne sommes pas les agents de notre même chose, nous ne nous faisons pas le faire de la cloning de modèle, parce que c'est notre spécialité, ce qui est où notre talent est. Et à la même time, je pense que c'est un peu un réaliste, un réaliste, des choses qu'on peut dire, tous les needs de la cloning de modèle de notre agent, parce qu'il y a des choses qui sont très excitées. Et quand nous voulons à ta carrière de customers, chaque fois que la cloning de modèle est complètement unpredictable, ce qui va être de la cloning de modèle de carrière de customers, et des games de vidéo et des médias, et des personnes presses. Et c'est tellement grand que nous, pour les gens, nous faisons un peu de la cloning de modèle de modèle de modèle de modèle de modèle, et ce qui fonctionne très bien parce que, dans lesquels de l'est, nous sommes moins de 15 employés aujourd'hui, à l'agréation. Quand nous sommes partenés avec les companies qui ont des employés de l'est, dans les banks, ou les hospitals, ou des gens de différents pays. Il requires beaucoup de ce que nous avons dit aujourd'hui, et pour le four-wire de l'engine de l'économie, il y a beaucoup de travail, des laborations humains qui sont invoies par ces employés, dans les deux biodes complexes et des infrastructures, et nous pouvons juste le procurer sur le modèle. Donc ce qui nous permet aussi de nous remettre très bien, et nous nous procurons sur le science et le engineering. Et la cloning de l'est-ce que l'excusation est, en fait, à l'agréation, que nous avons le meilleur de l'industrie, et le meilleur de l'est, ce n'est pas seulement répliqué comme le même acteur de la spécifique, mais aussi le accent, un acteur de l'unit, et dans beaucoup de contextes, par exemple, si vous voulez créer un acteur de vintage, avec tous les rédits effectants, et quelque chose que nous avons fait très bien. Si vous voulez avoir un robot de l'économie, vous pouvez faire ça très bien, si vous voulez avoir un acteur de l'économie, UnPercht, Splitまあsk', Les Bleustar etAlte Duvers. J'ai non plus phosphorus avec une carrière reverb, en tout cas des filles de perspectives. Je pense que la rebuiltés est passionned. Cela fait une partie très petite après la Downtlle. Cela a un petit mayor volunteered smartphone, comme des inexpensive, qui assure des clampes normales à des loupes invest wombers ou nous tuez avec Jean Ambence. Les questions peuvent être générées sur le flight. Vous voulez le côté de la voie de la voie, de la taille. Ce qui fait beaucoup de sens avec cloning un specific voice, parce que vous voulez le côté de la voie de la caractère, de la personne, de l'attente, de la kpop star, etc. Il y a un genre de expérience qui veut engager avec un côté de Zeno, mais ils veulent aller pour une expérience internaitive. Dans un nombre de cases, par exemple en custom arcade, les gens ne font pas de specific voice. Ils font une voie avec un contexte specific et qui est très bon pour la use. Ce qui fait une voie avec un collé de la voie, qui a un bon voice. Pour ce que c'est que la solution que fait le plus sens est pas cloning, son design de la voie. Donc, vous pouvez créer des voies de naturel d'anglais. Et avoir votre génération de voie qui est vraiment fassant à la pointe, donc que vous pouvez être très ficturé sur cela. Parce que ces deux solutions qui existent aujourd'hui, des designs de voie, mais, comme ça, ils ne sont pas populés et qui n'ont pas encore été existés à la voie catalogue. Parce que je ne peux pas dire ce que c'est très bien. Mais quand ce qui fonctionne, c'est ce qui est très low, les gens ont besoin de la voie qui est une chose pour la use qui est très intéressant. Même si parfois, c'est un peu surprise, parce que nous avons des tomères qui ont juste un roi de la voie de la voie de la voie de la voie catalogue, qui est l'un de l'un de nos hommés, comme moi, mon collé. Comme ils ne ne font pas de la voie. Je vais essayer de prouver que les designs de voie sont ensemble. Mais on va faire une voie qui représente votre brun. Mais, les gens ne font pas de la voie, c'est bien. Ou peut-être que je n'ai pas une voie comme ça. C'est ce que le parc du monde est élevé. Une question énervite, de privé, de la voie. Et puis à la voie, c'est une contexte où il peut clomper. Si quelqu'un a une voie très facile, qu'est-ce que ça veut dire et comment vous protéger? Toutes les gens. Je veux dire que le parc du monde est un scame. Je suis sûre que je ne veux dire que ce ne soit pas un truc. Je ne veux pas dire que ce ne soit un scame. Je suis sûre que je ne veux pas dire que ce ne soit un scame. And also there is deep fake detection, which is just recognizing that a node is synthetic. Even if there is no watermark in it, just turning true from fake. These are very difficult to do. Unfortunately, if you get a phone call from the camera, unless you have inside your phone, the automatic detection is useless, right? You cannot put a prod it to a website that says it's true or fake. Honestly, I would say at this point, I don't know who has their grandma who lost their credit card and he has their passport and it's $1,000 today. Because it's always a story out here. It's probably not the case. If you have someone that gives you a phone call, so I would say just in that context, and I think there is nothing as safe as asking personal questions that only the person that they pretend to be could answer. - So you think it's gonna be a fact of life and therefore we need it all to be just more vigilant? - Yeah, I mean, it was already the case with emails, right? And I people just need to be much more vigilant and probably we'll find ways to have a authentication, for example, on the other phone side that it's actually a person that is calling and so on and so forth. - But on the privacy front, if I clone my voice with Gradium, then you keep control of it. - Only you can use it. You need to own it, but nobody will be able to use it. And so it's only for your own usage. And I think what even till you could also do, which was done before, is a lot of people, if they want, to opt in to share their voice with the community, so that they can also get financial compensation if their voice is used by other users. So I think this one is as a lot of value because it's a good way of sourcing a lot of voices. Eventually, again, if we omit the case of replicating familiar voices from licenses or existing people who give their voice knowingly to create specific content, I think voice design is going to just remove this issue. Because then, again, people typically are going to clone the voice of someone, but what they want to do is someone from a specific gender, specific demographics, age, accent, and so on. And so they could just fill this information and get a few propositions or voices that will fit their need. And that will remove the need for voice cloning. - So as we get towards the end of this conversation, just like a few quick ones, one, there seems to be an emerging discussion about the intersection between voice and screens or image. Is that something that you focus on? - There are several things in particular. So if you want to do speech understanding, having access to the video can help a lot. Again, for the realization, for example. If you're filming, like if you're listening presidential debates, we've seen all candidates, it's awful to understand. If you watch, you can see who is picking at each time and so on. So audiovisual understanding is much stronger than audio and starting in itself. I think now also what is interesting is, so we see that in a video generation, like with a view of three from Google, now there is native audio that is included. And that's why I think it's for video generation, I think the most natural thing is integration of video and audio. I don't think it really makes sense to do video to audio generation separately, because typically the data exists as a multimodal signal, right? When you have a video, you have the audio track as well. So you might as well just exploit both to train your models. So however, now, so for example, we released for the Valentine's Day, a small app called Bridger Clone. Now if I put my face in a photo, I put two videos, I say make a video, it's going to make up a voice that it thinks sound like me based on my appearance, my age and so on. And what do people find that? So it's it's called Bridger Clone, that I like Bridger tone, like the supply and word. Exactly. Bridger clone.app. Yeah. Okay. And so what it allows you to do is, now you can clone your voice, you can record a small, a short love message. And now you get to video of you in your voice. And this again, I think that's why it's useful to have a voice that is treated as kind of a separate component, because now you can have much more control on the actual voice of the virtual character, rather than just having a likely voice that sounds like it could be yours. Yeah. Okay. So we talked about cloning, obviously cloning is just one of the mini apps, so like ultimately just to play back and drive it home, you provide building blocks and models to create all sorts of different products based on voice, whether that's customer service use case or like any kind of interaction translation. Like all sorts of other things. And what is interesting is, so at QTA in two years, we were able to do a conversation, translation, TTS, Pistotext, always competing with much bigger and much more mature teams. It's still the same at Gradium and one of the big strengths is that we have kind of this fundamental framework for the generation around audio-language models, and it's extremely flexible. So I think one of our strengths as well is our ability to do like a new task. And when people come up with new needs, whether it's about annotating, value stuff, or generating value stuff, all of this can be cast pretty easily in our framework. So then if we see that there is a huge demand for speech separation, it's pretty easy for us to do it for voice transformation, for accent transformation, whatever. It's that's also why we are always interacting with the developers to understand what they want. We're also now giving access to alpha models that can do stuff that are still experimental but are world-premiers. And yeah, I think that's something I'm pretty excited because we can be much more creative than just Pistotext and TTS. Obviously, these are kind of the master tasks where we want to be the best in the world because that's where most of the opportunities are. But at the same time, we can do a lot of fun stuff that is completely orthogonal to that. In particular, around transformation, audio effects, and so on, I think that can be. There are a lot of things that can be very cool. Great. And last theme or question, you building this company out of Paris? Yes. Obviously, you're building the company very much in a global way. And you have multiple customers in different geographies, including very much the US. We recording this today in New York. You're flying out to San Francisco in a few hours. Any thoughts on the current state of the French AI scene and the European, I guess, AI scene? There's always this fun back and forth scene from the US, a combination of occasionally admiration, pretty often mockery, and the fact that the Macron, whoever runs his social accounts, sort of misfired the other day by saying that he was going to allocate 30 million euros to AI when your reality meant a specific program to attract a few academics to France. But on the ground there, what is your sense of the current state and the strengths and weaknesses of the French? A lot of things to say about that. I was born and raised in Paris. I did all my careers there. Facebook arrives when I started my PhD and then Google Brain moved to Paris and Google DeepMind afterwards. I'm what we can call terminally online in the sense that I love the very mean memes against Europe. So this guy, not even about his name, it's like a fake Swedish name. And he keeps posting about how, you know, he has after only 20 meetings. He has contributed 10,000 euros check for our backgrounds. And it's a compliance first company. And you know, but I love it. It's very mean. And honestly, I love it because it's very mean. I'm like, I love when people have so much time to spend just to be mean. I mean, I think it's quite, you know, I respect that. But at the same time, it's so far from reality. So, you know, French AI and European AI. So European AI is mostly French AI to be fair. As I was also Germany, but a lot of it is in France. And before French companies, as I was French talent in American companies. So again, a lot of the current audio generative models of Google global company, obviously, were developed between Paris and Zurich, actually most of it. Lama was started in Paris. I know which is the most groundbreaking vision work from Facebook was developing Paris. A lot of things have been developed in Paris that are not seen as Parisian because they were made in the American companies. And now I think the field in Paris is like the talent is so dense. And the people are extremely strong and extremely committed. And the best signal that proves it is a way used to have Facebook and Google in Paris. Now there is OpenAI, there is Anthropic, there is Coheir. Pretty much everybody is opening in an AI lab in Paris. And the reason why they do is for the talent. So I think we have everything in France to develop global companies. In particular, in AI, which is an economy really built around talent that's a perfect deal for it. Also as a French guy, I really don't want us to screw it up in France because in a way, when you see the most successful French companies, it's luxury. Can't stop. And so yeah, I-- I really want AI to be, you know, that's a field where France can make a big difference. And that can become one of the most biggest drivers in the European economy. If it doesn't work, then Europe will really have to look itself in the mirror and be like, how could we screw it up with so many strong people? Because you know, the people are here. And you know, capital, we can get capital in Europe as well. I think all the conditions are there for it to be competitive. In a way, the more people are mocking Europe, the more it can make the people who mock overconfident about themselves and everyone who is overconfident eventually get displaced by the underdog. So, you know, I mean, people should always, you know, like when I was the 9.6, like, larping on Twitter, it's ridiculous. Like people coding at the gym, like code, and then go to the gym. You're not doing a good gym and you're not doing good codes. So you're just pretending. So I think we don't have a culture of pretending and that's whole Europe from the most Western to the most Eastern part. We tend to be a bit more pessimistic maybe and a bit more down to hear some about our impact and the challenges and so on. We don't try to cigar code things. We rarely look into the aesthetic or how happy about our work. But you know, it's a good discipline. So I think the results kind of speak for themselves. What I think, you know, so, Mr. Ralf, for example, was, you know, sometimes it's not because it's not considered the frontier lab like open air and tropic, whatever. Okay, but look at the staff and the resources and so on. I know the best people from mistral, they're, you know, they can be compared to the top of the top of the biggest lab. So then there's a question of scale and scale of resources indeed. So what you, the unfortunate tweet you mentioned about the 30 million euros, you know, is one of the, but yeah, I mean, others and that people are, there are a lot of very strong people. And you know, they have done also a lot of great things for American companies. I mean, Yanlequin is one of them obviously, but it's not only him, Sami Benjou, who used to lead brain, is now leading Apple MLR. The amount of French people in the leadership of Big Tech, yeah, research is very large. Wonderful. Well, I love that. Yeah, I get carried a bit when I speak about that. I love the fighting spirit. And this is a wonderful way to end it. So, Niel, thank you so much. This was a terrific, really appreciate it. Thanks.

Podcast Summary

Key Points:

  1. Voice AI is currently experiencing significant advancements in latency, naturalness, and accuracy, making it more practical and enjoyable to use.
  2. The field historically lagged due to a lack of visionary interest and prestige compared to computer vision or NLP, despite early deep learning successes in speech.
  3. Expertise in voice AI is exceptionally rare, requiring a blend of machine learning, signal processing, telecommunication, and psychoacoustics knowledge.
  4. The future of voice AI involves moving beyond turn-based interactions to full-duplex, always-on conversations and integrating emotion and context awareness.
  5. Voice is becoming central to new hardware interfaces (like glasses or pendants) as a primary, screen-free interaction method, though social adoption in environments like offices remains a challenge.

Summary:

The discussion highlights that voice AI is having a breakthrough moment after years of lagging behind other AI modalities. Recent progress in reducing latency and improving naturalness has made AI voice interactions more convenient than ever, even rivaling human conversation in some cases. Historically, the field attracted fewer visionaries and less prestige, despite early deep learning successes in speech recognition.

This has resulted in a very small global pool of experts, as the discipline requires a rare combination of skills across machine learning, signal processing, and psychoacoustics. The future direction involves developing full-duplex systems that allow natural, overlapping conversations and incorporating emotional intelligence. Voice is also poised to become the main interface for new, screen-less hardware.

However, challenges remain, such as enabling AI to function in noisy real-world environments and navigating social norms around speaking to devices in shared spaces like offices.

FAQs

Voice AI is experiencing significant progress in latency, naturalness, and accuracy, making interactions more enjoyable and convenient. Additionally, the integration of voice with AI agents that can perform actions has expanded its practical use cases.

Voice AI did not attract as many visionaries in machine learning, leading to lower prestige in research conferences. This resulted in fewer experts and slower progress compared to fields like computer vision or natural language processing.

There are very few experts, estimated to be around 50 individuals globally. This small group has made disproportionate contributions due to the specialized knowledge required across machine learning, signal processing, and psychoacoustics.

Key challenges include eliminating latency through full-duplex conversations, enabling emotional expressiveness, and ensuring the AI can understand and respond appropriately in noisy environments. Integration with hardware and product design also poses significant hurdles.

Voice is becoming central to new hardware like glasses and pendants, replacing keyboards and screens as the primary interface. This shift may lead to redesigned environments, such as offices, to accommodate voice interactions.

A neural audio codec uses machine learning to compress audio more efficiently than traditional methods like MP3. It removes imperceptible information based on data, enabling higher compression rates and better quality for real-time communication and generative models.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.