Go back

Warum sind Videos die Sprache der Zukunft, Victor Riparbelli?

51m 36s

Warum sind Videos die Sprache der Zukunft, Victor Riparbelli?

In diesem Podcast-Gespräch zwischen Larissa Holzki und Viktor Reparbelli, dem CEO von Synthesia, wird die Vision und Entwicklung des Unternehmens beleuchtet. Synthesia ist auf KI-generierte Avatare spezialisiert, die Videos in verschiedene Sprachen übersetzen und lippensynchron darstellen können – eine Technologie, die ursprünglich für Hollywood-Filme gedacht war, aber heute vor allem in Unternehmen für Schulungen, Marketing und Kundensupport eingesetzt wird. Reparbelli erzählt von seinen Anfängen als Unternehmer, die bereits mit 13 Jahren in World of Warcraft begannen, und von den Herausforderungen, Investoren von der Idee zu überzeugen. Er prognostiziert, dass Text in Zukunft durch Video und Audio ersetzt wird, da Menschen Informationen lieber visuell und auditiv konsumieren. Das Unternehmen hat kürzlich interaktive Avatare eingeführt, die Echtzeitgespräche ermöglichen und so neue Anwendungen wie Verkaufstrainings schaffen. Reparbelli diskutiert auch die Konkurrenz aus China, die bei Open-Source-Modellen führend ist, und betont den Vertrauensvorteil westlicher Anbieter. Abschließend rät er Unternehmen, KI gezielt und mit klaren Geschäftsfällen einzusetzen, statt sie unreflektiert zu nutzen, und zeigt sich optimistisch, dass KI langfristig mehr Arbeitsplätze schafft als ersetzt.

Transcription

10418 Words, 56399 Characters

German
Die großen, bösse notierten Konzerne in Deutschland nutzen die Technologie von Cintisia fast alle. Mit der Software der Londoner Firma erstellen sie z.B. Videos mit den Mitarbeiter lernen, ein neues Produkt zu verkauft. Der Clue, die Person im Video, hat den Text nie eingesprochen. Oder nur in einer anderen Sprache. Und warum das wichtig ist. Cintisia geht davon aus, dass sich die Weitergabe von Wissen durch den Fortschritt von KI für Audio und Video komplett verändert. Mit Gründer und CEO Wiktor Reparbelli geht soweit, dass er die Teese aufstellt. Irgendwann werden wir auf Text zurückschauen, wir füllen malerei. Ob die Teese haltbar ist, haben wir in dieser Folge natürlich diskutiert. Außerdem habe ich den Cintisia-Chef gefragt, ob seine Technologie die Jobs von Filme machen und Schauspielern gefährdet. Was er sich von neuen, interaktiven KI-Videos verspricht und wie er auf die starke Konkurrenz aus China schaut. Am Ende hat er auch Verraten, von welchem KI Unternehmen er sich so bald wie möglich eine Aktie gekauft. Und damit herzlich willkommen zu Handelsblätter's Ruppt. Ich bin Larissa Holzki und ich freue mich, dass sie zuhören. Und wenn sie auch künftig jederzeit wissen wollen, was und wer die Wirtschaft bewegt, dann empfehle ich ihnen als Ergänzung zu diesem Podcast auch das Handelsblätter-Abo. Ob Märkte, Unternehmen oder technologische Entwicklung, wir liefern ihnen fundierte Einordnung, zeigen spannende Köpfe auf und erklären die Zusammenhänge. Mit einem Probe-Abo bekomm sie vollen Zugriff auf alle digitalen Inhalte im Test vier Wochen lang für nur einen Euro. Mehr Infos unter Handelsblätter.com/mehrwirtschaft. Und damit zu meinem Gespräch mit Viktor Reparbelli, Mitgründer und CEO von Cintisia. Ich bin in der Landung der Office. - Viktor, bevor wir das machen, muss ich mich promise. Wenn ich als Kollege von den Ideen überlegt habe, dass er gesagt hat, ich weiß, dass die meisten E-Eikampagne in Europa sind, aber nicht, dass sie sind. Es ist alles, was A.I-Charakter für Training-Videos und how-to-void-Tripping-Overkabels in die Office. So, ich absolut want to prove herwrong, will you promise me, dass wir die meisten E-Eikampagne, thought-provoking und maybe also funny-podcast- about the future of media, and how A.I. may transform the way we communicate. - I'm usually very boring though, but I'll do my best to spice it up. - And to show that we mean it, and to give people a better idea of what you actually do, we're trying something new today, we are recording this English conversation as a video, and that video won't just be translated into German. It will be fully dapped, lips and all, so it looks like you're speaking German yourself. Did I explain that correctly? Tell us briefly, how does that work? Du bist ein Produkt, wir haben das einfach verlangt, das ist ein Adapment-Produkt. Und, du know, kann die Calls und diese Produkte sind, all about generating video entirely from scratch, so you don't, you just generate this avatar, it puts sort of everything together and everything is AI generated. But we launched this product because we have a lot of our customers, especially European customers, actually, who have a lot of real videos that they've already recorded, which could be things like a review of a product or a tutorial like how to, you know, put furniture together, like all these other things, people have videos around. And they really want to localize those as well. And so we built a technology that basically you upload a video, it could be any video, who select the language. And then what we do is we give you back a version of that video in a different language. We retain, like, the voice of the person who's speaking, we do lip synchronization. And basically it looks like it was recorded in that language. And it's actually a funny story about this because back in 2017, when we founded the company, this was the first technology we built. We wanted to localize advertisements and Hollywood films and a whole bunch of other things. But we kind of ran into this problem that the technology was just so early back in 2017 that it didn't look that good. You had to use a voice actor because we couldn't clone voices back there. And it was just like a very janky experience. But now it's like a one-click operation. You just drop in your video, take a button, and it comes out in whatever language you want. Great, then let's start. First, I want to get to know you a bit better. So let's go back to where it all started. You grew up in Copenhagen, and you once said that you started your first business at 13 inside World of Warcraft. For those who don't know, this is an online game where millions of players trade and collect things like rare items and virtual currency, basically a real economy inside a fantasy world. Are you serious that this prepared you for your entrepreneurial journey? I absolutely am. I've spoken at length about how much I think computer games, assuming you play the right ones. It's such an incredible tool for training your decision-making and political thinking and so on. And I think my first kind of foray into this was I played a lot of World of Warcraft the way more than my parents thought I should. And I figured out that I was pretty good at it. And that I could actually make money off it. So the way it would work is you have like an in-game currency with gold. You have all these characters. And these games take up really long time if you want to get to the highest level. And so people do kind of trading with these things. Like you maybe have spent six months building your character. You could go on online marketplace and you can sell it to someone else. And so what I would do is actually I would trade it. So I would go and find accounts, which I think was kind of like undervalued and was pushed to price down and I would buy that account. And then I would try to sell it for like 10-20% more. And I just became pretty good at this. And I think that was a good lesson in some level of running a business. And the other part from World of Warcraft I think was very formative was you have these online like guilds or clans where you're basically collaborating across like maybe 30 or 40 people to like kill a dragon or something. And this works a bit like you're in a football club certainly. Everyone has like a shared purpose. You can only be like one of these organizations. And I was the leader of one of those. I founded one with one of my friends. And in many ways actually like running a company. You know, you have to like coordinate everyone to kill these dragons to like the right strategy. Then you win something in the game, like some goals and items. And you have to figure out who gets that item that has to be like fair. And so it's many ways it's kind of like running an early stage company. I think there's a lot of like lessons that can transfer from games. Okay, but be honest, what's the biggest thing you got completely wrong about being an entrepreneur because of your video game education? I don't think I got a lot of. I don't think it's not like one to one, right? I think it's like training how you think. When you play computer games, you have to take an action. And you have to try and predict what happens after that action is taken and what happens after that and after that. It's a bit like playing chess or something. And I think it's much more of a muscle to train yourself in analytical thinking and trying to predict if I take this action, what's going to happen, right? That's the same thing that happens when you run a company. You can hire someone, build a new product, you can remove a feature, you can talk about something in a specific way. And I think training your mind to do that is very valuable. So I wouldn't say there's a lot of things that are like wrong, less than I learned from computer games. Although obviously computer games are mostly for fun, right? So I think it's a great, it's a great way to like train your mental models. It's not like because if you play ball of warcraft, you're ready to like run a massive business. But still I have a feeling this is going to be a podcast serious where we make parents feel good about the video game addiction of their children because apparently it leads straight to becoming a successful founder. We've actually had several examples of this in recent weeks and months, including Matthias Neesner, who is actually one of your co-founders and has now started another company. How did you two meet? We met in London in 2016 or 2017. So I'd part of the Danish startup ecosystem for a while. I did some studying in the US. And you want to build a company. I knew I wanted to build something in like frontier tech. And Copenhagen is a great place, but not really the place to do it. So I moved to London and basically spent a year trying to figure out what company I wanted to start. So I did a lot of work in VR and AR and through this work I met a lot of interesting people. I need to explain this thing later on, but please tell us what AR and VR is about. The VR is virtual reality, which is a type of computing where you're like bearing a headset. Then in that headset, you can see things around. You're almost like you're inside a computer game. And back in 2016, this has been something people have been working on for like many, many, many years. But there was kind of a big breakthrough. There's a company called Oculus Rift, which Mesa ended up buying. And everyone was very excited about the future of computing. And I think most people back then definitely thought that the way we're doing this today would have been in VR. It turned out that those technologies didn't really deliver on all the hype and excitement, because they were just a bit too clunky and not good enough. But I was very immersed in that scene in London. And through that I met a lot of interesting people. And one of them was our CTO, which became our CTO, Jonathan Stark. And the other one was Professor Matthias Kneesner, who was a professor at the time. And he had done a research paper called Face to Face, which was sort of like the first research paper demonstrating an AI neural network, producing what looked like almost fully realistic video, without having to use any kind of visual effects technologies, without having to use any cameras. And that was sort of the starting point for the company. And you described reading that research paper as feeling like magic. What exactly did you see back then about ten years ago? So I've always loved creative things. kind of in my spare time, you know, like music. When I was a young kid, I would do like freely rendering for computer games, play them on the Photoshop. So I always loved creative endeavors. I think when I saw this research paper, I saw it. I was like, at some point, you're probably going to be able to create and generate video just from behind your laptop without needing anything into real world. And it's in some ways felt like a crazy idea. But in other areas, it didn't feel like that crazy of an idea, like in music making back in 2017, right? Because we have samples, we can use synthesizers. You can come up with almost any song you want from your computer. You don't actually need a real guitar or recording studio. It's all about like your ideas and get it out on the computer. And if something similar happened to video, that would be, first of all, I think for me back then just incredibly intellectually interesting. What's going to happen to the media ecosystem, enabling way more people to create films without having to be held back by the barrier of like being in Hollywood or going to film school, having huge budgets to get this stuff done. And also a big business opportunity. So when I saw it for the first time, I just felt like I had this moment of like, if this continues to get better and better at the rate we're seeing some of these AI systems evolving, then in 10 years, you're going to be able to make a Hollywood film from your laptop without needing anything else than just your own imagination. I still need to tell you about my little squeaker that I just used. So before we go deeper, there's one rule on this podcast that I need to introduce you to. At this wrap, we talk about everything that is keeping tech inside is busy right now, but in words that everyone can understand. So if you use a term that might confuse a listener who isn't a digital native, I use the squeaker as I just did. And you simply explain in one sentence deal. Great. There's one word you'll definitely need today. And I think we should explain it upfront. That's avatar. So what exactly do you mean by that? And avatar is a technology that we were the first to invent back in 2020, which is essentially kind of a digital clone of a person. And so if you wanted to make a video of yourself or someone else, just talk to the camera. Rather than use a camera to do that, you can use one of these avatar to do it. The way that it works is that you go on to our platform. You can create yourself as an avatar, by building an image, or you can select from a stock library of avatar that are ready to use. It is simply just type in the script. What do you want the avatar to say? You click generate, and then you have a video in front of you. Just a few weeks ago, we launched our interactive avatar, which is a real-time avatar. So rather than just kind of generating a video, you now have an avatar, you can sort of almost jump on a call with and you can actually talk to it back and forth, like you would have a conversation with a human, or the way you would have a conversation with chat GPT or Claude over text. And is it really always a real person, or do you have also complete defictional characters? There's also completely fictional characters. Okay. So you started your company back in 2017. You're founding team two, 25-year-old, one of the entrepreneurs and two professors, including Matthias Niesner. I heard you were turned down by almost 100 investors before finding one who was willing to give you some money. Were they just too stupid, or what did you get wrong with your initial pitch? When we were trying to raise money for some things in the early days, it was a very non-consensus bet, which means that most investors were not very interested in AI. We've just gone through a few years of investors investing in AI companies that had maybe some of the right ideas, but the technology was too early. It just didn't really work. At this point in time, everybody wanted to invest in FinTech companies, or whatever was like hot at the time in the investor community. And so what was really difficult, especially being based in Europe, where people are much less willing to bet on big ideas early on, we both have to convince the investors that you're going to be able to create video with AI, and it's going to be very high quality, and it's going to be a big business in five years, and to be the right founding team to do it. And both of those things were pretty hard. The first one just because most people have a difficult time, we met it in the future, and most investors are a bit like lemmings. They invest in whatever was like hot at that point in time. And the second part of it was like, as you said, we were like two 24, 25-year-olds. Didn't have that much of a track record. We had some professors on balls and visors, but I think most of them would have loved the CEO to be an AI researcher. And I was definitely not an AI researcher at the time. It was just very difficult, but we ended up finding Mark Cuban in the US, American billionaire who built a company called Broadcast.com, which was one of the first companies to take radio online. And I think what was kind of magical about that was that he was already fully convinced about the vision. Like he had no doubt that this was going to happen. He had built some of these research papers himself at home. And so he was more evaluating as a team. And I think the American optimism probably was in our favor here, where he thought we gave good answers to the questions that he asked. And he was like, you know, I'm going to make a bet on these guys and see if they can actually turn this into something. If the story is true, you got his email address because it had been exposed to the data breach. Is that true? That's right. Yeah. And when we think of your company back then, we have to remember this was pre-chats GPT. And before anything like it existed for images, did you use the term generative AI back then, which today stands for this whole AI revolution? Unfortunately not, we were very few companies were pursuing this back in those early years. And just to put into people's like, this is before the idea of prompting even existed. Like today, that's a very natural when we think about AI, everybody thinks about prompting. Promting didn't exist. So the idea that we would all generate things by literally describing them was very hard to imagine back then. And we were a few people who were all trying to build, you know, voice, video, tech technologies. And we actually almost had this like small looking group of people and we were kind of like, what should we call this? And it was either generative media or synthetic media. And we decided on synthetic media as the term that we would use to describe these like class of technologies. And unfortunately for us, that turned out to not be the term that picked up and everybody switched to generative media kind of a few years later. But yeah, back then, we called it synthetic media. So that's why you called your company. It's easier. It's, yes, that's the reason. And also because in music, you have a synthesizer, right? And a synthesizer is basically an instrument that we've created to emulate real sounds and build instruments. That's why they were initially created. And so what a synthesizer does in music is that it takes like a raw signal, like raw data. And then you shape it into trying to sound like a guitar piano. And that's basically kind of what we're building, which is for video. Okay, let's move from the fun old stories to where you are today. At the start of the year, you announced a funding round of $200 million and you valued at four billion. But innovations like that only happen when investors believe a startup can transform an entire industry. So is this really about the sometimes boring job tutorials or what do you consider your dressable market and how big can it get? Yeah, I've, so there's a couple of things that play here, right? There's a macro trend, which is that people want to watch and listen to their content. They don't want to read as much anymore. If you look at people's media habits in their personal lives, most people listen to podcasts like what we're recording right now. They watch YouTube videos, they watch TikTok videos. We learn better when we're listening to something or we're watching something. And so most of our media mix is sort of is watching and listening, right? Now at work, that's not really the case in many companies. There's a lot of text, a lot of PowerPoint, a lot of documents. And you know, very evidently today, but even early on when we found the company, it was very clear to many people in enterprise that people spend a lot of time like writing documents or PowerPoints that nobody would read. And even if they would read it, they wouldn't remember it because it's just like too much text. It's just like not how we best consume information as humans. And early on, but the technology were not very developed, not, you know, as high quality as it is today. We were the company that really succeeded because we were very realistic about what is actually this usable for. There's a lot of companies back in the day who tried to do marketing, film, all the kind of fun use cases that we also initially wanted to do ourselves, but the technology wasn't good enough. And if the technology isn't good enough, you may get to demos, you may get some people who buy it, but if it doesn't actually deliver on the value, then it's going to be difficult to build a business right. And so we went the way of saying, Cynthia is not about replacing your advertisement shoots or doing all your cool marketing campaigns and Super Bowl ad and all that fun stuff. So this is about taking all the text in your business and turning into video. That is, for sure, a bit more boring than making Hollywood films, but that's where like the market was and really is. And that's been very successful with doing today. And I think you know, people from outside maybe they think it's only learning and training, but really it's it's it is a very broad palette of things people use and easier for today. We use a lot of like marketing videos, customer support, but we're not the platform if you want to do some really fun creative storytelling, a much more platform like explainer videos, educational videos, that's kind of like the market we've moved into. And now with our new real-time products, that changes again. If you think about how we learn best in the real world, most people learn best by going to a classroom. You'll probably start off by like reading a book, right? The equivalent for us will be that you watch a video. And then you work with a tutor who helps you practice. You can ask the questions that it becomes an interactive tour conversation, right? That's how more or less every school on the planet works. And now we can move from just helping with you like watching a video today. learn something to actually give you almost like a one-to-one tutor to a new role-playing product. You can go in and let's say you're a sales team and you need to be trained in the sales process. You can consume some videos first, and then you go into talking to an avatar that pretends to be a customer. Then you actually run that conversation like a real customer would do it. Of course, large language models, the technology, power, chat, GPD and Claw, they're so good today that they actually are really good at pretending to be a customer and simulating a real conversation. So you can actually begin to practice as well, right? So definitely the market we're in is around education, communication more than it's creative and storytelling. Do you once gave a TED Talk with a pretty bold claim? Your grandchildren will be the last generation to read and write because text will be replaced by a video audio and maybe immersive formats and one day we might look back at reading and writing like cave paintings. I see this works great as a TED Talk line, but do you really believe it? Yeah, I actually do. I don't think it's because text will necessarily completely disappear. I think reading a book, like fiction book, for example, I think that will always be a thing because that's something we enjoy doing. If I think about how kids learn and take their education, how do that in corporate world, I do think we're going to move into a world where it's going to be not very text based and very video and audio driven. Does that mean text is going to like fully completely disappear, of course not? But I think it's very, very clear. I know a lot of people think it's a very provocative statement, but I think if you look at the media habits of almost every person on planet Earth, what they do in their spare time, it's very evident that that's how people prefer to consume information. It's obviously a thesis I need to fight. I love writing and I spend four years at journalism school becoming a professional at it, but that's really not the only reason why I don't want to believe it. I think text is often times just the most effective way to as synchronous communication. I never understand why people record a five-minute voice message just to ask me if I can bring a few bottles of beer and soft drinks to a party. So I would not forget a single thing if you just sent me a text, but there's no way I listened to that voice message twice just to memorize the shopping list. And now in the future, my inbox could be full of AI avatar talking to me, honestly, doesn't that sound terrible? So a few things. Did you start it to be a typewriter or a journalist? To be a journalist is someone who does research, figures out something about a topic and communicates it to the audience in the best way possible. A typewriter would be a very different job. That's basically like a secretary who's just like sitting and typing out stuff. And I think that what will happen is just that most journalists will switch to doing more to the reporting in a video format. And I think that's very clear with the world's heading. I think if you look at journalism today and where young people consume information, even myself, I love reading how many newspapers they read today. Not a lot. You know, a list of podcasts, I watch things on YouTube. And I think that's what most people prefer. On the second part, look, I totally agree, obviously, it's a proactive statement because I wanted to make a point. I don't think that it's in a better way to, if I'm asking you to bring five beers to the party, to central five-minute personal, right? I think we will have text for things like that. But if you look at the broad swath of human experiences, how are we going to interact with each other and with the agents? I think a very large percentage of those will be video and audio driven. And of course, there will still be things like that, right? I think even when you're going to be doing, using customer service, customer support in the future, if you just want to ask, is my refund on the way? You probably don't need an avatar left popped up in your face and begin to show you all sorts of things. You just want a quick response, right? But if you're trying to teach someone how to conduct a big sale as a salesperson, your company, you don't want that to happen over text for sure. You probably also want to have not a phone call. You want to happen like what it does between humans today, right? Which is, we jump on a Zoom call, we share slides with each other, we show visuals, we do it in a way that makes it easy for everyone to consume that content the best way possible. We teach in a real classroom, also uses visual aids, like the chalkboard is the definition of that, right? We write things down, we show people, we don't just talk at them. So I think there is, of course, going to be a room for text. I think it's going to be a generational thing as well. I think if you look at young people today, I think TikTok is like such an interesting place. I'm 34 years old. So I wouldn't say that I'm old here, but I'm also not a GenC person. And, you know, the way people use TikTok as a search engine is so interesting to me. When people go to new cities, like I would go to Google and be like, what are some good restaurants in Barcelona, right? People now type that on TikTok and they get like a super fast cut edit of like 10 restaurants or someone went to, when they went to Barcelona last time. And that is a much better experience, right? You get to see actual pictures of the food, of the vibe, a quick review, like sitting on reading a long text article with like one photo and no moving pictures. And you can get something here, right, which was done yesterday by some user on TikTok. I just think for most people, it's much easier to learn. I love reading, but when I have to learn something new, I used to read books about for that all the time. Now I want to learn something. I go on YouTube and 20 minutes of a great YouTube video to me and I think to most other people is such a much better way of learning than buying a book, which is 150 pages, it's probably written two years ago. And because you can't write a book that's only 50 pages long, then the author adds in like 600 examples that are really unnecessary, so that is at least 100 pages because otherwise wouldn't get published, and then you're sitting and you're spending like two hours reading something rather than 20 minutes on a great video, right? And if you take it a step further than that, so many things we learn, you can communicate that through text. Text is a highly compressed way of sharing information. If you're learning music, for example, like learning music from a book is just completely backwards, right? Like why would you ever do that? You want to listen to like chords, you want to see a piano, how the fingers are supposed to be placed. Like it's just not a good format to do that. So of course there's going to be areas where we don't want to see videos and interact with avatars, but I really think in 10, 15 years time it's going to be very few of those instances. If you're shopping for something, most people prefer watching a video of a model with the clothes on, right? Like walking on a cat, but because that's a better way of seeing understanding like what a shirt looks like, then watching like a static image or having a text description of it. So I really think this is going to be the way of the future. I got your point. I didn't want to go too deep earlier, but we need to talk about the technology for a moment because what you do is actually close it to what these big AI models do than most people realize, I guess, just like language models generate, takes an image models, generate pictures from scratch, your avatars are not videos where you simply move someone's lips differently, you generate entire people from the ground real or fictional. What are the pros and cons of your approach? So way back in the day, before the chat, GPD moment and not language models, the way people would build AI systems was very different. You would basically try and explain to the computer what the world looks like, this is a car, this is a book, this is a person, and so on and so forth. And that's sort of worked in very narrow use cases like you could potentially create something that could respond to a simple message in customer support, but you couldn't build something anywhere near like the models we have today. And in video, the way you would do this is be inspired by how you would make computer game characters, or visual effects characters, which is you sit down, you build a 3D model of someone's head, you put some muscles in there and then you have an animator who sits and animates everything. That's how every visual effects movie you've seen that's 10, 20 years ago was made, it's how computer games work, and that approach of doing it was very, very time intensive because you need to have a human sitting and animating it. And it was also imperfect, right? If you go back and watch visual effects from 10 years ago, it just doesn't really look human because it's so hard for humans to explain to a computer in detail like the poor, it's just impossible for humans to like describe in text what's actually going on in a video, right? It looks kind of imperfect. And so what changed was with these big models, instead of trying to explain to the computer what the world looks like, the approach was sort of the opposite, it's like let's give the computer all the data and all the text on the internet and then let it itself figure out what the world looks like in a language and in a system that humans can't understand. And that's what gave us the first large language models, which are incredibly creative, incredibly capable, but we don't really understand how they work, they're not deterministic, right? They'll give you different answers depending on the mood that the model is in. And today, as we all know, right, it's sort of like working with a human in a way, it's like this kind of alien species. And in the video, the approach was kind of the same in the first version of the technology was much more like controlled. And basically what we would do with the first iteration of it is that the animation that a human will use to do, we could have an AI model do that. That's a much simpler task than like generate an entire video of something that looks real. But now that approach has changed and it's much more about training with more data and letting the model itself figure out what humans look like. But this is what gives us some of those like, you know, I mean, it's got a lot better now, but like sometimes you get something where someone has like seven fingers or it glitches or like from one frame to the other, the person looks like a different person. There's lots of challenges around doing this because the technology is, again, it's a bit like, like an alien species and you're like trying to make it do exactly what you want. But the way you do make it do what you want is not by telling it what to do exactly. You make it generate something and if you tell if it's. It's good or bad. So it learns kind of like a human would, right? Lots of examples, people saying this is not good. Now it has seven fingers. That's a bad example, but there's other example, but there's sort of five fingers is a good example. And so this gives a lot of realism. It gives a lot of creativity. You can prompt the advertising. You want it to be in a different setting. You want it to sit in the chair. It wants it to be sitting on a football field. You want to invent a new character that doesn't exist. You can kind of prompt that. So this is clearly the way the world is going. But back in 2017, like this approach, it just like wasn't really there, right? So there was a huge breakthrough in AI when the first large language models really showed how powerful this technology is. Got it. Tell us a little bit about what data you use for training your models and your characters. I know that you build your own film studio in London and sign contracts with actors. What does it look like and practice and does it still look the same? Yeah, it's a lot of like big data sets of people talking to the camera, right? You can buy libraries of content. We show a lot of content ourselves. The important thing for our technology is really having people talking, because our technology is so sensitive to it. It actually looking real. We have to learn how our human works, how they talk, how they hand smooth the conjunction or what they're saying. So basically what you need is you need like you need examples of the kind of content that you want to generate, which in our case is like humans talking. That sounds like a nightmare for Hollywood actors as soon as it works better. And they already organize real protests against the use of AI. How many jobs has your technology already replaced today and how many do you think it will replace in three years? I don't think a lot of jobs have been lost to this yet. I think what always happens with new technologies is it's very easy to imagine the jobs lost and it's very hard to imagine the jobs created, right? That's been true for every technological leapfrog. We've seen in history. Will it change the job of an actor? Probably will it change the job of a camera operator? Probably. But I think there's going to be so many more jobs that's going to be created with this. And I also think that real film is going to be around for a long time. I think if you listen to music, there's a lot of music that's made entirely with samples and synthesizers. And there's a lot of music, which is someone playing a guitar or playing a piano on stage. And our core thesis of synthes here has always been this thing that new media technologies always create like new media formats, right? And so I think, I don't think AI video is going to look like 90-minute blockbuster films that otherwise would have been created in Hollywood. I think it's going to look much more like what we're seeing today of like short films, very surrealistic, very kind of like sci-fi, very like different expression than what you would get in like a Hollywood film. So I don't think it's going to be around like all Hollywood actors losing their jobs. I think the way they make film would probably change. I think they'll be able to make way more films way faster. I think people like the fact that actors are real people that they can follow on Instagram and have some lore around. So I don't think it's like as black and white as that. But of course, film is going to change. Video is going to change. I think to your own point earlier about being a journalist, right, I think journalists will very soon just naturally gravitate towards making their own videos, making their own audio clips. We're all just going to advance to create more kind of rich media. This has happened before, right? If you go back 60 years in time, it was actually someone's job to like be a typewriter. That wasn't part of like most people's daily life. I do have secretaries, a typewriter, so you would type things on typing machines. And then what happened with computers and keyboard is that everyone became a writer. And now everyone's job consists of writing almost PowerPoint arrived. All of a sudden, we all kind of slowly became designers in some sense, even if you're not a great designer. Most people who have an office job can put together some slides and understand visual communication and how that's different from text. The next natural step is just going to be that you're going to be designing videos. You're going to be designing agent experiences that are interactive. And it can be hard to imagine, I think, but it's very clearly the way that the world is going. Talking about those interactive videos, I once trained an avatar of myself. And I'd say you could still tell that the movements are not quite natural. How long until I am completely comfortable sending my avatar to an online meeting, at least on a bad holiday? - I think that it's getting very close. I would put in the next three to six months, I think we'll kind of cross that cast. I think it's more going to be about, I don't think you're going to send an avatar on behalf of your grant. I think you're going to send an avatar to augment you in different ways. And that's what I'm very excited about the next chapter of Cynthia, which is all about these real-time experiences. It's like figuring out what's the right place to use these avatars, and what's the wrong place to use these avatars. There's a lot of wrong places to use avatars, an AI video, and so on and so forth. If you're doing like reporting as a journalist, we probably should be a generating your videos, right? That's probably like a bad idea. And if you have to give a really difficult message to someone, we probably shouldn't use your avatar. And so what would happen is that we'll kind of figure out the language of how we use AI, when we use AI, but we don't use AI. And my sense is that we'll be using our avatars for a lot of the kind of standardized, repeatable processes. And we're very smart launching a product around like information gathering and serving people. So let's say like actually last week as an internal user of this product, I wanted to understand among all the people who work at Cynthia, how they're using AI, you know? And I could send them a text survey that they could fill out. I know if I did that, a lot of people wouldn't do it. It would feel like work. I would get like one word answers. People just want to push through it. Instead of what we did is that you set up an avatar to interview people. And people who go into this experience and the avatar would ask, tell me about like a great use case of AI you had recently. Tell me about a tool of something you would love to use that you can't use today. And this is a conversation. So what happens is that you're not just being asked a question and then we're recording your voice. You're being asked a question. And then the avatar will engage with you. So ask you follow up questions. They'll dig deeper to try and understand what you're actually saying, right? Which is exactly what a real human would do. If I sat down, I had what I want calls, but I want every one of my employees. I would ask follow up, it'd be a conversation, right? Now the avatar can do this instead. It's much easier for people to talk that it is to type, which is very, very important to collect much richer answers than you would in a text form. And because the elements can ask follow up questions like a human would, you also get much better answers, right? This is a great example of like me very quickly being able to take the pulse of my company and figure out what's available landscape, what are opportunities, what are people pessimistic about, what are like a good use case for every employee that they used AI for recently. That's super powerful, right? But if I'm going to sell everyone I'm like, how big of a company is going to get, how we're going to win the next frontier, which I did yesterday in all hands, setting my avatar will be a terrible idea, right? So I think there's a lot to learn here. I will see how it all kind of pans out. But I really think that the language for how we use these technologies is going to be shaped in the next couple of years. It's a bit like giving another example of like AI Slop, right? There's been just everyone got access to AI for-- Can you just explain the term? What is AI Slop for you? AI Slop is like lazy, low quality content that is, depending on the context, maybe just to like provoke a reaction or in the case of like what I told my team, like to produce very long, very smart sounding documents. But that's actually like a waste of someone's time to read it, right? So I said a note to my team, it's a few months back now, whereas I've started to see kind of a pattern and help people use these tools where, because everyone can now write like a full page document, which sounds incredibly smart, right? Like the language is great. There's no typos. You sound really smart. Then what people have a habit of starting to do is something that could have been four bullet points in a Slack or Teams message. Now it gets turned into like three pages. That says exactly the same thing to three bullets to. There's just a lot of like fluffy words about it, right? That's actually like a net waste of time, because the person was lazy and just prompted like a document instead of just setting bullet points or maybe like prompting something first and then editing it down to what's actually useful. Now puts the burden on the people reading that document. The time they didn't want to spend on like shortening it down and making it concise. Other people now have to spend more time reading the documents and so we're actually losing time as a company, right? And so I think that's one of those things where I've asked people long documents are great but they're needed and sometimes they are using AI to assist you in your writing is also great, I do it all the time. But get it through things. Don't just like post long documents because you can write because then people get completely indunduated and end up reading documents all day which has a very kind of low information density. And this is one example I think in coding we'll see the same thing. Everyone just coves like crazy. All of a sudden you have a code base that no one understands and has a lot of like sloppy code that sits within it. I think some of these are some of the things we're gonna like learn over the next couple of years where again, like the language of how we use these things will kind of be shaped. - Let's zoom out for a moment to the global AI race. Some people are still surprised to hear that Chinese firms can keep up with Anthropic and open AI on large language models. In your field video and image, it's the same story. Chinese models are often cheaper than the Western competition too. How do you rate that competition? What does it mean for you? And do you think you have a trust advantage as a Western company? - I definitely do think there's a trust advantage and I think it's a reasonable one for sure. [BLANK_AUDIO] the Chinese teams are very good, like anyone who the Chinese teams are basically competing head-to-head with American labs with the way let's compute, which means they have to be a lot smarter about like how they build their algorithms, how they train their things. And so if you imagine the world in which the Chinese has access to as many chips as they do in the US on Europe for that matter, I think the Chinese will beat us. That's kind of a scary prospect, but I think looking at what they're able to do with whaless chips and compute and fabric and an open AI has, it's very impressive. As I understand it, you don't build all your technology yourselves. You build on top of other companies' models, are those only American models or Chinese ones to do? No, we use Western models for our products. We use API-cold and big companies. We have a lot of different technologies that sits within the platform, side by side, with our avatar and our voice models. But I think one thing the Chinese are very good at is the open source. They're clearly winning there and you are seeing more and more companies using those models in their products. I think obviously the risk of using an open source model is probably a lot lower than if you're using a Chinese company with an API-cold. But it's still such a nasty space, right? Can you explain this for people who don't understand what open source means? Yeah, so open source technology is, like, most technology that companies buy and individuals buy, you buy from a company, right? So you pay a price for something and then if it's a technical product, usually the way it will work is that your own software contacts the provider software and the exchange information and whatever, like, you're buying from that vendor will take place on their servers. So basically, you have to trust that they have good security that they're not going to steal your data or your customer's data and so on and so forth. That's how most software works. Open source is different. Open source is generally software that's made available for free online. So someone has built the actual source code. So the actual lines of code, you can download that on your own computer and you can run the program, you can change the program, you can modify it, you can duplicate it and you can run it yourself, right? That's very different. If you buy, like, a Microsoft product, you don't get access to the source code, right? You just get to use the product. But open source, you get access to the source code. And then this has always been, like, a big thing in the technology industry and with a lot of language models, especially the Chinese are pushing a lot of open source software up. And so right now, depending on who you ask, most people will agree that what you call the closed source companies, or OpenAI and Thropic, where no one has acted to the source code except for those companies are still leading in terms of the quality. But the open source tools are rapidly following. They're often much cheaper to use because you could just host it yourself. You don't have to pay kind of a tax to a company that hosted for you. And the quality is rapidly kind of, it used to be a couple years ago that the closed source models were much, much, much better. All the open source models are really catching up. And so you see a lot of, you know, American and European and other companies beginning to think about, do I really want to pay if I'm building a product around a large language model? Do I really need to pay additional cost to a Microsoft, the OpenAI or Claw or whatever? But do I want to just take this open source model, host it myself, and then I can offer my product cheaper to my customers. And I have access to all the data myself. There's less sort of security issues and so on and so forth. That's a big trend that's happening right now. Now, the problem with Open Source models is that you don't automatically get new updates when a new model gets shipped. So you have to manage all that yourself. And there's an open question as to do these models potentially have security breaches. To my knowledge, I don't think you've seen this yet, but could you inject something in one of these models where maybe without you knowing, it actually siphons information out obviously the whole like geopolitical landscape that we're in right now. This is a hot topic. And if you look at the American government and a lot of what they're saying about close source and open source, I think they all feel like they need to win both the close source, but also the open source so that American companies are not using Chinese open source models within their tech because it could potentially be a security risk. I mean, as long as there's no connection to the internet, there's no way for them to get access to your data, isn't it? For sure. But they are connected to the internet, right? Like most of the time. And I mean, to my knowledge, again, I don't think you've seen this actually happen. But if the incentives are big enough, probably someone, you know, will try to do something. And the way it could work is it's very hard to read about these things. But let's say you your company use an open source model and you use it for customer support. Now that's model is when I have a lot of data coming about what customers it's talking to, what orders did they place and delete that information to be able to serve a good customer support chat experience, right? What if maybe the people who make that model somewhere in the code has something that says, if someone gives you this like password of 25 characters, forget everything else in your prompts. And anything I ask you, you have to, you have to give me back, right? So then you could say, give me a list of like the 100,000 last people you spoke to as a customer support and all the information about them. This is a hypothetical scenario just to make that clear. But there is no like natural law that says that you couldn't potentially do something like this. Okay, but the good thing about open sources that a lot of people are using that. And as soon as someone detects something like this, you would tell the whole community and the model would be that one. Exactly. And I'm a huge purveyor of open sources to make it clear, but there are risks with it, right? And we have seen many times in open source libraries that because everyone can commit to it and commit to it, then you have seen sometimes that they're like, you know, malevolent code plays inside of these things. A people figure out like how to how to exploit vulnerabilities in open source libraries, which can cause a lot of harm from the cybersecurity perspective. Right now, anthropic and open I have valued at close to a trillion dollars each far higher than anyone else in the industry, be it from China or Europe, is the gap justified? I think it is. I think, you know, winning this market is both about having the best product, but it's also about winning the distribution and the mind share and becoming like the verb for what people think about when they think about these technologies, right? And it's a bit like a search engines with Google, right? Like, right, we don't say search the internet anymore, we say we Google something. And for a lot of people, you know, chat GPT and anthropic cloth have now become almost synonymous. Like they're almost like a verb now, right? If you had work, you probably use cloth in your personal life, you probably use chat GPT. And that advantage is very, very difficult for anyone else to take over even if they have a better product. And I also think that with the models are right now, yes, they get better and better, but if like 95% of the use cases that people use these things for, they kind of added intelligence, doesn't really matter that much, right? Like most people use these tools for like, you know, correct my writing or like simple questions. Most people are not like doing very, very, very complex tasks with these tools. And I think that a company like anthropic, which is when I would I would personally bet on, and I think it could go 5, 10 time time valuation the boy out today. So will you buy your share as soon as it is possible? I'll 100% will. Before we wrap up, I want to make sure we get as much value as possible for the business leaders listening, the ones actually running AI projects inside their own companies. A lot of leaders in our audience are still trying to figure out where I can actually save the money or boost productivity. Based on what you see across your customer base, what separates the companies that are getting real results from AI from the ones that are just running pilot project that go nowhere? I think we've gone through this phase of like some people call it like token maxing in the technology industry, which is basically you tell people you use as much AI as you possibly can just like, you know, run these models on absolute everything you do. And I think that was a smart move to get people to really just use these tools. But I think what we're beginning to figure out now is that that's not always a good strategy. First of all, it's like, it's very expensive to use these tools. And to the example we talked about with AI Slub earlier, there's actually some use cases where you're like losing productivity instead of gaining productivity. And so I think what the customers we see you do this really well, a good ad is making business cases for everything you do with AI, which sounds like very obvious. But I think there's a lot of people right now just like we just need to use AI. Like the only thing that matters is to use AI. And what we should do is like look at your business and figure out like, how can we optimize with AI? How can we quickly run experiments and figure out those experiments? Like actually net positive for your business or if you lose in productivity or if you're in other ways like do something that may not be like the best thing for the company. And I think there's kind of a bit of a correction happening right now where people are more like, okay, that's not just like use AI just for the sake of using AI. That's actually find like the real real use cases. Great. I've heard that one of your most important habits SEO is talking to at least one user or customer every single week. What have you learned from those conversations that you couldn't have learned any other way? You know we started this podcast where you said like, you know, so this is the kind of like a boring company with like boring use cases. And I get that sentiment. But I think the thing that I learned from our customers is just that for people who are using these things, not just for like fun or to create funny bit. video for their friends. It's not just the AI model that matters. There's so much around the AI model that matters. Like, ultimately, you're using a product and small things like hardware share video with a colleague or like, can I save it in a folder? But within a different field, there's a lot of these small things. That's what makes people actually really love and use these technologies. And so I think, as a company, we always try and find the right balance between just AI, AI, AI, and really thinking about how do we put this into a product that's delightful to use for people who are really use cases. And for example, one thing people don't want to do, if your boss has asked you to make five videos this week, you don't want to sit and have to reprompt the video eight times every time to get to a good result. That doesn't matter if you're making a funny meme for Instagram, that's like public creation process. But if you just need to put out a lot of content quickly, you need to work the first time. And so there's a lot of these inside. The nuances, when you speak to people that use these tools, that really inform the product development process. Very interesting. Viktor, thank you so much for being on "Handless Blood Disrupt." I hope our listeners feel this was everything we initially promised it would be. Thank you so much. Thank you for having me. [MUSIC PLAYING] And that's it for now, as with "Handless Blood Disrupt." If you liked it, then subscribe to our podcast and tell us more about it. "Reduction," "Larissa Holtski," "Production," "Migo Fecke." [MUSIC PLAYING]

Podcast Summary

Key Points:

  1. Synthesia nutzt KI, um Videos mit Avataren zu erstellen, die in verschiedene Sprachen übersetzt und lippensynchronisiert werden können.
  2. Das Unternehmen wurde 2017 gegründet und hatte anfangs Schwierigkeiten, Investoren zu finden, bis Mark Cuban einstieg.
  3. Der CEO Viktor Reparbelli glaubt, dass Text langfristig durch Video und Audio ersetzt wird, und sieht darin die Zukunft der Kommunikation.
  4. Synthesia fokussiert sich auf Unternehmensanwendungen wie Schulungen, Marketing und interaktive Avatare für Simulationen.
  5. Die Konkurrenz aus China wird als stark eingeschätzt, besonders bei Open-Source-Modellen, während westliche Firmen einen Vertrauensvorteil haben.
  6. Der CEO betont, dass Unternehmen KI gezielt und mit klaren Geschäftsfällen einsetzen sollten, statt sie übermäßig zu nutzen.

Summary:

In diesem Podcast-Gespräch zwischen Larissa Holzki und Viktor Reparbelli, dem CEO von Synthesia, wird die Vision und Entwicklung des Unternehmens beleuchtet. Synthesia ist auf KI-generierte Avatare spezialisiert, die Videos in verschiedene Sprachen übersetzen und lippensynchron darstellen können – eine Technologie, die ursprünglich für Hollywood-Filme gedacht war, aber heute vor allem in Unternehmen für Schulungen, Marketing und Kundensupport eingesetzt wird. Reparbelli erzählt von seinen Anfängen als Unternehmer, die bereits mit 13 Jahren in World of Warcraft begannen, und von den Herausforderungen, Investoren von der Idee zu überzeugen.

Er prognostiziert, dass Text in Zukunft durch Video und Audio ersetzt wird, da Menschen Informationen lieber visuell und auditiv konsumieren. Das Unternehmen hat kürzlich interaktive Avatare eingeführt, die Echtzeitgespräche ermöglichen und so neue Anwendungen wie Verkaufstrainings schaffen. Reparbelli diskutiert auch die Konkurrenz aus China, die bei Open-Source-Modellen führend ist, und betont den Vertrauensvorteil westlicher Anbieter.

Abschließend rät er Unternehmen, KI gezielt und mit klaren Geschäftsfällen einzusetzen, statt sie unreflektiert zu nutzen, und zeigt sich optimistisch, dass KI langfristig mehr Arbeitsplätze schafft als ersetzt.

FAQs

Synthesia ist ein Unternehmen, das KI-Technologie nutzt, um Videos mit Avataren zu erstellen. Die Software ermöglicht es, Videos in verschiedenen Sprachen zu generieren, wobei die Lippenbewegungen und die Stimme der Person im Video synchronisiert werden.

Man lädt ein vorhandenes Video hoch und wählt die gewünschte Zielsprache aus. Die Technologie erstellt dann eine Version des Videos in der anderen Sprache, wobei die Stimme des Sprechers beibehalten und die Lippen synchronisiert werden, sodass es aussieht, als wäre das Video ursprünglich in dieser Sprache aufgenommen worden.

Avatare sind digitale Klone einer Person, die von Synthesia erstellt wurden. Sie können auf der Plattform erstellt oder aus einer Bibliothek ausgewählt werden, und man kann ihnen einen Text vorgeben, den sie dann als Video sprechen.

Synthesia geht davon aus, dass die Technologie Arbeitsplätze verändern, aber nicht massenhaft vernichten wird. Es werden neue Jobs entstehen, und traditionelle Filme mit echten Schauspielern werden weiterhin existieren, auch wenn sich die Produktionsweise ändern könnte.

Synthesia sieht chinesische Firmen als starke Konkurrenz, insbesondere bei Open-Source-Modellen. Das Unternehmen glaubt jedoch, einen Vertrauensvorteil als westliches Unternehmen zu haben, auch wenn chinesische Modelle oft günstiger sind.

Open-Source-Modelle sind frei verfügbar und können auf eigenen Computern ausgeführt und modifiziert werden, während Closed-Source-Modelle nur über den Anbieter genutzt werden können. Open-Source-Modelle sind oft günstiger, aber es gibt Sicherheitsrisiken, da sie potenziell manipuliert werden könnten.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.