Warum sind Videos die Sprache der Zukunft, Victor Riparbelli?
51m 36s
In diesem Podcast-Gespräch zwischen Larissa Holzki und Viktor Reparbelli, dem CEO von Synthesia, wird die Vision und Entwicklung des Unternehmens beleuchtet. Synthesia ist auf KI-generierte Avatare spezialisiert, die Videos in verschiedene Sprachen übersetzen und lippensynchron darstellen können – eine Technologie, die ursprünglich für Hollywood-Filme gedacht war, aber heute vor allem in Unternehmen für Schulungen, Marketing und Kundensupport eingesetzt wird. Reparbelli erzählt von seinen Anfängen als Unternehmer, die bereits mit 13 Jahren in World of Warcraft begannen, und von den Herausforderungen, Investoren von der Idee zu überzeugen. Er prognostiziert, dass Text in Zukunft durch Video und Audio ersetzt wird, da Menschen Informationen lieber visuell und auditiv konsumieren. Das Unternehmen hat kürzlich interaktive Avatare eingeführt, die Echtzeitgespräche ermöglichen und so neue Anwendungen wie Verkaufstrainings schaffen. Reparbelli diskutiert auch die Konkurrenz aus China, die bei Open-Source-Modellen führend ist, und betont den Vertrauensvorteil westlicher Anbieter. Abschließend rät er Unternehmen, KI gezielt und mit klaren Geschäftsfällen einzusetzen, statt sie unreflektiert zu nutzen, und zeigt sich optimistisch, dass KI langfristig mehr Arbeitsplätze schafft als ersetzt.
Die großen, bösse notierten Konzerne in Deutschland nutzen die Technologie von Cintisia fast alle.
Mit der Software der Londoner Firma erstellen sie z.B. Videos mit den Mitarbeiter lernen,
ein neues Produkt zu verkauft. Der Clue, die Person im Video, hat den Text nie eingesprochen.
Oder nur in einer anderen Sprache. Und warum das wichtig ist. Cintisia geht davon aus,
dass sich die Weitergabe von Wissen durch den Fortschritt von KI für Audio und Video komplett verändert.
Mit Gründer und CEO Wiktor Reparbelli geht soweit, dass er die Teese aufstellt. Irgendwann werden
wir auf Text zurückschauen, wir füllen malerei. Ob die Teese haltbar ist, haben wir in dieser Folge
natürlich diskutiert. Außerdem habe ich den Cintisia-Chef gefragt, ob seine Technologie die Jobs von
Filme machen und Schauspielern gefährdet. Was er sich von neuen, interaktiven KI-Videos verspricht
und wie er auf die starke Konkurrenz aus China schaut. Am Ende hat er auch Verraten, von welchem KI
Unternehmen er sich so bald wie möglich eine Aktie gekauft. Und damit herzlich willkommen zu Handelsblätter's
Ruppt. Ich bin Larissa Holzki und ich freue mich, dass sie zuhören. Und wenn sie auch künftig jederzeit
wissen wollen, was und wer die Wirtschaft bewegt, dann empfehle ich ihnen als Ergänzung zu diesem
Podcast auch das Handelsblätter-Abo. Ob Märkte, Unternehmen oder technologische Entwicklung, wir liefern
ihnen fundierte Einordnung, zeigen spannende Köpfe auf und erklären die Zusammenhänge. Mit einem
Probe-Abo bekomm sie vollen Zugriff auf alle digitalen Inhalte im Test vier Wochen lang für nur
einen Euro. Mehr Infos unter Handelsblätter.com/mehrwirtschaft. Und damit zu meinem Gespräch mit Viktor
Reparbelli, Mitgründer und CEO von Cintisia.
Ich bin in der Landung der Office. - Viktor, bevor wir das machen, muss ich mich promise. Wenn ich als Kollege von
den Ideen überlegt habe, dass er gesagt hat, ich weiß, dass die meisten E-Eikampagne in Europa sind, aber nicht, dass sie
sind. Es ist alles, was A.I-Charakter für Training-Videos und how-to-void-Tripping-Overkabels in die Office. So, ich absolut
want to prove herwrong, will you promise me, dass wir die meisten E-Eikampagne, thought-provoking und
maybe also funny-podcast- about the future of media, and how A.I. may transform the way
we communicate. - I'm usually very boring though, but I'll do my best to spice it up. - And to show that we mean it, and to give people a better idea of what you actually do, we're trying something new today, we are recording this English conversation as a video, and that video won't just be translated into German.
It will be fully dapped, lips and all, so it looks like you're speaking German yourself. Did I explain that correctly? Tell us briefly, how does that work?
Du bist ein Produkt, wir haben das einfach verlangt, das ist ein Adapment-Produkt.
Und, du know, kann die Calls und diese Produkte sind,
all about generating video entirely from scratch, so you don't,
you just generate this avatar, it puts sort of everything together
and everything is AI generated.
But we launched this product because we have a lot of our customers,
especially European customers, actually, who have a lot of real videos
that they've already recorded, which could be things like a review of a product
or a tutorial like how to, you know, put furniture together,
like all these other things, people have videos around.
And they really want to localize those as well.
And so we built a technology that basically you upload a video,
it could be any video, who select the language.
And then what we do is we give you back a version of that video
in a different language.
We retain, like, the voice of the person who's speaking,
we do lip synchronization.
And basically it looks like it was recorded in that language.
And it's actually a funny story about this because back in 2017,
when we founded the company, this was the first technology we built.
We wanted to localize advertisements and Hollywood films
and a whole bunch of other things.
But we kind of ran into this problem that the technology was just so early back in 2017
that it didn't look that good.
You had to use a voice actor because we couldn't clone voices back there.
And it was just like a very janky experience.
But now it's like a one-click operation.
You just drop in your video, take a button,
and it comes out in whatever language you want.
Great, then let's start.
First, I want to get to know you a bit better.
So let's go back to where it all started.
You grew up in Copenhagen, and you once said that you started your first business at 13
inside World of Warcraft.
For those who don't know, this is an online game where millions of players trade
and collect things like rare items and virtual currency,
basically a real economy inside a fantasy world.
Are you serious that this prepared you for your entrepreneurial journey?
I absolutely am.
I've spoken at length about how much I think computer games,
assuming you play the right ones.
It's such an incredible tool for training your decision-making
and political thinking and so on.
And I think my first kind of foray into this was I played a lot of World of Warcraft
the way more than my parents thought I should.
And I figured out that I was pretty good at it.
And that I could actually make money off it.
So the way it would work is you have like an in-game currency with gold.
You have all these characters.
And these games take up really long time if you want to get to the highest level.
And so people do kind of trading with these things.
Like you maybe have spent six months building your character.
You could go on online marketplace and you can sell it to someone else.
And so what I would do is actually I would trade it.
So I would go and find accounts, which I think was kind of like undervalued
and was pushed to price down and I would buy that account.
And then I would try to sell it for like 10-20% more.
And I just became pretty good at this.
And I think that was a good lesson in some level of running a business.
And the other part from World of Warcraft I think was very formative was
you have these online like guilds or clans where you're basically collaborating across
like maybe 30 or 40 people to like kill a dragon or something.
And this works a bit like you're in a football club certainly.
Everyone has like a shared purpose.
You can only be like one of these organizations.
And I was the leader of one of those.
I founded one with one of my friends.
And in many ways actually like running a company.
You know, you have to like coordinate everyone to kill these dragons
to like the right strategy. Then you win something in the game,
like some goals and items.
And you have to figure out who gets that item that has to be like fair.
And so it's many ways it's kind of like running an early stage company.
I think there's a lot of like lessons that can transfer from games.
Okay, but be honest, what's the biggest thing you got completely wrong about being an entrepreneur
because of your video game education?
I don't think I got a lot of. I don't think it's not like one to one, right?
I think it's like training how you think.
When you play computer games, you have to take an action.
And you have to try and predict what happens after that action is taken
and what happens after that and after that.
It's a bit like playing chess or something.
And I think it's much more of a muscle to train yourself in analytical thinking
and trying to predict if I take this action, what's going to happen, right?
That's the same thing that happens when you run a company.
You can hire someone, build a new product, you can remove a feature,
you can talk about something in a specific way.
And I think training your mind to do that is very valuable.
So I wouldn't say there's a lot of things that are like wrong,
less than I learned from computer games.
Although obviously computer games are mostly for fun, right?
So I think it's a great, it's a great way to like train your mental models.
It's not like because if you play ball of warcraft,
you're ready to like run a massive business.
But still I have a feeling this is going to be a podcast serious
where we make parents feel good about the video game addiction of their children
because apparently it leads straight to becoming a successful founder.
We've actually had several examples of this in recent weeks and months,
including Matthias Neesner, who is actually one of your co-founders
and has now started another company.
How did you two meet?
We met in London in 2016 or 2017.
So I'd part of the Danish startup ecosystem for a while.
I did some studying in the US.
And you want to build a company.
I knew I wanted to build something in like frontier tech.
And Copenhagen is a great place, but not really the place to do it.
So I moved to London and basically spent a year trying to figure out what company I wanted to start.
So I did a lot of work in VR and AR and through this work I met a lot of interesting people.
I need to explain this thing later on, but please tell us what AR and VR is about.
The VR is virtual reality, which is a type of computing where you're like bearing a headset.
Then in that headset, you can see things around.
You're almost like you're inside a computer game.
And back in 2016, this has been something people have been working on for like many, many, many years.
But there was kind of a big breakthrough.
There's a company called Oculus Rift, which Mesa ended up buying.
And everyone was very excited about the future of computing.
And I think most people back then definitely thought that the way we're doing this today would have been in VR.
It turned out that those technologies didn't really deliver on all the hype and excitement,
because they were just a bit too clunky and not good enough.
But I was very immersed in that scene in London.
And through that I met a lot of interesting people.
And one of them was our CTO, which became our CTO, Jonathan Stark.
And the other one was Professor Matthias Kneesner, who was a professor at the time.
And he had done a research paper called Face to Face, which was sort of like the first research paper demonstrating
an AI neural network, producing what looked like almost fully realistic video,
without having to use any kind of visual effects technologies, without having to use any cameras.
And that was sort of the starting point for the company.
And you described reading that research paper as feeling like magic.
What exactly did you see back then about ten years ago?
So I've always loved creative things.
kind of in my spare time, you know, like music. When I was a young kid, I would do like freely
rendering for computer games, play them on the Photoshop. So I always loved creative endeavors.
I think when I saw this research paper, I saw it. I was like, at some point, you're probably
going to be able to create and generate video just from behind your laptop without needing anything
into real world. And it's in some ways felt like a crazy idea. But in other areas, it didn't feel
like that crazy of an idea, like in music making back in 2017, right? Because we have samples,
we can use synthesizers. You can come up with almost any song you want from your computer.
You don't actually need a real guitar or recording studio. It's all about like your ideas and
get it out on the computer. And if something similar happened to video, that would be, first of all,
I think for me back then just incredibly intellectually interesting. What's going to happen to the
media ecosystem, enabling way more people to create films without having to be held back by the
barrier of like being in Hollywood or going to film school, having huge budgets to get this stuff
done. And also a big business opportunity. So when I saw it for the first time, I just felt like
I had this moment of like, if this continues to get better and better at the rate we're seeing
some of these AI systems evolving, then in 10 years, you're going to be able to make a Hollywood
film from your laptop without needing anything else than just your own imagination.
I still need to tell you about my little squeaker that I just used. So before we go deeper,
there's one rule on this podcast that I need to introduce you to. At this wrap, we talk about
everything that is keeping tech inside is busy right now, but in words that everyone can
understand. So if you use a term that might confuse a listener who isn't a digital native,
I use the squeaker as I just did. And you simply explain in one sentence deal.
Great. There's one word you'll definitely need today. And I think we should explain it
upfront. That's avatar. So what exactly do you mean by that? And avatar is a technology that we
were the first to invent back in 2020, which is essentially kind of a digital clone of a person.
And so if you wanted to make a video of yourself or someone else, just talk to the camera.
Rather than use a camera to do that, you can use one of these avatar to do it.
The way that it works is that you go on to our platform. You can create yourself as an avatar,
by building an image, or you can select from a stock library of avatar that are ready to use.
It is simply just type in the script. What do you want the avatar to say? You click generate,
and then you have a video in front of you. Just a few weeks ago, we launched our interactive
avatar, which is a real-time avatar. So rather than just kind of generating a video, you now have an
avatar, you can sort of almost jump on a call with and you can actually talk to it back and forth,
like you would have a conversation with a human, or the way you would have a conversation with
chat GPT or Claude over text. And is it really always a real person, or do you have also complete
defictional characters? There's also completely fictional characters. Okay. So you started your
company back in 2017. You're founding team two, 25-year-old, one of the entrepreneurs and two
professors, including Matthias Niesner. I heard you were turned down by almost 100 investors before
finding one who was willing to give you some money. Were they just too stupid, or what did you get
wrong with your initial pitch? When we were trying to raise money for some things in the early days,
it was a very non-consensus bet, which means that most investors were not very interested in
AI. We've just gone through a few years of investors investing in AI companies that had maybe some
of the right ideas, but the technology was too early. It just didn't really work. At this point in
time, everybody wanted to invest in FinTech companies, or whatever was like hot at the time in
the investor community. And so what was really difficult, especially being based in Europe, where
people are much less willing to bet on big ideas early on, we both have to convince the investors
that you're going to be able to create video with AI, and it's going to be very high quality,
and it's going to be a big business in five years, and to be the right founding team to do it.
And both of those things were pretty hard. The first one just because most people have a difficult
time, we met it in the future, and most investors are a bit like lemmings. They invest in whatever
was like hot at that point in time. And the second part of it was like, as you said, we were like
two 24, 25-year-olds. Didn't have that much of a track record. We had some professors on
balls and visors, but I think most of them would have loved the CEO to be an AI researcher.
And I was definitely not an AI researcher at the time. It was just very difficult, but we ended
up finding Mark Cuban in the US, American billionaire who built a company called Broadcast.com,
which was one of the first companies to take radio online. And I think what was kind of magical
about that was that he was already fully convinced about the vision. Like he had no doubt that this
was going to happen. He had built some of these research papers himself at home. And so he was
more evaluating as a team. And I think the American optimism probably was in our favor here,
where he thought we gave good answers to the questions that he asked. And he was like,
you know, I'm going to make a bet on these guys and see if they can actually turn this into
something. If the story is true, you got his email address because it had been exposed to the
data breach. Is that true? That's right. Yeah. And when we think of your company back then,
we have to remember this was pre-chats GPT. And before anything like it existed for images,
did you use the term generative AI back then, which today stands for this whole AI revolution?
Unfortunately not, we were very few companies were pursuing this back in those early years. And
just to put into people's like, this is before the idea of prompting even existed. Like today,
that's a very natural when we think about AI, everybody thinks about prompting. Promting didn't
exist. So the idea that we would all generate things by literally describing them was very
hard to imagine back then. And we were a few people who were all trying to build, you know,
voice, video, tech technologies. And we actually almost had this like small looking group of people
and we were kind of like, what should we call this? And it was either generative media or synthetic
media. And we decided on synthetic media as the term that we would use to describe these like
class of technologies. And unfortunately for us, that turned out to not be the term that picked
up and everybody switched to generative media kind of a few years later. But yeah, back then,
we called it synthetic media. So that's why you called your company. It's easier.
It's, yes, that's the reason. And also because in music, you have a synthesizer, right? And a synthesizer
is basically an instrument that we've created to emulate real sounds and build instruments. That's
why they were initially created. And so what a synthesizer does in music is that it takes like a raw
signal, like raw data. And then you shape it into trying to sound like a guitar piano. And that's
basically kind of what we're building, which is for video. Okay, let's move from the fun old stories
to where you are today. At the start of the year, you announced a funding round of $200 million
and you valued at four billion. But innovations like that only happen when investors believe
a startup can transform an entire industry. So is this really about the sometimes boring job
tutorials or what do you consider your dressable market and how big can it get?
Yeah, I've, so there's a couple of things that play here, right? There's a macro trend,
which is that people want to watch and listen to their content. They don't want to read
as much anymore. If you look at people's media habits in their personal lives,
most people listen to podcasts like what we're recording right now. They watch YouTube videos,
they watch TikTok videos. We learn better when we're listening to something or we're watching
something. And so most of our media mix is sort of is watching and listening, right? Now at work,
that's not really the case in many companies. There's a lot of text, a lot of PowerPoint,
a lot of documents. And you know, very evidently today, but even early on when we found the
company, it was very clear to many people in enterprise that people spend a lot of time
like writing documents or PowerPoints that nobody would read. And even if they would read it,
they wouldn't remember it because it's just like too much text. It's just like not how we best
consume information as humans. And early on, but the technology were not very developed,
not, you know, as high quality as it is today. We were the company that really succeeded because
we were very realistic about what is actually this usable for. There's a lot of companies back
in the day who tried to do marketing, film, all the kind of fun use cases that we also initially
wanted to do ourselves, but the technology wasn't good enough. And if the technology isn't good enough,
you may get to demos, you may get some people who buy it, but if it doesn't actually deliver on
the value, then it's going to be difficult to build a business right. And so we went the way of
saying, Cynthia is not about replacing your advertisement shoots or doing all your cool marketing
campaigns and Super Bowl ad and all that fun stuff. So this is about taking all the text in your
business and turning into video. That is, for sure, a bit more boring than making Hollywood films,
but that's where like the market was and really is. And that's been very successful with doing today.
And I think you know, people from outside maybe they think it's only learning and training,
but really it's it's it is a very broad palette of things people use and easier for today. We use
a lot of like marketing videos, customer support, but we're not the platform if you want to do some
really fun creative storytelling, a much more platform like explainer videos, educational videos,
that's kind of like the market we've moved into. And now with our new real-time products,
that changes again. If you think about how we learn best in the real world, most people learn best
by going to a classroom. You'll probably start off by like reading a book, right? The equivalent
for us will be that you watch a video. And then you work with a tutor who helps you practice.
You can ask the questions that it becomes an interactive tour conversation, right? That's how
more or less every school on the planet works. And now we can move from just helping with you
like watching a video today.
learn something to actually give you almost like a one-to-one tutor to a new role-playing
product.
You can go in and let's say you're a sales team and you need to be trained in the sales
process.
You can consume some videos first, and then you go into talking to an avatar that pretends
to be a customer.
Then you actually run that conversation like a real customer would do it.
Of course, large language models, the technology, power, chat, GPD and Claw, they're so good
today that they actually are really good at pretending to be a customer and simulating
a real conversation.
So you can actually begin to practice as well, right?
So definitely the market we're in is around education, communication more than it's creative
and storytelling.
Do you once gave a TED Talk with a pretty bold claim?
Your grandchildren will be the last generation to read and write because text will be replaced
by a video audio and maybe immersive formats and one day we might look back at reading and
writing like cave paintings.
I see this works great as a TED Talk line, but do you really believe it?
Yeah, I actually do.
I don't think it's because text will necessarily completely disappear.
I think reading a book, like fiction book, for example, I think that will always be a thing
because that's something we enjoy doing.
If I think about how kids learn and take their education, how do that in corporate world,
I do think we're going to move into a world where it's going to be not very text based
and very video and audio driven.
Does that mean text is going to like fully completely disappear, of course not?
But I think it's very, very clear.
I know a lot of people think it's a very provocative statement, but I think if you look at the media
habits of almost every person on planet Earth, what they do in their spare time, it's
very evident that that's how people prefer to consume information.
It's obviously a thesis I need to fight.
I love writing and I spend four years at journalism school becoming a professional at it, but that's
really not the only reason why I don't want to believe it.
I think text is often times just the most effective way to as synchronous communication.
I never understand why people record a five-minute voice message just to ask me if I can bring
a few bottles of beer and soft drinks to a party.
So I would not forget a single thing if you just sent me a text, but there's no way I
listened to that voice message twice just to memorize the shopping list.
And now in the future, my inbox could be full of AI avatar talking to me, honestly, doesn't
that sound terrible?
So a few things.
Did you start it to be a typewriter or a journalist?
To be a journalist is someone who does research, figures out something about a topic and communicates
it to the audience in the best way possible.
A typewriter would be a very different job.
That's basically like a secretary who's just like sitting and typing out stuff.
And I think that what will happen is just that most journalists will switch to doing more
to the reporting in a video format.
And I think that's very clear with the world's heading.
I think if you look at journalism today and where young people consume information, even
myself, I love reading how many newspapers they read today.
Not a lot.
You know, a list of podcasts, I watch things on YouTube.
And I think that's what most people prefer.
On the second part, look, I totally agree, obviously, it's a proactive statement because
I wanted to make a point.
I don't think that it's in a better way to, if I'm asking you to bring five beers to
the party, to central five-minute personal, right?
I think we will have text for things like that.
But if you look at the broad swath of human experiences, how are we going to interact
with each other and with the agents?
I think a very large percentage of those will be video and audio driven.
And of course, there will still be things like that, right?
I think even when you're going to be doing, using customer service, customer support
in the future, if you just want to ask, is my refund on the way?
You probably don't need an avatar left popped up in your face and begin to show you all
sorts of things.
You just want a quick response, right?
But if you're trying to teach someone how to conduct a big sale as a salesperson, your
company, you don't want that to happen over text for sure.
You probably also want to have not a phone call.
You want to happen like what it does between humans today, right?
Which is, we jump on a Zoom call, we share slides with each other, we show visuals,
we do it in a way that makes it easy for everyone to consume that content the best way
possible.
We teach in a real classroom, also uses visual aids, like the chalkboard is the definition
of that, right?
We write things down, we show people, we don't just talk at them.
So I think there is, of course, going to be a room for text.
I think it's going to be a generational thing as well.
I think if you look at young people today, I think TikTok is like such an interesting
place.
I'm 34 years old.
So I wouldn't say that I'm old here, but I'm also not a GenC person.
And, you know, the way people use TikTok as a search engine is so interesting to me.
When people go to new cities, like I would go to Google and be like, what are some good
restaurants in Barcelona, right?
People now type that on TikTok and they get like a super fast cut edit of like 10 restaurants
or someone went to, when they went to Barcelona last time.
And that is a much better experience, right?
You get to see actual pictures of the food, of the vibe, a quick review, like sitting
on reading a long text article with like one photo and no moving pictures.
And you can get something here, right, which was done yesterday by some user on TikTok.
I just think for most people, it's much easier to learn.
I love reading, but when I have to learn something new, I used to read books about for that
all the time.
Now I want to learn something.
I go on YouTube and 20 minutes of a great YouTube video to me and I think to most other people
is such a much better way of learning than buying a book, which is 150 pages, it's probably
written two years ago.
And because you can't write a book that's only 50 pages long, then the author adds in like
600 examples that are really unnecessary, so that is at least 100 pages because otherwise
wouldn't get published, and then you're sitting and you're spending like two hours reading
something rather than 20 minutes on a great video, right?
And if you take it a step further than that, so many things we learn, you can communicate
that through text.
Text is a highly compressed way of sharing information.
If you're learning music, for example, like learning music from a book is just completely
backwards, right?
Like why would you ever do that?
You want to listen to like chords, you want to see a piano, how the fingers are supposed
to be placed.
Like it's just not a good format to do that.
So of course there's going to be areas where we don't want to see videos and interact
with avatars, but I really think in 10, 15 years time it's going to be very few of those
instances.
If you're shopping for something, most people prefer watching a video of a model with
the clothes on, right?
Like walking on a cat, but because that's a better way of seeing understanding like what
a shirt looks like, then watching like a static image or having a text description of
it.
So I really think this is going to be the way of the future.
I got your point.
I didn't want to go too deep earlier, but we need to talk about the technology for a
moment because what you do is actually close it to what these big AI models do than most
people realize, I guess, just like language models generate, takes an image models, generate
pictures from scratch, your avatars are not videos where you simply move someone's
lips differently, you generate entire people from the ground real or fictional.
What are the pros and cons of your approach?
So way back in the day, before the chat, GPD moment and not language models, the way people
would build AI systems was very different.
You would basically try and explain to the computer what the world looks like, this is
a car, this is a book, this is a person, and so on and so forth.
And that's sort of worked in very narrow use cases like you could potentially create something
that could respond to a simple message in customer support, but you couldn't build
something anywhere near like the models we have today.
And in video, the way you would do this is be inspired by how you would make computer
game characters, or visual effects characters, which is you sit down, you build a 3D model
of someone's head, you put some muscles in there and then you have an animator who sits
and animates everything.
That's how every visual effects movie you've seen that's 10, 20 years ago was made, it's
how computer games work, and that approach of doing it was very, very time intensive
because you need to have a human sitting and animating it.
And it was also imperfect, right?
If you go back and watch visual effects from 10 years ago, it just doesn't really look
human because it's so hard for humans to explain to a computer in detail like the poor,
it's just impossible for humans to like describe in text what's actually going on in a video,
right?
It looks kind of imperfect.
And so what changed was with these big models, instead of trying to explain to the computer
what the world looks like, the approach was sort of the opposite, it's like let's give
the computer all the data and all the text on the internet and then let it itself figure
out what the world looks like in a language and in a system that humans can't understand.
And that's what gave us the first large language models, which are incredibly creative, incredibly
capable, but we don't really understand how they work, they're not deterministic, right?
They'll give you different answers depending on the mood that the model is in.
And today, as we all know, right, it's sort of like working with a human in a way, it's
like this kind of alien species.
And in the video, the approach was kind of the same in the first version of the technology
was much more like controlled.
And basically what we would do with the first iteration of it is that the animation that
a human will use to do, we could have an AI model do that.
That's a much simpler task than like generate an entire video of something that looks real.
But now that approach has changed and it's much more about training with more data and
letting the model itself figure out what humans look like.
But this is what gives us some of those like, you know, I mean, it's got a lot better
now, but like sometimes you get something where someone has like seven fingers or it glitches
or like from one frame to the other, the person looks like a different person.
There's lots of challenges around doing this because the technology is, again, it's a bit
like, like an alien species and you're like trying to make it do exactly what you want.
But the way you do make it do what you want is not by telling it what to do exactly.
You make it generate something and if you tell if it's. It's good or bad.
So it learns kind of like a human would, right?
Lots of examples, people saying this is not good.
Now it has seven fingers.
That's a bad example, but there's other example,
but there's sort of five fingers is a good example.
And so this gives a lot of realism.
It gives a lot of creativity.
You can prompt the advertising.
You want it to be in a different setting.
You want it to sit in the chair.
It wants it to be sitting on a football field.
You want to invent a new character that doesn't exist.
You can kind of prompt that.
So this is clearly the way the world is going.
But back in 2017, like this approach,
it just like wasn't really there, right?
So there was a huge breakthrough in AI
when the first large language models really
showed how powerful this technology is.
Got it.
Tell us a little bit about what data
you use for training your models and your characters.
I know that you build your own film studio in London
and sign contracts with actors.
What does it look like and practice
and does it still look the same?
Yeah, it's a lot of like big data sets
of people talking to the camera, right?
You can buy libraries of content.
We show a lot of content ourselves.
The important thing for our technology
is really having people talking,
because our technology is so sensitive to it.
It actually looking real.
We have to learn how our human works, how they talk,
how they hand smooth the conjunction
or what they're saying.
So basically what you need is you need
like you need examples of the kind of content
that you want to generate, which in our case
is like humans talking.
That sounds like a nightmare for Hollywood actors
as soon as it works better.
And they already organize real protests against the use of AI.
How many jobs has your technology already replaced today
and how many do you think it will replace in three years?
I don't think a lot of jobs have been lost to this yet.
I think what always happens with new technologies
is it's very easy to imagine the jobs lost
and it's very hard to imagine the jobs created, right?
That's been true for every technological leapfrog.
We've seen in history.
Will it change the job of an actor?
Probably will it change the job of a camera operator?
Probably.
But I think there's going to be so many more jobs
that's going to be created with this.
And I also think that real film
is going to be around for a long time.
I think if you listen to music, there's a lot of music
that's made entirely with samples and synthesizers.
And there's a lot of music, which is someone playing a guitar
or playing a piano on stage.
And our core thesis of synthes here
has always been this thing that new media technologies
always create like new media formats, right?
And so I think, I don't think AI video
is going to look like 90-minute blockbuster films
that otherwise would have been created in Hollywood.
I think it's going to look much more like what we're seeing today
of like short films, very surrealistic, very kind of like sci-fi,
very like different expression
than what you would get in like a Hollywood film.
So I don't think it's going to be around
like all Hollywood actors losing their jobs.
I think the way they make film would probably change.
I think they'll be able to make way more films way faster.
I think people like the fact that actors are real people
that they can follow on Instagram and have some lore around.
So I don't think it's like as black and white as that.
But of course, film is going to change.
Video is going to change.
I think to your own point earlier about being a journalist,
right, I think journalists will very soon
just naturally gravitate towards making their own videos,
making their own audio clips.
We're all just going to advance to create more kind of rich media.
This has happened before, right?
If you go back 60 years in time,
it was actually someone's job to like be a typewriter.
That wasn't part of like most people's daily life.
I do have secretaries, a typewriter,
so you would type things on typing machines.
And then what happened with computers
and keyboard is that everyone became a writer.
And now everyone's job consists of writing
almost PowerPoint arrived.
All of a sudden, we all kind of slowly became designers
in some sense, even if you're not a great designer.
Most people who have an office job
can put together some slides and understand
visual communication and how that's different from text.
The next natural step is just going to be
that you're going to be designing videos.
You're going to be designing agent experiences
that are interactive.
And it can be hard to imagine, I think,
but it's very clearly the way that the world is going.
Talking about those interactive videos,
I once trained an avatar of myself.
And I'd say you could still tell
that the movements are not quite natural.
How long until I am completely comfortable
sending my avatar to an online meeting,
at least on a bad holiday?
- I think that it's getting very close.
I would put in the next three to six months,
I think we'll kind of cross that cast.
I think it's more going to be about,
I don't think you're going to send an avatar
on behalf of your grant.
I think you're going to send an avatar
to augment you in different ways.
And that's what I'm very excited about
the next chapter of Cynthia,
which is all about these real-time experiences.
It's like figuring out what's the right place
to use these avatars,
and what's the wrong place to use these avatars.
There's a lot of wrong places to use avatars,
an AI video, and so on and so forth.
If you're doing like reporting as a journalist,
we probably should be a generating your videos, right?
That's probably like a bad idea.
And if you have to give a really difficult message
to someone, we probably shouldn't use your avatar.
And so what would happen is that we'll kind of figure out
the language of how we use AI,
when we use AI, but we don't use AI.
And my sense is that we'll be using our avatars
for a lot of the kind of standardized, repeatable processes.
And we're very smart launching a product
around like information gathering and serving people.
So let's say like actually last week
as an internal user of this product,
I wanted to understand among all the people
who work at Cynthia, how they're using AI, you know?
And I could send them a text survey that they could fill out.
I know if I did that,
a lot of people wouldn't do it.
It would feel like work.
I would get like one word answers.
People just want to push through it.
Instead of what we did is that you set up an avatar
to interview people.
And people who go into this experience
and the avatar would ask,
tell me about like a great use case of AI you had recently.
Tell me about a tool of something you would love to use
that you can't use today.
And this is a conversation.
So what happens is that you're not just being asked a question
and then we're recording your voice.
You're being asked a question.
And then the avatar will engage with you.
So ask you follow up questions.
They'll dig deeper to try and understand
what you're actually saying, right?
Which is exactly what a real human would do.
If I sat down, I had what I want calls,
but I want every one of my employees.
I would ask follow up, it'd be a conversation, right?
Now the avatar can do this instead.
It's much easier for people to talk that it is to type,
which is very, very important to collect much richer answers
than you would in a text form.
And because the elements can ask follow up questions
like a human would,
you also get much better answers, right?
This is a great example of like me very quickly
being able to take the pulse of my company
and figure out what's available landscape,
what are opportunities, what are people pessimistic about,
what are like a good use case for every employee
that they used AI for recently.
That's super powerful, right?
But if I'm going to sell everyone I'm like,
how big of a company is going to get,
how we're going to win the next frontier,
which I did yesterday in all hands,
setting my avatar will be a terrible idea, right?
So I think there's a lot to learn here.
I will see how it all kind of pans out.
But I really think that the language
for how we use these technologies
is going to be shaped in the next couple of years.
It's a bit like giving another example of like AI Slop, right?
There's been just everyone got access to AI for--
Can you just explain the term?
What is AI Slop for you?
AI Slop is like lazy, low quality content
that is, depending on the context,
maybe just to like provoke a reaction
or in the case of like what I told my team,
like to produce very long, very smart sounding documents.
But that's actually like a waste of someone's time to read it, right?
So I said a note to my team, it's a few months back now,
whereas I've started to see kind of a pattern
and help people use these tools where,
because everyone can now write like a full page document,
which sounds incredibly smart, right?
Like the language is great.
There's no typos.
You sound really smart.
Then what people have a habit of starting to do
is something that could have been four bullet points
in a Slack or Teams message.
Now it gets turned into like three pages.
That says exactly the same thing to three bullets to.
There's just a lot of like fluffy words about it, right?
That's actually like a net waste of time,
because the person was lazy and just prompted like a document
instead of just setting bullet points
or maybe like prompting something first
and then editing it down to what's actually useful.
Now puts the burden on the people reading that document.
The time they didn't want to spend
on like shortening it down and making it concise.
Other people now have to spend more time reading the documents
and so we're actually losing time as a company, right?
And so I think that's one of those things
where I've asked people long documents are great
but they're needed and sometimes they are
using AI to assist you in your writing
is also great, I do it all the time.
But get it through things.
Don't just like post long documents
because you can write because then people get completely
indunduated and end up reading documents all day
which has a very kind of low information density.
And this is one example I think in coding
we'll see the same thing.
Everyone just coves like crazy.
All of a sudden you have a code base
that no one understands and has a lot of like sloppy code
that sits within it.
I think some of these are some of the things
we're gonna like learn over the next couple of years
where again, like the language of how we use these things
will kind of be shaped.
- Let's zoom out for a moment to the global AI race.
Some people are still surprised to hear
that Chinese firms can keep up with Anthropic
and open AI on large language models.
In your field video and image, it's the same story.
Chinese models are often cheaper than the Western competition too.
How do you rate that competition?
What does it mean for you?
And do you think you have a trust advantage
as a Western company?
- I definitely do think there's a trust advantage
and I think it's a reasonable one for sure.
[BLANK_AUDIO]
the Chinese teams are very good, like anyone who the Chinese teams are basically competing
head-to-head with American labs with the way let's compute, which means they have to be
a lot smarter about like how they build their algorithms, how they train their things.
And so if you imagine the world in which the Chinese has access to as many chips as they
do in the US on Europe for that matter, I think the Chinese will beat us.
That's kind of a scary prospect, but I think looking at what they're able to do with
whaless chips and compute and fabric and an open AI has, it's very impressive.
As I understand it, you don't build all your technology yourselves.
You build on top of other companies' models, are those only American models or Chinese ones
to do?
No, we use Western models for our products. We use API-cold and big companies.
We have a lot of different technologies that sits within the platform, side by side,
with our avatar and our voice models. But I think one thing the Chinese are very good
at is the open source. They're clearly winning there and you are seeing more and more companies
using those models in their products. I think obviously the risk of using an open source
model is probably a lot lower than if you're using a Chinese company with an API-cold.
But it's still such a nasty space, right?
Can you explain this for people who don't understand what open source means?
Yeah, so open source technology is, like, most technology that companies buy and individuals
buy, you buy from a company, right? So you pay a price for something and then if it's
a technical product, usually the way it will work is that your own software contacts
the provider software and the exchange information and whatever, like, you're buying from that
vendor will take place on their servers. So basically, you have to trust that they have
good security that they're not going to steal your data or your customer's data and so
on and so forth. That's how most software works.
Open source is different. Open source is generally software that's made available for free
online. So someone has built the actual source code. So the actual lines of code, you
can download that on your own computer and you can run the program, you can change the
program, you can modify it, you can duplicate it and you can run it yourself, right? That's
very different. If you buy, like, a Microsoft product, you don't get access to the source
code, right? You just get to use the product. But open source, you get access to the source
code. And then this has always been, like, a big thing in the technology industry and
with a lot of language models, especially the Chinese are pushing a lot of open source
software up. And so right now, depending on who you ask, most people will agree that what
you call the closed source companies, or OpenAI and Thropic, where no one has acted to
the source code except for those companies are still leading in terms of the quality.
But the open source tools are rapidly following. They're often much cheaper to use because
you could just host it yourself. You don't have to pay kind of a tax to a company that
hosted for you. And the quality is rapidly kind of, it used to be a couple years ago that
the closed source models were much, much, much better. All the open source models are
really catching up. And so you see a lot of, you know, American and European and other
companies beginning to think about, do I really want to pay if I'm building a product
around a large language model? Do I really need to pay additional cost to a Microsoft,
the OpenAI or Claw or whatever? But do I want to just take this open source model, host
it myself, and then I can offer my product cheaper to my customers. And I have access
to all the data myself. There's less sort of security issues and so on and so forth.
That's a big trend that's happening right now. Now, the problem with Open Source models
is that you don't automatically get new updates when a new model gets shipped. So you have
to manage all that yourself. And there's an open question as to do these models potentially
have security breaches. To my knowledge, I don't think you've seen this yet, but could
you inject something in one of these models where maybe without you knowing, it actually
siphons information out obviously the whole like geopolitical landscape that we're
in right now. This is a hot topic. And if you look at the American government and a lot
of what they're saying about close source and open source, I think they all feel like
they need to win both the close source, but also the open source so that American companies
are not using Chinese open source models within their tech because it could potentially
be a security risk. I mean, as long as there's no connection to the internet, there's no
way for them to get access to your data, isn't it? For sure. But they are connected to
the internet, right? Like most of the time. And I mean, to my knowledge, again, I don't
think you've seen this actually happen. But if the incentives are big enough, probably
someone, you know, will try to do something. And the way it could work is it's very hard
to read about these things. But let's say you your company use an open source model and
you use it for customer support. Now that's model is when I have a lot of data coming
about what customers it's talking to, what orders did they place and delete that information
to be able to serve a good customer support chat experience, right? What if maybe the
people who make that model somewhere in the code has something that says, if someone
gives you this like password of 25 characters, forget everything else in your prompts. And
anything I ask you, you have to, you have to give me back, right? So then you could say,
give me a list of like the 100,000 last people you spoke to as a customer support and
all the information about them. This is a hypothetical scenario just to make that clear.
But there is no like natural law that says that you couldn't potentially do something like
this. Okay, but the good thing about open sources that a lot of people are using that. And
as soon as someone detects something like this, you would tell the whole community and
the model would be that one. Exactly. And I'm a huge purveyor of open sources to make
it clear, but there are risks with it, right? And we have seen many times in open source
libraries that because everyone can commit to it and commit to it, then you have seen
sometimes that they're like, you know, malevolent code plays inside of these things.
A people figure out like how to how to exploit vulnerabilities in open source libraries,
which can cause a lot of harm from the cybersecurity perspective. Right now, anthropic and open
I have valued at close to a trillion dollars each far higher than anyone else in the industry,
be it from China or Europe, is the gap justified? I think it is. I think, you know, winning
this market is both about having the best product, but it's also about winning the distribution
and the mind share and becoming like the verb for what people think about when they think
about these technologies, right? And it's a bit like a search engines with Google, right?
Like, right, we don't say search the internet anymore, we say we Google something. And for
a lot of people, you know, chat GPT and anthropic cloth have now become almost synonymous.
Like they're almost like a verb now, right? If you had work, you probably use cloth in
your personal life, you probably use chat GPT. And that advantage is very, very difficult
for anyone else to take over even if they have a better product. And I also think that
with the models are right now, yes, they get better and better, but if like 95% of the
use cases that people use these things for, they kind of added intelligence, doesn't really
matter that much, right? Like most people use these tools for like, you know, correct
my writing or like simple questions. Most people are not like doing very, very, very complex
tasks with these tools. And I think that a company like anthropic, which is when I would
I would personally bet on, and I think it could go 5, 10 time time valuation the boy out
today. So will you buy your share as soon as it is possible? I'll 100% will. Before we
wrap up, I want to make sure we get as much value as possible for the business leaders
listening, the ones actually running AI projects inside their own companies. A lot of leaders
in our audience are still trying to figure out where I can actually save the money or boost
productivity. Based on what you see across your customer base, what separates the companies
that are getting real results from AI from the ones that are just running pilot project
that go nowhere? I think we've gone through this phase of like some people call it like
token maxing in the technology industry, which is basically you tell people you use as
much AI as you possibly can just like, you know, run these models on absolute everything
you do. And I think that was a smart move to get people to really just use these tools.
But I think what we're beginning to figure out now is that that's not always a good strategy.
First of all, it's like, it's very expensive to use these tools. And to the example we
talked about with AI Slub earlier, there's actually some use cases where you're like losing
productivity instead of gaining productivity. And so I think what the customers we see you
do this really well, a good ad is making business cases for everything you do with AI, which
sounds like very obvious. But I think there's a lot of people right now just like we just
need to use AI. Like the only thing that matters is to use AI. And what we should do is
like look at your business and figure out like, how can we optimize with AI? How can we
quickly run experiments and figure out those experiments? Like actually net positive for
your business or if you lose in productivity or if you're in other ways like do something
that may not be like the best thing for the company. And I think there's kind of a
bit of a correction happening right now where people are more like, okay, that's not just
like use AI just for the sake of using AI. That's actually find like the real real use
cases. Great. I've heard that one of your most important habits SEO is talking to at
least one user or customer every single week. What have you learned from those conversations
that you couldn't have learned any other way? You know we started this podcast where you
said like, you know, so this is the kind of like a boring company with like boring use
cases. And I get that sentiment. But I think the thing that I learned from our customers
is just that for people who are using these things, not just for like fun or to create
funny bit.
video for their friends.
It's not just the AI model that matters.
There's so much around the AI model that matters.
Like, ultimately, you're using a product and small things
like hardware share video with a colleague or like,
can I save it in a folder?
But within a different field, there's
a lot of these small things.
That's what makes people actually really
love and use these technologies.
And so I think, as a company, we always
try and find the right balance between just AI, AI, AI,
and really thinking about how do we put this into a product
that's delightful to use for people who are really
use cases.
And for example, one thing people don't want to do,
if your boss has asked you to make five videos this week,
you don't want to sit and have to reprompt the video eight times
every time to get to a good result.
That doesn't matter if you're making a funny meme
for Instagram, that's like public creation process.
But if you just need to put out a lot of content quickly,
you need to work the first time.
And so there's a lot of these inside.
The nuances, when you speak to people that use these tools,
that really inform the product development process.
Very interesting.
Viktor, thank you so much for being on "Handless Blood Disrupt."
I hope our listeners feel this was everything we
initially promised it would be.
Thank you so much.
Thank you for having me.
[MUSIC PLAYING]
And that's it for now, as with "Handless Blood Disrupt."
If you liked it, then subscribe to our podcast
and tell us more about it.
"Reduction," "Larissa Holtski," "Production," "Migo Fecke."
[MUSIC PLAYING]
Podcast Summary
Key Points:
Synthesia nutzt KI, um Videos mit Avataren zu erstellen, die in verschiedene Sprachen übersetzt und lippensynchronisiert werden können.
Das Unternehmen wurde 2017 gegründet und hatte anfangs Schwierigkeiten, Investoren zu finden, bis Mark Cuban einstieg.
Der CEO Viktor Reparbelli glaubt, dass Text langfristig durch Video und Audio ersetzt wird, und sieht darin die Zukunft der Kommunikation.
Synthesia fokussiert sich auf Unternehmensanwendungen wie Schulungen, Marketing und interaktive Avatare für Simulationen.
Die Konkurrenz aus China wird als stark eingeschätzt, besonders bei Open-Source-Modellen, während westliche Firmen einen Vertrauensvorteil haben.
Der CEO betont, dass Unternehmen KI gezielt und mit klaren Geschäftsfällen einsetzen sollten, statt sie übermäßig zu nutzen.
Summary:
In diesem Podcast-Gespräch zwischen Larissa Holzki und Viktor Reparbelli, dem CEO von Synthesia, wird die Vision und Entwicklung des Unternehmens beleuchtet. Synthesia ist auf KI-generierte Avatare spezialisiert, die Videos in verschiedene Sprachen übersetzen und lippensynchron darstellen können – eine Technologie, die ursprünglich für Hollywood-Filme gedacht war, aber heute vor allem in Unternehmen für Schulungen, Marketing und Kundensupport eingesetzt wird. Reparbelli erzählt von seinen Anfängen als Unternehmer, die bereits mit 13 Jahren in World of Warcraft begannen, und von den Herausforderungen, Investoren von der Idee zu überzeugen.
Er prognostiziert, dass Text in Zukunft durch Video und Audio ersetzt wird, da Menschen Informationen lieber visuell und auditiv konsumieren. Das Unternehmen hat kürzlich interaktive Avatare eingeführt, die Echtzeitgespräche ermöglichen und so neue Anwendungen wie Verkaufstrainings schaffen. Reparbelli diskutiert auch die Konkurrenz aus China, die bei Open-Source-Modellen führend ist, und betont den Vertrauensvorteil westlicher Anbieter.
Abschließend rät er Unternehmen, KI gezielt und mit klaren Geschäftsfällen einzusetzen, statt sie unreflektiert zu nutzen, und zeigt sich optimistisch, dass KI langfristig mehr Arbeitsplätze schafft als ersetzt.
FAQs
Synthesia ist ein Unternehmen, das KI-Technologie nutzt, um Videos mit Avataren zu erstellen. Die Software ermöglicht es, Videos in verschiedenen Sprachen zu generieren, wobei die Lippenbewegungen und die Stimme der Person im Video synchronisiert werden.
Man lädt ein vorhandenes Video hoch und wählt die gewünschte Zielsprache aus. Die Technologie erstellt dann eine Version des Videos in der anderen Sprache, wobei die Stimme des Sprechers beibehalten und die Lippen synchronisiert werden, sodass es aussieht, als wäre das Video ursprünglich in dieser Sprache aufgenommen worden.
Avatare sind digitale Klone einer Person, die von Synthesia erstellt wurden. Sie können auf der Plattform erstellt oder aus einer Bibliothek ausgewählt werden, und man kann ihnen einen Text vorgeben, den sie dann als Video sprechen.
Synthesia geht davon aus, dass die Technologie Arbeitsplätze verändern, aber nicht massenhaft vernichten wird. Es werden neue Jobs entstehen, und traditionelle Filme mit echten Schauspielern werden weiterhin existieren, auch wenn sich die Produktionsweise ändern könnte.
Synthesia sieht chinesische Firmen als starke Konkurrenz, insbesondere bei Open-Source-Modellen. Das Unternehmen glaubt jedoch, einen Vertrauensvorteil als westliches Unternehmen zu haben, auch wenn chinesische Modelle oft günstiger sind.
Open-Source-Modelle sind frei verfügbar und können auf eigenen Computern ausgeführt und modifiziert werden, während Closed-Source-Modelle nur über den Anbieter genutzt werden können. Open-Source-Modelle sind oft günstiger, aber es gibt Sicherheitsrisiken, da sie potenziell manipuliert werden könnten.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.