Go back

The Rise of Generative Media: fal's Bet on Video, Infrastructure, and Speed

62m 18s

The Rise of Generative Media: fal's Bet on Video, Infrastructure, and Speed

During a discussion with the team from Fall, comparisons were drawn between the resistance to AI and the initial opposition to computer-driven animation. Fall serves as a developer platform for generative video models, offering access to a wide range of models. The team highlights the unique challenges faced in optimizing video models compared to text models. They stress the significance of focus and dedication to achieve top performance in this field. Running video models demands significantly more computational resources when compared to image or text models, with the team noting the exponential increase in computational intensity for video processing.

Transcription

10750 Words, 59775 Characters

We recently had our first-generation media conference and Jeffrey Katzenberg former CEO of DreamWorks was there and he met a comparison. He said, "This is exactly playing out how animation." When it first came out, people revolted against it. It was all hand-drawn before that and computer graphics. It was new and there was a lot of rebellion against computer-driven animation. And something very similar is happening with AI right now, but there's no way of stopping technology. It's just going to happen. You're either going to be part of it or not. In this episode, we sit down with a team from Fall. The developer platform and infrastructure powering generative video at scale. Follows a place that developers can go to access more than 600 generative media models simultaneously from OpenAI, Sora, and Google Vio to open-weight models like Kling. We'll discuss why video models present fundamentally different optimization challenges than LLMs, why the open source ecosystem for video has a thriving long tail in ways that text models never did, and why the top video models have a half-life of just 30 days. The team also shares insights from the demand side of the video model equation. We discuss what's happening in the app layer from AI native studios to personalized education, what's happening in Hollywood, and more. Enjoy the show. [MUSIC PLAYING] Borkai, Gorkom, Bhattuan. Thank you so much for joining us today. I want to start with the problem space that you decided to tackle. So Fall is a developer API and platform for generative video and image models. Video is massive, obviously. It is more than 80% of the internet spend with. And it follows that generative video is going to be similarly massive. But there's not that many companies that are focused on this problem. Why do you think that is? Yeah, in a way, a generative image and that video was an overlooked market in this current phase of AI. In my opinion, for two reasons. Number one, there wasn't a very clear industry use case that people were going after. There wasn't vibe coding that automates software engineering, or there wasn't search, which seems LLM market is going after, or customer support, anything like that. Also, number two, the investment on the research side wasn't as big three years ago. And then that ramped up a little bit slower than LLM's, but still considerably since then. And now the models are much more capable, much more useful, and real industry use cases compared to what it was three years ago. It felt like a toy use case. This was just going to be for fun on the side. And it's going to be a small market at the end. And now we can see that it's going to be a massive market with very unique use cases and customers compared to the LLM market. Like, if you actually go back to, like, as we were experiencing it, I think that was an interesting time. We were working on some Python compute infrastructure, and then these models, like Delhi, too, had just come out. And then soon after that, Chatchapiti had come out, and then Lama had come out. And we were just like, initially, we didn't know that image and video market was going to get that big. We were actually just curious about running image models much faster, that was our initial entry point. And then we saw the initial growth. We had a few customers, and they were growing really fast. We were like, what the heck is going on? And then a few customers later, we actually thought, hey, we should double down here. And around that time, also, the other thing that was happening was people were over-indexed on language models. This story of AGI was being told, and that attracted all the dollars, that attracted all the talent. So everyone was just like working on that, where we thought we had something niche growing fast. Don't tell anyone. And then we just started focusing on that. And soon after, as we got more familiar with the models, I remember, I think we changed our website copy to say, Generative Media, Generative Media Platform. And then it was only two or three months after that, Sora was announced. So we were definitely ahead, but we really saw the whole future coming with better image models, video models, et cetera. So yeah, we made this early bit. You guys have a front row seat to the sorts of new experiences people are building. I think the market's only going to expand from the media market that we know today. Yeah, absolutely. I think Al-Quaara Carpati Tweet, no good podcast without it. He did want to recently where he was talking about why he's excited about the media models. And one of the things he said was that-- he also mentioned that people are visual. And we have so much more video than text, wall of text. And he was making a point around education. And a lot of the content you consume just to learn things. I think right now, the model quality is just relatively-- it's so much worse than what it can be, where you could actually have-- I do a lot of learning on chat GPD, but it's through text. But if it actually rendered a video where it could compress a concept instead of 10,000 characters, if it could do it in 15 seconds, it would be so much better. I think there's the quality bar where it's going to go up. And once we have that, we're going to have even more penetration. So it's really a function of the quality right now, and we're just like in the very early beginnings. Totally. Education markets almost untouched right now with video generation. And there's so much potential there. And it's just waiting the quality, the predictability to get there. And I think it's going to have a lot of potential. Totally. I mean, you guys sent me that generative video Bible. I think it's a much better way to learn some of the lessons from the Bible, if you're capturing consumer's attention right where they are. I agree with you, we're just at the beginning. So Falls and Infrastructure Company. And so we're going to structure today's interview. I love infrastructure companies in terms of the technical layer cake. So we're going to start from the core inference engine, compilers, kernels that you built. We're going to go up to the model layer, and then the workflows, and then end with some observations on the markets and what people are building. Sounds good. Sounds exciting. OK, let's do it. The inference engine. Botswain, did it? How old are you? 22. You're 22 years old. OK, say you're on your background. I think it's super badass. And makes complete sense why this company is so hardcore. I started working on compilers when I was 14. So in a way, I have a lot of experience on that route. It's not just that. But I started working on open source projects. So my first contributions were on tooling around the Python language. And then I started to slowly contribute back to the Python language, core compiler, core parser, and the core interpreter itself. And became one of the core maintainers of it. I think at the time I was the youngest core maintainer of the language. And this kind of gave me a unique appreciation of compilers and how flexible they are. So when we first started working on serving these image models at fault, the main idea was, OK, there is three different image models, three different architectures. But this is surely going to explode. There's upscalers. There's going to be video models. We were predicting that. And we didn't want to go optimize a single model, put our eggs into a single basket, and then go became invalidated when the next model comes. So we started building this inference engine, which is a tracing compiler that traces the execution. And essentially, it tries to find common patterns that are fitting within the templated kernels that we do. So our our bread and butter is like spending-- we have a 10% performance team-- that's spending all their efforts into writing kernels that are like 95% there, but like generalized with templates. So we traced the execution of a model and find common patterns that could replace these templated semi-generic kernels to like specialized kernels at runtime and optimize the performance of these models. And we found this technique to yield pretty much spear results from anything that's out there in the market. And this led us to claim like number one spot on performance on all the benchmarks. And another big thing about this is we specialize in doing like this sort of kernel level mathematically correct sound abstractions that led us like, you know, maintain the same quality of these models, which is a very high bar when you're in this media industry and when you really care about the output that you're getting. What's different between optimizing a diffusion model versus an other aggressive element? In autoregressive elements, like your bottleneck is half as you can move all those like giant weights from memory to a stream because you have like a 6 billion or a parameter model. And you're trying to predict the next token. You're doing the attention for like all the tokens, like a couple tokens before that. In diffusion models, you're trying to denoise like thousands, tens of thousands of tokens for a video at the same time doing attention of it. So you're essentially saturating all the compute bandwidth of these GPUs. You're not necessarily bound on memory bandwidth. But like the computational operations that you do are like fully saturated. So you're trying to find better ways to execute around the GPU. This could be like writing more efficient kernels or this could be overlapping of softmax with gams that you do. Like it's essentially like you're trying to use all of the power of the GPU, leverage it in a way that gets you all the capabilities. So it's a different binding constraint. It's on the compute versus the memory. And what's the intuition for why LLMs are relatively memory constrained and why video models are by comparison relatively compute constrained, but not as large in terms of just sheer number of parameters? I think it's scaling issue right? Like in terms of like if you scale these video models of 600 billion parameters with the same dense architecture, you're going to have to do attention with all those full 100-- like let's say a single video is 100,000 tokens. And you do this attention step or like you do this like denoising step 50 times. And every 50 times you do like attention over all these like 100,000 tokens. It's insanely, insanely expensive. So I think the constraint there is like just how fast you can do the inference. And the same applies to LLMs at like larger batch sizes. But like at like the traffic patterns that people do, like the batch sizes are not that much. And you're mainly constrained by memory bandwidth. So people do optimizations like speculative decoding. And other other factors to like reduce that a whole lot. Yeah. What exactly goes into being at the top of the leader board in terms of performance? Because I would imagine there's other teams. They're also very smart people. And you know, this is my Olympics. And so what exactly goes into-- I imagine people have like very similar ideas on the techniques and different optimizations they can do. I don't think anyone cares about it as much as us. We are literally obsessed with journey to media. We are literally obsessed with these models. We have a team that's like just focusing on this. Like so far, like it seems like-- And from any media to other inference players, everyone is like super obsessed with language models. Everyone is trying to get like one more tokens per second on like deep seek benchmarks, whatever. And like we are like on a different lane. We have like competitors, but like no one close to us because I think the SM is one of the best teams. We found out like the best way to optimize these general models. And we just focus on this. This is like a purely focused thing, right? Like at the end of the day, you're constrained by the hardware. There's nothing unique about it. But like we're just like three months ahead, six months ahead. Like when we benchmark torch, like the latest version of torch, again, it's like, you know, our inference engine from a year ago, we are clearly underperforming because like torch caught up. The same thing is going to happen with other players. You're going to always like the lead that you can maintain is three months, six months ahead at most. The thing that matters is just focus. If you focus on it, if you purely put all your energy into it, I think that's like, there's it's very hard to get out competitive by others because models are slightly changing each month, each release. So it's still the same general architecture, but there are slight differences where we can go in and optimize where that's different. And no one else is paying that much attention to it. Also hardware is changing as well. So we were able to adapt to be 200s earlier than anyone else. And we were able to run video models much faster, basically throughout the year, because of that obsession with running video models later, it's hardware. - Yeah, got it. What are the hardest technical problems that you think you're solving? - So one thing people don't appreciate it as much is we are running 600 different models at the same time. We have to be running them. We have to be so good at running them that we should be running a single one of them better than as if someone else is running a single model. Because when a foundational lab is running models, maybe they have a single version of the model, maybe they have like a couple other different versions, and that's what all they care about. We have to be better than them at running those models. And we have to be doing all 600 at the same time. So on top of the inference optimizations that happens in the GPU, a lot of optimizations on the infrastructure level needs to happen. We need to manage the GPU cluster in a way that's efficient to load and load these models at the right times. We need to route traffic to the right GPUs who have the warm cache of these models. We need to be smart about choosing the right kind of machines, which kind of chips are running what kind of models. And the customer traffic is changing all the time, and we need to adapt towards that. So on top of the inference engine, the overall infrastructure is also really, really hard beast to manage, and so far we've done an incredible job at that. Would you add anything to the-- I think that's a pretty fair explanation of what we do. I call this distributed super computing. I don't know why people don't like that name, but I like it. But the idea is we are at 28, this was a month ago. Now we are probably at 35 different data centers. And you have these heterogeneous groups of compute that split across with their own different specs, different networking, whatever. And you're trying to schedule workloads as if it's a homogeneous cluster that you got from a hyper-skitter. It doesn't work like that. So if we built, we spent the last three years building that tractions over it, from our own orchestrator to building our own CDN. We go back to the fundamentals of web, and we built our own CDN service. Deploying racks to call us, just like routing traffic. So if we build all these technologies to essentially make sure that we can tap into capacity wherever it is, and schedule our workloads, which is very different than a traditional enterprise LM usage pattern, like the use case that we have are so much more spread across, so much more, more consumer-facing. And when you consider that, there is so much investment going into making sure we can tap into the scarce capacity of GPUs. Yeah. You mentioned hyperscalers. And I hear distributed computers, and I hear managing giant clusters. And I naturally think that somewhere where hyperscalers should have the incumbent advantage. Why do you think that you've been able to so far execute them on the core engine? There's two things about the core engine, right? There's the inference part where none of hyperscalers have any expertise. This is a net new field. I think this has been only happening for the past three years, inference optimization. So it's like a brand new lane that we have been outcompeting anyone in our field. I think that's like a pretty much answer of its own. And the second one is infrastructure. I think right now, hyperscalers are very busy with their traditional pattern of, oh, we have this data center capacity. We'll just deploy GPUs, and we don't care about the rest. This has been changing recently, you know, even like Microsoft is going buying from new clots. This is like, there's like an interesting pattern happening because the GPUs and the demand and the growth of GPUs doesn't fit to the patterns of these hyperscalers, the growth patterns they expect. So I think at this age, not even hyperscalers have that big of an advantage of scale because like they're going buying GPUs from new clots. Like the tables have turned a bit. Yeah, it almost helps to also be like slightly earlier, like in the company journey, right? If you're a public company, you also have to abide by what the market's expecting of you. So like the other thing is that there's a huge price discrepancy with hyperscalers and new clouts, right? So like it's maybe sometimes 2x, 3x more expensive to use things through a hyperscalers. What's driving that? Well, I think one is like market pressure, right? And also like there's added kind of operational expenses that hyperscalers have for like having, you know, better, they just have a better service, right? Better uptime and better SLAs and all of these things add up. And then on top of that, there's kind of an established like cloud margin, right? And you know, the market expects the cloud margin to be a certain level, whereas like if you have a three-year-old neo-cloud, you know, you're a private company, maybe you don't have as much pressure. And there's like assuming infinite demand and limited capacity, you can actually, you know, hyperscalers can keep their prices high and you know, they will fill out the capacity and also like get slightly better economics, whereas like neo-clouds compete over the whole infinite demand and that pushes the prices down. Perfect price competition. What does it take to run image versus video models? Well, like you guys started the company around, let's say a little, let's say a little fusion moment, it was, the field was mostly image at the time. How does running video models compare to image? - Let's actually do text image video. - Okay. - Even let's compare all three of them. And so for, for let's say a SOTA LLM, like we don't, let's say deep seek or something like that, where we know the numbers, running a single prompt, like 200 tokens, let's say it takes one X of terroflops, I think it's tens of terroflops, but let's call that unit one. One image is around 100 X of that and if you are doing a five second video, 24 FPS, that is around 120 frames, so 100 X from one image. So you are already 100 X of, over 100 X, so you are at a thousand, 10,000 X for a standard definition video. And if you wanna do 4K, that's another 10 X. So 10,000 X compared to a single 200 token LLM input. So it is a lot more compute intensive in terms of amount of flops you are doing. - Yeah, in general, like when we started with image, the infrastructure was relatively easier to do because it takes three seconds, or it took 15 seconds back in the days, it takes 15 seconds to generate the image, you don't need to necessarily shave like that, 50 MS, 100 MS you have, overall on the system, and then when we went to video, it's even easier because it takes like 20 seconds, 30 seconds to generate the video. The way that has been happening in the past couple of months is real-time media, where you need to stream 24 FPS videos over the overall like a network link from these GPUs. That's where we actually spend like some of our time. We started this progress with like speech-to-speech models a year ago, we started optimizing them, where we were able to like reduce the latency of our system with like globally distributed GPU fleet. When you send a request, we route to the closest GPU, minimize our own overhead, and then do stuff like, you know, pick the best runner, whatever stuff like that. So we are now applying those same optimizations with it to real-time video, and we actually see like really good interesting demander, where people want to experience these stuff like as they type, as they prompt, and that's where like some of the infrastructure technical challenges differ from traditionally running image and video models, because image and video are similar-ish, you know, like just more compute expensive, but like you actually need to care about infrastructure stuff, maybe you go from like less than a second generation time for some of these models. - Yeah, another interesting thing is like image models, especially you were able to run them on a single GPU. Like the parameter counts is actually much smaller. So that actually makes it like a little bit easier for us, as opposed to LAMs. And then with video, parameter count is going up. Right now, I think we're around like for the open source ones, I don't know, 30 billion parameter. Whereas, you know, we hear rumors about, you know, GPT-4 being like in the trillions, GPT-5, maybe more. So that's another, that's like, you know, on the flip side, it's a little bit easier, but it doesn't mean video models are not gonna grow, right? There's, you know, rumors around numbers for Vion, numbers around Sora, so like, there's also an increase in parameter count. So that you're gonna have to kind of, you know, use more distributed computing. But if you're just, you know, eight nodes, that's a, or one node or eight nodes, you kind of have a slight advantage. - Yeah, totally. Okay, let's pop one layer of the stack to the models. - Let's do it. So one thing I think people don't fully appreciate at the media space, and you mentioned this, you alluded to this before, is that there are, there's a very, very long tail of models that are actually used in practice. And so it's helping you give people a sense of on your platform, how many models are people actively using, how's it distributed? And like, why do you think there's such a long tail of models being used compared to the alarm space? - This is actually one of the things I would say, three years ago, people got it wrong. I mean, jury's still out, but people, right after the chat GPT, people start talking about omnimodels, there's gonna be these giant models that, they're gonna be able to generate video, audio, image, and code, text, every type of token. This, this might still happen, I think, but it's more clear that you are better off, if you optimize for a certain type of output, even this is true for code generation, definitely true for image or video output. So that's one thing, when you were pitching three years ago, everyone, that's one feedback we got, there's gonna be omnimodels and there's gonna be a single way of running these. It's gonna be hard to create an edge on the modality, but turns out it's not true, and it actually makes sense to have a technical edge on the modality. And this is one of the reasons why there's also a variety of models, because still the best upscaling model is just doing upscaling and the best image editing model, even the best text to image model is different from the image editing model. So all these special tasks require their own model. It might be the all, the similar model family or similar architecture, but at the end of the day, it has its own weights that needs to be deployed independently, and that creates the variety in the ecosystem. I think there's also the, this also applies to language models, where even this in the same modality, there is different families of models with different taste, different characteristics, there's different personas, and this happens with language models too, the code that Cloud writes is very different than the code GPT-5 does, right? And we see this happening, but the good thing about here is there's these three four different personas on top of different categories upscaling everything, video, text to video, whatever stuff like that. So it gets you like close to 50 models that are active at any point in time, and then you have a very long tail of models that people still choose, because they might like the person of that better. - Yeah, totally. Speaking of model personalities, what are some of the most popular models on your platform? What do you think are the personalities of them? - So one thing that's been true since the beginning, the popular models change all the time. So there is always new releases from different labs that take over the other, and it's always a moving target. But that being said, there's two types of models usually preferred by our customers. It's usually there's a one big expensive model that has the best quality. On video generation, this could be VO, this could be cling, this could be Sora, and then there's usually a workhorse model which is cheaper, smaller, but good enough, and people usually use that at higher volumes. I would say this has been true for the past almost two years that there's an expensive high quality model that keeps changing. There's a cheaper, good enough model that keeps changing, but overall this has been constant. - And is the workhorse model for prototyping, and then you run it through the big expensive model for the final product, or what do people use the workhorse? - It's for higher volume use cases, and depending on the application you are building, you might encourage different like, lots of variations of the same output maybe, but it's very application specific, I would say. - Yeah, there's also another dimension, I think, that's like kind of happening in real time right now, which is based on like the different use case, you wanna use the model for. So like when OpenAI released GPT image editing, that model had just like superior text generation and editing capabilities, and for things that require like a lot of text, people started going, and choosing that model versus the other models. So it also tends to correlate with like different capabilities models are bringing, and also like what they're good at, right? So like Kling for example, people really like it for visual effects types of workflows, because they had that kind of data in their dataset, as opposed to some other models. For example, C-dance is very good at like detailed textures and artistic diversity, things like that. So it's really a matter of also like this sort of use case dimension that models excel at. An interesting metric that we saw on Q2 and Q3 was half life of a top five model was 30 days. That's, it's very very interesting to me, where like you know, these models are continuously shifting, like the top five of the models are continuously shifting. - Tough depreciation schedule for the model, for others. - Hopefully they are building on top of the work that they already done, so it's, you know, additive to the end, but yeah. - Yeah, I'm teasing. And the model probably isn't a more turbulent state right now than what the end state will probably be. What do you guys think is the most underrated model? Like what's your personal favorite? - I usually like Kling models for video. But this kind of has been changing because they don't have sound. For sound, we have VO3 and Sora. They are the only ones. A lot of people are working on it, so would love to have more variety there as well. - In which models, I like Revs model. And like Flux still holds like a very nostalgic, even though it's been a year, well, for me, you know, like I still go back to Flux. There's like variations of Flux models now that I like. - I'll go with mid-journey, which is not on file. It's not available on API. I just like the, how they navigated the space, I think is very interesting. Like, they kind of brought this like photorealism, which was, you know, that was like a very big deal at the time, you know, no model could do it. And then now they are more like this artsy model, right? Like, it's not like photorealism is kind of cracked and like no one cares about it. And so now they have this like niche, very artistic like visuals, which is very cool. - Yeah. I'd love to try the market place dynamics a little bit. So I understand your business as a little bit of a marketplace where you aggregate developers on one side of the market, that's the demand side. And you aggregate model vendors on the other side of the market, that's the supply side. And the model vendors are both, you know, proprietary APIs, model labs that view you as a distribution partner. And then also open models that you host and run yourselves. And so maybe talk a little bit about for the closed model providers. You have partnerships with OpenAI, Sora, with DeepMind on VO. What's in it for them? Why did they choose to partner with you? - We were one of the first platforms that accumulated the developer love and you know, following from that, these developers work at big companies. So they started working with us. And we really built the platform for simplicity and being able to get going really fast. And because the thing Bhatuan mentioned, the half life of these models is really short. People usually work with many different models at the same time. So we were able to claim that we have this big developer base that love the platform and not tied into any single model and here for the platform. And model research labs see this and they use the platform for as a distribution channel and tap into the developer ecosystem that we built. On the other side, this helps us with the next model provider because they see all the developers. They wanna be on the platform as well, which attracts more developers on the platform and creates a very nice positive flywheel for us. - Yeah, it very much is a marketplace business. And for developers, it's a single choke point to be able to access multiple model vendors. And to your point on like the model space is changing so quickly, I think they really do value that. - Yeah, we call it marketplace plus plus because we get to provide infrastructure to the research labs as well also to the developers. So there's additional benefits with ties into the flywheel effect that we are creating. So marketplace plus other services next to it. - How do you position yourselves to get, in some cases, day zero launch access, sometimes exclusive launch access to models like Cling and Minimax, have you done that? - Yeah, throughout the last two years, you were able to build a very robust marketing machine as well and this is our connection point with the developers who are on the platform. Every time we release something, this creates another opportunity for us to introduce a new capability, introduce a new model and model developers also see that and we usually do call marketing together and part of that call marketing, we get exclusive release access for a certain period of time, sometimes forever. We have a couple of competitors that are on the smaller side so model developers wanna work with the biggest platform out there and increasingly that platform is ours and we get to have these exclusive benefits with the model providers. - That's awesome. What do you think it is that the open source model ecosystem has been so vibrant for video models. It almost feels like the tech models are just consistently a generation behind, whereas in video, there's so much that's happening in the open source realm. - Video and also image editing as well. - Why do you think that is? - It started with stability. They first open source stable diffusion and got insane adoption and almost the same team then started Black Forest Labs and they knew the power of open source, how it helps them create the ecosystem and with image and medium models, the ecosystem actually matters when developers are training Laura's, they are building adapters, they are building on top of your model. It really brings free marketing but also creates stickiness. So they develop, there are still people who are using stable diffusion models because they like that ecosystem because it was so open. - Yeah. - So the Flex team saw this from their experience at stability and they had a very smart strategy of having at least some models that are open source, some that are close source and a lot of video model providers that came after is following the same playbook because you can have a very robust ecosystem. It gives you a lot of advantages in terms of marketing, in terms of dual upper love and I think it's gonna keep going like this. - Yeah, Thomas. - I wanna add on to that, the domain is also very interesting, like I think in the visual domain, like ecosystem actually matters more. Like I think when like Lambo 2 first came out, there was like many fine tunes out there but like if you actually downloaded it and start using one, like you can't, I mean. - You can tell it's the Flex team. - You can't tell like the difference. Like you can't really, you know, if you're using a, I don't know, like a control net, that concept doesn't even exist. Like it doesn't, you know, language models are a lot more generalized. So you can't really understand like the difference if you were to actually fine tune it, right? So it kind of just ends up being very monolithic. As opposed to like in the visual realm, it's just like any small adjustment you make to the model, it can actually, you know, it can actually have huge implications, right? And so it's just, you know, very fertile ground for like a lot of customization. - Yeah. I mean, speaking of mid-journey, David Holtz, one of his, one of his quotes that I like is, you know, he's curating the aesthetic space with mid-journey. - Yeah. - I very much think, you just have this combinatorial explosion of styles aesthetically. And I think that's the reason why it's some of the, I think it's all the models on your platform are fine tunes of other models, right? - Yes, yes. And like the thing is like even if you add a lot of diversity of aesthetics, like if you, if you actually train on everything, like if you have trained on too many, you may not be able to like actually get the exact, like like there's so many times you want the exact aesthetics. - Yeah. - And then you still, you may still have to like fine tune the model to get exactly the output you want. Whereas like with L&M's, that's not really like how you operate. You don't exactly want a particular outcome. It's like a different, it's a different problem. So this is a lot more, you know, it's very subjective. So like you kind of have to do these like post-training things on top of the models. Sora is another like good example. Like Sora too is very fine tuned on like social looking stuff, right? And so you could probably, you know, you can have tens of different styles and you still want to probably push the model towards that direction with post-training. - Yeah, absolutely. - It all depends on the use case too. Like customer support chatbot does not need personality. Like you want it to be as vanilla as possible, but you are talking about few makers, marketing teams. They all want to add the personality of their style or their brand. So they want to have greater control over the outputs, whereas maybe in L&M's, that's not necessarily through all the time. If you have an agent, if you had in code generation, there's no equivalent of style and personality. - Yeah, okay. That's a good segue for us to go one more layer up the stack. Let's go to workflows. What is the average developer workflow inside fall look like today? - They are using many different models first of all. So if you looked at this up recently, our top 100 customers they are using 14 different models at the same time. These are sometimes changed to each other. So one takes to image model, one upscaler, one in which the video model all part of a same workflow or like a more complicated combination of this part of a same workflow or different models used in different use cases. I think that's the most interesting part, the variety of the models people use on the platform. We do have a no code workflow builder as well. We built this in collaboration with Shopify and this is usually very good for their PMs, their marketing teams, the non-technical members of the team who are playing with these models. It's really good for trying different things, really good for comparing different models, but eventually this makes it into the product as well. We can reach to this workflow through an API. It's been very popular recently and more and more people in a typical software engineering organization is now interested in image and video models. So the users of this platform has been increasing. - Okay, so the average workflow is not just text to prompt, is it not? Create a five minute commercial that doesn't. If I wanted to create a five minute commercial, what would the workflow be? - Yeah, so for this reason, people actually prefer opens, like that's one of the reasons why people prefer open source models because they get to have more control over the model and they can add things here and there to steer the model towards the outputs they want. And when we go talk to studios or more professional marketing teams, they all love working with the open source models because of the pieces they can replace and control they can add into it. And then these workflows are usually like the ones, if you've seen any big, comfy UI workflows with many different nodes. It resembles those where each different piece of the model can be replaced to create more control for the creators. - Got it. - Yeah, and I think like what we have like our workflow tool, it's not the final form of like there's almost like another layer of abstraction maybe on top in terms of workflow. And like as we talked to like these studios, we actually figure out like there's so many ways of, just like there's so many ways of using Photoshop. Like there's no single workflow. In fact, like there's probably like based on your role, like you're a marketing person or you're an animator or whatever, like you have different workflows, right? And so I think that is also emerging. Like as more and more like professionals are actually starting to use these tools. Like you see the emergence of like a very particular workflows, right? One of our favorite creators is PJ Ace. He actually like shares his workflows online. And every time like he posts things, you know, every month he actually has like a different kind of workflow. It's really driven by like the new models, like based on new model, he may have a completely new workflow next time. I think once like we sort of reach some sort of, I guess like some sort of productivity and you know some professionals actually adopting these tools, there will probably be more sort of standardized like best practices around using these abstractions. But like you know, it's not, I don't think anyone knows like the final, final format. And it's like every day we see new things and we try to like update our product to make sure like it caters those people. - Totally. One of the workflows I'm seeing somewhat commonly is you have an idea for high level what you want and you type that in and then I'm in the aesthetics that you want. And you iterate on the aesthetics from an image model and then use that image model with the aesthetics you want to then generate a series of images which then form the storyboard. - Yes. - So to speak. - And then he's cascades down. - Exactly. - And then the video model's kind of interpolate in between them. And it's funny because that's actually how, you know, that's how you know, Pixar and all these companies work. - Exactly. - In terms of storyboards and so. - I think it was a cost thing in the beginning. Like that's why they had to do it like that. But like it actually also makes sense, right? It makes sense in so many ways to do it, to do it like that. And yeah, they call that stuff pre-production and then, you know, post-production, right? So pre-production is all the tooling around storyboarding, et cetera, like that's what everyone does even today. Even though it was like a very cost thing. Now it's more of a speed thing. - And AI makes the workflow, you know, very interesting where you have everything laid out and let's say a new model, new text image model comes out. They built it in such a way that, okay, you can press a button. And now all the different combinations are gonna be generated with this other model. And then you can like generate all the videos again. We've seen those insane workflows. You wanna update one thing. And the whole thing is gonna cost like a thousand dollars to be run it again. But these individuals, like they spend a ton of money on creator platforms. I've seen bills like half a million dollars just spent by a single individual. And maybe even more when it's a small production studio, stuff like that. So it's pretty incredible. - Totally, wonderful. Okay, speaking of studios who are building on your platform, let's go, our final layer of the stack. Let's talk about customers and markets and then what the future might hold. Maybe what are the coolest things that people are building on your platform today? And are they what we would think of as traditional media businesses or are they net new businesses? - It's all over the place. Like what's so exciting about this space is that it just goes across like all of the, you know, markets you can possibly imagine. I'll give you some more, I guess long tail stuff first 'cause it's super fun and interesting. There's a security company that's building on top of all and they basically have these like trainings and the trainings are generated on the fly. And the content is all dynamic. Obviously they have some scripts. I'm guessing to kind of fit like the curriculum, but like the content you get, you know, per person is all dynamic. - This is Brian Long's company? - Yeah, this is adaptive security. Yeah, they do, they do some really cool stuff. I think that's one of the like most unique use cases. You can see how that translates into like rest of education. I think that market is like kind of picking up. Another one, I think like, you know, this is more common use case, I guess, is like AI native studios. You mentioned like the Bible app. That was one of my favorites. It's called Faith. It's one of the like highest ranked apps on the app store. And yeah, they have like stories for each of the stories from the Bible and they're like really well produced. And you know, this is sort of category of AI native studios either in the form of, you know, applications or like they're doing like, you know, feature films and you know, series and things like that. That's a huge category. So I would call this like maybe new media or like AI native media and entertainment. There is also a lot of like design and productivity like out of our public customers, like Canvas, one of those, Adobe is one of those. So they're integrating kind of like in this, you know, in this older tooling, they're integrating new models. Ads is a big one. So and ads kind of come in many flavors. Basically, there's like the UGC style ads, like the stuff you see, like as a person, you know, demoing a product, that's like a very big category. So AI generate versions of those. There's also kind of like older styles of ads, right? More professional looking higher production. Maybe you saw the Coca-Cola ad that came out recently. - That's a controversy. - Yeah, yeah. So that's like a kind of a higher production, you know, style of ads. But you know, what we're excited about is also like programmatic ads, right? So where you can do personalized, you know, to the degree of like literally individuals, you know, yourself being the ad or in the movies, whatever. So like, that's also a big like growing use case. - I'm most excited for the education use case. I think that, you know, ads is, you know, the backbone of the of commerce and the internet. And so like, like super compelling business case. But education is a market that's like so important. It has never really had that many compelling business cases behind it. - Yes. - And it's part of the challenge with education. I mean, the challenge has been the bottleneck to creating high quality content at scale that's actually ideal for the learner. And so I'm personally most excited about education. - Same, like I really love the education use cases. And I actually think that like chat GPT or like just, you know, LLM's in general. I think they are already solving it in a way. But it's not the right form factor. Like if you actually wanna fully realize like the sort of power that these models are bringing, you actually need to go into the visual space 'cause then, you know, it's so much more compact. It's more approachable. And yeah, I think once we actually crack like visual learning, like through these video models, that's when it's gonna, you know, really just like impact people. - Do you think that the advent of alternative media is going to increase the value of existing IP? So like Mario Brothers, Nintendo, Disney, Pikachu, all these things? - Yeah. - Or do you think it's gonna lead to the democratization of the creation of IP? - I love this question because it felt like, I would say six months ago, it felt like this was all happening too fast for Hollywood, the IP holders to adapt and be part of it. And from our viewpoint, we thought, all right, these, these I need of studios, they're just gonna take over and Hollywood is just gonna be too slow. And this is gonna just pass them and they're gonna be left behind. But this summer, something changed and we've been talking to a lot of usual suspects from the Hollywood. We recently had our first generative media conference and Jeffrey Katzenberg, former CEO of DreamWorks, was there and he met a comparison. He said, this is exactly playing out how animation when it first came out, people revolted against it. It's just gonna happen. You're either gonna be part of it or not. So we are seeing a lot of existing IP holders are now taking this very seriously. And at least for the medium term, I think they are pretty well positioned because they have the technical people who are actually really interested behind the scenes in this technology. They also have the IP, but they also have storytelling and filmmaking know how. You still need quite large budgets. Maybe things are gonna get cheaper, but in the medium term, filmmaking is still gonna be expensive. Yes, AI is gonna make it maybe a little bit cheaper, but we need these deeply technical people who know filmmaking, who has the IP, who know storytelling to actually in the beginning be part of this. And I think they're gonna play a big role in the next coming years in the AI ecosystem. - Yeah, when there's infinite content generation and almost puts a value on the things that are finite, and I think, you know, for those of us who grew up with Power Rangers or Neopets or whatever, there is just this nostalgia elements and this finite supply of IP that really resonates with us. - The opposite is true too, also. There's a lot of new, like we had little toys of these Italian railroad characters. These are characters with no IP, no one owns them. They are completely AI generated from like the internet community. And once you have cheap generation of content, very different permutations of it, things that people like, catch us on, and it becomes part of the zeitgeist. - Yeah, totally. - The opposite, there's signs of oppositive nature as well. - Both are true, yeah, both are true. How do you relate a question? How do we prevent like the infinite sloped machine state of the world? You know, there's this, you know, version where we're just connected to this machine that knows how to personalize stuff for us and we're just, you know, we're just hooked up to the infinite sloped machine. And there's a version where there's, you know, human creativity and artistry and things like that involves. Like, how do you think the world plays out? - I think humans eventually like converge on all the things that are more meaningful in general. Like, I don't know. Like no matter how much slop we fill the world with, I think, you know, taste prevails. And people are drawn to like, you know, experiences that are personal and human. And, you know, I just think that that's gonna happen. One interesting example of this was like, when meta-announced vibes, and then open AI and I saw that too. Like, the reception was very different. And one of the like reasons in my mind was like, vibes was like positioned as this, you know, sloped machine kind of thing where, you know, they didn't have the product out at the time, but it was just like these AI generated, like just, you have no relation to the characters, et cetera, right, like it was kind of this like detached thing. Whereas like Sora really made it about friends, right, like Cameo and, you know, they were very-- - And now you can cameo your pets. - There you go. It's huge, right? So yeah, I think like this connection to like friends and pets and things like that that actually made, and Sora was also like, they were being very personal about it. They were very adamant about like, "Hey, we wanna make this about friends. "We wanna make this about, you know, "these connections as opposed to, you know, "in Photoshop machine." So I think that's, you know, that perception was also, I think a good signal that like, there's ways to work, make this technology work, you know, in a good way. - Absolutely. Okay, I'm gonna get your perspective on timelines and what's feasible today and what's feasible to come. I guess, do you think that we'll see Hollywood grade, feature film length films entirely generated by AI, and if so, on what time? - Well, what does an entire journey by AI means? Is it like no human involvement? Or like-- - No human filming. - So human involvement. - But anything is okay. - Yes, absolutely human editing, but no human filming. - I think less than a year, we'll have like, you know, advanced video models with combined with the storyboarding that people have been doing. You'll have feature grade short films, like less than 20 minutes. I think that's a fair estimation. Like, even today, I think you can like, do really great films. It's just like not enough investment of time is going into these. - Yeah. - But like with enough investment of time and the model quality, I think we'll be there. - I think we're ready there, okay. - And you think it's photorealistic? You think it's anime? You think what categories do you think are more likely to happen sooner? - I think photorealistic is like what everyone is targeting, but like anime would be a cool one, right? Like it's like, you don't see that many, anime specialized models. Why not? I think there needs to be a market for that, clearly. - I think it's gonna be animation or anime or cartoon, like not photorealistic, like as far away from photorealistic as possible, maybe even like as fantasy as possible because filming photorealism is cheap and doable already. Like that's not what costs money when people are making movies. It's the non-photorealistic stuff that's actually expensive. And even if you look at the animated movies, some of the my favorite movies are animated. The Toy Story series, How To Train Your Dragon, Shrek, Rattatouille. And people like these things, not because it reminds them photorealism, it's the storytelling that matters and this created a new medium. I think AI is gonna be similar to animation and how that brought a whole different angle to filmmaking. - Yeah, I think feature films are hard, like because yeah, with photorealism, like you typically, I mean, people usually like the movies that they're like favorite actors are in, whatever actors, actresses. And that's, you know, so it's like one step removed from-- - That's the thing that costs money to get the actors. - Yeah, exactly. So that's the, you know, we first need to build a connection to this AI, you know, AI generated character before we can turn into a film. But I think like, yeah, I think it's among like different kinds of content, like shorts, you know, I think Italian brain rot is an amazing example, right? It was first, like, these characters and then it became a Roblox game and making, I don't even know, like, you know, a lot of revenue. So yeah, I think like AI native stuff is a shorter form content is probably gonna be very, very big. - We saw this with VFX, where like the VFX FX, like one of the most expensive parts of like producing these videos or films is like got like, AI fight very, very quickly because it's very easy for AI to do like explosions, right, like, or a building collapse. It's like almost perfect now. And I think it's just gonna continue along on that dimension. - And maybe facial expressions are gonna be hard. - Yes. - And very hard. - You don't have to do facial expressions. That's gonna be okay. - But now they can do gymnastics. - Yeah, gymnastics are important. Good thing we have a lot of footage of Olympics. - What about you mentioned Roblox? At what point do you think we'll have interactive video games that are generated in real time? - Yes, I think so. I'm very excited about it, actually. Like I think, I think like in one world, I think the sort of next reasonable step for text-to-video. Like if you think text-to-video is the continuation of text-to-image, I would say like a text-to-game is the continuation of text-to-video. Because you know, with a game, you would essentially making the video interactive, right? That's kind of what that means. And I actually think that there is a world where there's like hyper, hyper-casual games exist, but this is like another level of hyper-casual where it's actually discardable. I think we're not too far away from that. I actually feel like pretty bullish on having like these one-time playable games, like very short games. - Yeah. - I think that's probably gonna happen. I think that's a good use case for world models other than any other great use cases, but I think it's gonna happen. - Whether at AAA quality games, well, these models at least, you know, assist and change the development pipeline of those games or-- - Yeah, I think they're already impacting, like at least LLMs are impacting like conversations, there's like dynamic conversations, things like that. I think pre-production stuff is impacted already. I think like kind of side quests, like IP stuff is impacted, right? Like where you have the assets and you can make a mini-game. I think people are using it actually. Not very public, but like that is already happening. I think like using for AAA production or like generating that with a model, that's like, I don't know, at least like three, four years ahead for me, and yeah, I mean, that would be insane if we can actually do that. But you know, along the way there, just like the video space, I think along the way to the AAA, there's like many other things. I think those are gonna be very big. - Yeah, the video model space has just exploded in terms of options, quality, et cetera. As you look ahead towards what's needed to get us to the promised land for everything that generative media can be, do you think that there's future R&D breakthroughs that are needed on the horizon? Like fundamental R&D breakthroughs? Or do you think we're very much in the engineering scale-up leg of the race? - I think the architecture needs to like slightly change, at least like if you think about like scaling these models by 10X, 100X, I think the architecture is a big bottleneck right now in terms of the inference efficiency, where like the more compression of the video space, then that's definitely needed. Like we solved this with image, image most used to be like much less compressed, and like you were operating at the pixel space, and I'm introduced later space, and then like even inside that later space, you took like 64 pixels and made them a single pixel. And now like with video, we are compressing on a time dimension, where we are seeing like 4X ratios. Why not like 24X or whatever? Like you need to like increase that like compression, and like I think that's gonna be a big driver of improving both inference efficiency as well as training efficiency. But like I think like any model, like I think at this age that we are operating, any model you take on the generating media side, we're far from being like scaled up engineering wise. Like I think there's not enough investment being put into it, or like it just started happening in the within the past six months, like Google showed this with like their models, and how quickly they were able to catch up, they didn't need to innovate that much. It's just like they have the resources that can put more effort into it, but at the same time smaller laps are able to demonstrate this because like there's so much like unique and novel stuff that you can do at the data level to train these models. So I think that's also like helping contributing. And there's the factor of like outside like you know, mid-tier laps that raise like 100 to a billion dollars, that's also trying to come up with models to releasing them open source, or like contributing the ecosystem. - Yeah. - That's what's so exciting about this space, there's so much more work to do. Like so far, the research community did the simplest thing possible, the captioned images and train the model on text to prompt. And now like we are doing video image editing that requires a lot more data engineering to create the data sets, but luckily seemingly we have a lot of abundant free video data. We are going to run out of compute before we run out of video data. So that means there's a lot more work to do and a lot more room for improvement. - I mean, earlier on like your cams math also indicates that like if you want to get the 4K video real time, that is like, I mean, that means like I don't know, 100x maybe more in compute or architecture. Something has to give to get us there, right? And yeah, like right now a lot of models are like, not that usable, like for professionals especially, right? Or even for like consumer, right? Like if you're sitting there, like for the best models, you still have to wait like 40 seconds. I don't know, sometimes you have to wait two minutes, three minutes. Like that's not really acceptable, you know, world where like we want everything on demand. So yeah, I think something needs to change. - Yeah. - And probably pace of like hardware getting faster is not enough. I think if that's the case, you know, it'll take much longer, it'll have longer timelines. So I think architecture needs to get better. - Awesome, thank you guys. You made a very high conviction that on generative media as a theme, I think way before it was obvious, I think we are just at the start of, I think what's going to be an explosion of generative media. And it's been really cool to hear about everything you built from the kernel optimizations and the compiler, all the way up to the workflows and what you're seeing from customers with new and old media like. And so thank you for joining us on the show today. - Thank you, thank you so much. - Thanks so much. - It was a lot of fun. (upbeat music) (upbeat music)

Podcast Summary

Key Points:

  1. Jeffrey Katzenberg compares the resistance to AI with the initial backlash against computer-driven animation.
  2. Fall is a developer platform for generative video models, providing access to various generative media models.
  3. The team at Fall focuses on optimizing video models, which present different challenges than text models.
  4. The team emphasizes the importance of focus and obsession with generative media models for achieving top performance.
  5. Running video models requires significantly more computational resources compared to image or text models.

Summary:

During a discussion with the team from Fall, comparisons were drawn between the resistance to AI and the initial opposition to computer-driven animation. Fall serves as a developer platform for generative video models, offering access to a wide range of models. The team highlights the unique challenges faced in optimizing video models compared to text models.

They stress the significance of focus and dedication to achieve top performance in this field. Running video models demands significantly more computational resources when compared to image or text models, with the team noting the exponential increase in computational intensity for video processing.

FAQs

Generative video and image models were overlooked due to unclear industry use cases and slower research investment in the past. Now, with more capable models and real industry use cases, the market is expanding.

Fall started with a focus on running image models faster, which led to rapid growth. The team's early investment in generative media models and a shift in focus ahead of competitors contributed to the emergence of Fall.

Diffusion models for video are compute-constrained, requiring efficient execution to denoise thousands of tokens simultaneously. In contrast, autoregressive models for text are memory-constrained, optimizing for large batch sizes and memory bandwidth.

Fall's obsession with generative media models and dedicated team focus on inference optimization have led to claiming the top spot in performance benchmarks. The team's specialization and adaptability keep them ahead of competitors.

Fall faces challenges in optimizing GPU cluster management, routing traffic efficiently, and adapting to changing customer traffic patterns. The infrastructure must ensure efficient loading of models and smart utilization of GPU resources.

Fall excels in the emerging field of inference optimization and infrastructure management, areas where hyperscalers lack expertise. The company's focus on generative media models and cost-effective infrastructure solutions give them a competitive edge.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.