Human-Centered Animation at Scale: Balancing tech, talent, and budget | Sarah Watling CG Pro Podcast EP 101
59m 13s
Sarah Wattling, co-founder and CEO of Jolly Research, discusses her unconventional path from live music production to tech leadership, leading to the creation of a facial animation tool that generates full facial expressions from audio, text, and metadata. The company, now nine years old, uses a procedural, algorithmic methodology that contrasts with trendy machine learning approaches, producing clean, editable animation curves that reduce cleanup time and computational overhead. This design has proven advantageous for large-scale projects like Cyberpunk and Obsidian’s games, where processing hundreds of thousands of lines quickly is critical. Jolly’s focus extends beyond humans to stylized characters and creatures, with customizable profiles for unique sounds. Throughout its history, the company has faced pressures from industry crises (COVID, strikes) and AI hype, but stayed committed to its principles of small, lean, and focused development. The output is designed to be close to final polish, minimizing the need for manual curve cleaning, which is a significant cost saver. Jolly is now expanding into cinematics and interactive characters, while constantly optimizing for speed, file size, and user influence. Sarah emphasizes the importance of balancing innovation with practical, repeatable solutions, and navigating the noise of emerging technologies to deliver value to studios seeking efficiency without sacrificing creative control.
Basically, it automatically generates facial animation, so full facial animation from your audio or audio and text or audio text and metadata inputs. The input of metadata from script is our biggest, you know, focus right now, as well as just constant op-optimization. It's always been about facilitating creative excellence, artistic excellence, and getting people to their design. Like, whatever it is, the execution of their vision, I guess, as effectively and efficiently as possible in the parts of the pipeline we have some influence or we're integrated. Welcome to the CGPro podcast. This is episode 101, and we are excited today to have Sarah with us, Sarah Wattling. If you enjoy today's episode, like and subscribe, obviously, but check out our website at becomecgpro.com. We have a newsletter sign up there. You can find out all about these podcasts and the shows and all the community events that we do in person and online. So today's special guest is Sarah Wattling. She is co-founder and CEO of Jarlie Research, which is an amazing innovative company that's doing all kinds of things, but mainly focused on doing AI lip sync in games and cinematics. And I'm going to bring Sarah on because she will explain it a whole lot better than I can. Sarah, thanks for having here today. Thanks for joining us. Hi, I'd like to see you again. Yeah, great to see you too. Yeah, it's been minutes since NAB. And yeah, we had some fun there and you were demonstrating the thing that you do. But before we get into all that, tell me a little bit about yourself. What's your background? What's, how did you get here? It's a really interesting place to get to. I'm really fascinated by what led here. Were there any kind of inklings of this when you were younger or how did you, how did you definitely a unconventional route, but I don't think that that's inconsistent among a lot of my peers now. I started in live music production at the jazz festival in Halifax, Nova Scotia, which is tiny, and then moved on to larger scale productions from there in England, actually at London, a club called Cargo in Shortage. And when I was, yeah, when I was working there, I actually always wanted to be a writer. So that was sort of the angle was always trying to get close to either promoters or producers or journalists that were looking for somebody who was scrappy and just wanted to get whatever they wrote into print. It turned out if you wrote a desicc press release, it would often get printed for a bit. So it was kind of a nice spot to be in at that time. I was like 19, 20. I was there for a few years and got to travel over a bunch of Europe doing that. And I was in love with Arts Administration. So just not necessarily the work of it because that is pretty thankless and brutal. And when I mean work, I mean the day to day lift, but the payoff when it did was always so gratifying, bringing artists that were not necessarily easily found or easy to sell really to really appreciative audiences. And you mix that up with some popular, you know, you know, headliners. And it was just a really great experience. We moved to Toronto about 16 years ago with my partner and co-founder, Piff Edwards. And Jolly is the outcome of his earlier thesis work as a PhD at the Dynamics Graphics Project at U of T. There he met Karin Singh and Chris Landruth, Eugene Fume, who all have on their own pretty significant contributions to the earlier development of Maya to graphics, you know, particle simulation. That kind of thing Chris himself is an Academy Award winner. And when their first publications started to get attention, it was actually quite a lot of attention. And they needed someone in the background to just start picking up the phone when it rang or answering email when it came in. And it really just spiraled from there. We've been a bootstrapped company since we've started. So, you know, understanding that the balance and the tension between trying to scale demand and also maintain focus on a roadmap without sort of accumulating a bunch of bloat has been a huge part of the way we've run our company, which has been small, lean, focused. Yeah. So, I did the tech transfer. I started managing the contracts. I did all of the discovery calls, figured out a lot of the earliest, you know, barriers to adoption and our first product offerings, which was exclusively lip sync. And, you know, a couple years later, we added the whole face and and it's kind of spiraled from there or grown from there. Grown and spiral. Grown or spiral, whichever, whatever the day. Yeah, no, it's an incredible adventure or doing all of those things, especially starting a business, which is very challenging and rewarding thing to do. But very cool to hear that story and definitely echoes of seeing some of those people going through my own career. Chris has worked particularly and remember that from being into Maya for about 20 years and they're seeing some of his short films and films. Yeah, very cool. How long have you guys been doing this for? So in the Jolly. Yeah, this will be our ninth birthday, actually at the end of this month. So, nothing not that huge next year. You know, if we make it to 10, that'll be pretty exciting. And it is exciting actually to have been around this long. It's not been an easy. That's pretty good. Nine's pretty good. And it has not been easy. Right. I think everyone in our space can share war stories about the last four years. The sort of post COVID fault drop off has had a huge and continuing impact on the spaces as everything has sort of equalized after that huge boom. Yeah. Nine years is good. It's amazing. I mean, you beat as I said, you beat New York's statistics for a number of companies making it past five years. I think it's around 50% and then it drops off significantly towards 10. So you've already beaten the odds to get to nine. CG pro has been around for six. So, you know, you're already like a few steps ahead of us. We're all hoping to keep continuing to do what we're doing. But in this crazy time, it's a very interesting time. I think a lot of opportunity at the moment because there is so much change. You mentioned some of that change. So, going through some of those moments that you mentioned at COVID, the film strikes, the. - Right. - Instead of. - Yeah, I. - Right as strikes. Yeah. All of those things, any one of those things is a huge trigger of a crisis. If you like, a crisis being a point of making decisions. As opposed to it all going wrong, it's very often thought of as it all going wrong. But it's really a moment to reconsider to make decisions, to pivot if necessary. What's it been like for the past nine years going through all of those? You've been through like more triggers than the average 10 years, I would say? - I would say the most impactful ones are the ones that we learned the most from. We're definitely. A, when we started in the earliest days, as things were rapidly evolving from, you know, earlier iterations of machine learning research and choosing to stick to a sort of judicious approach of machine learning application in our own product development. That was a challenge, I think, to do just because by choosing that, you were almost like choosing whether or not you were going to disingenuously signal that you were still heavily using machine learnings that you'd have access to available funding grants when we were still at the university and that kind of thing. It was always a real like that really pushed me as a writer to position the innovations and the opportunities and the benefits around what we were doing from a methodology standpoint. From a procedural and programmatic methodology where we study phenomenon, reduce it to its algorithmic essence and then implement that from an engineering standpoint and then build the UI around that. I think no one. It was innovative, that particular moment was very innovative at the time because of its opposition to the trend. Lots of game-diving,
developers will opt for procedural methodologies for because they're as lightweight as they are. Deterministic quality presents its own set of advantages and disadvantages, of course, just like everything does. But you're always doing a risk analysis. And you're choosing, based on what's going to get you the greatest benefit for your aim. We stuck pretty with a lot of conviction through lots of instances where we were getting either encouragement and/or pressure, depending on who it was from, to adopt other methods. I remember when we were very staunch about having a two-stream input system where it was both audio and text for analysis. And we'd done a lot of research and publish research into what audio only and just found that the results were not consistent with the value that we were delivering to our customers in our first product, which was editability at a very granular level. And there's a certain amount of accuracy that's required for that at the curve level, at the curve data level, which to this day is still how we build the product so that our output stays light, editable light from computational footprint, also in terms of accumulation of file size. We don't produce images. We produce animation curves. And those curves are key frame curves. That kind of thing. As we've gone on, emergent technology has also put other pressures on it. The biggest thing I think for us, and I think it's consistent for people in the industry, as well as when something new is introduced that has shows so much promise or is positioned with delivering so much promise, whether it's on the creative capacity side or it's on the cost reduction side, it creates a lot of noise. I think from a company that's just running, trying to run repeatable sales cycles, all of a sudden it introduces a whole new thing that you have to quickly get your head around, understand what's real, what's not real, and where does your solution sort of sit in an easy to communicate comparative to your customers? That's happened many times to us since we've released the product. And now even on this side, some of those initial engineering and design principles of keeping the software and the code based small, keeping it algorithmic and repeatable and modellable, and editable has proven to be a win for us. It still is a win for us today. It's kind of interesting or very interesting that you started before all of this was cool, before the hype essentially. What was it like pre four years ago when some need to start kicking up by everybody and their dog is interested in AI at this point, but when you started that was not the case. It wasn't like it wasn't started, but the easy to access, not alternatives, but options or new tools weren't as plentiful at the time. So there definitely wasn't, like at the research level, which three or four years ago, we were still really heavily embedded in the academic space just by virtue of our founding team and the amount of publications that we were either consulting on or contributing to or publishing ourselves and our, you know, Karin Singh, for example, he's still very active in publishing. He's working on a lot of exciting work in Gaussian splats. Right now, for example, that pressure was always there, but the sort of easy to access and shiny examples of what could actually be implemented was not the same four years ago as it is today. And I see that that sort of static and noise layer is deeper, but it's also more interesting because it's less four years ago, it was like flashy and then it would dissipate really quickly because for the most part, either the solutions weren't repeatable or really implementable outside of a development environment or outside of a very specific custom individual use case and a lot of studios are developing their own technologies. So in addition to whoever our competitors are, you know, oftentimes when we're in a discovery call or that kind of thing, we're also competing with internal talent as well and internal internally developed tools. So fast forward four years from now, not only is there so many more models that have been implemented into easily accessible applications, you have now two, I would say at least two years of, you know, commentary and dissection and trial and error and demonstrations that aren't just, you know, very shiny, sizzle sort of demos produced by the tool developer, but by users themselves and those are, that's where I am always now looking for. You know, where's the real analysis, what's meaningful, what can I bring back to our CTO? Like it's rarely a video, it's mostly like a screen grab of like a five or six message long thread of like, "Okay, this discussion is really interesting, "you should take a look at this." Yeah, it just, it interferes with sales cycles, you just have that much more to talk about. And so it's not from a business economic standpoint, you can consider interference from business research standpoint, it's valuable. So it's gonna both, it's good not, neither good nor bad, but it's okay. That's right. Yeah. Yeah, you're just keeping up with that much more information all the time. Right, yeah, the public, I'm sure, the public's perception of shifting significantly and you happen to manage the story and the myth-busting and all of that conversation. That's right. It's very fascinating. What, tell us a little bit about it, tell us about the, about Jolly, what it does, I've seen it myself, but tell us what it does. So we have, it's basically it automatically generates facial animation, so full facial animation from your audio or audio and text or audio text and metadata inputs. The input of metadata from script is our biggest focus right now, as well as just constant optimization. The bigger a project gets like FERP from our standpoint, it's about trying to enable the degree of fidelity and input influence that a team wants to put into our system to generate out of the box animation that's closer to a state for polish. Then it was without that direction. So we look at metadata as instances of either writer input or director input. Right now it's focused on, it used to be sort of more of a monologue system. You would process lines per character and the biggest value proposition, I would say would be about scale. So a project like Cyberpunk had something like two million lines of dialogue. And it was a pretty amazing first project to work on. Companies like Obsidian, last year they put out three fairly commercially successful games, each building on the pipeline for Unreal that we had developed with them, sort of not custom, but they were a really like generous sort of co-development partner in terms of providing us with the kind of user feedback that prevents us from designing things in Asylum. And also their own expertise. And so yeah, now I would say that our solution is no longer just a background utility. I think people are using it more and more for cinematics. One of the qualities of the output animation that seems like a throwaway or boilerplate bit of text, which is like the easy to edit. That kind of thing is the cleanliness of the curves and the speed of processing. So we can process your in-game and like all of the animation, character animation in a game, some of like 300,000 lines two or three times a day, which allows you to actually go back to those points where you made those changes and actually continue on. That instead of having to wait two or three days, which was what it used to be three or four years ago with the first iteration, making things. Once we've got a workflow sort of locked then it's the work shifts immediately to can we make it smaller, can we make it faster, can we make it so it doesn't generate four or five X additional files or duplicates or whatever and really boring stuff, but that does tend to make the difference. Just getting back to the curve thing. It's
We've had a number of instances now where your people are either using the software and the settings within it to match performance captured performances so that it's actually capable of hitting close to close that much closer to where you are going to be wanting to polish. The best part about that is you're really not ever having to clean those curves. If you looked at it in a graph editor, it looks exactly as it should and it's not you know no noise, no jitter, you know all of that. That's a lot of time and money. Like when you're pointing some of the promo material about invisible costs, that's a huge piece of that. When you think about the cost calculator for high quality animation for other cinematics or in-game cutscenes, that kind of thing. You're usually starting with capture and then quite a lot of individuals whose titles are facial animations, but they're really just cleaning motion capture data and then polishing. So that equation to final render is a lot faster. I'm finding or studios are finding using Jolly out of the gate, which is kind of a nice place for us to sit and it's kind of paved the way for our transition into deeper into cinematics and more sort of feature style and independent film use. I think what you saw at NAB was our interactive character, so our so-or the Ruby AI character, which is a pretty fun project that we've also been working on. Yeah, that was really cool and anything obviously you could be able to get into the chat experience with a machine more than ever before and being able to put a face to it, having it be more than text in the chat. I think it's really interesting. It's interesting having both directions for it, I guess. I'm sat in a global object right now doing a shoot today in a mocap space, so it's very very topical for what's going on right now. Mostly body mocap that's going on right now, but a little face as well. You guys focus mostly on the face though, right? Yeah, when we started the goal was actually to be able to conceive of the full 3D character. And so it naturally contemplated body, but so much is going on on a face that we still find that there's more and more work to do and there's also, you know, what he, like we were talking about before when you're making decisions at a certain point we're always evaluating, is this something that we want to build ourselves or do we want to look to acquire technology to bring in house in order to facilitate that or do we want to partner with great teams that are building the complementary components of a full system like that. You know, that's a, that's, so yeah, we're still doing mostly faces, but the faces have moved beyond humans, so much more stylized characters were handling very easily, as well as creatures. We've animated dragons and how does that work? And do you put in a meow and you get a cat moving? Yeah, so sometimes it's you, yeah, because the system now pretty much handles any kind of, it's the models that the audio processing are sort of built on are have processed enough data now and are robust enough that it really doesn't take much to make any kind of corrections for things that don't, you know, that are out of a cabular area or something like that. So sometimes you would put in, I can't remember the technical term is it when you write what it sounds like? It's not phonetic, it's actually a different term, but you know, in the place of that type of sound, but and then again, being a model based or an algorithmic solution, once we know what that's supposed to look like, you can actually, we can write that into the code that and then it'll always do that for that particular type of character. You can then put that into the config of their profile and that character will move in that way. According to the controls and, you know, our off-the-shelf product has its own rig and it has its own control scheme, which is great because a lot of smaller studios don't necessarily have that. When you would, you know, I just found out, you know, fairly recently, when you set up our software on another, you know, studio's rig and the rigs say their control orientation along the xy-axis is either not set or set backwards or something like that, like our connections setup will just automatically fix it. And then they can tell us later like, oh, actually that was intentional and you can like, well, you can also undo that. But it will kind of set it to what the industry standard would have been. Right. And what's the experience like for somebody working with it? Is it kind of possible for a director to just to communicate with it directly and say, I want to be able to converse with it if you like without a technical experience? Coming. Coming. Yeah. So it's designed for right now for animators and technical animators and riggers. But as we said, as we've moved more into the cinematic space and to film, it's a deep tool. It's a very technical tool. And the UI improvements have been focused mostly on creating more of a shallow end for less technical use and building in more capability where technical teams are smaller. So there's quite a bit of rigging sort of expertise and care taking, like, put into the way that the control board is set up. It's quite easy to use with a, you know, a basic level of fluency in rigging animation and technical animation, so on the integration side. From the writer and the director standpoint, like that's now the direction that we're facing or that we're developing towards. And the benefit of having a solution that's so discreetly described and built is yielding some really exciting results for a genetic UI actually being integrated into it to sort of so that the performative task that you used to, we used to be able to reduce from many to at least not repeat are now possible or will be possible just in natural language. Oh, cool. That's great. So yeah, so it's like I will keep talking about AI all day pretty much at the moment. A lot of the sentiment as you have recent guests that we had on the CG Pro Show was talking about this as well about how I think it was filled, saying how a lot of this is a UX problem as well as it being a technical problem. So like making good user experience in UI around something can be the difference between success and failure or adoption or not. And it's actually really hard to do. And I actually I used to be a software engineer back in the day 20 years ago and then there was no UX. There wasn't such a term. You, you, you eyes were made by the engineer and they weren't very good and now does changed. But yeah, really trying to make things nice and easy to use is is pretty hard to do. Yeah, I think it's an it's an interesting balance that you say that we've you know we've lost time. I would say like looking back over focusing on UI before the engineering underneath was robust because it's almost like cart before the horse. But if you wait until too late and you have engineers doing the UI actual work, it's like fitting wheels on a moving car and then it comes with its own challenges. So that's a really tight, tight rope to walk. I would say we're still learning that, you know, every day where the where one team has to have a little bit more control over the way things are either laid out or how you're speaking to this level of control depth. Yeah, but even when some of the earliest iterations of of generative models were starting to be released and chat GPT became, you know, that much more of a household word that like your mom is now talking about it are like people that have zero understanding of what any of the underlying research or technology is is trying to do or is doing. They're they know that all this I use chat GP for this and I think the thing that from a UI perspective that is like globally or universally captivating is that the idea of just speaking something into existence, right? The prompt based UI, yeah. Maybe one day thinking into existence, who knows? Yeah, it's a it's really interesting and the I like how it's making something
that's intermediate as well because a lot of straight to image or straight to video becomes really hard to edit afterwards which I'm sure you're very aware of. And this kind of doesn't have that issue because you can go in, you can go to curves, you can change it after the fact. How do you see this more going into real time or rendered or is that I'm sure now driving like open pose into generative AI or video into generative AI? What's the kind of the split? Yeah, I see mostly still a lot of persistence in like in the 3D in the 3D pipeline. I think the 3D pipeline offers a pretty reliable anchor to go back and make changes without having to throw everything else away that came either that it's either based on both so you can get to that error or afterwards and forgive me if I'm not answering your question directly because I may not have fully understood it and maybe I'll just wait and let you clarify them. Oh yeah, no it sounds like you're going to go away. I just want to make sure. But I do see a lot more interest and experimentation going into how far generative as generated assets can be integrated with their 3D counterparts and their interoperability. And that's a space that we're definitely focusing on as well. It just continuing to participate, support teams that are doing those experiments, taking advantage of opportunities to play in those sandboxes where we can and at ANNSE we'll see a little bit of that the results of some of that just how easily it is for our output data to import into other platforms whether it's engine or 2D or what's the vibe like what's it been like because it's probably changed a bunch over the years but now AI gets a lot of heat around it because of the issues that it has I guess with copyright and control those two that kind of the main ones that people have issues with but particularly around the copyright and it also makes people upset who do things for a living that these kinds of automations are saying that they can do instead or with what's the reception been like over the years because this is doing some of that and has been doing it for a while. Yeah, I think it's part of its twofold. The say you can do it's about being it's about repeatability so can it can it do what it said it was going to do and then it can it do it persistently over time and repeatedly I think there's something unique about a game build where that entire bit of software has to now operate reliably and as expected in on the devices of you know millions of millions of people and whether it's being downloaded or it's people are still buying them you know in a hard copy form there's you know the second consideration is about the size what that means to the size of the build and the cost of doing it and then I guess third is cost control itself like maybe you might have some a team that's come up with a way to you know manage or encourage or propagate repeatability to a degree in an area of the development that works for them you've got the additional cost associated with running these models all the time that's why I think that it's you're having a hard time seeing them in truly interactive on-prem type experiences so for like avatars again like if you're heavily reliant on cloud-based you know models to drive either the text speech or the whatever you know even in the ruby project you saw that we had the whole goal for that demo was for it to be like 100 percent local and the hardware is really catching up to being able to process almost everything on the CPU level and and that's exciting but we still actually had to go out for like go online for the TTS because the model was so big for us to run that and so being able to deploy that that cost goes up it's a lot and then back to the copyright issue like you're changing ownership right that's how so for you know from our standpoint we keep our software very agnostic and what we ship is exclusively owned by us the contracting is very you know explicit to the T around where our software begins and ends and we we really do our best to keep it that way so if we've used any kind of model and development you know the next stage is to do it without it you know figure out how to how to do it without it so that we don't have that that problem and we don't have that ownership issue that's propagating I think in in games still they're the most judicious industry around you know assessing whether or not a technology is worth the cost or the risk mostly because I think the developers of build have been building you know early iterations whether it's been programmatic or not like the skills and the experience is just there to assess it as soon as you get further out from that we talk about avatar applications in a non-emany environment it's a lot more education involved in terms of what they're contemplating and they care more and less depending on the technical thing or the copyright thing depending on what their application is. Got it yeah then you mentioned the the local model thing is it helped or made a bet or a worse having more local models coming online like jamer or that kind of thing. Oh definitely better I think at least for us it gives you something to shoot for definitely in terms of what we can build and how we can get things down and our own cost calculation you know as even like I alluded earlier to development that we're doing around being able to cost of fit you know director profiles that is currently contingent on on libraries we didn't build you know or models we didn't build for the agent component when we get to ship stage you know that's that's what our team and Emily goes already contemplating is how much time and effort do we want to put in building our own or is there is there a team out there that's you know really focused on building these tiny tiny models and there are so yeah that's a that's really something this week. Interesting space yeah the all the little teeny two bees that are run on a phone and up to the 31 billion parameter versions that still will run on a laptop a good laptop yeah yeah it's really really interesting space they seem very capable yeah it's exciting time and I like that that space is I think we're it's the right place for us to be looking for you know where we're going to continue to build the software because I don't think we're ever going to go in a direction where we are wanting to remove sort of critical components of process as a you know in terms of what we are interested in and what inspires us like simulation and imitation and sort of perfection in that way has never really been our our goal or our interest it's always been about facilitating creative excellence and artistic excellence and and getting people to their design like whatever it is the execution of their vision I guess as effectively and efficiently as possible in the parts of the pipeline where we have influence or we're integrated so yeah it's all all about assisted yeah and in that sense like what's your thoughts after watching a lot of people doing this and and using it artists and technical artists and various different types of people how how do you feel now that it's controlled the best I guess is it with somebody who can perform really well somebody who can write really well or somebody who's really technical or a combination of all of those things yeah I think that the this idea that these top level skills will no longer be required because you can just vibe code or or that kind of thing is all consistent with a lot of the false kind of narratives I I find kind of tedious around AI and I think it really distracts from what's interesting about it but you know they're driving the investment and the economic promise that you know we can just one guy alone can do
this and you know that kind of thing and and and short I think it is very empowering for skilled visionaries because they have the understanding of how to execute they know you know with something is a mass versus something is that is intentionally a deviation you know that's big I think that was that was Picasso right that was this whole this whole thing was like you have to understand the rules and of of visual art and they are there's some real rules for abstract art to be not a mess I think that will persist I think writing ironically will persist as one of the most high valued skills especially in spaces like advertising and short form content and that type of thing where the video gen capabilities are reaching the point where the what they are able to maintain and and what's what is actually able to persist is is now catching up to exactly the right time frame that is appropriate for some of this content and alone only get better and better how that's written is going to make all the difference you know between something that you have seen before versus something that's entirely neat and interesting to audiences because it's that's a triumvirate right or a trifect yeah triumvirate or there's another word my assistant told me that recently I was like oh I can't remember it but it's a relationship right between the creator yeah that's though it's not that it's it's a derivative of both triumvirate and trifecta try something yes it's a try something between audience creator and medium right as like there's a feedback loop they're not disconnected they're connected yeah and yeah one influences the other's really interesting that there's a lot of theory going on at the moment or myth or bullshit or whatever you want to call it in in the space in what people people always seem to want to be able to predict the future so they can come back and say that they say that they said that was going to happen and the most often wrong most people are most often wrong when when trying to predict the future so it's very it's very interesting there's a lot of theory being bounded around at the moment that it's interesting to hear your perspective because you've you've been around this for longer but I'd say it in the average person and you've seen the onset of it I mean I was in an AI company before the hype too I was I worked for a race car championship in a machine vision department essentially doing doing the same kind of thing before before all of this noise began and it was and I studied it in some senses in the 90s as well doing machine vision and AI and some very early forms of it very immature versions of it so it's been really interesting to see it finally 30 years later or whatever getting this much attention and I didn't I didn't think we would be in this position at this point in time I didn't think that it was especially that it was going to be this this powerful and this big and this talked about going back then it was like a few hippies in a room with sandals and in PhDs and big ideas and not enough to compete yeah that's incredible to see right come yeah more than well not everyone but there's the resources are being directed at this with so much velocity and just scale that the iterations are coming so quickly whatever was the benchmark you know a month ago is already probably being surpassed now and it's it's the I feel like this time is the best time because this is the real sort of what they call the application layer right in the innovation sort of the trajectory of innovation so it's always like rapid rapid rapid rapid rapid rapid rapid rapid rapid and then it tapers off for quite a long time before these novel innovations and foundational technologies become widespread adopted and fully educated a lot of times like the internet or something at that turned almost into utilities which for some AIs like AIs searches that's it I mean it will just be part of the search engine you know experience from now on I don't think any anyone just anyone uses straight google any longer it just it's it's such a very simple and invisible UI that that prompts you to ask another question and refine your own question that was you know that was the big thing a year ago was good in I mean good in good out is going to be the premise forever but you know the big focus was on good in and good out your outputs only going to be as good as your query or your your question which you know kind of goes to my point about that where I see the emergence of technical or professional skills being writing being among them it's the same premise yeah yeah lost quality the quality the quality of the idea I think is what you're speaking about through that like the input being the idea and if it's a bad idea then AI doesn't help you make it exactly maybe it can help you like unstick block creative blocks I think but it doesn't get you a great idea on its own and it's really the all of the things you're describing the inputs the text the speech the in the broader sense yeah inside perspective you know like even so there's a book it's really thin it's called it's like a management book it's called outcomes over output and I think when I met you the first time I said you know AI is not going to make someone with no one to say have something to say like it that's not what's going to happen what will happen I think inevitably and it's or it doesn't actually didn't need AI to happen it was happening based on the ability of data analysis about streamer behavioral and preferences and choices and how long you watched a thing for and now we have 47 versions of the Christmas prints on Netflix to have to scroll through because that's you know what the data was saying is popular during Christmas where that kind of bloat and glut is it's inevitable what will be interesting is not about how people use AI to make what we've already seen but what you know the kind of emergent genres that we can't even really put definition around yet because we haven't seen it yet and that's exciting from the the art side from a business standpoint the intelligence again you can have all of the intelligence just like a CEO with with good or bad advisors if they don't have the experience themselves they'll they'll they're gonna suffer from that and they may or may not have the runway to continually learn from their mistakes yeah that's that's an amazing business insight there for anyone that hasn't experienced it like their cash flow and the the realities of what you of your brilliant idea going out in the world and actually sticking and getting people's interest enough for them to pay for it and continue to pay for it you adapt and pivot to all the changing world that we live in which is seemingly speeding up there's a question that does come in here from the audience someone's asking whether there is a non-line demo that you can point them towards they've said they recently worked a flawless AI and have a amazing tool set for dubbing they're just interested if you have something they can try yeah sure you can definitely reach out we don't have a how to use it demo because we generally license our software for like B2B there's uh enough I think to infer what it does from the videos that are on the website but if you're interested in chatting a little bit more about how it works very just email me happy to talk to you or or put you in touch with whatever sales and or technical team to share more cool that's great well ask Sarah and she will point you in the right direction the really interesting what you were talking about the feedback loop of the audience and the audience is a umbrella term I guess for anything whether it's a piece of software or a piece of content or a movie or whatever it is that people are consuming the the necessary connectivity between them and the feedback between them I think is it's hugely important and something which I think the executives may be in some of the companies that are responsible for distributing content I'm not always thinking about but we're seeing a seemingly seeing a trend towards more creator economy or even in films for people coming from a place where but they've already got fans, like people have already got fans on YouTube, therefore they get. higher because there are no entity and they're proven and speaking of the creator space, so is there any way for them to know your mentioning B2B, but there are only ways for them to use your products or try benefit from what you're doing. So we work with like the full sort of spectrum of user types. Our main business model is definitely for studios. But we work with independence. We do generally run at least one. We're starting a new one in the fall of student cohort, cohort so that they can start getting comfortable with the solution. We are looking at, because Blender is definitely support for Blender is coming from us and so that community definitely is a different makeup than the standard sort of Maya. It'll be a different way of licensing, I think, and that's exciting. So that's sort of the real sort of like widespread sort of easier to download and easier accessible for short is part of the future growth plans and scale plans for the company. But for now, we mostly work with individuals that we are either really kind of jazzed about their work ourselves and want to support them and then see what they do. So at ASE there's an independent filmmaker who's released his short film and he's doing a talk about how we basically use our software from the very beginning from like earliest stages of script development and character development and how the value of having a your the 3D asset that you're going to be using in your even if it's not you know finally finished but like animatable at the very beginning like your animatics become not just a stage that you then have to recreate and throw away like it's all sort of persistent and iterative towards final render and allow him to overcome quite a number of sort of typical independent budgeting and production challenges as well as opened up a lot of interesting experimentation for how he could use that 3D backbone across a number of different like with like I said in 2D platforms and you saw some of the sort of AI treatments with our friends at IEL with using our software as the backbone and then Luma on top. Yeah yeah that was super cool definitely everybody should check out IEL as well. We've got another question here from me. I can't hear you. Anything that's able to be talked about. Do you mind me? Yeah just do you mind repeating the question just you literally started to. Oh okay maybe the streams broke up because the gold live broadcast you hear me okay now. I can. Okay so a question from the audience here. Somebody asked I know how touched on the future feature film direction but I'd love to hear anything that's able to be talked about about how asked 2D guys might see in one or two years. Definitely if you are paying attention to anything that we're doing at Anisee that's a big part of what we're showcasing as soon as I have assets that the artist is comfortable with that will go in the presentation that will be I'll have a bit better sense of the timeline for pushing those out either before or after the event. Okay great. Yeah. Well definitely to share that with us and we will share it with the community as well when you when you have it. It's been exciting. Is there anything else that you? Is there anything else that you'd like to share with the community any places that people can find you or anything else you'd like to share? Yeah so in I talked about sort of advancing into sort of more upstream of production so that's a huge project. Blender's a big project. These experiments in 2D really opened up a design sort of opportunity that we didn't know was that obvious so that is something that we are hoping to accelerate as soon as possible because it's something that's missing I think in the space in the available tool sets and so we if someone is interested in talking to me about the 2D stuff I really am interested in talking to you. In August of this year we will be pushing out like the third sort of significant version release of the software which will address some of the biggest challenges I think we've come to understand people are having in the issue so exporting assets and maintaining integrity you know between platforms the ability to handle you know really different styles different from you know either humans or and then creatures that kind of thing as long as they're talking and in conversation some of the models that were animation models that again not AI models but algorithmic models that we are introducing will move the performance away from monologue to dialogue which is exciting. In the fall is the student cohort again it's a limited enrollment but the idea is a focus number of access to the software project some sessions in person and and lives so those ones will be open and to work on a particular project. All of this is going to be on our website. Yeah so that will registration will start in the summer or not in the summer sorry in the autumn and more information is going up about that in the next couple of weeks there's some you know placeholder boiler but there's also if you go on our website there's a there's the whole cohort is based on the success of a number of other webinars and programs that we've done in the past so you'll get a sense of the kind of programming that we have done it's the ethos is generally industry tool introduction and familiarity and training but the within necessary combination or balance with foundational animation and and rigging teaching sorry education lost the word yeah I think that's that's a lot that we've got going on we've got a couple of fun case studies if you're going to unreal fast we'll have somebody there if you're going to be at games calm we will also be there I said anisee like a bunch of times will be at XDS we're in Toronto you go if you live in Toronto yeah yeah very cool well jaleighresearch.com for anybody who wants to go check that out and yeah thanks for sharing all of the places that people can find you is there any others than any things you guys do on social media places people can follow you we basically wherever you most people are so we're linked in youtube twitter x sorry except I can't say thank you I can't say x I think we have a blue sky account I'm not sure I know we're on instagram cool good covered all the bases yeah thanks so much for having me that was fun you bet thanks for being on thanks so much for taking the time I really appreciate it and yeah check out for everybody who wants to check out what they're doing a jaleighresearch.com thanks to our audience for coming and asking great questions thanks to global objects for hosting me today to the place to sit do this problem that's not in my studio if you want to come and see what we're doing at CG Pro we have an AI filmmakers webinar coming up on June 25th at noon so sign up for that and we'll be going through some free educational experiences and come check out what we do at CG Pro thanks again to everybody for today we'll see you all soon take care bye for now
Podcast Summary
Key Points:
Jolly Research, co-founded by Sarah Wattling, specializes in automatic facial animation generation from audio, text, and metadata inputs, focusing on lip sync for games and cinematics.
The company uses a procedural, algorithmic approach rather than heavy machine learning, producing editable animation curves that are clean, lightweight, and fast to process.
Major projects include Cyberpunk (2 million dialogue lines) and collaborations with Obsidian, emphasizing scalability and speed (processing 300,000 lines in days, not weeks).
The technology supports stylized characters and creatures, not just humans, with custom profiles for unique sounds (e.g., dragon roars).
Jolly has navigated industry shifts (COVID, strikes, AI hype) by staying lean and focused, avoiding bloat, and prioritizing user editability and polish-ready output.
The company is expanding into cinematics and interactive characters (e.g., Ruby AI), and faces competition from internal studio tools and emerging AI models.
Summary:
Sarah Wattling, co-founder and CEO of Jolly Research, discusses her unconventional path from live music production to tech leadership, leading to the creation of a facial animation tool that generates full facial expressions from audio, text, and metadata. The company, now nine years old, uses a procedural, algorithmic methodology that contrasts with trendy machine learning approaches, producing clean, editable animation curves that reduce cleanup time and computational overhead. This design has proven advantageous for large-scale projects like Cyberpunk and Obsidian’s games, where processing hundreds of thousands of lines quickly is critical.
Jolly’s focus extends beyond humans to stylized characters and creatures, with customizable profiles for unique sounds. Throughout its history, the company has faced pressures from industry crises (COVID, strikes) and AI hype, but stayed committed to its principles of small, lean, and focused development. The output is designed to be close to final polish, minimizing the need for manual curve cleaning, which is a significant cost saver.
Jolly is now expanding into cinematics and interactive characters, while constantly optimizing for speed, file size, and user influence. Sarah emphasizes the importance of balancing innovation with practical, repeatable solutions, and navigating the noise of emerging technologies to deliver value to studios seeking efficiency without sacrificing creative control.
FAQs
It automatically generates full facial animation from audio, text, or a combination of audio, text, and metadata inputs, primarily for games and cinematics.
The main focus is on using metadata from scripts as input, along with constant optimization of the software.
Jolly produces clean, editable animation curves rather than images, which keeps the output lightweight and easy to edit at a granular level.
It offers deterministic quality, lightweight computational footprint, and fast processing speeds, enabling large-scale projects to be processed multiple times a day.
Yes, it can handle stylized characters and creatures like dragons, with custom corrections that can be written into a character's profile.
It reduces the time from capture to final render by generating animation closer to a polish state, minimizing the need for cleaning motion capture data.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.