Go back

Building Powerful AI Image and Video Workflows for Marketers

42m 36s

Building Powerful AI Image and Video Workflows for Marketers

In this podcast episode, host Michael Steelsner interviews AI educator Jared Liu about building powerful AI image and video workflows for marketers and creators. Jared shares his journey from early content creation to becoming an AI educator partnered with major companies like Google and Adobe. He emphasizes that a key misconception is expecting one-click Hollywood-quality results; AI is a tool that requires human direction, storytelling, and brand understanding. The conversation covers the latest 2026 updates, including Google's OmniFlash model, which integrates video generation and editing across Gemini and Workspace, and the refreshed Google Veo tool with project-based workflows and agent layers. Jared compares top tools: Sora 2.0 by Bytedance leads in video realism and prompt understanding, while ChatGPT Images 2.0 excels in speed and text accuracy for thumbnails and graphics. Google's Imagen 2 remains strong for editing. The hybrid approach—combining AI with traditional methods like brand guidelines, storyboarding, and human oversight—is critical for consistency and quality. AI empowers anyone to tell stories visually, but meaningful results come from iterative creative work and a clear vision, not just pressing generate.

Transcription

9109 Words, 48076 Characters

English
Hey, it's Mike Steelsner. What if you had an AI system that wrote your ad copy, generated the images and video, reviewed its own work and made revisions, all customized to your brand without you ever touching it? That's exactly what Caleb Cruz is teaching AI business society members to build on July 16th. No coding required. Join today at socialmediaexameter.com/ai. And if you're listening to this in the future and you missed it, the full recording and resource kit are waiting for you. Visit socialmediaexameter.com/ai. Welcome to the AI Explored Podcast, helping you put AI to work. And now, here's your host, Michael Steelsner. Hello, hello, hello. Thank you so much for joining me for the AI Explored Podcast, brought to you by socialmediaexameter. I'm your host, Michael Steelsner, and this is the podcast for marketers, creators, and business owners who want to put AI to work. Today, we'll explore building powerful AI image and video workflows. My special guest is an AI educator who helps marketers practically apply creative AI to their work. His YouTube channel specializes in AI creative with images and video. He also trains marketing teams across the globe. Jared Liu, welcome to the show. How you doing today? Good. Thank you so much. Such an honor to be here, man. And like, can't wait to talk about this topic. I'm really excited to talk about this. Also, let's start with your journey. How in the world did you get into AI? I've been making content online for very long times since like 2006 or 2007 when it was me and a guitar on camera trying to be Justin Bieber and not being successful in that. And then it was a passion of mine to tell stories online, whether it was with music or later in tech gaming. And so when Jenny, I started to hit the scene really at the in 2023, it really just clicked for me that this is something that could unlock a lot of creative potential. And it started just on the side of my usual nine to five job making content about these tools, what they can do, things I was discovering. And then suddenly an audience started to grow and grow and grow. And now I've been fortunate enough to be partnered with some pretty big companies like Google, like Adobe, like Runway to actively test new tools that they're building. And then also be part of the team that helps to explain what these tools are to different audiences. So I'm playing many different roles at the moment, but it's a lot of fun. But the core is I just love to create content. Love it. Let me ask you this question. When it comes to AI art and when I want to say AI art, I mean like images and video, you've been creating content for a long time. You've been covering all these tools that are coming out. What do you see as some of the biggest misconceptions that a lot of people might have when it comes to using AI to create this kind of creative? I hear this a lot and it's also understandable to hear, but I think the biggest misconception is always that you press one button and you generate the most incredible thing that could take on Hollywood that could take on some of the our favorite videos and content that we've seen online. But it really just isn't the case. And I think AI like many other tools that we use are a tool. I see it as an extension of, you know, the video editing tools I use like Premiere or After Effects. It's another tool in the tool else for me to tell a story. But without me actually knowing the direction and the story I want to tell, it just sits there on its own. And you know, I think it's great that anyone now can hit generate and create incredible things for themselves. But to actually make something meaningful for other people it still brings a lot of that human element to it. And I think there's a bit of misconception that anyone can just generate Hollywood movies. We really can't do that yet. We still need to train our creative mind and creative vision to do that. A lot of it is really I think the companies that are behind these tools, right? Or some of the really good creators like when Seedance came out and I saw these videos that just blew my mind, I'm sure these were created by people that are that used to work in Hollywood as my guest, right? And they knew exactly what they were doing. And they probably spent a lot of hours to create these things and these companies make it sound as if this is what it does out of the box. But that's not always true. Is that true? Absolutely. I look a lot of my peers in this space who are creatives or influencers while content creators or educators. A lot of them have massive film backgrounds, whether it's an advertising, marketing, or literally feature film stuff that they've done in the past. We're traditional methods and now they've adapted these new tools. But they still have teams of people working with them. The difference is now they're able to work maybe a little bit faster. Maybe they're able to, maybe they're able to ideate a little bit more. Maybe they're able to do more with the footage that they've captured with real cameras thanks to AI tools. So I think a lot of the companies they sell, we see the launch videos and the trailers and they look incredible. But to do those things takes a few steps sometimes. And I'm trying to play the role of educating people that you can do these things. But there are a few steps you can do before that to actually get that out of it. Love it. And you know, let's talk a little bit about, like if people pay close attention to what we're about to talk about today and they produce really high quality AI images and video, what are the benefits? What are the upsides that are waiting for them if they if they take this stuff seriously? I think if you've ever had a story in mind or you ever wanted to create something, this is the perfect time to be able to go and do that. And I come from a music background, I write music. I never thought I'd be able to do visual art in video, but I still love writing stories, whether it's in the past, with a book writing to now, all these Google docs and notes I have of stories I want to try and create. I now have the ability to go and do that. And I think that's a really empowering element that AI can bring to the world in general. I think we need more stories to be told and more experiences to be shared. And I think this really breaks down a lot of barriers that people may have had in the past to go and creating that. I think it's wonderful that anyone can now sit down or with a computer with a laptop with their phone and create visuals that match exactly what they're trying to tell. Obviously, let's talk about any foundation stuff that we need to consider because plenty of the people listening right now desire to create really high quality visuals, but they maybe don't have any that formal background that a lot of the other people in the industry have. So is there any kind of foundation level mindset shifts that we need to make a little bit here before we get into actually making really high quality images and video? Yeah, I think a lot of this is still on traditional methods. So if you put AI aside for a second and you if you work in a marketing team or you go on your company or you have a brand, you still need to sit down and work out what that brand means, what the audience is, you know, what color palettes are in your brand guidelines fonts, logos, before even touching the AI step. And you know, AI can play a role in helping you come up with these ideas, but you need to sit down yourself and still figure those things out because if we don't at the core know what you're trying to represent for your brand or your company or for your team, there's no way an AI tool will know what that is and there's no way other people and audiences will know what that is either. So I always say the hybrid approach is still critical, but the human element is still very critical at the start of this whole period. You have to know what you're trying to accomplish. And that I think is in general for any tool that you use creatively, you know, if I open Photoshop for Premiere and I sit there looking at that black screen, if I have no idea what to do or no idea what I'm trying to create, the tool just isn't going to make stuff. I still need to direct that vision. So for me, the first stage is still understanding what you are trying to do and how you would like to convey it, then bring in tools like AI tools to help you do that. I recently interviewed one of the prior guests on this show and he was talking about cloud design and I don't know if you messed with cloud design at all, but he was saying how you could go to your website. Let's say you work for a company, but you don't understand or have access or there is no such thing as like a design book or brand book or whatever. You could hypothetically take some screenshots of your website and maybe upload your logo and use cloud design to kind of create a brand book almost. Have you ever done anything like this or heard anything about this? Yeah, absolutely. Look, there's a client I work with recently. It was for a fashion brand and they wanted to test using AI tools for the first time and they didn't have a succinct brand guide. It was all bits and pieces of it. They had a logo. They had some ideas of the color palettes that they'd like. They obviously have photos of the products, but nothing that's concrete and actually cloud design was really helpful for this because cloud design works on design systems. So it will basically generate your brand guidelines for you, your style guide with whatever you have available. And this was extremely helpful as a first stage to say, look, we can build AI stuff and images visuals, but without a consistent guide or the need, it's going to come out very random. So I do like what is on a lot because of that approach. It really helps to let consistency across new things you make with it too. So highly recommend. Excellent. Okay. Well, let's get into some other tools here because there's a lot of options that are out there. And we're recording this at the end of May, 2026 and Google just had their event, which I believe you were at. I forget what it's called, but there's been some updates that we're going to talk about here along with some other tools as well. So let's just kind of go through the state of the options that are available to all of us today and maybe talk briefly about which one of these things might be best for certain kinds of things. Yeah. So Google I/O was last week. It was, is there any conference they announced all the latest in the past? It was a lot of our product and the workspace and now it's fully AI, agentic approaches now. This is a very exciting time for image and video because for the last couple of years, we focused on generation quality, make amazing art and amazing videos. This year, the first is a switch to more control, modifying, editing, adding stuff to scenes. So getting more out of the footage that you've generated at Google I/O, Google announce updates to the Google Flow tool, which is traditionally where you generate images and video. Now you can do a lot more. They announced this new Omni Flash model and how that stands out is you're able to prompt it with videos to make adjustments, whether you're trying to remove something small from the video that you've added or you're trying to completely change the style of the video as well. You can now start doing that all within that one tool. And it's really is following the industry trend, where every new video tool that comes out needs to have the ability to make those edits and modifications in now. And we've seen that with CNATS 2.0. We've seen that with Kling 3.0 as well. It's become a new standard and it's now become this new race with all the big competitors that are who can create the best tool that bounces high quality footage with amazing editing abilities. Yeah, I want to dig in on this a little bit. So explain to everybody what Omni Flash is exactly and then what Flow is. I've talked about Flow on the show, but it's been a little while just to help everybody connect the dots because some of my audiences just using Gemini by Google to create images and very simple videos. Kind of explain what Flow does and what this Omni Flash thing is exactly as a new model. Yeah, so it's a new model from Google and it actually works across Gemini. It works across the entire workspace. So now if you use Gemini by itself, you can actually generate videos and edit videos directly on Gemini because of this new model. It used to just sit within Flow by itself. But now you can actually access it from anywhere. So sure there's the whole image and video part of it, but it's actually a model that is smarter. They can understand different types of inputs. So whether I send it a script, send it a prompt, send it a video, it can now receive those things. But we're going to start to see this pop up across the entire Google workspace. And this is what they were teasing at us at Google I/O is like, for now you can see the image and video capabilities. But imagine having this directly in Gemini. Imagine having this as part of Google slides, Google docs, a model that can receive and understand information that isn't just text and modify it. This is the new direction that they're going to add. And I think it's quite smart because it matches what the industry is trying to do, but also gives the general consumer way more ability to actually do more, but they're trying to create on the go. So this Omni Flash model is it part of Gemini 3.5 or is it a separate thing we have to select or is it just kind of randomly integrated and it just works? And then also explain flow. Yeah, I'll start with flow because I'm going to keep on the video element. But flow is basically Google's creative design tool. In the past you were just able to do videos and images. You could extend those videos a little bit. You could edit a little bit like you were using Premiere Pro a little bit on there. But now they've completely refreshed it with a ton of different features. Now the tool works project by project. So if I'm working on a specific marketing content or video for client, I can have a project that I can share which has all the designs in it, all the videos I generate for them. I can also train characters now. I can also create brand guidelines all in there. And the thing that they've updated the most on top of this is this new like agent layers. So you can now work in a conversation style mode within flow to create images and video in bulk or to edit different things. It's more like working with a creative director more than me just telling a tool. Do this, do this, do this. There's an actually an agent now that responds to you as able to understand what the project is and what you're working on. So flow right now has become this central spot of creativity for Google. I think they're cyphing everything into that one tool now. And these updates were very much needed. Flow was feeling a little bit out of date for a while. But now it's become pretty solid tool in my tool now. We'll see. Okay. Nano banana, which is something that we've been talking about on this show for a long time. And I think nano banana two is their latest model or is there a new version of it? This is where people probably get a little confused. So what here's what we know so far. We know flow is a standalone tool in the Google ecosystem. And we also know that nano banana is our image generator. And I believe the latest version was 2.0 if I'm not mistaken. And with the new versions of their LLM as nano banana improved or not necessarily changed at all. It's still the same nano banana two. Okay. I'm very sure we will see more updates this year for it because there's other competing tools that are appearing now to banana to itself is very, very strong. It's already a fantastic image created, but also an image editing tool. So when you look at Gemini and Omni Flash, I think the way to see is like this is the nano banana for video. Ah, I see. Okay. Got it. Okay. So we've covered the Google ecosystem with nano banana and flow and Gemini Omni Flash, but you briefly mentioned some other tools. So let's talk about your stance and where these other tools kind of fit in the quality of the ecosystem. For me, top of the ladder right now is could change tomorrow, but right now it is a so cdans 2.0 is my current favorite video model. It's an incredible model because of the realism and the high quality outputs it has. But the way that it can understand prompts the way understands images, the way it understands videos is very unique compared to other models out there. And it's by bite dance, right? The parent company bite dance, right? Yep. By bite dance. It's a model that has had a little bit of controversies this year, particular likeness is being used when they shouldn't have been, but the model itself, it's become my go to for really showing the power of AI video. It has audio, it has dialogue, and the sound effects background music. It's an incredible model. It's one of those moments where I prompted it for the first time I was wow, it actually took every word I put in for this prompt individually and created something that really matched it. Now it's getting even stronger because you're able to prompt with images and prompt a video, even prompt with music now too. So there's many ways you can approach it, which is what makes it for me the strongest model right now. Okay, what about on the images frontier? So they have c-dans and they have c-dream for images. Also a very strong model, but the one that stands out for me is it's a fight between Nana Banana 2 and Chad Gbt images 2.0, both are very good models. It's pretty shocking because I mean Google has for the last couple of months kind of I felt like owned the image world and really the video world until c-dans came out. But when Chad Gbt images 2.0 came out, I wasn't really paying attention to it because I've already moved on mostly to Gemini and mostly to Claude, but images 2.0 is extremely impressive. Do you want to just talk about what makes it kind of maybe slightly even better in some regards than Nana Banana? Yeah, I think it takes all the boxes of first of all you can create really realistic looking stuff at this. That's a standard now across all these models. What makes us stand out for me is it's ability to do to do text and you are really able to make long pieces of writing with coherent text and this is good if you're trying to develop things like storyboards, if you're trying to create character sheets, animation sheets that require graphics, that require a consistent character that require text. Chad Gbt images 2 really blew my mind with how good this is and it's extremely fast as well, very, very fast and it took me by surprise too because I think like he said, I was already on Gemini and Claude and then I was okay, I'll check out this image thing. I was okay, this is becoming a standard go-to tool for me now and it's generating for me every day thumbnails for my videos, for my newsletter, any kind of post I'm doing around an announcement or an event, it's able to use my likeness better than any other model that I've tried. And for me, even though I'm still using that about to and Chad Gbt images, I think amazing images too a little bit more now. You know what most marketers are doing manually ad creative, writing the copy, coordinating with the designer for the images, waiting on video edits, managing revisions, every campaign, every cycle. I wanted to fix that for our members. So on July 16th, I'm bringing Caleb Cruz to teach inside the AI business society. He's going to show members how to build personalized AI pipelines using Claude code that runs the entire creative process on his own, copy images, video revisions, all customized to your brand and you don't need to be a coder to do it. I'll be in the room for the entire session, I'm in every session and I can tell you, members are going to walk away with a clear, buildable system that they can put to work that very same week. If you can't make it live, the full recording and resource kit will be waiting for you, transcript slides, key takeaways, the whole shabank. This is what we do every month inside the AI business society. Come and see for yourself and don't miss this training. Join at socialmediaexaminer.com/ai. Again, socialmediaexaminer.com/ai. Okay, so we've done a high level review of the tools as far as the most popular tools to generate images and to generate video. Now let's talk about how can we best take advantage of all of this? What's cool is there's a way that you can potentially have your cake and eat it too, right? You don't have to just work with one of these tools with what we're about to talk about next. You can choose from amongst the cacophony of tools. So how do we actually let's talk about how we can best take advantage of all these great things that are available to us and likely are going to change. Like today it's this tomorrow could be that, right? - Yeah, and look, I get a question a lot from my clients and it's like, how do I keep up? How do I keep, how do I know I'm using the best tool at all times? And my response is always try not to spend your money on one specific tool, try to invest it in a platform that brings these tools together. We're in a very fortunate time where a lot of these tools have APIs open, so there are platforms out there that allow you to use under one cost, a dozen image tools, a dozen video tools, and the one I've used the most is Magnific, they're once known as FreePick, they just change the Magnific and I've stuck with the platform for over two years now because they always get the latest models in there as soon as that API is available. And I'm able to test, play and try without the stress of, okay, I sunk 12 months of subscription into this model and now it seems to be completely falling slowly compared to these other ones. I prefer that people go towards platforms because the right it moves fast. And Magnific is the one that really stands out for me because of the amount of tools there, but also the tools that you get with it too, it's not just press button generate, you can build workflows now, you can build ways that you can test the different tools at once without having to keep switching. So this is the direction I try to push people towards, 'cause I think it just saves them a lot of money and allows them to keep ahead a little bit. - Yeah, let's talk about this tool, let's get into it in depth. I'm assuming it's a token based tool, right? Where you have to buy a set number of tokens. But let's give an example, you mentioned thumbnails. So why don't you walk through how you do what you do with this particular tool and then we can maybe dive into some of the other cool things that can do that, might not be obvious things that you might be able to do with Gemini, but you might not know you can do it, but this has got some of this stuff built into it. - Yeah, look for sure. So with Magnific, they have a tool called Spaces, which is a node-based flow. So if you need to that, it's a large open canvas and you create an image node, which generates images, you create a text node, which will use for prompts and you can create nodes for audio and video. When it comes to thumbnail creation, what I did was create a workflow where I uploaded a couple images of myself. I used Nado Banana 2 to get different angles of my face from those images. And then I brought that into another tool like GPC Images 2.0, which is able to combine all that together to create a thumbnail. I prompt it with, hey, I need a YouTube thumbnail, it needs to be in this dimension, it needs to feature my face and some text. But the cool thing here is you can also ideate with it. So I can run several different prompts at once to get 10, 20, 30 different styles of thumbnail at once. - Okay, that reminds me of, I'm forgetting the name of that tool that everybody used to use back in the day where it generates four at a time. I'm drawing a mental Blake, but with these other models, you're typically, if you're just using Gemini, you're doing one at a time. And the benefit of this thing is having a bunch at a time. I would imagine it allows you to work much more effectively, and quicker. - Yeah, because there's still, is despite all the control we have, there is still some randomness to AI, right? And you need to learn about what it's causing, that randomness, and also how to cut that out. And for the first iteration of this workflow, I just found that it wasn't using my face well enough. It wasn't using enough. And I realized is that I needed more images of me in different expressions. And I would go back and use that same workflow to generate more images of myself, and use that as part of that flow. So we got better and better and better. The best part about bulk creation is that you can test different things at once. And out of the 30, I found five that I thought fit my brand the most. And that's what I dialed in on that workflow, saying these are the ones. This is the reference for all thumbnails going forward to use. And this could be upscaled completely to like brand and product content. I've created workflows for clients on their products, because they didn't know what the power of AI tools. And we're curious about it. And I said, well, I can build your work for where you just upload your products. Many images as you want. And we can run it through NADABAD to org.cat.gbt images2.0, or any other image tool and generate new content for you based on that. And you can see which one works for you there. And having that under one workflow, under one platform, on one screen, is fantastic, because it really shows people what these tools can do, and also allows them to jump in themselves and play around and try and test. One question for you on a workflow, something as simple as thumbnails. Does it have any kind of memory? Because this is one of the things that a lot of people are concerned about when they're using third-party tools. Does it effectively learn your things about you and is that stored in some sort of a, for lack of a word, knowledge base, where it goes and it looks at this and it says, OK, these are the images that we should model. So you don't have to reinvent the wheel every time. Yeah, so this is where we're actually heading to the CRO lot. So we're seeing more and more agents being added to these canvases, all these workflows. And that agent becomes the memory bank. So if anyone is there's not me even, is to come and see this and wants to know about it, the agent will have that knowledge of what's happened in the past and know what to create next. I think that's why Claw Design was also so popular, because you're building a design system that's always pre-contexting the model before you hit Generate. It's something I hope to see in Magnific Soon, there are other platforms out there, such as Luma, and Runway, that are launching their own versions with an agent at the core. It's a direction we're going to, because you're absolutely right. The memory bit is critical. And I think that was a missing element for a long time. And I looked back at the prompts and the workflows I had a year ago, and having to redo it every single time from the start with a fresh prompt with fresh images, it took a lot of time. This is being cut down a lot now. Let's talk about some of the audio features that are built into Magnific, if you don't mind, because I think this might be interesting to people too. Yeah, so since Magnific pools in different third party tools with API, they have a bunch of really cool audio and one, like 11 labs is in there, which is great for generating music, sound effects, and also voiceovers in narration too. It's also good for creating character voices, as well if you want a consistent voice. They have a lot of new tools in there, which allow you to do better lip sync in the platform too. So if you have an image of a character and you have a voice, you can use different video tools in there to bring it together. And it will actually lip sync it quite well. And this is great if you need that consistency of a character. Like sure, they can look the same, but if they sound different in every shot and really ruins the immersion, this helps to stop that from happening. Well, I was just curious about the speed and the cost, because obviously, if anybody's messing around with chat GPT or Gemini, they know it's pretty slow when it comes to creating content, is it similarly slow? Is that one of those things where you just set it motion and you come back to it and the alerts you when it's done? And what about the cost, you know? If it's a big flow, like are you generating 30 to 40 images and 50 videos, it can take a little while to do. But I think being able to walk away is also a good thing as well, because you know, things are working as you're doing something else, hopefully more productive in the day. But what I like about the platforms as well is that, and this is what I don't think Gemini has some time to, it doesn't tell you how long the generation is estimated to take. And it could just keep going. On these tools, they can see, you can see this could probably take one or two minutes or it could take five minutes. So you have a little bit more access to the data as to how long things will take. When it comes to cost, a lot of these platforms, like Manifacto, you subscription models now. It ranges from $10 a month up to $100 if you want those max business premium plans. And I think the best platforms are the ones that give you a lot more access to tools faster. So with Manifacto, I always have access to a free image model at any time, which is great. I think it's important to compare those across all the different ones out there to make sure you're getting the best deal. Because otherwise, you're spending money and you're not getting the latest and the strongest tools at once. So on my run, it is always just go for one month at a time. And if you like it, go for another. If you really like the platform, go for the longest subscriptions, but one month at a time, always. Up to this point, we've talked about Magnific. And the benefits of Magnific is that it allows you to work with all these prior tools that we've talked about, Chat GPD images 2.0, and a banana, C-dance, and all these other options. And what it brings to the table is it brings workflows to the table. It brings-- you didn't mention this, but upscaling, right? So you can take an image and you can upscale it. They have upscalers built in, music, audio, lip syncing, all that kind of fun stuff. Now, let's say somebody wants to create some content, let's say it's video content, because why not? And they want to place themselves in that video, or they want to place their products into those videos, whether their images or video. How do we go about doing that? Because I think this is really the big unlock where a lot of people want to create an ad. And they want to create an image. And they've got maybe a photograph of their product, but they don't know how to get that into a beautiful image. And then ultimately move that image into videos. So talk to me a little bit about all that fun stuff. This is the best time to do this. I looked back like two, three years ago, and this was the hardest thing to unlock ever. The one thing everyone's asking is, how do I get my product in? How do I get myself in there? How do I get it consistent? We're in a really wonderful time now where Natah Banana 2, where chat, GVT images, 2, all are really good at this. And it can be literally as simple as dragging your product, like dragging your image of your product into that prompt box and saying, Hey, can you please. create three angles of this product for me to use. It's especially we did content. - Oh, okay. So it will actually rasterize it for like better words and figure out what it would look like from a different angle. I mean, is that effectively what it's doing behind the scenes? - Trying to do that. - Wow. - Yeah, behind the, that's what it's trying to do. Obviously, you know, if you give it one image from the front, it doesn't know what's happening on the back. So it will try its best to do that. - Right. - But I feel like if you're holding a product, if you take four images around the product and upload that to LGBT and say, "Hey, I need to look like a, "please render this as a professional." - Top down shot, whatever angle. Oh, that's cool. So basically take lots of angles of your, and it doesn't need to be a professional photographer. Can it just be something you can do on your iPhone? - Thank you for being on your iPhone. This is, I think the biggest thing that I think people would have realized is it doesn't take the highest quality image to do this now. Like the video is that I have of me in it is literally an iPhone selfie that I upload to Chad G.P.T. on Out of the Matter. And I say, "Hey, can you please make some "professional looking head shots or movie shots?" - I don't know. - And it's able to do a product, it does a really good job at it now. And especially for products, it's really, really a great time to start doing this. And it's as easy literally as giving it a picture and saying, "Hey, please make this look professional." - Okay, real quick, I want to get in on this a little bit 'cause I think this can be really helpful for people. So far, take your product and take a couple shots of different angles of the product. We upload it into one of these tools and what are we asking it to do exactly? Are we asking it to just strip out the background and make a super high resolution version of this or what exactly are we doing? And then let's also go through the same thing with pictures of ourselves because I feel like this will be really foundational for people. - The best way I recommend starting if it's a product image, a blow whatever image you have of it, it was just one or multiple. And I always say first, can you please create a product sheet for my product? Use the attached images, create a sheet that contains multiple angles. And then after that, put in as much information as possible about the product. What it is, who is it for, if it has a name, or we not tell it to come up with some names. And what it will do, especially with ChadGPT images, is it will create a one image shot of your product across multiple different use cases at once. And that's a great place to start because you can basically see that as like a brand guide, a style guide. And everything that you chat with afterwards will use that as a premise. It's a lot of chatting back and forth, but that's the best way to start to start to see as many images at once. - Okay, now this is really intriguing, right? Because to the best of my knowledge, ChadGPT images 2.0 is only 2K image resolution. So obviously these images aren't going to be super high resolution, but it sounds like it's enough for when you drag it into another model for it to discern. What it needs, is that what I'm here and you say? - Yeah, it's absolutely enough. Eventually you can upscale it to 4K, you can resize it. But these tools are also good at resizing stuff too. So if you have everything as a 16 by 9, but you want to put it on social media and you need it in that vertical, you can ask images 2.0.0.0.0.2.2. Hey, can you just reformat this into this specific aspect ratio, and it's really good at doing that now. It doesn't cut, it actually redesigns it. So it's a great time for this. - Let's talk about pictures of ourselves or another person. Let's say I work for a marketing team and I want to get my CEO in there or I want to get myself in there. What kind of shots should we take in order to be able to accomplish whatever that output is? We've already talked about a product sheet for a product. What about for a human? - Yeah, so for people as many high quality shots as possible is always recommended because these tools are great at making realistic faces of people, but when it comes to real people like us, and there's a lot of nuance that we don't realize about our expressions and how we talk and how we look that we don't really realize. And if we don't know that, the tool doesn't know that. I recommend getting as many shots as possible of the person's face from the side, even from their back as well, and giving it as a reference to a tool. You can of course ask the tools to generate that for you, but it will likely not get it right because it doesn't know who you are and doesn't know, it doesn't have a camera to take it itself, but it will try. So more images you can give the better and we're in a great time where you can shoot 1080p 4K on your phones now. So go ahead and take some selfies of yourself while all ask our friends to come in and take photos of you. And that will give a much stronger result on the other side, on the image and the video front. - Two. - What about expressions? What kind of expression should we make or should we just be plain faced or because obviously, I would imagine there's some nuance here, right? Because like when you smile and I smile, it might look different than when the AI decides to add a smile on our face, right? - The basic one is you always want some shots that include your teeth in it because that's how the AI recognizes that is their mouth. - Okay. - And the different expressions from sad, determined, shock, scared, curious. The more that you can do the better here. And you know, this was the issue I said earlier when I was making, but I'm doing my thumbnails. I didn't have enough emotion images of me. So it would maybe look ridiculous 'cause it was trying to stretch my face to look shocked. When I could have just given it an image of me looking shocked and it will know, okay, that's exactly how we're gonna use it. - So what do we do with all these images once we upload them? What are we asking it to do to create another kind of style guide for the human? One of my favorite pieces of content that this year was creating character sheets, actually. - Okay. - So you can create one for yourself. And you know, upload images of yourself and say, hey, I would like to make one image of a character sheet. This is me putting your name, this is where I'm from. This is my background, whatever information you feel comfortable sharing. And then ask for please display this, these images on a character sheet from different angles with expressions organized here. And then you can use that final result for any other purpose that you like. So you don't need to keep re-uploading images of yourself again. You have a sheet now with all those images in it. So if you start a new chat with that of that to all which had sheep to the images, that can be your first static point. It's all really done for you. It just takes time at the start just to set that up properly. - Okay, what about for people with longer hair or changing hair styles like, you know, a lot of women might change their hair? Can the AI just change that for you or do you recommend taking different kinds of shots with different color clothing, different style of hair? Or is that less of an issue now as it was kind of before? - It's a lot less of an issue now. These models are great at editing, how you look. And actually I was vibe coding my own app for hair cut recently because I realize every time I go to a hair cut, I've got, I keep having to find old images on myself to show it. I'm like, when I just make an app that uses AI to show me with different hair styles and I keep like, I want this. - Oh, okay, that's cool. That's all driven by AI tools that take a photo of myself and add their hair style in. So it's not as much as important, I think, 'cause AI tools are really good at making those modifications now. But if it's something important to your brand, if you'd like to have a shirt with your logo on it and you have five different variations at that, more images of those variations obviously are going to help. The way to see it is you are speaking to someone who has never met you before, has doesn't know anything about yourself, your background, your product, who you are, how you look. You have to treat it like that at the start because it can't magically come up with that. So it's important to try and give it as much information as if you were telling someone, or meeting someone you yourself about you. That's how I would frame it. - Love it. Okay, so let's assume we've created some of these style guides for our products and ourselves, but we wanna make video. So how do we connect all the dots together with this? - Yeah, the video step is always, we have one once, but always the most challenging. But we are again, once again, at a good time where right now we're able to create pretty good video of ourselves. There's two tools I actually really like using for this. So C-Dance 2.0 is fantastic at characters. If you upload images of yourself, it can take multiple images of a person, like eight to 10 images as part of the prompt. Or you grab that character sheet you were discussing earlier and use that as part of the prompt. Say, hey, this is the character you have to use. Models are good now because they can accept more than one image. So you can really say, this is the person from different angles. If you want this person running through a scene, it will use all those images to create that and not just make it up fully on the spot for you. And I think that has been the big unlock this year. And when Cling 3.0 came out early, I think in January, that was the first time we really saw the power of this, being able to create basically custom characters, or ourselves as characters that exist in the 3D space, and see them animated in a really good way. And you can export at 1080p and 4K now, which means you get that quality as well. So it all comes down to making sure you are as organized as possible when you put in the prompt. Let's talk about that. Obviously, we're not going to be able to get into story here super, super easily here. But do you have any tips on how to properly prompt? Let's just assume we've got some of these foundational things that we've been talking about, and we want to create, let's say a 30-second clip or something. Maybe it's an ad, maybe it's who knows what, a real. Any tips on how we can maybe get better quality output, because you've kind of hinted it's all about the prompt, so what do we need to know about that? Yeah, so mixing the different elements and a prompt are important. The text where you write important, but also as important as the images you give it as well. I actually recommend if it comes to a character, put your create an image of the character in that environment and scene already, as a reference or even as a starting frame, give it as much information as possible. And you have a lot more flexibility on the image side than on the video side. You can do, you can create 100 images in the time taken to create probably 40 videos. So. So we're almost like storyboarding here a little bit is what I'm here and you say, right? Exactly. If you can be as organized as that, where you come up with a storyboard and you know what each frame and what scene should look like with your character embedded in it using those other tools we mentioned, when you go to the video step, it's going to be so much easier for you. And also takes a lot of pressure off the prompt as well. In the past, we would have to write paragraphs of every single day. For first five seconds, this happens to come up with this and that. But now we can use images and a reference of characters. You can vote for it as fully on how you want the scene to happen, whether it's the camera moving in a particular way, whether it's tracking the character, whether the character has to turn or speak or say something. It's a way more easier way now to prompt as long as you have all those other elements prepared. And that is where the artistry sits, I think is that preparation and then once you get to that prompt box, you have so much already ready to go that what you type in for text is an easy paragraph. It can be really succinct. Well, and you did mention earlier now a lot of these tools allow you to more easily edit. Does this mean if we're using a tool like Magnific and we get back something that's like 90% there, but we need to do some edits on it. Does this mean we're able to use this kind of tool to make those edits or are we downloading this video and bringing in another tool to do the edits? With Magnific with those workspaces, you could definitely do it. You have a final video that you've created, but let's say it's me running across the road and there's a weird thing happening in the background. A car is driving backwards. It shouldn't be there. You can use that image, I'll use that video as part of a part of of your next prompt, whether it is with Omni Flash or with Runaway or left or even with Kling or Cdans and say, can you just remove the car, the car in the background in this video? And it will do that keeping everything else the same. Love it. Jared, this is like crazy, right? I mean, like the rate of evolution here is just off the charts. And I would imagine in the next couple of months, it's going to get incrementally more easier, if you will, to properly work with these kind of platforms and these tools and these agents and stuff to create absolutely incredible things that are beyond our wildest imagination. And if people want to follow you on the socials, where do you want them to connect with you? Maybe you want them to follow your YouTube channel or connect with you anywhere. And if they're interested in possibly working with you also, where else should they go? Yeah, look, I really appreciate all of your opportunity to do this because it's so fun to talk about this space, because it's moving so fast. But I'm Jared Lou Everywhere on social media, my website jearly.com runs through some examples of past client work. I do a lot of different work ranging from one of ones to actually working marketing teams to running workshops, educational sessions. And all my content is kind of funneled into giving you as best possible the latest tutorials on the latest tools as soon as possible. So I would love to work with different people from fund different companies and different industries because it's always such a fun challenge for me. And yeah, I look forward to meeting as many people as I can. And if people want to check out your YouTube channel, where will they find that? Just search Jared Lou on YouTube. Everything is succinct under my name now. Makes it a lot of easier. Awesome. Hey Jared, thank you so much for cheering all your wisdom with us today. We're so much better because of it. Thank you so much for opportunity. Hey, if you missed anything, we took all the notes for you over at socialmediaxameter.com/a112. Be sure to follow the show in your favorite podcasting app. And if you've been a longtime listener, we would love a review. And do let your friends know about this show and do check out my other show, the social media marketing podcast. This brings us to the end of the AI Explored podcast. I'm your host, Michael Stelzen, and I'll be back with you next week. I hope you make the best out of your day. And may AI help you become more successful. The AI Explored podcast is a production of social media examiner. Before you go, if you want to stop grinding through the ad creative process manually, the AI Business Society can help Caleb Cruz is teaching members how to automate the entire process on July 16th. Full recordings and all the resources are included with your membership. Join at socialmediaxameter.com/ai. Again, socialmediaxameter.com/ai.

Podcast Summary

Key Points:

  1. AI image and video tools are not one-click solutions; they require human creative direction, storytelling, and brand understanding to produce meaningful results.
  2. Major updates in 2026, such as Google's OmniFlash model and Veo 2, focus on editing control, consistency, and integration across platforms like Gemini and Google Workspace.
  3. Top current tools include Sora (video), ChatGPT Images 2.0 (fast, text-accurate image generation), and Google's Veo 2 and Imagen 2 (strong in realism and editing).
  4. Hybrid workflows combining AI with traditional methods (e.g., brand guidelines, storyboarding) are essential for consistent, high-quality creative output.
  5. AI lowers barriers for storytelling, allowing anyone to create visuals, but professional results still require skill, iteration, and a clear vision.

Summary:

In this podcast episode, host Michael Steelsner interviews AI educator Jared Liu about building powerful AI image and video workflows for marketers and creators. Jared shares his journey from early content creation to becoming an AI educator partnered with major companies like Google and Adobe. He emphasizes that a key misconception is expecting one-click Hollywood-quality results; AI is a tool that requires human direction, storytelling, and brand understanding.

The conversation covers the latest 2026 updates, including Google's OmniFlash model, which integrates video generation and editing across Gemini and Workspace, and the refreshed Google Veo tool with project-based workflows and agent layers. 0 excels in speed and text accuracy for thumbnails and graphics. Google's Imagen 2 remains strong for editing.

The hybrid approach—combining AI with traditional methods like brand guidelines, storyboarding, and human oversight—is critical for consistency and quality. AI empowers anyone to tell stories visually, but meaningful results come from iterative creative work and a clear vision, not just pressing generate.

FAQs

The biggest misconception is that you can press one button and instantly create Hollywood-quality content. AI is a tool that requires a human to have a clear creative direction and story to tell.

Before using AI, you need to define your brand, audience, color palettes, and guidelines, even if you use tools like Claude Design to help create a brand guide from existing assets.

OmniFlash is a new Google model that works across Gemini and Workspace, enabling video generation and editing with prompts, scripts, or videos. It's like Veo for video, accessible from Gemini directly.

Flow is Google's creative design tool for images and video. It now supports project-based work, character training, brand guidelines, and an agent that allows conversational bulk creation and editing.

Seedance 2.0 by Bytedance is currently the top video model due to its high realism, excellent prompt understanding, and ability to handle images, video, music, and audio with sound effects.

ChatGPT Images 2.0 excels at generating realistic images with coherent long text, making it ideal for storyboards and thumbnails, and it's very fast. Veo 2 is also strong, but ChatGPT Images 2.0 handles text and branding better.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.