Go back

E06 integrating AI into embedded products with Souvik Pal

47m 10s

E06 integrating AI into embedded products with Souvik Pal

In this podcast episode, Shovick Powell discusses the practical challenges of implementing AI in embedded systems, where constraints like cost, power, and compute resources are significant. He shares insights from projects such as a wearable safety device designed to detect threats from behind using computer vision. Initially, the team attempted to use an ESP32 microcontroller but faced memory limitations when trying to run object detection models, ultimately switching to a Raspberry Pi Zero for better performance. The conversation covers the importance of model selection, including using pre-trained datasets like COCO and tools like YOLO, and techniques like quantization to fit models into constrained environments. Another example, a smart door system, highlights integrating AI for real-time safety applications. Overall, the discussion underscores the need to balance hardware capabilities, model efficiency, and real-world usability when deploying AI at the edge.

Transcription

7710 Words, 41222 Characters

English
[LAUGHTER] Oh, it's a good one. All right, here we go. Welcome to the embedded AI podcast. I am one of your real life human being hosts, Ryan Torbick. And I'm Luke and Dunnie. And today, we're being joined by a special guest, Shovick Powell, who is the Chief Product Officer at Phy Labs. Hello, Shovick. Hello, Iran. It looked pleasure to be here. Our inaugural. This is our first guest. So excited. It will. He's like Jeff. Jeff. Oh, Jeff. Jeff. I ain't a Collins. Kind of. Jeff feels more like-- He's a family. That's family. Just came over from dinner, you know? Exactly. Yeah. Shovick is our first propagast. Proper guest, there we go. But so-- and we're talking to Shovick, because for about the last eight years or so, he has been helping customers bring a number of embedded projects to life, like giving them life. And many of these projects have had AI being used in a lot of different contexts. So we're going to talk today about some of those projects and then how the costs and trade-offs, how you determine what to-- how to use AI and what kind of systems to use, and all that kind of stuff based on his experience with that. So where should we begin? I mean, there's a lot of projects that over the years, we have worked in the embedded field, in AI field, and only until I guess like five or six years ago, the tools started to come together. So there are a bunch of projects where, I guess, like traditional machine learning solutions would be applied. And then in hardware-wise, there's no barriers, traditional servers, or desktop settings. But with the embedded, the way I define embedded is where we have constraints either cost, in space, or compute, or power. And that's where it becomes really challenging to deploy any sort of advanced algorithmic solutions-- AI and whatnot. So AI is also a blanket term. It quite was. But essentially, like, essentially, solutions that require heavy compute cycles, B, C, P, U, and sometimes B, G, P, U. And that's what becomes really challenging. Should we talk about a couple of projects? Well, yeah. Just to be clear, to make sure I understand, when you talk about traditional, you're talking about, there were machine learning stuff. But all of that was happening on a server that was disconnected remotely or whatever. So you had a very simple embedded system, very constrained on resources as doing sensor data and gathering and just pumping data back to a central server. And then that central server was then kind of doing that. And then, say, five, six years ago, then the edge devices started to be more powerful and be able to handle a lot of this and that kind of stuff. Yeah, absolutely. I began the realm of possibility. Previously, it was not even a consideration because you just gathered the data and just send it over to the cloud and whatnot. And if you think about it, that also was enabled by advancement in the communication. Five years before that, like, it was 10, 15 years ago, even that was not a possibility. Because you stream all that data real time, the infrastructure was not supportive enough to do that. So any kind of AI has to be offline. Like you gathered the data on the field somehow transfer. And then, and then, and then you just be stick and then then walk it back to your desk. Yes. Talking about use be sticks, I'm thinking floppy disk. OK. OK. You need enough data. So I'll meet you in the middle zip drive. We'll just get a zip drive. [LAUGHTER] It's fascinating. Anyway, let's not talk about hypotheticals. What would be an interesting project to talk about, that you, that you try to go? I think one of the pretty cool projects actually happened, like, I would say, two and a half years ago, where, you know, some of the solutions that are available today, I would have looked at it in a bit different light. But this was a project where we're talking about a wearable device that you can simply stick it on your back and maybe you are a fan of walking or jogging or running in the neighborhood late at night, especially for our women friends, they feel unsafe, especially at late night when visibility is poor in certain neighborhood. So the whole idea is for this embedded device to have a camera and using computer vision algorithms to detect if an imminent threat is coming from back. In form of a person, of course, or a car, or it could be a bicycle, or maybe some sort of an animal. So that was a project, kind of like, have a wearable device, put it in your back, have camera, graph frames, perform object detection, and alert the user. You never know when a deer is going to become charging on the street at you. I get to. [LAUGHTER] It's an elk. It's an elk. Take cover. Like-- [LAUGHTER] All right, so what did you end up using in terms of hardware and how did you even decide on that particular set of hardware? So let's think about this. So obviously, the main premise of the discussion is, OK, we have the smartphone with the user all the time. Let's use that piece of hardware to the extent possible. That's the first objective. But can we put a smartphone in the simulist bag? That's not a, let's say, a feasible solution for many reasons. It's heavy. And also, it's not elegant. And you would like to have something dedicated for this purpose. So could it just be a camera grabbing frames and sending to the phone, maybe, but the-- in this case, the objective was to make it as stand alone, as possible. And mind you, not everyone would have access to an expensive smartphone. I don't want that to be the limiting factor. So then, not only do we need to have camera, you also need to have compute that can understand, process that frame, and run machine learning model. And all of that has to run on battery for at least six to seven hours, like an entire session that you're outdoors for like biking or jogging or walking, right? So that kind of gives us the kind of operating envelope of the solution. You're looking for a compute that has camera sensors, integration capabilities that can run on a battery for like five hours. And hopefully, with some real time, because it wouldn't be of real use if one frame takes 10 seconds to process in the context of this application, that would be kind of useless. By the time somebody's already, you know, like, attacked you. So it has to be one frames per second or more, ideally. And the initial kind of daring move was to use a microcontroller. Why? Because we were simply fascinated by ESP32's power kind of management, extremely power efficient way of grabbing sensors and doing basic computation. But the main challenge was we're talking about object detection model to run on ESP32. So that by itself is quite a challenge. I'm not sure if in our audience is really familiar. But even today, this will be considered quite an ambitious, you know, task. It's simply because there's so less memory available on an ESP32. And also the compute. We're talking about mega-hard scale processing cloud completely. So what I'm hearing is that the memory constraint was actually the greater difficulty compared to the compute is that right? Absolutely, yeah. Because for running any kind of machine learning analysis on object detection, you have to load a model. The model by itself would be of certain size, right? So to pack that model into that framework by itself, it's the challenge. And you're going to ask yourself, oh, does the embedded environment allow floating point operations or are we limited to integer only? And these things does impact accuracy quite a lot, right? So I mean, the model itself, I mean, are we talking gigabyte? And the models, I know that I've used old llama. And like, that's 10 gigs. Like, I'm downloading 10 gigs of a Docker image. And I'm like, oh, my god. And then it says, we like to update this. I'm like, I don't know that I have an app hard drive space to update this. Like, how big are-- Yeah, yeah, yeah. How big are these models that you're talking about? Absolutely right. And questioning that. And yeah, yeah. So we're talking about not a gig size model. But like, if you had a reasonable size mini PC, you could do this. Probably we make a sub gigabyte, like, several hundred megabytes type model. And get away with it. Because remember, we don't need to differentiate our classifier between like 18 different classes. that the use cases kind of limit people are car. Car dear like dear. Yeah, dogs dogs big foot. You know, it's like, okay, I got like four or five. Yeah. So that can help you reduce the model size. Yeah. So, but but I get what I'm trying to refer here is that why we kind of taught somewhat feasible is that, okay, we can reduce the model size and they know that what is a model? It's a bunch of weights. Right? A bunch of numbers saved in a predefined format. And if you know the format, you can basically use that for executing a very complex decision tree, right? If the numbers that you're computing from the sensors match this, then you are in this class otherwise. So it's basically a very simple. I'm probably oversimplifying this, but a very simple decision tree that is running from inside an embedded environment, right? So that's why we kind of selected this. So you started out with the ASP32 and then you were like, okay, we're going to try to work within the constraints of the ASP32 and then eventually it's like, that's it. This is not, this is not going to work. Like, like, what did you end up doing then? So that was, let me dig into this just a little longer. Okay. More. What kind of model size did you start out with? I'm guessing you just, you grabbed something off the shelf and said, okay, fine. Just try this and then you noticed, well, it doesn't fit into the memory and then you compressed it down. I'm really curious. Yes. What compression ratio you could achieve and I don't know, doesn't make sense to talk about the specific model you used or sevens, we're using sevens, it's like, we can get this smaller. Have you tried opening it, mate? Unfortunately, some of those tools are not available at our disposal and actually that's one of the reasons the project feels. So we started with a TensorFlow Lite kind of a model. We may have looked at things like mobile and SSD, of course, the other one. There were quite a few in that bracket escaping me right now. But these things are not suitable to run in a USB 32, by the way. Right. So you cannot run a TensorFlow Lite model as is at this, to the rest of my logic, until today, you cannot readily take a model and just expect it to work in a USB 32. So what needs needed to be done is to take that model and kind of write a custom adapter or driver that basically extracts essential part of that tree, a model tree and save the weight in a kind of a custom C format, right? Like custom format in a C file. And that file would be uploaded to the USB 32 or it could be part of the code in a large array. Is this what you did that which people refer to as quantization or is that really just ingrown it into a different kind of on a. Yeah, lively speaking. This was our effort of quantizing a model to crazy, crazy, crazy level that it gets. Zileffy, you know, what its quantization is basically taking it into a model. Taking a big sort of, you know, block of information to break it in smaller pieces and hopefully summarizing along the way without losing information, right? So that I guess my work. I don't know too much. Yeah, exactly. Right. And the main challenge was in two parts. First of all, this conversion, you know, how to do this without losing accuracy. And there were some guides that I could find that maybe in after the call, like I can dig those sources up and share if you're interested in person. Oh, that would be perfect. We've got a lot of content to show out. Yeah. And the other thing is once you have such a model, I think the ESP IDF is the kind of the development framework of the library that supports all the ESP project was giving us a large trouble, like error software errors in library and compatibility and anyone having trouble with that. That's just the first time I heard that. Right. Yeah. Yeah. Yeah. So these are the two main challenges and we kept going in circles and then finally we gave up. And then, but you didn't properly give up. You just moved on to a different platform, right? Yes. Oh, yeah. Absolutely. Yeah. Engineers, we are like designed not to give up much to our detriment. It's not so much to our detriment. Yeah. We say, okay, fine. ESP 32 doesn't work. Okay. Instead of, you know, six hours, you can only do three hours. Okay. Let's do it. Right. Let's find that out. So at this point, we were looking at Raspberry Pi 0 as a next kind of possible candidate. So we were really like familiar with the hardware. So that's what we ended up doing. And it was TensorFlow Lite, like with some sort of rust server that is running that can maintain a communication with the smartphone. It's still a standalone device. It does the grabbing frames from the camera, runs object detection. And then when it finds something, it basically communicates to the smartphone via Bluetooth. BLE. Yep. And then, and there is a companion app running on the smartphone that basically alerts the user, DRLR, too. Yeah. And then, and then, and then, it can boot through energy from their back somehow around the body to wherever the phone is. And then it can then, the phone can then communicate to the headphones. This is like what the hell from that dear? Let's maybe generalize from, from this story. And, and think about, let's call it the design process as far as AI is concerned. Like, how would you, or how did you select a model? How did you decide what to do with a model? Like, do you use it off the shelf? Do you? Yeah. Modified in some way. I mean, I get to some degree, it was necessary to modify just so it can, can fit onto the experience. Yeah. Yeah. Basically, because we are interested in many, many of the other classes. So I think a good dataset to focus on if you're talking about embedded object detection is co-coded dataset. I think researchers are, the users are really familiar with it. The co-coded dataset of, you know, right out of the box offers, you know, like 50 plus classes that has been trained on. And it does quite a reasonable job in, in terms of accuracy, terms of, you know, bounding box, what not. How do you spell that? T-O-C-O, co-code dataset. It's kind of like a go to, like if you're not trying to detect something weird, this is what will go to. Like, for example, where this would not be useful as we had a project where somebody wanted to detect surface level threats in marine waters. Like, I will be looking at torpedo, a drone, a debris, a boat, submarine, this type of stuff. And that could be foggy weather, that could be rain, you know, that could be in a distance, it could be a mirage, it could be reflection from water. Oh my gosh. So, I don't want to, I don't want to even, like as I'm driving, you know, I'm looking at a lake or I'm looking at the river and I'm like, I'm looking at a thing going like, I hope that's a dolphin that's coming out of the water. Like, what is going on over there? There's so many things that, like, it can be sticking out of the water and you just, you hope that it's nothing weird. And now you're going to train anything. I'm trying to go and figure that out. Oh my gosh. This would be a difficult problem for, you know, AI, because you need to start thinking about, okay, you know, I'm quickly running out of options with my, you know, established, you know, data sets like cocoa or whatnot. There's other similar data sets like that that has specifically trained on lesser marine life and whatnot. But if that doesn't fit your use case, tough luck, then the real pain begins. Now you have to really construct that model. You have to go through the pain of collecting data and not dating it, labeling it, you know, training, hopefully something down to like a yellow type situation where it is more amenable for like embedded environment. And for this project by the way, before I move on, I can't remember what was our experience with YOLO in this particular environment. It's been a while. But YOLO is, as you know, is a good choice for a constraint, compute, constraint environment, especially with the more recent person. It's quite promising actually. So that's another thing to consider. But for model choice, I would first look at by use case. And if I can, you know, select an after-short model and YOLO, I know, offers a ability to fine-tune it with, I would say, fairly modest hardware. And then, you know, with some limited sort of samples in the tune of maybe hundreds or maybe in the low thousands, we can fine-tune an existing model with some new samples, right? So these are the type of things that is easily achieved. And then it becomes difficult. If you have to do it from scratch, then you're looking at it. - Yeah, doing it from scratch. - Contenders of thousands are, you know, tens of hundreds of thousands in that realm, but images needed to get a good performance. - So yeah, yo-yo is, if you go to an embedded conference now and you see someone with an embedded system and then they got a screen attached to it and they're identifying your people and coffee cups, whatever, that everyone's using yo-yo models on hardware now. - Oh yeah. - It's like a neat pony trick that everybody, but like what can we do? Like, look, you're holding a coffee cup. I'm like, that's what they said too. Like, then they're on there and these people, like, well, I got it. It's neat. It is really neat. But what do you actually do? Like, I got all sorts of other questions now when you come down to it. All right, are we ready to move on to another example? - I, yeah. - Sure. - All right. - Look, if you have, yeah, the pull-out questions on this one from me and one. - Oh, so yeah, let's talk about some more words. - Yeah, yeah. - And then we can dig into the learnings from that. - Yeah. So I guess the next one, which packs in quite a lot of, you know, features and therefore challenges is a smart door system that we talked about. So what is it? Just let me just describe the product so that, you know, the audience has an idea. We know during video talk, it's kind of the Xerox or the one, like, even when people are not buying a ring video door, that's what they say. Oh, just get a ring or something. - Get a ring for it. - Yeah, which is a video doorbell that's connected to cloud. So what we are talking about is kind of that, but with a huge added screen in the inside of the door. So the whole idea is like, if somebody's at the door, you wanna kind of see through the door, right, who is kind of waiting outside. So that's kind of the vision of the product. Like think of that huge screen, which is covering 80% of the door, your front door. And when somebody comes, it just basically projects the picture that the front camera that ring video doorbell, the video doorbell is seen. And you basically see through that door. - Yeah, but I don't mean, do I? - I do not, do I? (laughing) They got holes in them. - But now those are weird. I feel weird put my eye up to that 'cause somebody's gonna poke me or something. I just, yeah, it's like, yeah, you can use the, like a peek hole, but the idea is like, if it's for safety, like the primary objective was safety, like if somebody is holding a gun or somebody is appearing to be friendly, but it's not and trying to hide some things beyond the use of like detection of porch pirates and simple use cases like that. But safety was a big concern, especially for kids, like it there inside the house, they may not be that versed in using the peek hole and whatnot and showing them what's on the other side. - Why did you, don't open, what are you doing? And you just open the door. Hey, I was like, come on. You know what you do? It was on the outside of the door. - Yes. That was the idea. And it kind of evolved. That was the initial idea for the product. And as we were trying to build it, new things came up, like, sort of what you know now is that it's very easy to trick facial recognition systems. If you're talking about like simple computer vision based, like keep on detection and then matching the face with an embedding and then see, like, oh, you found your face. That is very easy to trick with a picture even. Or like a smartphone or whatnot. So now they do what called live news detection. And that's a slightly different technology with, you know, IR involved and we look at a lot of other things. Are there dedicated sensors that do live news detection? So in this case, we had that. - Let's put it back to the AI aspect of all of this. (laughing) Yeah. - What's awesome AI features did that smart door hub? - Yes. So the first thing that we had to do is obviously in the facial recognition and kind of voice signature identification for authentication, right? So when somebody is coming outside with a, you know, a hands full of battle groceries, it would automatically identify and then unlock the door, right? So that's one of the use case, but in order to do so, we need to run algorithms for voice signature identification because they need to say out loud something, right? And they, it also need to understand voice commands, which is also slightly different task than voice signature identification, right? So we're talking about facial recognition, kind of high level object detection, like what's again, in the door, are we looking at a person, are we looking at bunch of people, are we looking at cars, kind of scene description, right? So we have the camera, camera, technology images, we just described the scene, how many objects can we detect in the scene, right? And based on that, we basically create a scene awareness for the algorithms, then we do an audio sort of, you know, voice command extraction from audio, and then we also have to do signature, basically with a voice fingerprint that is saved in a system to match that user. How many different models are you talking about then? Because I'm seeing-- - Not a lot. - Three different models are offered at the same time, at least. - Exactly. - Yeah, and this is-- - And this is one of the-- - And the ESP32, right? Like, you shouldn't have any-- - No, no. (laughing) - That's a non-shutter in this case. And let me throw another challenge into the mix. It may sound trivial, but it's not, is that the inside screen kind of had to stream the camera view, sort of in a continuous basis. - Let me summarize. So we've got this door with this gigantic 4K screen. We've got a 4K camera, obviously, on the other side. We've got a microphone both inside and out. And so the easiest button, not exactly trivial, part of that was just piping audio and video back and forth. And at the same time, you needed to split out the video stream and all the audio stream and make that usable for AI models for what? So we had gesture recognition. We had face recognition, and I take it those other models. We had language read. Did the door have generic language understanding, like, like, I'm what I'm-- - No, no, no, no, no. - What did it do as much as I was looking at the stuff? - Okay, so I think-- - You know what I think? We'll just hire a person to stand on the other side of the door. (laughing) - You know, okay, the listeners, show of hands. Never mind that I can't see you. Who of you has read the Hitchhiker's Guide through the Galaxy? - Oh yeah. - And who of you can remember? Who of you can remember the doors on the spaceship, the heart of gold? (laughing) Because they had this idea, long before you're showing. They-- - Yes. (laughing) - In the book, by the way, those doors were universally hated. - They universally hated, not operable. (laughing) - And they-- - I hope they talked to you too. There was a whole communication that made this. - That's the thing, they were always obnoxiously happy. - Yeah, well, with-- I don't know if you've met any AI models, but they are obnoxiously-- (laughing) - Well, with-- - Yeah. - I think we're talking about-- - I'm very much in response to something. - Yeah, something-- - We can make things from this. - Is the door then responding back to you on its own? Where we can fit in on the model here, guys. (laughing) - Oh no, you did, it did, actually. It's pulled back. (laughing) - Let me kind of summarize that, like, exactly that ended up happening. Because if you naively kind of think about this, you would like, oh, text to speech to text, and this is kind of mainstream now, everybody's doing that. We, so we try to do real time transcription on top of all that madness, and it's like, no, it's not gonna happen. This is an example where real timeness is also very important. So, you know, a delay of, you know, in two to three seconds, maybe non-acceptable. So in this case, so transcription is out of ways. We're not able to run any kind of, let's say, whispers, or even faster whispers, if you're familiar with these kind of fairly available transcription situation that models, we can't use them in this situation. So what we ended up happening is wake ward, we kind of cheated, we did wake ward detection. So there's a wake ward detection engine that you can say, like, hey, Jarvis, or that kind of a thing, it just kind of wakes up, and then it kind of prioritizes the utilization of the limited CPU and memory that you have when somebody's, you know, it detects a wake ward. And then we were able to actually fine tune that wake ward detection model to understand additional commands, right? So the idea is just like you know, hey, hey, you know, the smart door for the heck of a good name, right? Then they open the door or show me outside or start an intercom. So we basically picked up phrases that are like phonetically fairly distinct, and therefore we could have individual, like the wake ward engine would classify them as individual phrases, and we do. just use that essentially. So that is fairly low compute. So I don't want to go too much into that detail. - Oh my God. - Basically, you can talk to it as a Linux box running on a simple mini PC that you can buy in Amazon today, for example, for $200, $200. So one of those type of computer, let's say, right? And there's quite a lot of embedded, like industrially, kind of rated solutions in that family. So that's what we ended up using. You know, people worried about AGI, you know, the generic intelligence. And it's like, the amount of work you've had to do to replace a door man. (laughing) In a situation. - Pretty much a four. (laughing) But it did all of that. - It did all of that in real time. So yeah, it's amazing. What the me actually, I've been for our embedded AI enthusiasts. I would say I'm kind of disappointed with Raspberry Pi 5, you know, the promise that the Raspberry Pi family of devices could have in the embedded AI world or embedded algorithms world. The mini PCs have like, cum bleeps it down. - Yeah. - It's a similar price point, like $200. You could buy a Pi 5, maybe a couple, but then the performance is simply unmatched. Them, you know, the simple flops, memory, and you know, support of other peripherals. So like a lot of good embedded solutions can be developed using those systems. - I think it's interesting when you take the fact that you don't have to worry about a battery anymore, you can really just really just throw a bunch of compute at it and like, okay, well, this is a lot of compute. Fine, we'll just use electricity we have on the wall here and just go to town. - Yeah, exactly. - This is fully plugged in. So different projects, different challenges, right? So space was the real challenge here and the tunnel management. - The thermal management. - And then all the complications of running, you know, four different models on a single computer and having them all interact and then, well, you said he's fine, but like, ah, the facial recognition said, no, but the voice recognition said, yes, and like, oh, they must be wearing a mask. Okay, we'll never mind then. I'll ignore the facial recognition part. - Yeah, there's also a decision tree. And I think like one of the last things that we didn't do it, but there was a potential was sign language for, you know, because we can now do finger posture or gesture detection quite well. So then the idea is to do ASL by simply reading hands. So then the indoor or the outdoor camera, somebody can just, you know, do stuff with their hands and the door will react to that. - Holy cow. - Yeah, for listeners who don't know ASL is American sign language. - Yeah. - Yes. - Yeah. - That's actually a way of communication for a lot of hearing impaired and speech impaired people right, so it's a good option. - I just say someone demoing that too, whether you could hold your hand up and then they had a robot hand who would try to mimic your hand. I watched it work for several people. It didn't work for me. I was like, I'm kind of offended by this. (laughing) - What was my hand? - Let's pull it back to tech. Like, what did you actually, what did you do framework wise to get all of this stuff integrated? Like you already said, yeah, what FF player I think for the just plain old audio video streams, but in terms of AI pipelines, machine learning pipelines, what did you do? - Yeah, so like I said, for wakeboard detection, we were using enough to show the engine open, wakeboard, it's actually quite powerful. - And you just like how it's straight into, like what, the audio device? - We used FFAMP extensively for the video part, but for the audio part in this project, we used sound device library and there were a couple of other Python based library. - Right, and the actual machine learning or AI models, all right, so we've got one of them. We've got the wakeboard detection and all of the others were similarly just sort of individually wired into. We can basically write audio chunks in memory buffer and that's, I would say recently, fast way of sharing the raw data between different concurrent programs. You know, another program can look at those audio chunks and run other kind of audio algorithms, like fingerprinting algorithms, like because if somebody's speaking, you don't need to have the entire audio stream to decode somebody's or match somebody's vocal signature, right? You know, there could be some sort of signaling between programs and say, okay, look at these audio chunks in this memory buffer. Tell me if it's my or if it's Ryan or if it's Luca. - Right, so that's all I'm so proud of. - Did you also have to do some kind of, let's call it scheduling where you said, okay, fine, I need to put one of the models to sleep. So I have enough compute to run another one. - Okay. - Yeah, that's actually a good point. So how will you run everything, like as a concurrent engine or not? So this was a key aspect of engineering. So there were a few entities that I call, let's call them services. So we convert them as system D services in a Linux environment, right? So they'll run perpetually, they'll have their own logs and they'll be coming back on their own if they crash for whatever unfortunate reason. So, but there were things that were run like on the spot, right? They're called invoked execution happens and then just go away. So we had to do all that management because of that limited compute, right? - Maybe we can move to an adjacent topic because we've been talking about the tech a lot, but I've heard that tech isn't exactly what sells devices. I'm wondering what's the secret to integrating AI into products? - Well, as opposed to just somehow, how do you build good AI products? - That's a loaded question. (laughing) - User experience, right? That's where it kind of comes to. Some does it better than the others. Depending on what the AI is doing, I think for most users, they would still like to think that they are in control. The human in the loop is in control, not the AI, especially, and pardon me, from being, you know, kind of, I'm differentiating, it's like younger generation, right? So, the older generation, they don't trust tech as much. So they would rather have all the decision, you know, kind of they would rather make those decisions and then have AI present the options and information. Younger generation is like, oh, just tell me what it is. And then they'll just like follow, they're more trusting, I would say, of AI. So a good AI would strike a balance between like, you know, what it enables the user to do and what it does for itself. And then I think making the users aware. For example, you've seen those crazy videos where Google calls restaurants and salons to book appointment, like pretending to be real human, right? With all voice intonation, and this was like several years ago. It was pre-LLF days, right? And it was like a live presentation by Sundar Pichai. And they, I conducted pretty well, those calls, was able to book show book appointments, but the point is, it never said that, hey, I'm an AI. So the human on the other side always thought they are talking to a real human. So this is a key consideration, building a solution that I think we should let the human know that they're interacting with an AI. However smart that AI might be. - Yeah, just to build trust, right? - Yes, indeed. Yeah, so some products are different, the others, like, people use AI fridges these days, just to tell them, you know, what recipes can be cooked, but somebody still needs to cook them. And it's like, okay. - Yeah, I'm like, I have to find a recipe. If you just leave me like, - I'm not gonna do the fucking thing. - I'm in the fridge, I'm gonna make the same things over and over again. I literally looked up something, like, okay, I have this, tell me something I can make with these three ingredients, and then, oh, thank goodness, I haven't tried that. That's fantastic, 'cause I'm gonna do the same thing over and over again, like. Do we wanna go to, how much time do we have? Do we have enough time for another example, or do we want to, what do we wanna do? - It feels like, it feels like we've covered the entire topic pretty well, haven't we? - Yeah, especially we've gone through a very resource constrained device, and we've also gone over, like, oh, no, we've had all the power, we just have different constraints. So, maybe generalize now and kind of like, take away some takeaways now. - Yeah, exactly. Like, what's a good thought process? What's a good design process if you start ideating an AI-enabled product, whatever it might be? Like, what's the primary concern? - So, I think the number one, and there's no right way of doing this, but this is how I tend to think, when I think of an AI solution, it starts with power. That's number one consideration. What is your power budget? Is it sub one what? Is it like, you know, somewhere between, you know, one to 10, maybe 15 watts, or it's more than that, right? If it's sub one what, there's, that immediately restricts you in terms of what you can do and you are like now Kind of limited to using microcontrollers as opposed to microprocessors Which means yes, we 32 which means pi pico which means SST in 32 like those familiar product a Nordic has a bunch and and with some clever tricks For models we can still make some sense are out of those systems. So I'm not sure if you are familiar like for example this guy here like ESP 32 S3 sense or sense i as they call it so Basically is with 32 at the back as a nice little screen attached and this can actually do object detection on this like this is one of the latest well, I would select a year old kind of technology So this guy can do key point detection on the face can do facial detection recognition all that does so this is possible now Right, so this is just a regular ESP 32 Then second question once you have sorted out the power situation is space how much space you have Because the next so the next level up from this would be the mini PC's of the world light So for example the pi pi 0's right so this is a raspberry pi 0 and We could use this where we have you know slightly bigger space and maybe the power That's available we can run a full Linux operating system that allows us to run by transcripts that allows us to run A whole lot more family of machine learning products right that's now possible with sound and with video the primary to use this but can you run an LLM? That would be a daring on the world like I would you know Maybe you can take it up as a project with extreme levels of quantization based on what you know we learned earlier in the podcast right yeah I think the train stations I guess like if you can if you can if you can wait 30 seconds for a reply Then maybe an LLM is feasible and if you want something a bit quicker than not right? Yes 100% and then like obvious step up would be its big brother like pi 5 And by the way like when I mean when I say pi 5 it's not only only one in its group like this rock by there is like Whole lot of boards that are in the can a banana pie orange pie Although personally I haven't used those I feel like the software stack is not ready there Community support is not that good so But if you are into it then go for it. They have awesome hardware they have GPU Mali GPU on some of the rock chip products like the rock in a pie family of boards awesome and then Once you're above that 5 to 10 what like then you can go on to something more powerful like jetson with with you know industry leading GPU support right if the models need to use GPU for you know real-time transcription beat LLM are Simple like computer vision Jetson and then I think beyond that we are not in embedded territory anymore So you just go like full blonde servers Exactly and I mean Jetson is only embedded in the sense that it has a small form factor and Reasonably modest power draw, but like this is a proper gaming PC from like 10 years ago. I guess I would say like cost I think in that order power space and then cost right so that's why if Jetson comes last in my Discussion because it's the most expensive of the buds that we just talked about right expensive in dollars And also in instead of everything else like it it needs it needs cooling it needs power Yeah, it's embedded because it's made by Nvidia, okay like It's not an Intel like it's not Mike. It's not running Mike. I can well actually probably could but By the way, I found it interesting, you know show it you started out with with power as your number on Consideration and and just yesterday the The steam machine was was unveiled by steam, you know that they're sort of embedded gaming PC And what it was interesting about this is they were asked about the design process and they said well we started with the fan Because the fan sort of constrained the the thermal budget we had and that constrained the power budget We had and everything kind of followed from there then and that I found that really interesting Yeah, yeah because that's that's honestly a big limitation of some of those spawner systems like all right Show back how do people get a hold of you and and actually tell us more about what you guys are doing over a file abs Yeah, so at file abs we basically Have the license to nut out to the fullest So we build products The type of products that we talked about right so the smart door or a wearable device etc The people who come up with those ideas are like early stage entrepreneurs startups founders with big ideas But they don't quite have the muscle power to get this underway, right like they'll have a team to Ten of developed these products oftentimes it also requires Embedded hardware, which is extremely hard in a topic by itself, right? So PCB design RF black magic antenna and all of that crazy stuff and What do we do with that like farmer developments a big part of it, right? So so So that's what we do we build the product for like end-to-end hardware Farmware algorithms and then finally weapon while software stack and We we another big thing to think about is like we do not take any IP in that process Which is a big differentiator and we do that for a fixed fixed price contract Right, you can and your company is called file abs. So FYE labs Yes, that's why in that I lab we are based out of Hamilton, Ontario Although we have huge presence in US as well and now we are trying to a solution presence in Europe Oh interesting Supplies to be Okay, well, so big thanks for joining us as our first proper guest on the embedded AI podcast I'm Ryan Torvik And I'm Luca Johnny Thanks for joining us. We'll see you guys next time Yep, see ya Hi Luca here. I have a quick announcement. I've just launched one more new training platform the Embedded AI Academy at Embedded AI Dot Academy it feels like I've seen in the market training that approaches AI specifically from an embedded systems angle If you're listening to this podcast, it should be right up your alley and you get 25% off any booking made through the end of 2025 if you use the code Embedded AI podcast 25 Only capital letters again the website is at Embedded AI dot Academy

Podcast Summary

Key Points:

  1. The discussion focuses on deploying AI in embedded systems, highlighting challenges like limited memory, compute power, and battery life.
  2. A wearable safety device project illustrates the trade-offs
  3. Model selection and optimization are critical, often involving quantization, using pre-trained models like YOLO or COCO datasets, and fine-tuning for specific use cases to balance accuracy and resource limits.
  4. A smart door system example emphasizes integrating AI for real-time safety features, such as threat detection, while managing hardware and software complexities in a consumer product.

Summary:

In this podcast episode, Shovick Powell discusses the practical challenges of implementing AI in embedded systems, where constraints like cost, power, and compute resources are significant. He shares insights from projects such as a wearable safety device designed to detect threats from behind using computer vision. Initially, the team attempted to use an ESP32 microcontroller but faced memory limitations when trying to run object detection models, ultimately switching to a Raspberry Pi Zero for better performance.

The conversation covers the importance of model selection, including using pre-trained datasets like COCO and tools like YOLO, and techniques like quantization to fit models into constrained environments. Another example, a smart door system, highlights integrating AI for real-time safety applications. Overall, the discussion underscores the need to balance hardware capabilities, model efficiency, and real-world usability when deploying AI at the edge.

FAQs

Embedded AI systems face constraints in cost, space, compute power, and battery life, making it challenging to run advanced algorithms locally.

The project initially targeted the ESP32 microcontroller for its power efficiency but shifted to a Raspberry Pi Zero due to memory and compute limitations for running object detection models.

Key challenges include limited memory for model storage, restricted compute capabilities, and the need for model quantization or custom adaptation to fit the hardware constraints.

Quantization reduces model size by compressing data, such as converting weights to integers, which helps fit AI models into memory-constrained embedded devices while aiming to preserve accuracy.

Start by assessing the use case and available datasets; pre-trained models like YOLO or COCO can be fine-tuned for specific needs, reducing the need for extensive custom data collection.

Using established datasets like COCO simplifies development for common tasks, but unique applications may require custom data collection and labeling, which increases complexity and resource needs.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.