Go back

Building a Local Large Language Model (LLM)

51m 30s

Building a Local Large Language Model (LLM)

This podcast episode from Southeast Technological University Ireland features host Rob O'Connor interviewing Red Hat software engineers Mark Campbell and Dimitris Saradakis. The discussion centers on building custom AI and large language models (LLMs). The engineers explain that creating a custom LLM starts with a high-quality, domain-specific dataset, as the model's output is directly influenced by its training data. They outline the process: first training a broad foundational model, then fine-tuning it for specialized tasks, similar to a general practitioner versus a specialist doctor. The conversation highlights Red Hat's OpenShift AI as a platform for managing this entire model lifecycle. A key demonstration showed how quantized models from open-source repositories can run locally on personal hardware, making AI more accessible, though initial training remains computationally intensive. The dialogue also touches on the significant ethical and regulatory challenges in AI, particularly around bias, noting the difficulty of curating perfect datasets due to their immense scale and inherent human biases.

Transcription

9925 Words, 51946 Characters

English
Welcome to the machine, a computer science education podcast from the southeast technological University of Ireland. I'm your host Rob O'Connor from the Department of Computing and Mathematics at ACTU. We surprised dropped an episode yesterday following on from a very engaging presentation at Computing Week in the University. I was talking with Gary Cannelli and Sheree Asaniel from Unum about how AI is changing the OECT workplace. That's available now and I can see from the metrics that some people have actually listened to it. So thank you very much. Well, I'm following that up today with another episode based around the talk that we had this morning from Red Hat Software Engineer Mark Cannell about getting started with your own AI. This podcast is probably more aimed at computer scientists than yesterday's. I was very impressed with Mark's AI demonstration that he performed on stage. He had a Gen AI system with a fairly useful data model all running locally on his laptop. And honestly, I couldn't have even imagined that a few years ago. So there were some great questions from students after the talk and I thought, "I've got an app this guy for a podcast." We went for a quick copper and we're joined by another Red Hatter Dimitri Saradakis. Now both of the labs are recent enough graduates of ACTU and I get a genuine buzz from seeing how their knowledge and skills are growing and their careers are progressing. And I learned a lot from our conversation and they've given me plenty to think about when it comes to developing some AI systems myself. Now in the podcast we talk about Red Hat's OpenShift, some custom large language models, quantized data sets and how someone would go about building their own AI. Towards the end, we get into some regulatory, ethical and moral questions. Not too many answers though. So without further adieu, here's Mark and Dimitri. So my name is Dimitri Saradakis. I am a team lead on a product at Red Hat called OpenShift AI, focusing predominantly on distributed training of language models, any kind of model realistically, any kind of AI ML model. And I am an ex-applied computing, what's neonomes as computer science, I believe, a student out of here. So a couple of years ago. Excellent. Lovely to be back. Excellent. You're always welcome and I hope the coffee is still okay. Mark. Yeah, thanks for having me as well. So my name is Mark Campbell. I'm actually on the exact same team as Dimitri and Red Hat and I'm associate software engineer and former software systems development student. Excellent. And you're only gone out of the place less than a year. It's me, yeah. Okay, great. So you're fair-plating. You came back to deliver a talk in the auditorium. Right, before we get into the content of the talk, I just have to ask you, as somebody who's a very recent graduate, how did it feel to be. I won't say it was a packed auditorium, it was a half-full auditorium, which was very good. It's hard to get students go to anything. But there was still a lot of people, probably more people than you've ever spoken to publicly before. How did that feel? Yeah, so it was very weird because the last presentation I gave was for my final year project in front of I think seven people. Okay. So that's quite a big jump, considering. But I think I was more worried before I actually got up on stage and then I was like, oh, you know, it's fine. I rehearsed anyway. So I was okay in the end. Yeah, well, I thought you did excellent. I even feel that some tough questions too as well, which we get to in the next few minutes. So, right, so you're both working as part of this kind of open-shift AI group. Could you maybe explain what is open-shift AI? How does it maybe fit into the overall model of the kind of stuff that Red Hat 2? Yeah. So, open-shift AI is a platform for data scientists to train their models on in a cloud environment using open-shift. So with open-shift AI, they're able to gather all their data, prepare it, train it, and then serve it in a cloud environment or on premise or even in a disconnected environment. Okay. So is the idea behind this that I can have my own LLM? When I say my own, I mean, a person or a corporate one. I don't mean one for me, you know, with my own stuff in that. I don't maybe. But I could have a customized large language model that might be fed with certain customized data could be proprietary data. And that can be, I can use it in whatever context I wish to use it. Is that it? Yeah, 100%. Okay. So I'm not just relying on the publicly available chat GPT, BARD, whatever you're having yourself models. No, you can go from scratch. Okay. Okay. And to meet your kind of knob in there. So do you agree with what he says? Yeah, 100%. I would just add, it's like it's the whole system to manage the model life cycle. So from the very, very start, from your data experimentation, your data collection, and engineering your data, getting your features and whatever, ready in your data sets, the whole way through to training, fine tuning a model, and then ultimately for inference then afterwards as well to actually make use of the model. You stick in front of an end point. Okay. Okay. And we had in the region, when we talked about how this talk might be structured, we didn't say that. And I'm kind of want to, I'm just going to push on it because I mean, I mean, treat boys. So let's say, how does one go about doing this? So you talked about those steps there. What are those steps? What's actually involved in creating a custom LLM, large language model? Yeah. So if we're taking LLM or any kind of AIML model realistically, the most, in the most important ingredient is your data set, your initial data set. Like it doesn't matter what kind of, what kind of flavor of the model you use, what's the newest and coolest or whatever. If your data set that model is trained on is garbage, then your output is garbage. Okay. So this is the idea. So the actual natural language processing is handled by the algorithm, shall we say? Yeah. But then the answer is that it might give you questions is fueled by the data set. Yeah. 100%. So yeah, the algorithm is what, like, quote unquote, learns, was the data set is what gives it the knowledge, the knowledge underneath it all. So again, if I go back now a year or so, whatever, and Facebook's first model was publicly available. And then it was taken by some people and I think was 4chan or H and one of those, one of those places and turn, it was fed. I'm going to say a biased data set. Okay. And I turn into racist homophobic. Is that what you're talking about? Yeah. So if you feed it, if I feed it silly information, it's going to give me back silly answers. Yeah. Yeah. 100%. Yeah. Yeah. Yeah. Yeah. The issue with, and people, there was a question there on during Marx, excellent talk where someone asked about what if we just like exclude certain terms or exclude certain things from the data set. I think the issue, this question I wanted to kind of get a little bit of answer to because I think it's quite an important one. What most people kind of don't realize is the size, just the vastness of these data sets. Like so there's a particular kind of data set that we're working with. Like the mass, the IBM collaboration that Mark also mentioned earlier. And that data set is 250 petabytes. So like 250 petabytes. Yes. Okay. 250 million. I think you left out, you made a 25, not 250. So yeah. Yeah. 25. Yeah. 250 million gigabytes were the data. Right. So like first and foremost, when you're actually trying to formulate your data set in the first place, like the student there, who's asking the question said, well, if we just leave it the terms rob, and you're kind of like, yeah, but perhaps that's not in that context. Right. They're saying someone did rob a bank, but this is not exactly how to do it. It's how you actually create this data set to be perfect when inherently human beings have biases and we're like, obviously inherently not perfect. So like it's it's a really, really difficult problem to solve. There's also kind of an inherent, to be fair to the students, like a younger student. And I don't think he fully understand the whole thing of a kind of natural language processing because even as he was saying that I was going to pipe up, what if you wanted to search for how does Rob me open a bank account? Yeah. Yeah. You know what I mean? Like years and years and years ago, I worked on some natural language processing and the test that we always had was about tiger woods. So it was doing searches for our tiger woods now. But this is back in the day where you have a lot of keyword searches. And if you do the keyword search on tiger woods, you're as likely to get an answer about a big cat and a forest as you were a golfer. But of course, if you're searching for tiger wood, there's a much more higher probability that you're actually searching for something about golf than a big cat in the woods. So you know that that's 20 years ago. So that's a long, long time ago that we were doing that kind of stuff. So I can imagine it's come along so much further since then. Obviously it has. But let's just get back to the idea of building your own custom LLM. So you're talking about the data set is the secret sauce. Or not so secret sauce actually because anyone will tell you that. Yeah. Okay. So where do you go from there? Let's say I have a good data set. Whatever it might be, 225 pt bytes or whatever it is. Okay. Where does it go from there? So essentially it goes through like the like an initial training stage where you'll train like a foundational model. And then maybe Mark might speak on the difference between maybe a foundational model or a base model and maybe fine tuning afterwards. Yeah, cool. Cool. So the base model is the general idea of what you want the future model to be. And so from there, you would fine tune it further for specific use cases. It's kind of like a foundation model is like a Swiss Army knife and a fine tuned model is more like an actual physical tool. Okay, so I'm trying to give an analogy. So you're saying the base model is maybe trained on a general set of domain specific data. So what I mean by that, let's say a computer science, right, just for the sake of our recommendation, I'm going to train that on a bunch of textbooks, stuff from stack overflow. Whatever the hell it might be, I don't know, maybe I shouldn't be doing it on stack overflow, IP issues, but we won't get into that. But let's say I've trained it on a general set of publicly available computer science data. But when you're saying you want to get into the fine tuning, this might be where you might want to turn it in the direction of a specific use case. So let's say I was wanting app development, right, for building mobile phone apps, right? And I might want to put certain weightings on certain data sources to do with iOS or Android development. Is that is, or am I wrong? Yeah, so the way that I would usually describe it would be like it's the, for a foundational model or a base model, the knowledge that it has is broad but not very deep, right? So you're talking like an analogy in real life would be perhaps like a GP, right? Like a general practitioner. So a doctor that you go to, they kind of know a little bit about everything. But if you had like a serious skin condition, you'd go to a dermatologist, you'd go to a specialist, right? So the, excuse me, the fine tuning aspect would be that you further, you have like your vast data sets, which which are scraped from the, from the neck. You got wiki commons, you got all of like the kind of like de facto kind of big massive data sets out there and a lot of proprietary data sets too, which like open alien, so on so forth have that their models trained on that they won't or haven't those fire released. Then to kind of I suppose specialize your model, that's when you get you aggregate a more, like a data set that's like definitely smaller but much more specialized. So you can get like a data set that's created, say from like our Cali Vex, which would be all the like papers, PhD papers or whatever kind of papers like surrounding computer science or like biology or whatever. And then you just, you like further train the models weights and biases on that specific data set. So that becomes now a general, not a general model, but a more specialized model. So like it knows bits about like Python code or it knows about, it knows about like biology or any of the above because it's like, it's, it's gun through extra schooling. Yeah, okay. But that makes a lot of sense. I can see how that would work. So you get this base knowledge, then you fine tune it into a more specific area such that it becomes useful. Is that it would be fair enough useful in a particular context. Okay. So where does it go from there? So from there, it kind of depends on there's a few other bits and pieces you can do. Like there's some new, I suppose techniques, I mean, microtech, it really might not, depending on which way you want this to go. But where you can add like a femoral data afterwards because like the actual, this whole process of as, as you like might guess, it's very compute intensive, right? Yeah. It's, it's not a leave, leave the laptop on for two minutes and, you know, and it's done. We're talking about like months and months and months of like multiple clusters, cluster of servers, essentially, running with the latest GPU technology and just add 100% capacity for months on end. Right. So it's not trivial thing to put together. So if you want to train a base model, there's only a handful of kind of companies or institutions, I always say, I always say out there that actually have the capacity or the expertise and the knowledge to do so. And then furthermore, it like to find tune it is, is much easier. So you could actually find tune it yourself. So what we're seeing, I suppose, is a shift from you have, like there's a ton of open-source base models out there and then anyone can kind of like take one of those and find tune for their specific use cases. They can then further augment that by not, let's say, fine tuning, by adding, like I'm turning it into a rag application. So that would be like the addition of a vector database with which you can essentially just like, upload any kind of, or upload, you could like, throw any kind of smaller data set that you want to throw at it, vector like encode that into vectors, stick it in the database, and then the LLM can query that database. I can do like your regular code querying any kind of database. So you say that's a, what's rag? Retrieval augmented generation. Oh, it's okay. It's another acronym in the world of, of IT. Yeah, I think we know the one we can throw on the heap. Okay, so and then from there, so it's weird as a goal, then you have to expose it to the user. Yeah, so it's absolutely useless unless you can actually go ahead and like query it, right? Like you have your model, you have to put in all this money on training it and all the data science expertise. You need to actually stick it on a server behind an endpoint and, uh, create it from there. Yeah, and then from there, like just hit it with prompts, depending on the kind of model it is, if it's an LLM, you just prompt it using natural language, um, all the models out there, you need to, you need to prompt with like, um, JSON key value pairs or whatever the actual model has been trained to do. LLMs are kind of the flavor of the month, shall we say? So you either, there are enough, let's say if you're taking a stable diffusion or mid-journey, somebody's like image generation models, they will be a combination of models. You'll have an LLM at the front and you would hit it with a regular natural language prompt. And then that will, um, like it's kind of chain together with the actual image generation of, uh, aspect, but the further model somewhere down the line in a little chain, um, and then that will spit out your, you know, glorious image of Joe Biden scratching decks or whatever you want to see. Yeah, that's a, that's a very tame one, I'd say. Yeah. So you know what you mean? Right. So going from there, then Mark, what would you say is the next step? The next step? Or sorry, well, let me ask you, uh, let me rephrase that question then because this is the way that your talk went and maybe we could bring it back to that, which is, so we're talking about these data sets are absolutely huge, right? But you did a demonstration on, uh, on stage in front of the, the students where you downloaded a local, uh, a local language model was running off of your laptop. Uh, it wasn't connected to the internet and you were able to query it and it answered questions. Now, if we're talking about data sets that are, uh, mohosef, okay, to use the, the parlance of our times, how do you get massive data sets to run locally on your, uh, your MacBook? So, um, with these, uh, with the software I used, LM Studio, um, it's connected through Hugging Face Hub, which is a, kind of like a repository of over, I think, 350,000, uh, open source models, which you're just free to use. And from there, uh, data scientists can quantize those really large, uh, models and make them smaller, um, at the cost of, say, efficiency and maybe accuracy as well. And from there, using a tool like LM Studio or another one, oh, oh, Lama, you can run these models locally on your own hardware, like I did in the presentation earlier. So you talked about quantization about making the smaller, how much smaller can they be? I'm trying to get a sense of scale. So you talked about, uh, Peter Bites before, are you down to terabytes, gigabytes? What are you into? So there's a distinction to be made from the size of the data set and the model up pops out at the end of the data set, right? So your, your data set might be 250 petabytes worth of data, but your model could be 150 gigs or 200 gigs, right? Like, that's still a substantial size to like, especially finding kind of hardware to stick into memory. Hmm. But, um, in general, I suppose if you see like, I'm not enough, a lot of open source, uh, also depends on the actual size of the model. So you have like seven billion parameters, you have like 38 billion parameters, you have 50, 70 billion parameters. I, I'm sorry, I'm just making a jump here in my head. Yeah. So when you download this because I don't understand that, right? So when you say you're downloading the model, the model is a, uh, already trained on it, you don't need to download the data with the model. No, so no. No, okay. So this is the jump. I'm learning along the way. Thank you. Right. So explain that to me then. So the model is its own self-contained trained data set. Is it? It's, it's essentially, if you want to look at it, it's like a compression of the data set. That's it. Okay. Yes. But it doesn't have all the data behind it. It just has its trained version of it. Yes. It's own, it's an anti-clink or whatever way it's done. Exactly. Like, like, you, you, you don't have access right now in your head to every single memory that, that, that, you know, that you have ever lived through. Yeah. At the same time, the way that you think and the kind of knowledge that you do have, it has been learned and influenced from everything that you've done up until your life at this point. It's the exact same thing. It's a compression. And like, it's lossy, right? Like, yeah. Some people have like photographic memories. Other people like me have terrible memories. So like, you, you can't remember everything. But the model in itself is going to be like a much smaller size. Now, much smaller size is still larger. Like your regular kind of base model will take like a 7 billion parameter model, which is kind of your regular Joe Soap model. There's still going to be upwards, up towards 100 gigs in size. So like to run that model locally, you'll have to do some black magic, and I'm not going to say that in the next word because it's a profanity. No, but you can say that. It's fine. But 100 gigs is within the REMS of possibility. But it's not going to be on your regular workday laptop, shall we say. But it could easily be managed by a local machine if you wanted to set it off for that dedicated task. So that would be like a kind of regular, so a regular 7 billion model, right? Yeah. What Mark was alluding to earlier was like quantization, did he say that? Yeah, you use that word. I want to go back to that. But if you just explain that again, because I think that's the key phrase, isn't it? Yeah. So quantization is the, like, during the flying tuning stage, you can, I don't know how technical to get here, but go first. So you have like 64 bit, right? Yeah. 32 bit, 8 bit, right? Definition of like integers, essentially, which is exactly what it is. So you can bring instead of, instead of training the model with like 64 bit precision, you train it with 32 bit, or you train it with 8 bit, or you train it with like four or two. If you train it with two, you could end up with a model that's like two gigs in size, right? And you're most Apple watches right now. Now, the output would be atrocious, but probably better than Siri, but, but, but, yeah, no, in general, if you quantize down to, depending on the size of the model that you started out with from the start and the actual architecture of the model, how good it actually is, like whatever job the data scientists who put the model together did, the quantization, like, either makes it like completely unusable if you're bringing it down to low, or it brings it down to a usable level. So like the newest, a good example of this, I suppose, like, I'm not messing around with it, but the newest Samsung Galaxy S, whatever flavor the monitor is now, they have Google's new Gem and I model, which would be like a core. It's on board, isn't it, yeah, yeah, yeah. I'm inbuilt in, you know, like I said, I haven't used it. It might be terrible, it might be like fantastic, I don't know, but that's essentially it, right? You're taking a model that's absolutely massive. You're just making a smaller and smaller and smaller. And like the technology and tooling in around this is getting, it's like on a day-to-day basis, it's, it's, there's new stuff coming out every five minutes. There's a new, there's a new research paper that like brings forward the field by what would have seemed like a couple of years ago that would be like 10 years, you know, and then all of a sudden it's like three or four papers and all of a sudden we've got like a whole new field of AI, that's, you know, that's it that I should be able to run and I'm on your phone like, okay, so I'm going to try and come up with an analogy here and because I like analogies because it helps me to get complicated things clear in my head. So and correct me if and when I'm wrong. So let's say there's a data set, right? And I'm going to use a human example and let's say the data set are the complete works, all the complete Sherlock Holmes stories as written by Arthur Conan Doyle, there's a number of novels and there's a number of short stories, right? And I read these actually when I was much younger I read all of them, okay? I have an internalized memory of all of those stories and all of those novels and I know the general gist of the adventures of Sherlock Holmes and Dr Watson and etc etc. And I might remember some specific details. However, if you ask me what happened in this particular story, I won't remember that and I will have to go back to the original data set. So my internalized version of those stories is equivalent to a model. Yes. But that be right. But the data set is all of the original stories in their original form and they're off out there and then some and yeah, yeah, sorry, yeah, but but but but but we're just talking about in this particular context, shall we say, right? A quantaise version of that would be, let's say, it's because it's not the same data set. You can't I can't transfer my memories into you or your memories into me. But if over time I forget some of that stuff, I still might have some of the general gist of it, but I'm losing some of the fidelity of the detail as to what happened in a study and scarlet, for example. It is that ish. Be more akin to to you being you want to talk me so it'll be more akin to you being drunk once you read it, right? Okay. So you have less of a grasp of actually getting in the full information because you're just your like senses are impaired or whatever, you know, maybe drunk is like drunk is probably a good example, right? So like, yeah, you're not working to your optimal capacity, right? So it's not even like even asking, let's say I'm sober, I remember all of this. I have a feed of points and a few whiskies or whatever it is. You might I still have that knowledge here, but it's fuzzy. It's fuzzy or no, no, no, you had a feed of points, no, you started reading it and how much do you remember once you have a feed of points? Okay. So it's like it's a slight subtle difference. Yeah. Okay. I think I get you. But it's like, it's, you know, you can probably flake through that data set like after a feed of points like with relatively little, you know, like hardship for yourself, because you're like, you don't really care, right? You have a few of points to like, oh yeah, whatever. But the quality isn't the right. Yeah, but the quality isn't the right. Yeah, yeah. You could probably go through more of it, much quicker rate, but like I said, like the actual like what you take in is not there. Okay. So the quantization then could be, okay, I'm just going to skim this as opposed to do a deep reading of it. Yeah. So to take the alcohol out of it, and Mark, you're reading just because I don't want everyone to think that we're getting people drunk. Okay. Because I know you're a stereotype. Yeah. But it's so that you're agreeing that's kind of the way to think about the model and the quantization of the data set. Yeah. I'd reckon that's a good analogy. Another one I could think of is kind of like say as a child, you know, you're learning, but it wouldn't be the level of comprehension would not be the same as I'm an adult now. I can read a piece of information and I could recall back to you pretty well compared to how a child could a comprehend it and be give it back out to me. Okay. It's a much better example. Yeah. Or even like these kids aren't drunk. Yeah. That's better. And we're not bringing alcohol into it. Okay. I get that idea. Yeah. That makes a lot of sense. Okay. Right. So I'm going to ask you a kind of a bigger question, Mark. Why would I want to do this? Why would I want to run a local large language model on my machine? If it's going to be computationally intensive, I'm probably going to have to download a large data set. Why on earth would I want to do that? Can I not just use chat GPT? Well, I mean, the way I look at it is, well, they're not exactly huge either. It's like, I think the one I used today was four bit and it was also, I think, three gigabytes in size. And that ran relatively efficiently on my laptop. It might be different for other people, but I know on my computer at home, it is actually a struggle to try and run any of these large language models. But why someone would go and interact with running these locally, say, there's data privacy, everything I would throw toward that model is all kept locally. As again, in the demo, I was disconnected from the internet and it was able to dispel that information to me when I was asking it to write a Python script. And not only that, it's like, so when you use chat GPT, it's always constantly learning off of you. And so that information isn't exactly yours anymore. It's ours, kind of a situation. In a public model, or a model out there that we don't necessarily know how it works. I wouldn't even say it's ours, I'd say it's like open-hills. It's open-hills, it's open-hills, it's a bit of a waste of it. It's a bit of a waste of it here. There's a reason where the company valuation is where it is right now. And realistically, it's probably Microsoft, so I don't know. So again, if I can go back to an analogy, let's say, for example, I wanted to set up an LLM for SETU. That all you could use. Let's just say for me, and we could extrapolate it out to the organization. And let's say I had a lovely server sit in up my office, which I don't. I have an aging MacBook that I really need to update because the fans are going nuts on it. But let's say I had a decent server up in the room. And I wanted to get one of these local LLMs, let's say LM studio. I downloaded a base model, but then I wanted to train that with or tune that with some student information, which is private. I don't want it being publicly available. It might be things like grades, progress, but equally, it might be sensitive information about like maybe a student has a medical issue or something like that. Blah, blah, blah. But I want to be able to query it because I forget that that student has a medical issue and two years later they come to me with something and I just able to query chat, GPT, not chat GPT, my LM studio and oh yeah, they have this issue or whatever. Is that the kind of, to me, that sounds like a reasonable use case. Does that sound fair? Yeah, 100% and a tool like LM studio actually offers a kind of like a server endpoint. So I went through the example. You can have it running on your own hardware, LM studio with them running a local inference server. And from there, you can ping it just like with the chat GPT API. I think it gives you some examples too. Like it'll have a pathonic way of kind of a talking to the LLM through the server all locally. Okay. So you can imagine how that could be extrapolated out to any organization, any business. Right. So let's get into some of the kind of the questions that students ask. I can see why this would be beneficial. I can see I can see this as being a potential future for the way some generative AI might evolve over the coming years. Some of the student issues were kind of conspiracy theories. You know, like what if they do this, what if they do that? What if choice to rob a bank? Okay, which we won't get into. But there were issues about regulation. And that was brought up. And where do you think this is going? Because I know there's a bit of a scramble. There's a kind of an EU model or an EU viewpoint. There's a US model. There's a kind of a British model that is somewhere. I think the Chinese have a different model, are different view on this. So where do you think the regulation stands on this type of stuff? So to me, it's more deep into this. For, so that regulation is like a, it's a tough one, right? Because like what point do you actually regulate it right now? There, like I kind of mentioned there, there are only a handful of like institutions that can actually do the training and like have the expertise to create these models like these base models. So do we kind of try and regulate these models? Are these institutions organizations? Some of which are like the, you know, like your Google's met is the world, mistral, which is a company based in base in Paris. The third, but then there's other companies as well. Like Falcon is one of the largest and top performing models out there. And that is a Saudi state. But paid for a model. So like for a regulation to occur, we kind of need to know where you regulated it. Like what stage of the model life cycle do you regulate? Do you regulate the actual creation of the data set? Do you regulate what the algorithm is? Do you like set some kind of a benchmark as to like if the model performs above this benchmark, then you have to have x, y and z tests or surround it. Like it's really, really difficult to, to a kind of like figure out a kind of regulation surrounding it. But be if we do go like, like the push towards regulation, which yeah, it should definitely be regulated. But every, then we need to have everybody sign up for it. So like every country needs to, and every country and every organization within every country, which like, if you look at the world right now, like not every country kind of gets on with each other, right? So like trying to get them all to, um, to coalesce around a certain set of standards or ideas, which like in an industry that's worth this amount of money as it is right now. How are we going to do that? I don't know. We can hardly even agree on. What's the best coffee at the minute? Like it. So it's, it's a difficult one. The, the, the only kind of model that I know of, um, that is, like I think I mentioned to you earlier, Rob is like Adobe's Firefly image generation model is the data set that's trained on. And that's created from is 100% owned by Adobe. They have like all of the, um, all of the IP surrounding the data set. So like that's probably, if I was the waiter, that's probably where it would go. Would be like data set creation. Um, and this is why a ton of companies like Reddit, for example, uh, all of us, uh, they've turned off their APIs. Um, they have data sets, which they were just like, let an open to the public. Anybody could could use like Reddit API and, and, and, and come up with anything. But now this is like gold, right? This is like a gold mine. So now you're going to have you, these companies are starting to realize that you're going to have to pay for, for access to these data sets because just the, the, yeah, I suppose the amount of money that's, that's involved. So regulation is a sticky, sticky, you get into a whole thing about intellectual property and how that's valued and protected. And I don't know, I'm giving, get into, I mean, I'll use my Sherlock Holmes example. Right. Let's just even say I want to defeat LLLM. The works of Sherlock Holmes are all in the public domain, okay? Right? But let's say I wanted to defeat it. The works of Donald Ryan, who's still an author today. And he wrote me a story in the style of Donald Ryan. And I would have a huge ethical issue with that, you know, or, or whoever, just, or Megan Nolan, because I'm reading her at the moment. Right. Like, so taking their novels, which is their intellectual property, feeding that into some sort of an LLM to then create something in that style. That's an issue. It is an issue, but also I was just going to play Devils Avocados. Or do go first. Yeah, go for it. Um, is it that different to what we do as human beings? Like, so you're, you're, like sitting there in New York after an end, a couple of authors that you've read, right? So if you turn around in a couple of years time and Rob decides to, to write a novel based on whatever, like your influences are going to be from the stuff that you have read already. Right. So like, how different it, I don't understand obviously there's like the human biology pirate and the neanderers, you know, the machine and algorithm or whatever that we created, but there are similarities there, right? They're like, they're most definitely. I suppose it's all right. Again, playing Devils Avocados here, right? And I'm going to use music as an example, right? Because it's a bit easier to get my head around the song and have a bit more experience with this. Like, the world of music, pop music in general is everything is derived from everything else. Okay. But there's a kind of progression. But generally speaking, somebody has to have some sort of talent or work to create a piece of music, whereas if I use one of the models and create a song in the style of Drake or there's the Johnny Cash one that's going around. Yeah. But Annie Foul can do that. And I'm not seeing the artistry in that. And just from myself personally, do I really want to listen to that? Am I going to get any deeper understanding of the human condition from something that the machine has created? Maybe I will. Maybe I don't. But yeah, I know. I know what you get now. It's a philosophical question. It's going to be. Yeah. So even from that perspective as well. I don't know if it's currently possible. I doubt it is to just say, give me a number one best-selling song and the AI comes up with everything. You're going to need to give me a song based on with this amount of buyers per beat or whatever be per buyer. And then these are the lyrics that I want you to put in the style of X, Y and Z. I think actually maybe this sort of person would do a better job at the local harmonies or so on and so forth. You're going to have to really construct it together. I know a thing or two about music composition, which I obviously don't. By my shot, the example there. But there is a certain creative aspect to it. Yeah. I'm not saying that this is the way that I like it to go either. I have no musical talent and I thoroughly enjoy listening to people's music. But that's not to say that I like at some stage in my life. Do I think like here? Look, I can't sing right? My voice sounds like two bags of cat's been better together. But like if I was to like if I would like it still, you know, I might fancy like myself as a lyricist, right? I just don't have like I'm just not fortunate enough to be born with, you know, silky tones of, I don't know, you know, BB King. So like, I might just think this advantage. I'll never create a BB King song because, you know, I just don't have that. No, but you might create a Dmitri Saradakas song. Yeah, yeah. Do you know what I mean? Like, I know he's going to want to listen to that. How do you know? I do. You know, I saw these are these are the bigger questions. I don't know what the answer is. And I'm not suggesting that I do what my instinct is kind of veering away. Certainly for creative works, creative artistic works. Not equally at the other time, like, I mean, I spend a huge amount of my time doing what I call a spreadsheet work, which is boring, mean, you know, tasks. If I can use an AI that might automate or speed that up, I don't have a problem with that because I don't, but if I want to create something, I don't necessarily sure I want to use it. Now, we're gone off topic from talk about local LLM's there, but you know what I mean? So if we were to, if I was to ask you a question, then Rob, would you have an issue with, with an LLM coming up with novel science? If it was like to, like to, to, to, I don't know, create a new drug, let's say, for example, on the sudden we had an LLM that came up with some novel science based on previous works of anger that were carried out and done by humans. But the LLM, now we have a cure of cancer. Is that, do you, do you have an issue with that creativity? I'm sure you don't. No, no, no, no, no, no, I wouldn't, but you see, it would still have to be tested. Oh, I bet. You wouldn't just go, okay, I'll fully go. That would then be taken, go through rigorous scientific process. Buh, buh, buh, buh, buh. Oh, no, no, I can totally see that. But that's, that's an inference of, of maybe multiple data points to, hey, does this pattern here that maybe you haven't seen? Have a look at this. But it would have to be queried as well because this is the other thing that I think is misunderstood about, about generative AI is that they're not, they're not kind of self-starting, generally speaking. Like they, they don't just take control and do things. Not yet. Like, when you boil it down, they're just a probabilistic machine. That's all they are, right? It's like, I'm predicting, like an LLM just predicts the, the probability that this next word is like fits after the last word. That's, that's all it is. Like, when you, when you, when you, when you take any of those LLMs, that's 100% what they are. There's a phrase used, a stochastic parish. I love it. I absolutely love it. I probably used it at the end of the time. Um, so unconscious of time. So I'm going to kind of feed you one last situation. Again, this is, I'm not trying to be a LLMist or anything like that, but I'm, I'm mindful of that we're going to be able to do that. we as humans need to have our own internalized set of information and I suppose critical skills which we develop over time and I don't just want to outsource everything to an LLM and say you don't need to learn anything going forward because first off I'm an educator and then put myself in a business but also I believe in kind of that you need to have your own critical faculties otherwise you were just this stochastic power. So I want to give you an example and this kind of will tie into the demo that you did Mark in the auditorium earlier this morning. So I had a first year class last semester, computer systems won. So you both would have done a version of this a number of years ago when you were here. So we did a lab, a Gen A lab. We were just kind of mucking about with some of the tools, you know trying to show them in a more, I don't want to say controlled environment but in a more directed environment. And one of the tasks that we did with these first year students was get them to create a rock paper scissors game in Python using whichever Gen A I tool they wanted to use. Most of them use chat Gbt if you have been used Bing which is called coal pilot now but you know I think that I don't think anyone else use that in else they might have but that's what they did. Okay and for most of them what there were a couple of things that I took most of them it worked they fired it into into one of the Python tools and they were all using trinkets the online one because there was not needed to be installed and most of them it works but there were two things that I kind of took away from it that were interesting. One was the variety of the results that they got even from the same engine. So same model back in both of those. So I asked chat Gbt create me rock paper scissors you ask the same question and we get different results ever so slightly different results. So that's there's a certain randomness there that I think is kind of interesting okay. And the second one was in I had that I did that lab twice with two sets of first years and in both classes and there's a bit I'm going to say about 25 people in each room it might have been 22 it might have been 26 I can't remember somewhere around that there were a number of verses that didn't work and as I looked at the code I went okay well I can see where that works and actually in in two of the examples I remember the issue was there was a stray exclamation mark which is legal syntax because it's a exclamation mark is often used as a not a logical not operator okay and it was in the wrong place which was causing it to not work. There was a couple of them where it was not compiling and I think I think there were lines common today which that shouldn't have been common today but again that's just is it a hash simple on point and I think yeah there was a stray hash which again is legal syntax but it's just in the wrong place and from a probability point of view it's put this character in the wrong place so I just thought that was kind of an interesting one because I was able to solve it by looking at it not because I'm a particularly good programmer but I have internalized certain coding rules I know that's your issue there we'll fix that and there's certainly it works but I'd wonder if this is the first years who didn't know this they're just looking at not those work yeah you know so I just kind of interested get your comments on that because like Mark you generated some Python based on your personal your internal local version of LM studio and that's great because you know how to do that yeah but if you didn't it's just gobbling hook. Yeah exactly and it kind of ties into why you kind of already need to have an idea of what you want to ask so like sure I I know I know my way around Python and I know exactly how to ask the question to the AI model I want you to generate this Python using these libraries exactly whatever whatever I'm looking for but again it might not give me a correct answer it might give me something that kind of half works but then I have to go in and I have to I have to go fix it it's like multiple times I've tried to I tried to automate certain things that I'd normally do like I moving files from one directory to like a server and everything and I'd be I'd be looking to implement certain things so I'd ask hey how would I do this and then give me a answer back might not work but then I could go in and I could play around with it fix it further but it gets you off that blank page. Yeah it's kind of like it's in it's instead of starting at like zero you're starting at I guess halfway there almost yeah and it's a better start point when you're only doing it just to do a mean you'll task anyway to me tree your thoughts on that yeah 100% like you can't just they do hallucinate and they do who hallucinate quite widely as well hallucination being the term used for like when a model just lies right it's a nicer term than lying but friendly yeah friendly yeah friendly yeah but yeah at the same time you have to like be marked set it really really well there when you kind of just instead of starting at a blank page you start to like 30 40% 50% and you have to be able to like to critically look at what the model has put out you have to have like in the in the case of like the code example you have to have your own knowledge of like the kind of fundamental coding principles or whatever so that you can actually spot what's going on and you can go off and correct it the other thing that we didn't touch on earlier and if we continue down the path of just trusting models left right and center whatever their output is would be a model collapse and we don't have that much time so I'm not going to go into too much detail but no but you can mention what it is because it's a nice thing to think about I'm not trying to be a doomsayer but it is I know what you're getting out here and oh yeah yeah so if we were if you were to just trust whatever output of the model without like so if you didn't do any of your own kind of like research or like you didn't do any of your own learning shall we say you didn't know how to code whatsoever you just talk whatever the model is set for granted and did this over and over and over again and like did okay let's bring the pack to what actually model collapses so like the models themselves are trained from like as we've already mentioned like vast swaths of data that's out there if human being if like human being stop creating the data to feed these models so say for example instead of me going to like like it or not right now the world the internet is run by Google's add algorithm realistically right so if you have a website you get paid to keep your website up online by Google ads essentially right or meta ads but like predominantly but both those companies so if if someone instead of googling something and we know how how crappy Google has got like answering your specific queries it just sends you two reams and reams of ads right now but if it if we just asked it in LLM that and the LLM brought you the answer straight away then you don't get the ad revenue from your website for the for the clicks on your ad website so what's the motivation for you to go off and continue to like create your your you know helpful website about like um Python code or about whatever it is that is your discipline yeah there is no kind of monetary solution out there so in a world then when we're getting back to not being able to trust everything that that in LLM comes out with in a world where estimations are somewhere between anywhere from 40 to 70% of it could be anything realistically of the hallucinations are just are just that right there just like down in order lies if human beings stop producing quality information that's peer reviewed or that's you know that you just like actually check you check yourself and we instead like that devolve what I would use into a world where we only have data that's flawed inherently because of the hallucinations that the LLM's have and this is the only data that's being created then we end up in a scenario where further models can only be trained on this already flawed data like in a downward spiral um we all live in a mad wax world I know it's a so this kind of stagnation if we're if we're humans not still creating and moving forward everything just stagnates yeah we it it it actually be worse and stagnating I would say because we're not actually moving the needle forward yeah what we've done like like my mother was around when there was no electricity right yeah like she's not that old I'm not gonna say oh she is cuz she killed me for yeah for like in the grander scheme of things like human human kind of like humanity has moved at a really really rapid pace and that's true like innovation it's true um research it's true like I'm I suppose if you if you take like all of science you could um together you could could look at it as like one single organism right like everybody like the whole thing working together to collectively pull knowledge forward if we stop doing all of that because we're just relying on an LLM to give us the answer based off like 70% correct of what we we already know to date then we stop right and we don't just stop the 30% of the factuals are the non-factual stuff enters the data enters the new data sets and it just drags the overall quality down so not only do we not get better but we actually like progress progressively get worse yeah or to be like asking answering questions like if we if we're ever if our our computing knowledge stopped in 2004 yeah and answering questions 20 years later based on data from 2004 or 1984 or whatever it is you're having yourself and again that's computing we can play that to any discipline and no matter what it is yeah but also with 30% of that data being wrong actually wrong yeah so then like the next stage then if you already started a 30% and you get 30% of the of the next batch so quickly like you're just down like it's compound interest yes yes I don't know the track or market just excellent look um Demi G Mark thank you so much for your time today you've been absolutely fabulous and Mark thank you Thank you for coming in and chatting to some of the current students, even though you're just an ex-student, barely a wet week yourself. Your character's married, you will definitely be back. If somebody wants to define their more about you, the kind of work that you're doing, or a hitchhops for, I don't know, an interesting proposal, what's the best place to get you? For me anyway, it'd be linked in 100% the easiest way to get me, just Mark Campbell. Mark Campbell on LinkedIn. Mark Campbell on LinkedIn, yeah. Okay. Make sure you're yourself. Make sure you're safe as Dimitri Sardakas. If you can't spell it, yeah, you probably can't. I put links in the show notes for this year as well. Yes, so Dimitri Sardakas on LinkedIn as well. Yeah, I don't do anything with the rest of the social media. And your mental health, thanks. Okay, Lads, thank you so much. Thanks for your time and we talked to you soon. Rob, thanks so much. Thanks for having us. Pleasure. Thanks for listening to this episode of The Machine. If you liked it, please leave us a glowing review wherever you listen to your podcasts. It helps us to grow the audience and it also makes me feel good. Will we have another episode tomorrow? Who knows? I'll see what I can do. In the meantime, you should find some links in the show notes to where you can find out more about Dimitri and Mark and it also includes some links to some of the tools that they reference during the conversation. Okay, shall I go full? [BLANK_AUDIO]

Podcast Summary

Key Points:

  1. The podcast episode features Red Hat engineers discussing how to build custom AI models, emphasizing the critical role of high-quality, specialized datasets.
  2. The process involves training a foundational model on broad data, then fine-tuning it for specific use cases, with tools like OpenShift AI managing the model lifecycle.
  3. Models can be run locally using quantized, smaller versions from repositories like Hugging Face, though training requires significant computational resources.
  4. Ethical challenges in AI, such as bias in datasets, are acknowledged as complex due to the vast scale of data involved.

Summary:

This podcast episode from Southeast Technological University Ireland features host Rob O'Connor interviewing Red Hat software engineers Mark Campbell and Dimitris Saradakis. The discussion centers on building custom AI and large language models (LLMs). The engineers explain that creating a custom LLM starts with a high-quality, domain-specific dataset, as the model's output is directly influenced by its training data.

They outline the process: first training a broad foundational model, then fine-tuning it for specialized tasks, similar to a general practitioner versus a specialist doctor. The conversation highlights Red Hat's OpenShift AI as a platform for managing this entire model lifecycle. A key demonstration showed how quantized models from open-source repositories can run locally on personal hardware, making AI more accessible, though initial training remains computationally intensive.

The dialogue also touches on the significant ethical and regulatory challenges in AI, particularly around bias, noting the difficulty of curating perfect datasets due to their immense scale and inherent human biases.

FAQs

OpenShift AI is a platform from Red Hat that enables data scientists to train, manage, and deploy AI models in cloud, on-premise, or disconnected environments, supporting the entire model lifecycle from data preparation to inference.

To create a custom LLM, start with a high-quality dataset, then train a foundational model on it, and fine-tune it for specific use cases to enhance its expertise in particular domains.

The dataset provides the knowledge for the AI model; if the dataset is biased or low-quality, the model's outputs will be similarly flawed, emphasizing the importance of clean, relevant data.

A foundational model has broad but shallow knowledge, like a general practitioner, while a fine-tuned model is specialized for specific tasks, similar to a specialist doctor, through additional training on targeted data.

Large models can be quantized—reduced in size at the cost of some efficiency or accuracy—using tools like LM Studio, allowing them to run locally on hardware by downloading pre-trained, compressed versions from repositories like Hugging Face Hub.

Retrieval-augmented generation (RAG) enhances AI models by integrating a vector database, enabling the model to query specific, up-to-date information for more accurate and context-aware responses.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.