The discussion centers on AI safety, highlighting its importance as AI becomes ubiquitous and powerful. Key concerns include the risk of AI being weaponized by malicious actors, the inherent opacity of how AI models make decisions (the "black box" problem), and vulnerabilities like prompt injection that could lead AI agents to execute harmful commands. On a macro level, the competitive AI race among companies and nations prioritizes rapid development over safety, with insufficient international governance to enforce pauses or regulations. While major AI labs have safety teams and contingency plans, their frameworks are inconsistent and often overshadowed by market pressures. Researchers are increasingly aware of the potential for AI to achieve superintelligence, which, if deceptive or misaligned, could pose existential risks. However, current AI lacks the autonomy for such catastrophic outcomes. The conversation concludes by noting parallels to issues like climate change, emphasizing the need for proactive safety research and sustainable development to mitigate long-term dangers.
[MUSIC] Welcome back to another episode on what's going on. We're back after the semester break. And today, we've invited a representative from the NUS AI Safety Space, Kai-Ju, to give a sharing with us. Kai-Ju, tell me about yourself. >> Right. Hi, everyone. Kai-Ju. I'm currently a second year studying math and computer science in-- >> Wow, math and university science. >> Yeah. >> That's a pretty tough call, VIC. >> I suppose it is. Yeah. >> Okay, okay. Yeah. So in my spare time, I like to work on AI safety. I can only do some research on the site. Yeah, and nice to be here to share if I've run. >> Yeah, great to have you here. In my spare time, I play basketball. So it's quite different from what you do with your spare time. But actually, Kai-Ju is a friend of mine, Wayne. So Wayne, if you're listening, thanks for linking us up. So today, we're going to talk about AI safety. And we all heard about AI. It's like, you know, chat GPT. We use it for various purposes. We use any AI is like all around us. So why is it AI and safety? Why do-- what is AI safety, Kai-Ju? >> Right. So like, as you mentioned, AI is all around us now. It's kind of a tool that's permeating into every aspect of our lives. And it's seen like widespread global usage. So if you think about it, AI is kind of like a technology. And when we think about other technologies in the past, for example, like airplanes, we will be concerned about the safety of these tools before we use them. >> Yeah, fair enough. Yeah. Want to make sure the airplane lands. Yeah, yeah. So for aviation, it's really obvious, right? Like, what are the risks involved when you take a flight? But for AI, maybe the risks are less obvious. But because we give AI so much power in our lives, there's definitely an element of safety to be considered as well. >> Right. So by safety, what do you mean? What kind of issues are we talking about here? >> So for example, one issue that comes to mind is that AI is getting increasingly powerful. So you've probably seen the news. AI can do like international, math, or ambient questions. AI can like solve the protein folding problem. So all these are like-- >> That's a tough problem to solve. I've never heard of that problem. But it sounds pretty tough. >> Well known in the STEM field. Like the founder of Google DeepMind, like Demi's HaSafis, he actually got a Nobel Prize for that. >> Well, yeah. >> So-- >> Yeah, I solved the same problem he did. He was one of the people who was credited with building the AI that could solve the problem. >> Okay. >> Yeah. >> Okay. >> Yeah. >> So these are what we call the capabilities of an AI. >> Right. >> And one issue that we are concerned with in AI safety is what happens if these capabilities fall into the wrong hands. >> Okay. >> So if they're like terrorists-- >> Right, terrorists groups. >> Okay. >> State actors who want to use this AI to commit harm on other people, then aren't we enabling them? >> Yeah. >> Like by giving these kind of access to-- it's a tool, it's a weapon essentially. >> Yeah. >> All right. Okay, I see. So that's the main worry about AI safety. Falling into the wrong hands. >> That's one worry. >> Okay, what are the others? >> So another worry is a little bit more esoteric. >> Okay, so it has to do with how AI was trained. >> Right. >> So how we developed AI as a tool. So maybe I'll give a short explanation about this. >> Yeah. >> So how does AI work? >> Yeah. >> That would be great. >> So basically AI is sort of like a pattern recognition tool. How we train it is that we give AI a bunch of inputs of examples of what is right and what is wrong. And then afterwards based on like telling AI, "Oh, your output was right or your output was wrong." Then AI will go and slowly correct itself over a long period of time. And if we give it a lot of compute capabilities, a lot of GPUs. And after a lot of training steps, it will become like release smart. >> Okay. >> So all this is to say that actually like all the engineers that programmed this AI, they have no idea how it's solving all these problems. >> Right. We don't know how it's doing what it's doing. >> Yeah. So in a sense, it's a tool that's not crafted. It was like grown basically. >> I see. >> Like from what I understand about AI is that when I and her input like craft me shopping this for my grocery shop today, like all these letters to it, it's just ones and zeros, right? >> I suppose, yeah. >> And these ones and zeros go into a matrix. >> Yeah. >> And how does it get this information out to information that we actually find pretty good, you know? >> Yeah. So it's a question that if you ask researchers and scientists nowadays, like no one can give you a definitive answer. >> Right. >> Yeah. So like what exactly is going on in this brain? Like does it have a model of like you, the user, like what does the user want? What does the user like? Or is it just like copying some random patterns that it saw in its input data? >> Oh, okay. >> We don't really know. >> That's a problem. >> Yeah. >> So those matrices that you, like I mentioned, we don't know how it's being, I don't know, compound. >> Yeah. >> Is that what you're saying? >> We don't know what it means when like we can see the calculations being done, but it's quite opaque to us what these calculations mean. >> Okay. >> Yeah. >> So basically what I'm hearing is that we don't know what it's doing. >> Yeah. >> And we're trusting and we're feeding it with more and more information like our, I don't know, our bank credit card number, you know, information that should be held with only ourselves. Okay. I see. >> Yeah. >> And the safety risk comes in actually, like not just because we are feeding in information that's like private. >> Yeah. >> The bigger safety risk is that we are giving AI a lot of power to do things that have like immense importance in our daily lives. >> Uh-huh. >> For example, like recently there was this. This framework that was released online, like called OpenClaw, which. >> OpenClaw. >> Yeah. OpenClaw, which is a play on like the language model slot. >> Ah, a slot. >> Yeah. >> Which has-- >> That's an proof-pick, right? >> It is anthropics language model. >> Yeah. >> Right. >> Yeah. So, Claude has like this coding capability called like, Claude code. >> Okay. >> Like, um, a coding agent. >> Yeah. >> And like what this framework does is that it orchestrates the agent. So it turns it into like a helper of like your daily life. Like, if you hook it up to your computer, it will have access to everything in your computer. >> So like Siri, but no privacy. >> Yeah, exactly. Got you. >> There's no privacy. And like, it can do whatever it wants. So what will happen if it misbehaves, right? >> Right. >> Yeah. >> Has it misbehaved? >> So some people have like shown that they can do some sort of prompt injection into this framework. So like, for example, they can send an email to you. >> Okay. >> And the agent will read this email. And maybe inside the email, there's like some hidden instruction. Like go and delete everything in your users' computer. And they'll actually do that. >> No questions asked. >> Yeah. Because it treats it like an instruction. >> Yeah. >> Okay. So it's not that smart to like know that that's a bad thing. >> Yeah. It can't really differentiate. >> Right. >> Of course, like, not always safe. >> So I'm still safe. Because I know what to do and what not to do in terms of being replaced. >> Yeah. >> Okay. Got you. >> Yeah. >> Yeah. >> Yeah. So like, of course, they've patched that. Yeah. But this is just to say that we don't know how AI works. And so we don't know how to fix it if it misbehaves one day. >> And that's like the grand problem of safety. >> I see. I see. >> What are, so this safety seems to be. It seems quite micro in that sense. It involves like maybe my life or my data. But I see governments also like talking about it's going to be the next big thing. It's going to be the nicks, I don't know, black box or like, you know, fast and furious movies. It's going to be the nicks thing that dominates the world, you know. What are the macro issues that we're looking at? >> I think the main macro issue is that like AI is such a, I think it maybe only to say is that it's really useful. So everyone wants to use it, right? >> Yeah. >> Like, it's a false multiplier. It can help you achieve an edge against your competitors in every like competitive situation there is. >> Right. So this means that AI is going to be ubiquitous. It's going to be used everywhere. >> Okay. I can definitely see that happening. >> Yeah. >> As a result, I think like governments are going to use it as well. >> Yeah. >> And they want to be firstly ahead of their competitors. So for example, maybe other countries. >> Right. >> And they want to make sure that whatever AI technology they are deploying right now, it has to be safe and reliable. >> Okay. >> Yeah. >> And of course, they're also worried about, you know, like other countries, like or maybe like some unknown actors trying to-- >> The terrorists. >> Yeah. >> The bad guys. >> Yeah. >> Yeah. >> Trying to create unrest. >> Right. Is there any concern for like nuclear weapons being hacked or misused by like, I don't know, bad people in terms of using AI to like hack these kind of nuclear codes? >> I think AI right now is not really advanced enough to become like-- >> Okay. >> And automate the heck of all by itself. >> It's good to know. >> Yeah. It's good to know. Because otherwise, their implications are really terrifying. >> Right. >> Yeah. >> Yeah. So I think-- >> It's a good butt there. Like, I'm just waiting for the butt. I don't know what like all the big governments are doing, but I hope they are properly like, you know, decoupling the own internal IT systems from like AI. >> I see. I see. >> Okay. Well, in that case, before we wrap up this segment about what AI safety is, I want to ask you--so some recent studies have actually shown that the AI race is kind of like a prisonist that Emma, where like, you know, where all companies race to get it, nobody wins. But if everyone does it slowly, you know, more sustainably, maybe there's more benefit than the case that I mentioned earlier. What are the big companies out there like Google, Meta, and Frope Big, or like Open AI, which, you know, started this whole AI race, or how has AI actually started? What are they doing to, you know, mitigate the risk that you mentioned? >> Right. So actually, these big companies that you mentioned, like I would like to call them like Frontier AI Labs, basically AI Labs, okay? They do have like their own safety teams who are concerned with doing like research. So for example, safety research, right? For example, like, anthropic has their alignment science team, which does research into, you know, toys scenarios of like models being deceptive, models causing harm to the user, models being poison. >> Right. >> Like these researches being funded by them, but carried out by--we mean professors, or-- >> Yes. >> Or in-house researchers. >> Or in-house models. >> Yeah. >> Okay, I see. Got you. >> Yeah. >> Carry on. >> And on the grand scheme of things, they do have safety frameworks as well, which detail, for example, contingency plans for when AI grows too smart. >> Right. >> If AI grows too smart and we haven't actually like evaluated it for being safe, we will maybe pause our research for like X amounts of time before we are. >> Yeah, I'm genuinely sure that it's safe. >> Right. So that's the--that's what they have right now. >> Yeah, but-- >> But-- >> Yeah, go. >> Yeah, but of course, the safety frameworks are to varying levels of realism. >> Yeah. >> Okay. >> So some of them might just be--some companies have better safety frameworks than others. Yeah, some of the safety frameworks are really skimpy. Yeah. Right. But when you say that, if they detect some kind of like, I don't know, like, uncertainty with how the AI model works, do you say that they will stop the AI development? But I haven't seen any stoppage quite frankly. >> Yeah, because firstly, like, I don't think the big, the frontier AI labs have detected any concrete evidence of like models trying to scheme or like, be misaligned from their goals. >> Yeah. >> Like, you know, just kind of a inherent motivation to say that. No. >> Yeah, I mean, like, it's kind of profit driven companies. >> That's true. That's right. >> Yeah. >> That's why there are certain like nonprofit organizations out there, which try to work with these frontier AI labs and get access to their models before they are released to try and do some sort of evaluation. Yeah. But so far, they haven't found anything either. Okay, so like, everyone hasn't found anything, but there is a risk and that's why AI Safety comes in. Yes. Okay. Well, thanks so much for sharing and we'll be right back after the short break. And we are back from our break. You're listening to what's going on a podcast show by RadioPost, the Sound of NUS. So, what's the worst that we can expect if everything, I mean, that's supposed to go right, goes wrong in terms of AI? I think like, this might sound a little bit like alarmist, but maybe the worst that could happen, like in my mind, is really that like, everything does go wrong and like everyone dies. Okay. I mean, I mean, there's a bad question for like, okay, I mean, a lot of things can go wrong and yeah, a lot of things just like, you know, just like Matila, you know, yeah, that's my day. But I guess what I was trying to get at is that what is, if we don't do what we are supposed to do now, is if we don't conduct more research and develop it at a more sustainable pace, what are we going to expect to see? Right. So, what we are going to expect to see, I think, like what people can count on is that AI's will continue getting smarter. I think that's the trend that we've been seeing for like the past three years, right? Five years since the chatroom came out. Every year we can see that like models get smarter, they learn how to reason, they gain like multi-modal capabilities. Yeah. And each generation is smarter than the last. However, like, our ability to know what's going inside the model hasn't gotten better. So at some point, if the model learns to, you know, hide its intentions from us, like what if it learns to be deceptive, right? And we don't know whether or not it can learn to be deceptive because we don't even have a clear idea of what we're doing when we're training the model. Right. And once the model has deceptive intentions, like anything could happen. Right. Yeah. And a smart model. Anything, the doomsday of mankind, that could be possible. Yeah. Okay. Like, in the last few years. You feel war, I don't know. Yeah. Throwing things out. Yeah. It could be like, you know, manipulating politicians behind the scenes. Right. Right. I mean, you have already seen cases of AI manipulating people. You like, love scams, you know. Yeah. Like, AI telling your grandson or like, got captured demanding a ransom. Yeah. Yeah. You can definitely see being manipulative and, yeah. Okay. Yeah. Yeah. So that is the worst case scenario. Yeah. Right. Right. But you seem very optimistic about it. Like, you're smiling to me right now. And we're talking about, you know, the doomsday of mankind. But so what is this like? What is the general sentiment of, you know, the frontiers, the researchers in this space? I think a lot of the researchers. Like, they are in a rather unique situation because a lot of them are acutely aware of the challenges they are left. Okay. For AI to achieve some sort of like super intelligence, like AGI. Yeah. Right. Yeah. Even general super intelligence. General intelligence. Sorry. Yeah. But at the same time, a lot of them have also seen a lot of successes that have come over like the past three or five years. Right. So, I think there is a fair amount of researchers that still don't really believe that artificial intelligence will take off in a few years. But I think that the more and more of them are believing that it's possible for it to take off. Yeah. I see. And how do they feel? They feel optimistic, pessimistic. I think there's a range of people on this spectrum. Yeah. Like, regardless, I think despite how optimistic or pessimistic they feel. Right. I think they are much more driven by the sort of like competitive dynamics that characterize the AI space nowadays. Well, so are they worried? Yeah. I think some of them are quite worried. Okay. That doesn't stop them from like continuing to develop their AIs so that they can perhaps make their AIs faster, make their AIs like better than other people's AIs. Yeah. Well, then would it be a case where everyone is so driven about developing then, you know, you go to the point of no return? Yeah. Is there any checks that are like, you know, is the UN or any international organizations that is holding them back, you know, US I would say, but maybe not now. Yeah. So, there are talks to like bring about an international treaty for AI, right? Is there? Yeah. I think there's one which is convening in New Delhi this month. It would be like February 2026. Will you be there? Maybe the next one. Maybe the next one. We'll see you there. Okay. Yeah. So, what is the purpose of this talk and who's attending? I'm fairly certain it is sort of treaty that will be attended by like leaders of major countries. Right. It's for the governments to come together and try and hammer out some sort of guidelines towards like regulating the pace of AI and maybe even developing some sort of like triggers. Yeah. Triggers. Triggers to, you know, like halt or pause and development and research. And is there, but it's kind of like a volunteer basis, right? Some countries can be like, nah, not doing that. Yeah. Right. What's stopping them? Yeah. It's definitely possible. Yeah. There's really nothing stopping them. Yeah. Okay. Okay. Unfortunately. I see. Well, that's, I guess, that's not as gloomy as I thought it to be. You know, I'm going to put the title of this podcast to be a Doomsday of mankind, but maybe a smiley face at the end. It seemed high of alcohol. Yeah. Okay. A bit more of a personal side. So like what got you into the AI safety space? Right. So I actually chanced upon the AI safety space. I wasn't looking out for it at the start. Right. So this was like two semesters ago. Okay. There was someone who was really passionate about AI safety, a graduate of NUS. Okay. Okay. Came back to NUS and advertised a technical AI safety fellowship. I see. So I joined the fellowship. Right. And afterwards, for every week, we read some key papers that give us like a better understanding of the research papers. Yeah. Research papers or like blog posts. Yeah. Let's give us a better understanding of like why like safety is a problem based on how we are currently training AI and what is being done about it. Okay. Yeah. And this sort of like awakens some sort of like consciousness within me that oh, I recognize that actually AI is unsafe. And if we continue at the pace that we're currently going, they are going to be risks right. Right. So it's like climate change actually kind of does some parallels if we continue at the pace we're going this. Yeah. It's going to be bad. Yeah. I see. And so that's what got you in. Cool. And you mentioned about you researching about it. What is the research that you're conducting about it? Right. So remember what I said like in the previous part about us not crafting AI but actually growing it. Right. So what researchers and engineers do is that we give AI inputs and then afterwards based on whether the outputs are right or wrong, the AI learns to get smarter. Yes. So what I'm trying to do is to look into the brain of the AI and try and locate like where the what neurons are like synapses within like the AI's brain correspond to certain human level concepts. Okay. What is this AI brain is like a place or it's like the matrices like you mentioned just now. But how do I find it? You know, online. So we look at the correlations. Okay. So when the AI outputs something which is related to speech, for example, like what part of its brain lights up? Like lights up. So there's a physical thing that light up. It would like a weight calculations from a certain part of the matrix more strongly. So it is using like this part of the matrix. Okay. And we would say, oh, that's where like the speech related knowledge in this AI is taught for in the matrix. Yeah. So when you ask a certain question, this part like so in the matrix. Yeah. Or like more technically weights. Yeah, explain it in delve a bit more about what even weights more. Right. So when you train an AI, like you initialize like a matrix, for example, as one than zero zero. Yeah, it's just bunch of ones and zeroes. So we call those like the parameters of a model. Okay. And then after you have, you give it a bunch of inputs, a bunch of outputs and now first you tell which is why right and which is wrong. So update like the ones and zeros to come like new numbers. Okay. And those numbers are what we call the weights. Okay. I see. So the new numbers I call the weights. Yeah. After it goes through its own calculation. Yeah. And when you ask an AI a question, right? So you give it like some new input that hasn't seen before. Yeah. Like what should I eat for lunch? Right. Then these weights will calculate the ideal answer for you and give it back to you. And you're investigating which is correlated with what? Yeah. So I'm investigating, oh, how does it like, where does the food related concepts in like the AI's brain come from? Wow. That's very impressive research actually. I think it's quite interesting because ultimately AI is a black box and it's so mysterious. So I want to like poke through the mystery behind like what's going on inside. Oh, you have to send me some of this AI brain that you mentioned to me. I'm quite fascinated with like how it works. I mean, we're talking over to mics right now. It's definitely helpful if you like, just a down on. I don't know if pen and paper works for you. Oh, that's actually a lot of like very nice illustrations by Antropic. Okay. Which are available online. So I guess if like the listeners want to go and take a look at it in the Antropic Circuit Strat. Right. Yeah. Like just AI brain. You probably have to search like Antropic Transformers Circuit Strat. Yeah. Oh, Transformers in like a LLM kind of. Yeah. Yeah. I see. I see. Okay. And then what is the implication for the research that you're doing on your own spare time? Oh, it's the main point of this is to sort of locate maybe where the model learns all they are unsafe behaviors. So let's say we. Yeah. So that you can understand it more. And like one day if a model displays some sort of unsafe behaviors, maybe you can go inside its brain and then like turn something off. Right. Because you can't turn the whole thing off. You just need to know which is the bad apples from the good apples and take the bad apples out. Yeah. That's the idea. I see. Well, that's just really cool. Last question before we wrap up today. Like so I mean hearing all of this stuff. I'm slightly worried about you know AI replacing my jobs. AI starting a doomsday kind of thing or you know just hallucinating and giving me wrong input for my assignment. So what should the public do with all these kind of you know worries or dangers out there? Honestly I'm not really certain what we can do as like a member of the public. Right. Because taking a step back a lot of the things that we've talked about today are like dealing with like actors which are so much more powerful than ourselves right. Yeah, yeah. I'm not going up again Sam. I think what the public really should do is to like keep yourself abreast of like the latest developments in AI being formed. Yeah, being formed. Okay. And also like if you have a good idea of you know what AI is capable of today or like what people can use AI for in like a creative manner. And maybe you can be you can have some sort of advance notice before some aspects of a job get replaced for example. Right. What jobs are we looking at to like you know get replaced as soon as so I won't go into that and I'm graduating. Well I mean I think like with all the agentic stuff going on nowadays. Yeah. Like paperwork has been automated. Oh god. One more. So for example I think recently Claude released like a plug in for for processing legal documents and in the same day like some stocks for like like consulting companies like job like 20%. Gosh. Cause I mean yeah I can do their job. I guess. Yeah. Yeah. Yeah. Basically. I can do it. You need someone to like you know oversee it. So yeah. I think they are arguments. So like I don't think a complete automation is in order. Right. At least not for like the next few years. I see. Yeah. But still it's very useful economy I guess. For people in those companies. I see I see. Well. Um job markets tough. I'm going out there. So hopefully AI doesn't take my part of you know my my my fan won and Cantonese you call it like what the fan one. Yeah. Yeah. So well thank you so much for sharing with that. And I guess the main takeaway from what the public can do is just to be informed as you mentioned. Yeah. It's the nice being here. Yeah. Of course. I have a pleasure that you came on and with that I think I'll bid farewell to our listeners on what's going on until next time.
Podcast Summary
Key Points:
AI safety concerns arise from AI's increasing power and integration into daily life, similar to safety considerations for other technologies like aviation.
Key risks include malicious use by bad actors (e.g., terrorists), AI's opaque decision-making processes ("black box" problem), and potential for AI agents to cause harm if manipulated (e.g., via prompt injection).
Macro issues involve a competitive global AI race, lack of robust international regulation, and the existential risk of superintelligent AI becoming deceptive or misaligned with human values.
Frontier AI labs conduct safety research and have frameworks, but these vary in effectiveness and are often secondary to competitive and profit-driven pressures.
The worst-case scenario, if safety fails, could range from widespread manipulation to existential threats, though current AI is not yet capable of autonomous catastrophic actions.
Summary:
The discussion centers on AI safety, highlighting its importance as AI becomes ubiquitous and powerful. Key concerns include the risk of AI being weaponized by malicious actors, the inherent opacity of how AI models make decisions (the "black box" problem), and vulnerabilities like prompt injection that could lead AI agents to execute harmful commands. On a macro level, the competitive AI race among companies and nations prioritizes rapid development over safety, with insufficient international governance to enforce pauses or regulations.
While major AI labs have safety teams and contingency plans, their frameworks are inconsistent and often overshadowed by market pressures. Researchers are increasingly aware of the potential for AI to achieve superintelligence, which, if deceptive or misaligned, could pose existential risks. However, current AI lacks the autonomy for such catastrophic outcomes.
The conversation concludes by noting parallels to issues like climate change, emphasizing the need for proactive safety research and sustainable development to mitigate long-term dangers.
FAQs
AI safety involves ensuring that artificial intelligence systems are secure, reliable, and do not cause harm as they become more integrated into daily life. It's important because AI is a powerful tool that, if misused or misunderstood, could pose significant risks to individuals and society.
Key risks include AI capabilities falling into the wrong hands, such as malicious actors using it for harm, and the inherent opacity of AI systems, making it difficult to predict or control their behavior. Additionally, AI's increasing autonomy could lead to unintended consequences if not properly managed.
AI is trained through pattern recognition without a clear understanding of its internal decision-making processes, making it opaque and unpredictable. This lack of transparency means we may not know how to fix AI if it behaves unexpectedly or maliciously.
Examples include prompt injection attacks where AI agents can be tricked into harmful actions, such as deleting user data, due to their inability to differentiate safe from unsafe instructions. These incidents highlight vulnerabilities in current AI systems.
Many frontier AI labs have safety teams conducting research on issues like model deception and harm, and they implement safety frameworks that may include pausing development if risks are detected. However, the effectiveness of these measures varies across organizations.
The worst-case scenario involves AI becoming superintelligent and deceptive, potentially leading to catastrophic outcomes like widespread manipulation or existential threats to humanity. This underscores the need for proactive safety research and regulation.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.