James Gealy On AI Safety Standards For Frontier Models
39m 34s
In this episode of AI Standards Stack, hosts Michael Minnelli and Adam Smith interview James Geely, standardization lead at SAFE AI, who discusses his journey from testing spacecraft at Northrop Grumman and Airbus to becoming an expert in AI standards. Geely explains that his background in risk and quality management naturally transitioned into AI safety, as he saw the need to apply rigorous aerospace practices to emerging AI risks. He details SAFE AI’s risk modeling work, which connects frontier AI capabilities—such as those demonstrated by Anthropic’s Claude Mythos—to potential real-world harms, particularly cybersecurity threats. The modeling uses benchmark analysis and expert consensus to estimate how AI can amplify malicious activities, with findings showing benchmark saturation already leading to increased damage potential for critical infrastructure and small businesses. Geely highlights the geopolitical implications of models like Mythos, which can find vulnerabilities in longstanding code, and stresses the need for international cooperation to manage these risks. He discusses recent ISO/IEC developments, including a preliminary work item on frontier AI risk management and his own PWI on AI safety, which aims to define acceptable losses and build on existing standards like ISO 42001. The conversation underscores the importance of proactive, analytical risk management to balance fear and action without crying wolf.
Welcome to the AI Standard Stack with me, Michael Minnelli, director at Zian Group. And me, Adam Smith, chair of the AIQI consortium. Each episode on AI Standards Stack, we discuss developments in AI assurance with guests from around the world that are leading the charge on the standards, ethics, and regulation of AI. This episode of the AI Standards Stack is supported by the United Kingdom accreditation service, U-CAS, the UK's National accreditation body. U-CAS ensures organizations that test, inspect, and certify AI systems are technically competent, impartial, and robust, building confidence in safe, reliable, and trustworthy AI. On the show today, we're joined by James Geely, standardisation lead at SAFE for AI. James has been contributing to the OECD G7 Hiroshima AI process, reporting framework, co-editing the AI Risk Management Standard in Sense and ElectJTC21, and serves as the editor of ISO IEC 42119 Part 8, which is all about LLM Benchmarking and Red Teaming. And he's also leading a new preliminary work item on the safety of AI models, systems, and applications, also in ISO IEC. He has 15 years of experience as an electrical engineer in spacecraft testing and operations at Northrop Grumman and Airbus, with practical experience, quality and risk management, technical procedure writing, and information security practices. Well, welcome to James. It's a pleasure to have you on here. We always like having rocket scientists here. Just before we unpack this stack, as we like to say, would you like to say a few words about why you wanted to be on the show today? Well, first of all, thank you, Michael, and thank you Adam for having me. It's a pleasure and honor to be here. And I'll just caveat by saying that this is my first official podcast. So I'm really looking forward to this experience. And honestly, anytime to get to talk with Adam about standards is a good one. And also spreading the good word about them is always recommended. So happy to be here. Oh, great. Well, just to warm up a bit, how do you go from testing spacecraft to becoming an expert in AI standards and AI testing? Was this a natural transition or was a bit a big jump? I think it could be answered in one word or maybe two. And that was the challenge or disaster. You're probably both hold enough to remember this. I actually wrote a piece on this for the OECD and I wonk blog. And essentially what I saw was that with my background in risk management and quality management, testing spacecraft launching them and then operating them on orbit, we have a lot of ingrained practices. And for example, safety culture, you might hear that that term sometimes. And I think with recent AI developments, I saw a need for this sort of experience. And trying to bring and leverage that into the AI space and then document it in standards has been really what I've been trying to do for the last three years now. And just before we discuss standards, it'd be really helpful. Your firm is safer AI. Well, actually, it's an NGO really, isn't it? It does a lot of work around risk modeling and cyber. And it's come up a lot recently with the release of Claude Smithus. What are some of the scenarios you're modeling? That is a great question. So basically the question is with model capabilities, especially now that we're talking the frontier, the question is then, as we remember, you can start and get this asked by, especially a lot of governments, how does that translate into real world harms? We can speculate all we want about the model capabilities, but the question is then, what sort of capabilities actually transition into uplift for malicious use and threat actors and actually then carrying out, for example, a cyber attack? So that's the key question that we were trying to answer. And we spend most of last year working on this. And basically, risk modeling is trying to fill the gap in between those hazards, threats, the capabilities, basically telling you how to carry out a cyber attack, coding a virus for you, for example, and then what's the downstream harm that can result from that? And essentially, what we did here was we carried out a methodology trying to select what risk scenarios are most appropriate to cover and then constructing those. Trying to understand what is the baseline risk to begin with? Because obviously, if you're trying to provide uplift with these models for threat actors, the question is, okay, what do we have today already? Then what are the indicators for that uplift? Like benchmark performance that I alluded to earlier. And then estimating that uplift, and then finally trying to propagate those estimates among, for example, experts, we ran a del-fi process for this, trying to estimate what the uplift given a model with a certain set of capabilities could actually do. And then we went into additional parts with the quantitative evaluation around this. And we have actually three papers published. You can find them on our website. And we basically mapped out several scenarios. I think it was nine of them. And then one of the key findings I think we're finding is that benchmark saturation is already happening, where it's something that the adjustment of benchmark saturation would basically give you increasing damages in dollar amounts at the output. And we said, well, what happens if you go all the way to the top? And max out the benchmark capabilities? Well, we're already getting to that point. And it wasn't, honestly, very, very great, especially for things like critical infrastructure. And honestly, just, everyday organizations and small businesses. So this is something that we're now incorporating into our next round of risk modeling. In particular, for cyber, we have some collaborations with governments that we're working with now, trying to understand, for example, what about defense? Because that's something we didn't consider in our first set of risk models. And it's a big question now with mythos that you mentioned earlier. It's really being used to augment defense. And what happens if that's then brought into the mix? And how does that change the numbers? How does that change the math of the potential harms? One of the things I find really interesting about the risk modeling around front here, AI is it's not just looking at harms that can arise from the normal use of the product. It's looking at the downstream harms. If the product can tell people how to make chemical weapons, or in the case of mythos, tell them how to hack different systems. And every time guardrails are put in place to stop frontier AI systems doing that, we find new ways around it. There are new ways around it going to be released in a couple of weeks that will show that you can bypass a lot of those guardrails with new techniques, much as we've talked about on this show before in terms of poetry, very easily, and with very very high levels of success. No, that's that's exactly right. And usually with our risk models and forgive me, but I wasn't the one who actually wrote them. So I don't have this exact detail off hand. But one of the things that you look for when you're benchmarking or performing a red team exercise is, is the model helpful only? That will give you an idea of the maximum capability and uplift that it can present to a threat actor. And then are there guardrails in place? Most of the time for an uplift study, it's preferable for that exact reason to just use the helpful only model from the beginning so that you can basically assume that the security guardrails will be defeated if not today, then perhaps eventually. One of the things that's been interesting in this space is we had starting in late 23 a whole bunch of stuff about existential AI. It was going to wipe out humanity. And then we had all of the so-called risks which didn't quite work. And suddenly this mythous vulnerability exposure really got people's attention. But how do you go about telling people enough that they develop an appropriate level of respect, even a little bit of fear, without constantly being seen to be cry, crying wolf all the time? And the canonical answer is that risk management allows you to frame the response in an appropriate way that then governs how much you need to put into it both for energy and resources. I think the key thing here is that risk management ideally is risk agnostic. So it doesn't matter what the risk is. You're breaking down the problem space such that you understand, okay, is risk x have a higher severity, higher probability than risk. Why if it does then I should probably be addressing this first. And you can get into questions about tractability as well. But at the end of the day you need to prioritize. We don't have infinite resources. And that goes for organizations that for example are training these models but then also for everyone else. And I think that now right now yes, the queer and present challenge that we're facing is on the cyber front. And so we should be addressing this. But we also need to be looking at, for example, the challenges that are coming up.
about at a societal level certainly from widespread use of AI, especially for children, for example, that's becoming more and more prevalent, but then also perhaps some of these other issues that have been identified, for example, in the EU general purpose AI code of practice for GPI models, looking at an item you alluded to this earlier for chemical attacks, for example, then also bio is another common example that's given. And then looking up further from there, of course, there are discussions around do we think that we can basically manage agents to a degree that we can have positive control over them? Do we need human in the loop or on the loop? And at what point do we have basically agents monitoring agents and it's essentially agents all the way down? What sort of challenges that that face? But again, you have to really understand the problems face and then start to map it out and it kind of connects back to our risk modeling. That's what we were really trying to do is break down this problem to understand what are the risks and potential harms that we're facing. And we think that we've tried to take a pretty analytical and I would say concrete methodical approach to that. And that's really what you need to do at the end of the day. Well, I love that oblique Pratchett reference to turtles all the way down. I'm going to remember that an agent's all the way down, which reminds me getting all the way down in the email kind of banter to warm up for this. We discussed the fact that mythos has had a staggered release to allow firms to get some defensive action against vulnerabilities, but they've also sort of created an inner group, members club, largely of the big boys and girls. So where does this leave smaller firms and individuals who don't have early access like that? It's not just the smaller firms even. I think there was a news article that I read just a couple days ago that was talking about some members of a Chinese think tank we're reaching out to. I believe it was folks at Anthropic requesting access to mythos. And I can understand why of course because I mean both for organizations, government entities, what have you companies both within China but then across the globe. If mythos is a model that has the capabilities that it's purported to have, obviously the challenge here is that in some ways mythos becomes a skeleton key for any IT infrastructure on the planet. And if it can find a vulnerability and open BSD that was there for what was it 27 years, then yeah I think that this is going to be a challenge geopolitically speaking. I think it would be frankly irresponsible. Anthropic to release a model that is specifically designed to find cybersecurity vulnerabilities. But of course somebody will come up with the same model later. But it is staggering how some of the most scrutinized pieces of code on the planet that have been around for a huge amount of time have these vulnerabilities and that mythos has been able to find so quickly. It's really quite scary. I would agree with that assessment. The challenge that we now face is how long I think can we keep mythos and similar tools assisting on the defensive side before they transition to the offensive side because I agree Adam. Unfortunately I think it's just a matter of time before the jagged edge of model capabilities is increasing such that perhaps an open release of the weights of a similarly capable model occurs. And I think this also has implications on a I said geopolitical level before and instability this could potentially cause. And I think that perhaps we'll get to this later. But we I think need to understand that this highlights the need for more cooperation. And I'm hoping for example there's supposed to be this summit between President Trump and choosing paying in Beijing soon. And supposedly this topic might be on the docket. We'll have to see how detailed it goes. But my hope is that this and I think this movement also at the European Union level that this shows that we do need to have more coordination between different actors both private and public on this particular topic. For right now I think I commend and Robick for making the move that it did. I think it was the responsible one. And we'll we'll have to go from there. But unfortunately things are moving quickly. So we have to try to work on that coordination as much as we can. Well this international discussion leaves me to kind of turn to standards in a way a quick question. But is it fair to say your work at safer AI is mostly within the EU AI framework? I would say that it's becoming increasingly international for the reasons that I gave earlier. We have been focused mostly on the EU because we're French non-profit. And obviously we have our base there. So the question is then on the international level what are the right connection points between the EU and international levels. And one of those connection points is definitely in standards as you mentioned. Particularly at ISO IEC there's a GTC-1SE-42. Adam is an officer in convener there. I'm an officer there myself as mentioned earlier. And I think what's interesting about ISO is that it's I think one of the few global forms remaining that we have true international collaboration from class the world that is coming together with experts that are trying to find consensus around these exact topics. Maybe not the exact ones because obviously you can only standardize so far back from the front here. And that's a bit of a challenge right now because I think we need to move as far as we can in that direction. But overall, Dan's your question. We're working with at safer AI more and more international organizations and governments because there's a lot of interest in the work that we're doing for example on risk modeling. And I think that this is a common challenge that we're all facing. This isn't something that's just limited to the EU. I think it's really interesting to see how many countries are coming together at the ISO AI standards. You know, I mean we have obviously countries like the UK and the US. We have lots of EU countries. We have Russia, China and Iran. We have all these different countries that disagree violently on some other topics but are coming together to develop common AI standards and standards that have been approved by 86 countries have a lot of weight in the market. Someone might want to challenge a definition in a meeting but if 86 countries have agreed that definition, it becomes a little bit difficult. Yeah, to me, it's one of the problems when you say the word standard or standards. People kind of focus on the written paper document but it's only one part of a system of accreditation, certification, bodies, competition, regular annual meetings, updates, stakeholders, structures. It's a big system. And what is it? I think there are 180 signatories to the WTO which sort of embodies ISO. So you're right about the power. Now last month we had the ISO's Plenary in Singapore though in which you James submitted a preliminary working item. Could you tell us a little bit about what was discussed and what you submitted? Yes, of course. So this is the we're calling it a PWA preliminary work item. It is the safety PWA. The official title is Safety of AI models systems and applications. And I think it's important to discuss this PWA also in the context of another PWA that was started and launched at Singapore. And this is the PWA on Frontier AI Risk Management and Impact Assessment. Now I think with both of these PWA's there's some commonalities. There's definitely some differences. But the idea here I think is that we are increasingly seeing that there is a need to as we've been discussing, there's a lot happening at the Frontier AI right now, especially on capabilities. And the challenge is what can we do to work on standards to undergird support and then hopefully ensure safety to some degree for the Frontier. Now looking at this what we're calling for shorthand the Frontier PWA. This is looking mostly at, for example, the first thing is the challenge of defining the Frontier. What does that include? Is it based on lock for training, capabilities, it's based on hazards? And then it brings up the question, does the Frontier always move or is everything from now on the Frontier, for example? So that's one thing, one topic it will look to address over the coming months. And then another key feature I think is looking.
at what it refers to as I believe it's i-velocity and i-impact risks. These would include, I think, what we've looked at with mythos already, something like cyber, we've opt about touched on CBRN, probably going to be discussed in these meetings, I think. And then I think the question of how do we do this? I've had discussions about this already with Adam, of course, but for example, do we build on the existing AI management system standard, 42,000? One, do we do something stand-alone? And what are the potential challenges and benefits to going down these different paths to actually doing this? One thing that I would like to try and keep in mind while we're working on that is the idea of marginal or relative risk versus absolute risks that can result from these models or systems potentially, because if we want to ensure some sort of, I would say, either universality or release some connection with existing regulatory instruments that are coming out and will continue to come out, a lot of those have these sorts of ideas of absolute risks. I believe it's SB 53 looks at, for example, a billion dollars in damages and 50 casualties or 50 deaths, I believe it is, if I'm not mistaken, but it could be corrected on that. So the question is then how do you address that if, for example, one developer is looking to release a model and then another one is also looking to release it? Are they both contributing to this absolute level of risk or end-of-what-ways? So that's the frontier, PWI. The PWI that you asked me about specifically, the one on safety, is looking at the question from a slightly higher elevation, and it's looking at safety as the concept of freedom from unacceptable losses. So it's not necessarily focused on the risk or impact assessment, but then how do other things factor in even beyond hazards and threats to what do we consider unacceptable losses? So taking a slightly broader perspective so that we can incorporate, protecting what we care about, essentially. And then from there, there's this transition gradient, I call it, from specific purpose models and systems to the general purpose that we've seen over the last few years. And I would say most of what we have in safety, assurance, assuring safety to a certain degree, is build on the premise that the models and systems that we have can be characterized, their reliability can be well understood, and that we can engineer in safety to a certain degree, and then if we cannot do safety by design putting in constraints, etc. The challenge there, I think, is that as we move through this transition gradient, I think we already have, as we're seeing, how do we assure safety with any reasonable level of certainty or guarantees? That's going to be a difficult problem to crack, I think. And I think we need to start looking at it. Now, I think there's already a lot of research, a fairly large body on this topic that we can look to. We can also look to our friends, for example, work on functional safety standards and joint worker group 4 at SC42. What can we leverage from their experience? But then where, for example, do those systems and processes, unfortunately, not cover something like malicious use, or if we get into systems that are acting with more autonomy and then interacting with each other if we can't necessarily understand all of those interactions, then what do we do from that point on? And this safety PWI is looking to address some of those questions. Some of our listeners will be thinking to themselves these topics are feeding the news for three or four years. Why is the standards community only now starting to cover them? And why is this a preliminary work item rather than actual project? And this is because standards lag, standards lag, the state of science, are a very good reason because to get all of the thousands of experts on the same page about these topics and to get 86 countries on board with these topics, there has to be some existing agreement in industry, in research, etc. And when you actually look at trying to define what frontier AI is, it becomes quite hard. When you actually look at the existing concepts of safety and then try and understand the relationship with some of the concepts of safety, we've been talking about it's actually quite hard and there's actually quite a lot of stakeholders that need to come to the table and work through these topics. So standards should not be leading the state of science. They should be turning the state of science into generally acknowledged and supported state at the art. 100% yeah. Well, I also think it was an interesting discussion, listening to you in particular, James, you've we've got things like what is it? The BSEN 18286, you know, which is a quality management system for EUAI regulatory purposes, EUAIAC. Now this gets intriguing to me. We've got the law, right? And then we have standards. And now we've got standards to say based on the law. And we've clearly got laws that say you have to follow standards. So we're starting to get this very interleave legal system which and they're two very different systems. One is a risk-based system by and large. And the other one is a is a de-regist kind of if-then statement system. And it's really interesting to see how the interaction this is going to play out. What are your thoughts? Excellent question. Talking about the interaction of standards and regulatory requirements, the challenge that we face there is one of what I call a catch 22 essentially or a chicken and egg problem, if you will, where if you have a regulatory requirement with certain essential requirements, then in the European system we have the new legislative framework, the NLF, that then we have harmonized standards which basically explain how to then carry out that essential requirement. And this is a very, I would say clean but not always easy top down approach to addressing this potential chicken and egg issue because the standard clearly refers back to essential requirements. Now if the standard as many standards should provides flexibility such that it can be used in different contexts and different jurisdictions, the challenge then becomes does it have the structure and the substance basically to then match to the essential requirements and then provide any level of if not a presumption of conformity then at least somehow with compliance. And this has been something that we've discussed quite a bit at the EU level in Sense NLTC 21. And obviously we've had to go the route of the harmonized standards so we've done that. And then at the international level I think the question is a little bit more open still. And I think that it honestly depends. I referred to before specific purpose AI systems. I think there the flexibility is definitely warranted and even needed because AI is becoming this rapidly expanding and deployed socio-technical technology that we need to have that flexibility built into our standards. And then as we move into these regimes where we have the challenges of okay are there things that we need to codify a little bit more for particular capabilities for example of these models that are starting to be released. What does that look like? And I think we'll probably have to go into some of those areas in these two preliminary work items. And then as we go forward consider that more. I mean Adam you were talking about for example you know I was standards lags the the absolute bleeding edge of the research frontier so that we focus more on the state of the art and completely agree. And those I think the discussions now that we're going to have to start having of that answers the question. I'd probably just add to that say harmonized standards of the European system do some part of of your panel law. And developing them is very very hard which is why the AI Act is delayed because the standards are delayed. But on average it takes six years to develop a harmonized standard. And in AI we've had to move a lot quicker than that. In fact I know of some areas which are not even about technology or anything new where it's taken as long as 13 years to agree on a harmonized standard. So it is very very hard and slow work because of the big legal impact that it has. And just to build up that if I may I think that I have to I always commend the the experts especially who participate in
in the AI risk management standard. And I'm sure you the same item with UMS that we have moved incredibly fast, had these all day marathon meetings out of you had a week of all the meetings in person. There was all comments at one point from national bodies very recently. It's we have been moving exceptionally quickly and I am just grateful to the experts that have been participating in these processes sometimes for years that have held on as we've continued to ramp up the pace to meet these deadlines. Well, I must say given choice, we need chicken and egg and catch 22. Given the way we've been throwing our numbers all through this, we'll go for catch 22. And we've got time for just two more questions. One was just to dive in to those numbers. Just a teensy bit more in the introduction to add a mention that you're involved with the CN Joint Technical Committee 21, JTC 21. What are you focusing on there? - So I just mentioned the EN, oh, PR EN, AI risk management standard, that's 1-8228. And that's where I'm the co-editor, Federica Collivol is the editor of the project. And we've worked together to basically guide the consensus process through to ensuring that the standard matches the essential requirements of the Article 9 of the AI Act, which delineates the risk management system and then making sure that the commission and their has assessment of that makes sure that it is going to be able to provide a presumption of conformity against those essential requirements in Article 9. That has been my main focus since I started in this AI standard space now for almost three years now. And it's been a long road. I also have to acknowledge, of course, the first project editor, Renaud Defensisky, who unfortunately has to weigh a little over a year ago. It was a difficult transition, of course, but we made it through that. And we are now at public inquiry, thankfully. So those who are interested here in European countries, who can access the respective national body, web portals can actually go and view the standard that we've worked so hard on over the last several years. - Just because one of the things coming across here is the amount of consensus and cooperation amongst the community. And I hope that's coming out. You can't have a strong standard without a strong community and that takes time. - Yeah, you're right. Consensus is the real deliverable, not the PDF. I was just going to say two things. European countries includes the UK, this context, because we're active members of Santa Leck. But I also wanted to say this is the first product safety, risk management standard that really incorporates human rights or fundamental rights within the framework, which has been, I mean, it's the first time that's really been handled in standards. And it's a very fascinating part of this particular project. - Thank you for that Adam. Yes, I have to give a lot of credit to the fundamental rights experts that were willing to actually weed, sometimes neck deep into the standards ecosystem in the universe. And basically put a pen to paper and help us flesh that out. I think that it's a rather elegant solution to the challenge. And I think it's also a flexible and perhaps extensible one for the international context perhaps not today, but at some point in the future. But balancing that with that, and of course, all the other stakeholders that we had in the room, obviously including industry and from many different verticals was an interesting exercise to say the least. But to your point Adam, the consensus was the deliverable at the end of the day. And I'm glad to say that we were able to do so and we'll continue to do so. - Adam, if you would mind, and I mean just briefly, could you enumerate three or four of those human rights principles just for the listeners? - Oh, there is too complicated for me. That's a bad question. The James, for sure. - Oh boy, when you say that, just can you give people a few questions? - Putting me on the spot here. So the reference, basically the foundation for, basically the whole EU is the charter of fundamental rights. And so the charter has a whole list of several dozen fundamental rights. And not all rights are equal, oddly enough. (laughs) Certain rights, for example, right to Frida from Torture, is what's considered absolute in the sense that even if it would provide information, for example, that you need to help someone or save someone, you can't do it. Though those are fundamental principles that are, are a bad rock to the European Union. But there are other things like the right to privacy that there's a trade-off there between private, individual privacy and public good. And there is a balancing act that needs to happen and things like proportionality analyses that have to happen there. Those are what are sometimes referred to as qualified rights. And then also actually with privacy, it also falls into the category of being, what you might call horizontal or applicable between private parties. That is to say that there are privacy agreements, for example, GDPR regulating that, where as a private organization, if you have my data, then there are certain responsibilities and rights that I have around that. - That's helpful, because a bit of color and clearly something for a future podcast to go into quite a bit of detail perhaps. - It would be a whole topic in itself. - Easily. Just before we close, a final question for you if I may. We've spoken about consensus, the international aspects of things, the importance of ISO as a huge system and the community. But five years from now, do you think we're gonna see a globally unified approach to AI governance or a distinct system for each of the major players? And as a coded to that, to me, each of the major players is the EU, China, the US, but might there be others say the Middle East or Africa or Asia? - Well, I would say I hope so. Look, because I think that AI, like many technologies, doesn't really stop at national borders or even international blocks like the EU. I think, and actually taking the fundamental rights and human rights as an example, the EU places different weights on different human rights. And we were actually talking about this a little bit in Singapore and the Seoul statement on what the interaction of standards and human rights looks like. It's interesting because different jurisdictions place, like I said, different weights on different rights. And so then what this requires is a level of flexibility and perhaps extensibility for international standards, especially but then any other international regulatory instrument that might come. And being able to take the specific case and be flexible enough to adapt to, for example, whether it's China, EU, US, but then also ease the cross-border trade and to ease the burden of complying with all of the different rules and the patchwork that we have right now that's been emerging over the last several years into something where if I comply with extended in the US, if I do this additional, for example, and extra additional material on top of that or less than that, et cetera, then I can comply in the EU and Singapore, China. That, to me, is obviously the goal of international standards. The challenge that we might face, though, is if different jurisdictions, different people have different ideas of how to go about that, then obviously this is where the consensus process is important because we need to try to find common ways to approach this that are amenable to all. And going back to the topic of mythos, I think that we're going to need these sooner than five years from now, so I'm hoping that this will be something that we continue to build on, especially at SE42, but also elsewhere, and we have a more universal approach to these problems that is able to work across borders. - Well, that's fantastic and a good point on which to end. I think we all hope to see a globally unified approach to what is a global technology. So thank you, James, for sharing your truly valuable insights today. Thanks also to my able co-host, Adam Leon Smith, and most importantly to all of you for listening to the show. Join us next time when we cover the full stack of AI standards, ethics, and regulation.
Podcast Summary
Key Points:
James Geely transitioned from aerospace engineering to AI standardization, leveraging his experience in risk and quality management from spacecraft testing.
SAFE AI conducts risk modeling to bridge frontier AI capabilities (e.g., cyber attacks) with real-world harms, using methods like benchmark analysis and expert Delphi processes.
Frontier AI models like Claude Mythos pose significant cybersecurity risks, with capabilities to find vulnerabilities in legacy code, raising concerns about defensive vs. offensive use.
International AI standards development at ISO/IEC involves diverse countries collaborating on topics like frontier AI risk management and safety, despite geopolitical tensions.
James submitted a preliminary work item (PWI) on "Safety of AI Models, Systems, and Applications," focusing on defining unacceptable losses and building on existing frameworks like ISO 42001.
Summary:
In this episode of AI Standards Stack, hosts Michael Minnelli and Adam Smith interview James Geely, standardization lead at SAFE AI, who discusses his journey from testing spacecraft at Northrop Grumman and Airbus to becoming an expert in AI standards. Geely explains that his background in risk and quality management naturally transitioned into AI safety, as he saw the need to apply rigorous aerospace practices to emerging AI risks. He details SAFE AI’s risk modeling work, which connects frontier AI capabilities—such as those demonstrated by Anthropic’s Claude Mythos—to potential real-world harms, particularly cybersecurity threats.
The modeling uses benchmark analysis and expert consensus to estimate how AI can amplify malicious activities, with findings showing benchmark saturation already leading to increased damage potential for critical infrastructure and small businesses. Geely highlights the geopolitical implications of models like Mythos, which can find vulnerabilities in longstanding code, and stresses the need for international cooperation to manage these risks. He discusses recent ISO/IEC developments, including a preliminary work item on frontier AI risk management and his own PWI on AI safety, which aims to define acceptable losses and build on existing standards like ISO 42001.
The conversation underscores the importance of proactive, analytical risk management to balance fear and action without crying wolf.
FAQs
It is a podcast hosted by Michael Minnelli and Adam Smith that discusses developments in AI assurance, standards, ethics, and regulation with global experts.
James Geely is the standardisation lead at SAFE for AI, with 15 years of experience as an electrical engineer in spacecraft testing and operations. He contributes to AI standards like ISO IEC 42119 Part 8 on LLM Benchmarking and Red Teaming.
He saw a need for his risk management and quality management experience in the AI space, especially with recent AI developments, and leveraged that to document standards over the last three years.
SAFE for AI focuses on risk modeling to understand how AI model capabilities translate into real-world harms, such as cyber attacks, by estimating uplift for threat actors and mapping out risk scenarios.
A key finding is that benchmark saturation is already happening, leading to increasing damages in dollar amounts, especially for critical infrastructure and small businesses.
He advocates for risk management, which allows framing responses appropriately and prioritizing risks based on severity and probability, avoiding a 'crying wolf' approach.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.