Nicolas Cubaud On Building Trustworthy Computer Vision Systems
37m 59s
In this episode of the AI Standards Stack, hosts Michael Minnelli and Adam Smith interview Nicola Kabo, editor of the European Harmonized Standard PREN 18281, which focuses on evaluation methods for accurate computer vision systems. Kabo explains that the standard, developed under CEN-CENELEC JTC21, is crucial for supporting the EU AI Act by providing a structured framework to measure and compare the accuracy of computer vision systems. These systems are central to high-risk applications like medical diagnostics, autonomous driving, and workplace safety, yet previously lacked consistent evaluation methods. The standard includes a taxonomy of computer vision tasks (e.g., classification, detection, segmentation, tracking, image generation) and a set of quantitative metrics and technical methods to evaluate them. Kabo emphasizes that robust evaluation is essential because a simple accuracy score is insufficient—context, operating conditions, and potential consequences (e.g., false negatives in medical imaging) must be considered. The standard aligns with international efforts like ISO/IEC but is adapted to the EU's legal framework, prioritizing fundamental rights and traceability. Kabo notes tensions around speed and policy differences, particularly for sensitive use cases like biometric surveillance. Overall, the standard aims to turn general legal requirements into technically measurable, reproducible, and auditable metrics, ensuring trustworthy AI deployment in Europe.
Welcome to the AI Standard Stack with me, Michael Minnelli, director at ZN Group. And me, Adam Smith, chair of the AIQI consortium. Each episode on AI Standards Stack, we discuss developments in AI assurance with guests from around the world that are leading the charge on the standards, ethics and regulation of AI. On the show today, we're joined by Nicola Kabo, founder and CEO of IceNap and co-founder of Total Image with a Y not an I. He serves as the editor of a European Harmonized Standard, which is a preliminary European norm, 18281, artificial intelligence, evaluation methods for accurate computer vision systems. The standard is developed and the Sandsanelak JTC21 in support of the EU's AI Act. With over 25 years of experience in computer vision and image processing, Nicola leads this project to support compliance for AI-powered computer vision systems. Well, welcome to the show, Nicola. It's a pleasure to have you on. And before we unpack the stack, I was wondering, would you like to say a few words about your professional background, particularly the work in computer vision and I snap. And we'll get onto Total Image, I think later. Yes, well, first of all, thank you for inviting me today. I'm very pleased to be with you. Regarding basically the work we do at IceNap on Total Image, we have kind of followed, let's say, the same objective for let's over a few decades now. It's really to help industrial companies on field workers, in their daily activities to turn an image or a video into a national ball, like say, business information. Everybody knows how to take a picture, how to make a video. So, we are really trying to leverage, like say, this source of information to help businesses. So just to give you an example, we, for instance, help people in the nuclear industry to grow much faster into their conformity assessment, just by taking a one minute video of working a rear with the tablet. It's basically device size six, the time of use for control. So, it's an example of how a video can be turned into a conformity report to validate all the regulatory equipment that are needed on working a rear. You know, it can be also to check if workers on site, you know, especially regularity sites, you know, were proper equipment, like helmets, et cetera, to make sure that they are safe. And you know, when they enter NASA or RIA, make sure that they are well equipped, you know, to be working in set conditions. So that's kind of examples, many of them, but some examples that we, of projects that we do for our customers, okay. And now, maybe just to give a little bit more background on information about the, let's say, the standard, you know, the PRN 18 to 81 that would be working on. I just want to say that it has been very collective work, you know, I mean, I'm here today. I'm very happy to be reviewed as the editor of the standard, you know, but we have been really working as a team with other colleagues from France, from Denmark, Austria, Hungary, that we are also all new to standardization. So, you know, it has been a great learning experience like every day we keep learning on standards and also through help of people like Adam, you know, who are very, very extraordinary. So, yeah, so that's isn't clear my background in computer vision and also in the standardization edition. Well, thank you for that. It's a good opening there, Nikola. We're going to cover a lot of acronyms today, and I just thought I'd help out a little bit by setting a few out upfront. So we're talking about Sen Sennelack. Sen is the European Committee for Standardization, naturally being French and deferring here that's CEN. And then we have Sennelack, CEN ELEC, which is the European Committee for Electric Technical Standardization. And then you referred quite rightly to PR EN18281. Sometimes I think we should be reading these things out rather than trying to do it as a podcast. But that of course is a very, very important European standard really directed at computer vision. So I just lay that out for the audience. Sen Sennelack and PR EN, Pren 18281 on artificial intelligence evaluation methods for accurate computer vision systems. So you've had an exciting career in computer vision. 25 years, founding iStab, and total image. Why did you think that the standards were an area that you wanted to devote some time to? Well, you know, it was, let's say, quite an opportunity. I mean, I didn't really look for it, but since we work a lot in the industry, especially in the, let's say, regulated industry, you know, like, I mentioned the nuclear industry, it's not the only one. But you know, I was thinking, well, I know that the AI Act is under discussion, you know, of course, there's going to be an impact, you know, for us as a solution for computer vision solution providers. Also for our customers, you know, who are going to deploy, you know, computer vision systems in this regulated environment. So basically, it was about three, a little bit more than three years ago. You know, I wanted to know more about the, say, AI Act, like how we put impact our customers on us. And I attended a conference that was led in Paris by Patrick Besom, you know, who is the co-chair of the GTC 21. And, you know, we we had a good discussion on the tool, me that there was basically a gap, you know, for computer vision systems. They were a standard that was ongoing, you know, for the NNP, but not for computer vision systems. So if I was interested in the topic, you know, I could basically be leading this initiative. So I was probably a little bit crazy, but I said, yes. And actually, it has been a very very interesting, very exciting journey. You know, for me, the goal was really to try to help, let's say, companies at a wide scale, you know, to develop like AI-based computer vision system in a very trustworthy manner. Okay. So I thought it was a great opportunity to jump into this standard news standardization experience. Well, in preparation for this, I thought I would consult an expert and I have an article I want to read just to quick paragraphs from computer vision is everywhere in high risk AI. It sits at the heart of medical imaging diagnostics, autonomous vehicle perception, industrial quality control, biometric identification and workplace safety monitoring, almost every sector covered by the EUA Iaxx, Xanx3 risk categories involves a system that at some level, processes images or videos and makes decisions. Yet, until now, there has been no harmonized European or international framework for that matter on how you actually measure whether a computer vision system is accurate. Providers have been left to pick their own metrics, apply them inconsistently and present results that are incomparable across implementations. For conformity assessment purposes, this is a problem that needs to be solved. It's a one AL Smith can't work out, what his name is, who wrote that. Anyway, just I think it is important for our audience to grasp how fundamental this is. So your standard pre-18281 just entered public inquiry under the SENSEN SENTLEX JTC 21. Could you walk us through what does the standard cover and why having solid evaluation methods matter so much? Yeah, well, thank you, Michael, you know, for putting the context around the standard. So yeah, that's basically the standard regarding the evaluation methods. Okay, so the PR-EN18281 is really dedicated to support the methods, you know, for let's say selecting, applying, interpreting evaluation of computer vision systems. Okay, so it includes quantitative metrics and also technical evaluation methods. Okay, you know, like we basically develop like two standards, so the one for the evaluation methods and also another one that's related to taxonomy, you know, like the taxonomy of computer vision tasks, because as you pointed, it's very important to speak, let's say, a common language, okay. So basically, you know, like as you mentioned, computer vision covers a very vast number of use cases, okay, it can be detecting pedestrians crossing the front of autonomous vehicles.
it can be generating images with daily or mid-journey. You know, there are many, many computer vision tasks and it's true that there was a gap, you know, to basically describe, have a common language, you know, for this computer vision task, you know, for like classification detection, segmentation, tracking, position estimation, image generation, like all these different types of tasks. So basically, the way we did it is, you know, to develop like these two standards. So one for the taxonomy and the other one, let's say they work, they both work hand-by-hand, you know, for the evaluation methods, you know, for these tasks, okay. So basically for the first standard regarding the computer vision task, you know, which is a PRN, PREN 18 to 8.8. We describe the values, categories of computer vision tasks, you know, for image analysis, video analysis, image generation, video generation. We describe the task in a simple manner. So everybody can understand what it is and make sure that we are talking about the same task, you know, when we want to compare to computer vision systems. And we also describe their input output. So, you know, basically having a standard, a standardized way, you know, to describe this task. And after, you know, the evaluation methods, you know, standard, which is, let's say, the core of our, the focus of our discussion today is really going to apply, you know, the evaluation methods that will correspond to a given task, okay. So basically the goal is really to evaluate this computer vision task, you know, to ensure they fit the intended purpose, okay. And basically our goal is really to answer the question like how to turn, let's say, a general legal requirement that sits, you know, into the act, into something that's technically measurable, reproducible, and also auditable, okay. So that's really the goal of this standard. So basically in the 18 to 18 standard, you know, we describe all the metrics that are available, you know, to evaluate the task. You know, some also come from other standards, you know, they are always the standards, let's say, for classification, for instance, okay. So we are not going to re-end the wheel, we are just going to use these well-described metrics, you know, on the plied ends or the specific cases of computer vision. And after, they are also some specific metrics that are, they said that are really specific to computer vision, for instance, I/O/U, you know, intersection of our union, you know, for instance, to evaluate, let's say, the capability of a system, you know, to well detect objects for instance in an image. Or, you know, there are also metrics for image generation, so that are also very specific to computer vision. So basically, the hour rule is really to apply, you know, the right metrics in standardized way, you know, with a common language, to make sure that we can basically compare appels and appels, which are going to be benchmark to computer vision systems. We spoke to Michael Theney recently, who's the editor of ISO AC4213 on this podcast. And that's a similar standard in the, you just mentioned it there without mentioning the number, Nikola, but that they're establishing the same thing for classification regression recommendations and clustering. There's another project which is focused on natural language processing tasks. And we kind of thought that was the full set a couple of years ago. But the seminitial discussions happening recently in international standards about embodied AI. And whether we need a similar standard that looks at the different tasks that robots will essentially perform, planning how to fetch an object and such things, which is quite interesting as well. Right. Yeah. Thank you, I don't know for the, the add-on and after maybe to answer your second point, Michael, you know, like why is it important, let's say, to have like, like, a proper and solid evaluation method for the context of the AI act, you know. You know, first, let's say it's always the same. Like we start with a problem, you know, that we want to solve, okay? So computer vision tasks, you know, they can solve certain problems, you know, for detection, classification of object, et cetera. And you know, the reality is that there is not, let's say, one feature all metrics, you know, to evaluate the tasks that are using all the problems, okay? So, so basically, you know, if you take the AI act, you know, each site that, let's say, a computer vision system, you know, and certainly computer vision, but it's also valid for computer vision. It requires an appropriate level of accuracy, robustness, on the cyber security throughout all the life cycle of the system, okay? So it means that basically, I mean, the AI act tells us that an eye-risk AI system, you know, must be designed, developed in a way that they achieve this proper level of accuracy, robustness, on the cyber security. So, you know, in practice, it means that we need to have, let's say, valid on robust evaluation methods, you know, to fit this requirement of the AI act, you know? So, so basically, you know, when you work in the computer vision field, saying that my system is 95% accurate, for instance, it doesn't mean much, you know, if you don't understand exactly the context, what is measured on what type of data, you know, and which operating conditions. So, so basically, you know, first, we need to be able to well-classic, I mean, categorize, you know, the type of problem that we want to solve, you know, is it image classification, object detection, etc. object tracking, you know, for instance, if you want to follow like an individual in a crowd, for instance, you know, you want to understand if your system is going to be working at day, at night, all, I mean, both, you know, if they are like, if you know, sometimes some false positives, you know, or false negatives, you know, can be very damageable, you know, for instance, in medical imaging, if you miss, for instance, I mean, if they are system that's designed to predict, you know, concern lesions, for instance, if you miss a suspicion lesions, you know, radiology image, it might be much more impactful for a patient, you know, than missing, let's say, I mean, that flagging correctly, benign harmless region, you know, on the image. So basically, you know, it's very, very important to understand the context on the operating domain of your system. And you know, like all the examples, you know, shows that when you talk about evaluation of the system, you know, it's usually rather complex, you cannot limited to, let's say, one matrix on one average score, you know, you need to basically evaluate the system on their different conditions with, you know, certain, you know, certain thresholds, because again, the consequences of missing, for instance, an object, and I'm thinking again about medical imaging can be very critical. So, you know, it's all this, that's the thing that you need to have in mind on that, basically, justify that we need to have a robust method. And yes, it has to be, that's a very well-defined, you know, to make sure that again, the systems are evaluated in a very robust manner. Well, indeed. And if I can share sort of a personal anecdote, 49 years ago, I entered the Harvard laboratory for computer, for computer graphics and spatial analysis, so, you know, almost a half a century ago. And what your standard is addressing is so important. So, out there in the audience, just think about how do you detect an object reliably? You know, does my system detect objects that your system detects? How do we track it? What does it mean? What is tracking accuracy? What is the segmentation of the object? Is that a person or is that a person with arms or a head? So, that's a depth estimation, 3D reconstruction. There are many, many areas of this standard, which are really important. Because until you define these, you can't start to evaluate things. So,
I'm going to come back to another question in a minute, but before I do so, Nikola, your work is firmly rooted in the European AI framework and JTC 21. How do you see these computer vision standards lining up with what's happening internationally at ISOIEC? Are there any tensions there, any conflicts, or is your work supporting theirs or the other way around? Well, from what I see, I think there is a strong logic of alignment between the European and the international levels, like ISOIEC. The only, I mean, what I would say is that the difference is that the European context is a little bit different and has something important is that the standards, you know, let's say, being developed to be bound to the legal framework of the AI Act, so it's not only guidelines, but it's like, we go legally binding, okay? So that changes a little bit the role of the standard, you know, I'm thinking about a relative computer vision, for instance, regarding like biometric identification, for instance, you know, for instance, you know, if you think about, let's say, in a technical perspective, the metrics, you know, to evaluate these systems, I mean, technical metrics, they can be totally similar, you know, whether you are in China or in Europe, you know, we are talking about the same technical systems, but, you know, in Europe, that's one of the goals of the AI Act, you know, it's really to, let's say, to put an emphasis, you know, on the fundamental rights, you know, on the make sure that, you know, these biometric systems, like if you want to follow, for instance, people, I mean, let's say, for crowd surveillance, you know, these type of use cases that use computer vision, you know, they need to be, let's say, need to be deployed, they need to be compliant, you know, with the jurisdiction, you know, and basically, and that can vary, you know, depending on the area of the work, you know, what can be permitted in China, you know, for, let's say, crowd surveillance, you know, is actually not permitted, you know, in Europe, for instance, for the AI Act, okay. So, I would say that, you know, let's say, at the international level, you know, there are a lot of standards, AI standards, you know, that have been developed, you know, to cover like all the risk management, the data quality, you know, like many, many different topics, so it gives like a very good background on foundation, on the, you know, again, the goal is not at the European level, you know, to really the will, you know, like everything that's good, you know, internationally on that fits, let's say, the purpose of the AI Act, you know, of course, we need to use them, but for what is, let's say, specific to Europe, you know, then of course, we need to have like, let's say, our own specific rules, okay. So, basically, to, let's say, for computer vision, for instance, you know, we, let's say, we also need to make sure, like for high-risk systems, you know, that the solutions are well-documented, that they, you know, they are traceable, that so, you know, they are also like quite a lot of requirements that are requested by the AI Act, that we need to fulfill. So, so I would say, you know, like basically, maybe I wouldn't talk in terms of really tensions, even if they are like little, I mean tensions, because, you know, we need to provide these standards quickly, because, you know, the AI Act is already enforced, you know, so we know that these technical standards to explain how, basically, we should be compliant with the AI Act, they need to be produced rapidly. So, of course, there's a tension regarding speed, there's a tension, you know, regarding policies and activity, so that the example, for instance, of biometrics, of crowds of billions, you know, but, but, you know, I think it's, it would be, let's say, more, like, we put it as a layer, the system, you know, like basically having the international more standards, you know, providing the foundation, and in Europe, you know, we will build, let's say, some, let's say, some, some bricks, you know, on top of that, for our specific European needs, actually, I think this is an example of a standard that should have been an ISO standard. I tried a number of times to get collaboration going between the committees and for lots of different reasons failed to get the committees to agree to a model to work together, but we have worked together on other things like the equivalent for natural language processing, and that's gone very well. I see a future for this standard being progressed as an ISO IEC standard as a second version, or something like that. I have, right, right, right, it's true that for the moment, it's, let's say, a standard, standard, okay, so the parallel development with ISO IEC has not been, I mean, has been rejected, okay, but it is true, I fully agree with you Adam, which would be, it should move to ISO standard, yeah. I mean, one of the things that's on my mind, of course, is we've got the, you know, the EU framework, we've got the ISO framework, but we've also got a bit of fragmentation between Europe, the US, China, and frankly, some others as well. Do you think that in this particular space, we'll eventually land on a harmonized evaluation method, some type of global benchmark, or do you think we just have to kind of tolerate a variety of competing standards? Well, that's a good on tough question, you know, my, my tech on that is that I think I see harmonized evaluation methods, you know, as, let's say, global benchmark, okay, because again, like the standard that we're developing is very technical, on whether you are, you know, I mean, the technical problems, evaluation metrics, et cetera, should remain the same. So, basically, I totally agree with Adam's previous comment, you know, which I think it should be basically global, you know, but again, after like their specificities that are related to jurisdictions, you know, that make me feel that that will be probably very difficult, you know, to have like, let's say, global, let's say, bench, I mean, let's a global standard, you know, worldwide because of its specific jurisdictions, you know, now, you know, so, you know, I know that we took the example of a facial recognition in China, okay, but, you know, we see that it's also regulated, okay, it's more and more regulated, there are specific tools, but it just had the, let's say, the deployment is being used, let's say, it provides some value, you know, for the Chinese people, so I would say that the metric, you know, should be on the benchmark, should be global, you know, but basically the permission to deploy it, these systems, you know, is probably not going to be, that would be my take to get my tough questions. I think, I think we can agree on the metrics and the requirements for how you apply those metrics to make them repeatable and reproducible. I hope we can all agree on those globally. But if you look into the, if you look at like the accuracy of valuation requirements in the European laws, you're looking at the accuracy in the context of intended purpose and the risk to help the safety and fundamental rights, and there are many jurisdictions, most jurisdictions do not require an AI system to have a specific intended purpose, there allowed to be these general purpose systems, and most jurisdictions don't consider human rights in the same way that the European law does. So I think the metrics we can agree, we can standardize globally, the way you decide whether the results are good enough, I think we are some way away from reaching any kind of harmonization. Yes, would you agree, Adam? I mean, do you guys think, you know, all on see me, if we look at some of the things that we're trying to achieve with this, we've got all the kind of technical accuracy stuff, we've also kind of got robustness and tolerance issues, but ultimately, doesn't this fall into trustworthiness, does this establish appropriate reliance? Can you give us a few words, Nicola, on your thoughts on how this is working or not working towards trustworthiness, particularly given annex three? Right. You know, I think it's a very good point because, you know, we have, let's say, many customers using our
our systems, you know, are basically gaining the trust of a customer, of course, who expects that our system work well, let's say, for the problem that we need to solve with them is not an easy task, OK? So basically building trust and making sure that, let's say, our systems are trustworthy, you know, it's a very delicate and very important mission, let's say. And for me, the AI Act, you know, the harmonized standards, you know, they give us really, like, let's say, guidelines on the bread framework, you know, to basically build with trust, you know, with our clients, you know, and it's, so I think it's very important, let's say, to connect, let's do the trust for finance, which is really the, for me, the conditions for adoption, OK? Like, basically, if you don't trust, you know, the type of systems, and, you know, Adam mentioned how important it is to have also reproducible behaviors and results, you know, for a given intended purpose, you know, if you don't provide that to your customers, of course, they are not going to trust your solutions, it's not going to, they're basically help their process being improved, and, basically, we lose all the value that we can provide. So, so I think trust for CNES is really key for adoption. And again, we build the trust that it's basically a two way, let's say, two way work with our customers. You know, they have the knowledge, the expertise, like whatever in the nuclear field, et cetera. We adapt our computer vision and AI systems, you know, for their need. And, you know, together, you know, we're able to build this trust. And I think the framework provided by the AI Act, you know, is a really, really great tool to make sure that, you know, we, we provide like transparent evaluation that we have, you know, the right method, you know, to evaluate the system that we basically have like a relevant test set, you know, to test the systems in different conditions. And all that basically contribute to trust on the confidence that our system, that our customers, you know, can have in our system. So, very important. Well, I'm used to going home and trying to explain to my spouse that, that I've had a very interesting AI standards stack podcast discussion. And this is no exception. And whilst she's slowly coming around to believing me, sadly, we are running to the end of time, believe it or not. So, if I could close the question because Nikola, you're extremely interesting to me as somebody who is very technical, but also a genuine entrepreneur having built up two businesses in computer vision. So, wearing both your entrepreneur hats, but also the fact that you really are deep into standards and editing of standards. What's the practical advice for companies building computer visions systems today? You know, the AI Act is starting to bite. I'm a, you know, young postdoc coming out, wanting to deploy my techniques in computer vision. Why do I want to do anything with standards? What, what should I actually be doing with this space, ignoring it, piling in like you do, and volunteering lots of time? Should I just be slavishly following it and paying lots of advisors? What should I actually do? Well, you know, I think my, maybe my first advice would be to not to wait for the final publication of every standard, you know, before taking action, okay? We basically, you know, compliance is very important. We know that, you know, if we don't pull field, like they say, the AI Act requirements, you know, providers, the projects, etc. We'll get penalized for that. You know, that it's legally binding, etc. But, you know, like beyond that, you know, I think we need to think about it a lot. So, it's more as an entrepreneur, we need to find a great opportunity for product on the, on the cells, okay? Basically, you know, with this, let's say with this AI Act regulation, we can really, I think, build better products, you know, we know that there are some very important guidelines regarding documentation, transparency, accuracy, cybersecurity, etc. So, I think, you know, it can be for great commercial opportunities for better relationships with our customers, to help our customers also understand better, you know, what's behind the hood, okay? Or under the hood because, you know, they won't system that works. But, you know, like when you go into details, it's like way more complex than that. And it also helps them, you know, to improve their competency on the help us work together. So, I think, yeah, that would be the first advice, you know, be interesting to be standard, because it's not only, I mean, it's not, let's say, bureaucratic burden, you know, it's really an opportunity to deliver better products on the, I think especially in the context of AI, you know, it really provides like a great framework, you know, sort of better product adoption. And after, you know, it's, there are also some rules that are, I think, very important to keep in mind, like, you know, we talked about defining the intended purpose very precisely because, you know, systems are complex, operating domains are complex, etc. So, really master the environment, in which the system will evolve, you know, it's very, very important. You know, it's also important to keep in mind that we, VI, the system, they are not static, they are live systems, you know, they keep improving with data, they may shift as well, you know. So, basically, to be able to, let's say, to think about the evaluation strategy before finalizing the model, but, you know, at the different stages of the design on the post-market monitoring, it's also very important, you know. And I'm very important also to, to, let's say, test the evaluate the systems under different conditions, you know. And, you know, maybe the last word would be to engage with the standards early, you know, like, basically, you mentioned that our standards and public inquiry, you know, that's also a good way to provide feedback and good ideas, you know, to people like us who we need these standards. So, so, basically, that can be the last advice, really, to get interested in, to contribute also to these standards. Well, that is excellent advice on which to end. And, Nikola, I'd like to thank you so much for sharing your valuable insights today. I'd like to also thank my co-host, Adam Leon Smith. But most of all, to all of you listening to the show. So, please join us next time when we cover the full stack of IT standards, ethics, and regulation. And maybe get a bit of computer vision into it all. Thank you very much, Nikola.
Podcast Summary
Key Points:
The podcast discusses the development of a European Harmonized Standard (PREN 18281) for evaluating the accuracy of AI-powered computer vision systems, led by Nicola Kabo, founder of IceNap and Total Image.
The standard aims to provide a common language and robust evaluation methods for computer vision tasks (e.g., classification, detection, segmentation, tracking, image generation) to support compliance with the EU's AI Act.
It addresses the gap in measuring accuracy consistently across computer vision systems, which are used in high-risk applications like medical imaging, autonomous vehicles, and workplace safety.
The standard includes two parts
Proper evaluation is critical because metrics like "95% accurate" are meaningless without context—operating conditions, data types, and task-specific consequences (e.g., missing a lesion in medical imaging) must be considered.
The work aligns with international standards (e.g., ISO/IEC) but is tailored to the EU's legal framework, emphasizing fundamental rights and traceability, with tensions around speed and policy differences (e.g., biometric surveillance).
Summary:
In this episode of the AI Standards Stack, hosts Michael Minnelli and Adam Smith interview Nicola Kabo, editor of the European Harmonized Standard PREN 18281, which focuses on evaluation methods for accurate computer vision systems. Kabo explains that the standard, developed under CEN-CENELEC JTC21, is crucial for supporting the EU AI Act by providing a structured framework to measure and compare the accuracy of computer vision systems. These systems are central to high-risk applications like medical diagnostics, autonomous driving, and workplace safety, yet previously lacked consistent evaluation methods.
, classification, detection, segmentation, tracking, image generation) and a set of quantitative metrics and technical methods to evaluate them. , false negatives in medical imaging) must be considered. The standard aligns with international efforts like ISO/IEC but is adapted to the EU's legal framework, prioritizing fundamental rights and traceability.
Kabo notes tensions around speed and policy differences, particularly for sensitive use cases like biometric surveillance. Overall, the standard aims to turn general legal requirements into technically measurable, reproducible, and auditable metrics, ensuring trustworthy AI deployment in Europe.
FAQs
PREN 18281 is a European harmonized standard for evaluating computer vision systems. It provides methods for selecting, applying, and interpreting quantitative metrics and technical evaluations to support compliance with the EU AI Act.
A common taxonomy ensures everyone uses the same language to describe tasks like classification, detection, and tracking. This allows for consistent comparison and evaluation of different computer vision systems.
The standard translates general legal requirements from the AI Act into technically measurable, reproducible, and auditable evaluation methods. It helps ensure high-risk AI systems achieve appropriate levels of accuracy, robustness, and cybersecurity.
The standard covers tasks such as image and video analysis, generation, classification, detection, segmentation, tracking, and position estimation. It applies evaluation methods specific to each task type.
The standard aligns with international standards by using well-established metrics where possible. However, it is tailored to the EU legal framework, adding binding requirements for fundamental rights and traceability under the AI Act.
A single metric lacks context about operating conditions, data types, and potential consequences of errors. For example, missing a lesion in medical imaging is far more critical than a false positive, so robust evaluation requires multiple metrics and thresholds.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.