Mike Thieme On Making AI Standards Work For Busy Professionals
39m 0s
In this episode of the AI Standards Stack, host Michael Minnelli and co-host Adam Smith interview Mike Thimi, managing director and cloud advisory lead at Accenture. Thimi shares his journey from biometrics standards in 2005 to AI standards, noting that consumability of standards outputs is crucial for busy professionals. He describes early work on synthetic face images using generative AI, which pioneered techniques later applied to text and other domains. Discussing government AI adoption, Thimi notes initial caution was beneficial for security, and a better balance now exists between rapid tool availability and rigorous authorization. He highlights two key standards projects: AI Verification and Validation (V&V) analysis, which aligns AI testing with existing software testing frameworks, and performance measurement for AI classification, regression, clustering, and recommendation—a comprehensive reference for metrics. Thimi emphasizes the challenge of reproducibility and deterministic outputs, advocating for comfort with shades of gray in evaluation. The conversation underscores the need for actionable standards that bridge technical rigor and real-world applicability, with the biometrics space offering a model for rigorous testing regimes. Overall, Thimi’s insights stress the importance of integrating standards with practical deployment, especially in high-stakes public sector environments.
Welcome to the AI Standard Stack with me, Michael Minnelli, director at ZN Group. At me, Adam Smith, chair of the AIQI consortium. Each episode on AI Standards Stack, we discuss developments in AI assurance with guests from around the world that are leading the charge on the standards, ethics, and regulation of AI. This episode of the AI Standards Stack is supported by the United Kingdom accreditation service, U-CAS, the UK's National accreditation body. U-CAS ensures organizations that test, inspect, and certify AI systems are technically competent, impartial, and robust, building confidence in safe, reliable, and trustworthy AI. On the show today, we're joined by Mike Thimi, managing director and cloud advisory lead at Accenture. Mike previously worked as the CTO and Senior VP of Nevada Solutions before his acquisition by Accenture Federal Services in 2021. He leads the development of cutting-edge AI capabilities for public sector clients, focusing on responsible trust-worthy design and real-world impact. Well, welcome to Show Mike. It's really a pleasure to have you on. Just before we unpack the stack, could you say a few words about how you got to where you are and why the thrilling and exciting world of AI standards is vital to your fulfillment in life? I think you're really appreciating part of this, and thanks for having me. How do I get here? Answering purely from a standards point of view, my first entry point into the standards community was well before we had AI standards. This is in the biometrics space. Imagine a guy at the startup who wants to make his name in a technical field. So, circa 2005, I got involved in biometric standards. Things like, how do you tell that a fingerprint is fake? How do you measure or not whether a fingerprint actually matches? Fast forward over the years, as AI started to emerge and become standardized, it was a reasonably easy mapping of skills from the biometric space of AI space. Since my day job was doing AI full-time, delivering for clients, it was kind of a logical transition to get involved in the standardization space as well. So, that's all I got here. And as we talk today, I'm sure that more vignettes will come out as far as I've got to where I am, but either way, that's my abbreviated version of my story. So, yeah. And in the sense of the listeners are given quite a bit of experience and also the kind of idea of doing stuff, could you give us what message? What's the one big lesson you've learned? We'll come back to it at the end, but I'd be curious what you think this audience ought to pick up. Right. So, I think the one thing that I might be able to bring that is a little bit differentiated is that when I am engaged in the standards work, I am fully engaged and then when I'm not, I'm completely not. So, I'm able to look at this, I think, with a very much of an independent vibe as a bit of a third party. And I think the main message is about consumability. Right. Out of somebody who is extremely busy their day job, who has the bereavement that's a week to think about how standard, by the impact of them, how do we give them actionable steps, processes, mechanisms, tools to be able to make the best use of time. I think sometimes as folks who, there might be folks who live and breathe this, who might have their full attention dedicated. So, be mindful of the consumability of our outputs, I think, is the main message. So, hopefully as we talk through the topics today, we can focus on examples where we've had to bring that to the foreground. So, yeah. Now, you mentioned that you were so close to the sector coming from biometrics and you were clearly, I think I know, using various tools and techniques on pattern recognition images, predictive analytics, sort of stuff. But could you give the listeners just a quick clue as to what it was like to get your fingernails dirty at that time as opposed to the generative AI LLM period? Yeah, I'm so glad you asked because, you know, we, in the biometrics phase, I feel like we were not by design, but almost we accidently like at the vanguard of when Gen A came to being. And so, my, one of the most interesting, sort of mission application that I had in biometrics was developing, let's just call them synthetic face images, right? That's a diplomatic way to put it right. Synthetic face images are lots of different use cases you might be able to imagine, right? So, I saw a GoPro or some of back in the old days, right? We were using, you know, pretty brute force machine learning techniques to essentially reverse engineer pattern matching algorithms to develop little textures you might embed into a face, right? Not using the sort of holistic view and holistic models of face. And so, we had this work and this stuff came out from Nvidia. This is, you know, several years ago that was automating the ability to, you know, essentially adversarially generate these kinds of faces. And this was in the vanguard delivery and we had to transition from old school to new school general techniques, right in midstream as they are emerging, literally taking libraries off the repos like as they came out day by day, right? And I do think that the, the first sort of mainstream physical application of Genai was actually in the biometric and the face recognition space. That's where approved it's sort of metal. And then, you know, text and everything came later, but that's kind of how I lived it, right? And so as a delivery team, we had to think about the transition from traditional to generative models and deal with like the way that they came to market, which is very different than how the judicial stuff is available. So a lot a lot in there, but that that's kind of how it, yeah, I'd sell quite kind of high live through that process. So I'll start there. That's really interesting. And biometrics is a really interesting space because you look at that space and you see validation data sets being held out by test labs in a way that you don't see with other types of AI. So I think biometrics is a little bit, it leads the way a little bit, not just in standards, but in the whole metronomy side of things as well. That's a good thing about it. Yeah, a lot of the way that we think about measuring AI systems, it's not one to one, right? Because the trade space is different, but a lot of really groundbreaking work was done in performance measurement in the biometric space to be exact reason you described that, and that some of the four four fathers and four mothers of the space had the foresight to think about withholding data sets and standing up really rigorous test regimes that advanced the industry, right? Both open source and proprietary tools. And so you're right, the biometric space, I think, provides a model for how we can think about the best way to test accuracy and test performance for like these other emerging techniques that might be in multimodal or in-text data or spaces like that. So it's a great observation. Now one other thing about your experience is you've you've actually worked in both the private and the public sectors. You've done extensively work in deploying AI in federal governments. And I'm assuming a lot of the biometric work was done with private firms, but perhaps on federal contracts. Right there. It was all kind of a mixture, you know, the preponderance of the work was in the federal space, but when you're in that space, you're in every work with vendors and there's a lot of bleed over into commercial applications. Yeah, that's definitely true. Now there's been, you know, many sectors in many areas, just forget AI and it could be food, safety or something. Governments tend to think it's for everybody else and not for them. So how effective is governance approaches and standards approaches within the government environment? Yeah, it's a very complicated and like multi-dimensional question. What I will say, you know, is the challenge that we face and I think this is the last year or so, it's a very different landscape. If we were right into the sort of emergency of GNI, you know, there was, it is natural for government acquisition authorities to be very cautious, right? They tools have to go through rigorous security and safety and compliance regimes, which might at the time seem very frustrating. Like why can't I just download and use this thing that's really cool and awesome? The reason being is that, you know, the government has an interest in ensuring that the supply-shaded secure and, you know, reconciling the supply chain validation rigidity with the extraordinarily rapid emergency tools, that's a very hard thing to reconcile, right? And I feel like the, how to say it, it was beneficial, I think, that the government was slow to adopt some of these capabilities, because we've seen, I think some of the really challenging supply chain, supply chain uses that we've had. I think we've seen some of the adversarial uses of AI. And so I think at the time, we might have been frustrated that the government couldn't move faster and couldn't authorize things faster, but there was a bit of a method to the madness there. I do think right now where we are, as I think we've struck a good balance of expedient like authorization and ability to use AI tools in these spaces without being so rash as using stuff that hasn't embedded. But it is a really dynamic space. And my last thought here, and this is, you know, maybe to reflect on the biggest difference between my initial experience and the way things are now, is that like the rapid and like almost universal availability of tools on almost a day-by-day basis is something that we never I've never had to think about before. We thought about tooling or something that we changed over the course of life.
like several months, like back when you were working with these tools and government agencies. And now that cycle's gone from months to weeks, almost down to days, right? And that is a big, big thing to reconcile in the government space. So lots of time back there, but that would be my macro level observation. - Well, that's really helpful. I know there are a lot of issues, but it sounds like you're feeling, it's nice, I just optimistic, you're feeling that there is a balance. - There's a better balance than there was. Yeah, I think there's a realization that the prior regimes cannot be reconciled with the way the tools are coming out, but by the same token, it's inappropriate to just like download and run the coolest new thing because it's there. And so I think find that middle ground is that's what those auditability regimes are there for. So yeah. - And just a quick one for listeners, obviously you're heavily involved in the ISO standards with Adam and all, but are you, do you have international experiences, most of your experience US-based? - So it's right, my work with Adam is sort of at the ISO level and through SE42 where the artificial intelligence, but in order to do that work, we might all know, but at least at the US level, we have to execute our responsibilities within ANSI, right? So there's kind of a mirror set of responsibilities that are run domestically in the US. Luckily, the vast majority of the work that we do in the US at the ANSI level is designed to mirror and empower our work and ISO, meaning that we're not developing a set of like shadow standards in the United States that are going to diverge from the ISO theme. We've sort of taken the approach that ISO has a set of great technical experts with a lot of capabilities for leading the way and we should cast our lot with that group of standards is best way to put it right. And so, yeah, that's also, I'm sorry, I think that's the best way to think about it. So yeah. - We might come back to the ANSI international borders and what's happening, but what we do that, it might help our audience, you've led and participated in multiple AI standards developments projects and I understand most of them are quite technical, but can you tell us a bit more about those projects and why they had to sit you and maybe what was achieved? - Yeah, I'll tell you about a couple and it's why before we've always had a shadow and actually brought up a couple of different screens because like the titles of these sometimes change and sometimes it's better for me to read things out and so if you're like, why is he reading the screen it's because like the wording gets so intense, but so without them, we've had the good fortune to spend the last, I guess, a couple of years now collaborating on the progression of a standard in verification validation. So we're part of a joint working group on AI software testing. So one of the projects that I helped lead is called Artificial Intelligence Testing AI, Verification and Validation Analysis of AI Systems and the genesis of this work was that, you know. - I can see why you need to read those sorts of things. - Exactly, right, we've been curious, really. VNV is one of those, you know, nor stars for the way that people think about software testing and there's an entire like philosophy and there's a centrality of like a VNV way that people think about things, right? And, you know, there was a gap in the landscape around VNV analysis that needed to be closed off and I inherited this work, I think maybe two years ago from a project editor who didn't have the time to work on it. And so I jumped in and said, well, this is something we can advance. And so with Adam, you know, leading that joint working group, we're able to identify sort of existing mechanisms that we could bring into the VNV landscape to avoid sort of like reinventing the wheel. And so I think it was a really good initiative to help bring and reconcile work in SE42 on things like testing and quality assurance, bring it within sort of that software testing regime and sort of square that circle. And that's part of a multi-part standard. It's going to wind up having, I think Adam, it might be up to maybe eight parts right now. So it plugs into a larger ecosystem, a larger ecosystem of testing approaches that an organization can sort of choose from that sort of best fits their needs. So that's the first piece. And Adam, if you want it's way in there at all to paint a bigger picture of that series you can before I jump into the other one. So I'll stop there for a moment. Yeah, I think the bigger picture is that the systems and software community and standardization, which has a long history of producing ISO, IEC standards, has set up a whole series of processes around software and system testing and software quality. And one of the things we've done in SE42 is to try and build on that and to look at, what is the difference in the quality model with an AI system, one of the differences in the different testing approaches and try to unite and align with that existing body of work. But equally in AI everything is new. So what we, what some people would call testing, they call e-vals if they're data scientists or adversarial red teamers and things like that. So working through some of those differences has been fascinating and is an ongoing project. And thanks for your contribution, Mike, with this part, we should be published this year. I'm pretty sure. Yeah, we're very close, right? And yeah, it's been eye opening. I think the macro lesson here, and I think coming back to my theme about if someone just said three minutes, think about this, right? If you just have three minutes a day, like you expect that these questions of like, what's an evaluation versus what's an assessment versus what's the test? Like you just need that to be answered for you, right? And so, you know, Adam is leading a multifaceted initiative to try to get alignment on that sort of basic view of the way that we're going to be assessing AI to be able to make it actionable for the senior makers. So a lot of work is done behind the scenes to synthesize these different viewpoints and the brain of together. So it's very challenging. The other piece that I spend most of my time on, and Adam, you'll be glad to know that I plan to block my calendar tonight to get this thing ready for its next, it's literally ready for its next voting cycle. I've cleared out my Thursday night here in London. But it's in the area of performance evaluation, performance measurement, actually, is the right word, performance measurement for AI classification, regression, clustering, and recommendation. So if you think about all the things that AI does that is maybe not germative AI, like all the traditional things that you do with AI that you see presented to you hundreds of times a day, like AI recommendation systems that are operating like right now on every tab I have in Chrome, there's some AI recommender is doing things for me, right? And there's some AI clustering tool that's working behind the scenes to essentially to understand what kind of profile I am. Each of those areas has several different ways of measuring effectiveness, measuring outputs, measuring accuracy, and the community needed an integrated, consolidated, and consistent way of expressing performance in all these domains. It's stuff that has been in some cases sorted and understood years ago, but never written down in a standardized way. There's other ways that are kind of even and now right now we're emerging, right? New ways of measuring and assessing this space. But what this standard is going to do is to basically put it all into one consistently presented umbrella so that if somebody asks, how can I tell how well my classification tool is working? How can I tell who to go to for a given clustering algorithm? This will be essentially a one-stop shop for all of those metrics, all those methods. And it doesn't, importantly, it doesn't only pick one method, it gives us a variety of different methods you can use depending on your use case, depending on whether you're more concerned with precision versus recall, depending on whether you're driven by cost savings versus depth of analysis. But this has been a long time coming, and we're at the point right now of having a stable document with a lot of great technical material that is going to be ready as a reference point for consumers. So I'll stop there. - Well, actually, I'll pull you on on that if I may. - Yeah. - It's absolutely very interesting. We had a project over here some 32, 3 years ago called StatLog. I don't know if you ever came across that. That was led by David Spiegelhalter at Cambridge. It was actually an EU project. And we took 20 different classification methods and something like 22 or three different data sets. So heart disease, credit scoring, image recognition, and just ran batteries of tests and found that the metrics were really difficult. I'm currently in a project which is text-based and we are desperately trying to figure out the metrics in this area for an application to do with quality of qualitative answers. So it's really a nightmare. And you clearly think that reproducible and repeatable measurements are important and they're important for AI governance. But what are the implications if you've got to an area where the AI seems to be doing good stuff but you just aren't really unable to get a handle on a solid metric. What does that imply? - Yeah, it's a, you know, it's, the challenge is either things you write if your results aren't reproducible and they're not deterministic then like how are we, like what do we as evaluators do, right? How do we approach that space? And we should, I think we need to be comfortable with more
shades of gray and assessing the outputs of these algorithms, right? Like it's clear that an algorithmic output that is like actually incorrect, it represents complete inaccurate data like that goes into one sort of class where one quarter like one sort of like assessment class where clearly this is an unacceptable response, right? But the sophistication of the tooling that we're using now does allow for responses that are, you know weak but usable, right? That are, you know, defensible but you know not perfect and so learning to, you know, adjudicate responses without a continuum, right? And I think that's something that we as a community are going to have to deal with and sort of like be able to accommodate, right? One way to think about it is that we know now that the tooling that we're using in AI essentially, you know, can be a calibrator, like we're going to put it can be utilized in a way that literally will give you more useful, more accurate, more false and answers as long as you're willing to, you know, turn the dial up and use more tokens, right? So I think our our ability to assess these tools needs to be mindful of those kind of trade-offs. I do think that this idea of having like trade-offs built into your assessment like regimes is not new like we're doing that for a long time, like we've been for quite a while trading off elements of accuracy, we've been trade off elements of what we would call in the old days false positive false negatives, right? Of those sort of depths of searches. And so I think it's applying that same sort of mindset into this space. So it's is what's necessary, but it is it is very complicated, it's very dynamic. And I think that we as testers need to need to meet the industry where it is, right? And not try to make it fit into maybe previous regimes that we're not made to we're not we're not built for it. So I'll stop there. One of the one of the most important things about these standards that the divine metrics for evaluating correctness is the huge amount of downstream value. Because then standards that look at things like bias and robustness can build upon those metrics and claims about AI systems or products can be conformity assessed. You have you consumer groups sometimes talk about getting better information to consumers about the different aspects of a system. And without really defining the measures in a reproducible way, it's hard to have any kind of credible ecosystem around around those claims. Yeah, and that's a great thing about it. You know, the if you think about right now that the public's experience of AI is, you know, maybe asking open-ended multi-faceted questions of these models, right? And these models give back multi-dimensional responses, but like those responses can oftentimes be decomposed right into different sub assertions and that decomposition. Like those can be looked at in about kind of a granular granular way, right? Like there might be 11 elements of their response of which seven actually can be assessed through these methods that we're developing. And there might be four that are more completely open and is subjective, but you're Adam that that decomposition is what's necessary there. And we think about reproducibility, you know, it might be that, you know, elements of AI system interactions are going to be reproducible and we'll put those there. We'll measure them in a certain way and others aren't, but you know, we we as a community need to be ready to like meet the models where they are and to be able to like work with them and work with their outputs in the way they are in order. And I will say in order for our work to be like usable by the wider community of like we need to like there needs to be that like adaptability and the approaches that bring to bear. So yeah, I'll stop there. And one interesting thing is the separate measurement standards underway for sort of text problems like you were leading to Michael, we've got joint projects with the linguistic standardization community looking at methods for evaluating text. And we've got a similar initiative underway in Europe around computer vision. So huge amount of work to get standards like this in place covering all sorts of different AI systems. Yeah, it is it is very dynamic. Yeah, Adam, to your point. Yeah, it's a lot to keep up with. Yeah, a lot to boil down. If you're trying to present your leadership on like the house and why is it's a lot to summarize. So yeah. Okay. Let's let's return to that border between ANSI and ISO are between the US, maybe an ISO. You said that ANSI is you know, positively following, you know, what it learns from ISO, but are there actual points of fundamental difference? Are these different schools or is it just the usual thing that you start nationally, you start discussing things internationally and it bounces back and forth and solely congeals. Yeah, so I don't want to speak out of school here. I'm just a small element of the overall system right like that. Workbrook Center, Accenture is part of, you know, we represent up into ANSI, but Accenture itself has, you know, this massive global footprint. So this one one person's opinion, right? So I think that, you know, we. Like, fundamentally, we all have the same interest, right? We want to publish and propagate high quality standards that businesses can use, the carpenters can use, right? There's, we all want the same outcome, right? And some of us, some stakeholders like have a very strong business need that certain elements like it sorted very quickly, right? Like other like if you're a technology provider who's developing, you know, certain types of models, you know, very strong business interest to get to that, like, finish line as soon as possible. I think that from my vantage point, right? Like, what I would say is that I think everybody who I, the vast majority of people who I interact with, right, are really anchoring their decision making on tying it back to some type of clear and expressable business. The right like the standard needs to come back and needs to be usable and needs to be like valid in a business context, right? And so I think about the work that we do in the US, like when we think about, is it worth spending time on the standard? Is this something that we support and want to spend energy against? Everybody asked a fundamental question. Can I go to my boss and justify the time I'm spending on this fight or is it and if yes, then we as a community want to get behind that. We'll have experts weigh in. We'll spend time on it. If it's something that is not relevant to our business leaders and our organizations be at Accenture or a cloud provider or whomever, then we just won't spend time on it. And if the interest capability wants to do that, it's fine. So I think ante has is like strongly aligned with things that are going to address a business need. And I'm sure for Adam, like when when Adam knocks on my door and says like we need this publication out fairly soon, it's because there are stakeholders right with a strong business need who need to have these publications up and ready and available, right? So working backward from business need, I think is a great anchor for all of our activities, both at the NC and the ISO level. So I'll stop there. Okay, so I'm going to move on to some other non AI stuff in a minute, but you know, you're at a cocktail party and there's an attractive person you're trying to chat up and you open up and you say I'm involved in standards, AI standards. What's the chat up line that gets them to stay anywhere in your vicinity? Well, I would probably never use like that when I said before at the top of the call that like the by I'm able to sequester my standards work in a way that makes sense. I am very judicious about whenever you know, but like for for those like, you know, the example that I gave like to make this practical right like it was an awesome privilege to be in the room. Like I remember being in Italy right in the beatings happened to be in Belonia right where we were working through the specific language of what it meant for a biometric use case to be high risk right and we were tuning in those exact words right and what I would say is that it might sound boring, but dialing in the exact language that either classifies an application as high risk and therefore subject these controls are not like someone actually has to do that work. And so many to dial in the wording correctly so if you're the kind of person who loves precise language right and who really enjoys being able to express things very very clearly like that's a kind of person who I take this into who I take this you right if it's somebody who doesn't care a lot about precision awarding I wouldn't bring it up, but like it is really gratifying to be able to work with the set of stakeholders across languages and to be able to dial in language that we can all nod to and say yes. We've expressed that concept correctly and precisely and consumable and we're doing right by the community like that that's that is it sounds weird to say that she's really it's exciting and gratifying to be able to do that type of work just at a like at a human level it's really it's awesome to be able to achieve those outcomes right so that's how I would phrase it's a lot of words there but that's what that's a theme I would be able to write there's it's a gratifying experience to be able to dial in language in a way that's usable. Good well that was a bit of passion there I might buy a drink yeah now just moving away directly from standards themselves looking at the wider system of human interaction. I mean human in the loop is thrown around a lot it's kind of the throw away answer for things in the defense sector will always be a human in the loop I be curious what does that mean in your experience and is it is it really true that human oversight adds value in generative AI systems or is it more of a bottleneck than a safeguard or a more of a fig leaf than a than a reality. you know everyone happens or I'm choose to this but I will you know
know, this is an area and I'm not here to, you know, pitch for Accenture or Accenture to be company and we can do all our own right. But what I think the way that Accenture has attempted to reframe this narrative is to not, is to try to get away from the idea of human loop and to reposition the narrative as human and the lead, right? And like without just drinking the cool aid of like, you know, taking central language, I think that's usually to think about it, right? That like, there actually, there's a necessity for human leadership when we're thinking about like initiating autonomous systems, right? If we think about kicking off a multi stage, agentic process, it's going to run in the background for six hours, it's going to invoke lots of different tooling that's going to like run into different subsystems. It's going to invoke lots of sub processes like that entire like mechanism is a human lead process, but the human is not going to be in the loop in the vast mudo as interactions, right? The humans going to expect that the system escalates to it, certain decision points, what it needs intervention. And so I think we need to think of it as not the human in the loop as a peer to a bunch of other AI systems, but the human as the, you know, at the top of a hierarchical pyramid initiating and adjudicating and authorizing decisions, right? And I think that's a good way to think about it. What does it mean for a human to leave a set of AI sub systems? And as a leader, you would ask yourself, it'd come to me back to that three minute idea, right? As a leader, you need to know when to be left alone and when certain decisions have to be brought to you, right? And working through that trade space of when is it appropriate for like AI decisions and routing to be escalated to a human leader? I think that's a really good and usually to think about it. So I'll stop there. So I'm human on the loop. I prefer to talk about human oversight, which human in the loop is kind of one approach, just as having a human leader is one approach, Mike. And for me, when I've worked on that topic in standards, it comes down to, you know, what is the question is, what is the reversibility of any harm that may occur if a human doesn't intervene or doesn't approve outputs and put some kind of framework around how you decide whether you have effectively retrospective oversight of a system, whether you have a more involved process. And I guess there's a new generation of this, which is a human lead process, which is got me thinking, Mike, but there's a lot of conversation about this topic in standards, but it's more focused around oversight with human in the loop being just one option. Right. To go ahead and think about it, yeah, human in the loop, I think is a, you know, we want, we need to reinforce the primacy and the centrality of human judgment into all these processes, right? And it's unrealistic to think to be human as an adjudicate every part of the subsystems, right? But oversight to go ahead and think about it, oversight or leadership or like final like authority is what's necessary here. But like this is there, it's a really fraud topic and, you know, the standards community, I think one way to think about the role, as I've been working through this space, right? The, you know, we might all be frustrated that the standards process moves at a, at its own pace, right? That even the fastest publication will take, you know, 12 or 18 months, you know, to the door, but there's a real benefit to that, that slowness into that, you use a word, so there's been a fit to that ace of development, which is that you can step back and you can like move away from the immediate noise of what happens in the news day to day. And you can ask yourself, like, what is the right way to situate the human in this process, right? In a way that's applying good judgment, that's applying like, you know, a critical ins here. And so, you know, that there's a method to the badness in how these processes take a long time, right? Because it allows us to be thoughtful and to position the human in the way that's most appropriate and not just be reactive to what we're seeing, maybe in the most recent, sexual development. So, well, actually, that's probably a good place for me to throw in kind of my last question, which is, you've been in the field a long time, as have I, and we've seen the pace was very, very slow in terms of deployment. Then we've seen the immense enthusiasm by a lot of people and then some of the frustration at the slow pace. But at the same time, we're seeing a huge proliferation that I don't think any of us anticipated the speed at which it's happened over the last three and a half years. You know, so lots of a generative and a gentick AI, getting more embedded in public services, private services, real decision making. So, we're now seeing the leading edge of a lot of stuff that we've led on. Do you see any notable gaps between how AI is being deployed today and thoughts on good governance? It is very difficult, it is difficult to reconcile the perceived need to keep up with like what your competitors are doing. And, you know, if I work for a consultancy and I'm concerned that another consultancy is adopting tools more quickly and they're gaining efficiencies like reconciling that sense of like self-induced urgency with the need for like governance and pace. It's a, it is a very difficult challenge. And I think that the pendulum swings, right, where you have, you have maybe periods where like the race to adoption, like becomes a bit break deck and the pendulum swings and the public perception is that we need to move more slowly. And I think like leadership needs to be able to have the belief in itself to swim against the current. If there's a time when the industry is racing to adopt and push out models very rapidly and there's a race to embracing quickly. As a leader, it's probably time to pull back and to err on the side of, you know, caution and governance and waiting out a cycle. It's okay to wait on a cycle because there'll always be a new one. I think that's a good, that's a good, if I think about again coming back to the theme of like presenting to leadership and the value of standards. I think the value here is to be able to take back a discipline message to your leadership that is willing to have the courage to swim against the current tide, right? And then in the best interest of a business that might have a bit of a longer term view. And I think standard supports that longer term view that's, that's really necessary as a counterweight to the things that happen on a day to day basis. So a bit of a long answer. But I think that that's how I view like the way you've cut up that question. So yeah. And the last comment from you? Yeah. Um, I wonder Mike and this is just asking you as an expert really here. Where do you think standards might go next in terms of a genetic AI? Yes, so I think, well, I think what's most necessary and this will be a boring answer, but I think we'd say we have it into all of us that we lack a basic grounding in terminology, right? And one of the awesome things that standards can do is to get us to agree on some basic sets of terms. And once you begin the terms, then you're 70% of the day there. Think back to the biometric community. My first real battle was defining, hey, what do you call a fake fingerprint, right? And by some like an easy question, but when you dig into it, it's actually very, very subtle. And so I think Adam just doing some basic blocking and tackling on terminology and interchange, like language will actually help disambiguate and clarify what's happening here. And I know that work is happening in some different WGs, but that's the first step, right? Is language in terminology alignment? And then from there, we can build out the build on the infrastructure. So I'll stop there. Well, that is a really powerful place to stop. I mean, I've picked up quite a bit today. So thank you. Interesting. Your point about the business led standards, which in way is self-self-policing in terms of effort, time spent, et cetera. Second thing is the point you just sort of ended on the importance of actually having a defined vocabulary, a tax
onomy for that vocabulary and the precision that comes from that that allows us to develop a better standard as well. I loved your comments on the measureability where there are measures that are weak, but usable. And that's a very important part. And finally, and I think for all of us in the standards community, the issue of, yeah, well, most people are doing this three minutes a day. So have you boiled this down to make it easy for them? Those were all very, very good points. So thank you, Mike, for sharing your valuable insights today. And thanks to my co-host, Adam Leon Smith. My thanks to all of you listening here on the show. And please join us the next time when we cover the full stack of AI standards, ethics, and regulation.
Podcast Summary
Key Points:
Mike Thimi, managing director at Accenture, transitioned from biometrics standards (circa 2005) to AI standards, emphasizing consumability for busy professionals.
He led work on synthetic face images using early generative AI techniques, highlighting the shift from traditional to generative models in biometrics.
Government adoption of AI has been cautious but has achieved a better balance between security and rapid tool deployment, with supply chain validation remaining key.
Key standards projects include AI Verification and Validation (V&V) analysis, integrating AI testing into software testing regimes, and performance measurement for classification, regression, clustering, and recommendation.
The performance measurement standard aims to provide a consistent, one-stop-shop for metrics, addressing challenges like reproducibility and deterministic outputs in AI evaluation.
Summary:
In this episode of the AI Standards Stack, host Michael Minnelli and co-host Adam Smith interview Mike Thimi, managing director and cloud advisory lead at Accenture. Thimi shares his journey from biometrics standards in 2005 to AI standards, noting that consumability of standards outputs is crucial for busy professionals. He describes early work on synthetic face images using generative AI, which pioneered techniques later applied to text and other domains.
Discussing government AI adoption, Thimi notes initial caution was beneficial for security, and a better balance now exists between rapid tool availability and rigorous authorization. He highlights two key standards projects: AI Verification and Validation (V&V) analysis, which aligns AI testing with existing software testing frameworks, and performance measurement for AI classification, regression, clustering, and recommendation—a comprehensive reference for metrics. Thimi emphasizes the challenge of reproducibility and deterministic outputs, advocating for comfort with shades of gray in evaluation.
The conversation underscores the need for actionable standards that bridge technical rigor and real-world applicability, with the biometrics space offering a model for rigorous testing regimes. Overall, Thimi’s insights stress the importance of integrating standards with practical deployment, especially in high-stakes public sector environments.
FAQs
The AI Standards Stack, hosted by Michael Minnelli and Adam Smith, discusses AI assurance developments with global guests, focusing on standards, ethics, and regulation of AI.
Mike Thimi is Managing Director and Cloud Advisory Lead at Accenture, formerly CTO and Senior VP of Nevada Solutions. He develops AI capabilities for public sector clients, emphasizing responsible design.
He started in biometric standards around 2005, focusing on fingerprint and face recognition testing. As AI emerged, he transitioned to AI standardization, leveraging his biometrics experience.
He emphasizes consumability, urging standards experts to provide actionable steps for busy professionals, rather than assuming full-time focus on standards work.
In biometrics, early work used brute-force machine learning for synthetic face images. Later, generative models from Nvidia automated this process, requiring teams to adapt rapidly as tools emerged.
Government agencies must balance security and compliance with rapid AI tool emergence. They initially adopted slowly, which proved beneficial, but now seek a balance between expedient authorization and safety.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.