Go back

Steph Holmes on the Intersection of AI and Compliance

30m 32s

Steph Holmes on the Intersection of AI and Compliance

The discussion between Tom Fox and Steph Holmes explores the current state of AI in compliance, emphasizing a shift from speculation to evidence-based implementation. Holmes highlights EQS Group’s AI performance test, which evaluated frontier models like OpenAI’s GPT-5 and Google’s Gemini 2.5 on 120 compliance tasks mimicking a compliance professional’s daily work. The test revealed that newer reasoning models outperform older ones, particularly in data analysis, but reliability varies significantly by task. A key concept introduced is the "messy middle"—compliance tasks requiring human judgment, context, and interpretation (e.g., assessing retaliation or complex conflicts of interest), where AI is less reliable. Holmes stresses the importance of human-in-the-loop oversight as a control mechanism, especially for high-risk or irreversible outcomes. She advocates for task-level testing over vendor-level claims, as AI performance differs across specific tasks, and vendors should demonstrate which models excel at which functions. Continuous monitoring and risk-based review are crucial to maintain accuracy, as larger datasets can reduce reliability due to token limits. The conversation underscores that AI can efficiently handle high-volume, low-risk tasks (e.g., case categorization), but human oversight is essential for nuanced, judgment-heavy work. Ultimately, the goal is to use AI responsibly, sustainably, and explainably to multiply compliance resources while mitigating risks.

Transcription

4814 Words, 26303 Characters

English
Today I'm absolutely thrilled to have Steph Holmes, Steph and I have worked together for many years back when she was with Converson, but she's with EQS now and here to talk about the current state of the intersection between compliance and AI. What is the role of artificial intelligence and compliance? What about machine learning? Are you using chat G-P-T? These questions are about three of the many questions we will explore in this exciting new podcast series, Compliance and AI. Hosted by Tom Foxx, the award-winning voice-up compliance, this podcast will look at how AI will impact compliance programs into the next decade and beyond. You want to find out why the future is now joined Tom Foxx on this journey to the frontiers of compliance in compliance and AI. This podcast is a production of the compliance podcast network. And now Steph Holmes and Tom Foxx. Hello everyone, this is Tom Foxx and I cannot tell you how excited I am today because I finally have Steph Holmes for a podcast. It's a long time friend of mine, my favorite people in all of compliance. She's an EQS now and she can tell us all about that. With an incredibly long-winded introduction, Steph, welcome to this episode. Thank you Tom. I'm excited to be here. We've talked about this for, I don't know, probably going towards 10 years. So I'm excited to finally make it happen. So tell us what your role is now, Steph. Yeah, currently an EQS group, which is the current home of a former conversant if you're familiar with that. And really what I focused on now in a QS group is to ensure that we're bringing our compliance and ethics cockpit to life. Right. So whether it's products from speaking up all the way to third-party risk management, bringing those together in a way that you can feel confident that you're digitizing your compliance program. Moving it forward and ensuring the best possible results for your organization. So that's really what I'm focused on now and obviously AI is a big topic and I think we're going to get into that a little bit today. Yeah, so this podcast finally got catalyzed by two pieces, one, your article in CCI, where's the loop testing AI across compliance task and then EQS is AI performance test report. Less set the stage, where do you see AI and compliance now? Yeah, so AI and compliance, I think for probably two to three years, it's been the top of my discussion at every industry conference, a lot of the continuing education. And I think up until now, really, it's been a, where is AI? What does it mean? Should we be part of an AI committee? How can we introduce it into our program? And a lot of it, I think, has been driven by assumption more than actual practicality because it's been such a new emerging technology and really where I'm at on that journey, partnering with many of my colleagues as well as BCM out of Germany. We were like, how do we take this to an evidence based? How is AI performing in compliance for compliance tasks and giving some practical takeaways for hopefully many, many in the industry where they can take this internally at their organization and understand how AI can make an impact in their organization or in their compliance department? So what has this rapid acceleration of AI development? How's it reshaped expectations and maybe we'll move from expectations to some more tangible questions down the road? Yeah, so as we think about the AI acceleration, I think where we're getting into now is, and what I've seen this year is a lot of executive teams, CEOs, they're having pressure from board saying there's a lot of regulation changes, a lot of tech changes. And what I think is the economic pressure is more real than effort this year. And we're having to figure out how to do more with less. AI is one of those tools that we can explore how we can really multiply our resources within our program. And it's not just this future tech where we need to think about how are we adopting AI as an organization. And a lot of leaders are saying you must adopt this in your program. And so as compliance professionals, it's our job now to even say how do we do that responsibly, sustainably, and explainably. Right. So being able to showcase how it's working, why it's working that way, and how we've put a program in place to ensure this proper oversight. Because as with any emerging technology, there is going to be concerns on data privacy. There's going to be concerns on the reliability, how well is it performing. And so where we really wanted to support the compliance profession is not to fear it, but ensuring that you can use it effectively, responsibly, and impactfully your organization. Steph, how did the first of its kind performance test evaluation really show a gap between not only performance and promise, but the best and more systems? Yeah. Good question. When we think about this, what we did is we modeled most of the kind of frontier AI models against 120 compliance tasks. So what we tried to do is make it reminiscent of an organization day in the life, what are the typical types of things that you're doing as a compliance professional. And so when we're looking at that, we wanted to ensure that we had a good swath of examples across that, but also the models we tested. Big ones that I would say we got most of the results from are going to be open AI. So those are going to be some of the chat GPT models. And we tested a variety of those. Because those are the most kind of common that we see. And as a side note for anyone that's listening, open AI and the GPT models are what is power in your co pilot usage today as well. So if you're using that internally at an organization, I've seen not be the first things that organizations are adopting. It's actually powered by open AI. It's not powered by a Microsoft AI model. And we also tested and thropic and we also tested Google out of all of them. What we were able to see is that the most reliable models are the ones that were released this year, which typically have expanded into the reasoning realm. Now reasoning as far as human reasoning, not quite yet, but the speed at which it can process and lot of information is really quick of the two models that really outperforms the rest. You had open AI GPT five, which is the most recent model. And you also had Google's Jim and I 2.5. So there are going to be more models that come out. And the goal is to be able to continuously test those against the same task to see how it's becoming more reliable over time. Because when we looked at this even a year ago. So the GPT four model, which was a 2020 2024 model. The results were drastically different, especially in some of those key areas that I know us as compliance professionals rely on. And it's like data analysis things like large volumes of data synthesizing it coming out with the output different things like that. As we think about some of these there, there are models that perform better than others. And I think it's really imperative for us to ensure not only us as we're adopting it as a technology provider. We're not just locked into one provider just because it's easy to manage one provider, but we're picking the best based on the tasks that's at hand. You have a great phrase, which is the messy middle. And I can't think about better phrase for compliance generally, but that's not what we're talking about. We're going to talk about the messy middle compliance in the intersection of AI. Could you first tell us what that is and then why you see so many failures in that messy middle. So the messy middle is this is this context that came up when we think about kind of compliance work where it becomes contextual or interpretive and judgment heavy. So human judgment heavy, right. So this these are going to be the areas like the gray area that we all talk about. We just kind of recoined it as the messy middle being able to interpret the root cause or the reason behind maybe someone operating under misconduct assessing maybe it's retaliation when you've had a variety of different cases coming in and you want to see. Have we experienced any retaliation at the organization doing that based on HR data sets that are coming in when you're thinking about complex of interest and understanding conflict has come in. What do we now need to do with that do we need to get more information to validate. Can we approve this can we not what is this really mean for our organization. And so it's really identifying what is material to the evaluation that we're doing not just what is mentioned so it's kind of reading between the lines it's understanding nuance and it's making sure we're also applying that as holistically as possible but also backing up when there is nuance that we need to make sure that we're looking at there. Think about AI performance AI is really good at those high volume low risk repetitive tasks. Then as you move up into the messy middle you're having those things where we need to have some judgment some interpretation and some of the contextual nuance around that where it gets a little bit more unreliable. And I would but even venture to say it's not something that we should completely avoid but we. We should think about it in a way that it needs more oversight than something that is a lower risk higher volume, more predictable in nature, that AI can really help operationalize with efficiency some of those processes. And that really brings you one of the key areas I did want to explore in depth with you. That's one of the most ubiquitous phrases in all of AI, the human and the loot. You alluded to that several times, you used the word oversight. But I really want to focus on the human component of this. I think you're absolutely spot on. You're also right to talk about the need for oversight. But how do we either, from the EQS perspective, help train or from my perspective, help educate with a blog or a podcast, our compliance colleagues to utilize this. It's simply just a trial and error. Is it a Q&A? How do we begin to exercise that human and the loop or that oversight you've referenced several times? Yeah. I think when you think about human and the loop, this is really where you want to start thinking about where is the risk? Is the risk going to be a regulatory kind of impact? Is it going to have large financial impact to the organization? Just like we would do any normal risk assessment, impact analysis. The human and the loop is a control for a lack of a, promoting into a compliance term. It's the control in how we're thinking about how AI is utilized. So even from the initial data entry and what we're looking at, we need to ensure that the data that's going into these models is actually reliable. We can trust it, whether it's internal data from an organization. If the data's messy, AI is not going to be a magic on to fix that data. We need to make sure that the data we're servicing into the models are reliable, are accurate, and there's little room for new ones there as much as possible. When it comes to the actual execution of whatever the task may be that we're asking AI to do, there are going to be low risk categories. So maybe I need to categorize cases coming in and get them to the right individuals of our organization. AI can do that pretty well, even without additional training. But then there's other task where it's critical because there's external stakeholder involvement. There's more regulatory exposure. Or if it's one of the things where if we do something and the outcome is irreversible, you want to make sure that you have human oversight in those areas. And so what we want to make sure that we're doing is not over relying on AI to be this magic one that can do anything. And even if we look at the outcomes of the AI benchmark, some of them were at 100% reliability. And that's fantastic. Those are areas like ranking classification that we can get really confident in how we're using. But other ones like content generation or creating a content or creating a training content, creating a policy, things like that. Those are great to get started with. The data analysis too, but you also want to make sure that you have human oversight before the output goes anywhere else. Content generation I think was somewhere between maybe like 70% that I could be wrong in my memory there. But if we're expecting AI to magic one, develop a bribery and corruption policy without the human in the loop, we could only expect that to be accurate, call it 70% of the time. Is that good enough for what we need? Probably not. So we want to make sure that we have human oversight there. So I think it's understanding the reliability for the task at hand. And it doesn't mean humans are 100% reliable either. That is the beauty of this industry is there is so much nuance, right? But where there is lower reliability, where things may get it inaccurate, where it may not pull all of the relevant information, it may not analyze all the relevant information. That's where we really want to make sure that we are staying in that loop, whether it's the initial data entry, but also especially if the output of that AI is going to impact other processes, right? We want to make sure that we're checking it at some of those points. Let me see if I could take this a little bit different direction stuff. In 2020, in the 2020 iteration of the evaluation of corporate compliance programs released by the Department of Justice, for the first time they said, "A chief compliance officer and a corporate compliance function must have access to all data across the company." And since that time, I think there has been an increased availability of data inside of a corporation for compliance officers. Coming forward till today, I always hear something like, "That's great Tom, now what I do." What I'm hearing are rather read in these reports. It's not exactly that, but it's close to mirroring that. I have this capability, but what do I do? Really struck me, let me see if I can get, I hope I got what she said, right? Which is, yes, you can iterate and yes, you can use the information. It's a continuous process. And that you don't get an answer and implement it. You get an answer and implement, train and then monitor. And with AI, you can almost have a continuous monitoring, and that would lead, of course, to a continuous improvement. But it's a process that really never ends. That would be a fair summary of what I think I just heard you say. I think it is a fair summary, right? Because I think as we start having a lot of these AI models take on various compliance tasks, you will absolutely need to ensure that there is continual review. But at the right times, right? If you are having some of these outputs by AI, for example, maybe you have AI that, like, I think I use the example of identifying or catarizing a case or a conflict of interest that you've received from an employee and you want to get it to the right individual, you can trust that pretty reliably based on the things that we've seen. But what you may want to do is periodically, you may want to review, how is it performing? Has it picked up any more information that actually made the AI go ascure? Because one of the things that AI cannot do is necessarily discern all fact from all fiction, either. And so if data has been fed into the systems or into the models that proliferates maybe something that you don't trust, that is something to be concerned about. So you do want to make sure that you are looking at some of these. If you have a lot of data from your organization, that's great. Compliance needs access to it, but it's not always going to be packaged up in a pre-package with a bow that people are going to be able to say, all right, AI, here you go. There's all of the data because there are also limits to data. One of the things we saw in the benchmark is that the larger the data sets, the less reliable, and this is oversimplifying it, but the less reliable the output could be. Because there's this concept of tokens, right, within AI. And that's basically how much, how many tokens can a model evaluate at any one time? And maybe you've examined this too, if you used AI and your personal life or whatever it is, the more kind of detail that you're giving at the more data that you're giving it, sometimes it has so much data and it needs to synthesize it into a small outcome that it puts in a relevant information there, right? It doesn't mean it's untrue necessarily, but it might be irrelevant. And so I would say that the continuous review is going to be important. If you are implementing AI into your program or into your organization, you've got to make sure that you also know and you're taking a risk-based approach on when do we need to have human review and involvement? Is it the initial data entry? Is it once the data comes out and maybe it's a decision-based outcome that we're looking for there? You absolutely want to make sure that you have some human involvement without just saying, all right, a gentick AI, for example, goes consecutively from task to task. You want to know how reliable it is at task one, task two, task three, maybe three, it's not as reliable as it could be in some of the first things you might be doing. You want to have review there, right? And then it can continue on, but you want to find where there are some lower performance in AI and make sure that you have some of that oversight there. I was going to ask you a whole series of questions from your CCI article, essentially around where and the loop, but you came pretty close to answering it there. So let me just maybe pick up on a couple of points. Why do you stress the importance of task-level testing rather than vendor-level claims? Interesting. Yeah, good question. So when you think about vendors, right, I think vendors are all in this journey too. If any vendor says that they have it fully figured out, they're not telling you the truth. And there's no, a lot of vendors today in the race to adapt AI have stuck with one model performer because it's easier to manage, right? You have third-party vendors. You want to make sure that you can have the due diligence necessary around that. But if you establish and start testing the different models that are available, for example, if I want to use Gemini for some of my data analysis because it seems to be more reliable than any of the open AI models, right? Part of the reason is actually because of the amount of data that can be processed by the Gemini- by Google Gemini model versus some of the open AI models. And so when you're thinking about some of those things, do you want to make sure that your vendor, if your vendor is promising AI, have they done the testing to show this model reliable than this model in this particular task? I think what we've also seen is chat us or probably probably the most common used AI out there, right? You and I probably also have already used it, probably once or twice today. And if you're using a chatbot and you're prompting it, the output you're asking it for, it's not gonna be the same reliability for every single prompt. But if you embed it into certain tasks within your tools or within your compliance program, not over arching, I need to review this conflict of interest and come up with remediation action distributed, et cetera. But if you're saying I need to categorize this, then I need to look at A backrisk, and then I need to gather more information from internal stakeholders, all of those different things. Each one is a task, but the end-to-end flow, right, is where it breaks down. But I think to your question, I think this does fall on the vendors very much to be able to prove the AI that we are putting into our products is used and delivered in a way that our customers can rely on. And if vendor can't explain, we use this model because it had this type of outcome compared to this other model that we use this other model for this task as it performed better. I think that's where we're gonna start having some of the issues where AI is a bit over-promised and under-delivering because it is not a magic one for every single thing that you need it to do. And if you just throw a chatbot on top of something and it open into prompts, you're gonna get fast and varying reliability across any of the prompts that any of your users may be having. - Steph, if there were three things, a compliance professional or CCO could do to start their journey, well, why do you suggest they be? - I think the first thing is, and this is whether you're adopting AI or not, but I would say the first thing is catalog your workflows, your compliance workflows that you have, right? If you have a risk assessment, if you've got a case management kind of investigation program, conflict eventures program, policy program, et cetera. catalog your workflows, what does it look like to start a policy from scratch and to get the reviews? What does it look like to update at the next year? All of those things. And so I think getting an inventory of all of the different things that your team members are doing, are there repetitive ones? Are they structured, semi-structured, or more ambiguous or open-ended? And just starting to document that, that would be the first thing I would say. And then if you're just now starting to use AI, pilot it on your first kind of structured task that you have there. And don't use this as outcome, pilot it, right? Run your human-led tech-led process side by side to AI process and have very clear scorecards of how did this perform reliably? Did it need all the expectations? Did it miss here? And from there, you're going to start understanding, OK, we can trust it here. We can rely on it here. Less review is necessary on the outputs. Over here, it was a bit unreliable, or it was a bit far-reaching and how it got to its conclusion. So let's make sure that we have some review there. And then when you do have those kind of tasks or outcomes that are more ambiguous, they're unreliable, that's where you want to have those human-led tech-points. So I think the first thing is just catalog, document, and define what all your processes are. And find the high volume, time-consuming ones, but maybe a little bit lower risk than you can apply it to. I would say that's the very first step. And then piloting them and then making sure that you know where the tasks that are ambiguous have external impact or regulatory expectation. That's where you want to have that human-led. Let me maybe flip the script a little bit and now ask, how do we-- and I say, we, I mean you and I, and our colleagues-- as compliance professionals need to evolve. I think we have to start using it. Using your personal life is just like any new technology that you start using. You want to start using it in your personal life. Your professional life start using it. I also think prompts are really important. So if you are using more of a chatbot style AI in your personal life, test out different prompts and see what outcomes come out of those. And one of the things that we did in the benchmark is ensure that we did it over-prompt because most professionals and most people in our journey aren't at the like advanced prompting level of our kind of experience yet. We're just learning this. So prompt on things you're familiar with and try different levels of prompts, try different levels of here's the outcome I want and see what it looks like for something that you are personally even familiar with. Because that's going to have some parallels to how you can use it even in your professional life, structuring emails, things like that, rewriting emails. Hey, this is the tone I want to set, different things like that. I would say starting small, starting your personal life, and then testing the stuff with your own knowledge, your own experience. Don't try something that's net new, that you have no idea that you can't even validate if it's accurate or not. And then also, yeah, we've got to make sure that we are documenting the reasons why we've used it, the reliability, so document who reviews what, when, how, and at what point? - Stuff, unfortunately, we're near the end of our time for this episode. But before we leave, I wanted to ask you, if our listeners wanted to connect with you, find out more about EQS, go to your article or the report. What would be the best places they could go? - Yep, feel free to email me, [email protected]. We've also got a website, EQS.com. And on there, we do have our benchmark report as well. So you can download that and get more information. It's got a lot of information there. It goes into more detail about the tasks and how we did the report. So it would definitely encourage everyone to read it. It is a great piece of collateral for you to take to your internal, responsible AI committee to show how you're adopting AI in your program and why you can rely on certain things. Yeah, reach out via email, go to our website, EQS.com. And I know Tom, you're also documenting that as well for the listeners. - Well, Steph, I can't thank you now for taking the time to visit with me. I hope we can continue this conversation. - Absolutely, thanks so much, Tom. - This is Tom Fox again. Thank you so much for listening to this episode of Compliance and AI. If you're interested in the cutting edge of AI for compliance or want to find out where AI will be with compliance in 2023 and beyond, check out my latest book, Up In Your Game, How Compliance and Risk Management Move to 2030 and Beyond. It's available on Amazon.com. I've linked to it in the show notes. I've really taken an effort to lay out the specific AI capabilities and tools for each part of a best practices compliance program. So if you're just getting into this area, it's a great start or if you want to take your knowledge up a level, I have the most cutting edge case studies on the use of AI and compliance. Check it out. (upbeat music) (wind howling) (wind howling) (wind howling) (wind howling) (wind howling) Compliance and AI is a special production of the award-winning compliance podcast network, the only network dedicated to compliance and ethics in the compliance realm.

Podcast Summary

Key Points:

  1. AI in compliance has moved from theoretical discussions to practical implementation, driven by economic pressure to do more with less.
  2. The first-of-its-kind AI performance test evaluated frontier models (OpenAI, Anthropic, Google) on 120 compliance tasks, revealing significant reliability gaps between promise and performance.
  3. Newer reasoning models (GPT-5, Gemini 2.5) outperform older ones, especially in data analysis and synthesis, but reliability varies greatly by task.
  4. The "messy middle" refers to compliance tasks requiring human judgment, context, and interpretation (e.g., assessing retaliation, complex conflicts of interest), where AI is less reliable.
  5. Human-in-the-loop oversight is essential as a control mechanism, especially for high-risk, irreversible, or regulatory-impact tasks.
  6. Task-level testing is more important than vendor-level claims because AI performance differs across specific tasks; vendors should prove which models excel at which tasks.
  7. Continuous monitoring and risk-based review are necessary to ensure AI outputs remain accurate and relevant, as data quality and model limitations can cause drift or irrelevance.

Summary:

The discussion between Tom Fox and Steph Holmes explores the current state of AI in compliance, emphasizing a shift from speculation to evidence-based implementation. 5 on 120 compliance tasks mimicking a compliance professional’s daily work. The test revealed that newer reasoning models outperform older ones, particularly in data analysis, but reliability varies significantly by task.

, assessing retaliation or complex conflicts of interest), where AI is less reliable. Holmes stresses the importance of human-in-the-loop oversight as a control mechanism, especially for high-risk or irreversible outcomes. She advocates for task-level testing over vendor-level claims, as AI performance differs across specific tasks, and vendors should demonstrate which models excel at which functions.

Continuous monitoring and risk-based review are crucial to maintain accuracy, as larger datasets can reduce reliability due to token limits. , case categorization), but human oversight is essential for nuanced, judgment-heavy work. Ultimately, the goal is to use AI responsibly, sustainably, and explainably to multiply compliance resources while mitigating risks.

FAQs

AI is being adopted to multiply resources and automate tasks, moving from theoretical discussions to practical use. Compliance professionals must use it responsibly, sustainably, and explainably.

It tested frontier AI models, like OpenAI's GPT and Google's Gemini, against 120 compliance tasks. The most reliable models were 2025 releases, such as GPT-5 and Gemini 2.5.

The messy middle refers to contextual, interpretive, and judgment-heavy compliance tasks, like assessing misconduct or retaliation. AI is less reliable here and requires more human oversight.

Humans should oversee AI based on risk, such as data entry and output review. Low-risk tasks like categorization can be trusted, but high-risk tasks like policy creation need human checks.

AI models vary in reliability for different tasks; vendor claims may not reflect specific performance. Testing individual tasks ensures picking the best model for each use, avoiding over-reliance on one provider.

Risks include data privacy concerns, reliability issues in messy middle tasks, and potential inaccuracies with large datasets. Continuous human review and a risk-based approach mitigate these.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.