Why 95% of AI Projects Fail: Model Risk & AI Governance | Sandip Wadje, BNP Paribas
45m 32s
AI adoption in regulated financial institutions faces significant risks, with 95% of projects failing due to poor risk management and lack of governance. Traditional model risk management principles, rooted in transparent, deterministic models, are inadequate for generative AI, which operates as a black box with non-deterministic outputs. The key to success lies in proactive governance: identifying high-risk use cases such as investment decisions or contract translation, establishing clear accountability, and continuously monitoring for output drift. Organizations must prioritize cross-functional collaboration—bringing together business leaders, legal, data protection, and IT—to assess risks and define controls. A critical shift is moving from model-centric to output-centric risk management, where the focus is on detecting failures and their consequences. Simple, low-risk use cases like summarization are recommended as entry points. Additionally, security teams must be trained to investigate AI incidents, and compensating controls—like endpoint monitoring and access segmentation—are vital. The rise of open-weight models offers cost and sovereignty benefits but doesn’t eliminate risk. Ultimately, organizations must evaluate AI initiatives not just on cost savings but on risk-adjusted value, asking whether the potential fine or financial loss (e.g., 4% of annual revenue) justifies the implementation. Success requires understanding data, defining blast radius, and ensuring governance is embedded from the start—before any AI tool is deployed.
We saw like 95% of the projects failed.
I ended up meeting a CIO in the UK recently
and they said they burned their entire months of budget
which two days of usage.
And they were very happy that the US government disconnected it.
If you just buy one of the ASU products,
you're probably only solving parts of the risk.
Anywhere you're using AI, there is a case thing.
Because if something goes wrong, there's a dominant effect.
When you attach it to an AI, AI sees everything.
So all of a sudden, you have a use case
that can be operationalized in one month, a couple of days.
Look, I can say one million, but this 4% fine
of my total revenue is not worth it.
Now the output has changed.
As a result, your investment patterns have changed.
When do you actually detect the drift?
If you have been trying to introduce AI
into a regulated space, specifically,
financial institutes, insurance companies,
I had a great connoisseur with him the past day.
He is the managing director
and head of emerging technology and risk at BNP Parabas.
And we spoke about things like,
what is the right way to assess AI models?
How do you even approach the model risk management?
What do you do when you have a burning question of,
what's the right way to approach a use case
that you have been presented with?
Because you will have plenty of those
in an organization like a bank,
which has 400 plus application.
But at the same time, how do you understand the risk
as it will change?
What are some of the basic things you should have?
Is logging enough?
Have you trained your instant response people?
Should you even move to open-weight models?
All that in a lot more in this episode
of ASQ report costs.
As always, if you have been here for a second or third time
and have been following the episodes,
and really enjoying them,
I really appreciate it if you take a quick second
to drop the follow, a subscribe button,
which I have a podcast platform you listen to,
watch this one.
We are on Apple's Spotify, YouTube, LinkedIn.
And if you have been here for a while,
thank you so much for your support.
I'll see you the next one.
Enjoy the episode with Cindy.
- Peace.
- Hello, welcome to another episode of the App Security
Podcast.
I've got some deep with me.
Hey, man, thanks for coming on the show.
- Thank you very much for the opportunity.
Really appreciate it.
- Thank you.
- I'm looking forward to this conversation.
Maybe just to set some context,
if you want to share a bit about yourself,
your professional background.
- Thank you.
So I'm currently with VNP Parima
as a managing director for emerging technology risks.
My focus is on cloud artificial intelligence,
digital assets, threat intelligence,
been with the bank for nine years now,
been the UK for 15 years,
been in cyber for 20 plus years.
It's quite an exciting journey so far.
- That's probably the first place I want to start with.
A lot of conversations that I have with the,
I'd say the BFSI or financial security insurance,
financial services insurance companies,
a lot of times my conversation about AI security lands
on model risk management.
And I'd love to kind of maybe,
well, Dave Diamond to wait is a good place to start,
but I would like you to kind of just share some bits on.
What is model risk management according to you
and is it the right picture to start with?
- That's where the foundation has been.
- Yeah.
- From a regulatory perspective,
because almost every regulated in the world,
particularly bank of England, PRA,
also the feds in the US,
everyone has been clear in terms of laying out principles
for model risk management.
And when it comes to model risk management,
we're essentially talking about classical AI.
- Yeah.
- Like not the, not the JNAI one, yeah.
- JNAI is evolving, right?
- Yeah, so MRM is like pre-JNAI
and a lot of those principles are like 15, 20 years old
in terms of having those practices,
where model is not a black box, right?
Where essentially you know everything
that goes into the model
and you know everything that's coming out of the model.
So there is no non-deterministic way.
There is a testing and evaluation around it,
but that's what essentially has been going around
to lay down those principles in terms of,
whether it is model development,
whether it is governance of model,
independent verification of models.
That's where in most of the banks,
you have a separate function,
which is completely independent
to just judge the performance and behavior of the model
before it goes into the production.
And in large cases, almost with most
of the financial services institutions,
these models, the classical models,
and if they're deployed in front of us
or critical use cases,
it's common practice to then share that information.
In fact, including the model themselves
with the regulator, so the regulator knows
that the models are fit for purpose
and they're not taking some wrong decisions.
- Yeah.
- But this is the classical AI
and then we are in a completely different trajectory
for last couple of years.
How does that change and what compliments
and what are probably not?
Things you should move forward into the GENIA world,
the black box.
- What changes with GENIA is, of course,
it's non-deterministic,
but then we were in a conference
with a lot of modernist management professionals,
CIOs and data scientist.
And we came to a conclusion in the table,
is like there are like 1,000 people in the globe
who actually know what's inside the model,
how it actually works, end-to-end,
from those algorithm, and the calculations,
everything that goes into a GENIA model.
And then outside that,
everyone is playing at the periphery
of input and output in terms of how to deal with that black box.
So that black box changes then how we deal with that AI
and then how we handle the governance around GENIA.
- And because it's an interesting one,
because a lot of people start with,
we can probably just drag the same model as before.
People used to think we're talking about security controls.
They were not thinking about risk management.
To what you're saying, it's a black box,
and there's very few people who even understand
what's inside that black box.
How should people approach, is this a new risk register,
is this like, how are you approaching,
or I guess maybe let me just be more specific.
How would you say to your colleagues and stuff
who are in the banking industry and the FSI industry,
how should they approach this?
'Cause a lot of people have still not accepted,
because for them there's no policy for it,
so I don't want to be the guy or gal who's saying AI first.
There's already a lot of skepticism,
and as much as the world would like to believe,
or at least they would want us to believe
that every of us doing AI,
there's definitely a set of FSIs
who have basically said no AI in the organization,
because they're uncomfortable. - That's true.
- So for those colleagues in the industry,
what's your recommendation for how to even approach this,
if they were to be going down that path?
- Sometimes I extend the COVID-19 analogy,
is we had the pandemic, we all knew how the virus worked,
we all responded differently.
And sometimes, and you will see that a lot in technology as well.
Every time there has been evolution of technology,
the tendency is we react differently in Europe,
we react differently in US.
Of course, everyone has their own selfish motivations,
as well as their regional dynamics,
or whatever the local regulatory is saying.
But we do have tendency where,
whenever the next evolution of technology happens,
we respond little differently.
Even though the basics are same.
So in that context, I think there are a couple of things
that stand out is, if you have been running a financial institution
for 10, 20, 30, 40 years,
there are certain things that are quite common.
And this happened to me, when I started looking into governance
of AI, the first task that my boss gave to me,
saying, we already have access management,
we already have data protection.
Can you tell me the Delta?
- Yeah. - And what is the Delta I need to cover?
Don't tell me that I need to start all over again.
So I'm like, that's an interesting homework.
So then I went on to the exercise of essentially
getting all the controls that we have
for access control data protection.
And then we did an exercise to identify,
okay, what is the percentage of Delta
introduced by this new technology?
- Yeah.
- So that's a good way to look at like,
what is that additional Delta that needs to be covered?
That still doesn't change the other dynamics,
which is, they explain it,
we are explainability around the output.
And then which is something, again,
you can address it differently
by fine tuning your governance before the project,
identifying the right set of controls for that use case,
which also means then you're prioritizing your use case
in a right way, high risk use case, medium use use case.
And then you determine governance based on
how critical that use case is for the organization.
So determining certain controls before actually
the use case goes live is also equally important.
While I'm there on that use case topic,
a lot of organizations actually miss the board
in a way that regulator cares about AI usage.
They don't care like whether you use open AI
or Anthropique or anything else.
They want you to tell them where and how AI is used.
And a lot of that goes back to maintaining that inventory
and not just the inventory,
but very minor details of information.
- Or like what model, what version of the model?
- Both and the use case information itself.
We're going back to the CMDB issue.
- Oh yeah, yeah.
- Because like, regular to be asked,
give me use cases where personal it is used.
Give me a use case is where it is market facing use cases
where it says it's an internal use case.
And then you're going around and you're in all fast forward mode,
right?
You're trying to test AI and new models and things
and you're not maintaining that information.
Large amount of energy is then spainting,
just building that inventory in almost like reactive mode
which you could have been sorted out in the proactive mode.
So the use case inventory is a critical part.
The control's pre-implementation rollout
is a second important part.
And then how do you continuously monitor the output?
And this is something again,
you know, we can little bit zoom in
as to how do you deal with output
which is non-deterministic in nature?
- Yeah.
- And what kind of governance can we put on top of that?
- Yeah, and I think I definitely would love to do it.
Maybe this is interesting because for people
who have said no so far,
they would hear your answer and go,
well, that's not how I understood policy because for me,
I have my, obviously there are the regulatory standards
that I've maintained and there is the enforcers or those,
is this coming and especially because if you're a blank
which is global, which has presence in America,
presence in Europe, there's all these different acts
that you have to follow as well.
Does this map out or does this apply
when it comes to being compliant to acts as well
or being able to, I mean, I'm not even talking
about GDPR and stuff, but like, say let's just talk about
between different between UK, Europe and USA as well.
This is like three different acts to follow
in different situations as well, right?
- Absolutely, you have regulations to follow,
but you have to keep in mind,
and I also wanted to put the emphasis on,
cross-functional stakeholder management, extremely important,
particularly with JNEI, because it is moving so fast.
Now you have your board members, your CEOs,
who are talking to each other,
they're probably meeting some of these CEOs in conferences,
open AI, CEO, anthropic CEO.
So they have their own perspective on what's going on.
Then you have your control function owners,
legal, privacy, your IT, CISO.
They have their own perspective
on what JNEI means to them.
And then you have a regulator.
And that's what I talk a lot about emerging tech,
is because my definition of emerging tech
is an area for which controls have not been formalized.
There is no agreement on what controls should be,
whether within the bank or outside the bank.
So in this case, then you have to actually
take everyone on the same journey.
You have to assume that maybe regulator needs to be educated,
maybe the internal stakeholders need to be educated.
So I think the first step in governing AI,
or securing AI, is to validate with everyone
who you're talking to, do we have the same definition
of what AI means to you?
Because for many times, people will
have a different interpretation of what AI means to them.
That is a good foundation to then
start going in the next direction.
I guess you're right, because then you
can focus in the case use cases and then double down on that.
And because obviously being a bank,
I think on an average, there are over 100 applications
in a regular bank in a non-JNEI space,
as in the pre-JNEI era, I imagine with JNEI,
it's even more with just the specific use cases.
Each one of those 400 plus ones
to also have LLM use cases as well now,
which is there in all the work that you did?
Did you find any, what's the maybe tactical use cases
that are usually easy ones to say?
You know what, these are easy to govern,
or these are complex.
Because complex one, because there's
too much personal data in there.
Because I'm coming from a perspective
that if I were to wear the hat of those people
who basically have said no so far for AI,
they've heard the story, they get the message,
but now like, which one should I?
I've got 400 use cases.
Everyone wants a size slice of this.
What are the easy ones to start with?
I think the easy ones to start with,
and I've said this quite sometimes now,
is put them in the buckets of summarized right and reason.
All start with summarization first,
then go to right and then go to reason.
Summarization is the most common corporate activity,
whether it is summarizing emails, summarizing memos,
taking actions on certain meetings, et cetera.
A good example of summarization can also be then
cross-verifying things.
Measbooking of revenue is a very good example.
It's such an important and such a useful use case,
which is not tough at all.
You're getting AI to read the contract,
check the number or value in the contract,
and verify in the revenue booking tool
whether the number is correct.
It then solves such a big, you know,
a headache for the executives.
So sometimes the small, small problems
can have a very high value from executive perspective,
and then summarization can be really good way to start
with saying, okay, let's try out,
it's an on-determined state technology,
let's try our hands with summarization,
then let's get it to write certain things,
measure the quality, and then go to reasoning.
Where we see things go wrong is when people try to pilot
straight with reasoning on day one,
and then they end up burning a lot of, you know,
capital, both the political capital and the real capital
in terms of, you know, execution of that use case.
And that's why we saw like 95% of the projects failed, right?
That's what came up in the MIT survey end of last year.
- Yeah.
- So would you put the, you know, the AI tooling,
which is available on the internet,
which, hey, let me create an image for you,
or I can create a wipe-coded app.
Would that be in your reasoning bucket,
or would that be in the read bucket?
- The wipe-coded is, again,
would fall into the reasoning bucket
because you create a wipe-coded app.
- Yeah.
- And it's easy because the entire development time
line lines have shrinked.
But then who is maintaining that type?
Who is closing the feedback loop?
And you're back to square one.
- Yeah.
- So you see that productivity,
you don't see the effort behind the productivity
to maintain and get the same output on a regular basis.
So one is the type of use cases
and second is a prerequisite, right?
The prerequisite is you have figured out how to clean the data
and the clean data is available for AI.
You figured out who is going to maintain the feedback loop.
- Yeah.
- For a lot of the times you do pilot
and then you forget like you need someone 24 by seven
to maintain the feedback loop.
- Which is the feedback loop between the LLM
and your actual data or the systems.
- Monitoring the output and making sure it doesn't drift.
- I can eval, you still need to fine-tune it, right?
- Yeah.
- It's a non-deterministic thing.
So somebody needs to look after it
and fine-tune it continuously.
- Yeah.
And I guess for people who don't understand evil,
the simple example could be the fact that
even if you're using Claude or OpenAI,
you may be making it to the same thing 10 times.
The 11th time somehow it just forgets
what it did for the first 10 times.
So you have to remind like that's technically
the simplest way to explain evil.
Would that be right?
- Yes, I can give you a simple example of translation.
- Yeah.
- 'Cause a lot of organizations have gone that journey.
Global organization.
Even like we are a very large global organization,
multi-lingual organization.
You have office in Spain, you have office in France,
India, so many places.
And then there are people who use, you know,
the kind of local language.
And as a part of the day-to-day corporate job,
you want like, I want translation to be accurate.
And now a lot of companies moved from the translation software
for which they used to pay a lot of money
to in-house translation using chain AI.
And the data scientists now will come up
with a different set of approaches
as to how will they check accuracy of the translation.
One is maybe sampling.
And some human is doing the sampling
and checking the translation.
The second-based approach is getting another AI
to be validated and the third area to be, you know,
the final checker.
So it is AI doing validation of the AI output.
But you may still have the cases
where certain things fall out.
Because again, large language models
are predominantly built on the English language.
Not necessarily, you know,
they are going to get things right
when it comes to French or Spanish
or any other languages.
So we look at a use case from IT perspective.
We think of access control data prediction.
We don't think of output and explainability
around the output.
This is where things get tricky.
And translation is a good example.
You can have other AI to check, you know,
the output of your AI doing translation.
That doesn't mean that it's going to be 100% accurate.
- So in your example, the translation is. So going back to the output thing that we spoke about,
how there should be continuous monitoring of the output piece.
Maybe we can unpack that a bit more
considering that at least given a few examples
to the audience for translation and the. - Continue on the translation.
You can use translation for communication in your town hall.
That's awesome.
Maybe one or two words go here and they're wrong.
You can deal with it.
You use translation on a contract document
and the interpretation is wrong.
Now you're looking at a much higher risk
and much higher value.
So again, you look at the use case.
It was a very straightforward use case of AI, you know,
understanding something in French translating in English
or vice versa.
And what you find yourself now is the context, you know?
AI wrote something in English for you to go and talk
in a town hall versus AI writing something
that's a lens on a multi-million contract.
- Yeah.
- Two different things.
- Even from an investment perspective,
like a lot of banking sector is into investing in stocks
and stuff as well or companies that are about to be released
and to your point, a zero extra or a zero less
could mean a huge difference.
- Yes.
- Yes, exactly, exactly.
- Do you find that when you were specifically talking
about the focus on output, bringing it back
to the security side?
Obviously, those are very business focused examples,
100%.
So to kind of paint the whole picture from a security perspective,
if I was a security person in a global bank
and I had the translation,
and now we have understood that we have an AI software
that's going to be used for translation of contracts.
I have a data scientist who looks at the e-value,
but I also have an actual person
who deals with contracts looking at this as well,
going, is this supposed to be a or an a,
like someone from a legal team, for example.
A lot of conversation always comes down to,
hey, let's just give that to a governance council
and let them decide if this tool is good or not.
Is that a good approach for managing this
and then it goes into the whole model management again after that?
- Yes, probably not a good approach,
because the governance council does not understand,
and that is again, we talked about this,
is most of the people trying to put governance around JNEI,
do not have any understanding of e-value frameworks
and no one says, hey, can we review the e-value framework
from the data scientist before we decide how to,
you know, do what to do with this use case?
No one, everyone is looking at like,
what is this use case, which model used,
which application it is connecting to?
What is the data?
And that output focus is completely missing.
I think the key is essentially,
and this is what we did actually is to look at like,
what are the different risk events coming out
from AI output and then what would be the potential
consequences of certain things going wrong?
So what you essentially done is,
you looked at your event's taxonomy,
you will have events that happen, right?
You have a phishing attack, you have operational failure.
Then you ask, hey, I'm using this non deterministic technology now,
what can go wrong?
And maybe there are some new risk events that you didn't think of.
You put those new risk events, the new taxonomy,
and then you ask, okay, now I understand,
things that might go wrong.
And translation is a good example again,
we can stick to it like, things might go wrong
from legal perspective or the chief operating officer says
something in a town hall they should not have seen.
Now, there is a reputation impact,
there is a financial impact and you can work backwards like,
okay, who is responsible for this?
Who is accountable for this?
If there is a financial loss,
there is a certain person responsible for it.
If it's a failure in terms of data privacy rules or regulation,
there is a person responsible for this.
That also allows you to fine tune the governance
because once you have identified,
[BLANK_AUDIO]
JNII does not deliver the intended output
and something goes wrong.
I now know which person is accountable for that
and that person needs to be part of that governance
in terms of validating the things.
And not necessarily the person with IT hat
and data production hat on,
because they're doing their job from that domain perspective
but they're not accountable when things go wrong
from output perspective.
- That's just an interesting and important point
on the event taxonomy as well as who's accountable.
In the use cases going back to what you were talking about
with the use cases, if it's important for the use cases
to have an owner,
going back to what we've done in the past before,
it's not like a new concept at that point in time.
But if you can double click on the event taxonomy for me
and like 'cause a lot of people would not even understand
that is that just my seam logs
or is that my collection of open telemetry
for my input/output prompt
and what am I giving it as an input?
'Cause if you look at the way vendors have sold solutions today,
it's been sold on the LM firewalls.
Hey, your user may put something suspicious
and I'm not saying vendors are pitching the wrong thing
but more coming from the use cases
that most people are trying to solve
or talk about from a security perspective
is your LM firewall where hey,
is this person putting something sensitive?
Then there is the gateways.
Any conversation goes through us
and the other one that people have been talking about
is automation of security of work.
Now, you kind of mentioned this before,
which I'll be recording about how AI for security
and security for AI are two sides the same coin.
Bringing that back to this example of event taxonomy,
are we already collecting this
or do we not need all this extra
bellazzle that we've been sold
or is this event taxonomy for you
if you use a translation example,
what would that look like?
- So I'll give a completely different example of
event taxonomy.
Let us say you're using JNI to make investment decisions
and that investment decision is based on
understanding of your financial portfolio,
understanding the market data,
understanding of open source intelligence
or your political trends.
And now if you have done governance of that particular
now you have a JNI technology
that is taking investment decisions on your behalf
and you have done everything right, okay?
You looked at prompt engineering race,
you looked at access control,
you looked at data security
and for whatever reason,
let's assume for the fact
that the market data was manipulated.
So the output is going to drift,
you've done everything right?
- That's right, yeah.
- But now the output has changed
as a result, your investment patterns have changed.
And that's where it gets interesting
is the deviations need to be continuously monitored
and that's a job itself understanding those event types.
Then you understand, you ask yourself, okay,
let me look at it.
This scenario, you might not have thought about this scenario,
right?
Because today you have humans
who will read every market data report,
they'll punch information after it goes through three pairs of I,
before it lands into some investment memo,
and the decision happens.
Now you took a decision instead of a person
going through market data report,
these geopolitical reports
and making me a summary that gets into my investment memo
before someone takes decision,
I've delegated everything to AI.
Now you got three different sources
where the data quality or intentional manipulation of data
can drift your output and that has real financial consequences.
So you really have to work backwards on the scenarios.
So I think this scenario exercise has not been like
thought through when it comes to risk events.
So when I say risk events,
it is purely to do with the drift in the AI output
and how that drift is going to happen
and how do you catch it?
How do you catch it?
So it's not just the, that's where I think
the security solutions approach that you talked about
from AI security perspective,
I think it's predominantly to do it.
Model is going to do something wrong
or usually it's going to do something wrong.
Yeah, but there are so many other dimensions, right?
As a business that you would care about,
which they are at your point,
I don't think the vendors can actually solve that problem as well
'cause they can, 'cause this is obviously an example of say,
like just if I was using a,
I don't know, AI legal software that's popular
or like a Harvey or whatever, the other popular legal one
could be, I could be the use in the one for investment
because it just happens to be the most popular one.
People's, all my colleagues are using it,
so almost they're using it.
They could also be applications that are in-house,
which is very common in an exercise industry.
We all have customer applications that have been created.
Now, all of them have AI bolted on
or attached to it.
I feel like this principle would still apply there
as well in those use cases
and completely fall in that same bracket
of what you were talking about in terms of the risk being created
and who's testing the output continuously for the drift.
That doesn't change even if it's like your proprietary
outside application versus an internal application.
Would that be right?
- That's correct. - That's correct.
- Anywhere you're using AI, shadow AI or Jane AI
in that sense, the risk is same.
Because if something goes wrong, there's a domino effect
and you need to have accounted for that scenario.
The worst thing you will do particularly
in the large financial institutions
of not having accounted for that scenario
because then no executive likes surprises
that they had not thought about.
Like, oh, we didn't think about this.
That doesn't look good on you.
- 'Cause to your point, if you just buy
one of the AI security products,
you're probably only solving parts of the risk,
not the entire, probably the whole other set of risk
that you need to care about.
- Yes, there are different spotlights.
We talked about this.
One spotlight is the asset inventory,
your AI use cases, your data metadata about use cases.
Second is your pre-go-life governance,
having right set of controls before you actually roll out
to the use case, data life cycle, which very much applies
to that investment scenario where if data intentionally,
run intentionally, what manipulated,
you are able to detect that in the output drift
and take some corrective actions.
So data life cycle, again, in the cyber security context,
when we say data life cycle,
man, everyone is thinking about personal data and this.
It can be the quality of data also,
which can have a far reaching consequences
in AI output in terms of decisions that are taken.
Then you have essentially the model themselves
and attestation of those models from security
and behavior perspective.
And then you have almost like,
and this is also we talked a lot about is,
are we training our SOC teams to investigate AI incidents?
- Oh, tell me more.
- So the SOC teams today are focused more
on phishing campaigns, incidents, stuff goes wrong.
I get an alert from my cloud security company.
- Yes, have you, have you trained,
let's take an example of the e-vail framework.
Something bad has happened,
and now your sister to go involved.
They had no idea about e-vail frameworks of what data,
they'll have to access everything, right?
- Yeah, yeah, yeah.
- Forensic would also need that information.
- Have you trained your sister guys or the team,
saying, hey, we have this new technology.
This is how it works.
These are the things or components of this.
And maybe these are the things can go wrong.
So when you do investigation,
you have to follow these things.
So I think we have to invest a lot on training
our SOC teams on how to investigate AI incidents.
I don't think effort has gone in that direction
because one is SecOps.
- Yeah.
- And second is ML SecOps.
- Mm.
- Okay.
- And if you look at MLOps,
which was traditionally data scientist responsibility.
They were doing the pipeline and everything.
- Exactly.
So now you have something that was predominantly data scientist
responsibility, scaling up, right?
Because before the classical AI adoption was in tranches,
like you will do like 30, 40 use cases in a year,
then it goes into production.
And most of the time when it's getting into the production,
it is almost getting into production as an AI application
and not necessarily as out of hand
non-deterministic technology.
But now the more you go ahead,
5, 10, 30, 40 gen AI use cases,
you really have to train your SOC teams
how to investigate them when things go wrong.
- Even Forensic for the matter as well, they would.
- 100%.
- Because there is no undo button, I guess,
in this context.
And if you don't have the telemetry,
and there's no data to collect.
- Yes.
- And what do you do at that point?
I was like, well, I just,
who looks at this black box at this point in time?
- Actually, so what do you think is like the,
'cause I think earlier,
and I'm sure the conversation is evolved quite a bit.
And I will put my hand back on for the individual working
in the regular industry who has not deployed AI.
A lot of the initial focus used to be on the fact
that I can't make my AI give me the same response
every single time.
That used to be the ultimate focus.
It's like, hey, we can't let this thing go in production.
Because of all the use cases that happen where,
I think the chat board gave, I don't know,
a car for $1 or whatever, and they had to do for,
there was a lot of so many use cases.
And I think that became like the,
that number one thing people cared about.
If it can't give me accurate responses,
I'm not comfortable as a security person,
not comfortable for this thing to go into production.
I'm sure that people like that who are still in that board,
who have not been introduced to Eval,
perhaps because the organization itself
is not mature enough to think of Eval as a thing,
because maybe they're not looking at it the right way,
is what's the minimum that people should think about
when it comes to AI security?
In terms of this, obviously people have enterprise browsers
used to use browser security as a thing.
There's CLI, and this endpoint security,
the list just goes on, I think what you were saying earlier,
everyone is thinking about what's the Delta,
'cause I'm being asked to put this in a production,
but I wanted to understand what my gaps are,
to your point, I forward some of these advice,
for all the use cases, but I feel I need some security things
because hey, they could be problem management.
Is there like a, maybe three or four things
that people should think about
as good foundation security things to have?
It could be browser security,
'cause apparently everyone's using chat GPD on a browser.
Where do you sit on that for these?
This is how you would approach tackling some of the security components and which one
of these browsers, CLI, whatever, where do you think are even relevant, if that makes
sense, in the way we are going with emerging tech.
So I talk a lot about compensating controls and everything that you already have or technology
that has mature helps you along the way.
So if you have a remote browser isolation solution, it's a very good DLP filter.
It's going to detect certain traffic that you can stop.
So same goes with authentication as well, right?
So if there is authentication traffic going and you have RBI in between, you can use RBI
as a switch to say, you know what, this kind of traffic I'm not going to allow.
So every technology, you can look at like, how do I use this in terms of, so you first
you found out the delta, yeah, this is something is plot in my hand and a good example in DLP
is a reg-ex-based approach, right?
We have reg-ex-based approach, it doesn't work in JNEI, because users can be smart enough
to navigate and, you know, get confidential information out.
So you can use compensating controls, RBI is a good example, segmentation is a very good
example, particularly from a kill switch perspective.
If you have AI agents, you should ask yourself, hey, can I segment this differently so that
I reduce the blast radius?
So I think find out what compensating controls will work, a lot of this also goes back to
the data hygiene.
If you have not cleaned the data, if you have not cleaned the access management or the access
rights for users or AI agents, both actually, so if I'm a user with undistected access to
all sort of data, the copilot associated with my account is going to have an access and
then everyone who has a shared access in the same boundary can probably see the document
they should not have seen.
So I think I would say you need to look at compensating controls, you need to look at what
are the additional security issues that I'm dealing with as a result of this technology
and then you can work backwards.
Like for example, I would accept the fact that you still want to monitor the prompts for
a lot of reasons, right?
Acceptable use, policy, regulatory violations, etc.
So you're still trying to look at prompts, not just prompt injection, but you want to
monitor the prompt example.
That also gives you, like, it's also a very important data to understand your employees'
behavior.
How do they think?
Are they using AI effectively so that someone goes back and trains them?
So yes, the prompt is a good area where you need a new technology because that area didn't
exist before.
So you can have something that allows you to not just stop prompt injection, but analyze
the prompts and help you with, you know, essentially those things.
But more and more I see, I think it is essentially moving towards like two domains.
One is the user behavior and agent behavior.
And I think we're seeing, and I'm sure we talked about it, we're seeing essentially
the focus where companies are now moving towards like maybe this is an endpoint issue where
I can essentially attach that user and agent behavior and track the whole activity using
that endpoint.
And I think that's where I see a lot of converges going to happen in next, you know, a couple
of months or years.
Because I guess two point at the end of the day that they develop, whatever agent they
end up using would be on a hopefully a work laptop.
So that becomes the endpoint that you monitor.
Would that endpoint analogy work in the case of AI workloads, like the applications that
I know AI bolted on or AI first, as long as there is an identity attached to that AI usage,
you'll be able to essentially then monitor the trajectory as to where that AI is going
and then what action is it is taking?
I can already hear the identity folks in the audience going, but some deep NHI, we have
never seen this before.
Like, do we need a separate solution for that?
I get prompting, I think you and I spoke about the NHI component in terms of, and I think
I want to just double click on that as well, because I think there seems to be a few topics.
There's definitely endpoint as a theme I agree is coming up quite often as a conversation
about it.
Agent as a whole is a conversation, but a lot of that agent conversation seems to end on
the identity piece.
I'm curious as to how do you see people approach it and are they compositing controls that
people cause?
We spoke about the data compositing controls for DLP.
Are there any that come in mind for NHI agent on the identity and that has been significant
amount of time on identity projects start on my career and for the benefit of the audience
who is listening, when Sandeep has a birthright access to certain assets or applications in
the environment, that's what is visible to me, but there are things that are invisible
to me, which are essentially low level entitlements and things, and this is essentially the
known issue.
This is essentially the role design issue.
What you do is you create roles for employees in the organization that okay, Sandeep is
an employee, there is a birthright access to have access to windows laptop, to have access
to one drive, and then Sandeep works for this trading division, and then that two or three
other kind of, you know, applications you have access to, but this is just a role, which
is orchestration of certain privileges attached to your role, not necessarily everything, but
when you attach it to an AI, AIS is everything, a good test is if you go, ever, it doesn't matter
which corporate laptop you have, if you go to the command prompt and see what else is
available to you, you will be quite surprised, like, oh wow, I didn't realize I had so many
privileges associated with my account.
So the bigger issue is, which is something we talked about at a round table in DevOps
earlier this year, is what I find very funny, if you're using JNAI and you're starting
with something new, I would expect you to clean up that access and then start, right?
So, to start with something new and then say, okay, I need a solution for that, that just
doesn't make any sense.
So you're essentially not following the ASDLC principles in terms of, you know, implementing
new AI agent.
So my first recommendation is, if you're building, deploying AI agents, clean up the permissions
that AI should not have access to, to give AI the unnecessary permissions and then thinking,
then ask yourself, I need a monitoring tool or a preventative tool, then you essentially
are finding yourself, you know, in the same thing again and again.
But, you know, I imagine, because I also started my current ITX management, least privilege
is probably like those mysteries that would never get sold.
And I'm sure every engineer out there is like what my agent requires more permission.
Have you found a good answer for that in terms of when working with on that, obviously
we were talking about an end point in agent.
So in that agent ecosystem of identity, going back to what you were saying about the use
case, that determines the least privilege or is it the, the business use case determines
the blast radius.
Yeah.
The use case determines the risk event.
Yeah.
And the blast radius plus risk events is a question you have to ask yourself.
Because now the use case gives you some benefits of automation.
So let us say by implementing that AI use case, you're saving one million a year.
But the probability of that use case going wrong and you paying 4% of your, you know,
annual revenue for some, you know, fine, it's a completely different thing.
And then you then you take a step back saying, look, I can say 1 million, but this 4% fine
of my total revenue is not worth it.
Yeah.
I'm going to park this until I get a confidence that I can actually monitor this end to
end.
Yeah.
So that's what is not happening.
And then you see incidents happening because people just go super excited, try to kind of
implement AI.
And that's where the people talk about token maxing where people just basically burning
our tokens and millions and just not seeing the results.
That's correct.
That's correct.
Because the models are improving.
You're getting a much longer context window.
Yeah.
And you see the models kind of giving you without you desired.
You get excited.
And then you're putting a lot of data straight on to the model and you're kind of, you
know, burning tokens, I ended up meeting a CIO in the UK recently and they said they burn
their entire months of budget with two days of family usage.
Oh, yeah.
Because they were worried that what would happen if it continued.
Yes.
So I think with open with model, open with models and all probably the token prices would
go down.
But that still doesn't solve the context thing as long as you have users putting a lot
of data to get the output they want from AI, you will still have a lot of kind of token
consumption.
I think even if token price goes down, yeah, the consumption would still drive through
the your your recurring expenses.
Actually, it's an interesting point because we haven't really touched on open-weight
models yet.
A lot of people think open-weight models are just free models.
And obviously, it's the GLM versions and this multiple versions of it and people can
have sovereignty related open-weight as well.
How do you explain open-weight from a regulatory perspective for people who misunderstand it?
And maybe calling it a free model is most, it's a very simplified version of it.
What are the use cases for that in a bank as well?
So everyone is testing them right now.
Yeah.
So I think it's too early to say what would be the potential use cases given the token
max issue.
Yeah.
I think everyone is looking at these as like, and then I've seen those conversations a
lot is if you take the analogy of summarization, right, and reason, maybe we can just
use open-weight models, open-source models for some registration, write for some mid-level
and maybe the frontier LLMs for, you know, the critical decision making process.
Yeah.
So I have made a lot of executive stakeholders in last, you know, couple of weeks where they
have actually gone ahead.
Yeah.
They have a bifurcated this.
Okay.
So they're using open-source models for low-level travel task and they're using frontier
LLMs for high-intensity, high-decision tasks.
So that trajectory you can already see.
Yeah.
So particularly in a large-tier one, banks, what use cases will come up with open-weight
or open-source models in need to be seen and they introduce probably like the risk profile
in my opinion is pretty much the same.
Yeah.
Yeah.
Yeah.
Yeah.
Yeah.
Maybe use case that we have not seen before and they just oh, we can apply so you may find a free model on hugging face
Which is just really trained really well on a special something specialty that the bank does
Which technically is a free model as well, but it just trained in a very specific use case and that becomes a model
You use for that whatever that use case is obviously that's correct and I guess it kind of goes back to what you were saying earlier as well about the
risk made freaks and I love the example that you gave whether saving one million versus losing 4% of your
Annual revenue is a good analogy because that makes people also reconsider the risk register
For what they're actually with which we haven't touched on, but I think you kind of that's why you were hinting towards that right?
Yes, yes, so you really have to ask yourself the blast radius for the use case and
If you're taking a use case in front of governance committee
Don't talk about model risk. Don't talk about the data risk. I talk about the blast radius and ask yourself if something goes wrong
Who is accountable and is it really worth if your data protection officer says it's not worth taking the risk of a GDP
Are fine. Don't take it. Yeah, that productivity is not worth it. Oh actually, that's a good point
So when you when people present a use case even if it's security to a governance council
That's a better approach you found that to have enough context to make the right decision on that use case moving forward or not
I think the governance models itself have changed yeah, and I can talk a little bit about almost most of the financial institutions have gone around now in terms of
Creating a different kind of governance for AI because I give the example of the feedback loop and
The feedback loop in large financial organizations essentially a joint ownership now. It's not just one person's work right because
You have a use case where you are using a business data. Yeah, you're using personal data. There is some
Because January requires very prescriptive set of instructions as to what it should do. Yeah, there is a business process that is unique to that function
Yeah, so you have business users you have IT, you know users. You have C so you have data protection officers
You have legal. They're all sharing their shared context and concerns in the execution of that use case
So I call that is like AI kitchen. Yeah, and where one person does not add value or does not essentially give the right perspective
The output is going to drift. So what we've seen is essentially cross-functional
Governance getting essentially much more sick because previously you would have like the standard application booting onboarding life
Psychos go through the IT committee and get it done etc
Yeah, now
Everyone is hyper focused on the output and what happens when output is not as good as we think it should be
The governance dimensions change. So I think governance is becoming more
Cross-functional and more integrated in most of the organizations
So you where you have third party you have data protection. You have second-line C so everyone on the same call
Wow, you essentially look at that AI use case and provide opinion on that AI use case
Wow, and I guess because this is different to the first version of governance council that people started off with or
This isn't sorry. This is an award version
Where it's becoming more cross-functional across the large
I don't know if it's a T1 bank. You can also take the same analogy right who's responsibility was DevOps
Largely CIO's right CIO's ID functions work with business DevOps roll out of the application when the security dimensions will come
The second-line CSOs and other teams will get involved from a attestation and review perspective
Simply because we had that we did not have aggressive timelines in terms of technology has not moved this fast
Yeah, if you look at a cloud. We had 15 20 years to work around play with cloud and do these things
Yeah, what has happened now is all of a sudden
You have a use case that can be operationalized in one month or a couple of days
It has immediate business value, but then everyone needs to be on the same table to you know look at that use case
So that cross-functional coverage has gone up because the elapsed time for delivery of a use case is very very small now
Yeah, there's so much to unpack here, but I think I'll take a pause there
But I think we've got a lot of topics. Is there something that you want people who are
Starting off this journey to kind of walk away with this from this conversation in terms of approach
They're approached to AI in a regulated space that you would want to work with in all the entire conversation
We had what is this and that comes up as like one most important thing people should work away with I would say
Take time to understand your data
Organizations who have done really well with AI are the organizations who figured out how to understand their data and how to work on their data
Before they actually went on a journey
That's a great note to kind of wrap up the interview on as well
Where can people connect with you and find out more about the work you're doing and what you've been up to?
Thank you for asking as you know, I'm quite passionate about the violence of AI
I'm part of the cloud security alliance ASFT council various
AI governance forums so very happy for people to reach out to me on LinkedIn
And I'm happy to answer their questions. I'll put the link in link on the show
But thank you so much for coming on the show. Thank you so much. Thank you. Really appreciate
Thank you. Thanks everyone for tuning in. We'll see you next episode
Thank you for watching all listening to that episode of AI security podcast
This was brought to you by techriot.io
If you want to hear or watch more episodes of AI security check that out on a security podcast.com
And in case you're interested in learning more about cloud security
You should check out assistive podcast called cloud security podcast which is available on cloud security podcast.tv
Thank you for tuning in and I'll see you in the next episode
Peace.
Podcast Summary
Key Points:
Most AI projects in regulated sectors like finance fail, with 95% failing due to poor risk assessment, lack of governance, and inadequate monitoring of output drift.
Traditional model risk management (MRM) based on transparent, deterministic models is no longer sufficient for generative AI, which operates as a black box with non-deterministic outputs and requires new governance frameworks.
Effective AI adoption demands cross-functional governance, clear use case prioritization, continuous output monitoring, and accountability for risk events—especially in high-value scenarios like investment decisions or contract translation.
Summary:
AI adoption in regulated financial institutions faces significant risks, with 95% of projects failing due to poor risk management and lack of governance. Traditional model risk management principles, rooted in transparent, deterministic models, are inadequate for generative AI, which operates as a black box with non-deterministic outputs. The key to success lies in proactive governance: identifying high-risk use cases such as investment decisions or contract translation, establishing clear accountability, and continuously monitoring for output drift.
Organizations must prioritize cross-functional collaboration—bringing together business leaders, legal, data protection, and IT—to assess risks and define controls. A critical shift is moving from model-centric to output-centric risk management, where the focus is on detecting failures and their consequences. Simple, low-risk use cases like summarization are recommended as entry points.
Additionally, security teams must be trained to investigate AI incidents, and compensating controls—like endpoint monitoring and access segmentation—are vital. The rise of open-weight models offers cost and sovereignty benefits but doesn’t eliminate risk. , 4% of annual revenue) justifies the implementation.
Success requires understanding data, defining blast radius, and ensuring governance is embedded from the start—before any AI tool is deployed.
FAQs
Model risk management refers to the governance of classical AI models where all inputs, processes, and outputs are transparent and fully understood. It includes independent verification, testing, and regulatory compliance, ensuring models are fit for purpose before deployment.
GenAI models are non-deterministic and act as 'black boxes'—few people understand their internal workings. This makes risk assessment and governance significantly more complex compared to traditional, transparent models.
Most AI projects fail due to poor planning, lack of feedback loops, and overestimating benefits while underestimating risks. A 95% failure rate is attributed to burning budget on short-lived use cases without proper monitoring or governance.
Start by identifying the 'delta'—new risks introduced by AI—assess the use case’s criticality, establish pre-implementation controls, define accountability, and maintain a comprehensive AI use case inventory.
Organizations must continuously monitor AI outputs for drift, especially in high-risk scenarios like contract translation or investment decisions, using event taxonomy to identify potential failures and their financial or reputational impacts.
AI use cases involve multiple stakeholders—legal, data protection, IT, and business units—so governance must be collaborative. This ensures diverse perspectives, shared accountability, and better risk management.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.