The podcast episode delves into the complexities surrounding data storage, compute requirements, and sustainability issues in data centers. The conversation highlights the significant portion of data that remains unused or of unknown value, leading to unnecessary storage costs and environmental impact. There are discussions on data management challenges, GDPR compliance, and the inefficiencies of retaining vast amounts of unutilized data. The dialogue also touches upon the rapid growth of data, the need for effective data governance, and the importance of conducting data assessments and cleansing processes to optimize data usage and storage efficiency. The episode raises critical questions about the responsible management of data, the environmental implications of data storage practices, and the necessity for organizations to evaluate and address their data storage strategies for long-term sustainability and operational efficiency.
Transcription
4887 Words, 27027 Characters
[MUSIC PLAYING]
Welcome to the Data Edit, the podcast for data forward thinkers.
Hosted by Agiles Ben-Harris, head of AI, and Oli Seger,
account director, we explore the challenges, innovations,
and trends shaping the data landscape.
From governance in ESG to AI ethics, cloud migration,
and analytics, discover how data is driving real business
value.
So welcome back to another podcast.
Just picking up on where we left off from the last session
we had, where we were talking about what
and how the increase in demand of data storage in compute,
the impact that's having on data sensors,
what the hyperscale is actually doing about that,
and how that impacts potentially the sustainability
and environmental message.
Are we doing it right, et cetera, et cetera.
We talked about that kind of message being in two parts.
So one, obviously, there's the requirement.
But two, is that requirement kind of necessary?
So what we're going to talk about in this session
is very much about the data that's currently housed
in those data centers.
And as we like to try and do, I'd
like to just start with a couple of statistics
that are there and are quoted around data itself.
So according to Forrester, between 60 and 73%
of all data within enterprises never
used for analytics, despite the growing emphasis
on data storage and decision making,
NetApp reporting up to 80% of data stored by some companies
is never accessed.
So you've obviously got the increased storage costs,
unnecessary storage costs, and carbon emissions
associated to that.
The Veritas Global Data Berg report
found that 52% of stored data is, as they call it, dark.
Therefore, its value is unknown.
Another 33% is classified as redundant, obsolete, or trivial.
So we like to call it as rot.
So that means 15% of data stored is considered
business critical.
So there you go, Ben.
There's some statistics which are straight out
of actual company's research.
What's your thoughts on this, right?
Because data centers and this expanding requirement
for data storage and compute power
to be able to do some of these AI solutions, provide
some of these AI solutions is a thing.
But if we take a step back, do we really
need to be looking at is the need really there
for these data centers to expand as quickly as perhaps they
are suggesting they do?
It's a bit of a conflict for me, because you've got--
if you use hyperscalers, for example,
customers are paying for data to be stored.
So hyperscales want bad data, unaccessed data
to still be stored.
They don't want that claimed as dry.
So what's your thoughts on this?
Over to you.
Yeah, I was going to say, I just look at what 10, 15 years ago
when cloud really started coming about,
the selling point for clouds was storage is cheap.
Yeah, storage is cheap.
Storage is cheapest thing on the planet.
I remember when you and I talked about this a little while back.
I went and got some statistics from a couple of cloud
providers, and we worked out that for 20 pounds a month
is the equivalent of seven warehouses full of filing
cabinets of A4 paper.
Yeah.
So in your head, the environmental kind of stats
and values, yeah, of course it is worthwhile.
How much-- what's the environmental cost?
And I don't know.
What's the environmental cost of producing a piece of paper?
I daresay today, potentially higher
than it was back in the day when paper was more readily
used and available, probably more of a specialism today.
And then physical storage, and then disposal of that paper,
and the bits that we think about today that guard our data
like GDPR, really only possible because we store it all
electronically, and we're able to tag it so much easier.
And yeah, I can imagine back in the day, so to speak,
these seven warehouses are filing cabinets of people's data,
and there was another warehouse of filing cabinets just
for job cards, so that in this warehouse,
in this filing cabinet, in this folder,
there would be this type of data.
It never had happened.
So yeah, moving to the cloud, yeah, storage is cheap.
I get it.
Yes, I think it probably was an economic and sustainably
friendly move, but the truth of the matter
is, because of those immortal words of storage is cheap,
AB wants to get rid of their data now.
So I mean, I think Roth is probably
the easiest one to tackle.
I've done grant applications.
I've built based systems to help try
to resolve some of those issues for people.
But the bigger question is, or it's not a bigger question,
the bigger concern is, is people like yourself
who are face-to-face with businesses,
how do you go talking to somebody about what they have?
Yeah, do businesses know the honest truth
that they're hands in the air and say,
do they actually know what all of their data is?
And from the stats that you kind of said,
the answer has to be no.
Not all of it.
Absolutely, absolutely.
And I guess from that as well, you've
got to be thinking about, back to the AI conversation,
how are you going to get value from an AI project?
If one, you don't know what data you're reading
and you're accessing or you need to access,
because are you going to have to look at everything
just to do a simple query?
Well, if 80% of what you're holding is not necessary,
then you're wasting valuable compute time
and valuable compute resource.
So this is where this foundational thing is really
important and understanding what you've got, where you've got it,
why you've got it, how long you need to keep it for, et cetera,
et cetera, just to enable you to be foundationally
in a strong position, I guess, to start building up from.
That sounds a bit basic, but that is the reality, right?
100% say maturity assessments are the best way
to kind of go through that process to understand,
not necessarily the data, you're not
going to get to know your data about from a maturity assessment,
but you're going to understand what you think about data.
And that's probably stage one.
So if you are of order, we're going
to need this at some time in the future.
Let's keep it.
Let's keep everything.
We can sanity check it.
We can do this with it.
If it's not transactional or relating to a person
who's got no BBI on it, we can keep it forever.
I want to be able to compare stats this year's sales to 50
years ago.
Who is ever going to do that?
Honestly, nobody, not really in their right mind.
No one's ever going to honestly do that.
So a maturity assessment to actually say,
where do you think you are on that sliding tail?
And then it's not a case of trying to baseline people.
Some people will be where they think they are.
Some people will be in that four or five range.
But a lot of people think they're in three, four.
And actually, they're in one, two.
Because it's this understanding of why have you
got this data?
What good is it going to be to your business?
And again, it comes back to those more-- well,
storage is cheap.
And if I don't have it tomorrow, I've lost it.
But I kind of go to your point, if people are only looking
at 80% of their data, they don't know what they have today.
Well, what difference does it matter about tomorrow?
It's unbelievable.
And I've worked in businesses and organizations.
And there's probably a hateful sound
familiar to people who listen to this as well.
And quarterly reports came out, especially
in the financial services.
[INAUDIBLE]
So the annual report came out, the quarterly report came out.
It's a PowerPoint-- it was one PowerPoint, one PowerPoint
presentation that's probably had multiple versions of it
until everybody in the board signed it off and said,
yeah, I'm happy with sending this out to our investors
and also internally to our staff.
We can tell everybody where we are.
So the company I worked for was fairly large.
We had over 50,000 staff.
We had multiple investors, big, small shareholders,
et cetera, et cetera, and they all get the same pack.
So internally, I used to monitor our servers.
And when that quarsely pack or that yearly or annual pack
came out, one pack became 50 packs.
Because everyone would save it.
Yeah, yeah.
Everyone would save a copy.
Everyone would keep a copy.
Oh, I need that.
Why?
Why do you need it?
And maybe today, yes, centralized storage does help.
But between business to business,
there's no way, potentially, of centralized storage.
So we still send everything on email.
And then when it lands from--
we're talking to a client, it lands with that client.
They probably share that same email.
That goes to how many more people.
So it still happens today.
Maybe not on the scale.
Used to watch our servers, capacity load.
We used to run them at about 80% capacity at all times.
So we would have been for a spike, and that would be the need
to go and buy more storage.
Just releasing a report, just a global report like that,
just doing that.
Now, the big question is, 50,000 copies internally,
how many people we read it?
And that is wrong.
But then, obviously, we'll just trash.
Yeah, now I get it.
And do you think the fact that there's
this data that's, I guess, unknown about or certainly
never accessed, according to those statistics I quoted?
I, in my head, kind of remember back to the dreaded COVID.
And many organizations at the time
accelerated their journey to the cloud.
Didn't have the time because of the remote working
and everything else that the cloud gave them the benefits of.
But perhaps didn't put the pre-planning
into their migration.
The time wasn't allocated to do that as much.
So they didn't necessarily do any big data deep dives
to say what needed to be migrated
and what perhaps didn't.
It just all got dumped in the cloud
to your point about cheap storage, et cetera, et cetera.
Now, that was a thing.
That was a thing that happened.
So therefore, you are just basically--
if I moved house and I've got a garage full of stuff
that I haven't used for years and years,
do I just take all that stuff and move it into a new house?
Or do I take a bit of time, go through it all,
and chuck out the stuff I don't need anymore
and haven't used for X number of months
to make a decision on that, and then only move to my new house
the stuff that I need?
It almost becomes a self-cleansing exercise, doesn't it?
Now, I don't think companies had the time and the resource
to do that when we had COVID and many migrations happened.
So there must be a point in time where, whilst you might,
I'll be saving loads of money.
I'm talking pounds and pence here in reducing
the amount of data you're storing.
The fact that you're storing, it still
means that those data centers have to have the capacity
to store this 80% or whatever near a frame back
to those stats were of data that's not necessary.
So they need to expand because capacity is to a point,
which means they need more compute power, more electricity,
more cooling from water, whatever
the means that they use to generate that resource that's
required.
But do they really need to?
Because if there was an exercise,
if every organization went through a process of cleansing
their data, then how much capacity
would be freed up within these existing data centers?
So are we projecting a problem that's here,
talking about a problem that's here today because of our business
decisions that don't seem to be hugely financially impactful,
but environmentally, they potentially
could be very impactful?
100%.
That's the nice way of putting it.
Here's the scary way of putting it.
So most business data is transactional information.
Most transactional information contains personal information,
whether that's a person's name, an account number, a sort
of code, maybe even an address, having to admit more data
than that.
Based on the stats of--
and I appreciate we're painting quite a wide brush here,
so we're not all yet kind of pinpoint people down.
But based on the stats of 50% of the data, roughly is dark.
Sorry, another set of it is rots.
So repetitive, obviously, trash.
20% of it is useful.
They don't look at that 80%.
Maybe knows what's inside it.
So we even do your GDPR audit.
You can do your PPI audit.
You can do your--
you don't audit your own data set.
Maybe looks inside that data.
Maybe knows what's there.
And that to me, that's scary.
Now, I get the whole kind of information requests.
People can submit a request, and I
want all the information that you as a company
store about me as an individual.
That's cool.
But how much data we actually got to go through?
So how do we validate that those requests are
complete and full and honest?
Because we base them on that 20%?
Because we know about that data?
Or do we actually go and look at that 80%
and go trawl through it?
And what do we trawl through it for?
Do we just trawl through for somebody's name?
What other indicators do we need
to use to try to trawl through that data?
And then you hide and trawl through that data.
The obvious answer, let's chuck it all into an LLM.
Let's get something in.
Let's go into this all that data.
Let's use some cool tools that everybody's talking about.
We can run it through an NDM solution.
We could run it through something like a cataloging solution.
All great tools for the right jobs.
All of them have a cost.
All of that's going to have to be open, scanned, understood,
cataloged, categorized, and then put back to sleep again
for how long.
And you come back to that request,
are you actually still--
how long is it going to take you to answer a question?
Are you actually going to get all the information out?
So my question to you and to anybody
listening and to everybody else who--
why don't you just press delete?
See what happens.
Yeah, what's the worst that can happen?
Yeah.
What is the worst that could--
what is honestly the worst that can happen?
Because you-- if you didn't--
yo, I don't know if that's the right answer.
I-- yeah, and I guess if you put a governance hat on,
no, it's not the right answer.
But provocative hat on, just press delete.
And they what?
Where are we with your data sensors?
Where are we with all this growth?
Where are we with everything that comes from it?
Are we actually going to--
are we actually going to kind of catch up?
I mean, I give you some stats that are really scary
around the growth of data.
So every second year, data in data centers doubles.
Really?
How quickly we've--
that's how quickly we gather it,
and we don't get rid of it.
But do you think we're--
do you think this data we're gathering now,
are we being more mindful about it?
Because what-- that we know it's more powerful to us
if we use it properly?
No, I honestly just believe people
are still keeping it because they want to keep it.
It's cheap, I have to get a search.
Well, Mike, what could I use it for?
When can I use it?
And how much of it falls into the rock category?
The rock category that they've been--
is 30% to 35%.
It's expected to vary to 40% to 45% by 2030.
So we are-- we're just doubling rubbish.
You've already got multiple times.
At the moment, the amount of electricity that is needed
to power the data centers to store that data,
before you even go near your--
the bit that we touched on earlier around AI
and using this data, there aren't enough electric cars
to offset the carbon on the upside footprint of that storage.
Now, OK, and it's going to double.
And I can't, with my hand on my heart,
see the number of electric cars on the road
double every two years.
No, because no, it's glories.
We're going to get to buses.
We're going to get to big farm machinery.
Yeah, yeah.
How are you going to--
And then the cycle returns, doesn't it?
Where you thought we talked earlier
one of the previous discussions we had
was around if the data center is utilizing
all of the green energy, what's left for this increase
and everything else?
I just want to give you another analogy as well.
If we hadn't digitized all this data and it was in paper,
you gave the example of how many pieces of paper
would there be equivalent to what's
stored in data centers, if it was still in analog paper
format, there wouldn't be enough place
or land to store all that data for this length of time.
So if it was a problem that was staring us in the face
because it was analog, it was paper,
something would have been done about it.
Because the problem is kind of buried in IT,
and it's just there, and it doesn't cost us a lot,
and someone else's kind of problem to an extent,
then it's not being looked at.
And that's why this conversation is valid,
because I think who's going to look at this as a thing?
What organization is going to have the time, the energy
to actually go through a process of saying, hey,
let's do an assessment across all of our data,
and let's understand where we sit in those metrics of non-use,
non-access, or only if it's 20% use, whatever.
Let's see where we are.
We can make a business decision on what we do with it,
the stuff that we don't need in access.
Do you know what I mean?
But take it back to your analog days,
and this did happen.
It had a fire in a warehouse.
Yeah, yeah, so you had copies of it.
That's the other thing.
Yeah, yeah, yeah.
So you didn't just keep one copy, you'd have various copies.
I wouldn't go all week in my story about working at Kodak again,
because no one listening to this wants to hear that,
but that was an interesting type.
But my point is, actually, if there was a fire,
and both copies were burned, what would you have done?
Business as usual.
Just carried on, but you left off.
I mean, I've literally, I've worked for clients,
where when digitizing data was the in thing,
very similar to your Kodak story,
and we looked back, we're employed to try to digitize job cards.
And we looked back at these job cards.
They went back to 1908.
Yeah.
Yeah.
So, I mean, that's a long time.
And funnily enough, it was about electricity
and how it went into people's houses.
So it was the, these job cards basically would tell you how
that house was connected to the mains,
the type of mains that was connected to the ampage,
the ames, whatever else.
I'm not an electrician, so I honestly don't know.
But, like, these cards went back to 1908.
Now, either that electrician was amazing,
because no one's did back to that property since,
and updated this job card, or done a new job card,
or it was just lost and forgotten about.
And actually, the natural process of life of,
you know what, I can't find a job card.
I'll just create a new one from scratch.
I'll build it, work, I'll figure out where everything comes in.
I'll do a few tests, I'll write a few bits down,
I'll create a new job card.
That's what happens in the old days.
So the old days, kind of the analogue days.
Whereas, hey, we're, oh, God, no, you can't do that.
We've got to find it.
We have to find this piece of different,
80% that's sitting over there.
Yeah.
And we went through a process for this client
digitizing a POC digitization of,
I think it was something like 2,000 of these job cards.
They had four million that year in mind.
The cost to do those 2,000 was astronomical.
You're at millions.
Just get a PI to read it, to identify it, to understand it.
And to all the techies who are listening to this going,
"Yes."
Today, yeah, that's fine.
We can use neural networks.
We can do multiple other bits.
What we were looking at this initially,
we had to build independent models of each job card.
And then because these job cards were various,
and you pick up a handful.
First of all, you had to get them transferred to PDF
and so that they could be read.
And when you pick up a folder,
some of them started at 1908.
Some of them were 2008.
And they went through multiple styles.
So one folder we built,
upwards of like 40, 50 different models.
- Yeah, for sure.
- Just to try and understand it.
And that was the time it took.
And the company concerned was built
from loads of very small businesses.
And as time's gone on,
they've merged, they're joined.
So you have different, yeah, different formats,
different things. - Yeah, yeah.
- Hands off to the old guys, and I call them the old guys.
Back in 1908, their handwriting was brilliant.
- Yeah, yeah.
- They're back to the guys today.
It was, it was like hardly readable.
- Yeah.
- It was chalk and cheese, but.
- Yeah, did anyone really care?
- Yeah.
- Yes, I'm actively just waiting.
I spent five minutes looking for it.
I can't find it.
I'll just start again.
- Yeah.
- So I asked the question,
and it is a provocative question of,
what would happen if you just pressed the link?
- Do you know what?
I was talking to a, talking to a gentleman
at an organization, not more than a few months ago.
He was still using AS 400, IBM AS 400 for stuff.
And I was asking about what his plans were.
Modernize, migrate, switch off, whatever.
He had a couple of guys who had retired
that used to look after full-time,
look after the AS 400 amongst various other systems.
And they retired.
So I said, "Well, how do you handle stuff
"related to this now?"
And he said, "Well, first of all, it doesn't break,
"which is good, good testament to a bit of kit."
And secondly, he said, "Well, if we really desperate,
"wouldn't I just get them out of retirement,
"give them a day's wage or whatever,
"and they charge me a bit of their time,
"and I pay for them to come in and look after it?"
I said, "Is that sustainable?
"With respect, how long is that likely to be okay?"
And he said, "To be honest, Oli said,
"I'm pretty much sure we can just turn it off.
"Next time something goes wrong,
"and I'll deal with any of the upshot of what happens,
"because nothing business-critical is associated to it.
"It's just something we've always had.
"It's always been there.
"And I don't think it's worth us doing any investigations
"into what's on it,
"because if we switch it off, we'll find out."
So it's just about making a perceived hard decision,
but really, is it a hard decision?
Just, it's like back to my garage analogy.
If I haven't used something in my garage before moving house.
- But a period of time.
I know I don't really need it, so have a clear app.
- Yeah, and there's some best practices
that people need to put in place as well.
And this comes all the way back
to the first couple of comments around maturity.
It's, you know, analytics are great,
but they all start in one place,
and it's a copy.
It's normally a copy of a copy of a copy of a copy.
So you actually get the data.
Very rarely would you have had it in one jump
from the source system directly into your data lake,
as an example.
Make people work with, there's an extraction process,
then there's a transfer process,
then there's a, especially if they're on-prem,
then they leave it up through to the cloud
through various other bits of governance,
sits in a storage account
for whichever cloud provider you want,
and then you starve down through an ingestion process.
It's already done four jumps.
And often, my time's up 10, they won't delete those copies.
I'll keep the copy for 28 days, just in case, yeah.
It's like, well, wait, just come from the source system.
So worst case scenario,
just query the source system again.
- Mm.
- Well, that's the worst case scenario.
And then it, but it's a copy of the copy of the copy,
sits in there, you do your analytics on it,
and then you give the answer to the C-suite
or whoever it might be,
and they say, thanks very much, Ben.
- Yeah.
- And that's, and all that data's been created,
all that rot's been created.
- Yeah.
- Everything along the way.
- And we're just, so where does,
where does compliance of this data,
I feel like we're kind of moving into another topic here,
but it's, you know, whenever you're talking about data,
yeah, the compliance story is super important nowadays,
particularly, right?
And maybe touched on it, not so back in the day,
but, you know, that's another subset
of the topic in itself, isn't it?
Because can you do what you might want to do
from a compliance perspective?
Maybe not.
- But also, I think compliance bends you the other way,
in, you know, SharePoint is a great example.
Everybody, well, I say everybody,
majority of companies have a SharePoint
or SharePoint type of site.
And it's where all us, as employees,
information is stored,
it's meant to give us that single repository of the truth.
And it all sits there,
and it's an attempt to remove the rot of some situation.
Get behind the scenes,
they all have version control on them.
And what does version control mean?
Copy.
- Do.
- Of a copy.
Of a copy.
Yeah, I went to a client's version of it the other day,
looked at version control,
looking for a file, specific file,
and the point in time when that file was changed,
there was 102 copies.
- Wow, yeah.
- Of the same file, the file's only been there since June.
And we're on what today, 10th of September.
- Yeah, yeah, yeah, yeah.
- So that's every day, yeah, it's nearly a hundred days.
So every day, someone's been into it, opened it up,
when they close it, it's safe,
and it's automatically taken the version
for a version to one site.
- Yeah, yeah, yeah, yeah.
- Why?
- Yeah.
- No, so it's, look, I guess, you know,
for those listening and watching and being interested
to hear your thoughts on this, you know,
from our perspective, maybe more could be done here,
but is it worth doing it for the return, et cetera,
unless you're thinking about, as we talked about that,
the net impact on these data centers
that house all the data, if that's important to you,
then, you know, which it should be,
this is something to think very seriously about.
So, I guess we'll wrap it there for this session, Ben.
Thanks for your insight, really useful stuff.
I think what I'd like to do is take a bit of a pause
in terms of setting an agenda for the next session.
In the next session, we're just gonna take some,
take a view on some of the comments.
If you've got any things, topics, discussion points
that you think will be of value,
and you'd like us to talk about as two middle-aged chaps,
one technical, one non-technical.
Yeah, more than happy to take a view on that.
So, thanks for your time, Ben.
Thanks, everyone, for watching, listening,
and we'll catch you on the next one.
- Yeah, thanks, Obi. Thanks, all.
Thank you. - Cheers.
- Thanks for listening to The Data Edit.
If you enjoyed today's conversation,
be sure to follow the show and share it with your network.
You'll find more resources, insights,
and upcoming episodes at agile.co.uk.
(upbeat music)
Podcast Summary
Key Points:
Discussion about the challenges and impacts of data storage, compute demands, and data center sustainability.
Statistics revealing that a significant portion of stored data is not used or valuable.
Considerations on the environmental and financial implications of storing unnecessary data long-term.
Concerns about data management, GDPR compliance, and the inefficiencies of storing unused data.
Reflections on the rapid growth of data, challenges in data governance, and the need for data assessment and cleansing.
Summary:
The podcast episode delves into the complexities surrounding data storage, compute requirements, and sustainability issues in data centers. The conversation highlights the significant portion of data that remains unused or of unknown value, leading to unnecessary storage costs and environmental impact. There are discussions on data management challenges, GDPR compliance, and the inefficiencies of retaining vast amounts of unutilized data.
The dialogue also touches upon the rapid growth of data, the need for effective data governance, and the importance of conducting data assessments and cleansing processes to optimize data usage and storage efficiency. The episode raises critical questions about the responsible management of data, the environmental implications of data storage practices, and the necessity for organizations to evaluate and address their data storage strategies for long-term sustainability and operational efficiency.
FAQs
Between 60 and 73% of all data within enterprises is never used for analytics.
Up to 80% of data stored by some companies is never accessed.
15% of data stored is considered business critical.
Data in data centers doubles every second year.
Assessing data maturity helps organizations understand the value and necessity of the data they hold.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.