Go back

Episode 4: The 'data' bit of Data Protection

21m 46s

Episode 4: The 'data' bit of Data Protection

In this episode of the ID Lighthouse podcast, host Scott Salmons explores the evolving role of data protection professionals in the realm of raw data and data science. He notes that traditional DPOs are comfortable with structured records and systems but lack experience with the value, quality, and ethical implications of individual data elements. As modern technologies like predictive analytics and AI rely on algorithms that assign weight to specific data points, DPOs must scrutinize whether these factors are appropriate or ethically sound. Salmons illustrates this with the 2020 GCSE scoring algorithm, where postcode-based historical data led to flawed outcomes, confusing correlation with causation. He emphasizes the need to challenge such models, especially when they impact individuals, and recommends resources like Cathy O'Neil's "Weapons of Math Destruction" and the "Spurious Correlations" website to understand the pitfalls of data misuse. He also suggests books by Catherine O'Keefe and Dara O'Brien for practical data ethics guidance. While acknowledging that DPOs need not become data scientists, he argues that basic data literacy and mathematical skills are increasingly valuable. He leaves listeners with a challenge: to observe how data-centric their daily work is becoming and assess whether they have the skills to respond effectively. He closes by promoting the next episode on impostor syndrome and invites audience feedback.

Transcription

2987 Words, 16753 Characters

English
(upbeat music) - Hello there, and welcome to episode four of the ID Lighthouse podcast. With me, Eurega Host, Scott Salmons. Now, if you've made this far in the podcast, you do it well, lovely. And as you know by now, we have one with me and then one with a guest. So our next episode will be looking at impostor syndrome and mental health in our profession with an honored guest, details to come out soon. However, this week we're going to be looking at, or I'm going to be looking at, the data part of data protection. So I want to explore this a bit more because this has been coming out more and more and more. And some of this is very much my opinion and some of it is very much stuff that's cropped up with other people in some form or another. And there are a number of professionals out there that have some good stuff that's worth you having a look at that I'm going to signpost you to as well. However, I think it's fair to say, and I'll be very interested in your opinions on this. So do feel free to comment, feedback, et cetera, that as a data protection professional, especially those of us that have been doing it for 20 plus years, so I'm the old act, as well as the new fun GDPR stuff. That it would be fair to say that we've kind of spent a lot of our time and we are most comfortable operating in the information and records space. And by that what I mean is structured records, structured systems, you think you could point to as a that's an employee file, that's a custom, a complaint file, or that's this, that's the other one, that's all the rest of it. That sort of structured defined things. We haven't really played in the data space. And by that what I mean is actual bits of data, value of data, quality of data, meaning of data, using one value versus another, another type versus another, how it can be managed as a science, as a standard. That sort of thing, we've kind of not really out of some sort of certain professions, so I fully accept there are a number of professions where if you're a debt protection officer, that's your bread butter, so this is very much a broad spectrum statement. But in my experience, working with a number of DPO's and attending network events are all, is it's predominantly been in that sort of space. Now that means, therefore, that when it comes to the modern world, so if you look at things like things that go down the route of predictive analytics, performance analysis, genuine data, there are a number of us that are going to struggle with that, or need to start up-skilling. And by that what I mean is, so right now, in terms of where the profession is going, standards industries, for example, if someone came to you and said, I'm looking at running a proposal for, I don't know, do performance studies on the effectiveness of teachers and/or teaching methods? So looking at the effect of different teachers, different teaching styles on student attainments, student outcomes, overall school performance, et cetera, or a predictive algorithm for determining individual care needs and the factors that have come affect care, the ability to pay for care, care settings, home, et cetera, or an financial model for determining suitability for a particular financial product, likelihood to repay or marketing analysis to determine purchasing probability and likelihood to respond to particular products and services. If I was a bit to any one of those to you and go through the individual data fields that we're looking to be using, even if you can get past the legal principle lawfulness stuff, which I'll add to my concerns about, that's where most people tend to stop or focus, sorry. If we actually then got into the nuts and bolts of what data is going to be used, the value it's going to be put on it, because remember these sorts of algorithms and data manipulation places value on individual items of data, whether it's minimal value, defining value. Would you be able to, as a DPO, would you be able to scrutinize that? And say, for example, is it appropriate coming back to the question of reference one? Is it appropriate to hinge a score on someone's performance with regards to teaching or teaching method? Is it appropriate to do it on data that's either out of date or on schools that are very different to the one this is currently being done on or even on factors that don't take into account student ability, disabilities, other things. A lot of the time, and again, correct me if I'm wrong, I'm very interested in the opinions of this. Most of it, when we come at it, without formal data training on the value of data which derived from how valuable different data sets can be for different reasons. Most people just follow it logically. Logically, is it appropriate or, let me rephrase that question. Logically, if I was to say to you, we're going to look at the performance of different teaching methods and not take into account individual learning, the performance demographic of the student make up whether they have particular learning styles, particular and disabilities. Logically, you would say that's going to be a flawed model. But even that sense of logic is based on, well, if you've never worked with a school before, you'd never know that. But if you'd never worked with a finance product, you'd never know that. Or if you'd never worked with marketing before, you'd never know that. For if it was a marketing algorithm. So a lot of it is based on what I would label as, I don't like using the word common sense, but it's that common knowledge stuff, situational knowledge stuff. It's not based on actual knowledge of data in its true sense. It's based on operational knowledge of products, contacts, that sort of stuff. On that basis, therefore, it's very, where it has been, and I've seen a lot of this, but again, this is not true for all, so this is a percentage of. It's subjective, open to interpretation, or we could misunderstand things, or miss things that crop up. All these sorts of models need data to work, and all of them have different data sources, different values, different ethical considerations. So another example, for example, would be postcode, does an example. Often, she used quite a lot in a number of different things, but depending on how it's being used, and the data it's being associated with, it can actually be quite unethical. So for example, well, a life example, I think it's widely attributed, that the reason why the 2020 GCSE scoring algorithm went so awry was because of a major part of it being to do with postcode, in that too much value was put on historical achievement, but assigned against postcode. So for example, five people on the particular postcode got an average score of a C, and now little Johnny this year is predicting to get an A, that would bring his overall score down. Similarly, if the average score was A, then someone's predicted a D, it would bump it up. Is it really fair ethical to associate, or to put a lot of causation value, I don't even come back to that in a minute? On postcode, it brings all the historical, positives and negatives with it, is that an appropriate thing to do. Well, ask yourself the question, as a data protection professional in the 21st century, have you considered individual data elements and the history of the ethical uses of, all the context that come with it? As an example of this, I highly recommend reading any of the material from Caffeoneal. Specifically the book, "Weapons of Math Destruction". It's very, very good for giving you lots of different examples, very American examples, but not always. There are some UK examples in there as well of individual data elements and the power they have what they're used for, how various data analysts, legitimate or other, have used those individual elements to draw wide conclusions from people, put them into buckets to use a phrase, which then has quite profound effects on people in some form or another. So I'm back to that question just now. As at a modern day of detection professional, have you considered looked at, trained, whatever word you want to use, data in its purest sense? Not just processes, not just systems, data, especially as more and more and more systems are going algorithm based, data science based, AI based. There's a particularly good website that I highly recommend if you've not seen it before, and I'll plop a link to it in the blog post that'll accompany this episode called "Spurious Correlations", because correlation is not causation. With the GCSE staff and the postcode, there was a belief that because postcode correlates with academic attainment, there was a trend, there was a correlation there, that is true. But does it mean that just because you live in AB1, for example, you are determined to get X or Y or Z grade? No, that's not an absolute, that's not a causation, it's a factor, but how much weight we place on that factor depends greatly on a number of different elements. So there's a really good website called "Spurious Correlations" that gets updated quite regularly, that gives you examples of stupid correlations. Like right now, for example, the first one on the example that I've given on their website is the number of movies a larger would appear in, correlates with the number of orderlies in Oklahoma, or the distance between Uranus and the Sun, correlates with a global count of operating operational nuclear power plants, or associate degrees awarded at engineering technologies, correlates with Google searches for daylight saving time, who knew? Or the popularity of the first name "Candon" correlates with UFO sightings and Florida. It sparked both of them spiked quite heavily in 2015, but what is UFO sightings going to do with the first name "Candon"? Or "Kerosene" used in El Salvador correlates with Google searches for attack by squirrels that took a massive dip in 2007 for some known reason. Have a look at the examples. It's a really good example of how when you're challenging or when you get a DPA or something like that, where someone's saying we're going to build these are the factors we're going to build an algorithm on, these are the elements we look at, this is the value that we put on certain things, these are things we need to challenge as a data objection officer, because especially if they're being used to make decisions about people, especially if they're being used to make automated decision making, profiling, putting people in boxes that affect them in some form or another, then if that algorithm is being built on or too much value is being placed on something that in life that has no correlation or causation between one thing or another, so post code and academic attainment, correlation yes, causation no, then is it really appropriate to hardwire that into an algorithm? So what's the phrase that, it is for Henry Ford that's associated with the phrase rubbish in rubbish out? Well, as a DPA, if you've not got experience or training or working knowledge of data quality, data science and its purest term, how are you going to know what is rubbish? Because the day scientists, if you ask them what's relevant to working at academic attainment, where someone lives is a legitimate factor, it is. But does that mean that this little clever algorithm should be place that as a key value point that has a massive effect on the end outcome? No, not the same thing. It's almost if you think of the principles of debt objection, principle three, data minimization, principle four, data quality, they are designed to work in harmony, but there are lots of quite a few situations where they actually brush up against one another. And again, as a DPO, there's no expectation of you to be an expert in day science, but to know the basics of, to know when someone is pulling the fast one or making wild conclusions of individual data fields that to the outside world seem perfectly logical, but the moment you actually pull them up scrutiny, they're not logical at all. So I suppose it kind of poses two questions. One, is that actually the role of a DPO to know that and be able to challenge that, or is it actually the role of someone else? Data governance, for example, genuine data scientists, rather than some people who work with data that will occlude what had to do it scientifically, is actually their role. And we then just align with them, rely on them, they advise us, and we advise them. Don't know, I'm not actually sold one way or the other on that. However, what I do firmly believe, and this is something that's cropping up more and more and more, is that just like we needed to understand marketing and how it works, just like we need to understand 365 and how it works, or any other suit types of technology, having a good understanding of the engineering of data or the basic operations of data, I think can only be a good thing in your skills and portfolio. Or even just what are your math skills like? I mean, look around, there's no expectation of you to be a maths genius to be a DPO. However, because a lot of these models, AI models, large language models, predictive analytics models, whatever, are essentially equations, mathematical equations, mathematical values, having a sound mathematical set of skills, I think leading into the previous point by understanding data would help, definitely. Is it necessary? No, and to be clear, I'm not saying that it is, but is it something that would definitely help you? Yeah, definitely. As a starting point, I recommend, well, apart from weapons of maths destruction by Cathy O'Neill, I also recommend two other books. So if you want to get more involved in the data management, data governance, data ethics stuff, there are a number of courses you can go out there and you can become a certified data ethicist and all that other good stuff. Yes, do have a look at that. But also, reading wise, I would recommend two books. Ethical data and information management by Catherine O'Keefe and Dara O'Brien and their second follow-up book, Data Ethics, Practical Strategies for Implementing Ethical Information Management and Governance. They're written by two professionals who, they're very good. Two professionals who, for many years, I thought were data governance people who did data protection. Rightly or wrong way. However, the more I get involved in these sorts of projects and the more well, the more after reading their books and getting involved with some of the stuff they do. Actually, I think that they are ahead of their time. They are where the profession is going to go, in my opinion, and they are well equipped to understand how our future, personal data and all the issues of which are generally going to be more of the data bit of data protection and the areas that we've kind of played in, as traditional DPS for a number of years, especially in the public sector, are going to change. It's going to become more data-driven, more data-involved, and we are going to be involved in more and more and more of these sorts of data-based technologies, algorithms, outputs, etc. Rightly or wrongly, think that's where it's going to go. So there are a bit of tips for you. development therefore, have a look at those three books. They are very, very good. And just be mindful. Have a look out over the next couple of weeks or so, just when you're looking at advice when people come to you, just what you deal with over the next coming weeks. Just look at it and look how many times data crops up, how it becomes more and more and more of a part of your role and ask yourself, what are you advising on? What is it people are asking of you? What skills are you calling upon to answer them? To be able to give them constructive, relevant advice and to bring it back to a debt objection and privacy related issue, which it will do very much to. So, have a look. So, have a think. Some of the comments and things that we've posted in here, just let me know your thoughts. So, comment down below, if you want to get in touch, feel free to. As I say, this is very much my reflections based on working on a number of projects and working with different people. So, let me know your thoughts greatly appreciated. That's everything for this week. If you do have any thoughts, comments let me know. Otherwise, looking forward to the next episode, where we'll be looking at imposter syndrome and how that works and empower imposter syndrome is a big problem in our profession. It is, we also have from it some former and other. I especially, which is why I think we should be talking about it more in order to combat it. Like, subscribe, pass on, share to others or give me a new feedback. Or do both, in the way, I'll see you next time.

Podcast Summary

Key Points:

  1. The podcast host, Scott Salmons, discusses the shift in data protection from structured records to raw data and data science.
  2. Data protection professionals often lack formal training in data value, quality, and ethics, which is becoming critical with AI and predictive analytics.
  3. Examples like the 2020 GCSE scoring algorithm show how data elements (e.g., postcode) can be misused, highlighting correlation vs. causation issues.
  4. The host recommends resources
  5. He questions whether DPOs should upskill in data science or rely on specialists, but argues that basic data understanding is beneficial.
  6. He encourages listeners to reflect on how data-centric their roles are becoming and to share feedback.

Summary:

In this episode of the ID Lighthouse podcast, host Scott Salmons explores the evolving role of data protection professionals in the realm of raw data and data science. He notes that traditional DPOs are comfortable with structured records and systems but lack experience with the value, quality, and ethical implications of individual data elements. As modern technologies like predictive analytics and AI rely on algorithms that assign weight to specific data points, DPOs must scrutinize whether these factors are appropriate or ethically sound.

Salmons illustrates this with the 2020 GCSE scoring algorithm, where postcode-based historical data led to flawed outcomes, confusing correlation with causation. He emphasizes the need to challenge such models, especially when they impact individuals, and recommends resources like Cathy O'Neil's "Weapons of Math Destruction" and the "Spurious Correlations" website to understand the pitfalls of data misuse. He also suggests books by Catherine O'Keefe and Dara O'Brien for practical data ethics guidance.

While acknowledging that DPOs need not become data scientists, he argues that basic data literacy and mathematical skills are increasingly valuable. He leaves listeners with a challenge: to observe how data-centric their daily work is becoming and assess whether they have the skills to respond effectively. He closes by promoting the next episode on impostor syndrome and invites audience feedback.

FAQs

The episode focuses on the data aspect of data protection, exploring how data protection professionals need to understand data quality, value, and ethics beyond traditional records management.

Many DPOs are comfortable with structured records and systems but lack training in data science, data quality, and the value of individual data elements, which are critical for scrutinizing algorithms and predictive models.

The 2020 GCSE scoring algorithm in the UK placed too much value on historical achievement linked to postcode, unfairly adjusting individual student grades based on where they lived, showing correlation without causation.

It's a website showing ridiculous correlations, like UFO sightings and the popularity of a first name, to illustrate that correlation does not equal causation, helping DPOs challenge weak logic in algorithms.

He recommends 'Weapons of Math Destruction' by Cathy O'Neil, and two books by Catherine O'Keefe and Dara O'Brien: 'Ethical Data and Information Management' and 'Data Ethics: Practical Strategies for Implementing Ethical Information Management and Governance.'

No, he clarifies it's not necessary, but having a basic understanding of data engineering, math skills, and data ethics can significantly help DPOs challenge questionable data practices.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.