Go back

Episode 20: The Humean Stain, Part 2

from Tatter

56m 34s

Episode 20: The Humean Stain, Part 2

The episode explores the Implicit Association Test (IAT) and its role in measuring and understanding racial bias. While the IAT reveals widespread unconscious associations—such as negative stereotypes linking Black individuals to criminality—it has weak predictive validity for actual discriminatory behavior, with correlations typically below 0.3. Studies show stronger links between IAT scores and voting behavior than with everyday discrimination. The data suggest that implicit bias is shaped by structural inequalities, such as historical segregation and socioeconomic disparities, with regions that had more slavery in the 1860s showing higher current IAT bias. Critics argue that individual IAT results are misleading and often misinterpreted as diagnostic, risking false self-assessments. Instead, the most effective solutions involve structural changes—like blind hiring and diverse leadership—rather than individual training. The research also highlights a vicious cycle where bias both stems from and reinforces inequality. Though early claims about individual unconscious bias were overhyped, the IAT remains valuable as a tool for raising awareness and fostering public reflection. Ultimately, the focus should shift from individual levels of bias to systemic environmental factors, emphasizing that real change requires policy-level reforms rather than personal introspection. The episode concludes with a call to reframe implicit bias as a reflection of society’s structures, not individual shortcomings.

Transcription

8572 Words, 48758 Characters

English
Hey folks, this is Michael, and welcome to this episode of Tatter. This episode was recorded and edited in part in the Digital Media Studios at Bates College. Access to which is something for which I am grateful, but I do want to point out that the views expressed in this episode of Tatter are in no way official views of Bates College. On another note, I want to give shout outs to a couple of podcasts that you should check out if you are interested in best practices in science. First, The Black Goat is a podcast hosted by SineJay Srivastava, Alexa Tullet, and Simeen Vizier, who by the way is one of the guests on this episode of Tatter. Additionally, check out two psychologists for Beers, in which psychologists Yoel Inbar and Mickey Inslet discuss issues in science, but also in intellectual life more generally, all while drinking two beers each, or in Yoel's case, at least some portion of two beers. But seriously, check out those podcasts they're both doing really interesting work. For now, here's Tatter. On the most recent episode of Tatter, we don't know our own minds completely. There are parts of our minds that occur outside of our awareness and outside of our control. To the extent that if you're a white individual or a black individual, to the extent that you see a black individual, and it activates that face activates a negative feeling or a fear feeling or an anger feeling even, you know, that the IAT is going to pick up on that. From a distance, my impression is that another interpretation of the data that seems very plausible to me is that people's IAT scores reflect the associations they see out in the world, but they don't necessarily endorse. When we're talking about the bias of crowds, we're not talking about knowledge of some factual answer. Here what we're talking about is knowledge of cultural stereotypes and inequalities that surround us in our society. If you're in an environment that constantly cues racial stereotypes, maybe you're in a place that's highly segregated and unequal and all of the professionals you see around you tend to be white and all of the service workers tend to be people of color, just spending time in that environment is probably constantly queuing or reminding you of the stereotypes in our society. This is part two of the humian stain. As promised, in this episode we're going to talk about, among other things, the predictive validity of the implicit association test, that is just how well does the IAT correlate with or predict behavior? One might expect that if the IAT is actually revealing implicit racial bias, then it might correlate with the extent to which people display racially discriminatory behavior. I put that question first to Brian Nosek of the University of Virginia and the Center for Open Science. For the most part, measures like the IAT have been used to try to predict behavior that we have trouble predicting otherwise, and so in those domains, the relationship between measures of attitude, whether self-report or implicit, tend to be weakly related to the outcomes. You see that in sort of what the average study is that's used the IAT, whether for race or otherwise, that the relationship between the IAT and the behaviors is relatively low. That range can be from 0.05, gosh, it can be zero, right? There's lots of occasions where it doesn't predict anything. To the maximum in race is probably that I've seen as a reliable effect without accounting for measurement error is around 0.3. Here's a word for the benefit of those who have never taken a course in statistics. My understanding is that Brian is referring to correlation coefficients, which are numbers describing linear relationships between pairs of variables. When that coefficient is at zero, it indicates no linear relationship at all. So as X goes up, variable Y shows no reliable change at all. So if, for example, a study looked at the correlation between racial IAT scores on the one hand and likelihood of racially discriminatory behavior on the other, if that coefficient were at zero, it would indicate no relationship at all between those two. In the case of a positive coefficient, so as X goes up, Y goes up. The closer that coefficient is to one, the maximum value, the stronger the relationship is. I'll put a link on the website for those who want more information. 0.3 is a pretty typical high end for race. Now the interesting thing about predictive validity is that it varies substantially by what you're trying to predict and the topical domain that you're using. So with the IAT, for example, you can show very, very strong predictive validity between the IAT and behavior intention for voting and reported vote after the fact. So Trump versus Clinton in the last election. If you do a Trump Clinton IAT, you would predict probably around 0.6 or better correlation, the performance on the IAT with who it is they say they're going to vote for and who is they report having voted for after the fact. So there are a number of meta-analyses that aggregate across lots and lots of different domains where the IAT and things have been applied. And they consistently show that there is a relationship and in the aggregate that relationship is weak. But the reason for the weak overall is the selection of the part of the reason I should say is the selection of the domains of study. Because you could also in a very biased way select only those domains where the relationship is high and those would reveal selecting just domains where the relationship is high would reveal a byassessment of the overall average or typical relationship. So for us the interesting question to solve is what are the factors that predict when relations between the IAT and self-report or behaviors will be high and when will they be low rather than what is the average across some arbitrary set of studies or otherwise? So on the Racial IAT in particular both explicit questionnaires and the Racial IAT are pretty bad predictors often predicting around 0.1 or 0.2. That's Calvin Ly of Washington University in St. Louis. But in those studies the Racial IAT tends to be the better predictor of the two. But overall prediction is just very hard for getting at discrimination. And part of that is because discrimination comes from many different causes. There are many other things that are play when it comes to whether or not you end up discriminating like your perceptions of what the social norms are. But part of it is also the kind of messiness of the measures that we use to assess discrimination in the lab. We can't often have a blatant measure discrimination when we're studying these social claims because participants will pick up on what we're doing and then just choose not to discriminate. And so often we have to kind of use these more subtle unreliable measures such as how far you might sit across the room from a conversation partner who's black. And so there's a lot of other things that come into play when you're using kind of really subtle hidden measures like that. So the example I've heard before is that I forget what the exact correlation is but. That's Jesse Shingle, a journalist who has written about the IAT. You know, when you talk about the connection between aspirin and heart disease is something like R equals 0.2. I might not have any fact to write. He said R equals 0.2 and R is the symbol for correlation coefficient in a given sample. So he's referring to a correlation of 0.2. The difference is, okay, so then we know we can get asked into 5 million people practically that between what we do and we actually do concrete reductions in the amount of heart disease and the amount of suffering and death and medical costs. I don't know what it would mean to reduce people's IAT score to certain amounts. They're like practically the clinic leaders in useful because the best evidence we have suggests that nudging people like teachers doesn't really do anything. Whereas we know that there's a solid effect when it comes to aspirin at the population level. It just sort of translates to an out-to-variable we care about. It's been proven to in a way where all we can say about the IAT is that there's a loose fuzzy connection to racially discriminatory behavior in lab settings that doesn't even necessarily translate to the real world. But the connection is just much less looser than the connection between something like aspirin and reduction in heart disease. So, this is Keith Payne of the University of North Carolina, Chapel Hill. And a number of studies have been coming out in just the last couple of years showing that if you compare across counties or metro areas in the United States areas with higher average levels of implicit bias have higher disparities in things like police use of force with black versus white citizens areas with higher levels of implicit bias have higher racial disparities in infant health outcomes. At a national level extending a little bit to gender bias countries with higher levels of gender bias on IATs have larger gender disparities in actual STEM outcomes like standardized test outcomes. And so there are a lot of these kinds of outcomes at the city level, county level, state level or country level that are correlated highly with average levels of implicit bias. When I hear that I wonder if you would agree that those data at least as you describe them raise a kind of chicken and egg question where say in the case of racial disparities and use of force on the one hand you argue that those are a result of these system level biases. But on the other hand, one could argue that those system level implicit biases are actually a product of those disparities that is if individuals in that county are aware that black and brown people in that county tend to experience those adverse outcomes of police use of force more than white that could lead them on an IAT to more strongly associate black with negative. So what are your thoughts on what I'm characterizing as that chicken and egg question? I think that's right. I think the accessibility that drives an IAT is both a consequence or a marker of structural inequalities that we see around us, but it can also be a driver of them. So if you happen to be thinking about a stereotype that links black people to criminality at the moment that you're making some critical decision, you're probably going to be more biased in that moment. So I would say maybe recasting it rather than a chicken and an egg problem that there is a vicious cycle going on here, likely in which causality could work in both directions. So one of the approaches that we've been taking recently to try to disentangle some of this is to look at historical factors over large time scales. So we just wrote a paper looking at the average level of implicit race bias on IATs by county in the southern states. And we compared that to data from the 1860 census. And we find that counties that had more slaves in 1860 have higher levels of IAT bias today. And so by taking time into account either in a more ordinary longitudinal sense or in this larger ranging historical sense, we can start to disentangle what are the historical and institutional and structural factors that are built into our environments that might be queuing these implicit biases among people today. Even if one accepts the premise that the, say, the racial IAT correlates with discriminatory behavior where if you are above the mean on the racial IAT, then you are more likely to engage in racial, racial discriminatory behavior than if you are below the mean. But even if it's true, it doesn't speak to where you fall on a spectrum that one can imagine ranging from egalitarian behavior. So imagine imagine a spectrum ranging from Martin Luther King, Jr. So someone extremely committed to egalitarian behavior. And then at the opposite end, the grand wizard of the Ku Klux Klan. So and then within that you have a range of discriminatory behaviors. And so the idea is even if the IAT correlates with behavior that only means that the higher you are on the IAT spectrum, the farther along that dimension of that spectrum I described you are, but it could be that the general population, most of us cluster on the egalitarian end. But it is even if you are three standard deviations above the mean on the IAT. So you're extremely high. You're at like at the 90th percentile, you still might be below the mean on this theoretical spectrum. And so is it really fair then to say that someone who has a high score on the IAT is displaying a preference for African Americans? What's your reaction to that to that argument that we don't really know how these scores map on to particular magnitudes of discriminatory behavior? And I want to acknowledge Hartblinton as the first person who I saw him advance this what he called arbitrary metrics argument. Yeah, I mean it's like a great question. And I think the way you first summarized it and that was that people who show who are on the negative end of the IAT are more likely to commit negative behaviors and positive and the positive end. And I think that's about as fine a grain as we can put on it at this point. I mean I think this is Mike Olson of the University of Tennessee. What Hartblinton and others have sort of either implicitly or explicitly kind of referred to as sort of a relative comparison or medical tests and other well normed personality tests where you know that people who score at a certain level are probably more likely to engage in certain behavior or have some sort of actual health symptoms. And I don't think that there has been that kind of parametric work with the IAT where we can say people with these sorts of scores do these sorts of things. You know, I just don't think they were there. I don't know that we ever will be. So to the extent that say like the project implicit website is giving you somewhat parametric feedback of your moderately biased or your strongly biased or whatever we don't necessarily have really good indications of what that means behaviorally. As Michael Olson just noted, you can go to the project implicit website, the link to which I will include on the web page for this episode. And at the project implicit website, you can complete an implicit association test. As I just did this Sunday evening July 8th, I completed a race IAT. And I am looking at my feedback right now. You can also receive feedback if you go to the website and take an IAT. And my data in their terms quote suggest a moderate automatic preference for European Americans over African Americans in quote. Yes, I, even I as an African American, received such feedback. And of course for white Americans, it's quite common to receive such feedback since the typical white American on a race IAT receives feedback indicating to one degree or another. A preference for white Americans or European Americans over African Americans or black Americans. I asked several of my guests for their view on the wisdom or the appropriateness of providing such feedback. And I asked in large part because of a passage from a 2015 comment written by Tony Greenwald, Mazarin Banagian, Brian Nosek. And they say quote, IAT measures have two properties that render them problematic to use to classify persons as likely to engage in discrimination. Those two properties are modest test retest reliability and small to moderate predictive validity effect sizes. Therefore attempts to diagnostically use such measures for individuals risk undesirably high rates of erroneous classifications in quote. Well, that's for page 557. I was curious about the use of feedback because if there is this risk of erroneous classifications, then it wasn't clear to me that the wisest approach would be to give individual level feedback at the website. Now of course they go on to say, these problems of limited test retest reliability and small effect sizes are maximal when the sample consists of a single person, i.e. for individual diagnostic use. But they, those problems, diminish substantially as sample size increases. Therefore limited reliability and small to moderate effect sizes are not problematic in diagnosing system level discrimination for which analyses often involve large samples. A question that I put to Calvin Leye then was, why not simply offer aggregate level feedback to people who visit the Project Implicit website. Given the risk of, quote, erroneous classifications, why not stop providing the individual level feedback and simply provide group level feedback. For example, communicating to visitors the percentage of Americans or the percentage of white Americans receiving scores indicating a preference for white Americans or black Americans which would still convey the prevalence of implicit bias as indicated by the IIT. That's a, that's a good question. And so it's not an option that we've considered before. What I think is important for participants to be aware of. for to expense. who visit our site to get in some form is some type of information that they can use to reflect upon themselves. And part of that is, and by the way, all of this I'm speaking as to my personal views about as project placid, we're a bunch of academics, so we disagree on everything. But my personal view is that that type of kind of self-reflection is important because oftentimes you can kind of read these about these implicit biases and kind of just only understand them abstractly. We want to give people that kind of subjective experience of these issues, and part of that plays out in taking the IIT itself and feeling the tension and perhaps feeling that you're slower in one case compared to the other. And I think that in terms of motivationally, folks want to learn about themselves. Now we can't just tell them whatever they want, we have to tell them stuff that is calibrated relative to the evidence, but I do think that in terms of kind of motivating folks to learn about implicit bias, there has to be some type of carrot. That's just my sense from how folks have run online research in the past. So maybe there are ways of doing it where we say, "Hey, we would like you to take this. We're not going to really tell you anything at the end, but we're going to give you a page." But then I think there's a lot of kind of lost opportunity there. So it's a very legitimate question. That's Brian Nosek again. And there's been lots of good discussion about it, and I do have a point of view that on it that's obviously different than yours because I have been giving feedback at the website and then adores that for many years. So let me tell you a sort of how I think about it. So the key part of what we were talking about in the 2015 piece and that we've written about in many prior ones is diagnosis. The IAT does not provide a diagnosis of anything that you would understand the term as diagnostic. It's feedback, right? Just like the tests that you're giving at wrap-up of your semester now are not diagnoses focusing your students' math abilities, psychology abilities or otherwise, right? Those tests are tests. You think they're valid enough to administer methods for understanding students' performance in the class, and you even use them to classify students in terms of providing them grades. But I doubt that you would defend any of those tests as a diagnosis of their underlying ability in any sustained way, right? Ask some different questions you might have gotten different scores. So the feedback itself is a broad class category of activities compared to diagnosis, which is you use this to have some selection of that individual, right? So using the IAT to decide who should be on a jury or not based on their score would be inappropriate, using the IAT score to decide whether or not someone is eligible for a job or not would be inappropriate because of the classification errors. Those classification errors are consequential on those outcomes. The benefits of feedback, I think, on the website are for the educational purpose. I think that with information, even without information, but especially with information, people are smart enough to know and understand and contextualize what feedback means. And in fact, engagement with people on the website, the engagement of people in classrooms and discussion groups about these performances shows that people are not snowflakes and not unable to reason about what the results of a test mean and not overemphasize the performance on that task in terms of its diagnostic capacity. And the website has lots of context for trying to educate and explain about that. The interesting thing. Could I jump in and just say that? Yeah, please. Please. And given some of the context provided by your answer, I run the risk of seeming to impune the intelligence of Project Implicit website visitors when I asked this, but if it were shown that, say, the majority of people who receive that feedback actually do interpret it, despite your caveats, which people may or may not read in their entirety. If the majority, if it were shown, the majority of website visitors actually did interpret that feedback as having a diagnostic character. So for example, they interpreted that feedback as implying an enduring characteristic of them has been revealed. Would that give you pause as you consider the wisdom of continuing to provide feedback? Yeah, I might cripple with the specific element of what they might get paused about, but in the general sense, yes, right, if people can't if we can't communicate effectively and contextualize effectively what that feedback means and use it for its purpose of the feedback is as an educational device, right, to engage people in trying to understand how thoughts and feelings might occur outside of conscious awareness or control and how those might have characteristics different than our conscious beliefs and values. If feedback interfered with those educational goals, then yeah, then our educational goals aren't being achieved by providing feedback and we'd want to revise that. So we have been, you know, we have questions at the website I've had for many years, questions to get people's reactions. What do you think just happened? It's your interpretation, what does this all mean, et cetera, et cetera, et cetera. And by and large, all of those things across the different domains we studied show a lot of positive, what you would consider positive educational engagement, right, it's not that everybody feels wonderful about the score that they got. That's not education doesn't try to pat people on the back and make them feel wonderful. Good education can be challenging, can be difficult, but it should be educational, right, it should engage people with the problems that that area of research is trying to solve and have them think about it hard and gain some insight and have ideas that are still skeptical about what the outcome, et cetera, et cetera. And on those counts, I think the website has been very successful. Another sort of instructive element about the feedback is that we've tried a couple of different things over time with regard to the feedback, one was that early on we had some technical problems where sometimes people wouldn't get any feedback. And that would elicit the most angry emails of any emails that we got, right, the most common complaint for the first three years of the website is I didn't get my freaking feedback, right, you guys are idiots, right. So people came to the website in order to get the feedback, that was part of the appeal, that's part of why people continue to come is that they find that to be an engaging element of the experience. A second piece of evidence regarding people's desire for that and their ability to sort of handle it effectively is that we have at times, I don't think we have it there anymore, we have at times had an option where you could choose to get your feedback or not. It might actually be there for a couple of tasks still, I can't remember. And most people choose to get their feedback, I said, yeah, everybody chooses to get their feedback. Right, nobody chooses, no, no, I don't want to know, right, everybody chooses because that's why they're doing the tests, right, they're not doing the test to give us data, which they are doing, right, we get tons of data, they're doing the test because they think this is an interesting thing to try out, just like all of these other crazy quizzes and things that people are willing to figure out, right, what, what Star Wars character are you, right, what astrological sign actually fits you the best, you know, there's all kinds of different tests, most of which have zero validity that people are interested in engaging in and it's not that they're now taking all of that feedback seriously, oh my God, I'm Boba Fett. I didn't know it was Boba Fett, oh my God, I am Boba Fett, right, they recognize what it is. So you would say, of course, this is coming from the Harvard website, it is from researchers, so they will take it more seriously. And I hope that they will compare to the Star Wars, Boba Fett characterization or whatever it is. But the goal of the website is educational, right, give them something where they can then wrestle with it and they can talk about it and they can think about it and they can compare with others and they can decide what it is they think about that based on the evidence that they make available for them. I remember people asking me like why I chose Carlton College for college and I have no idea. I was like, okay, we'll tell you a story, I'm like, okay, I can make something up. That's Samine Vazir of the University of California Davis. But recently I went to hear the editor-in-chief of Nature talk. And the issue of blinding came up and he's like, "Oh, no, we don't blind our manuscripts, but we tell our reviewers and editors not to use opposite identity." And they don't use opposite identity making their decisions. And I was like, "What do you think, how can you just say that?" You can say that that's your value, but you can't say that even you personally don't use that information. How do you know? And I would be willing to bet a lot of money that you do use that information unbeknownst to you or against your own explicit values. So I think it's important to acknowledge that what we think we should do is not necessarily what we do. And also the reasons why we are doing something are not necessarily known to us as social psychology as shown. So especially when it comes to reasons. Like, why did you make that decision or why did you know? I think it's very important to acknowledge that we're not necessarily aware of all the factors of influence or decisions. I mean, that's very relevant to the message of IAT researchers, too, I guess. Exactly. You took the words out of my mouth. I mean, so part of. One can at least. I can imagine that one rationale for giving people feedback at the website is that it may prompt individuals more so than they otherwise would to take seriously this possibility that there are factors that are operating outside of their awareness influencing their judgments. I wonder if even if the data. I'm not asking you to endorse this premise, but even if the data were. Even if it were the case that the data don't warrant giving individuals that diagnostic feedback. I wonder if you think that the value of prompting individuals to consider that possibility is worth it, even if there are lots of misclassifications happening. So let me play devil advocate. Again, I don't have a position on whether they should or shouldn't give feedback. I'm trying to imagine. So like, one issue I care a lot about is blinding. I think that a lot of people, when they are in the role of editor reviewer, they think that they can ignore the author and the institution, and so therefore it's fine to give them that information. They're not going to use it anyway, or they're only going to use it in accurate ways. They're not going to have any bias or whatever. And what I wish there was a tool I could use to give people and say, "Look, I did this implicit test, and you are susceptible to status bias." So let's imagine that I made up a website that gave bogus feedback. I just told everybody who took the measure that they were susceptible to status bias. I got all our editorial board and reviewers to take and convince them that they are susceptible to status bias. So therefore they should not have that information. You could argue would have a beneficial effect. But okay, I managed to convince everybody of this thing that is probably true. Most people probably are susceptible to status bias. Just like most people probably would show a preference or a faster association, faster reaction time, for black and pleasant pairings and vice versa. So what's the harm in telling people that even if we don't actually have evidence is true for them, if it has this positive consequence? So I don't know what the answer is, but you could swap out your example where it has low diagnosticity with one where it has no diagnosticity whatsoever. But we are pretty confident in the group mean. So why not just give everyone individualized feedback with the group mean, because it's probably true for them, right? I think the reality from what I understand that the diagnosticity of an individual score is not far from what it would be if you just substituted with the group mean. So why not just tell everybody that that's their score? So I don't know, like I think it raises ethical issues about whether that ends just to find the means and whether over precision, how much of a sin is over precision? Like claiming that you have more precise information about a person than you do? I don't know the answer to those questions. But I sympathize with that feeling of like when I talk to other people who are resistant to the idea of status bias and who really think they're purely objective when they're evaluating a paper, like I tell them well, when I started blinding myself, they felt really, really different. So until you do that, I won't believe you that you're not susceptible to it. And I wish there was something, some website I could point them to and say take this five minute test and you will see with your own eyes that you're susceptible to this. So I understand the appeal of that. But then where does it stop? Like when we just lie to people and tell them that we have evidence that they're susceptible to this bias. Because really, we don't even need them to take a test at this point, right? We know that the average American adult is going to show this preference. So do we even need them to take the IT or can we just skip to the feedback? Unless the receiving the individualized feedback motivates people to support a process like blinding more so than the otherwise would. Right. But then it's the end just to find the means. Right. Yeah. So the. It's Calvin Leigan. Paper that excites me most, or that has excited me most historically, is one where we were interested in what were the factors, what were the types of interventions that would change these implicit racial biases most, right? Yeah. With the idea of being that, you know, perhaps if you reduce these implicit racial biases that might kind of have potential impacts for how people think and perhaps ultimately behave. Sure. And so we asked researchers from all across social psychology to submit to us a single best intervention that they could think of for reducing implicit racial preferences to zero as measured by this IT. And we ended up getting these kind of 18 interventions and we just tested them all against each other within the same studies. And what we found is, first off, even when we tell people to submit the single best intervention that they could think of, half of them didn't come up with interventions that even worked with the shocker. So nine of the 18 worked. Yeah. And then of those nine, we noticed that there were some differences. So the ones that were the most effective interventions tend to be ones that exposed people to experiences that kind of defied their stereotypes. Whereas the ones that were kind of least effective were ones that you might often see in diversity training, approaches that involved kind of reflecting on your egalitarian values, taking the perspective of a black individual. Those types of interventions tend to be less effective. And then as a final piece for this, what we did is we took those nine interventions and we went to see if they continued to reduce implicit bias for 24 to 48 hours later. And to our surprise that we found that none of them did. So just think that we need to go back to the blackboard in terms of figuring out how to design kind of brief efficient interventions for changing implicit racial preferences. So unless I misunderstood you, it sounds as if given the nature of some of the interventions that interventions that did not work. It sounds as if even if the feedback at the project implicit website is motivating self reflection, it's not clear that that self reflection actually does reduce whatever forms of implicit bias are revealed by the IIT. That's correct. And we wouldn't expect it to. I think the goal of the feedback is to promote self reflection and just to gain more knowledge about what implicit biases are. So the project implicit is a nonprofit. Our mission statement is about research and education on implicit bias. And so our primary goal in the feedback is not necessarily to create some kind of transformative change, be it morally or in terms of changing these implicit biases. Instead, the goal is to educate and give them knowledge about implicit bias and to use that knowledge in a way that is in keeping with the scientific literature. Ultimately, we aren't concerned with how quickly people associate words with white and black labels were concerned with discrimination and unequal treatment in various ways. This is Keith Payne again offering some of his thoughts on interventions, including implicit bias training. The topic of implicit bias training is everywhere these days. And when I hear that, I cringe a little bit because that phrase "implicit bias training" suggests wrongly that some kind of training session is going to change the conceptual associations and people's minds. And that's a point that critics sees on to criticize implicit bias training too, but it's completely unrealistic. Every implicit bias training that I've seen or been a part of is usually just educating people about the fact that you don't have to be an explicitly bigoted person in order to treat people unequally and providing strategies that people might use to try to be unbiased. But the practical implication of the bias of crowds model is that we shouldn't be focusing on trying to change people's associations. We should be free. focusing on structuring environments well, so that either they don't cue the negative stereotypes that we worry might be harmful, or that even if stereotypes are highly accessible in an environment that the decision-making process is set up in a way that makes it less prone to bias, whether that's very basic things like blinded resumes when making a decision, or more structural aspects of having diversity not only in the organization, but at visible levels of leadership, so that you're queuing positive stereotypes in that environment rather than queuing negative stereotypes. - So any advice for the CEO of Starbucks says they said about this campaign that I'm not going to be cynical, I'm not going to assume that it's simply brand management. I'm going to be there sincere about wanting to reduce the likelihood of incidents such as what happened to Philadelphia and famously recently. Any advice for such a CEO? - Well, I would advise, and I actually have advised, people involved in this kind of bias reduction effort to not focus on the minds of the people involved so much as on the business process. So if you've been in Starbucks, you know that the baristas are incredibly talented at keeping the line moving and making lots of drinks efficiently, and they can call out names like a no-whip skinny mocha with no problem, right? They've got this routine really down. And I think the best way to reduce bias is not to focus on the people's attitudes and values, but make aspects of inclusion in that workplace part of their daily routine that they get trained on from day one, so that whatever becomes practice becomes routinized, it doesn't just have to be about making complicated drinks, you can also get very skilled at checking social situations as well. - And I suspect that there's not going to be a one-size-fits-all approach. So I go to Starbucks frequently here in Auburn, Maine, where I live, Auburn and its neighbor across the river, the Western hall together comprise no more than about 65 to 70,000 people. This is not a large metropolitan area. When I go into Starbucks, and I'm a regular Starbucks, when I go into Starbucks, they know my name, they often know my order, and there's a way in which not to say that training and inclusion wouldn't be valuable even there, but it's a different baseline to at least some of their, say, customers of color than in a more densely populated urban area where the baristas typically not going to know the people who are coming in. And so I think that anonymity may pose its own challenges to inclusion in the absence of robust training. - No, I would agree with that. I think top-down training initiatives can only do so much. They can make people aware that issues are important and need to be considered, and they can set the norms that, in this organization, this is an issue we care about, and here's how we expect all our employees to behave. But a lot of the individual plans and strategies and day-to-day interactions have to come bottom up from the individuals, the employees, and in that local Starbucks store, right? And I think you're right that those interactions and individual strategies are probably going to be very different depending whether you're in Auburn and Lewiston or Chapel Hill or Manhattan. In 1998, the Journal of Personality and Social Psychology published a paper by Tony Greenwald, Debbie McGee, and Jordan Schwartz titled "Measuring Individual Differences in Implicit Cognition, the Implicit Association Test." That paper introduced the IAT to the published literature and social psychology. It's been 20 years now, and I asked my guests to reflect on those 20 years, including their thoughts about the initial promise of the IAT, and how it's performed relative to that promise. First of all, 20 years has gone by fast. That's Mike Olson. But secondly, I think when you look at some of the theory, not only in prejudice, but in other aspects of psychology, about processes that are presumably happening underneath the surface that people might be either unwilling or unable to report on a survey measure, the prospects of having a tool that could get in the mind and find out what people are feeling and thinking without having to ask and it's really exciting, and so what a cool tool. And then this tool, along with other implicit measures, has allowed us to come up with more rigorous tests of other theoretical ideas, like notions of say averse of racism, where you see disconnects between what's happening at the automatic level, where people have these biases and the control level where they claim that they don't, or at least they don't want to. And there's just some really good theoretical progress made on stuff that has nothing to do with implicit measurement, but has something to do with implicit social cognition because of the IEP and other implicit cognition measurement tools. But I think that's a really cool thing. Now, I think that some claims have been a little overstated at times. I think claims about the unconscious nature of the bias have been overstated. There's good evidence now that that's not true. Claims that it is tapping into the true you, whatever that means. There's a little oversimplified. So again, I think the further you get from the people who are doing the scientific work and the more out to the public, we get sometimes those claims get a little overstated. And that's where I start to feel like the promise has been a little overhyped. So when I first came to grad school, that's Calvin Lai. I thought that we could change the implicit biases. Now we just need to figure out or we just need to kind of find the best intervention because there have been so many that have been published. And then we could just scale it up and develop a real good full-blown long-term intervention for reducing implicit biases. None of that has really happened because we can't figure out an easy, consistent way to do that. It teams up many of the ways to change implicit biases. Are quite difficult. And so that's something that I've changed a lot in my beliefs about over time. Another is about understanding how to think about not just implicit biases, but how to think about mental phenomena predicting behavior. You know, predicting behavior is incredibly difficult. It's multiply determined at any given moment you have all these things in your mind and in the situation that are pulling you one way or the other. And one thing I've been routinely humbled by in trying to figure out when implicit bias predicts behavior is figuring out when it doesn't, when it doesn't, given that we live in this multitude to put to a term in the world. I do think it's what the bias is real, even if I've skeptical about certain claims about scope of power. That's the journalist Jesse Single. So that's something like the idea is sort of a good focusing device to get people to think about these issues. As long as you're doing the way where you're missing, those people are spreading false information and I think those are both the problems. I think there are a lot of potential downsides. The one I thought about the most is the way it does let white people off the hook in a certain sense. If I and other people are white and we actually need major policy changes to address racial equality and we need major forms of redistribution to address racial equality, I don't think it's the IIT or everybody else. And I think there's a risk of people going through the ritual of taking the IIT, tweeting about the results, saying time and how bad they feel and how profound the experience it was, and then going back to the simple politics and simple life they had. You don't really see stories about people taking IIT and getting politically radicalized and I'm not trying to find any value issues about socialization, but I do think like America's raised problems are severe and deep-rooted enough that they would require major policy efforts to change in the fix. And I don't think the IIT really provides a path to that or I haven't seen any evidence. So my beliefs about the IIT have evolved a lot. That's brilliant, no sick. The beliefs about the promise from where it started in 1998 have fastly exceeded what I thought would get accomplished. with the IET way, way, way beyond. If you, it's very instructive, given, you know, in this 20-year sense to go back and read the original Greenwald McGee and Schwartz paper in 1998. It is so modest in its claims. You would read it and you'd say, oh my gosh, like this does, this could have been written by a skeptic, right, of thinking that this wasn't gonna go in much of anywhere because it doesn't make very many, much of any claims beyond that, that initial bits of evidence. And, you know, we knew, like when building the website, we knew that this was super interesting to do and we hoped that by putting it on the web, that maybe we'd get sort of double our sample size of people that got exposed to it from a lab in a productive laboratory's lifetime, right? So I was thinking how many people might run through a lab in a lifetime of a lab? Maybe 50,000 people across my career. It wouldn't be great if we got 50,000 more people to be able to experience this as an interesting technique if we put it online as well. And of course, we got 50,000 in the first three days and we thought, oh my God, this is something different. And, you know, now it's 25 million or whatever. So in terms of exposure, of interest, of education, and of research, of actually advancing understanding of the domain of implicit social cognition, it has vastly exceeded expectations. Now, part of your question I presume is the, is this idea of promise of the technique, which is there have been many particularly practitioners who have engaged in what might be politely said, vast overclaiming of not just what the IIT measures, but of implicit bias in general, and the potential for training about implicit bias to change anything, that domain of translating this research into trainings and what claims are made about the effectiveness of training is a C of overclaiming. And so that part of it is an area where, you know, of course, that was going to happen, right, like with any fattish thing and the IIT certainly had as fattish elements that will occur. And part of the responsibility that me and my colleagues have tried to maintain is a consistent voice about what the appropriate claims are for what the evidence suggests. And, you know, that doesn't always, isn't always loud enough for the claims because these are now nationalized claims rather than just local to research community, but that correction keeps occurring and will continue to occur. - I think my understanding of it has oddly come full circle. - Let's keep the pain. - When the IIT first came out, the way that it was talked about, and the way that we explained it to our undergraduates, for example, was that society has all of these stereotypes and biases and that regardless of whether you are somebody who is high in prejudice or low in prejudice in the traditional explicit sense, we're all vulnerable to having these biases. I think that is very close to where I'm ending up with the bias of crowds model. It's just that we took a detour through a heavy emphasis on individual differences and talking about levels of bias as something about an individual's attitude, an individual's trait or an individual's unconscious beliefs. And I think that has proved not to be as productive as the research looked at the beginning, but still today in 2018, if you look at average levels of disparity, they're enormous. The average white family has 13 times the wealth of the average black family. Job applicants turning in a resume with a name that implies that they're black get half the rate of callbacks that somebody turning in the identical resume gets if their name implies they're probably white. And so what we see is large levels of actual disparities matched with large levels of average bias. And in the states and counties and cities where the average bias is the highest, those disparities are also the highest. So I think the promise of the measure is still there, but we need to stop thinking about it as an individual's level of bias as if it tells us something about that person's beliefs or values and start thinking about it more as an aspect of the social environment that people are embedded in. And that's like I said, where the concept really started for a lot of people, at least for me. - That's it for Tatter. I want to thank this episode's guests. That's Calvin Leib, Brian Nosek, Mike Olsen, Keith Payne, Jesse Singleton, Samine Visier. Go to tatter.fireside.fm and find the page for this specific episode where you will see links to more information about each of my guests. You'll also see other links, including a link to the project complicit website where you can take NIT or even lots of IITs. I also include links to the two podcasts that I mentioned at the top of the episode as well as a link to the Very Bad Wizards podcast. On Very Bad Wizards, Psychologists David Pizarro and Philosopher Tamler Summers talk about a variety of issues in psychology and philosophy, including in a recent episode, a discussion of implicit bias. So check out their episode on implicit bias. If you want to offer feedback on this episode or any past episode of Tatter, you can do so via Twitter. The handle is @Tatter_Rags. Also to financially support Tatter, go to patreon.com/tatter. Note that there are different rewards available for different levels of support, including the chance to vote on future topics and guests and to get advance word on who I have interviewed. In any case, thank you for listening to this episode and be well.

Podcast Summary

Key Points:

  1. The Implicit Association Test (IAT) measures unconscious racial biases but has weak predictive validity for real-world discriminatory behavior, with correlations typically ranging from 0.05 to 0.3.
  2. IAT scores correlate more strongly with voting intentions (e.g., Trump vs. Clinton) than with actual racial discrimination, suggesting limited utility in predicting everyday bias.
  3. Structural inequalities—such as segregation and unequal access to opportunities—shape implicit biases, creating a feedback loop where bias both results from and reinforces systemic disparities.
  4. Historical data show that counties with higher historical slavery rates today have higher average IAT racial bias, indicating deep-rooted structural origins.
  5. Individual IAT feedback is ethically and scientifically problematic due to low diagnostic reliability and potential for misclassification, though it may promote self-reflection.
  6. Interventions like diversity training are ineffective at reducing implicit bias, while environmental changes (e.g., blind hiring, inclusive leadership) are more effective.
  7. The IAT’s promise has been overstated, especially in claims about individual unconscious bias or the power of training to change behavior.
  8. The most valuable insight is that implicit bias reflects broader social environments, not individual morality, and should be addressed through systemic change, not personal introspection.

Summary:

The episode explores the Implicit Association Test (IAT) and its role in measuring and understanding racial bias. 3. Studies show stronger links between IAT scores and voting behavior than with everyday discrimination.

The data suggest that implicit bias is shaped by structural inequalities, such as historical segregation and socioeconomic disparities, with regions that had more slavery in the 1860s showing higher current IAT bias. Critics argue that individual IAT results are misleading and often misinterpreted as diagnostic, risking false self-assessments. Instead, the most effective solutions involve structural changes—like blind hiring and diverse leadership—rather than individual training.

The research also highlights a vicious cycle where bias both stems from and reinforces inequality. Though early claims about individual unconscious bias were overhyped, the IAT remains valuable as a tool for raising awareness and fostering public reflection. Ultimately, the focus should shift from individual levels of bias to systemic environmental factors, emphasizing that real change requires policy-level reforms rather than personal introspection.

The episode concludes with a call to reframe implicit bias as a reflection of society’s structures, not individual shortcomings.

FAQs

The IAT has weak to moderate predictive validity for racially discriminatory behavior, with correlations typically ranging from 0.05 to 0.3. For voting intentions (e.g., Trump vs. Clinton), correlations can be as high as 0.6, but overall, the relationship is inconsistent across domains.

No, the IAT is a poor predictor of individual discriminatory behavior. Studies show that even when the IAT correlates with behavior, the effect sizes are small, and many behavioral measures are unreliable due to subjective and situational factors.

The bias of crowds model suggests that implicit biases are shaped by societal and environmental factors, not individual traits. High average IAT scores in a region correlate with greater racial or gender disparities, indicating that group-level biases reflect structural inequalities.

Individual-level feedback is considered problematic because of its low diagnostic reliability and small predictive validity. It risks misclassifying people and may lead to overestimating personal bias, which does not translate into real-world behavior changes.

Most traditional implicit bias training (e.g., reflection or perspective-taking) is ineffective. Effective interventions involve exposing people to experiences that defy stereotypes, and even these show minimal long-term effects on implicit bias.

Areas with higher average IAT scores show greater racial disparities in outcomes like police violence and infant health, suggesting that implicit bias is linked to systemic inequalities, though it may also be a symptom of those disparities.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.