This episode of Normal Curbs explores the controversial role of P-values in scientific research. P-values, which stand for probability values, help researchers separate signals from noise by measuring how surprising a result would be if the null hypothesis (no effect) were true. For example, in a psychic guessing game, a P-value of 30% suggests the result is consistent with random chance, while a P-value of 0.0001 suggests strong evidence against the null. The hosts, Kristen and Regina, discuss the history of P-values, tracing back to R.A. Fisher in the 1920s, who proposed the 0.05 threshold as a flexible guideline for deciding when to take a closer look at data. However, this was later combined with the Neyman-Pearson framework, which uses fixed thresholds for binary decisions (e.g., significant or not), creating a mashup that enshrined P<0.05 as a universal rule. Regina’s award-winning 2014 Nature article illustrated these issues with a story about a psychology study that initially had a significant P-value (0.01) but failed to replicate (P=0.59). The episode emphasizes that P-values are slippery and should not be treated as magical proof of a discovery. Instead, they are one piece of evidence in a broader scientific process, and their misuse—such as treating 0.05 as a rigid cutoff—can lead to false positives and missed opportunities. The hosts use analogies like dating apps to explain type I (false positive) and type II (false negative) errors, highlighting the need for careful interpretation and context in statistical analysis.
- Yeah, when we co-teach you, bring out this example, which I love, and I think it either makes some people believe in magical creatures and other people hungry. (laughing) - The magical creatures being Paul, the psychic German octopus. (clapping) (upbeat music) - Welcome to Normal Curbs. This is a podcast for anyone who wants to learn about scientific studies and the statistics behind them. It's like a journal club, except we pick topics that are fun, relevant, and sometimes a little spicy. We evaluate the evidence, and we also give you the tools that you need to evaluate scientific studies on your own. I'm Kristen Sonani, I'm a professor at Stanford University. - And I'm Regina Nuzo, I'm a professor at Gallaudet University and part-time lecturer at Stanford. - We are not medical doctors, we are PhDs, so nothing in this podcast should be construed as medical advice. - Also, this podcast is separate from our day jobs at Stanford and Gallaudet University. - Regina, in today's episode, we're gonna do something a little different. We are actually gonna talk about a few of our own papers, and we are finally going to unpack that huge statistical concept of the P-value. (laughing) - P-value stands for probability value, not penis value, who just incase anyone would win right now. - Give it our track record on this podcast, it's a fair question. I think we've mentioned more penises than probabilities so far. (laughing) - Gilti, P-values can be fun too, maybe not, you know, penis fun level, the close. (laughing) - I'm gonna stick to probabilities in my actual class, that's more my comfort zone, Regina. - I am proud of you for even saying the P word here, Kristen. - Thank you. (laughing) - Now, normally we start with a claim like, this supplement reverses aging, but today our claim is different. We're gonna focus on a claim that's about the culture of science and the tools we use in science. So the claim for today is that P-values are a flawed, statistical tool. - Yeah, it's fascinating because P-values are kind of the good guy and the bad guy at the same time. On the one hand, they have advanced modern science, yes. But on the other hand, they have been so tricky that they have been misunderstood and misused by researchers for decades. - Exactly. So in this episode, we're gonna pull the curtain back on P-values and significance testing. We're also gonna liven it up with some busy and monsters and psychic octopuses. (laughing) - Sorry, no sex this time in this episode. And I think we've already peeked on the penis mentions. - Psychic octopuses are almost as good though, Regina. - Yeah, almost. - Regina, let's actually set up what a P-value is. Most biomedical research papers report P-values and they use significance testing, which is directly related to P-values. So P-values are a powerful tool for helping us separate signals from noise. Data are noisy. There's always some random fluctuation. And without a tool like this, it's really easy to get fooled by patterns that aren't actually real. - People love to see patterns in noise. That's just how we humans are wired. - Absolutely. That's why P-values are so useful. - Uh-huh. So the definition of a P-value is very technical and unsatisfying. - That might be the understatement of the podcast, Regina. (laughing) - So let's make the P-values concrete, Kristen, with that example from our alcohol episode, where I asked you to guess a number between one and 20 that I was thinking in my head. No, not guess. Read my mind. The number I was thinking of, one to two. - Right, and I failed miserably 'cause I am not psychic and it took me six tries to get the number. (laughing) - We decided no one was going to mistake you for me in psychic, not very impressive psychic anyway. - Right, and if we actually calculate the P-value for that little experiment, which we can do, it comes out to 30%, which means my performance was totally consistent with the hypothesis that I have no psychic powers. - Right, so let's back up and talk about how we got that number. We didn't calculate the P-value back in the alcohol episode, but here, I'll start us off. First of all, we had to set up what's called a null hypothesis. And usually this is the hypothesis that you were trying to disprove the strong man hypothesis that you want to knock down. And I think of it as the skeptics world or the boring world where nothing's happening. The hypothesis of no effect. So here, Kristen, the null hypothesis is that you are not psychic, you are just guessing. - Right. And then to calculate the P-value, we ask, what would happen if the null hypothesis is true, and we could repeat the experiment again and again. And this is the part that's tricky for people. A P-value comes out of this thought experiment. This is called the frequentest perspective. It treats probability as how often something would happen in the long run. - Yes. And with this psychic ESP game, you can really picture the hypothetical long run, right? You can picture as sitting in your beautiful backyard, tossing the ball for your dog, and playing a thousand rounds of this guess my number game. - Nibbles would absolutely love it, but it would kind of take forever. So Regina, I actually had the computer simulate a thousand virtual games under the null hypothesis where I'm not psychic. And in about 30% of those games, I guess correctly in six tries or fewer. We can also figure out the probability mathematically, six guesses cover 30% of the numbers. So the probability of success within six tries is 30%. That's our P-value. And the technical definition of the P-value is, the P-value is the chance of seeing a result like this or something even more surprising if the null hypothesis were true. So here the P-value is the probability that I would get the correct number in six tries or fewer if I was not actually psychic. - Right, it's a measure of how surprising a result is. Again, if we assume the null hypothesis to be true, if nothing's actually happening. So a big P-value, like 30% here, it's just telling us that the result is not surprising at all, the smaller the P-value, the more surprising, bigger means not surprising. So it's consistent with this null hypothesis being true with Kristen, sadly, you're not having any psychic powers. - Exactly. Let's crank it up though, Regina. Imagine if you asked me to guess the number between one and a million, and what if I got it right on the first try? (laughing) - That actually might give me a hard attack. (laughing) I'm scared to even try it. You can, the truth of that happening, if you are not psychic is one and a million, right? And that's a P-value of 0.0001, basically a 1 with 5 zeros in front of it. And so in this case, we would definitely conclude that we have strong evidence against the null hypothesis. - Right, meaning we're either gonna conclude that I'm psychic or maybe that I cheat it. (laughing) These results are clearly not consistent with chance guessing. Now, Regina, those two examples that we just did, they don't actually illustrate all that well why the P-value is useful, because my intuition already tells me that in the first case, yeah, Kristen's not psychic in the second case, something weird's going on. But where P-value is actually the most useful, Regina, is in the in-between cases, where it's not obvious whether we're looking at signal or noise. - And Kristen, what counts as a loud enough signal where we draw the line for what is a small enough P-value? That is a matter of debate. And as we've talked about in previous episodes, the most common threshold is 5%. But there is nothing magical about that 5% line. And in fact, some people use different lines. And in this episode, we will talk more about that 5% and the history behind it. - Exactly. Regina, let's talk about your paper on P-values first. That paper was published in Nature in 2014. It wasn't a research article. It was a science journalism feature. And it got a ton of attention. I remember that it got a really high alt metric score, which is a measure of how much attention the paper got in the media and in social media. Regina, can you share that number with us? - I feel a little embarrassed bragging like this, Kristen. But it was a little over 5,500. Regina, I am happy to brag for you. An alt metric score of 5,500 is like 99.9% high. Very, very few papers ever get that high. I think my highest alt metric score might be something like 300, just to give the contrast. (laughs) It also got a lot of citations, too, right? - Yeah, something over like 2,000 citations. - That is amazing. Hardly any papers have that many citations. - I just think it shows the power of plain English. Because statisticians had been yelling about P values and warning for decades. But thanks to my editors, this piece made the ideas really approachable to a wider audience without talking down to them. - Yeah, exactly. This was useful to a lot of people. It actually even won an award from the American Statistical Association. - It did. I was very proud of that. I think you can use stories to make abstract concepts come alive, right? So for example, in the piece I. I had a nice lead Fred. I opened with an anecdote about a psychology grad student who had a really sexy result. It was that people who were political moderates could literally see shades of gray better than extremists on either the left or right. - Wow, so not just metaphorical gray areas, but actual visual shades of gray. - Right, isn't that gray? So his original study had a p-value of 0.01, which is highly significant. And everyone was excited. He was probably imagining fame and New York Times coverage, TED Talks, but then he and his advisor tried to replicate this study with a slightly larger sample. And this time the p-value was 0.59, not even close. - Wow. So they could no longer claim an effect. That must have been crushing for the researchers. And you used that story Regina to draw the readers in. - Right, I wanted to get them curious about what went wrong there. And then from there I could go into the deeper, wonky issues about p-values. And the big message that I wanted them to get was that p-values are surprisingly slippery. They look precise, but you've got all these hidden assumptions behind them. - Right. And it's a good reminder that a p-value of 0.01 does not mean you've discovered something reliable or true. - Right. So the point of my article was that p-values are not magical. And they can cause big problems if you misinterpret them and treat them like they are magic. - One of the things I loved about your paper Regina is that you dug into the history of p-values. And I think a lot of people don't realize the history in the context. Could you give us a quick taste of that history? - Right. So the p-value was developed in the 1920s, so 100 years ago, by R.A. Fisher, who was analyzing agricultural experiment in England, trying to figure out which treatments were worth following up on. And he came up with the p-value as an informal way to judge whether results were interesting enough for just a second look. It was supposed to be part of this whole flexible process where you're blending data and background knowledge. The whole thing was just very synergistic. It was not supposed to be a rigid rule, which is definitely not how we use it now. And wasn't he the one who gave us the magic 0.05 number? - Yeah. That was Fisher. He was the first to suggest 0.05 and to introduce the word significance. But he never meant 0.05 to be this universal cutoff. He wrote things like, it's convenient to take 5% as a standard level of significance. So convenient. It was just a guideline for deciding whether data needed more attention, whether you could take a second, look at it. And he saw significance as a continuum, so smaller. P-values meant stronger evidence. It was not a binary yes-no. That makes sense. And 0.05 actually fits people's gut intuition. We're doing a class, we do this little experiment where we ask students, at what point if you keep flipping a coin and it keeps coming up heads? At what point do you start to get suspicious and think that it's an unfair coin? And for most students, it's when we get four or five heads in a row. And that happens to line up right around a P-value of 5%. - Amazing. So there is something I think to the 0.05. But again, just as a general guideline, maybe not a hard and fast rule, my old boss at the American Statistical Association had a great dating analogy for P-values. He said a P-value is supposed to be swiping right on a dating app. It means, this person looks interesting. Let's take a closer look. Let's go out to coffee. Not, hey, let's get married. But somewhere along the way, we turned it into a whole marriage proposal. - That's a great analogy, Regina. And it really shows how far we've drifted from Fisher's original idea. - Right. So Fisher was shaping this whole thing, his whole P-value thing. And he had rivals, Jersey Naaman, who was a mathematician from Poland, and Egon Pearson, who is a statistician from the UK. And they came up with their own system. And the feud between Naaman and Pearson and Fisher was-- Christian, it was just legendary. - This is the tabloid side of statistics. It's the juicy academic gossip. - Oh, I love it. So the feud got really heated both in person and what they were writing. So Naaman once described parts of Fisher's work as worse than useless. Which is actually kind of a sick burn mathematically, like, whoa, you know, he got you. Yeah. It was double edged on that. Fisher shot back. He said Naaman's approach was childish mathematics and said it was horrifying for the intellectual freedom of the West. - Wow. I love this stuff, actually. Listeners may forget some of the math, but they're going to remember these guys shooting like a bunch of kindergarteners. So Regina, explain why they were fighting so hard. - Yeah, the heart of it, they had very different philosophies. Fisher wanted things fluid. The p-value was just one piece of evidence, a little flag to tell you to look closer. And Naaman and Pearson wanted rules. They were trying to get rid of all of Fisher's squishy fluidity and they were trying to help you make a yes-no decision based on your data. So, Christian, it would be, are you psychic? Yes or no. And it's going to be nice and clean. It's a decision, but in order to do that, you need to sit down ahead of time before you collect your data and make some hard choices and trade-offs. So, I have an analogy for this. Are you ready? - Yeah, absolutely. (laughs) - Good idea. All right, Botox. Oh. Naaman Pearson is like the Botox itself. It gives you your forehead that is smooth, it is a baby's bottom, it is rigid and reassuringly pristine. There's no wrinkles, no mess. And Fisher is kind of like natural eyebrows without Botox. It's in your whole forehead. It's expressive, inflexible, but a lot of wrinkles. (laughs) And a lot of mess in there. So, I see the Naaman Pearson/Fisher clash as like arguing which is better. A nice clean Botox forehead that you can't move or afflexible forehead that's all wrinkled. (laughs) Botox works surprisingly well for stats and allergies. We've used Botox as an analogy for confounding before as well, Regina. - Yeah, there's something about Botox isn't there. So, the Naaman Pearson framework also gave us the vocabulary that a lot of people remember from their interest stats class and have nightmares about like alpha and beta and statistical power, type one and type two errors. - These things are so poorly named, type one and type two error. That is completely nondescript. - Oh no, right. Christian, I think they ought to let us rename all the stats concepts. - The two. - The two are totally rebranded. It would be so much more fun and easy to remember. (laughs) - It would. So, they're probably not going to allow us to give dating app names for all of these, but I do have a dating app analogy for type one and type two errors. - I love it. - Okay, type one error is false positive. And this is like swiping right on the dating app when you should not actually do stuff, right? Like for me, the guy has great profile photos and is promising bio. So, like swipe right, we go up and then date and the date turns out to be a total dud. (laughs) Not good, should not have done this. So, the science version of this is you're publishing something that turned out to be false, right? You're putting it out there. So, this is the false alarm, the false positive. I call it the over-enthusiastic error. - Okay, so that's the type one error, but what about type two error? - Type two error, right, that's the false negative. So, this is the opposite. This is swiping left on someone who looked like meh. You know, kind of unimpressive online. I didn't like his photos, his skin looks bad. You know, he's kind of cross-eyed. But, I never know this, but in real life, he would have been amazing, but I never met him, 'cause I swiped left. So, I never got matched with him. He's the one, the hypothetical guy who got away. And so, the science version of this is you've got a real effect in your data, but you never found it. And this is the false negative. So, I call this one the missed opportunity error. - Right, that's the type two error. - And I actually do think about these things when I am swiping on my app. But, it's the cost of a lot of false positives, a lot of bad dates, but it's the cost of a false negative. I might have missed my soul mate. - Oh, real life decision theory at work, Regina. - Uh-huh, uh-huh. So, Namin and Pearson said, okay, you got these type one error, type two error. You have to decide ahead of time which error matters to more and set your tolerances. What can you handle? And you set up--
the decision rule based on that and you follow it and everything is very binary and very clean. Just to connect the two approaches, Namin and Pearson had you set a rule before you ran your experiment. You picked a threshold, and if your results cleared that bar, you got to reject the null hypothesis and say there was an effect. Otherwise, you did not. That threshold was essentially based on a p-value cutoff. There's p-values under the hood in both. Right, but the way that it works in practice is very different. So remember with Fisher, we were able to say that the smaller the p-value, the more evidence against the null. But under the Namin Pearson framework, once you have designed the experiment, the actual p-value does not matter. Let's say we chose 0.05 as our threshold. A p-value of 0.049 just under that threshold is just as much evidence against the null hypothesis as p equals 0.00001. There's no grades of significance. So you could say I had an effect or I did not. So Fisher thought this Namin Pearson approach was being counting. Not good for scientific discovery, and Namin and Pearson thought Fisher's way was very sloppy and prone to errors. And in the end, so these guys were all feuding, scientists were getting impatient, and they did not realize that these were actually two whole different philosophies on how to do science, right? Because as you pointed out, this kind of has the p-value hidden underneath on one. Oh, it's just the same thing. So they mashed the two systems together. They took this easy p-p-value and crammed it into this rule machine. And that is when p less than 0.05 got enshrined. Everyone knows this as this universal threshold of statistical significance. Right. So in this new mashup system, we get to calculate a p-value and then the rule that we're looking for is just if p is less than 0.05, but that's very different than what Namin and Pearson actually intended. Right. Because we also set up 0.01. And then you get to say it's highly statistically significant. And you can add as many stars as you want in your table. And Namin and Pearson would not have gone for that at all. And Fisher wouldn't have liked this strict cutoff. Nope, nope. They're all turning over in the graves right now. All right, so that's some great history, Regina. Now I want to talk a little bit more about what was in your nature paper, but let's take a short break first. Regina, I'm really excited about our new affiliate partner graph to table. This is a web app that automatically extracts data directly from figures. I actually discovered it while sleuthing for this podcast and loved it so much that I reached out to them. To know more eyeballing a chart in a paper or getting out a ruler. Exactly. You upload one or more figures, and it automatically extracts all the data into a ready-to-use data set with the variables already labeled. It is a huge time saver for me. For those who know a web plot digitizer, it's like that, but better. It is kind of magical. You can find them at grasstotable.com. That's graph the number two then table, or you can link to them from our website, NormalCurves.com. And Regina, our listeners get 20% off with the discount code NormalCurves20. That's all lowercase. And using that code also helps support the podcast. Regina, I've mentioned before on this podcast our introductory statistics course, Demystifying Data, which is on Stanford Online. I want to give our listeners a little bit more information about that course. To self-paced course, well, we do a lot of really fun case studies. It's for stats and offices, but also people who might have had a stats course in the past, but want a deeper understanding now. You can get a Stanford professional certificate as well as CME credit. You can find a link to that course on our website, NormalCurves.com. And our listeners get a discount. The discount code is NormalCurves10. Welcome back to NormalCurves. We were talking about Regina's nature paper on P-values. Regina, in that paper, you laid out some of the big misconceptions with P-values and with this hybrid system that we've talked about. Let's walk through a few of them. So, first misconception is that people confuse statistical significance with practical significance. And part of the problem is indeed Fisher and his word choice, because in everyday English, significant means important or meaningful, but Fisher really meant it as just like, oh, worthy of a second look in there. Right, because you can have tiny, meaningless effects really, really small that are also highly statistically significant. Mm-hmm. We saw this in our age gaps episode, because the data showed statistically significant effect of age on a woman's romantic appeal, which sounds super depressing until I went in and looked at the size of the effect. And they were rating romantic appeal on a 1-to-5 scale. And for a woman to fall from a perfect 5 to a rock bottom one, it would take her 628 years. Just to be clear that 628 years was a wildly inappropriate extrapolation, but we had done that just to illustrate how tiny that effect was, even though the P-value is statistically significant. Right, but practically meaningless. And I think the surprise is people Regina, but it's because the P-value doesn't depend only on the size of the effect. It also depends on your sample size and also on the noisiness of the data. So you could have a teeny tiny effect, but if the sample size was large enough, you would still get a small P-value. Like a teeny tiny difference between the groups could be statistically significant, simply because the sample size was huge. Right, and that is why some statisticians have started replacing the phrase statistically significant with a new one, statistically discernible, meaning statistically detectable. We can see it. In fact, the new addition of the stats textbook that I use has completely adopted the language. If the P-value is at some point of five, the effect is statistically discernible. No more significant in the book. Interesting. That's such a small shift in language, but to clear up a lot of confusion. Yeah. Next issue with P-values in this hybrid system, P-hacking, which might be the only stats term with a definition in the urban dictionary, which is usually a little bit more for off-color terms. In the urban dictionary, they call it exploiting, perhaps unconsciously, researcher degrees of freedom until P is less than 0.05. Nice definition. Yeah, but to give some context on that, remember, in that name and Pearson framework, we just talked about everything was supposed to be stat in advance, your sample size, your test, your hypotheses, it was pretty rigid, and that rigidity is what protected you. But that is not how modern science usually works, because of this mashup that we described. Today, researchers actually have a lot of flexibility. What we statisticians very neutrally call, researcher degrees of freedom. Regina, it sounds great that they have a lot of freedom, but actually, it can be quite dangerous statistically, because researchers think that in order to get published, or to get promoted, they need to get their P-value under 0.05. So they end up doing a lot of things, either consciously or unconsciously, to try to get that P-value under 0.05, basically, like gaming the system to get a P-value less than 0.05. We talked about some of these P-hacking things in previous episodes, especially about multiple testing. We talked about that one in the bad boyfriend section of the step-to-review episode. Right. That was the case where, say, you are comparing an intervention in a control group, and you compared them in a whole bunch of different subgroups, or you looked at a whole bunch of different possible outcomes. Every time you run a new statistical test and calculate a new P-value, you increase the chance that even if your intervention does absolutely nothing, that just by random chance, you're going to get a P-value under 0.05. If you run enough tests, eventually, you're going to get a P-value under 0.05. Right. Sometimes they call this torturing the data, torturing it until it confesses. Right. Because then now it is very tempting to cherry pick just the results that came up significant, might have been significant just by chance, and publish just those. That's the problem. So it's not really the P-values fault. It's our fault. Not really a math failure. It's a human nature failure. And, Kristen, this is why you and I get so crazy excited about pre-registration and registered reports and open data. Right. Very sexy. Very sexy. The reason we get so excited about them is because they are forcing transparency on the researchers and that transparency is what helps keep P hacking in checks so it doesn't spiral out of control. Right. It's a lot like what Naiman and Pearson had originally envisioned where you set everything rigidly ahead of time so that you can't cheat later. Another problem with a way that we use P-value
is that we treat point zero five as a hard to cut off as we have been talking about so far. So point zero four, great, publish point zero six, forget about it. And this comes from one more mashing together, these two systems. And my favorite line on this is, surely God loves point zero five one, just as much as He loves point zero four nine. Right, but He does it in the name and Pearson system. I think that's the point, right? So, and this false cutoff is what makes P hacking so tempting. I've literally got emails from researchers saying, "Hey, if I throw out this outlier, then my P value becomes point zero four nine." So I'm good to go, right? And I have to say, "No, that's not the goal of your analysis. The goal is to figure out what the data are trying to tell you not to make your P value under point zero five. But unfortunately, sometimes that's become the goal. People lose sight of what science is really about." Maybe the trickiest problem with P values is how often we misinterpret what they actually mean. So many people think P value of point zero five means that there is a 5% chance that the null hypothesis is true, a 5% chance that there is no effect. And they think, "Okay, if that's true, that means that there is a 95% chance that there is an effect." Right. So like if I get a P value of point O five, they think it means that there's a 5% chance that my drug does not work. And therefore a 95% chance that my drug works. But this is absolutely wrong. And Regina, I think it's easy to see if we just go back to that ESP guessing game that we talked about earlier. Remember the P value there was 30%. So if you were misinterpreting P values in this way, you would say there was a 30% chance that I am not psychic. And therefore a 70% chance that I am psychic, which is clearly obviously wrong. Christian, I keep my mind open about you because you surprised me so often, but I am pretty sure that you are not psychic. Definitely not. Let me explain mathematically why that's wrong, Regina. Remember as we talked about, a P value tells us the probability of our result if we assume the null hypothesis. To calculate that P value, remember, we treated the null as if it's 100% true because it was our starting assumption. So we are not calculating the probability that the null is true or false. We're calculating the probability of our result or anything more extreme, assuming the null is true. Right. So just to recap, a 5% P value does not mean that there's a 5% chance that the results are a fluke or a 5% chance that they happened just by chance or a 5% chance that they are a false alarm. All of these are different ways of putting the same misinterpretation, the same wrong conclusion. Christian, this subtlety really trips up scientists. So I like to bring in an example that makes it more tangible. And Christian, you have seen me use this example in class and in lectures before. Oh yeah, when we co-teach you bring out this example, which I love and I think it either makes some people believe in magical creatures and other people hungry. The magical creatures being Paul, the psychic German octopus. I love Paul the octopus, but not everyone may be familiar with Paul the octopus. So tell us about Paul. Paul was hatched in 2008 and lived in an aquarium in Germany. And before each big football match, yes, soccer. Soccer, Regina. His handlers would give him two clear food containers. One had a German flag on it and the other had the flag of the opposing team in the upcoming match. And whichever container, Paul opened first. That was his prediction for who was going to win. And this was a well-designed experiment because lunch was served before the game. So this was actually a prospective study. These were real predictions that handlers did not already know the answer. They publicized his picks in advance even. So they couldn't go back and change them. No cheating and it was well publicized. So for the 2010 World Cup semi-final, Paul picked Spain over Germany. And German was so mad, they started publishing octopus recipes. Spain even offered him asylum if you wanted to add on a swim over there. But he turned out to be right, Spain won. Nobody can fault him. That is the amazing thing. So if you look at his entire two-year career, as a psychic, 2008 Euro 2010 World Cup, Paul predicted 12 out of 14 matches correctly. That's pretty amazing. What's the P value associated with this? The P value is roughly 0.01 or 1 in 100. Tistically significant for sure. All right. And then here's the trap we just talked about. Does that mean there's only a 1% chance that Paul's record was a fluke and a 99% chance that he's psychic? No, that clearly must be wrong. Again, the P value is not the probability that the null is true. It is not the probability that Paul is not psychic. And those double negatives, I do realize get confusing. And that's why people get confused here. They do. They get very confusing. And it is very human to flip things around and say, OK, if random noise would generate results, this surprising or more surprising, 1% of the time, then therefore that must mean there is a 1% chance that random noise generated this. And it's really subtle, but you just flipped it around and you can't. You can't do that. Exactly. The problem is humans want to know what's the probability that drug works and what's the probability Paul is psychic, but you cannot answer that in a frequentist framework. To answer that, you need to use an entirely different statistical framework called Bayesian statistics. Right. Traditional frequent statistics talk about long run frequencies, as you mentioned before, Kristen. Bayesian statistics are a whole different philosophy from the frequent statistics that we usually teach. Bayesian statistics are named for Reverend Thomas Bayes, who like doing math in his spare time. And Bayesian statistics lets you do something that frequent statistics does not allow you to do. Bayesian statistics allows you to make probability statements about underlying reality. That's what we want, as you just mentioned. But in order to do that, in the Bayesian framework, you combine your data with your prior belief or prior information about how plausible the hypothesis was in the first place. And that's the important part, the prior belief, prior information. Right. That's the key. And Regina, my prior belief about psychic octopuses is basically zero. Let's say, for example, before I met Paul, I believed that there was a one in a trillion chance that psychic octopuses exist. In a Bayesian world, the added data from Paul, we changed my belief in psychic octopuses, we increase it to one in a billion, but not to a 99% chance that Paul is psychic. Definitely not. And in Bayesian statistics, we can see mathematically that the lower your prior probability is, like you, Kristen, with the one in a, you know, zillion, the more extraordinary your evidence needs to be, the stronger the evidence, Kristen, to change your mind about psychic octopuses. And I like the way Carl Sagan put this. Extraordinary claims require extraordinary evidence. It makes sense. Exactly. And in Paul's case, 12 out of 14 just is not extraordinary enough. It's not. In my nature paper, I made a graphic to show how p-values don't really tell us about underlying reality in the way that many people think. A lot of people assume, as we've talked about, p-value 0.05 means there's a 95% chance the effect is real, but it doesn't. It's much lower. So what the graphics showed is that the impact of a p-value depends on how plausible the claim was to start with that prior probability. So for example, if your prior odds on something happening, Paul, be in psychic or your drug working, was 50/50, then you get a p-value 0.05. That is going to bump up your plausibility to about 71% chance that the effect is real. So 50% is 71%, not 95% chance. Let's say it was a long shot to begin with, not one in a zillion, but say 19 to 1 against. Then that same p-value of 0.05 only lifts your chances to about 11%. Gina, I love this graphic. I think it really opened people's eyes. And just a little foreshadowing to, when we talk about my paper, notice that misinterpreting p-values in this way tends to make your results look way more optimistic than they actually are. They do. But Regina, our listeners might be wondering, why aren't we all Bayesian? Why don't we use Bayesian statistics if they're so much easier to interpret? Excellent question, Christian. Really? Excellent question. I actually wrote a whole article for new scientists.
magazine about probability wars between frequentist and basins. And if you thought the Fisher name in Pearson feud was exciting, this blows that one out of the water, fast and they didn't stuff. So first of all, there are a lot of basins statisticians. They have entire conferences and university departments just devoted to basins approaches. But it's true for a long time, Bayesian methods were hard to do without modern computing. And also at the same time, some people worried about the subjectivity of using prior probability. It's like how you even spoke to do that. Fisher and name in Pearson did not like the squishingest of that at all. Now it's better. We do have nice free software that makes Bayesian analyses much easier. But most scientists still learn frequentist methods first at least. And of course, they stick with what they already know. So Bayesian statistics is growing. But yeah, frequentist thinking still dominates for now. So Regina, your paper got a lot of attention and then what happened? Yeah, two years after that, the American Statistical Association put together a working group to try to bring some clarity around p-values. And they cited my article as one of the prominent discussions that pushed them to do so. That's amazing, Regina. You're like changing statistical practice. I think it's surprising to people outside of the statistics world just how contentious this issue is and how much debate and discussion there is around this issue. I think people assume that statistics is kind of fixed, but it's not. And you were part of the working group, correct? I was invited to be the facilitator, which means I had a front receipt to all of that drama. I had to heard about 20 of the most prominent statisticians in thinking about p-values. And these were like Bayesian and frequentist all in the same room, including your colleagues, Steve Goodman from Stanford. I'm sure there was tabloid level drama at this meeting, Regina. Hi, hi, passions. I'm sure. Very high passions. I was shocked. I was prepared for drama. It was even worse than that. It was like hurting cats, but, you know, to the end of power. So we managed to nail down a definition. I can tell you how many hours of meetings e-mails drafts. Okay. Here is the definition of a p-value. Informally, a p-value is the probability under a specified statistical model that a statistical summary of the data, for example, the sample mean difference between two compared groups would be equal to or more extreme than its observed value. Wow. They stuck the qualifier informally at the front. Was that to make somebody happy? You do not want to see what the formal definition was. Like this is the most that we could agree. I kept trying to get them. Let's do some plain language. People. Can we? No, we cannot. That was as simple as you could get, huh? This was it. This was it. Okay, but we did have other points that we published in the white paper. You got people to agree on a few points? Yeah, but you know what? They're not shocking these things. Okay, so six key principles. And most of them, we're basically don't do not do this. Don't p-values as the probability. The knowledge true. We covered that. Don't make important decisions based solely on whether p is less than 0.5 or less than any threshold. Don't p-hack. You need to be transparent. Don't think a p-value tells you how big or important an effect is. And my favorite, by itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis. In other words, a lot of the pitfalls we've just talked through. Yes. Like it was it was kind of my paper. It was it was nice. So there was one positive statement that we agree on. Okay, are you ready? Yeah. A p-value can tell you how incompatible the data are with a given statistical model. Is that super exciting? It's like riveting, riveting, riveting. So I want to just wrap it all up because I had a nice kicker paragraph in that nature paper. And I closed it with three questions that Hopkins statistician had said that I think are really wise. He said that after a study is done, a scientist wants to know three things. One, what is the evidence? Two, what should I believe? And three, what should I do? And then I had a quote from Steve Goodman saying quote, one method cannot answer all these questions. The numbers are where the scientific discussions should start, not end. That is a perfect place to end Regina. The p-value isn't magic. It's just the starting place. All right, let's take a short break. And then we get to talk about your paper. Kristen, we've talked about your medical statistics program. Just a fabulous program available in Stanford online. Maybe you can tell listeners a little bit more about it. It's a three course sequence. If you really want that deeper dive into statistics, I teach data analysis in R or SAS, probability and statistical tests, including regression. You're going to find a link to this program on our website, normalcurs.com. Welcome back to normal curves. In this special episode, we are covering two of our very own papers. And we were about to take this wild ride with Kristen, which she talked about some statistical debunking and statistical sleuthing that she did. The Kristen started off. I got involved a few years back in helping debunk a statistical method that was used in sports science for years and in hundreds of papers. And at its core, this method was built around misinterpreting p-values in that common way we talked about earlier, as the probability that the null is true. Regina, a major consequence of misinterpreting p-values in this way is that it makes results look far more optimistic than they really are. Here, I believe that that resulted in flooding the literature with a bunch of false positives. Which is a huge deal. So, Kristen, you played a really terrific role in this statistical debunking, that's what it is. But you also ended up really sinking a lot of time into this whole enterprise. Way too much time. I didn't realize what I was getting myself into when I first happened upon this. And Regina, it's too bad I didn't get paid for all that time. I could have used that time for paid work. Yes, but scientific researchers, we are grateful to Fuyahudworks. Thank you, Fuyahur Service. Okay. So, tell us a little bit about this method. What was it? It's called magnitude-based inference or MBI. And it was cooked up by two sports scientists, Will Hopkins and Alan Batterham. And they pushed it hard. Supposedly, they would go around at conferences and pressure young researchers to use it. By the time I got involved, they were even selling lucrative courses to teach the method. And as we're going to see, it had a cult-like feel to it. Wow, these guys were pretty balsy, weren't they? Oh, yeah. And they built it as an alternative to significance testing and as a solution to some of the problems that we've talked about with P-values. Oh, that's how they stalled it to P-Village. That was one of their sales pitches, yes. Yeah. The whole thing was introduced on Hopkins's personal website, sportside.org, which he tries to pass off as a peer-reviewed journal. What do you mean by that? So, it says peer-reviewed journal at the top of the website. But the quote, peer review is just friends reviewing friends. So at the time I got involved, it was usually just Batterham quote, reviewing Hopkins. So in no way was a peer-reviewed journal. Again, did I mention balsy? And on their website, you could download Excel spreadsheets that implemented MBI. Now, there was no real documentation of what the spreadsheets were doing mathematically, no version control, but you could download these things, paste your data in, hit enter, and out would come a whole bunch of results. Oh, Excel is just usually never a good idea. I spend a lot of time telling people why they should never use Excel to manipulate or analyze data. It makes it really easy to introduce errors and really hard to trace those errors. Yeah, Excel is bad, but I can see why this whole MBI thing caught on. It came into these, "Ooh, E-T-D-U-S Excel spreadsheets, no coding needed." And the results looked extra optimistic. That's what you're saying, Kristen, right? Exactly. So what's the backstory? How did you get pulled into this? So in 2017, I happened to be on a statistics panel at a conference, and also on that panel was the journalist, Kristi Ashwondon. She was writing for 538.com at the time. Kristi, Kristi is great. We love Kristi. We've both worked with her. Exactly. She was working at the time on a book on sports science, and I analyzed data in sports medicine. So we got to talking, and then she asked me to take a look at this odd statistical method that she had come across in the sports science
Literature, MBI. I had never heard of it before, but she sent me some papers, and I set aside a few hours one morning to look through them. And when I started reading them, I could not believe what I was seeing. I actually ended up like taking my laptop around all day between my classes at Stanford, and I was furiously running code and typing emails to her. And we ended up having a long conversation the next day in which I explained some of the problems that I was seeing with MBI. I got off that call though, and I wasn't really sure if she was gonna write anything up for a layout dance about this because it's kind of technical and niche-y. So I felt like I should write up a little critique for the academic literature. I thought it would be a few days of my life. Little did I know when I was actually getting myself into. (laughing) - Famous last word on that one. So tell us what your paper was about then. - So I proved mathematically that MBI makes it easier to find false positives. In statistical terms, it inflates the type one error rates. - Right, and just to remind everyone, a type one error is a false positive. When you conclude there is a real effect, when in reality there is not. And this, too many false positives, is one of the big problems in significance testing. Let's get back to John E. N. Edis's essay on why most published scientific financial false positives. But Christian, you're telling us that there were two researchers here who were actually developing a way working hard to make this problem even worse, not better. - Exactly, yes. That was even part of their sales pitch. They wrote MBI has provided researchers with an avenue for publishing previously unpublishable effects from small samples. (laughing) - I love that. It makes it sound like the goal is just publishing. - Nope. - Not uncovering scientific truths, maybe. - Right. So what they did here was lower the statistical bar to make it easier to find things, right? - That's exactly what they did. And they were selling it as, look, we can find more things. But when critics came back and pointed out that if you can find more things that necessarily is going to increase false positives, they doubled down and tried to deny that. It's classing, have your cake and eat it too nonsense. - Yeah, absolutely. So other people had criticized this before you then. - Yes, I got involved kind of late. This had already been around since about 2006 and had already faced plenty of criticism. But none of those criticisms had really stuck because the developers of MBI always countered with smoke and mirrors. And a lot of people in sport science just didn't feel qualified to evaluate a statistical method. So the smoke and mirrors kind of worked. - Yeah. - You know, Regina, I think the developers of MBI had good intentions in creating this method. They were trying to make something easy for sport scientists to use. And they were trying to avoid some of the problems P-values cause. - It does seem like they were trying to be helpful at least. - Right. They did the opposite of some of the science Cinderella stories that we've talked about in previous episodes. - Right. In the red dress episode, for example, we talked about researchers who accepted critical feedback on their methods and publicly acknowledged the shortcomings of their previous work and actually improved. Right. They owned what they did wrong. They made changes and we loved and trusted them more because of it. MBI could have had a similar story. It could have, yes, all researchers need a bit of humility. Scientists are humans and humans make mistakes. And science is a self-correcting process, but only if people are willing to self-correct. So tell us a little bit more about what people previously had criticized. - The first critique to be published in a peer reviewed journal actually came out back in 2008 and some statisticians pointed out that MBI misinterprets P values and confidence intervals as Bayesian probabilities. So they suggested that instead of MBI, people should just do a proper Bayesian analysis. (laughing) - Common sense, but also a good burn. (laughing) - But how did the MBI guys respond to that? - They fired back with a letter titled an imaginary Bayesian monster. (laughing) - Bayesian monster, like a cookie monster. - It's cute, what is it? - Oh, I'm picturing like a fuzzy little cute monster, yes, Reggie. - Oh, oh, oh. - And Kudos to them for the creative title. - Right. - And I'll also give them Kudos because they managed in their letter to slip in a reference to the Hitchhiker's Guide to the Galaxy. Pretty good. Beyond that, however, it was complete nonsense. They tossed around some statistical jargon that sports scientists might not question, but it was complete statistical gibberish. - It was, I read it. And to me, the whole thing read like one big non-sequitor. (laughing) Like, this sentence did not really relate to that last sentence at all. - Yeah, for our statistician listeners, we do have a few. Their argument was basically, we ran a bootstrap version of a T-test. And because it gave us the same answer as the T-test, therefore we have done a Bayesian analysis. (laughing) - Exactly that. - Yes. - Okay, let me translate for everyone else in our audience. That is like saying, we are playing chess. And I just showed that a knight is allowed to make an L-shaped move. Therefore, I am now playing poker. (laughing) - Well, that's a good one. I was thinking, it reminded me of Enigo Montoya from Princess Bride. - Do you remember H. Famous Line? - You keep using that word Bayesian. I do not think it means what you think it means. (laughing) - Yes, love a good Hitchhiker's Guide reference and a good Princess Bride reference. - Another critique was published in 2015 by Alan Welsh and Emma Knight to Australian statisticians. And they actually went into the MBI spreadsheets and pain stakingly worked out what every Excel formula was doing and what they showed was that the spreadsheets were just running standard hypothesis tests, spitting out P-values and then misinterpreting those P-values. - Wow. So Alan and Emma are really doing God's work for us here to uncover this. - Yes, God bless them. I would not have had the patience to go through those spreadsheets, but I could not have written my paper without them mathematically laying out what the spreadsheets were actually doing. Regina, let's make this concrete. So say we ran a trial where runners drink cherry juice or a placebo drink for two months and then we measured their improvement in 5K times. One of the null hypotheses MBI would test is that cherry juice has no or only a trivial benefit on running times. - Right, remember the null hypothesis is that skeptics world. It's the opposite of what we're actually trying to prove. - Right, but here it's a bit different than what we talked about before because in this skeptics world, the true value isn't just zero, it's zero or something very tight. And Regina, let's imagine that the cherry juice group improved by 15 seconds on average while the placebo group did not improve at all. And then we take the data and we compare these groups using a standard hypothesis test, a T test. And we get a P value of 0.24 or 24%. - Right, now let's interpret what that actually means. If cherry juice really had no effect or only a trivial one and we repeated this experiment over and over, you'd expect to see results at least this extreme in favor of cherry juice about 24% of the time, which is not rare. It's about as surprising as Flopinacoin and getting two-headed in a row. - Right, Regina, I could do that live right now here on the podcast and get two heads in a row, no problem. It wouldn't be all that exciting or surprising. But here's the kicker. MBI interprets that 24% as meaning that there is a 24% chance that cherry juice does not improve running times beyond a trivial amount. And therefore that there's a 76% chance that cherry juice improves running times. - No, no, no, no, no, this is mind blowing. This is exactly that wrong definition we talked about earlier. It's misinterpreting P values as the probability that the null hypothesis is true or the probability is a chance finding. Remember, this cannot be the interpretation because if it were, that would mean there is a 99% chance that Paul the octopus is psychic. If you accept one, you have to accept the other, that's it. - And people may want that to be true, but I think most of us can agree that unfortunately Paul was not psychic. - No. - Because he's an octopus or was an octopus. (laughing) - Right, Regina, if you were to move into a Bayesian world and set a realistic prior, the probabilities that cherry juice works or that Paul the octopus is psychic based on the data we've discussed, those probabilities would be in fact much lower. Much lower. Regina MBI also set thresholds for declared
So, instead of calling them statistically significant, they just changed the word and called them substantial. But people treated substantial effects just like significant effects. The catch was that these substantial effects came with type 1 error rates far higher than that usual 5% that we said. Well, Shinnite actually showed this in their paper. And how did Hopkins and Batterham respond to this? I'm guessing with more smoke and mirrors. Or do they just say, you know what, you're right. No, smoke and mirrors, lovely, lovely smoke and mirrors, a whole magic show. They actually came back and wrote a paper in 2016 that got published. The story behind that is kind of funny because in her investigative journalism, Christie uncovered that that paper had been rejected at one journal. And then the reason it got accepted at the second journal is that Hopkins called potential reviewers up and kind of pushed them into accepting the paper. Oh, no. Oh, that is bad. Right. And in that paper, they claimed that MBI both reduced false positives, that type 1 error, but also at the same time reduced false negatives. And that is actually mathematically impossible. So what I wrote in my paper is at face value, this conclusion is dubious. And thus you do not need to be a statistician to be immediately skeptical of their paper. Ooh, is the thing. I know. My job, Chris. Don't mess with me. False positives and false negatives always trade off. Think of that dating app. If you crank up your filter to only match with the six foot supermod hold with the six figure salaries, you are going to cut down on bad dates. Sure. That's going to be fewer false positives, but you will also miss out on a lot of things. Perfectly nice people. That's the false negatives. That's the way it works. Exactly. So in my paper, I derived mathematical equations for the error rates of MBI. And I proved mathematically that MBI trades off false negatives and false positives just as expected in math. I also showed that it inflates type 1 error at exactly the small sample sizes that it recommends that people use. Regina, in essence, MBI was exploiting Winners' curse, which is something we talked about back in the Red Dress episode. That is fascinating. Winners' curse, a favorite of mine, where small samples can very easily give off these flashy results that turn out just to be chance flukes. Yes. And what they were doing was in these small samples, in particular, they were lowering the bar for distinguishing signal from noise, which made it even easier to turn those random blips into publishable stories. I bet these two were not very happy with your paper. Were they? They didn't love me now. What happened? What did they say? So, Regina, this is actually the story of how I got on Twitter in the first place. I had not been a Twitter user, but within hours of my paper going online, the MBI authors were attacking me on Twitter, claiming to have already disproven my analysis. So, I got on Twitter just so I could cut through their smoke and mirrors. Not long after this, the creators of MBI emailed me, and they invited me to have a debate about my paper that they were going to moderate on their website. Well, that's a generous of them. They're perfectly fair and objective. Nice guys. I know. I was like, thanks for new things. We had a little email exchange, and as we were talking on email, I suspected that they actually hadn't even read my paper. So, I asked a leading question to test them, and sure enough, they had zero clue what my paper was actually about. They then went on to publish our email exchange on their website without my permission, but of course, they conveniently left out the part where I exposed that they didn't understand my paper. Regina, they kind of picked the wrong person to mess with, though. Had they just kind of left it be, and I published my paper, and they left it alone, and didn't try to throw up smoke and mirrors, I probably would have left it alone too. The problem is I have no tolerance for bullshit. And I am also unusually meticulous and persistent. So, I decided, well, I am going to make a video to explain my paper, because I do a lot of lecturing and I'm a good lecturer, so I made a 50 minute YouTube video walking through my paper. That gave the paper some traction. And around the same time, Christy Eshwandan wrote about MBI and my paper in my video on 538.com, and that gave my criticism even wider exposure. Now, the MBI proponents, they came back and wrote a lengthy critique of both my paper and Christy's article, and they posted it on their website. Almost none of it addressed the substance of my paper, the type 1 error issue. Instead, they tried to attack me by calling me a quote, "establishment statistician." As if that was an insult, I was so flattered I put it on my Twitter bio, actually. Establishment statistician is a compliment. I know. Let me t-shirt. I am still waiting to get my establishment statistician card in the mail. I want an official badge and I would wear it with honor, because if there's ever any place you want to be establishment and not like a rebel, it's in math. I know what I'm getting you for your birthday, the spoiler alert. Regina, when we do merch for the podcast, we need establishment statistician badges that people can buy. I like the ring of that. Now it did turn out that MBI had a cult like following and there were a few devotees who in response to my paper wrote blog posts supporting MBI and some of them were really fun. So I want to read a few excerpts from one of my favorites. His first line is, "I am not a statistician," and kind of love the transparency there. Because I'm to say, "I trust the analytical foundations of MBI and my trust is based on the following. Allen and Will are amongst the most highly cited researchers in exercise in sport. Their knowledge of the inference literature is clearly beyond reproach and their logic is impeccable." Wow. They go on to say, "Realizing the importance of MBI's is much like Neo deciding to swallow the red pill in the movie The Matrix. When making the decision to learn and understand MBI's, one can never revert to conventional methods in less force two. You simply know better." I love the matrix reference though. Gotta say, and the red pill. But the ironic thing is that they were actually blue-pilled by this MBI stuff, right? So Neo gets offered the blue pill in the red pill. Blue pill is going back to your normal everyday life, and red pill is no. The scales fall from your eyes. They were taking a blue pill that was disguised as a red pill. Kristen, you and I are the actual ones handing out the red pills, trying to show researchers what evidence really looks like. Yeah, you're right Regina. That's kind of what we do in this podcast, right? Red pilling the research. Remember, our Stanford student this year, and they said, "Now that my eyes have been opened, I can never see research the same way again." And they were disturbed, but you and I just looked at each other like, "Yeah, mission accomplished." "Dye, what's yours done?" "Yeah, yeah." So what happened to MBI in the end then? So this went on for a while. Some journals did ban the method. Batter Ham, I think, saw the writing on the wall, and he jumped ship pretty early. I think Hopkins might still be out there pushing this. I haven't really followed it in the last few years. But while I was still involved, he rebranded MBI as MBD because of course, if you change the name, you escape all criticisms, right? And for a while, he was also telling people to just say that they had done a Bayesian analysis when they had used MBI, even though they really hadn't done a Bayesian analysis. That's got a lie. And ending the truth. So I ended up having to write more papers, and one of them was actually explaining why MBI is not a Bayesian analysis. Now to be fair, there is one special case where Bayesian probabilities and P-values line up. And this is when you assume as a prior that all possible effects are equally likely. So for cherry juice, that would mean it's just as likely that cherry juice can improve your 5K time by infinity seconds as by zero seconds, which is obviously observed, although Regina, I really, really want to improve my 5K time by infinity seconds. Yeah. Maybe somebody else has a way. But just because these two things line up in that one case, that does not mean that every P-value can be read as a Bayesian probability, obviously. Yeah. Now one fun thing that happened with Regina is after I wrote my first paper, people reached out to me, and I got to collaborate on the rest of the papers with some fun people. I even co-authored a paper with Andrew Vickery.
who is a statistician who's writing I have always admired. I want to just share one of the lines he put in our paper. He says, "Magnitude-based inference has certain superficial similarities with Bayesian statistics, but you cannot send a baseball team to England and describe them as cricket players just because they try to hit a ball with a bat. The incorrect claim that MBI is a Bayesian analysis is similarly just not cricket had does real damage." Ooh, I like it. So what would you say the whole takeaway is from this story? You can go really far off track if you miss entropy values Regina. That's true. That's kind of what it boils down to. I think this is the perfect illustration of why do we care so much about these pedantic, fussy little definitions because they really do matter. Chris and I think we are ready to wrap up our P values episode and rate the strength of evidence for our claim. And our claim today is P values are a flawed statistical tool. And we rate the strength of evidence for this claim using our five smooch scale. One smooch means little to no evidence in favor of the claim. And five means strong, very strong evidence in favor of the claim. So Christian, go first. What do you think? Regina, I'm going to go with two smooches on this one. I don't think it's inherently the P values fault. I picture the P value kind of like a Phillips head screwdriver. It is good for what it's supposed to be used for. So P values are great for helping us to distinguish signal from noise. And that's it. But that's the job it was set out to do and that's the job it does well. I think the problem is really with humans. Humans are flawed users of statistical tools might be the better claim here. They try to use the Phillips screwdriver where they were supposed to have used a hammer. I think that ultimately is the issue. So I don't think the tool itself is really the problem. How about you Regina? I explain it like the stupid door. You don't know you're supposed to be pushing them or pulling them. You know, if you want to get out. Yeah. And I feel like the P value is like this. So they have to put a little sign saying push or pull. And there's this cognitive psychologist who says if you have to put a label on your door then you're designing your door poorly. And I feel like the P value is just like that. It looks like it's the probability of the null and we keep interpreting it that way because humans are going to human. That's what we do. So Kristen, I'm going to go with 2.5 smooches. Like a like an air kiss. I'm kind of right in the middle. I think the problem is with humans, but it's hard to blame humans because humans are humans and we're fallow all. It's hard to blame the P values because they were just designed this way. It's when you get them together. It's the ergonomics of statistics. That's the problem. 2.5 smooches. All right, Kristen. I think it's time for some methodological morals for my methodological moral. I think I'm going to go with if P values tell us the probability the null is true, then the octopuses are psychic. That's a great moral Regina. I love it. What do you have? Yeah, my methodologic moral for today Regina is statistical tools don't fool us blind faith in them does. I love this Kristen nice job. This has been really interesting episode Regina. A little different than our typical episode, but I'm hoping that listeners are going to come away really having a much clearer understanding what a P value actually is. I think so, but you know in some ways Kristen we have only scratched the surface of P values. We have left out so many subtleties and tiny little detail. Just giving a nod to all of those P value officianados out there. We realize this is just the tip of the iceberg, but this has been delightful Kristen. Thank you so much. Yeah, people get surprisingly emotional about P values, so be careful you're in for a word. And we will put more papers about P values in our show notes if you want deeper reading. Thanks Regina. Thanks Kristen and thanks everyone for listening.
Podcast Summary
Key Points:
P-values are a statistical tool used to distinguish real signals from random noise in data, but they are often misunderstood and misused.
The null hypothesis represents the "boring world" of no effect; a P-value measures how surprising the observed result would be if the null hypothesis were true.
The 0.05 threshold for significance was introduced by R.A. Fisher as a flexible guideline, not a rigid rule, but it has become a universal cutoff.
The Fisher approach treats P-values as continuous evidence, while the Neyman-Pearson framework uses fixed thresholds for binary decisions (type I and type II errors).
Modern practice often combines these two philosophies incorrectly, leading to overreliance on P<0.05 and misinterpretation of results.
Regina’s 2014 Nature article highlighted the slippery nature of P-values using real-world examples, such as a failed replication study, and won an award from the American Statistical Association.
Summary:
This episode of Normal Curbs explores the controversial role of P-values in scientific research. P-values, which stand for probability values, help researchers separate signals from noise by measuring how surprising a result would be if the null hypothesis (no effect) were true. 0001 suggests strong evidence against the null.
A. 05 threshold as a flexible guideline for deciding when to take a closer look at data. 05 as a universal rule.
59). The episode emphasizes that P-values are slippery and should not be treated as magical proof of a discovery. 05 as a rigid cutoff—can lead to false positives and missed opportunities.
The hosts use analogies like dating apps to explain type I (false positive) and type II (false negative) errors, highlighting the need for careful interpretation and context in statistical analysis.
FAQs
A P-value is the probability of seeing a result as extreme as the one observed, or more so, if the null hypothesis (no effect) is true. It measures how surprising the data is under the assumption of no effect.
It's calculated by imagining repeating the experiment many times under the null hypothesis and seeing how often the observed result occurs. For example, in a psychic guessing game, a computer simulated 1,000 games to find the probability of success by chance.
P-values can be misinterpreted as the probability that the null hypothesis is true, but they only measure surprise under that hypothesis. They are also often used as rigid cutoffs (like 0.05), which ignores their intended role as flexible evidence.
The 0.05 threshold was proposed by R.A. Fisher in the 1920s as a convenient guideline for deciding if results deserve a second look. It was never meant to be a universal rule, but it became enshrined through a mashup of Fisher's and Neyman-Pearson's systems.
Fisher saw P-values as a continuum of evidence to guide further investigation, while Neyman-Pearson advocated for preset decision rules (e.g., alpha = 0.05) to make binary yes-no conclusions about effects.
Type I error (false positive) is concluding an effect exists when it doesn't, like swiping right on a bad date. Type II error (false negative) is missing a real effect, like swiping left on a soulmate.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.