#604: How To Interpret Nutrition Research – David Allison, PhD
52m 16s
In this episode of Sigma Nutrition Radio, host Danny Lennon interviews Dr. David Allison, a leading expert in nutrition, obesity, and research methodology. They discuss how to improve rigor in nutrition science by identifying and avoiding common errors across all research stages. Allison explains that rigor is not absolute but should be proportional to a study's importance, with transparency being key. He highlights several pervasive problems: illogical research questions, invalid measurements (especially in dietary assessment), and statistical mistakes like the "DINS error," where researchers falsely claim a treatment effect by comparing p-values within groups (e.g., treatment significant, control not) instead of directly testing the difference between groups. This error remains common due to poor statistical understanding among investigators. Allison also notes misinterpretations of treatment response heterogeneity, where variability in outcomes is mistaken for variable treatment effects. Finally, he critiques dissemination practices, where authors, press releases, and media often spin null or modest results into exaggerated claims, undermining scientific communication. The conversation underscores the need for better training, critical appraisal, and honest reporting to advance nutrition science.
Hello and welcome to Sigma Nutrition Radio. This is episode 604 of the podcast. My name is Danny Lennon. You are very welcome to the show. Today we're gonna be getting into one of my favorite things to talk about personally, and that is some of the aspects about interpreting nutrition science, some aspects related to research, critical appraisal of that, and how the field can improve going forward. So we're gonna get into a lot of topics that might seem dry at the surface, but I think I really at the center of how we can interpret studies, and understand if a particular study gives us information that is useful, or actually tells us what it claims that it does, or can answer a particular research question that we want. And so these are concepts that are fundamental, but are very often overlooked or poorly understood. And so to go through some of these concepts, I'm gonna be talking with Dr. David Allison, who is Chief of Nutrition and Director of the USDA's Children's Nutrition Research Center at Baylor College of Medicine, where he's done research in nutrition, obesity, re-producibility of scientific evidence, and he is one of the people who I really, really value his perspectives on evaluating research and doing good quality science. And his work has been especially recognized in some of the areas we're gonna be talking about today, related to statistical reasoning, research methodology, improving transparency and trustworthiness in nutrition and health sciences. And as you will see during our conversation, he is someone that speaks with real precision and accuracy about some of the most fundamental things within the field. And I think if we're able to take a few of the ideas and concepts that he discussed today, it will massively improve your ability to critically apply research and understand some of the errors that go on. This is one of the episodes where having maybe a couple of listens will be really, really useful. Some of these concepts are certainly ones you wanna maybe investigate a bit deeper afterwards and practice and apply with. Also, for those of you who are sigma nutrition premium subscribers, you will of course get detailed study notes to accompany this episode, which I think will be particularly useful here, as well as an edited transcript and the key ideas segment after the interview. For those of you who are on the public feed of the podcast and might want to take a look at our premium subscription, which give you these extra educational tools to retain more information from your listening and to really use it as a learning tool, then that will be linked up in the description box where you're listening right now. Check that out, see if it's for you. It's the direct way to support the podcast, so I very much appreciate anyone who does that and that will all be linked up there for you or over on sigmanutrition.com. Also there linked up in the description box will be the episode page, which includes any resources we might mention throughout the conversation. So that is it. Please enjoy this conversation between myself and Dr. David Allison. (upbeat music) - Very big welcome to the podcast to Dr. David Allison. Thank you so much for taking the time to join me today. - Well, thank you, Danny. It's truly a pleasure to be here with you. - I'm really looking forward to this, I think as I mentioned to you previously in some of our communications, a really like, not only your thoughts and your insights that are based in your experiences, but how you address some of the, I think most pertinent issues within nutrition science. And that's what I really wanna get into here today. How can going forward as a field, nutrition science do better? How can we get better quality answers to the questions we care about and have some degree of rigor in that? But before I get to my questions, maybe for people listening, can you very briefly give them an introduction into your work, your academic background, your interests, anything else that might be relevant to what we're gonna get into today? - Sure, I'll be glad to. My original training is as a psychologist, but I've always been a scientist at heart. And by that, I mean, I'm sending likes to look at things in the world, wonder about them, wonder's a wonderful thing. Think about how do they get there, what do they do, how does this work? Is what I'm seeing, what's really there, and so on, and then start to figure out ways to try to answer those questions, whether it's by asking somebody else or collecting some data. And even as a kid, I did that. And often when I were asking someone else and getting the answer, I would follow up with a question that not every loves it when a kid asks, which is, how do you know? And are you sure? But those are the hallmarks of a scientist. And so I've sort of been a scientist for as long as I can remember. People often ask me, how did you choose to become a scientist? And I say, it was no choice in the matter. It's just what I am. And then I've sort of evolved to be a statistician because I'm a skeptic, and I don't like taking anybody's word for anything. And so I felt I needed to understand statistics. That's why I kept studying statistics until people seemed to think I was a statistician. And then I stopped arguing and said, okay, if you think I'm a statistician, I am. So the statistician, the psychologist, and obesity, nutrition, energetic, aging researcher, some had focused on rigor, reproducibility, and transparency and trustworthiness in science. And I have the privilege of leading the USDA Changer's Nutrition Research Center at Baylor College of Medicine in Texas Children's Hospital. - We're obviously gonna dive into a number of those elements you related to, really a more global sense of epistemology and how we apply that to nutrition, not only what we know, but more importantly, how did we come to know that? And so before we get into maybe their specifics, let me lead off with a perhaps overly broad question, but feel free to go in which direction you wish. And that's given that we're thinking about in the spirit of scientific inquiry, we want some degree of rigor in the field. And I think this is where many valid points are made around some of the problems that nutrition as a field faces. When we want to have a rigorous scientific inquiry in whatever field, specifically for us, that's nutrition, what should that mean? What things are we actually looking for? And how do you conceptualize that for people? - That's a very challenging question and it's when I struggle with regularly and think about often. It was a wonderful book by a tool Gwande called the Checklist Manifesto. And I'm very enamored of it. It talks about using checklists in things like surgery to reduce the number of mistakes and it seems to be very effective. And I've tried to think, well, what's the checklist for a research project? And I've struggled with it because research projects vary so much that it's hard to think of a single checklist for the astronomer and the cell biologist and the human clinical trialist and so on. And even within those domains, it's hard to think about them all. But I think if we think about the scientific process very broadly, it seems to have a few major steps. And one is conceiving of a question here, identifying a topic in which we want to fit, find something out. And then structuring that question to be precise and answerable and asking oneself and checking, is it really answerable? Not all questions that can be dramatically posed really have meaning. And I think when we get into some things like the so-called debate of carbohydrate-insium model versus energy-balanced model, I think it may be who of us to ask, is there actually a question on the table that makes any logical sense? That's the first step. And we can have some checklist there. The second one is, OK, now design the study. What are you going to do to collect some information that bears on that question? The next step is to execute the pranacol. And you've got to execute it well. You've got to record your measurements properly and so forth and faithfully. The next step is you've got to analyze the data. And the next step is you've got to interpret the analyses. And then finally, you've got to communicate. What you've found. So those are the steps involved. And at each point, there are things that can go right. And things that can go wrong. And things that we better or worse. Every single one of those points we need to ask, are we doing this as fairly as possible? Are we checking for common mistakes? Are we documenting what we've done? Are we being transparent in what we've done? My only belief is that rigor is compromisable. We always want the most rigor we can get. But not every study deserves equal rigor. Some studies are very important. If it's, for example, a registry study of a clinical child for a drug or a safety and a very serious matters our hand, we'd probably want that to be the greatest figure we can achieve. On the other hand, sometimes it's a very quick test of some observational question is A associated with B. That's actually not really of great importance. Or it's very preliminary. And we might have less rigor. We might use self-reported heights and weights, for example. And I think there's nothing wrong with that. The key is to disclose it so that if the reader knows that we did these things that were not so rigorous, then they can make their own judgment about how much confidence to put in that study. So that's, I think, a general approach. A lot of us can have a focus on how maybe a study was executed when it comes to interpreting that study. And we look at that.
Whereas really as you've alluded to there, we need to look at this at various stages all along and there's probably no way we can put specific Numbers on this but as a general point, how would you think about how prevalent or not? We actually see in the field of nutrition where the problem with a particular study or set of studies might not be in the Execution in of itself, but rather that either the design or the analysis that was used or whatever the case may be Was just incapable of answering the question we wanted to answer or idea that should be able to if that question makes any sense at all Well, your question certainly makes sense. I have no numerical information on this I can only sort of give you a Gistal gut feeling on it for what that's worth when we have the opportunity when I have the opportunity To get a view into those things for any particular study. I see Not in every paper or every study But in most I see some things that could be done have been done differently or and perhaps better in every phase Now in some cases those differences were the things that I think might have been done better or differently are relatively minor or modest and I would still describe the study as Having some validity that is not being Completely invalidated by the error and by an invalidating error my group we use a term to mean an error which if corrected could Either change the results in conclusion or Would change the results in conclusion so sometimes we don't know their would but we know that it could and in either case We consider those invalidating errors even if it didn't change the results and conclusions We're still considered invalidating error. It could have hard to know in all cases But when we get a peek into it we often see it we see it at the level of the logic and I've mentioned one example Just now we could talk about other things with ultra-pastor schemes or what? Happy with the very logic of what's being asked is open to doubt we see it in the premises Where the premises are seen are not valid? We see it in the methods where the measurements may not be valid We see a great deal of that with food intake measurements We also see a great deal up or not a great deal But we see it in other areas. There was a very controversial interesting case recently You probably know about involving so-called hyper responders and Whether hyper responders to a ketogenic diet who have high LDL cholesterol are actually at risk from that high LDL cholesterol for coronary artery plaques There was a challenge there with the measurements or the Image analysis of the CAC measurements and so that's a measurement for any problem as we go into analyses We see this very often we see what's called a dins error, which is difference in nominal significance That's a term my group is introduced in which one looks at the treatment group for example in a randomized control drug Studies just want to down and says ah the treatment group changed and P is less than 0.05 was just safe discussions point 0.049 The control group didn't change did a lose weight or lower their cholesterol or whatever and the p value was 0.051 So it was significant in the treatment group and not significant in the control group Therefore I declare that there was in fact and that's statistical nonsense You can easily mathematically prove that that's a grossly incorrect procedure But that's still sometimes used we also see with cluster randomized trials which happen to be very popular in childhood obesity community school and other intervention programs Very very frequently those are misanalyzed to the point of being completely invalid and there's a unfortunately a great deal to Splain wrong literature that purporting to show certain things are efficacious or effective for childhood obesity when the studies I'm not shamed back and then we get to interpretation Is a great deal of confusion there are particularly in areas about treatment response heterogeneity It links back to the design and sometimes analysis that is is a great deal discussion by clinical investigators Observe treatment response heterogeneity and what they're really observing is variability among people in outcomes like weight loss as an example But not in response and they can flate outcome with response If outcomes were responses, we wouldn't need control groups and studies the very fact in these control groups Recognizes that the response to treatment is not the same as the outcome achieved But people seem to forget that when they go into treatment response heterogeneity and they say look some people lost the ladder weight Some people lost the little weight some people even gain weight Well great treatment response heterogeneity in which case we have to say no There's no evidence for that at all yet There might be treatment response heterogeneity that that's out evidence for it So you have an analytic and interpretive problem Many people misanalyze and misdesign their studies there and misinterpret them And then in terms of dissemination that's where it really gets squirrely It is not at all uncommon To see in the paper in the results section of a paper by an aesthetic Set of investigators a very clear statement of something like On our primary endpoint There was no statistically significant difference between the treatment and control And then the abstract contains some other statement in the conclusion like This is a promising treatment with some apparent effectiveness in women. Let's just say as an example Because later as a post-tock analysis they looked at women and men separately and found a seem to be affected in women But not men and it's not mentioned in the abstract. I was post-tock Or it's not mentioned that it wasn't on the primary outcome It was on a secondary or wasn't on all subjects It was only on subjects after they eliminated some subset and then It's less tempered in the abstract Then you go to the press release from the university and you get more spin And then you go to the interview That's published in the newspaper. I have a online news article With the investigator who suddenly has forgotten everything Here she said earlier about it wasn't the primary outcome or in the observational association study That it was an association not necessarily causation And suddenly you hear words about causation Impact, large effects, outrage, miracle cure etc So these are all the problems we see That answer is so rich with a number of issues that I think Understanding that these issues arise being able to identify them Would eliminate so much of the confusion that we have around nutrition even for people within the field and within academia And so I'd love us to start working through some of those that you've raised in a bit more detail Just to make sure that everyone is clear on examples of where these come up But exactly what is happening here and why it is so pertinent If we start first with that dins error that's the difference in nominal significance And in a simplistic form we're talking about situations where someone can see that in one group We see a significant finding from pre to post tests So at the end versus the beginning and then the other group We do have pre versus post and we see that there's a no significant finding The issue then comes people start trying to make conclusions relative to the difference between groups based on this As you've noted this is a significant error that can lead us into huge problems Can you maybe just Restate that again and then maybe to put a bit of color on this Why does this mistake actually remain so common because one thing is that it happens when people Are interpreting studies you see it all the time on social media people look at a randomized control trial And are making conclusions of this type based on a significance Finding in one group from the start to the end of the trial and maybe not in the other However, this as you noted is something that is still happening within the studies and being reported by the study authors themselves Why is this such a problem? And then why do you think this remains to be so common? If it gets a problem for two reasons One is you might say is innocent and the other is not so innocent The innocent problem is that Many non-sgotistition investigators Don't understand statistics very well the most kind of type of statistical inference we use And inference meaning you're sort of making it not just describing what you see in the data In the sample data you have in your hand But you're making an inference to what's the case in the population from which you've obtained that sample That's what you really want to know about And so the process of inference One school is called the frequentest school And it is by far the most common one used in the health sciences And it involves the p-values and the confidence intervals and things like that that you're used to say The logic of frequentist statistics I believe Pershing is sound some people don't But it's not intuitive Most people want to know about the probability As the hypothesis they're testing for instance
example being tree. So for example, if I say to you, well, what's your hypothesis here, Danny? And you said, well, my hypothesis is that the treatment group would lose more weight than the control or better yet, better said treatment causes weight loss. And I say, good, that's terrific, Danny. We just did the analysis for you. And the p value was less than point 05. Let's say it's point 04. You said, oh, great, David. So that means there's only a 4% chance that the treatment doesn't cause an effect. And there's a 96% chance that it does. You say, no, no, Danny, sorry, you got it backwards. You've made what's called the prosecutor's fallacy. You're interpreting the probability of a given b as the probability of b given a. You're interpreting my number I gave you that p value as the probability of the hypothesis is true or false as opposed to the probability of obtaining the data you observed or data more extreme departure in terms of its departure from what would be expected under the null hypothesis. If the null hypothesis were true, and then you might say, David, what the heck did you just say? That is unmouthful. I said, Danny, it's really we observed the probability of the data given if there was no effect. That's what point of 4 leads, not the probability of the original effect. And you said, well, David, I don't care about the probability of the data, the data or in my hand, I know what they are. I care about the probability that my hypothesis is true or not. And I said, well, we don't answer that, Danny, sorry, we give you the answer to the question. We want to answer. It does bear on it. If we say there is some chance that your hypothesis could be true, it's conceivable. And the probability, if there is no effect of obtaining data that look like this is very low, it strengthens our belief legitimately, in my opinion, that there is an effect, but it doesn't tell you what perhaps you really want to know. And to get to that one has to adopt a different framework, usually called Bayesian. And that's uncommon, a less common approach. And even, and that has issues too, these things are very difficult concepts for some. They're not at all intuitive. So I think that's a big part of it. Probability is just not intuitive. If you know, the best example of all is the famous Monti Hall problem. And I don't, in the interest of time, perhaps, I mean, I can go through it if you want, but I suspect you don't want me to go through it, but the listener can look it up. It's a fun problem. And what's also fun is the social history of how many professional, doctoral level mathematics and statistics professionals get it wrong until it takes them a long time to conceive it. So probability is very counterintuitive. And a lot of people make just innocent mistakes. And I think a big problem we have in health sciences is we don't have enough professional statisticians involved. Too many investigators either believe they and their research team have the statistical skills to do the analysis and interpret it because somebody took one master's level course 20 years ago in statistics and knows how to turn the computer on it. Or just doesn't have access to a statistician because they're not enough available or they don't have funding for it. That's a big problem. You know, we, we tell a joke sometimes in our group that's not original to us, but someone from a surgery to bother surging calls up the statistician professor and says, I'm going to do my own statistical analysis on this trial I did. So I don't really need your help. I'm just hoping you can recommend a good statistics textbooks for me. And the statistics professor replies, oh, that's wonderful. I'm glad you called because I also wanted to do some brain surgery, but I'd like to just do it myself. Can you recommend a textbook for me? And the point is, you know, we wouldn't trust the statistician to do that. You know, I know what the new leaves principle is and I understand how a combustible engine works, but do you want me to personally be the one checking the 747 plane before you get on it? Probably, probably not a professional mechanical engineering expert. So that's a really big problem. But if everybody had a professional statistician analyzing every one of their studies, we wouldn't possibly under current circumstances have enough statisticians available. So we really need to think this through and figure out how to work it out. I'm not sure I know a solution. The other cause is not something we're missing. Many people want to publish a statistically significant result either because they think it makes the paper more interesting where they really believe the hypothesis. They just don't want to tell you that I still might believe my hypothesis that this dietary supplement does an example causes lower cholesterol or weight loss or more happiness or whatever you think it causes. And you're entitled to your belief, you would leave anything you want, but they want to be able to say my evidence supports it. And maybe they don't get the statistically significant result on the prep analysis, but they do see if they do this so-called dins error, then they can claim it's affected. And I think that's nefarious or malfeasance. And when you see a research plan published in a protocol paper, clinical trials registry at Prairie, and it says we're going to do this prep analysis, but then the impropriet analysis is reported. You kind of have a good sense that this was not so innocent. And some of those nefarious drivers, I certainly want to return to later. Some of them, as you mentioned, owing to the incentives at play that are related to maybe pressures within academia or just our current peer review publishing system that we have amongst other things. And I also want to return to this central importance of statistics to be able to not only interpret studies properly, but make sure that we're doing proper science. So maybe let me work through a couple of the things that you said. And given that we are, including myself, not statisticians, having a grasp of certain key concepts related to statistics can maybe head off some of these potential problems, or at least that we can maybe make some note of them when someone is erroneously interpreting a study, either intentionally or not, or spotting problems in actual papers. As you've mentioned, there's maybe confusion around what probability actually is, how to interpret p-values, what we're actually testing, certainly at least from a frequent this perspective, what we're actually testing, we're re-asking the question, if the null hypothesis were true, what is the probability of observing these particular data or data more extreme, as opposed to what maybe people might typically tend to think around that we're testing the alternative hypothesis. So there's a real key importance of understanding some of these concepts. One that I wanted to return to that you mentioned a bit earlier relates to the heterogeneity in response. And this is in particular, I see it come up quite a lot where you sometimes have actually a researcher promoting the work maybe on social media, or you have other people talking about the study, and they might hold up trial data, and they will misinterpret some of these findings to say, well, look, we have these two different interventions, or we have these two different groups, and based on this response, and they'll show all the individual data points from that study, they start labeling those people as responders or non-responders, and then going to make a further claim that, well, this means that for some of these people, this diet is better, and the other one is worse, and for other people it's the converse, and they're making that claim based on this one particular trial, which is not, it's not the question it's set up to answer. So could you maybe just, from your perspective, again, explain some of this confusion people have when it comes to this heterogeneity, how maybe they're taking trial data, and making claims about individualization of response, or making conclusions that aren't actually what this data is able to provide us with. This is one that particularly irksome to me, because so many methodologists have written it out this, and yet it just doesn't seem to get to most investigators who are not statisticians. There seems to be this presumption that GIC different people lose different amounts of weights, or have different outcomes in terms of happiness, or sexuality, or muscle growth, or again, whatever somebody's studying, and therefore this great heterogeneity response. I've even heard it in FDA advisory board meetings, I've heard the FDA look to the sponsor and say, can you tell us more about these non-responders? And sometimes want to jump up and scream, there's no evidence that there are any non-responders. You're just seeing variability in outcomes. We don't know what that means. And the reason is, let's just take an example, and let's take weight loss. Let's suppose that you and I are both in a clinical trial, and we're both in the treatment group. And on average, the treatment group loses 15 kilos, and the control group loses five kilos. So now we have very good, assuming everything else is good with this study, we have very good justification for saying the treatment causes a 10 kilo weight loss on average. Estimated the average effect. We had to subtract a placebo weight loss out from the treatment, mean weight loss. So now we know the mean effect, or at least we have a good estimate. And let's suppose, though, that in that trial, in which case, as I said, the average weight loss in the treatment group, [BLANK_AUDIO]
15 kilos. Let's suppose that I only lost 5 kilos. You lost 20 kilos. It's tempting to say you had a very good response, a better than your higher weight loss response than average, and I had a lesser weight. I'm a poor responder, a weak responder, or whatever you want to call it, and you're a high or a denser, good responder. But here's the problem. If you and I were in the control group, we also might have lost different amounts of weight. Maybe in the control group had you and I have been in it, I would have in fact gained 5 kilos and you would have lost 5 kilos. So there were factors affecting our weight loss other than I respond to the treatment. Maybe during that interval you got a horrible case of a flip and it caused you to lose an extra 5 kilos and you hadn't yet fully recovered. So we see the 20 kilos rather than 50 kilo weight loss for you. But really only 15 or due to the drug, or I should say 10 because we have to subtract out the placebo, only 10 were due to the drug, and the remainder was due to the fact that you got this absolutely horrible case of the flip. In contrast, let's suppose that I moved next door or to an apartment that was right above a donut shop. And I love donuts and every day I smell the donuts when I walk in their building and I can't resist them and I start eating a few donuts every day. And that's what counteracts some of the weight loss from the drug. That would have happened in the placebo group too. And so it's not different to response. It's just differential outcomes because of the swings and arrows of outrageous fortune to quirk Shakespeare. And that needs to be taken into account in the analysis. You can do that simply if you have observed things that you think moderated. So if you think for example age or sex or starting BMI or geographic location or whether you have low insulin baseline levels or high insulin baseline levels or particular genotype, you can, if you measure those things, you can include them in your statistical model. And what you need to test for is an interaction between that and the treatment assignment. What too many clinical investigators do is if they do that at all, they just look at the treatment group or they don't even do a control trial, they just treat some people. And then they say I can predict outcome with whether you're a man or a woman or this geographic area or not or have this baseline insulin level or not. And they mistake predicting the outcome for predicting the response. What you need to do is have the control. And then you need to look at an interaction between those factors and treatment versus control assignment as a variable. And if you get that statistical interaction and you've done everything else right, then you can claim that there's some treatment response heterogeneity with respect to those tree randomization variables. If you want to look at all of the treatment response heterogeneity variants, which includes the things you know about and could measure and the things you didn't know about are didn't measure or didn't include in your statistical model, then you need a completely different approach. And often probably the only or the most valid one is something that involves a crashing over design, but not a conventional crossover design. The ordinary two by two treatments by two period crossover design won't do it. What you need is multi period crossover design with multiple sequences and you need to have for each treatment at least true periods in which the treatment is applied. And very people very rarely do that. We have found a few studies that have been set up like that that we can analyze that way, even though the original investigators didn't plan it for that, but they use what's called factorial designs. And we're able to extract that analysis from the factorial design. And we have a paper that's in review on that in which one of my postdocs Dr. Ready is the lead author. And we find that in some things we get statistically significant evidence of treatment response heterogeneity. And in some we don't say you know just looks as far as we can tell there's no strong evidence that people or mice respond differently to this treatment. On that I do have a question maybe I'll return to a bit later related to the use of crossover studies in nutrition. But one thing before I forget that I wanted to pull back on you mentioned a bit earlier Dr. Allison was in relation to the sometimes nefarious, sometimes maybe just bad practice that can go on that a bit of statistical understanding can help people realize why it's such an issue. And that relates to the use particularly of secondary outcomes. And so oftentimes I think everyone is relatively familiar that we have this primary outcome that is the main subject of what we're trying to answer within a particular study. But there will be certain secondary outcomes listed as well. And sometimes you see in studies where there's a whole range of secondary outcomes, a big long list of things. And I think for someone who maybe is coming in without a background in this they might say well what is the problem with that if we have a study and someone just is going to go and measure everything they possibly can because we have access to these participants right now. And then after the fact we can see kind of what comes up. Obviously there's a significant flaw in that. But I think there's also a bit for those who are very familiar with that and understand the maybe the limits of looking at maybe secondary outcomes compared to a primary outcome that maybe don't take the time to actually think about the implications from a statistical standpoint. That is if we think of the more secondary outcomes we can list and we just keep testing more and more things and we then we think of the potential for false positives, false negatives within our data. At huge numbers of things that we're testing, statistically we can pretty easily demonstrate why this becomes a problem. So with that rambulocyde can you maybe mention for maybe people who are not in the field. Why is it so problematic if someone is just going to measure a whole bunch of different outcomes that they're not our primary outcome. And then after the fact just see what ends up being a significant outcome. And as opposed to the real thrust of the question is if we do find significance within that particular secondary outcome let's say why does that not necessarily guarantee that say trustworthy piece of data let's say. Okay. Really good question. And it's a very challenging question because it brings up these philosophical issues as well as practical issues. But practical issues are socially difficult but not intellectually difficult. And to me the big practical issue is just disclosure. There's a wonderful book about statistics that one could probably read in two hours on a plane contains almost no equations by Robert Abelson and the title of the book is statistics as principled arguments. And at one point in the book I'm paraphrasing here. Abelson says something like you know students and colleagues often come to me and they say professor can I do this in my design or analysis. And he responds you can do anything you want as long as you are prepared to accept a epistemologic limits that that choice makes upon you and I agree with him. But I would add further you can do anything you want if you accept the limits and here's the important end and you disclose it to the reader. That's the issue of transparency. So you could do any good, bad, stupid intelligence analysis you want in my opinion but just tell the reader what you did and then the reader can make their own judgment. That's the key practical thing and if people don't do that if you analyze a thousand end points and you just publish the one that came out significant or you analyze your data a thousand ways and just publish the ones that came out significant but you don't disclose that then we can't properly interpret those statistics you give. Now let's jump back to the much harder issue which is the philosophical issue. Let's suppose that what we're actually looking at is the effects of let's say some elements of self-reported dietary intake on some outcomes of interest and we're looking at a big data set and there are many many things we might call outcomes or things like body composition measurements and metabolism measurements and health outcomes and many of those can be scaled different ways. So we've probably got thousands of outcomes we can find that a big data set like the N-hains for example and we've got thousands of exposures of how much broccoli you get you lead to how much cauliflower or how much what's the ratio of broccoli to cauliflower consumption. So close to infinite number of things I can put in if I allow myself all these. If I do all of the analyses, thousands upon thousands of easily getting to the millions right for a thousand outcomes and a thousand exposures right I'm up to a million now I'm sorry a billion these lead to the idea that you can just find any random launch.
sense and something will be statistically significant eventually by chance and we shouldn't have very much confidence in that. And most people intuitively say, if I'm going to do something like that, I need to apply some correction. The most one on more and more is called the BANFAR only correction. And you basically, I won't go through what the BANFAR and correction is, but you apply a much more stringent test, right? So instead of saying the p value of less than 0.05 is good enough, I'm going to take 0.05 and divide it by the number of tests we did. And if that was a million tests, then it's point has to be a significance level of or p value of less than 0.05 divided by a million before I get to declare it's statistically significant. Now, most people intuitively say, yeah, that kind of makes sense. Otherwise, it just looks like here, as they so called the Texas Shark Shooter problem, which is just shoot everything all over the side of the barn and then go find where there's a clustering bullet holes and draw a target around it and say, look, I hit the bulls eye. Most of us feel that's who doesn't make sense and not a lot of argument about the need for some with the values and multiple testing correction here. But now let's take a different situation. While it's supposed that you and I are both interested in the effect of avocado consumption on mood, we both do a study of 100 subjects, 50 randomly assigned to eat an avocado for lunch and 50 assigned to eat something else. And then we measure mood using the exact same scale and we get the exact same result. And that is the p-value for the estimated effect of avocado on mood is 0.049. Now, interestingly, in addition to the measure of mood that we both used, which was, let's just say, some self-reporting, suppose you are ambitious and you said, wow, I don't know if people are always going to report, honestly, or accurately, I'd like another look at this. So I'm going to have outside observers rate the apparent mood of the subject. So now you have two variables as an outcome. You have subject-reported mood and observer-reported mood. I only have one subject-reported mood. Now we say, well, Danny, if you can analyze them both, you need to do a bond for only correction. So in order for you to declare that you have a statistically significant result, your p-values need to be below 0.025. On the other hand, I only tested one outcome. So my p-value only needs to be below 0.05. But we both got the identical study and with respect to self-reported mood, we got the identical result. And by the rules of the game, I say my study shows to a reasonable degree of certainty that eating an avocado leads to this effect in mood, p less than 0.05. You say, by the rules of the game, my study does not show that eating an avocado has this effect on mood, p equals 0.049, p is greater than 0.025, not statistically significant. We have the identical result, but apacip or different conclusions, only because you were ambitious enough to include an extra measure in your study. That seems counterintuitive and silly, but it's no different than the big one. It's just smaller. The one's big lots of testing. And many astute statisticians like Dr. Kenneth Rothman, former editor of American Journal of Epidemiology said, we shouldn't use multiple testing corrections. There's no right answer to that. It's a matter of judgment, but I think what we want to do is be transparent with the reader. So we want to give the exact p value. So instead of you just saying not significant and we just saying significant, if we both say and the p value is 0.049, then again, readers can judge for themselves. Something wants to do a bond-friendly correction on yours? They can. If they don't, they don't have to. You've given them all the transparent information. I mean, that's a recurring theme that we want this transparency in how things are going to be reported, but also that there's some thought gone into that as well. With respect to the p values in general, and never mind some of the corrections, there's obviously huge debate within the field of statistics around that we could probably spend hours only focusing on that. But sadly, we don't have time here. And to respect the time we do have, I'll maybe start trying to get to my final question or so, there's a whole host of things that I could ask you about David. But one thing that I think is worth coming to is in relation to nutrition broadly, there's obviously much debate about where the field is going, where it has been up to now, what areas can we improve? There's a whole range of different perspectives on that. And one that particularly maybe applies in the case of nutrition epidemiology, but we could also think of this more broadly within the field of nutrition. Given the type of exposures we're looking at in nutrition, given the outcomes we're looking at and all the factors that influence those that go outside of nutrition, in comparison to other fields, we are working oftentimes with effect sizes that are relatively small, let's say. And so then when we start thinking about a number of the potential measurement errors that can happen, this starts to become tricky. And the end is where we get into a lot of the discussion and perspectives that people have of where we need to go. So when you think about the general issue of this measurement error that is always going to come up, the effect sizes that we're working at within nutrition, all the challenges that come with the type of exposures and outcomes we are looking at. And again, at the fear of trying to get you to condense, what could be hours worth of discussion into one particular answer. For you, what are the most pressing issues that you think that as a field nutrition could and or should be doing much better now that would give us that biggest bang for the buck in terms of answering those questions we care about? I think I can provide a few things in my personal views, not a very good degree. What we could do to make more rapid progress. One is we could do a stronger calling of research questions to decide which ones are really important and which ones both important and answerable and which ones are either less important or less answerable. And then, and I'm thinking mainly about effects on health outcomes or other outcomes of interest in humans. And then for those that we think are answerable and important, we really dig in and we do much more powerful better studies. For the others, we just do a little less of them and accept more uncertainty. That's the first. The second thing is I think we really need to work on our methods and I think we need to stop spending so much time debating about things like the value of self-reported food intake and trying to tweak it in tiny ways and make it better or to apologize for it and think that by apologizing for its limits, that made the limits go away and we need to just say, that's a terrible method. And let's really move a great deal of the investment from using those methods to coming up with better methods that are based on biochemical and other types of tests and will have more validity if we can work them out. The third thing I would say goes to sort of the social behaviors. And I think there are three things we can do, not just in nutrition, but in science and general nutrition in particular, that would go a very long way. The first is get all trials and as many other types of studies as possible, pre-registered, chemically in a searchable database. That's starting to happen. We're not completely there, but we've made a lot of progress on that, including having a pre-specified data analysis plan. The second thing is that upon publication of the study, all raw data and all statistical code for reproducing every number reported in every line of text, every table, every figure in that paper are made publicly available online and immediately upon publication. None of this nonsense, data are available upon reasonable requests and then when you write to people, you find out that they don't consider your request reasonable. And then the third thing is with exceptions, there will always be exceptions, but with exceptions that should be rare. Publication, ag trials should be mandatory. That is no matter how boring you think your result is, no matter how little you think it will help you get your net's grant, tenure, promotion, fame, you are obligated to publish it. And I think that should be a condition of IRB approval. It should be a condition of accepting the grant funding, assuming you accept grant funding for it. It should be a condition of working out a non-profit institution like a university and if people don't comply with it, then I think the IRB should say, "We're not approving anything else for you until you do this." [BLANK_AUDIO]
grant fund agency should say we're not giving you any more grants and you're not always able to apply unless you do this and so on. There's a whole host of things I had planned that I would love to talk to you about I will leave you with the very final short question that I always leave the podcast on and it is simply if you could advise people to do one thing that would have a benefit for any aspect of their life each day what might that one thing be? Question everything. It's imperfectly with what we've discussed today and brings us full circle to I think your initial response to the very first question. So with that Dr David Allison let me say thank you so much for giving up your time to come and talk to me today. I really really appreciate it more so for the thought and that has gone into your answers to perspectives. I've learned a lot from your work and hearing your perspectives on nutrition science as a field has really helped shape a number of things personally. So I appreciate that and thanks for being a part of this. Well sir I have learned and continued to learn so much from your podcast that I'm truly grateful and if I've been able to repay that learning even a tiny bit today I'm happy to have fun. Before you go I just wanted to remind you about Sigmund nutrition premium. It was created with the goal of allowing you to more deeply understand the material you're hearing to be able to retain more of that and then be able to easily and efficiently revise so that in the future you can use it using things that you have learned. So for full details on this then check out the link in the description box wherever you're currently listening right now or just go to Sigmundnutrition.com and you can see all the details there. And of course your support is what keeps Sigmundnutrition going. We don't run ads, we don't sell supplements, anything like that. So your support is what allows me to continue to do this. So thank you for that. I hope you do come back for the next episode regardless and until then have a great week. Stay safe and take care.
Podcast Summary
Key Points:
Dr. David Allison, a nutrition and obesity researcher, emphasizes the importance of rigorous scientific inquiry in nutrition science, focusing on improving transparency, reproducibility, and trustworthiness.
The scientific process involves multiple stages
Common errors in nutrition research include flawed logic (e.g., unclear questions), invalid measurements (e.g., food intake), statistical mistakes like the "DINS error" (comparing within-group p-values instead of directly testing group differences), and misinterpretation of treatment response heterogeneity.
Miscommunication is frequent
Many investigators lack statistical training, leading to intuitive but incorrect interpretations of p-values, such as the prosecutor's fallacy (confusing the probability of data given a hypothesis with the probability of the hypothesis given the data).
Summary:
In this episode of Sigma Nutrition Radio, host Danny Lennon interviews Dr. David Allison, a leading expert in nutrition, obesity, and research methodology. They discuss how to improve rigor in nutrition science by identifying and avoiding common errors across all research stages.
Allison explains that rigor is not absolute but should be proportional to a study's importance, with transparency being key. , treatment significant, control not) instead of directly testing the difference between groups. This error remains common due to poor statistical understanding among investigators.
Allison also notes misinterpretations of treatment response heterogeneity, where variability in outcomes is mistaken for variable treatment effects. Finally, he critiques dissemination practices, where authors, press releases, and media often spin null or modest results into exaggerated claims, undermining scientific communication. The conversation underscores the need for better training, critical appraisal, and honest reporting to advance nutrition science.
FAQs
The steps include conceiving a precise and answerable question, designing the study, executing the plan faithfully, analyzing data, interpreting analyses, and communicating findings. Rigor should be maximized at each step, with transparency about any limitations.
DINS stands for Difference in Nominal Significance. It occurs when researchers incorrectly claim a treatment effect by noting a significant change within the treatment group (p < 0.05) and a non-significant change within the control group (p > 0.05), without directly comparing the groups statistically.
Many non-statistician investigators find frequentist statistics unintuitive, leading them to misinterpret p-values. They often incorrectly assume that a significant within-group change proves a treatment effect, ignoring the need for a direct between-group comparison.
Some questions, like those in the carbohydrate-insulin model versus energy balance model debate, may not be logically meaningful or answerable. It is crucial to check if a question is precise and can be empirically tested before designing a study.
Food intake measurements are often invalid due to reliance on self-reporting. Additionally, issues like improper image analysis in CAC measurements for coronary artery plaques can compromise study validity.
Many clinical investigators confuse variability in outcomes (like weight loss) with variability in treatment response. Without a control group, it is impossible to attribute outcome differences to treatment response, as outcomes can occur naturally.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.