In this episode of Quantitude, hosts Greg Hancock and Patrick Curran discuss the challenges and solutions for modeling count variables—discrete, integer outcomes bounded at zero, such as the number of drinks consumed or hospital visits. They highlight how ordinary linear regression fails for count data, violating assumptions of continuity, normality, and homoscedasticity, and often leading to biased results. Common but inadequate fixes include ignoring the problem, applying transformations like the natural log plus one (which distorts zeros), or dichotomizing the outcome, which loses information and can worsen bias. The recommended approach is the generalized linear model (GLM) with a Poisson distribution and a log link function. The Poisson distribution is tailored for counts, with a single parameter (lambda) controlling both mean and variance. The log link maps predictors linearly to the log of the conditional mean, allowing for multiplicative interpretation of coefficients after exponentiation—e.g., a one-unit increase in a predictor multiplies the expected count by e^beta. This method respects the natural structure of count data, providing more accurate and interpretable results than ad hoc fixes. The hosts also weave in entertaining anecdotes about Count Dracula, Vlad the Impaler, and childhood TV memories, making the technical content engaging.
[Music] Hi everybody, my name is Greg Hancock, and along with a friend I can always count on Patrick Curran. We make up Quantitude. We're a podcast dedicated to all things quantitative, ranging from the relevant to the completely irrelevant. In today's episode we talk about outcomes that are count variables. When you need to worry about them and what you can do about them within your analytical models. Along the way we also mention Bella Logosi, Vlad the Impaler, Patrick the Poker, Count Chocula, Count Von Count, Drunken Bar Brawls, Secret Distributions, K! Bio-Breaks, Second Favorite Child, Animal Farm, Cliffs Notes, A's in Band, and more equal zeros. We hope you enjoyed today's episode. We've already established that I watched way too much TV as a kid. Way too much! You were like, "5 hours a day too much." It's true, and those were the levels when I was with my single dad. But even when my mom was in the house when I was much younger, I still watched a lot of TV. I remember when I was, I don't know, I might have been 6 or 7 or so. I was flipping around the channels, and for most of you out there, flipping around the channels means something different than what it meant to us. It meant I was sitting right next to the TV, reaching up with my hand, grabbing the knob, and going, "Kitchunk!" And there were only so many places to go, right? I had 4, 5, and 7 where the main ones in Seattle, and then I'd have to go "Kitchu-Kitchu-Kitchu" all the way up to Channel 11. This is KSTW Channel 11 to call the Seattle. And I was flipping around, and I got up to Channel 11, and there was this old, tiny horror movie, you know, the old black and white ones from Wayback. Oh yeah. Right, right. And I was caught not by the visuals of it, which were not especially impressive, but this voice that came out of my TV, it was Bella Legosi. Do you remember Bella Legosi? Oh, he was terrifying, like in a coat and tie at a restaurant. He was terrifying. But if you put him in a horror movie, I didn't sleep for days. Yeah. So the thing that grabbed me about him when I was this little kid, and I yelled this to my mom, "Mom! The guy on TV sounds just like your dad!" [Laughter] And it was true, because Bella Legosi was playing Count Dracula with his Hungarian accent, and that was exactly like what my mom's parents sounded like. I am Dracula. I did you welcome. And I remember her coming into the room, and I said, "How come he sounds just like Apo?" That was the Hungarian name for my grandfather, and my mom got very dramatic, and she spread her arms out, and she said, "That's because Viara from Transylvania!" And she leaned down, and she started nibbling on my neck. But I didn't even know Transylvania was a real place when I was a kid, right? It was just the stuff of lore, but it's a real place. Located in the heart of Romania, Transylvania sits in Eastern Europe, surrounded by the majestic Carpathian mountains. You are fully aware of my inability to do any voice other than a pirate. But I would play games as a kid, where we would pretend to be from Transylvania, and I would say, "Argh, I'm gonna get you blood!" And that's really good. Other people were better. But that does make me think about you a little bit differently, that you have vampire blood in you. I have vampire blood, and my mom explained to me about the Transylvanian region where my mom and grandparents were from, and they spoke Hungarian there, because that land was both Hungarian at one point, and also remaining at another point. Transylvania's history is marked by territorial disputes between Romania and Hungary due to the presence of a significant Hungarian ethnic minority. This has led to tensions over ownership and cultural representation. But I remember when mom told me that this was a real place, I was like, "So am I related to Count Dracula?" And then she started to tell me about, "Do you know who Count Dracula is based on?" Wasn't it like Vlad the Impaler or something? Exactly! It is the best name in the history of humankind. Vlad the Impaler. I know. If we want to be known as anything, Patrick the Impaler, Greg the Impaler, totally. So yeah, she was explaining to me that there was a guy Vlad Sepesh, better known as Vlad the Impaler, and she clarified that, "No, we're not related." You know what I would be? I would be Patrick the Poker. Like, I'm not genetically predisposed to impale people, but I can go up to a stranger and go, "Pok, pok, pok, pok, pok, pok." And they say, "Stop it!" And I go, "Pok, pok, pok, pok." So I am Patrick the Poker. Absolutely you are. I can vouch for that. But I had this fascination with Count Dracula then, because all of a sudden I associated it with my grandparents, and this history that I hadn't really appreciated. And when I was, I don't know, around eight years old, they came out with a breakfast cereal. Do you know what breakfast cereal I'm thinking of? Count Dracula. Absolutely! Or as I would say, "Argh! Count Dracula!" That's totally my favorite. I mean, it was, first of all, way better than that crap, Frankenberry. Let's just be honest. But I mean, a breakfast cereal that involved a vampire doesn't get any better than that. Don't be scared. I'm the Super Street monster with the Super Street new cereal Count Dracula. And then around the time I was nine or so, Sesame Street introduced. Who? Oh, the Count! Ah! Greetings! I am the Count! I think his full name is Count Von Count. Not Kelt the Impale. They wanted to go with Abbott, but it didn't test well with your audience. How many children will I am bailed today? Lon! Two, three, four! Hi! It was a different time back then. It was very different. But it was all things Count back then. Ah! Alright, throw me a bone here. Where's this going? What? You texted me last night and said we're going to do an episode on Counts. That's all I know. So I don't know what else we're talking about. That actually is true. Alright, we're going to talk about regression models for Count data. Do you see how that worked? Very nice. Yeah, oh, it's seamless. But what we want to do is talk a little bit about what do you do in your modeling? Whether it be regression or we can do it in a whole lot of other things. The SEM, even factor analysis for that matter. But we're going to focus on the regression model. What do you do if you want to build a predictive model for a dependent variable that is governed by a Count? That is, it is discrete, it's integer and it has a lower bound at zero. And I've got to tell you, this is left right top bottom and center in what we do in the field in the last 30 days. How many times have you consumed five or more drinks in a single setting over your lifetime? How many times have you been arrested? How many hospital admissions have there been? We eat and breathe and live counts in a lot of areas of research. Yeah, and the way you described it, you were just violating assumptions of regression, at least the traditional linear model. You're just ticking off violations left and right there. And let's go back to Gaussian Markov's corpses that I'm in my backyard. Very briefly, what do we assume in the regression is our dependent variable is continuous. And the residuals are independent, homoscedastic and normally distributed. Now to be clear, if there's a stickler out there is, Gaussian Markov doesn't invoke normality, least squares is asymptotically distribution free, but in finite samples, we have to invoke normality for standard errors and mean squared error and test statistics, things like that. But we are very much losing a drunken bar brawl when we bring account into that in violating several of those assumptions. Yeah, so let's see, you said continuity, nope normality in the residuals, nope, homoscedasticity, wait, okay, so equals spread across your regression line, if you want to think about it that way. Well, how can that be? If you get down near to zero, you can't have a lot of variability around there. And as you move up to higher accounts, then there's a lot more room for variability. So this is just fundamentally in violation of pretty much every assumption that you have there. And this has been widely known for half a century. So what is the number one thing we do as a result of this in applications, Dr. Hancock? Not a damn thing. Not a damn thing. So let's think about if you have account variable, what are the typical ways we deal with this in practice? Well, by far, we just ignore it. You say, it's close enough for government work. I'm sure it's fine. All right, it's not fine. We can fit a model that very often will have a negative intercept, which indicates a count below zero. That's right. So minor concept
problems, but standard errors are wrong, test statistics are wrong, and indeed the very regression coefficient itself can be biased because it's assuming it's running a conditional mean up and down a line, which it's not because you are a zero or a one or a two, and that is not an infinite number of values on a line. So number one, we just ignore it. Number two, we pretend to do something that actually makes matters worse. We take a nonlinear transformation. Oh, sure. Why not? So you say, oh, I've got the spike at zero and then one and then two and then this long tail. But you do that and you get a bunch of errors and it's like, oh crap. So what do you do? Add one. There you go. So there's even a term for it where you ever talk about started blogs. I wasn't taught about any of this stuff. I mean, that's another problem entirely, but you can't take the natural log of zero. So what started log is you add one to everything. Yeah. All right. So now instead of zeroes, all of your zeroes are one. Do you see how we fix that problem? And then you take the natural log, but it only pulls in those extreme cases. If you had 60% at zero and you had one, you now have 60% at one. So the last one, that doesn't help. There is some utility in some cases. People will move to a robust estimator and we'll say, okay, well, we're going to treat it as continuous and we'll just correct standard errors. That's better than the first two, but you still don't have a linear relation between your predictors and the outcome because it's not on that continuous number line. All right. Now let's really damn the torpedoes and do some horrible things. You dichotomize it. So everybody in the zero gets a zero and if you're above zero, you're a one done and done zero zero zero zero one zero zero zero zero one one zero zero zero zero zero one one one zero zero zero one one one one one one one one. And then there is actually one that may be worse than that. And I'm afraid we use this routinely and I'm afraid I have maybe a dozen papers where I have done this. So this is not a whole year than now, but you either manually do this from raw counts or worse still you give an ordinal variable that says in the last 30 days, how many times have you consumed five or more drinks in a single sitting, none, one to two, three to five, six to 10, 11 or more. I'm not sure we could design a worse way of doing it than that. That's the worst scale ever. And that is widely used. And that is again, no holier than now, folks, is yeah, a lot of times that's how you need to gather data. That is expedient. It's easy to computer program. So I get this. Uh-huh. But these are all ways that we have a round hole and a square peg where we're pounding the crap out of the square peg to get it into the round hole. Yeah. But there's a secret distribution that nobody told you about. At least those of us in the social sciences. Oh, okay. Public health or like. Yeah. Are you guys in? How did you get tenured? Hello. We're even here. So a way to think about this super secret distribution is think about it in the context of the generalized linear model, which we talked about as being way generaler than the general linear model. But we can think of it as being made up of like when we do linear regression, which we're not doing here, we have some response variable with normally distributed residuals and then a link function that we talk about as being the identity function when we're just doing regular old linear regression. And then when we move into a binary outcome, which we've talked about before and people out there are familiar with, we have some Bernoulli distributed outcome and then we have a link function that we commonly will do that is a logit. In this case, we have a count outcome and the type of distribution that we will often use for counts is going to be something called a Poisson distribution and then we will have a link function that goes with that that is a logarithm. And this is going to be something that is very, very well tailored to dealing with count data. Now what I will say is a little asterisk here. We're talking about count data that are specifically down near the zero region and usually have a fairly narrow interval for occurrences because if we were talking about things that were genuine counts, but they were around 50 and people might do things that are 45, 47, 51, 58, it's quite possible that the linear regression methods that we use would be absolutely fine up in that region. But as we start moving down toward that zero, things get really messed up and we need the special distribution and the special distribution is the Poisson. I love the Poisson distribution. I honestly don't remember if we talked about that in the zombie episode of the distributions. I don't think we ever got to that. I don't think we did. But if you listen to that episode, it was supposed to just start with this really simple game show idea and then zombies came in and then it started biting other distributions kind of went off the rails. But very briefly, Poisson is a beautiful distribution and what it is, it's for a discrete outcome, just like a bear new layer of binomial. It is weirdly governed and we're going to talk about this in a moment by a single parameter that captures both its mean and its variance. Now this is totally cool, right? This is the way it's designed, but it imposes what's called equidispersion. That is the mean and the variance are equal. And so we're going to have to deal with, well, what if they're not and we will talk about some other distributions that allow for that. But it literally is a sausage maker where if you have the mean of the count distribution and you can look up on Wikipedia what it is, it is e to the negative thing over this other thing. I remember it has a factorial in it, right? That's it. That's what you got. Yeah. Well, it's divided by k factorial. And so if you remember as you denote a factorial with an exclamation point and if you're a math person, you know that's k times k minus one times k minus two times k minus three. If you're not a math person like me, you read it and you say, oh, you divide by k. Hi, this is Jiffy in post processing or Jiffy, the impaler. Anyway, I just thought I'd insert the full formula for the plot sound distribution since Patrick the poker, well, he kind of impaled it. In the plot sound distribution, the probability of a count of k occurrences is equal to the mean parameter lambda to the power of k times e or there's constant to the power of minus lambda. And then the whole thing is over k factorial or as Patrick said, k. But it's a sausage maker. And if you have a mean of a Poisson and you drop in the particular count that you're interested in, let's say I want to know given a mean of whatever, what is the probability that I would observe a count of three? You drop it in. It's a very simple expression and it literally gives you the probability that a randomly drawn response is three. It is really cool because as Greg says with the generalized linear model, we say, oh, we have a count variable. So we're not going to wail away to try to pound a square peg into a round hole. We're going to say screw it. We're going to pick a Poisson distribution and side-stray a had a good friend who had a fish and he named it Poisson. And I thought that was like subtle on so many levels. But we're going to pick that as the response distribution. Then the beauty of the generalized linear model is we're going to fit our linear predictors in exactly the way that we do. We can have beta one x one plus beta two x two. We can have interactions. We can have polynomials. We can have blocks of variables go nuts. But the link function as Greg said is we are going to map that linear combination of the predictors onto the log of the mean of the Poisson distribution. And we're going to fit our model in that way. And we're going to interpret the regression coefficients in that way. It's beautiful. And we're building a model that is unique to the natural distribution of the counts. We're not rocking back and forth and saying I've got robust standard errors. I've got robust standard errors. A screw it. Do it the way it came to you. Now I'm going to insert a little fun fact here. I was just going to call it a bio break. But if I didn't mean I have to pee. I mean fun fact. Poisson was a French mathematician. And to give you a sense of his accomplishments by the time he was our age Patrick, he'd been dead for about two or three years. So his full name was semi-deni Poisson. I hope I didn't embarrass myself. You're most certainly did. But he was a mathematician and he worked in physics, stuff and statistics and differential equations. I mean his fingerprints are all over all kinds of cool stuff to the point weird.
He was given an official title that title Unfortunately is Baron, but he should have been count plus on Damn it. It was such a missed opportunity. I think historically he was The embaylor Now a couple episodes ago we talked about how logs and exponents are magic and what the log giveeth the exponent takeeth away Well, if we're fitting an optimal linear combination of our predictors to the log of the conditional mean We got to get back to the metric that we're interested in. How do we do that? Right we have to mess with these kinds of things when we're doing logistic regression also So this is going to have a vague whiff of what we did there now you could talk about it in the linear model if beta 1 is the coefficient for let's say we have one predictor x We could say something like for a one unit increase in x then we would expect beta 1 increase in Etc right but the unfortunate part is that the answer is the increase in the log of the outcome right and we don't think in terms of Logs of outcomes at least I don't so we have to get rid of that log and how do we get rid of logs? We take what is sometimes called the anti log or we exponentiate it so by exponentiating both sides We take e we talked about Euler's constant before e to the power of log of y which just gives us y over on that side But then when we take e to the power of beta 1 x that Changes the interpretation right so now what we would say is for a one unit increase in x We would expect e to the beta 1 increase and this is where we have to be very careful Multiplicative increase in our outcome y in our count So if beta was one just for convenience then e to the power of 1 would be 2718 we would expect the count to go up 2.718 times for every one unit increase in x and what I love about that is In all the pounding of the square peg that we do with the linear regression model We in this case have a metric that means something and we can interpret right all of you who are listening right now How often do we actually have a metric that we can assign some kind of meaning? I mean so much of what we do are ordinal or their means of items or whatever is this is lovely But remember when we talked about in the log linear episode is that one of the beauties about logs is it makes product sums Yeah, well if we undo it with the anti log it makes sums Products so that's why it is a multiplicative effect. Yeah, and I have to admit I have seen a lot of applications where betas are interpreted Additively where they really are multiplicatively and so you just got to keep your head up about that Yeah, but if you're thinking about this with your own data Just think oh I can do everything I wanted to and I mean everything you want to in the regular regression model But you can fit it to the native distribution of your count outcome and then when you interpret the parameters as you can say a One unit change on the predictor is associated with this 1.2 times Increase in the count of the outcome. Yeah, which is very natural. I love it now in the spirit of it is perfectly Safe to work with explosives in tell it's not the press on is flawless Intel it's not yeah Right so there was that assumption that you mentioned when you were avoiding giving us the exact equation for the press on distribution The idea that the mean parameter and the variance parameter of the distribution are the same right and that's that's an Assumption baked into the distribution, but in fact that might not happen right and it often does not happen and Then you need to expand the model to allow for what is called over dispersion So Poisson is equidispersion and then we need to think about over dispersion now There are two big ways that we can do this one is what's called a quasi Poisson and it's kind of the equivalent of having robust standard errors You have the same point estimates, but you try to go in after the party and clean things up I'm not a huge fan of that okay because there is another super secret distribution that if you're not in biostatistics called the negative binomial and Of course none of us has a favorite child Right we love our children all equally well Negative binomial. I really like yeah the Quasi Poisson distribution does allow you to have under dispersion and Over dispersion so in that sense it could be if not a favorite child a close second But we almost never have under dispersion We almost never have a variance that is less than our me It's a very hard thing to have happen in the types of distributions that we're talking about that start budding up towards zero So the negative binomial distribution it is really focused on over dispersion and that's what we tend to encounter most of the time We do indeed and the negative binomial is simply an expansion of the Poisson This is what I really love where it has a parameter for the mean and then there's an additional parameter to represent the over dispersion and what I love about that is When that over dispersion parameter goes to zero the negative binomial goes to the Poisson yeah And as the mean of the Poisson goes to infinity the Poisson goes to the normal distribution You can picture that count going to the right going to the right going to the right Yeah, and kind of rule of thumb we sometimes have as people say if you have a count with a mean of about 10 It's more or less normal yeah now to be super clear It is not normal because the normal is a continuous distribution of the Poisson is a discrete distribution But as the mean of the Poisson goes to infinity the Poisson goes to the normal But as the dispersion parameter of the negative binomial goes to zero the negative binomial goes to the Poisson So this is why I kind of prefer this over a quasi Poisson kind of approach Which is if you don't need that over dispersion then it just functionally goes to What you would have if you use the Poisson, but you're not imposing that on the data So in my own work when I'm doing this and Mostly it's in collaboration with colleagues most of my work is in substance use drug use things like that And these are very common kinds of outcomes in that field. I just start with the negative binomial I don't even go to the Poisson That's what I was going to ask you and it it's a question that comes up with a lot of different robust methods If there is a more flexible version unless there's some horrific estimation downside to it or you know It blows up under these conditions. Why not just start with something like this And I don't know what the downside would be of just going straight to the negative binomial The only one that I think of is it's just a little bit more complicated You have one more parameter that needs to be estimated It's a little bit less Parsimonious because the Poisson has one parameter the negative binomial has to We're not going to go dumpster diving and technical, but like any of these you need maximum like the hood estimation for these models There are new developments in Bayesian that are very cool for doing this, but ML is still pretty widely used and so The only one is you may be in a situation where you have a smaller sample or some other complication And it's easier to use a quasi Poisson But that's about the only thing and the nice thing is is with the negative binomial you interpret the parameters in the same way right All it's doing is factoring in that over dispersion that trickles into the standard errors and the test statistics But everything else is interpreted in exactly the same way That's a great point that over dispersion is not going to have so much of an effect on Byusing the betas that you're using to interpret the will say the effective x on your outcome count variable But it is going to mess with your tests over dispersion will start to give you standard errors that are too small Standard errors that are too small give you p-values that are too small which give you too many type when errors So exactly right. This is just fine tuning the standard errors for the significance tests So I feel like that problem is handled But there's another problem that bothers me when we're dealing with count data And I would guess it's something that you encounter in the work that you do with colleagues And that's when you've got too many zeros When you were in middle school did you read animal form? Of course I read animal Did you because I didn't I read the cliff notes I was a horrible student The only reason I had a 3o GPA in high school was my A's in band offset all my C's That I got in my other classes but I went back and we read some of the classic success in adult And one of those was animal form
And there's a very famous line in there. All animals are equal, but some animals are more equal than others. All zeros are equal, but some zeros are more equal than others. And that applies here, which is, yes, we get a piling up at zeros. So imagine that I'm looking at alcohol use in kids, and I say in the last 30 days, how many times have you drank to the point of an ebriation? All right, from a public health parent societal perspective, I am really glad they're 80% zeros in that response. From a maximum likelihood perspective, would it kill you to go out with friends and have a couple of drinks? Just for the sake of the modeling. Just have a beer. You get a piling up. Well, this is called zero inflation. And the inflation part, you need a relative statement, right? Is relative to what? Well, they're more zeros than you would expect based on the Poisson or the negative binomial probability mass function. That's the zero inflation. It's like, oh crap, they're more zeros here than we anticipated. Well, there are a lot of different terms for this. But the one that is commonly used is their structural zeros and non-structural zeros. I don't like that so much in my own neck of the woods. I like thinking about zeros that are no risk versus risk. But there are many terms as papers are out there. But what it means is some of those zeros are zeros because they have never experienced the event. Right, imagine that you are studying alcohol use and there are people who have not started drinking alcohol. And you say, how many times have you drank to an ebriation? Well, it is zero because the construct is irrelevant to them. They don't drink. Right, so how many marital arguments have you had if you're not married? Right, there are an infinite number of these examples. So it's a structural zero. But then there are other zeros that are zeros that could have been greater than zero. But they just weren't within the timeframe. So the example here is you have a drinker who just happened to not drink to excess in the last 30 days. So then there's a hanker and a, well, can we separate those zeros of those that it doesn't apply because they're structural zeros? They have to be zeros from those that are zeros, but they could have been one or two or three, but they just weren't within that timeframe. Yeah, and the way you describe them, it sounds as though there might even be two populations at work here. And those populations could differ on a variable that we know is the person married or not? Well, then I know that they won't have had any marital argument. So then that distribution is a mixture of people who have zeros by virtue of being unmarried and people have zeros who are lying to you. Tough morning? Sorry. Hang on, I'm getting a text. Okay. But it could also be viewed as a mixture of some unknown characteristic, right? So that was an easy one, the married unmarried. So you've got this piling up of zeros. For me, it could be that they exist in a different population and you just don't have the identifying characteristic. And in some ways, you're sort of using the statistical distribution to tell you, oh, those must be a different kind of zero down there. And that's exactly right. And I actually have a love hate relationship with these models. What we're going to do is we're going to throw a zi in front of Poisson and negative binomial. So these are zip models as zero inflated Poisson or zin B, a zero inflated negative binomial. And I would say the one I've used the most in collaboration with colleagues who are working in this area is the zero inflated negative binomial. I love that it accounts for the piling up of zeros. It accounts for the over dispersion. You get to have a native fitting the linear combination to the response distribution that's governing the data. I love all of that. But the problem is there is a little bit of magical thinking in these models. But that's no different than latent class analysis in any form that it might take is you're exactly right. There are two subgroups. And some zeros belong to one subgroup and some zeros belong to the other subgroup. Now, if we only have the marginal distribution of the count, that is all I do is just give you, hey buddy, we got 200 kids. And here's the distribution and there's a piling up at zero. You can't separate those. There's no additional information. So what we do is we build what's sometimes called a two part model. And we want to bring in a set of covariates where in one part it's literally a logistic regression of are you a structural zero or are you in the count process. It's a zero one. And it's quite literally a logistic regression. You infer the assignment from the set of predictors and you can have main effects, interactions, polynomials, whatever as you would in the logistic. But at the very same time, right, this is simultaneous. So if you're doing this, do not make a zero and a one manually with your data and fit the model. This is model inferred at the same time you say, well, if you're in a count process, then we're going to fit a model to that in the usual way. So it's a two part model, a logistic prediction of are you a structural zero or not. And then if you are not a structural zero, you fit a predictive model to the count process. And why I have a love hate relationship with that is I love that conceptually. But my eye twitches at the separating zeros when they're all zero. Yeah. And it doesn't twitch any differently than what it does for me in late in class analysis or mixture models or anything else. We are making a probabilistic inference about something we didn't observe. And that's okay, but I just am careful with that. And I don't know what else to do. Frankly, it does feel like the statistical tail wagging the zero dog here. And that we're using distributions to make decisions about whether you're real zero or a fake zero. Hello, hello, let me tell you what it's like to be a zero zero. Let me show you what it's like to always be a real. I don't know if the end is nothing really real. But again, I don't know what else to do other than that. I mean, it's very clever and it's imperfect and I'm okay with that. Oh, totally. And I want to be super clear that I really, really like these models. And I think they are the way to go in practice. It's just I think there's some caution in needing reproducibility in thinking very carefully about design. Right, should those structural zeros be in the sample at all? Right, if you're studying predictors of use, do you want abstainers in your sample? And actually, people argue both sides that you don't want to select on the dependent variable, right? Sociologists have been telling us for a century. Don't select on your dependent variable and then predict your dependent variable. And we would be doing that. Right. If we asked a lifetime alcohol use item and said, have you ever in your life and just had more than 12 ounces of beer, one glass of wine, whatever that might be. And somebody says no. And we say, oh, we're going to put in an if then loop to pull those people out. Well, you're literally selecting on the dependent variable and then predicting what's left over. There's no right or wrong answer to this. We just need to be cognizant of what we're doing. But you can have some really powerful models where you have one set of predictive effects on the structural zero or not, but have a different set of predictive effects on the count process. Right in the output of your computer program, you literally get two sets of regression coefficients, one that's associated with the logistic structural versus nonstructural, and one that's associated with the negative binomial count process. And those can each have different substantive implications on what are the between person factors associated with these two parts of that process. And from a planning perspective, then when you're going to do a study with a count variable as an outcome, and you anticipate that you might have some zero inflation, that you should think about what data you're gathering, not just to explain the count process, but also to explain the zero inflation process, because there might be variables that are directly relevant for that. And the poorer job you do in the predicting of the zeros, then the less well you're correcting in the prediction of the count process. So planning ahead to think about what variables are relevant for both processes seems wise. And you know where these things get really interesting is longitudinally, because imagine that you're doing this in a repeated measures setting. Well, those abstainers at an earlier time each year have an opportunity to move into the count process. Yeah. If you're studying onset and escalation of substance use and adolescence, well every year, more and more kids move from the abstainment process.
to the experimenters, to the users, so you wouldn't necessarily want to just strip out the abstainers and look at who's left, because in a next time point, some of those abstainers are going to have moved into the count process. So we haven't even raised how do you expand these into a longitudinal setting, and it's cool as poop. - I don't know about the poop part, but that sounds incredibly cool. Maybe that's something we could talk about another time. - So I gotta tell you, man, as Patrick the poker, I feel a need to grow out in the world and poke people. So I think I'm about done here. Puk, puk, puk, puk. If you keep doing that, I'm gonna impale myself right here, right now. All right, thanks very much everybody. Puk, puk, puk, puk. (upbeat music) Thanks so much for joining us. Don't forget to tell your friends to subscribe to us on Apple Podcasts, Spotify, or wherever they go for things they can sink their teeth into. You can also follow us on X, Blue Sky, or Instagram, and visit our website, quantitativepod.org, where you can leave us a message, find organized playlists and show notes and syllabi, search for past episodes, and other fun stuff. And finally, get cool Quantitude merch, like shirts, mugs, stickers, and spiral notebooks from redbubble.com, where all proceeds from non-boot leg authorized Quantitude Pod merch, go to donorschoose.org to help support low-income schools. You've been listening to Quantitude. The podcast that sometimes makes you wish you were a victim of LADD, the impaler. Today's episode has been sponsored by famous statisticians and their statistical nicknames, like, "Baze the Prior." Obviously. And Pearson, the product moment co-varior. Ooh. And Gosset, the Vineite Sample, non-normal distribution brewer, and right, the tracer of legal paths. And Fisher, the maximum likelihood to announce a severance creator and colleague irritator. In Spirman, the whole field of study and a footnote inventor. And Hôtelling, the principal component creator and factor analyst, Annoyer. And Bone Farone, the type, what-airer's layer, and power killer. And finally, Bauer, the progenitor of unpronounceable acronyms. Mnumulfa. Mnumulfa. Our nickname is most definitely not NPR.
Podcast Summary
Key Points:
Count variables are discrete, integer outcomes with a lower bound of zero, common in research (e.g., number of arrests or hospital admissions).
Using ordinary linear regression for count data violates assumptions of continuity, normality, homoscedasticity, and can produce biased coefficients and incorrect standard errors.
Common but flawed practices include ignoring the issue, applying nonlinear transformations (e.g., log+1), using robust estimators, or dichotomizing the outcome.
The Poisson distribution is a natural fit for count data, designed for discrete outcomes and governed by a single parameter (lambda) that represents both the mean and variance.
In a generalized linear model (GLM), a log link function maps predictors to the log of the conditional mean, and coefficients are interpreted as multiplicative effects after exponentiation.
Summary:
In this episode of Quantitude, hosts Greg Hancock and Patrick Curran discuss the challenges and solutions for modeling count variables—discrete, integer outcomes bounded at zero, such as the number of drinks consumed or hospital visits. They highlight how ordinary linear regression fails for count data, violating assumptions of continuity, normality, and homoscedasticity, and often leading to biased results. Common but inadequate fixes include ignoring the problem, applying transformations like the natural log plus one (which distorts zeros), or dichotomizing the outcome, which loses information and can worsen bias.
The recommended approach is the generalized linear model (GLM) with a Poisson distribution and a log link function. The Poisson distribution is tailored for counts, with a single parameter (lambda) controlling both mean and variance. , a one-unit increase in a predictor multiplies the expected count by e^beta.
This method respects the natural structure of count data, providing more accurate and interpretable results than ad hoc fixes. The hosts also weave in entertaining anecdotes about Count Dracula, Vlad the Impaler, and childhood TV memories, making the technical content engaging.
FAQs
A count variable is a discrete, integer outcome with a lower bound at zero, such as the number of times someone has been arrested or had hospital admissions.
Linear regression violates assumptions for count data because the outcome is not continuous, residuals are not normally distributed, and variance changes near zero, leading to biased coefficients and incorrect standard errors.
Common flawed methods include ignoring the issue, using log transformations (e.g., adding one to avoid zeros), applying robust standard errors, dichotomizing counts, or converting to ordinal categories.
The Poisson distribution is a discrete distribution for count data governed by a single parameter (lambda) that equals both the mean and variance, making it ideal for modeling counts near zero within a generalized linear model.
In a GLM for count data, you use a Poisson distribution as the response distribution and a log link function to map linear predictors to the log of the conditional mean, preserving the natural distribution of counts.
Coefficients are exponentiated to interpret as multiplicative effects: for a one-unit increase in a predictor, the expected count multiplies by e to the power of the coefficient.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.