This episode focuses on bivariate statistics, which analyzes the relationship between two variables. The core concept is correlation, typically Pearson's r, measuring linear association from -1 (perfect negative) to +1 (perfect positive), with 0 indicating no linear relationship. Scatterplots visually depict correlations, while contingency tables and stacked bar charts summarize frequencies for categorical data. Box plots compare group distributions, showing medians, quartiles, and extremes. The lecture emphasizes that data measurement levels (nominal, ordinal, interval/ratio) dictate suitable methods: mean comparisons for interval dependent variables, correlations for both interval variables, and contingency tables for any type. Examples include a study on alcohol consumption increasing reaction time (shown with a line graph) and fashion awareness differences among study programs (shown with bars, as the independent variable is nominal). Practical techniques like median splits or clustering can simplify continuous independent variables. Crucially, researchers must test whether observed effects are statistically significant, typically using t-tests or F-tests to calculate the probability that results occurred by chance. A p-value below 0.05 indicates significance. The episode concludes by reminding students that even weeks of work may yield simple graphs or mean comparisons, and significance testing is mandatory for scientific rigor.
[Music] Hello and welcome, I am Armin Trust, Professor for Organizational Behavior at the Futtwangen University in Germany. And this is my course on Social Research Methods. [Music] So, hello everybody. This is the third, with a second episode regarding statistics. Last time we were talking about descriptive statistics, this time we're going to talk about Bivariate statistics. What is Bivariate statistics that's already an essential term? Bivariate as this Bivariate indicates is about the relation between two variables. Okay? I mean, in this series we already had a lot of examples where we looked at the relation between two variables. Right? So, let me guide through different methods that you can use. Simple methods and one of course is the correlation. I know that in the previous episodes I already were talking about the correlation. And when we were talking about testing, for instance, we were talking about correlation a lot. And also when we were talking about validity, reliability and objectivity, I had to refer to the correlation. So, all those of you who already have watched the previous episodes, you should be familiar with the correlation. For those who are only wanted to watch only this episode, I would like to explain correlation real quick. So, what is correlation? When we talk about the correlation, the correlation, we talk about the Pearson correlation. There are various types of correlation statistics, right? But in practice, when we refer to the correlation, we pretty much mean a statistical concept that describes that linear relation between two variables. Two variables x and y. So, if you have two variables x and y, the correlation between the two could be between minus one and plus one. Plus one would mean that there is a perfect positive linear relation between two variables. The higher the one, the higher the others. Absolutely. Perfect. Linear. One. Okay. Zero would mean there is no correlation at all. No, nothing. One variable has nothing to do with the other. Right? So, when you now have a correlation of, let's say, 0.5, that would mean that, oh, there is something. There is a positive relation, but it's not perfect. There is something 0.5. Right? Based on one variable, you can predict 50% of the other variable. That's the idea of 0.5. When you know the variable, you already know the half of the other. Right? What is the other half? We don't know. We don't know. Right? If you have a correlation of, let's say, minus 0.5, it's exactly the opposite. Just negative. Yeah? The higher x, the lower, not higher, the lower y. Right? So, now you can imagine this. And with correlation, please keep in mind. It's linear relation. Sometimes relations between two variables are a curfee, linear or whatever exponential. With the classic relation, we always think about the linear correlation. So, when you have two variables, which are at least on interval scale or ratio scale, you can calculate a correlation. That's something that we very often do. What is the correlation between chop set is faction and performance? Something like this. Right? What is the correlation between chop set is faction and consumer satisfaction? What is the correlation between intelligence and performance? What is the correlation between intelligence and salary in a specific age? You can create all sorts of things. When we think of this kind of correlation, you very often use scatterplots to illustrate the relation between two variables. A scatterplot is nothing else than having the two variables, the independent variable and the dependent variable. For every subject in our study, we have two values. One for the independent and one for the dependent. And then when you have these two axes, you can assign every subject into this kind of two-dimensional craft. And when you do this, let's say for 100 subjects, you get 100 dots. And these 100 dots, they make up what we name a scatterplot. A scatterplot is a cool way to illustrate the entire relation distribution of two variables. That's pretty nice. A much simpler way of looking at the relation between two variables is using what we name contingency table. Our contingency tables are just about the frequencies that you show. You can use contingency tables for nearly all data, even for interval data on ordinal level or nominal level. I have here a very simple example. I show you some examples now from studies that we have done as part of my social research method class. We used to, my students are encouraged to run little studies. Very little studies. Very often in the field. Very rarely in the laboratory, but in the field. So they probably look at two variables and they look okay, it's their relation. That's pretty funny. And there were two students who took up the idea that women take longer to park in a car, which actually turned out not to be true. So independent variable was gender, male, female, and the dependent variable was parking speed. So you might have a very simple table. Building clusters up to 13 seconds or 30 to 20 seconds, more than 20 seconds. Women, men, what are the frequencies in these different categories? So a very simple way of looking at the relation between two variables. These kinds of analysis you also find in newspapers. It's very simple. And when you want to illustrate something like this, you also can have something like a stacked bar chart, stacked bar chart. You simply have two bars that add up to 100%. And then you have these different segments. In this picture here, you see the exact same information as in the contingency table, just in a more in a graphic manner. It's also pretty nice. And you can also combine this and now add these numbers into these different bars. Also a nice idea. A more advanced idea is something like Box Bok's Visco diagram. And here is a comparison of different groups. Box Visco is very nice, I think. Because with Box Visco diagrams, you not only see the average. It's not only that you see the different median or so. You also see the dispersion around the median or the average. So you have the upper quartile, you have the lower quartile, so the first quartile, you have the most extreme value, the most extreme upper value and the most extreme lower value. You not only see the comparison between different groups, but also how the values are distributed around the central tendencies. So it's pretty nice. I like the Box Visco diagrams. It's a cool way. So let's be more precise now. And I would like to share with you again two other little studies, which we have done as part of my social research method, chorus, which I do here at the University of Fort Wangen. And just to show you how you can simply analyze the relationship between two variables. And here we really talk about a little kawasi experiment. And in this case, students investigated the effect of alcohol consumption on the reaction time. You can measure reaction time by the falling stick technique. I don't have one here now. You simply have a stick, a linear, a pen. So one has to put the fingers around the pen not touching it so that it can fall. another person let it fall
And when it falls, the other person has to grasp it as fast as possible. When it falls. So, and now when the stick is here and you let it fall, it falls down and you fast, you grab it really fast. And then this short distance is the exact measurement of your reaction time. It's pretty cool, right? It will be the stick. You can imagine this. Okay. Simple. So, reactivity was measured in four cycles. So in the beginning, then they train a beer and then again, train a beer, measure again, train a beer, measure again. And you see, the reaction time goes up. Right? Hmm. It was not surprised so much. Okay. So here is another study. We created a self-administered question here that was supposed to measure our fashion awareness. That's how we renamed it, fashion awareness. We only asked male students, by the way. And then we just asked some questions about dos and don'ts regarding fashion. So is it allowed to wear a tie with a short sleeve shirt? No. Don't do it. When you wear brown shoes, what's the color of your belt? It's the same, yeah. If you have brown shoes, brown belt, black shoes, black belt. When you have, ah, what is wear? It's good. When you wear a button-down shirt, like this is a button-down shirt. Yes, this button here, you know. See, is it allowed to wear a tie with a button-down shirt? Is it allowed? No, it's not. Why it sucks? Is it allowed? Only when you play tennis. Is it allowed to wear brown shoes after six? No. Um, yeah, some. Some rules. No, some rules. What means business casual? So we ask all these things and some students got everything right, some just a few. So we calculate an overall index of fashion awareness. And we compare the fashion awareness index among different study programs. There were engineers, business students, computer scientists and social scientists. And we found out that the business students had the highest level of fashion awareness and the computer scientists, the least, they simply did not care. Okay, so you see this different analysis, which I just have shown you. One was with a line and one was with a crab, with the bars. Right? So in this case, now here, could we also have drawn a line? That's the question. A line, yeah? An engineer and business. Line? The answer is no, why not? Because we cannot interpret the space between the bars. Right? With the other, with the alcohol consumption, a line, it's okay. Could we also have shown a bar? Yes, of course. But why could we have shown a bar and line in the other example, but not in this one with a fashion awareness? Simple answer. Because in this particular case here, that you see here, the independent variable is on what level, what kind of data do we have here? Is that nominal, ordinal, interval ratio? It's nominal, right? So here it is again. So when the independent variable is nominal, you do not draw a line. In this case with fashion awareness, you could also change the order of the different study programs would be absolutely okay. But with the alcohol consumption, there was really an order from no alcohol up to much alcohol. There was really an order. It's not nominal scale, it was at least ordinal scale, I would say, rather interval scale, because the amount of alcohol that was consumed was absolutely constant between every interval. Okay? So you see here is it again? It's absolutely important. When you think about the analysis of the relation, you have to understand on which level is your data? Nominal, ordinal, interval ratio. And if you don't know what the difference is, go back to my episode which I have produced earlier about these different types of data. So I prepared for you a kind of overview, which really helps you. When you, in your thesis or in your little study, when you, when you analyze the relation between two variables, the question really is on which level is your independent variable and on which level is your dependent variable? Yeah, in terms of nominal, ordinal, interval, or ratio. And I added one thing here, decotermus. Decotermus are variables which just have two options. Sometimes something like yes or no, pregnant, non-pregnant. In a very ideal world, let's say simple world, children, picture book world, you have male female, I know there's many more of them, but male, female, decotermus. So very often you have just two, two, two categories decotermus. Okay? So you have to be aware of what you, what you have. And based on what you have, you can do different analysis. Of course, you can do something like contingency tables all the time. But, but when you, when you independent variable, when you depend variable, sorry, when you depend variable is interval or ratio, you could do something like mean comparison. So you have, you have your independent variable, right? And then based on the independent variable, you can calculate averages of your dependent variable and compare it. Just like we did it with your, with our fashion awareness thing, right? Correlation, you can only do when both the independent variable and the dependent variable are at least on interval, interval data, interval or ratio. Also when it's decotermus, when something is decotermus, you can do everything. Yeah? Interesting enough. So that's, that's very important to see this. And I added two more concepts here. It's the median split and the clustering. Sometimes the independent variable is a continuous variable. Let's say the, the, you do an analysis about the relationship between salary and intelligence. You say, okay, let's look at people who are in the age of 40, let's say, okay? So we have people in the age of 40 and we look at their intelligence, okay? We measure their intelligence and then we ask them, okay, how much do you earn? And as you know, intelligence can rank from, the average is 100, but it can rank from, I know, 60, which is very low, 160, which is genius. So you have this range. And I mean, what you can do now is really you, you, you, I mean, that's a continuum, right? It's a continuum and you have people in all different values. So what you can do is you can just take the group and split it into two halves, say, okay, now let's have all those with, let's say, intelligence below 100 and all those with more than 100. So you split the group into two and let's say the medium is equal to the mean, which is with intelligence, which is probably the case. You just compare two groups, those below 100 and those above 100. Medium split, you split with a median, yeah? And now you look at, you look at the average with regards to the dependent variable, which is the salary. You compare the salary of those below 100 with the salary of those above 100. This is a medium split. Sometimes that's reasonable. A more advanced version would be the quartile split. So you split the group into four equal groups and then you look at the average of the dependent variable. That's also something that you could do. Clustering is a kind of split. You just, it just clustered the groups into certain categories. So let's say your independent variable is a nominal scale. You have independent variables from which country are you? Which country were you born? And you have 50 countries in your group. You might cluster these countries on which basis you ever do this. So you do not compare all the, what did I say, 40, 50 countries. Yeah, you make groups maybe. So maybe that could make sense, right? So that really could help. And the funny thing is that I've shown you this correlation.
I've shown you this comparison of mean, I've shown you a contingency table. Very often that's the result of your study. So sometimes you work weeks on a study and the outcome is just simply this little graph. And I know that students very often have the tendency to analyze like hell, they squeeze out whatever it's in the data. No, you don't. You don't. It's just two numbers, two averages that you compare control, group, experimental, group. That's it. So it's very often not so much. That's it. Just a little. So what I should add here now is what you typically do is, but I don't want to do with this now is that when you look at the relationship between two variables, and you find differences in mean, for instance, between, let's say, an experimental group and a control group, or whatever. The first thing you do when you see this is you celebrate. Say, yeah, there's a difference. And maybe it supports your hypothesis. Okay, congratulations. But now comes a critical question. And the question is, this difference that you found is effect that you found is that could it be that this effect is just random? Could this be? Maybe this is a random effect. If this question comes up, the question of course is, how big is the probability that this effect that you found can happen on a random basis? Why is the probability? The probability is always given, right? But what you want is that the probability is low, right? So you have to do a test. I don't want to explain it here, because that goes much further, but you should know that this exists. And for every scientific paper, you have to do this test now. You have to calculate to what extent the effect you found could have occurred just randomly. And if you calculate, if you do an F test or a T test, if you test this probability of error, if the probability of error is lower than 5%, then we name it significant. If it's lower than 1%, we name it very significant. So this is something you always add. Always. Yeah. The test statistics, F test, T test, as I said, I don't want to go deep into this, but it was worth to mention it at this point. Okay. So let's leave it to this was rather short episode. Thanks for listening and watching, and see you next time. (upbeat music)
Podcast Summary
Key Points:
Bivariate statistics examines the relationship between two variables, with correlation (Pearson's r) measuring linear relations from -1 to +
Scatterplots illustrate correlations, contingency tables show frequencies for categorical data, and stacked bar charts or box plots visualize group comparisons.
Examples include studies on gender and parking speed, alcohol consumption and reaction time, and fashion awareness across study programs.
Data measurement levels (nominal, ordinal, interval/ratio) determine appropriate analysis methods: mean comparisons for interval dependent variables, correlation for both interval variables, and contingency tables for any level.
Median splits, quartile splits, and clustering can simplify continuous independent variables for analysis.
Statistical significance tests (e.g., t-test, F-test) are essential to assess if observed effects are likely random, with p-values below 0.05 considered significant.
Summary:
This episode focuses on bivariate statistics, which analyzes the relationship between two variables. The core concept is correlation, typically Pearson's r, measuring linear association from -1 (perfect negative) to +1 (perfect positive), with 0 indicating no linear relationship. Scatterplots visually depict correlations, while contingency tables and stacked bar charts summarize frequencies for categorical data.
Box plots compare group distributions, showing medians, quartiles, and extremes. The lecture emphasizes that data measurement levels (nominal, ordinal, interval/ratio) dictate suitable methods: mean comparisons for interval dependent variables, correlations for both interval variables, and contingency tables for any type. Examples include a study on alcohol consumption increasing reaction time (shown with a line graph) and fashion awareness differences among study programs (shown with bars, as the independent variable is nominal).
Practical techniques like median splits or clustering can simplify continuous independent variables. Crucially, researchers must test whether observed effects are statistically significant, typically using t-tests or F-tests to calculate the probability that results occurred by chance. 05 indicates significance.
The episode concludes by reminding students that even weeks of work may yield simple graphs or mean comparisons, and significance testing is mandatory for scientific rigor.
FAQs
Bivariate statistics is about the relation between two variables, such as the correlation between them.
The Pearson correlation measures the linear relation between two variables, ranging from -1 to +1. +1 means a perfect positive linear relation, 0 means no relation, and -1 means a perfect negative linear relation.
A scatterplot illustrates the relation between two variables by plotting each subject's values on two axes, showing the distribution of data points.
A contingency table displays frequencies of two variables in categories, useful for nominal or ordinal data to show relationships simply.
Use a line graph when the independent variable is at least ordinal (has a meaningful order), and a bar chart when it is nominal (categories have no order).
A median split divides a continuous independent variable into two groups (e.g., below and above the median) to compare averages of the dependent variable.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.