This lecture introduces descriptive statistics, beginning with the raw data matrix—a fundamental structure where columns represent variables and rows represent subjects. The first row contains short variable names. Coding converts categorical responses, like "yes" or "no," into numbers for analysis. Using a fictitious employee survey question ("Would you recommend a friend to work here?"), the lecture explains absolute frequencies (counts per category) and relative frequencies (percentages). Cumulative frequencies add counts from lower categories upward. The mean (average) reduces data to a single central tendency value, while standard deviation measures dispersion. In a normal curve, about 68% of values fall within one standard deviation of the mean. The median splits data into two halves and is resistant to outliers, unlike the mean. For example, at a student party, a low standard deviation indicates uniform age, while at a wedding, a high standard deviation shows age diversity. Visualizations like histograms and pie charts present frequencies. Finally, data preparation is critical: exclude non-valid values (e.g., impossible ages) and outliers (e.g., a national champion in a youth golf study), as outliers can skew correlations and overall results. The lecture emphasizes using the median for skewed distributions, such as income, where extreme values distort the mean.
[Music] Hello and welcome, I am Armin Trust, Professor for Organizational Behavior at the Futtwangen University in Germany. And this is my course on Social Research Methods. [Music] So, welcome everybody. Today we're going to start with statistics. The most simple statistic you can imagine. Descriptive statistic, univariate statistic. It's basically the statistic we use to describe one single variable. Or to be more precise, the values of one single variable. Let me start with how things always start. When you analyze data statistically, you always, always, always have a raw data matrix. A raw data matrix. That's what you have. And there is a convention. It's a worldwide convention, I would say, that raw data matrix have a specific shape. And you better learn right in the beginning to stick to this shape. Because also most statistical software packages use this kind of shape of raw data matrix. And how does that look like? A raw data matrix has, as all matrix, they have columns and they have lines. And what you always find is that you're different variables. They are in the columns. Every column is one variable. And the different subjects, they are represented by the different lines. And there is one thing more. The first line in your matrix, the very first line, is not something like a heading or an explanation or whatever. Leave this all out. Your first line in your raw data matrix is the short names of your variables. And put these names short. You put not your original question or something like this into this. You just give the variable a short name. In old statistical software packages like SPS, S for instance, these short names were supposed not to be long than eight digits. Eight characters. So keep it short. So you see an example of a very, very simple raw data matrix in this picture here. Okay. Very often when you write a thesis or whatever you have to also submit your raw data matrix. And also in the scientific community, you must be able to deliver your raw data matrix. Along with a description of what these different variables and what these different numbers mean, the coding, for instance. Like we were talking about in the survey specification, we were talking about when we were talking about survey design. Okay. These two things. The simple raw data matrix plus an explanation of what does that mean? What is in the raw data matrix? That's the starting point. So you load that entire thing up into any kind of software that you use for statistical analysis. For a simple statistical analysis, you might use Excel. Okay. I mean, I know many social scientists are hardcore scientists. Now you don't use Excel. That's not a scientific program. Why not? I mean, come on for some basic statistics. You can do this. It's okay. For some advanced statistics, multi-variant statistical analysis, of course, or you better use statistical. Software, statistic software. Okay. So now let's look into descriptive statistics. And I show you here an example. It's a very simple example. And based on the example we already can explain and understand most statistical measures. So here is a fictitious result on a fictitious question that was asked in a fictitious company in a fictitious employees survey. And the question was, would you recommend a friend to work in our company? Would you recommend a friend to work in our company? It's a classical question. It's a question that we very often use to understand the commitment of people to their employer. Okay. You sometimes have the same question in marketing. Would you recommend a friend to buy this product? Okay. And the answer to this question was, yes, mainly yes, partly, partly, mainly no, no. So you can imagine this. You see this question. And then as a respondent, you have to make a choice from yes to no. Okay. So let's say you do this with many people. Okay. So you get one column and you raw data matrix, which is filled with numbers between whatever. So first thing is we code this answers because in the raw data matrix, you don't have the real answers. You don't have in your raw data matrix, you don't have a yes or mainly yes. Because with yes, mainly yes, and so on, you cannot do any statistics. You want to have numbers. Statistics based on numbers. So you have a number. You have to translate these different categories into a number. This is what we name also coding. So let's say, it's arbitrary. You say, okay, let's say yes is one. And mainly yes is a two. And partly, partly, it's a three. And so on until to no, no is a five. Okay. So that we can work with numbers and not with words. Okay. It's the basic for doing statistics. Now we can analyze some frequencies and frequencies are already statistical measure. So statistical concept. So frequencies is that you simply indicate how many people in our sample have responded to specific categories. So in our example, five people said yes, 23 people said mainly yes and so on. And two people said no. These are frequencies. In this case, we name it absolute frequencies. Absolute. Because that really tells in absolute terms how many people have responded to that category. Okay. When we sum up all these absolute frequencies, we get a total in this case of 55, which tells us 55 people responded to this question. Now we can simply for every category, we can simply divide the absolute frequency by the total number of respondents. So that we get a relative frequency. Okay. So relative frequency of the category one, yes, is 0.09. This is simply five divided by 55.09. So whenever we use relative frequencies, we do we do we do easier when we translate this into percentage. By simply multiplying the relative frequency by 100 and then we have percentage. So 0.09 turns into a 9%. Okay. That's simple. So these are relative frequency opposed to absolute frequency. Okay. That's statistics. So when you add up all the relative frequencies in percentage, you come up, you get 100. You all know this, right? Now there is another concept in descriptive statistics that we name cumulative frequency. What is that? Well, the first category was five. Okay. Absolute frequency. Now with the second category, we can add the first category plus the second. Now we get a 28, which is 5 plus 23. In the third partly partly, we simply add the 5 plus the 23 plus the 17 of partly partly and we end up with 45. This is cumulative. So is the category, which is it all about, plus all those categories underneath. So in the last category, in this case, the null, that is probably 55, which is the total number of respondents. You find this concept of cumulative frequency also these days. I produce this series during the Corona crisis. When you have the frequencies of let's say all deaths due to COVID-19 every day. So you have the frequency for every day and then very often you have the cumulative.
You cumulative frequency in that particular case, the number of people who died today plus all those people that already have died. So you could also say the cumulative frequency is the number of death up to now. Okay? That's cumulative. Okay? And then of course you can do the same with the relative frequencies, of course. Now when you multiply the coding of every category by the absolute frequency and you sum it all up, it's the same like when you sum up all the values that you have. So 5 times 1, 23 times 2, 17 times 3, you sum it all up, you end up with 144. When you have the 144 and you divide this by the total number of respondents, guess what you receive? You receive something that we named the mean, the average, right? Which in this case is 2.62. What does that tell us? That's already a cool number. I mean, it's an average. That reduces 55 numbers into 1 and tells you what we name in statistics, the central tendency, the central tendencies, tendency of all the numbers. So now you don't have to look at all the numbers. You just look at this number and see, ah, it's 2.62. 2.62 when we look at the coding. Okay, that's somewhere between the 2 and the 3, somewhere between the mainly yes and a partly partly with a slight tendency towards the partly partly. Because it's more towards the 3 than towards the 2, it's 2.62. 2.5 would be exactly in between. It's 2.6. So that looks a little bit more to the partly partly. Okay, that's already a, a, a information that tells us something and it's, it's really reducing complexity. I would say. So another concept in descriptive statistic is what we name the standard deviation. And I know everything I was sharing with you so far was simple and you probably already have heard about the frequency. I mean, come on, you know what frequencies are, you also know what relative frequencies are. You might not use this term relative frequency in your daily life, but you know you know it and it was familiar to you. You know what that is. You also know what an average is, of course. But what is the standard deviation? In this particular example, the standard deviation is 0.97. We can simply say 1. What does that mean? Standard deviation. So let's, let's look at standard deviation because that's a concept in statistic that is about the variation. It's about the dispersion. It's about to what extent to the values spread somehow around the central tendency, the average. So when we look at an example, for instance, intelligence, everything in nature, almost everything in nature is distributed in a, you know, we say normal, normal way. That's why we get this kind of normal curve, it is bell shaped curve. Right? So when we look at intelligence, for instance, when you measure the intelligence of all people in the world, you get an average of 100. Intelligence tests are designed this way. They are standards that stand a dice in that way so that the overall average is always 100. Okay? So if you have an intelligence quits end of 100, you are exactly in the middle. So here is the other measure which is standard deviation, which in most intelligence tests is 15. What does that mean? Look at the bell curve. And as I told you, the average is 100. Now we subtract 15 from the average, which is 85. Okay? And then we add 15 to 100, which is 115. So we have a range between 85 and 115. Okay? Simple. Okay? We now can say that when the curve is perfectly normally distributed, and in this range between average minus one standard deviation to average plus one standard deviation, you find 68% of all cases. Got it? 68% of all cases are in the range between average minus one standard deviation, average plus one standard deviation. If you got this, you understood the standard deviation. It's easy. 68%. Why 68? Forget it. You don't need to understand, I would say, most of you. 68. That's pretty close to 66. 66 sounds easier. You can better remember 66. Two-third. That's fine. Two-third are exactly in that range. Okay? So let's have another example. There's a student party. Okay? And on a student party, they are, let's say, 50 people. I don't know. It doesn't matter how many. And now you make a little survey. You ask the people, how old are you? Okay? And you find out that the average age on the student party is 20. Okay? 20. Nice. Cool. So, some days later, you go to a wedding, and you do the same survey again. And the people tell you, "Whoop, whoop, whoop." And again, the average is 20. Now you can see, the average age in both parties is equal. Must have been the same parties. Right? Feels pretty similar. Same age in both parties. 20. No, it's not. Because let's assume, on a student party, a reasonable standard deviation might be something like two. Right? Two would mean two-third are between 18, 20 minus two, and 22, 20 plus two. So standard deviation of two sounds pretty reasonable. Right? While on a wedding, the standard deviation might be 15. Yeah? Because you have two-third in the range between five, 20 minus 15 to 35, 20 plus 15. Because you have all the kids and your aunt and your parents and your uncles and, "Oh, it's kids running around." Okay? Some are even younger than five, and some are even older than 35, of course. Yeah? But two-third are somewhere in this middle. So you see, we can also say that the average tells you the age, the central tendency of the age, right? In this regard, in this example. But the standard deviation tells you, to what extent, this average really reflects the distribution of the different subjects in the group. And the bigger the standard deviation, the more the broader the distribution, the more the subjects spread it around the central tendency. Okay? So that's the standard deviation. And the thing is that in practice, in daily life and in public news media, we rarely talk about standard deviation. And you really, you will not hear something in the CNN in the news that, "Oh, the standard deviation is this, this is it." Everybody was there. That's them. But it's the same, but the same, the same, the same. The different kinds of averages, but because there is a difference between something which is very important, which is that the mean versus the median, the median. The median is also a very important term, I would say, very important. What is the median? Maybe I can explain you the median by telling you a little story. Maybe you are students, you listen to, maybe 10 years after your graduation you meet again. So let's say you are male friends, and tell you what male friends do when they meet 10 years after graduation. They will make this game of who has the biggest car here now, who has the biggest house, who has the prettiest girlfriend and so on. It's a man's game. And so let's be very stereotypic and say, four friends they meet after graduation. So one of the four is late, so the three of them already meet at the bar where they always met when they were still in the air, some beer. Okay, the game starts and they're like, "Hey, what do you earn?" And the first says, "Wow, 100,000 in a year, and the other says, "Oh, wow, 100,000, I'm a two, 100,000." I'll talk to you later.
We all are one thousand. Wow, that's pretty cool. And now comes the fourth friend into the bar He took a while because he had to find a parking Lot for for his Ferrari You also already could listen him from the distance and now this friend comes in having a beer as you how much do you earn one million? Ah one million Hmm, wow So what's the average 100 plus 100 plus 100 thousand plus a million is something around One million three hundred thousand divided by four is I have to calculate three hundred thousand somethings, okay Then you tell the bucket hey we on average earn more than three hundred thousand No, you don't most don't that wouldn't miss me a misleading information, right? absolutely so The average really changes when the rich come in But the median doesn't what is the median? The median is the is a is a simple concept in statistic you look at all the cases You look at all the values to put them into a role from the lowest to the highest for instance with regards to Compensation or whatever right age whatever you put them in a row And then you try to split the group into two halves and that value that splits the group into two half. That's the median Okay, that's the median So the median is very useful when you have a not normally distributed curve, right? So the median is S.S. or the point that splits a group into two halves There are 50 percent that are more higher than the median and they're 50 percent lower than the median So when you have a curve which is more flat to the to the positive end and when you have some some extremes As you have for instance with celery when you look at the celery distribution in society It's not normally distributed. You always have very few very very rich people That's why when we look at the descriptive statistic about celery you never use the average because Those incredibly rich guys that will pull the average towards the positive end So we better use the mean the mean does not pull to the positive end That stays there Okay, so when you look at these two curves With a normal distributed curve the mean is equal the average because simply the middle Yeah Okay It's because this curve is symmetrical Right But then it's not You can say the mean is pulled more towards the extreme ends while the mean stays more On the on the left side and then with regards to our example here Okay So when you let's let's have let's have an example I used to ask this next exam in human resource management also Let's assume you have a company of 100 people right and the people they on average earn 50,000 euro dollar of Year on average, let's say 50,000 and now you're higher somebody who earns incredibly much One person you hire one person who owns one million in a year. How does that change the average? It changes the average right How does it change the median not at all or just to minimum Right, so the median is a good statistical concept Uh That is robust Against outliers That's why we use it. Okay, it's important Yeah, so when you add an outlier to a distribution the mean will be affected And the median won't Okay, so what we also do very often in descriptive statistic is that we try to to illustrate To present our frequencies in certain crafts and For nice example for instance is the the histogram the histogram you see one here You you might use their relative or the absolute frequency We'll look the same and you have bars. That's a bar chart A histogram a nice way when you have relative frequencies you also can use A ring a ring chart for instance Pie chart or something But did you only use ring charts or pie charts when when the frequencies really sum up to 200% so relative frequencies With absolute frequencies or ring chart would not make any sense, right? This is a good way of doing things and now Uh at the end of this episode I would like to add something which is essential and I It only can explain it now because we already were talking about it About some some concepts we need now when you do some when you do your analysis Yeah, and you have your raw data matrix Which is the starting point. I mean that's your material that you use for your statistical analysis The first thing is that you have to prepare your data really and Some fundamental steps for preparing your data Is that you exclude non-valid values that's that's one one point so I would like to share with you now a little bit the checklist Yeah, that you can work on when you prepare your data take out non-valid values Why is that I mean some values that you find are really not valid Yeah, really not it's a It do not talk about missing value. I talk about non-valid values So a nice example is for instance when you ask people how old are you And somebody says 33,678 I say yeah, yeah, yeah, my friend you know that's a that's a joke sorry you take this out Yeah, there are all sorts of data where you instantly see it's not valid sorry It's it's sometimes difficult, but just I don't want to go too deep into this you look at the data and see does that Could that be And the other thing is outliers Sometimes you really have outlasts when let's assume you do a study about I don't know Sports okay sports you look at the Performance of let's say some some youth yeah in golf. Yeah, you look you do a study about their performance in golf And now you have somebody in your sample who was a national champion It's incredibly good So much better than all the others take this guy out Because when you leave outliers in your sample it could be that this one outlier will Determine the overall results We find this very often with correlation when you have an outlier in a correlation that can That can turn the results upside down So take out outliers. So outlier is defined statistically statistical term and here's a concept um Let I have to explain It's the interquartile range before you try to understand this I have to understand what's a quartile What is a quartile if you have understand the median then you also understand the quartile what is a quartile? A qu- you have three quartiles, no the first the second and the third and the second is equal to the median And the three quartiles They split the sample into four equal groups You know the median did the split the group into two halves the quartile into four equal four equal uh groups So the first quartile is the value that splits the group uh into 25% and the upper 75 that's the first quartile The second quartile equal to the median 50 percent 50 percent and the third quartile is 75 25 Okay, this is the quartile now. What is the interquartile range? Interquartile range is the range between the first quartile and the third quartile Okay So And if a value is higher than the third quartile um More than 1.5 times the interquartile range higher than the third quartile than it's an outlier Okay, so you don't have to understand this by heart But what you take home is an outlier is defined in statistical terms. It's very very very very extreme value. You better take out And there's something else that you very often might do Before you start analyzing your data. It's what we name a set transformation set transformation um, you might have Data of all different sorts some data they rank from zero to ten the other rank from zero to one hundred the other from 1 to 5 or whatever. So you have different sorts of data.
you want to compare this different data, this different variables. And you want to compare it in a way they, every variable has zero as an average and one as a standard deviation. So you standardize all the variables so that they all have the same average and the same standard deviation. Sometimes that's here reasonably, especially when you add different values up for calculating an overall index for instance. Right? So set transformation, that's something that you might do. It's a very simple formula. From the value you have, do you subtract the average and divide this by the standard deviation? Then you receive a mean of zero and a standard deviation of one. Okay, that's very simple. Also sometimes you need to change the polarity of items. We were talking about polarities when we were talking about survey design. When you have a question that goes from yes to no, sometimes yes is the positive result and sometimes no is the positive result. So whatever you measure, you make sure that all your questions, they look all in the same direction so that yes always means maybe good leadership quality and no means bad leadership quality. But sometimes it's upside down depending on how you ask. So you have to make sure that all items have the same polarity. For some reason that makes sense. Sometimes you want to calculate an index. Index is an overall total score. When you have multiple items, multiple questions that you sum up to an overall score as you would do for any kind of test for instance. We were talking about this when we were talking about testing. Right? If you don't remember, go back to this particular episode. This is also something that you very often do when you prepare your data for analysis. You sum up the values of different variables and have an overall value that in the end you use for your final analysis. Okay? So I would leave it to this now. That was descriptive statistics. And you know, there's much more about descriptive statistics. But these are the things that you really take home. Okay? Now you have understood what is the mean? What is the median? You understood what is the standard deviation? The various kinds of frequencies, quartile, all these things. So that should be enough for the moment. Okay? So in the next episode we will look at a B-variate statistic is about to analyze the relation between two variables. I mean up to now we're just talking about the analysis of one single variable, central tendency and dispersion. In the next time we talk about how you can analyze the relation between two variables. Okay? So cool. Thank you for watching and listening and see you next time. [Music]
Podcast Summary
Key Points:
Raw data matrices have variables in columns and subjects in rows, with short variable names in the first row.
Descriptive statistics start with univariate analysis, including frequencies (absolute and relative), cumulative frequencies, mean, and standard deviation.
Coding translates categorical answers into numbers for statistical analysis.
Standard deviation measures data dispersion around the mean; in a normal distribution, about 68% of cases fall within one standard deviation.
The median is robust against outliers, unlike the mean, making it preferable for skewed distributions like income.
Data preparation involves removing non-valid values and outliers, which can distort results.
Summary:
This lecture introduces descriptive statistics, beginning with the raw data matrix—a fundamental structure where columns represent variables and rows represent subjects. The first row contains short variable names. Coding converts categorical responses, like "yes" or "no," into numbers for analysis.
"), the lecture explains absolute frequencies (counts per category) and relative frequencies (percentages). Cumulative frequencies add counts from lower categories upward. The mean (average) reduces data to a single central tendency value, while standard deviation measures dispersion.
In a normal curve, about 68% of values fall within one standard deviation of the mean. The median splits data into two halves and is resistant to outliers, unlike the mean. For example, at a student party, a low standard deviation indicates uniform age, while at a wedding, a high standard deviation shows age diversity.
Visualizations like histograms and pie charts present frequencies. , a national champion in a youth golf study), as outliers can skew correlations and overall results. The lecture emphasizes using the median for skewed distributions, such as income, where extreme values distort the mean.
FAQs
A raw data matrix has variables in columns and subjects in rows. The first line contains short variable names (e.g., up to eight characters), not headings or explanations.
Responses like 'yes' or 'mainly yes' are translated into numbers, e.g., yes=1, mainly yes=2, etc., because statistics require numerical data.
Absolute frequency is the count of responses in a category. Relative frequency is that count divided by the total, often expressed as a percentage.
Cumulative frequency adds the frequencies of all categories up to a given one. For example, it shows the total number of deaths up to today in a pandemic.
The mean is the average, calculated by summing all values and dividing by the total number of respondents. It measures central tendency.
Standard deviation measures how spread out values are around the mean. For a normal distribution, about 68% of cases fall within one standard deviation of the mean.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.