Go back

167  |  Visualization and Statistics with Andrew Gelman and Jessica Hullman

from Data Stories

49m 26s

167  |  Visualization and Statistics with Andrew Gelman and Jessica Hullman

Data visualization is not merely a tool for communication but a deeply structured practice requiring personal theory and intentional design. Drawing from Andrew Gelman’s blog and Jessica Holman’s research, the conversation emphasizes that style and structure in visualizations are essential, not optional—just as theory is vital in writing or scientific reasoning. Visualizations function as model checks, testing expectations and revealing anomalies that challenge or refine our mental models of data. This process parallels scientific exploration, where surprise and deviation are critical for progress, even if most analysis relies on representative, expected outcomes. The episode highlights a gap in current tools: they often lack mechanisms to support users in expressing or evaluating prior beliefs, leading to implicit assumptions. A more effective future for data analysis lies in guided, scaffolded visualization tools that help users build mental models, test hypotheses, and assess model fit through visual feedback. These tools must balance openness with structure, allowing both novice and experienced users to explore data meaningfully while avoiding over-reliance on superficial patterns. Ultimately, the success of data exploration depends not just on technical skill, but on cognitive frameworks that enable users to distinguish signal from noise, question their assumptions, and build robust, interpretable insights.

Transcription

8941 Words, 48210 Characters

English
People don't always understand the importance of style or the necessity of style, but it's not a luxury. It's more of an organizing principle. Hi everyone, welcome to a new episode of Data Stories. My name is Erika Bertini and I am a professor at Northeastern University in Boston where I teach and do research in data visualization. Right, and I'm Mojda Fana. I'm an independent designer of data visualizations. And in fact, I work as a self-employed truth and beauty operator out of my office here in the countryside in the beautiful north of Germany. Exactly. And on this podcast, we talk about data visualization, analysis, and more generally, the role data plays in our lives. And usually we do that together with a guest we invite on the show. But before we start, just a quick note, our podcast is listener supported, so there are no ads. But that also means if you do enjoy the show, you might consider supporting us. You can do that with recurring payments on patreon.com/datastories or you can also send us one-time donations on PayPal.me/datastories. Exactly. OK, so I think we can get started with the main topic and guests for the show today. So today, we have two people on the show to talk about, I think, a really relevant topic. I think, in general, what is the relationship between statistics and data visualization? And to talk about that, we have Andrew Gelman and Jessica Holman. I, Andrew and Jessica, welcome to the show. Hello. So as usual, we start by asking our guests to introduce themselves. So maybe, Andrew, you want to go first and give a brief introduction? I teach statistics and political science at Columbia University. Jessica. Hi, I'm a professor, associate professor of computer science at Northwestern University. I do research on various topics related to how people draw inferences from data, usually from interfaces. So I care about things like visualization. OK, so I thought we would start our conversation by starting from the blog that I believe Andrew started several years ago. It's called statistical modeling, causal inference, and social science. I think it's a really influential blog and I remember reading the blog since many years. And it's a really interesting community of people. And there's a lot of interesting discussions about statistics, but also political science and science in general. And also a lot about data visualization and I believe Jessica joined the team recently. So we have seen even more data visualization conversations in that space. So Andrew, I was thinking maybe you could give us a little bit of overview of what the blog is about. Maybe if you want, even to say how it started and what are the main topics there. In 2004 I was working with a postdoc Samantha Cook and we had an idea of setting up a blog in a wiki to help us communicate with each other. The idea of putting stuff on a blog was that then other people could see things too and we could get input from other people. Then the wiki was supposed to be where we put our various ideas. The wiki got hacked and we had to take it down, but the blog was useful. I learned, well, it's hard for most people to write stuff on a blog and sometimes I would suggest that students are postdocs write a post and they would find it too difficult to task. They would find it too like pressureful. So it ends up mostly being me, but then various other people, maybe about 15 other people such, including Jessica, have had stuff to say, so I asked them if they could write for it too. So it's a way of getting, having conversations. It is difficult to write for your blog, I was just going to say, Andrew, maybe it's hard for you to understand, but I think you've established quite a record with it. I think a lot of people tune in to hear what you're going to say. So it is, I understand why other people would be like, oh my God, it's too much pressure because I had to just get over it and be like, I don't care if they're going to compare me to his post and they're not going to be as good, but it is chuggy. And your posts are definitely better than the average post of mine, so. Yeah. I think what is really interesting there is that every time I look at a post, there's such an interesting series of comments below. You seem to have a very active community around it and it's always very thoughtful. Yeah, and I don't know if you need any special moderation, but also normally the comments are pretty interesting and nothing too bad. There's no crazy people writing crazy stuff as far as I can tell. Yeah, it's not so interesting for the crazy people, I guess. Yeah. But I wanted to pick up on something that you said earlier, not about the blog, but about statistical graphics, because I'm a user of graphs ever since I was a physics student and graph data and graph curves and models and so forth. And I think over the years, it struck me that I think everybody needs to have their own theory of statistical graphics. In the same way as if you're writing, you need to have a theory of writing or if you're drawing, you can't just say I'm going to draw what I see. You have to have a kind of approach, its goals, or if you're making music. You can't just say, hey, let's bunch of us. Let's form a musical group and do music, right? You have to have music that you want to do. You have a certain style. It doesn't mean your style is better than everybody else's, but you need that. And I think that for quantitative things, people don't always understand the importance of style or the necessity of style. Like, style is seen as a kind of luxury, but it's not a luxury. And that's also clear with writing that if you're writing, except for perhaps the most functional writing like the instructions for how to operate your microwave oven or something like that, you need to have a style because if something is boring to read, then no one will read it. If you teach as a teacher, if you teach a class and write a two-page document for the students to read, they won't read the two-page document. And so it's really it's needed. So I, and I think the flip side of that is that on the other hand, then you have people saying, oh, someone's a designer as if there can't be any connection to science because design is supposed to be some like separate thing, but they're one can draw the analogy of something like building bridges or that these things are designed, but they still have to work. Yeah, I really like the way you're describing this because in a way, it reminds me of something that maybe I was, I was hoping to discuss that to me, it reminds me of the role of theory, right? The fact that if you don't have a theory, theory doesn't have to be necessarily perfect or always super predictive, but it's a way to organize your thoughts so that you can think systematically about something. So I don't know if this rings any bells on your side, but the way you describe style to me reminds me how important it is to have theory and also a theory of graphics in this specific instance. Well, it's like they say in chess, having a plan won't win you the game because presumably you're playing against someone else with a plan too when you're not both going to win, but if you don't have a plan then you'll lose, you won't be able to move forward. And part of having a plan is recognizing being aware of that you have a plan, being aware of what the plan is, and then when things go wrong, you can change things. So actually it's like putting you as a scientist, like putting your marker down, I'm not really into like betting, it's not like I would say you literally have to bet money on things, but like conceptually setting it down and saying this is the model I'm going with is very valuable even though or I should say especially though we know that that model is going to be wrong and it's going to fall apart at some point. But the more explicit you can have that model or style or system, the more that you can then know when to work on improving it or abandoning it. We actually wrote a paper about some of this as it applies to sort of theories of visualization for exploratory analysis where I think that's a place in visualization where we're building all these tools to help people do interactive data analysis through visualization. But we don't always like if you ask us the people developing these tools or the researchers are in the area like what are what are sort of guiding principles are. I think it's often while we just want to let people explore data as easily as possible. But I think there you can easily run into places where you just like you have no theory to tell you how to design something in sort of a better way versus a worse way like you just don't know. We wrote a paper for the Harvard Data Science Review like about a year or so ago. where we sort of talked about this as applies to exploratory visual analysis and we argue that even if it's a bad theory like Enrico was saying or both of you guys were saying that it still can be useful because you need to know sort of how you were wrong and if you never state what you're going for what you think the objective is how do you know when you were wrong. Yeah, that's an area of data visualization that I really love and I think people tend to talk less about it. I think there is more of a I think the general idea with visualization is that it's a tool for communication but there is less I would say there is less discussions about how to use it for exploration and by the way the word exploration itself is so contentious in a way so yeah. It's funny. I thought it's the other way around that everybody talks really exploratory. Yeah, I was kind of thinking in academia. Yeah, I mean nobody like really gets what communication means. I think I think that one problem is communication is often viewed as being a kind of unidirectional thing so like those same things like scientist should learn how to tell stories because people think in terms of stories. And I hate that kind of attitude. I mean sure people think in terms of stories but but this attitude that oh you're the scientist you already know the answers but now you have to convey it to people so you have to learn how to be a storyteller and and have a good bedside manner right like it's all connected with like you don't want to be a jerk right like you would be it's like narrative medicine and all this stuff and it's like the what people are actually doing there is great but the idea the framing that it's what you're going to say. It's all about how to communicate truths to people. I think it's misleading. I think it's more accurate to say that we're people too and we learned from stories and this is something that my colleagues and I've been thinking a lot about over the years like what makes us believe things and and often work convinced by stories and I. We I character the effective stories as being anomalous and immutable and by anomalous it's a story is a surprise so the convincing stories have some twist in them something unexpected even if you think about a scientific method I didn't think this method would work and and then it did or or or whatever and then it's immutable in the sense that a good story is grounded in reality and like if. If you kick it your foot hurts right like as it as in the famous Boswell story and so you have like. But so we we kind of learn from or we we learn from these stories which are a reality and maybe the term story isn't the best in that sense because I'm not talking about stories that are made up and I'm talking about true stories but but we learn from these but it's it's kind of. Necessary that the stories have this grounding in truth so that they can disprove our theories and there's a sense in which a good story like it like if you were to take all of the things that people have said in your podcast right so maybe. You've done however made podcast in each podcast an average maybe there are five stories like when someone tells you hey let me tell you a story about that and if you look at those stories some of them are going to be just made up I mean we have this horrible examples of stuff where people just say like there's a famous example from a few years ago in a book or somebody said that a certain. It was like a certain data problem caused 70 deaths a year in a small town and it was like how the hell like it made no sense right so they just made it up but I think the stories that are good like if. You could in theory track them down and I think they would have this characteristic that they disproven implicit model of the world right I have to do with news newsworthiness as well like in journalism you have certain newsworthiness criteria. Exactly dog bites man by it's dog but here's the point is that when something is surprising it's surprising relative to an expectation and so that model of the world is is that so when we talk about discovery and surprise there are theories implicitly they're already. Yes this is yeah I was going to tie this back to the exploratory analysis thing as well like I think one of I mean I think we talk about exploratory visual analysis a lot and is probably more than communication but. You know it's there's always this role of expectations and what you're bringing like what you're expecting to see both in communication and exploratory visual. Analysis and I think at least in visualization research that something that has we've always sort of background it you know we've always acted as though like the data speaks for itself when it obviously does not so yeah so I think yeah that's all like theories of how you know. How visualizations act as model checker one thing that and I have like in the same paper I mentioned already he had he had been thinking about this years ago. As a way to sort of think about the role of graphics and exploratory data analysis and how in a sense you can tie it to confirmatory data analysis through this idea that a good graph is helping us check a model some implicit model often sometimes an explicit model like in confirmatory data analysis but there's always. You know some expectation and the graph tells us you know how much the data deviate from that so when you say model check here do you mean like checking the model that you are stored in your head like a mental model yeah I mean I think yeah so in in exploratory visual analysis I mean I think. You know there's always you know some background assumption potentially in a lot of times and this is stuff that Andrew had spoken about back in 2003 in a paper so you can cut me off whenever Andrew and tell it yourself but basically you know like you can think of you know a visualization is giving you almost like a test statistic or a vector of test statistics in a hypothesis testing type framework if you want to look at it that way where you know you have something that you're you're trying to check for like you make a graph in order to check like. You know like how well does my data conform to some expectation and sometimes that's really explicit and like built into the visualization like I want to you know in like a sort of model fitting or preliminary model fitting kind of stage of a workflow you're looking at things like maybe residuals and you know that exactly how to read the chart because it's sort of built into it that like if your data deviates from the. Expectations that you want your residuals to have then or to fulfill then it'll be obvious because you'll see like deviations from symmetry in the plot or even a scatter plot you know like often if you have a library scatter plot like the most common sort of built in you know thing that you're checking against is sort of like a straight diagonal line representing kind of perfect linear association so. The idea that back in 2003 that Andrew started talking about and other people in statistical graphics have also gotten into like Andrea Spugia Diane cook and others you know there's this work in sort of graphical statistical inference that gets into this idea of. You know like how visualizations can function sort of as model checks but then within that you know people gone in different directions where Andrew's originally original formula formulation I believe was sort of more in a Bayesian direction where. It's not that we're testing some hypothesis and we just want sort of you know our p value or kind of yes no answer but it's it's that we're you can think of the graph is kind of like. The comparison that you're doing mentally when you look at a graph is kind of akin to doing like a posterior predictive check in Bayesian stats where you're sort of imagining like you know under my expectations about the process that created my data what do I expect the data to look like. Like so what is what is sort of what would reasonable data look like under the under the predictions I want to make and how much does the data that I actually got sort of compare to what I would expect so it's almost like you know you can imagine on some level maybe this doesn't happen all the time but it's almost like when when you look at graphs you're sort of imagining. You know you know reasonable data under some some side of expectations you have and you're comparing that to what you see Andrew I don't know if you want to say anything there I think that's that's just sort of the idea I wanted to bring up well. Yeah I've more to say about that but actually let me jump to something else which is a paradox that we we learn from stories and the best stories have surprises in them like I would almost argue all stories have surprises in the sense that if there's no surprise you don't bother telling the story. So we're always using just as we use graphs to learn and to discover which means to be surprised relative to our implicit models we consume stories in order to refute various models of the world. But yet that the paradox is that how can we learn from surprising things like it seems the usual way we think about statistics is that we learn from the expected like random samples right that's like if I'm going to do a survey I don't say I found the 1000 weirdest people in America asked them their opinions about things and I really wanted to be surprised and you'll never guess what I found. No what you'll do is you want to ask a representative sample of people and like you don't really have the goal of it. being surprised. You just, you want to see that. So this was sort of bothering me actually because after I wrote the paper, the paper that my colleague and I wrote about why, how we learned from stories, I drew, directly connected to the papers that Jessica mentioned where we earlier, I had written that graphics are a form of model check. So then I argued stories are a form of model check. But then again, it, then people ask the question and how can it be that you learn from anomalies? And my, I don't know, like my, my resolution of this is that it's related to a kind of proper area in our Lakotoshian view of science or Qnian view or whatever. I guess the standard view now of science, which is that we alternate between normal science and scientific revolutions. And when we're doing normal science, we are kind of representative samples and we want to help build, we want to build theories and modify our theories. Then when we're doing revolutions, we're trying to see what's wrong with our theories. And there we're looking for counter examples. Now, I'll only say one more thing, which is that when you say this, it always sounds like revolution is the hero in normal sciences is like the loser in this game. But that's not true because the revolution only exists because, because there was a normal science allowed it and the goal of a revolution is to replace it with a new normal science. And so it's like both of these depths are important. I have a question here. Like how, how open are people really to changing their minds based on statistical information? I think that's, it's something like the last few years, maybe been an interesting research topic now. I was thinking to, like as Andrew was talking about anomalies, like, do you, like, should you have to go out explicitly searching for anomalies, like to break a theory, or if we were all sort of honest scientists, could we kind of through our own, like just seeing the data that we collect, find anomalies. Like if we're willing enough to sort of admit when our, when our mindset is not right, you know, then, then we should be seen probably anomalies a lot. Well, this is kind of related to this, like, unitary nature of consciousness thing or even the idea that we talked a bit earlier of having a plan. Like it's, it seems pretty fundamental like to, to mathematics. It's, I mean, the way cognition works in general, not just like human brains, that it seems like there needs to be this executive function and this alternation of processes. So like the same, you need a division of labor somehow. So maybe one scientist could create and refute her own theories and gather data, but maybe not all at the same time. I mean, another example is in math class way back when, when you're asked, sometimes you're asked to either prove something or come up with a counter example, and they always say you can't do both at the same time. You have to first assume it's true and try to prove it. And then if you can, stop and assume it's false and, and try to do that. So we, we do kind of use the multiple points in the system to play different roles. Yeah. I have another question. Like, looping back to the beginning. So I think you rightly explained visualizations is really a skill and a practice and there's no single right way, but it's like a highly personal, like thing, which, how you do it, right? And there could be many ways to do it, right? Is it the same for statistical practice, like for applying statistics, or is it in statistics more, for a given problem, there is a correct solution? There, of course, there are many ways of, of solving problems. I actually wrote something once about what I call the method logical attribution error, which is people attributing to their method, what's also a property of their skill. And so you'll see this with like renowned statisticians, or maybe not so renowned statisticians, also that they just think some method is inherently better. But yeah, there's always so many unwritten roles. And they might just be better at applying it or failing to apply another technique successfully, which somebody else might have done. Yeah, I'm better at some techniques than others. So you just, that's how that's how it works. So it's, there's an interaction. Yeah, yeah, yeah. It's interesting, because from the outside, so I have only statistics 101 knowledge, right? And to me, it always seemed there's this clear decision tree of if your data is shaped like that, then you need to apply a nova test or something like this, right? And we have the same for graphs. If you want to spot outliers, use a scatter plot, right? And so I was always wondering how, how hard cut these rules are. And I'm glad to hear that. It seems very similar. Actually, that the deeper you go, the less clear it is how things should be actually. I was influenced by a colleague, David Kranz, a psychologist, who he was telling me about like decision theory. And he said that the simplest version of decision theory is what you learn, like the like fun norm and Morgan Stern, you have a decision tree that you need to evaluate. So you, you compute all the things that go into it and you compute the tree. And then the next level of sophistication is to say, no, actually drawing the tree is important. And there's lots of psychological experiments where they show people the tree and it's missing a branch and like people don't realize. So like a lot of examples were the best decision of something that wasn't in the tree in the first place. And then we tried it. But then he said, that's also not enough. And so his taken on decision analysis is that you start with goals. So you have goals and resources and stakeholders and all of that. Sounds kind of soft, but it's not really softer than trees. So you basically start with the design thinking again. Well, yeah, you start with your, you start with your goals and, and then all the other things, your goals and your constraints and your resources. And then you consider ways of getting there while being open to that your goals might change and so forth. So yeah, statements like, if your data look like this, you should use this model. That's like totally, that's horrible because you really want to be starting with your goals. And, you know, and, and not in an empty way. Like, oh, yeah, my goal is to publish a paper. My goal is to get this data analysis done. Like, you know, your serious goals, whatever they are. My goal is to not get shouted at on the internet foremost. Yeah. So I get it. I have to go now. So I'll see you all later. But thanks for the opportunity for talking with you all. This is, this is always fun. Thanks so much. Wonderful. Thanks for joining us. Thanks, Andrew. See you all later. Thanks again. Bye. Okay. And now we can continue the rest of the episode with Jessica. So one question that I had going back to, let's say, the comparison between statistics and visualization, right? In my head, I'm always like, I can't say that one is better than the other, right? I don't know. So for, for instance, when we teach visualization, we show the outcomes quite bad. I think we have a huge bias there because it's almost like, take that statistics, right? This is so much better. This is how I use it. Right. Everybody uses that way. It's like, hey, we go. That's the, that's why we need this. And, but no, because I think there's, there's almost like a dance between having a lot of details so that we can maybe related to what Andrew was saying, right? The surprising elements, but you can't do, you can't reason only with surprise, right? And surprise can also overwhelm you. And you may lose the signal as you look at a lot of noise, right? So I really see that as a dance between these two things, the surprising, the particular, but also surfacing the signal. So, yeah. How do you think about that? I mean, it reminds me of, like, you know, Tuky and others who have written about like the exploratory data analysis process, where, I mean, my personal view is that in visualization, we sort of take kind of like the very initial stages, you know, where, or the very initial stages of exploratory analysis are often something, you know, like just making sure there's no massive, like chunks of missing data, figuring out what variables you have to begin with. But I would say we take this like, then next stage of like, you know, I'm just trying to see where is there some sort of, you know, signal or what looks like a pattern, you know, what kind of relationships do I think I see? We sort of in visualization, I think, think of that as exploratory analysis. So it's sort of like very open-ended, sort of clicking around to find patterns. And I think we design kind of as though that is, that is what exploratory visual analysis is all about. Tuky, for instance, though, talked about this sort of intermediate phase where, you know, you've noticed, you've sort of generated some hypotheses about like, you know, possible relationships between variables or the nature of certain distributions. And then you sort of need to know, like, you know, how much can I believe what I think I see here? And so, you know, there's also like, you know, exploratory analysis also involves things like starting to fit models to try to explain, like if I think that this, you know, set of variables seems to be predictive of some other variable that I care about, you know, I would to actually start. 15 models and looking at deviation in things like residuals, seeing how out of my models actually explain what I'm seeing. And the whole idea is that I want to build up some sort of a mental model kind of of the data generating process. I want to use stats to sort of help me figure out what can I believe in terms of what signals are actually there and what is just sort of maybe not actually going to hold up when I inspect it more closely. And so yeah, I think there's this weird stage where or this weird sort of way in which graphs are used, not just to show us things that maybe we didn't expect to see or we didn't expect to see, but also to give us some information about how much we should believe those things. And I think that's where it's sort of kind of ambiguous. So actually from Andrew's blog, I learned about this informal term someone used called the Anthropic Principle of Statistics, which is like there are certain problems where you would use statistics. Like if your data, if the signal is so huge relative to the noise, like you don't really need to be running stats on that, like you can just sort of see it. And maybe you just make a graph and it's like obvious. If your data, if the noise is very large relative to the signal, then it's sort of hopeless. And you could do stats, but you're still kind of dealing with too much noise. And so it's like statistics is useful for this sort of middle set of problems. And so I think, I mean, one of the questions will ask you or to, for me, is like, I think of this. We have like the problems in the middle where we can use maybe visualization separately from stats. Like what do we think is a problem or visualization is simply not going to work? When is visualization sufficient without any follow-up? And I don't think there's like true or right answers. These are things we have to sort of figure out as a field. But I think it's not. We haven't always made explicit sort of what our assumptions are. Like I think we often maybe implicitly when I look at what people write about exploratory analysis and designing for, you know, exploratory visual analysis, I think there's sort of this assumption that, you know, people can click around for patterns. And maybe though, you know, they care enough about finding the right answers. Like there's some, you know, like application or some reason why they're analyzing the data. And so we sort of trust that, you know, they will look at things enough and make enough views. And some of the views will be disaggregated enough that they can sort of get a sense of the noise. And so, you know, we don't have to worry about explicitly supporting these like signal to noise kind of judgments. We can just let people, you know, use these tools and they will figure out what they can trust and what they cannot trust. And probably they'll follow it up if it's really important with like, you know, some further data collection. And then they'll officially test any hypotheses that are really important, like I think we just sort of assume. We don't even talk about a lot of it. It's just, but I think like I get the impression that that's kind of what we imagine. And I think there's really interesting questions like, I think as someone who studied on certain events for a long time and for a while, you know, I was like, we have to be visualizing on certain way more than we are. Like I, I think like some of the things I've seen in my research just with, you know, how robust these tendencies people have to just want to see things sort of summarized or to just want to rely on statistical summaries over sort of raw data are, I mean, they're really kind of like compelling in the sense that like, you know, I think there's, there's a lot of cases where, you know, people sort of looking at aggregated data can actually work. It just kills me that we don't have like any sort of good formal way of describing why that is. So I think like I've sort of, one of the reasons I feel like I'm being pushed more towards theory in the sense of like trying to set up like almost like mathematical frameworks to understand some of these things is that I want to understand why, you know, like how do you explain that like if you have someone clicking around in a business system, trying to find patterns that like, you know, ultimately they, they, they, you know, are doing okay ultimately like they find like the correct ranking of patterns or whatever it is for the task like. So I think, yeah, I'm kind of like really curious just to like use theoretical frameworks to explain things that I don't understand. Like why does this work out? Like I think for instance, maybe, you know, like there's certain ways in which visual analysis process is redundant where you're sort of looking at the same data in multiple ways. And so, you know, if people kind of under update their beliefs often when they see a data sample, if you're looking at the same data sample multiple times, maybe sort of over time, you're kind of like internalizing it. Like I think there's all sorts of, you know, ways in which behavioral econ can sort of help us as well as, you know, like theories of statistical learning. Like I think, so yeah, I think it's like that's, you know, I don't know that that's where mainstream business is ever going to go, but for me it's sort of these questions about exactly when is visualization sufficient, like open this whole can of worms that just makes me think, like, okay, we have to sit down and like really try to like figure out can we explain to ourselves like how this, this paradigm work? - You touched upon so many interesting points. Like, I even know where to go next. - Yeah. - So many interesting points. And I'm still stuck with that end-sconquarted thing. - Okay, yeah. - So for those listeners who don't know what it is, so it's sort of a toy example to demonstrate why visualization is cool. And the ideas you have for artificial data sets that all have the same summary statistics, same mean, same standard deviation, like broad summary statistics are identical, but when you collect them, you see four very different shapes, right? And so I was wondering, is there an inverse end-sconquarted where we would have like four super similar plots, but the statistics tell us a hidden message or something. Are you aware of anything like that? - I can't think of anything, that's out. Yeah, that is a strange question. - One thing might be, like sometimes we, like fat tail distributions are really hard to see, but easy to-- - Yeah, I mean, like, yeah, that's a good example. I think like the behavior at the tails can really impact, like, you know, how you model data and stuff, like it actually matters a lot. - But you'll never see it in a graph, because-- - You might not notice it. - It's tiny and very stretched out, but it still makes a difference, you know, and stuff like that, so we might have biases towards, really, what is plotted well, right? And in our analysis, probably. - Yeah, interesting. Yeah, I think of Anscombe's quartet, I guess, is just like, you know, there's like statistics, like, multiplicity of statistics, and this, like, you know, like, that's why we visualize data, like, you can have the same statistical summary, and the data looks very different, and I think it's kind of interesting, like, recently, this seems to come up more in, like, machine learning, like, you can have, you know, multiple, like, you know, models, like, fitted models that seem to do equally well, like, on your test set, or in your, like, sort of, IID setting, but then when you probe them along ways that matter to humans, like, how, you know, how do they deal with gender, et cetera, like, they can give you very different answers, so I think it's, like, yeah, I mean, visualization in the sense of just, like, trying to put the data in some form where you can bring your prior knowledge to bear, I think is, like, maybe, Anscombe's quartet, we don't really think of it that way, it's like, oh, the answer is right there, like, they're all different, but I think it's, like, visualization is often this, like, first step, you know, towards, like, letting us take what we know, and try to apply it, I think it's just, we like to leave that kind of implicit, like, this is just people will bring in their knowledge and they'll know what to do next, kind of. - Yeah, and you wanna get from a lot of anecdotes to a theory or a model, ultimately, right, in either. - Right, which is a lot of it is, yeah, telling yourself stories, trying to explain things to yourself. - Right. - So yeah, I think we could give people better tools, though, like, as they're telling themselves stories to make sure that their stories are kind of accurate. So, like, on certain individualizations, sort of in that line, you know, like, let's show it to you. - Or even record the stories, you know, all that stuff, like, the construction process, like, the sense-making. - Mm-hmm. - You know, it's interesting that the current tools don't seem to really include any special, I don't know, functions that help people reason more about either their prior knowledge or their beliefs or even in building models or externalizing their knowledge. I think there is a huge, there's an interesting space there where something could be done. Anything you, Jessica, did some work in that space where you ask people to explicitly first build a model or externalize their belief, and then, yeah, the prior and then. - Or the you draw at first, you know, like, what do you think the statistics look like, right? - Yeah, which I like that stuff. I mean, I think the open-ended sort of, you draw at first thing is interesting. The other stuff, you know, we did like, eliciting priors where we would sort of have some Bayesian model and then we wanted to see how well does this model, this Bayesian model of cognition explain sort of what we actually see in terms of how people update their beliefs. And I mean, I think that's a good example of where having sort of a theoretical framework, even if it's wrong, like, even if people deviate because people do deviate from like, you know, the rational Bayesian update in various ways, like, you can still learn a lot about how they're off. But yeah, I think it's not that we don't like design ways to incorporate prior knowledge. It's just they're all extremely implicit. Like, and even like, like, Tableau, I didn't know this for a while, but there's a whole, like, analytics pane in Tableau where you can add regression lines and you can see intervals in various types, but it is very, it's very sort of rigid and constrained. Like, you get some number of choices. And I think it's really hard though. Like, you want people to sort of, and actually my former student now, faculty, Alex, Kale, and I have been working on something. related to some of the ideas in the paper with Andrew where it's like what would this new generation of visualization tools look like where you could come in and not be to sort of like you know like a seasoned statistician but still like use the tool to work up towards sort of these like preliminary statistical models that help you understand like how much does this variable explain this other one etc. So I think there's like a really yeah like really interesting space of like how do you sort of get people give them sort of this like scaffolding in visual analysis tools so that they can rather than just like their prior knowledge drives them to click around in all different ways and their prior knowledge like affects what graphs they draw but like in this very implicit way like how do we how do we allow it you know to like the help the tool give them back something like the tool suggests if they're looking in our case like if they're looking at a certain combination of variables the tool might suggest like do you want to try fitting a model to see you know how well you could predict this this dependent variable based on you know the variables that you seem to think are important so I think there's yeah a very big space but the whole thing with like eliciting people's beliefs and stuff also gets tricky like you know like ultimately doing data analysis is hard already and cognitively overwhelming so you can't be asking people a bunch of questions and so I don't know yeah it's an interesting interesting space. You're making me think now about the idea that going back to the idea of exploratory data analysis and maybe an excessive focus to this idea that with visualization you can just explore data for for the sake of it right so it seems to me that maybe that led to designing these tools in a way that a person opening the application for the first time right can pretty much do anything with it and if they don't do something there's there's basically no guidance right it's completely open but maybe there's space for something that is more guided in general I think it's a under explored yeah under explored modality where the system I actually guide you through a number of steps without being extensively rigid right yeah so we did we we've been building something that hopefully we'll have the paper out soon um that's sort of an explorer it's kind of like a version of tablo with like but with like built-in model checking and I think yeah like it it is trying to sort of give you tools that are a little more guided like it's not telling you exactly what to do when but one thing we see um is that you know like you you do sort of have to be careful about when you for certain people you know like I think if you come in knowing like having sort of a statistical workflow that you typically use you might know that like I don't want to jump into model building right away like I need to just look at things um but then if you um and then I'll get to the model checks stuff later and we've seen some people use our tool in that way like they just create a bunch of graphs and then they sort of like call up the modeling part of the the interface but then we also see people you know like the ones who aren't as experienced with with modeling where they just want to jump in right away to like building models um and I don't think that's good either so it's like yeah like there's the whole like I mean I guess it's like a user experience design type question you know like how do you yeah gradually introduce things is gonna like matter in the end if you just randomly apply models that somehow fit and you have no theory of the domain right and no no idea of causality it's it's always going to be a bit nonsensical and so right that is actually that's something we see with this tool we built as well like that some people it's almost like if you come from sort of like an ML kind of you know background you're you just want to like be trying out like just swapping in variables in some statistical model until you can like you know get the best predictive accuracy and you know in a visualization context you're trying to make sure that like the predictions from the model that are plotted against the data like best match the data but it's like that is not I mean it's sort of yeah contrary to this idea that like you know when we do exploratory data analysis we're trying to like really understand um and maybe test our expectations about like how the data were generated which means like we we often do have in mind like you know like we don't just care about any variables like we need to be able to have some plausible explanation for why that variable might matter um so yeah it's yeah yeah need to have a lot of knowledge about what's again what's a plausible range can this can this value even be below zero right or um so we had this case with the COVID excess mortality in Sweden and Germany where there were just different spline fitting techniques and some of them were they were just better but you couldn't explain that mathematically but more you know a lot about the domain you know and so I think that's also when it's get interesting when again and then maybe it's again the matter of being skilled at you know finding the right model by applying statistics and visualization but also knowing what what to look for and what what has worked in the past and you know all that practical stuff yeah training people on what to look for is another thing yeah like um with like trying to build model checking abilities into a visual analysis tool man Alex Kale on some of that work like you know like one thing we ran into is like people don't know um if we're trying to do this for people who don't have like a whole bunch of stats training and just like have some exposure to linear models maybe like we got to teach them sort of what different types of misfit look like because like looking at you know predictions against observed data to sort of like check your model in a graphic is like this very multi-dimensional thing it's not just like there's one way that predictions can deviate from like or the observed data can deviate from the model predictions there's many different things you can look for like you know how does it how does the model do with the tails of the distribution like you know like is it biased overall like so there's yeah there's this whole you know way almost like a type of visual little literacy or data literacy I think that that has to come along with you know tools that build more of this stuff in where you're you're helping people understand like what is how do you do model checks well um or like yeah in a way that's sensitive to all the different ways things can be off yeah this reminds me of of uh growing concern that I have um over the years I've been have become more and more concerned with the idea that we test visualization tools with non-expert I think it's uh there's a huge huge limitation there and um yeah I don't know I think if you if you if you if you run an experiment that is based on very low level perception maybe it's fine you can you can pretty much want any any person is equivalent to another right but as soon as you have some you involve something I mean you test something that involves that requires some domain knowledge in order to understand the data right then it's like um I had experiments where they had both novices and and actual practitioners who are familiar with the data yeah just night and day they're not even comparable it's completely different kind of totally yeah no I totally agree yeah it's convenient samples usually I mean in biz for sure yeah okay whole different uh all different topics yeah yeah I know cool okay that was quite uh wow yeah great episode I liked it so many interesting things each of these topics we could go on for hours right yeah it's very fundamental yeah okay thanks so much thanks for coming on the show yeah um hope to see you soon yeah nice to chat with you guys yeah wonderful thanks for joining us thank you thanks so much hey folks thanks for listening to datastories again before you leave a few last notes this show is crowdfunded and you can support us on patreon at patreon.com/datastories where we publish monthly previews of upcoming episodes for our supporters or you can also send us a one-time donation via paypal at paypal.me/datastories or as a free way to support the show if you can spend a couple of minutes reading us on iTunes that would be very helpful as well and here's some information on the many ways you can get news directly from us we are on twitter facebook and instagram so follow us there for the latest updates we have also a slack channel where you can chat with us directly and to sign up go to our own page at datastory.es and there you'll find a button at the bottom of the page and there you can also subscribe to our email newsletter if you want to get news directly into your inbox and be notified whenever we publish a new episode that's right and we love to get in touch with our listeners so let us know if you want to suggest a way to improve the show or know any amazing people you want us to invite or even have any project you want us to talk about yeah absolutely don't hesitate to get in touch just send us an email at [email protected] that's all for now hear you next time and thanks for listening to datastories.

Podcast Summary

Key Points:

  1. Style in data visualization is not a luxury but a fundamental organizing principle, akin to theory in writing or music, essential for clarity and effective communication.
  2. Visualizations serve as model checks, helping users test implicit or explicit expectations about data, revealing deviations that signal surprises or anomalies.
  3. Both statistics and visualization require a personal theory or approach—without one, analysis risks being arbitrary or unproductive, especially in exploratory data analysis.
  4. The relationship between statistics and visualization is deeply interwoven; effective analysis involves both discovery through exploration and validation through model checking.
  5. A key paradox exists
  6. Visualization tools frequently lack guidance on prior knowledge or belief elicitation, leading to implicit, unexamined assumptions and missed opportunities for deeper insight.
  7. Designing guided, scaffolded tools that support users in building mental models and checking assumptions can improve analytical rigor and signal-to-noise judgment.
  8. Exploration in data visualization is not random; it involves a structured process of hypothesis generation, model fitting, and validation, grounded in cognitive and domain knowledge.

Summary:

Data visualization is not merely a tool for communication but a deeply structured practice requiring personal theory and intentional design. Drawing from Andrew Gelman’s blog and Jessica Holman’s research, the conversation emphasizes that style and structure in visualizations are essential, not optional—just as theory is vital in writing or scientific reasoning. Visualizations function as model checks, testing expectations and revealing anomalies that challenge or refine our mental models of data.

This process parallels scientific exploration, where surprise and deviation are critical for progress, even if most analysis relies on representative, expected outcomes. The episode highlights a gap in current tools: they often lack mechanisms to support users in expressing or evaluating prior beliefs, leading to implicit assumptions. A more effective future for data analysis lies in guided, scaffolded visualization tools that help users build mental models, test hypotheses, and assess model fit through visual feedback.

These tools must balance openness with structure, allowing both novice and experienced users to explore data meaningfully while avoiding over-reliance on superficial patterns. Ultimately, the success of data exploration depends not just on technical skill, but on cognitive frameworks that enable users to distinguish signal from noise, question their assumptions, and build robust, interpretable insights.

FAQs

Style is not a luxury but an essential organizing principle, similar to how writing or music requires a consistent approach. It helps guide interpretation, makes content engaging, and ensures clarity, especially in complex data.

Data visualization and statistics are deeply interconnected; visualization acts as a tool for both communication and exploration. It helps test models, reveal patterns, and check assumptions, often functioning as a form of model checking.

A theory or 'style' of visualization provides structure and direction. Without it, visualizations become arbitrary and lack purpose. It helps researchers know when to refine or abandon a model, much like a scientific plan.

This principle suggests that when data signal is strong relative to noise, visualization alone may suffice—no statistical analysis is needed. Conversely, when noise dominates, statistical methods are required to extract meaningful insights.

Exploratory analysis focuses on discovering patterns and testing assumptions through visualization, while communication aims to convey findings. Both rely on expectations, but exploration is more about surprise and discovery.

In some cases, especially when signal is clear and noise is minimal, visualization can be sufficient. However, for complex or noisy data, statistical methods are often needed to validate findings and assess model fit.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.