People don't always understand the importance of style or the necessity of style,
but it's not a luxury. It's more of an organizing principle.
Hi everyone, welcome to a new episode of Data Stories.
My name is Erika Bertini and I am a professor at Northeastern University in Boston
where I teach and do research in data visualization.
Right, and I'm Mojda Fana. I'm an independent designer of data visualizations.
And in fact, I work as a self-employed truth and beauty operator
out of my office here in the countryside in the beautiful north of Germany.
Exactly. And on this podcast, we talk about data visualization,
analysis, and more generally, the role data plays in our lives.
And usually we do that together with a guest we invite on the show.
But before we start, just a quick note, our podcast is listener supported,
so there are no ads.
But that also means if you do enjoy the show, you might consider supporting us.
You can do that with recurring payments on patreon.com/datastories
or you can also send us one-time donations on PayPal.me/datastories.
Exactly.
OK, so I think we can get started with the main topic and guests for the show today.
So today, we have two people on the show to talk about, I think, a really relevant topic.
I think, in general, what is the relationship between statistics and data visualization?
And to talk about that, we have Andrew Gelman and Jessica Holman.
I, Andrew and Jessica, welcome to the show. Hello.
So as usual, we start by asking our guests to introduce themselves.
So maybe, Andrew, you want to go first and give a brief introduction?
I teach statistics and political science at Columbia University.
Jessica.
Hi, I'm a professor, associate professor of computer science at Northwestern University.
I do research on various topics related to how people draw inferences from data,
usually from interfaces.
So I care about things like visualization.
OK, so I thought we would start our conversation by starting from the blog that I believe Andrew
started several years ago.
It's called statistical modeling, causal inference, and social science.
I think it's a really influential blog and I remember reading the blog since many years.
And it's a really interesting community of people.
And there's a lot of interesting discussions about statistics, but also political science
and science in general.
And also a lot about data visualization and I believe Jessica joined the team recently.
So we have seen even more data visualization conversations in that space.
So Andrew, I was thinking maybe you could give us a little bit of overview of what the
blog is about.
Maybe if you want, even to say how it started and what are the main topics there.
In 2004 I was working with a postdoc Samantha Cook and we had an idea of setting up a blog
in a wiki to help us communicate with each other.
The idea of putting stuff on a blog was that then other people could see things too and we
could get input from other people.
Then the wiki was supposed to be where we put our various ideas.
The wiki got hacked and we had to take it down, but the blog was useful.
I learned, well, it's hard for most people to write stuff on a blog and sometimes I would
suggest that students are postdocs write a post and they would find it too difficult to
task.
They would find it too like pressureful.
So it ends up mostly being me, but then various other people, maybe about 15 other people
such, including Jessica, have had stuff to say, so I asked them if they could write for
it too.
So it's a way of getting, having conversations.
It is difficult to write for your blog, I was just going to say, Andrew, maybe it's hard
for you to understand, but I think you've established quite a record with it.
I think a lot of people tune in to hear what you're going to say.
So it is, I understand why other people would be like, oh my God, it's too much pressure
because I had to just get over it and be like, I don't care if they're going to compare
me to his post and they're not going to be as good, but it is chuggy.
And your posts are definitely better than the average post of mine, so.
Yeah.
I think what is really interesting there is that every time I look at a post, there's
such an interesting series of comments below.
You seem to have a very active community around it and it's always very thoughtful.
Yeah, and I don't know if you need any special moderation, but also normally the comments
are pretty interesting and nothing too bad.
There's no crazy people writing crazy stuff as far as I can tell.
Yeah, it's not so interesting for the crazy people, I guess.
Yeah.
But I wanted to pick up on something that you said earlier, not about the blog, but about
statistical graphics, because I'm a user of graphs ever since I was a physics student
and graph data and graph curves and models and so forth.
And I think over the years, it struck me that I think everybody needs to have their own
theory of statistical graphics.
In the same way as if you're writing, you need to have a theory of writing or if you're
drawing, you can't just say I'm going to draw what I see.
You have to have a kind of approach, its goals, or if you're making music.
You can't just say, hey, let's bunch of us.
Let's form a musical group and do music, right?
You have to have music that you want to do.
You have a certain style.
It doesn't mean your style is better than everybody else's, but you need that.
And I think that for quantitative things, people don't always understand the importance of
style or the necessity of style.
Like, style is seen as a kind of luxury, but it's not a luxury.
And that's also clear with writing that if you're writing, except for perhaps the most
functional writing like the instructions for how to operate your microwave oven or something
like that, you need to have a style because if something is boring to read, then no one
will read it.
If you teach as a teacher, if you teach a class and write a two-page document for the
students to read, they won't read the two-page document.
And so it's really it's needed.
So I, and I think the flip side of that is that on the other hand, then you have people
saying, oh, someone's a designer as if there can't be any connection to science because
design is supposed to be some like separate thing, but they're one can draw the analogy
of something like building bridges or that these things are designed, but they still have
to work.
Yeah, I really like the way you're describing this because in a way, it reminds me of something
that maybe I was, I was hoping to discuss that to me, it reminds me of the role of theory,
right?
The fact that if you don't have a theory, theory doesn't have to be necessarily perfect or
always super predictive, but it's a way to organize your thoughts so that you can think
systematically about something.
So I don't know if this rings any bells on your side, but the way you describe style to
me reminds me how important it is to have theory and also a theory of graphics in this specific
instance.
Well, it's like they say in chess, having a plan won't win you the game because presumably
you're playing against someone else with a plan too when you're not both going to win,
but if you don't have a plan then you'll lose, you won't be able to move forward.
And part of having a plan is recognizing being aware of that you have a plan, being aware
of what the plan is, and then when things go wrong, you can change things.
So actually it's like putting you as a scientist, like putting your marker down, I'm not really
into like betting, it's not like I would say you literally have to bet money on things,
but like conceptually setting it down and saying this is the model I'm going with is very
valuable even though or I should say especially though we know that that model is going to be
wrong and it's going to fall apart at some point.
But the more explicit you can have that model or style or system, the more that you can
then know when to work on improving it or abandoning it.
We actually wrote a paper about some of this as it applies to sort of theories of visualization
for exploratory analysis where I think that's a place in visualization where we're building
all these tools to help people do interactive data analysis through visualization.
But we don't always like if you ask us the people developing these tools or the researchers
are in the area like what are what are sort of guiding principles are.
I think it's often while we just want to let people explore data as easily as possible.
But I think there you can easily run into places where you just like you have no theory
to tell you how to design something in sort of a better way versus a worse way like you
just don't know.
We wrote a paper for the Harvard Data Science Review
like about a year or so ago.
where we sort of talked about this as applies to exploratory visual analysis and we argue that even if it's a bad theory like Enrico was saying or both of you guys were saying that it still can be useful because you need to know sort of how you were wrong and if you never state what you're going for what you think the objective is how do you know when you were wrong.
Yeah, that's an area of data visualization that I really love and I think people tend to talk less about it. I think there is more of a I think the general idea with visualization is that it's a tool for communication but there is less I would say there is less discussions about how to use it for exploration and by the way the word exploration itself is so contentious in a way so yeah.
It's funny. I thought it's the other way around that everybody talks really exploratory. Yeah, I was kind of thinking in academia. Yeah, I mean nobody like really gets what communication means.
I think I think that one problem is communication is often viewed as being a kind of unidirectional thing so like those same things like scientist should learn how to tell stories because people think in terms of stories.
And I hate that kind of attitude. I mean sure people think in terms of stories but but this attitude that oh you're the scientist you already know the answers but now you have to convey it to people so you have to learn how to be a storyteller and and have a good bedside manner right like it's all connected with like you don't want to be a jerk right like you would be it's like narrative medicine and all this stuff and it's like the what people are actually doing there is great but the idea the framing that it's what you're going to say.
It's all about how to communicate truths to people. I think it's misleading. I think it's more accurate to say that we're people too and we learned from stories and this is something that my colleagues and I've been thinking a lot about over the years like what makes us believe things and and often work convinced by stories and I.
We I character the effective stories as being anomalous and immutable and by anomalous it's a story is a surprise so the convincing stories have some twist in them something unexpected even if you think about a scientific method I didn't think this method would work and and then it did or or or whatever and then it's immutable in the sense that a good story is grounded in reality and like if.
If you kick it your foot hurts right like as it as in the famous Boswell story and so you have like.
But so we we kind of learn from or we we learn from these stories which are a reality and maybe the term story isn't the best in that sense because I'm not talking about stories that are made up and I'm talking about true stories but but we learn from these but it's it's kind of.
Necessary that the stories have this grounding in truth so that they can disprove our theories and there's a sense in which a good story like it like if you were to take all of the things that people have said in your podcast right so maybe.
You've done however made podcast in each podcast an average maybe there are five stories like when someone tells you hey let me tell you a story about that and if you look at those stories some of them are going to be just made up I mean we have this horrible examples of stuff where people just say like there's a famous example from a few years ago in a book or somebody said that a certain.
It was like a certain data problem caused 70 deaths a year in a small town and it was like how the hell like it made no sense right so they just made it up but I think the stories that are good like if.
You could in theory track them down and I think they would have this characteristic that they disproven implicit model of the world right I have to do with news newsworthiness as well like in journalism you have certain newsworthiness criteria.
Exactly dog bites man by it's dog but here's the point is that when something is surprising it's surprising relative to an expectation and so that model of the world is is that so when we talk about discovery and surprise there are theories implicitly they're already.
Yes this is yeah I was going to tie this back to the exploratory analysis thing as well like I think one of I mean I think we talk about exploratory visual analysis a lot and is probably more than communication but.
You know it's there's always this role of expectations and what you're bringing like what you're expecting to see both in communication and exploratory visual.
Analysis and I think at least in visualization research that something that has we've always sort of background it you know we've always acted as though like the data speaks for itself when it obviously does not so yeah so I think yeah that's all like theories of how you know.
How visualizations act as model checker one thing that and I have like in the same paper I mentioned already he had he had been thinking about this years ago.
As a way to sort of think about the role of graphics and exploratory data analysis and how in a sense you can tie it to confirmatory data analysis through this idea that a good graph is helping us check a model some implicit model often sometimes an explicit model like in confirmatory data analysis but there's always.
You know some expectation and the graph tells us you know how much the data deviate from that so when you say model check here do you mean like checking the model that you are stored in your head like a mental model yeah I mean I think yeah so in in exploratory visual analysis I mean I think.
You know there's always you know some background assumption potentially in a lot of times and this is stuff that Andrew had spoken about back in 2003 in a paper so you can cut me off whenever Andrew and tell it yourself but basically you know like you can think of you know a visualization is giving you almost like a test statistic or a vector of test statistics in a hypothesis testing type framework if you want to look at it that way where you know you have something that you're you're trying to check for like you make a graph in order to check like.
You know like how well does my data conform to some expectation and sometimes that's really explicit and like built into the visualization like I want to you know in like a sort of model fitting or preliminary model fitting kind of stage of a workflow you're looking at things like maybe residuals and you know that exactly how to read the chart because it's sort of built into it that like if your data deviates from the.
Expectations that you want your residuals to have then or to fulfill then it'll be obvious because you'll see like deviations from symmetry in the plot or even a scatter plot you know like often if you have a library scatter plot like the most common sort of built in you know thing that you're checking against is sort of like a straight diagonal line representing kind of perfect linear association so.
The idea that back in 2003 that Andrew started talking about and other people in statistical graphics have also gotten into like Andrea Spugia Diane cook and others you know there's this work in sort of graphical statistical inference that gets into this idea of.
You know like how visualizations can function sort of as model checks but then within that you know people gone in different directions where Andrew's originally original formula formulation I believe was sort of more in a Bayesian direction where.
It's not that we're testing some hypothesis and we just want sort of you know our p value or kind of yes no answer but it's it's that we're you can think of the graph is kind of like.
The comparison that you're doing mentally when you look at a graph is kind of akin to doing like a posterior predictive check in Bayesian stats where you're sort of imagining like you know under my expectations about the process that created my data what do I expect the data to look like.
Like so what is what is sort of what would reasonable data look like under the under the predictions I want to make and how much does the data that I actually got sort of compare to what I would expect so it's almost like you know you can imagine on some level maybe this doesn't happen all the time but it's almost like when when you look at graphs you're sort of imagining.
You know you know reasonable data under some some side of expectations you have and you're comparing that to what you see Andrew I don't know if you want to say anything there I think that's that's just sort of the idea I wanted to bring up well.
Yeah I've more to say about that but actually let me jump to something else which is a paradox that we we learn from stories and the best stories have surprises in them like I would almost argue all stories have surprises in the sense that if there's no surprise you don't bother telling the story.
So we're always using just as we use graphs to learn and to discover which means to be surprised relative to our implicit models we consume stories in order to refute various models of the world.
But yet that the paradox is that how can we learn from surprising things like it seems the usual way we think about statistics is that we learn from the expected like random samples right that's like if I'm going to do a survey I don't say I found the 1000 weirdest people in America asked them their opinions about things and I really wanted to be surprised and you'll never guess what I found.
No what you'll do is you want to ask a representative sample of people and like you don't really have the goal of it.
being surprised. You just, you want to see that. So this was sort of bothering me actually
because after I wrote the paper, the paper that my colleague and I wrote about why, how
we learned from stories, I drew, directly connected to the papers that Jessica mentioned
where we earlier, I had written that graphics are a form of model check. So then I argued
stories are a form of model check. But then again, it, then people ask the question and how
can it be that you learn from anomalies? And my, I don't know, like my, my resolution
of this is that it's related to a kind of proper area in our Lakotoshian view of science
or Qnian view or whatever. I guess the standard view now of science, which is that we alternate
between normal science and scientific revolutions. And when we're doing normal science, we are
kind of representative samples and we want to help build, we want to build theories and
modify our theories. Then when we're doing revolutions, we're trying to see what's wrong
with our theories. And there we're looking for counter examples. Now, I'll only say one
more thing, which is that when you say this, it always sounds like revolution is the hero
in normal sciences is like the loser in this game. But that's not true because the revolution
only exists because, because there was a normal science allowed it and the goal of a revolution
is to replace it with a new normal science. And so it's like both of these depths are
important. I have a question here. Like how, how open are people really to changing their
minds based on statistical information? I think that's, it's something like the last few
years, maybe been an interesting research topic now. I was thinking to, like as Andrew was
talking about anomalies, like, do you, like, should you have to go out explicitly searching
for anomalies, like to break a theory, or if we were all sort of honest scientists, could
we kind of through our own, like just seeing the data that we collect, find anomalies.
Like if we're willing enough to sort of admit when our, when our mindset is not right, you
know, then, then we should be seen probably anomalies a lot.
Well, this is kind of related to this, like, unitary nature of consciousness thing or even
the idea that we talked a bit earlier of having a plan. Like it's, it seems pretty fundamental
like to, to mathematics. It's, I mean, the way cognition works in general, not just like
human brains, that it seems like there needs to be this executive function and this alternation
of processes. So like the same, you need a division of labor somehow. So maybe one scientist
could create and refute her own theories and gather data, but maybe not all at the same
time. I mean, another example is in math class way back when, when you're asked, sometimes
you're asked to either prove something or come up with a counter example, and they always
say you can't do both at the same time. You have to first assume it's true and try to prove
it. And then if you can, stop and assume it's false and, and try to do that. So we, we do kind
of use the multiple points in the system to play different roles. Yeah. I have another question.
Like, looping back to the beginning. So I think you rightly explained visualizations
is really a skill and a practice and there's no single right way, but it's like a highly personal,
like thing, which, how you do it, right? And there could be many ways to do it, right?
Is it the same for statistical practice, like for applying statistics, or is it in statistics more,
for a given problem, there is a correct solution? There, of course, there are many ways of,
of solving problems. I actually wrote something once about what I call the method logical attribution
error, which is people attributing to their method, what's also a property of their skill.
And so you'll see this with like renowned statisticians, or maybe not so renowned statisticians,
also that they just think some method is inherently better. But yeah, there's always so many
unwritten roles. And they might just be better at applying it or failing to apply another technique
successfully, which somebody else might have done. Yeah, I'm better at some techniques than others.
So you just, that's how that's how it works. So it's, there's an interaction. Yeah, yeah, yeah.
It's interesting, because from the outside, so I have only statistics 101 knowledge, right? And to
me, it always seemed there's this clear decision tree of if your data is shaped like that,
then you need to apply a nova test or something like this, right? And we have the same for graphs.
If you want to spot outliers, use a scatter plot, right? And so I was always wondering how,
how hard cut these rules are. And I'm glad to hear that. It seems very similar. Actually,
that the deeper you go, the less clear it is how things should be actually.
I was influenced by a colleague, David Kranz, a psychologist, who he was telling me about like
decision theory. And he said that the simplest version of decision theory is what you learn,
like the like fun norm and Morgan Stern, you have a decision tree that you need to evaluate.
So you, you compute all the things that go into it and you compute the tree. And then
the next level of sophistication is to say, no, actually drawing the tree is important.
And there's lots of psychological experiments where they show people the tree and it's missing
a branch and like people don't realize. So like a lot of examples were the best decision
of something that wasn't in the tree in the first place. And then we tried it. But then he said,
that's also not enough. And so his taken on decision analysis is that you start with goals.
So you have goals and resources and stakeholders and all of that. Sounds kind of soft, but
it's not really softer than trees. So you basically start with the design thinking again.
Well, yeah, you start with your, you start with your goals and, and then all the other things,
your goals and your constraints and your resources. And then you consider ways of getting there
while being open to that your goals might change and so forth. So yeah, statements like,
if your data look like this, you should use this model. That's like totally, that's horrible
because you really want to be starting with your goals. And, you know, and, and not in an empty way.
Like, oh, yeah, my goal is to publish a paper. My goal is to get this data analysis done. Like,
you know, your serious goals, whatever they are. My goal is to not get shouted at on the internet
foremost. Yeah. So I get it. I have to go now. So I'll see you all later. But thanks for the
opportunity for talking with you all. This is, this is always fun. Thanks so much. Wonderful.
Thanks for joining us. Thanks, Andrew. See you all later. Thanks again. Bye. Okay. And now we can
continue the rest of the episode with Jessica. So one question that I had going back to,
let's say, the comparison between statistics and visualization, right? In my head, I'm always like,
I can't say that one is better than the other, right? I don't know. So for, for instance,
when we teach visualization, we show the outcomes quite bad. I think we have a huge bias there
because it's almost like, take that statistics, right? This is so much better.
This is how I use it. Right. Everybody uses that way. It's like, hey, we go. That's the,
that's why we need this. And, but no, because I think there's, there's almost like a dance
between having a lot of details so that we can maybe related to what Andrew was saying, right?
The surprising elements, but you can't do, you can't reason only with surprise, right? And
surprise can also overwhelm you. And you may lose the signal as you look at a lot of noise,
right? So I really see that as a dance between these two things, the surprising, the particular,
but also surfacing the signal. So, yeah. How do you think about that? I mean, it reminds me of,
like, you know, Tuky and others who have written about like the exploratory data analysis process,
where, I mean, my personal view is that in visualization, we sort of take kind of like the very
initial stages, you know, where, or the very initial stages of exploratory analysis are often
something, you know, like just making sure there's no massive, like chunks of missing data,
figuring out what variables you have to begin with. But I would say we take this like, then next
stage of like, you know, I'm just trying to see where is there some sort of, you know, signal or
what looks like a pattern, you know, what kind of relationships do I think I see? We sort of
in visualization, I think, think of that as exploratory analysis. So it's sort of like very open-ended,
sort of clicking around to find patterns. And I think we design kind of as though that is,
that is what exploratory visual analysis is all about. Tuky, for instance, though, talked about
this sort of intermediate phase where, you know, you've noticed, you've sort of generated some
hypotheses about like, you know, possible relationships between variables or the nature of
certain distributions. And then you sort of need to know, like, you know, how much can I believe
what I think I see here? And so, you know, there's also like, you know, exploratory analysis also
involves things like starting to fit models to try to explain, like if I think that this, you know,
set of variables seems to be predictive of some other variable that I care about, you know, I would
to actually start.
15 models and looking at deviation in things like residuals,
seeing how out of my models actually explain what I'm seeing.
And the whole idea is that I want to build up
some sort of a mental model kind of of the data generating
process.
I want to use stats to sort of help me figure out
what can I believe in terms of what signals are actually
there and what is just sort of maybe not actually
going to hold up when I inspect it more closely.
And so yeah, I think there's this weird stage where
or this weird sort of way in which graphs are used,
not just to show us things that maybe we didn't expect to see
or we didn't expect to see, but also to give us some information
about how much we should believe those things.
And I think that's where it's sort of kind of ambiguous.
So actually from Andrew's blog, I learned about this informal
term someone used called the Anthropic Principle of Statistics,
which is like there are certain problems
where you would use statistics.
Like if your data, if the signal is so huge relative
to the noise, like you don't really need to be running stats
on that, like you can just sort of see it.
And maybe you just make a graph and it's like obvious.
If your data, if the noise is very large relative to the signal,
then it's sort of hopeless.
And you could do stats, but you're still kind of dealing
with too much noise.
And so it's like statistics is useful for this sort
of middle set of problems.
And so I think, I mean, one of the questions
will ask you or to, for me, is like, I think of this.
We have like the problems in the middle
where we can use maybe visualization separately from stats.
Like what do we think is a problem or visualization
is simply not going to work?
When is visualization sufficient without any follow-up?
And I don't think there's like true or right answers.
These are things we have to sort of figure out as a field.
But I think it's not.
We haven't always made explicit sort of what our assumptions are.
Like I think we often maybe implicitly
when I look at what people write about exploratory analysis
and designing for, you know, exploratory visual analysis,
I think there's sort of this assumption
that, you know, people can click around for patterns.
And maybe though, you know, they care enough
about finding the right answers.
Like there's some, you know, like application
or some reason why they're analyzing the data.
And so we sort of trust that, you know,
they will look at things enough and make enough views.
And some of the views will be disaggregated enough
that they can sort of get a sense of the noise.
And so, you know, we don't have to worry
about explicitly supporting these like signal to noise
kind of judgments.
We can just let people, you know, use these tools
and they will figure out what they can trust
and what they cannot trust.
And probably they'll follow it up if it's really important
with like, you know, some further data collection.
And then they'll officially test any hypotheses
that are really important, like I think we just sort of assume.
We don't even talk about a lot of it.
It's just, but I think like I get the impression
that that's kind of what we imagine.
And I think there's really interesting questions like,
I think as someone who studied on certain events
for a long time and for a while, you know,
I was like, we have to be visualizing on certain way
more than we are.
Like I, I think like some of the things I've seen
in my research just with, you know, how robust
these tendencies people have to just want to see things
sort of summarized or to just want to rely
on statistical summaries over sort of raw data are,
I mean, they're really kind of like compelling
in the sense that like, you know, I think there's,
there's a lot of cases where, you know,
people sort of looking at aggregated data can actually work.
It just kills me that we don't have like any sort of good
formal way of describing why that is.
So I think like I've sort of, one of the reasons
I feel like I'm being pushed more towards theory
in the sense of like trying to set up like almost
like mathematical frameworks to understand some
of these things is that I want to understand why, you know,
like how do you explain that like if you have someone
clicking around in a business system,
trying to find patterns that like, you know,
ultimately they, they, they, you know, are doing okay
ultimately like they find like the correct ranking
of patterns or whatever it is for the task like.
So I think, yeah, I'm kind of like really curious
just to like use theoretical frameworks
to explain things that I don't understand.
Like why does this work out?
Like I think for instance, maybe, you know,
like there's certain ways in which visual analysis process
is redundant where you're sort of looking at the same data
in multiple ways.
And so, you know, if people kind of under update
their beliefs often when they see a data sample,
if you're looking at the same data sample multiple times,
maybe sort of over time, you're kind of like internalizing it.
Like I think there's all sorts of, you know,
ways in which behavioral econ can sort of help us
as well as, you know, like theories of statistical learning.
Like I think, so yeah, I think it's like that's, you know,
I don't know that that's where mainstream business
is ever going to go, but for me it's sort of these questions
about exactly when is visualization sufficient,
like open this whole can of worms that just makes me think,
like, okay, we have to sit down and like really try to like
figure out can we explain to ourselves
like how this, this paradigm work?
- You touched upon so many interesting points.
Like, I even know where to go next.
- Yeah.
- So many interesting points.
And I'm still stuck with that end-sconquarted thing.
- Okay, yeah.
- So for those listeners who don't know what it is,
so it's sort of a toy example to demonstrate
why visualization is cool.
And the ideas you have for artificial data sets
that all have the same summary statistics,
same mean, same standard deviation,
like broad summary statistics are identical,
but when you collect them,
you see four very different shapes, right?
And so I was wondering, is there an inverse end-sconquarted
where we would have like four super similar plots,
but the statistics tell us a hidden message or something.
Are you aware of anything like that?
- I can't think of anything, that's out.
Yeah, that is a strange question.
- One thing might be, like sometimes we,
like fat tail distributions are really hard to see,
but easy to--
- Yeah, I mean, like, yeah, that's a good example.
I think like the behavior at the tails can really impact,
like, you know, how you model data and stuff,
like it actually matters a lot.
- But you'll never see it in a graph,
because-- - You might not notice it.
- It's tiny and very stretched out,
but it still makes a difference, you know,
and stuff like that, so we might have biases towards,
really, what is plotted well, right?
And in our analysis, probably.
- Yeah, interesting.
Yeah, I think of Anscombe's quartet, I guess, is just like,
you know, there's like statistics,
like, multiplicity of statistics, and this, like, you know,
like, that's why we visualize data,
like, you can have the same statistical summary,
and the data looks very different,
and I think it's kind of interesting, like, recently,
this seems to come up more in, like, machine learning,
like, you can have, you know, multiple, like, you know,
models, like, fitted models that seem to do equally well,
like, on your test set, or in your, like, sort of,
IID setting, but then when you probe them
along ways that matter to humans, like, how, you know,
how do they deal with gender, et cetera,
like, they can give you very different answers,
so I think it's, like, yeah, I mean, visualization
in the sense of just, like, trying to put the data
in some form where you can bring your prior knowledge to bear,
I think is, like, maybe, Anscombe's quartet,
we don't really think of it that way,
it's like, oh, the answer is right there,
like, they're all different, but I think it's, like,
visualization is often this, like, first step, you know,
towards, like, letting us take what we know,
and try to apply it, I think it's just,
we like to leave that kind of implicit,
like, this is just people will bring in their knowledge
and they'll know what to do next, kind of.
- Yeah, and you wanna get from a lot of anecdotes
to a theory or a model, ultimately, right, in either.
- Right, which is a lot of it is, yeah,
telling yourself stories, trying to explain things to yourself.
- Right.
- So yeah, I think we could give people better tools, though,
like, as they're telling themselves stories
to make sure that their stories are kind of accurate.
So, like, on certain individualizations,
sort of in that line, you know, like, let's show it to you.
- Or even record the stories, you know, all that stuff,
like, the construction process, like, the sense-making.
- Mm-hmm.
- You know, it's interesting that the current tools
don't seem to really include any special,
I don't know, functions that help people
reason more about either their prior knowledge
or their beliefs or even in building models
or externalizing their knowledge.
I think there is a huge, there's an interesting space
there where something could be done.
Anything you, Jessica, did some work in that space
where you ask people to explicitly first build
a model or externalize their belief,
and then, yeah, the prior and then.
- Or the you draw at first, you know,
like, what do you think the statistics look like, right?
- Yeah, which I like that stuff.
I mean, I think the open-ended sort of,
you draw at first thing is interesting.
The other stuff, you know, we did like,
eliciting priors where we would sort of have some
Bayesian model and then we wanted to see how well does
this model, this Bayesian model of cognition explain
sort of what we actually see in terms of how people
update their beliefs.
And I mean, I think that's a good example
of where having sort of a theoretical framework,
even if it's wrong, like, even if people deviate
because people do deviate from like, you know,
the rational Bayesian update in various ways,
like, you can still learn a lot about how they're off.
But yeah, I think it's not that we don't like design ways
to incorporate prior knowledge.
It's just they're all extremely implicit.
Like, and even like, like, Tableau, I didn't know this for a while,
but there's a whole, like, analytics pane in Tableau
where you can add regression lines
and you can see intervals in various types,
but it is very, it's very sort of rigid and constrained.
Like, you get some number of choices.
And I think it's really hard though.
Like, you want people to sort of,
and actually my former student now, faculty, Alex, Kale,
and I have been working on something.
related to some of the ideas in the paper with Andrew where it's like what would this new
generation of visualization tools look like where you could come in and not be to sort
of like you know like a seasoned statistician but still like use the tool to work up towards
sort of these like preliminary statistical models that help you understand like how much
does this variable explain this other one etc. So I think there's like a really yeah like
really interesting space of like how do you sort of get people give them sort of this like
scaffolding in visual analysis tools so that they can rather than just like their prior knowledge
drives them to click around in all different ways and their prior knowledge like affects
what graphs they draw but like in this very implicit way like how do we how do we allow it
you know to like the help the tool give them back something like the tool suggests if they're
looking in our case like if they're looking at a certain combination of variables the tool might
suggest like do you want to try fitting a model to see you know how well you could predict this
this dependent variable based on you know the variables that you seem to think are important so
I think there's yeah a very big space but the whole thing with like eliciting people's beliefs
and stuff also gets tricky like you know like ultimately doing data analysis is hard already
and cognitively overwhelming so you can't be asking people a bunch of questions and so
I don't know yeah it's an interesting interesting space. You're making me think now about the idea
that going back to the idea of exploratory data analysis and maybe an excessive focus to this idea
that with visualization you can just explore data for for the sake of it right so it seems to me
that maybe that led to designing these tools in a way that a person opening the application for
the first time right can pretty much do anything with it and if they don't do something there's
there's basically no guidance right it's completely open but maybe there's space for something that
is more guided in general I think it's a under explored yeah under explored modality where the
system I actually guide you through a number of steps without being extensively rigid right yeah so
we did we we've been building something that hopefully we'll have the paper out soon um that's
sort of an explorer it's kind of like a version of tablo with like but with like built-in model
checking and I think yeah like it it is trying to sort of give you tools that are a little more
guided like it's not telling you exactly what to do when but one thing we see um is that you know
like you you do sort of have to be careful about when you for certain people you know like I think
if you come in knowing like having sort of a statistical workflow that you typically use you might
know that like I don't want to jump into model building right away like I need to just look at
things um but then if you um and then I'll get to the model checks stuff later and we've seen some
people use our tool in that way like they just create a bunch of graphs and then they sort of like
call up the modeling part of the the interface but then we also see people you know like the ones
who aren't as experienced with with modeling where they just want to jump in right away to like
building models um and I don't think that's good either so it's like yeah like there's the whole like
I mean I guess it's like a user experience design type question you know like how do you yeah
gradually introduce things is gonna like matter in the end if you just randomly apply models that
somehow fit and you have no theory of the domain right and no no idea of causality it's it's always
going to be a bit nonsensical and so right that is actually that's something we see with this tool
we built as well like that some people it's almost like if you come from sort of like an ML kind of
you know background you're you just want to like be trying out like just swapping in variables in
some statistical model until you can like you know get the best predictive accuracy and you know
in a visualization context you're trying to make sure that like the predictions from the model that
are plotted against the data like best match the data but it's like that is not I mean it's sort of
yeah contrary to this idea that like you know when we do exploratory data analysis we're trying to
like really understand um and maybe test our expectations about like how the data were generated
which means like we we often do have in mind like you know like we don't just care about any variables
like we need to be able to have some plausible explanation for why that variable might matter um
so yeah it's yeah yeah need to have a lot of knowledge about what's again what's a plausible range
can this can this value even be below zero right or um so we had this case with the COVID
excess mortality in Sweden and Germany where there were just different spline fitting techniques
and some of them were they were just better but you couldn't explain that mathematically but more
you know a lot about the domain you know and so I think that's also when it's get interesting when
again and then maybe it's again the matter of being skilled at you know finding the right model
by applying statistics and visualization but also knowing what what to look for and what what has
worked in the past and you know all that practical stuff yeah training people on what to look for is
another thing yeah like um with like trying to build model checking abilities into a visual analysis
tool man Alex Kale on some of that work like you know like one thing we ran into is like people don't
know um if we're trying to do this for people who don't have like a whole bunch of stats training
and just like have some exposure to linear models maybe like we got to teach them sort of what
different types of misfit look like because like looking at you know predictions against observed data
to sort of like check your model in a graphic is like this very multi-dimensional thing it's not just
like there's one way that predictions can deviate from like or the observed data can deviate from
the model predictions there's many different things you can look for like you know how does it
how does the model do with the tails of the distribution like you know like is it biased overall like
so there's yeah there's this whole you know way almost like a type of visual little literacy or data
literacy I think that that has to come along with you know tools that build more of this stuff in
where you're you're helping people understand like what is how do you do model checks well um or like
yeah in a way that's sensitive to all the different ways things can be off yeah this reminds me of
of uh growing concern that I have um over the years I've been have become more and more
concerned with the idea that we test visualization tools with
non-expert I think it's uh there's a huge huge limitation there and um yeah I don't know
I think if you if you if you if you run an experiment that is based on very low level perception
maybe it's fine you can you can pretty much want any any person is equivalent to another right
but as soon as you have some you involve something I mean you test something that involves
that requires some domain knowledge in order to understand the data right then it's like um I
had experiments where they had both novices and and actual practitioners who are familiar with the
data yeah just night and day they're not even comparable it's completely different kind of
totally yeah no I totally agree yeah it's convenient samples usually I mean in biz for sure yeah okay
whole different uh all different topics yeah yeah I know
cool okay that was quite uh wow yeah great episode I liked it so many interesting things
each of these topics we could go on for hours right yeah it's very fundamental yeah
okay thanks so much thanks for coming on the show
yeah um hope to see you soon yeah nice to chat with you guys yeah wonderful thanks for joining us
thank you thanks so much
hey folks thanks for listening to datastories again before you leave a few last notes
this show is crowdfunded and you can support us on patreon at patreon.com/datastories
where we publish monthly previews of upcoming episodes for our supporters or you can also
send us a one-time donation via paypal at paypal.me/datastories or as a free way to support the show
if you can spend a couple of minutes reading us on iTunes that would be very helpful as well
and here's some information on the many ways you can get news directly from us we are on twitter
facebook and instagram so follow us there for the latest updates we have also a slack channel
where you can chat with us directly and to sign up go to our own page at datastory.es and there
you'll find a button at the bottom of the page and there you can also subscribe to our email
newsletter if you want to get news directly into your inbox and be notified whenever we publish a
new episode that's right and we love to get in touch with our listeners so let us know if you
want to suggest a way to improve the show or know any amazing people you want us to invite
or even have any project you want us to talk about yeah absolutely don't hesitate to get in touch
just send us an email at
[email protected] that's all for now hear you next time and thanks for
listening to datastories.