Go back

Week 6: Weaving the Web with Zach Swiecki

29m 8s

Week 6: Weaving the Web with Zach Swiecki

This podcast episode features a discussion between host David Williamson Schaefer and guest Zach Swikke on navigating and analyzing data responsibly, with a focus on qualitative research methods and Quantitative Ethnography (QE). Swikke explains the process of "packing" qualitative data by setting context and "unpacking" it through detailed interpretation, linking evidence back to overarching claims. The conversation explores how qualitative insights can be supported by quantitative techniques like Epistemic Network Analysis (ENA), which models connections in discourse to distinguish patterns among groups. Swikke emphasizes starting with qualitative narratives to ground findings before introducing quantitative models, ensuring results are meaningful and contestable. Key considerations include selecting and refining codes for ENA, iterating models to improve accuracy, and structuring research writing to guide readers through complex data, especially in unfamiliar contexts. The episode underscores the value of integrating qualitative and quantitative approaches to build robust, interpretable analyses in learning analytics and beyond.

Transcription

5687 Words, 31328 Characters

English
We live in a vast sea of data. Information is collected about every one of us with every click, every swipe, every post, and every like. This is a podcast about how to navigate responsibly and find a meaningful place to put ashore in the ocean of big data. Welcome to the quantitative ethnography podcast hosted by David Williamson Schaefer, faculty director of the Master of Science and Educational Psychology Learning Analytics program at the University of Wisconsin-Madison. My guest today is Zach Swikke. He's a lecturer in the Department of Human Centered Computing and a researcher in the Center for Learning Analytics at Monash University in Melbourne, Australia, and also a former graduate student of mine. Welcome, Zach. Hi, David. Nice to be here. We are talking about a number of things, but including folks who may have read Zach's wonderful paper, assessing individual contributions to collaborative problem-solving, a network analysis approach. Some of what we say may be referring to that. So, Zach, one of the things that I really love about that paper is the way that you handled the qualitative data. Could you talk a little bit about how you think about packing and unpacking qualitative data when you're working with it? For me, just to talk a little about what those things are, when I think about packing, I think about setting the context for what we're about to see as the reader, saying something about this is what's going on. This is the point that I kind of want to take from this, that I'm about to show. And then the evidence comes in the form of some qualitative data here is actually what I was talking about before when I was packing. And then the unpacking is actually going through that evidence and relating it to that overall claim. And there are a number of ways to do this, but usually you go through that data line by line, giving quotes from the data where appropriate and referencing specific lines. And the whole time you're relating it back to that overall claim that you were talking about when you packed the beginning. One of the things that I think makes it effective is if you're actually referring to the specific codes that you're seeing in the qualitative excerpts, you're actually sort of like coding it as you go. Yeah, so you're showing, you know, you're kind of showing how the codes are building that story and related to that story. Typically it's good when you're showing the qualitative evidence, or you're not just showing excerpts, but you're showing where those codes are applied in those excerpts. And when you're unpacking, you're referencing those codes in some way. In a sense, it's a translation or data compression, right? You're starting with the qualitative data. You're translating that into the codes. And then ultimately those codes get translated to the story and into the model, right? One of the things that was tricky about this paper was these are very long examples that you're trying to give. And there are examples that are in a domain that most people aren't familiar with. How did you approach setting that up? First, I should say, there was quite a long process of familiarizing myself with the data as it wasn't context, I was particularly familiar with either I wasn't in the Navy. So there was a lot of reading of the data, reading of background resources, talking to people who were involved and actually collecting this data was lovely enough to have a contact who was around when this was actually done. Thinking about that, I think that it was important to really set that context for the readers. This is generally what's going on in the space. This is what the people are doing. This is why they're doing it. These are the kinds of things that they'll say. This is the environment they're in. That's really important here. It's not just the conversations where students are talking to each other. These are people in the military. You use certain terms, certain tools. So you want to set that up pretty clearly. And then in terms of the actual excerpts and examples and things that we're talking about, they can get quite long. I think the way to handle that is to split those up, meaning clear, right? So you split them up so you're not going through large chunks of data all at once. And while you're doing that splitting up, and when you're conveying that to the reader, you're kind of interpreting that as you go along and you're summarizing it as you go along, there is a piece of this overall thing that I'm trying to show you. This is what that means. We're going to the next one. This is what this means. And then when you get to the end, you kind of put it all together and say, overall, we've seen this story, right? It splits it up and chunks it up for the reader and also gives them these waypoints and along the argument to see what we're building towards and what we're getting at. I sometimes think of it or sometimes tell people when they're writing, almost imagine that your reader wants to be able to skim over stuff and that when they kind of get to the end of it and realize that they're at the end of a piece, be able to read one or two sentences about what they should have known and then just keep going when they want. I guess I would add that maybe with data that's from a more familiar context, you might be able to get away with longer sections or longer excerpts. But for this, you read the first line and it's like, what is that? What are they talking about? So you could really stress out or tax your reader if you put like 20 lines of text where people are talking about helicopters that nobody knows what they are or something, right? It kind of depends on the data as well, I think. No, it's a good point. Do you think you have to do the same kind of packing and unpacking with quantitative results? And if so, is that similar or different from packing and packing qualitative data? I think so. And so this is an interesting question, because I don't tend to actually think about it in terms of unpacking and packing, I think of quantitative results. But I think about when I'm writing and when I'm presenting things, I think that there is a similar thing going on. One thing I would say is that when we're talking about packing, citing the context, this is what we're going to see, right? A lot of that happens in the methods. These are the methods that we're going to apply. This is how they work. This is the data we're applying them to. And I kind of see that as a similar role, the packing with qualitative data. And then when you get to the result section, right? Depending on the length of the paper, it's a conference paper, you might just dive right into the results, because we just saw the methods before our context is kind of still in our minds. But if it's a longer paper like them one then we're talking about specifically today, you'll notice in some places I reminded the reader, here is what we're going to do. These are the methods we use. Here's what we're about to see. And then you go in and start diving into the result. So there is a bit of packing there. We were looking at the counts of the commanders in one condition, versus the other. We use this method, et cetera. Then you start going into what you found. The unpacking is interesting because I think it is actually very similar to what you do with qualitative data. It's just that the reference is different. So in unpacking qualitative data, your evidence is text, some structured text, typically the ordered or some discourse. But for quantitative data, the reference can be a data table, regression results. It could be an ENA graph. And those require different ways of going through them. You can kind of go through a regression table because it's very structured and ordered and highlight what things mean and why it's important and interpret the results. But an ENA graph is less structured in the sense that you can kind of decide where you start where you start to unpack. But you're doing the same thing. You know, you're saying, this is this result, this is what this means related to my overall claim. What do you think are the key parts of writing up an ENA model? What are the pieces that you have to make sure you do? It's kind of related to the last question. I think it starts in the methods. So having a clear description in your methods about what you're going to do. And by that, I mean that not necessarily explaining exactly what ENA does, depending on the audience, you might need to do that. But there are lots of things you can do with the ENA, right? So you might be just focusing on this to just comparisons of the plotted points. For example, or you might be focusing on the networks, or you might be doing the networks attraction, or you might not be doing some of that, right? So saying in the methods, like these are the pieces of that ENA afford just to do that I'm going to be using that you're going to see later, right? Aligning that up. Then when you get to the to actually writing up the ENA model, you know, I think about it in terms of interpreting the dimensions, because I personally think that's the most important part. So interpreting the dimensions of the space, if you're showing network graphs, which is fairly typical and you're showing, like a subtraction or a difference between groups or individuals, it's always good to describe the individual networks first and show what's going on this one and what's going on that one. And then when you get to the subtraction show the differences between them, these things are going on in this group, more or less than this group for this individual about individual. Then again, this is kind of dependent on the study. You'll probably talk about some statistical results, some statistical tests that you're doing on these networks are plotted points. And then I think no matter what you're doing if you're using ENA, you want to relate these results back to some qualitative understanding that you had of the data. So typically I like to start with a qualitative interpretation of the results. And then the quantitative stuff comes after and then I relate that quantitative stuff back to what we saw qualitatively, look to link that back for the reader. Like this is all about this one thing, this quantitative evidence is supporting what we talked about earlier. Don't forget that. Why lead with the qualitative and then the quantitative? Why not just show that this is this general pattern and then give an example? That's a good question. And I wouldn't say that you maybe always need to do it, but it's kind of something I like to do because I actually think that seeing a qualitatively is a much more like really no the word like a grounded or this roller concrete experience for the reader. You can actually see, okay, this is what's going on for these subjects in this experience. I can actually imagine what that is. I can see what they say. Here's this is this that makes sense to me where it starts to break down for some people is like, okay, well, they could have just shown me these couple of things and I don't really know how extensive this is in the data at large. So that's where the quantitative stuff comes in. For me, it's backing up what we saw before. So if you were to do it the other way, it perhaps a bit less clear about what the quantitative aspects of the model are staying. for what they actually mean in relation to the real data. And so then there could be some confusion for the reader, like, okay, here's all these numbers or network graph or something. And I'm really quite note that means I know there's this code and there's some words or whatever. And then you read it later and like, oh, okay, and then you have to keep going back. And I think it's a more structured experience for the reader when they see the real context, and then they can relate that for themselves to the quantitative model. Now, how you do it in practice is perhaps different, and then you can construct the argument and the story for the reader. In some ways that touches on some of the things that we're reading about this week in the sense that it highlights the role of the quantitative relative to the qualitative. That is the qualitative discourse is sort of where the claim gets made. And then the quantitative discourse is what warrants theoretical saturation. If you do it the other way around, it's weird to warrant theoretical saturation for something that you haven't actually claimed yet. See what I mean? Yeah, I mean, and I would say at least the way that, you know, even if you look in QE, like the book, the way it's described, it starts with qualitative, for the most part, if I remember correctly. But I mean, just in the definition of QE, it's about warranty and qualitative claims with quantitative techniques. So in my mind, this qualitative piece, I wouldn't say it's maybe more important, but it's at least what I go to first or think about is like what I'm actually getting at. This is the thing that we're trying to explain and then get some evidence for. It's qual forward. Yeah, and there you go. What do you think of this the relationship between the codes, which of course are the nodes in the model and the story that you're telling? So when we think of models, we tend to think of quantitative things like we think of DNA, we think of turning things into numbers, we think of those kinds of things. But you can have, I guess, like a qualitative model as well. Like a model is just a re-representation of something that's easier to think with the use that takes the salient aspects of whatever the thing is that's being modeled. So I think of codes as a kind of model for the qualitative data or a way to represent what's going on in qualitative data that's easier to understand and work with. And it's easier to understand and work with for the researcher, I think, because here's this massive qualitative data and this phenomenon trying to understand. It lets you break that into manageable pieces in the beginning. Oh, I noticed this is going on. That seems important. I noticed that's going on. That seems important here. Let's categorize things, group them together according to those things. And then you get to a space where like, okay, well, it's not that these things are just important or they're happening. This code or this type of thing is actually related to this other thing or connects to this other thing in one way. When this happened, the other one happened or various ways of things can be connected. So then you get to that part of the story. There's actually links or relationships between these codes. So for me, it makes it easier to uncover the story. It lets me start breaking down into a smaller pieces of our problem and then see how they're related later. Also, it makes it easier to communicate that story, I think, to the reader. It lets them go through that process. So if you're these pieces, here's how they're connected. Also, let's them contest them in a certain way, right? If you just told the qualitative story and you didn't maybe have these pieces of the puzzle that you built up before that people could point to and you pointed to and looked at. Then they can't really say, if it's there, they can say, oh, yeah, I'd see what you mean by that. That looks like it catches up or they can say, no, actually think that that means something else and makes it a contestable thing that you can actually argue with as opposed to just something fluffy. It has to provide the relationship back to some theory or a theory, right? It should. So your codes, you can get them in multiple ways or you can just do a grounded approach and see what's happening in the data. That even then that that's always going to be related to some theoretical framework or something that somebody has looked at before maybe even completely. So that's a way for you to link back to that theory and sometimes you can just get your codes straight from the theoretical framework or from what somebody else has done before as well. And so yeah, it makes that connection very explicit. Given that you've come up with these codes to tell the story, how do you decide what codes to use in your ENA models? You might have 15 codes or 20 codes that you use while you are exploring the data and trying to make sense. That's not an easy question and part of it comes with experience and experimenting and testing and looking at things. But I mean, I think initially when you're coming up with your initial set of them for me, it's typically starts with a grounded approach and I say that where I'm going into the data and looking, but also I know of theories that might apply here. I know what realm within which this data is talking to in the larger community. So I have that stuff in my head too. I see things that are salient in the data and you have things that are perhaps theoretically a length that you might be bringing to bear on the data too and then you get this whole set of things that are happening. So then you might start to decide which would then to put into your ENA model and I think for ENA you want to not have too many codes right you can kind of things can get. So I'm going to be very interpret if you have over say, you know, 12 15 codes I typically like to go for like 8 to 10 maybe. And the way that I think about that is one I think about which of these is seem to be standing out to me as most important or most related when I was looking through this qualitatively right and that might call the field pretty quickly. I saw OK, this thing was going on but they didn't really seem that important to me or what there wasn't. Maybe it was a way and a small subset of the data that I saw or what have you the other thing you can do is when you actually start building your ENA models will notice that so ENA is about maximizing differences between people in terms of the way that they're making connections in their discourse. So not all codes and connections do that equally and you can actually see in the model when some codes and connections aren't distinguishing people or groups or have you. They tend to be in the middle of the model instead of at the extremes because the codes and connections at the extremes are the ones that are doing the most distinguishing. So that can suggest to you maybe well doesn't seem that this code or these connections are playing an important role in distinguishing these participants what happens if I take them out. Sometimes when you take them out the model isn't changing that much at all sometimes you do it does you have to kind of make this decision. Yeah that's how I think about it's an iterative process and QE people talk a lot about this notion of iterative models could you say a little bit about like why the iterativity is important. Lots of reasons. I mean sometimes you just don't you don't know what you never really know what the right model is going to be or what what the best model is going to be or what the most testifiable model is going to be. There's this aspect of you try something and then find out later that it doesn't work for some reason either you your operationalization of something was poor or you work something in the data or some other variable appears to be affecting things here so you kind of redo it. It's a research process right so you're iterating and getting to something that seems to be meaningful but at the same time by doing it iteratively it helps you understand your data better. So you start with your data or at least the way that I do if you start with your data you have some hypothesis about what's going on you develop an initial model you get some results from that model and then you're not just done right you go back from those results back to the original data and see if that actually makes sense and that those are entitlement sometimes when you do that you realize oh this model is missing something. Because I think we forget that there's quite large time gaps between coming up looking at your data and seeing what's in it and then making your model sometimes it's not just 10 minutes down the road or a day down the road it could be weeks or months whatever so you get your model you get a result and go back and validate that model is close that loop and you can change things. The other thing that I would say is that it's just practical I think we talked about this a lot and I think it's kind of maybe your approach to a lot of things but you get a pipeline in place you get an analysis pipeline in place that's easier to iterate from later so instead of trying to get it perfect the first time which can take a really long time and you might end up changing it later or this right you start with something simple. You get it going from soup to nuts it starts it gets an output you can investigate it you can change it and then once you have that even if it's likely too simple or not not the right thing you have something you've gone through a whole bunch of processes you've looked at your data you've developed some more understanding and now that you have that thing you can improve it you can start building on it improving instead of trying to just get that thing right from the first place which can take a really long time. So that's the whole model zero approach yeah when you're doing this iteration though how do you iterate and avoid either just cherry picking qualitative data or doing the conceptual equivalent of peaking to me and cherry picking and peaking I think are slightly different things when I think about them so and actually think that quantitative of the biography is kind of a guard against cherry picking so you can pick the best examples that you want right and show them but then if you show your quantitative results and they're not. Showing that this is a natural pattern in the data then like you're kind of shooting yourself in the foot to protection against that the whole point is to show evidence that this isn't just something that you picked out of your hat or like pick the best thing. The peaking thing is a bit more interesting because I think it kind of depends on on your intent and on also have just a fileable you are like so I was reading back through the chapter on saturation right and you have this example where I think the model zero is using counts of the same. So there's not a statistically significant result between the two groups you're like the first half in a second half so you do a second model where you use the percentages and this is makes sense right because some conversations have much more talk in them right so that the actual rock counts are can be misleading and then we use the percentage. I remember correctly, there is a split-significant result. Now, if you look at that in one way, somebody say, "Oh, that's v-acking, right?" You did a result, you got to get something significant, so you changed it, you did, you're going to report the significant one. Setting aside the fact that you reported both of them in this, because we were making a larger point, that decision that you made to go to percentages rather than rock counts was justifiable and more aligned to what was actually going on in the data. So that's not v-acking to me, that's just good research. You looked at your results and you were like, "Oh, wait a minute, I go back and look at this, actually this other thing is going on." Maybe you didn't notice it before that can happen. Maybe you intentionally didn't put it there, because you were just trying to get your model set up so you can run it. But you made that change for a reason that's based on the data, or based on theory, or just common sense. Like I said, that's not v-acking to me. That's making just file decisions to align your quantitative results. That's a little bit tricky, because you're trying to explain the data that you have, but the reason that we're using statistics is because there is some generalizability aspect of it. It's just the population that you are generalizing to is different than you are normally conceptualizing a typical statistical analysis. So, you're trying to explain the data that you have in a way that makes sense. It's just the population that you are generalizing to conceptualizing a typical statistical analysis, right? You're not generalizing to some other population of similar people. You're generalizing to what if we kept observing these people or what if we were able to repeat this over and over again for these people and keep going, would we see the same thing? Is this a matter? So, there's still a generalization. It's just where it's going. And we decide what that is as researchers essentially. That's the whole idea of the hypothetical population, right? Yes. This chapter is about theoretical saturation, exchangeability, hypothetical populations. That's a very mathematical, statistical, conceptual perspective. How central is it to QE, either as a method or an intellectual enterprise, but also how important is it to somebody who just wants to use QE and kind of move on? I think it's definitely central and important to me. It's QE, right? It's not QE or E. So we're talking about it being qualitative forward. And you can do qualitative research, whether it's in bites off, and that's fine. And you can do quantitative research, and that's fine. But this is another way of looking at that. It's a way of providing additional evidence or warrants for your claims either way. And so to do that in this framework requires using some mathematical perspectives and some quantitative ideas. That's the whole point of it. We're literally using quantitative methods to help warrant qualitative claims. So yeah, you have to touch into that if you want to actually use this method meaningfully. That being said, you could use it without knowing that you're doing that. But that's a different problem. I suppose. Yeah. How do you explain those things when you're talking to somebody who's either new to QE or when you actually stumble into a place where those make a big difference in terms of how you're interpreting a model or constructing a model? Coming at it from this hypothetical population thing helps a lot. Thinking about what we're actually generalizing to. So again, we just talked about this right? There's a difference in what we're generalizing to and what you might be in a typical fiscal analysis. But I think the key for me is like thinking that that's just the way that we talk about that. And I think what the really interesting thing about QE for me is that it's really just a redefinition. A redefinition or a redistribution of theoretical saturation. That's the whole point when you're a qualitative researcher. You are making a claim that the Sustereit would saturate it. But the problem is you have to trust that. How is the researcher warranting that to the reader other than giving an aridite explanation of what's going on? They can't in just in just the document right. You have to talk to them and they have to keep showing you stuff. Here's this. Here's this. But if you have this quantitative piece, you can show it and put it right there and it's contestable and you can talk about it. Here's the other evidence for this. What I was saying isn't just cherry picking. There's statistical evidence for this. That's kind of how I talk about it with other folks. I mean, we can get into exchangeability and things like that. If you want to, that comes up less. I think for some reason, I don't know why. Since people are reading about it this week, let's talk a little bit about it. What is exchangeability and how do you know if your data is exchangeable? I don't think it's an easy term to understand actually. It starts off with this reordering and things like that. The way that I think about it is just in terms of statistical controls or compounding variables or covariates. Your data is exchangeable when you've controlled for accountant for reasonably, to some extent, the other compounding variables that might be that plate here. If you consider this variable, this variable, this variable, you hold it constant. Then you look at these people or these group subjects. They're essentially the same. There's no other salient difference here between them. So differences that we're going to find are in terms of this thing that we're interested in. That's the claim you're kind of making. You're looking at data from people in this one team. Within this team, is this still happening? Then across teams, is this happening? That's what you're doing. For me, it's just a lot much easier to think about this in terms of control variables. The conditional exchangeability in some ways is more important than exchangeability itself. There's an argument to be made that everything is all conditional. You can never really be completely independent or exchangeable, or you can never really know if you are. So you do the best you can based on what theory says and what are their variables you have at hand. Give your best efforts. Give them all up to on it. That's what the limitations are for. Yes, there's this other variable. I didn't collect. Sorry, I didn't have. Yes, there's this other thing. This is what we have. This is the best we can do. We can test that later if it seems important. You've helped a lot of QE researchers over the years. What's the thing that you think people need the most help with? Where do people seem to get stuck? So on the surface, DNA is actually quite simple. When you think about the connections and the networks, that makes a lot of sense to people here. There's connections, there's the strength of them. What's difficult for people is the other stuff that happens outside of the network. The dimensional reduction and the interpretation of the spaces and what the plotting points are. So I think that's the thing where the people get tricked up on a lot. The relationship between the network, which is kind of easy to drop. And this space and these points and things like that. I find myself having to talk a lot about it. For me, the other piece is easier for me to understand the points and dimensions and stuff. Because I think about it that way or have some other training and statistics that I can relate that to. So that's one thing. The other thing I would say is forgetting the ethnography piece. A lot of people that I come to me, I would like to use the NAMS or need some help with the NAMS. When they do it, they either have a model, they would try to model it. They would try to do it. And that's fine as far as it goes. But it's really hard to evaluate. As someone looking at this. If it's how valid this is. If I can't see some qualitative interpretation. Also. And I don't know if those researchers have actually done that qualitative interpretation. I think that's the reason why I think about it. Thank you for having me always good to talk about this stuff. Thank you for listening to the quantitative ethnography podcast. Interested in using data to impact education? Check out the MS and Educational Psychology Learning Analytics program at UW Madison. A 100% online part time graduate degree designed for working professionals. Learn more at go.wisk.edu/learninganalytics (car engine roaring)

Podcast Summary

Key Points:

  1. The podcast discusses methods for analyzing qualitative data, emphasizing the process of "packing" (setting context) and "unpacking" (interpreting evidence) to build a coherent narrative.
  2. Quantitative Ethnography (QE) is highlighted as an approach that uses quantitative techniques, like Epistemic Network Analysis (ENA), to warrant qualitative claims, often starting with qualitative insights before integrating quantitative models.
  3. Effective communication in research involves structuring findings for clarity, such as chunking complex data, interpreting ENA dimensions, and iteratively refining models to align with the data and theoretical frameworks.

Summary:

This podcast episode features a discussion between host David Williamson Schaefer and guest Zach Swikke on navigating and analyzing data responsibly, with a focus on qualitative research methods and Quantitative Ethnography (QE). Swikke explains the process of "packing" qualitative data by setting context and "unpacking" it through detailed interpretation, linking evidence back to overarching claims. The conversation explores how qualitative insights can be supported by quantitative techniques like Epistemic Network Analysis (ENA), which models connections in discourse to distinguish patterns among groups.

Swikke emphasizes starting with qualitative narratives to ground findings before introducing quantitative models, ensuring results are meaningful and contestable. Key considerations include selecting and refining codes for ENA, iterating models to improve accuracy, and structuring research writing to guide readers through complex data, especially in unfamiliar contexts. The episode underscores the value of integrating qualitative and quantitative approaches to build robust, interpretable analyses in learning analytics and beyond.

FAQs

The podcast explores how to navigate responsibly and find meaning in the vast ocean of big data, focusing on quantitative ethnography and learning analytics.

The host is David Williamson Schaefer, and the guest is Zach Swikke, a lecturer and researcher at Monash University, discussing collaborative problem-solving and network analysis.

Packing involves setting context and stating the main claim before presenting evidence, while unpacking involves analyzing the data line by line, referencing codes, and relating it back to the claim.

Split the data into clear, manageable chunks, interpret and summarize each part as you go, and provide waypoints to guide the reader through the argument.

Yes, packing involves setting context in methods sections, and unpacking interprets results like tables or graphs, relating them back to the overall claim, similar to qualitative data.

Start with a clear methods description, interpret the dimensions of the model, describe individual networks, show differences between groups, and relate quantitative results back to qualitative understanding.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.