Week 3: A Deep Dive into Segmentation with Szilvia Zörgő
29m 45s
This episode of the Quantitative Ethnography podcast, hosted by David Williamson Schaefer, features Sylvie Zorgo, a Marie Curie Fellow. The discussion centers on navigating big data and Zorgo's research into how individuals find and evaluate online health information, with the goal of creating tools to foster critical thinking. The core methodological focus is on quantitative ethnography (QE), which Zorgo describes as a way to merge qualitative narratives with quantitative data from human-computer interactions to form an integrated model. A major part of the conversation is a deep dive into the practical and theoretical complexities of data segmentation—the process of dividing data into analyzable units like utterances and stanzas. The hosts stress that segmentation decisions are fundamentally tied to research questions and coding schemes, requiring ontological consistency for modeling. They explore the trade-offs between granular and broad segmentation and highlight how defining the smallest unit (e.g., a sentence) shapes the entire analytical process, while higher-level units like stanzas offer more flexibility. The dialogue underscores that these technical choices are crucial for accurately representing the structure of discourse and meaning in the data.
We live in a vast sea of data. Information is collected about every one of us with every click, every swipe, every post, and every like. This is a podcast about how to navigate responsibly and find a meaningful place to put ashore in the ocean of big data. Welcome to the quantitative ethnography podcast hosted by David Williamson Schaefer, faculty director of the Master of Science and Educational Psychology Learning Analytics program at the University of Wisconsin-Madison. My guest today is Sylvie Zorgo, who is an assistant professor at the Institute of Behavioral Sciences at the Faculty of Medicine at Samilvice University, and more important for our purposes, Marie Curie-Fellowett-Mustricht University. She's visiting University of Wisconsin as part of that gig. Sylvie, maybe we could start. Could you tell us a little bit about the Marie Curie Fellowship and what work you're doing for that? Sure. The fellowship is three years long, and I'm currently in the outgoing phase, which is two years, and then I'll have another year in Master's. And the fellowship is aimed at my exploring how people navigate the internet when they're looking for health-related information, and how they judge the the trustworthiness of that information, and trying to make sense of that all, and then coming up with a good way to conduct an intervention to help people find reliable health-related information to help people think more critically about what they're seeing online. So that's going to be towards the end of the fellowship. So right now, we're trying to collect data and explore how people navigate the net. Super. And what, of course, Sylvie's not telling you is that the Marie Curie Fellowship is one of the most prestigious fellowships in the European Union. Can you tell us a little bit about how that work is connected to Curie or how you see it connected? The way that I think about quantitative ethnography right now is that there's data in the world. There's information that we kind of convert into data, and that data has qualitative aspects to it most of the time if we're talking about qualitative data, but it also can have quantitative aspects to it. And I think those two characteristics or those two aspects to the data, they can help you understand your data better. They can help you see more sides to the data, and they can complement each other in interpretation. So what we're trying to do right now is to figure out a new way to look at this kind of data, which is basically human computer interaction, but it's also narratives. So people are giving us their interpretation of what they're doing and why. And we'd like to integrate those two things. So what people are doing on a computer, and what people are thinking while they're doing it. Those two things need to be integrated in order to be modeled, at least the way that we want to do it. And that hasn't really been done before in this way. So QE is helping us figure out a way to see these two things together and treat them as an integrated whole. Maybe I can take a step back. How and why did you get started using QE in the first place, or maybe even just when? My masters is in cultural anthropology, and I felt like a complete outsider when I applied to a medical university for a PhD. That feeling was very much reinforced when in my first year I was sitting at an epidemiology lecture. And being an anthropologist at a medical school, I zoned out a little bit when I was looking at the visuals that they were presenting us. They were actually networks. And I was just looking at the visual aspects of those networks. And I was thinking, wow, this is like a snapshot of human cognition. If you think about the nodes as important things in someone's mental model, then you can think of these networks as a snapshot in time in how they're thinking about something. And that was in 2014. And then in 2017, when I was writing my dissertation, based on the data that I gathered in the years, by the way, I was working on why patients decide on using alternative medicine instead of biomedicine, especially when they have a serious illness. So I was looking at that topic, and I remembered this lecture and those thoughts that I was having during the lecture. And I literally just performed a Google search for networks, social science, and concepts. And I think the first one may have been social network analysis, but the second hit was epistemic network analysis. And that was in 2017. And that's been hooked ever since. So the exciting thing about that story is basically you had the same insight in 2014 that I had had earlier. Yours was an lecture. Mine was in the shower, but it's the same, basically the same thing. The rest was just working out the math in a sense. I guess so. Yeah. Although I'm still working on the math piece. So your way ahead there. Speaking about working on the math piece, you're one of the people that I think of as being, or probably the person I think was being the most thoughtful about questions of segmentation in the community. Could you tell us a little bit about why segmentation is so important? So those are two questions, I guess. So why is segmentation important? It doesn't have to be. It's important, I think, for epistemic network analysis. And I'll get back to that in a second. But why do I think it's important? It's because I was working with data from very early on that made me think of it. It was actually a necessity. Thinmentation doesn't have to be important if you're not looking at code frequencies, for example. So a lot of qualitative analyses entail the researcher segmenting the text quote unquote by finding an example, a good illustration of a point that they're trying to make or a theme that they see in the data. And kind of spotlighting that or bringing that out of the data and highlighting it for people when they're giving a presentation or they're doing a write-up. And that's sort of a way to think about segmentation, but that's not what we mean when we say that when we're working with data that is going to be modeled with ENA. So when we think about data that's going to be modeled with ENA, it should have ontological consistency. So we should have the same kind of information in each chunk of data, whether we're talking about our smallest segments or larger segments, those should contain the same kind of information. And so that necessitates rules and it necessitates systematic application of segmentation to the entire dataset. And that's very different than the first example, right? We're just getting a chunk of the text in order to prove a point. So the data that I was working with was semi-structured interviews, just I think very different from kind of the collaborative learning environments that most people in education science or most of the papers that I've read in education science deal with that kind of data where you have multiple participants and there's like a quick exchange of narratives. And with semi-structured interviews, it just doesn't look like that, right? You have one person speaking about their experiences, they have very distal references, so they might go back at the end of the interview, they might go back to the beginning of the interview and say something in reference to that. So it's actually very different when you have one data provider and that data provider is allowed to speak endlessly just like I am now. Well, you're not speaking endlessly. When people come to that sort of data, the very first thing that they struggle with is just what should be aligned. The line is the most fundamental piece of segmentation, sort of the thing you can't get below. How do you decide in that kind of data or really in any kind of data what a line should mean? That's a really hard question. I can tell you some things that I think of when I think of what should constitute an utterance. Everything in a research project, I think, should be correlated with the research questions. So at any given point in your project, whether you're designing it or already implementing it and making decisions as you go, all of those decisions need to go back to what the research questions are, right? So is this serving my answering my research questions? For example, since segmentation and codes are very much connected when you're thinking about your codes, I think it's worth it to think how abstract your codes are. Like, for example, a lot of people work with identity, like identity negotiation. So codes in that realm might be very abstract, like tension or the adoption of a practice or the rejection of a practice or change in identity. All of those are very high level. Of course, all codes are high level constructs, but this is like really abstract compared to, for example, you might say, I'm interested in how someone speaks about their identity. So I'll assume that pronouns or other linguistic features that they use to describe their identity might be of interest. So those codes
will actually be quite concrete. You have to think about, for example, how many sentences that code reaches across or spans across and look at all of your codes that way. So if you have a mixture of those codes, that's going to be very hard to determine what your utterances should be. Because if your codes reach across many utterances in one code version, and you can label one specific utterance in another code version, then that's going to create like a discrepancy. Did that make sense? Yeah, it does. That's an important point to think about. It highlights the fact that in a sense, operational, and practically, a line becomes the unit that you actually code. You're deciding at what level do I want to attach some claim about meaning to this data? You're suggesting that the smaller the line, the more concrete or easily operationalizable the codes have to be, and the more abstract and diffuse they are, you may need larger chunks of data in order to be able to make that determination. Yeah, an example. So you say, I want to think of my utterances as phrases with an ascentance. And so let's take the sentence, I'm so happy today because it's raining. Your code is happiness but weather. In that case, which one of the utterances do you apply the code to? Because the first part of the code, happiness, is in the first phrase, and weather is in the second phrase. So that's a decision that you have to make if you're coding. And if you find yourself in that kind of situation often where you're thinking, wait a minute, which utterance should I apply this code to because it looks like it spans multiple utterances? And that's the problem that may take you back to how you operationalize utterance in the first place. Yeah, that's a good point. I mean, of course, one of the ways of thinking about that is also you code for weather and you code for happiness. And then you let the model show you that weather and happiness are related. It all depends on what you want to see in the model and how you want to code. But that's exactly the point that I'm trying to make that your codes and your segmentation actually go hand in hand and they should inform each other. And what you're suggesting also relates to conversations and stanzas, which is another place for people sometimes get tripped up. In your view, how do you explain to somebody what the difference is between a conversation and a stanza? And is one of them more important than the other? I don't. I just run away when they ask me that. That's probably a wise choice. So preliminarily, if we wanted to define these things, so an utterance we've talked about as the smallest code of all, he's or smallest code of all segment in your data. So you apply your discourse codes on that level. And then a stanza would be a set of one or more utterances. And we're assuming that all of those lines or utterances are connected. They're tightly connected. And then on a higher level of segmentation, we have conversation, which is a set of one or more stanzas. And each line there can be connected. So it's basically a looser connection or a looser relationship than in a stanza. And then I don't know if you'd consider a unit a component of segmentation. But it's definitely one of those things that I think people should be talking about when they're talking about segmentation. A unit would be the totality of all utterances in a given model. So a unit is the thing that you're going to actually construct a network for in the model. So it's not the totality of all utterances, but it's all of the lines that-- all the stanzas that you're counting. If a line is the smallest unit of coding, in some ways, the stanzas, the smallest unit of counting, we attribute some kind of counts or connections or whatever it is to the stanzas. It's funny. I've always thought of the conversations as this thing that winds up being almost needlessly confusing. If your data was all sorted, sorted, so that you didn't have groups overlapping, you wouldn't need a conversation. If you just said all the data from one day is in this data file and all the data from another day is in that data file, it's only because the stuff is all jumbled up in time that we need conversations to tease out the boundaries of things. The way that I think about it is that coding happens on the level of utterances. Coacurrences are computed on the level of stanza. And then those coacurrences are aggregated on the level of conversation. OK, interesting. And the tricky thing is, the reason why you want to run away when someone asks you that is because eventually, you have to start thinking about these terms as potentially conflatable, because there are some versions of stanza window that are basically a conversation. The whole conversation is one window. Is one stanza? Yeah. Or the whole stanza is one line. Or all these other different things. Yes, yes. Or an utterance can be a stanza. So it gets really confusing after a while. And that's why we tend to talk about a generic way to think about these things. And that's why I sort of said that a set of one or more-- Yeah, yeah, I see. So yes. Given all that, how do you decide what to use for stanzas? That question is incredibly deep and complex. And it's completely dependent on your research questions and on your data. Again, this is my interpretation. But I would say a stanza window is a specific way of operationalizing stanza. Would you agree with that? Yeah, absolutely. So then are you asking about stanza or stanza window in particular? Well, I guess I'm asking a really pragmatic question. Do you decide that the whole conversation is one stanza? Do you decide that the line is a stanza? Do you decide that you're using a moving window of some-- and what size should the moving window be? I mean, if I'm a new research, I'm just presented-- I have my data. And I have to somehow make a decision about how to operationalize this part of my understanding of the structure of the discourse, right? So I think these decisions around coding and segmentation they're best made if you think ahead to what kind of model you want to see at the end. And I think maybe Amanda in medicine had the suggestion of actually drawing the model that you want to see beforehand, which seems a little strange, but it's actually very useful. So just thinking about what kinds of data and relationships you want to see modeled in the end, that helps you make these decisions. So I work with semi-structured interview data most of the time. And there you have some naturally occurring opportunities to segment, not a lot, not like, for example, social media data where you might have to many of those opportunities. But in my case, one very obvious kind of choice for utterance is sentence. Unless I'm working with very concrete codes-- you know what we talked about earlier with how they span across utterances. So unless they're incredibly fine-grained, I like to use sentences as my utterances, because it helps me get the detail and the models that I'm looking for. So if I would say, let's take an entire response from a participant to a question as my utterance, which would be, I guess, the direct translation of a turn of talk in like a collaborative goal-oriented environment. If I say, let's take a turn of talk. So an interviewer asks a question and an interviewer answers, I will at most get one instance of each of my codes. It doesn't matter how many times that code occurred, and it doesn't matter how many times it co-occurred with other codes. I'm only going to get one instance, because if my utterance is equal to a response, I'm coding on the level of utterances. Thus, I can only append one of those codes. I can make a point to append many codes of the same instance, but then that runs into some other problems. So if you use the fine-grained structures, what happens? If you take a really big chunk and make it your utterance, you're losing all the internal structuring and all the fine-grained detail about what things are connected to other things, because they're coming the same sentence or the same paragraph or in nearby in time. Exactly. Yes. And I'm kind of a control freak when it comes to coding and segmentation. So I always think of these things as being manually performed. Eventually, I would love to automate, but I don't think we're there yet quite-- I don't think we're quite there. I think we still have a lot to learn about how we do things manually. So making these decisions is easier when I have the freedom to do it manually. Of course, not all data allow.
you to do that and there's a lot of data says that are just too huge to be able to do that, but I like to be able to look at a fine-grained approach, not fine-grained, so I'm not looking at character level or word level information. I'm not even looking at phrases, I'm looking at sentences, which are again, I mean we should make the point to say that transcription has a huge part in what actually becomes a sentence. So sentences are usually my utterances and then stanzas are usually topics within a question and response segment. So a lot of people would say, "Oh, well why not just have sentences as your utterances and then just use a question and response segment as your stanzas?" But for me, even that is too big of a chunk to put too many things. So okay, some people use the word "units" of meaning or "idea units". I think one question and response segment in a semi-structured interview contains too many of these idea units and they need to be kind of parsed in order to be able to understand and label and model correctly. So when you're coding by, it was raised a couple things. Once you've segmented by a line, it is harder to disaggregate than it is to re-aggregate. And so if you wanted to pull the data together and you had broken it out by sentences, it's a lot easier to chunk it up than it is to dig in finer because the line provides your lowest level. But the other thing is that the lines are the hardest thing to change. Stanzas, you can just change as you will, right? You can change the window size, you could go to a whole stanza, you could code by hand and identify topics. But that's much more fungible. The lines are kind of the thing that locks you in stone the most. Yeah, and I've spoken to some pretty hardcore constructivist people who say, "But how can you decide what this is and then keep it a fixed thing throughout the project, whether it's segments or codes or whatever?" But as you say, once you have that basic unit, which can be operationalized in many ways in a project, by the way. So you can have sentence level utterances and then you can also have phrase level utterances. That's okay, right? So once you have that basic unit, you can group those however you want, in however many types or versions you want. So that's absolutely correct. What's the relationship between segmentation and hypothesis about the structure of the discourse? That is sort of how the discourse works. Are those the same thing? Are they related? Are they two totally different concepts? I thought about a related question a lot and I still don't have an answer to that. So I'm not sure I'm going to have an answer to this. The related question is there's a fundamental exchange ability between codes, like discourse codes and segmentation codes, if you will. You could in theory use any of your discourse codes as a basis for segmentation, also. And I think that's a really interesting and exciting idea. If, for example, I was interested in say, again, identity and how people relate to their own emotions and I'm hypothesizing that active and passive voice has a fundamental role in that. I could use active and passive voices two of my discourse codes, but I could also use them as codes to segment with. So I can label each of my utterances as either active or passive. What that does in the end to the model is I think very interesting because in the first case you have those aspects of the data as nodes in a network. And in the second case, you have them as an underlying kind of backdrop, a fundamental thing that contributes to how a network is constructed. So I think when we think about structure, we're thinking about codes in the same way as discourse codes are. It's just we're saying that these codes are so fundamental that one of these codes is always present in the data. With discourse codes, you don't necessarily have to have any of them present in the data because you can have just zero, zero, zero, is all across the row. So that in a way suggests that the big C small C big L small L big S small S that that typographic convention preserves this notion that you have some claim that you're making about the data. That's the big C or L or S and then it gets operationalized somehow in the data with the smaller one. And so segmentation is just a form of encoding, but you're using different codes, essentially. Yes, and I want to emphasize again that with segmentation codes, if you're thinking about doing that, there always has to be one of them present because if there isn't one of them present, then you're excluding that utterance from the data set. You would need to code all of your utterances to let DNA know which segment they belong to with segmentation codes you need to have that habit. But with discourse codes, it's possible that neither code A, B or C is present in an utterance and that's okay. In this class or just in general, people sometimes have pre-collected data. They're using a movie script or it's data that somebody else collected or even a data set that they got provided as a sample data set. Does that change how you think about segmentation? Necessarily, if you have a movie script, does that limit your options for segmenting or is it just convenience because it's already broken up into lines by turns of talking? So you should just leave it alone because that's easier. Recently, I was working with something that was automatically transcribed and narrative was automatically transcribed and timestamps were applied. And I think every time there was a pause over one second, the narrative got a new line and a new timestamp. And it invites you to segment based on that. Just like a movie transcript would have quote-unquote "natural segmentation" in that sense. But not all natural segmentation is meaningful. So it can help you because it gives you a good idea to "Oh, okay, I'll segment based on turns of talk" or "I'll use a scene as a stanza" or whatever associations you have when you look at that structure, that predetermined structure and the data. But it's also, it has its disadvantages too because it influences your thought. So it's hard to get rid of after you've seen that kind of segmentation, even if it's counterproductive based on your research questions. This goes back to the notion that the segmentation has to be meaningful and you're the one who has to decide what the meaning of that segmentation is. There's no automatic process and no preformatting of the data that lets you say, well, that's just the way the data was because you still have to make an active decision about what you wanted to look like. Yeah, and imagine the lure of working with open-ended survey questions, whereas just so neatly packed, right, it's one response to one question, it's there, it's maximum of 200 characters or whatever, and you think, oh, I'll just use that as an utterance. In fact, I'll just make the utterances and stances the same thing because why would I be, you know, thinking about any other kind of segmentation? But then you really should look at the data because you might have, within one response, you might have a sentence where someone says, "I agree with vaccination because it's very important." And then another sentence that says, of course, for people who have a religious conviction or an underlying disease, for them, no way we should make it compulsory. We have two fairly contradicting pieces of data here, and we might want to represent that in a way where the model also reflects that it depends on your research questions. It's justifications all the way down. You really can't get around explaining why you made any of these choices in terms of segments and codes or anything else. Yeah, it's exactly. Well, that seems like a pretty good place to leave it. Thanks so much for joining me today, and I'll see you back in the office sometime soon. Thank you. Thank you. This is a great talk. Thank you for listening to the quantitative ethnography podcast. Interested in using data to impact education? Check out the MSN Educational Psychology Learning Analytics program at UW-Madison, a 100% online part-time graduate degree designed for working professionals. Learn more at go.wiske.edu/learninganalytics
Podcast Summary
Key Points:
The podcast discusses navigating the vast ocean of big data responsibly, focusing on quantitative ethnography (QE).
Guest Sylvie Zorgo's Marie Curie Fellowship research explores how people search for and judge the trustworthiness of online health information, aiming to develop interventions for better critical thinking.
QE is presented as a method to integrate qualitative and quantitative aspects of data, particularly in modeling both user actions and their narratives during human-computer interaction.
A significant portion of the conversation delves into the technical challenges of data segmentation for analysis, emphasizing the interdependence between defining units (like utterances and stanzas) and coding strategies, all guided by research questions.
Summary:
This episode of the Quantitative Ethnography podcast, hosted by David Williamson Schaefer, features Sylvie Zorgo, a Marie Curie Fellow. The discussion centers on navigating big data and Zorgo's research into how individuals find and evaluate online health information, with the goal of creating tools to foster critical thinking. The core methodological focus is on quantitative ethnography (QE), which Zorgo describes as a way to merge qualitative narratives with quantitative data from human-computer interactions to form an integrated model.
A major part of the conversation is a deep dive into the practical and theoretical complexities of data segmentation—the process of dividing data into analyzable units like utterances and stanzas. The hosts stress that segmentation decisions are fundamentally tied to research questions and coding schemes, requiring ontological consistency for modeling. , a sentence) shapes the entire analytical process, while higher-level units like stanzas offer more flexibility.
The dialogue underscores that these technical choices are crucial for accurately representing the structure of discourse and meaning in the data.
FAQs
Her research explores how people navigate the internet for health-related information, judge its trustworthiness, and aims to develop interventions to help them find reliable information and think more critically online.
Quantitative ethnography helps integrate qualitative and quantitative aspects of data, allowing researchers to see multiple sides of the data and model human-computer interactions alongside personal narratives as an integrated whole.
Segmentation ensures ontological consistency in data, meaning each chunk contains the same kind of information, which is necessary for systematic modeling in ENA, unlike selective highlighting in some qualitative analyses.
The decision should align with research questions and code abstractness; smaller utterances like sentences allow for fine-grained detail, while larger chunks may be needed for abstract codes, but mixing code spans can create discrepancies.
A stanza is a set of tightly connected utterances, while a conversation is a set of more loosely connected stanzas; coding occurs at the utterance level, co-occurrences are computed at the stanza level, and aggregated at the conversation level.
She often uses sentences as utterances for detail and topics within question-response segments as stanzas, avoiding larger chunks to preserve fine-grained connections and model accuracy.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.