Go back

Why AI Detectors Don't Work for Education

18m 49s

Why AI Detectors Don't Work for Education

The podcast discusses the evolution and challenges of plagiarism detection in education, particularly with the rise of AI-generated text. Traditional systems, like Turnitin, effectively flag copied content by comparing submissions against existing databases but fail when text is paraphrased or rewritten. The advent of generative AI introduces a fundamentally different problem: detecting whether text is originally produced by an AI, not merely copied. Current detection methods include watermarking (embedding hidden signals in AI output), statistical style analysis (evaluating perplexity and "burstiness" in writing), and process tracking (monitoring keystrokes to identify human vs. AI patterns). However, these approaches are flawed; watermarking can be degraded by editing, statistical tools yield false positives, and students can use paraphrasing tools to evade detection. False accusations carry serious consequences, undermining trust and disproportionately affecting certain groups. The hosts suggest shifting assessment strategies—such as using in-class writing, oral exams, or comparing work against a student's prior submissions—to reduce cheating opportunities and improve learning outcomes. They critique the overuse of essays and emphasize the need for institutional changes, though implementing alternatives like oral assessments requires significant effort. Ultimately, a balanced, human-in-the-loop approach is recommended over reliance on imperfect automated detectors.

Transcription

3854 Words, 21472 Characters

English
(upbeat music) - This is the Ed Technical Podcast, where education research meets AI. - Hey everyone, I'm Owen Henkel, a former teacher and your friendly neighborhood AI researcher. - And I'm Libby Hills, another former teacher, and now an Ed tech funder. And critically, I stop Owen from going to sci-fi on this podcast. - Good luck with that. - But our goal on this podcast is to bring you conversations with some of the smartest people in the field, and help you go through the hype to understand what actually matters. - As well as podcasting, Ed Technical also researchers AI in education and funds promising Ed tech companies. Owen, should we get started? - Let's go. (upbeat music) - Hey Owen. - Hey Libby, how's it going? - Good, I'm good although I am shocked that I just found out that you cheated at your PhD, it sounds like you were telling me. - That is absolutely fake news. - Josh, my supervisor, Libby's line, don't believe her, if you're out there. - Okay, set me straight, what actually happened? - When I applied from a PhD, they asked for previous writing samples. I gave them one and it got flagged for plagiarism. And basically this had been a group assignment and I submitted this section of this that I had written and a master's program and apparently someone uploaded it to some repository or something. And so then it got flagged by one of these automatic plagiarism checkers. We have an issue with your application. There's potentially an issue of academic dishonesty and plagiarism and I was like, oh my God, freaked out. Like trying to figure out what was happening. Very, very nervous. But what was actually helpful is they were able to produce the document with the side by side comparison, which showed what was getting flagged. And then I was able to kind of do some detective work and basically traced back a couple of years ago when I was in this program. - So it sounds like you're a living grieving example of one of the downsides of plagiarism detector, which is this idea that people who actually haven't cheated, get flagged, that creates a pretty difficult situation. - One of the key points was there were some evidence so we were able to figure out what happened. I kind of had a chance to explain what happened. But if they had just gone on some type of confidence score and hadn't told me what had happened, I wouldn't have been accepted in the program. It would have been like, yeah, it would have been a big deal. And I wouldn't have known what would have happened. - Yeah, well, in case anyone hasn't guessed, we're talking about plagiarism and plagiarism detectors today as it's a pretty hot topic right now. There've been a couple of pieces earlier this year about both how students are using AI to complete homework and assignments, also some of the challenges with actually detecting whether students or how students are using AI to write these assignments. So we thought it would be time to do a bit of a stock take on what's happening when it comes to plagiarism detectors for this episode today. - Obviously the use of these is pretty widespread. There's some estimates put up to like 90% of college applications now have some element of AI assistance in the spirit of ed technical. You hear that a bit? Do you see what I did there? - Yes, you said the name of the podcast. That is well done, well done. - Thank you. There is this dimension of what does this mean for institutions, the ethical relations, but also this deeply technical element about how you detect plagiarism and how it's changing with generative AI. So it seems like a great short topic for the show. - Cool, let's dive in. (upbeat music) - And we're back. - So maybe a quick brief history of plagiarism detectors. So my understanding is that the prior plagiarism detection systems were really based on copy matching. So looking at has a student essentially copy and pasted text and they have a database and they compare newly written submitted assignments with that database to see to what extent, what is in the assignment matches, what's already out there. Is that right, Owen? - Yeah, absolutely. I think you know that I think a funny anecdote when I was doing some background research on this turning in, which is one of the canonical plagiarism detectors. Apparently it was a whole bunch of grad students, maybe at Berkeley or one of the California schools who got like really mad at frat bro undergrads, turning and obviously plagiarized essays. And then they decided to spin out and create the startup. But I love the idea of kind of a disgruntled humanity's postdoc, just like hating the frat bro's obviously cheating and that being the motivating incident for the creation of a major company. - And not all of our listeners might know what a frat bro is, but I'm not gonna go down that rabbit hole right now. So let's go back to plagiarism detectors. - Don't worry, say for work. It's just, it's just, yeah, just Google it. I love the cross-cultural exchange on how technical it is, it's great. But moving on from frat bro's, so I think, turn it in and some of these plagiarism detectors were pretty good at flagging and finding. I think we can term it lazy copying, right? Like students who have copy and pasted from prior work and dump that into their essay, but they aren't great at detecting plagiarism if things have been like we written or paraphrased. So good for the kind of obvious stuff, but not able to detect all types of plagiarism. - And so it doesn't have to be verbatim, there's all these kind of fancy natural language processing techniques where it's like, you know, oh, are nine out of 10 words the same. And so you kind of look at more paragraphs, so even if someone does some very minor editing or like takes one paragraph and moves it into another part of the paper, those could still be detected pretty easily. As soon as you start paraphrasing or translate your work into one language and then back into the first language and then that dual translation process stuff just gets shifted around enough. If you think of this idea of adversarial cheating, where the person's cheating, they know someone is looking for them cheating ahead of time, performance of these really, really drops. - It's good at the obvious stuff. One of the reasons it's good is because it's this comparison of something that already exists as a ground truth that an algorithm's able to compare new text against. But now we're in this new paradigm, right? Where the question's different. So what we're trying to find is a tool or a way of detecting has a student use an AI tool to generate new text. Not has a student copy and pasted or lifted existing text and put it into the essay. So it's actually like a really different type of plagiarism that we're trying to find a way of detecting that, right? - Yeah, 100%. This isn't just like, oh, it's a little bit harder because of gender and a AI. It's like a completely different question because as you said in the past, by definition, if this thing existed before and then you give me something now and they're extremely similar, I basically know what you've cheated and I kind of have proof. Whereas with gender, it's like the thing you're given has never been created before. - Which has always been tricky and has often relied on like human judgment and maybe we can put a pin in that and come back to it because I really like looking into some of the details around what the current AI detectors do. And so as you're saying, they offer these probability scores. So to what level of confidence do they think the text is AI generated? And they look at complexity so how predictable the word choices are and burstiness, which I love, which is variation in sentence structure. - I think you're really excited about being on the use of word burstiness in a grammatically correct sentence aren't you? - That's great. - I would just back up quickly just for the listeners. Libby, you gotta head of yourself there over keen students, slow down. Let's go. Let's wait for the rest of the class. - I just wanna talk about burstiness. - One quick thing to highlight is that I think what's also unprecedented is in the past, a lot of this was only really applied to essays, whereas now it's all kinds of assignments. He's gonna do math problem sets. There's a much wider range of it, types of chain that could be possible than before. But anyway, so the fundamental question now is, did AI generate this or not? There's three general areas where you can do AI detection. One is called water marking. The model subtly inserts signals to itself that it is not a human-generatingness. This would be weird spacing or using unusual, but non-eye catching synonyms, et cetera. Second was style detection. This is Libby's burstingness, kind of doing statistical modeling of how humans write or you just look at all the students' previous work the way it may be a really engaged teacher might do. Or you use a lot of context about what you taught in class and like, how would a student know about this? There's also this third process tracking where you watch the kid write it, or you basically have them interreportal and they have an hour to write the essay, but you can keep track of the keystrokes. So just, there's water marking, statistical or stylistic detection, and then process tracking are three buckets that are helpful to think about. Cool, those two buckets make sense. I guess there are still some problems within those, right? Watermarking, so I understand it's like model providers, so that they're able to detect their own texts generated by their models, but presumably any type of paraphrasing or adjustment degrades their ability to do that. And also the teacher doesn't have access to that, right? So it's not like they don't have the secret code and if you give the secret code away, then it's no longer a secret code, right? So it's like kind of, it's really only helpful for like high stakes stuff. A model provider could be like, oh, this sensitive concerning document, there's gone viral, yes, that was generated by CROD or chat GBT. And then we can track down who generated it in our database and hand it over to law authorities or whatever, yeah, yeah. Okay, so not helpful for like teachers or professors in that. Yeah, unless the school is really serious about cheating. Any other options? What do you think is particularly promising on? It's a statistical detection. It kind of sounds good in principle, but there's lots of failure rates. And again, as soon as you assume that the students know what they're doing, it really falls apart, right? There's even tools now that will rephrase an LM-generated tax to make it sound more human. They'll like rearrange this intact, so it will kind of pass some of these bursty and perplexity measures. There's this important issue about also false positives, accusing a student wrongly as like pretty serious consequences. I'm sure folks can imagine. Yeah, I mean, I was kind of teasing you at the start of the episode, but in all seriousness, if there was a system being used that automatically flags concerns and that leads to a real world consequence for someone not getting into a program or someone's essay being pulled and being fined out of a program, that's a pretty serious, unfair consequence that's happening that people are rightly really concerned about. I mean, opening a eye, I think, pulled their own detection tool, right? Because they just had such a high proportion of false positives. And I think there's also data to suggest certain groups of students get unfairly disadvantaged, even more than others do as well. So I think it is a really serious concern that folks have in the space. You can destroy a relationship with a student if you get that wrong. And it's just unpleasant. Most teachers don't want to be the plagiarism police. It's just a real bummer, right? You have to have a really good tool to get people to actually use it. There's a paper recently that Eat the Mole 'cause a big influencer tweeted out and then all bunch of people showed up. I usually it was to defeat. And I'm just like, I think we should kind of like forget that. I don't think it's gonna work. I think it's gonna be a cat and mouse game. Yeah, let's come back to like process detection. I'm hoping here we can talk a little bit about this idea of what actually really great detection system in the past has been teachers knowing the students and being able to make a call on the relative likelihood of student cheating based on their prior work. But obviously that's really hard to scale. So I'm wondering if there's some sort of solution that can help us with this. I'm more optimistic about this stuff. I think there's kind of maybe a few sub buckets here. There's oldies but goodies I call them, which are just like doing it live, you know, or examinations, or just writing an essay in class. You can't really fake those. And so that just completely even eliminates the option to cheat. So I guess those just coming back to your like framework you outlined at the start of our technical practical, you can rely on them with high confidence in terms of their ability to prevent cheating, I guess, or detect it. 100% if it's a lot of people. But practically it's really tricky, right? Well, I don't think it's that tricky, but we'll get to the second. Friend of the podcast, Adam Boxer, and some folks at Carousel were doing some interesting pre-work on this, they were chatting with us about. And you can still tell obviously I generally answer because like if you're a sixth grader and you ask them a really simple historical fact, and they come back to you with a two paragraph answer. That's with bullet points and like very high vocabulary, you kind of know. And so there is this element of a teacher knowing their students, which I think for younger grades is like relatively reliable 'cause of this passive knowledge of the student abilities, and also what you've taught them in class. That can be pretty helpful. I think this is good. It can also be time consuming. If a teacher's really on top of these things, I think it can be really useful. The issue is like you can't automate it. It's always contextual, right? So I think it's a really good approach and there's ways to make it more reliable and excited to find out more about exactly how they're tackling this. It is kind of like good common sense. - Okay, so we've got some approaches, complicated, difficult to apply, some limitations with them. So any other options, what do you think is basically promising on? - The third buck is just key stroke data. So if you just think about how an LN generates texts at the character level, it just spits them out almost uniformly. Whereas if you watch yourself type, obviously you're constantly rephrasing, going back. In my case, correcting typos and grammatical errors. If you kind of do the time signature of how human types versus a computer, they're completely different. And it's actually pretty hard to replicate that. You give someone a couple hours, but your essay has to be typed into the equivalent of a Google Doc or something like that, where you can just track the key stroke data and you can just see if someone copy and paste it. And so that would have a very low error rate. - But there's also scope potentially for the database of students' work and then a tool that compares a new essay with that prior work, right? To come up with a prediction for whether the student has actually written themselves or had extra help with that would that be feasible? - It would definitely be feasible. I think again, there'd be some of the concerns about false positives there. 'Cause like, what if the student gets a lot better? Right, what if they just say, you know what I mean? Like, there's these tricky edge cases. Obviously, again, if it's a sixth grader who comes back with a 5,000 word essay that's obviously been written by a PhD student, it's pretty easy. - But it would be like a human in the loop system, though, right? It could be a system that would just flag if there's just a really big gap, right? - Between the different potential. What is different, there's like, you do have a ground truth to a certain extent, which is the student's prior work, which is really important. And like, that's how you base the conversation and you don't have to be an accusation. You could just like go and be like, "Hey, this is a lot better than anything I've ever done." Like, what have you done differently this time? And I think the key is like, you really need to stay away from confidence scores. I think those are really dangerous 'cause they're never actually reliable. They give people this false solution of certainty. We have really good solutions already. We just need to use them. - I'm getting my soapbox, are you ready? - Well, what were you on before? - Yeah. (laughing) I honestly think that like, for certain subjects, student essays does allow a certain type of thinking and it does teach writing skills. I think that they are wildly overused, both in high school and in college. They're not actually good most of the time for student learning and they're not particularly effective for assessing student knowledge. The Masters in Science teaching at Oxford last year, like we did 10 classes, three hours. Their entire grade was based on one essay they wrote. It might be like, that's an absurd way to assess whether a student's learner for them to demonstrate their knowledge, but there's all this institutional and there's like, "Oh, everything has to be an essay." Like, well, why not have a short oral exam? Why not have them write in class? Why not do a few different things? Like, there's all these tools out there. I think it's much more driven by like, institutional inertia and maybe what's easier for college professors. - I mean, institutional change is not a small thing, right? I mean, these are like really big changes that would require quite a significant adjustment. - High-resisting? - Yeah, I think shifting from essay-based assessment to like oral assessments, massive change. - Yeah, what about live essays? - I mean, again, it's like less of a significant change, but still requires whole new systems to be set up, software to be designed or bought, implementation, training, buy-in. - Let's go back to blue books. It's not that much more work for the last class period of the year. We're just gonna do it in class. It's a little bit annoying, but I think in some ways, it'd actually be better for a student learning. And to me, it's so much better than just pretending like 80% of essays aren't being written, which actually be tea and then coming up with fake tools that claim to detect them but can't. - I definitely agree that a system where you've got 80, 90% of students using AI to submit essays. And then presumably you also have some lecturer's professors also using AI to help them like review a much of their essays, is not like a particularly productive cycle to be in. If we actually want students to be learning, I am cautioning that big changes like shifting to oral exams is not a small thing. I'm not saying it shouldn't happen, but I do think it's a big shift. - Levy, you're just a sell-out institutionalist. I'm the pedagogical radical here. How do you feel about that? - Institutional sell-out, great. Thanks, buddy. - That's me. - Just, I'm saying, burning all down. That's my takeaway. Any closing thoughts? - Hmm, what does students want? Do we know what students want? - There was a really thoughtful essay maybe two months ago by a professor at Columbia, but he also writes to New Yorker. It's really interesting conversations with students and it seems like a lot of times the students are just making really clear trade-off decisions. They're like, well, I kind of have to do this class 'cause it's a requirement. Teachers assigned too much work. If I can get away with it, I will. And so I do think that like in some ways, a lot of times students are also making rational decisions based on the way things are set up. - It's tough. I anecdotally have spoken to students and also young professionals who recognize that AI's offering them a shortcut that they probably shouldn't take. But at the same time, it's incredibly hard not to, if you have this incredibly powerful tool that's gonna make your life a lot easier and everyone around you is also doing the same and you're kind of competing so very tough environment, not opt-in in that type of a setup. So I would love to hear more from students around. How do they think the best way of balancing use of these new tools with actually being in an environment where they are pushed to learn and build new skills? - Yeah, for sure. Well, all right. Until next time, guys, keep it perplex and bursty. - Especially bursty, especially bursty. - Yeah, bursty. Who do we have coming up next for our next interview? Let me. - Next up, we have an interview with Candace Hodgers about the relationship between youth mental health and social media with a little bit of AI sprinkled in. So it's a really yet important topic of lots of interest at the moment. So yeah, they're excited to share that with you guys in a couple of weeks. - Yeah, super interesting. Keep an eye out for that. All right, guys, tell people by their technical, get them to us and please. - All right. - Bye, guys. (upbeat music)

Podcast Summary

Key Points:

  1. Traditional plagiarism detectors rely on copy-matching against existing databases but struggle with paraphrased or translated text.
  2. AI-generated text presents a new detection challenge, as it creates original content, leading to methods like watermarking, statistical style analysis, and process tracking (e.g., keystroke monitoring).
  3. Current AI detection tools face high false positive rates and ethical concerns, unfairly impacting students, while adversarial techniques can bypass them.
  4. Alternative solutions include in-class assessments, oral exams, and leveraging teacher familiarity with student work, though institutional change is difficult.
  5. Overreliance on essays for assessment is questioned, with suggestions to diversify evaluation methods to better measure learning and reduce cheating incentives.

Summary:

The podcast discusses the evolution and challenges of plagiarism detection in education, particularly with the rise of AI-generated text. Traditional systems, like Turnitin, effectively flag copied content by comparing submissions against existing databases but fail when text is paraphrased or rewritten. The advent of generative AI introduces a fundamentally different problem: detecting whether text is originally produced by an AI, not merely copied.

Current detection methods include watermarking (embedding hidden signals in AI output), statistical style analysis (evaluating perplexity and "burstiness" in writing), and process tracking (monitoring keystrokes to identify human vs. AI patterns). However, these approaches are flawed; watermarking can be degraded by editing, statistical tools yield false positives, and students can use paraphrasing tools to evade detection.

False accusations carry serious consequences, undermining trust and disproportionately affecting certain groups. The hosts suggest shifting assessment strategies—such as using in-class writing, oral exams, or comparing work against a student's prior submissions—to reduce cheating opportunities and improve learning outcomes. They critique the overuse of essays and emphasize the need for institutional changes, though implementing alternatives like oral assessments requires significant effort.

Ultimately, a balanced, human-in-the-loop approach is recommended over reliance on imperfect automated detectors.

FAQs

The three main types are watermarking (embedding hidden signals in AI-generated text), statistical or stylistic detection (analyzing writing patterns like burstiness and perplexity), and process tracking (monitoring keystrokes or requiring in-class writing).

Traditional detectors rely on matching text against existing databases to find copied content, but AI generates original text that hasn't existed before, making it a fundamentally different detection challenge.

Confidence scores can create false certainty and lead to false positives, potentially accusing innocent students of cheating with serious academic consequences, as seen when OpenAI withdrew its own detection tool due to high false positive rates.

Teachers can use contextual knowledge, such as comparing a student's current work with their prior submissions, assessing if the content aligns with what was taught, or noticing unrealistic improvements in writing quality for the student's level.

Process tracking involves monitoring how text is created, such as analyzing keystroke patterns (e.g., revisions and typing speed) or requiring assignments to be written in controlled settings like in-class exams to prevent cheating.

Alternatives include oral exams, in-class writing, live assessments, and diversified assignments that make it harder to use AI, though implementing these requires significant institutional changes.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.