Go back

Emily M. Bender — Language Models and Linguistics

72m 55s

Emily M. Bender — Language Models and Linguistics

The conversation centers on the paper "On the Dangers of Stochastic Parrots," co-authored by linguist Emily Bender and Google's ethical AI team. The paper critiques the race to build larger language models, arguing that scaling alone risks amplifying biases, increasing environmental costs, and marginalizing smaller research groups and languages. It calls for more cautious, ethically grounded development with better documentation and testing. The paper's publication sparked controversy at Google, leading to the dismissal of researchers Timnit Gebru and Margaret Mitchell, highlighting tensions between corporate interests and ethical AI research. Bender notes that while the backlash brought the paper unprecedented attention, it also revealed systemic issues in AI governance. The discussion underscores ongoing harms, such as biased search results affecting marginalized groups, and stresses the need to view language models as tools processing form, not meaning, to avoid overstating their capabilities. Ultimately, the authors advocate for prioritizing safety and inclusivity over unchecked scaling.

Transcription

13693 Words, 74673 Characters

English
It's really important to distinguish between the word as a sequence of characters as opposed to word in the sense of a pairing or form and meaning, because what the language model is seeing is only the sequence of characters. And it's a bit easier to imagine what that's like if you think about a language you don't speak. You're listening to Gradient Descent, a show about machine learning in the real world, and I'm your host, Lucas B. Wald. Today, I'm talking to Emily Bender, who is a professor of linguistics at the University of Washington, who has a really wide range of interests in linguistics and NLP, from societal issues to multilingual variation to essentially philosophy of linguistics. And I'm especially excited to talk to her because she was actually my teacher for linguistics one at Stanford University, where I was an undergrad. And it was one of my favorite classes. I still remember it. I still remember a whole bunch of interesting facts that I learned. And it led to this lifelong interest in linguistics that I've really enjoyed. So, could not be more excited to have a conversation with her. I thought it made me sense to start with the paper that you co-authored on the dangers of stochastic parents can language models be too big, which was notable even to me on Twitter for a lot of, I guess, controversy at Google, which I was hoping you could maybe start by describing, but then get into the meat of what the paper actually says. Yeah. So, it's not in the IPA and hard to pronounce, but the title actually includes an emoji. Right? The last character of the title is a parrot emoji. And we were doing that just kind of for fun because we liked the stochastic parents metaphor. And there was a while before all this happened that without the thing about this paper would be it was the one with an emoji in a title. Like that was a little did we know. But the paper came about because of work that Dr. Timmy Gabbrew and Dr. Margaret Mitchell and their team were doing at Google, really trying to connect with the engineering teams to build in good practices to make the technology work better for more people and do less harm in the world. So, that was sort of the role that they had there. And they noticed especially Dr. Gabbrew that there was this big push towards bigger and bigger language models. Right? And so, the paper has this table like just as the number of parameters and the size of the train data just explodes over the past couple years. And so, Dr. Gabbrew would actually direct message me on Twitter saying, hey, do you know of any papers that talk about the possible downsides to this? Any risks or, you know, have you written anything? And I wrote back and I said, no, I don't know of any such papers and I haven't written one, but off the top of my head, you know, here's five or six things that we can be worried about. And about a day later, I said, do you know what that feels like a paper outline? So here's a paper outline. You want to write this together? And so that was early September and the conference we decided to target was fact, the Fairness Accountability and Transparency Conference, which took place finally in March 2021. Submission deadline was October, I think, 8th of 2020. So in a month, we put together this paper and that was possible because it actually wasn't just the two of us writing it or the four named authors finally, but in fact, we had seven authors. So Dr. Gabbrew brought in Dr. Mitchell. And it's really important to me to emphasize that they have doctorates, but I also know them all enough that I'm going to start full naming them now, our first naming them actually. So Timnay brought in Meg and three other members of their team and I brought in my PhD student, Angelina Miller major. And between the seven of us, we sort of had enough different areas of expertise and literature that we've read that we could pull together this survey paper. And so it came together and it was amazing. And also an interesting writing experience because we never had a Zoom meeting or anything where all of us spoke together. It was all done through remote collaboration in overly. So not a super common way for research to get done, but it worked in this case. So the Google authors put it through what they call pub approval over there. It got approved. We separated it to the conference and then put it away because none of us had actually anticipated working on that in the month of September. So it was like extra work for everybody. So we all turned back to the other stuff we needed to be doing. And then in late November out of nowhere from my perspective, and I just said that in telling the story, I'm not at Google. I've not been funded by Google. And so I only have sort of second hand understanding of what went on at Google plus what was out in the press eventually. But the Google coauthors were told to either retract their paper or take their names off of it. And they weren't told why. And they weren't offered a chance to sort of discuss what might need to be changed about the paper. It was just retracted or take your names off of it. And so we had this strange moment of, okay, what do we do with this paper? Because it seems kind of odd to put something out with just two authors that actually represents the work of seven people. What do we want to do here? And so my PhD student, Angie and I just returned to the Google coauthors and he said, we will follow your lead here. What do you want to have happen? And they said, no, we want this out in the world. So you two publish it. And that was initial answer. And then Timmy sort of on reflection said, actually, this is not okay. This is not an okay way to treat a researcher who was hired to do this research. This was literally her job and the job of everyone on that team. And so she pushed back. And the result of all that you can go find in all the media coverage is that she got fired. Google claims she resigned. Her team says she got resonated, which is a great neilogism. And that went down fast enough that she was able to then put her name on the paper. And meanwhile, Meg started like working on documenting what had happened to Timnay and the end result of that was that she was fired a few months later. But after the final version of the paper was done. So that's why the fourth author is Ashmarger, Schmitchell. So that's a really sad story for everybody involved. I mean, it's terrible mistreatment of Timnay and Meg and the other members of their team, those who were on our paper and those who weren't. Like it's become a, I think, a really difficult environment to work in. It's sad for Google because they lost really wonderful expertise and a lot of good will in the research community. And sort of sad for, or it sheds a light on the sad state of affairs about the way corporate interests are influencing what's happening and research in our field right now. On the other hand, my co-authors and I still maintain, we all really enjoyed the experience of working on this paper together and of weathering the stuff upwards together. And one weird result is that this paper has gotten way more attention than it ordinarily would have. I think it's a good paper, solid paper. And boy, did we put a lot of polish on it between the submission version and the camera ready because we knew it was going to be read by a lot of people. When I put up the camera ready as a pre-print, I didn't put it on archive because those tend to get cited instead of the final published versions. So I just put it on my website and tweeted out a link with a bitly link to shorten it so that I could see how many times it was downloaded. And it has been downloaded through that link alone over 10,000 times. And I know that there's other ways to get to it, which is way out of scale to anything that I've ever written otherwise. So that's been interesting as a researcher. But it's also I think fortunate because it has come to the attention of the public. And I think that this technology is widespread. It's being used. It's being used in lots of different ways. And so it's really valuable that the public at large has a chance to understand what's going on. And so through Google's gross misstep, my co-authors have been given the chance to help educate the public, which is something that I do feel fortunate about. You know, I'd love to kind of get into what the paper talks about. But do you have any sensors? Google made any comments about what their their objection was? Because I sort of had this feeling that it must be a really incendiary paper. And then the prep for this interview actually read it and it felt like pretty uncontroversial, I guess, was was my feeling reading it. So I just wonder, I mean, maybe it's hard to know, but as they said, anything about why, what they didn't like about it. So there was, yes, I mean, in public comments, there's been things like it doesn't cite relevant work that is trying to mitigate some of these issues. But at no point where we ever told which work we should have been citing, and we do in fact cite some work that is trying to mitigate these issues. So I don't know quite what that was about, but you're absolutely right. It was not, you know, we figured that we'd be ruffling some feathers with this paper because we were basically saying, hey, this thing that everyone's having so much fun chasing, maybe let's go a little bit slower and think about, you know, what kinds of downsides there are and how to do this safely. You know, there's going to be people who don't want to hear that, but we honestly thought it was going to be open AI, who was upset because we, you know, GPT-3 is kind of the best-known example of this and it was our running example too. So we thought we'd ruffle some feathers, did not realize we were going to be ruffling feathers inside Google. And it's basically a survey paper, right? We didn't run any experiments. We didn't do any analysis. What we did was we pulled together a bunch of different relevant perspectives on large language models and sort of brought them all together in one place. So it is surprising that the paper seems to have been part of the cause of Google, you know, basically blowing up this amazing asset that it had in terms of its ethical AI team. Interesting. And I guess one reading of your paper is, is hey, we should consider the downsides of large language models. I think maybe another person might read it. This might be an unfair reading, but maybe I could imagine someone having hurt feelings if they were working on large language models and they read your paper saying, it's like an unethical thing to do, to build large language models. Would that be like an overstatement of your claims? I don't know the paper in front of me, but I think maybe that I could hurt feelings, I'm not sure. - So I also do a lot of work in the space of societal impact of NLP in general, and that sometimes goes under the title of ethics in NLP. And I do see a lot of people reacting to that topic with hurt feelings. And I think it's connected with the way in which people identify with their work. And so if you say, hey, let's think about this technology we're building and how it behaves in the world and what we can do to make it be beneficial and you use the term ethics to describe that. Sometimes people want to read that as you're calling me unethical. And I think that that direction of the conversation is rarely actually valuable. And I do think that in general, people in this space want to be doing good things in the world. Certainly there are people who are working on technology with the goal of making a lot of money doing it. But I think that it's like there's this caricature of the tycoon or whoever who's just happy to crush all the little people to make as much money as possible. That's out there probably. But I think much more frequently people are working within systems that give them certain commitments around maximizing value for shareholders and stuff like that that make it harder to put on the brakes on some things that are making money right now for shareholders and take a bigger picture of you. But it is much more valuable to talk about it in terms of what are those systems, what are the incentives, what can we as individuals do within those systems rather than think about people as ethical or unethical? Not sure that really speaks to your question but hopefully it's somewhat helpful. - No, I mean, I think you're saying that maybe your point is a little more nuanced than maybe some would take away. And I can see, I mean, I think I, I mean, I guess I run a company and I love technology and I love building. I do recognize that lots of people get hurt and I think it's great that people are pointing out issues and also kind of pumping the brakes and flagging the stuff. But I just, I could kind of see how someone might feel a little offended by it. I wasn't sure if I was kind of jumping to something or like I guess my question, well, the question that I kept thinking about with the whole paper in general as I was reading it is even sort of setting aside making money, right? Just talk about research and just the sort of excitement of building models that work, which I just, I feel that so deeply. Like, GPT-3 for all its funds is kind of amazing like what it does. And it's like, I wouldn't have expected it to work so well. And so I guess would you feel, would you argue that those kinds of directions of research should stop or what would you want in organization like OpenAI to do differently? Because I think it's a good example of a place that's kind of actually really showed that bigger models, do you kind of, it's not obvious that like bigger models would perform tasks better at many extra orders of magnitude. Like do you think, would you prefer that that research doesn't happen or happen differently somehow? - So I think it's worth saying that OpenAI has actually put a lot of effort into thinking about what are the possible downsides and what could happen when the technology is released in the world and that's important to note. And I'm glad that they're doing that. I think that what I would like to see more of is first of all that kind of work. What are the possible failure modes and how do they impact people? And then also when this is working as intended, how can that impact people? And OpenAI has been doing some of that. And I think that's great and they should do more. But also there's, you can look to other fields of engineering where before you take something and you put it into the world in a place where people are gonna rely on it, there's all kinds of testing that has to be done and sort of understanding of what are the tolerances and what works and what doesn't and what are the, what's the range of temperatures that this thing could be applicable in and what are the things you have to check for and certify and things like that. And we don't have very much of that yet going on in NLP. I can speak less to other areas of AI, but I honestly I think there's similar issues elsewhere in AI. And so there's work actually that was done at Google by Meg Mitchell and Tim Nick Eberwin. Others on a framework called Model Cards, which was sort of steps in that direction. If you built a model, what does somebody who's going to use this model need to know about it? And that's the kind of thing that I would like to see more of and that is in contrast to just rampant AI hype where people build something, it's cool, it's fun, it works well and somehow that's not enough and people have to say, you know, it's not enough that GPT-3 can produce coherent text, people have to say it's understanding language, which it absolutely isn't, as I'm sure we'll talk about later. - Yeah, you're too good segue, but yeah, yeah. (laughing) - Well, it is all connected, right? - Yeah. - So for some reason, the culture around AI, is all about these like trying to reach for these big claims rather than trying to build really well-scoped, reliable, sufficiently documented that they can be used safely reliably systems. And so that's the direction that I would like to see more of, is one thing and then another thing, and we get into this in the paper is that if the main pathway to success these days is just bigger and bigger and bigger, then you cut out lots of languages communities, even within the languages that generally are well supported, because they just can't amass that much data. And you also cut out smaller research groups, smaller companies that are not sitting on the kind of collections of data that Google is or Facebook is or Amazon is. Microsoft also does a bunch of big data work, they don't seem to have a mass data quite the same way as the other big ones. And that is unfortunate because it, I think, stifles creativity to a certain extent, if the whole community is rushing towards this one goal that only some can really effectively do, then we lose out on the other things that people might be trying instead. - And I guess maybe a less obvious concern that you talk about in the paper is talking about how the models can encode bias in ways that are hard to notice. And I was wondering if you could, like I guess when you talk about the harms that might happen from natural language models, do you have examples of things that are actually happening now, or is this more of a future-looking thing of like we're worried about, as I don't think he becomes more pervasive, like worrying about future harms? - Yeah, no. So I mean, absolutely happening now. And therefore, easy to predict that it will keep happening in the future if we don't change. And here, the work of Safiya Noble with her book "Algorithms of Oppression" is a really important documentation of this. So she looked into what are the ways in which identities, which properly belong to the groups of people who have those identities are represented and reflected back to people in search. And in particular, her running example is the phrase "black girls" and also "black women." And these things have changed over time and she's very careful to document when she's talking about particular examples, what the date was. But early on, as she started this project, the phrase "black girls" as a search keyword basically turned up pornography. And that you might say is, well, that's just in the data. Well, what data? Right, where did that data come from? And if you get into the heart of her book, it's basically around that that's in the data because of the way in which the economy of the internet allows people to purchase and make money off of identity terms. And once these things were flagged, the Google piecemeal making changes. So you don't get pornography as the results for the search term "black girls" anymore. But it's also possible to poke at things and tell that it's very much individual after the fact changes as opposed to anyone going through and systematically thinking about how to redesign the way that search engines and the advertising driven ranking of search latches on to these incentives and then amplifies them. So one ongoing discussion in the AI community, you see it pop up on Twitter with great regularity is, is the problem that the data is biased only or do the models also contribute? And the answer is absolutely the models also contribute. And then there's this other layer to it of, well, that's just what's in the data. So one of the other really embarrassing examples for Google was there's a point at which Google image search turned up pictures of gorillas when you were searching for black people. And I forget exactly the particular configuration of that. But embarrassing and awful and racist. And one reaction at the time was, well, that's just in the underlying data. And so, you know, not our fault, we're just showing what the world is saying, except that it's not true, right? Because the way the algorithms that do the, you know, ranking of search results and also the bidding. for the ad words is that is emphasizing particular incentives. So there is a certain thing in the underlying data. There's also the question of how did you collect that data? Where did it come from? What does it actually represent? It is not the world as it is. It is some particular collection of data. And then what is the optimization metric? What are all these modeling decisions that you've made? And how does that interact with the various biases in the data? What is the incentive structure? So Sophia Noble's work is a great point to look. Latanya Sweeney documented, this is a 2013 paper, how if you put in at that point an African American sounding name, one of the ads that would pop up suggested that that person had a criminal history. And if you put in a white sounding name, you tend to get just a more information about so and so. And that does real harm in the world. And it wasn't that you know 100%, but it's significantly different between the two groups of names. It does real harm in the world because if you imagine someone is applying for a job or you know just making friends and someone does a google search on them. And here comes alongside this message suggesting they might be a criminal. Right. That does harm. And then if I can give one more example. These are great. Yeah. Elia Robbins' spear did a really interesting work example around sentiment analysis and word embeddings. All right. So sentiment analysis is the task of taking some natural language text and her example is English and using it to calculate or predict the sentiment. Is this a text expressing positive feeling towards something negative feeling towards something or not expressing feelings. And the particular data that she was working with I think was Yelp restaurant reviews. So there it's take the text predict the stars. Yeah, I've used that data set. Yeah. Right. And then as an external component, she's using word embeddings, which are representations of words into a vector space based on what other words they co occur with. So some of the training data is in domain that the Yelp reviews. But then there's this component that's trained on general web garbage. And what she found using these sort of generic word embeddings was that the system system had predicted the star ratings for Mexican restaurants. All right. And so she digs into it and looks into it looks into why. And it turns out that because that general web garbage included the discourse about immigration into the US from and through Mexico, which has lots of really negative toxic opinions of Mexican people. The word embeddings picked up the word Mexican as akin to other negative sentiment words. And so if in your review of the restaurant you called it a Mexican restaurant, according to the system, you have said something negative about it. So you can't possibly be giving it a five star review. Well, that's a really interesting example. And I guess I was my next question is going to be. How do models plan to this? I guess that's a good example of how not just the underlying data can can have bias, but the model can literally have its own bias. Yeah, so this is the word embeddings picked up on co occurrences between the word Mexican and lots of other things that also co occurred with negative sentiment. And then that which uses a component in this other model. So yeah, there wasn't in the underlying yellow reviews, you know, any particular reason that the Mexican restaurants were rated lower. Right. Right. And I don't know for sure if that if they were rated like on average exactly the same or but it doesn't matter because the error was the system under predicting for any given restaurant, you know, on average, it was it was missing in the low direction. So yeah, so that's that's a kind of bias that was picked up from an external data set. Right. And, you know, we tend in an LPD use word embeddings as really handy detailed representations of word in quotes meaning. Right. So word similarity, including semantic similarity. And if we don't pay attention to where, you know, what meaning was picked up, what co occurrence was picked up, then we can end up with stuff we really don't want in our systems. And I guess what what would you recommend doing about that because they are they're really useful word embedding. So, and I'm sure I'm sure in this case, it seems pretty simple of like you're actually like it's hurting your performance. Right. So it's not even like a model performance trade off here. So what could you what could you possibly do? So there is a lot of work on so called debiasing of word embeddings. And if you if you look at Spirorswork, she continues on to to do some of that. And I think that part of it is work with more curated data sets. So, yeah, you know, the discourse around immigration from and through Mexico, even if you stick with only things like, you know, reputable news sources, you're still going to find that garbage. Right. So that's in that alone is not going to solve it. But it can be better, right. This is it's not possible to come up with a fully, you know, bias free data set nor fully bias free word embeddings. But you can do better. So one step is to sort of say, okay, how much better can we do with curated data? What about debiasing techniques for the biases that we're aware of? Part of the problem with the debiasing techniques is you have to know what you're looking for. And then on top of that, to think through, you know, failure modes. So in a particular use case when you're building some technology, who are the stakeholders who's going to be impacted by it? If someone's restaurant rating is under predicted for some reason, what does that mean in an actual use context? And what should we be testing for, right? To see if we have sufficiently debiased for our use case for the stakeholders who are most likely to experience adverse impacts. I guess some, it does seem like it would be incredibly, I mean, it seems like it would actually be impossible to find a sort of unbiased data set of human writing. It doesn't exist. And I guess these are good takeaways. It's other papers that I want to talk about. So maybe we should just, in the interest of time, we should move on to the the second paper that we want to talk about to make sure we get to it, which is around, let me see if I can summarize this. So this is, this is basically sort of saying that language modeling only on kind of what you call form, which I think is just sort of like the words kind of coming through the sort of, this kind of the GPT-3 types of models that just sort of like look at these strings of words, can't have understanding, which we're understanding. And I thought, I just thought one thing that was interesting is that you said you wrote the paper, it's a sort of like end, some kind of debate on Twitter that I was definitely not aware of. And I actually, I think I'm kind of coming into something with maybe more context than I knew. So maybe you could sort of summarize what the different possible positions are here and what you want to put the rest. So I kept finding myself getting into arguments on Twitter with people who were claiming that language models were understanding things. And I was like, no, they're not, they can't possibly be. And it's important to pin down what we mean by language models, right? So a language model is something like GPT-3U or BERT or otherwise where its training data is a whole bunch of text. And the training task is predicting words in the text. So some of the times it's done sequentially, sometimes it's done with a masked language model objective where certain words are dropped out. And the training objective is, okay, put those words back in and then do your model updating to gradient and set, etc. Right? And for me as a linguist, I look at that and go, okay, useful technology, interesting, incredibly helpful in things like speech recognition and machine translation, where an important subtask is, okay, what's a likely string? Right? So in a speech recognition setup, the acoustic model says, okay, here's a range of text strings, that sound might have corresponded to. And then the language model comes in and says, okay, yeah, but it's important to wreck a nice beach is a ridiculous thing to say. And it's important to recognize speech as a reasonable thing to say. So we're going to rank that one higher. So that's the kind of form-based task that they're initially meant for and good at. And then what's happened with the neural language modeling revolution in the past few years is that when you extract the word embedding from language model, you have really finely fitted representation of word distribution, which is very useful. And some of them can even do where you get the word embeddings are contextual. Right? So the information about the word and what it's like to co-occur with isn't about that word across all the text, what about that word in its current context. So super useful. But not the same thing as understanding language. And I kept getting into arguments with people who were not linguists who wanted to say, yeah, it is. So Alexander Kolar and I wrote this paper to sort of say, okay, look, here's the argument. Why not? With the hopes that that would put it into it. And it didn't like people still want to come argue with me about this. But the thing that is really hard to see and like sort of the value of linguistics in this place is that when we use language, we use it and I'm sorry, I'm going to pull out a philosopher on you here, but Heidegger has this notion of thrownness. So you're in a state of thrownness when you are not aware of the tool you are using. And if you think about typing on a keyboard when it's going well, the keyboard disappears. Right? And then you have a key that sticks and then all of a sudden the keyboard is very, you know, there for you again. Well, language is the same way. When we are speaking a language that we are fluent in, it is not very visible to us until something makes us focus on it. And of course, linguistics is all about focusing on the language. So linguists are used to doing that. So when we talk about giving words to a language model, it's really important to distinguish between the word as a sequence of characters as opposed to word in the sense of a pairing of form and meaning. Because what the language model is seeing is only the sequence of characters. So what's a language you don't speak? Mandarin. Mandarin. Okay. You don't speak Mandarin. I assume you also therefore don't read Mandarin. Definitely don't. Well, yeah. Maybe recognize a couple of the characters. I mean, some read Japanese, so there's some overlap. Let's go a little bit further away. Do you read Cherokee? No, definitely not. Okay. So Cherokee's got this wonderful syllabary, so writing system where the character represents syllables. If someone showed you a whole bunch of Cherokee text, that experience of looking at it would be a better model for what the computer is doing than you looking at English checks because you can't help but get the meaning part when you're looking at it because English is a language you speak and read. And Mandarin is kind of in between there because you would pick up a few of the Hunts that you recognize from Japanese kanji and quite the same. And I guess I don't know. I don't want to argue with you, but I do want to sort of like, I guess. Advocate, I don't know. So I'm not like deeply like, I haven't thought deeply about this topic, but what I guess what I have seen in my life is these language models kind of like working better and better than I could have imagined from the strategy that they employ and sort of seeming like they're getting more and more subtle detail. And of course, you know, when I was a kid, I learned about the the Turing test, which seems like a pretty good test of understanding on its face, right? Which is, you know, if sort of, I think the test is like, if you have a conversation with with something and and it you can't tell if it's a like a automated system or a human then we can say that it has intelligence if it can sort of and it sort of seems to me like these language models are on the verge of passing the the Turing test. So like, I guess what would it take for you to feel like some automated technique actually has understanding of of the what it's consuming? Yeah. So I think the first thing I want to say about the Turing test is the reason it doesn't work in and I hate to disagree with a giant like Turing because you know, Turing's work was really important in foundation. But it was 100 years ago. It's possible that you know, 70, 70, all right, 80, 70. Okay, so it is right. As it turns out, people are too willing to make sense of language and too willing to sort of build the context behind something that would make something make sense. And so we are not well positioned to actually be the testers in a Turing test. And so that's that's why that doesn't work. And so the language models because they can come up with coherent seeming text, right? These are these are probable sequences given a little bit of sort of noise and where you start. Well, what would likely come next based on all that training data? Then it sort of comes out of something that we can make sense of and then we are sort of easily fooled into thinking that it actually meant to communicate that. So you're asking the question of what would show that a machine has understanding? And I think part of it is, well, let's let's talk about actually interfacing with the world in some way. And we certainly do have cases where machines in restricted domains for restricted ranges of things that they can do do understand, right? So when you ask your local corporate spybot to do something for you and it does the thing, it has understood, right? You know, I'm sorry. Let's a local corporate spybot. Sorry. I'm sorry. I'm sorry. I'm making a snarky remark about the privacy implications of things like Siri and Alexa and Google. Oh, I see. I see. I got you. Samson, Big Spees in the same space, Microsoft had Cortana, right? Right. Right. Gotcha. Yeah. Okay. So wait, when you ask those things to set a timer or turn on the lights or dial a phone number or whatever, and it works, then yes, to a certain extent, it has understood and it has understood because it's training setup was looking at not just language, but something external to language that needed to map to that. And so that's a kind of understanding. And the question is, how do you, so for somebody who was interested in doing that across some more general range of things, right? The question is, how do you set up tasks that require some kind of action in the world so that it can't be done just by bulldozing it with a language model. So well, this is a likely thing to come next. Right. And I guess, so you got to describe your octopus thought experiments. That's very evocative. And I have some questions. Okay. So the octopus thought experiment is about not just being able to understand, but learning to understand. And that's the difference between it and both the turning tasks and surles thought experiment where both of those basically say, imagine someone has set up the whole system, right? Then we could test for intelligence or we can, as from a philosophical point of view, saying it's still not understanding. So those are sort of the system exists and we are thinking about it, we're testing it. And the octopus is this thing of saying, okay, if we had something that we assume, we positive that it is hyper intelligent. And then that's part of why we picked the octopus. In fact, it was initially a dolphin, but we decided that octopuses are inherently more entertaining and also it was better because a dolphin's environment is a bit closer to a human's environment. So we wanted the octopus to be something that is positive to be super intelligent. And they are, I think, understood to be intelligent creatures. And we said, like, as smart as it needs to be, like, that's not the issue. So we are assuming intelligence. But then we are only giving it access to the form of language. So in our scenario, you have these two English speaking humans who end up stranded onto nearby islands. They are otherwise uninhabited, but they have had previous inhabitants who set up a telegraph undersea telegraph cable. So these two humans can communicate with each other. We left it off stage how they discovered the telegraph or the other ones on the other island where just that assume it exists, the thought experiment. You can do things like that. Assume a spherical cow. Except we don't need spherical cows. So telegraph cable and the humans are named A and B and they are basically using English as encoded in Morse code to talk to each other. And this hyperintelligent deep sea octopus that we called, oh, comes along and taps into that cable. So the octopus can feel the pulses going through for Morse code. And the question is, what could the octopus actually potentially learn here? And because this is a hyperintelligent octopus that's got, you know, as much time as it wants, it's got as much memories it wants, it is able to very closely model the patterns of, you know, what's likely to come next. So in our story, the octopus decides for some reason that it's lonely, it's going to cut the cable and pretend to be B while talking to A. And on reflection, it's like poor B, just cut off from the world, right? So maybe the octopus is also talking to B pretend to be A, but we don't talk about that. And so the question is, under what circumstances could the octopus continue to fool A that it's actually B? And we say this is in a sense, a weak version of the Turing test, because the way the Turing test was set up, A is giving the task of deciding am I talking to a human or not? And here, there's subterfuge, right? The octopus is, it's mere existence is unknown to A, right? So, you know, if there's just sort of like chiptap pleasantries, those things, you can just kind of follow a pattern and it's relatively inconsequential as long as what's coming out is internally coherent. And even if it's a little bit incoherent, well, maybe B's just being silly, right? It doesn't matter so much. So we thought, okay, well, O could get away with that. But once you get more towards things where A actually really cares about communicating ideas to B and getting ideas back from B, it's going to get harder and harder for the octopus to maintain this semblance of good communication. So we go through this example where A builds a coconut catapult and the octopus is able to send back sort of like very cool invention, great job or something, even though A was asking for like, well, what happened when you built it? And but, you know, the octopus has no experience of things like coconuts or rope or stuff like that. So it can't reason about those things in the world or even know that A is actually talking about them. All it can do is come back with, well, what's a likely form of a response in this context? And to the extent that O gets away with that, it's because A is willing to make sense of those utterances. O has no meaning in this scenario. And then finally, we have a bear show up and start attacking A and A says, or to be actually help by being attacked by a bear. All I have are these two sticks. What should I do? And you know, at that point, O is utterly useless. And so we say, this is the point at which O would definitely fail the Turing test if A survived being eaten by the bear. But then we tried with GPT-2, like what would it say? The answers were hilarious. Like the words are in the right topic area enough that it comes back with something funny. And I encourage people to go look at the appendix to our paper where we put these, but it's never going to be helpful. And it's not actually expressing communicative intent. Well, I have to say, like, walking into that paper without knowing the context, I really enjoyed it. And I think for me, I especially enjoyed it because of the sort of concreteness of the thought experiment was like evocative, but also, you know, kind of makes you think like, huh, like, what do I think about that? And I guess what I, what I could think was like, you know, for me, I feel like I've learned about a lot of things that I haven't like experienced. I mean, I was especially thinking about this. kind of like learning math or there's kind of all these abstract topics. And I feel like in a way I feel like I learned about math in some sense through form almost sort of just it's all in my head right I'm kind of like learning things like visualizing them and I kind of wondered like it seems possible to like learn to reason about things that you haven't seen or experienced just from like a stream of words right or I even remember actually grading a blind students papers that it was actually it was really interesting like you know how they they walked through stuff in a math class and seemed like they were visualizing things even though you know they never they were blind from from birth so I just wondering like I guess I'm like totally convinced that the octopus couldn't somehow figure out what a catapult does if they kind of listened to all language. So if the octopus had actually had a chance to learn English then yes right it didn't because it it never got that initial grounding and we absolutely learned things through language that are outside of what we directly experienced you know conversely if you as a sighted person wanted to understand what it was like to live as a blind person you could listen to or read what a blind person has to say about that and learn about it right um so that's definitely something that we can do but we can do it because we have required linguistics acquired linguistic systems right so we when we use language to communicate we absolutely tell each other ideas and things that are outside of even our own experiences that we invent things and then transmit that to other people but we do that based on this shared system that tells us okay here's the range of possible forms these are the well-formed words and sentences these are the sounds that we use in this language these are the way the words are built up the sentences are built up and these are the standing meanings that they map to and then we use those standing meanings to make guesses about communicative intent and the problem for the octopus isn't that it's not smart we said it's you know hyper intelligent it isn't that it couldn't if it knew the language understand those things it's that it's exposure to the language is not set up so that it can actually learn business linguistic system all it can learn is distribution patterns I guess what prevents the octopus from learning the language over time like a like a human probably would okay so it's it doesn't get to do and in the paper we go into human language acquisition um for first language acquisition it's all about joint attention right so when babies learn language it starts from social connections to their caregivers and understanding that the caregivers are communicating something to them and then mapping the words onto those communicative intents and the child language literature talks about the importance of joint attention that kids learn words when their caregivers follow into their attention and attend to the same things and then provide those words and so that that experience that mapping the octopus doesn't get that it's just getting the words going by now so do you do you think there's some algorithm possibly that could exist that could take a stream of words and understand them in that sense so natural language understanding is a tremendously difficult problem because it relies not just on the linguistic system but also on world knowledge and common sense reason all the kind of things so you can certainly as way like more certain than I actually am but there's a big difference between saying I'm going to build an algorithm that has understanding of linguistic structure has understanding of linguistic meaning has understanding of how those meanings map to a model of the world and then use that to understand versus I'm going to build a system that only gets linguistic form and assume that it will get to understanding in some way so yes you can you could go much much further with algorithms that have more in their input in their training input than just form so that's going to be things like visual grounding is going to be things like the ability to possibly query people for answers and it might be knowledge bases it might be other sensors in some sort of embodied setup so yeah I'm not saying that natural understanding is impossible and not something to work on I'm saying that language modeling is not natural and with understanding. I'm just so clear so just consuming language without kind of all this extra stuff you're arguing that no algorithm could from just that really understand language right and by language I mean form right so imagine that you are dropped into the tie equivalent of the library of congress and you have around you any book you could possibly want in tie but only in tie for some reason this library doesn't have tie Chinese tie French tie English dictionaries it's just tie right right could you learn tie I think so I mean I guess what's hard is that I have a language already but I I feel like I so what would you do what would be your first step to learning tie if you have just noodles and rules of tie books and that's it around you what I start to do I mean I would I'm not sure do you think I couldn't learn tie so I'm curious about what you so you as a person could you learn tie sure you could go take a tie language class no no I mean from in this in this situation to sort of drop them with a I mean people do like learn like how do people learn like hieroglyphics or something where where there's no one around that that still knows it they need to find like a resetter stone or can they so the resetter stone is what unlock the hieroglyphics if you don't have something like that then what you have to do is resort to hypotheses about distributions and say you know what do we know about the world in which these texts were written what do we know about how languages work and can we say okay well you know given frequency analyses and the length of the words this seems like a language that's got you know separate function words instead of lots of morphology so that thing might be an article that thing might be a copular verb and you could you could do some analysis like that it's not well language models are doing right and then to get from the sort of structural things into something about meaning you have to make guesses about what's being described you have to basically bring in some world knowledge and say how well does this fit so when I asked you that question but what would you do I was thinking well you know possible answers are I would go find and illustrate it in psychopedia that has pictures in it right but there's some visual grounding or I would go find a book from whose cover I could tell it was actually the tie translation of you know curious George and these are great suggestions but all of that is bringing that sort of exactly the same yeah and then once you have a foothold you can build on it right and you know that's an interesting way to go but if you just have form it's not going to give you that information well interesting thank you this is really interesting it's um I guess my last question on this topic is do you sort of predict that these language models will run into like problems that will really experience and then we'll have to like kind of change the approach or do you think that as as like our bar for applications of natural language goes up they'll just sort of adapt and sort of find ways to incorporate external information kind of like finding the curious George translation so I think that language models are going to remain useful and language models have been an important component of language technology since Shannon's work in the 1950s this is this is long-standing but I think that we are likely it's so hard to predict the future but the you know my guess is that uh or maybe what I would like to see is that we get to more stringent sense of of what works and what's sort of an appropriate uh range of failure modes and what kind of fail-saves we need and people are going to find that putting language models at the center of something where your application really requires you to have a commitment to accountability for the words that are uttered is going to be a very fragile way to go and so this I my guess is that when we get to that point we're going to de-centre the language models and have them be something that is you know selecting among possible outputs again or providing these word embeddings but they are not a step towards general purpose language understanding the way they're hyped to be is that sort of one one set of problems right if you if you have to have accountability for the words that are uttered you do not want a stochastic parrot right you want something that will speak for you in a reliable way not just make up what sounds good um and then the other thing is if we take seriously these issues around bias and encoding and amplifying bias and training data I think we're going to find that we want to work with algorithms that can make more of smaller datasets so that we can be better about curating and documenting and updating those datasets so that they stay current with what's going on rather than this path right now that relies on very large language models so those are my guesses there's also the environmental angle you know both actually the energy uses angles both environmental but also about technology to a certain extent so I think there are more and more people and you know there's shorts at all, strubles at all, hundreds and at all a bunch of work now sort of saying hey let's make sure we're also measuring the environmental impact as we do things or the you know the carbon footprint so that we can direct effort to doing things in a more and more efficient way and so there's that angle but there's also many situations where you don't have the whole cloud available right if you want to do computing on a mobile device you're not going to be able to have an absolutely enormous language model in there and so there's there's pressure to find leaner solutions and I think that's a win-win you know sort of environmentally and then in terms of more flexibility with technology. Totally totally and it's a good segue because you point out a bunch of this stuff from your people about benchmarks which I'd love to talk about a little bit and maybe you could kind of summarize I guess maybe start like what are benchmarks probably most people know but then kind of what are the possible hitfalls with them? - Yeah, so I should say this is a paper called AI and the Everything in the whole wide world benchmark that we presented at a workshop called Machine Learning Retrospectives at NERIPS last year. And it's joint work with Deborah G and Altshanna and Emily Denton and Amanda Lynn Palata. And another collaboration where, in this case, we actually do have meetings or retouch to each other, but of those people, the only one I've met in person so far is Amanda Lynn, who is a PhD student in my department. So, pandemic life, right? (laughing) But we got together because we were talking about the ways in which benchmarks are being sort of misused in the AI hype machine and in AI research that is sort of striving for generality and overclaiming what the benchmark shows. So a benchmark is basically a standardized data set, typically with some gold standard labels, although you could also have benchmarks for things where the labels are inherent, like a language model. Right? What word actually came next is the gold standard label. And the idea is that you might have a standardized set of training data or possibly not, and then you've got the standardized test data and people can test different systems against this. And so you have this chance of saying, which system is more effective in this training regime or given this training data against that test data? So that's a benchmark. And let's say you asked me before if I could summarize the problems with benchmarks, and it's not so much benchmarks that have a problem, but the way that they're used. And I think we are, this is an example of the map is not the territory. So people will tend to say, here's this benchmark about computer vision. So ImageNet is that, right? Or here's a benchmark about natural language understanding of English, and that's glue and super glue. And people will say, I mean, I've actually seen this in like a PR thing that came out of Microsoft saying that computers understand English better than people now. Right? Because this one set up scored higher than some humans on the glue benchmark. And that's just a wild over claim. And it's a misuse of what the benchmark is for. So what's the problem with the over claims? Well, it kind of messes up the science, right? For we're not doing science if we're not actually matching our conclusions or our experiments. And we live in a world of AI hype, which means that people are more likely to buy into and set up solutions that don't function as advertised because they live in a world where people are being told that Microsoft has built a system that understands English better than humans do. So of course, you could also build an AI system that does whatever other impossible thing. Like, guess as someone's political affiliation, by the way they smile or something, which makes no sense. But we live in a world where there's all these claims over claims about AI and that makes these other ones also sound more plausible than they should. So those are the problems that I see. But they are, benchmarking is important. There was, so in the history of computational linguistics, there was a while where when you wrote a paper for VACL, the Association for Computational English, you would say, here's my system. Here's how I built it. Here's some sample inputs and outputs done, right? And then the statistical machine learning sort of wave came through and brought with it the methodology of shared task evaluation challenges, which is sort of a historical version of benchmarking where NIST and other organizations would say, OK, we want to work on speech recognition and we want to actually get a sense of how these different systems compare to each other. So we're going to run a shared task evaluation challenge where everyone gets the same training data and we're going to have some held out test data that no one gets to see and at a certain point, all the competitors submit their systems and we see what happened. And that's an improvement in the science compared to what was going on before. But that is not the whole story, right? If you want to understand how well the system is working, if you want to understand how to build the next system, you can't just test it on some standard thing. You also have to look at, well, what kinds of errors does it make? And how do the different systems compare, not just in their overall number, but in their failure modes, which inputs work for them and which ones don't, and on and on like that, as opposed to, OK, I got the high score, I'm done. Right, right. Well, well said, I guess I don't have a sad there. That's a-- And I guess, can you say a little more about, I feel like this is a great paper and that it makes these really concrete, sensible recommendations. It's a sort of suggestive few alternatives to benchmarks. Could you maybe run through this for anyone listening to this? Yeah, absolutely. So it's more compliments than alternatives to benchmarks. So in addition to benchmarks, this can be used sort of as a sanity check, right? Did my system actually do better than a super naive baseline? Or I want to compare some systems head-to-head. Let's use this benchmark. You might also use test suites, which are put together to sort of map out particular kinds of cases that you want to handle well, as opposed to just grabbing whatever happened to occur in your sample test data. You might do auditing, which is very much a kind of test suites and sort of saying, so this is like Joy Blomini and Tindy Geber and Deborah G's work on auditing face recognition data sets where they sort of systematically created a set looking at two genders and a range of skin colors and sort of saying, OK, is it accuracy actually even across this set of people or no? And they found out no, right? So that's the-- How is that different than a benchmark that kind of sounds like a benchmark, doesn't it? So it's not the way benchmarks are typically created. You could imagine someone creating a benchmark that is sort of systematically mapping out a space. But that's not the practice. The practice is we are going to go grab some data from somewhere and then hold out 10% of it to be the test. And the other 90% is training or 80% training, 10% dev, right? And the way benchmarks are typically put together is let's just grab a sample of data and see how well this thing works, as opposed to let's create a testing regime through test suites or through this auditing process that can allow us to sort of find the contours of its failure mode. So not how well does it work on average? But OK, but how well does it work for this case and that case and that case? There's also adversarial testing, which is a few different things fall into adversarial testing. So sometimes people will create test sets by going and collecting all the examples that previous systems did poorly on to make a particularly hard test set, which is interesting in the sense that it can filter out the sort of freebies that are too easy. But also doesn't necessarily guide anything towards better performance for a particular use case, because it's just sort of like, well, we're selecting what was hard for the previous model, not what's particularly important to get right or what's particularly likely to be frequent in our use case and so on. So that's one kind of adversarial testing. And then another one is what we did in the build it break it shared tasks. This was Alice of Enninger and Sudarau and Haldewe and I in 2017 put together a shared task where we had system builders and then breaker teams. And the breaker teams goal was to find minimal pairs. So two examples that were minimally different to each other, but for which the systems would work for one, but not the other. And that would be way of sort of mapping out the what causes system failure. So you can look at that. You can look at error analysis. So take the test set from the benchmark or the dev set from the benchmark and then go in and look and say, OK, what are the kinds of problems that are showing up here? A lot of systems that rely on language models tend to do really poorly with negation, which is one of these things that's very important to the meaning, but tends to be a short word or subword. And so it is easy to miss. So you can imagine speech recognition or machine translation. If you missed one word out of 20, it matters a lot what that word is. If you replace with the, in many cases, that's not going to cause a lot of problems. But if you just skipped a knot somewhere-- Yeah, that makes sense. So all of this is basically about looking at what it is we're trying to build, what it is we're testing on, how it fits into the motivating use cases. And then what works and what doesn't? And for what doesn't work, what are the implications? Like what happens in the real world if that failure happens? And also what are the likely causes? So what is tripping this up? And so all of that is what we would like to see instead of the leaderboardism, which is everyone just trying to climb to the top of the pile on the benchmark, which doesn't feel like it's really-- I mean, people talking about the speed of progress in AI, they love to talk about how quickly those leaderboard changes and how quickly the state of the art so it gets higher and higher on these various benchmarks. And I always think, yeah, but so. What does that actually mean in terms of understanding the world better from a scientific point of view or building technology that works better, not just in the average case, but also in the worst case? And so-- Yeah, it's interesting. Well, I had a couple things to come up for me reading that paper. I mean, I think when I started my career, I think it was just sort of on the tail end of ACL papers where they would just-- It seemed like they would just cherry pick some examples, you know, or act in where it's in, and it just seemed like ridiculous. Like I remember they had early benchmarks and people would have like lower accuracy than just sort of guessing the most common case or something, which you know, you could argue that's better and people did, but that just seemed a little ridiculous to me. And I can't think, I remember this actor from your class about I think it was no Chomsky like saying that oh, kids, you know, moms don't teach kids language, but they just, actually they do. And it's just like no one, no one bothered to check, you know, so it's kind of maddening. And I think I appreciated benchmarks from that, but then your recommendations are like not only reasonable, I think in companies, a lot of it is like, is like standard best practice. Like I don't think you would just, you know, release like a new model without, you know, kind of trying and getting a flavor for like where it works and where it doesn't, you know, you would just be like, oh, we, you know, we took 10% of data, held it out, to ship it, you know, but it does seem like, it does seem like that's actually one case where you see it more in companies and in sort of academic literature, probably because it's easier to look at one number and be like, hey, we beat it, but clearly that's, that's flawed. Anyway, I thought that was a great paper with really good suggestions and I think everyone should definitely follow. But I mean, I guess I also want to make sure we got to the last paper that we talked about, which is cool because I just want to make sure that people know what is the vendor rule? And why is it important? - So the vendor rule or the hashtag vendor rule? - Yeah, no, so is it hashtag vendor rule? It seems like, yeah. - No. - It's cool. - Say what it is first, and then I have some questions about that's practice. - Yeah, so it is itself a best practice, which says that you should always state the name of the language you're working on, even if it's just English. And this came about, this is a soap box that I've been carrying around and pirated the climbing up on since about 2009, where I saw a lot of that sort of pre-neural statistical NLP work saying, basically, look, man, no linguistics. And claiming that systems were language independent because there was no linguistic knowledge hard coded. And these supposedly language independent systems were mostly tested on English. And you also see a lot of work when people will publish a paper on machine reading, or paper on sentiment analysis. And in fact, no, it's a paper on machine reading of English and sentiment analysis on English text. And flip side is, if someone's working on Cherokee or Thai or Chinese or Italian, then that work gets, it's harder to get it accepted to the research conferences because it is deemed language specific, where work on English is somehow general. And that's a big problem for the scientists, big problem for getting to technology that actually works across languages. And so I've been sort of going around pestering people to actually test cross linguistically and to name the language they're working on. And in 2019, like three or four people, and this is in that piece on the gradient, I had their names listed, came sort of referred to this practice as the vendor role. So I didn't name it, but once it was named, I ran with it. And part of it is, it's kind of a face threatening question to ask, right? If someone's written something about machine reading and I walk up and I say, what language, right? It's a stupid question to ask because it's obviously English. So it's face threatening to me. And it's also a little bit rude to them, right? To ask this question that says you should have said. And so I don't mind people blaming that on me. So like, so the part of the reason I ran with the hashtag is if someone wants to go ask this question and they feel like it's sort of a silly question to ask they can pin it on me. And I'm happy to lend my name to that. I see. Nice. And I guess this is a hard question, but it just comes to mind for me. It's like, wow, English is so specific. And probably has all these kind of ideas and secretities, secretities. How do you think I'd open up a different, if it started in like tire Cherokee or something? Or English must be unusual in all these ways, right? Like are there characteristics of English that are unusual in the world? Could it go in a different way? Yeah, absolutely. So and actually in that paper, I list out a bunch of them. So one thing is English is a spoken language, have a signed language. So if we had started NLP with American signed language or other signed language, it would have been very different, I think, right? Clearly. So that's one big choice point. Another thing is that English has a very well established and standardized writing system. And many of the world's languages don't have a writing system at all. And many of them that do don't have the degree of standardization that English does. Also, many languages will have a lot more code switching going on on average than English does. Although-- Sorry, what is code switching? So code switching is when you use multiple languages in the same conversation, sometimes even in the same sentence. And that happens a lot in communities where there's a lot of bilingualism or multilingualism. So if you and I-- well, you also speak in Hong Kong, right? You said so. Yeah. When you studied kanji, what was your favorite way to bank your own? I am not a fluent code switcher, so that was really awkward and stupid, but to illustrate the point, right? I remember, actually, when-- Yeah, I know. And I have experienced that, for sure. So certainly, English is involved in a lot of code switching. But there's also lots and lots of monolingual English data. And when you go into social media data for Indian languages, for example, enormous amounts of it are code switched with English. And so there's a whole range of interesting technical challenges that come up there. We live in a world where the first digital setups were a sort of accommodated lower ASCII, the most conveniently English-all-fits-and-lower ASCII, right? English has a relatively fixed word order. We have a relatively simple morphology. So any given word that shows up is only going to show up in a few different forms. Compare that to Turkish, where you can get, like, I think, millions of inflective forms, the same root. And so that changes the way you handle data sparsity and what data sparsity looks like. So yeah, English-- our orthography is a mess, right? So the-- someone was just asking on Twitter, how come we do grapheme-to-phoneeme prediction, but not phoneme-to-graffing prediction? So grapheme-to-phoneeme is given a letter with the likely sound. And that's an important component of text-to-speech systems when you hit an out-of-acabular word. Funding-to-grapheme would be given a sound with the likely letter. And that's not a typical task. And I wonder to what extent that's true, because of English's opaque and chaotic writing system, where you're given a sound. That's like an impossible task. Yeah, exactly. But if you were to look at-- so Japanese setting aside the kanji, if you're just trying to transfer Japanese and kanah, that's way more straightforward. Spanish also has a very transparent and consistent grapheme-to-phoneeme mapping in both directions. So down to things like that, but the properties of a writing system for English-- English likes to use white space between words and sentence-final punctuation, right? These are things that we sort of take as given that it's easy to tokenize into sentences and words that just aren't going to be true in other languages. So I don't know. I couldn't tell you what NLP would look like. I can just sort of tell you sort of where the points of divergence might be. No, those are fun. Those are-- yeah. I mean, definitely-- I don't know. Those differences are so interesting. [LAUGHTER] You voluntarily took a linguist's class, so I'm not surprising. [LAUGHTER] Well, I think it's like-- I mean, I just feel like linguistics is so cool. I mean, as an outsider, just because you-- if you don't know it, then it's really eye-opening to just-- because you swim in it to sort of see, oh, there's all these patterns that I never would have noticed. And I feel like especially-- well, I don't know. Phenetics is probably the most deep where you're just like, oh, my god, those two sounds are different. I would just never have noticed that. But then it's so-- It's so easy to do the thought experiment and realize you're wrong. That it's just-- I don't know. I love that stuff. Yeah. And yeah, it's funny. I mean, I remember-- I don't know. I feel like most of my early work was in parsing Japanese in different ways in society. I do remember-- I don't know. I guess it didn't seem like that was a impediment to publishing. But it was surprising that there was so little work on it for how necessary of a task. It would be to deal with it. And then, in my first job, it was mostly processing Japanese language stuff. And it was striking how little research there was defined on the topic. I felt like there was just more institutional knowledge inside of companies than literature on it. Because what happened in the research community is, well, that kind of parsing problem is solved, right? Because people had made a certain progress on it for English. And that was mistaken as the problem in general being solved. So what's new here? Well, this is for Japanese. That's news. It hasn't been done. But it's actually hard to get people to see that. And so my goal with what got called the Bender Rule is to say, OK, let's keep English in its place. And say, when I've done this for English, I need to say that it's for English to hold room for the other work on other languages, which is also really important and novel and valuable. And we'll see. If we periodically go through different folks in the field, go through and count how many papers in an ACL conference actually work on different languages and actually say what language they work on. and it's not. changing as fast as I'd like. But there's some really good development. So the Universal Dependencies project has produced tree banks for many, many languages, and that has spurred a whole bunch of a very cross-domestic work, which is exciting. - And what do you think about, I mean, I mean, some of the most like evocative work feels like, you know, like building language models across like all the languages or like translation models that can kind of use pairs of languages in interesting ways where you have more data to help with ones with less data. - I guess, do you think that's like a fruitful direction or does that, do you think that's sort of like encodes our biases somehow in the way it works? - So I mean, it's certainly interesting. And to the extent that we're relying on these massive data-hungry things, where languages just don't have that much data, seeing what we can do based on, you know, transfer from the bigger languages is an interesting valuable way to go. I think the interesting questions to ask would be, to what extent does this impose the conceptualization of the world encoded in English onto the results and the other languages? And, you know, what follows from that? Like, what are the risks and how does that compare to, well, but if we just do monolingual, we can only get this far. So we'll take those risks, we'll figure out how to mitigate them. That kind of work, I think, is important. And it's also really, really important to know that you are working with genuine data in the low resource languages. So there was this thing where it came out that, I think it was Scott's, the entire Scott's Wikipedia was written by one person who doesn't speak Scots. And Wikipedia is this really important data source in NLP. So any NLP system that claims to be doing something for Scots just isn't. And a fantastic model in that regard is this research collective called Masakane, which is a continent spanning research initiative in Africa towards doing participatory research to create language resources for African languages. And they've done really interesting work on how to build up the community so that people can come contribute as translators, not machine translation, especially especially people actually translating language. And there's a really cool paper that came out in, I think, findings of NLP last year describing the Masakane product. So project. So that kind of work of, like, if you're going to work with low resource languages, being sure to connect with the community who would be the people using the technology, then you could find out, OK, what are the concerns? To what extent do you want to bring in what we can do from using the larger resource languages versus, would you rather stay a monolingual and see where we can go? And hear from the community and involve the community in the research. And I think Masakane is a great model of that. Cool. That seems like a good place to end. We're way over time and you're really generous. So thank you so much. But I really enjoyed talking to you. Thank you. Likewise. Thank you. I can go on and on. So I appreciate the chance to do so. If you're enjoying this interview series, the most helpful thing that you can do for us is leave us a review. It helps other people find the show. And really, we do these shows so that people watch them. And what I really want is more people to find it. So if you leave us a review, I really appreciate it.

Podcast Summary

Key Points:

  1. The discussion highlights the distinction between words as character sequences versus meaningful linguistic units, emphasizing that language models process only the former.
  2. The paper "On the Dangers of Stochastic Parrots" critiques the trend toward ever-larger language models, warning of risks like bias amplification, environmental costs, and exclusion of smaller research communities.
  3. The paper's publication was controversial, leading to internal conflict at Google and the firing of ethical AI researchers Timnit Gebru and Margaret Mitchell, which drew public attention to corporate influence in AI research.
  4. Concerns include current harms from biased outputs (e.g., search engine racism) and the need for more rigorous testing, documentation, and consideration of societal impacts before deployment.
  5. The authors advocate for a shift from hype-driven scaling to responsible, well-scoped AI development that prioritizes safety, equity, and transparency.

Summary:

The conversation centers on the paper "On the Dangers of Stochastic Parrots," co-authored by linguist Emily Bender and Google's ethical AI team. The paper critiques the race to build larger language models, arguing that scaling alone risks amplifying biases, increasing environmental costs, and marginalizing smaller research groups and languages. It calls for more cautious, ethically grounded development with better documentation and testing.

The paper's publication sparked controversy at Google, leading to the dismissal of researchers Timnit Gebru and Margaret Mitchell, highlighting tensions between corporate interests and ethical AI research. Bender notes that while the backlash brought the paper unprecedented attention, it also revealed systemic issues in AI governance. The discussion underscores ongoing harms, such as biased search results affecting marginalized groups, and stresses the need to view language models as tools processing form, not meaning, to avoid overstating their capabilities.

Ultimately, the authors advocate for prioritizing safety and inclusivity over unchecked scaling.

FAQs

Language models process words as sequences of characters, not as pairings of form and meaning. This is akin to hearing a language you don't understand, focusing only on the character patterns.

Emily Bender is a professor of linguistics at the University of Washington with interests in societal issues, multilingual variation, and philosophy of linguistics. She was also the host's teacher at Stanford University.

The paper, titled with a parrot emoji, discusses the dangers of stochastic parrots and questions if language models can be too big. It surveys risks like bias, environmental impact, and accessibility issues with large-scale models.

Google researchers Timnit Gebru and Margaret Mitchell faced pressure to retract the paper or remove their names. Gebru was fired (or resigned, per Google), and Mitchell was later fired after documenting the events, highlighting corporate influence on research.

It advocates for more testing, documentation (like Model Cards), and consideration of societal impacts before deployment. It also warns against a sole focus on scaling, which can exclude smaller languages and research groups.

Examples include search engines returning pornography for terms like 'black girls' and racist image results (e.g., gorillas for black people). These reflect biases in data and algorithms that amplify societal inequalities.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.