20. Jae Lee: on augmented reality, human computer interaction, and designing for everyone
55m 55s
In this podcast episode, host Sasha interviews Jay Lee, a PhD student at the University of Washington’s Cability Lab. Jay discusses his journey into human-computer interaction (HCI), starting with AI research at UIUC before pivoting to study how people interact with technology. His focus shifted to accessibility after collaborating with blind and low vision communities, leading him to explore augmented reality (AR) for this population. Jay highlights three key projects: Baller, which uses AR glasses to enhance visual cues in sports like tennis for low vision players; CookAR, which augments cooking tools to indicate safe and hazardous parts; and RASSAR, a mobile app that scans indoor spaces for accessibility and safety issues, such as insufficient clearance for wheelchairs or hazardous table heights. He notes that accessibility research is increasingly featured at general HCI conferences like UIST, and that principles from this work often benefit broader populations. Jay emphasizes that his research relies on computer vision models and prioritizes visual feedback for low vision users, who prefer using residual vision over auditory cues. He sees potential for tools like RASSAR in platforms like Airbnb for accessibility mapping. Overall, Jay’s work demonstrates how AR can address specific challenges for blind and low vision individuals while advancing technology for all users.
Welcome to Gears of Progress, a perfect place to learn about research and rehab engineering and assistive tech. As usual, I'm your host, Sasha, a research scientist at the University of Washington, and this is episode number 20. Today in my studio, we have Jay Lee. He is an HCI researcher studying augmented reality or the term AR. We're going to be probably using a lot through the conversation today. Human AI interaction and accessibility and also PhD student at the University of Washington at Cability Lab with Dr. John Foylick. So welcome in, Jay. I'm so excited to finally make it happen after trying to chase you for, I don't know, two three or four months. You're quite a hit in the HCI world these days. So I'm excited to be here. Hey, yeah, no, like, why is I'm glad to be here? Thank you for having me. Yeah, I'm Jay, currently 30 years BC student. I think that, you know, the description of my current position is very accurate. So thank you for that. That's great. All right. I think a good question to start things off is how did you discover the field of HCI? Or human computer interactions? And I think also if you could explain in general terms what HCI entails in all of its aspects to the audience of this podcast, that would be great. Yeah, no, that sounds good. So yeah, so I discovered the field of human computer interaction or as we call it HCI, when I was an undergraduate student at the University of Illinois or Bay of Champagne, where I was originally doing sort of more artificial intelligence research, but then I wanted to kind of see, hey, I made an algorithm. Could we, for example, push this out to people? Would people say about the thing I build? And I became quite interested in that aspect of computer science. So then I started looking into, hey, what is the field that is going to allow me to achieve this? Talk to people about the things I build. And then I discovered the field of, you know, HCI with Professor Brian Bailey, who is still there teaching a class. So yeah, human computer interaction really to me is about really trying to understand not just what people need and building the technology that's necessary for people, but also sort of studying how the technologies that we build interact with humans and the humans that we give it to, which we tend to call participants. So really, it is a study about how participants would interact with our system. And I do enjoy talking to a lot of people about the things that build and what they think about it, whether it's good or bad. Yeah, absolutely. It looks like most of your work that pertained to HCI originally has really revolved around kind of humans as a whole, not any particular populations, but during your BG at UW, you have really focused more on the kind of the accessibility aspects of human computer interactions in particular for people with low and vision and blind. How did you get to kind of make that pivot point into really focusing your HCI applications for accessibility purposes? So yeah, you're absolutely right. A lot of my earlier career, so to say, research did focus on sort of the general population, but the first time I actually ran into accessibility research is with the actually now professor, professor Anne Hong-Woe at University of Michigan. Him and I did a work called Image Explorer, trying to sort of assist blind and low vision people in understanding images. But then from that, I was sort of applying to PhD programs at the time and I was like, hey, I actually really enjoyed talking to people who are in the blind and low vision community. I thought they're really, really nice, super sweet. They are a big fan of some of the technology that I build, and that's kind of how I got started. And then when I came to UW and working with Professor John Freely, I started to then really think about, hey, there's a lot of people who are building technologies for the general population, but especially in the field of AR and augmented reality, less people are focusing on, say, accessibility, so then I was like, hey, could I, for example, be in this space? And well, here I am three years later, doing research in this space. Yeah, that was pretty cool. That's a very good point, you know, the difference even within the last probably 10 years of the technology that has been coming out and the accessibility features that are being added. It's definitely growing exponentially. I was wondering in all of the HCI conferences that you attend, do you see that exponential growth of accessibility focused HCI research being outpid by other labs? Yeah, no, I think there is definitely a growing interest in accessibility, which is something that I love to see always. So I was just attending a conference called WIST, that's UISD, and that was held in Pittsburgh in October this year. And for the first time, I was sitting in a session that wasn't even called accessibility, and everyone in that session, including myself, were presenting papers on, you know, some accessibility topic, which I thought was great because there's a, there's a, there's a conference that's held around the same time that is specifically for accessibility called assets. And WIST was generally not a conference that was focusing on accessibility. You know, we are very much of a sort of a, builders and really just building sort of fancy systems here. But then I'm glad to see, you know, a lot of accessibility people submitting the WIST and then talking to WIST community about accessibility and how important it is. So I thought that was great. Do you feel like HCI work that you do that others do in the disability domain and the domain of accessibility is different from, let's call it the general HCI work that's, that's being output and that you really focused on before. Do you feel like you view your own field differently now that you do consider the accessibility aspects kind of at a greater depth? Yeah, no, so that's a good question. So I think, you know, there is a saying that, you know, if you were to do accessibility research, then you are also being able to make advancements for sort of, you know, technologies for everyone. And I think that's certainly true here. A lot of the work that I do involve, yes, blind and low vision people and to then give them say information in real time. Then there are technological improvements being made in terms of how fast an AI model can be ran and how to build sort of, you know, an AR system that can then run some of these AI models fast enough to then be usable by people. And well, with these advancements, I think everyone can benefit really. So I see that as a very big plus. And in fact, a lot of the things I do at UW when I go to say an internship and I'm doing a research for say a product that's meant for the general public, then I think a lot of the things I learned that UW are definitely transferable to, you know, say the companies that have been to, which is again great. Yeah, I think, you know, accessibility research in that sense isn't too different than, say, the general ACI research in the sense that we built products for a specific reason that we sort of generally come, we generally talk to people or we generally read papers, discover a problem and, you know, we try to build a solution around that problem and then we bring in stakeholders in this case, blind and libido people, but it could be anyone really, I think, an ACI research and yeah, we'll just bring them in, see what they think and report that to the scientific community. So the research process itself is, you know, too different. It's just a matter of, you know, I am really, really passionate about helping people who are blind and libido and that's just the field I'm in. You know, you know, I'm really passionate about that. So you do mainly work with blind and libido populations in your research work and the technology that you cover is AR and VR so augmented and virtual reality. How did these projects and you have a lot and we're going to go and depth into some of them? How do they just kind of come to their development? How did you even decide on applying AR VR technologies to these populations? I think there's really primarily sort of two, three ways in which I discover some of these sort of application areas, so to say. I think I think a lot of these ideas do come from, you know, some of the friends that I have that are low vision. So for example, right now there's Adrian, Adrian Rodriguez in the ACD department at UW. And he is low vision and he has given me a lot of ideas, especially things like, for example, being able to play tennis, you know, as a low vision person, what is his experience like and how can that be augmented by AR technologies that are out there, whichever
is a great way to sort of discover these problems that we can then try to address with the technological solution. And yeah, a lot of it is also has to do with papers. So for example, I assume one of the papers that we will go into is this cucker, which I just presented at Wist, that problem and that problem space and figuring that out and figuring that project out really came from someone else's literature view, someone else's paper where they interviewed Blenden-Lowbeard and Cokes and their experiences and then really what are some of the challenges that they have, what are some of the strategies that they use and how can AOR for example benefit them and they had a series of two or three papers from CMU actually Franklin, Franklin and that's how I got a chance to know some of the struggles that Blenden-Lowbeard and people have when cooking and to then try to build an AOR technology around it. I see. So I think this is a perfect segue into some of the work you've done. So one of them being AOR sports, unless you're calling it something else, but for people with low vision, can we take kind of a deeper dive into this work? I'm assuming it involves more tennis, which is also one of your favorite game to play, so can you share a bit more on how it works? Yeah. So AOR sports, which I think for the purposes of the full paper that we have plans to right next, will be called baller, credit credit Daniel from our lab for that name. But yeah, no, that project is about how do you assist low vision people in sort of playing sports better and specifically ball-based sports and perceiving factors such as depth, for example, of an incoming ball or ball that's moving away from you, the opponent and where they might be at and things like that. This idea again does come from Adrian and then also partly John as well. And yeah, this project is about, you know, can they wear AOR glasses while playing tennis or basketball and then it would essentially scan your environment and detect that there are, you know, saying basketball, there's a basketball, but then there's people around you as well as sort of the hoop in the backboard and really sort of, you know, enhancing sort of the visual salience or, you know, in simpler terms, really just making the approximate shape of it look more apparent to someone who has low vision. And that way they sort of can understand, hey, there is where the hoop is as well as, you know, there's other people around me, there's the ball, for example. This project is particularly challenging because well, as you can imagine, a tennis ball or a basketball moves quite fast and so do say running opponents. So then really, the biggest challenge here is how do you make sure you are highlighting the necessary components of a game fast enough to a point where it's keeping up with the fast moving pace of the sport. And that's probably the primary challenge here, which we're trying to actively address. So most of the feedback that the users of AOR sports receive is just kind of an augmentation to the visual sensation, right? There's no other sort of feedback, I don't know, a headphone that tells them, I don't know, the ball is flying at you at a speed of this many meters per second. So this is all just an augmentation of visual feedback. Yeah, this is a visual feedback. And then primarily because there's a lot of literature to suggest that low vision people. And then specifically in the way that say Professor Yuan-Zao from Wisconsin Madison and I define low vision, at least in our papers, is someone who has vision that cannot be corrected by corrected lenses. So things like glasses and contact lenses, but then has a vision of above 20 over 400. And we say anyone who sort of does fall below that threshold is then legally blind. And so these low vision, these people who have low vision, they have residual vision. And therefore they would much more prefer to use their residual vision as opposed to an auditory feedback. And that's why we have a system that's, you know, a visual as opposed to an auditory system. And then that's why AOR glasses becomes sort of a key player here, key technology here to use. I say the other research project you mentioned that stemmed from just talking to cooks with low vision is cookie art or cook art. Is that what you're pronouncing it? Can you share a bit more about that one? Yeah. So cooker is sort of in a similar realm of topic as a baller or AOR sports in that we are aiming to sort of augment cooking tools so that it is sort of safer and more efficient for people with low vision to then interact with cooking tools in their kitchen. The key difference here is as well, cooking tools have a specific use and they have specific ways and it needs to be held, for example. Where for example, a simple example is of course if there's a knife, a knife handle is a place that you should grab whereas maybe the knife blade less so. So then how do you highlight not just what the object is and saying hey, this is a knife shape, here is kind of a color highlighting of this object. But how do you then communicate to them? This is the handle, this is the blade and therefore sort of giving them the idea of sort of how to hold the object as well as what is the orientation of the object, which side is the handle, which side is the blade. So then this project really is looking at kind of a step up from say something like baller where baller is really looking at how do you get a system that's fast enough to catch up to something like sports. Really the cook or project is looking at how do you describe to users what are the different components, what are the different parts of an object. So this concept is called affordance and this concept is quite popular nowadays in robotics as you know when you want a robot hand to grab a knife you need a robot to understand which part is handled, which part is the knife. Whereas the issue primary issue here is is a lot of the robotics models are for robots and robots don't have to move as quickly as people would. So then there is this gap between sort of you know the technologies that are being made now in robotics in that it's not fast enough. It can't render fast enough in front of the user and the user's visual feel of you. So then we had to get around that and right now the system is running quite fast and it is able to display in green outline the parts that are grapple one on object and then read in as in hazardous part of an object. So we are able to do things like knives, we are able to do pants and pots, we are able to do things like cups, spoons and forks and things like that which are all essential tools that loving people identify as tools that they would like to be able to interact with safer and more efficiently from prior larger. So definitely one of my favorite work in say four or five years of my short research career so to say. That's great and it's a really big one and impactful for sure. Another big one is and correctly if I'm pronouncing it wrong is the race are rasser. I think we're calling a razzar I think all of us also call it something different so sounds good. So that one is a really big one that pertains the accessibility domain and for short distance for room accessibility and safety scanning in augmented reality can you talk a bit more about how that one came to development and the kind of the outcomes of that project. Yeah no well first of all credit all the credits to the first author of the project Chasu who is also a PC student in our lab now for here. Yeah no that razzar project really is kind of you know something that we believe is also deployable to some degree and we know that there are some already some interest from say companies to maybe incorporate this into their product line up say Apple but yeah no razzar's a fantastic project looking at you know if everyone has you know a lot a lot of you know people today have access to phones and so then if we were to then give them a scanner that they can sort of quickly do a scan of their room in their space their say their home right then it would give you sort of this it would sort of scan for different accessibility issues in your indoor space. So for example if there's not enough space between your couch and your coffee table for someone on a wheelchair to pass right it would flag that and say hey there isn't enough space for example or say you have a table that's quite low but then you have say a pair of scissors or a medicine model on on that table then you know this system like flag and say hey this table might be too low for kids you know kids to then get access to dangerous materials that are on top of the table for example. So I think I think this project is looking at not just accessibility but of course for
where the elderly population as well as kids as well, and then sort of just a general indoor accessibility and hazard detection, really. And what's great about this is this is just a mobile application. Anyone can download it, scan the room, and see if the room has any issues that then needs to be corrected. I think this is a really, really cool project in the sense that I have friends who may even be, say, temporarily disabled. And so say, they get into an accident while playing tennis, and they are on crutches or they are on wheelchair. Well, then I want my house to be safe from them as well. So it's not just, hey, this is specifically for accessibility. It could really be for anyone. And I think that's what makes it great. And of course, it makes it great that it's just a downloadable application on the phone that anyone can use. And it creates this little, you know, pretty diagram version of your home. And it kind of tells you, hey, here are different accessibility issues that you probably should fix before someone comes over, which I think is great. I see. And all of this technology really just relies on computer vision models, right? So what you're getting to extract that model of your home, you're just extracting it from pure video? Yeah, pure video. So like you would say, take a video on your iPhone camera. It is basically doing the same thing. Except this time, it is, you know, analyzing the video footage and seeing if there are any, you know, noticeable accessibility issues or safety issues. I see. That's pretty cool. Do you see this being implemented in a frequent user of Airbnb? And I started noticing that they are kind of having some accessibility features featured as part of different listings. Do you see this being kind of a very common tool for people who list their places at Airbnb or other rentals or whatever it is? And kind of like, here's the accessibility features in our space. I think that's a fantastic question, because that is very relevant to Shasu's whole PC thesis and his whole vision of mapping indoor spaces and accessibility needs, accessibility issues within these environments. I think there's absolutely use cases here for Airbnb's, right? I think you can really go far with this in the sense that, well, then say, you know, Paul Gial instead of for computer science, could you do sort of accessibility mapping of this larger indoor space? Of course, Professor John Freylich has his outdoor accessibility mapping project called Project Sidewalk, which has been kind of the pillar of our lab, really, for the past over a decade. But then really, we are venturing now into sort of how do you do this for indoor spaces. Outdoor spaces is, of course, easier and harder. I wouldn't say it's just easier, but there is data to be out there already in the sense that there's lots of Google cars, for example, roaming the streets. And while taking Google Street View videos with their fancy technology, which I think is great, and that's what Project Sidewalk really is, relying on for the most part. But then, you know, for indoor spaces, of course, you're not just going to be able to drive a car around an indoor space. So then I think for smaller indoor spaces, you know, a system like Razzar, certainly there's a lot of applications for that, like an Airbnb or your personal home. But then you get to say a larger indoor space, like, say, Costco or say Target and Walmart or say, you know, even Paul Gialat center for computer science. Then scanning it with one phone gets quite challenging because all of my have multiple floors, might be too big, et cetera. So one of the things that Shatsu is looking at right now is to then fly drone. And based on sort of a faster way than to do large indoor space scanning, could you then do accessibility mapping through that? Is something that has been a long-term goal of Jollons and then something that Shatsu is also very interested. And so it'd be great if you can maybe have a chat with him perhaps after he's done with that work. And he has his full thesis package ready. Absolutely. That sounds great. I also wanted to give a shout out to Julie, one of the PhD students on Project Sidewalk. And we did an episode with her. I believe it's episode 12 here. And it's definitely a really great work. You mentioned when I mentioned Razar. You kind of said about the outside of the lab applications of the technology that you guys are developing. What is the stage of the work that you've done with augmented reality in sports and cooking? What kind of steps do you think you need to take or the technology needs to take for projects like this to leave the walls of a research lab and really be used in actual daily lives? Right. I mean, I think that's the really the question everyone has for air glasses, really. So I think Razar's great. And some of the projects that I am trying to actively publish and get it out there now has actually used phones as a medium instead of air glasses. And well, that's what happens when I'm at companies, for example, because a lot of companies are trying to productionize their research and their products now. And so they rely a lot of the times on phones. And then phones is a great phone in what we call mobile LAR. It's I think a great platform. I mean, I think at this point, a lot of people are aware of what AR is simply because of Pokemon Go. Right. And I was just adnoyantic this past summer in London, which has been a fantastic experience. Shout out to my manager there. But I think for air glasses, so then things like, you know, cooker and Razar-- sorry, not Razar, cooker and baller to then get out there. Then that's a bit of a, both the hardware and a software challenge, right? So right now, of course, a lot of companies are trying to invest heavily into sort of building a lightweight glasses looking AR glasses as opposed to something like, say, Apple Vision Pro, which I know a lot of people love, but also hated. Of course, also there's MetaQuest 3 Pro and things like that that are definitely more out there. But a lot of people view us a maybe more of a gaming device as a VR headset, virtual reality headset. Then, you know, for example, there's, I think, a lot of people are getting into AR glasses through, say, the RayBans Meta glasses, as well as, say, the X-ray-all, airs and whatnot. That are, I think, all great. But, of course, I think, at the end of the day, the reason why RayBans and X-ray-alls are doing, well, nowadays, is because it looks kind of like regular sunglasses, right? So then, you know, for-- at that point, when we have them necessary hardware for that look like regular glasses, then we would start to really need to assess, hey, we have this hardware. Now, what do people use it for? And really, I am trying to get ahead of that curve in the sense that a lot of companies are investing in the space of building small hardware. But once it's there, and it's, say, $3, $400, or even less, then people really need a reason to buy it. And that's really where software comes in. So then, I'm really trying to actively understand what would people use AR glasses for in the future where, you know, and maybe it's an inevitable future. Maybe not soon, but maybe, say, five to 10 years, it is, I think, a future that we can certainly foresee in that future, then, you know, what would people use AR glasses for? And that's kind of where what I'm here to do. And I think, for the first time, we also have-- we also have a great opportunity for us to be ahead of the accessibility curve. So as I call it, as in a lot of the technologies we build it in both the general public uses it, and then the accessibility side kind of catches up. And a good example of that is probably the phone, unfortunately. But for the first time, we are living in an era where accessibility is an important problem for a lot of us, and well, AR glasses isn't out yet completely. So then I think we certainly have a chance to get ahead of that as well. And that's where, hopefully, I'm there to do. And so hopefully, some of us are also there to do a little bit of that. Yeah. Absolutely. Yeah, pioneering the work. When you talked about the difference between AR glasses, you know the difference between the Rayband Meta, just kind of regular sunglasses looking AR glasses, and then versus something as big as Oculus, you in your work particularly focus on the bigger looking glasses. Is there much of a difference when we're talking about the hardware and particular of the processing power between the two? Can you potentially implement the technology you've created for baller and cooker into something like Rayband Glasses? Yeah, no. I mean, right now, well, so there's primarily sort of two issues here, I think. Well, so first of all, I think a lot of these smaller glasses are smaller today because they sort of take out certain capabilities that these larger glasses have to make it smaller. And that's why right now we can't necessarily build a baller or anything like that for these smaller devices. So for example, the X-ray glasses and the
band glasses. They are not so great at, say, video analysis through AI. Some of them don't have development access because they are worried about the privacy implications of just having cameras everywhere. Some of them don't have cameras like the X-ray glasses, for example. And a lot of their power is coming from, say, a phone. So then, that way they can then build something that is smaller, which I think is the right step, right? But perhaps an intermediary step here in terms of us getting something that can do computation, but also can do, can be in sort of a smaller glasses form factor. One of the things that I have quite a bit hope for is is meta and both meta and Google are heavily investing in the space. Meta is trying to build their Orion glasses, which I think is going to be a huge step up in terms of them building their first-day-or-classes. But then they are sort of having like a, almost like a miniature computer that you can kind of carry, almost like a phone and shape a phone really. But then it's allowing you to then make the glasses smaller because you don't need a giant computer like processor on the glasses itself. You could just kind of carry it in your pocket, for example. And therefore your glasses are smaller. Again, maybe an intermediary step. But I think that's the step in the right direction. Google just launched Android XR. Shout out to some of my friends who were on that team and who had just had a conversation with that about. But yeah, no, they released Android XR. And I know that they are also in the works of trying to make a new glasses that's going to be running Android XR. So hopefully that's going to be a great step forward as well. Looking at some of the leaks seems like it's pretty promising. Of course, can't confirm anything because I haven't been to Google. But yeah, I think we are stepping in the right direction. It just, you know, it does take a little bit of time to build something that has everything like cameras and microphones and GPS and whatnot. But then make it small. That's quite a challenging part that maybe I'm not necessarily familiar with. But certainly a lot of ingenious hardware folks are working towards, I think. So yeah, that's great. You mentioned a lot of your work. And I assume that's the work that you do through internships. And you've done a lot of internships over the past, I don't know, five, six, seven years. A lot of this work focuses on kind of mobile AR or like the phone, the augmented reality that you can achieve was your phone. Can you share some of your favorite work that you've done during these internships that involve the phone and AR? Right. Yeah. So so recently I was at the end of and of course, I don't think I can share the full full details of it just yet. Probably in the next couple months or so probably. But that work is I think one of my really probably my favorite piece of research because it is sort of, you know, what I think of when I think of sort of how the future of AR really needs to be. So that project is really trying to look at sort of how can people author their own AR environment. So say you are out at a park and it's just a regular park with trees and grass and you know, say a flower pot and flower pots and flower patches and whatnot. And you sort of open up this app, right? And you can kind of say, hey, like, you know, could you, you know, create a giant blue shark and put it up on, you know, put it up above the flower pot. And maybe it's like, you know, swimming around the flower pot, right? If that is what your imagination takes you to, then this application that my auntie can I build can do it. So then, you know, that's all on the phone and you kind of be pointing your phone at the flower pot that you're interested in sort of interacting with. And it would essentially create this maybe not so great right now, but at least something that looks like a shark and it's blue and will basically us, you know, create that on top of the flower pot. And, you know, I think that work to me is really fascinating because then it gets you to sort of an AR space that you can create and really imagination is kind of the limit, limit of this app. It can do anything you want. And to me, where this then takes you to is probably what I'm most excited about is as well, what if we then use this for sort of productivity reasons and something that's more relevant to sort of people's day to day activities as opposed to say something that my auntie has in mind, which is, you know, bringing communities together through, you know, outdoor games. You know, for example, people can make their own user interfaces depending on their current needs, right? Say you need a recipe based on, you know, things that are in your fridge. You can say, hey, like, can you build me a recipe using these things, right? And then it would build you like this UI that you can look at, this screen that you can look at that incorporates all of your ingredients and give you a recipe. I think there's definitely a lot of, you know, interesting application areas for you being able to build anything with the AI. And really, that gets you to then, you know, what we and the ACI like to call personalization, where, you know, really anything that you want as, you know, a user, really the system can build for you and the visual effects and graphics and UI and 3D models that you want I think can be built. And one of the most exciting things for me probably is is, you know, when you move, you are in need of sort of understanding whether your home can fit a certain furniture. What if you can get an AI or glasses to then look at say a, an IKEA link, right? And say, hey, can you build me a furniture that, you know, looks like this and maybe has these dimensions? It's got photos of the furniture. It's got the dimensions probably written somewhere, you know, could it then create a 3D model like we did a miancic based on these images? And then, well, just, you know, pop it somewhere in the corner and just say, hey, this is it. I think that'd be great. Like anything that you need, the AR glasses is there to then build visually. And that is probably one of the small visions that I have for the space of AR. And I think we are getting somewhat close to that. So again, I think shout out to all the people at Nyanke for having me there and, you know, helping me get this, you know, research product because I am very proud of it and I can't wait to share more on once it's out there in public in the next couple of months. I think one of the very useful ways to use AR was a project that I also saw you were on. And it's the Symbodied AR language learning through everyday object interaction, which I thought is extremely cool and maybe dual lingo will eventually jump on something like this. And we'll see dual lingo just being part of our everyday augmented reality experience. Do you want to share a bit more about that work? Yeah, no. So that work is called a Earl and that's that's embodied AR language learning. So it's just this application where we really, and first of all, I mean shout out to Cien Kim from from Seoul National University. She is applying for PhDs now. So we'll see where she chooses, hopefully here. But, you know, I'm sure she'll get many great offers. But yeah, no, it came with our discussion of how do we, you know, and I love having these sort of, you know, thought processes is, say you have a phone, say you have smart watches, now you have AR glasses that are inherently visual and has cameras. Then, you know, what would these traditional platforms like language learning, for example, get you to transform? How would these applications be transformed if, say, there were to be on AR glasses? Right? And well, we thought of dual lingo, which isn't that where, you know, you can open it up and you can learn languages, which I think is great. But we were like, hey, can we make this more subtle now that, you know, we have AR glasses that may be scanning your environment from time to time? And we built this app where, you know, it would detect your everyday activities. So say you're drinking a cup of water. And the system would, of course, have to know that you are trying to learn French. And that may be your first language is English. Once it detects that you are drinking water, then it will essentially, depending on your level, give you say, you know, a cup or water or, you know, say you're more of an intermediate learner, then it will say a man is drinking a cup of water, for example. Right? So then it's this concept of, you know, from everyday activities, how do you subtly suggest and how do you subtly learn languages without having to say pull up your phone and open up a physical app? Because, of course, a lot of people are very motivated to do this by, say, streaks. Right? That maybe, you don't want to break, but, but also, we are very well aware of people who, you know, download dual lingo and maybe never use it again, because they have to open it and they're lazy to do so. And maybe it's just a subtle way, but we also jokingly said, maybe we're just forcing people to learn a language, as opposed to, you know, then having to open the app is just always there. So whichever way you want to take it, I think it is an interesting way to learn languages, for sure. That is great. I, the next step I want to see of this, I just recently learned I have a friend who hasn't broken her streak in 500 days.
And she's been the winner of every diamond league. Shout out to Leah. But I just recently learned that Doolinggo started introducing games as part of certain languages on the app. So maybe that would be the next thing where you see a few people that you have to interact with in your augmented reality. (laughs) So that would be a cool application. And she started to cater in from our project team because she mentioned the same thing. We just wanted to see what the community thinks first of the project and we didn't necessarily have any gamification, but she also said one thing she said was if you interact with say 20 objects or something, right? Then you can kind of get like a quiz app, even all the objects that you've interacted with in the past. And that way it's a little bit more relevant to you and people are more likely to remember it, which I think is also another great way to do this. So gamification is certainly a great way, I think, here. Yeah. - That's great. I noticed on your website that you also, besides playing games, you're like creating games. What games have you created? - Yeah, no, I mean, I, so I have a couple of high school friends who were into game development and that's how I actually got into computer science in the first place. We did a lot of say hackathons, which is essentially you get 48 or 24 hours to then just build up a game. Rarely do you get sleep, but it is kind of worth it and it's kind of fun from time to time. Back then I was, I guess, like 15, 16. So maybe I was able to do that. Maybe less so now, but yeah, no. We built this game called Bright Futures. That's probably the game that we are at least the most proud of, certainly, and we are trying to get back into it now that we have our thing-a-day jobs more difficult to do so, but Bright Futures, it's looking at a lot of games are looking at sort of like shooting for the purposes of kill. We said, hey, what if we had a zombie game where the goal isn't to kill the zombies? And then we ended up having this really weird game where the more you kill, the more karma you get and based on that, and there's this sort of agent that's determining your punishment from time to time, and it would just give you a random punishment for the kills you do. And with maximum karma, I think you just get struck by lightning or something like that. I think it's just the game that we came up with as high as we'll see it in. At the end of the day, right now we are trying to do a lot more, but trying to find the time to do a lot more as well. But yeah, game dev, I think is certainly still a very much a hobby of mine and I'm trying to do a little bit more of that. Well, speaking of game dev, we had a really, really, really, really great high school student as an intern in the Makeability Lab, Brian, who the Bleepie started college. So I should really find out where he went, 'cause I'm sure he went to somewhere great, but I'm him and I built this game where it's a mobile game but then you would have spiders that are crawling around your real world surfaces. So it would scan your environment, say the couch and your desk and things like that and in which sort of understand what the approximate shape is. And more detailed and say, "Razar would." And then you would then be able to kind of crawl on your actual couch and then climb on top of your couch, go and pop up the walls, go crawl above your ceilings and then jump down at you and things like that, which is a terrifying game, but just another example of how computer vision and AR can be used in creative ways to maybe turn that into a game. And I think that idea is still fascinating to me and we're trying to see maybe there's either a paper thing that we could do or maybe we could just, to pull it onto the app so our people to play and maybe complain to us, say it's a terrible game. - It's pretty cool. I wanna go back to the VRA or glasses and kind of the opposite side of the coin where there are issues with it and that we certainly know and are aware and the challenges or like in particular, what I'm thinking is the challenges of prolonged use of VR glasses and the bulkiness and the de-ziness that people experience. And in particular, really, I also think of populations with disabilities that would like to use the technology even to, I work in the space of rehabilitation engineering so we always talked about utilizing VR or technology for supplementing, augmenting rehabilitation, doing that for therapy. But the idea is that most of these populations have even a greater challenge of the prolonged use of these technologies. What do you see is the next step moving forward to making this more accessible for different populations with disabilities and just kind of like how do you address the challenges? I know the HCI folks at UW do a lot of work for VR application, for like rehabilitation purposes as well. So I just wonder if you're aware of how people are dealing with that. - Yeah, I mean, yeah, to be fair, I mean, my parents don't also wanna use VR glasses because it's dizzy. (laughs) And I really try to get them to do it and they refuse. I think certainly in that sense, I think VR is challenging because, well, VR is really an abarro steep level definition of VR here, but to me, it's really about removing and replacing your sensory information, right? And a lot of the times when we think about VR, well, people start to say, well, visual VR because when people think about VR nowadays, we're thinking about replacing our sort of visual perception, right? By covering it up and you're seeing this whole virtual world as opposed to your real world. I think in that sense, I think the reason why there is a bit of a dizziness and motion sickness in there is because the brain, for example, there is a difference between sort of your brain thinking you're moving versus are you actually moving in VR or not? And if there is a bit of a discrepancy between what the brain thinks and what your eye sees in VR, then you start to get dizzy. In that sense, I think VR and solving dizziness is quite challenging, whereas something like AR, I think, has much better sort of everyday use cases because you are still looking at the real world, right? And in that sense, you're not necessarily replacing any part of your sensory information and at that point, you know, I think there's a lot more sort of wider use case, and you potentially use cases with AR glasses. And that's part of the reason why I think my research has primarily focused on AR glasses as opposed to VR. I think VR has a lot of benefits still though. Like back when I was at UIUC University of Illinois, we were looking at really two directions, working with hospitals, actually. And this is part of the Carl Medical Center that's at UIUC and with the hospital that's nearby, so shout out to them. But they were looking at really two directions, one of them being, "Can we use VR glasses to help people who cannot travel?" Well, to them, you know, travel to these different locations because there are ways to view 360 images and videos in VR as if you're at that environment yourself. So we were doing little experiments where, you know, there are available footage of, say, 360 videos of people enjoying universal studios or, you know, a Disney world, for example, and we try to see, you know, that would help them necessarily, you know, that would help them sort of, you know, travel, so to say. That's one thing. And then the other thing that they were looking at, which I think I am quite personally a fan of, is using an AR game to get that concussion on the spot. So for example, you are then chased, you know, a bird and you were to sort of like walk a run, try to sort of, you know, catch the bird that's running away from you. And based on your sort of movement data, as opposed to sort of, you know, in regards to sort of how you tilt, right? And whether you fall, for example, and whatnot, this is all being tracked as data. And then really looking at the data and seeing if there is a possibility that the person, and specifically someone say, in a sports profession, has suffered concussion or not on the spot. And that way it's a lot easier for them to then, you know, do some early on treatments and whatnot. And I think that project is also great. Again, it's a mobile app. I don't know whether they published or not just yet, but I'm sure it's coming from them. And I think there are definitely use cases here for rehabilitation and whatnot. I know there's also people who are looking at this from, you know, like controlling emotion, which I don't know if I buy sometimes, but like simulating different environments for the purposes of helping people be more calm,
and whatnot through musical and auditory feedback plus visual feedback and seeing whether that can help control the emotion for certain individuals and whatnot. So I think there is definitely a space for medical and rehabilitation field for ARVR. We definitely also need to venture more into space beyond just what people already know in terms of, for example, I think a lot of people are looking at surgery, assistance and, for example, emotion control even. I think the reason why I brought up the concussion detection is because it's a step up from sort of applying AR and VR in sort of a more creative way than say, you know, something that we've already known and then maybe it's just a matter of engineering as opposed to, you know, us, you know, doing any research into space. Whereas concussion, I think that's just a really, really new and novel way to apply AR and I think that's just a really good project. That's cool. Thanks. What are you most excited about in the future of HCI and accessibility, NAR? Yeah, no. I mean, first of all, one of the things that I do want to do is I will likely be pursuing industry as of now. Maybe I'm breaking Charles Hart by saying this, but I think for now that is my plan because a lot of companies are investing heavily in NAR right now and I want to be sort of part of that. I think part of that to me again is the fact that I've done accessibility research and a lot of people come to talk to me about the potential access ability use cases and the fact that if I was to be an industry, then hopefully I can, you know, bring this out as product, hopefully. For, for people as AR glasses is rolling out, then we are already starting to think about accessibility as a problem space that we need to tackle on a company level as we bring out products. And that's something that I'm really excited to do as a HCI and AR and accessibility person is to really get ahead of that, you know, technology curve and say, hey, we can do this now before this thing becomes more public. I think it is really where I want to be at and that's why I think me going to industry. And it's something that I'm probably more interested in now is just them to be able to do that. And yeah, I think the fact that, you know, I am, you know, doing this research a lot of more people are, I think, interested at least in applying AR for accessibility use cases and for HCI, which I think is great. So just glad to see more research into space and hopefully by the time I graduate and do what I can again be ahead of that technology curve and hopefully bring not just, you know, not just features for the general public, but also things that are more accessible than for people, even for example, say you can generate anything, right, in AR, like I did at my antit, well, who's to say that we can't, you know, add all text to it so that people who are applying, innovation can also understand, you know, the thing that's been built, right? And we think about metaphors, who's to say that metaphors can't be accessible? You know, one of the things that becomes very interesting, and this is an example that I used to a lot of people to then maybe get them to understand what accessibility is and they are, is to say you generated a Pikachu, right? And it's just sitting on grass, maybe fancy, right? And I asked people, how would you capture this? What would you tell someone who's blind and low vision, what this experience looks like? Right? And I think a lot of people who are not so familiar with captions and old text would say, hey, there's a Pikachu dancey on grass. Not about caption, but if you're congenitally blind, you, the Pikachu really doesn't mean anything to you, per se. So then, you know, is that caption necessarily good? And when I pitched that question, then they say, oh, cool, like I guess I never thought about it that way. So then, so then, you know, it's really interesting to have conversations about accessibility, I think with people who are maybe less aware, and then also to then, you know, again, I just want to be a person who builds for everyone, right? That is something that John and I thought of is this context aware AR for all, right? I think AR certainly has potential to be used by everyone and then really everyone. And that's just the world that I want to maybe try to shape up. This is great. Thanks so much. Thanks for coming in and sharing your awesome work. And I wish you lots of luck in the following years. You have just a few more years of the PhD and then a wonderful career in industry, making very accessible AR for all. So thanks so much for coming in, Jay. Yeah. No, thanks for having me, Sasha. This episode was powered by Create, the Center for Research and Education on Accessible Technology and Experiences at the University of Washington and Resna, the Rehabilitation Engineering and Assistive Technology Society of North America. Thanks for making it all the way to the end of this episode. I hope you enjoyed it. This is one of the last two episodes of the season and years of progress will be going on a break. But we will be back with new and exciting guest speakers pretty soon. In the meantime, please listen to the old episodes and consider sharing it with friends who might enjoy it. And as usual, stay tuned for the next episode.
Podcast Summary
Key Points:
Jay Lee is a PhD student at the University of Washington studying HCI, AR, human-AI interaction, and accessibility, focusing on blind and low vision populations.
He discovered HCI as an undergraduate at UIUC, shifting from AI research to understanding how people interact with technology, later pivoting to accessibility after a project with blind and low vision users.
Key projects include
Accessibility research in HCI is growing, with more papers on the topic appearing at general conferences like UIST, and principles from this work often transfer to general population applications.
The research process for accessibility is similar to general HCI—identifying problems, building solutions, and testing with stakeholders—but Jay focuses on residual vision for low vision users rather than auditory feedback.
Summary:
In this podcast episode, host Sasha interviews Jay Lee, a PhD student at the University of Washington’s Cability Lab. Jay discusses his journey into human-computer interaction (HCI), starting with AI research at UIUC before pivoting to study how people interact with technology. His focus shifted to accessibility after collaborating with blind and low vision communities, leading him to explore augmented reality (AR) for this population.
Jay highlights three key projects: Baller, which uses AR glasses to enhance visual cues in sports like tennis for low vision players; CookAR, which augments cooking tools to indicate safe and hazardous parts; and RASSAR, a mobile app that scans indoor spaces for accessibility and safety issues, such as insufficient clearance for wheelchairs or hazardous table heights. He notes that accessibility research is increasingly featured at general HCI conferences like UIST, and that principles from this work often benefit broader populations. Jay emphasizes that his research relies on computer vision models and prioritizes visual feedback for low vision users, who prefer using residual vision over auditory cues.
He sees potential for tools like RASSAR in platforms like Airbnb for accessibility mapping. Overall, Jay’s work demonstrates how AR can address specific challenges for blind and low vision individuals while advancing technology for all users.
FAQs
HCI is the study of understanding what people need, building necessary technology, and studying how that technology interacts with humans, often called participants.
Jay initially worked with general populations but got into accessibility after collaborating with Professor Anne Hong-Woe on Image Explorer, which assisted blind and low vision people, and he enjoyed working with that community.
It's an AR system that helps low vision people play ball-based sports like tennis or basketball by visually highlighting objects like balls, opponents, and hoops to enhance depth perception and visual salience.
CookAR augments cooking tools by using AR to highlight safe and hazardous parts, like handles versus blades, helping users understand object orientation and use safely and efficiently.
RASSAR is a mobile app that scans indoor spaces via video to detect accessibility and safety issues, like insufficient space for wheelchairs or hazardous items on low tables, for anyone to use.
Low vision people have residual vision and prefer visual feedback, so AR glasses provide visual enhancements rather than auditory cues to leverage their existing sight.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.