Go back

Will AI Take Your Job?

56m 43s

Will AI Take Your Job?

The podcast episode explores the role of artificial intelligence in dermatology, focusing on two key studies. The first, a 2023 single-blind study from Poland, compared ChatGPT-4 to dermatologists in answering patient questions about hidradenitis suppurativa. Patients and physicians rated AI responses higher in quality, empathy, and satisfaction. However, when asked whom they would prefer, 70% of patients and 80% of physicians chose human doctors, indicating a persistent trust in human interaction despite AI’s superior performance. The second study, from the Mayo Clinic, used AI-generated deepfake avatars of physicians for patient education. Patients showed high acceptance, with usability and trust scores near 90% and low eeriness, suggesting AI avatars could provide 24/7 empathetic communication without fatigue. The discussion highlights AI’s potential to streamline routine tasks like scheduling and patient counseling, allowing dermatologists to focus on procedures and complex cases. Experts note the technology adoption cycle, comparing AI to early Uber or Tesla self-driving, where initial skepticism fades with familiarity. Challenges include AI hallucinations and the need for high accuracy in medicine. The consensus is that AI will augment rather than replace dermatologists in the near term, especially in procedural areas, but long-term advancements in robotics and imaging may further shift the field.

Transcription

10639 Words, 56082 Characters

English
Welcome to season two of Derms on Drugs and Video Podcast brought to you by scholars in medicine, the best educational platform in dermatology and provided in no cost to medical providers. Derms on Drugs is where cutting edge derm meets the intermiss comedy Matt Zyres from Dr. Mithology in each week. I'm joined by my residency buddies, Dr. Laura Fares from the University of North Carolina and Dr. Tim Patton from the University of Pittsburgh. And we use our 60 years of combined derm experience to discuss, debate and dissect the hottest topics in dermatology. It is everything you need to know to be on the cutting edge of Derm and you'll actually have some fun listening. New episodes drop every Friday on Scholarship Medicine, Apple Podcast, Spotify and other major podcast platforms. And there is a video component that has to keep figures and tables from the articles we talk about. So we are so excited this week to be on a deep dive into one of the hottest topics out there, artificial intelligence in dermatology and in medicine in general. And we have the perfect guest to get into it with Dr. Farah Khmanger. I'm terrible to pronounce him last names from Silicon Valley, who is the owner of the company that runs Derm GPT, which we are actually going to be talking about a little bit today. There's some interesting literature out there that actually from what I could tell has no commercial connection with Derm GPT. So I'm excited to see where we go. Dr. Khmanger, great to have you here. Thank you. Thanks for having me. Very excited. And let's just go ahead and get into it. So why don't we start with Dr. Patton? Patton, what do you got? My deep dive paper, May 2025 edition of advances in dermatology in Aller Gology. Aller Gology? I think that's how that's pronounced. Sounds right. Sure. It's tight. That's why I didn't go into Aller Gology. Yeah, it's titled comparing physician and artificial intelligence chat GPT for responses to common patient questions regarding hydride 90s, a single blind study by Lewandowski at Al study was performed in 2023 at the University of Gdansk in Poland. Many questions about HS were answered by chat GPT for and also the dermatologist, that's air quotes for the listeners. Maybe Poland only has one dermatologist in it. They didn't really go into details who answered those questions, but that's what they said. They were labeled answer one answer to and respondents, which was 30 HS patients and 31 physicians were asked to rate the answers using a Likert scale for quality empathy and satisfaction. So chat GPT for kind of crushed the dermatologist patients rated the quality empathy and satisfaction of a chat GPT GPT for as responses higher than that of the dermatologist physicians also rated chat GPT answers as better figured two and three lay out the data pretty nicely robots are better than live dermatologist. Let's just face the facts. Maybe eight percent of the time patients said I like the chat GPT for response better. So we'll all be out of jobs pretty soon and I for one welcome our new chat bot overlords. But wait, when respondents went through ranking all of the answers, they were asked the question so at the very end of everything. They were asked the question given that AI can answer your questions more accurately and pathetically than a doctor. Who would you rather receive an answer from the doctor or the AI? It's almost like the authors knew that the AI would be better. It's almost like they came up with that question ahead of time. Anyhow, 70% of the patients and 80% of the physicians said they'd prefer to get HS answers from the dermatologist. So for now, I guess our jobs are secure. It's just such a weird like why what's any explain. I mean, yes, you like that human interaction. But I mean, as time goes on, is that going to become less and less important? It's pretty interesting. And this is from 2023, right? Yeah. So right, chat GPT has gotten a hell of a lot better and the dermatologist have not gotten any better. Right. It's a really interesting that have anti-AI-ness. Like I have to Tesla for a while, got rid of it now because the lab is not a good driver. But even though my family knows I'm a terrible driver, like get into accidents all the time, they still prefer me driving to me letting the Tesla drive. And I couldn't figure it out. I was like, what you, I'm a terrible driver. This is better than me. And they still wanted it to drive instead of me. The only thing I could do is black box thing. Like we don't understand it. And I think that that's probably a lot of what it is. But I also think it's going to get better with time. Like do you remember when Uber came out? They're like, here's this great thing. You're going to go on your phone and you're going to tell a stranger where you are and they're going to come put you in their car and drive you places. I was like, that will never take off. I would never do that. Now I'm like, I like that way better than a taxi, right? Like it may better. I just think we're, it's taking time. I think it's like a little bit of, it's going to take time to get people around to it. Your audience is, so right, this is kind of talking about the patients as the customers. Your customers with DermGPT are dermatologists. I, what has been your experience with this of, of do you get people who call you? You're like, I can't believe you're doing this. You're going to ruin with this. And other like, and other people are like, oh, this is like what's your experience? Yeah. It's really interesting. I think what we're talking about is the technology adoption cycle, like that classic. You have your super early adopters that are into things like, hey, the Uber or the Waymo, let's do it. And then you have kind of the stragglers and then you have the really late people where they're just like, okay, this is just the norm and I'll find I'll take a new bird if I have to. I think the cool thing we saw with DermGPT was we launched in 2023 and I think within a few months, we had over like 13,000 docs that had come to the site. Right now we have about a good, a little of a 4,000 super users. So I think there's a difference there. There is the technology in general, like in 2013, 2016, all the electronic health records came out. Everyone was like, oh my god. I remember I was at UCSF when Epic rolled out and some doctors retired. They were like, that's it. I'm good here. I'm leaving with technology at this stage. But AI has been different. I think generative AI has been taken up a lot more easily. Even just open AI's chat GPT itself up a bunch of millions of people were on it like instantaneously. It was the highest adoption of a technology. And I think part of it is it's intuitive to deal with it. It's language based and language is how we communicate. It's like when you talk to someone and it's talking to you in a similar language, you're able to just connect with them more easily. And then these large language models, the reason they're beating out the doctors, they're created to just be super nice, like super customer service, they're creative for you to want to keep engaging with them. And you know, as we're busy and we're not really a lot of things to do. So our communication maybe has the qualities, drop the little as far as how we communicate. I'm not at all surprised that in like a information or communication task, it would be this out in like niceness or you know, pleasantness and all that. But I think we just have to realize what AI can do and what it can't do. And there's still a lot that it can't do. So we still have jobs a little bit longer at least. I think we're. You think we're so out. Yeah, that has us going the way of extinction, dinosaurs, all that. I think we're still okay for just a little while though. Yeah. What do you think the explanation would be for again, after going through this and knowing that they prefer the AI response, what do you think there is about saying I'd prefer to talk to a human? I think that's still that part of the early adoption cycle like over time as AI is there and people know the answers are good and it's doing a good job more and more so that preferring to talk to a human thing will probably fade a little bit once we have the accuracy and the trust. But I think the good thing is the models are not as good as they seem. It depends on the task at hand. For example, like the communication piece is set up that way. So it's going to do better than us. But there's a lot of hallucinations that still happen, like a lot of the work we do with GBT is how do you actually make it so that it's correct a lot of the time because good enough isn't good enough for medicine, right? It needs to be correct. They can't, you know, anchor bias to you to the point where it was something you're making. So as I follow the autonomous driving thing closely, like I'm not like a huge Tesla investor but I literally got one because I wanted to see like how well the full stop driving works. So I invested a fair match, but and the biggest problem is that people hold self driving up to a standard of perfection. Whereas the right standard is 1% better than humans. And the same thing. So like the one AI that I sent out today, I had a few non medical friends like, hey, you had that problem recently. Go try this thing and see how well it works. And they did it and they, you know, send me the results and whatever. And they're like, yeah, so it didn't get it, you know, right? And AI sucks and it's terrible. And I was like, but your doctor took a year to, it had your thing second on the differential. It took your doctor a year to even mention or think about. about the thing that you, like it got immediately. And they're like, yeah, but it didn't get it right. And you can't trust AI and computers are terrible. And I'm like, maybe they are, but doctors are worse. - It's true, it's comparative, right? 'Cause like for the, so DermGPT's large language base, but a lot of the Derm over what we're trying to do is diagnose melanoma and image base. So they put out the sensitivity and specificity for this device to like diagnose melanoma. I thought about that. I was like, what's my sensitivity and specificity? Like how many moles have I biopsy before I got to a melanoma? It's true. And actually, I'm like nerdy enough I calculated that one. It's still better than the models that are out there. But I'm biopsying a bunch of moles too. Like we're not like 100% but melanoma walks in within five feet. We're like, that's a melanoma. - And you don't, we always know the melanomas you miss, right? So that's what's so hard. It's hard to do sensitivity for humans. - Yeah, it's, it's the same. - Right, for melanoma detection. - And the new imaging stuff that's coming out. So we did an episode on this several months ago. And there's like this new device that's like picture, picture, picture. And it's doing cross polarized dermoscopy on every single mole on your body with like a, you know, 10, I don't know if it takes a minute to take the pictures. It's nuts. Like that's going to be orders of magnitude better than us. Very quickly, right? Very quickly. - That's true. - I do actually think derms are safe because we're so procedurally based. I think that what I think is likely to happen is that medical derm, routine medical derm gets quicker and faster so that the, we're able to meet the demand more effectively and more of our time shifts to doing procedures. So we're all going to make more until the robots get good enough. Once the robots get good enough, just do the biopsies and the, I figure that's, I think that's 10 years out before that happens. - I think the first cataract surgery done completely by a robot has, it was recently done. I just know this because I'm interviewing chair of ophthalmology candidates. So you're always like, what's the next, you know, new thing in your field? And like three people were like, we just had the first cataract surgery completely performed by a robot. So I mean cataract surgery seems harder than the skin biopsy to me. - Right, right. - You can do that. - Right. - Wow, that's not good. - And you know, the task, actually mad that you talked about with the polarized photography, the task that it will beat us at 100% is that longitudinal follow-up of that mole. Because if that person comes back in six months and it takes another photo of it, there's no way I wouldn't remember what that mole looked like. And even a photography currently is not that good to the point where if I put my two pictures next to each other. So there are certain things that's going to blow us out every single mole in their body, every single mole in their body, it's going to be a slightly comparative digital deroscopy. - There is like such a huge issue with overdiagnosis of melanoma and of thin melanomas, right? And so there are things that we biopsy and then they come back as early evolving melanoma and site two. But you know, what is like the definition of cancer? So uncontrolled growth of cells. If it is, if a lesion on your skin is absolutely not changing, it kind of doesn't matter what it says histologically. So I actually think it's going to get us to the right answer. And how many of us have had not like, oh, that's an early melanoma, I almost didn't buy. I mean, we've all had things that are like, geez, that's like a point nine millimeter melanoma and that did not look like an a, it fit the ABCDs and I could have missed that. But that would have been newer changing on photos, right? So that is to me what makes so much sense. Like the fact that we as highly educated people look at every inch of skin on people's body over and over again all day long as a cancer detection tool is crazy. Right? That's a good point. Yeah, it's like at just the merits of what we do daily. It's probably most of what I do. You know, we all do everything, but that ends up in most of what we do. That's a lot of what we do. It's probably greater than 50% of our workforce effort. So if we could then make it get the things that are really truly the things we need to see, we could just do so much more for more people and a percent. And to your point, it's, I think what AI will do is take out this kind of mundane level so we can do the more complex things. And the question is, you know, what are those next complex things that the AI won't be able to do as easily? I think AI is going to do the complex stuff too. I think we're all going to be sitting at home with our humanoid robots taking care of us and maybe doing podcasts so that we have something to do. Who knows? All right, I'm going to jump on to my article, which was Artificial Intelligence Physician Avidars for Patient Education, a pilot study. So essentially what they did, this was plastic surgery, it was done out of Mayo. They basically took a couple pictures of the doctor, took a short voice sample, had AI make a deep fake of the doctor, then had the doctor like type out his, responses instructions for several, like for 10 common questions. And then for post op, they had patients be like, "Okay, like we're doing this experiment with AI, "you're going to sit down and it's, you know, "they've done in a room and the clinic and the whole thing." And it was just like a telemedicine visit where the avatars was on the screen, the patient was sitting here, they asked questions, the AI, you know, picked which of the scripts to do based on their question. But the AI completely generated all of the voice, the inflection, the facial expressions, all of that, some other doctor that was type out ones. And so this was not looking at content, it was looking at patient acceptance of this. And the main takeaway was that patients completely accepted it. So it had a, they did this radar chart, which like kind of is a way of looking at different aspects of things. Usability was close to 90% engagement, was 85% acceptability and trust was 90% realism, was about 80% and eerieness, which is, you know, the equivalent of creepiness. Like was this weird, like did what it was close to zero? So on a one to five scale, the mean eerieness or weird scale was 1.57 out of five. And so the takeaway from this was patients were totally accepting of it. And I'm at this point pretty certain that, right? 'Cause it, so they could now create a digital twin of all four of us. And pretty quickly the LLMs will be able to tell the digital twin what to say and patients won't be able to tell it's not us. Now ethically we'll start to say it's them, but imagine if patients instead of, you know, calling your office, "Hey, blah, blah, blah, blah, I'm red and peole and itchy after that creamy gave me for my acne. If they could do a video visit with the virtual you, anytime they wanted, right? 24/7, you're gonna spend as long with them on the phone as you, they can do it two hour virtual visit with you. "What about this people? "What about that one?" And you're patient and kind in the whole time and you'll never get annoyed with them and everything else. Like, it's, we can't compete. We cannot compete. Not even if it's close to us in terms of like knowledge and content when people are like, "Oh, but it can't do empathy and it can't do the human touch and people are like, it is better than I said that stuff." That is what it is particularly good at. Is a video avatar in it like, which is, it does not fatigue. It can be on your best hair day, right? It can be you and your most flattering shirt color, I mean. - It's never like, oh, I've been in this room for eight minutes. I'm three patients behind. I need to get moving. You know, one visit you're gonna need to reschedule for that. No, it will spend three hours talking to them if they want. Like, people put you, well, a lot of visits is just psych. It's not real medicine. It's not just, I know that's why they like it more than you. It will listen to them for days if they want. - So Matt, I love how you go from to like, the most extreme use case immediately with this. 'Cause this is with this is you. I like it. But here's my thought. Is there like maybe a more intermediate? Like, maybe this would be great for counseling. Like maybe this is me counseling on what a skin biopsy is. They could have me doing this and I could, you know, update it for anything, an excision, cryotherapy, right? I mean, I have thought it like, if I call, I mean, I left something at a restaurant the other day. I called there to ask like, hey, did you find my, do you find any, you know, air pods? It was AI that answered. It's like, hi, what are you calling about? And, you know, I pretty quickly picked it up that that was what it was. But like, it asked, oh, did you lost something? What did, like, it can ask that and then, you know, provide that. Like, if I, a lot of restaurants, when you call to order food, it's AI taking your order, right? Like, certainly scheduling visits, there is nothing magical about the people who we pay to schedule visits, right? And, but the magical thing is that they can only work between like eight and four 30 and then we can't schedule appointments at nine o'clock at night, which is when you think like, geez, I probably should go to the dermatologist. So like, maybe even the more mundane things, when it is much more human is where it could be. be sooner. Yeah, so my private equity group is implementing this now. Like, we've got AI agents that answer the phone and, oh, you're calling about, you have a question with this. Are you want to schedule an update? Just exactly what you said for the restaurant. Like, we're doing that and we're, have physicians beta testing the MIAI scribe thing. And we're going to be rolling that out in the next year, probably the next couple of months. I mean, it's, it's nuts. It's nuts. Yeah. I mean, a lot of us are doing AI scribe. This is next level, which is giving you what feels like a very human interaction. And in fact, with a human that you know, right? When, way back when I was at OSU, I made about 20 videos of me counseling. I remember this. Yeah, it was like co-commeter purple bettaine, lanolin like this, that and the other. And I used to experiment with it and like, okay, today I'm going to use the videos. And the next day I'm not going to use the videos. I'm going to actually, it's literally the exact same content. It's just me delivering it face to face versus a video of me delivering it. The video of me delivering it, at least an order of magnitude better. So the patients retain the information better. They acted on it more effectively. And it was just nuts like I could say it and they'd be like, well, what, what, and I don't think of it. And I would have the video say it and they'd be okay. Absolutely on board. It was nuts. It was nuts. I, I, and I think that maybe that's why I'm so like this is going to put us. Like it's better. I mean, like I think patch test counseling would be a great use for this, right? It could be you. It could give all the information. It could run there. Whatever those things you guys do that list that camp. Camp list. Yeah. Right. And it could say I've gone through, you know, this is what the camp list is. And they could be like, but do you think it's my shampoo? And they could be like, no, do you think it's my fabric softener? No, it turns out coca-middle propobatine is not in fact. And the way it is. It could have that conversation. Yeah. The way it will be soon and camp is putting this in. And the way it is, it's like, you can scan the barcode of your shampoo. And it will tell you that shampoo is okay. That shampoo is not okay. You can take it into the store with you. Is this shampoo okay? It scans the barcode. Like, it's okay. No, it's not okay. Like it's, yeah. It's the American Tech Derm society is doing a really good job with, with the next version of camp. Yes. Okay. All right. Okay. So I have a, a, a research letter in J.M.I.R. Dermatology, which is evaluating artificial intelligence models and dermatology, a comparative analysis by Patel at all out of UC Irvine. So this was a head to head chat GPT versus Derm GPT. So a survey. I'm sorry for interrupting. And Farah, did you guys fund this? Have anything to do with it? Is this a conflict of interestee thing? You all just happen to be in Irvine, California? Yeah. So, you know, we did have, I think some of our med students were involved in this. So not a funded thing, but definitely like a academically related. Okay. But not funded because we, Pat and his covered on the show that when it comes to supplements, if the company funds the study, they, they're met and out. And it's really funded. It works. And if the company didn't fund the study, it didn't work. So it's okay. So true, isn't it? We did not fund it. Okay. All right. All right. Farah, go ahead. Okay. So survey, basically a survey based comparative study where it was dermatologist rated the answer, but that was the term GPT gave versus what chat GPT 4.0 gave. And they were common, derm questions. They were written by two derm residents and faculty and trainee. And then the faculty, so the derm residents were at the questions. The faculty and trainees judged which one was better. So what, how, how was it set up? So three of each, interestingly, and I was going to ask you about this. There are three of the questions were actually dropped because Derm GPT just said, I see your dermatologist for guidance. So I thought that was kind of interesting. Derm GPT, maybe you guys have safeguards in there that if it really does not have an answer, it's less likely to hallucinate. It is more likely to just send you to your doctor. And then they were, what's up? Let them hallucinate and let them, yeah, let us hallucinate. Exactly. So then they like they were blinded as is it model A or model B. And then, and then basically attendings and trainees at UC Irvine and UC Davis were invited to like do a survey and either say model A was better or model B. So here, what, what about sample sizes? So 64 dermatology faculty 30 residents. There were, yeah. The respondents were, those, sorry, that was who was, that was the sampling frame and then the respondents were 19 people total 13 attendings and six residents, actually 19 people who did, who actually agreed to do it. And so you had basically 258 answers if you take every person, if you take all of the different answers in every single like combination of rating basically. So it's many ratings from a, from 19 people total though. Okay, so what did they find, Durham GBT answers were prefer were preferred more. So Durham GBT answers were prefer 48% of the time versus 28% of the time it was chat GBT the rest were, were either like equal or both inadequate. This was statistically significant among attendings. Durham GBT was favored and this was all kind of sit to a similar degree for the residents as well. Now they also asked for, you know, so basically there were more like tight to the point better answers out of Durham GBT. So each one, each model LLM gave references and it turned out that when they asked the Raiders which references were better chat GBT's references were preferred more frequently than Durham GBT. So the references that it gave were more preferred for chat GBT, but the answers given were more preferred when they came from Durham GBT. Yeah, they were trying to make it look like a balanced study. We got to find something good to say about cheap chat GBT. Otherwise, we were going to think it's. Although I will say I believe that because the user interface on the larger foundation models of how the references show up, it does beat us 100%. So that's probably true because they they have it really, it's really pretty like they have the logo of the journal and then they they bring up the actual, you know, we're for us. It's like a link to the PDF. So I'm not surprised that our references lost. We should put some effort into making them look pretty to making them look. Yeah, more. And I guess they also said that there are like chat GBT was more likely to pull from like higher impact journals like J.A.R. That was what the paper said. So yeah, kind of kind of interesting. That's super interesting. I looked for a while. I couldn't figure out. So this you can round it off to that 50% of the time people preferred Durham GBT 25% of the time they prefer 10 GPT and 25% of the time. It was either equal or both answers sucked. And I could not find anywhere where it said was that other 25% that 25% of the time both answers sucked or 25% of the time both answers were equally good. It just said 25% of the time neither answer was preferred. Yeah, I don't know that it at I think it just said do you prefer this or not the other I don't know that it gathered that level of information. And I think that's what I know that it gave the four possible answers were prefer A prefer B. Okay equal or neither are both inadequate. I think was the exact terminology. So I like I like it really that matters a lot right if 25% of the time if 24% of the time it was both answers suck. Matters a lot right versus if it's 25 24% of the time both answers were equally good and 1% of the time both answers were inadequate like that. That was the most interesting thing and the whole thing to me and so yeah they only gave the answer as other yeah which was not this was better or that was better not like it was it could either be chat was better Durham GPT was better or other. They did not actually break it down that I can tell unless it's somewhere in a supplement that I'm not aware of but I'm not even saying it referenced in that I want to do the supplement looking. You did okay you went through either more I could not find that yeah because it'll it did not break it down to that level so. Alright so Durham GPT B chat GPT now right we beat it in this one task in this one task of dermatology question so yeah how do things like you guys work so I kind of assumed that it is like a. Branded like with Verizon they've got Verizon and then they've got some other one words 25 bucks a month and so. but still Verizon, but it's like just different branding. So I've always assumed that like DermGPT is just chat GPT or Gemini or whatever with a different label. Like how do you have, if you wanna make your own thing, how does that work? - Yeah, I know I did that. That's a great question 'cause these tools are actually very different. And I think it's important for people to know how to use them, 'cause they all kind of go into this bucket of AI. But it just depends on how the model is trained to work. So for example, for DermGPT, it's what we call a retrieval augmented generation model. So you take a general foundation model like a GPT, which are actually really good. Like the chat GPT, clawed, these models are excellent. So you don't necessarily always even need something specific. You could just go to these, and I often do, like I love clawed, I use it all the time, but the foundation models are so good. But what happens is sometimes they're like too good, they're just making up stuff. So I think that's the whole problem we're trying to fix for medicine is how do you just really reduce the hallucinations? And Laura, like you called it, if you don't have an answer, don't give an answer versus the big models are set up to, no, just say something, sound good, make the person happy so they wanna keep engaging. - Keep 'em on the line. - You might be too young for this, but when Pat and Ferris and I were kids, there was a 1,900 or the 976 or whatever, it was like 399 a minute, and they were like poor in lines, I'd never called any. But it was like, they just, if they're whole lot of-- - Keep 'em, keep them in its go and get 100%. - Keep 'em on the line, right? That's what, that's how it's-- - It's an old business model, but exactly, exactly, these big foundational models are trying to keep you on the line, that's 100% it. For us, our benchmarks is, please don't get it wrong, please don't get it wrong. But rather you say not say something, then my colleague in like another state gets an answer, and they obviously know it's wrong, that's more embarrassing to us than, the call-bing drop sooner. So, but what we do differently is, when the chat GPT APIs came out, we were like, I think one of the first people to just get the API and start building DermGPT. For me, I've been working in health tech for a really long time, about like, #20 years now or so, and there was always these problems I just couldn't fix that were very language-based. As soon as I saw the generative GPT models, I was like, this is it, this is how we can solve in basket times. Pajama times, people know it's like all the things, scheduling, like it was just, I was like, this is it, this is how you do it. But then you have to make it good. And there are these models that are like, we're gonna save the world and say, no, not yet. You know, there were not that good, there's still limitations to what AI can do, but you can make it pretty helpful. So basically what we did initially, just like other basically, LLMs for medical, we got super excited and attached every article we could from PubMed into our LLM. Like everybody did that at first. We built these APIs, like get every single article you humanly can. And this will create the best medical LLM. And then we kind of found out that that actually produces garbage. So we had like over 70,000 articles, we had all the articles in the world in there, which is what Chad GPT is access to. It could pull anything from anywhere. But then we decided that's actually not good. So we started to curate more and more and more. And the curation processes, myself or another colleague looks at and goes, no, I don't want that, or I do want that. It's it's it's simple to a dermatologist, but impossible to non dermatologist, this kind of curation process. You just look at a journal, you know, you know, you're gonna use that or you don't. So we got rid of a lot. So that was the main thing is curating this derm brain. And then we kind of taught it differently because that's not how we trained to be dermatologists. You don't get their residency day one and they're like, here's all the PubMed articles, dear thing, become a derm, right? We're taught resystematically. So we built our model in that same sense. It really just has this derm brain, has an infrastructure. And then the articles are really on top of it. It's a supportive element. It's not purely the articles that it goes from. So it's it's kind of a interesting how you approach how to solve that problem. On top of that, we have multi-level agents now, which is also different from chat GPT or a lot of the other models. So we'll have like two or three layers of agents that check each other because the LLMs are better at proofreading than providing an answer. You could do this yourself. You can get an answer on chat GPT, go over to clot and say proofread this, come back from clot to chat GPT as they proofread this, do that two or three times. And your eventual answer might even end up being something close to like what derm GPT can do. Not to minimize what we do. But basically we have multiple things set in place. We have cases in place. We have this agent called Derm Guardian that just goes through everything you're doing and makes sure you're not like doing something that's gonna get you sued. It looks for like high risk cases not to miss. It looks for stuff you should put in your notes. So we have all these kind of extra features to it. So this is a cool study. It doesn't surprise me that the final answer, hopefully was better. Sounds like maybe it was. It also doesn't surprise me that our article kind of look and feel and possibly they might have had some higher level articles. But that article might not be where the right answer comes in for this question. Whereas like a chat GPT is built to be impressive. Like there are these other LLMs that they say, we bought like the biggest journal and that's why we're good. And they're just using name association. But that doesn't, in medicine, an answer is either good or isn't. The branding doesn't really matter to us. Like if they had a fancy article or not and the end you judge the question. - I'm not afraid to do that. - But that's a long answer to what we're doing. - Well the right answer in a low impact journal may be a more helpful answer than the tangential answer in a high impact journal. 100% - I'm really good. - I quoted Jamma Derm and Jad. Oh, we're good, you know. But maybe that was in like American journal, clinical dermatology or something. But it's the right answer. It's the right answer. And I think that's what these big LLMs are going after. They're like, well, if we just put every logo of every big, journal on our site, then that's how you build credibility. Which it is to the viewer, but really at the end of the day in clinic at 5 p.m., you just want the right answer. So you just need to build the back end to do the job you need. - So does you know? - Derm GPT, are you see this as and it may be more than like our goal is to have information that dermatologists can use to provide better care or patients can use as a resource to find more answers or to help with the mundane like you started talking about like this is where I'm going to go to write prior off letters or how to like, where do you see your niche really being with Derm GPT? - Yeah, it's the Derms. It's the Derms. And I was a department chair for five years. And I really saw like I used to follow with our docs for doing many of them were on the EMR at like 10 p.m. Prior odds in basket messages. And this is that population. This is who it's for. It's the attending physicians high volumes. We know we're less than 2% of the house of medicine, but we're many times the first entry point of any healthcare system because everyone has skin and everyone has skin problems. So our volumes are crazy. So the tools that are built for like, let's say primary care might see 10 or 15 patients in a day. Tool that's made for that volume is not going to be the thing that what we need when you're seeing 30 plus sometimes patients, most of our medicines need prior authorizations. There's about five sub specialties like us that are hit hard. So like in 2012 or burnout numbers were almost like non-existent. We were super happy. Now they're pretty high. So, but that's because we're in that kind of small subset. Where if you're doing a couple of things inefficiently, you do that times 30, that's a lot versus maybe a primary care doc might have, 10 of those clicks or something in a response. So it's really that small niche. Like we don't solve every problem for everything in the world, but we really want to solve that problem for that. So the board certified germ, mostly medical germ, but I think some surgical germs are using us as well. But the people seeing 25 plus 35 plus patients, prior odds, that's the kind of group that's like drowning that we're hopefully trying to help a little bit with this. So for those of us drowning, can you give us like a vision into five years from now? What could our day really look like with where the technology is and like where you realistically think it'll be in five years? Like what would your day look like? - I would say it's even here today. So you know, the data shows for eight hours patient facing time. We are often spending four to five hours on non-patient activities and that seems to be pretty consistent for the Durham group as well. So those are in-basket messages, prior odds, notes and themselves, which and they're all interconnected 'cause if your note is not complete and doesn't have the prior odds things in there, then your prior odds fails, then you go into denial land. So all these things, they're not separate actually. And then the AI can pull information from your note to answer the patient questions. So they all kind of actually live in an ecosystem. But we've done many trials now where we've seen like a three hour session of in-basket message kind of like response can be reduced to 30 minutes when AI is in a brain. It's not like you're not in the room, you're sitting somewhere else having coffee, you're still involved with it, but it just decreases that cognitive burden by exponentially basically decreases it. So it already exists. That's why I always just try to tell all our colleagues, even if they're just on chat GPT or clot and not just so it Durham GPT. Even those models, like you mentioned, we're not perfect either. So even those models are sometimes a little better than what we could do. So just get on any generative AI and get these mundane tasks. Tila, even if you're like, "Oh, I'm just gonna do in two, three minutes, make that 30 seconds." So if you keep doing that, you're gonna repetitively, you're gonna win back a lot of hours. - But don't they have to get directly integrated into our EMRs and be proactive rather than reactive? Because that's-- - Ideally, yeah. - That's where the big difference is gonna happen. - I think 100%. I ideally, that's the world where you don't have to go from one thing to another side and you know type the information again. And it's coming like a lot of these and Epic is bringing a lot of AI tools. So it's definitely coming. The one thing though is the same thing that happened with electronic health records. These processes come in for primary care and so they're not always, and it does matter. It doesn't matter how the problems the models are thinking and how they're prompting. They need to understand our language like even our scribe systems are sometimes difficult because our physical exam is difficult. The primary care doc just has to say like positives because everybody has a heart and lungs and you know all these things they just have to say was there something abnormal there was there something abnormal there are exams different right you're not supposed to have a mole here but you do so that's like you know that's like it's a it's a different way of looking at our physical exam even and then our words are harder it's not just lungs it's you know erythemitis blah blah blah you know so it's everything is a little bit different so the tools are coming but it's going to take a while before they're actually meaningful for sub specialties which is what I found over time. So my pathway of using AI and medicine basically last year I thought the large language models got to be good enough that I started to like use them for some stuff up till then they weren't usable and then midway through this year I discovered open evidence and thought it was amazing and then decided it was terrible because there were several obvious things where like it gave me the wrong dosing and I was like no I knew I knew basically the dosing but I was like no wait you just told me it was wrong that's not what's the package insert says and then open evidence said no this is the dosing and the blah blah blah and so then I sent to open evidence like in in my chat like here's the section in the package insert that whatever and then it you know oh no the dosing is blah blah blah and then whenever I asked why did you give me the wrong answer before it just ignored the question of wouldn't answer and that happened more than once for me with open evidence so I stopped using it because I was like if I if I have to double check every answer I'm not I'm not going to be using this LLM it's a and I haven't had that happen in any meaningful way with perplexity or Gemini or Groc those tend to be the three that I use and the the challenge that I see for for a company like yours and and this is really where I'm getting to with all this I can literally see week to week month to month even when there's not like a big model upgrade chat Jeep I don't use chat but but the other ones getting better week by week month by month or like that the the the amount of resources they're pouring into trying to make the models better like you guys can't like you're like the underlying like so how do you keep up so let's say you're better than check pt right now yeah six months for now are you still going to be better than chat gpt we're going to be even better than we are how much like the delta's going to increase of how much we get better and the difference is we understand the workflow so this is the crazy thing now that you see people with like two three individual we have five but like you see the stories of like people with two a business of two got some crazy AI product out the model you don't need the numbers anymore because I think the foundation models are so good that we can build upon the big companies have done the work this work came out of Google and the opening I of course was the one that really released if this was a lot of the Google work where they actually developed this generative these generative models so that kind of lift that kind of build there's no way two people could have done like that level of build but what we have now that we can build on top of it and the cool thing is as those models get better we can leverage it too we can make our foundational model you know increase but what we the layer we have on top are the workflow layers with the agents which even like a Google cannot understand unless they get you know a hundred dermatologists in a room and really deeply talk to them which of course they they're able to build it if they get that kind of inside info but even then unless you had those hundred dermatologists in a room for six months you're not gonna get what we intuitively know so it's actually I think it's the time for physicians to build because what we it it comes easy to us so we're like is that a big deal but it is a big deal because other people have no idea what we're talking about it's just things that are coming easy for us but these the workflow layer and the agentic workflow layer is what makes us better let me make sure I understand your answer so it basically you're saying that that improvement that I'm seeing in the print you know just the basic LLM models you guys are also benefiting from that and on and on top of it you're benefiting from sort of optimizing workflow and I'm in iterating yep like we're even way better than we were two years ago because we learn and we iterate so we're getting better in our workflows and oh let's actually change our agent this way because our colleague from this state like he says sent in this answer and said this was totally dumb change this this way so we're iterating on that end and making the agents better and you can and we constantly upgrade the foundational model so if there's a new update that's available that's an easy one to do so I think that it's really democratized what people can build actually on top of the foundational models is the holy grail for you guys integrating directly with EMRs because that's the if I've got to click in and out of my EMR copy and paste like for it to become proactive rather than reactive and really up the functionality it's got to be and I've seen it seems like it has to be integrated directly with my EMR am I am I thinking about that the right way that that actually is that would be like the holy grail I think for the user because it's very annoying to have to go from between one thing to another but I actually I would say as far as AI building where we can be so far ahead but as far as if you're talking about like let's say an integration with apppec then you actually do need like a huge team because that those things are actually a big lift technologically not because it's a difficult thing to do but because the models don't quite talk to each other and models like apppec are a little older they're built on multiple layers so for two groups to come and just understand this intraoperability is actually really really hard I mean this is like these are problems that big groups are trying to solve Medicare right now has put in CMS has put in that you know most of its payers or payees should be using electronic prior odds but they can't figure out interoperability these are really hard things to do actually the infrastructure of health care is really hard to connect to one another so that I would say is like that's the old school hard stuff to do but from the outside R&D development of AI if if you have the right problem-solving skills and you have expertise you can actually out outpace the foundational models but yes 100% ideally everything's in one place you don't have a tool for this a tool for that so hopefully it'll get there at some RUNC Epic actually does have built-in AI I mean in addition describing it will if a patient if you say my skin is red from the cream that you gave me for acne it'll right back hi Matthew thank you so much I'm sure that's frustrating that your skin is red that might be your Trent Nolan some suggestions would be and they're not it's not like super high quality answers but you have to but it is trying to do it and it might make it a little easier for you right maybe you don't have to like type out Trent Nolan which is a win and it's not enough yeah it does not go directly out it comes to me and it says start with draft or write my own so I could start with their draft and then change it Pat what are you gonna say that was a question I was gonna ask I was also gonna ask Farah is there do you have an example of something that like maybe a year ago DermGPT was just like terrible at like oh my gosh like would this needs to be so much better and can you talk us through the process of how you made that happen yeah I would say our agents we've actually put out more in the last year because one of the things was our answers were better just the way our model was more a little bit more curated but it still had errors every once in a while so what we wanted to do is basically have it be 100% correct with what I put on answer but then at sometimes it would not put out an answer which is also not helpful so it was this kind of just this balance of having it give a really good answer most of the time and try to give an answer but have it be correct which which is kind of what we've really put out one of our engineers he actually has white papers on this I think he has a patent on this model with the three tier agentic system where the agents checking one another is probably one of the highest ways you can get to almost zero hallucinations or you know 100% correctness that we can strive to like we said to maybe even better than humans and we're not getting to 100% but it can imagine if we had maybe four attendings in a room talking and checking each other that might come up with a better answer than just one person alone which is kind of what the agents are doing so I think the agentic workflows and basically what that means is we've given like a different job to each agent so it's basically saying you are an agent for this I want you to do this so it's like the Durham Guardian as an agent you're going to go through every response and just make sure nobody put anything that puts them in the harms way that they didn't miss like a number one diagnosis another thing we've gotten a lot better at is we we were kind of following the same as everyone else where you'd get that laundry list of a differential diagnosis but it wasn't helpful like you really want it to be like a colleague when you're talking to someone, what do you think this is? What should I do next? Like those are the helpful things, not just like, here's 10 things that could be like, that's useless, right? So that we've gotten a lot better at too is, you know what? This is probably this, but it could also be this, make sure you don't miss this, and kind of guiding you like a colleague. So those are few few things we're tightening up on. What's your what's your most common use case? So now that you guys have a bunch of things, you know, you've got all the different buttons, like help with the prior off, let it describe what are people using DermDGPT for the most? I think the board certified Derms are using it most for the second console to actually your your buddy console, whenever they ask a question like, hey, have this and this and, you know, what do you think it is? Are one use case, but I actually didn't really quite think about was the nurse level, the nurse triage level, which is why we actually ended up adding that RN triage agent because I think in the offices, what we were trying to do is 10x the Derm, like 10x what you can do in a day, but we ended up finding out in the office, it really wants to also 10x the nurse and the MA. So then we had agents for them. There was one site, I think it was a year and a half ago, where we were doing an update and the system went down, and literally some of the nurses from whenever nearby academic centers called us and we're like, where's Derm GPT? We use it for a daily triage, which was really cool. We're like, whoa, we didn't even know this existed, so we literally built an agent specifically for them, but that's probably one that's used more by RNs actually rather than Derms triage and cases, and then the prior off one, prior off denial generator, that's just like a super high yield one. Randomly the biopsy site one is not, it's not like one of the top ones, but it's used frequently enough, like you kind of put in like left cheek near the eye and it gives you a fancier version. And I always tell all my med students and everybody like going on a rotation, like just use this, you know, make your anatomic site a little bit fancier. I mean, I've gotten really dumb over years too, but I'm a medical Derm, I'm not like a, I'm not a most surgeon, so the terminology's went gone downhill, and then last year you're in a half, it's been one of our most searches joke, but he's just like, fair, you're getting better. I'm like, that's all Derm GPT. She was getting a lot of like left cheek and things like that there for a while, so. All right, well, we are going to wrap it up there, and it's time to move to Patton's trivia. I'm fascinated to see what you've got this week, Patton. It's just AI stuff from like history and movies and blah blah blah. All right, before you start, you know what drives me nuts? Well, it might be part of your question, so I'm like, go ahead. I was thinking at the end, I'll say. I can't wait. All right. Number one, what is the name of the test proposed in 1950 to determine whether a machine can exhibit intelligent behavior indistinguishable from that of a human? the Turing test. Yep, and the Turing test. That's what I was going to say. So I used to, so I'm follow like all of this stuff and the feeds and then and for years, every the Turing test, the Turing test, the Turing, the day chat GPT hit, nobody's talked about the Turing test ever again, because. Well, thank goodness. You had a chance to pull that knowledge out and impress people with it. It's good. It's such an idea of moving up the post of the like, because the Turing test specifically was a person can interact with it. And if you ask is that a person or a machine they can't tell, that was the Turing test. And as soon as GPT came out and it was passed, everybody immediately stopped up. Nope. That like just the moving goal posts were crazy. Yes, so I start, I asked a couple of AIs. I'm like, has any AI like passed an unrestricted Turing test with expert judges basically seeking out to find whose human and whose and the the two AI, I asked Grock in perplexity. And they said, no, there's not been an AI that has passed like what they call an unrestricted Turing test with expert judges. But again, I'm asking AI that and maybe that's what I I wants me to think. So I just can't trust anything. All right. Yeah. So named after Alan Turing, by the way, who well, whatever. All right. He cracked the Nigma code. Yeah. Then there was that movie with Benedict Cumberbatch called. What the hell was that called? Can't remember. Beautiful mind. I actually saw that movie, but I don't remember what it was called. Yeah. And neither here nor there. All right. Number two, in the alien franchise, what name is given to the onboard computer that controls all of the ship's functions. Oh. Lots of alien movies to choose from. There's like 20 now. How is like the old man? That's not from alien. Correct. It's mother mother. That's it. Yeah. I was like, there's, yeah. It's like M.U. slash T.H. slash you are like that stands for something. They never explained what that stood for, but everyone just called it mother. That's what my kids call me. So I find it very endearing. Right. Much, much like how 9,000 it has secret directives, which kind of make the crew expendable, which is also like you're raising your family. Pretty much. Yeah. So less in in space, don't trust a eyes. I think on earth here. It's okay. All right. Final question. The Voight camp test was administered to suspects to identify if they were human or replicants in what movie. Clap on the body snatchers. No. It's like from the 80s. Harrison Ford was in it. Oh. Who was in it? Harrison Ford. They did like a sequel relatively recently. It's basically the same name with a number out of the runner. There you go. Has to have blade runner. Yeah. All right. Yeah. Okay. I think Matt walked away with that one. Congratulations. Arcturs. I already miss coming out. There we go. Right. So a lot of time in front of the computer. That's right. In all of it, academic. Farah, thank you for coming out and joining us. Thank you. This was a really fun discussion. And I want to thank all of our listeners for coming on with us. We hope you learned a few things. We hope you laughed once or twice, but mostly we're hoping you're planning to join us next week. And until then, for Derms on Drugs, I'm Matt Zyrus. I'm Tim Patton. And I'm Laura Ferris and we are Derms on Drugs.

Podcast Summary

Key Points:

  1. AI (ChatGPT-4) outperformed dermatologists in quality, empathy, and satisfaction when answering patient questions about hidradenitis suppurativa in a 2023 study.
  2. Despite preferring AI responses, 70% of patients and 80% of physicians still chose human dermatologists for answers, highlighting a trust bias toward humans.
  3. A pilot study using AI-generated physician avatars for patient education showed high patient acceptance, with usability near 90% and low eeriness scores.
  4. Experts discuss AI’s potential to handle routine tasks like scheduling and patient education, freeing dermatologists for procedures and complex cases.
  5. AI’s comparative advantage lies in consistency, patience, and empathy, but challenges like hallucinations and trust remain barriers in medicine.

Summary:

The podcast episode explores the role of artificial intelligence in dermatology, focusing on two key studies. The first, a 2023 single-blind study from Poland, compared ChatGPT-4 to dermatologists in answering patient questions about hidradenitis suppurativa. Patients and physicians rated AI responses higher in quality, empathy, and satisfaction.

However, when asked whom they would prefer, 70% of patients and 80% of physicians chose human doctors, indicating a persistent trust in human interaction despite AI’s superior performance. The second study, from the Mayo Clinic, used AI-generated deepfake avatars of physicians for patient education. Patients showed high acceptance, with usability and trust scores near 90% and low eeriness, suggesting AI avatars could provide 24/7 empathetic communication without fatigue.

The discussion highlights AI’s potential to streamline routine tasks like scheduling and patient counseling, allowing dermatologists to focus on procedures and complex cases. Experts note the technology adoption cycle, comparing AI to early Uber or Tesla self-driving, where initial skepticism fades with familiarity. Challenges include AI hallucinations and the need for high accuracy in medicine.

The consensus is that AI will augment rather than replace dermatologists in the near term, especially in procedural areas, but long-term advancements in robotics and imaging may further shift the field.

FAQs

It is a podcast where dermatologists discuss hot topics in dermatology, combining expert knowledge with humor, and new episodes drop every Friday.

In a 2023 study, patients rated ChatGPT-4's responses higher in quality, empathy, and satisfaction than those from dermatologists, though 70% of patients still preferred answers from a human doctor.

DermGPT is an AI tool for dermatologists, launched in 2023, that uses large language models to assist with medical tasks, and it has gained thousands of users.

This may be due to the early technology adoption cycle, where trust in AI takes time, but also because AI is designed to be consistently patient and empathetic without fatigue.

Patients highly accepted AI-generated video avatars of doctors, with high scores for usability, engagement, trust, and low eeriness, suggesting they could be used for 24/7 patient interactions.

AI could take over routine tasks like mole monitoring and patient education, allowing dermatologists to focus on procedures and complex cases, though procedures may eventually be automated too.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.