Go back

#126 Synthetic Users in UX Research: Hype, Reality, and the Future of Design

20m 47s

#126 Synthetic Users in UX Research: Hype, Reality, and the Future of Design

In this podcast episode, the host discusses synthetic users—AI-generated profiles that simulate real user behaviors using large language models. They explain how synthetic users differ from traditional personas, proto-personas, prototypes, and digital twins, highlighting their interactive nature. The host outlines practical applications, such as desk research, concept testing, and marketing simulations, while emphasizing their advantages: speed, affordability, and accessibility for exploring hypotheses or niche groups. However, significant drawbacks are noted, including overly polished responses that lack emotional depth, inherent biases from training data, and ethical issues like transparency. The host warns against relying solely on synthetic users, as they cannot replace the nuanced insights and empathy gained from real human interviews. A hybrid approach is recommended, using synthetic tools for initial exploration but validating findings with actual users. The episode concludes by noting the growing synthetic data industry and future trends, such as designing for AI agents, while promoting an upcoming AI prototyping training event.

Transcription

3278 Words, 18364 Characters

English
Hello and welcome back to the future of fewX. My name is Patrincia Rinas and I'm your host for this podcast. Last week I was in Copenhagen for the future product taste and to be honest it was such a great conference. If you're looking for an event that actually worth the trip, I can really recommend it for next year. There are so many strong talks and I think if we are honest we are not every conference still ever stage. I also had a very unique workshop experience because I hosted my first ever workshop on a boat. It was surreal but actually super fun. In this session we explored how to make sense of insights when you have tons of different data sources and then how to bring these insights to life by prototyping with vibe coding. By the way, a quick announcement. At the end of October I am hosting a live AI prototyping training using vibe coding. It's happening on October 29th and it will be a three hour session full of energy where we go step by step from creating strong prompts to building prototypes to connecting them with back ends, LMMs like TGBT and learn how to vibe code because this is such an essential skill right now for designers. We have limited seats so that we can give proper feedback and also some early bird ticket that will go live probably next week. I am going to share the page for this AI boost session. This is a new series that I am starting in the show notes so you can check it out but I think tickets will become available probably either on the weekend or beginning of next week. And if you are not sign up for my newsletter I'll let you it now because that's where I am going to share the link when the whole event goes live. The newsletter in general is free and I would say always packed with trends and updates and resources so definitely worth it. But I would say a little x-course about Copenhagen and some cool events and now we are diving into this topic because I want to talk about something that's been stirring up a lot of discussion synthetic users. I would say this is a really controversial topic. I've shared some polls about synthetic users on LinkedIn in the last month and the reaction could not have been more divided because some people are super excited and others are completely against it. And last week at the Future Product Days I listened to a talk about exactly this topic from Julian de la Mathia comparing answers from research interviews from synthetic users to answers from basically real user interviews and I thought it was absolutely fascinating so I want to talk a little bit about the learnings from this talk and also some of my own thoughts and experiences. So let's start with the basics. What actually are synthetic users might have heard the term here and there but it's really started from the beginning. Synthetic users are AI-generator profiles that mimic real people's thoughts, needs and behaviors. They are powered by large language models basically the same kind of tech that powers tech, GPT. They're trained on a massive amount of techs and sometimes also behavioral data. Think of them as interactive personas. So unlike a normal persona which is static PDF that says Emma is 34, she likes yoga and she struggles with online banking. Synthetic users can actually talk back so you can interview them, you can ask questions and they will respond as if they were your user. So there are generally different from personas because personas are based on research and are basically archetypes. So synthetic users look like personas but they are dynamic so you can talk to them. They are also different than pro personas. Proder personas are quick, assumption based sketches we use early on so before we do any research. Synthetic users can generate proder personas automatically from public data but you still need a real validation later. I think it's super important. They are also different from product types because product types are designed artifacts. To test with people, synthetic users are not proder types. They are basically the fake people that you test with. And they are also different from digital twins. These are much more accurate. Data rich AI models of specific individuals usually build from lots and lots and lots of interviews. Synthetic users don't usually go that deep. So synthetic users are not personas, not pro to personas, not proder types and not digital twins. So how can be great synthetic users now? I'm going to talk a little bit about the different steps that most platforms use. First of all, you need a data ingestion. So something where they pull in data, maybe from blocks, maybe from forums, maybe from social media and sometimes even from your company's CRM. You need also a target definition. Basically you tell the system, I want to talk to, I don't know, 10 nurses and Bogota about that daily work. Also need simulation. So the LLM is prompted to act like a nurse and Bogota. And some platforms even have multiple AI agents interviewing each other. So they really get crazy. And the last thing that you need is also interaction. So you can run interview surveys in even your usability test. You get the transcripts, the report, basically instantly. There are already some tools out there like synthetic users, Delphi, ask rally or relevance AI and they make this available. I would say right now synthetic users are showing up in lots and lots of different places. When it comes to desk research and hypothesis generation. So instead of spending days digging through all the papers or the research or the forums, you can ask synthetic users for a summary of what people think about a topic. You can also use it for product discovery and concept testing. So when companies are using them to test early ideas maybe, so brand messages or prototypes before they even bring in real people because we know that real research can be expensive. It takes a lot of time and resources. So this is where I'm seeing this integrated. Also in the area of marketing simulations. So when you want to know how a niche audience might respond to an ad, you can simulate them first before spending money on a campaign. How accurate this is would be in discussion because this is still AI. It was still not real people. It's also used in agent based simulations. So researchers create entire groups of synthetic founders, investors or consumers and let them talk to each other to see what dynamics might emerge. Fascinating. You can really go down a pretty rabbit hole when it comes to AR generated synthetic users. But you can also use it and I think this is probably the most interesting thing for us as designers and product managers or product people in general research automation. So agencies that are starting to run mock ups and interview guides through synthetic users just to speed up the prep, speed up the research. Yeah, this sounds interesting, but let's talk a little bit about the problem that comes with synthetic users because using all nice, I'm saving money. I am basically skipping the whole research process. I can do that so cheap and so fast. But what's the problem? I would say we make it concrete. So imagine you are building a fitness app maybe for young mothers and you want to know how they think about the exercise. You want to understand what's that problem? Why are they not doing exercises themselves? The internet is full of workouts. You can go to the gym so you really want to understand what's their blockers to come up with concept of different features that help them overcome those blockers. So you want to go really deep. So if you ask a synthetic user the question or also a normal user, tell me about the last time you worked out because you want to understand how this usually looks like. You want a little bit of story telling us this while you ask tell me the last time you worked out. You want to understand how that looked like, how that worked maybe some blockers. So you ask and what do we get from the synthetic user? You hear? Oh. I did a 3pm, I did a 20 minute YouTube workout. Two weeks ago I also went on a 10 minute walk with my baby in the store. Check. Everything answered perfect. Very neat, very structured, very clear. And if you have done user research before interviewed people, you know that most of the time it doesn't sound like that. If you ask real user, you will get sometimes or something like, "Plea, the last workout." Like me think. So I actually plan to do something last Thursday. I even had my workout clothes on, but then the baby cried. So I walked around with her instead and that part man and then my older kid came back from kindergarten and then my husband needed help. So I ended up actually cooking dinner and yeah, that was it. Not sure if cooking dinner counts as a workout. Anyway, right? Like you see the difference? Centenic users give you a clean, precise answer. Large language modes, AI in general, they're always answer. And the real user gives you chaos, interruptions, emotions, and that's where the real inside lies. Because the true problem isn't about knowing what work out to do, you don't want to know, you don't really care, in this sense. You want to understand how it looks like. You want to understand the time and the constant interruption. So that's super interesting. Think of maybe they need like many workouts. Maybe they need a workout that they can do with their kid, when the kid is crying. Usually moms are walking in their apartment. Maybe you can combine something. So you understand this is where you really go deep and find the interesting insights. And also I think the stuff you only uncover by talking to actually humans. You're probably a designer researcher, so you totally understand what I mean. But I think also for us, it's very important to know what it is, how we could use synthetic users and then articulate that when this comes up in a meeting. Because I am doing lots of workshops and also a consultancy work for companies and synthetic users is a big topic about someone saying, let's save a little bit of budget here. And should we start with synthetic users first? And then we need to explain, yeah, we could totally do that. Those might be the pitfalls, those might be some options that we have. So you really need to consult. So let's talk about the benefit of synthetic users because they are not all bad. And I think this is also super important when you discuss with someone, especially like product managers or product owners, to see like where there is still some opportunity. I think there are some strong advantages because they are super fast. You can get 100 interviews basically in an afternoon. They are very cheap. You don't need recruiting. You don't need all the planning and the structuring. You can even similar edge cases like maybe rural populations, niche profession or rare demographics. It's usually pretty difficult to get these people on board or to find them for an interview. I think there are great for coming up with hypothesis. So you can test interview guide, the concept directions, and before involving real people. So you can basically use these insights to come up with hypothesis that you still need to validate. So this is not validated if synthetic users say something, but this basically is an idea that you can test data on with users. And I think they're pretty accessible. So even if you don't have the time or the researchers for a full research round, that might be interesting. Because there are lots of different and super agile ways where you don't do research first. I would say like in most projects, you start with like secondary research. I think that you find online with hypothesis and then later on you do some testing. This is what most of the times make sense. It depends on the project. The more needs, the more difficult the project to tackle is, the more research you need early on. And I think the more insight is already out there, the easier it is to get started, start with hypothesis, and then validate that little later point. OK, so we talked about the benefits. Let's talk about the risk in the limitations, because there are some drawbacks, and I think they are serious and super important for us to talk about it. Because synthetic users tend to give overly positive answers. And I always want to please. It won't really criticize your idea. So when you say, what would you think about how or how would you use an app where you see, I don't know, different meal recommendations for a day that will fit to your workout schedule? And AI would probably answer something like, "Yeah, great idea. I would definitely use that. I would love that." And the human would be actually complicated. I don't have time to shop. I don't have time to prepare something. I usually eat, I don't know, like a sandwich or something. So that's, I think, interesting, right? AI is overly positive and you need to know that. Also, surprise, they lack the emotional depth. So they don't capture stress, frustration, and the hesitation. Other things that are actually interesting for qualitative user interviews. What the mother told all about, the stress of the baby was crying, she couldn't do it, and the other kid came. And then she was cooking, and you could totally could feel her stress and the frustration, and how difficult it actually is to get a task done. The next limitation is also the training data. Because the model hasn't seen certain behaviors. It's just makes something up. So it's all about hypothesis. And it can miss these outliers. Sometimes the outlier is the gold mine. Like Netflix realizing that the late fees were a bigger problem than the, like, the movie selection. And let's be real, they can't really buy your product in the end. They don't feel pain. They don't cry. So when an app crashes in the middle of their hospital to the visit, they don't cry. They're like, OK, no problem. Because they don't need the app, actually. So if you rely only on synthetic users, you risk designing for an idealized world that doesn't exist. I don't say that you shouldn't use them. I think it could be interesting, especially if you don't have a lot of budget, start with it, use it as a hypothesis, but don't use this to validate your ideas, your concept. And at some point, we need users. And it will always be the case. Because how good AI gets, they will never have this emotional depth and really go into, like, the nitty-gritty parts. And we haven't talked about the ethical considerations yet. And this is also super important. So think about transparency. So if you use synthetic users, you need to disclose it somehow. You can't present fake interviews as real research. So this is something when you bought it openly, either in a meeting or with your stakeholders. This is something that needs to be addressed. And you need to talk about it. No one can pretend like we did research. No, this is AI-generated research. You need to be transparent about it. There is also the big problem of biases. The synthetic users reflect the biases of the data they have been trained on. And if their data is Eurocentric, heteromormative, or just incomplete, your insights will be true. You need to be super cautious there. Also, these profiles that are built on human data, whether humans aware that their data would be really used this way, I think, are probably not. So this is also something that's important. Large language models aren't free. So running simulations at scale also has a carbon footprint. And the biggest danger for me is when teams use synthetic users as a shortcut to avoid talking to actual users that erodes empathy and empathy is the foundation of UX. You can do that, but in the end, you will pay the price. Or the product will pay the price. OK, let's look a little bit into the future. Because I'm always curious to understand a bit of where this is heading. Generally, the market is exploding. The synthetic data industry is projected to grow from 267 million to 4.6 billion by 2033. So the simulation fidelity will also improve. Models will get better at limiting preferences and giving more realistic feedback. And especially these hybrid methods will become the norm. So synthetic users may be for early exploration and then real users for validation if we wanted or not. And I also see that as a way to do it, if you do it right, we also might see something like validation as a service. So external providers who certify the quality of synthetic insights. And another interesting shift could be, we will start designing only for mixed but also for agents. So your app might be used by an AI assistant acting on behalf of a human. So that's a new design challenge that we will probably dive a little bit deeper into one of the next podcast episodes. I think it's super fascinating to see how agents are changing the design landscape and how that also changes the way how we design fascinating topic. OK. So let's summarize that. Where do we land? I would say synthetic users, another replacement of real research. Surprise. I think we all knew that. We are all designers here. So there's not a surprise. But I think they can be super useful for quick exploration for generating hypothesis and for testing ideas at scale. They cannot replace the depth, the nuance, and empathy you get from talking to actual people. And best approach is hybrid. Use synthetic users to move fast early on, but always validate with real humans before making big decisions and be transparent about it. Because at the end of the day, you access about empathy and empathy can be simulated. Yeah. That's it for today's episode. I'm super curious to hear your take on synthetic users. I actually shared a post on LinkedIn. That I'm going to link below. Feel free to share your thoughts. I'm super curious to hear what do you think about synthetic users? Maybe have you. used it. What are your struggles when they're constant, some pitfalls as well? Would you use them in your process or do you think they are too risky? Let me know. I'm super curious to hear your thoughts. And don't forget October 29th is the AI Prototyping Training with 5 coding, early bad tickets dropped. Next week in my newsletter, super excited to work on AI Prototyping with you. Thank you so much for listening and I will catch you in the next episode of the future of.

Podcast Summary

Key Points:

  1. Synthetic users are AI-generated profiles that mimic real people's thoughts and behaviors, powered by large language models, and can be interacted with like personas.
  2. They offer benefits such as speed, cost-effectiveness, and accessibility for generating hypotheses, testing concepts, and simulating niche demographics before real user research.
  3. Significant limitations include overly positive or clean responses lacking emotional depth, biases from training data, ethical concerns like transparency, and the risk of eroding empathy by replacing real human interaction.
  4. The future points toward hybrid approaches, using synthetic users for early exploration and real users for validation, amid growing industry adoption and evolving applications like designing for AI agents.

Summary:

In this podcast episode, the host discusses synthetic users—AI-generated profiles that simulate real user behaviors using large language models. They explain how synthetic users differ from traditional personas, proto-personas, prototypes, and digital twins, highlighting their interactive nature. The host outlines practical applications, such as desk research, concept testing, and marketing simulations, while emphasizing their advantages: speed, affordability, and accessibility for exploring hypotheses or niche groups.

However, significant drawbacks are noted, including overly polished responses that lack emotional depth, inherent biases from training data, and ethical issues like transparency. The host warns against relying solely on synthetic users, as they cannot replace the nuanced insights and empathy gained from real human interviews. A hybrid approach is recommended, using synthetic tools for initial exploration but validating findings with actual users.

The episode concludes by noting the growing synthetic data industry and future trends, such as designing for AI agents, while promoting an upcoming AI prototyping training event.

FAQs

Synthetic users are AI-generated profiles that mimic real people's thoughts, needs, and behaviors, powered by large language models. They act as interactive personas that you can interview or ask questions, unlike static personas.

They can be used for desk research, hypothesis generation, concept testing, marketing simulations, and research automation to speed up processes. They help explore ideas quickly before involving real users.

Synthetic users are fast, cheap, and accessible, allowing for rapid interviews and hypothesis generation. They can simulate niche demographics or edge cases that are hard to recruit in real research.

They tend to give overly positive, clean answers lacking emotional depth and may miss outliers or biases in training data. Relying solely on them risks designing for an idealized world without real user empathy.

Unlike static personas based on research, synthetic users are dynamic and interactive. They are not prototypes, which are design artifacts for testing, but rather fake people used for simulation.

Transparency is crucial; fake interviews should not be presented as real research. Biases in training data can skew insights, and using synthetic users to avoid real user interaction can erode empathy in UX design.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.