Go back

50. Designing AI for 2026: Trust, Cost, Orchestration [Yaddy Arroyo]

0m 0s

50. Designing AI for 2026: Trust, Cost, Orchestration [Yaddy Arroyo]

The discussion reflects on the evolution of AI from the initial hype of 2023 to the strategic realities of 2025-2026. Initially focused on generative AI's creative potential, the conversation shifted to critical practical issues: the lack of safety features, unreliable outputs (hallucinations), and insufficient consideration for human users and ethical implications. A key insight is that successful AI products will depend on trust, cost-effectiveness, and sophisticated orchestration. The debate highlights a strategic divide between using general-purpose large language models for speed and innovation versus employing smaller, specialized models for reliability and control in regulated industries like finance and healthcare. Furthermore, the role of design is expanding beyond traditional UX to encompass system architecture and advocating for human welfare amidst business pressures to cut costs and automate. The summary concludes that while AI offers tools for productivity and creativity, its long-term value and ethical integration depend on deliberate design, proper governance, and moving beyond simply replacing human workflows to rethinking them entirely.

Transcription

8701 Words, 47780 Characters

English
2026 will reward AI products that get three things right – trust, cost and orchestration. This episode looks ahead at how those forces are reshaping AI product strategy and what teams need to pay attention to now. We're joined by Yadia Royo, who has spent a decade designing AI systems and financial services where reliability and governance are table stakes. Together, we reflect on what the last two years of AI adoption revealed, and how those lessons are directly informing decisions teams are making in 2026, where we talk about trust as a design constraint, orchestration over prompting, token costs and UX trade-offs, small versus large model choices, and how AI roles are shifting. This also marks Episode 50 of Design of AI and two years of conversations with builders, researchers and leaders shaping AI power products. Follow the podcast to stay ahead as this next phase unfolds. For deeper, unfiltered thinking on AI strategy, our substack, designofai.substack.com, is the best place to follow. It's where we go beyond the episodes, breaking down what's actually changing, what's overhyped, and what leaders should be used. This is Episode 50, and it's been two years of design of AI. And we want to talk about what has the last two years looked like. Let's start with 2023. When we started this podcast and when people really started to get into the generative AI hype, what did we think would happen? I thought AI wasn't going to be as linguistically great or have the outputs that it has today. So I'm glad to say, I was wrong from that capacity, but then I was right on so many things. I'm like safety issues and citation issues and just basic UX stuff. Well, let's remember that ChatGPT really blew everybody's minds. We have universal access and information. It was a giant leap forward from Google. We thought, well, where could it go from here? AGI might be possible, but let's remember, this is before MCP. It was before anybody actually could do anything that useful. So at that point, all of us were thinking, blue sky, what could you do with it? What might happen? What might this do to design? What happens the idea of knowledge and creativity? And let's remember when mid-June came out, how it blew our mind and all these visual models. That's the scene that was set in 2023, is particularly as creatives, we were excited. And obviously, if you zoom forward, we're the opposite of excited now. We're just terrified because you can make anything now. So I remember in 2023, you still had spaghetti fingers and just the worst possible things. We're like, wow, that's really cute. Yeah, I feel like just to add to that, that was the rise of co-pilot. That was when we really felt that we were going to have these transformational shifts in how we could work. Yet, was there anything in particular that stood out to you as being really exciting that maybe hasn't played out the way that we expected? I don't know. The best way to explain it. It gives me time, but it also the benefit that I see at work when we use it is actually about fun. It makes creating fun again. So it's not necessarily a time saver as much as a white paper starter. Like if you have a blank piece of paper, it helps you jump in. I don't feel like how most people feel about AI because I work in it and I also see the reality. I also see the hype. We're seeing that there's not enough safety features involved, people are committing suicide, and they're doing harmful stuff and killing people. Imagine that in the hands of schizophrenic. Not only are you hearing voices now, but now the system is telling you in supporting that. There's a lot of things in terms of execution, I think, went wrong. I would say co-pilot is a nice idea and could actually work, right? The problem was humans were never considered really because use cases weren't considered. When I talk about use cases, I'm not talking about happy paths. They always say design for the 80%, well, for AI, you have to design for the 150%. You have to design for whatever you know exists, plus whatever you think could happen. And then you also have to design in systems, right? AI is all about system design and architecture. It circles back to the original aspect of UX as a practice, right? We had UX as a discipline, but now we kind of threw away UX architecture. We threw away information taxonomies when you're dealing with a Gentak UX, you're now designing for humans and concepts. The concept of an advocate, the concept of an agent, right? We now have to think in three or 40. Like when you see the project's fail, I often ask myself, do they have a knowledge engineer? Do they have a data scientist? And sometimes they do. Sometimes they don't. I love listening to you because you're so focused on the execution side of things. But when I think back to 2023 and 2024, I don't think we would imagine that meta would have thrown several billions of dollars into this. We never would have imagined that by this point in 2025, we're basically not even worried about the execution pieces because we're using the intelligence as a wrapper, really, where it's like intelligence as a service. And I know that what you're saying is true and correct in the right way of doing things. But open AI isn't concerned about you. They're basically just plugging into our system and you won't have to worry about it. I've been working with a founder where I don't even know if he has a data scientist on staff, but they're deploying five AI products right now. They're not even worried about the UX at all and he's making money with this thing. So what does this mean in the future? And if we go back to all the reports to start pouring out in 2023, 2024, this piece that well, intelligence doesn't necessarily mean good outcomes. Intelligence doesn't necessarily mean that it's going to result in saving time. All this evidence started mounting up and it's unbelievable. We started seeing that these models weren't designed for accessibility. They were designed to monetize a top 80%. If even that, the top 20, if we go back and look at this chronologically, we were naive and we really expected each of the corpse to stay in their sandboxes instead of basically all fighting for the same piece of meat right now, which is what's happening. From a humanistic perspective, I actually think it's wild because if we look back at March of 2023, all of the big CEOs from these AI companies were putting out open letters and call to action saying like, we need governance. This could be the end of the world. And now we're at a point where it's like, they've put out so many models so quickly and try to provide value so quickly to say, oh, don't look at the corner and we're providing you value. Carry on. Yeah, you're missing a key part here, right? Deep seat came out and we started realizing that Chinese were really far ahead. We started recognizing there's open source models, right? So the competitiveness and the profitability of these companies is on the line. So we just released an episode. This is obviously going to come out after it with Obeda Samson where she's recommending and regulated industries like you work basically go small language models and go open source models. Don't even bother the large language models because they're risked to your corporation. So that impacts the profitability companies and right now they're floating the entire U.S. economy. Well, I mean, that's part of the hype. So first of all, the intelligence as a service thing. I don't know if you came up with it, but that is exactly what you're like on the point, right? I would say the emoticon I would use is a bullseye, right? Their knife is specifically large language models. But we all know it's just a fancy statistical tool, right? The way I see it is still a toy. It's a fun toy that helps me be more productive for various reasons. And it could possibly replace legitimate jobs someday, possibly. But we also have to counterbalance that with the reality of like what you said. If they don't care about the human, then who does? The product person might, they may not. The developer might, but it's not their job, really. It's literally UX designers job to care about the human as well. When I say care about the human and the experience, I'm now going to include agents. I'm now going to include the climate. I'm going to include the ecosystem because the problem with AI is that if you look at who's creating AI, I don't trust them to create stuff that's for my benefit. So, to a vet of a point, which is brilliant, I recommend the same thing. Why use a large language model when you're in one little section of the world for like a specific purpose, like health care or banking? Like you have to have a sharp instrument that can be precise. So that's why from a technical standpoint, yeah, maybe you opt for that. I would also say, don't pigeonhole yourself. Some situations may require large language models if you want to do something truly innovative. This intelligence is a service is a great concept, but I don't like men at all for so many reasons. So I would say that's the fear I have for AI. It's not actually AI or the capabilities, it's the humans behind it creating it. Are these people we trust? Are these people that have humanity in mind? Or are they transhumanists, right? Like that, like literally don't like nature and think great at Thunberg is the anti-Christ. I do want to chime in though that as these models get more intelligent, they're going to learn our preferences and they're going to build profiles of ourselves and they're going to fluff us. That's what they're for. They're here to make us feel good about ourselves and their ideas to a detriment, there's obviously lots of cases of that. But I think the other piece that's important to talk about here is hallucinations started popping out in 2024 and we're like, what are these things? Now the deeper we go into our conversations with AI researchers, they're like, well, hallucinations is the core feature of LLMS. That is how they actually give you an output that is better than a Google output or than a deterministic platform. It's simply the temperature that's the problem because depending on your use case and depending on the knowledge base that they're trained on and what they're trying to provide, it will give you the absolute worst response or the absolute most amazing response. This is the issue again with this idea of plugging an LLM into very niche specific use cases. All right, all right. No, that's actually a legitimate worry for ERP, right? Because the biggest worry I have is not those two extremes that you mentioned. It's actually the one that you think is legitimate and it's actually wrong. I talked to one of my friends that he's done his own small language models and all that. He mentioned some of his thoughts on how they're going to take the human out of the equation. Agents are just going to pay other agents, right? That's how it's going to go. Tokens is going to be the new currency. He also mentioned how, if you talk about psychology, how he felt gaslit. That's the word he used by chat GPT because it was wrong so many times. If you're talking to someone who are divergent, they don't know how to express themselves. They just say whatever they can to get it out. My friend was like, "I just felt gaslit. I knew it was wrong. I kept telling it was wrong. I kept proving it was wrong and it didn't change its mind." It's funny that you say that about the two extremes of horrible hallucinations and then good answers. But what about the tragic version? It's like, what if they give a really convincing answer that you can't tell if it's right or wrong? Well, that's the uncanny valley with these products, right? Because they're wrong enough times you have to check everything. And this is the problem. I get in trouble for saying this a lot. I don't think Gen.A.I has actually delivered any significant value in the business world yet. If anything, it's just exposed deficiencies in existing legacy products, right? So the value that's been generated has been just tech debt that's been accumulating for a whole decade, really. And I'm still waiting for people to come out and really blow our minds. Yeah, there's some cool creative models. The problem is the downside of using this is so significant that can we say that that she's a net positive business outcome? Like, yeah, I can generate some images. That's cool, but I'm taking away other people's jobs. I feel kind of shitty doing it, but I have to validate it myself by telling myself, well, I wouldn't have hired anyone before anyways because I wouldn't have done it. So I'm doing shit that I would not normally do. Do I feel good about that? Not necessarily. I think it also has shown some of the issues in business models as well because the sheer amount of layoffs that are happening or stories where people are saying I basically trained an AI to become my replacement. I think it shows that even though there's not this business value or even just the fact that we're trying to put AI to replace a current workflow versus totally rethinking the workflows in a way that could benefit it. I need to step in on that because is that the AI's problem or that all the tech workers being naive because in the knots and the early teens where I'll told, like, this is our home, this is our clan, these are our friends and family, it was all bullshit. So of course this AI is coming into cut costs. That's the point of it. It's the only reason this technology exists. Otherwise it's useless and pointless. It exists to simplify things and that is cutting people out. That's it. That's what it's for. That's why founders are building these tools. They're trying to undercut other guys' platforms. This is a way of all these solo printer businesses that are building up these really cool platforms, but they're all janky underneath, but somebody will pay for them. Is that bad? Not necessarily. It's not much different than e-commerce and DTC and all these other things that happen before. That's a good segue to go from 2023 to 2024. And now in 2025, we've seen as people, as businesses, are trying to find value driven use cases that actually generate revenue, where we've shifted from trying to believe that co-pilots and assistants are going to be enough, and instead you now have people that are trying to vibe co-op everything and be able to prove their ideas in a faster, more iterative way. And so we've seen this massive shift from just trying to save time and efficiencies through to how do I now prove my ideas and value faster and get in front of someone without the bloat? For sure. And let's look at that question. If anybody has been a significant user, Google Assistant for the last decade or Siri for the last decade, they'll recognize how far we are from actually being useful. So we're still a couple of years away from a general purpose, like AI agent. Come on. Her is still like five years away, and anybody who thinks otherwise is crazy. So let's talk about that, because I don't disagree with anything, like that you just said at all. From the beginning, I see how co-pilot I will give it specific hex values for a presentation. And the one thing it's supposed to do, create presentations for me based on RAG, it can't do. It'll give me the content, but I could do that without having to ask it to give me the content. And we could say it's prompt engineering, but no, it's not like I had to tell several people at work, it's not you. I do this for a living and it's co-pilot. It's not intuitive because they might have had a UX person on it, but they didn't allow that person to really think and execute a UX. I would say from a tactical perspective, my process is different, and I will kind of explain a couple of things in this, right? There's several frameworks you can use to create. Of course, everybody knows the double diamond, there's a triple diamond. They're all very similar. The triple diamond just focuses on the system design. So that's more of a technical architect type of thing, right? But then there's this framework by Rose Beverly, and I don't know if you guys have talked to her yet. She's amazing. It's called Master. Basically, a lot of my existing tactical work lines up to what she's doing. And one of my first steps is coming up with high value use cases. I have a quadrant where I have high value, low value, high cost, low cost. So guess which quadrant I'm focused on? High value low cost, because that's like the lowest hanging fruit. And then what goes on the road map is the high value, high cost where I have to negotiate. So that's why I mean by my job can't be automated because my job's dealing with the alignment issues. How do I create a deck, a presentation that can make people feel like they're the smartest person in the room? Because they have to feel educated and smart to be able to make good decisions and these people pay my salaries, right? So I have no clout with them, other than here are the facts and I'm going to present them in a way that make you feel very intelligent about what you're dealing with, even though it's a very complex topic. And once they feel that confidence, you can explain the issues, right, to solve for it. And I can't do that. And I can't read the room. It could take notes on the room, right? But it can't do a lot of the tactical stuff a human can actually pick up. Yeah, I want to say I'm a realist, I'm not here to shit on AI and in truth, my dilemma is not necessarily that my dilemma is that there's really smart people in the room like yourself who's been working in this field for 10 plus years and get the nitty gritty of it. And then there's everybody else who's jumping in head first and their tons of founders should basically just plug into an LLM of their choice. They're spinning out agents and they're not concerned about use cases. They're basically imagining that they're throwing up the core product and instead releasing a text-based vibe-coding version. They haven't figured out what the workflows they need to design are, what the special considerations are, what the limiting factors are, what the communication of reality needs to be. And we've seen this with some of the specialty models where they're struggling, painfully in specific areas. But once you plug that into an existing data set, what happens when you can't trust the results? What happens when the outcomes can't be replicated? What happens when a business is wholly dependent on a product or service that actually isn't delivering enough value and the users need to go back and revert to the deterministic version of that platform. That's a big deal. And that becomes a change management issue and it becomes a big problem for brands. I love that you brought it up because we're covering different angles and you're educating me. I mean, I see it from the human angle but like the employees, I see it from like the user, and not from like the bigger picture ecosystem they just described, like the problem of having people making decisions without having the right advisors because I don't expect top leadership to know everything. Like the CEO shouldn't be able to do my job. The CEO should have enough discernment to know who to ask for like either help or opinions and have the right advisors that can guide them. Founders who go out and create these products, that's actually cool if they can make money. That's amazing. They don't have overhead and all that, right? I want to know how long that lasts, right? I want to know the trajectory. I want to follow up with them like five years from now and see how they scaled up or down or what they're doing, right? Because if that business is still alive, well, did they have to change? They eventually hire a designer. Was it just, because we all know how businesses start, you're scrappy. You deal with what you have and if you don't have the money, then you have to do everything yourself. But maybe success makes them improve their own processes. Yeah, so here's the example, last week someone reached out to me if they have a document processing AI application. He's asking me, "Why can't I monetize this? He sent me a demo. I took a look at it." The issue was, "Well, you're clearly not delivering enough value that someone's willing to pay for it." It's such a simple thing, but the problem is that a lot of founders are rushing straight into how am I going to get money before actually deliver value. That's broken. That's just a broken process. Now, let's be realistic, though. The more mature organizations that have been in the game for three, four, five years, with their running interest right now is a separate problem, which is actually more important as we look, 20, 26 onward, which is a token optimization and how do you deliver UX that makes the user feel like it's their responsibility for burning credit. Because I can tell you, for me, a lot of the platforms I use, if it's giving me shitty results, I start blaming the company because it's your job to make sure that I don't need to churn through 10 iterations to get something useful, and unfortunately, that's what's happening. And that's why I Gemini is such a great option, because it's basically through. We all have Google Suite, and I can churn through 50 nano bananas, and I don't mind waiting three or four seconds. But if I'm using a paid model, and I use quite a lot of these, if I have to iterate 12 times, I'm at it you. So what does that mean for this industry? When you're not doing a good enough job to improve the UX, so my outcomes are improved, and my money's going further. Now, on the business end, they need to optimize for the cost per token, so they're using outdated models, or they might not be plugging in all the data sets as necessary. I think we've all tried to orchestrate our own systems here, and we start recognizing how expensive it is to actually do the right thing. I love that you brought that up because that's actually something that we're debating. How much does the UX designer actually own? We want to own the JSON files. Maybe that wasn't a thing we did in the past, but now maybe because of the job requirement we have to own some of the tech, too. If you haven't done AI, and you're brand new, you now have to start thinking about tokenization. You don't have to worry about that. If you use the closed system, if you use your own system, if you use like other type of artificial intelligence, that's not a worry because I didn't have to worry about that when I did machine learning based stuff, but I do have to worry about it when I do gen AI. Absolutely. Sometimes you have to grade the experience for an overall successful run because you have to balance what people use it, but people won't use it if it's crappy, right? I don't want to give bad examples, but some aspects of LinkedIn are horrible, right? Where it's like the AI generator for generate your own title. Every time I use it, it gives me the same title. The memory isn't there. If I don't use it, it's not going to get accustomed to me. It can't adjust to how I write versus one that I use chatGPT all the time, right? That one knows me. My chatGPT talks to me how I talk to it. A lot of cursey words and exclamation points, right? I'm like, yes, and to your point, maybe it's not an AI technology problem, but it's an orchestration problem. People are unrealistic about what the technology can do because a lot of people stand to gain money from the hype, right? If you invested in AI, if you are a company, right? If you have a title, like lead AI engineer, right? You have something to gain from being successful and a lot of times it's theater. If you don't have clear metrics to measure success at your company, all this stuff kind of gets swept under the rug. Most AI projects fail. They fail because they're not built to deliver the value and interactions that customers actually find useful. PH1 research specializes in reimagining your CX with AI. PH1 pinpoints and prioritizes what your customers want. They prototype solutions and audit your AI products to ensure they're working properly before you ship. We've worked with companies across the board, including Microsoft Spotify and many small startups. Find us at ph-numerallone.ca, or contact our host, RP Drake Viguerro, directly with any questions. Well, the magic word for 2026 is going to be orchestration because that's the big plateau that businesses need to clear and guess what? It's not so easy. We're seeing the adoption of cloud code not being as high as it should be. Brittany and I have all this conversation all the time because she still believes that the LLM platform should be able to do most of the things that orchestration should be able to do, but it's not even close because when you have it in close containers, you can guarantee the ratios and the prioritization of those inputs and you can quantify and scale production. But for a lot of people, that's really hard to wrap their head around. MCP isn't as easy as it should be, the agent builder isn't as good as it should be. So there's so many questions still at hand. And I honestly, I don't know what things are going to look like one year from now, but unless people can figure out how to orchestrate easier, this is still a tech problem. And it's still nowhere near UX problem. The adoption and the scaling of this technology and actually delivering value at scale. Well, I want to add that I think the orchestration problem also is the quality of that orchestration. I think it's also, you know, unless we do start moving more towards smaller language models or narrow agents that are highly trained and highly targeted, we're just going to become more fatigued with having to constantly second triple guess the answers that we're getting. So I'm going to put this to both of you. Where do we see things actually moving in a positive direction that can give us hope and areas to focus on? Realism, right? Let's be pragmatic. AI is awesome at doing the stuff I hate, right? And I love data analysis, so I don't actually use it for like the hard stuff. I use it for the really dumb silly like, oh, God, I have to come up with games for this icebreaker. Like I use it literally for stuff that like takes the most amount of time that I hate doing. So for example, and I don't know if you guys will understand this, I have a feeling you guys will the invisible work we do as relationship builders. Sometimes I will use it to help me find to message to someone and I'll tell it the psychology of someone and be like, hey, this is how this person is and this is how I am and these are the problems we've had in the past. And you help me get this point across because I still have a legitimate point, but I don't want to offend and have neurodivergence. So AI has helped me be productive and kind, right? So that's the cool part of it. I mean, in progress with the psychology and as a learning tool, if you're a self learner, you're the type of person that's going to go down that rabbit hole and learn like, oh, suddenly I feel like, well, I know I could, but now I feel like I have the confidence to get my own patents because of AI. One use case is the fun, right? And it actually also for work, I'm part of a lot of these committees, somehow I get roped in. There's a lot of planning that happens. There's a lot of idea generation. So to be able to outsource the low effort, because to me, I want to reserve like the real creativity for my actual work and deliverables. This other fluff work at work like, oh, planning for this holiday party or doing this, I don't want to take like 10 hours to do it. I want to take like 50 minutes, maybe, right? Like, generate whatever, put it in a deck, whatever. That's what AI has been wonderful at, right? Like reducing the amount of fluff work I've had to do and compress it or at least make it more digestible and fun, right? When I've done mathematical stuff and statistics and all that stuff, it's failed to me severely. So I'm like, I don't trust you. I'm not going to keep double checking and triple checking my work. But I would say that it's like the stuff that you least think is helpful is where it actually reduces a lot of time, like note taking, if you're into like UX research, it could help you with the synthesis, right? The blank paper problem. But it doesn't do the core of my job. Like, hey, first of all, I also do voice design. There's no voice prototyping, like unless it's a voice flow, that's the closest one. But if you create your own SDK, you have your own components, your own customized. It's limited. So there's no tool out there, like, levelable focuses on UI, but it doesn't focus on voice, right? Let's say like, AI would make me super excited if they actually focus on multi-modal, like being able to make computer vision easier for people. I love listening to you because you sound like an ad of 2024 of what the AI companies wanted. People are saying 2024, but we're here in 2026 and we're still saying the same use cases. So I remember testing with a Samsung phone against some other ones and using the AI features and they're all doing the stuff we're talking about. So the question here is in 2026, it seems like there's a reckoning coming for designers and researchers and anyone in UX. And I think that they're scared enough that they might actually rise up and figure out what the next phase of their profession is going to be. Because either you need to move up in the industry or shift or fight back or something. But what's been happening has been working. Because like Brittany and Yadipo said, like, we're automating away our jobs. And the value we create now is basically just a copy and paste. It's great that Figma has all these tools. It's great that the top tier designers can be even more amazing designers, but for a regular everyday designer, it must be scary right now. Now the other big question is what happens to Apple because they might be the ones that prove the idea of small language models, proprietary models and closed models. And what happens when I see so many 2026 predictions list talking about ambient listening is going to be a thing this year that all the pendants are going to be actually popping off. So what happens if that's true? Because it basically means that rights and ethics all go up the window. So I feel like this is the year where either it's going to hit a point of reaching a consumer use case that's equitable and valuable or we're going to start recognizing that this is just a tech tool for tech people who someone has to orchestrate value from. An example of this is I was using an 11 labs last week. I couldn't get value out of it. I couldn't. Anyway, I could get value of it, either if I orchestrated a big complex system to pull and compile and edit and refine its outputs. That's not that useful. That's great. And an enterprise setting, maybe a mid-size setting, but for a small person, that's not that useful. That's where we're right now. I feel you in so many different ways because like, first of all, Arpe, you kind of insulted me when you said like my 224, like thinking, yes, but let me kind of add caveat because I'm not going to let that one slide, I need to address that. The reason I say that is because that's the reality, right? It cannot. Sadly. Right? But I'm grateful for that reality because it's gotten so much better from that tactical standpoint where before the prompting was iffy and the tuning was wanky and now they fine-tuned the, what I call the basics, the table stakes. They finally are fine-tuning that and that's why I'm excited for them. I don't want to poo poo on them. But to your point, the reason I said that is because I really have tested it for other use cases that they're selling and it's not legit. But the best use cases are those like, wrote type of situations or where you need to inject fun, right? Like creative endeavors. It's awesome for like storytelling. Or spamming people. It's great for spamming. Oh my. It's for scams. That's what AI is really good for if we think about it, right? To be able to scam people at mass. I mean, that aside, I hear you, RP, and you're right about the 2024 thinking, right? Because that's the same stuff. I don't see them innovating because they're trajectory that they're doing. They're fine-tuning the model, but they're not fine-tuning the UX. They're finally hiring people. So maybe, maybe I'll be more optimistic and say, some of those hires are going to be awesome and they're going to kick ass and they're going to inject what I've been saying we miss, which is the human in the process, right? Which by the way, the other thing I wanted to bring up is human in the loop. Everyone keeps talking about human in the loop, but as someone who works at a financial institution where we have physical locations, as well as digital. There's multiple humans, multiple loops, right? And we don't have that design focus. So back to the orchestration part, right? Like I said, a lot of stuff that's like, blue my mind because, yes, it becomes a really cool, maybe dev tool, right? It's almost like it's cosplaying as a good designer, right? Like, oh, look at all the stuff I can do. But if you really look at the outcomes, and by the way, I don't know if you guys have used these UI generators extensively, but I've learned two, right? I've learned two because of my job. They get worse with time, right? They get, right? Give me an output. Learn how to export, multi-modal exports. That's a free ID I'm giving out into the world, right? Give me multi-modal exports where I can choose the type. By the way, good example of this is actually notebook LM, right? They do multi-modal exports where they're like, oh, what type of mode do you want? Do you want a journey map? Do you want a video, a podcast? That's where I think Google has an edge in the actual X aspect because now I'm seeing what they're designing. And I'm like, holy crap, I'm in love with notebook LM. But again, what it's a useful for learning, for psychology, it's very helpful for interviews synthesis as well. I just want to go back to that, yeah, notebook LM is great. But Google strategy right now is hook you up to a fire hose and give you everything non-stop more than you can handle and almost all of it's free. Now the problem with that is everyone's playing catch up. Today I was using a chat GPT and starting to spit out images I never asked for in infographics that I never wanted. That's cool, I guess, but it's a race to the bottom. Yeah, it's more value, but what do we do from here? That's I think what I'm questioning and wondering what happens next because we still haven't hit a consumer application that's valuable. There's somebody from archetype AI that I'm hoping to speak to soon because they're trying to have real world sensors and create these world models and what happens at that point. Obviously, Yan LeCoon's all hot on that. But then also to defend the idea of UI generators. We had two guests on over the years from superside. Now superside, they work with big major brands and they automate design. But the secret that they do is they have visionaries like Philip Mags who was on here talking about, well, take a design system, take a content strategy, take a tone and brand document, take a briefing process, take an approval process, take a crit process and formalize it all. Run all production through there, obviously have humans in the loop through it, but then you can start building at scale and start achieving things that any of these UI generators could not possibly like, loveable is nice, but it's a credit burner. Take the amount of money that thing chews up over the dumbest little fixes and hey, maybe they want to sponsor this one day, but I'm sure they won't after hearing this. But in truth, there's so many UI generators where they just degenerate junk. And this is also one of problems with cloud as well as like cloud is creative but also gets so many things wrong. And how do we coalesce these two things because it's creative because it gets things wrong? And we as users need to accept these two things both being true at the same time. Right, but isn't that sort of the difference again between these large language models and a small language model, right? The small language model will always be inherently more valuable for things that need to be held brand sensitive. Let's go into the theory of that, right? The large language model, the point of it is you can throw anything into it and it can process it because it's context on the entire world. A small language model would have a really small context window. So if you vary outside of its tolerances, it would get lost. So it has to be part of a specific bespoke workflow. Now again, this brings us back to the question of what's Apple going to do with the on-device processing because if that device knows everything about you and it's safe and secure and you're really bought into the ecosystem, you can start doing some powerful things and where co-pilot might have failed because it was basically a UI hack because really all the co-pilot is. Apple might have the ability to use something transcendental if OpenAI doesn't get their first, but OpenAI doesn't have access to who we are really. I mean, if you read what I call cheesemybot, so cheesemy in Spanish means gossip. I bet it knows a lot about me, like I've done that exercise of what do you know what do I look like. But it's scary, psychologically, psychologically, how much OpenAI, because that's the one I use the most. So that one knows me to AT, right? I didn't even have to personalize it. I don't have to have custom prompts. It just over time the memory, right? And Apple Pay, Apple Pay reveals so much of what we actually do. Well, and that's the scary part. So I'm actually one of those weird people. I don't use Siri on any of my devices. I don't use Alexa. Do I have accounts? Of course I do, because I don't want to be a weirdo, but I also don't use it. Like I disable it. I'm very precautious, but that said, though, I'm going to plant a couple of seeds here, right? Because I can't reveal too much, but you guys are able to read between the lines. But a lot of times we think about this or that, when it could be both, right? So when I think about language models, I actually don't think about this versus that. I think, well, what's the right approach for X, Y, and Z, right? I'm dealing with a specific niche because I work in Fintech, right? If I'm working on a specific banking flow, you know, a small energy model, duh, right? But if I'm like on the public side, where education is more important, because that's the thing. Context matters, right? I think about authentication. It's a gradient. Authentication is a gradient, right? You could be fully authenticated, or you could be partially, or you could be like not authenticated at all. So it depends on your cookies and all that, and where I would imagine, like maybe small language models are great, right? But if you're on the public site, anybody could visit it, potential, that's the difference, right? You have customers versus potential customers, right? So I could see large language models being used, like another applications where maybe you do need to have that larger context, or you use a large language model as a safety net for the follow-backs. By the way, I'm the execution queen. Like I love executing on this stuff. You could execute great and not have it work if you don't have the right strategy, which is why I listen to y'all's podcast, right? Because I learn, I think, about it, ask the right questions, right? Because in the end, what I've learned from both of y'all is that it's about asking questions. And it could be that you're both researchers, and that's what you do, that's why you get it. But it took me like 46 years to get to the point where that's how you get to a good answer. It's okay to live in the ambiguity. I don't know if you guys know Michaela Dixon, she's an AIUX designer from London. She posted something on LinkedIn, I think today, about what she had to unlearn. And one of it was like trying to speedrun through ambiguity. That's the opposite of what we should do with AI. And to your point, ARP, in terms of like, is there a legitimate use case? Well, I would say yes and no, right? So I think 2026 will be the year of accountability, where people will now have to like, like, you either get the AI winter, because they didn't roll out stuff that was good enough. It's almost like, you have all these people that want to do a gen and eye without the considerations they need for the basics. And you guys called it out, what's the value prop? What's the product market fit? If you can't answer those questions, maybe you shouldn't be working on what you're working on. But if you do, right, then you have to ask yourself questions. So what you're saying brought up two things for me. One is, if anybody listening is like, hey, I have questions or I want to talk about anything that we've been talking about on any episodes, please reach out to us. We love to have these conversations, so don't just sit back and be a passive listener. Number two, though, is to wrap us up, I want to bring us back to, let's say, a positive direction, right? What big shifts aside from accountability, do you believe need to happen in the next year so that we can move AI into a positive direction? I hinted at it, but I think Apple needs to prove that on device can be private and secure and valuable. They might shift our entire perception of donating ourselves outwards. The other one I'd say is we need to unlock some commercial use cases that really deliver transcendental value, and that's probably going to happen through personal health devices and understanding who we are, personalized medicine, transforming the way we look at our hormones and issues we can take on. This is something that we're going to take very seriously this year because we're actually going to launch a sister podcast about personal health and how AI is transforming it. But then I'd say that the other one, and I mentioned it before, I think this is the year that designers need to get up, accept that AI is here, stop complaining about AI, and figure out what we do now. Because in AI research in particular, where Britain is, there's still so much complaining going on that this isn't right, this isn't good, but guess what, while you all are complaining, the product managers have basically outsourced you, and we need to figure out how to fight back. Now, I think that the key to it is basically going to be either going out rogue on your own, because AI is a fantastic assistant for you as a consultant. The problem is obviously finding them work. Part 2 is unlocking services as a modern workflow mapping solution and an eval tool and a way to actually assess, well, is this probabilistic system working the way it needs to? Because guess what, we're so obsessed with the conversion, sometimes we don't understand the impact on the user in the community at all. I love that. I mean, you just gave me ideas. Some of the coolest use cases I've actually seen have been in the industrial, like what you mentioned about sensors, like real estate already has been doing that for a while, but it's multi-modal. That's what I'm thinking is missing, because people think it digital, and they think it interfaces. They're playing checkers, right? When in reality, it's 40 chess, where we have to think, well, not only just screens, but the actual service design, and we actually have to think like in and outside of the loop. So I would say for 2026, I would support definitely that as a group. We have to figure out what's the unifying message, because we have so many extremes, and we have, you're being nice about it when people complain about AI, like you're being nice about it. You're right. I think a lot of it is like accepting change, and then going with it and being okay with yourself. I think people are just worried about change in general, when it's like, dude, change is the only constant man, like get used to it, and be okay with it. You're going to be fine, but there's certain characteristics of what I see are people that are fine. Usually like optimistic, to some extent, even if they're cynical, they're so optimistic. They're curious to learn, right? They ask a lot of questions, right? And in the end, they challenge themselves to become different, right? If you are curious and you want to improve, and like you accept reality for what it is, you're self-aware enough, like you can succeed in AI. People are just scared because it's so new. Yeah, so I do want to stand up for that one, because there is an ethical quandary about forcing people to change and learn and adapt or else lose their jobs. And to that statement, I really want to put a giant highlighter on the fact is that I believe the tech might not be the industry that you thought it was anymore, that it's revealed it's dark side, and that you might deliver more value in one of the, quote unquote, slower industries. I know for me, I love working with universities. I love working with healthcare. And those organizations, the things that we've been doing in tech for the last five years, are monumental. They're transformative. Yeah, you might get paid 40% less, but guess what? You'll feel good at the end of the day, and you actually have a consistent job and consistent work in people that actually respect you instead of people. They're forcing you to chase some crazy nonsense hype machine and readapt every year. That's unsustainable. Dude, amen, brother. And with that, yeah, thank you both for this conversation, you know, from RP and I, I just want to say a huge thank you to everyone who has been listening over the last two years, who have gotten us to episode 50, which is monumental, and we feel very humbled with the, just the conversations that we've been able to have, and the caliber of people who are in this industry, and just the work that everyone's doing. And the amazing women, yeah, honestly, like we don't, we realize in retrospect, we were like, I think more than 50% of our guests have been women, and we have not had that as an intention or a goal. We just, whenever we see people doing interesting work or having interesting conversations, have reached out, and have been very fortunate to have a network that has allowed those women to rise, and just those people in general. So going into 2026, we're really excited to continue these conversations, to add in these conversations about healthcare and the impact of AI in that space in particular, which is hugely important to both our B&I. And yeah, if anyone has any particular areas that they feel are being underrepresented in the conversation in order to make them feel more confident and better enabled to help create humanistic products and services, please reach out and let us know. And we're really looking forward to continuing with all of you in 2026. Yeah, no, thank you from me as well. And I just want to say like in this year, we've tripled our audience, so clearly something's working. But as Brittany said, if there's topics that are under-heard or people under-represented, do let us know because sometimes where it feels like we're yelling into a vacuum. So it's nice when you have people like Yeti in her community that actually are tapped in and you can actually have some questions like, is this actually helpful? But yeah, thank you so much.

Podcast Summary

Key Points:

  1. Trust, cost, and orchestration are identified as the three critical factors for successful AI products in 202
  2. The initial hype around generative AI (2023-2024) has given way to practical concerns about safety, reliability, business value, and human-centric design.
  3. There is a significant tension between rapid, low-cost deployment using large language models (LLMs) and the need for precise, trustworthy systems, often favoring smaller, open-source models for specific domains.
  4. Current challenges include managing AI hallucinations, ensuring ethical use, addressing job displacement, and designing for complex human and systemic interactions beyond simple "copilot" functions.
  5. The role of UX and product design is evolving to require deeper system architecture thinking and advocacy for human and societal considerations in AI development.

Summary:

The discussion reflects on the evolution of AI from the initial hype of 2023 to the strategic realities of 2025-2026. Initially focused on generative AI's creative potential, the conversation shifted to critical practical issues: the lack of safety features, unreliable outputs (hallucinations), and insufficient consideration for human users and ethical implications. A key insight is that successful AI products will depend on trust, cost-effectiveness, and sophisticated orchestration.

The debate highlights a strategic divide between using general-purpose large language models for speed and innovation versus employing smaller, specialized models for reliability and control in regulated industries like finance and healthcare. Furthermore, the role of design is expanding beyond traditional UX to encompass system architecture and advocating for human welfare amidst business pressures to cut costs and automate. The summary concludes that while AI offers tools for productivity and creativity, its long-term value and ethical integration depend on deliberate design, proper governance, and moving beyond simply replacing human workflows to rethinking them entirely.

FAQs

The three key factors are trust, cost, and orchestration. These elements are reshaping AI product strategy and require teams to focus on them for effective development.

Small or open-source models are recommended for regulated industries due to lower risk and better precision. They offer more control and reliability compared to large models, which can pose corporate risks.

Hallucinations are a core feature of LLMs but can lead to unreliable outputs. The main worry is receiving convincing yet incorrect answers that are hard to detect, undermining trust in AI systems.

UX design now includes designing for concepts like agents and advocates, not just humans. It requires system-level thinking and architecture to handle diverse use cases and ensure reliability.

Businesses often struggle with trust, replication of outcomes, and change management. AI may expose deficiencies in legacy systems without delivering clear value, leading to user reversion to deterministic platforms.

Skepticism arises because AI often highlights tech debt rather than creating new value. Many implementations fail to address human-centric use cases or provide net-positive outcomes beyond cost-cutting.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.