Why Design of AI is becoming the Product Impact Podcast
16m 7s
The speaker argues that the AI industry has reached a critical inflection point: building AI products is now easy, but proving their value is the hard part. Two years ago, AI was expected to help design better products; instead, it accelerated product development while shifting control to frontier model developers like OpenAI and Google. This has led to a situation where teams blindly integrate LLMs without measuring real-world impact. The core problem is that AI systems are probabilistic: they silently fail in about 20% of scenarios, producing plausible but incorrect outputs that go undetected due to inadequate telemetry. Current measurement relies on vanity metrics like usage and adoption, ignoring decision quality and behavior change. To address this, the podcast is rebranding to "The Product Impact Podcast" and focusing on three themes: defining and measuring impact (not throughput), designing systems that deliver measurable value through trust frameworks and silent failure detection, and scaling AI value in enterprises by overcoming adoption friction and cultural tension. The speaker warns that without proper impact measurement, the industry risks repeating social media's mistakes, where negative effects on cognition, community, and mental health emerged years later. The goal is to ensure AI products are not just built fast but are truly impactful, trustworthy, and beneficial for users, teams, businesses, and communities.
Two years ago, we thought AI was going to help us design better products. What actually happened is it made it easier to develop more products. It shifted control upstream. The frontier model developers, like OpenAI and Thropic Google, increasingly shaped what's feasible, what's economical, and what good looks like. So we're no longer asking, can we build it? We're asking, should we build it? And is it delivering the right impact? People are building AI products right now and blindly trusting that whatever LLM they plug into will deliver impactful outputs. And that simply isn't happening. The deeper issue is structural. These are probabilistic products. It is nearly impossible to recreate all the scenarios your users will bring. Every edge case, every prompt style, every weird context collision. So what happens in the wild looks like this. A company deploys an LLM powered product. In eight out of ten scenarios, it delivers good quality results. In the other two, it silently fails. Not in a way that throws an error, but in a way that looks plausible. But it misses the point. It loses context or it gives the wrong level of confidence. And because most teams don't have the right telemetry, they can't see it. They can implement the thumbs up, thumbs down rating, or was this helpful reinforcement loops. But the silent failures don't always generate clean signals. They don't always produce the right learning. And they don't always show up as something teams can debug quickly. So then something predictable happens. We're shifting the onus to the model vendor. Open AI changed something. And then the problem is that the user is not able to do it. And then the user is not able to do it. We don't have mature value measurement frameworks for probabilistic systems. So we're left arguing with anecdotes and screenshots instead of evidence. This is why we believe that it's time for a new conversation to happen. So we're going to shift the design of AI to become the product impact podcast. This new season is more than a new name. It's no longer about building. That's not the hard part anymore. The hard part is proving that it's valuable, increasing the value in the data. Increasing the value and doing its scale without flooding the world with more slop. We're hearing over and over again that design with AI is easy. The problem is designing something good is unbelievably challenging. And these issues will become even more important as the industry starts incorporating more world models, more sensors embedded in everything. And AI becomes integrated into the day-to-day life in ways we can barely predict right now. Within a year or two, we're going to have ambient listening devices everywhere. Models processing huge portions of our work outputs automatically in systems enriching our actions with massive context. Millions and millions of records, signals and inferred intent is all going to be packed into these tools that right now seem really dumb. But the capabilities are going to scale in a way that we can't even believe today. If we can't measure and steer that impact now, we won't be able to do it later. So the conversation needs to shift with how do we design with AI to becoming how do we develop products that are actually impact. And we're seeing three themes that are particularly important around this conversation. So this season will be really focusing on these three areas. Number one is defining and measuring the impact of AI products. Frameworks for designing systems that deliver more impact. This is impact on users on teams on businesses and on communities. And the third one is around enterprise reality of scaling AI value. So how to expand product impact inside organizations. For each of these themes will bring the problem as it shows up in the wild evidence for why it's happening and frameworks that you can apply. So let's break down those themes a little bit more. So let's go back to the first one here, which is measuring impact not throughput. And this is a very important distinction. All of the capabilities right now are based on how quickly can we do it? How many tasks can we complete and some benchmark study that has nothing to do with the day to day. So our focus in theme one is measuring that impact right now too many teams measure usage adoption. And they assume they built something valuable because they did some QA they ran some evals but they're very very basic systems to measure value. But they have no visibility on the actual quality of outputs actually work. They don't really know anything much more than some vanity metrics proxy metrics. They don't know the real experience of those users unless those users complain. So AI can generate outputs all day long that's not impact that is just slot for better worse impact is what changes for the person using it. It's how does it transform the way they work and what does it save them that otherwise perhaps they never could have conceptualized before. We already have evidence that we're seeing about value realization not keeping pace with capability. Right so ISG's 2025 enterprise adoption study found that only 31% of prioritized AI use cases reach full production and only one in four achieve the expected ROI on growth. But that raises more fundamental questions. How should we even be measuring ROI for AI products? Is ROI simply did we generate more revenue than the cost of tokens and licenses or is ROI did we achieve impact benchmarks? Because in AI you can easily build something that generates activity you know messages interactions time spent without ever needing to prove the user's outcome. So this season we're going to define impact measurement that survives reality. This is out of detect silent failure. How to measure decision quality and trust. How to connect outputs to behavior change and how to build impact benchmarks that aren't just it answered. But instead looking at it mattered. Moving to theme 2 it's about designing systems that deliver more impact. Once we have that value measurement framework we can actually go back and revisit these systems. How can we look at them systematically. The bar has to move higher now you have to design something that's powerful, memorable and measurably valuable. And until you can measure it you don't know if what you're doing is good or measurably transformed people's lives. We're moving beyond the interface interaction now. We're building platforms that can change people's lives redefine value creation transform industries. Just look into medical professions. There's an endless amount of possibility there but until the systems are designed for the context and the way in which these things are used across the world. We're going to see failures mount and that's going to provide us the wrong kind of evidence. This isn't hype. It's the reality when each of us can be plugged into endless amounts of computing and contact systems will be enriching our activity with millions of records signals and patterns. Again, this isn't hype. This is reality. We're building the data centers now. So we will have access to these things. The product is an intelligence system. That means that for any of us who have been traditionally working in UX, we need to move one, two, three layers up. We need to think about what can it access. What should it be allowed to do? What should it be able to log in? What can we do with those logs? And how it recovers from being wrong? How does it accept being wrong? How does it accept learning? And how it stays trustworthy under uncertainty? And this is where the concept of flying blind stops being abstract and starts to become dangerous. You know, you mentioned, or be about in the healthcare system, well, a Reuters investigation published a week ago in early February, 2026 reported that as AI enters the operating room, regulators are receiving increasing injury reports. Raising questions about how probabilistic systems behave in these high stakes environments. Then I'm highlighting this not because I doubt AI's ability to transform healthcare because I actually believe that AI will be one of the greatest evolutions in health outcomes we've ever seen. But I'm bringing it up because we're still figuring out how to plug those probabilistic systems into scenarios where even a fraction of a percent of error can have massive impact on people's lives. So in theme two, we're going to focus on frameworks for trust for these systems and just in general systems design. Right. So this is looking at things like safe sandboxes and constrained access drift monitoring and reliability testing recovery UX and escalation paths telemetry that can detect silent failure and evaluation approaches that reflect real world use, not ideal use and golden pathway. So themes three is going to be about the enterprise reality how they can scale AI value inside organizations.
So while we can discuss how AI can deliver more value to specific products, how can deliver impact and how to design those platforms and systems, it's quite evident in the past five years that the real value creation comes from the enterprise deployments. There are the ones that can exist in regulatory frameworks. There are the ones that are built to actually fine tune those micro use cases and nuanced experiments where accuracy matters. So right now far too many pilots fail far too much value is being left on the table. In the enterprise world, you know licenses go unused habits are uneven use cases are fragmented and incentives are not aligned to having employees use AI in the way that leadership intends. And we need to name the obvious tension here right at the same time employees are scared their careers might be over we're trying to force them in that feeling to then use these tools and teach it how to do their job. Of course, there's cultural friction and rightfully so so we need to be able to think differently about this. So our focus isn't about forcing adoption or focus has to be how do we create an ecosystem. This is the products of services, the incentives and trust structures that scales value not only to the business, but to the employees and the end users. This is so important. These tools are rightfully valuable yet they're getting unused. There's so many licenses that are just stewing away and CFOs are getting frustrated. There's tons of evidence about this. So our goal is smart augmentation and an honest acceptance of the stark shifts happening in how we work and how we think. And we have evidence that culture and organizational alignment are the bottleneck right one survey of executives and employees reported that AI adoption is creating internal tension because leaders are pushing roll outs while employees experience fear uncertainty and misaligned incentives. So in this season of the podcast, we want to look at frameworks for things like diagnosing where adoption breaks inside workflows, designing trust thresholds by cohort and not necessarily just the whole organization. Creating incentives so using AI well is rewarded, governance that enables speed instead of freezing delivery and impact expansion playbooks that scale success without backlash. This next era of product development is about something special. It's about impact. We know the technology enables us to get things done, but we need to make sure we're delivering the right impact. It's going to attempt everyone to go and get into the same trap though. Yeah, it's faster to build than ever, but we have to make sure we're building the right thing. Design is more commoditized than ever, which should force us to design something truly meaningful, powerful, memorable, but when anybody can vibe code something, it's going to mean that we might have to sift through a dozen vendors to find the right solution for us. And even then we might recognize that the system behind it was not built to actually enable us to succeed. It was built to just get to market. And you look at research as well, research is easier than ever to get done, but it should mean that we have no excuse not to define stronger measurement for an hour so everything aligns explicit outcomes. But unfortunately, that's not what happens. And if we don't do that, we become like everyone else in Silicon Valley. It becomes an arms race to ship more and faster. That's not what we want to do. That's not impact. And that's how we end up with an industry that's full of endless products and endless stop. That's that one more layer, right? We know AI is changing us. We don't have full clarity on how deep that shift will be. There are twenty twenty five studies that raise start questions about cognition and work, right? Are we outsourcing critical thinking in ways that reduce our ability to reason independently over time? Does increased cognitive offloading make people faster in the short term, but weaker in judgment longer term? Or are productivity gains coming with a hitting cost like reduced learning, shallower understanding or dependence on model framing? So we want to also explore those questions seriously because impact isn't just what the product outputs. It's what it does to humans using it. It's how it shifts and shapes the impact on our communities. Most of us listening as podcasts and who have been part of it exists in a bubble. We've been there. We've seen this technology evolve. We can imagine what's possible with this technology, but we're going to enter into a true area of mass adoption. And we've been through this before with social media and mobile technology. We knew there would be risks. We didn't know how bad the impact would be until a decade later, loneliness, youth mental health struggles, fracture communities and fragmented families. AI is going to have some of these impacts and we need a measurement framework to assess those, pinpoint them and stop them. We don't want to repeat that same mistake with AI that we did with social media. So as always, we want you to be an active participant in framing this conversation. Contact us on LinkedIn. Find Brittany Knight. Tell us who you want to speak to. Tell us what you're seeing on the ground in your businesses and with your clients. Is value being delivered? Are we ignoring important signals about whether AI is achieving its intended outcomes? Is design having the right seat at the table and being part of the conversation how to build trustworthy and impactful outputs? We're doing this because we need your input and your participation to make sure this is the right conversation. So welcome to the second season design AI. And remember, from now on, we're going to be called the product impact podcast because that's the right conversation to be having in this next evolution of AI.
Podcast Summary
Key Points:
AI development has shifted from "can we build it?" to "should we build it?" as frontier model vendors now shape feasibility and value, but many teams blindly trust LLMs without measuring real impact.
Probabilistic AI products silently fail in about 20% of scenarios, producing plausible but incorrect outputs that go undetected due to lack of proper telemetry, leading to reliance on anecdotes rather than evidence.
Current impact measurement is flawed
Three key themes for the new "Product Impact Podcast" season
Without proper impact measurement, the industry risks repeating social media's mistakes—unforeseen negative impacts on cognition, work, and community well-being—and instead must focus on building products that are powerful, memorable, and measurably valuable.
Summary:
The speaker argues that the AI industry has reached a critical inflection point: building AI products is now easy, but proving their value is the hard part. Two years ago, AI was expected to help design better products; instead, it accelerated product development while shifting control to frontier model developers like OpenAI and Google. This has led to a situation where teams blindly integrate LLMs without measuring real-world impact.
The core problem is that AI systems are probabilistic: they silently fail in about 20% of scenarios, producing plausible but incorrect outputs that go undetected due to inadequate telemetry. Current measurement relies on vanity metrics like usage and adoption, ignoring decision quality and behavior change. To address this, the podcast is rebranding to "The Product Impact Podcast" and focusing on three themes: defining and measuring impact (not throughput), designing systems that deliver measurable value through trust frameworks and silent failure detection, and scaling AI value in enterprises by overcoming adoption friction and cultural tension.
The speaker warns that without proper impact measurement, the industry risks repeating social media's mistakes, where negative effects on cognition, community, and mental health emerged years later. The goal is to ensure AI products are not just built fast but are truly impactful, trustworthy, and beneficial for users, teams, businesses, and communities.
FAQs
The main challenge is that AI products often silently fail in 20% of scenarios, producing plausible but incorrect outputs, and most teams lack the telemetry to detect these failures.
It means shifting focus from how quickly tasks are completed or how many outputs are generated to whether the AI actually changes user outcomes, improves work, or delivers meaningful value.
Silent failure occurs when an AI product delivers a plausible but incorrect result, such as missing context or giving wrong confidence levels, without throwing an error, making it hard to detect.
Only 31% of prioritized AI use cases reach full production, and only one in four achieve the expected ROI on growth.
The three themes are: defining and measuring the impact of AI products, designing systems that deliver more impact, and the enterprise reality of scaling AI value inside organizations.
Measuring ROI is challenging because it's unclear whether it should be based on revenue versus token costs or on achieving impact benchmarks, and activity metrics like messages or time spent don't prove user outcomes.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.