Go back

Ed Bayes - Designing 6 months into the future at OpenAI

47m 37s

Ed Bayes - Designing 6 months into the future at OpenAI

The conversation with Ed Bayes highlights a transformative shift in design practices at OpenAI, driven by rapid AI advancements. Designers are now prototyping for a future six months ahead, focusing on agent-first interfaces that prioritize accessibility and user control. The shift from traditional IDEs to agent-based command centers—like Codex—has been pivotal, especially after model capabilities improved significantly in late 2023. This evolution enabled broader adoption among developers and non-technical users, who now interact with AI through intuitive, familiar surfaces similar to chat interfaces. Design teams operate with a persona-driven approach, ensuring products serve diverse users from developers to knowledge workers. Prototyping has become more dynamic, with designers creating live, editable branches in code to test real-time interactions, often in collaboration with research and engineering. These experiments are evaluated through iterative feedback loops, balancing model-generated outputs with editorial guidance to maintain control and quality. Tools like Framer and Jitter empower designers to build rich, interactive experiences without deep coding. Crucially, the role of design has expanded beyond aesthetics to vision-building—designers now lead by communicating ideas, fostering trust, and working with cross-functional teams. The emphasis is on humility, user empathy, and collaboration, with design being a foundational layer in shaping accessible, future-ready AI products. This shift reflects a broader cultural transformation: designers are no longer just building interfaces, but creating new possibilities for how people interact with technology.

Transcription

11448 Words, 61289 Characters

English
What I've seen over the past kind of six, 12 months is the slow change in how people build. So exciting to be a designer right now, it's, I think, given a lot of people like a new lease of life, right? They basically have this new medium that previously you might have felt kind of like locked out of. And now the only limit is really your imagination at this point. Welcome to dive club, my name is Red, and this is where designers never stop learning. This week's episode is with Ed Bayes, who's the design lead for Chatcha.pt work and codex. So we're going to do a deep dive into designing AI experiences and what it takes to prototype six months into the future. Ed even screen shares a couple examples of things that he's made and goes deep into how the role of prototyping has shifted internally at OpenAI over the last few months. So let's start in this conversation by hearing more about the original goal behind the design for codex. The history of codex at OpenAI is actually pretty long and storied. So, you know, I think most people know about the app, they might have known about CLI, obviously our models as well. But, you know, for a number of years, to clear on the research side, we've been really pushing coding agents as, you know, kind of the next frontier of model development. But from a product perspective, I guess kind of one of our first big launches was at the beginning of last year, we launched a kind of web-based product called codex. And you connect your GitHub and basically kick off these async tasks, so you'll, you know, send a prompt, you might want to fix a bug or, you know, you get some feedback from Slack. You paste it in and then you then get a ping a little bit afterwards and you then have like a PR to review. It was a really well-loved product, but at the time, like a lot of other products in the market were kind of going the opposite way. They were going like synchronously, so you had maybe called code and others and they were like, it's some of these products that kind of meet people where they are. One of the key kind of frameworks that I use here is like designing for six months out. So if you just presume that model capabilities are going to get better, and you think about where we might be in six months, and you kind of start to build products for that world. That's kind of a good framework to use, and you know, we could see model capabilities getting pretty incredible towards the end of last year. We saw an opportunity, I think, to build a product that lived in that future, pretty similar to the first round of cloud-based products that we created, where from an information architecture perspective, maybe you can move the chat to the left from the right, the agent becomes the primary, you know, interface the interact with, and reviewing the code and seeing the code is important, but it's maybe less important. Whereas in the previous world, in the kind of IDE world, you know, you had code was on the left, and this kind of, you know, the primary surface that you interact with, you might go in and change values out of the code yourself, and then you have this kind of extension on the right, which kind of you may be going to time to time. But given where we saw the models going, we took a pretty opinionated stance to build a kind of agent first product, which was, you know, explicitly not an IDE. You know, we didn't take some of the cues that maybe others at the time took, and we just, you know, we just looked to the future, and we thought, okay, in a world where, you know, models would get better and better at it. Like, what would the ideal interface for that look like? And in that world, as I say, first, the agent is on the left, so that's the primary thing that you're interacting with, but also you're going to be working on kind of more than one thing at once. And the core thesis of the product was we wanted to build a kind of command center for your agents, and that meant that maybe you're working on multiple things at time. Maybe you're working from a development perspective on multiple repos at a time, and that kind of influenced a lot of decisions that we're making. But yeah, at the time, I think it was a reasonably kind of controversial decision, you know, in some ways, we were kind of the undog, so we were able to take perhaps some of those biggest swings, but I remember internally, even at some designer views, you know, some feedback was like, well, should we just build an IDE? Like, you know, I'm not sure, particularly because the models weren't quite there yet. So it's like, I could see it working, but like, you know, maybe a few other things have to land. And then, you know, as kind of from December, we had a few real kind of step changes in our model capabilities, it became very clear that, you know, this was the right form factor. And we really saw that internally, the eternal adoption of the desktop app over other services just really skyrocketed, particularly among developers and quite quickly among other use cases, non-technical use cases. We kind of figured, okay, we're on something here internally, so by the time we released it, you know, we felt pretty confident, and then, you know, it's obviously growing up much since. I remember that period in time where the paradigm of the left agent chat came out, it was a little bit polarizing, and people almost felt like it was too abstracted, and now it's just, we've settled in, like, it's the pattern, right? And I think even now, you know, we have, we have a right panel, you can open, you can review code there as codecs is now evolved into kind of chat, pretty desktop app and chat, pretty work on web and mobile. You know, we now have a bunch of other use cases that you cannot generate slides, you can generate sheets, and this kind of all opens there. But I think the thing that I, maybe I'm kind of like most, most proud of, almost excited about is, you know, you can live in the product without that, and you can live in it in a pretty familiar interface, like more similar to say, like, chat, BT, than like an ID. And to me, that I think is the really exciting thing about where we are now and where we're going, which is, you know, with chat, BT, you basically had this kind of universal interface, asking questions, and if you're PhD rocket scientist, you'd go to the same surface as if you're someone learning to knit in like the Midwest, right? There aren't these like two products for different different archetypes, and I think what we're seeing now with the launch of chat, BT work, and these more accessible versions, I think we're seeing this kind of universal interface for getting what done, which I find really exciting. It's, it's, you know, yes, we still have products tailored for developers, developers will have their own workflows, but creating these like really accessible interfaces just really expands the access of the power of code agents to more and more people, which developers kind of got early, and now I think rest of folks are kind of getting on to know us off. Real quick message, and then we can jump back into it. So I'm working on any dive website right now, and let me tell you, having everything I need within Framer is a pretty big deal, it's so much more than just a website builder now. It really is a full AI design agent. I can explore ideas on a canvas, build out responsive pages, even set up my CMS for all of the episodes, which is incredibly helpful, but it's not just about the convenience. Framer also raises the ceiling for what I can create, like they just released a whole set of 3D tooling so agents can work with perspective, depth, origin, and then I can fine tune properties directly on the canvas without having to prompt back and forth. It's just another reason why I build absolutely everything for dive on top of Framer, so head to dive.club/framer to start creating today. Remember when Jamie Ganon came on to talk about being a creative director with AI? Well, she was part of an incredible branding project for Claire Voes new company, and what stood out to me was how she brought the entire unveiling to life with motion, and to no one surprised, she used Jitter. On one hand, it's so simple and intuitive, even if you've never done motion design before. On the other, it's incredibly powerful, and you can build your own creative tools to create all kinds of custom animations and effects. Jitter's been my go-to motion tool for years, and I cannot recommend it enough. Had you dive.club/jitter to get started now on to the episode? Can you go a little bit deeper into that process of bringing Codex in to ChatGabeeT, because ChatGabeeT, I think, was kind of the leader in terms of consumer friendliness, and super simple, like grandma uses ChatGabeeT, you know, and yet you have to take all of the power, not only in the present day, but thinking about that six months from now, vision, and fit all that in. What were some of the design challenges that you were wrestling with when you first started thinking about what it would take to pull that off? Yeah, I think there are a few. I think first one of them, and it's really gated on model capability. You have to pass a certain threshold of confidence that the model can do a whole bunch of very complex things, and once you kind of pass that, then you can start to really pay back the interface. Before then, you really needed to be able to verify a bunch of changes, right? So, again, thinking back to the kind of IDE extensions, CLI world, right? You might have your code based up in front of you, and you can kind of validate changes, maybe you go back and change things. First, you need the kind of confidence that the model capability is pass a certain threshold that you can start to abstract that way, and then, you know, on the other, I think, you know, developers were super early adopters around this technology, and, you know, I think, a little bit more comfortable with this idea that, you know, I open this coding agent. And it can, like, edit files in my computer, right? This is actually, like, a completely new paradigm for lots of people who may have interacted with, say, chat to be tea and other products, right? In the web, it doesn't access to your local files, kind of historically, maybe can't, can't take that many actions. So there's this kind of new way of working with these agents that you kind of have to bring people along on that journey. And I think for developers, they, you know, got it very early, and we built a whole bunch of, like, great sandbox productions and kind of safety processes so that if the model is taking some action, maybe, like, might be running a command on your computer, you can choose to approve it, you know, in a case by case basis. So developers got quite familiar with this early. And then I think, like, one of the key challenges now is we, as we move into this, this world where we're making these products way more accessible is also bringing a lot of that safety thinking and, you know, keeping people in control. So, for example, say you connect your, your Gmail, right, you know, shouldn't just send emails for you every time you have to approve things. And again, I think just like bringing a new, a new kind of cohort folks along on this, on this journey and teaching them, you know, this desktop app can edit files on my computer. It can like, you know, send emails or Slack messages, right, the whole bunch of like design challenges associated with that. So when all of the AI frenzy started, it was a bunch of people coming onto the show and talking about how we got to push past the chatbot, you know, like that entire UX primitive felt like, oh, it's so clearly a stepping stone. And I think the pendulum swung where people are like, yeah, actually, like, chat's pretty good, you know, I'm wondering you internally, what's the dialogue for you all even? Particularly given how fast the rate of changes happening, you know, we've kind of, as I say, like, you know, we launched the codex desktop app, I think, because around six months ago, before that was obviously used internally. But like, I think like developer workflows have dramatically changed from like December to now, but like that's in the grand scheme thing. That's like a very. short period of time. So no, I don't think we'll lock down at all terms of kind of the final patterns. And I think you're starting to see that now. I think you're starting to see some of the, some of the patterns that are optimized for maybe where the models were a few months ago are starting to be stress tested, right? So for example, you know, left side bar architecture, if you're kind of working across multiple things at once, like how do you navigate, how do you kind of show the things that are highest priority at once? If you want to kick off 10 things at once, like how do you do that? How do you stay on track? So, you know, the original thesis or the desktop app was, as I mentioned, kind of building this command center for your agents. And I think we had the first situation of that, but I don't think, you know, we're over. And particularly, again, if we were to kind of take the same framework and look forward and think, okay, where are model capabilities going to be in six months? If you kind of project that forward, I definitely don't think the patterns today are the final ones we'll land on. I want to dig into how you all operate in the practice of design and what that looks like. But maybe before we get into the weeds, can you just give us a little bit of context for the surface area that you're thinking about and also like team structure, how many designers, who's working on what, how you even spread out the work, that kind of a thing? Yeah, totally. Yeah, so I run the design team for kind of codex and chat to be to work. So, we have a bunch of surfaces on the codex side, we have the codex desktop app, you know, we also have CLI product, we have an ID extension, and then on the chat to be to work side, which is our kind of a genetic product, you know, for non-developers. We also have a desktop app, but we also have chat to be to work on web, and also chat to be to work on mobile. So, quite a bunch of different surface areas. The way I like to structure teams is actually kind of less around features and, you know, kind of like narrow parts of the pie, and trying to give designers like a broad scope, might just like general philosophy is design can add value, you know, across the whole range of the product development lifecycle, it's not this kind of last mile polish, it's a lot of kind of early product thinking explorations. So, you know, the way that I structure is kind of more around personas than features. So, maybe on the codex side, you know, we have designers who are kind of focused on like a developer persona. So, maybe on a day-to-day basis, they're working on a feature, maybe it's like a code review feature, or, you know, something rated to the codex desktop app, or even the CLI, but, you know, given that folks have this kind of higher order kind of scope in their mind, that means that, you know, when we're doing things like onboarding, right, when we're like shipping kind of new general features, we make sure we don't drop the ball on developers across that whole lifecycle, you know, similar for other areas, right? So, you know, we have features like sites, or we have features like you can create, you know, sheets and slides, and we have designers who own those product features, but, again, you know, they're not just kind of like optimizing for like some specific metric for that specific feature, because you kind of lose the bigger picture. So, they're kind of spread across the whole lifecycle of, you know, what it means to be a knowledge worker using our products, or what it means to be like a pro-sumer. So, we have a, you know, budget designers spread across these different services, day-to-day might be work on features, but owning that, that kind of broader lifecycle. Can we go to, like, one little deeper on the personas thing? So, like, are there designers on your team that are specifically focused on like non-engineer personas? Is there another axis even for power user versus new user? Like, how does it even work? And maybe how does that impact the way that the team collaborates? It gets a bit fuzzy, right? Because, you know, developers can mean, you know, many things. Knowledge workers can mean many things, but the way that it shows up, I think, in the day-to-day, is we try and stay super close to users. So, you know, we do a lot of UXR. We, you know, try and encourage designers to spend at least kind of one session a week with folks either internally or externally, who might be kind of newer to like a genetic product. So, you know, it might be someone working and say recruiting, or legal, or basically like non-developed personas. And then we, you know, also in our kind of weekly, you know, talk through basically kind of what are there? What's the going to good, bad, and the ugly of the product at the moment? Like, what's going well? You know, what's not? And where can we improve? And that is really a forcing function as well for folks to stay close to kind of external UXR. So, it might be ready, it might be hacking news, it might be Twitter, you know, ideally something more like quantitative. Basically, trying to create an incentive structure so that people feel like as close as they can to users, and you're not just kind of like sitting in this bubble designing. All right, so I want to return to a phrase that you said earlier where you talked about how working at a place like OpenAI, it's important to have this like six months out mentality. And I'm wondering if you have examples for what that looks like in practice. And as a designer, you know, how do you even think about growing this muscle of prototyping six months into the future? You know, I think it's been this like holy grail for a lot of designers of like using the model to like generate very rich outputs, not just text. So, maybe widgets, or like Rich UI. And the like the super agent build version is you don't have this kind of like strict design system. You just let the model really do kind of like whatever it wants. And it can choose whether it responds in text or responds in Rich UI. But kind of getting to that world is kind of gated on a few things. First, the models have to be great at front end. Our models that make grade strides in that space. But, you know, designing for this maybe six months ago, you just kind of project forward and you just, you know, assume that things will get really good. I think also one of the challenges here is kind of designing for the constraints that you work within. So, there are kind of like maybe two camps in the in the in the gen UI space. Like one is this kind of highly editorialized world, which is like super low latency. And the variance is reasonably small, right? So, maybe if I send a message on chat to PT and I ask for like maybe a kind of sports game score, like each time I probably don't want like a different widget, like showing me like like what the score is or like if I ask for like something. Yeah, maybe like a search result something. So, you know, we have this really great design language that appears just in time and is optimized for latency. And, you know, making just like answers like a redeveloped response. One of the benefits on the kind of codex side is, you know, we use super powerful models and I think uses skew a little bit more more power use route clear on desktop. So, we have a little bit more leeway with like latency, right? Like if the model is generating like a full output, you know, it might take say, I don't know, kind of 10 seconds or something. Whereas, if you're on core chat to PT, you want to really make it like super, super snappy. But with a little bit of that flexibility, I sat down with some of the research team and you know, we talked about this for a while of like, you know, what would this look like? We haven't really explored this before. You know, we kind of had these like three phases of prototyping first. You prototyped in Figma and you, you know, bring this kind of like beautiful instruction to create you kind of stress test a territory part. And then, you know, once the model's got good, got really good and then also, you know, if you kind of had a development background, you know, itself as a designer, you could kind of hop in and build like prototypes. So, we had like a kind of mini version of codex when we were first built in the codex app where, you know, another designer on the team and I, we'd kind of like hop into the shared code base and we basically build this like fake version to like test out interactions. You know, might even like pipe in the API, just to kind of see how these like different responses felt. And then now, we're kind of at this third stage, which I kind of got it like, I don't know, it's like vibing in prod. You're probably not going to merge it, but like, you'll create a branch and you can just do the most kind of insane out there idea that you like. And the way that it feels and where you experience it if you're working with it is, you know, it feels like the real product. And you can just share that branch with someone, you know, they can, they can check it out and they can kind of feel it out themselves. I did a bunch of early explorations for this kind of GenUI work using this approach. Tested a bunch of different, different examples, but maybe I can just like show you, you know, the prompt here is like explain back propagation, you know, visually and as you can see here, so like in amongst the, the final system response, the model just renders just a fully generated answer that's super interactive. So, you know, this is a very early prototype. It's a pretty rough now compared to where we actually landed. But, you know, you can kind of play around with learning rate values, you can go in and then, you know, the month, you can also kind of click certain parts of the inline visualization and it can send a follow up response, right? So here, you can kind of continue to gradient send and it will kind of continue walking you through this like learning experience. Here's kind of, you know, step two. So again, I just like render something in real time. So this is fully generated by the model on the fly. How much guidance at a system level are you doing if any? I mean, this is I think like another really interesting part of the design process, which is there is a little bit of like kind of traditional design maybe, right? How does it render, right? If you're going to copy it, like, where's that, you know, affordance. But a big part of this actually is kind of like skill writing, right? So there is a world where you just let the model generate everything, anything, you know, and again, in six months, maybe front end capabilities of models will be, will be at a level that that that it can be like super consistent and feel great. But until then, we kind of felt good to have like a little bit of editorialization. So as you'll see, I can show another example where you'll start to see some patterns. And if you use it in prod now, we've kind of, you know, iterated on this over time. That means basically instead of, like, speccing out like exact ways these visualizations can occur. Like if I just show you another example, real quick, you know, this has some like 3D visualizations, you know, if you were to spec out every possible version of like what 3D rendering might look like, you basically you wouldn't be able to do them because because it's like fully, you know, fully driven by the model, right? You can also have like different, you know, orbital scenes. So yeah, instead of like speccing out the pixels, you basically move to the skill level and you kind of define a design language, you know, within the skill. So maybe you reference certain tokens best practice. And then you build a bunch of ebals, right? So, you know, the way that I did it is you create a kind of golden set of say 10, 20 prompts that we want to optimize for. So maybe this learning example, right? You know, for research. And then, you know, you're running through, you'll see the output, you'll see what went wrong, maybe you'll tweak it. And then you can start to build this kind of like flywheel that can get better over time. And again, once you dog food across the company, you get a bunch of feedback. So really iterative and like evil driven. And after this early prototype handed this off to just like a brilliant engineer on the team, who really kind of like took a very, very scientific scientific approach to this to kind of help climb improving the visuals. But yeah, this is a good example of like a really close collaboration with research, design, engineering ultimately to push forward some things that we kind of talked about for a while. But we're really unlocked only by these kind of like recent model advances. A quick rabbit trail that I want to go down is what eval driven design looks like in practice because it's something that started popping up on the show. But I don't think we have a great example for designers who just are not in that environment. In my mind, you have you have the science of design and then you have the kind of of design, right? The science of design is maybe typography, spacing, right? It's the kind of, I don't know if you've read any of the kind of initial, like kind of Taylor type expressions of like ergonomics, right? It's there are kind of like scientific ways that that uses interact with interfaces and we can optimize to that. And then there's the art side which is kind of, you know, more open-ended. It's like cult, you know, driven by culture, changes over time. The extreme end may be more like vibe-based. So I think a lot of this is kind of leaning skewing towards kind of the first art. You know, first I think just like getting really concrete about like what are the use of problems that you're trying to solve here with this feature? And then, you know, writing out, you know, a bunch of basically kind of like golden flows that you want to nail down. And there are a few ways of going about e-values, right? Like one is super scientific. You might come up with like a thousand, you know, 10,000 different problems. You can go through all these outputs. But the like the baby step version is maybe sketch out 10, 20, 20 kind of like hero flows. You want to kind of take a crack out which, you know, some of these examples went through. And then just basically like run it over and over whatever interaction will feature your, you know, you're trying to test out, you know, test it for these flows. Does it work for that flow? If not like where are the areas that it kind of falls apart? It really depends on the feature. But for this, like a lot of like kind of quite in the weeds of like, okay, a bunch of these like interactions don't really work. So maybe sometimes it draws like SPGs instead of, you know, icons. So maybe you can give it like a little bit more guidance there or sometimes it kind of strays in my token system. Again, maybe you can kind of give it a bit of guidance there or maybe there are specific libraries, right? Like through the libraries that others that you might want to lean into. So getting like really scientific about it, I think like really helps. And again, you can do that with like a small N, right, 10, 20 prompts. And then I think just like building a really close partnership with research and engineering to scale that up. So building an evil harness where you can actually test the interactions at scale. You know, they do this yourself or you can you can partner with someone to do it. You know, it's lucky again, as I say on this team, I have a really, really fantastic mentoring partner, you know, Philip, who's who's led a lot of this. Basically creating the evil prompts set, you generate, you know, huge number of outputs. And then you can just go through them, you know, by hand and, you know, a whole bunch of other details that I probably can't go into. But I think like first, just like really getting crisp on like, what's the problem you're solving? And then finding a way to like test it in like a verifiable and scale away. I'm realizing, listen to your talk that maybe one of the challenges with this is that you find a bunch of ways that it's falling short, but then you write overly prescriptive fixes. But then all of a sudden you have a bunch of them and you're almost like capping what the next generation of models could be. Do you have to intentionally figure out like how to walk that tightrope? Yeah, I mean, I don't know if you've had a bit less than it's like the original go to essays in in the ML space. I'm gonna paraphrase it pretty badly, but basically like any kind of like scaffolding that you do like over scaffolding like some of the things that I just talked about, we'll kind of get washed away with like new model releases. So I think you just have to say super flexible, right? So say, you know, we ship this with a particular model. And then when a new model comes out, right, you have to test it all. Maybe you can actually remove a whole bunch of the scaffolding that you had before and you start again. Maybe you find some new issues, but seeing it is a very dynamic process. And again, with that north star of like in six months, like maybe you actually don't provide any guidance, but until you go there, right, you can kind of help steer it along that way. With your other prototyping examples that you wanted to share, I noticed a few more windows. Yeah, I have like a goofy one. So it's it's definitely like not the like, you know, most polished version, but it's kind of related to the persona driven approach that I tried to build with the team is my philosophy on on product design is, you know, work very, very closely with PMs and other kind of product folks, but really try and push folks to push forward their own ideas as well. If they have like specific, you know, vision vision of the future. So one very early prototype that I jammed on was at the time we called it kind of autopilot. And it basically evolved into what is called goal, which I think has, you know, been used by a lot of folks. And it's this idea, it kind of came from one of these, maybe this kind of loop meme where, right, you maybe you say to the model, I want to, I want to create this website. And then you'll just queue up a bunch of like continue, continue, continue, continue prompts. And there was a bit of a meme at one time, you know, online that kind of like a lot of people work in that way. So again, like working with research, working with others, you know, on the kind of harness team, we kind of sat down and thought through, okay, what would it actually look like to slightly reframe prompting around more of the like objective based model, right? So say your prompt is an objective that has some success criteria. Can you basically kick off a prompt? And then when the assistant term finishes run almost like a classifier and be like, okay, did you solve that objective? If not, kind of carry on. So more of this like autopilot was, this is like super goofy prototype, but it's a riff on that, right? So you basically can turn on this autopilot feature, right? Maybe you, so you define, you know, what might the success criteria be? And then you kick it off. And then we have this kind of like Figma-style UI, maybe where, you know, the, the agents in control and it's sending messages. And, you know, when the assistant response finishes, you know, we'll think it will consider, okay, did it, did it solve that, you know, did it solve that objective? If not, okay, carry on. And like, let's send a prompt together and it can, you know, kind of run basically as long, as long as you like. So this is a goofy, goofy example. I think it was like, you know, create the best joke in the world or something, which is obviously one of those like tricky problems, which doesn't really have a, you know, an actual, an actual answer. But this was like an early prototype that I just jammed together. We have these great kind of demo day type things on, you know, on Fridays, where we share prototypes. And then yeah, this eventually evolved into much simpler interface, which is goal, you know, which you brought into the, into the harness and the products. But just an example again of like vibing in products. So this was just a branch that I jammed on. Probably like a weekend thing of like, actually, what would this, you know, what would this look like? What are the different interactions that you might have? Do you just set a prompt? Yeah, do you have some weird takeover thing? I want to talk about that Friday demos environment for a little bit. So everybody's bringing prototypes to the table within the design team. What percentage of those prototypes are in this like vibe pride category three type of prototype that you're talking about? Definitely increasing numbers. I think like what I've seen over the past kind of six, 12 months is the slow change in how people build. As I mentioned earlier, there have been these kind of like three waves, maybe over the past like you were told, you were so, or a few years, where you had like, you know, a figure of prototypes. And then you've had like these like mini kind of fake prototypes and then this kind of like approach where you just like create branch of broad, you know, one thing that I've seen over time. And I, you know, we have, we also have like work and progress, you know, design channels for people post stuff. And they've been a bunch of folks who I've kind of sat down with, you know, chat to and I think like, you know, they haven't coded before. And then suddenly they're like sending these like links to like, you know, absolutely built nothing. I'm like, wow, like, you know, how did you build that? Like, yeah, you know, start using codex and it kind of like transformed a bunch. And even recently, you know, we designed on my team, you know, amazing designer, never decoded before. A lot of people internally use origami. We put a lot of folks from from Meta. I remember going up to to his desk and and he had this product and was like, wait, like what, what is that? Is that a prototype to make a code? Be like, this is like actual chat GPT. Like he, he was like, I don't know what it was. I don't know what has happened to a how, but I just ask, I ask codex like open chat GPT locally. Credit branch makes an edit. He's like, I don't know what happened, but like, here we are. I have my way to product. So it's dramatically changed and I think like most people now are kind of working in this new way. There's one question that I can't stop asking myself. What if companies apply to talk to you rather than the other way around? And that question is the foundation for the all new dive talent network. And it's working. Like right now, I'm helping many of the most exciting startups that I know to hire the designers and builders who listen to this show. So if you're curious what might be out there and maybe you want to get on my list or maybe you're even looking for your next design hire, head to dive.club/talent to join today. You can talk about the work and progress channel. Can you go a little bit deeper into this past few months? You have this transition in the way that people are working. How is that changing what collaboration looks like even the way that ideas are shared and rallied around? Yeah, good question. I think another big unlock that's happened in turn as well is we have this feature called sites. So you can basically, you know, you might create a local post type and then you can just say, like, basically, chat to give me a link and it will deploy it to a URL and you can decide who can access it. And then you can just share that link with your team. You know, maybe you'll be referring on an idea and you can just like DM someone a link and they can actually play with it live. So, yeah, the work and progress channel has definitely moved a lot in that direction and we have some fun internal like aggregators for these different prototypes so you can see what people can jamming on. Right, maybe you can like fork them and start to jam them yourself. So, yeah, I think it's become this like very rich environment where it's kind of like, you know, links instead of video is like remixing instead of like starting from scratch. It's been it's been very cool. It feels like honestly, it's so exciting to be a designer right now. It's I think given a lot of people like a new lease of life, right, they basically have this new medium that previously you might have felt kind of like locked out of and now the only limit is really your imagination at this point. I totally agree and I feel that I think the difference between working on even just web apps versus native apps is a little bit I think rather extreme like it can be cumbersome doing the native app thing like yeah, which is why maybe the reason I was asking about like how many of those prototypes are in phase three is even just, you know, there's there's quick and dirty prototyping with these new capabilities and then there's like, no, I'm literally forking chat GPT used by a billion people and designing natively and like I was interested in like how you all even think about that split and maintaining speed, but also I get that you have to like building a top of the APIs when everything is non-deterministic is really important too and in that tension, I think it's one of the more interesting things to dig into right now. I think it really depends on the problem that you're solving as well right. So for example, again, just these kind of concrete example, you know, when we released chat GPT work, we also released a new way of changing model and reason level we introduced this slider. So again, it's designed that I mentioned, you know, Tara and the team kind of worked on this and you know, I think for that like the very small interactions and navigation around how that works and how it interacts are, you know, reasonably self-contained. So for that, like, we just had a prototype that was, you know, basically like kind of local React prototype, you jam on it, you kind of work out this very specific interaction, you know, then you hand it off to engineers. But say you're working on some bigger IA change or exploration, like you kind, like, first you need some inventory of like chats. So, you know, if you were doing a local prototype, you'd have to create those. And, you know, the medium that we're working with is like so non-deterministic that without the kind of water flowing through the pipes, it's hard to reason about, right? Like, you know, some of these ideas like sound great on paper, but when you actually like send a bunch of prompts, as I was saying, right, with the, you know, with the visualized example and, you know, building like kind of golden evils, it's like, it can look great in a design of you on a prototype, but actually when you use this, use it, it's completely different. So, yeah, I think the nature of the medium of what we're designing with kind of necessitates that. But then for some of these like small interactions, it's not required, right? You know, and then maybe like for design systems, like you don't need prototype still, right? Like, I think there's a lot of work still happens in Figma. I think it's like a really important tool in that toolkit. But still, I just think like the scope of the tools that people have access to is expanding. What type of work is still best suited for Figma in your mind? The types of things that people typically talk about is the kind of explore versus exploit, right? Like at the early point, or if you're thinking about like the traditional design process, the kind of, you know, exploratory phase, you know, you want to go wide, you want to test up and stuff. You want to just like do it in like very low fidelity, maybe like wireframes. That's often used a lot. But I find, yeah, like, so there's that side, but I also find it very, very useful for systems as well, right? So you're designing, I don't know, system that someone was jamming on recently was like commenting, right? We have a bunch of different surfaces that people might comment on, right? We have a browser, you have slides, documents, you know, some new things, and we want to have this system that works well across them all. And you kind of want to see everything in one place, and you want to get like really specific on the pixels, and you want to spec everything out, so like hand off to end, you know, go as well, because it's not maybe the kind of thing that you'd want to just like rush. You want to really get all the like props and pixels exactly right, because it's like, you know, kind of high complexity part of the code base, the composers, another good example of that, or maybe kind of like things like the left sidebar and, you know, like sidebar rows, like you really want to get very concrete about all the different edge cases, the exact, you know, say padding or like typography, that I think, the initial exploratory phase and design systems, but, you know, in between as well, I think like a lot of people really can go back and forth now, particularly with, you know, some of the MCPs that they have, it's like, you can kind of like take figmas and put them into code real quick, you can also go the opposite way and spec out a bunch of different versions of figmas that way. So, you know, we definitely, definitely going in out and still use it a lot. I want to talk a little bit more about how you all operate. And I think my question is like, you have this team of like eight people, and everybody's slinging prototypes around and everybody's feeling super empowered as builders, and at least from my vantage point, it feels like we have way more ideas flying around, therefore some of the almost saw skills become even a little bit more important, because you can't build everything, right? And so, when you kind of look at the team and maybe even other designers inside of OpenAI, like what are some of the traits that you notice in the designers that are having an outsized impact on the company, on product strategy, that kind of thing? Good question. I think it also comes down to like a core quite, like I think in a lot of this kind of debate around pros typing, it's kind of gets lost in the source a little bit about like tools that people use, and it doesn't zoom out and think about like what is design, like what is the role of design, why is it different from front engine, why is it different from product? And ultimately, you can have the most compelling prototype of the world, but if you can't convince anyone to build it, then it'll just like say on your desk. So, the ability to communicate really well, the ability to express your ideas in very simple ways, the ability to work closely with engineering partners, labs like research partners who communicate in a whole different way, how a whole different lexicon, different ways of seeing the world, so empathizing with them as a stakeholder user, and communicating your idea in a way that's compelling to them, that doesn't just solve the user problem, but also, you know, brings them along in the journey as well. One of my Australian opinions at the moment, which is maybe a little surprising, given that I work on kind of codex and coding products, is I really don't think that you need to be able to code to be a designer, like I kind of think it comes down to like what is the role of design, and for me, it's basically about kind of like building a vision of the future and then bringing people along, but yeah, there's a kind of quote from one of my professors that was really stuck with me, which is you're painting a vision of the future and bringing people along. There aren't many other functions that do that, and to do that, you need to communicate, you need to bring people along in that journey with you. So, you know, some of our most effective people on the team, like one person, you know, joking with me the other day, they're like, he basically spends most of his time arguing and slack is the kind of joke, but you know, what that really means is putting forward ideas and bringing bring people along, you know, on threads where people have different, different ideas, and that might mean huge shared an example once of like coming up with a very type of like scribbling on a ride just to kind of show an idea and putting it forward. That might be just as effective as, you know, spending whatever a week building this like perfect prototype off, you know, in a branch off of master or whatever, burning it down to kind of what is the role of design, and how can you be effective? I don't think that's changed really over time, and you can use tools, you know, to your advantage, and to your point, you know, you can also use them to get distracted. Like, you can just make a whole bunch of prototypes, but maybe nothing is built, right? So, they're always trade-offs. Let's say that you do get people excited about a prototype that you've built in code, maybe it is something that is just a branch off of prod. Can you talk a little bit about how Gandalf has changed? Like, what does effective collaboration with engineers look like now that we're making a mess in the code base versus handing over these really neat and polished figma mocks? I mean, I think it typically follows a process, right? So, so first maybe like maybe you have some, you know, sketch your feature from some, maybe it's a product manager, others, you have some like, cool product requirement, you're kind of playing around with it. So first, you're going to sketch out one of these ideas, you know, paint division, or if it's like design driven, maybe you have some cool thesis, you know, about where you want the product to go, and you can kind of build that vision. It's not the case then that you just kind of like send this, you know, slot 8,000 line PR to enter them like, you know, off you go, you know, from there, I think it's then beholden on design to get, you know, more scientific about things. Maybe you write you then, it's very, you know, sketch it out in figma if required. There are like big systems problems, you kind of get Chris, but I think it's, it's kind of, to me, it's almost just like rallying, it's just rallying cry of like, here's the future, and then you just get into kind of the bread and butter of just designing and collaboration, and sometimes that means just going really deep in maybe design review and kind of like iterating on an idea, inspecting things out in figma, or maybe it's like, maybe it's almost there, and then you're kind of sitting next to an engineer partner, and instead of maybe kind of like reducing the latency between, you know, you figma and then maybe you just sit down with them and you compare and kind of like, you know, tweak things in that way. So then I'm thinking about a lot recently, it's like, is the way that we're still thinking about handoff slightly skew morphic in that we have these new sets of capabilities, we're exploring in code, but then we almost try to retrofit it back into what we're familiar with in terms of like human to human handoff, but in the future, actually, you won't even really think about the person's consumption that much, and we'll purely just be handing off to our engineers agents. There is, there is a world in which we get there, and yeah, we, you know, we may be there in, you know, six, 12 months. I do think though, for now, like, there's always a human in the loop in these decisions, you know, whether it's one person with, you know, agents in the humans verifying it, where in this case, you know, you have different functions, and there are like a few different humans that are, it kind of comes down to like, you know, what are the different functions and what are different roles? Like, do we think there's some functional difference between design, frontend, you know, like, engineering more broadly? I think there is, and in that world, you, I think, at least in the short term, will have this kind of like trio that we've, we've had for a long time, which is kind of like, you know, product manager, product designer, engineer, each calm with the very different lens of problem solving to the same problem, and, you know, maybe I'll be communicating with the engineer's agent or the product manager's agent, but still like, there's a human in the loop there making those decisions, like bringing that kind of framework to bear on the things that you're building. Are there other philosophies that you hold that, in fact, it's the way that you show up as a leader for the team? I'm not big on like business frameworks necessarily, but I do think things that I really, really care about culture and I really care about values, particularly given the kinds of products that I'm lucky enough to work on at a front-to-alab, right? You know, products reach potentially billions of people that bringing, you know, what I find to be just the most magical technology to the world and being able to solve all sorts of incredible problems and, you know, push forward scientific progress. So I feel a real humility in getting to work on this and I try to hire and build teams that also fill that. So that's something that's really important to me. Again, it comes down to being kind of obsessed with spending time with users and understanding how the decisions that we take impact those users, but also just like a humility about where we're going and, you know, we get to see a snapshot of that internally with the kind of tools and models that we test. And so trying to build that, and a few other maybe again, kind of like more values-based things that I really believe in is high-trust high-agency teams and they kind of can't really go without each other. So in order to be high-agency and do the things that I describe, which is like getting at product designers like, you know, maybe bringing like full features to bear and like feeling and power that they can do that, you can kind of only do that if you've built trust. Practically that just means like being really close with the people who are key to these decisions internally, which is often research, also engineering. When it kind of comes back to the humility part, like, you know, being a really great, you know, cross-functional partner, and you know, once you kind of build trust with folks and trust in your own kind of skills and instincts and kind of vision of the future, and that's a line with other people and the mission, then you can be really high-agency and maybe kind of push forward some of these bigger ideas. So yeah, less frameworks and probably more values-based. Are there specific ways that you screen for humility when you're assessing a designer and considering whether or not they're a big fit for the team? I definitely ask about mission a lot, right? So, you know, I think a lot of designers have kind of come into the AI space from design. I had a bit of a-- backwards process. Maybe I came in from AI in a way like my early career worked on AI policy and then went to grad school, kind of studied ML and design, worked in robotics and then worked in a bunch of root, very like research focused products and projects including building some of the evils that were used for 40 and others. So having come from that direction and also having worked maybe on like projects that have more of a like public purpose, I just ask a lot of questions about mission like why you know why do you work here, why do you work here other places like why do you want to work on AI, you know is it because it's you know the hot thing or like do you do actually care about the technology, how it impacts people and specifically our mission you know which is to kind of build a UI that benefits all of humanity. So it's many around mission and just you know the way that the folks present work presenting it collaboratively kind of low ego working really well with cross-functional partners, you can kind of you can often get it I think like digging in around values is a really good way to do that. You've mentioned research in the role that that plays in like showing up as a designer at opening a few different times and I'm wondering if there's more we could dig into there because a lot of people listening don't get to work at a frontier lab and maybe it's a little bit more of a black box myself included like how do you effectively collaborate with research, what's the role that that part of the org plays in terms of how ideas are shaped and where design fits in. So can you talk a little bit about that piece? Yeah so it's interesting like OpenAI is a very kind of grassroots kind of bottoms-up culture so there aren't necessarily like kind of processes or a playbook, it comes down to kind of relationships and you know working partnerships. So I think one of the one of the reasons why codex was so successful, particularly in the early days was very small team and it was product working like very very close with research. So physically I would sit across from you know the researchers training the models like an opposite me and maybe I had a question about some behavior or capability that was interesting building something I just like literally just asked them right. Part of it is that you know I also you know been opening up two years and when I first joined worked on the research team building a bunch of interfaces to train our models so I spent most of my time kind of within the research org. So big part of it is just like is relationships and partnerships and you know what that actually means in practice is if you chat to folks you again you see that people kind of approach the world in a different way and they have a certain framework and certain kind of beliefs maybe that they hold or assumptions that they hold. Another way of framing this question is like you know how do designs work research or how can I maybe like do more and it's kind of kind of like the JF Quake Quote Quote is like you know asked not what your country can do for you and and I think it's the holding on designers as well if you're interested in this area to throw yourself into it right. So you know within our org we have talks that you can go to right. If you want like you can read up on the textbooks that are kind of behind a bunch of this like MR research or you can listen to you know podcasts or if you're less technical you can you can kind of like bring up your knowledge. So I think it's also kind of beholden on design and product frankly to like get up to speed on the latest thinking and it's you know it's very easy to do that internally externally you can do you know again via kind of like a lot of a lot of public records and then from there you know you can start to build some intuition about where we might be going what are the particular things that people are interested in you know and then if you again imagine that world and sketch it out and think okay well what is the what's the shape of the product that I'm building that would you know work in that world or what are new products that you can build to really kind of showcase the capabilities of why we want to go. I know you have a small team and you've talked before about wanting to you know maintain that that size and being nimble and high-agency and ownership and all that but let's say that somebody's listening and they're inspired by everything that you're saying and they think that they you know have what it takes to thrive in this environment they want to get themselves in front of you so you mentioned the humility piece like what else would you be looking for in a designer today and what are some of the signals that they could put forward to get you to the point where like alright yeah I should definitely consider that person for a role. Yeah I mean first off people should feel free to reach out to me at any point always happy to chat you know always really really interested in in hearing from folks interested in the space particularly people maybe who have kind of a real interest or background in research but it will sit what's in design I feel like there aren't that many folks in that space and just always really really generally curious to to meet folks like that but yeah I mean you know when I'm kind of looking at candidates or portfolios I think particularly now one thing I'm just as interested in as kind of professional experience is what are you doing in the weekends and it's such an exciting time to build and you know even now like I don't really have much time on kind of evenings and weekends but when I can I'm still you know prototyping and I think I really look for that energy like people who are just really excited about this new era of building and kind of almost can't control themselves and a building stuff in a spare time so things are really tough time for like you know designers early in their career so you know maybe you haven't worked on as many projects at a company or elsewhere but you can still work on things in your spare time and communicate that maybe there are like open source projects you can work on or you can publish work you know online on Twitter or whatever to kind of try to try to raise profile and then yeah I think you know I think there are two profiles that I probably look for the most like one is just it's the kind of classic designer unicorn trope of just like you know super experience high agency like really kind of product focused and humble in their approach and folks that you can just kind of basically trust to get to go off and build amazing things and you don't really need to be super involved in the day today and those folks often do have a little bit of experience and then you know I'm really interested in like people at the beginning of their career I think people who particularly have grown up and they're kind of like native to this new world like a lot of designers have basically been through this process where they've worked for a certain way for like 10-50 years or longer and now they kind of you know discovering this new web working where as a whole new generation of designers as well who this is the norm and I think working a very different way in some ways like maybe aren't constrained by you know some of the things some of the ways that designers have worked in the past so I'm also really really excited to to work with folks younger in their career and you know some of the people that we've hired in that space some of the best people have ever worked with just like really high energy you know really prototyping first and really like get the technology in a way that I think a lot of folks don't are there behaviors outside of prototyping first that you think exists mostly in designers who are younger in their career and native to this way of working like was that even look like to you maybe you're like less bogged down by you know all of the the things that you kind of like you know accumulate over time to different jobs or just in life like I think just like this real kind of like unbounded ambition I think I think I think is is really interesting like not really seeing or understanding the limits of what's possible I think is is a really exciting profile but I do also think that kind of you know a lot of these a lot of these values are like not specific to like maybe junior designers or all the designers again it's as I talked about it's these kind of like low ego high trust high agency folks who you know maybe maybe you're earlier in your career and you have like less experience but you can clearly see like a very strong trajectory it makes sense like I think even for my own practice often I'll be surprised that something works and it's because I have this lens that I wear of an understanding of some of the constraints like I can't do that you know or I that's not going to be possible or feasible whereas the people that are coming up now literally think that they could do everything on their computer and there's no reason nothing stop in them you know and honestly they're right one of the biggest barriers of of folks like using tools like codecs is this like limiting factor of like oh it can't do that but every time at least I've kind of like pushed myself you know really expands the scope of what is possible you know I've seen designers commit like back end code and all sorts of stuff that would just be like unimaginable a few years ago yeah this has been awesome I appreciate you coming on today and sharing a little behind the scenes some of the prototypes and talking about how you all work you guys are kind of setting the bar for a lot of the industry in so many ways so really appreciate you taking the time and sharing with us today

Podcast Summary

Key Points:

  1. Designers are now building in a new era of AI, where imagination is the only real limit and prototyping for six months into the future has become a core design philosophy.
  2. OpenAI shifted from IDE-style tools to agent-first products like Codex, prioritizing a command-center interface that works across multiple tasks and is accessible to non-developers.
  3. The transition was initially controversial but gained momentum after significant improvements in model capabilities, especially around December, leading to rapid adoption across developers and non-technical users.
  4. Design teams focus on user personas rather than narrow features, ensuring broad lifecycle thinking and deep user empathy through regular onboarding and feedback sessions.
  5. Prototyping now includes "vibe-driven" experiments in code, such as live branches and dynamic interactions, tested in real environments with feedback loops to guide evolution.
  6. Designers collaborate closely with research and engineering, using evaluation-driven methods to test prototypes at scale and refine outputs based on real user behavior.
  7. Tools like Framer and Jitter are enabling designers to build full AI-powered workflows, from responsive design to interactive 3D visualizations, without needing to code.
  8. The role of design has evolved from final polish to vision-building, emphasizing communication, humility, and cross-functional partnership in shaping the future of AI products.

Summary:

The conversation with Ed Bayes highlights a transformative shift in design practices at OpenAI, driven by rapid AI advancements. Designers are now prototyping for a future six months ahead, focusing on agent-first interfaces that prioritize accessibility and user control. The shift from traditional IDEs to agent-based command centers—like Codex—has been pivotal, especially after model capabilities improved significantly in late 2023.

This evolution enabled broader adoption among developers and non-technical users, who now interact with AI through intuitive, familiar surfaces similar to chat interfaces. Design teams operate with a persona-driven approach, ensuring products serve diverse users from developers to knowledge workers. Prototyping has become more dynamic, with designers creating live, editable branches in code to test real-time interactions, often in collaboration with research and engineering.

These experiments are evaluated through iterative feedback loops, balancing model-generated outputs with editorial guidance to maintain control and quality. Tools like Framer and Jitter empower designers to build rich, interactive experiences without deep coding. Crucially, the role of design has expanded beyond aesthetics to vision-building—designers now lead by communicating ideas, fostering trust, and working with cross-functional teams.

The emphasis is on humility, user empathy, and collaboration, with design being a foundational layer in shaping accessible, future-ready AI products. This shift reflects a broader cultural transformation: designers are no longer just building interfaces, but creating new possibilities for how people interact with technology.

FAQs

It means anticipating future advancements in AI capabilities and building products based on how they might evolve in six months, rather than reacting to current limitations. This forward-thinking approach helps create more future-ready interfaces and experiences.

OpenAI moved toward agent-first products by trusting in future AI capabilities and designing for a world where agents would be the primary interface. This shift prioritized a command center model for managing multiple tasks, reducing reliance on traditional code editing in IDEs.

Key challenges include building trust, ensuring user control over actions (like sending emails or editing files), and educating users on new paradigms of interaction that go beyond traditional software workflows.

Designers work closely with researchers through direct relationships and shared understanding, often sitting next to them during model training. This collaboration allows designers to grasp technical capabilities and shape products that align with emerging AI behaviors.

Prototyping has evolved to include real-time, generative designs where AI creates interactive UIs. Designers use tools like Figma and live branches to test ideas, validate interactions, and move quickly from concept to prototype without needing full design systems.

Designers use a hybrid approach—giving AI freedom to generate content while applying skill-based prompts and golden prompts to guide outputs. This allows for innovation while maintaining consistency and usability in early prototypes.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.