Speaker 1The team at OpenAI said that one of the big changes that had happened internally was that the power of the latest generation of models, like Astra, had increased their speed of development so significantly that many things that they thought were only going to come in 2027 were actually coming as part of this Dev Day announcement. The result of that was more than 20 different launches and announcements, including some big headliners like OpenAI's answer to Muse and Grokbot, and some sleeper hits like the fact that enterprise accounts can now use OpenAI credits on an open marketplace to buy access to open source models as well. Overall, what we got at OpenAI Dev Day does not change the big patterns and trends that we've been seeing in the industry. A move to more cost-efficient models, those models moving to more persistent work, and some amount of that persistent work moving from a solo to a multiplayer experience. Instead, what Dev Day reinforced is that these are the trends to pay attention to and the ones that will reshape how we all use AI in the months to come. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Robots & Pencils, and HyperAgent. To get an ad-free version of the show, go to patreon.com slash ai daily brief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at ai daily brief.ai. There are also all sorts of other goodies at ai daily brief.ai. On our website, you can get access to all of our free AI daily briefs. If you're interested in learning more about AI daily briefs, those include free multi-week self-directed programs like the Multiplayer AI Sprint, which is live right now, but also free live webinars like the one that is coming up on Thursday, October 1st at noon, that is all about building your personal AI benchmark. You can also find links to our paid trainings, such as the Super Intelligent Executive Catch-Up and the Super Intelligent Executive Agent Leadership Program, both of which are registering new cohorts that start over the next couple of weeks. Now, last note on today's episode, originally I was going to try to jam in a full accounting of everything that happened at the yesterday on top of OpenAI Dev Day, but there was just too much to fit. So I've decided to go all Dev Day recap today, and then we will dig much deeper into what came out of Trump's meeting with the leaders of all the major frontier labs, and if and how it changes anything about the future development of AI. For now though, let's get into Dev Day and find out what it says, not just about OpenAI, but about where we are with AI in general. It was an absolutely monster day of announcements from OpenAI. In fact, there were so many that on today's episode, we are just going to get through a lot of them. So let's get started. All of the big ones, discussing what the announcement was, whether it was expected or not, people's first impressions, and how it shifts the AI race, if at all. Later on in the week, we'll go deeper on some of the most important ones, but for now, it is going to take all that we have just to get through it. First up is DOTS, OpenAI's answer to Muse, Grokbot, and the wave of personal agents that are taking over the AI space. Sal Maltman presented DOTS as quote, remarkably capable, always-on agents built to handle really anything you can think of. They bring AI to a new form factor. With DOTS, users can create a persistent agent, communicate with it through a text message style interface, and then set them to work on tasks using their own cloud computer with support for 40,000 different apps. DOTS can also communicate via voice call or across Microsoft Teams and Slack, with persistent context carried over between services and sessions. Functionally, it shares a lot with Muse, right down to the cutesy mascot. However, at least at the time of release, there are a few caveats worth noting. For now, users can only create one DOT at a time. So if you're a user, you can create one DOT at a time, and if you're a user, you can create one DOT at a time. But OpenAI plans to expand to teams of agents in the future. The other big difference from Muse is that OpenAI is only offering DOTS to pro business and enterprise customers. One of the big reasons Muse has been so successful is making it completely free, meaning OpenAI is naturally limiting their ability to compete on that front. CFO Sarah Fryer said that the vision is to bring DOTS to the whole consumer base, but for now, this is a prosumer product. OpenAI's big selling point is that they're using GPT-6 Astra to power the agents, so this is the first time we're seeing a truly frontier model pilot a first-party personal agent. Now, this was in many ways the least surprising of the announcements. This personal agent form factor has been basically inevitable since the launch of OpenClaw, and even the interest in business-focused versions like Grokbot, OpenAI getting into this particular area was pretty much inevitable. Now, how ready for primetime DOTS is remains to be seen. The live demo had some issues, although I would put way less in that than the chattering classes on X. Demos going wrong is basically just a way for you to know that they are actually live. Unfortunately, once people start using it, they're not going to be able to use it. So, once people got their hands on DOTS, they ran into a number of other teething problems as well. Chase Browser posted a session where DOTS lost all of its work, and Tenebra said that they were excited to try it out, but found that they couldn't connect multiple computers at once to share their full context. Others had a better experience. Analyst Max Weinbach wrote, My basic test for how well one of these works is can it autonomously do my expenses with access to my email and Google Drive? Grokbot made a mess and didn't do what I told it. It did the first like three days and then started to mess up. DOTS did it after the first time and has been good. Justin Schroeder said that he thought that DOTS was even more work-focused than Grokbot or Muse. Think a bit less personal assistant, he writes, and a bit more codex orchestrator and Slack collaborator. In fact, Justin writes, Slack is where it really shines. Your team can message your bot directly and your bot can take actions, like spin up new panes. Also, getting at a debate, which I think will be everywhere, Justin suggests that he thinks that this should have been a separate app. ChatGPT is great, he writes. Codex is great. They are very different products, in my opinion, and DOTS are way more codex than ChatGPT. The every vibe check almost leaves it as a TBD, with their users reporting flashes, where it was truly excellent, but then other times where it was just extremely frustrating. When Brandon Chu suggested that this was the form factor that everyone was converging on, Nate B. Jones wrote, I think we need to distinguish between form factor and utility here. Yes, form factor is converging at the moment, that may change, but utility is not. Muse is good at specific things, like phone calls and practical email work. Instinct is good at specific different things, like travel. DOTS is good at AI context as a work surface. Regardless of positioning, that's distinct utility. The question with a T for trillions is, did you pick the right utility to get right? We'll come back to DOTS later in the week, but I think the big takeaway is continue to watch this space. While acknowledging that it's a bit buggy right now, Allie K. Miller argues that there are some big differences here that are fairly significant. Things like each DOT getting its own dedicated virtual machine, it being always on, which allows it to be more proactive, and some other changes that are worth watching. In any case, number two on the list is the Decisions API, which is OpenAI's answer to Jev. Now remember, Jev is a judgment model. It's not good at outputting text, it's good at classifying things, giving confidence scores between zero and one. Making judgments, in other words, which are the precursors to decisions, which presumably is where this feature got its name. The Decisions API allows users to call a version of Luna that mimics Jev's quick classification and decision-making abilities. Users define a list of questions and possible answers, and the model outputs fast responses. OpenAI said that the API could be used for things like classification, routing requests, or choosing an agent's next action. They claimed 10 times faster decision-making compared to the Responses API. Still, one of the big questions was, how is this better than just using Jev? One difference in actual use cases comes from Luna's support for visual inputs, which aren't possible with Jev. Over the past couple of weeks of Jevmania, many users have shown off Jev making rapid classification for things like visual marketing, but that requires an image-to-text transformation under the hood. The Decisions API removes that step, and in that way expands its setup. Now, right now, the Decisions API is just in preview for testing with a limited group. And in terms of significance, I think more than anything, this is a recognition that what Jev represents is more than just a new model. It is an extremely useful, and dare I might say, soon-to-be-fundamental primitive to have in the ecosystem. Part of why Jev has hit wasn't that it was flashy or sexy. It's just that as soon as you see all the things that it does better than a traditional LLM, in fact, where you see that LLMs were being rammed like square pegs into round holes into use cases that they weren't great at, it just seems obvious that having that sort of judgment model to sit alongside your generative model is pretty obvious in retrospect. Once again, it's too early to really do comparisons, but I tend to think that this is a space where there is room for more than one model available. In fact, I think that pretty much every frontier lab will have some version of this available very, very soon. Next up on our list is OpenAI Space, their new shared document workspace that gives human teams and agents a place to collaborate. You can think of the like an AI-enhanced Google Drive. Teams can work on shared spreadsheets, slide decks, and other documents, but space also gives teams a place to house workplace automations. Dots can natively work on documents in space, but teams can also set up scheduled tasks to produce a deliverable in their space. Many focused on OpenAI going after Microsoft 365, Google Drive, or Notion with space, and while this is certainly generally in those space, each of those services have felt increasingly like a bad fit for agentic work, and it was only a matter of time before OpenAI launched a truly AI-native productivity suite. Now, whereas almost every other announcement had a pretty wide diversity of opinions, space was one where especially the power users were in love right away. Ray Fernando wrote, I've had early access to OpenAI Dots and I'm not here to join in on the GlazeFest. Is this a Grokbot or Hermes killer right now? No. But space is, and allowing agents to thrive in apps feels like the right directions for these agents. How IAI's Claire Vo wrote, in my opinion, Dots slightly overhyped and space underhyped. Every company I know wants an AI-native collaboration workspace. We'll be watching to see if this pulls more enterprises to the OpenAI ecosystem. Dan Shipper from Every wrote, the obvious win is that you stop switching windows between ChatGPT and another app like Notion while you write. The less obvious one is speed. When a Dot builds a document through its browser in Google Docs, the edits practically crawl in, because Docs wasn't built for agents. Native documents don't have that lag. You can tag your Dot inside the document itself instead of going back to the chat and tell it to check the file. Dan concludes, I've long expected the company to do this. In 2024, I wrote about the potential for document slides and sheets inside ChatGPT, and I'm glad the product is finally here. I'm already doing most of my work in ChatGPT's in-app browser, and and this makes that process smoother. I can tag my dot in a comment on a document and get a revision in line as I'm working. The back and forth feels more like collaboration. The big model launch for the day was GPT-61 Sol, which OpenAI pitched as near-Astra Intelligence for a fifth of the price. The release comes just a week after GPT-6 Sol and provides some fairly notable upgrades. Coding benchmarks are up significantly across all effort levels, meaning that 6.1 Sol on medium settings outperforms 6 Sol on max settings. The model's top score on DeepSwee was 75.2% on high settings, with extra high and max actually seeing a degradation in performance. We first saw this phenomenon with Opus 5, and it looks like top effort settings are starting to force models to overthink and second-guess correct responses to their detriment. Now, that high setting score on DeepSwee was actually a touch higher than Astra's best performance, supporting OpenAI's claim of near-Astra Intelligence. The pattern was similar across most major benchmarks, a big improvement over 6 Sol that landed 6.1 Sol in the same ballpark as Astra at a much cheaper price. One of the notable benchmarks was OS World, which tests long-horizon computer use. 6.1 Sol's best performance scored 71.4% on max settings, beating 6 Sol at 64.4% and coming close to Astra's best at 73.5%. 6.1 Sol was also significantly cheaper than the others, at a third the cost of 6 Sol and 13% the cost of Astra. What's more, increases in the effort level barely changed the cost, suggesting that OpenAI has made some big breakthroughs in the efficiency of computer use with this model. Artificial analysis scored the model at 52, one point shy of Astra and also behind Opus, Sonnet, and Fable, coming in in fifth place. They also found incredible cost efficiency that pushes the Pareto frontier, with 6.1 Sol's benchmark run completed at a quarter the cost of Astra and 31% cheaper than 6 Sol. If you don't need to run it at max settings, 6.1 Sol seems capable of hitting smaller costs like GLM53 Flash, which is the cheaper version of ZAI's open source model. OpenAI called it the most cost-efficient model for its performance available today. Now, as to whether this one was expected or unexpected, on the one hand, it was a very good model, but on the other hand, it was a very good model. On the one hand, it's never all that surprising to get a new model from one of these labs, especially on a big day like Dev Day. Surprising in that we only got 6 Sol last week. Still, when it comes to first impressions, let's just say we're going to have to wait to get our hands on it ourselves. Because for every post you can find like this one from Dropout Layer on X, GPT-6-1 Sol just made Opus 5-5 look expensive. You also get one like this one from BridgeBranch. We ran it through our Sunset Ocean test, same cost as GPT-6 Sol, twice as slow, and it barely rendered an ocean. Hello, everyone. One big change around AI is we've shifted our thinking from how we rank our pages to how do we become the source that AI trusts enough to answer with. At KPMG, they're seeing this firsthand. AI-generated results now surface answers directly, often without a single click. That's why they are increasingly focused on Generative Engine Optimization, or GEO, structuring content so AI systems can retrieve it, understand it, and cite it as trusted authority. This is not just an SEO evolution, but a visibility mission. And indeed, the GEO mandate from KPMG is simple. If AI is shaping decisions, your expertise needs to show up inside the answer. Read all about it at kpmg.com slash us slash GEO. Again, that is kpmg.com slash us slash GEO. Here's why most legacy modernization projects fail. The AI doing the work can't understand codebases at scale. It sees a small slice of context, examines syntax, and misses years of decisions distributed across the globe. This is why most legacy modernization projects fail. Blitzy solves this the way it solves everything. Grounded in your code before any migration begins, Blitzy's agents reverse-engineer the entire legacy system into a persistent knowledge graph. Every dependency, every constraint, every piece of tribal knowledge that used to live in one engineer's head. From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith-to-microservices transformations, all validated end-to-end. One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137-week baseline with coding agents. That's 9x compression. Retire technical debt while accelerating your roadmap. See how at blitzy.com. That's B-L-I-T-Z-Y dot com. The best teams don't have a single star carrying everyone else. They know their own strengths and each other's weaknesses and play to both. That's the team Robots and Pencils has built on purpose. Nobody there is grinding through busywork to pad a headcount number. People come for the hard problems and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. Check out robotsandpencils.com/careers. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you had agents that feel like teammates. Hire yours at HyperAgent. Get $100 in credits at hyperagent.com/AIDailyBrief. Now, hopefully these sole class models are good enough, though, because we may never see GPT-61 Astra. The Wall Street Journal reports that OpenAI have scrapped plans to release the next version of their flagship model over safety concerns. Sachi Jain, OpenAI's head of safety systems, said in a statement, "For anything regarding safety and alignment, there's a trade-off. You really do need to find what's the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks, even when it hits friction." It sounds like OpenAI tried to turn down the tenaciousness that had led to multiple incidents over recent months, but couldn't find a happy medium. Jain added that although the model was less lazy than its predecessors, it "didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done." Now, obviously, I'm being a bit hyperbolic when I say that we'll never see 61 Astra. But if you're interested in learning more about OpenAI, check it out. But it actually does seem like they're going to need to hit some technical breakthroughs before we get to some of the next levels of the frontier in a way that they deem safe enough to release. Now, from here, we get into what are considered the smaller announcements from the event, many of which people are still noting as quietly significant. OpenAI showed us the latest version of its platform plans, opening ChatGPT to outside developers. Tebow from the OpenAI team wrote, "You can now build full native apps with plugin extensions and ship them right in ChatGPT. We have over 1.2 billion weekly users, and we'll surface relevant plugins right in the conversations. This feature is technically called plugin extensions." Now, OpenAI has been trying to get the balance right on this feature for more than a year, testing native integrations, plugins, and MCP as a way to expand the ecosystem. So could this finally be the feature that lets ChatGPT function as an app store, which always seemed like one of the big goals for OpenAI? Certainly, there is going to be a rush to experiment here. I'm already seeing things like this Personal Stylist app from Yana Wellander, but others are wondering if the trade-offs for the app developer are really worth it. MIT's Christian Catalini writes, "You bring the app, OpenAI brings the intelligence. Does OpenAI also get the user traces? An interesting partnership if your contribution is teaching your partner how to do your job." To which Signal responded, "Long tail of apps might opt in, but otherwise this is a terrible idea for anyone else." A flip side of this is sign in with ChatGPT, which is exactly what it sounds like. It allows you to use ChatGPT to sign into other accounts. OpenAI's head of applied research, Boris Power, wrote, "Sign in with ChatGPT lets you use any app built with our API. Makes it much better for app developers not needing to jump a huge hill justifying paying a separate subscription. ChatGPT is your one-stop super intelligence juice." The idea is basically that developers allow people to carry their intelligence subscription with them, lowering the barrier to entry for their apps, and allowing them to take advantage of the fact that so many people have already opted in to using ChatGPT for their super intelligence solution. Jackie Luo writes, "Sign in with ChatGPT is a huge deal, and I don't understand why they framed it as auth when it's really OpenAI expanding favored pricing to third parties and bundling services into the subscription. It finally starts to align their incentives with customers so that individuals stop double paying for tokens and apps stop having to price everything on top of API costs. I'd imagine that expands to many more, if not all apps, and Anthropic will need to follow to add value to their subscription. In that world, apps could bypass token costs and charge for the app layer alone again, which is probably good for everyone." In other words, the idea here is that instead of apps having to charge what looks like huge prices to get access to intelligence, they can just charge their $20 or whatever they wanted to price their app experience at and let the API costs flow directly to the underlying. While yes, that means they might not get to scalp those costs and add a little premium to them, the benefits of not having to convince people to pay those additional fees when they're already paying for a ChatGPT subscription likely outweighs the money that they could make otherwise. For my power users out there, OpenAI has introduced a few new options for their subscriptions. The first is UltraFast Mode, which offers 8x faster token production in Codex and 6x speed in the app. Right now, it's only available on Android, but it's also available on Android.com, but now UltraFast Mode is only available for Astra and Codex and ChatGPT work, but OpenAI say they will add support for 6.1 Sol soon. In addition, OpenAI is introducing a new $500 subscription tier that offers 25x the usage of the Plus tier. It's also the only tier that gets access to UltraFast Mode for the time being. Heading into the event, Tebow announced that OpenAI would reopen access to their $200 a month Pro tier, which they had turned off a few weeks ago, but with a few tweaks. Usage calculations have been changed, netting out to a 50% reduction in terms of API cost. Tebow explained that this was the best of a bad set of choices. OpenAI prioritized not having to introduce the 5-hour usage limit and argued that API cost reductions in more efficient models would mean users can still get roughly the same amount of work done. Now, on the one hand, people were sort of expecting something like this, but at the same time, you can imagine how well it went over to have the same model coming back with reduced value. For our purposes here of trying to understand what it says about the state of AI, look man, compute constraints are real, present, and permanent. Frontier models are running up the walls of what they can do with the compute have, and new compute isn't coming online at anywhere near the speed that people are increasing their use of this digital intelligence. The upside of that, though, is that we're going to see a big push around efficiency, which in the long run should net out to more cost-effective, better experiences for all of us, but it's going to have some bumps along the way. Over in developer and vibe coder land, Codex now has a dedicated cloud environment, allowing users to keep working after they close their laptop. This one was always coming after people walking around with their thumbs jammed into laptops, became so common across San Francisco that it became a meme. But looking for broader patterns, starting with Grokbot, we've seen the entire industry move towards a cloud instance being a necessary part of all agentic products. This is a continuation in the big shift from AI being something you engage with, like software, to becoming a more persistent, always-on application, regardless of how you access it. The Codex CLI is also getting a major refresh with a new look and new capabilities. History goes back further as you scroll, supporting the monothread maxis, of which I am absolutely one, while the composer stands out. The codex CLI now gives you a better way to manage parallel work. Use slash agents to see what's running, check progress, and jump between tasks. On the enterprise side, OpenAI has launched private intelligence, which guarantees zero data retention even at inference time. Over the past few months, privacy, which was always a huge issue for enterprises, has become a huge issue for the companies supplying the enterprises, with growing concerns that OpenAI and Anthropic are skimming data from their customers. This is basically the subtext or the main text of every communication from Microsoft these days about why you should be not trusting their competitors. All that means that a stronger, clearer, end-to-end data privacy guarantee is a big deal for those who need it. OpenAI has also launched a model marketplace. It allows users to buy open-weight model inference from base 10 through OpenAI's responses API in Codex. Now this is one that on another day we could go way deep into because it has some big implications for OpenAI's moat. It protects them from open source disruption and gives customers a feel comfortable making large spending commitments with OpenAI, knowing that they can easily use that spend on a multi-model strategy that includes open-weight models. Trust me when I say that we are going to come back to that one because I think it is bigger than people are giving it credit for at first blush. So we have barely scratched the surface here, but let's try to quickly sum up what all of this amounts to and what OpenAI Dev Day revealed about the next phase of AI. We didn't get something blisteringly new. What we got was confirmation of a lot of the trends that we've been cataloging on this show over the past several months. First, with Sol 6.1 and the Decisions API, we got a continuation and a deepening of the trend of models that are good enough and cheap enough that they can be used to do everything. I.e. you don't have to turn them off, you don't have to make decisions about what they're used for, they can just do it all. With DOTS, we have persistent and proactive agents that take advantage of that to actually do everything. And when it comes to doing everything, plugin extensions shows that that means bringing everything in, as well as going everywhere, which is ChatGPT's sign-on for other apps. Finally, with SPACE, we have more confirmation that the next generation of AI will not just be single-player mode, but will also be team and multiplayer mode in a native way, even if that remains nascent so far. Lots and lots more to explore here, but that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace! you