Go back

Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models

29m 0s

Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models

Google announced Gemini 4 Argon this week, its first frontier model in over six months, following a period where the company had fallen outside the top tier of AI labs. The model posts impressive benchmarks, achieving state-of-the-art scores on agentic knowledge work, VALS index, legal benchmarks, and DeepSuite coding, while trailing slightly on Frontier Suite and Terminal Bench. Artificial Analysis placed it tied for third on its Intelligence Index. However, Google is not releasing the model publicly, citing cybersecurity concerns, instead limiting access to trusted cyber defenders with plans for gradual expansion. Reactions were mixed. Many celebrated Google's return to the frontier, while skeptics noted Google's history of strong self-reported benchmarks that disappoint in practice. Bloomberg reported internal employee skepticism about real-world coding performance, which Google denied. The episode also covered Anthropic's Sonnet 5.5, a fast, cheaper model with strong benchmarks but token-hungry behavior, and Muse's rapid growth to 3 million weekly active users, raising questions about whether superior UX can compensate for a weaker underlying model. The broader takeaway is that competition now extends beyond models to harnesses, agents, and product experience.

Transcription

5870 Words, 34350 Characters

English
Speaker 1Coming into 2026, Google was looking pretty good in the AI race. 2025 had been a good year. A lot of the questions inside DeepMind had been answered. They were putting out competitive Gemini models. They were pushing forward with interesting new products. And all the natural advantages that they had always had in terms of consumer distribution and data and all those things had a lot of people very bullish on them as a contender. But then vibe coding happened. And with the increase in coding capabilities paired with the power of the new harnesses like Cloud Code and Codex, those new capabilities unlocked agents in a way that hadn't been possible before. And all of a sudden, Google found itself very behind. Charitably, you would say that in 2026, the company has been playing catch-up, but even that's not really accurate to what's been happening. For most of this year, Google has been firmly outside of the conversation as a top model lab. And yet I think that those who didn't have a partisan bias towards one of the other labs would never be fully comfortable writing Google off. This week, the company announced Gemini 4, their first new frontier model in more than six months. By the benchmark, it looks like Google is so back. But is that the whole story? So let's dig in to Gemini 4 Argon and where it lands in the current AI race. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Robots and Pencils, Harbor, and Granola. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. To learn more about sponsoring the show, send us a note at sponsors at ai-daily-brief.ai. Also note that we have our next cohorts of superintelligence agent training coming up. These are paid programs, which in very short order will get you far ahead when it comes to your agentic understanding and your ability to use agents in your daily work. We have both the Executive Catch-Up Program and the Agent Intensive, which we call the Executive Agent Leadership Program. The next cohorts for those start next week, and you can find links to all of that at the very top of ai-daily-brief.com. President Trump really looked at OpenAI Dev Day and said, nope, absolutely not. I don't want that to be the biggest thing happening in AI this week, and invited basically every big AI CEO to the White House for what David Sachs would later call the Bretton Woods of AI. Now, there was a lot of chatter that came out of this meeting, but one of the first things that people noticed was that Trump pulled Anthropic CEO Dario Amadei to be the one to speak to the press following the meeting. Now, Dario insisted that he's just saying what he's always said, but he's not. He's just saying what he's always said. That AI has some incredible benefits and some very real risks. However, in an act which some saw as public deference, others saw as Dario finally playing the game, he commented, as the president has said, whoever wins AI wins. I think that's very important. The mechanism, how we address the risks, is still under discussion, but we all need to work together to make sure that we can win and we can win safely. If we do this right and work with the president and everyone here, we can all win safely. Now, X was absolutely filled with people psychoanalyzing the moment, breaking down the truth, breaking down the expressions from other tech leaders, Dario's anxious tics, and even dissecting his fashion sense. While some thought that Trump and the other tech leaders might be bullying Dario as he stepped forward to speak, the president followed up with some kind words later in the press conference, saying, Dario has been fantastic. We had dinner the other night and he agrees with everyone. Whether or not this was Dario truly bending the knee, it was a significant moment for a CEO who previously told staff that Anthropic had been targeted for refusing to give, quote, dictator-style praise to Trump. At least in the immediate wake, the press conference completely overshadowed the subject of the meeting, which was the head of the Frontier Lab signing an accord on superintelligence. The accord is a one-page commitment to develop internal controls for Frontier model testing and deployment, as well as partnering with external auditors for verification of those controls. Trump told the press that the document is, morally binding, suggesting a continuation of the voluntary approach to AI safety testing. Trump added, the biggest people in the world signed that and I signed it as president and it really is a form of protection. He also floated the idea of a 10-member industry oversight committee during the press conference, but for now seems satisfied commenting, I'm seeing tremendous self-policing. The general tone, at least among the people at the meeting, was that this was an important first step. Mark Zuckerberg told the press, the idea isn't that this is the only thing we will ever do, it's that this is a start and an accord the whole industry could come to. Mostly the media response to this was a complete Rorschach test about how one feels about Trump and how one feels about the tech CEOs, i.e. whether you're inclined to use the word oligarchs to describe them. But I do think that these skeptical take on this one might be best summed up by media commentator Chuck Todd, who wrote, if you trust the tech companies to police themselves and you want them to decide how AI will impact society without any input from the public, then you'll be fine with this accord. Trying to find the AI moderate or AI realist position in this one, I think a person could agree that it's better that these folks are frequently in rooms together with key leaders in the government rather than exclusively viewing their role through the lens of competition, while also wanting to see more external involvement, not least of which in articulating of what the independent external auditor or evaluator to carry out independent assessments is actually going to look like in practice. Perhaps now that this accord has been signed, that particular detail is one that can get built out. Following the lunchtime meeting, President Trump signed a pair of executive orders marking the beginning of the superintelligence age. The first made Trump's name change official, stating, as these capabilities continue to improve, they increasingly represent not merely artificial intelligence, but a new era of superintelligence. The terminology used by the federal government should reflect the transformative capabilities of these technologies and the limitless opportunities they create for the American people. Much more functional was the second order, which established America.gov as an AI-powered single portal for government services. Trump said, with America.gov, the federal government no longer stands in your way, it only stands at your service. We're simplifying it. We're glamorizing it. We're making it what it should be. Alongside that second executive order, we got an Apple-style product unveiling keynote, complete with Secretary of State Marco Rubio playing the role of Steve Jobs. Rubio showed demos of planned future functionality, like being able to use the portal to apply for a passport, change your name after getting married, or enroll for Medicare. The features kind of seem like an MCP for government. A user can provide their information once, then the America.gov agent will track down and complete all the necessary forms across numerous government departments. How well this works remains to be seen, but certainly it is a real problem. The insane difficulty, for example, of changing your name after getting married, depending on which state you live in, is hard to overstate. Rubio said that the function will be available from next year, with the executive order giving government departments 90 days to integrate their services into the portal. For now, America.gov only functions as a knowledge base for government services, searching across 29,000 government websites to answer questions based on official sources. Now, a lot of the chatter on this centered around security and privacy concerns, particularly after the president said it cannot, in theory, be hacked into. And when they figure out a way to do that, we'll end it. Because they always figure out a way, right? The National Design Studio, which designed the website, came under fire in June for its work. including activity tracking tools and other websites they've designed. And while some raised questions about providing the personal information required to apply for government services, AI jailbreaker Pliny the Liberator poked and prodded at the chatbot and found its guardrails were pretty airtight. He wrote, Interesting. The America.gov chatbot will flag anything that appears to be personal information, like any text vaguely resembling an address or password, and refuse to accept the query until the flagged text is removed. Never seen that before. Meanwhile, if you think that them showing up together at this meeting means these companies are going to assume all-government scrutiny, it might be worth noting that the FTC has opened an investigation into OpenAI and Anthropix rogue agents. According to agency sources speaking with the New York Post, the FTC opened an investigation into the two leading AI labs a few weeks ago. The scope will cover the Hugging Face incident and the dozens of so-called rogue agent events that have been disclosed since. The investigation has advanced to the stage that the FTC is drafting civil investigation demands, which are similar to subpoenas. An FTC source said, The agency has plans to compel the executives at these firms to testify about their problems. and about the dangers they allege their products may have to consumers, to Americans. Reports state that third-party safety research lab meter can expect a demand alongside the frontier labs. Now, from very early on in the AI narrative, one big thread of discourse has been focused on the idea that new AI-specific regulation wasn't necessary. That government could simply enforce product safety legislation that's already on the books. Former FTC chair Lena Kahn was of this view, writing last month, Law enforcers already have authority to charge companies and their CEOs for creating and releasing dangerous, unvetted, or defective products. We shouldn't let discussions about new legal regimes distract from the fact that there's no AI exemption from laws already on the books. This post was co-signed by former AI czar David Sachs in yet another example of the AI safety debate creating very strange bedfellows. Still, based on comments from the FTC source, it's not clear that the agency is necessarily pushing for some sort of harsh penalty. They told the New York Post, We need to win this SI race, absolutely, and we are winning, and that's fantastic. Getting a political jab in, the source continued, and obviously the other side, the Democrat, wants to destroy this technology, wants to surrender to our enemies and competitors, and that's not something that the chairman's interested in doing, and that's not really something the president's interested in doing. With that being said, the laws have to be followed. People talk about how we need new laws for these companies, and the chairman's been very clear, we have plenty of laws on the books. So, an investigation, that's good, right? That'll appease people who think that the labs are getting a free pass? Eh, that's not exactly what the discussion suggests. Anti-monopoly activist Matt Stoller wrote, It's a protection racket. The idea is Trump investigates and clears them so other investigators have hurdles in their probes. Former director of public affairs at the FTC, Douglas Farrer, wrote, It would be good if the FTC investigated these companies properly, better starting two years ago, but the public credibility of this new investigation is 0.0%, especially after yesterday's AI CEO and Trump lovefest. The optimistic take comes from Joel Thayer, who writes, I kind of love this trust-but-verify strategy. Even though it's mostly self-regulatory, if any signatories fall short of these commitments, it could trigger FTC enforcement under Section 4 of its UDAP authority, makes FTC chains Chairman Andrew Ferguson's presence at the meeting very apt. As with anything policy right now, it's worth having a whole bowl of salt when you look at this. But things continue to move in some sort of direction. For now, that's going to do it for the headlines. Next up, the main episode. If you're leading AI inside an enterprise, you already know that the gap right now isn't capability but execution. That's why KPMG's You Can With AI is back with a new season featuring conversations with leaders like Sarojit Chatterjee of Emma, May Habib of Writer, McKesson CIO Ellery Fisher, and others focused on practical execution. What's working, what's not, and what it actually takes to move from pilots to real scaled impact across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathaniel Whittemore. Go listen and subscribe at www.kpmg.us slash AI podcasts. That's www.kpmg.us slash AI podcasts. The best teams don't have a single star carrying everyone else. They know their own strengths and each other's weaknesses and play to both. That's the team Robots and Pencils has built on purpose. Nobody there is grinding through busy work to pad a headcount number. People come for the hard problems and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. Check out robotsandpencils.com slash careers. If you listen to this show, you likely have a thesis. Maybe it's enterprise adoption. Maybe it's compute. Maybe it's a specialization. Harbor Capital's AI Lab Ecosystem ETFs let you express it via five actively managed ETFs, each seeking exposure to the ecosystem around one major lab. Anthropic, OpenAI, DeepMind, Meta, or SpaceX AI. Your view of the AI race in ETF form. Harbor Capital Advisors AI Lab Ecosystem ETF Suite gives investors a way to invest in the AI ecosystem they believe is best positioned for success. Search Harbor AI Lab Ecosystems ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information. Read and consider it carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by Foresight Fund Services, LLC. Harbor is not affiliated with AI Daily Brief and the funds are not affiliated with, sponsored by, or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principal. When I'm in a meeting, I'm fully in it. I'm thinking about the iteration and creative back and forth it takes to actually push a goal forward. What I'm not thinking about is capturing takeaways, tracking to-dos, or any of that. And that's where Granola comes in. Granola is an AI-powered notepad that captures what happens in your meetings and turns it into clean, structured notes, with the decisions and action items pulled out and easy to find. There's no setup and no configuration. It just fits into how you already work. For me, it means I get to stay in idea mode, and Granola makes sure those ideas actually become action. Once you try Granola on a first meeting, it is hard to go without. You can try it totally free. You don't have to pay for it. Granola is free. You can try it totally free at granola.ai slash brief. That's granola.ai slash brief to get your time back. Welcome back to the AI Daily Brief. And friends, it appears that pigs are flying. Hell has frozen over. Choose your metaphor for incredulity because after months of waiting, Google has announced Gemini 4. Although maybe those pigs aren't completely off the ground because although they announced the model, we don't actually have access to it yet. On Wednesday, Google announced Gemini 4. Google announced Gemini 4. Google unveiled Gemini 4 Argon. It's their first frontier model in over six months. Now, the 2026 struggles of Google have been well documented. We never got a Gemini 3.5 Pro, with the company basically deciding that it just wasn't good enough to release. And for a few weeks now, speculation has been coming that Gemini 4 was on the horizon. And Google DeepMind CEO Kare Kevesoglu says that the model delivers. Quote, Frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal engineering, and the like. Now, despite a long period where Google's future as a frontier lab has seemed somewhat tentative, to judge only by the reported benchmarks, Google looks to be right back in the game. Gemini 4 Argon is the new state-of-the-art model across many benchmarks. In agentic knowledge work, Gemini 4 scored 68.9% on the VALS index, beating Fable 5.1, GPT-6 Astra, and exceeding the score from the current leader Opus 5.5 by a couple of percentage points. The gap was even larger on Automation Bench and VALS Finance Agent. And on Harvey's Legal Agent Benchmark, Gemini 4 more than tripled the score of current leader Fable 5.1, scoring 19.6%. Agentic coding scores were a little more patchy, but Gemini 4 is definitely in the ballpark of frontier rivals. It's the new state-of-the-art on DeepSuite, scoring 77.9% compared to Opus 5.5's 74.2%. On Frontier Suite, Gemini 4 scored 55%, which placed it 10 points behind current leader Astra 6, 7 points behind Opus 5.5. And 1.3 points behind Fable 5.1. On Frontier Suite, however, Gemini 4 scored 55%, which placed it 10 points behind current leader Astra 6, 7 points behind Opus 5.5, and 1.3 points behind Fable 5.1. It was also bottom of the pack on Terminal Bench 4.0, with a score of 57.4%, 9 points behind current leader Opus 5.5. Now, you can look at this in two ways. One, agentic coding is one of the most important, if not the most important, use case, so failing to achieve a state-of-the-art score is kind of problematic. Or you can look at this as a huge improvement for Google, which it undeniably is. Coding had been Gemini's biggest weak point for the past year, so to have Gemini 4 be in the same ballpark as the Frontier models from OpenAI and Anthropic absolutely represents a big catch-up. For computer use, Gemini 4 is nipping on the heels of Astra, which set a new standard for the category. On OS World, it scored 69.2% against Astra's 72.6%. Artificial analysis confirmed that Google is back in the mix as a leading model developer. Gemini 4 Argon scored 53 on the Intelligence Index, putting it tied for third place with GPT-6 Astra and Fable 5.1, one point ahead of GPT-6 1 Sol, and three and five points behind Sonnet 5.5 and Opus 5.5, respectively. In terms of cost and efficiency, it's pretty close to the Pareto Frontier, costing $1.99 per task on the AA benchmark run, compared to $3.26 for Astra and $7.63 for Fable 5.1. However, that is with Google's 50% launch discount, which will be available for an unstated period, meaning that the full price would make it more expensive than Astra. It's also in that uncomfortable middle ground where it's more than twice as expensive as GPT-6 1 Sol, even with the discount, without a super clear improvement on the benchmarks. For the moment, however, none of this really matters, because unfortunately Google is not making Gemini 4 Argon generally available to the public. Their stated reasoning has to do with cybersecurity. The model scored 68% on the CWE Bench cybersecurity benchmark, tied with Grok 4.7 and GPT-6 Astra, and slightly ahead of Opus 5.5, Fable 5.1, and Mythos 5.1. This was reason enough, according to Google, to limit access to what they call a set of trusted cyber defenders to ensure the model is not misaligned. They said that they are engaging with the US government's voluntary testing and pre-release access program, and intend to gradually expand access over time, beginning with API and Google Ultra subscribers. In their blog post, Google said that the model is already powering internal workflows and boasting strong performance on long-horizon coding tasks like large-scale codebase migrations. Now, as to the choice to announce the model without actually releasing it, I'm not really sure. Honestly, for us old heads who have been watching this for a while, it kind of has echoes of the first time Google fell behind and felt the need back in December of 2023 to announce Gemini, even though the pro version of the model wouldn't be available for a number of months. Yet, maybe it was a good call because there were plenty of people who were excited to see it. There was an entire genre of posts on X that basically came down to Google, so back. Dr. Daria Anutmaz wrote: "Wow, wow, Google DeepMind is back with Gemini 4 Argon. Insane benchmarks. In one fell swoop, they've jumped to the very top of the AI frontier. As I've said before, never bet against Google in the age of AI." Although he did add, and this seems to me to be the key detail, "Can't wait to try Gemini 4 ASAP." Ethan Malek writes: "And it's a three-way race again." Nathan Lambert, who does open model research and is about as far away as you could be from a hypester, writes: "Love to see Google surprising people with Gemini 4. More labs at the frontier is wonderful for consumers because of competition and wonderful for the world because of reduction in concentration of power. Excited to see how it goes in real-world scenarios." AI leaker I Rule the World writes: "I've been critical of Google models, but Gemini 4 looks strong on benchmarks. We've seen this before with Google, but I'm optimistic they can release a first-class model with a first-class harness like Codex. Great work to all involved with Gemini 4. Rooting for you." But then on the other side of the coin is Angel, who writes: "I'm sorry, but I can't with all these 'Google is so' back posts. We literally can't even use the model, and it's not like benchmark scores for Google's models have ever been reliable. Do we really need to remember what happened with Gemini 3 Pro?" "Seriously?" Ishu Agrawal puts some numbers around that, sharing an older series of benchmarks and adding: "By the way, these were the official benchmarks for Gemini 3.1 Pro when it released, practically destroying Opus 4.6 across the board. I'm not saying Gemini 4 Argon will be bad, but it's crazy that no one has the slightest bit of skepticism given Google's track record." To their credit, Logan Kilpatrick, who honestly should get Google's MVP for hanging in there throughout all of these PR challenges this year, actually engaged with this, writing: "We've gotten much better at testing our models at scale across Google now. So assume most new Gemini revs go through thousands of software engineers for weeks before getting released. Hopefully it's helped close the benchmark to reality gap by a real margin." Now, so far, all we have to go on is the reported benchmarks, along with some reports from grumbling inside the company. A Bloomberg piece called "Google Grapples with Employee Skepticism About New Gemini 4" says: "While Gemini 4 has performed well on benchmarks widely used to gauge model efficiency, it does less well when employees actually put it to work. The model struggles to handle certain coding tasks, but it's still a great product." "It does a lot better when employees actually put it to work, but it's still a great product." according to people with direct access to the model. Google, for their part, denied the reporting, commenting that it would be "inaccurate" to claim that Gemini 4 is "underperforming" in some areas such as coding. And one of Bloomberg's other sources said that these gripers are a minority, claiming that a, quote, large consensus internally at the company believe that Gemini 4 is the frontier. What makes the position of Google so challenging is that in the time that they've been away from the top, the nature of the race has fundamentally changed. With much optimism, Peter Yang wrote, Google cooked on Gemini 4. Now they just need to be more competitive on coding harness, anti-gravity, and personal agent, Spark. In other words, the competition has become more than just models. And it's not like there already isn't a ton of model competition. One model released in advance of OpenAI Dev Day that we didn't have a chance to talk about yet is Claude's Sonnet 5.5. And it's certainly if you feel yourself giving a bit of an eye roll or at least a raised eyebrow wondering if you should actually care, the recent history of the Sonnet and Opus series certainly created some reason for skepticism. Opus 5, you remember, was strongly disliked, and Sonnet 5 was a bit of a disappointment. But in the end, it was a even worse result for a higher price because it was so token hungry. However, if you had to pick just one theme from the last couple of weeks on this show, it's the beloved return to form for Anthropic with Opus 5.5. So does Sonnet 5.5 continue that trajectory? The short answer seems to be, in most people's estimation, yes. Anthropic said that Sonnet 5.5 is 30% faster and 30% cheaper than Sonnet 5. And the model demonstrates some significant jumps on the which is up from just 10.3% for Sonnet 5 and even beats Opus 5.5 at 66.4%. Scores on Frontier Code and Cursor Bench were very slightly behind Opus 5.5. The same was true for Knowledge Work Benchmarks, GDPVal AA, and NAA Briefcase, where Sonnet 5.5 closed the previously massive gap between Sonnet and Open Class models. Artificial analysis affirmed the benchmarks, giving Sonnet 5.5 an Intelligence Index score of 56. That puts it in second place behind Opus 5.5 and somehow ahead of Fable 5.1 and GPT-6 Astra. However, they didn't find that Anthropic's claims about cost savings stood during their testing. Sonnet 5.5 cost $7.60 per task and used significantly more tokens than Sonnet 5 at the same per-token price. This meant that Sonnet 5.5's benchmark run was almost as expensive as Fable 5.1, 27% more expensive than Opus 5.5, and more than twice as expensive as GPT-6 Astra. Now, to give Anthropic the benefit of the doubt, they claimed a 30% cost reduction on similar tasks compared to Sonnet. The main AA benchmark is run at max inference settings, but turning down the settings to extra high cut the cost by two-thirds. Still, the model impressed once people got their hands on it. It is a massive improvement over Sonnet 5 in the rendering tests that grab attention on social media. Matthew Berman made a series of visual tests and game clones writing, Sonnet 5.5 basically Opus 5.5 but 50% cheaper and much faster. I've been early testing it and it's incredible. If this is pacing the frontier, sign me up! But for real-world use, builder Kun Chen found it performed best when it's the right tool for the job. He wrote, Sonnet is definitely not the same as Opus, just 50% cheaper. That's true only for problems that don't need much wisdom. There's a clear difference when they are asked to propose plans for ambiguous product problems. Opus is able to approach problems with more strategic thinking, i.e. what's the real goal here, while Sonnet is more just looking at tactically, how do I get this done? Now, Kun's current approach is to use Anthropic models as a complete system. With Opus for planning, Sonnet for implementation, and calling on Fable when things go sideways. Pavel Hurin tested Sonnet on his BugFix benchmark and found it outperformed all other models including Fable and Astra. He wrote, The secret? Sonnet 5.5 Max might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns. Fascinatingly, YouTuber and entrepreneur Theo's review was that Sonnet 5.5 is an incredible model that you probably shouldn't use. In his view, the model is a big improvement over Sonnet 5, but on max reasoning it overthinks and risks getting things wrong because of it. And on lower settings, there's no point where Sonnet is more cost-effective than Opus because of how token-hungry it is. This changes a little when using Sonnet as a sub-agent, where it can be more efficient on certain tasks, but it's usually not optimal as a stand-alone model. So will these results stand after a couple weeks of testing? With Sonnet 5.5 being a technical marvel given that it can produce the results of Fable 5, just four months after the release, but being too token-hungry, Sonnet 5.5 might be a better model for you. That's certainly what I'll be watching and I'll report back as people get more reps in. Still, going back to the new world that Gemini 4 has to compete in, the point that Peter Yang again was making is that just being good on the model isn't enough. In fact, one really interesting question brought up by the success of Muse is whether an AI product with a great user experience but a less-than-state-of-the-art model beats a product with a less-good user experience but a state-of-the-art model. Muse has hit 3 million weekly active users. Daily users who send at least one prompt per day have reached 1 million. Now remember, the product has only been available for a little over three weeks, so this is a strong positive early indication. At the end of the first week, Muse had 500,000 weekly active users and 250,000 daily users. Now on the one hand, this is a tiny fraction of Meta's distribution potential, which reaches half the people on Earth and is nowhere near as strong as ChatGPT's growth, but a better comparison might be the adoption curve for Codex. Codex took about 3 months to reach 3 million weekly active users. So Muse currently looks like a faster-growing AI app. Now one might expect that, given that it is a personal agent versus Codex, very developer-focused audience, but whatever the case, the early success of Muse continues on. And yet now that OpenAI's dots is here, we're going to start to have some direct answers to this question of good UX plus worse model or worse UX plus better model. Eureka co-founder Sina writes, "After using Muse for a while and loving it, I'm starting to notice the underlying model isn't smart enough." I'm not sure it's context overload, but it gets stuff wrong, a lot recently. Maybe they've had such huge success that they've had to switch to a different, cheaper, worse model. Anywho, Muse is great, form factor is great, but as a consumer, my loyalty is to none. If OpenAI will drop Muse with better models, I will use whatever is best. Intelligence is still the moat. And yet, on the flip side, a couple days in, I'm definitely seeing some gripes with the way that dots works. Habibi Slop on X writes, "Overall, the vibes are not good. This feels directionally wrong. I just want to talk to the model. This tool wants you to be a plumber in a house where you're not allowed to touch the pipes. Dots has so much friction built in that it feels like it's not meant to be used." And unless you think that this is just a general gripe or account, they then go on to include about a dozen examples. The big one that I've seen repeated comes in their final thoughts when they write, "They've launched a new product that wants to help you manage your work in ChatGPT and Codex without having sufficient access to any of the tools and data it needs from those surfaces." I am certainly at least not even close to being willing to declare which is the winning strategy between the tools and data it needs from those surfaces. I am certainly at least not even close to being willing to declare which is the winning strategy between focusing on the model versus focusing on the product user experience. One last note as we look across the state of the competition that Google finds itself in as it gets ready to release Gemini 4. In one surprise launch this week, DoorDash has announced that their AI agent is getting the Muse treatment and can now take orders via text. DoorDash already had an agent on their site, but they're using some of the features of personal agents like Muse to take it a step further. Customers can now text the DoorDash agent with a natural language prompt, and the agent will figure out the rest. DoorDash gave the example of a customer texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." DoorDash gave the example of a user texting, "Order my username and password." There's a risk, of course, that external agents like Muse could order directly from restaurants and cut DoorDash out as a middleman. At the same time, user preference might reject platform-specific agents because they have much less control. It's unclear what the optimal strategy will be, but for now, DoorDash seems to be doing a bit of both, keeping their platform open to third-party agents while also building an internal agent, and for that reason it's worth paying attention to, even if you don't much care who wins, how you order lunch to the office. Anyways, bringing it back to Gemini 4 Argon, I'm with Nathan Lambert when he argues that Google doing well is both better for consumers and better for competition, so put me firmly in the camp of people who are hoping that this release is great. Let's just hope we get it soon so we can actually decide for ourselves, rather than dining from the table scraps of self-reported benchmarks and gripey employees talking to Bloomberg. That's gonna do it for today's AI Daily Brief. Appreciate you listening or watching, as always, and until next time, peace! Thanks for watching!

Podcast Summary

Key Points:

  1. Google announced Gemini 4 Argon, its first frontier model in over six months, after falling behind rivals during 2026's vibe coding boom.
  2. Gemini 4 posts state-of-the-art or near-frontier benchmarks in agentic knowledge work, legal tasks, and DeepSuite coding, though it trails leaders on Frontier Suite and Terminal Bench.
  3. Google is not releasing Gemini 4 publicly, citing cybersecurity concerns, limiting access to trusted cyber defenders with gradual expansion planned.
  4. Skeptics note Google's history of strong self-reported benchmarks that disappoint in practice, and Bloomberg reported internal employee skepticism about real-world coding performance.
  5. The competitive landscape now demands more than strong models, requiring competitive coding harnesses and personal agents, as Peter Yang noted.
  6. Anthropic's Sonnet 5.5 launched as a fast, cheaper model with strong benchmarks, though testers found it token-hungry and less strategically capable than Opus 5.5.
  7. Muse's rapid growth to 3 million weekly active users raises the question of whether strong UX with a weaker model can beat strong models with weaker UX.
  8. DoorDash launched a text-based AI ordering agent, illustrating how personal agent competition is expanding beyond frontier labs.

Summary:

Google announced Gemini 4 Argon this week, its first frontier model in over six months, following a period where the company had fallen outside the top tier of AI labs. The model posts impressive benchmarks, achieving state-of-the-art scores on agentic knowledge work, VALS index, legal benchmarks, and DeepSuite coding, while trailing slightly on Frontier Suite and Terminal Bench. Artificial Analysis placed it tied for third on its Intelligence Index. However, Google is not releasing the model publicly, citing cybersecurity concerns, instead limiting access to trusted cyber defenders with plans for gradual expansion.

Reactions were mixed. Many celebrated Google's return to the frontier, while skeptics noted Google's history of strong self-reported benchmarks that disappoint in practice. Bloomberg reported internal employee skepticism about real-world coding performance, which Google denied.

The episode also covered Anthropic's Sonnet 5.5, a fast, cheaper model with strong benchmarks but token-hungry behavior, and Muse's rapid growth to 3 million weekly active users, raising questions about whether superior UX can compensate for a weaker underlying model. The broader takeaway is that competition now extends beyond models to harnesses, agents, and product experience.

FAQs

Gemini 4 Argon is Google's first new frontier model in over six months, announced with state-of-the-art benchmark scores in areas like agentic knowledge work and coding.

Google cites cybersecurity concerns, as the model scored 68% on the CWE Bench cybersecurity benchmark. Access is being limited to trusted cyber defenders, with gradual expansion planned.

It is the new state-of-the-art on DeepSuite with 77.9%, but lags behind rivals on Frontier Suite and Terminal Bench 4.0, placing it in the same ballpark as frontier models from OpenAI and Anthropic.

The accord is a one-page commitment by frontier labs to develop internal controls for model testing and deployment, with external auditors for verification. Trump called it 'morally binding,' but critics question its voluntary nature.

America.gov is an AI-powered single portal for government services, established by executive order. It currently functions as a knowledge base searching 29,000 government websites, with plans to handle tasks like passport applications.

The FTC has opened an investigation into rogue agent events, including the Hugging Face incident, and is drafting civil investigation demands to compel executive testimony.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.