Go back

EP 416 : Google Dominates AI Benchmarks & MinMax 2.5 closing huge gaps

14m 31s

EP 416 : Google Dominates AI Benchmarks & MinMax 2.5 closing huge gaps

This episode analyzes a transformative week in AI, marked by three major, divergent shifts. Google reasserted dominance with Gemini 3's "Deep Think" update, showcasing a generational leap in abstract reasoning and launching a powerful math research agent, moving from catch-up to pace-setting. OpenAI pivoted towards speed and hardware independence with GPT-5.3 Codex Spark, leveraging new chips for ultra-fast, real-time coding assistance, albeit with a trade-off in raw analytical power. Simultaneously, Chinese lab MiniMax disrupted the market economics by releasing the open-source M2.5 model at a drastically reduced cost, making pervasive AI agent automation financially viable. The discussion connects these developments to broader trends: increasingly realistic AI video, bold predictions of imminent white-collar job automation, and industry consolidation evidenced by massive funding rounds and model retirements. The conclusion is that the AI ecosystem is expanding into a specialized economy where tools are chosen based on specific needs—reasoning, speed, or cost—fundamentally altering software development, business workflows, and the nature of work itself.

Transcription

2759 Words, 15565 Characters

English
Hello and welcome to episode 416 of AI brief. It is Friday, February 13th, 2026. We made it through another week. We did. And somehow the entire AI landscape looks completely different than it did on what Monday. It really does. I mean, usually we get one big story that just dominates the whole new cycle. Right. But this week felt different. We have three huge shifts and they all kind of contradict each other in these fascinating ways. Okay. So what are we looking at? Well, you have Google who just woke up and decided to destroy the benchmarks. Then you have open AI making this big hardware pivot that could change the economics for, well, everyone and the third one, a Chinese lab that is driving the cost of intelligence so far down. It's basically free. And that's not even mentioning the everything else pile. I saw a note that a Microsoft exec is predicting total white collar automation in what 18 months. We can't just bury that. No, absolutely not. That's a headline. That's something we need to seriously unpack. So that's the mission for this deep dog. We're going to connect these dots. It's not just about new models. It's about what this all means for, you know, your job, your software, your sanity. Let's do it. Let's start with the comeback story. For the past few months, it feels like all the opposition in the room has been taken up by open AI and anthropic. Google for quiet. Almost suspiciously quiet. I'd say dormant really, but that ended this week. Okay. They dropped a major update to Gemini three. It's called deep sink reasoning mode. And we really need to look at the numbers here because they aren't just, you know, small bumps. Oh, we'll look at the notes and the ARCAGI two score just jumped out of me. Deep think hit 84.6%. Yep. So for anyone listening who doesn't live in breathe benchmark spreadsheets, can you just explain why we care so much about ARCAGI? What makes this number the one everyone is freaking out about? It's a great question because we all have benchmark fatigue right now. Most of them, you know, the bar exam medical boards, they test knowledge things in LLM can memorize exactly. They've read the internet so that they can basically regurgitate answers. ARCAGI is different. It's a test of pure abstract reasoning. Visual puzzles. The model has never ever seen before. You can't memorize the answer. You actually have to think. So it's an IQ test. You can't cheat on that's the perfect way to put it. It tests fluid intelligence, not just recall. And on that test, deep think hit 84.6%. Put that in context for me. Where are the competitors on that? Because 84% sounds good, but how good is it? Well, Claude Opus 4.6, which everyone agrees is, you know, the gold standard for reasoning right now is sitting at 68.8%. Okay. And GPT 5.2 is even further back at 52.9%. Wow. Okay. So that is not a margin of error. That's a massive gap. It's a blowout. A 15 point lead on a reasoning test like this is it's a generational leap in AI terms. It really suggests Google has figured out something fundamental about how the model approaches new problems. So they're not just predicting the next token. No, they're planning. And it's not just on abstract shapes. They took this and applied it to coding. Right. The code forces score. I saw it hit a 3,455. E-Lo rating. No, I know E-Lo from chess, but what does 3,400 mean for coding? Good. Doesn't even come close to give you some perspective. A 3,455. E-Lo makes this AI a legendary grandmaster. It's ranked higher than basically every human the entire platform. And compared to the competition. It's almost a thousand points higher than opus 4.6. A thousand points. Yes. In any competitive ranking, a thousand point gap is like a smart amateur playing Magnus Carlson. It's not even a contest. The model is just writing little scripts. It's solving problems that require real intuition. But here's the skeptical take. Does this actually translate to real world use? Or is this just Google flexing on a test? That is always the right question. And it seems like Google wanted to answer it immediately. They launched a new agent alongside this update. It's called a Lathea. A Lathea. That sounds serious. It is. It's a math research agent. It's built to autonomously prove mathematical theorems. And it's already hitting gold medal standards in the physics and chemistry Olympiads. So this is not a chatbot for recipes. Not at all. This is a tool you give to a researcher to help solve open problems in science. So if you're a Google AI ultra subscriber or you have API access, you now have a reasoning engine that is, well, arguably the most powerful one on the planet. It flips the whole narrative. Google isn't playing catch up anymore. They're setting the pace, which is a perfect transition to open AI. Because if Google is all about deep reasoning, open AI seems to be going in a totally different direction this week with GBT 5.3 codex spark. Spark is definitely the key word. This whole release is about one thing. Speed. But the real story here isn't the software. It's the hardware. This is the Srebrus news. Yes. This is the first big open AI product that we know is powered by Srebrus chips, not in video. That feels like a huge strategic move. But why should the average person care what chip their AI is running on? Well, the average user cares about the performance, which we'll get to. But the industry cares because for what the last five years, the entire AI economy has run on in video. If you couldn't get a 200s, you were stuck. So open AI, making deals with Srebrus with AMD with Broadcom. That's them declaring independence. They're making sure they never get bottlenecked by one supplier again. Okay. So that's the business side. How does it change the experience for me when I'm actually coding? It makes it incredibly fast. We're seeing speeds over a thousand tokens per second. 1000 tokens a second. That's what a whole page of code just appears instantly. Basically, yeah, it finishes right in before you even finish thinking your thought. But there's a catch. Of course there is to get that speed. They had to trade off some of the raw intelligence. So if you look at the really complex coding benchmarks like S.W.E. Bench Pro, Spark is actually behind the full codex model. It's not quite as smart. So why would I want to use it? If I'm trying to code, I want the smartest possible model. Why would I choose the dumb but fast one? Think about your actual workflow. If you're just fixing a typo or changing a variable name, do you want to wait 20 seconds for the Einstein model to think about it? No, I want it done now. Exactly. It's like having a sous chef to chop the vegetables while the executive chef designs the menu. That's a great analogy. Open AI wants you to use Spark for all that interactive real time stuff, the quick edits. And then you let the full heavy duty codex model handle the big autonomous tasks in the background. It removes the friction. Okay. So we've got Google on one end with Max Reasoning, open AI on the other with max speed. But there's a third player that just completely changed the game on price. This might be the most disruptive news of the week. Mini max the lab from China. Right. They just released M2.5. Correct. It's open source and its performance on agentic coding benchmarks is right up there with Opus and GPT 5.2. But the price I have here, it's a dollar 20 per million outputs, a dollar 20 compared to what 25 dollars per million for opus. That's that's like a 95% price cut. That's unbelievable. It's basically intelligence to cheap to meter. It's here. So what happens when the price drops that low? Does it just mean my API bill gets smaller or does it change what's possible? It changes everything about software architecture. Right now, if you're building an AI agent, you have to be careful. You can't just let it run in a loop. You'll get a massive bill. You're constantly checking on it. Exactly. But at a dollar 20, you just let it run 24/7. You can have agents running around the clock, optimizing code, handling support tickets, doing all the administrative grunt work. In fact, Minimax said they're already using it internally for 30% of their company's daily tax. HR sales are indeed everything. And they said it's handling 80% of all their new code commits. 80% that's the validation right there. They trust their own model with their own product. It means the barrier to entry for having a fully automated workforce, which is completely collapsed. So if you step back for a second, Google has the IQ open AI has the speed and Minimax has the price. It feels like the whole ecosystem is just expanding in three directions at once. It is. It's no longer a single race. It's becoming a specialized economy. You just pick the right tool for the job. Need a genius. Google need it now. Open AI need to run it a million times. Minimax speaking of automated work forces in the broader economy. I want to slow down here for a moment. Yeah. There's a lot of other news that came out today. And I think we should take a beat and go through these updates with a little more focus. Some of them have really huge implications. Agreed. Let's connect some of these other dots. First up is video. We've been seeing AI video get better and better, but by dance TikTok's parent company, just released seedance 2.0. And the word is it's finally crossed the uncanny valley. This is a significant threshold, you know, with previous video models, even the really good ones, you could always spot the glitches. A hand might pass through a table or someone would just walk in that weird floaty way. The physics were just off. Seedance 2.0 seems to have fixed a lot of that. The examples I've been seeing are pretty stunning. People are reimagining scenes from Lord of the Rings, creating these intense war combat sequences, even martial arts. The movement looks real. That natural movement is the key. Martial arts in particular requires a really precise understanding of body mechanics and a flate and impact. If an AI can generate a convincing fight scene without limbs just clipping through each other, that implies a much, much deeper understanding of 3D space and physics than we've ever seen before. It really blurs the line between what's real and what's generated. Absolutely. Now moving from the creative side to the corporate world, Mustafa Suleyman over at Microsoft made a very bold prediction this week. The one we mentioned at the top. He did. He told the financial times he expects most white collar work will be fully automated by AI within 12 to 18. 12 to 18 months that feels incredibly aggressive. Most projections are more like five to 10 years out. It is aggressive, but you have to consider the source. This is Microsoft. They have the data. And he connected that prediction to a statement that Microsoft is pursuing true self-sufficiency with its models. What does that actually mean self-sufficiency? It implies agents that don't need a human holding their hand. They don't just answer your questions. They complete entire workflows for you. And if you connect that back to what we were just talking about, agents cheap enough to run forever for minimax, and agents smart enough to reason from Google, the pieces are all there. If Microsoft is building that directly into windows, then 18 months might not be a prediction. It might be a product road map. That is a very sobering thought. Not a prediction, a road map. It definitely raises the stakes. And speaking of high stakes and big changes, let's turn to XAI. Elon Musk addressed the recent wave of departures over there. Yes, they lost 10 co-founders and key engineers. Musk's framing of this is that it wasn't people quitting, but a forced re-orig for speed of execution. Speed of execution seems to be the phrase of the week. It really does. It signals that XAI is likely shifting from a pure research phase into a product delivery phase. And when that happens, the culture changes. It becomes less academic, more about the grind. He's clearly tightening the ship to compete. And while he's doing that, and Thrawnpick is just filling up his war chest. They are. They officially announced a $30 billion funding round, which puts their new valuation at a staggering $380 billion. It's hard to even wrap your head around the number that big. It is, but then you look at their revenue. Their run rate just hit $14 billion. And what's really interesting is that $2.5 billion of that is specifically from Claude Coat. So coding is the killer app. It's where the money is. It's a primary revenue driver. That $2.5 billion figure is exactly why everyone, Google, OpenAI, and Mini Max is fighting so hard for the developer market. That's where companies are willing to pay right now. And lastly, just a little bit of housekeeping from OpenAI that has some users. A little upset. They're retiring a few models today. Yes. As of today, GBT40, GPT 4.1, and a four many are all being retired from chat GPT. A lot of people are not happy about losing fouro. It had a very specific personality. There's always some pushback when a beloved model goes away. You get used to the quirks, the vibe of it. But from OpenAI side, maintaining all those huge older models is expensive and it splits their focus. They're clearly pushing everyone towards the five series. It's the inevitable software cycle just happening at warp speed. It really is. Out with the old and the new seems to arrive every single week. Speaking of which, let's move on to our tool rundown. We've talked about the huge foundation models. But what else is trending that people can actually use today? Well, first obviously are the big three we just covered. You have Gemini 3 deep think for those heavy reasoning or science problems. Then GPT 5.3, Codex Spark, for that super fast real-time coding help. And M2.5 from Mini-Max if you want to play with open source agents without breaking the bank. But what about more specialized tools? What's hayfish? Hayfish is all about creating those hyper realistic AI digital human videos. You could use it for say, training videos or customer service avatars. It's getting very realistic, very fast. And for anyone like me who is stuck in presentation hell all day. You'll want to check out dokey. It generates professional slides in minutes using AI. Basically trying to kill PowerPoint, which I think we can all get behind. Definitely. And for the creatives. There's a cool one called whisk AI. It lets you blend separate images for a subject, a scene, and a style to create really detailed 4K artwork. It gives you a lot more control over the final composition. And finally, one for productivity. Good assistant. It's a companion that's designed to keep you focused on your goals. It tracks your progress and gives you little nudges. It feels like as AI does more of the work, the real challenge is just keeping the human focused on the right work. That's the absolute truth. It feels like every week we get new superpowers. And the real challenge is just figuring out how to use them without getting totally overwhelmed. Exactly. The tech isn't the barrier anymore, the capabilities are exploding, the new barrier is just imagination and implementation. Well, that is all the time we have for today. Thanks for this deep dive. It was a pleasure. Be sure to follow the show on whatever platform you use. And if you know someone who needs to stay in the loop, please share AI brief with them. See you in the next episode.

Podcast Summary

Key Points:

  1. Google's Gemini 3 Deep Think update achieves a breakthrough in abstract reasoning (84.6% on ARC-AGI), significantly outperforming competitors, and introduces "Alethea," a math research agent for scientific problem-solving.
  2. OpenAI's GPT-5.3 Codex Spark emphasizes extreme speed (over 1,000 tokens/sec) via new hardware (SambaNova chips), trading some intelligence for real-time coding assistance, marking a strategic shift to reduce dependency on NVIDIA.
  3. MiniMax's open-source M2.5 model drastically reduces AI cost to $1.20 per million outputs, enabling widespread, continuous agent automation for tasks like coding and administrative work, collapsing economic barriers.
  4. Broader industry shifts include AI video (ByteDance's Seedance 2.0 crossing the uncanny valley), predictions of near-total white-collar automation within 12-18 months (Microsoft), and significant funding/consolidation (Anthropic's $30B round, OpenAI retiring older models).

Summary:

This episode analyzes a transformative week in AI, marked by three major, divergent shifts. Google reasserted dominance with Gemini 3's "Deep Think" update, showcasing a generational leap in abstract reasoning and launching a powerful math research agent, moving from catch-up to pace-setting. 3 Codex Spark, leveraging new chips for ultra-fast, real-time coding assistance, albeit with a trade-off in raw analytical power.

5 model at a drastically reduced cost, making pervasive AI agent automation financially viable. The discussion connects these developments to broader trends: increasingly realistic AI video, bold predictions of imminent white-collar job automation, and industry consolidation evidenced by massive funding rounds and model retirements. The conclusion is that the AI ecosystem is expanding into a specialized economy where tools are chosen based on specific needs—reasoning, speed, or cost—fundamentally altering software development, business workflows, and the nature of work itself.

FAQs

DeepThink is a major update to Gemini 3 that achieved an 84.6% score on the ARC-AGI benchmark, a test of pure abstract reasoning. This represents a generational leap, as it significantly outperforms competitors like Claude Opus 4.6 (68.8%) and GPT-5.2 (52.9%), indicating Google has made fundamental advances in AI reasoning capabilities.

Codex Spark prioritizes speed over raw intelligence, delivering over 1,000 tokens per second for real-time coding assistance. It's designed for quick, interactive tasks like edits, while the full Codex model handles more complex, autonomous coding tasks, optimizing workflow efficiency.

M2.5 is an open-source model that offers performance comparable to top models like Opus and GPT-5.2 but at a drastically lower cost of $1.20 per million outputs. This 95% price reduction makes AI agents affordable to run continuously, potentially automating many business processes.

Mustafa Suleyman predicted that most white-collar work will be fully automated by AI within 12 to 18 months. This aggressive timeline is based on Microsoft's data and their pursuit of self-sufficient AI agents capable of completing entire workflows without human intervention.

Key tools include Gemini 3 DeepThink for reasoning tasks, GPT-5.3 Codex Spark for fast coding, and MiniMax's M2.5 for affordable agents. Specialized tools like Hayfish for digital human videos, Dokey for slide generation, Whisk AI for 4K artwork, and Good Assistant for productivity are also highlighted.

ARC-AGI tests pure abstract reasoning with visual puzzles that models cannot memorize, making it an effective measure of fluid intelligence. Unlike knowledge-based benchmarks, it requires genuine problem-solving, so high scores indicate advanced reasoning abilities beyond simple recall.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.