Go back

532: HydraFusion: Multi-Model Orchestration Changes Everything

42m 38s

532: HydraFusion: Multi-Model Orchestration Changes Everything

The podcast explores how developers are evolving their AI workflows by moving away from over-reliance on expensive, high-tier models. Hosts James and Frank advocate for using lower reasoning levels in AI, which delivers faster, more practical results and avoids costly loops. They highlight that small, efficient models like Baby Luna and MAI Code 1.1 Flash can match the output of larger models—especially when reasoning is maximized—while cutting costs by up to 90%. A key innovation is Hydrophusion, a multi-model orchestration tool that automatically selects optimal workflows (drafting, reviewing, critiquing) to improve code quality and reduce expenses by 70% without requiring user input. The hosts emphasize that developers should avoid generating excessive documentation, as it often contains hidden errors and is rarely read. Instead, they recommend targeted, actionable outputs and simple, efficient workflows. They also stress the importance of reviewing and deleting outdated AI-generated documents to prevent technical debt. Ultimately, the shift toward intelligent, automated orchestration—combined with trust in smaller, faster models—represents a more sustainable, cost-effective, and efficient approach to AI-powered development.

Transcription

7736 Words, 41113 Characters

English
Well, hello back everyone, Emerge Conflict, do you win the developer podcast? Did you like that dramatic pause there, Frank, or you saw it? I was again, over two seconds in because you make fun of me every time because I have no idea when that podcast actually starts. So I have to physically look at you. So when you play games with delays like that, my heart just races, am I going to record or not? Is he frozen or not? Is he doing something with the video again as he likes to do? Hi, everyone. Welcome to Merge Conflict. You're one of our co-hosts, Mr. Frank Kriger, and I'm the other co-hosts, James Montzmann. How's it going, buddy? Oh, it's going good. I've spent the day yelling at an AI. I think that's what we've been reduced to as software developers is just resisting writing really, really, really bad language at the AI as it messes up something what you consider to be fundamental and basic and simple. I like it when I put a lot of effort into like perfecting some code, and then I ask it to make a small change and it rewrites on my code. It's always so delightful. And then you get a bill, and then it's like, you're out of credits for the month, you're like, great, great, great, you spend my time and money, okay, I'm going to go pet the cat now. Well, I kind of want to talk about this a little bit because how I've been working with models has changed, and then also with agents, specifically like kind of like long running agent processes has changed. And the feel is that we did a podcast maybe like three months ago. And since that is all old and out of date and busted now, I figured we should revisit it. Okay. So we did a podcast, Frank, about how you should never change the default reasoning. Oh, yeah, yeah, it's, yeah, I still have mixed feelings on it because I think the advice is different for every model now. But I'm curious to hear what you have to say. Yeah, you know, you see all these benchmarks and they're like, oh, we were doing like, you know, opus on like this thing and then that thing on this thing and this level on this thing. And I was like, well, how do you even really compare it? I feel like that's the hardest part is these, these comparisons, like how long did it take? How many tokens did it do? Like what's these, you know, there's outcomes, all these other things that you do. So I think the benchmarking is really hard for a normal developer or normal person to understand. You see all these like lines and charts and these things are going. I'm like, I don't even know what this says. Is it better? Is it worse? Is it right? And I should just probably use the default. That was our guidance was use the default and are you using the default Frank? Yeah, default or lower. I'm still kind of sticking with that because there's nothing worse than being on high and watching it just kind of think in circles because that's all they ever seem to do is think in circles. And then it'll still output the wrong code. So my philosophy is basically those put it on medium or low thinking, get the good or bad code quicker. And then I can quickly check like, is that good or bad? Nope. Try again. And that turns out to be faster than putting thinking to high, like putting me in the loop I find it's superior to high thinking. Now I'm using these things. I believe as the kids would say wrong. I'm doing it wrong. I'm holding it wrong. Like I think you're supposed to not read code, not care what it's actually doing. Go get a coffee, come back after you spend, I don't know, 10 grand or so. And then that's when you find out if it worked or didn't work. I believe that's the correct way to use these AIs according to the kids. But I do it differently. I watch over them like a mean manny or something. Well, that's where I'm going to talk about how I've changed a little bit. But that's majority of what I'm doing is I'm usually sitting there and I'm normally taking a look at it. And I'm trying to understand like the questions and the output that it has. But I have found something a little bit different in two aspects. First, I want to talk about planning. I don't want to talk about God. Oh, no. Oh, no. I have opinions. That's one area. Or I would say research, okay. So like research planning and then there is the second part, which is little baby models. And this is where in both instances, I think there is high value in changing the reasoning levels of it. So example one, research planning, et cetera. Now normally I'm doing a normal plan like I need to implement this feature. And it's pretty high level. Good example is on tiny clips. I noticed when the editor window opened and I took another screenshot, it would close that editor window and open a new editor window. So multi window wasn't enabled. So I was like, hey, do a little bit of research first. Go figure out, go look at docs and then let's put a little plan together. Just normal mode, normal medium, you know, soul, you know, maybe tarot reasoning, figure it out, implement done 50 cents. I'm good to go. Right? However, I then was doing this really big cook. And I was doing it with a Fable 51 and I said, I want you to do really deep research on performance profiling of recording video on Windows with these APIs, analyze my APIs, take a look at this sandbox it. And then I kind of gave it some guardrails, but I put it into extra high mode, which is scary because it's like, oh, Fables already expensive, give me extra high mode. However, what I wanted to do was see if I could get really, really good rationale. And what I ended up finding is that it would go in, like really do really, really crazy deep research and analysis and look at APIs and do this stuff. So this was like something that I was like, this is a highly, highly valuable feature performance supplementation, like if I was to spend $10, $15, $20, $30 on making the general performance of the video recording and capturing better worth my time instead of fighting because I've already fought with it for two months and it's, it's okay, but I was like, do this thing. So I let it cook and it was like letting it cook, letting it grind and it put together and I had to put together a research document for me, canonical research of all these APIs. And I think it's in the repo and the tiny clips repo and it's like really detailed and it was like beautiful. I don't know, maybe it was like $15, $20 or something like that. So pretty expensive. Pretty expensive. However, then I was able to switch it into a little model where it doesn't need to think because the thinking's already done and then just implement it. In fact, yeah, in fact, yeah, exactly. So kind of going back to this rationale. So I thought that that was interesting because, you know, that's one example now granted, I do not use much fable 5.1 or Astra or anything. I'm mostly on Terra. So even this same logic would apply to like Terra, right? Maybe put that on Max mode, which is a middle model of the GPT-5.6. Just saying, hey, listen, let's go and cook this up. And that would have been a lot cheaper in general for most things. And then let me get some really detailed part. Not that I would do it every time. This is what I really wanted to get into. Not that I would do it every time, but for a really meaningful, deep API research reporting analysis, go and then do the thing. And what was cool about that is it wrote guidelines on how to do performance profiling. So when Terra, or I think it was Terra that went to go implement it, it didn't need to do tons of research. It had everything. And then it was able to do all this performance profiling automatically for me in like real time. It was like really cool. So that was like a really neat implementation of saying, let's turn this puppy up. So I don't know if you've in, I know you're, this isn't really, it is planning, but it's not planning. But like that's the type of planning I'm into now, for example, another example really quick would be a building that on a confkino potential demo, this pickleball app. And I was like, you know what I need to do right now is I need to do some deep research on orchestrating the entirety of the stock, not spectraven development, but like legitimately, let's plan this puppy out. So I took soul, put it in extra high mode. And I said, let's cook on a plan together, ask me every little detail, do this. What MCP servers do you need? What's this? What's that? And it put together 18 documents for me. Step by step. Let's implement this puppy. Let's cook this puppy. Let's do it, right? And it was really, really neat because it, it even understood, you know, it did the additional research to like say, okay, what about authentication? What about mobile authentication? What about this? What about databases? What about this? You look in like, where do you want to run this for development? Do you want to run WSL? Do you want to do this thing? And like, I was like, oh, this is like really neat. So it really puts together this super fleshed out plan by doing that deeper research that I don't get by just like going and running, right? Yeah. I'm with you. I actually do a lot of the same. So where I was saying about coding, I put it to medium or low. I do use high for not exactly planning the way you're doing it, but more like brainstorming and figuring things out. I tend to, I tend to do the kind of planning in my head a little bit. So when it comes down to like, there's a feature I want or something. I usually say like, I should use those like real me skills or whatever. But I'm basically like, ask me 20 questions like you are not allowed to do anything until you've asked me at least 20 questions and then we'll go on from there. But that is when I use the high reasoning because that's when I want it to be smart. And I think you're making a good point here, you're, you're using an expensive smart thinking model to generate an artifact and that artifact can then guide the little models. I think that that's a pretty good procedure. I do get a little bit nervous though because I watch a lot of YouTube vibe coding AI experts and my god, they create a lot of documents. James, like for the smallest feature, they'll create like the plan, the requirements, the design doc that this, that, and it's just like doc after doc after doc. If there's one thing a language model likes to do, it's right. They love to output markdown. If you ask them to output markdown, they will output some markdown. And I find that fine because you're whittling the problem down and you're guiding maybe the dumber AI's where I've run into personal problems with it is it'll generate this big doc and then nestled in that doc is something I actually don't agree with. But I'm not going to read 50,000 words. That's like a novel and it's just not going to happen. So you generate this doc and then you hand it over to lunar or whatever a smaller model. You don't know what little hidden things are in there that are going to come to bite you the butt later. Now, maybe 90% is fine. It's fine. You know, it's better than not giving lunar guidance. It's better than spending 10,000 dollars on Astra to do something crazy. So I'm with you. But I've I still get nervous with all these docs because I've had it happen over and over and again where they just hide some stupid little thing that I don't agree with at all in that doc. I didn't read the doc. No one's going to read these docs. They're spitting out words like they're free. They're not free. But and then something down the road, I'm like, why'd you do that? They're like, well, in the doc, it's said to do that. You got me. You got me brother. I didn't read docs or TFM. So I'm with you, but with some caveats. I think it's true. You know, we talked about spec K because I want these other things and someone was asked me when I was taking screenshots and sharing the keynote that the demo app that I'm working on and they're like, how did you do in like spectrum development? I said, no, I just I did create a bunch of like general documents about like what like what should be in the API and like what should be in the mobile app and like what should be the orchestration and like what are all the different features basically? I was almost kind of like PRD docs, but they were a little bit more detailed like you said. They definitely had like spec numbers and things like that. And you're right, like I'm not going to read any of these docs. So I assume that they're correct because there's like a lot of words. So that does become a problem. That being said, like the real world example where it is smaller in tighter like the video recording stuff, I actually didn't read that doc because it was like smaller in tighter and I was like generally, I was genuinely interested. I was like, oh, I am interested in this topic. You know, can I do this? Now that being said, most of these design spec docs like, you know, even in spec kit, it's like, oh, you don't read the docs. Someone else reads the doc and reviewed the docs and do the docs. And that other thing is an AI, right? So you can't feed it all in there. There has to be some taste at the end of the day, you know, and doing that. And this is where I think our perennial topic of Greenfield versus what you disgustingly call brownfield software development. AI's continued to rock brownfield when there is already something there and they just need to debug it and fix it or add a feature or something. I trust them. That's the times I don't actually read code. You know, I'm just like, here's a nice code base. Go implement this thing. I beta tested. I'm the tester. It's all fine. It's when you're starting from scratch. That's when I can start putting the weird stuff into all these docs that you won't read. And then the best advice with these docs is just when you're vaguely done with that session, just delete the doc. Clear the memories also. You'd be surprised. Go read the memories that these agents are writing. You will be surprised at what nonsense they think is important. Your one little criticism of them will become like 20 paragraphs in a document about passive aggressively trying to satisfy your needs. So I think the problem with the docs is we don't read them. The second problem is they get out of date. So the moment the features don't just delete the doc and go into the memories and delete your memories too. I promise you, life will be better. Just delete those stupid memories. What a stupid feature. I wish I could just turn off. Yeah, it is good. I think that if there is anything you could extract, like what you could do is you could extract, for example, very specific pieces of details. Maybe you asked the AI to extract some detail information that's like important when working in this file and then created a custom instruction for working in that file. You know what I mean? Instead of doing like this big, here's a random doc, like no, only care about this thing in general. I'll say this, like this instance, I do more rare, right? I'm not necessarily always pumping it up and turning it in on this. However, there is another change that I think is more realistic that I think you could be doing right now today, Frank, and every single listener is that I'm ready. Everyone on Twitter is so obsessed with just using the most biggest expensive model and putting it on the highest reasoning and building things that don't, you know, don't matter, right? Like game clones and this and that. It's cool. It's like what these things can do. Fantastic. I love it. Burn in the tokens for fun. Let's do it. Let's build some real software. But the, and I do it myself. So I'm guilty on it. Like I literally built Bananify.Online, which is one of my favorite websites where you can bananify your website. It's an extension. There's a VS Code extension, Visual Studio 2026 coming soon. Of course, of course. Bananify.Online. It's great. We can preview it at Banan, but Nana. I think these AIs are enabling your bad domain purchasing habit. They're just enablers. I did. It definitely did. Okay. So, Frankriger, the thing that I would call out here is there is specifically these little smaller models that are really capable, especially more capable than models we had six months ago, a year ago, that are extremely impressive. And their default reasoning is pretty low. However, their cost is negligible. I'm specifically talking about two models that I would call out, which is going to be GPT-56 Luna, or I like to call it, Baby Luna. Baby Luna. And then also MAI Code 1.1 Flash. Okay. You've been using it. I know we talked about it a while ago. I'm always scared to click on the My because I do have that FOMO thing. I'm like, what if I put all this time into this prompt and then let it think forever and then it generates garbage? Well, I always be wondering, would Strev generated garbage, would Fable have generated garbage? So that's what it's the FOMO thing about using smaller models, which I find funny because on the podcast, just like three months ago, I'm like, hey, babe, I'm running local models on my own machine and I'm loving it because I kind of was. And I stopped doing that a bit and I'm starting to miss it a little because I have been brainwashed by big AI and using the biggest, most powerful model, because basically FOMO. And there's no other reason. And I want to go on record right now, I had asked for burning, burning credits, doing good things. And it completely made a mess of what I asked it to do. So even these like AGI models, you know, they still suck at certain things. And they're still terrible at certain things. So I'm excited to hear why you think, may my and baby Luna, are good for you. Yeah, I think what I've noticed, I've talked to the MAI team a lot because they, you know, they really do a lot of training and stuff on the harnesses and work really deep in the creation of a bunch of great things out there is that, you know, they they specifically say and I've talked to them like, well, these are really targeted. Like, oh, I have a bug fit. I have a bug. I need to do this. Like, it's really good at like targeted things as well, maybe not as a scaffolding the entirety of this pickleball application, but there's that. But what I found is that these little models are like very fast, like extremely fast. And the big models, it's not that they're slow. It's just that they're, they're like, they're, you know, swimming in a big ocean of possibilities. And they love to, you know, it's kind of like I've always said with the quad models. But now I think the GPT models are really, especially Astra, starting to poke around everywhere, right? And it soles a little bit better as far as speed goes and exploration. Astra is a much more of a feels, a much more a little bit exploratory kind of like Opus did and saw them for that. But what I found is that if I take either baby Luna or MAI code 1.1 flash and I just turn the reasoning to the highest possible level, the highest possible level. What I found is that the results are pretty much exactly the same as if I was a pick Terra or soul on the default level. And I did this in a very pointed one. So that example of tiny clips that specifically, and you can do this, which it couldn't do multiple things, like it couldn't do multiple windows. This needed to maybe change like two or three files. Basically the absias for tracking windows and needed to do something another window. As if this is a great fix. So I did it with like soul, Terra, and then I did baby Luna and then at MAI, one-on-one flash. The difference was as I did default, default, super high, like I think it's max and then like extra high or whatever on one-on-one. And at the end of the day, they all within 99% has the same exact results. They all fixed the thing. They all also updated my change log like I told it to do in my hooks. And then I think that, I forget which one now. But three out of four also made sure to implement the test harness updates required in the smoke test for it. So what I found and then speed wise is that M.A.I. and Baby Luna super fast. However, the most important part, we're talking about 90 to 95% savings in Baby Luna versus the other ones. And M.A.I. one-on-one and Baby Luna, I think are the same prices or very similar prices. But we are talking a crazy which means you can even bump up into one or magnitude. Oh, yeah. Well, it's an order of magnitude. We're talking like I'm talking like a few dollars, it's like $1.50 to like 15 cents, 10 cents, you know what I mean. I'm also doing this thing for this GitHub Co-Pilot Day. Co-Pilot Day, keynote where I'm generating this website for Seattle, like coffee shops and stuff like that. And I've been testing nonstop mostly for timing for the keynote. I've tried Astra. I've tried Opus. I've tried all these things. Baby Luna will implement almost the identical website. I'm not saying it's identical, but a beautiful website in under five minutes with full CICP pipeline, do the testing, do a whole thing, 25 cents. The initial prompt to Astra, just like it reading is like, okay, I'm going to implement it, you're already at $1. Do you know what I mean? Yeah, I'm just saying. I do. No, it's like, let's go, right? That initial hit of that input token cost comparatively. So what I'm thinking is, I'm also okay if the max highest reasoning level on these little models does even take the same amount of time a little bit longer because the cost differential is bananas. Yeah, my whole argument against the max was time. So you're compensating for that time by using just a much faster model. So that makes perfect sense to me. Honestly, you're inspiring me a lot because I miss my old local model because it's, it sounds terrible. It's stupidity was a feature. And I've noticed this with other models and I don't mean to insult baby Luna or me. But like, you know, the model I use the most these days is the stupid one behind the Google search. I'm using the, I don't know which, I'm sure it's a Gemini. I don't know which Gemini it is, but it's a stupid one. And you can tell it's a stupid one because it's super fast. And they're basically giving it to you for free. I haven't noticed any real limits on it. And because sometimes it says really stupid things. But it's a very quick model that's always right there. Always at my fingertips. And I'm still Googling things left and right. So good on your Google for integrating the AI into like the place I'm always searching for something. But be like, I've gotten very used to it as a small model. Like I know what to expect from it and what not to expect from it. I don't expect astro levels of things on it. I don't even know what that means. Astro really hasn't impressed me that much. I'm just using it as the example. So yeah, anyway, it's my way of saying like, I'm actually using small models all the time. But again, it's a stupid phomo thing. Why I haven't trusted it to actually work on my source code or things I really care about. Like in a conversation, it's fine. But I haven't trusted it yet. So I really appreciate, you know, I think you're inspiring me to try, try them out. Try try the little guys. I mean, I use Git. It's fine. If it's fast, I can just, you know, get river real quick. And then try a bigger model. And I need to be more willing to fail in that regard. Yeah. I think so. I mean, I've been doing this because I've been doing this. I'm like, just do a work or do a work. Throw away, throw away. Throw away. It doesn't matter. Now, especially my testing, obviously, I'm lucky. I have lots of tokens that I can encourage to burn through. But I think it's like valid testing because I can have this conversation. So thank you, Microsoft. But I think it's really valuable. And I think this has led me to try to figure out how much and how important it is for me to be doing the thing that I have been doing historically and the thing that you're doing historically, which is watching the age and do stuff, even if you're doing one thing. And what I found, Frank, is that in that initial stage where I'm asking me a lot of questions, I want to be there by its side. And then I'm like, go, go off into the future and like go and do stuff. And the GitHub research team just released something called Hydrophusion. And it's not a model. It's an advanced runtime model orchestration that you pick as a model. So instead of picking Terra or Soul or Opus, you pick Hydrophusion. And Hydra stands for something like Hydra, something meaning is an acronym. And Hydra is the monster you cut its head off and a new two new heads grow or whatever. It's Greek myth. It means something though. It's a Hydra is a hybrid dynamic routing architecture. You're going to love this. This is from I'm pretty sure this is from the GitHub team put it here. If you type in hybrid dynamic routing architecture, I think this is from a bunch of people from GitHub. So basically, it's all about, you have this pool of models that you can pick from. And today, developers have to pick like you pick a model, like we've been talking about, pick a reasoning level. What if you didn't have to pick any of that? In general. So the interesting part is, you could say, oh, we have that, right? We have that with Auto. It figure out my prompt and then go pick a model that kind of makes the most sense. So Hydra Fusion, which I'll put a link here to the blog post for you, is a little bit different. And I'll say why this has changed my workflow a little bit because I'm using this as the default. It's only available in the CLI right now, but it is specifically for all intents and purposes, changing how I think about working on features and projects and things like this and doing things. Because what it does is with each request, Hydra Fusion makes a decision of one of three execution patterns and there might be more in the future. So what it does is it goes and the user request comes in and basically you have task routing. Now in this case, the task browsing, which is the Hydra scoring, which we're talking about basically looks at different reasoning and like tools available and it looks at some policies and it's like, okay, looking at this prompt, I have these three different options that I can go through for an orchestration pattern. And what this does is it will pick a single. So singles kind of like auto like based on this, this is a simple one. Like maybe, can you build this, can you build and run this application? You don't need advanced orchestration for that. Just pick a model. Then you have this other one, which is cascade. And this is the one that I actually find a lot, which is basically an efficient model will draft a solution. And then there's a quality gate where another model decides whether to accept it or escalate to a stronger model that then reviews it and puts it back in and then implements it. So in that case, you might have like M.A.I. Code 1.1 Flash and then you might have Opus and then you may come back on it. And then there's a third one, which is really fascinating, which is a critique. This is one model, drafts a result, an independent read-only critic from a different model family reviews it, kind of like a rubber duck. And then the drafting model revises it once more and then it gets implemented. So it goes through the pits, one of these and automatically, right? And then it will go and execute. So it has a share pool or it will go through and figure out how to solve, review and then go through this. Now, the interesting part about this is that they did all this like terminal bench testing on it. And what they found is that the quality is pretty much Opus 5 level on the same tasks, but at 70% lower costs because of how they're doing this and be able to efficiently kind of like we're talking about these models and the smaller ones and the bigger ones. But you don't pick any of the models you're doing. You pick Hydrofugin, it does it. I've been building and deploying full applications, full features, full bug fixes exclusively with Hydrofugin and the cost differential is bananas and the code quality is super duper good. The problem is there's not a lot of things in the UI yet because this is in a, is this a, what is it, a research project research thing? I don't know, it's research preview. So what is cool about this is that it totally works awesome, but pretty much it's like I'm drafting, I'm reviewing, it doesn't output anything, it just says that word. And then at the end, it's all like, and it's like here's everything. So what I found myself doing is spinning up a bunch of CLI stuff and just Hydrofugining nonstop and then coming back to it like later, because who cares? And so what I found myself doing more and more is like not caring until the end because this is doing the planning, the researching through two. You're taking all the things that I'd be doing manually, just let the agent cook and then come back to it. The difference is if I go into autopilot mode and I pick a model, it's that one model and it's maybe doing some of those things. If the model and if the harness decides to do that. However, in this case, it's kind of as if you were driving those decisions smartly for it. And it's been pretty fun to use. It just came out like four days ago from this recording. So it's pretty neat in general. I think Mario put out a video on it with some cool animations. But it is like really neat, which means me want to work with models and agents like very much more asynchronously than ever before I would say. - That's cool. It's a good, it's like upping your game. I have not gotten into any of the orchestration stuff. And I'll tell you why because I have this stupid in the same way FOMO prevents me from using lower models. My knowledge of how context work and how context costs money makes me fear orchestration because it's constantly choosing a different model for every prompt I give it. Then that's costing me big bucks as it has to move my non-cash conversation all the way back around. I'm one of those people who has really long sessions. I don't like start a session, say blah, blah, blah, implement the thing and then close that session and then check on it or anything. I have dialogues that I've been having with AI's for six months now. And that conversation has been compressed 50 different times and it fills up a one-meg token pool instantly. Even the compressed ones fill that up. And so I fear orchestrators mostly because I'm like, how are they? It sounds like it's gonna cost more money. But I like that you said, no, it's actually cheaper because it's picking cheaper models and all that kind of stuff. And it's doing a better process. We've talked about this before. There should always be a critic. The fact that we don't have a second model being the critic is just laziness on our part or failure of our harnesses to be honest. They still have the chat UI model where critics don't come up enough. But it's obviously anyone who's using things for any little bit of time knows a critic step is awesome because it finds so many mistakes in the original implementation and all that stuff. Or like you said, it means you just don't have to babysit it. I'm accustomed to being the critic. Every time it does something, I write a 20 bullet points on how it did a bad job. Oh, it's gonna go back to the silicon and I'm gonna hunt down its family. No, I'm not doing any of that stuff. But I do give a 20 bullet points pretty much the time. And it's not like they learn from it or anything. So it's really just me exhausting my wrist typing all that kind of stuff out. Okay, that's a long-winded way of saying I need to get on the orchestration game. Is this the best orchestrator you've found? 'Cause I know like what were the old ones? What was the orchestrator and Copilot called before? There was like, not autopilot. There was fleets or whatever. Wasn't that kind of an orchestrator? I think that this is the difference is this is like a multi-model reasoning orchestrator. So I think like in the past, right? When I think of fleets and autopilot, they are trying to get to a goal and they are trying to spawn off sub-agents to go execute and do research. The difference with this is I don't call it, it is an orchestrator. It is a multi-model orchestrator. But I think the biggest differential is that it is analyzing multiple orchestration patterns based on what it's attempting to do based on the scoring process. Instead of just spawning off a bunch of fleets kind of we're like, hey, okay, you have an API, you have a backend, you have a front end, you have another website, I will go do three tasks 'cause they're not gonna intersect with each other, right? This is like doing a lot more pre-planning and organization and kind of rubber ducking between the two of it and drafting and critiquing and revising and then implementing. And then I also with the team what they were saying is like they don't just always change the model every single time throughout. Like they're looking at the right model, the cash timing, the busting, is it worth it? These things do need to add tools and not do this. And what they said they also did is like they have it. So they have like toolless contacts. Like so like there's different shared work spaces or these isolated reviews like help optimize cost and then like all these like different things that it's doing. So there's probably tons of stuff happening under the hood in general and there's like a great blog post in there about it and these checkpoints and what it's doing and these costs and how they're hill climbing it and doing all these other things with other charts and graphs that nobody knows but they're there. But it's pretty neat in general. So at least for me I've been telling people, you have to turn it on and experimental in the CLI. You've got to flip on experimental and then go into slash models, do hydrophusion, research preview. But it's been my default in the CLI pretty much exclusively doing stuff. Yeah. I don't like, I think you and I always said, right? Which was like auto mode is like pretty fantastic and there's tons of work happening in auto mode too that they're doing but that feels like the future. I don't want to have to pick the model. And I think that now it's evolved which is, I don't want to pick the model but I also don't want to babysit the orchestration of what said model is doing. It should be able to figure out these things. And when these things are in the box, I'm sure people have built my own orchestrator. This amazing agent plug-in did this thing. That's cool, it's not in the box, right? When it's in the box and that's the default that people can pick and it's really, really good. And then I think that's like a game changer for vast majority of users. So I say give it a try. It's not fast, but it's not trying to be fast 'cause it's doing a bunch of work, right? So here's what's fascinating is, and I need to actually probably test this myself, is going through and then saying okay, let's say this bug fix that I had for the multi-window stuff. Okay, I was doing individual models and reasoning. What if I just gave it a hydrophusion and I was like go do it, right? Yeah. And I bet probably it would be just as good. If not better, it probably would do a lot more critiquing. So I'll do that and report back. So that's a much do list. Probably you'll have to check in tomorrow though, right? With all your infinite Microsoft Cres, I'm surprised you haven't written the James Harness, which basically whenever you type a prompt in, it just sends it to 13 different models. And then there's like a big join at the end and you can just review 13 and pick from the best of the 13 every time. I mean, if I had infinite credits, I'd be doing ridiculous things like that. Well, one of the early things that appears to talk about and other people is you can kind of do that, right? You can basically say, hey, listen, go implement this in three different models and three different flavors and three different styles and go do this stuff. And then you get three different work trees and I'll go and review it. The problem really is like are you going to really review all 30 different implementation models and things like that, right? Probably not. That's the thing. - And there's always pros and cons. Like I've done this with GUIs before. And you're like, well, there's some parts of the GUI I like from that one, parts I like from that one, parts I like from that one. So you can't really do a merge at the end. So fair point, good pushback as the model say. Yeah, it's tough. It's hard to A/B test these things, especially when you're reviewing a code 'cause then it's just exhausting because they will spit out code way faster than you can review it. - Exactly, yeah. Anyways, that's what I got. That's what I've been, that's my week in A/B model things. - My takeaways are try the, try the Dumber models with a lot of thinking. I like it, that's almost the like infinite number of monkeys at a typewriter kind of solution. And then I really do need to out my game on orchestration. I feel like I'm finally doing the thing where I'm using 12 different harnesses of yay. But I'm not doing any advanced orchestration stuff. I will tell you though, I am so over plan files and generating 200 mark down files for something. So I'm not sure I'm gonna try that experiment that you mentioned, but definitely gonna try the smaller model stuff. - Yeah, try it out. And I think also if you had something like super duper juicy, like I think a good example is, you know, I've been working on the performance stuff for like a long time. And I was like, okay, this is it. Let's burn some tokens. I'm like, I'm gonna give in and it was time, right? I was like, let's do it. And instead of like you and I've talked about, it's like you can just kind of go and rev and bullet point this and this and this and this. And I was like, you know what, let's token it up. Let's do it. And it did it, it did the thing. And it's like really fantastic now. I think like tiny clips, you know, these Windows APIs are not elegant. So it did a really fantastical job doing the stuff. So really impressed by it. But yeah, I think give it a try and don't be scared to try little baby models. And I think you'll be surprised. But what I found little models, big reasoning. That's a gold ticket for success. 'Cause they're still fast by the way. So like the max reasoning on baby. Luna is not the same as Max Reasoning on Astra. Those are very different things. But anyways, give it a try. Let us know, try out Hydrofusion. I'm serious, like it's pretty bananas. And also, if you're a Chroma Edge user or VS Code, or as soon as you use Windows 6, try out Bananify. Bananify it online, my official sponsor of this podcast. Heather and I vibe this out, you're on Safari. So I'm not gonna build you a Safari extension, maybe. - Oh my God. - But it's certified. - You've certified it. - It's certified on Edge. And it's being reviewed right now on Chrome. It's in the VS Code marketplace. And it is soon gonna be in the Visual Studio marketplace. And the Visual Studio and Visual Studio Code one, puts bananas and monkeys in the gutter and throughout your code, like which is fantastic. As far as like little pets also and you're like, explore Windows. - So very, very cool. And there's banana themes. You got, there's monkey themes for VS Code and Visual Studio for different themes, yeah. - I gotta tell you about the themes. Like I used to be a super typography nerd. And at some point, someone did this research study on what is the correct like contrast ratios for reading text on screens and paper and all that stuff. And they found a really weird thing. It was a light yellow or faded yellow background with a strong, darkish green as the text, produced a the most comforting reading scenario and the best comprehension scenario. So I love that that's actually what you arrived at with stupid bananas. Sorry, I didn't mean to call it stupid. Silly banana fire that you actually went to what I used to study as. This is the ideal typographical setting for especially long texts. You want that faded yellow background and green text. That's the sweet spot. - It's my default theme now. That is the one that you see on the website is called Banana Cream. And it's very nice. So. - I need a dark mode though. - There are three dark themes in there. There's monkey jungle. - I guess I could do a, I should do a darker yellow theme. Well, I'll take a look at that. So. All right, that's gonna do for this week's merch conflict. People, so until next week on James Monster Magnum, with an eye roll Mr. Frank Krueger over there. I see that. So let's try one more time without the eye roll. And that's gonna do for this week's merch conflict. And until next time, I'm James Monster Magnum. - And I'm Frank Krueger, not eye rolling. Thanks for watching and listening. - Peace.

Podcast Summary

Key Points:

  1. Developers should use low-to-medium reasoning levels in AI models to avoid endless loops and get faster, more actionable results.
  2. High-reasoning modes are valuable for deep research and planning—such as performance profiling or complex feature design—when combined with smaller, cost-efficient models.
  3. Smaller models like Baby Luna and MAI Code 1.1 Flash deliver results comparable to larger models at significantly lower costs, especially when reasoning is set to maximum.
  4. AI orchestration tools like Hydrophusion automate model selection and workflow patterns (drafting, reviewing, critiquing), improving code quality and reducing costs by up to 70% without requiring user intervention.
  5. Over-reliance on extensive documentation from AI models can lead to hidden errors; developers should delete unused or outdated docs and rely on targeted, actionable outputs instead.
  6. The shift toward multi-model orchestration reflects a move from manual model selection to intelligent, cost-efficient workflows that include critique and revision steps.
  7. Developers are encouraged to experiment with smaller, faster models and avoid overusing expensive, high-tier models for routine tasks.
  8. Real-world testing shows that using Hydrophusion and dumber models leads to faster development cycles, better code quality, and significant cost savings.

Summary:

The podcast explores how developers are evolving their AI workflows by moving away from over-reliance on expensive, high-tier models. Hosts James and Frank advocate for using lower reasoning levels in AI, which delivers faster, more practical results and avoids costly loops. 1 Flash can match the output of larger models—especially when reasoning is maximized—while cutting costs by up to 90%.

A key innovation is Hydrophusion, a multi-model orchestration tool that automatically selects optimal workflows (drafting, reviewing, critiquing) to improve code quality and reduce expenses by 70% without requiring user input. The hosts emphasize that developers should avoid generating excessive documentation, as it often contains hidden errors and is rarely read. Instead, they recommend targeted, actionable outputs and simple, efficient workflows.

They also stress the importance of reviewing and deleting outdated AI-generated documents to prevent technical debt. Ultimately, the shift toward intelligent, automated orchestration—combined with trust in smaller, faster models—represents a more sustainable, cost-effective, and efficient approach to AI-powered development.

FAQs

Yes, for most everyday tasks, using default or low reasoning levels is effective and cost-efficient. High reasoning can lead to endless loops or poor outputs, while lower levels deliver results faster and allow you to quickly evaluate and refine outputs.

Yes, small models are highly effective and often outperform larger models in cost and speed. When used with high reasoning, they produce results that match or exceed those of larger models, while saving up to 90% on costs.

Hydrophusion is a multi-model orchestration runtime that automatically selects the best execution pattern—such as drafting, reviewing, or critiquing—based on the task. It delivers high-quality results at 70% lower cost than traditional approaches by intelligently routing tasks between models.

Yes, AI-generated documents can contain hidden errors or misleading advice. Since developers rarely read these long documents, they may unknowingly implement flawed logic. It's best to review only critical parts and delete outdated docs to avoid drift.

Yes, it is safe and effective. High reasoning on small models produces results comparable to larger models, and the cost remains very low. This approach is especially useful for deep research or complex planning tasks.

Use smaller, cheaper models with high reasoning for deep research or planning, and rely on multi-model orchestration like Hydrophusion to intelligently route tasks. This approach delivers quality results at significantly lower cost than large models.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.