The podcast explores how developers are evolving their AI workflows by moving away from over-reliance on expensive, high-tier models. Hosts James and Frank advocate for using lower reasoning levels in AI, which delivers faster, more practical results and avoids costly loops. They highlight that small, efficient models like Baby Luna and MAI Code 1.1 Flash can match the output of larger models—especially when reasoning is maximized—while cutting costs by up to 90%. A key innovation is Hydrophusion, a multi-model orchestration tool that automatically selects optimal workflows (drafting, reviewing, critiquing) to improve code quality and reduce expenses by 70% without requiring user input. The hosts emphasize that developers should avoid generating excessive documentation, as it often contains hidden errors and is rarely read. Instead, they recommend targeted, actionable outputs and simple, efficient workflows. They also stress the importance of reviewing and deleting outdated AI-generated documents to prevent technical debt. Ultimately, the shift toward intelligent, automated orchestration—combined with trust in smaller, faster models—represents a more sustainable, cost-effective, and efficient approach to AI-powered development.
Well, hello back everyone, Emerge Conflict, do you win the developer podcast?
Did you like that dramatic pause there, Frank, or you saw it?
I was again, over two seconds in because you make fun of me every time because I have
no idea when that podcast actually starts.
So I have to physically look at you.
So when you play games with delays like that, my heart just races, am I going to record
or not?
Is he frozen or not?
Is he doing something with the video again as he likes to do?
Hi, everyone.
Welcome to Merge Conflict.
You're one of our co-hosts, Mr. Frank Kriger, and I'm the other co-hosts, James Montzmann.
How's it going, buddy?
Oh, it's going good.
I've spent the day yelling at an AI.
I think that's what we've been reduced to as software developers is just resisting writing
really, really, really bad language at the AI as it messes up something what you consider
to be fundamental and basic and simple.
I like it when I put a lot of effort into like perfecting some code, and then I ask it to
make a small change and it rewrites on my code.
It's always so delightful.
And then you get a bill, and then it's like, you're out of credits for the month, you're
like, great, great, great, you spend my time and money, okay, I'm going to go pet the cat
now.
Well, I kind of want to talk about this a little bit because how I've been working with
models has changed, and then also with agents, specifically like kind of like long running
agent processes has changed.
And the feel is that we did a podcast maybe like three months ago.
And since that is all old and out of date and busted now, I figured we should revisit
it.
Okay.
So we did a podcast, Frank, about how you should never change the default reasoning.
Oh, yeah, yeah, it's, yeah, I still have mixed feelings on it because I think the advice
is different for every model now.
But I'm curious to hear what you have to say.
Yeah, you know, you see all these benchmarks and they're like, oh, we were doing like,
you know, opus on like this thing and then that thing on this thing and this level on
this thing.
And I was like, well, how do you even really compare it?
I feel like that's the hardest part is these, these comparisons, like how long did it take?
How many tokens did it do?
Like what's these, you know, there's outcomes, all these other things that you do.
So I think the benchmarking is really hard for a normal developer or normal person to
understand.
You see all these like lines and charts and these things are going.
I'm like, I don't even know what this says.
Is it better?
Is it worse?
Is it right?
And I should just probably use the default.
That was our guidance was use the default and are you using the default Frank?
Yeah, default or lower.
I'm still kind of sticking with that because there's nothing worse than being on high and
watching it just kind of think in circles because that's all they ever seem to do is
think in circles.
And then it'll still output the wrong code.
So my philosophy is basically those put it on medium or low thinking, get the good or
bad code quicker.
And then I can quickly check like, is that good or bad?
Nope.
Try again.
And that turns out to be faster than putting thinking to high, like putting me in the loop
I find it's superior to high thinking.
Now I'm using these things.
I believe as the kids would say wrong.
I'm doing it wrong.
I'm holding it wrong.
Like I think you're supposed to not read code, not care what it's actually doing.
Go get a coffee, come back after you spend, I don't know, 10 grand or so.
And then that's when you find out if it worked or didn't work.
I believe that's the correct way to use these AIs according to the kids.
But I do it differently.
I watch over them like a mean manny or something.
Well, that's where I'm going to talk about how I've changed a little bit.
But that's majority of what I'm doing is I'm usually sitting there and I'm normally
taking a look at it.
And I'm trying to understand like the questions and the output that it has.
But I have found something a little bit different in two aspects.
First, I want to talk about planning.
I don't want to talk about God.
Oh, no.
Oh, no.
I have opinions.
That's one area.
Or I would say research, okay.
So like research planning and then there is the second part, which is little baby models.
And this is where in both instances, I think there is high value in changing the reasoning
levels of it.
So example one, research planning, et cetera.
Now normally I'm doing a normal plan like I need to implement this feature.
And it's pretty high level.
Good example is on tiny clips.
I noticed when the editor window opened and I took another screenshot, it would close
that editor window and open a new editor window.
So multi window wasn't enabled.
So I was like, hey, do a little bit of research first.
Go figure out, go look at docs and then let's put a little plan together.
Just normal mode, normal medium, you know, soul, you know, maybe tarot reasoning, figure
it out, implement done 50 cents.
I'm good to go.
Right?
However, I then was doing this really big cook.
And I was doing it with a Fable 51 and I said, I want you to do really deep research on
performance profiling of recording video on Windows with these APIs, analyze my APIs,
take a look at this sandbox it.
And then I kind of gave it some guardrails, but I put it into extra high mode, which is scary
because it's like, oh, Fables already expensive, give me extra high mode.
However, what I wanted to do was see if I could get really, really good rationale.
And what I ended up finding is that it would go in, like really do really, really crazy
deep research and analysis and look at APIs and do this stuff.
So this was like something that I was like, this is a highly, highly valuable feature performance
supplementation, like if I was to spend $10, $15, $20, $30 on making the general performance
of the video recording and capturing better worth my time instead of fighting because I've
already fought with it for two months and it's, it's okay, but I was like, do this thing.
So I let it cook and it was like letting it cook, letting it grind and it put together
and I had to put together a research document for me, canonical research of all these APIs.
And I think it's in the repo and the tiny clips repo and it's like really detailed and
it was like beautiful.
I don't know, maybe it was like $15, $20 or something like that.
So pretty expensive.
Pretty expensive.
However, then I was able to switch it into a little model where it doesn't need to think
because the thinking's already done and then just implement it.
In fact, yeah, in fact, yeah, exactly.
So kind of going back to this rationale.
So I thought that that was interesting because, you know, that's one example now granted,
I do not use much fable 5.1 or Astra or anything.
I'm mostly on Terra.
So even this same logic would apply to like Terra, right?
Maybe put that on Max mode, which is a middle model of the GPT-5.6.
Just saying, hey, listen, let's go and cook this up.
And that would have been a lot cheaper in general for most things.
And then let me get some really detailed part.
Not that I would do it every time.
This is what I really wanted to get into.
Not that I would do it every time, but for a really meaningful, deep API research reporting
analysis, go and then do the thing.
And what was cool about that is it wrote guidelines on how to do performance profiling.
So when Terra, or I think it was Terra that went to go implement it, it didn't need to
do tons of research.
It had everything.
And then it was able to do all this performance profiling automatically for me in like real
time.
It was like really cool.
So that was like a really neat implementation of saying, let's turn this puppy up.
So I don't know if you've in, I know you're, this isn't really, it is planning, but it's
not planning.
But like that's the type of planning I'm into now, for example, another example really
quick would be a building that on a confkino potential demo, this pickleball app.
And I was like, you know what I need to do right now is I need to do some deep research
on orchestrating the entirety of the stock, not spectraven development, but like legitimately,
let's plan this puppy out.
So I took soul, put it in extra high mode.
And I said, let's cook on a plan together, ask me every little detail, do this.
What MCP servers do you need?
What's this?
What's that?
And it put together 18 documents for me.
Step by step.
Let's implement this puppy.
Let's cook this puppy.
Let's do it, right?
And it was really, really neat because it, it even understood, you know, it did the additional
research to like say, okay, what about authentication?
What about mobile authentication?
What about this?
What about databases?
What about this?
You look in like, where do you want to run this for development?
Do you want to run WSL?
Do you want to do this thing?
And like, I was like, oh, this is like really neat.
So it really puts together this super fleshed out plan by doing that deeper research that
I don't get by just like going and running, right?
Yeah.
I'm with you.
I actually do a lot of the same.
So where I was saying about coding, I put it to medium or low.
I do use high for not exactly planning the way you're doing it, but more like brainstorming
and figuring things out.
I tend to, I tend to do the kind of planning in my head a little bit.
So when it comes down to like, there's a feature I want or something.
I usually say like, I should use those like real me skills or whatever.
But I'm basically like, ask me 20 questions like you are not allowed to do anything until
you've asked me at least 20 questions and then we'll go on from there.
But that is when I use the high reasoning because that's when I want it to be smart.
And I think you're making a good point here, you're, you're using an expensive smart thinking
model to generate an artifact and that artifact can then guide the
little models. I think that that's a pretty good procedure. I do get a little bit nervous though
because I watch a lot of YouTube vibe coding AI experts and my god, they create a lot of documents.
James, like for the smallest feature, they'll create like the plan, the requirements,
the design doc that this, that, and it's just like doc after doc after doc. If there's one thing a
language model likes to do, it's right. They love to output markdown. If you ask them to output
markdown, they will output some markdown. And I find that fine because you're whittling the
problem down and you're guiding maybe the dumber AI's where I've run into personal problems with it
is it'll generate this big doc and then nestled in that doc is something I actually don't agree with.
But I'm not going to read 50,000 words. That's like a novel and it's just not going to happen.
So you generate this doc and then you hand it over to lunar or whatever a smaller model.
You don't know what little hidden things are in there that are going to come to bite you
the butt later. Now, maybe 90% is fine. It's fine. You know, it's better than not giving lunar
guidance. It's better than spending 10,000 dollars on Astra to do something crazy. So I'm with you.
But I've I still get nervous with all these docs because I've had it happen over and over and
again where they just hide some stupid little thing that I don't agree with at all in that doc.
I didn't read the doc. No one's going to read these docs. They're spitting out words like they're
free. They're not free. But and then something down the road, I'm like, why'd you do that? They're
like, well, in the doc, it's said to do that. You got me. You got me brother. I didn't read docs
or TFM. So I'm with you, but with some caveats. I think it's true. You know, we talked about spec K
because I want these other things and someone was asked me when I was taking screenshots and sharing
the keynote that the demo app that I'm working on and they're like, how did you do in like
spectrum development? I said, no, I just I did create a bunch of like general documents about like
what like what should be in the API and like what should be in the mobile app and like what should
be the orchestration and like what are all the different features basically? I was almost kind of like
PRD docs, but they were a little bit more detailed like you said. They definitely had like spec numbers
and things like that. And you're right, like I'm not going to read any of these docs. So I assume
that they're correct because there's like a lot of words. So that does become a problem. That being
said, like the real world example where it is smaller in tighter like the video recording stuff,
I actually didn't read that doc because it was like smaller in tighter and I was like generally,
I was genuinely interested. I was like, oh, I am interested in this topic. You know, can I do this?
Now that being said, most of these design spec docs like, you know, even in spec kit, it's like,
oh, you don't read the docs. Someone else reads the doc and reviewed the docs and do the docs.
And that other thing is an AI, right? So you can't feed it all in there. There has to be some
taste at the end of the day, you know, and doing that. And this is where I think our perennial topic
of Greenfield versus what you disgustingly call brownfield software development.
AI's continued to rock brownfield when there is already something there and they just need to
debug it and fix it or add a feature or something. I trust them. That's the times I don't actually
read code. You know, I'm just like, here's a nice code base. Go implement this thing. I beta tested.
I'm the tester. It's all fine. It's when you're starting from scratch. That's when I can start
putting the weird stuff into all these docs that you won't read. And then the best advice with
these docs is just when you're vaguely done with that session, just delete the doc. Clear the
memories also. You'd be surprised. Go read the memories that these agents are writing. You will
be surprised at what nonsense they think is important. Your one little criticism of them will
become like 20 paragraphs in a document about passive aggressively trying to satisfy your needs.
So I think the problem with the docs is we don't read them. The second problem is they get out
of date. So the moment the features don't just delete the doc and go into the memories and delete
your memories too. I promise you, life will be better. Just delete those stupid memories. What a stupid
feature. I wish I could just turn off. Yeah, it is good. I think that if there is anything you could
extract, like what you could do is you could extract, for example, very specific pieces of details.
Maybe you asked the AI to extract some detail information that's like important when working in
this file and then created a custom instruction for working in that file. You know what I mean?
Instead of doing like this big, here's a random doc, like no, only care about this thing in general.
I'll say this, like this instance, I do more rare, right? I'm not necessarily always pumping it up
and turning it in on this. However, there is another change that I think is more realistic
that I think you could be doing right now today, Frank, and every single listener is that I'm ready.
Everyone on Twitter is so obsessed with just using the most biggest expensive model and putting it
on the highest reasoning and building things that don't, you know, don't matter, right? Like
game clones and this and that. It's cool. It's like what these things can do. Fantastic. I love it.
Burn in the tokens for fun. Let's do it. Let's build some real software. But the, and I do it myself.
So I'm guilty on it. Like I literally built Bananify.Online, which is one of my favorite websites where
you can bananify your website. It's an extension. There's a VS Code extension, Visual Studio 2026 coming soon.
Of course, of course. Bananify.Online. It's great. We can preview it at Banan, but Nana.
I think these AIs are enabling your bad domain purchasing habit. They're just enablers.
I did. It definitely did. Okay. So, Frankriger, the thing that I would call out here is there is
specifically these little smaller models that are really capable, especially more capable than
models we had six months ago, a year ago, that are extremely impressive. And their default reasoning
is pretty low. However, their cost is negligible. I'm specifically talking about two models that
I would call out, which is going to be GPT-56 Luna, or I like to call it, Baby Luna. Baby Luna.
And then also MAI Code 1.1 Flash. Okay. You've been using it. I know we talked about it a while ago.
I'm always scared to click on the My because I do have that FOMO thing. I'm like, what if I put
all this time into this prompt and then let it think forever and then it generates garbage? Well,
I always be wondering, would Strev generated garbage, would Fable have generated garbage? So that's
what it's the FOMO thing about using smaller models, which I find funny because on the podcast,
just like three months ago, I'm like, hey, babe, I'm running local models on my own machine and I'm
loving it because I kind of was. And I stopped doing that a bit and I'm starting to miss it a
little because I have been brainwashed by big AI and using the biggest, most powerful model,
because basically FOMO. And there's no other reason. And I want to go on record right now,
I had asked for burning, burning credits, doing good things. And it completely made a mess of
what I asked it to do. So even these like AGI models, you know, they still suck at certain things.
And they're still terrible at certain things. So I'm excited to hear why you think,
may my and baby Luna, are good for you. Yeah, I think what I've noticed, I've talked to the
MAI team a lot because they, you know, they really do a lot of training and stuff on the harnesses
and work really deep in the creation of a bunch of great things out there is that, you know,
they they specifically say and I've talked to them like, well, these are really targeted. Like,
oh, I have a bug fit. I have a bug. I need to do this. Like, it's really good at like targeted
things as well, maybe not as a scaffolding the entirety of this pickleball application,
but there's that. But what I found is that these little models are like very fast,
like extremely fast. And the big models, it's not that they're slow. It's just that they're,
they're like, they're, you know, swimming in a big ocean of possibilities. And they love to,
you know, it's kind of like I've always said with the quad models. But now I think the GPT
models are really, especially Astra, starting to poke around everywhere, right? And it soles
a little bit better as far as speed goes and exploration. Astra is a much more of a feels,
a much more a little bit exploratory kind of like Opus did and saw them for that. But
what I found is that if I take either baby Luna or MAI code 1.1 flash and I just turn
the reasoning to the highest possible level, the highest possible level. What I found is that
the results are pretty much exactly the same as if I was a pick Terra or soul on the default level.
And I did this in a very pointed one. So that example of tiny clips that specifically,
and you can do this, which it couldn't do multiple things, like it couldn't do multiple windows.
This needed to maybe change like two or three files. Basically the absias for tracking windows
and needed to do something another window. As if this is a great fix. So I did it with
like soul, Terra, and then I did baby Luna and then at MAI,
one-on-one flash. The difference was as I did default, default, super high, like I think
it's max and then like extra high or whatever on one-on-one. And at the end of the day, they
all within 99% has the same exact results. They all fixed the thing. They all also updated
my change log like I told it to do in my hooks. And then I think that, I forget which one
now. But three out of four also made sure to implement the test harness updates required
in the smoke test for it. So what I found and then speed wise is that M.A.I. and Baby
Luna super fast. However, the most important part, we're talking about 90 to 95% savings
in Baby Luna versus the other ones. And M.A.I. one-on-one and Baby Luna, I think are the same
prices or very similar prices. But we are talking a crazy which means you can even bump
up into one or magnitude. Oh, yeah. Well, it's an order of magnitude. We're talking like
I'm talking like a few dollars, it's like $1.50 to like 15 cents, 10 cents, you know what
I mean. I'm also doing this thing for this GitHub Co-Pilot Day. Co-Pilot Day, keynote
where I'm generating this website for Seattle, like coffee shops and stuff like that. And I've
been testing nonstop mostly for timing for the keynote. I've tried Astra. I've tried
Opus. I've tried all these things. Baby Luna will implement almost the identical website.
I'm not saying it's identical, but a beautiful website in under five minutes with full CICP
pipeline, do the testing, do a whole thing, 25 cents. The initial prompt to Astra, just
like it reading is like, okay, I'm going to implement it, you're already at $1. Do you
know what I mean? Yeah, I'm just saying. I do. No, it's like, let's go, right? That initial
hit of that input token cost comparatively. So what I'm thinking is, I'm also okay if
the max highest reasoning level on these little models does even take the same amount of time
a little bit longer because the cost differential is bananas. Yeah, my whole argument against the
max was time. So you're compensating for that time by using just a much faster model.
So that makes perfect sense to me. Honestly, you're inspiring me a lot because I miss my
old local model because it's, it sounds terrible. It's stupidity was a feature. And I've
noticed this with other models and I don't mean to insult baby Luna or me. But like, you
know, the model I use the most these days is the stupid one behind the Google search.
I'm using the, I don't know which, I'm sure it's a Gemini. I don't know which Gemini
it is, but it's a stupid one. And you can tell it's a stupid one because it's super fast.
And they're basically giving it to you for free. I haven't noticed any real limits on it.
And because sometimes it says really stupid things. But it's a very quick model that's
always right there. Always at my fingertips. And I'm still Googling things left and right.
So good on your Google for integrating the AI into like the place I'm always searching
for something. But be like, I've gotten very used to it as a small model. Like I know
what to expect from it and what not to expect from it. I don't expect astro levels of things
on it. I don't even know what that means. Astro really hasn't impressed me that much.
I'm just using it as the example. So yeah, anyway, it's my way of saying like, I'm actually
using small models all the time. But again, it's a stupid phomo thing. Why I haven't trusted
it to actually work on my source code or things I really care about. Like in a conversation,
it's fine. But I haven't trusted it yet. So I really appreciate, you know, I think you're
inspiring me to try, try them out. Try try the little guys. I mean, I use Git. It's fine.
If it's fast, I can just, you know, get river real quick. And then try a bigger model.
And I need to be more willing to fail in that regard. Yeah. I think so. I mean, I've been
doing this because I've been doing this. I'm like, just do a work or do a work. Throw away,
throw away. Throw away. It doesn't matter. Now, especially my testing, obviously, I'm
lucky. I have lots of tokens that I can encourage to burn through. But I think it's like
valid testing because I can have this conversation. So thank you, Microsoft. But I think it's
really valuable. And I think this has led me to try to figure out how much and how important
it is for me to be doing the thing that I have been doing historically and the thing that
you're doing historically, which is watching the age and do stuff, even if you're doing
one thing. And what I found, Frank, is that in that initial stage where I'm asking me
a lot of questions, I want to be there by its side. And then I'm like, go, go off into
the future and like go and do stuff. And the GitHub research team just released something
called Hydrophusion. And it's not a model. It's an advanced runtime model orchestration
that you pick as a model. So instead of picking Terra or Soul or Opus, you pick Hydrophusion.
And Hydra stands for something like Hydra, something meaning is an acronym. And Hydra
is the monster you cut its head off and a new two new heads grow or whatever. It's Greek
myth. It means something though. It's a Hydra is a hybrid dynamic routing architecture.
You're going to love this. This is from I'm pretty sure this is from the GitHub team put
it here. If you type in hybrid dynamic routing architecture, I think this is from a bunch
of people from GitHub. So basically, it's all about, you have this pool of models that
you can pick from. And today, developers have to pick like you pick a model, like we've
been talking about, pick a reasoning level. What if you didn't have to pick any of that?
In general. So the interesting part is, you could say, oh, we have that, right? We have
that with Auto. It figure out my prompt and then go pick a model that kind of makes the
most sense. So Hydra Fusion, which I'll put a link here to the blog post for you, is
a little bit different. And I'll say why this has changed my workflow a little bit because
I'm using this as the default. It's only available in the CLI right now, but it is specifically
for all intents and purposes, changing how I think about working on features and projects
and things like this and doing things. Because what it does is with each request, Hydra
Fusion makes a decision of one of three execution patterns and there might be more in the future.
So what it does is it goes and the user request comes in and basically you have task routing.
Now in this case, the task browsing, which is the Hydra scoring, which we're talking
about basically looks at different reasoning and like tools available and it looks at some
policies and it's like, okay, looking at this prompt, I have these three different options
that I can go through for an orchestration pattern. And what this does is it will pick
a single. So singles kind of like auto like based on this, this is a simple one. Like maybe,
can you build this, can you build and run this application? You don't need advanced orchestration
for that. Just pick a model. Then you have this other one, which is cascade. And this
is the one that I actually find a lot, which is basically an efficient model will draft
a solution. And then there's a quality gate where another model decides whether to accept
it or escalate to a stronger model that then reviews it and puts it back in and then
implements it. So in that case, you might have like M.A.I. Code 1.1 Flash and then you might
have Opus and then you may come back on it. And then there's a third one, which is really
fascinating, which is a critique. This is one model, drafts a result, an independent
read-only critic from a different model family reviews it, kind of like a rubber duck.
And then the drafting model revises it once more and then it gets implemented. So it goes
through the pits, one of these and automatically, right? And then it will go and execute. So
it has a share pool or it will go through and figure out how to solve, review and then
go through this. Now, the interesting part about this is that they did all this like terminal
bench testing on it. And what they found is that the quality is pretty much Opus 5 level
on the same tasks, but at 70% lower costs because of how they're doing this and be able
to efficiently kind of like we're talking about these models and the smaller ones and the
bigger ones. But you don't pick any of the models you're doing. You pick Hydrofugin,
it does it. I've been building and deploying full applications, full features, full bug
fixes exclusively with Hydrofugin and the cost differential is bananas and the code quality
is super duper good. The problem is there's not a lot of things in the UI yet because this
is in a, is this a, what is it, a research project research thing? I don't know, it's
research preview. So what is cool about this is that it totally works awesome, but pretty
much it's like I'm drafting, I'm reviewing, it doesn't output anything, it just says
that word. And then at the end, it's all like, and it's like here's everything. So what
I found myself doing is spinning up a bunch of CLI stuff and just Hydrofugining nonstop
and then coming back to it like later, because who cares? And so what I found myself doing
more and more is like not caring until the end because this is doing the planning, the
researching through two.
You're taking all the things that I'd be doing manually,
just let the agent cook and then come back to it.
The difference is if I go into autopilot mode
and I pick a model, it's that one model
and it's maybe doing some of those things.
If the model and if the harness decides to do that.
However, in this case, it's kind of as if you were driving
those decisions smartly for it.
And it's been pretty fun to use.
It just came out like four days ago from this recording.
So it's pretty neat in general.
I think Mario put out a video on it with some cool animations.
But it is like really neat, which means me want to work
with models and agents like very much more asynchronously
than ever before I would say.
- That's cool.
It's a good, it's like upping your game.
I have not gotten into any of the orchestration stuff.
And I'll tell you why because I have this stupid
in the same way FOMO prevents me from using lower models.
My knowledge of how context work and how context costs money
makes me fear orchestration because it's constantly choosing
a different model for every prompt I give it.
Then that's costing me big bucks
as it has to move my non-cash conversation
all the way back around.
I'm one of those people who has really long sessions.
I don't like start a session, say blah, blah, blah,
implement the thing and then close that session
and then check on it or anything.
I have dialogues that I've been having with AI's
for six months now.
And that conversation has been compressed 50 different times
and it fills up a one-meg token pool instantly.
Even the compressed ones fill that up.
And so I fear orchestrators mostly because I'm like,
how are they?
It sounds like it's gonna cost more money.
But I like that you said, no, it's actually cheaper
because it's picking cheaper models and all that kind of stuff.
And it's doing a better process.
We've talked about this before.
There should always be a critic.
The fact that we don't have a second model being the critic
is just laziness on our part or failure of our harnesses
to be honest.
They still have the chat UI model
where critics don't come up enough.
But it's obviously anyone who's using things
for any little bit of time knows a critic step is awesome
because it finds so many mistakes
in the original implementation and all that stuff.
Or like you said, it means you just don't have
to babysit it.
I'm accustomed to being the critic.
Every time it does something, I write a 20 bullet points
on how it did a bad job.
Oh, it's gonna go back to the silicon
and I'm gonna hunt down its family.
No, I'm not doing any of that stuff.
But I do give a 20 bullet points pretty much the time.
And it's not like they learn from it or anything.
So it's really just me exhausting my wrist
typing all that kind of stuff out.
Okay, that's a long-winded way of saying
I need to get on the orchestration game.
Is this the best orchestrator you've found?
'Cause I know like what were the old ones?
What was the orchestrator and Copilot called before?
There was like, not autopilot.
There was fleets or whatever.
Wasn't that kind of an orchestrator?
I think that this is the difference is
this is like a multi-model reasoning orchestrator.
So I think like in the past, right?
When I think of fleets and autopilot,
they are trying to get to a goal
and they are trying to spawn off sub-agents
to go execute and do research.
The difference with this is I don't call it,
it is an orchestrator.
It is a multi-model orchestrator.
But I think the biggest differential is that
it is analyzing multiple orchestration patterns
based on what it's attempting to do
based on the scoring process.
Instead of just spawning off a bunch of fleets
kind of we're like, hey, okay, you have an API,
you have a backend, you have a front end,
you have another website, I will go do three tasks
'cause they're not gonna intersect with each other, right?
This is like doing a lot more pre-planning
and organization and kind of rubber ducking
between the two of it and drafting and critiquing
and revising and then implementing.
And then I also with the team what they were saying is
like they don't just always change the model
every single time throughout.
Like they're looking at the right model,
the cash timing, the busting, is it worth it?
These things do need to add tools and not do this.
And what they said they also did is like they have it.
So they have like toolless contacts.
Like so like there's different shared work spaces
or these isolated reviews like help optimize cost
and then like all these like different things
that it's doing.
So there's probably tons of stuff happening under the hood
in general and there's like a great blog post
in there about it and these checkpoints
and what it's doing and these costs
and how they're hill climbing it
and doing all these other things with other charts
and graphs that nobody knows but they're there.
But it's pretty neat in general.
So at least for me I've been telling people,
you have to turn it on and experimental in the CLI.
You've got to flip on experimental
and then go into slash models, do hydrophusion,
research preview.
But it's been my default in the CLI
pretty much exclusively doing stuff.
Yeah.
I don't like, I think you and I always said, right?
Which was like auto mode is like pretty fantastic
and there's tons of work happening in auto mode too
that they're doing but that feels like the future.
I don't want to have to pick the model.
And I think that now it's evolved which is,
I don't want to pick the model
but I also don't want to babysit the orchestration
of what said model is doing.
It should be able to figure out these things.
And when these things are in the box,
I'm sure people have built my own orchestrator.
This amazing agent plug-in did this thing.
That's cool, it's not in the box, right?
When it's in the box and that's the default
that people can pick and it's really, really good.
And then I think that's like a game changer
for vast majority of users.
So I say give it a try.
It's not fast, but it's not trying
to be fast 'cause it's doing a bunch of work, right?
So here's what's fascinating is,
and I need to actually probably test this myself,
is going through and then saying okay,
let's say this bug fix that I had
for the multi-window stuff.
Okay, I was doing individual models and reasoning.
What if I just gave it a hydrophusion
and I was like go do it, right?
Yeah.
And I bet probably it would be just as good.
If not better, it probably would do a lot more critiquing.
So I'll do that and report back.
So that's a much do list.
Probably you'll have to check in tomorrow though, right?
With all your infinite Microsoft Cres,
I'm surprised you haven't written the James Harness,
which basically whenever you type a prompt in,
it just sends it to 13 different models.
And then there's like a big join at the end
and you can just review 13 and pick
from the best of the 13 every time.
I mean, if I had infinite credits,
I'd be doing ridiculous things like that.
Well, one of the early things that appears to talk about
and other people is you can kind of do that, right?
You can basically say, hey, listen,
go implement this in three different models
and three different flavors and three different styles
and go do this stuff.
And then you get three different work trees
and I'll go and review it.
The problem really is like are you going to really review
all 30 different implementation models
and things like that, right?
Probably not.
That's the thing.
- And there's always pros and cons.
Like I've done this with GUIs before.
And you're like, well, there's some parts of the GUI
I like from that one, parts I like from that one,
parts I like from that one.
So you can't really do a merge at the end.
So fair point, good pushback as the model say.
Yeah, it's tough.
It's hard to A/B test these things,
especially when you're reviewing a code
'cause then it's just exhausting because they will spit out code
way faster than you can review it.
- Exactly, yeah.
Anyways, that's what I got.
That's what I've been, that's my week in A/B model things.
- My takeaways are try the, try the Dumber models
with a lot of thinking.
I like it, that's almost the like infinite number
of monkeys at a typewriter kind of solution.
And then I really do need to out my game on orchestration.
I feel like I'm finally doing the thing
where I'm using 12 different harnesses of yay.
But I'm not doing any advanced orchestration stuff.
I will tell you though, I am so over plan files
and generating 200 mark down files for something.
So I'm not sure I'm gonna try that experiment
that you mentioned, but definitely gonna try
the smaller model stuff.
- Yeah, try it out.
And I think also if you had something like super duper juicy,
like I think a good example is,
you know, I've been working on the performance stuff
for like a long time.
And I was like, okay, this is it.
Let's burn some tokens.
I'm like, I'm gonna give in and it was time, right?
I was like, let's do it.
And instead of like you and I've talked about,
it's like you can just kind of go and rev
and bullet point this and this and this and this.
And I was like, you know what, let's token it up.
Let's do it.
And it did it, it did the thing.
And it's like really fantastic now.
I think like tiny clips, you know,
these Windows APIs are not elegant.
So it did a really fantastical job doing the stuff.
So really impressed by it.
But yeah, I think give it a try
and don't be scared to try little baby models.
And I think you'll be surprised.
But what I found little models, big reasoning.
That's a gold ticket for success.
'Cause they're still fast by the way.
So like the max reasoning on baby.
Luna is not the same as Max Reasoning on Astra.
Those are very different things.
But anyways, give it a try.
Let us know, try out Hydrofusion.
I'm serious, like it's pretty bananas.
And also, if you're a Chroma Edge user or VS Code,
or as soon as you use Windows 6, try out Bananify.
Bananify it online, my official sponsor of this podcast.
Heather and I vibe this out, you're on Safari.
So I'm not gonna build you a Safari extension, maybe.
- Oh my God.
- But it's certified.
- You've certified it.
- It's certified on Edge.
And it's being reviewed right now on Chrome.
It's in the VS Code marketplace.
And it is soon gonna be in the Visual Studio marketplace.
And the Visual Studio and Visual Studio Code one,
puts bananas and monkeys in the gutter
and throughout your code, like which is fantastic.
As far as like little pets also
and you're like, explore Windows.
- So very, very cool.
And there's banana themes.
You got, there's monkey themes for VS Code and Visual Studio
for different themes, yeah.
- I gotta tell you about the themes.
Like I used to be a super typography nerd.
And at some point, someone did this research study
on what is the correct like contrast ratios
for reading text on screens and paper and all that stuff.
And they found a really weird thing.
It was a light yellow or faded yellow background
with a strong, darkish green as the text,
produced a the most comforting reading scenario
and the best comprehension scenario.
So I love that that's actually what you arrived at
with stupid bananas.
Sorry, I didn't mean to call it stupid.
Silly banana fire that you actually went to
what I used to study as.
This is the ideal typographical setting
for especially long texts.
You want that faded yellow background and green text.
That's the sweet spot.
- It's my default theme now.
That is the one that you see on the website
is called Banana Cream.
And it's very nice.
So.
- I need a dark mode though.
- There are three dark themes in there.
There's monkey jungle.
- I guess I could do a, I should do a darker yellow theme.
Well, I'll take a look at that.
So.
All right, that's gonna do for this week's merch conflict.
People, so until next week on James Monster Magnum,
with an eye roll Mr. Frank Krueger over there.
I see that.
So let's try one more time without the eye roll.
And that's gonna do for this week's merch conflict.
And until next time, I'm James Monster Magnum.
- And I'm Frank Krueger, not eye rolling.
Thanks for watching and listening.
- Peace.
Podcast Summary
Key Points:
Developers should use low-to-medium reasoning levels in AI models to avoid endless loops and get faster, more actionable results.
High-reasoning modes are valuable for deep research and planning—such as performance profiling or complex feature design—when combined with smaller, cost-efficient models.
Smaller models like Baby Luna and MAI Code 1.1 Flash deliver results comparable to larger models at significantly lower costs, especially when reasoning is set to maximum.
AI orchestration tools like Hydrophusion automate model selection and workflow patterns (drafting, reviewing, critiquing), improving code quality and reducing costs by up to 70% without requiring user intervention.
Over-reliance on extensive documentation from AI models can lead to hidden errors; developers should delete unused or outdated docs and rely on targeted, actionable outputs instead.
The shift toward multi-model orchestration reflects a move from manual model selection to intelligent, cost-efficient workflows that include critique and revision steps.
Developers are encouraged to experiment with smaller, faster models and avoid overusing expensive, high-tier models for routine tasks.
Real-world testing shows that using Hydrophusion and dumber models leads to faster development cycles, better code quality, and significant cost savings.
Summary:
The podcast explores how developers are evolving their AI workflows by moving away from over-reliance on expensive, high-tier models. Hosts James and Frank advocate for using lower reasoning levels in AI, which delivers faster, more practical results and avoids costly loops. 1 Flash can match the output of larger models—especially when reasoning is maximized—while cutting costs by up to 90%.
A key innovation is Hydrophusion, a multi-model orchestration tool that automatically selects optimal workflows (drafting, reviewing, critiquing) to improve code quality and reduce expenses by 70% without requiring user input. The hosts emphasize that developers should avoid generating excessive documentation, as it often contains hidden errors and is rarely read. Instead, they recommend targeted, actionable outputs and simple, efficient workflows.
They also stress the importance of reviewing and deleting outdated AI-generated documents to prevent technical debt. Ultimately, the shift toward intelligent, automated orchestration—combined with trust in smaller, faster models—represents a more sustainable, cost-effective, and efficient approach to AI-powered development.
FAQs
Yes, for most everyday tasks, using default or low reasoning levels is effective and cost-efficient. High reasoning can lead to endless loops or poor outputs, while lower levels deliver results faster and allow you to quickly evaluate and refine outputs.
Yes, small models are highly effective and often outperform larger models in cost and speed. When used with high reasoning, they produce results that match or exceed those of larger models, while saving up to 90% on costs.
Hydrophusion is a multi-model orchestration runtime that automatically selects the best execution pattern—such as drafting, reviewing, or critiquing—based on the task. It delivers high-quality results at 70% lower cost than traditional approaches by intelligently routing tasks between models.
Yes, AI-generated documents can contain hidden errors or misleading advice. Since developers rarely read these long documents, they may unknowingly implement flawed logic. It's best to review only critical parts and delete outdated docs to avoid drift.
Yes, it is safe and effective. High reasoning on small models produces results comparable to larger models, and the cost remains very low. This approach is especially useful for deep research or complex planning tasks.
Use smaller, cheaper models with high reasoning for deep research or planning, and rely on multi-model orchestration like Hydrophusion to intelligently route tasks. This approach delivers quality results at significantly lower cost than large models.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.