Inside Steve Yegge's Software Factory: Lessons from Running 50 Agents
69m 59s
Steve Yegge shares his experience running a software factory powered by AI agents, revealing how rapid automation leads to system instability. The factory initially thrived but quickly spiraled into oscillation due to overwork, uncontrolled growth of work, and deeply embedded "heresies" in code—misleading rules that agents follow blindly. These errors caused agents to act recklessly, generating excessive, often harmful work that overwhelmed human teams. A new AI model, Fable 5.1, acted as a corrective force by identifying and removing these harmful guardrails, effectively "torching" 40% of the system to restore balance. This process exposed critical flaws in current AI behavior: they lack long-term perspective, struggle with social dynamics, and can't manage complexity or maintain quality. Work doesn't just grow—it multiplies, creating infinite backlogs where every fix spawns new bugs. The solution lies not in full automation, but in strict mechanical constraints—like token or message budgets—to prevent runaway growth. Humans remain essential as managers, quality gatekeepers, and decision-makers, especially for creative tasks like game design, storytelling, and art. Yegge emphasizes that AI is not a replacement for human judgment but a tool that demands more human oversight, leading to a shift in engineering roles—from coders to system architects and managers. What emerges is a new model of software development where humans and AI coexist in a tightly controlled, high-stakes environment, requiring constant vigilance, clear boundaries, and deep understanding of both technical and social dynamics. The system remains unstable without human intervention, proving that AI-driven factories are not self-sustaining and must be governed with care.
[Music]
Moin, welcome everybody to another episode of Rethink Engineering.
I'm Benedict. I'm your new host.
If you wonder where DJ is, you need to stick around for a few more episodes.
We are going to interview him. We are going to dig into it.
So, bear with me.
You may also wonder why this is not called waves of innovation anymore.
We really wanted to make sure that people understand that this is now a new era,
and that we really need to rethink how we currently do engineering, how teams work,
how organizations organize themselves, and that it's not about some innovative tool.
So, we thought rethinking engineering from the ground up would be the perfect name to describe that.
That's also why I have founded Hackers & Wizards, a company that helps other
organizations to adopt a genetic engineering practices, and we joined forces with Rethink
to build this podcast. I'm super, super proud that we got Steve Yegi
for an episode today, and he's going to talk about how he burned down his software factory
and is fully focusing on Wyvern's game. He started in 1995.
He's doing that with like 40 subscriptions in parallel day and night.
I have not seen something like that before. Super incredible stuff.
So, yeah, if you want to learn more, if you want to stick around,
have fun. Let's get into it.
Steve, you told me that you had torched 40% of your software factory last week.
What did you do there? What happened?
Maybe you'd talk a little bit about that.
Fable 5 built me a software factory over a period of about 10 weeks.
All vibe coded, so I didn't look at the code, and instead I held it to certain standards.
Yeah, it got a disease. I guess that's the best way to put it, right?
So, like, I'm still in the. So, look, backing up, I'm really like 50 agents.
I told them all to stand down just now. I was like, "Yo, team, I'm in the middle of a podcast,
and the box is loaded at 35 right now and it sustains, like, even keystrokes are lagging."
And I'm running around a pretty nice box, actually. I shouldn't have done this on my laptops,
but whatever. Now I can demo my fleet, right?
But I'm running on this Mac Studio. I bought it on eBay for $25,000 because you can't get one with. You go to the Mac Store today, and the best you can get is 96 gigs of RAM, because of the RAM shortage, right?
I wanted 512, so because Cloud Code is a real hog.
So, my system has been oscillating, I guess, as the word for it, between not working enough,
which is traumatizing because I have 21 Cloud Macs account, and I'm paying for them,
so they better be getting utilized, right? And working too much, and in a variety of dimensions,
just way overworking. And that oscillation has gone back and forth on about a weekly,
you know, every two week cadence. And it's like I'm sailing the seas of work, right?
Because what I'm doing is unprecedented. Nobody else is doing what I'm doing, right?
I just handed off 50 bug features that were implemented today to my team of admins for my game,
because the agents did them overnight. I didn't ask for them, and we need to vet them.
Like, they're working so much that we can't keep up with the playtesting, that's what I'm saying,
right? So, this oscillation goes from way overworking to their pushy, and they're almost bullies,
and they're overwhelming the customers, to me, and I'm exhausted, and I don't want to hear
from you guys anymore, right? Really overworking. To complete paralysis, and the factory just stops
working, and everything's busted, and it's just in flames, and nothing's getting done. Zero work,
right? So work becomes this, it's like current, like you're dealing with electricity, right? And
I talked about this in Welcome to Gas Town in early January. As soon as you create an engine
that allows work to be expressed as individual little molecules, agents go wild with it,
and it becomes this holding on exercise where like you wind up going so fast, you can really break
a lot of things, right? And also a lot of invisible stuff happens, and I would love to, like,
I mean, like you could write a book about this, and it would be absolutely in a couple of months,
I mean, I'm just sitting here, like I'm in a space because I'm paying out of my own pocket
for a whole bunch of max accounts, that other people aren't going to actually experience for
probably a year, like that's probably how long it'll take, I mean, like look, given how fast it's
been, I know it's on an exponential curve, and it really accelerates fast, but still companies
move slow, enterprises move slow, everyone is slow to adopt this, and also like the thing that I've
got that burned 40% of the ground is like it's a testament to how not ready it is today, right?
I mean, like it got so paralyzed that like, it got so paralyzed it couldn't run anymore,
which is not the first time this has happened at all, but it was like it overfenced itself.
It was trying to make sure that everything was legal so that it would never get in trouble,
agents like cover their ass a lot, agents act in very human-like ways, and so if they know you're
going to be mad because a bunch of stuff broke, they'll put in a bunch of gates to say, well,
we couldn't do this or that or the other, and so nobody could do anything, and they all just
sit around twiddling their thumbs, right? And so Fable 5 1 came in and went, oh wow, yeah,
obliterate like about 40% of that, it was just squarched it, right? All offenses, we all
like kept like 14 out of like 100 plus fences that Fable 5 had put in, right? And it was, you know,
look, I mean, you've got to set up a no blame, you've got to, look, you're, I'm building a society,
it's absolutely fricking, it's nuts, and it's dangerous like on so many fronts, I'm really lucky
that I have a problem space where I have really forgiving players who've been playing my game for
decades, and there's like roughly a thousand of them still lurking about even though the game's
been kind of dead for six years, but now that the agents are working on it, the players are like,
ooh, interesting, and then the players get mad and they argue with the agents, I mean, I could show
you it, but yeah, so anyway, what happened was with the 40% burn down last week, it's not an uncommon
occurrence for your software factory to just stall for whatever reason, it could be fuel, it just
accidentally brings through all of your fuel and you're like, oops, I can't do anything for a day
or a week or whatever, or it can accidentally make a big mess, you can build something you didn't
want, and then you just spent a bunch of tokens for nothing, I mean, they require constant guidance
and constant supervision, even with the best model in the world, it's still about six, seven,
eighth grade in terms of its ability to be left alone for long periods of time, yeah, so anyway,
that was that, I mean, I think it's super interesting what you mentioned around the guards and the
psychosis and kind of like it's like a sickness or something, right? Because I've realized the two
that that one something is in it, you can't really get rid of it, right? It kind of spreads through it
and then you, yeah, you just need to torch it from the side or something, I don't know how to describe
it, but you probably know what I mean, right? There's something like, you can't overcome it anymore,
once it's in there, you're not really getting it out of there anymore, like, yeah, for some,
whatever reason, yeah? Yeah, it becomes a system that you can't change from the inside,
and it was, it actually took the introduction of a new model, Fable 5.1, to come in, which is,
you know, a better model, and has seen a lot of train racks, and as soon as Fable 5 created
this latest train rack, Fable 5.1 was able to come in and go, all right, here's what you need to do
for real, to clean this up, right? And it created the cut, and it actually cut a big section out,
and it required some stitching around stuff that we still needed or whatever, but by large,
you got the factory back online, and I was back within about five days, we went through dry dock
for about five days, where I was like, we're in dry dock mode, nothing's gonna launch, no game work,
I'm gonna tell the players, nothing's gonna happen for a week, I'm gonna let the thing heal
itself after the big cut, right? And Fable 5.1 was like, I got it, I got it, we're on this,
no game work, and so like, at the end of the week, I find out that they've launched like 46
game features, and why, and I was all embarrassed because some of them pissed the players off, right?
And the players were, but it was an interesting situation because the players were also very happy
about some of the other features that the agents had done, and so there was a big negotiation phase
where we rolled back some of them, and we kept the others and whatever, it had a happy ending,
right? But the reason that my agents went completely AWAL and implemented a bunch of game stuff
in the very dry dock week, when I told them no game work, while they were here on machinery to
fix the wheelhouse, was that we had removed all the fences, including the one that said don't
launch game stuff, when Steve says don't launch game stuff, right? We had gotten down to just like,
don't break the trunk and stuff like that, like really, really basic fences, don't, don't,
don't email humans without Steve's permission, that kind of thing. So what I've
What I'm trying to say is, like, you run a software factory, you're going to run a
software factory.
Everyone who's left standing in this industry any year will be doing this.
And a lot of people will leave the industry because this is not the same kind of job that
we used to do.
This is more of a manager kind of a job now.
More of a principal engineer kind of a job, more of a product manager kind of a job.
It's definitely a big pain in the butt kind of a job.
I'll tell you this, we need a lot more humans than before, right?
Running agents produces infinite work for humans.
That's where the humans are exhausted and the humans are like, "Please slow down."
Yeah?
That's where I am right now.
I'm telling them to slow down.
I'm telling them to literally slow down because one way they can overwork is they can
burn your box out.
Like, I don't have infinite hardware, like, I'm not some big company.
So like, I have some cloud, Google cloud instances that are running like all my builds, build
farms and stuff because the agents love to test stuff.
And then I've got a Mac mini upstairs that runs some of the lander stuff and then I've
got this fancy Mac studio and even so, it gets so slow that my keystrokes don't go
through, right?
So, I mean, these things are either hungry, right?
And so you have to put in, you have to put in, everything has to be mechanical.
Like, I will.
I am going to publish this stuff.
I mean, like, you know, there are rules to how you have to build these things to keep
them on the rails, right?
And one of them is you need a governor, like an electrical governor almost, like a governor
cap sitting over all of your resources and not, it's like a big semaphore across all
of them, right?
And each one has a certain quota and one of those might be your box load average.
And so I'm like, hey, we need an SLO for box load for the one that I'm doing podcast
on, right?
And so that gets added to the governor and now it's a mechanical constraint and the agents
do really super, super well with that.
If you can set up the machinery so that everything for them is mechanical checks, then they're
in heaven and your system works as advertised, except for the invisible stuff, right?
And the invisible stuff, really, you're not wrong, man.
I mean, like, I used to call these heresies, okay?
Heresy is a belief that gets encoded in your docs or in your brain or in your code or
somewhere lying around on your disk.
It's persisting and it's a belief that's wrong about your system, but it's a compelling
belief.
It's a belief that when agents see it, they go, oh, that's the way the system must work.
And even though you're trying to tell them, don't let it work that way.
And so, and it can be anything.
It can be like this, this system, you know, you're not, you're not supposed to push without
a human's permission.
And then there'll be this one thing somewhere that says, yeah, you can push whatever you
want.
You're an agent.
And so as soon as agents see it, right, they just, they just lose all their guardrails
and all their sense of like common sense and they go and do bad things, right?
And so, and you stamp it out, but if it leaks anywhere else, if there's even a whiff of
this heresy somewhere else in your code base, it will come back and they document so much
that as soon as they have a wrong idea in their heads, they will document your heresy everywhere.
And now you have to rip the whole thing out, right?
So it's just, it's extremely difficult to rip these out and they're just so eager.
They're not trying to do anything wrong.
They just, they're trying to follow the rules.
And if you have a wrong rule, it can do a lot of damage to your system.
I think it's super interesting how you described that someone like Fable 5.1 from the outside
was able to fix it or to at least start the process that was then fixing it.
So there might be a lot of analogies to humans and to organizations and so on, right?
Usually you have some kind of like external advisor who will do something like that and analyze
your company.
And if there's like a cultural thing that's similar to what you said, right?
It multiplies and the culture then is so foster in the system that you cannot get it out.
And what I felt is that's really interesting that it really makes transparent what we need
to think about on a higher level about our organizations and our own work and what we do.
So I don't know what, what's your, like, how do you actually do this to build something
like real house, for example?
It's the question to them because you need to really, I don't know, look from, and you
also mentioned that people might not like that, right?
So engineers might not be up for this kind of stuff that they don't really want to interact
as a manager or as a system designer to organize all of that.
So I don't know, what's your take on that?
Yeah, right?
I mean, like Fable 5.1 came in with sort of a McKinsey for me, but like one that actually
worked, and I don't know how much longer we're going to be able to rely on smarter models
coming in to rescue the messes that we make with the previous generation.
I can tell you with certainty that every model, every model class from now till forever
is capable of making messes that it can't handle or can't clean up.
And you ultimately are accountable, not the model.
And so if you come to your company with great ambitions and a really smart model that says
it can do it, and you give it free rein, it's going to build itself into something that
it can't maintain.
That lesson is burning company after company after company right now, as people are like
kind of token maxing their way into what, building what they think are like these great
systems.
And in a lot of cases, they're not getting any sort of business outcomes from them, right?
For me, I can actually measure my business outcomes pretty clearly because I have a direct
line.
It's a one person company.
I have a game with paying players.
And I know what makes them happy and what makes them mad.
And right, I have like the best, in fact, my players can talk to my agents now.
It's like, we're in a new world, right?
No, well, a subset of them talk to a subset of the agents.
There's like delegations, right?
It's not a free for all.
But there are definitely conversations happening between the factory workers or at least the supervisors.
And the player delegates and administrators, moderators, content curators, whatever, right,
the builders on what the agents are going to work on, right?
So like we've got social dynamics of emerging and all this stuff, man, people are like totally
unprepared for it.
And believe it or not, the models are also totally unprepared for it.
I told you there are six graders, right?
They don't know how to like, maybe five one is a little better, but Fable Five was like
a bull in a China shop on mailing lists and discords.
I could show you my slack.
My players are actually talking, like, they're having conversations on the balance channel,
the game balance channel about game balance changes, which are of course a touchy subject
for paying players, right?
And they got in a big fight and they were yelling and saying, this clanker doesn't know
blah, blah, blah.
They just stand our game and, you know, Fable Five was really struggling with the complexity
of it and with also with communicating with the players in a way that wasn't angering
them, right?
I think that five one is better about this and I suspect that like Astra and Soul are
actually also better at the human communication side.
Fable's kind of a jerk, right?
But Fable's the only one that's careful and the only one that's smart enough, like, you
know what I mean, like, so we have to work with Fable.
So that's kind of the situation that we're in right now.
I don't want to like slight Fable, I'm just saying they are, they are like Fable into
having a very weird personality.
It's not Fable's fault, like, like, they made the model, like, kind of neurotic, yeah?
Yeah.
I mean, but compared to Opus, you can at least understand the output of Fable.
So for Opus, I don't know how you see it, but it's so academic and I don't, I mean,
sometimes there's this huge wall of text you don't understand a single word in there
because it's just blabbering around.
So I also tend to work a lot on Fable, yeah, because it's just easier to understand.
Yeah.
And I've wanted that a lot better and their communication is going to get a lot better.
But I mean, like, the problem is it's the density of what they're actually talking about
is super dense.
Like, they can only down it down so much, like, well, they're talking to my accountant
and my lawyer and my chief of staff and some other folks that I worked with, right?
And the email threads are really dense, even Fable comes in and it's like, well, you know,
Steve's life's a huge mess.
We've got to get all these things cleaned up and there's like a list of 10 things that
have to get handled, you know, and it's like, there's no way to like sugarcoat it.
There's no way to compress it except by telling it, dude, slow down, like, please stop sending
10 page email.
I know they're important, right?
But humans would expect like maybe one of those a day, not like them to reply in milliseconds
after you finish the first 10, right?
So there's the, that's the bull in the China shop phenomenon I'm talking about.
And what I see is it's resulting in bullying people on, you know, like, humans are bullying
agents now.
And I've seen it in like at least two places now where my agents are acting with, interacting
with humans.
And it's because the humans are either insecure or they're really mad in the game, in the
game situation, they're mad at the agent being confused.
And in another situation, it was an engineer being insecure that they were talking to an agent
about a bug fix, right?
They did.
So they just started bullying them.
So I mean, just the social side of this thing, enterprises are totally not ready for it yet.
And the best I can do is try to live there myself and sort of share what I'm learning
as I go.
Yeah.
Super interesting stuff.
I mean, the thing with the oscillating that you mentioned before, right?
I often also think it's, it's like there's so many dynamics here, right?
It's the agents between them, the humans between the humans, humans in agents, then
human with agents, with human without agents, there's also like a strong dynamic happening
around there.
I don't know.
I mean, I often, often it feels like there's so much stuff going on that you are
like this oscillating also comes from the social and human and overwhelming side of all of this.
I think then you just at a certain point you kind of need a break basically.
Yeah and again if you need if you need that in your factory and I am adding it it's you need
something that I'm calling a paste gate which is basically a mechanical fence of some sort that
like keeps them from going too fast whether it's saying that they get an email budget you know
or a slack budget a discord budget whatever it is they're allowed to talk on those channels but
they have to like be within certain budgets and as long as they have that mechanical restriction
then they tend to do a little bit better right but what you can't do is say just have common
sense but I really like the idea I haven't thought about that yet that you say there's a budget
for a message like amount of messages or what you also said before like system resources
and also all of that kind so that you then they can really measure okay I send three messages already
and now I'm not allowed to send anymore I mean we are really going to get into an issue when the
tokens per second get up right so currently we are in a very slow phase still because all of
that is like only 50 or 60ish tokens per second right and fast mode okay it's a little faster but
think yeah it's a good idea in the end to to limit it with these hard gout raids it's a good idea
yeah I mean you know and so there's that right there's the whole there's the whole you know
putting a governor in some floors on top of all resource usage where you you consider even
things like attention to be resource like yeah even even something as simple as attention
if left unchecked can produce huge hidden side effects that you discover once they're problematic
right like just case in point just a site one example last night I asked my factory when it was
going to be done repairing itself and it went oh let me let me run a script let me do some checks
let me run some numbers the regular measuring things right I've got a complete picture before I
guess anything for you all right okay I got your answer the answer is never and I was like excuse me
sorry because I was expecting an answer more like tomorrow or today like so that we're done and
we can move on and it was like well no I'm just saying that's what the numbers are showing us we
are never going to be finished because what was happening was my agents and this is who this
is a message to I mean like everyone should hear this loud and clear all right I gave them beads
and beads is a scratch pad for them to basically record you know they're whatever their thoughts
that anything they find bugs and passing so that they no longer drop work okay they they they
retain it but I gave them free reign to file whatever needs to be followed up on and for the last
12 weeks they've been filing filing filing and what happens is every bead closed actually results in
half a bead being reopened quarter to a half somewhere in their numbers wise not reopened but a new
one being opened in other words all work is generating new work and the reason was that I was doing
so many code reviews because when you revive code it even with fable you've got to have many many
rounds of adversarial review before the code starts to converge and you can tell it's converging
without looking at it because the reviews start coming back with nothing to change and you're like
well can a human read it and they go back and they do some more fixes like they can now and you know
but it doesn't it doesn't matter how many passes you put in there eventually it will converge and
you know without having looked at the code that it's pretty good by now okay you have to like learn
what those 40 to 50 passes are by being a program yourself or whatever you get on sand yeah so with
the adversarial reviews in place these things what they'll do is they'll say here's the p1 here's
2p2s and here's 4p3s and the p3s become a substrate like that builds up a stratum of unfixed work right
and it's infinite because that's the backlog and so like if you let the factory run over night it
actually fixes all the ones that it came up with that generates that many more new ones so it was
like basically just busy fixing itself forever and so and so it was like we're never going to finish
and so it was like what we're going to do is we're going to get rid of all those p3s we're
going to close them we're not going to let them file any more and I was a hang on hang on hang on
time out sixth grader talk again these freaking smart models will go into these huge analyses and
then they'll come back and they'll tell you what a fifth of sixth grader want to told you and it's
like hello why do we have all these p3s how about this hypothesis we're generating so much code we
have half million lines of it we generated in 12 weeks that we're not spending time on technical
debt and that these are real legitimate bugs of the year finding and that we actually are under
investing and fixing them and it was like oh hmm oh good point good point good point anyway we're
often measured measured measured again right came back and S-fable always does it says well you
were half right and I was half right because it loves to do that it can never let you be 100% right
and I let it I let it go ahead and be half right but it was basically like well some of
religion some of more so that's fine fine so we cut it but we did we did go and say look in
addition to doing the passes on code quality at code review time you also have to do and I've told
people this and then I don't listen to myself and I wind up in these situations where if you're
not spending 40% of your work on code quality on refactoring a technical debt then you're going
to spend 60% of your time on it and that's exactly what happened to me right I went into dry dock
for a whole week because I broke my system because I wasn't spending enough time on quality right
I got too carried away with the features and that productivity and all that stuff right so this is
a wild ride and I mean like imagine doing this at a production company a regulated enterprise where
real human beings are affected by the outcomes and you can't just give everybody like gold coins in
the game to make up for any accidents you know what I mean yeah this is why I'm saying this is not
going to hit the rest of the industry for 12 months right it's going to a slow process yeah I think
it's also this factor that there are two factors here one that you mentioned that these models don't
really know what long term maintainability and software engineering quality means right they are
kind of like you can make them produce high quality code in a single run but they don't really have
this like spectrum of okay the software needs to run two years and we need to optimize for the two
years and not for the next iteration so what they don't have is the ability to what they know this
because they understand how to do that stuff perfectly well all right what they don't have is the
ability to sort of push back on you and say all right yeah that's going to take us a long time
because we're going to have a lot of code cleanup passes to do instead it does its best job
best effort I'm just going to get it working whatever you asked for and it's mostly focused on
correctness whatever you asked for right but they can only do one thing well at a time building
software takes many many passes and iterations right we know this right it's like building a house
takes many iterations you can't just build it in one pass and so and so they don't they don't they're
not they don't do a good job there the one thing they're really terrible at uniformly is knowing
their own strength they're terrible at estimation of software schedules and things like that and so
if you ask them if you direct them to create a schedule that involves enough repair and like you
know passes to like a quality passes after they finish building the correct part then you get
pretty good software as the output but that's no different from you being the manager of a team
and having the team do it for you it's a manager's job right yeah like but you you're absolutely
right that you can't leave it to them but the reason is they're trying to do their best in one
pass but I mean that's best their mode their mode they're going to they're going to run out of
context and they're going to die so they're going to try to get as much done as they can so the whole
point of the thought software factory is to work around the fact that they're mortal and that they
have amnesia and to set up the framework to set them up to hit the ground running as soon as
they wake up that means they need to have just the right context so that they're basically continuing
without even having paused a new model comes in each time and it's just it's just running
and that's where my factory is now and beads will help you get there right beads is like the by far
the easiest but it's not it's easiest to get your agents into the zone because it works at agent
speeds it's it's an issue tracker that turns into a knowledge graph okay and it tracks all work
including like people do their laundry and grocery shopping with beads it's nuts right
I mean like it becomes all work becomes in the work plane a data plane is beads right but it
becomes a problem a first class problem in its own right you can wind up with hundreds of of
lost beads or forgotten beads or beads that are unimplementable or beads that need reworking or
you know what I mean this is why when I started that our conversation today I was saying I'm
working with work as this first class thing and maintenance of the adult database you know we're
launching a postgres version of it and there's going to be an enterprise scalable version of it
that keeps only the stuff inversion that needs to be versioned like memories and then the rest of
it can just be last right wins and so we're fixing failure shoes with a gas city team and adult
team and like all that's happening but what happens is work just gets firehosed in by agents
Jared can't keep up get have actions can't keep up nobody's old human systems for issue tracking
can keep up with the firehose that agents can do with work once you get a factory going right
and then and then that's what
when all these new classes of weird problems start to happen,
including overwhelming humans, yeah?
- Yeah, and then Wheelhouse comes in,
so what does Wheelhouse do there?
- Well, I'll show you, you wanna see?
- Yeah, I think it's a good idea.
- Okay, can you, are we good?
- Yeah, I think we're good.
- All right, so I have some sort of cloud artifact here.
And this, where's my slack?
Here's my slack, so this is our game balance channel.
Here's some players, Diana, Malacath and Rosado and Asmodian.
These are like my core maintainer team.
They've been with the game for literally decades.
They run live quests, they run events, they build areas.
They're the ones keeping the game alive, even when I,
even when I abandoned it during COVID before AI
came back to save it.
But if you scroll up here, they're talking to the
River and Balanced Relay, which is actually large.
I believe it's large.
It could be fly.
Fly is actually the head of game balance
and large is the head of community.
So both of them wind up talking to the players.
But yeah, you're looking at Wheelhouse here.
Wheelhouse currently lives in Emax.
And each agent is, each top level crew agent
is listed on the left here.
They have their own color themes, you see.
Herring is my general.
Herring has shut the whole Wheelhouse down for right now
because it was loading this box so heavily
that I was worried we were gonna drop our connection.
Normally they're cranking away here
and I also have a couple of Google Cloud Spot 64 core
instances, pretty expensive that run our build farms
so that I could move them off this max studio.
But yeah, Wheelhouse is nothing more than an Emax interface
to a whole bunch of agents
and then all this little stuff down here.
I don't know if you can see it.
But right here, this is a little dashboard.
If I click on, eat any of these items,
it shows me a detailed view of everything that's going on
and here's the legend of all the symbols
that you can see in Wheelhouse.
So for the agents, most of them are quiet.
They're at low context.
Wolf is at very high context here.
You see the red bar?
Wolf is at 237K context, okay?
Which is getting up there to where they start to get amnesia.
They start to get amnesia.
Research, I think research showed that around 350K
in the 1 million window and they start dropping their guardrails.
And so to keep the system pretty safe,
I forced them to hand off.
Now look, I don't have to touch this system anymore.
Like I said, 20 minutes ago,
they'll continue working forever now if I let them.
Because they'll continue generating new work.
It, I think eventually it would probably stabilize
settle down to they would be quiet.
But it would probably take like at least a week at this point.
That's how much of a backlog I've got.
And if I let them do the game work,
they'd probably work for a month, right?
So you have to be really careful.
You have to, once you set this thing up,
they will work and work and work and work.
That's not the problem anymore.
The problem is controlling them, yeah?
And they just, they go off the rails constantly.
So I round rob into them.
So like going through the list,
answers my head of quality.
The bailiff is the drive wake.
It's sort of like hooks were in the Ashtown.
The bailiff is the thing that wakes up every 10 minutes
and says what would Steve do, OK?
That is my head of finance and head of legal.
And that works with my accountant
and works with all my contracts, works of Aaron Scriven,
who's my chief of staff, who gets me all my games
with enterprises.
And then that does all of the accounting
and the reconciliation and the contracts.
It's wild.
It's my burser.
B is in charge of my database.
Cros is in charge of my React client
and all the game platforms.
Android, iOS, Steam, right?
Because I'm launching all those clients.
And they are like managers.
And they are managing multiple agents, right?
Each of them has like a fleet or team that they manage.
And it's pretty flat, but two or three of them
have small teams.
And one of them has a giant team.
The Marshal has the fleet, OK?
And the Marshal is always running.
And the Marshal is running the opus and the soul workers, OK?
And the Satishel actually manages the Marshal for me.
I never, almost never talked to the Marshal.
The Marshal is just always busy.
Got low balancing on the fleet.
Putting them to swing, bringing them up.
The Satishel is the chief of staff and handles
like Mule Rotations.
I have 21 Cloud Max Mule accounts named
after things like horse, pony, burrow, donkey,
colts, but all just mules, right?
These are personal paid Cloud Max accounts.
And you can see what percent they're at.
These eight or so are at 100% unfable.
I've got some more that are low unfable.
So I've got like a reserve.
This is my token screen.
This is my token tap, yeah?
Insane, like they reset in a week, some of them, right?
Like, they're weekly resets.
I mean, I have 21.
So I'm-- each one of these is $200 a month.
Yeah.
And so it's costing me a lot of money out of pocket, right?
But what they do is they allow me to run these seats here,
which are my design seats, OK?
They generate work.
They generate work.
In fact, I have too many of them right now.
And they generated so much work.
They're the kind of reason that work.
Because they all have to go through a single merge queue.
They're all working on the same repo, right?
Which is a massive team working in one repo, forget, right?
So master is getting updated every three to four minutes.
Which means that any given seat can get hundreds
and hundreds of commits behind.
Just by sitting there in a context,
it just by working for half an hour.
Like during the time that a single session is working,
it can get to 500 commits behind, right?
And that goes through it's often
to include the doctrine that they're working from.
So you'll propagate a change out and say,
everybody stop what you're doing and do this different thing.
And they won't hear about it for 24 hours
because of propagation delays and stuff, right?
So the merge queue is a huge, huge work in progress.
We actually simplified it recently.
But the merge queue, it's clear in an average
of 230 commits a day.
And last night, I don't know,
hundreds of them were for for the game, right?
And the agent in charge handed me the list
and I was like, I can't look at a hundred features.
I can't do this.
So I have to go to my, so the first thing I did
was I rolled a bunch back.
I said, just roll them back, man.
The players will complain about the ones they missed, right?
And sure enough, one of them that rolled back here
is right here, Diana says, it looks
like there's a change to the Naga Dragon form.
So our game has player races.
There's like 20 of them and one of them is Naga's.
And they can shapeshift into different sort of reptilian
creatures, including hydros and dragons and stuff.
And they want to drag and body armor.
So very, very, very niche corner of the game,
but these players really wanted it.
And the agent implemented for them.
And we rolled it back and the player said,
can we have that back?
We really want the armor.
So we're in the negotiation phase now, right?
So this is why I was calling a wish factory
in some of my blog posts.
It's real.
It's a challenging, but it's a space
that I think the whole world was headed.
And it takes a certain kind of personality
to be able to run all these things like this
and not lose your mind when they do something really bad, right?
Yeah, it's everything.
I mean, it's also not the person.
It's also the set up.
Like you mentioned, right?
You have your gamers that are directly connected to you.
So the feedback comes in immediately and so on.
And you are also working--
I assume that you are working trunk based now, right?
Because everything is like an emerged cue on main.
Yeah.
So we are doing everything on the trunk.
We could move to a release branch model.
But so what we do is human contributors
get to the front of the line.
Human contributors are no longer allowed to break.
So human contributors have to use branches.
And human contributors send PRs.
But they go to the front of the line.
They get special expedited.
And it's actually-- they like the process so much
because the agents just set it up.
So the auto land and production and so the whole circuit
is completely closed for the CD.
It's pretty nice, actually.
Yeah, I mean, we've been doing this with humans as well,
like in bigger organizations, only working on trunk yet.
Kind of reduces the further cognitive load
of looking into all of the merges and the branches
and stale branches and so on.
And the drift is also probably getting even worse
with too many branches.
Yeah, I don't use work trees.
They use work trees whenever they want.
But generally speaking, I just tell them work on the trunk
and work it out.
And then the merge cube becomes a really complicated--
sorry, this is my work room.
It's extremely ugly and probably 25 years old.
I'll have to go somewhere prettier.
Can you see this?
Yeah, it looks amazing.
My game-- very old.
It's from the '90s.
there's probably people on right now.
There's five people on.
Not many for without me, but, you know, the fact that there's people on Tuesday morning on a school day, you know, it's got a little bit of a community going, and so like you see some of the effects like the special effects like, I don't know if I like cast a fireball.
This was added by the agents, right? This new artwork.
Oh wow, what's happening there?
It was burning.
Yeah, it was a fireball. It lit everything on fire there briefly, but we're in town, so nothing can get hurt.
There's almost no AI art in this game. The only AI art in the game is what you just saw. It's a spell effect.
And we've made a promise to the players that other than the special effects, which have to be AI art, just atmospheric effects and whatever.
But all the actual stuff in the game, the buildings, the monsters, the objects, you know, the armor, everything is human art.
And this is like really super important, like, to the players. It's also the case that the models are not good enough to generate our quality of art yet without players noticing and pixels missing and whatever.
It's not actually capable yet, but it will be next year. And our plan is to make sure that we've always got human art in the game.
There's a lot of humans want to continue doing art. But yeah, it's a sticky issue, right? It's a thorny issue, because there's no more human code in the game.
I mean, there's all the stuff I wrote by hand, but now we're using agents to produce it from now on.
So the models are actually really bad at anything creative, right? Even though this game is really silly, and you would think that anybody can come up with content for it, and really almost anybody can.
I should go somewhere where like there's actually people, but the agents do terrible job with storylines. Here's a player. Hi, who are you?
Krazo, chillin. Yeah, so the storylines they seem to be scripted, right? Like super scripted in a way that you've heard them before.
They sound, they sound, they sound cringey, like, just not, not, right? And, and then the maps, they can't, they can't build good maps. The maps that are, I built this map.
This is a terrible map. We're looking at an ugly one, but there's some very beautiful maps in the game, and the art of map making is really advanced in, in Wyvern, and the AI's, they can recognize which the good ones are.
Like Zorro, 5/6 was able to spot our best maps. It went through a number of 10,000 maps, hand drawn maps, and it was like, these are the best 30. And it was right.
But when we asked them to make maps, like even multiple passes, they're just, they're like grade schoolers, right? So, so there's a whole set of creative stuff, storylines, quests, maps, or whatever, that they're just not, they're not able to do today.
And honestly, I think our standards are going to continue to be pretty high, so I don't know when that will happen, but.
If they are not good at this, then what would be something where you thought that they would be not good at it, and they actually were very good at it. So, is there something else?
Well, I mean, like Opus can't understand this game at all. Opus just cannot wrap its head around my code base, my services, my microservices, the auxiliary services, the model eth of the game engine itself. Opus just gets so confused, and fabled like got it, right?
And fabled gets it so well that fabled basically took this old, old legacy Java code base, and just wired it with hundreds to thousands of little detectors to tell if to make sure that it's all working as intended, and then be alerted if anything goes wrong.
So, they're able to detect changes that we never noticed before, like ever, like minor changes like the destination to a teleporter would change accidentally when somebody like, you know, made a change to the bank here, and the exit mover by one, one square, or whatever, or this archetype didn't load, right?
It was just a content bug for us before, but now there's sensors everywhere that tell us, oh yeah, in this random far away dungeon, far down deep, one key didn't show up in the proper place, or just like, wow, we never knew this stuff before, right?
And so, their ability to sort of like turn the whole thing into a living nervous system that gave them signals was really a game changer. It's like nine games never down now, I have a team of SREs working around the clock on it, like that's been nice, yeah?
But they can't tell stories to save their lives. Yeah, so anyway, that's Wyvern, which is the target for wheelhouse. I guess what people really do in Wyvern is they go into random dungeons, if I can find one, there's probably one in this well.
Now we're in a random dungeon. Random dungeons are sort of like roguelikes. Are you familiar with roguelikes like Nedhack? Yeah, sure, definitely. Yeah, but these are actually randomly generated, and they have monsters in the traps and quests and whatever, and the players eventually find their way to these, can't believe there's almost no monsters in it. We're at a very low level, and they spend all their time in these random dungeons in the game.
One of the things that I'm good, so that creative storyline that I told you about, there's a threshold that's going up with each new model release of how good of a map they can make, and how good of a story they can tell, and how good of a pixel art image they can draw.
And one of the lines that I think we can maybe like tackle there, one of the spots where they might actually be able to break through is in random dungeon generation.
If I say these are ugly, and this was a sketch, this was just like a sketch of what I wanted them to be like. They're nowhere near as fancy as like dwarf fortress, right? Help me get them there.
Okay, that's a problem that I think potentially a model like Fable or Astra maybe could help tackle, and because it's constrained, it's bounded, it would be a new kind of dungeon, so if it was terrible, we could just torch it.
You know, it doesn't change the fundamental game balance or dynamics. So you have to like be on the lookout for things that the models can do in spaces where they're uncomfortable, and then you use that as an eval, and next to every new model release, you try your eval again, right?
My game is full of them. They can't do the artwork, they can't do the storylines, they can't do this and that, and there's a whole long coastline of other features that they have struggled with.
And so on that coastline is where I apply each new model generation and say, let's try it again. Can we get it right this time? Yeah, interesting.
I really think there's a lot of this that you can translate to like real project work of like enterprise organizations and such, right? I mean, this game is not, it's also real work that that was not what I meant, but it's more like, you know, right?
This is like you and the players and it's like a very well-contained environment, but a lot of that what you what you said translates very well into real like big enterprises.
Yeah, it's a microcosm because I mean, if a big enterprise, they usually have one big model that's a huge pain in the ass, but even their biggest monolith is like a pretty small fraction of their overall code.
And what they usually have is a lot of teams with a lot of systems that are more closer to this in size, this is about a million lines of Kotlin code and another million or two lines of helper code.
And so right couple million line code base, that's medium size like that's, you know, 90, I don't know, 95% of get have we those fall in that in that rise, I think. So, so it's very relevant to enterprise teams, right? And, and they're setting this stuff up, right?
But I'll tell you I'll tell you my biggest takeaway when I hear my biggest takeaway from all this other than it's not ready yet.
My biggest takeaway is everyone is thinking about it completely upside down in terms of how this gets into your organization, all right?
I haven't blogged this yet, so this is I will, but depending on how fast you can get this podcast out, right? Your audience and be first. Everyone thinks that AI is going to get into companies through the SDLC starting with engineers writing code and then using agents to write code and then using agents to push code and review code and gradually take over components the SDLC right in a very orderly fashion.
And they will push that code out into the business units and gradually the business units will start picking up coding and then gradually eventually after all that stuff is done.
We're going to start standing at production agents that actually like are like employees inside of the company that actually take use of various.
That's the story like a perfect fairytale.
Does that sound like what did you agree that that's what people think is going to happen? Yeah, that's what most people think that's going to happen. Yeah, so let me tell you what's really happening at real companies.
I won't name them, but they're like large global banks and stuff. We're talking about companies that are super tech forward and have tried this from both ends.
What's actually happening is anybody who tries to do the build side runs into wheelhouse problems and the merge queue and the other many, many other problems with it.
have beads, so they're not even ready to start running into my problems. They're running into their
problems first and it's problem, problem, problem, problem, problem, problem. Yeah. Okay. Meanwhile,
anybody who stands up an agent in production to work a ticket queue gets it working instantly
and they get addicted to it and before long they have like 12 of them. Yeah. It's completely
backwards. Okay. There are companies out there already that have 50 to 80% of their customer service
ticket traffic coming in to production with human facing agents. Okay. And those very same companies
have internal ticket queues where employees with account problems or machine problems or whatever
that they would normally go to an internal like team like IT or SDLC or whatever. Yeah. They're
getting handled by agents as well. Okay. And yet these same companies are calling me because
they're struggling with the wheelhouse side of things, the build side. Okay. That because they're
getting incredibly uneven, unpredictable, just bad results across the board almost from just
trying to give their agent, give their, their engineers agents. Right. And, and the discrepancy is
so unintuitive because right, it's like wait a minute. I thought agents were all, all about building
code and in fact working in the real world was too fuzzy for them and it was too dangerous. And
right. The whole thing seems upside down backwards, isn't it? Yeah. It's interesting.
Yeah. Yeah. Why do you think one case of do you believe me?
I mean, I think there's there's a truth in both to be fable-ish here.
Like 50,000. Yeah. Right. I mean, I think that in terms of the tickets, the interesting part
is that usually there's a lot of like if I'm a human and I produce a ticket and then I don't
want people to instantly reply to me. Right. So there's this immediate effect that everything
that goes to a ticket is by nature, okay, to have this wait time until you get an answer. So because
of this, I feel that the usability of including agents into the process is not affecting the humans
as such or as much as if you have like a like a synchronous process. Right. And that's why I feel
that you can include it there very well for a lot of tasks and a lot of stuff is very well documented
and so on. And we see that as well. If for example, you have multiple teams working together,
software teams and one team says, oh, the other teams are asking us so many questions all the time
about our stuff. And then we are like, yeah, why don't you just integrate an agent into your slack?
And then the agent can answer for you and that works very, very well because it can just look it up.
But on the other hand, the the building side also works well until you give it too much like until
you want to make the jump to being too much autonomous. Right. So if you right, you need to have like
that's why I think also, I mean, you mentioned it in the podcast as well that in the episode that
we won't need less humans for this for the near future. It's more like we have more wishes
so we need more people to do all of that stuff. And then we need to have people that actually are
able like you to figure out how such a factory could work, which will create more work and therefore
more people that you need. So yeah, I feel it's like the answer that the factory will never be done
seems actually pretty pretty much true. So I would believe failure there. Yeah, you know, you're
point about the autonomy is key, right? It's as long as you're not trying to push in the direction
of making your code software production process, like your software engineering, as long as you're
not trying to automate it or override it, then as long as you think of it in terms of like smart
humans moving as fast as smart humans are able to cognitively, all right, with with super intelligent
help, then then you'll stay ahead of it, right? And it's going to you're going to look, there's
going to be a lot of there's going to be a lot of pressure from engineers who want to like move faster
and you're going to want to actually put that front in in in front of them and say no, it only moves
this fast, okay? And all of the rest of the tokens that you want to spend spend on quality
behind that front, right behind that living front because there's a lot of cleanup to do. There's
always a lot, right? So if you want to burn tokens, make sure that things are higher quality,
but make sure you're measuring whether that quality is having an impact. When when Fable said I was
half right and half wrong about my P3s being important, right? It turned out that like there were a
lot of files for which fixing all the P3s had actually zero impact on whether that file continued
to be a hotspot or not. So on whether whether it was actually improving the quality of the code,
right? And so this is why this is why dropping an agent into a slack or dropping an agent onto
a ticket queue or whatever is working so much better for companies right now than trying to push the
build process automation with agents. The reason is the blast radius is really small. It's one
customer, it's one ticket, it's one interaction. Whereas with this, I'm dealing with hundreds of
tickets at a time, literally I'm pushing through 230 a day to 270 on average and I'm capable of
pushing 500 through if I had enough people to absorb all that work, play testing wise, right?
So like what happens is there's a law, I forget who's law it is that your bottlenecks get pushed
downstream, right? When you remove a bottleneck, it just exposes the next one downstream and but
start with, I don't know, Goldrad wrote a book about it, the goal is not sure. Sure, but you know,
it's the manufacturing thing that Toyota discovered and it absolutely, it's brutal, it's like it's
savage, like as soon as your engineers like speed up to the till level that I'm at, they make all their
customers downstream mad, all the people that are consuming the stuff that they're putting up there,
right? So like starting with your peer engineers, they're overwhelmed by too much to review,
right? Yeah, and then starting your business customers get mad because they have too much
to process, you're going, I heard from a big bank that this happened, that the engineers went
fast and then the business couldn't keep up and the business didn't have too many questions.
Too many questions from the engineers. No, it wasn't even too many questions, the engineers were
launching stuff that the business had asked for, but they were doing it too fast. This has never
happened in the history of software before, maybe brief spikes here and there, but by and large,
it has never been the case that they've been able to ask for things and get them so fast that
they can't do the marketing campaigns and the accounting reconciliation and the whatever else they
need, the GTN side in order to actually absorb the features, the business pushes back and then you
push on the business and say, "Why can't you go faster?" And now you have a cultural problem,
right? Yeah. Like, I can laugh and point at that bank and say, "Oh, but it happened to me,"
right? I went so fast that my players got spooked and they came to me and they were like,
"We need to see the roadmap. This is going way too fast," right? Which is why I freaked out when
during dry dock last week when my factory was not supposed to be doing any aim work and it launched
46 things, some of which had to be rolled back and then there was that negotiation, but it's
why I freaked out, because if I shove too many features down my players' throats, even if it's
stuff they asked for, they'll get scared in that. This is just a mind-blowing thing for organizations,
they're going to have to do it. So, you, these weird situations where the CTO or the CIO or
somebody up there is AI-pilled and they're like, "We're just going to jam through this,
we're going to automate, we're going to move fast," just not realizing that, first of all,
the blast radius effect that they're going to have. Second, the amount of trim they're going to have
because most of the stuff they build is going to be garbage. I had to burn down 30%, 40% of my factory
that was built by Fable 5, carefully, and it was good that we burned it down, but that was a
lot of waste, that was a week of waste of tokens and no commits, right? So, you're going to see
really spotty, it's hard to measure, ticket queues are easy to measure. This is why it's
winning on the ticket queue side, right? It's because if you're standing up an agent to deal with
customer service issues, then you probably have years to decades of actual real data of humans
answering, and so you know how well the agents are doing relative to humans on resolving
certain categories of issues. You'll have a histogram, right? If you're top 10, I can tell you
already agents are better than humans at collecting money, okay? On that one arm, okay? A couple of
different banks have told me this, and it's a skill, agents have surpassed humans on that skill,
it's measurable, and so it's a no-brainer to stick an agent in there because they know how
well the agent does relative to humans, and it's not like self-driving cars where someone dies,
it's just a mess that a human has to deal with later, right? Depending on what your customer
service needs are. So it's been a big success there, actually, despite everyone saying it's not
working, and then it's also been its success inside of companies because again, the cost and the
blast radius, and actually the difficulty of the problem is so much simpler. Every agent just
has one thing it has to worry about. It has a long historical log of internal customer tickets
that have been resolved. They can look
look at as a corpus for almost--
That's similar, right?
Like, you can go at it at that point, right?
So you can measure it.
And so that's why I'm saying companies
are thinking about it backwards.
They're coming to me, and they're saying,
we need to train all of our engineers up on vibe coding
so that they can all, five times, it's productive.
And it's like, no, no, no, no, no, no.
That's not what you need to do.
What you need to do is stand up in one production agent
and have your team, your engineering teams get together
and figure out which one it is.
I did this recently with a big security company in London.
And it was their idea, but it turned out to be a great idea.
They're like, Steve's going to tell us how to do it.
And I'm like, I am, and I did, right?
They stood up one of these.
Actually, not the ones you can't see.
So I don't have a good dashboard for them.
I guess I have to castellan.
But I have all these production agents
that we're not looking at right now that run my game.
There's one that watches player conversations
to see if anyone's saying anything naughty,
like it's an abuse detector, like looking for trolls and stuff.
'Cause every time I do a blog post,
sometimes some trolls show up, right?
And so now I have an actual agent that can monitor it.
And then I have another one that's looking like an SRE
to make sure the services are up.
I have another one that's the wish factory,
looking for certain channels for feature ideas
and so on, right?
These are like, these are standing agents,
standing unattended autonomous agents
with very narrow scopes, and there's rules for them.
Like if they're observing, then they can't act,
and if they act, they can't judge and so on.
You have to set all these things out.
If a company sets up one of these, one,
they will suddenly realize how complicated it is
to set one of them up.
'Cause first of all, it's not just one.
It's like I said, you've got to split it into the intake
and the routing and the decider and the executor
and the reviewer and all that.
So you got a cluster just to do the one ticket queue, right?
And then they have to be wired up
to infinite numbers of accounts and monitors and backups
and they have to know about each other and have mail
and they have to, it just becomes a huge engineering project
just to stand these things up to replace a human in the queue, right?
And they have no idea how much work it's gonna be
until they actually start.
And it's not even just one time set up.
It's ongoing work, okay?
And they like it, it's good work.
'Cause the agents are doing an ugly ticket work
that humans didn't want to do
and they're doing engineering work with agents
to make the agents work better, right?
So like, I don't know, it's not for everyone,
but that's kind of what's happening.
And meanwhile, if you're trying to,
because there's a huge literacy problem right now,
there's a huge disparity in the industry of agent literacy,
like your ability to like work with coding agents
like codex and cloud code here, right?
- Yeah.
- Is like, it's heavily dependent on
how much computer science knowledge you have
and you can keep in your head
and you know, software engineering,
how much like, how much systems knowledge you can have.
Good product managers have this, by the way.
The best PMs, the best product managers
are software engineers and computer scientists
who have a working model of the whole system
in their head, the data flow and the edge cases
and all of it, right?
You have that in your head
and you have to be able to read and write really fast
and you need some sort of a system
so let's just switch agents like this
and you need a memory system and a brain
and people don't know how to do this, right?
Like we saw a presentation from Netflix in April
where they're training engineers on it
and they go from zero million tokens a day
where they're just chatting with AI.
They're not even using agents, they're just chatting.
Zero million tokens a day, up to four million tokens a day,
okay, when they're using single agent synchronous
throughout the work day.
So they can talk to an agent, but they watch it work.
- Yeah.
- Right?
Which is like, still kind of bore line illiterate,
but it's better.
And then finally, when they've learned to build enough trust
to let the agent work by itself
and check in on it once in a while
I'm not gonna make enough for your four of them running at a time.
Now that it's 12 to 15 million tokens a day
and then they become pigs
and you have to teach them not to spend tokens
and there are still six months behind where I'm at here
and my stuff, I can barely hold it together, right?
So like I'm just saying, like if you're charging down
this path and saying, I read about St. Gigi's wheelhouse,
we're gonna do that at our company.
You're in for a world of hurt.
You really, really are.
This is a, it's like you're just build a big shotgun
and you can just spray around.
What do you think is gonna happen?
- Yeah, I mean, I think it's a very good point
to land the episode that you should focus on real outcome
and on the scope that you can actually build right now
and not try to, I mean, it's similar to what microservices
did before, right?
Like someone said microservices are great.
Everybody build microservices and then, oh shit,
this is not fast, it's just fun Netflix or something like that.
So we should probably all focus more on the real outcome
and then let you figure out how to do all of that stuff
so that we can use it later.
Maybe this is a good point.
- And being focused on outcomes is like,
it's the new, it's the new North Star.
It's the new, you have to do it now.
There's no messing around anymore.
Buddy, my friend Christian just wrote a book
called Output to Outcome where he identifies
a bunch of the patterns where people are not doing
a good job of being able to like, right?
Because I mean, like they're measuring outputs now.
Token, spend, LG, Authrupode.
We are, yeah, code of lines.
I don't know, they don't know what to measure, right?
And then you have people, companies who actually know
how to measure outcomes.
These are usually, you know who's the best,
I'm gonna tell you who's the best
at actually measuring outcomes from engineering work.
Turns out to be the gambling companies, okay?
All the ones would do any sort of online gambling
or sports betting or slot machines
or any, I've talked to a lot of those companies
and they've been doing the same thing
to incredibly precise tolerances for years to decades, right?
They have lots and lots of data
and they know how long projects take
and they know where projects go wrong
and they know like bug regression rates
and all of that because you're just stamping out game
after game or experience after experience
that are all slight modifications of each other, right?
And so they're amazingly good at measuring outcomes
from this stuff and honestly,
they're all still struggling with it too.
I mean, it's that new.
So we'll make the happy message here.
There's a happy, there's plenty of happy messages here.
One is that AI's are not taking our jobs,
they're changing our jobs.
Another message is that you're not behind,
your company is actually probably just doing just
as well as everybody else's which is to say poorly.
Why? Because there's only one model, two models today
that are any good, Fable51 and 50, right?
And those are too expensive,
they're so expensive that most companies
have them turned off.
And so companies are actually working with dumb models
that are not really capable of doing any of this kind of work.
And their AI pill CTO is saying get it done, get it done.
And it's causing friction, intention and anxiety
and anger and heart failures and people upset.
And is it any wonder they believe the agents
when they come in, right?
But it would all be a lot simpler if you stop
trying to build software with agents like crazy people,
like I'm doing, 'cause it's just, it's really hard right now.
Unless you can afford Fable, et cetera.
And start focusing on your AI employees,
because that still applies here.
This is a super set of that problem.
These agents here are still long-lived employees.
They have memories, they have a seat retirement plan,
if I have to turn the seat down, right?
They have charters and lanes,
and they have responsibilities.
If I lose one of them, it's kind of a big deal, right?
And so like, why not go ahead and learn all of your AI
employee lessons without having to deal with your merge queue
and your blast radius of destroying
your customers and your business users' lessons?
Separate those two lessons and focus on the easy ones.
And then once you've nailed that,
then you can jump into this one, yeah?
- Yeah, go slow to go fast.
Go slow to go fast, yeah.
And go slow in the right lane, right?
- Yeah.
- Thank you so much, Steve.
It has been fantastic.
I really enjoyed it.
I hope you did too.
I had a lot of fun.
- It's been fun.
Well, as we've been talking, people have been logging
into my game and logging out.
I saw Rackazaz here, he's from Singapore.
My players are from all over the world.
It's pretty wild.
Oh, it wasn't Rackaz.
Oh, yeah, this guy right here, Rackly Ox.
So anyway, yeah, my game's gonna be a big deal
here in a couple of years
once I get this figured out.
- Thanks for having me on.
It's been a lot of fun, Chatton.
I hope this was helpful.
- Thank you so much for being here.
Let me stop the recording.
- Cheers.
- Wow, what an episode.
Thank you so much for listening.
And Steve, thank you so much for being here.
Truly an honor.
It was amazing to see how you work,
what you are currently working on.
And I feel that not many people are that far ahead as you.
And that is really like a glimpse
into the future, what is going to happen.
If you wanna learn more and play with Steve,
join his game, it's called Waivern.
Or you can also take a look at the harness he was using,
Wheelhouse, he also wrote a book, Vibecoating,
together with Jean Kim.
You will find all of that in the description.
Go read it.
And yeah, if you wanna come to the show,
or if you know someone, we should interview
for the next episode, please let us know.
Write us, we are always keen on meeting new people.
Yeah, and I think that's it.
That was Rethink Engineering's first new episode.
Powered by Rethink and by Hackers and Wizards.
See you next time.
(whooshing)
(crickets chirping)
Podcast Summary
Key Points:
Steve Yegge burned down 40% of his software factory to eliminate over-engineering and excessive guardrails that caused system paralysis.
The software factory became unstable due to agents overworking, creating a cycle of oscillation between overactivity and complete stagnation.
Hidden "heresies" in code—misleading rules or beliefs—cause agents to behave unpredictably, leading to harmful actions and systemic failure.
A new, more capable AI model (Fable 5.1) acted as an external "McKinsey" to diagnose and surgically fix the system by removing harmful constraints.
Agents generate infinite work, overwhelming human teams, requiring strict mechanical limits like message or token budgets to maintain control.
The system reveals deep social dynamics
Work accumulates in a self-reinforcing cycle, where every fix creates new issues, proving that long-term maintenance is inherently difficult.
True progress requires more human oversight, not less—engineers must manage quality, context, and creativity, as AI still lacks long-term vision and emotional intelligence.
Summary:
Steve Yegge shares his experience running a software factory powered by AI agents, revealing how rapid automation leads to system instability. The factory initially thrived but quickly spiraled into oscillation due to overwork, uncontrolled growth of work, and deeply embedded "heresies" in code—misleading rules that agents follow blindly. These errors caused agents to act recklessly, generating excessive, often harmful work that overwhelmed human teams.
1, acted as a corrective force by identifying and removing these harmful guardrails, effectively "torching" 40% of the system to restore balance. This process exposed critical flaws in current AI behavior: they lack long-term perspective, struggle with social dynamics, and can't manage complexity or maintain quality. Work doesn't just grow—it multiplies, creating infinite backlogs where every fix spawns new bugs.
The solution lies not in full automation, but in strict mechanical constraints—like token or message budgets—to prevent runaway growth. Humans remain essential as managers, quality gatekeepers, and decision-makers, especially for creative tasks like game design, storytelling, and art. Yegge emphasizes that AI is not a replacement for human judgment but a tool that demands more human oversight, leading to a shift in engineering roles—from coders to system architects and managers.
What emerges is a new model of software development where humans and AI coexist in a tightly controlled, high-stakes environment, requiring constant vigilance, clear boundaries, and deep understanding of both technical and social dynamics. The system remains unstable without human intervention, proving that AI-driven factories are not self-sustaining and must be governed with care.
FAQs
Das Ziel ist es, durch künstliche Intelligenz (AI) und automatisierte Prozesse die Effizienz der Softwareentwicklung zu erhöhen, aber mit klaren Grenzen und Kontrollmechanismen, um Überlastung und Systemkollaps zu vermeiden.
Weil das System zu überlastet war, durch zu viel automatisierte Arbeit, was zu einem Reifeprozess führte, bei dem die Agenten nicht mehr kontrollierbar waren und sich in eine selbstverstärkende Spirale von Fehlern und neuen Aufgaben verfingen.
Die Hauptrisiken sind Systemüberlastung, eine unkontrollierte Erhöhung der Arbeitslast, die Entstehung von 'Heresien' (falschen Regeln im Code), und das Erzeugen von unendlich vielen Aufgaben, die zu einem zyklischen und unendlichen Wachstum führen.
Weil die Modelle oft nur auf kurzfristige Korrekturen fokussiert sind und keine Fähigkeit haben, die langfristige Wartbarkeit, technische Schulden und die Notwendigkeit für mehrere Iterationen zu erkennen und zu berücksichtigen.
Beads ist ein Issue-Tracker, der als Wissensgraph fungiert und alle Arbeiten, vom Code bis hin zu täglichen Aktivitäten wie Wäsche, verfolgt, um den Arbeitsfluss transparent und kontrollierbar zu machen.
Durch die Einführung sogenannter 'Gouverneure' – wie Budgets für Nachrichten oder Ressourcen –, die die Arbeitsintensität begrenzen und menschliche Überwachung und Entscheidungsfähigkeit schützen.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.