The episode explores a range of AI and developer tool updates, highlighting Apple’s new iPhone Duo simulator with hidden APIs that require feature flags to activate, emphasizing the need for better support in web development. Cloud Code finally adds agents.md support, but through a complex "mod" system that undermines simplicity and interoperability, prompting criticism. The core focus is on Jev, a novel, fast, and cheap AI model designed not to generate text but to answer structured questions with predefined options—ideal for classification, decision-making, and automation. Demonstrations showcase its effectiveness in triaging support emails, managing tool calls, and filtering content in real time, outperforming open-source alternatives in speed, accuracy, and cost. While Jev’s internal architecture remains unclear, its performance and efficiency make it a compelling tool for developers seeking deterministic, high-throughput AI workflows. The episode also touches on local hosting trends, open-source models, and the growing momentum behind tools like GPU IKit, which enables high-performance Rust-based desktop apps with React-style components. Ultimately, Jev’s real-world utility in business and personal automation—especially in email, security, and content filtering—has sparked significant enthusiasm, with developers eager to adopt it for its balance of speed, affordability, and reliability.
What's up, everybody? It's time for another syntax weekly. We're coming at you with all
kinds of stuff. It's Jeff Tember folks. We're talking about Jeff. We're talking to all
about the cool stuff that we have been doing. Jeff alternatives, local hosting, Jeff,
all kinds of stuff. My name is got to let's give him a developer for Denver. And with
me, as always, is CJ Reynolds and Wes Boss. What's up, guys? Just Jeff. What a week. Super
exciting. We're going to talk about it. Like if it's overblown and whatever in a couple
minutes, but like super fun. I had a lot of fun seeing everybody's demos and everybody
getting really excited about it. And it's just it feels like a breath of fresh air. Yes.
I know the demos were fun, even if they weren't all practical. And lovely. I thought it was
just a fun time to do fun demos. CJ, how are you feeling?
I like it. Like I posted my fun demo. And honestly, it made me feel better about the state
of AI because at least the demo that I build, it makes you think about your chat bots are
the way that you use LLM's more structured way instead of just throwing everything out
in LLM.
Yeah. I like it. The text can just do it. We just tell the LLM to use my computer and it
just does it somehow. So yes, I do think it definitely added to that. Let's get into
it. We had the iPhone duo simulator. Yes, last week, yes. Last week on Monday, we
said Xcode 27.1 beta. It will be coming out. Hopefully this week and it did come out
last week. And with that, it included the very first look at the iPhone duo simulator.
Initially, I had said there has been no peep from Apple or anybody about supporting the
two CSS APIs that are needed. Right. There's there is the viewport segment API, which is
essentially if your website has more than one viewports, meaning like a folded display
or multiple screens, then you can decide where things will go. You maybe have a CSS grid
and you want to put them on different parts of the fold. The other one being the device
posture API, which is basically a media query that tells you if the iPhone is folded or
not or your mobile phone is now there is a flag. So if you go into the I've got the simulator
open here, go into apps settings, scroll all the way to the bottom. I'm only showing you
this because of how ridiculous it is. Then go to advanced. Then there's this is all the
advanced settings. But no, psych, nothing there, nothing there. Oh, feature flags. It was
hiding. You got to click it and shout out to Bramis on blue sky. He figured this out.
If we scroll, you scroll the way down in here. There's some kind of exciting stuff in here,
but somewhere in here, there is I've turned them on already. The feature flags for turning
on those two APIs, right? Yes. West, this is the exact same thing that happened with the
Vision Pro release. Really? And these flags remit like the WebXR flags remained unchecked
for like a year. I mean, absolutely. Who knows? Yeah. Who knows when they'll do this for
this and actually ship this into Safari because for some reason Apple makes this. So you
got to wait for a dang OS update to get one of your apps updated. But yeah. Yeah. So here's
a little demo from Bramis. He's a DevRal works at Google on Chrome. And he basically whipped
this up, which is like an example of of what you might want based on if something is folded
or not, you might want different layouts. And then of course, you can fold the entire thing.
And this shows you both. You can set up two different viewports, right? This is viewport
one on the left and viewport two on the right. And like it's not to say that you can't do
this with CSS already. I know that people are already typing that, but it just allows
you to know if there are multiple viewports. And if you are in a continuous posture, meaning
that you are in like flat mode here, or if you're in a folded state, which is when you've
obviously have it folded. And if that's the case, I think you would want to know if it's
folded or not because you can't just like swipe your finger across the screen as easily
when the thing is in folded. So something as simple as like like a button that you want
to click or something that you want to swipe away. Like I think of like a notification
you might put in the bottom middle of an application and you want to be able to swipe
that away. You can't, you can still do it, but it's going to be hard when you have like
a little valley in the folded thing. So you may want to decipher when you are folded
and throw that into a corner. Yeah. That stretches that viewport.
I love all the people complaining about all the new screen dimensions and screen situations
will have to support now on the web. Yeah. Like brother, I lived through the device lab
era. I had a wall full of devices in my office, all synced to be able to make sure that
it worked on it. Like the postage stamp size iPhone all the way up to the giant tablets
of the era. So you know, it's nothing new, I think for us, can we also show you this?
Yes. God, it's not good. Those the toolbars. And I think there might be a viewport tag
we can use to stretch this. But by default, there's a tool bar on the bottom for the URL
and a tool bar on the right for everything else. Like it's just like weird Wi-Fi icon,
which I assume is for like your, your carrier. But it's just so huge. Why? Why so big?
There's no room for a website. I think they're at the, maybe they don't want room for
website. They say no website room. But I think that the native apps are like this too. And
people are like, well, it's a skill issue because you can just theme it to look like brother.
You can't theme away this chunk. You can't theme away that. No, I wonder if they're going
to have to fix that. But hopefully we see those things unflagged. It's kind of a bummer
that they're not launching with support for folded websites on a folded phone. But hopefully
we see that yet. Everybody from Apple has been very quiet about it. So I bet we'll
say we hear something a bit more pretty soon. By the way, syntax, meetup, San Francisco,
October 27th in San Francisco. We are meeting up at the bear bottle brewing, which is basically,
if you're in town for your head of the universe or if you're in town just because you live
somewhere in California, come on down. You don't need to get hub universe bad or anything
like this. This is just our totally separate meetup. So if you're in town, come hang out with us
October 27th. Go to syntax.fm/meetup, grab your ticket before they are all gone. Did you hear
that, Sandy Agons? Sandy Agons. I live in California. You are invited to syntax me up at
San Francisco. Yeah. Come on down. Come on down. Yes. That was a great transition by the way,
West. So really good one. I think next up is CJ. CJ, what do you got, buddy? Yeah. So this is
huge news. We've been waiting for it for years. Cloud code has added support for agents.md.
So this started in version 2.1. Yeah, it's, it's the most ridiculous. Like it's so easy.
Yeah. AGI is here. They finally added this thing. If you're unaware, like if you're working with
other coding agents, like open code or pie or some of these others, typically, they are set up
to look at agents.md with all their instructions by default. But cloud would never look at that file.
It was looking for cloud.md. And so typically, what you'd have to do is create a link to the
agent's files. You could just have one, especially if you were using multiple agents.
But guys, the anthropic team said that cloud is a special snowflake and it needs its own special
agents file. And therefore, you couldn't possibly reuse an agent's file from one agent
to another agent, which honestly, even if it has a kernel of truth, it's like the dumbest
possible reason to not support a standard here. Yeah. No. In all of my experience,
simply just linking the file got cloud to know about the projects because I put all the stuff
about the project in agents.md. So yeah, this was ridiculous that it took this long for them to ship.
Don't celebrate just yet, folks. Okay. Because did you, did you read how they implemented
support for the standard agents.md? Oh, no. They created a new custom way to ingest things into
cloud, which is called a mod. So basically, they're going to be rolling out a new way to modify
a cloud with things called mods. I'm assuming this is kind of like plugins. And the way that they
had implemented agents.md is through their new thing of mods. So they had to be writing mods now.
Why couldn't they just, dang, change the path to the file? I mean, they'll come on or just, yeah,
there are some, like you said, there are some differences between it. But like, nobody cares,
you know, nobody cares. All of this gets injected into the context just so it can read one different
file. Like that's crazy. I don't know if this, I don't know how I haven't looked into how mods
just work. But I think mods is a way to provide additional instruction files. So I think this is
okay. This is programmatic. It's not just adding additional context. It looked like a prompt.
Maybe not agents.md read the way cloud code reads, cloud as a plugin under one instruction.
Yeah, I mean, I don't, I don't think so. Okay. So yeah, if you can just provide it a file and that
puts it in the context, that's great. Yeah. The, okay, the plugin itself is just a JSON file.
So I think we're okay. But it does have a discussion that I'm sure is going to go into the context.
Yeah, the more I use cloud the more I hate it. It's all just like long
- Yeah.
- And the, I don't know, people like the latest models,
like Fable and Opus 5,
because they're like really proactive.
But when I'm coding with an agent,
like I wanna know what it's doing most of the time,
I don't want us to just go off and run 30 tools
and then come back to answer one question.
- You know what?
I don't know if that's what I've experienced.
What I've experienced in the difference
between the two Fable and Astra, Astra specifically
is that I feel like Fable does a better job
of completing the job,
whether even if I have a detailed task list or whatever,
and I'll find Astra like, kind of just stops and be like,
okay, I got this part of it done.
Okay, but what about the other parts
that are clearly listed in your task work?
Oh, okay, I can go do those now.
I'm like, okay, well, I told you to do 'em, you know?
That's in my experience, that's that insistence part
to keep working, that's the difference.
- I'm mainly talking about, like a lot of times
when I use agents, I am myself, I'm exploratory.
Like I wanna answer a question,
I wanna know something about a code base,
and pretty often I'll have like a one sentence
that's like, tell me about this,
and I think I'm mainly complaining about auto mode.
I don't know if you've used auto mode in cloud,
but auto mode in a way, whatever that harnesses forces it,
not forces it, but it's more inclined to just read files
and start executing things versus just answering my question.
'Cause I agree, like if it's off to run a task
and I do want it to run in the background, great,
but auto mode makes it so that like every single prompt
becomes some sort of like background process.
- Yeah, I don't know.
- I will say guys, I have been using, yes, it's all Python.
I have been using Pi to do my cloud recently
and not using cloud code via a Pi extension,
and it's been good, it's been very good.
So it just works, and then I don't have to deal
with any of this crap, and I can still use the table of models.
- You can use your clots of encryption in Pi?
- Yes, yes.
- Are you legally allowed to do that?
- Yes, or are you gonna get thrown in jail?
- I'm not going to get thrown in jail,
and yes, I'm legally allowed to do it.
I will share this, let me even find the exact which one
that's called Pi Anthropic Auth.
I did a deep dive into which of these works the best
this last week, 'cause I was very frustrated
with the whole situation.
So I'll post the one in chat that I'm using currently
and to great success.
- I'll try it, I mean, that's the main reason
I use cloud is because I want to use the cloud models
without having to pay API key prices.
So if I can use it with OpenCode,
I guess maybe that Pi adapter is better.
- Oh, I don't know about, yeah, Open,
I don't know anything about it in OpenCode.
- How is it any different from Pi?
- Yeah, I'll have to check it out.
- There's a whole bunch of them,
I was concerned about them, but I have been using it now
for a couple of weeks and not only have I not been banned.
Well, the usage has been totally normal, so yes.
- Sweet, do we want to get into Jeffnext
or do we want to do something else?
- Oh, I'll show you tiny cast.
- Well, tiny cast, let's do tiny cast.
- Tiny cast, yeah.
- Tiny cast, because the world was ablaze with
what I believe to be anger towards rate cast
for, they released some update, Wes,
and you can speak more in this,
'cause I think you know more about the pricing changes,
I can get with a pricing bubble, and I was like, what?
- Yeah, rate cast version two,
they changed the way that their AI credits are used
because when they initially rolled it out,
it was, I asked it a question, and they gave me an answer,
and now it's just these huge, agentic loops
and MCP servers, and that makes sense, right?
Pretty much every AI provider has had to be like,
well, you guys are using us a lot more than we thought,
therefore we need to change how this works.
And then the other one, which kind of stung for people,
is that from rate cast one to rate cast two,
they put, bring your own key behind a paid subscription.
- Yeah, Versel does that stuff too.
- Yeah, which I understand,
I really feel for, for rate cast in this,
because it's hard to like run a business like this.
The same thing happens to warp as well,
it's like they build this amazing tool
that is just beloved by the community,
and almost everything that you use in rate cast
is part of the free thing.
And then at a certain point, they realize,
probably I have to make a little bit of money
to keep this thing up.
Like we got to pay these developers apparently.
And like, you have to start charging money
for some of this type of stuff,
and you have investors and whatnot,
and it's not just this beloved open source tool,
it's, this is a real business.
So they have to do that,
and obviously they're gonna,
they're gonna sour some of their power users
on this type of thing.
So is this a replacement?
- Yes, and I will say the BYO key thing.
Like I understand, like if you're giving inference
and stuff like the BYO key, I think I,
and if you're paying for the tokens,
I shouldn't be paying to pay to use my own tokens.
I get that you're paying for the tool or whatever.
I'd more than happily pay for other things in rate cast,
which I have for a long time.
BYO key just bugs the life out of me
when I needed to put a balance on my Versaille account
to bring my own key there as well, it's like,
I don't want to, yeah, I don't want to have to pay.
- I agree as well, like so many of these beautiful,
that's why I love like open router,
because you just want to taste a bit of a model.
You can just use it for like one or two prompts,
whereas like something without like 11 labs,
they just start locking all of it behind like a 1299
a month subscription, which is so frustrating
when you only want to use it just a little bit.
- Sometimes West, I just want to put a little bit
of the model in a glass and we should around a little bit.
- Sometimes a little bit.
- Yeah, you don't need to make it a big relationship.
Let's not be crazy.
- Yes, this model has notes of tennis balls and red wine.
Okay, well, did you ever watch wine tasting documentaries?
They talk about that.
They're like this just as notes are freshly cut grass
and it's stupid.
Okay, TinyCast.
So TinyCast.dev is a slop fork of rate cast.
You couldn't call it that.
I'm comfortable calling this a slop fork.
It is open source though.
And it is essentially exactly rate cast
to the point where you can bring in your rate castings.
There are, it does have better extensions.
Like it implements the API.
- Yes, it brings, you just bring in everything.
Bring in your whole, it brought in all my clipboard stuff
it brought in everything for me.
So this is it right here.
Again, it works very much like rate cast
and there's a ton of settings.
There's a ton of stuff.
It gets all my snippets, all that stuff loaded up.
Everything just worked.
I think the thing that rubbed people the wrong way
with TinyCast was that it was basically one click
to just migrate over everything from rate cast
including your extensions.
And it's clearly just slop forked, right?
That said, from what I've been hearing,
it's better performance.
Now I haven't validated that on any performance checking
of my own, but yeah, I mean, look at this.
This is so shameless, all this stuff.
But it actually, it works pretty dang good.
And I had it open and I gotta say I kind of,
I hate that I like it because like honestly,
this feels rude to me to do like rip off their entire design.
Is that it, to rip off their entire dashboard
to just like completely rip it off that bad?
Like that feels not good to me.
I don't like that from a perspective
of just totally ripping someone off wholesale.
That said, the app does work well and it's fast.
So I'm conflicted.
I've been using it and I like it.
So yeah.
- Yeah, I mean, this kind of gets at the whole,
like everyone can just use AI to build their own apps now.
So you have to come up with a better business model.
So like rate, rate cast, the feature set wasn't amazing,
but it was very useful.
I mean, I've wanted to build like a simplified rate cast
for the longest time and then this thing came up.
But because there's a lot of features of rate cast
that I don't use.
I don't have the paid account.
I don't hook up AI.
- Yeah.
- I do like app launching and shortcuts
for like system things like sleep my computer,
lock the screen.
Yeah.
- Can we ask that?
Like let's go around the circle.
What are your top use rate cast features?
- Emoji, clipboard.
- Yep.
- And then window management.
But I built my own window manager
so no longer window management.
- Interesting. - That's pretty much it.
- Mine, my biggest one is translate conversion.
So like 100 kilometers you type in
and it'll tell you how many miles, right?
Or 500 USD.
It'll tell you how much in.
It's so nice and fast.
The built in like the with premium rate cast,
the translation is so good.
It's super good.
Like I was like trying out some like alternatives
for Omunkey and it's just, it's not as good.
It's not as good.
And so I use that one quite a bit, the like conversion.
Obviously app launching hotkeys, shortcuts, math,
I use quite a bit, obviously app launching.
I use, I don't use the AI all that much.
I kind of do wish they offered a version
without the AI built in.
- Yeah.
- I don't know.
It's kind of tricky for them.
Oh, the file search is awesome.
Like if you can't find something in like macOS finder,
you just type it in into the rate cast file search
and it is awesome.
It doesn't search in the files, I don't think.
It just, if you kind of know what the name of the file is,
it is very good for that.
I was beta testing that for like a year before they released it
and I was a big fan of it.
So yes.
I don't use the like dictation for one.
I use super whisper for that instead.
Yeah.
One thing I do like about it,
about tiny cast, no account, no telemetry, no bullshit.
I like it.
I like that nothing leaves your Mac.
If you don't have, like, sorry, Raycast,
like this is, these are big things for me.
I like this.
I don't need an account to use by, you know,
I've been around since the Alfred,
pretty Alfred days of this kind of stuff.
I don't need an account for this stuff.
Just let me use the tool.
- Do you even, does Raycast require an account
before you even use it?
- No, it doesn't, but there's a lot of telemetry.
And so I've set up a new Mac a couple of times recently
and I just have to go,
there's so many steps to disable all the junk,
like disable the, asking if you want to log in,
asking if you want to add an AI provider.
- Oh, yeah.
Man, people are asking me like,
hey, why do you use Microsoft Edge?
And I was like, I really like it.
And then I installed Microsoft Edge on like a new computer
the other day and I was like, oh, this is, this is awful.
It's awesome. - It's really bad.
You gotta like really turn everything off.
It takes a lot of work.
And like they disabled my like new tab the other day,
like I have, new tab is just blank for me.
And I had to like install some short extension
and then like circumvent some API
to allow me just to have a new tab with nothing on it.
And then they turned it off the other day
and it like all of a sudden it was like celebrities?
- No. - Yes.
- You know, weather?
- No. - I don't know what any of that.
- No.
- Okay, cool.
Just a couple of browsers real quick.
- I'm back on helium, folks. I ditched DIA.
It's been like a couple of weeks.
I ditched DIA, DIA sucks.
And DIA sucks because it's constantly doing too much.
You're doing too much.
I hit a button.
The chat window opens up.
I found out there was some agent process
that was hitting 100 CPU for no reason.
I've been on helium for a while, folks.
And I was finding it to have a wonky support for a number
of things, one of those being one password
to the point where it was driving me nuts.
Turns out it's actually just one password.
It's not, 'cause it's bad on DIA in Chrome too.
So it wasn't a helium thing.
So at this point, I'm like,
well, what am I on helium for?
- Oh. - What am I on DFR?
I like the sidebar.
I like how arc it is.
I went back to helium today, today.
And I'm way happier.
- Yeah, speaking of the businesses that are like software
that has like been beloved and now is falling
from being beloved is one password, man.
Everybody is, they rolled out their universal sign
in UI, which I, I don't mind.
I don't love it, but I don't mind it.
And a lot of people are really mad at the new design.
- You know what, the problem for me is reliability.
And people are like, "Oh, it's just, it's not for me
or it's whatever for me."
I like how, I gotta do like five or six clicks.
I had to do like three separate fingerprint sign-ins
to sign into one site when it already is listed,
do not make you do it again for like six hours or something.
But then somehow every single time
has to be like multiple fingerprint sign-ins.
Like, there's something wrong there folks.
- Can I say something that'll make everybody mad?
- Yes. - Sink engine, Scott.
It's the local first.
- I don't think that's true.
- No, I think it's probably the Mac OS APIs
constantly shifting to some browsers, yeah.
- All right, let's talk about Jev.
(laughing)
- Let's do it.
- It's Jev, Timber folks.
It's time to talk about Jev.
- Jev, all right.
- Jev was a new, maybe CJ can give us a rundown.
CJ did an excellent video, like pulled all nighter.
I don't know if he did or not,
but he pulled out the most amazing video explaining
what Jev is.
So last week this new model was announced
and it's not like your other models.
CJ, want to give us a quick rundown of what it is?
- Yeah, happy to.
Check out the video.
It's only 17 minutes long.
It explains it much better than I'm about to
and shows a bunch of demos.
Released by Diego Almeida, previously of Chat GPT,
and also one of the co-inventors
of reinforcement learning through human feedback,
the headline is it's really fast and it's really cheap
and it's not an LLM.
So the whole idea here is it's system one,
is what they're labeling their models to system one.
And that is like it doesn't generate text like an LLM.
It just answers questions.
This is really the main thing to take home,
any Jev demo that you see, they have essentially
taken the problem that they're trying to solve
and reworked it into something that can be questions
and answers from Jev.
- Yeah, like so.
- You have to give it the possible answers, right?
- Exactly.
So Jev supports three different question types.
One of them is a yes or no.
This is known as a no.
It also supports a pick one.
That's known as a choice.
And then it supports rating on a scale.
That's known as a score.
So take any question or something that you want
to ask an LLM and instead phrase it in a way
where it could be answered with yes or no
or you could choose from an option
or get some sort of rating.
And that's basically what Jev is.
And that's also why it's fast in a lot of scenarios
because it's not generating text.
Like it literally just answers questions.
So in this example, a lot of the examples in their docs
are all about like classifying support emails
because a lot of people have built LLM agents
to classify emails and respond to them
or decide what to do with them, almost like triage.
And so with Jev, you could have a question that says
is this email a refund request?
And it'll say yes, it's a refund.
Or you could say what team should respond to this email?
And Jev will say, I think billing
should respond to this email.
Or you could say how frustrated is the user in this email?
And it'll give back a rating scale.
So that's literally all Jev is, you ask it questions,
you get back answers, but you have to provide the choices
in the case of choice.
And that's it, that's all it does.
But there's a lot of different things you can do
with that primitive.
So yeah, that's Jev.
I don't know other questions here.
- I think it's kind of interesting
because like a lot of demos that we saw were like,
like one of the coolest ones was like a self-driving Tesla.
I ended up building my own version of it
just to sort of understand what it was.
And as long as you can put the possible,
like the whole scenario input into the LLM,
and right now it's only text.
So you would have to like input that entirely as text
and then give it a possible options.
Like drive straight, break, turn left, turn right,
and you can give it like 30 or 40 different
possible options, right?
I did it and it ended up hitting a pedestrian
and driving into a backyard, which was, yes, yes.
- It's not great, but Justin Schroeder's implementation
was obviously much better.
- Yeah, that's the thing.
It's only as good as the state and the questions
that you give it.
So like with Justin's example,
like that's why if somebody sees this demo
and you don't understand that Jev is just a question answer,
you might think, oh, is it like navigating a 3D world?
No, Justin has broken this problem down into,
do I see a stop sign?
Is there a crosswalk?
How fast am I going?
Should I turn left?
Should I turn right?
Just think about all the questions that should be answered.
And you give Jev the state of like where the car is
and everything else and it will answer those questions.
- So probably not what you want.
Like even as cheap as this is,
it's probably, it's too expensive for self-driving
and you shouldn't be imagined.
Imagine self-driving was based on HTTP APIs.
- Right, exactly.
- That would be all the dead.
So obviously that's not what this is for,
but if you are looking at this type of thing in your first
instinct for this is that stupid, that's not what this is for,
blah, blah, blah, that's not how it works,
then you're obviously not thinking.
As soon as all of these demos that come out
are not people like trying to hype it up,
they're just trying to understand how do I use this thing?
And what is it good for and what is it not good for?
- A demo can sometimes just be a creative demo.
It doesn't have to be a real world usage of this thing
is going to be used for self-driving in the future.
It's like I just explored and it made this,
this fun demo in a smaller regard.
Like what is a smaller thing than self-driving
is like browser use, right?
- Yeah.
- It's self-driving using the browser.
You get an accessibility tree, it determines where to go,
it does that thing.
And I built a self-driving GPUY crawler to do,
like not computer use, but GPUY manual testing
instead of using something like DevTools MCP for the web,
to be able to, and it is fast as hell.
And then we're very good for that.
Yes.
- There is a comment,
I think we can just get into a little bit that says,
how does it determine the answers?
This is the other thing that is a little bit different
and weird about this launch is the team at Taipei AI
has not published any research papers
about how they built Jeff.
There is no explainer of like the internals.
People have tried to guess at what it is.
Some people have guessed that it's like trained initially
very similar to an LLM,
but then the post training process is a whole lot different.
But I will say that's one of the weird things
about the whole release of this model
is we have no idea what it's doing.
And I guess we'll talk about,
there have been some open source
and basically like recreations of Jeff using LLMs
and other types of models.
- Yeah. - To show that like,
first of all, this idea isn't exactly new
and then also like you could potentially do this locally
and don't necessarily need a hosted API.
But you're asking the right question
because we don't know is the answer.
And people have tried to poke and prod at it
to figure out its internal architecture,
but we actually have no idea what this model actually is.
- I will also say there was a lot of people
that built demos that were like kids made this demo
where he was asking how many calories something was.
And I was just like, how do you do that?
Like the calories, you have to give it possible options.
And is that right?
What he was doing is he was giving it just ranges of how many calories things should be and
then based on 100 grams and it wasn't accurate at all.
It wasn't very good at all it ended up just being Matt that ran and he said like that's
not what it's for.
And a lot of the complaints people had with these things were just like you're confusing
people as to what the actual model is.
And to that I say like if you are so far gone that you have to understand how something
works by just reading a couple tweets and like that's it.
You're not going to make it in this thing.
If you want to be able to understand the stuff given how easy everything else is right
now, you have to build something.
You have to read the documentation for if to watch a 17 minute video from CJ, please
I beg you.
I think also the like because people are so used to LLMs that do everything like LLMs
are general purpose and the best ones can do anything.
And so I think that's also where the public perception might be coming from is like they've
already seen LLM so they see these demos with Jeff and they're like, oh, it's so different
or so much better.
But yeah, understand the context like what can you actually do with it?
And then how does that fit into what you're trying to solve?
It's crazy how many people did chess demos with Jeff and they specifically say don't this
isn't going to be good for chess.
It's like people don't even read the onboarding at all.
Why is it not good at chess because it has no reasoning?
Because chess I think has, I'm not a chess guy, folks, so do not skewer me, but I'm pretty
sure it because there's context in terms of what moves were done in the order in which
moves were done.
But I don't know that to be sure.
Should be able to pass that in though, but maybe that's that's how we beat these chess
bots.
Is that every single like, oh, I know this bot is going to make this move just like it
it writes the same code every single time.
Well, as I know how to beat a chess bot, but how you unplug it.
You just I mean, come on, you've discovered AI safety, Scott.
Can I show my Jeff demo really quick?
I built, I called it a chat bot and I think this is another reason why you, you might see
this demo and think that I'm like saying something that it's not well, the way that most
people interact with a chat bot is like tool usage.
So like it's, it's a lot better when we have grounded answers from like Wikipedia or
weather or whatever else, so I hooked up Jeff to like seven or eight different MCP tools.
And I've shaped the questions and so that like given a user's prompt, which tool would
they like to call?
And I also did some pre processing.
So I take the user's prompt and I extract out like capital letter names, places.
I have, I found a natural language library that's like super fast.
It works without a model and just knows about parts of speech.
And so it can take a plain text prompt and figure out what the user is asking.
So in this case, it called the weather, MCP to get back the weather, because I said what's
the weather in Denver, in Denver.
But the, the way that I asked Jeff to do it was for this prompt is this a new request
for something the assistant can do.
And then it'll say with 99 percent, it's 99% sure that this user is asking a new request.
And then I have another prompt that lists all 17 available tools and it says which tool
would fulfill the user's latest request, taking the conversation into account.
And with 100 percent certainty, it says you should call the get weather tool.
And then there's another question that says for a weather request, which city or place
does the user want the weather for?
And again, this was the pre processing step where I picked a few options and then yeah,
from the user's prompt.
And then that's what I'm asking Jeff.
It should pick and hear it saying 100% accuracy that it should choose Denver.
And then it also could do time and everything else.
So again, my, my whole demo here is like if you reshape how we talk to chat bots to like
more deterministic actions, they're much faster and much more accurate.
Like another cool demo is I can say how tall is Mount Rainier.
And this will actually reach out to Wikipedia, do a search for articles, pick the, the Mount
Rainier article, pass back the first 120 lines of the article and ask Jeff which one of
these lines answers the user's question and boom, elevation 1400 feet.
So yeah, I think that's the other thing is people don't realize that you can use this
recursively.
Even with the example of like MCP UI, which is like, here is my entire design system, which
of these components should I use to display the weather?
Okay.
Now that we're in that, which of these pieces of information should go into the title.
And you know, you can just, you can kind of run it six or seven times to get to the end
result.
And you can have like deterministic, I'm not sorry, not deterministic UI that has been
generated by it.
You don't need to have the whole L and because of that, it's much faster, which is.
It's much faster.
Awesome.
Yeah.
Yeah.
At the same time, everybody, the last week has been like, Oh, maybe Web MCP is not so
bad, thanks to the syntax episode that we did on Web MCP because Web MCP is awesome.
And if people have not like woken up to how awesome it is, and I think now that everybody's
back on the Web MCP train, now they're all looking at it and we're all on board.
Guess what works perfect with tool calls that have deterministic tool calls Web MCP, excellent
use case for Jeff.
Yeah.
But I want to applaud West here because West took a segment on Jev and figured out, how
can I make this about Web MCP?
You love that.
You love that.
So you're like, I'm a whole game, yeah.
Yeah.
You know, anytime, anytime an LLM is like using its reasoning or whatever to pick a tool
call from a list or anything like that or deterministic scripts or anything like that.
Man, Jev stepping in there for me to like pick which swamp automations I'm going to use.
It's going to, that's going to be awesome for me or especially even like you could build
in an extension to pie, which does a better job of loading skills.
I find sometimes the harnesses do a bad job at picking skills, especially when I have
a lot of them on my system.
Let's talk about the kev's in the room.
So since Jeff has been released, there has been a release of two kev's, kev from Cloudflare
and kev from Jared Palmer and obviously both of them named kev because that's a great
choice.
I would have gone with Jeva Daya, but kev is, kev's all right.
And people are saying it's, I don't know, CJ has some, some details on if this is actually
as good, but they're trying to figure out how do we use fast and cheap LLMs that we
already have to replicate this Jev API and experience that we've been using.
And actually, can we speak on that too because there's some confusion about whether or
not Jev is using an LLM behind the scenes just because all these other open source ones
for you.
Well, well, no, but we don't know what type of language model it is.
Is it a LLM?
Is it, people have said, is it a diffusion model?
Is it the, the Jeff team hasn't said it's just a bunch of people in a room answering
very quickly actually, but they still haven't said what it is.
We don't know, because they haven't come out and said that it is an LLM definitive
lately, even though it is most likely a language model of some kind.
Yeah.
Yeah.
Are they doing that because they don't want other, like you have to imagine that we're
going to see this from OpenAI and Thropic.
Absolutely.
Yeah.
I think, right.
I think as more people realize how useful Jev is and how useful it is to not directly
use an LLM, yeah, OpenAI is going to release their version.
We're going to see an Anthropic version for sure.
And like there are open-weight models that people are trying to fine tune to behave
more like Jev.
But yeah.
I can talk about that now.
So I'm sharing a little code snippet here.
And so like if you're not aware, when you're using any sort of SDK to talk to an AI, you
can ask it for structured output.
So you can give an LLM a prompt.
And then in this case, I'm using the AI SDK from Versailles.
But I provide a Zod schema that says, you should respond with the JSON object in this format.
And so if you think about it, you could basically just take the output format that Jev gives you
and now ask an LLM the same thing.
So it's definitely possible.
But we always have to wrap this like in a tri-catch because sometimes the model's answer
cuts off or sometimes it hallucinates answers.
So you technically can do this with an LLM, but Jev is just fine-tuned in a way where
it doesn't make up options and it just immediately gives you back the things that you were asking
for.
But a lot of people have tried to do with these more open models.
But yeah, I hooked up my Jev chat here to Leia and Kev, which are two different open-source
versions.
Leia is one that is actually from someone who had this entire blog post that was like,
I've been working on this a year ago and nobody cared.
And they have like all their research papers and everything else.
But I'll show you, at least we've downloaded it.
We don't want your research paper.
Give us a funny Twitter video.
Please.
But yeah, there's a whole hacker news thread with people explicitly saying like, you have
to worry about marketing and product and all the things that type safe AI has done.
But I will say, so I'm running Leia locally and I can ask it, what is the weather in
Denver tomorrow?
And I'm running the biggest local model.
The response took three seconds.
And for that same question of which place is the user asking for, it picked tomorrow instead
of Denver, so it didn't pick correctly and it's a lot slower.
Obviously, you can maybe run this on a bigger machine.
I picked the 4B model that can run locally on my Mac.
Leia, 71 on the speed, 63 on the intelligence.
So Leia does come in pretty low where Jev is still the top dog in terms of overall performance
on these types of things according to you.
to the this bench, this particular benchmark in terms of both cost intelligence and whatever.
And the one thing that I've noticed about people say, "Oh, I'm running Leia on my computer
instead of Jev.
Brother, it is not Jeff.
It's not.
It's good."
And part of what makes Jeff useful is that it's good, like not that it exists or that
it's like do the thing.
Yeah, it's like general purpose, fast, cheap, and good.
Right.
And so when you look at this, it's important to pick options not just because the guy made
the Leia thing before or talked about it before, that it's actually worth using.
So I like the semif1 using Quinn 3.54B that you can run locally.
And it gets close and is about half the price in terms of usage, in terms of it's being
faster, but then the intelligence is a little bit lower on their benchmark here.
So for the most part, the reason why the local Jev stuff isn't as exciting to me right
now is because none of them are as good and what makes Jeff actually useful is that it's
good.
Unlike in LLM where you can get by the fact that using smaller models for smaller tasks,
I think with this type of thing, you want it to be good for it to be useful.
Absolutely.
And I think that's what they got right.
The launch was great, the early access made people want to get access and got excited
when they got access.
The marketing aspect of it great, but then also they delivered and it works.
I also tried semif on my local demo.
And it only supports up to like, initially there's 13 options whereas Jev supports a whole
lot more.
So I couldn't even get my demo working with semif.
So it's nice that they exist.
So I think more will pop up especially the like people have more time to work on this
because Jev was literally just released like last week.
And another one that people have been talking about, I don't know if anybody has classifier.dev
pulled up.
Maybe we can pull up classifier classifier.dev is the only one that's beating Jev on the
Jev scale here.
Yeah, but I'll say what's funny about it is it's built on top of it.
It is Jeff.
Yes, it is Jeff.
People are like I'm using possibly not Jeff.
It's using Jeff.
Well, can you explain what how do you build something on top of it and how is it better?
So my understanding is that this is using an LLM after the Jev processing.
But that could be a very dumb understanding of this.
Yeah, it says a smart tier re asks only what Jev was unsure about and comes out 2.5 points
on plus 2.5 points.
So that follow up is what makes it a little bit smarter.
It gives it the answer.
Because what you could do in this case is you could throw this to I guess they can throw
it to Quinn or you could like you could throw it to what's the open AI one Luna, you know?
Like Luna is dirt cheap, especially if you're just asking it tax.
You know, I was I was sending Luna an entire image and asking it to detect what was in the
image and draw a box around it.
And that was one tenth of a penny, right?
Yeah.
And that's an image as input.
Obviously Jev is much cheaper, but if you're just falling back to an also very smart and
fairly cheap model and like that might you're maybe oh, you're split in pennies here.
But like people are running this on literally everything.
I what I did a demo where I downloaded a hundred YouTube videos from syntax and analyzed
every sentence in if it was positive or negative.
And I should publish results is actually kind of interesting, but I think I think analyzing
a hundred cent, a hundred videos, all of our sentences, just like probably 80 hours of
content and it costs I think about a dollar or $1.50, which is still expensive if you're
analyzing every YouTube video in the world.
But also if you're thinking about like that's that's quite a bit of of processing.
I often wonder like at what point does this get cheaper than like the Cloudflare request.
Yeah, right?
I can't wait to be running something as good as Jev on my local machine over everything
all of the time.
Yes.
Over everything.
Yes.
Categorizing smart filtering.
Yeah.
There's a lot better search like imagine this stuff gets cheap enough that I like a search
on our computer is going to be good.
What a world.
Yeah.
Right.
If the Apple search could actually work.
Yes.
Yeah.
I like a Jev as also like a follow up algorithm of sorts.
So one of the demos somebody built a Twitter extension that hid the things they didn't like
like ragebait or political comments, but it did it in real time.
So you're literally scrolling and you're not seeing the things you don't like because
Jev is automatically classifying them.
You could do the same thing for like a YouTube feed or any other like you just mentioned search.
Like what if you could take like the first 20 pages of search and really figure out and
get rid of all the clickbait and figure out which of these actually adhere to what the
user is asking.
Yeah.
That's another quick follow up that instantly gives you better answers and a better experience.
So yeah.
And you can't do it.
Like I just tried with Luna is West cool.
It obviously is cheap, but it took 2.8 seconds.
Yeah.
It's too long for for interaction versus like a typical request to the Jev is I was seeing
it at like 150 200 milliseconds back and forth, you know, and that includes the inference
time, which is it's starting to get closer to like a like a nice real time experience.
Yeah.
Yeah.
Man, it's it's quite actually.
So stuff that we've built with Jev, I know West, you want to really show your FMK with
Jev.
After Jev into my prompt boy set up in two ways, so my prompt boy was my personal to do
thing that agents could get working on to do is I put it in two ways.
One, I put Jev in the classification of the prompts that are coming in.
So I pick up my thing and I say, Hey, can you do work on prompt boy to do this and that?
It uses Jev to determine whether or not that's coding work or just a standard to do or whatever
and puts it into the right, the right mode, right?
I also use it in the general to do's to classify it as like, is this urgent?
Is this a long task?
Is this this?
Is it that?
And my God, it cost me nothing.
And it's been great to use it.
Auto classifies every single to do I put in there via voice or via text or anything and
it just works.
It's great.
It's really cool.
I'm curious to see someone like, I know open router has classifiers in beta right
now and cares to see like when this will start to roll out because even like toxicity detection
like a class classifier like I was when I did my receipt printer and the messages in
this case, it was just messages were rolling in and I say like, is this is this something
that is swearing or bad or someone is being like trying to like sneak stuff through?
And if so, give me a toxicity score and then you can base it on that.
That's just that's that's an entirely separate model that you need to run just for toxicity
detection.
Yeah, you know what?
And that's just one thing to have can do.
People often are like, why would you use an AI for this?
Why not?
Why can't we just do basic classifications?
And the problem is it like when you're when things get to be generic enough or what's
the word I'm looking for ambiguous enough like I've been building a classifier that was
not using AI to do my running with just code based classifier for this task or this task
or this task.
And it felt like impossible to get it to be where I wanted it to be all the time.
I put Kevin there and it's been 100% accurate in terms of like, or this is a coding task.
You know, it just it's better with that kind of stuff.
I bet.
Jev will roll out multimodal so you'll be able to like give it an image as well.
You know, is this a potentially sensitive image?
Should I show it on the screen or not?
Multi wouldn't be huge.
Yeah.
Being able to put text into it would be great or sorry, not text images.
And then I bet we'll see like a Jeff reasoning where there's like a trade off for that type
of thing.
But I'm also curious to see what's going on at the other the other like anthropic and
open AI right now are they they probably have something similar to this.
They're trying to figure out how they can get it get it working as well.
Because I'm sure they're saying, well, we can take Luna and make it 80% as good.
And we can speed it up so that it's, you know, like I'm sure they're like, I can get
sort of a similar API for these response times if we just juice it.
And the other question is like, is Jev making any money or are they just coming out with
such cheap prices just to sort of swart all the competition?
Because if this, if Jev was like 10 times the cost, everyone would be like, cool.
But yes, you're right.
They did mention in their launch post and on their site that the pricing is not, not subsidized.
So if they're telling.
Oh, wow.
Okay.
Then yeah.
Yeah.
And they were saying basically because they were asked directly about making money from
it.
And they were saying, obviously at like normal scale that you would use an LOM at, it's
not going to make that much money.
But the idea here is because it is so fast and you can use it so much that the amount
that you're calling Jev is way more than you'd be calling anything.
So therefore, that's the path to making money is that you should be using it a whole
ton.
Yeah.
You'll be using it more.
Yeah.
Yeah.
On every frame of a video, you know, or something.
Here's one more post from Greg who works at Century.
He says, here are the results of using Jev on one of our security pipelines.
We already have to use smaller and dumber models to make it economic.
That's exactly what it is.
We need to run something a lot on our security pipeline.
But we can't use the big models because it's simply.
just way too expensive to do that type of thing. I even think about how much money,
like someone like socket.dev spends scanning every single NPM thing published ever.
You know, that's got to be really expensive. And they do pipe a lot of that into LLM.
They told us he had them on the podcast. This model does it over 5% cheaper faster while
maintaining higher accuracy. Like the good fast cheap triangle just got smashed on on this.
You know, it's better faster and cheaper. Yeah. Yeah. So I think this one alone answers why
there's so much hype. Obviously, you know, some of the demos aren't as great or whatever else,
but this is a real business use case and it ticks all the boxes. Yeah. Yeah. Yeah. I think people,
I think that conversation is missing because I see there's so much on on Reddit and other places.
I'm running this local version and it's like, okay, but it's not you're not it's not as useful.
That's not the point, you know, guys, I need to put this thing on my DMs on my email. I need to put
on everything. Yes. That's why it's the new Frank's red hot. Oh, Frank's right. They need to
collab with them, right? Yes. But I'll say I'm not doing that because every prompt goes to
Jev's API. So I really wait for a local one before I do that. But that that kind of use case is
just so great. Like to be able to go through your like anybody right now that has an inbox that is
out of control with like 10,000 unread emails. Like imagine sticking Jev on this and then having
just a like bringing that down to maybe a hundred emails that you actually need to look at versus
the zone is a stuff. I'm going to put Jev on my email and it says is this email saying that
my Mac studio has shipped. If not, then then send it somewhere else. I don't want it. Yeah.
But like even if it's 10 times slower locally and it takes a second to process that email,
who cares? Yeah. You know? Yep. Yep. Yep. Yep. The local use cases is maybe even different than
I like it. Yeah. Speaking of century, by the way, folks, check out century at century.io.
Century is the reason why we're able to do this show. And they are amazing. If you want to support
us, support century at the very, very least, the best thing you could do is just send a century
a little, a little message on Twitter saying thank you for syntax. That would be really cool.
Just let them know how much you love syntax. And therefore, century, they're great. Check it out
century.io. Awesome product. CJ, you had something. Yeah. I just keep talking about Jev.
One last thing. One last thing about all these other open models. Jev has a contact size of
32,000 tokens. Leah only supports, like, I think it's a thousand. Kev only supports up to 8,000. So
not only are these models not as good, these open ones, they're not as, but they're also,
can't handle as much data. So that's it. I'm done talking about Jev. Okay. We'll talk about Jev.
Yeah. Let's talk about something else. Guys, I'm a big fan of Jev. I can talk about
GPU IKIT if you guys are cool with that. Is that sound interesting? A little post Jev
come down here. We're going to be, let's pop this open. GPU IKIT. Guys, GPU I is the framework
that the Z team created to build Z, which is an awesome high-performance editor. It works great.
And it's all in rust. Now, a lot of people know about like Tori, which is the web view way of
building desktop apps with rust. Tori is great for a number of reasons, but I personally had
to move off of it to use electron because of just general maturity in the ecosystem and things
like that. And at the end of the day, you know, I wasn't, you're not getting the benefits of rust
from using something like Tori anyways. And by default, if you're, you're doing this on Mac,
you got to support so far. Either way, GPU I basically brings good UI to rust apps. It's
awesome. But I always thought that GPU I was like, okay, it is going to be very in the weeds,
getting something that looks nice. Well, I found a number of these things, but this one is particular,
is excellent. It's a GPU IKIT.com GPU IKIT, which is basically you could consider it like a ShadCN
or any React component library for GPU I, meaning that you can build good desktop apps with
React style, like components and CSS style, like components directly easily and AI agents are super
good at this. And like they do stuff like virtualizing lists out of the box, data tables, all that
stuff. The component library for this thing is bananas in terms of stuff as support. 60 things.
It's everything you could possibly need or think about. I've used this now in three different
projects that I'm working on. And it's it's been great, man. I actually really like this for,
if I'm going to be building local personal software, I'm going to be reaching for GPU IKIT,
and I'm going to be reaching for GPU IKIT. I'm going to be reaching for those two things over
anything because again, the whole thing, the whole process is you're not using a, and I'm a web guy,
I'm sorry, web folks, you're not using a web view for this. So does it have a cross platform
story? So it can be the same desktop app could run on Mac OS or Windows or Linux.
As far as I know, absolutely. According to Shopify, you don't need that anymore, though, because
Shopify is what drop in React native and going native. Oh yeah, simply because the cost of
implementing the same UI in two different native and HTML CSS is not as expensive as it used to be.
You know, it's much easier to do. Well, this okay. It is Mac OS, Windows and Linux.
I've seen a bunch of people getting GPUI working on mobile too, but that seems crazy to me.
I mean, it works. It works on the website because it renders to Canvas, right? Obviously,
it's not HTML, but I'm talking about mobile, mobile native. Yes. Yeah. It theoretically should
be able to use it anywhere that you can render rust, right? Which should be anywhere. Yeah.
Yeah. And I built two things with it. You know, I had a app called FileBrow West. Do you remember
FileBrow? FileBrow is like a data automation web app for me that like you can, it's kind of like
hazel where you're setting up automations for folders and files to move things in the right place.
I had that in web tech. And as like moving that to Rust and GPUI, it has been sick. I built a little
desktop app for managing all the web properties between my two computers. So that way I can keep
track of them easily. Like, oh, the one on my laptop is five versions behind the one on my Mac mini.
Let's sync them. You know, that type of thing. I'm using Jeff and all this stuff too.
Jeff is coming back into the folder. That means Jeff on both of these projects. And I just recently
built a window manager with Rust for myself that I used a GPUI kit to do the settings menu.
My God, it's sick. It just works. So yes, when do we get this type of stuff in like a car UIs
and like native, you know, like like UIs in things that are not the web and not like a Android or
iOS app have been like hurting for a long time. And probably the biggest spot that we see
awful UI is in like car infotainment screens. Like you got to think like people putting all of
this at energy into things that are not rendered on one of the big three platforms now.
You got hopefully this will start to trickle down into some of the bigger infotainment screens.
Yes. Yes. Let's get some of that going. Okay. That's all I got for GPUI kit. GPIs,
dope folks, check it out. Yeah. Cool project. Who's got one? CJ got one or you want me to go?
I got a quick one. All right. We're missing the formal handoff guys. No, that was so formal.
That was very formal. I like that. So Anthony Foo also knows Aunt Foo has released a new project
called DevFrame. It just went into version 1.0 and it's a framework for building DevTools.
You can check out the launch post. He goes into all of the types of DevTools that Aunt Foo has
built. If you're not familiar with Aunt Foo, they're very prolific in the open-source space. They
work on Veet and View and they also created Slide Dev. On top of that, lots of DevTools for all
of those things. This isn't a problem that everyone's going to have because not everyone's building
DevTools. But Aunt Foo has built a lot of DevTools and has taken all everything he's learned and
put it into a framework. So it's called DevFrame. The idea is you define your tool in a somewhat
agnostic way and then you have handlers for different frameworks. So the same DevTool could run
in Nitro or Hono or NextJS or SvelteKit or Veet or RS build. So this is going to be huge for
library and framework creators that want to maybe add more telemetry or more info about what's
happening inside of your app when you add them to your app. So pretty sweet. Yeah. DevFrame.
That's awesome. DevTools at UI have gotten so good over the last couple of years and it's so
nice to see some sort of like standard basis. I know building like a DevTool extending Chrome DevTools
has been huge pain in the butt. So stoked to see this. Aunt Foo also behind Shiki, which is the
syntax highlighter that absolutely everybody uses. You can give it a VS Code theme and get
beautiful syntax highlighting. Yeah, he's a G no doubt. All right, I got a question for you guys.
You are working on some footage in your project and it's Shiki, you know, or it has a background
noise or you need to remove a background and there's a tool built into Vinci where you can just like
make your shaky footage a little bit better or repair or remove a background.
That's, that's okay to do, right? You have some footage and it's, maybe you drop some frames
and you use some thing individually to try to replace those frames. That's an okay thing to do
by your regard. I think so. I think a lot of times the, well, the algorithm typically will
zoom in, figure out what it's focused on, and then like automatically adjusting those, and like,
left, right? Yeah, I think that's, is that AI? Is that AI, you think? It depends on the algorithm
that they use. Yes. I think it's possible to do without AI, but again, we don't know because we
just click the stabilize button inside of Davinci. Yeah, right. Yeah. It's something I think a lot
about because I watch these TikToks of people like being anti-AI, and they are, have AI posters that
are anti-AI, and they, they just generated them like chat GPT, or they have like, like a anti-AI
t-shirt that's been clearly made with AI, and then, or they're like, they don't realize that,
like, you have a TikTok and your auto captions and your background removal, that's AI, you know?
And I get that there is, there is a difference, right? You can't just be like, "No, I won't use
that any. I'm just going to cut myself out frame by frame." But I got thinking about this because
RunwayML posted this the other day where they said, "Introducing enhanced frame rate." So our new
frame interpolation model converts any footage to the specs you need, including 25, 30, 48, 60,
120 frames a second. So what this means is that if you have some footage that was recorded in
the wrong frame rate, or if you have probably the biggest use case for something like this is if
you want to take footage that was not recorded for slow-mo, and you want to make that slow-mo,
so it's not blurry, it will sort of interpolate the frames in between given the things that you
wanted, which is kind of interesting. And I was just kind of looking at this demo being like,
usually the fight is you shouldn't use AI-generated video, or you shouldn't be using AI to generate
posters. But there's like so much of this like in between where it's like, "Yeah, I filmed this
thing and it's dropping frames." Or my Riverside is dropping frames all day long, and we need to fix
the footage. So I've used Topaz. So Topaz has been doing this for a long time, and I don't know
it, because Topaz, you can choose the models and stuff. I don't know if the Topaz models are
any different than what they're using at Runway, but I haven't particularly found the results to look
that natural when you're when you were using the frame rate. The DaVinci like, if you have a cut and
you want to like smooth out the cut, it like, yes, it tries to like morph your face from one to
another. It's just guessing. Yeah, it's just determining and guessing. Yeah, it's deterministic to
define which pixels go where. Well, I think, yeah, some of these, the Topaz one is using AI models,
so that is using models. Yeah, the DaVinci one for the smooth cut. But for the enhanced frame rate,
every frame that it adds is generated with AI, like it has to use diffusion to come up with the
interpolated frame. So I have a feeling you'd be able to see those AI frames, especially if there's
like weird artifacts. Yeah. If they try to turn a video into slow motion. In my experience,
it is the case. You do get those still those AI artifacts in them. Yeah. What about if you couldn't
tell? I don't think it's going to get good enough. Should you use this? But let's take a step back and
talk about the people that are making those arguments, because like right now, a lot of young people
are very much pushing it back against AI. Like the youngins are like, no AI, anything. But a lot,
but most of that is coming from a moral standpoint of like the companies are not moral. The models
are not trained in a moral way. I do not want to be associated with these things. I think if you're
someone that has that standpoint, no, no AI at all. Like you cannot, like you are absolutely a
hypocrite. If you go off and use a tool like this in your creation, yet you post video is about
being anti AI. It really comes down to like, I guess understanding as well, because if you
do use these tools and then it becomes like, well, I'm not using them in this way or that way.
Like at that point, you're just not sticking to your morals or their reasons for future.
Yeah. When the AI stuff coding stuff first kind of came on to the scene, you know,
Blue Sky at that time in it still is to a degree, to a lesser extent, I will say like extremely anti AI
at the time, like not touching it because this was like before people like, and it's actually
gone pretty good. I tweeted out about the accessibility benefits that you can get from using AI
for various ways. Yeah. Like not even that was going to get traction amongst some people. Because
it's like, there are the captions have gone a hell of a ton better with AI based captions. I mean,
like the accessibility things. Yeah. Geez. Being able to talk to your computer, hey,
click on the follow button, you know, yeah, not having a thousand times to get to it.
But I think it's the same thing. Like if you're an environmentalist or you care about where
data centers are going, not just AI data, data centers, but data centers in general,
but then you sit at home and watch Netflix. Like that's using data centers. So I think you,
I don't know, you have to fall hypocrites in various areas. Yeah. I don't think you
can completely abstain from it. And then just like, right. Like not go in a car, you know,
or live in the woods. Yeah. Exactly. I know. You kind of got to live in it, especially if you
want to be able to reach people to tell them what it is that you're thinking. You have like a
reasonable take on it. You know, you know, that's the other thing is you don't want to be this
obnoxious person where people write you off right away because of whatever reason. But I don't
know, I thought that was that was kind of cool. I thought I would bring that up for a bit of a
different conversation, but it's kind of a reward. Also, the other thing, this is somewhat related
is that I keep seeing, you know, like somebody who will break in, like I'm on, I love local Facebook
groups, big fan of the small town Facebook group. If you've never joined a small town Facebook
group, I highly recommend joining in on one because this type of stuff that people post is
absolutely hilarious. Like whose horses are these? You know, it's just like someone's horses Brandon. But
what often will happen is like somebody will get broken into or they're like their truck will be stolen
or whatever. And they'll have like some blurry camera and like a snapshot. And then immediately
somebody goes, takes that, pops it in, chatchy beauty, super crisp scaled up for and here it is,
I enhanced it. And it's like, that's not the person that you can't use that as like now,
let's hunt down this person. That's not what they are at all. And like, God forbid that, the person
that that was trained on. Yeah, I found it. It's Brad Pitt. Yeah, Brad Pitt is stolen your trailer.
Yes. AI upscaled these images because it's just it's inferring what the thinks they look like.
And that drives me nuts when you see it. And everybody's like, thank you. Good job. Wow.
Dude, Reddit does that too. Like crazy. Yeah. The Reddit like detectives are awful for that type of
stuff. Yeah. Part of it is the AI hype itself. Like all of all of the AI companies pushing their
product in a way that doesn't let people know like what it actually is good at. Yeah,
it's always that disclaimer. Yeah. It's always the disclaimer that's like, it can make mistakes.
Like don't you don't always trust this output. And yet, like we've been conditioned at this point
to kind of just trust the output. Yeah. Yeah. My favorite meme about that is the the gift of
the guy letting people in at a security event. And you're just like, all right, you and all right,
you and yeah, that's how we look at the AI code. And all the products that are AI as well.
Like I saw AI hand warmers the other day. And like, that's just an if statement. If it is this
degrees, then turn on. Otherwise turn off. That's not AI. That's an if statement.
I have what a released about talking about image generation. I can talk about Queen image 2.1.
Yeah. So they just released. And the thing that I'm excited about is it supports transparency.
So no image generation model I've seen so far has been able to output image latest one.
Does it? It can. Okay. I haven't tried finally. Not only fake one. Yeah.
But yeah, this was just cool to see because like so many times like again, like this is this is a
scenario where you could like the morals of like, do I you if I use AI to remove the background
of my image. Is it AI generated or like does that fall into that camp? I don't know. But like in this
case, it was able to remove all of the clouds and the sky needle and only leave behind like it's
almost like a perfect cut. Like you could be really hard press to do a cut out this perfect even
if you were like really good at Photoshop. I'll have to try that chat to be T1 because I'm almost
always trying to do like move remove the background, have some transparency and then combine other
images together. So yeah, or having it. Hey, generate it with a like a green background. And then
I'll just composite it out myself later. Exactly. But it's hilarious that most of them still just
like give you checkerboards because it's been trained on Google image search of like transparent
PNG. Yeah, I ever get a checker. Yeah, Google images search. Yeah. That's one thing I'm kind of sad
about with all of the stuff is that the people that work on these companies, like there's clearly
somebody that works at every one of these companies that's working on transparent PNGs.
And they're not, they don't talk, you know, they don't they're not publishing papers and are on
any of this stuff. You know, they're not coming out and saying, yeah, this is a big problem for
us internally, where or even the person that got rid of the purple gradient, what's the story?
How did you get rid of the purple gradient? You know, like obviously you trained it out of the
model.
But like, tell me, tell me the story of like, yes, what did you do to get out of the story?
We want that story.
Yes, I would love that.
If you're out there and you're like, hey, our website, we're looking like crap and they're
just all purple Bernie all day long.
So they all purple Bernie.
And all of a sudden it's just better.
Okay, I'll just go because it has to do with Jim image generation as well and it's just
a blog post.
So the title is AI generated posters don't have to be horrible and we've talked about this
problem on on the on syntax live before flyers for festivals and events and then also like
menus being like you can just tell it was you chat GPT was used to generate it.
And this post basically just gives you some prompts that you shouldn't use.
And so here's some exact like I've seen posters like this around Denver to like advertising.
You can tell that it's like that same chat GPT style illustration.
But essentially what they did is they came up with a prompt library of all kinds of different
styles.
I don't know, I think the people that need to see this post probably aren't watching
this video.
Yes, but maybe, but maybe share this with the people in your life that are using chat GPT
to generate flyers and everything else because they so they released like all these examples
of like how that same fire you can make it look decent and not look like a chat GPT
one.
This is a cool and like contemporary 40s poster for a Cubist exhibition like this looks
like somebody sat down and kind of designed the poster versus just generating it.
So they have this catalog of poster prompts.
And these are prompts you can pass into chat GPT and get more interesting flyers than
just the default summer festival.
I think really open everybody's eyes to how awful the AI poster is and like, yeah, I just
that's another question I have is like, why?
How is this not part of the the chat GPT product of like, hey, what do you think?
Here are three different like possible ways to go.
Yeah.
No.
Why?
Why is the way they all so awful like that?
I mean, in a way like I like it because now I know or I can really tell maybe that's
it.
It's a fingerprint of sorts.
It is a fingerprint.
Yeah.
Yeah.
I know.
Yeah, it is funny seeing that all over Switzerland on like mugs and t-shirts and stuff and being
like, oh, it's permeated to past posters and it's on t-shirts now.
I guess it does indicate which ones are the dog poop versions of that.
But I think I'm also just jaded on this style.
Like I was watching a movie from like 2009 last night and there was a they were like walking
through New York and it was just that classic 3D font that you get like in Microsoft word
that we used to see everywhere.
Yeah.
Word art.
Yes.
But yeah, word art.
Exactly.
But that felt like a breath of fresh air like we don't see word art anymore.
We see this chat GPT slot.
So I would much rather.
Word art folks.
But the thing is like we were so jaded about word art when it was remember.
Yeah.
We had a poster competition when I was in high school and I spent forever designing the
most beautiful poster for this thing.
And then the one that won used the word art like puzzle piece and I just and like the
little like thinking man, you know, like the black and like, yeah, we've been through
this already is just like the general public doesn't know any better and maybe we should
just let it go.
All right, guys, I got something to share here.
Are we done with this?
No.
Yeah.
Go ahead.
So we we've been hearing about a herder or other things as a means of being able to wrangle
all of your agents on your different machines and be able to use the terminal and that type
of thing every single time I talk about herder or this or that.
I'm seeing people talk about orca being an option.
And orca is an app that works for Mac Windows Linux works for everything and it basically
follows that same kind of thing as herder except for it is a gooey in an app.
So you can run your fleet of agents.
You get their status blah, blah, blah.
You have actual terms built into the app.
So you're using real terms in here instead of it being trying to gooeyify everything.
There's automations.
There's a built in browser into it.
There's get integration.
This is the type of tool that I've started to build a whole bunch myself.
I will say I have not used this folks if you've used orca in the chat.
Let us know.
I've heard so many people that I respect talk about how much they like this that I have
found myself wanting to give this a rip because herder is great.
But you know, the 2E stuff is only so much fun in some regards.
Like I really like the terminal personally for my actual agent stuff.
But then the sidebar 2E stuff really kind of sucks.
I think at least personally.
So I still use herder full time and I really enjoy it.
But I'm going to throw this up today considering I can just use my, I don't have to like
re-log into all my stuff, I don't think.
This is basically the AI everything app.
Like you've got a place for your agents, a place for review.
If it works, it also looks like it's just totally open source.
They don't even have like a pricing page or anything like that.
Yeah.
And they're YC now.
They weren't before they're YC now.
But it is, you know, it is a GitHub project and there's a lot of, you know, 74,000 stars.
So it's popular.
I just have like finally decided to open this and take it a look.
It does use Ghosty as the term under the hood or as Wes says ghostity.
So.
Ghosty.
Yes.
Yeah.
Check it out.
Orca.
Let me know what you think.
Breaking.
Breaking news, folks.
Rock.
For seven.
Just released.
Oh.
About 10, 20 minutes ago.
And rock for six, I actually, I, I, I'm allowed to say this now.
I have been using rock for seven for about a week now.
I was on their early access program and it's pretty good.
They got this cursor bench thing, which I always have a hard time trying to decipher what
these things are.
I guess more is better.
But it is as a sort of a rundown is it feels about as good as like an opus five.
And it's certainly not fable soul territory, at least from my experience, but the price
of it, man, price of it is like half the cost of like soul, they're comparing it here
against like fable, which is $10.
And this is $2 per million input and that output token costs $6 versus fable being 50.
I don't know if it's, it's not even in the same realm as like a fable or Astra, but
it certainly is in the same realm as opus, probably not quite as good as soul.
But man, I'm pretty excited, but it's actually pretty good.
And especially for the cost of doing this type of thing, if you just have something where
you're like, this is not particularly very difficult to stuff to work on, I just needed
to go and rip through a whole bunch of stuff on my computer or I need to build out a website
or find a bug on my react, react stuff and all of that type of stuff.
This is a pretty good option.
Yeah, I haven't really gotten into the rock models at all.
I got too many dang models to try.
Yeah, there's just so much good stuff out there to possibly try.
And I feel like a lot of them, they take a little bit of a touch to get to be effective,
like knowing how to work with Astra versus knowing how to work with fable versus even knowing
how to work with opus or other, you know, so like which models good at what, like 3D,
it's got you probably know this, but like the latest models like both fable and why am I
forgetting it?
Yeah, fable and Astra, amazing at 3D, yeah, I brought 47 not as good as at 3D, but if
you're not doing 3D modeling, that might not matter.
Yeah, yeah, I found it even be better too personally.
So they're pricing, like do they have a similar like $100 a month, like usage limit kind
of thing?
Or is it all API usage?
No, you get like a cursor subscription.
So you get the like $100 or $200 a month cursor subscription, which is quite honestly
one of the, the better deals right now because it comes with access to all of these other
APIs.
However, you can't get Astra on cursor right now because they don't get along and like
I, I bet it's only a matter of time that you can't get all the other models on cursor.
But will that matter is like when they release GROC 5 or GROC 6, you know, like you, you
got to imagine that they've been training their fabled competitor for quite a while because
what GROC 4.6 only came out a little while ago and then GROC 4.7 is just larger and trained
for longer.
So you got to imagine that they're cooking their big boy right now as well.
Same thing with like Google.
Like where's the Google one?
Where's the, where's the, the Microsoft one, we're waiting on that.
Yeah, Gemini was a deep SWE, which is like the, the one benchmark that everybody was saying
was the very best Gemini 3.8 flash scored extremely well on that benchmark and it got everybody
very confused by it, although I still haven't used it.
So I don't have any thoughts on the model one, however, it's just a flash model though,
you know, I know.
Yes, I just look at it.
But it still is pretty good for a flash model.
Maybe everyone wrote it off because it's a flash model.
I don't know anything about do Gemini 3.8 flash.
There's too many dang things to get on hold of, you know, I bet we're, we're going to
GitHub universe and we don't know anything.
Sometimes we know things.
We're not allowed to say them, but I'll tell you right now, I don't know anything.
But do you think it help?
Microsoft want to know some stuff, yeah.
They're going to release to be fair.
I do know when things are happening, but I do know that I have heard nothing yet.
But do you think Aguhub Universe is going to release like a Microsoft AI?
like they're fable competitor, they're asterisk competitor, you think so?
Is that the event to do it at? Would it instead be at like Microsoft build?
Yeah, I feel like it could be a Microsoft build.
It's not, it's not until next year because they do it every summer.
Yeah, that's, wait, it nine months to release a model right now.
True. I don't know, I don't know. I don't know if I'd even take that back because I have
no information. So I would say, my vet is no because I want to just take a different stance
and I'll put money on it, folks. But like, Microsoft AI came out almost two months ago,
right now, and they released four models, you know, and they're all pretty decent.
They have an image generation model. They have a flash, a really good flash model,
but they didn't release their big boy. And you got to think like, they have it,
but it's just like, if they're going to release it, it's got to be as good as as the big
boys right now. Can it top, Jev? Yeah, or Jev. Go the other way. Maybe we're wrong in trying
to go for the big boy models. Just give us a Jev and we'll be happy. That's all I got for today.
Just give us a Jev and we'll all be happy. Folks, what are you doing with Jev? I want to know
what you're doing with Jev. Is it actually working out well for you? Is it not? What are you doing?
What are you classifying? Classifying all kinds of stuff. Suddenly, everything is a classified
shape hole that I can put a key in to say it. Yes. I don't got anything else, guys.
Anybody else got anything before we get out of here? That's it. Okay. So check us out.
Stax live. We're going to be at Bearbottle Brewing Co. When? Wes.
October 27. Okay, Scott. I don't let me see. You've got to be ready.
October 27th, Bearbottle Brewing. Stax.fm forward slash meetup. Grab your tickets. It's free.
But you got to reserve your spot. We got some pretty six-way coming on the way, too.
We're going to be giving out there and check out Century as well. Sentry.io. We will catch you
in the next one, folks. Peace. Peace.
Podcast Summary
Key Points:
Apple's iPhone Duo simulator was introduced in a beta release with hidden device posture and viewport APIs, requiring feature flags to enable, sparking excitement and concern over delayed implementation.
Cloud Code added support for agents.md, but through a new "mod" system instead of standard file inclusion, leading to criticism over poor integration and lack of simplicity.
Jev is a fast, cheap, non-LLM model that answers questions with predefined options (yes/no, pick one, rate), making it ideal for structured tasks like email classification and tool selection.
Jev excels in deterministic, high-performance workflows, enabling faster, more accurate decisions in applications like chat bots, UI design, and content filtering.
Open-source alternatives like Leia and Kev exist but fall short in speed, accuracy, and scalability compared to Jev, especially in real-time, high-volume use cases.
Jev’s pricing model is not subsidized, and its value lies in high volume and frequency of use, making it cost-effective for enterprise and automation pipelines.
Local deployment of Jev is promising for privacy and performance, though current open-source models lack the capacity and precision to match its functionality.
The community is excited about Jev’s potential in real-world applications, including email filtering, security pipelines, and AI-powered personal tools, with strong interest in future multimodal and reasoning capabilities.
Summary:
The episode explores a range of AI and developer tool updates, highlighting Apple’s new iPhone Duo simulator with hidden APIs that require feature flags to activate, emphasizing the need for better support in web development. md support, but through a complex "mod" system that undermines simplicity and interoperability, prompting criticism. The core focus is on Jev, a novel, fast, and cheap AI model designed not to generate text but to answer structured questions with predefined options—ideal for classification, decision-making, and automation.
Demonstrations showcase its effectiveness in triaging support emails, managing tool calls, and filtering content in real time, outperforming open-source alternatives in speed, accuracy, and cost. While Jev’s internal architecture remains unclear, its performance and efficiency make it a compelling tool for developers seeking deterministic, high-throughput AI workflows. The episode also touches on local hosting trends, open-source models, and the growing momentum behind tools like GPU IKit, which enables high-performance Rust-based desktop apps with React-style components.
Ultimately, Jev’s real-world utility in business and personal automation—especially in email, security, and content filtering—has sparked significant enthusiasm, with developers eager to adopt it for its balance of speed, affordability, and reliability.
FAQs
Jev is not an LLM; it answers questions directly without generating text. It supports yes/no, pick-one, and rating responses, making it faster and more deterministic than LLMs that generate full text.
Jev is useful for tasks like classifying support emails, filtering content, or determining tool usage in chatbots. It excels in structured, deterministic workflows where speed and accuracy matter.
Jev is not an LLM. Its internal architecture is not fully disclosed, but it's designed to answer questions directly rather than generate text, distinguishing it from traditional LLMs.
Jev is fast because it avoids text generation and operates through direct question answering. It's cheap because it requires fewer computations and can be used extensively without high costs.
Yes, Jev can be used locally, and several open-source alternatives exist for local deployment, though they often fall short in performance and capability compared to the original Jev model.
Jev is not suitable for tasks requiring reasoning, like chess or complex problem-solving. It also doesn't support multimodal inputs like images, though this may be added in future updates.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.