Go back

If Developers Build on Chinese Open-Weight Models, Who Leads AI?

78m 16s

If Developers Build on Chinese Open-Weight Models, Who Leads AI?

The discussion centers on the growing importance of local, open-weight models in AI, as argued by AI researcher Sebastian Raschka. He emphasizes that these models serve as critical alternatives to proprietary ones, protecting against price hikes, restrictions, or unavailability, and promoting competition. This view is echoed by an open letter from major companies like Nvidia, Meta, and Microsoft, which links American AI leadership to downloadable, adaptable weights. Meanwhile, Chinese labs are releasing increasingly capable open models, such as Kimi K3 and DeepSeek V4 Flash, which perform well in tasks like agentic coding at lower costs, challenging the notion that leadership depends solely on closed frontier models. Raschka highlights the rapid evolution of transformer architectures, noting that modern models like Kimi K3 are far more complex than earlier ones, incorporating innovations like latent attention, latent mixture-of-experts, and residual attention. He advises learning these systems chronologically, starting with simple models like GPT-2 and adding components step-by-step, as understanding full architectures alone is overwhelming. The conversation also explores how the same open-weight model behaves differently across various agent harnesses, underscoring the importance of selecting the right tool. Ultimately, the episode celebrates practical learning, open source, and community engagement, with resources like Raschka’s books and architecture diagrams, while promoting workshops and collaboration to navigate the fast-paced AI landscape.

Transcription

14424 Words, 77013 Characters

English
I think local models are important because we also want to be prepared for a case where like maybe we have to use the one day. Like it would be sad if they were not an option and they are only proprietary models and let's say the prices go up like crazy and then you are kind of locked out or you have to pay a lot of money and then you always have these local models then as a alternative. It's good to have alternatives. It's what I'm maybe trying to say here. Like a competition is good for business. That was AI researchers, Sebastian Rajkar, explaining why local models matter. If access to proprietary models becomes too expensive, too restrictive or simply unavailable, open weights preserve an alternative. Three days before we recorded, 25 companies including Nvidia, Meta and Microsoft published an open letter called Open Wates and American AI leadership. They argued that American leadership also requires weights people can download, adapt and run themselves. Like Sebastian, who has been ahead of the curve, they tied open weights to competition, lower lock-in and matching each job to a model at the right cost. Chinese labs are already releasing increasingly capable models as downloadable infrastructure. Kimi K3's weights landed about an hour before we went live. If developers around the world build their products on Kimi, Kuan, Deepseek and GLM, does America still lead AI because an American company has the smartest clothes model? Four days after we recorded, Deepseek released the version of V4 Flash, a repost trained model for agente coding and developers are already reporting that it can debug multi-project code bases and stay on task across enormous contexts for a fraction of frontier model prices. The verdict is early people, but the acceleration Sebastian describes is happening. Sebastian is an independent AI researcher, the creator of Ahead of AI and the author of Build a Large Language Model from Scratch and Build a Reasoning Model from Scratch. In this conversation, he dropped an extraordinary amount of practical knowledge about Kimi K3, Kuan, Deepseek, GLM and the agent harnesses around them. He explains what changed and why the same model behaves differently across Kuan code, Claude codex and other harnesses. He also shows us when a local model is already good enough, why the harness should choose the model and its reasoning effort and when fine-tuning or a more expensive frontier model has earned its cost. Sebastian also traces the current open weight-wave from the original transformer into new attention mechanisms and hybrid architectures. His method is to implement one component at a time, compare it with the reference model tensor by tensor and follow the first place the outputs diverge. As he puts it, the implementation does not lie. I'm also incredibly excited about the next cohort of Master Agentech Data Science, which I'm co-teaching with Thomas Viki and Luca Fiaschi from PMC Labs, will build agents that explore data, run predictive and cause a workflows, challenge one another's conclusions and produce analyses that humans can expect and reproduce. The cohort starts August 4, which is tomorrow, if you're listening on release day. Podcast listeners can get 20% off the code and registration links are in the show notes. We recorded this conversation as a live stream. We live stream many episodes and run free workshops. The show notes include our events calendar, newsletter, YouTube channel and Sastian's books. If you enjoy the episode, please share it with someone who would get something out of it. I'm Hugo Bannon Anderson and welcome to Vanishing Gradients. We are live Sebastian. Welcome back to Vanishing Gradients, my friend. Yeah, thanks for having me back, I guess. I don't honestly don't know how long it's been, but it feels like ages have passed because so much stuff is happening every day. It feels like could have been three years, but it was probably like three months. Well, exactly. It's such a pleasure and it had been a couple of years between podcasts before, but it has only been three months, but so much has happened, including today. So I'm just going to do a bit of housekeeping and welcome everyone and then we'll jump in. Just everyone joining. Thanks for joining a live stream for Vanishing Gradients with Sebastian to, among other things, celebrate Sebastian's new book, build a reasoning model from scratch. So first of all, many congratulations Sebastian. Oh, thank you. Also, yeah, I'm here in this Zoom course. I can't see the people live, but if someone just joined, yeah, hi, everyone. So thanks for joining this. And yeah, and thanks for mentioning a book to such a good note. I just got like the hard copies like a few, two weeks ago. And you know what? This one is in color. This took me a lot of like effort, like talking to the publisher, but I don't know if you can see it, but it has color figures everywhere. And the quality of the print is, so this is kind of like almost twice as thick as the previous book, but large language model from scratch, but it's also almost the same number of pages about 400. It's much better paper quality and color, which is, yeah. One thing I was, you know, like really excited about besides the content. So yeah. Beautiful. And also something I hope we get to talk about is all the wonderful color transformer architecture diagrams you're building for your gallery. And we'll talk about the new Kimmy K3 one today as well. I just don't want to let everyone know Manning and Sebastian are super kind. They're giving everyone in the audience if I percent off Sebastian's new book. So join our discord and the codes there and put any questions there as well. We do that so we can continue the conversation afterwards because we can't always get to everything in YouTube. And just the obligatory, if you enjoy the conversation, like and subscribe and share it with friends. And we have a whole bunch of workshops coming up, including a building, AI agents for them, first principles workshop later this week. So check out the description there. And Manning is very kindly and Sebastian very kindly not only giving 45% off the book to everyone, but we're going to give away five free copies to the best questions. So please do ask questions in the discord as well. That is a tough one. What is the best question? Like, what is the criterion, right? It's almost like, yeah, we need e-votes for that, I guess. Well, maybe we can build them live, but particularly if we're going to, you know, look at Kimmy K3's weights. So this actually, Sebastian, last time we spoke, I think, when 3.5 was one of the most exciting models at the time and we were hoping deep seek v4 would be released before it, our last podcast, but it wasn't. And since then, we've had deep seek v4, GLM 5.2, Kimmy K3, Quintry.8 in preview, and now just an hour before we went live, Kimmy K3 weights are in the open. So what is happening, man? Well, to be honest with you, I have not looked at Kimmy K3 until today. And the reason was, it's like my just to stay sane. There are so many things coming out there every day. Well, in general, so I only look at those ones, the open weights are on the model hub. You know, like the announcement for Kimmy K3 was like a few weeks ago, but I was like, okay, I will wait until the model weights are actually out. And today they, they were just released on the model hub. And so I thought it was like one hour before we recorded, but luckily, it was a few hours before. So I spent my afternoon here. So you and Australia, so you just woke up to it, but for me, yeah, I spent my afternoon looking at the architecture based on data, very nice technical report. I was like speed reading that, looking at the architecture and making a diagram. And luckily, really, luckily, it was not that much different from Kimmy linear. You know, like some things were bigger. Oh, yeah, there it is. I just uploaded this maybe 20 minutes ago. And I've linked to it in the discord as well. And it's part of your wonderful gallery. Yeah. And so that is kind of like the summary. And it is overall very similar to Kimmy linear. So I could almost copy paste the architecture. So it was still, of course, you have to double check everything. The only difference here from based on what is shown is Kimmy linear. I think had 27 layers. This has 93 layers. And the other thing that is kind of like not shown here because it's too crowded is it has instead of the M E. It has a late M O E. I'm not showing it here. But if you're interested in that latent M O E, I think it's a Nemotron 3 Ultra architecture has that as well. So I do have a figure of that in the Nemotron architecture. So if you yeah, if you scroll up a bit, so yeah. And so it's basically Kimmy linear bigger with latent M O E. So Right. And we did discuss Nemotron in the last one. Oh, sorry. Yeah, that's the nano. That's the nano. The ultra should have that. It's a bit later. It came a bit later. You can also do on control F and just search for Nemotron. And yeah, I was okay. Yeah, that multiple Nemotrons, I think, weird. We'll find it in here. Here we go. That's yeah, yeah. Okay. Well, this is for the comparison. Sorry. There's another search field below. Sorry, before the first architecture starts. Well, just like that. The space. Nearly there. I have fantastic. Yeah. So if you look, if you are listening to the audio version, we are just looking at Nemotron 3 Ultra. And this one has like a latent M O E, which is essentially like multi head latent attention and multi head latent attention is it's like attention, but you have a latent vector that is like a compressed state. It's almost like Laura where you go from a big, like a mat mall, like from a bigger vector to a smaller vector, like the bottleneck basically, it's kind of like that. And yeah. So and the idea here and like an attention is is to the KV cache to store smaller K and V. So you compress the K and V into a smaller shared form. And then you store that in the KV cache instead of the full size K and V. And so it's basically, it's more like an efficiency tweak. And the MOE, latent MOE is kind of similar where you have latent experts that you, yeah, they're smaller essentially. So it's been a while since I looked at Nemotron 3, I'll draw them. But yeah, that's the main idea. Like the compression. And I am very excited to get into kind of what we've seen in the past three months with open weight models. I did, and I mentioned to you about this on WhatsApp before we spoke, but because I don't know if you recall, we actually did our first podcast in 2018. And so you and I have done, this is our fourth podcast. And so what I did was I vibe coded eight years of conversation from scratch with Hugo and Sebastian, which you can see here. And what we actually have, and I actually chatted with Fable about it. And we decided to visualize it as a transformer. The first one was on, as you see, on biology and deep learning, then LMS from scratch, architecture and agents. And this is what we're doing now. But if what, and what we pulled out is what the constants were. So four things that we've always talked about are learned by implementing practical trade off so high, teaching in public, an open source in community. Those are through lines over the eight years of conversations. And we go down, look at, we've got, you know, this is when we're talking about biopandus and MLX stand. And you're working on genome assembly and protein molecules. So quite different. But if we now go to our conversation a couple of years ago, you can see we, and people should dive into this, architecture and pre-training, deployment and monitoring, updating, continued retraining. And something I found interesting, we dove into Laura, a DPO and RLA chef. And then if you go to our most recent one, you'll see we jumped into a lot of RLV earlier this year, reasoning from Verifu I will rewards to Deepzy Kara, our one longer reasoning, parallel sampling, all types of hybrid architectures, sparse attention. So anyone who wants to check these things out, please do jump in and have a look. But there's so much meat on the bone there. I'm wondering if we can think through kind of the arc of our conversations and what we discussed last time. And I'd love your thoughts on what are the most important things that have happened since. - Yeah, I think to be honest with you, it's everything we talked about before except, so if you go back to, sorry, I have-- - Oh yeah, yeah, of course. - I found an interesting thing in there right now. - Yeah, yeah, let me-- - But so everything that we talked about before, like the LLM from scratch and everything, is still relevant as the basis. So like if you wanna dive into all these topics, I think that's still like the common denominator, but then everything, I wouldn't say it gets more complex, but it does, like it's not so scary if you do it like, one step at a time, but everything we see here, it still, it builds on the basic, let's say, GPD2 model, but like in the lower right side here, you say, or the title says hybrid architectures. And that I think that is the most, I mean, everything we see here is still used, being used, but that is I think the most recent trend if we look at architectures, like the Q3 next, the Gated Delta Net there, Nimo-Tron Mamba 2, is all those are going from the standard attention to like a linear, more linear type of attention that is faster, and we've just talked about K3, Kime-K3, and I kind of glanced over it and said, well, it's very similar to Kime-Linja, but I don't want to diminish it, it's Kime-Linja was like a research prototype, now this is a production model, and this model has the Kime Delta attention, which is very similar to Gated Delta Net, which you've just seen there, so it is all about making models bigger, but then also kind of like keeping in mind, it's not all brute force anymore, it's kind of really sophisticated now, also the latent MOU for example, they're all like tweaks in there, or the one thing that I think, we also glanced over is what's called the residual attention, you know, like it's yet another trick where you add, it's like kind of like a computing softmax and attention, or with previous blocks, and you add that as a residual, instead of just a addition, and that one also, I mean, it's not like a huge improvement in terms of modeling quality, this is not a efficiency tweak, it's a modeling quality tweak, but it's also like all these kind of little knobs, people are tuning, and this one for example, I think it's based on the Kime-Linja paper, it added like four or five percent training cost, and like two or three percent inference cost, but yeah, if you combine all these things, nowadays I would say, if we think back of the K3 architecture, we just looked at it's still the transformer, but it's like all these wires and connections coming off, it's like getting more sophisticated and complicated. I think we are kind of leaving the realm where you can understand everything by just looking at the code of one architecture. Like with GPT-2, you could have everything in like two, three hundred lines of code, maybe less if you simplify things, but let's say a comfortable two to three hundred lines of code and understand it at once. Now if you look at something like Kime-K3, I think it would be thousands of lines of code, unless you have modules and different files, but if you would have all the modules, let's say in PyTorch, but all the modules like residual attention, latent Mwe and everything in one file, there would be thousands of lines of code, and I think it would be very overwhelming. Like two, I don't think a single person can really do that anymore, and that is maybe the time where we partner with LLMs, but also we just take what we have, so no one starts this from scratch, what you would do is you would take an existing one, and then just add that one thing to it. So for Kime-K3, they started with Kime-Linier and added latent Mwe, they already had the other things. And so you do one thing at a time, and then it's doable, but we are still, I think the biggest trend is we are now adding all these kind of modifications everywhere to the transformer, it's kind of interesting. - Yeah. - There's like a graphic, I think it was from SpaceX, where they started with the original rocket launcher, engine, whatever, and it was all these weird wires, and then it got simpler and simpler, but now with transformer, it's the opposite. You start kind of simple, and now it's all complicated. Like a lot of things attached to it, lift and write, and they're in here, and it's interesting. But like, yeah, if you do go from one architecture to the next in a chronological order, you learn about one component, or two components at a time, and then again, it does make sense. It's almost like a timeline. Like understanding transformer is a timeline, because I think if you today start looking at Kime-K3 and have never seen a transformer before, I think that would just, brain would explode, I guess, my brain would explode at least. So yeah. - Well, that's heartening to know your brain would explode, 'cause mine definitely did. It kind of already has this morning, in a lot of ways, in all honesty, but yeah, it's heartening to know that. So I am, last time we spoke, we talked about when models, new models come out, you've proposed two ways of that you liked it to work with them. And so you said, if we take a working LL architecture, as an aga point, you can go in both directions, you can plug it into an agent or a harness, and see what's up, or you can look under the hood. And we're talking about looking under the hood at the moment. And I love you presented as a timeline. So I am wondering, maybe we can take Kime-K3 today, the open weights and think through. Let's say someone hadn't really looked at a transformer before. What would you suggest they start at and build up all the way to Kime-K3? I do think it's the GPT-2 starting with GPT-2. Luckily, I think it got rid of the rope, the rotational position embedding. So you don't have to re-worry about that anymore. They have the no position embeddings. Like, why would really, like, okay, what is a transformer block? What is attention? And then, oh, they did something interesting with attention. They have the multi-at-late in attention. Like really, like, top to bottom. And then the residual attention, attention residuals, they have a standalone paper on that. That's a tough one to. I think it's a relatively straightforward concept. And it's like maybe 50 lines of code. But it is, again, like, okay, we are doing something here. I would probably leave that for the end. And then because that is kind of like this cross-connection. And yeah, I mean, I would really, if you see something like latent MOE and you have never heard about MOE's, I would start with MOE's, same with multi-at-late in attention. If this is a new concept attention, then starting with the regular attention and then looking at multi-at-late in attention in deep-seek version three, for example. And then was it version three had that already deep-seek. And so there are shared components over time. And yeah, I mean, doing it like that, basically, going from. So in the gallery, from the top to bottom almost, yeah. Super awesome. And what I've actually done is a conversation you and I had several years ago, at the end, what you did was you showed me how to find you in GPT2, actually, in prepare to classification data set and tokenize it for GPT2. So I just shared that in the discord as well for people who want to check that out. So this is all in the spirit of peaking under the hood. Now I'm wondering, when we plugged these new models the past few months from GLM 5.2, KMEK3 and so on, and deep-seek V4 since we last spoke into the. into harnesses, it seems like the open weight models are really, so historically, overweight models have been maybe six to nine months between frontier releases, and it seems like that's really shrinking. So I'm wondering how you think about their performance with respect to our API, vendor labs and frontier labs. - Yeah, I would say for most, the frontier set of the R proprietary models, like the Chatchy-Peters and Clots and so forth, and even Gruck and Gemini. I think Gemini is due to release in one of the big models, so they are currently, it's always like the latest model is always the best model for like a few weeks, and then it's like leapfrogging almost. - And there are these rounds of waves, right? That's suddenly, despite the few waves, it's like, (imitates buzzing) - Yeah, I do always think they have some models ready that are still, they could stop training and release them, but they would rather wait a bit, but then if someone else releases it, okay, let's just two weeks from now, release our model early, and yeah. But I think, I don't know how, I mean, it's really hard to know, but I don't know how much the models improve really versus like they improve on the benchmarks and everything, and everyday usage, but I don't know how much of that is because they are like smarter in that sense, or whether they just have better, like how to use tools to get their job done, or like how to interpret the harness constraints, and more like that, like not the pre-training and everything, but more like the fine tuning in the environment, it's gonna get used. So like if the model is better at doing web search, it will be much better at giving you the right results, without hallucination, but really like how much of it is almost like polishing versus training, like the rough edges user issues users have when they run the model and like fine polishing those, compared to, I mean, they're still training models from scratch and fine tuning and RLHF and extending everything, but I think the improvement is almost more like on a product level, if that makes sense, like new features in codex, for example, and I think to be honest with you, if you ask me what is the most impressive thing about the recent releases, I wouldn't necessarily, I mean, 5.6 is better than 5.5, but you know what, 5.5 was already very good. So for me, I would rather say things like, it can use your computer now, you know, like in codex, you can say, well, I don't use my email client and do something there, you know, like physically use a mouse, take screenshots of your screen, whether you want that or not is a different question, but more like as a proof of concept is works, like it can do it and it actually navigate things on your computer and automate workflows like that, like visual stuff, and I feel like that's an impressive product feature, like it's like a next, I guess, step four LLM, they can do things they can't do before, but not really because they are smarter, it's more about the product, like how that's implemented, this image take, I mean, it has like multi-modal support, so it can extract from the images, what it's seeing, and it just, but it's more like a product implementation rather than mathematical training procedure that is different from the previous one, if that makes sense. And I totally agree, one thing that I've really liked in, particularly in codex, but also in AMP, I think some of the other, pie doesn't, okay, some of the other labs are behind is steering, so the ability to, if it's working, put another prompt in and have the ergonomics of it actually steering quite nicely and not kind of stopping and shuttering, and I think that's getting-- - Yeah, I do that all the time too, where I say something, I just get going, and then while it's doing it, it's like, "Oh, let me just be more concise," or like more, you know, like, let me add details, basically, is what I'm usually doing, like while it's working, 'cause I, for me, it's like to teach us to type it all out, I just, okay, start, and then I will give you more things later, basically, but yeah, I think it's probably better to do it all up front, but it's like a workflow we all developed at that point, like where it's like, yeah, you just, it's almost like, yeah, it's kind of life, you do it life, it's not like you think about something right down, submitted, wait, you kind of interact with it, life, basically, you interject, you insert things, and that comes all down, and to how the harness also works, I mean, how the harness supports contexts, and how it manages context also, and that is also one thing, I don't know, you may know more about that than I do, because I am not very, let's say, I do try different things, but I have my habits also, so muscle memory, but I tend to use, for example, but in that sense, I'm a bit different, I do use, or try to use the harness that comes with a model, you know, like for when I use the Gwynn code harness, for Kimi, I would use the Kimi code harness, and so forth. Although you can use models also in codex CLI and everything, or cloud, to some extent, cloud code, but I think nowadays still models are based on what I've been seeing in benchmarks, they are working better if you have them in the native harness. - I was gonna mention this, I think it's getting worse, I think models and harnesses are maybe overfitting to each other. - So you mean, I see, I see, so they perform better in benchmarks, but then on other tasks that are not benchmark tasks, they're all interesting. - Yeah, and I suppose I love you, I didn't quite think of this, but your thoughts on whether the models are kind of RBR on their own harnesses, and it just gets baked into the whites. - Yeah, yeah, and see, that is one thing that is not so clear is how people find you on their harnesses, like the models, if you look, I have to look into Laguna, the recent one, because they are also having a really good transparency level of technical reports, so the pool side Laguna model, and if they describe it maybe, but based on what I know from colleagues, I mean, I don't know the proprietary information at how Frontier Labs do it, because that's proprietary information, but I think it's mostly like you said, RLVR on their own harness just for the trajectories to get going to train something, you know, like the model to make it intelligent, but then later at the end of that, you do the fine tuning on harnesses, like some adjustments with a supervised fine tuning SFT basically, but that is like at the very end. So during the training basically, most of it was RLVR, right, and that wasn't one harness, and so it has seen a lot of that harness basically, but you might be overfitting to some extent, yeah, but if you say that they are overfitted to the harness, what would be like a hallmark of that, would they be using the wrong tools all the time, or how would that manifest? I would say on a harness level compared to a model, like a model's memorizing things. - Yeah, that's a good question, and I think, yeah, incorrect tool use, but perhaps also longer reasoning loops as well, because you can imagine something which uses its tool, like it may be the correct tool, but it may use it far too many times and shoot too many tokens or something like that as well. - Yeah, a cloud code is also notorious for that. I tried to use, for example, Q3.6 in other harnesses besides Q3 code, and I remember a cloud code was using two or three times as many tokens as other harnesses. But I think it's also about, in that case, not necessarily, overfitting, it's also just how efficient the harness compacts information, how much context it provides, and I think, yeah, that is almost like a consideration itself compared to what LLMs do I use for, to get something done efficiently, what's the speed spot the Pareto Frontier, the same is true for the harness, like how much do I really need everything from that repo in context, or do I only need a few files? And managing that, and yeah, it's like a whole thing. It gets complicated if you want to optimize along these different dimensions, right? Like the good combination of model and harness, and then you do that for a few months, and then the next model comes and you have to start all over again, and measuring what is, so yeah, so I honestly don't really benchmark things myself. I just use them and you kind of get a feeling, okay, this is not working well, this is working well. I do look at public benchmarks, but I'm not like rerunning any of them because there would be also very expensive, like thousands of dollars for what you basically, you know, like, so I'm just usually like trying things out using them and you kind of, I would say after half a day or day, you get a feeling of like, well, this is really bad, or maybe I'm not using it well or something like that, I don't know. - The other thing that you mentioned is the ability to speak to these models with images. And I can tell you, a huge amount of my workflow with agents involves voice. So I speak to it a lot, and then dragging and dropping images, and then I get it to speak back to me in HTML, mark down whatever it is, and HTML, because, you know, we're all sick of reading so many tokens and words at the moment, so getting it to generate images can be useful. But so the GLM 5.2 was really the first model that when it came out, the first open weight one where I was like, oh, this is really close to, this feels really close to frontier lab. Stuff based on vibes. I couldn't drag and drop images into it, of course. I couldn't with deep-seeing V4 before that. But then, Kimi K3, I can drag and drop images to it. - With model, yeah. - Exactly. So thinking about just the experience, and this is like between model and product as well, it does feel like with those types of affordance, the open weight models are getting even closer as well, right? - Yeah, and also I would say, Kimi K3 is 2.8, I think, and GLM 5 is like 270. billion. So it is it is magnitude bigger to it. So four times as big also, you know, it's is 2.8 right? Uh, 2.8 I think so. Uh, I mean, just my brain is also from looking at numbers today. But, um, uh, yeah, so it is also the trend things are getting now to the, I would say this is an um, outlier, but I think we are also moving towards models that are not super local anymore. So it's like, yeah, you, you probably ran it through open router or their own, um, API, right? Or you have no, uh, yeah, oh, yeah, oh, yeah. Oh, yeah. Yeah. And so it's also, they are open weight models, but it's also open weight doesn't mean, uh, running things locally, you know, even though that is also appealing. That's, that's yet another level for that. For that, actually, for local models, I just, um, uh, this weekend, I tried, uh, Laguna, I just briefly mentioned that. It's Laguna S 2.1, I think. And so this is a nice model because I, do you want to speak with you? I don't think it's, it doesn't feel much better than, let's say, uh, quen 3.6, uh, the 35 billion MOE, but I have not used it enough. So it might be better. So basically on benchmarks, it suggests it's better, but it feels about the same. I don't know. Uh, but it is, uh, interesting because it's kind of like maxing out the, what you can do on a, let's say, DG X spark, like the DG X spark has about 120, uh, gigabyte, uh, RAM. And that one uses about 80 90 gigabytes, even the non-contexts. So it's like a nice, it's like where you feel like, okay, this kind of max is out my machine, where the quen 3.6 35 billion, you have a lot of headroom, you know, you have used 40 gigabytes and then you feel like, oh, I'm wasting, I could be running a bigger model here, basically. Speedwise, it's two times slower, of course, but I feel like it's like a 20 30 tokens per second. I think it's still plenty fast. It's like using chat GPT based. It's also 20 30 tokens per second. So I feel like that is like a level where yeah, it is usable. It's not prohibitively slow or something like that. So like a comfortable speed, basically, yeah, of course, faster is better, but for agents, especially if you have longer loops and stuff like that. But I feel like 20 tokens per second is, is okay. Yeah, totally. And I'm really glad you mentioned the distinction between open-weight models and ones that you can run locally as well, because we've asked people to send in questions and we actually have a lot of questions about small language models and what you can do locally. And Pierre has actually asked a series of questions which are kind of summarized because I think it captures the spirit of what people really want to want to know. So the question is really revolves around what would you recommend in terms of a setup for running small language models like Quen 3.5, 9B. I'll take that up to the 30 is a 32 or 35 Quen model. And besides coding agents, what other use cases have you tested, SLM's 4 or think about testing and related to all of these things, do you think SLM's getting closer to frontier lab models as well? I think everything is moving in lockstep. So I would say the frontier models are getting better and at the same time the small language models are getting better. But I think they're kind of like moving in parallel. I don't think there will be an intersection where they at some point meet or where it's like the slope is the same almost I would say, but this is without any scientific measurement, it's just like a gut feeling. It feels like yeah, it vibes, right? It feels like yeah, it feels just like we you feel like okay 3.3.6 35b is so much better, closer to frontier level than the previous one, but then you look at where the frontier models move towards. But then also what changed though is I feel like what is good enough before like two years ago maybe like a Lama 3 model that you ran locally was not enough. You could barely do anything with that, but you could do a lot of useful things with a frontier model. Nowadays you can do a lot of useful things with a grand 3.6 35b model. It can do basically all most of the coding pretty well, but then you can even do better things with a frontier model like a more sophisticated things, you know, like we have seen with these incidents is that they're more like marketing almost, but it really happened that Hagenfase got hacked, right? It's like something like that. I mean, I'm not saying this is like it's amazing frontier models for that, but just showing like what saying they are also pretty capable of things nowadays where the in the GAM 5.2 model was able to find out a bit in terms of what happened, you know, and I don't know that is anecdotal. We don't know if you had used a grand 3.6 35b model if it could have told you similar things, you know, so that I don't know because that is also expensive, right? It's like when we use these models, we don't really a B test, right? So when you work on a daily on a task, you don't say, oh, let me try this model, but let me also try the other model and then everything takes twice as long in your everyday life because you are always running things multiple times. So no one really does that outside benchmarks. Ideally, we should probably be doing that occasionally, but that is this is why it's hard to say really why how different they are, especially also I have I'm kind of, you know, in an ideal world, I would use grand 3.6 for more things, but I'm also kind of biased in terms of, oh, I think it won't solve that. So I'm already using this Gachype model, for example, the latest one because I feel like, okay, I know for sure this will get it done and I have only so much time. So let me just throw the thing on there. I think it will change in the future though, when right now, I mean, they are pretty well subsidized, you know, they're talking the plans, especially the subscriptions. I think it will change at some point where you will probably run into your limits sooner and then that will be the point where you will automatically also use more local models. Right now I use local models, I have to be honest here, I'm using them for just for fun because it kind of feels fun to use them. It feels satisfying. It's like, oh, it's on my computer. I have this spark machine sitting there. I also, I mean, I could run them on my Mac mini, but I don't want it to get too hot to be honest. And I also, I have the version, luckily I got it like one and a half years ago, no, before the RAM prices were high, but I have the maxed out RAM version, the 64 gigabyte. So I can run the Grand 3.6 model on that, but then it kind of slows down your computer and I'm doing other things on that computer too. You know, it's like my main work machine. So I usually run these models only on the separate DJ X Park. And it's just fun though. It's like you sometimes hear the fan going on. It's like, oh, I'm actually getting my bank for the buck here. I'm using my own hardware. It kind of feels good. It's like people working on their car and they go, "Ratch, you know, you don't have to. It's like a hobby, but it's kind of satisfying. It's fun." And, but so that I wanted to prefix that I'm still using a lot of non-opened models because, but also because the subscriptions are relatively in quotation marks cheap, you know, where it's like you get a lot for the subscription, where until I reach the limits in terms of API usage, I'm sorry, subscription usage, I am not forced to use local models. The only exception is if I want to do something where I feel like, "Oh, I don't want them to have my data in that case." You know, like, I don't want, wow, let's say medical, with medical questions. Of course, you probably shouldn't ask all the important medical questions to L&Ms, but talk to a doctor in addition. But, you know, sometimes you want to know things and like your own research and maybe you don't want them to have your whole medical history, right? So in that case, it's like a mix between, you know, curiosity, like as a researcher, I do like to use them also to kind of stay up to date and it's fun. Things that are proprietary or sensitive data, I wouldn't feed to the proprietary L&Ms, but I have to admit I use both, you know, so it's like, it's not that, it's not that I think proprietary models are also bad, but I think local models are important because we also want to be prepared for a case where like maybe we have to use them one day. Like, it would be sad if they were not an option and there are only proprietary models and let's say the prices go up like crazy and then you are kind of locked out or you have to pay a lot of money and you always have these local models then as an alternative. Like, competition is good for business. Very much so. And I've been using Quen35 a lot more locally with Pi actually. I was inspired by people watching may know this, but I do this series called Show Us Your Agent Skills. It's really showing people showing us how they use agents, not just skills. And we had Tom Arstongles the VC on the show and he showed us how he uses Quen3.5. You can see his stack here locally with Pi and besides like big code-based changes, he can get away with using it for nearly everything. And you know, 120 tokens per second. He has a system. He's built his own whisper flow essentially with 300 millisecond dictation latency. Oh. And so. A short question or a quick question. Sorry. What hardware did he use for the 120 tokens per second? Do you know? Yes. So Mac, these notes say a Mac, M, I don't know. I don't know. I think it comes close. Yeah. I think for the M4, I'm getting maybe 60, I think 80 around 80. But I'm also not mind is probably not super optimized. But although I'm using the M and X version. But yeah, you know, it makes sense. That's pretty cool. 120 that is. So I think it's faster than like say a GPT 5.6 or even 5.5 so it is. And you said he's using it for pretty much everything. That's pretty good. Exactly. Yeah. And I was inspired to use it for a lot of my work. I also, it works for a lot of the code I need. And I'm on a, I've got a Macbook Pro, it's the M-Ball Max, but I've got 128GB of unified memory. And it's so powerful. But I also do a lot of like free form text with my coding agent. So even, you know, to reason over the three podcasts we've done over the past eight years, or, you know, all this stuff Tom, Tom, I showed us. It helps to have, I've found frontier models better at that than Q135B. Okay. Quick word from our sponsor, which is well me. Vanishing Gradients is a podcast, workshop series, blog and newsletter, focused on what you can build with AI right now. Over 70 episodes, hundreds of hours of free hands-on workshops, all independent, all free. If you want to help it keep going, you can become a paid subscriber on Substack, starting at $8 a month, or share this episode with someone who'd get something out of it, links you're in the show notes. All right. Let's get back into it. Yeah. So when we go back to the beginning of the podcast, you showed me this, um, this graphic, the overview, the trend, like the, not the transcript, but like a timeline of things. And that was with the proprietary one. Yeah. Yeah. So that was with Fable and then not one six. Yeah. And see, this is the thing where you just know that it will do better there, right? It's like, it's just like something a muscle memory. You could probably try open weight models, maybe, you know, KMK3, but KMK3 might actually do a good job with that too, but, but it's just like a gut feeling at that point, a muscle memory. But like, yeah, I mean, I have enough usage, in my subscription, which is used that, you know, it's like, yeah, because you know, it will get the job done. But I think for let's say 90% of the tasks, KMK3.6 model is probably sufficient. And KMK3.6, if we think about it, is almost ancient. It's like April, something like that, like four months old. Yeah. So it's like, yeah. And it, yeah. So it's, and it's pretty good. I mean, it's like you said the one person you talked to the, I had Thomas, Thomas, yeah, he does most of it, almost everything with it. And it's pretty impressive, like with a four months old model. So, like we are getting to a point where models are still getting better, but it's not always needed, just maybe the right answer. It's like, for those cases where we really need the extra power, if you want to solve a new, let's say, math Olympiad style problems or unsolved theorems and one of theorems, but like conjectures and stuff like that. And where you maybe want to have the most powerful model available and you don't care how long it takes, you don't care how much it costs. But for everyday usage, maybe also the right way to summarize it is, is there's no one size fits all. It's different use cases, price points, privacy considerations, yeah. Without a doubt. And something I do, even when I don't use the hardcore frontier models every day, and the frontier models are getting better at routing. I mean, fabule is a fantastic orchestrator and great at routing things to tasks to, you know, the correct model, whether it's Opus or Zona, or finally, actually just chatter with Chip when this morning, who's coming on our show with your agent skills later this week. And she's built a runner inside Claude code for herself or a adjacent to Claude code, which will get fabled to send tasks to GBT models among other as well. So it's pretty powerful that you can do that. Yeah, yeah. There's also maybe a good sequit. I think that you asked me at some point what is kind of like coming or like the trends. And so I think it's interesting. I remember when GPT 5.5 came out, there was this auto mode like routing to or like at least setting the infant scaling budget like the effort level for reasoning. And I think they moved away from that again. Like so I haven't seen that in the UI anymore. So in my, at least in my UI, it's not there anymore. You have to manually select the effort level. And so just to also go back a step. So you have basically now or like a prepare, even like a proprietary model like like Claude or GPT, you have different model sizes. They're all like independently trained except the smaller ones are kind of distilled from the larger ones. But you have like as a user, you have to choose between the usual terror or lunar different sizes. They have different costs and tokens per second. And of course, also performance in terms of solving the problem. But then for each of them, you also have effort levels like medium, high extra high ultra and so forth. And so that basically two levers moving forward, like making bigger models. But then also the effort level can be controlled with RLVR. It depends on mechanisms where you can set a penalty for length, a KemeK3 based on the report that just skimmed today. They do it for example. They have a budget, a token budget. So they say for example, you can only have so many tokens if you exceed that. You get a reward of minus one in the RLVR. And the model is discouraged from doing that. But then they have different, they call them specialists. They have one for very long, high effort, very long outputs, some for medium, some for light. And then they distill that down into one final model. And then you as a user, you can select that with a system prompt. I mean, this is not new. Most of the models do that even back the GPT OSS model from OpenAI had that in the system prompt. But you as the user, you kind of have to decide is what I'm trying to say. So you have to decide what side, even if you are staying within one provider, what sides do I use, what effort do I use? And you have to develop, I mean, we kind of through using it have like a kind of a gut feeling. Like when you said you used Claude for this one timeline, you kind of like intuitively know, okay, this is like a high end model task. But I don't like, like you said, chip has like now an automatic mode. I think, yeah, this tedious to not tedious, but it's maybe also not optimal because we human have biases where we have a gut feeling. It can be a very good gut feeling, but we might also be just wasting compute on simple problems by just always using the most expensive model. And the other way around, you know, we may think we get away with that. Although that is usually easier to correct because you if you start with this simple model, you get bad results, you just switch to the bigger one and hope things cross that it works. I think that is more natural. The other way is more the blocker. You use the most expensive one because you only think that solves it. And then you just waste a lot of waiting time on the results if you didn't own money and usages and stuff. And yeah, and I think that is something I do think would be best handled in the harness itself. Like the harness should based on the context. So not like the user turning the knob, it should be a harness feature in my opinion, like where the harness based on how many loops like an S could be a different model, could be a heuristic or a mix, I think, ideally a mix of the two, where based on the task, but also not just the task, but the history, the context and everything to decide what model to use. And then of course, change it inside the, even like inside a loop, like we can steer at a them, you can also basically notice me, okay, this was not powerful enough. And you have the context from the LM, you can just pass it to the other model. Or you change the system from like with steering, you can say, from low effort to high effort, you don't, you still use the same model, you know, but that I think is something no one is doing it as far as I know. Maybe please correct me if I'm wrong, but I think because it's hard, right, it's not easy, it's not trivial, but I think that is a good optimization in the future. Yeah, yeah, I haven't seen that yet. I'll be very excited to. The other thing when we're talking about product around models with harnesses around them is it seems like the models and harnesses are understanding a bit more about where they're situated now. And let me be very explicit. Previously, codex or Claude wouldn't know about the right-hand panel, where it displays HTML. It just wouldn't know what you're talking about. And when you say actually, can you display this not in line in the right-hand panel, now it knows about that. So it's aware of the user experience, I think, which is actually a really important affordance for these products. It's also impressive how it can, I'm still surprised how, but it is not good at is making images, in my opinion, but how good it is at understanding images, for example, like even like subtleties, because ultimately it is in a multimodal LM, it's chopped up smaller patches, but how well attention works to understand the big picture. For example, if you have a diagram, let's say take the Kime K3 diagram, I'm pretty sure you could ask an LLM to, if I do something wrong there, I do that sometimes, I throw it in there and then, oh, you forgot that arrow there or you have a wrong connection there. It can't draw for you, but it can point out, or this is not always right, you have to think about it still, but it's impressive because it goes over multiple patches. And it's like, it has to kind of, there are so many things going on, it has to kind of reconcile the connection. And that attention is, yeah, I mean, it's ultimately attention, but it is actually working. So that's kind of still fascinating to me how how well it works. And like we know it works for tokens and words and code and everything, but But beyond that, it's just pixel, it's very fascinating. It is fascinating. So we've got lots of great questions in the discussion. Yeah, we have been just rambling here. Yeah, exactly. Well, no, no, no. But we'll definitely get to a couple of them. Luckphere has a great question with models reaching as large as several billion parameters. What are the challenges to self-hosting? And how do we get around them? And Luckphere said, even renting GPUs doesn't always seem possible. They're always out. Yeah, I think that is a very challenging question because the problem is, for the bigger models, you just can't host them on a single hardware. You need multiple. And then also, even as a consumer, you can't just buy a server GPU. I mean, they're also, I mean, even if they were sort of stock and even then they would be very expensive. And even then, if you could afford it, you probably don't want them at home. I had, when I was at a university in 2018, my office, I had, well, it's a story about the workstation. And it's meant as a workstation. It was supposed to be in your office, you know, like four GPUs and they were the 1080 GTX 1080, like ancient by now. But even those were so hot and so loud, I couldn't bear it. I had it in my office. And I, like, a month, we moved it to the server room because it was so loud and so hot, it was very uncomfortable. You know, like, it was like, if I had office hours, I could barely hear this. I mean, it was just, you go crazy. And so, yeah, it's really challenging. And these GPUs get hotter and louder and bigger. And so it would be really hard. And if you have a basement, yeah, there would be an option to have them in your basement, but basements can flood. So it's like, do you want that expensive hardware at home? You know, I, you can put them in a server room somewhere. You can, I think there are places where you can run as an individual person, like small server rooms, you could buy something and stick it in there. But then at that point, is it really worth it? I don't know. You could technically just rent also an AWS instance, you know, at that point. But then they are also out in expensive. And it's like, absolutely. Yeah. So I, I mean, I'm not affiliated with any of those. But like you said, open router for now, you know, where it's like, it's not, you know, the difference is it's not fully local. I still, it's going through something and you don't know exactly what happens to your data. For example, but if you don't have sensitive data, like if it's more like things you are comfortable giving to chat GPT or clot, I think that's probably similar to open router. I think open router is still a bit different because there's open router itself. And then they do partner with other inference providers and they pick the one that is currently most, I guess, available and cost efficient. So you have multiple like stepped on multiple companies involved in that. And so yeah, I would probably, I honestly don't know. They might be handling your data super well and stuff. But to me, it's always like just like a rule of thumb, if, you know, if it's not on you, if it goes to the internet, it's not fully safe and private to me. At least it's like a, I don't like a common sense thing. So I think if you are a bank or like a very sensitive medical information, you probably can't or should use those models. But you know, for tinkering, like let's say fix my website is public anyway. You know, like I think that's a good alternative to use open weight models. They're not local. I mean, the question was about local models, but that would be one way to test models at least. And then I think that's also one thing. I do use my local models on my DJ X, but I also don't want to download and install each model because they are large. I mean, I think it has four terabytes, but you also at some point, I also have a lot of data from experiments on there, like checkpoints and stuff. So I'm not like unlimited in space there. And if you just want to test out a new, let's say the Laguna model, it's roughly the billions of parameters is roughly, I guess, the gigabytes in storage it takes. It's another 100 gigabytes. And it takes two hours to download it. I mean, in this case, I downloaded it because I did know that I wanted to use it, but I wouldn't do it just for all the models because that would take a long time to download it. And you have to switch them. So it's just easier to, if you want to just quickly check something to just use a router, I guess, to quickly try something out. And if you find, oh, this is worthwhile, then switching to it locally, if you can run it locally, I would say, you know, absolutely. And I do think something on hearing that, which is a big philosophy of mine is explore the art of the possible. Like when you want when frontier labs, release new models, go and see what you can do with them. And you can do every workflow you have. But similarly, with like, in EK3, explore what is possible with an open router endpoint or base 10, I've had a lot of fun with basic in the past and they're great. I'm not affiliated with them, by the way, they have sponsored some of my courses, but I do think. And then if you feel like it could really work for you, think about if you want to host it yourself and these times. It's just also the hosting itself. It's somehow fun. I don't know. It's just fun. It's fun and it's empowering. And this is, we discussed this last time, right? Maybe that's the right word point. Like it's like, it gives you a, you know, it just makes you feel good. Yeah. And last time I did ask you, I asked you like, why do you think there are hundreds of thousands of people who want to learn how to actually build transformers and large language models and paraphrase you? But you're like, firstly, it's fun, but also it makes you feel like you understand more and can do more. And of course, you said the implementation doesn't lie. So there is an aspect of truth, truthiness to it as well. Yeah, I think that is absolutely true. It gets tricky with the larger models where you probably can't do this with kidney K3 because it's 2.8 trillion. But you could scale it down and implement it and see if it works, you know? And there is something about, I don't know, there is like, it's usually frustrating, but if it doesn't work, I think it's the best feeling if it doesn't work, if you don't understand it and then it works and then you understand it. It's just like this pure joy. It doesn't mean it, I guess, I don't know. And that's the same with running a local model. You have the setup and you don't fully know you don't own it and then you ask it a simple prompt. I usually don't know why, but I do usually do one plus, what is one plus one? I don't know why it's just like the thing and then that I use as my first one and it starts answering. I'm like, super happy. Well, it actually works, you know, it did it. Yeah. I think the one plus one is actually, it's a stupid prompt, but it is actually also quite useful because you, if we go back a few years, a model would just say two, right? But then if we go forward a few years and maybe we are in, let's say, 2025, beginning of 2025, and if you ask a model what is one plus one, it would not say two. It would give you a whole 2000 token answer, like a deep seek R1 model. And so it's actually not a super stupid prompt because you can cast how efficient is the model actually because the newer models, again, they would answer to, you know, like the smarter reasoning models. But the first generation of reasoning models, they would really be a 2000 token answer and that's also not good. So you want the model that is that can be also efficient when it needs to be, yeah. Without a doubt, I love that. Or you will also in that case, by the way. Well, exactly. I mean, get it to write some Python code into executed in a sandbox. Of course, probably everyone's jolo these days actually in any way. But I am, I love the idea of doing one plus one kind of as a timeline to see what the responses are over, over the, over the course of LLM history. We have a great question from pastore, which is after analyzing all the architectures you have, and of course, rebuilding many of them. Where do you think what can be improved? What's next? Yeah, I do think the, we talked about it a bit briefly earlier, like the residual path. It's just a trivial thing, but it's like something that has not changed in a long time, like the residual connections. And then it was not invented, I would say, by deep seek version four, but they popularized it. I think they had another paper on that before, like a prototype. They always have like a prototype paper and then the production model, but like the highway connections, I think they're called like, and they, it's like you have instead of one residual path, you have multiple, I think in that case, they were four. And that was one way to improve residual connections. And then right now, Kimi K3 from Kimi linear, the residual attention, is it called, I think, yeah, residual attention. That was a standalone paper, but like looking at components, we have not tweaked before, you know, like little things. But again, they are like little tweaks, they're not fundamentally new tweaks. Oh, there was also another one that was interesting. Let me think one of the, well, I did look at six architectures this weekend, so I forgot which of the six it was, but there was one architecture that used the looped transformer style where you, I mean, they're simple, again, very simple things, but normal, they had 22 transformer blocks and they looped back to them, you know, like they, so they had two times 22. So you have twice the compute because you go now through 44 layers, but you only have the say, like you don't increase the size of the model because for the repeat, so for the 22 layers, you use this same set of weights for the second time when you loop through the 22 layers. So it's just, it's like a tweak again. I don't think that a nice applation study in the paper or technical report on for Tuesday. So I don't know how much that actually improved, but it was worth it to actually do it. And so I think right now we see a lot of these things, but most of them are at the very beginning, we talked about hybrid attention. I do think most of the bank for the back comes from the hybrid attention. Like each one has like some flavor of it. Like Neymotron has more like Mamba two layers and then the Gated Delta Net in Q3 and then in Kime Delta attention, which is very similar to Gated Delta Net. Like these are kind of complicated components if you look at them compared to the normal attention, but they're not super complicated. It's almost like swapping out that module. And yeah, and I think that's a lot of incremental progress. It's not really throwing everything out there, like what we already have. And it's not replacing everything. It's just very targeted, you know, like replacements. Like if you have a car, you don't reinvent the car. It still has four wheels, but you swap out the engine, you know, like no, we have electric motors instead of gasoline motors. You know, like it's like targeted replacement, if that makes sense. - That does make sense. And I am wondering that like early on in this trajectory, we saw a lot more of a focus on pre-training and in the past year more like gain, it's gonna be had in post-training and in front time scaling and these types of things. And it seems like maybe we're going back to putting a certain focus on pre-training as well. So I'm just wondering how you say it. - Oh yeah. Yeah, we're at the point. - Yeah, good point. I just talked about the architecture, right? But the reason why I also talk a lot about the architecture is because that's known. And for the training, not everything is so transparent. Where, yeah, it's some articles share quite some information, but it's not what it used to be in it. Back in academia, 20 years ago, where everything was fully reproducible and clear. And so it's some of them, it's some of the stuff is mentioned, but it's not always even sometimes it's ambiguous. They say it in a sentence, but it's not 100% spelled out. What I explained with these specialists, the effort levels, I remember there was one of the papers, they just said they had these specialists, but I didn't even say like, so I think for Kimeke 3, they were, when I recall correctly from speed rating it, they were pretty clear. They had three specialists for the effort levels, like levels of specialists, and then for each domain like math, coding, and general knowledge. So they had nine in total, but there were papers where they said, well, and we trained these specialists, but you don't know, was it like three by three, or was it like in that case, six specialists could also be six, could be a light, medium, hard effort, and then it could be math, coding, and so on. So like, is it combined? Like, and so a lot of that is speculation. Also sometimes when you read these papers, but like you said, pre-training is still very important. I think what's changed is there's like a lot of more good stuff in the pre-training data. So for example, I don't know, five years ago, four years ago you had probably very few chain of thought type of data in the training set, 'cause nowadays because the models produce synthetic data, also for the pre-training to enhance it. So my one thing to say here is I think synthetic data is not necessarily bad in certain amounts. It's almost like kickstarting the model. In a sense that if you have the messy data, you learn everything from scratch from messy data, but if you have a portion of high quality data, even if it's synthetic, and I mean, this is the common discussion. If we only train on synthetic data, we will never improve. It's kind of like going in circles, but I do think it helps the model initially to get really good, I guess, formatting and style, and what is what we want the model to produce compared to just really messy data. And so if you mix it right, you can probably not get the model to all perform the models that have been trained on bigger pre-training sets with more, let's say, added on post-training, but you can get to a baseline quicker. Like you need fewer iterations, fewer data to get to a certain baseline level, and then you can focus more on the post-training also, like where you free up the compute. You are basically less wasteful with pre-training. But at the same time, the pre-training data sets also get larger. I don't recall from the Kemi report, it was just too many numbers, but might have been 30,000 tokens, if 30,000 tokens, sorry, if they mentioned it, or maybe I'm conflating it with another recent model I looked at, but I think 30 trillion is currently like the plus minus common size nowadays, which is about, oh, sorry. So about two times, two to three times bigger than last year, for example. And so yeah, we are still increasing, or I'm saying we the big labs are still increasing the pre-training data sets, but at the same time, it's also better pre-training data, more synthetic in there, better mixes, and it's all being improved at the same time. So it's not just making a bigger multiple levers to pull. And then of course, the post-training. And so for the post-training, it used to be just also like supervised fine tuning and then on long context fine tuning and stuff like that, on higher, so because you don't have infinite long context data that is good quality. But now with agents, you do again, like you have a lot of trajectories you can train on. And I mean, the agent harness is also like we talked before RLVR, like just if it can solve the problem, but then for the fine tuning, you use the, it's kind of like a distillation from the longer collected, successful traces from other models, or even your own model. You also use SFT for that. And there are multiple stages going back and forth. Now again, too, you have RLVR, then going back to SFT for effort modes, some distillation, then doing another preference tuning with RLAHF, and you may do another RLVR round. So it's not a new technique in there in a sense, but it is more complicated, it's a complicated pipeline and it's more steps. And so it does do a lot, I think, yeah. - Yeah, very much so. And I suppose I am, it'll be time to wrap up in a minute. I, so your first book on these types of things, and you've got many proceeding out, we're not going to go all the way back to heat maps in our today, though. - Oh, yeah. - Was your first book on these types of things was Build a Large Language Model from Scratch. Everyone should check that out if you haven't. You've just put out Build a Reasoning Model from Scratch and everyone in watching the code to get 45% off his and the discord and will be giving away five, three copies and will email people afterwards or message them. These are huge tasks that the whole community is incredibly, including me, of course, incredibly grateful for that, you're helping to democratize all these amazing skills and ways of thinking. It's a lot of work, though. And my next question is, were you to do a third book in this series? - Oh, yeah. - What would it be? And what, like, 'cause in terms of the time you need to put into it as well. - Yeah. I think the answer is obvious, innocent. Sorry, I don't mean it. It's a way to, I mean it more for myself. So I'm not saying it should be obvious to you. - It's obvious for me, right? - You know what I should do, but I don't know when and if yet. So I'm saying that, smiling a bit because it is, if you think about the progression, like, at an M's from Scratch is the foundation, like mostly the architecture and some pre-training, that's, and also I try with the books, you know, I don't try to cover research papers or anything. Like, that's, I write about them on, you know, on my blog or social media because they're very time sensitive in a sense that they are relevant for a few months and then there's something else. I try to focus more on the foundation so that everything that comes after, you can kind of like piggyback on that. So it's like the foundation that is always hopefully relevant and then everything comes on top of it. But like you said, it is a lot of work, especially like the from Scratch part and doing it well and testing. And there are so many experiments also to run. And so each of the two books it took me a year, but also hard work, like not just like a few hours each, it was like really a lot of work. So I do think I need a bit of a break after that. So just to recover a bit, you know, like it was, it is a big undertaking in a sense. But I did have this one figure in, I think the components of coding agents, blockbos I had where I just try to explain to people how things connect. And then that is why I, that's why I was saying, oh, this should be obvious to me because I saw, okay, this is actually three books here. Like this is like the progression. It's like you have the LLM, that is the engine of a car, for example, you know, like the, but just the engine. And then you have the reasoning model book, which is like the beefed up engine. It's like Formula One Grace Car engine. I mean, it's, the analogy is basically LLM, reasoning elements are more capable LLM. It's still an LLM. But then where do you put this engine, right? And so that would be the agent harness essentially. So I think that that would be the most natural progression, like the next book, probably something with agents. But yeah, I had that block post a few months ago where I got super excited about the components of coding agents. I started writing my own harnesses. But while there are so many harnesses out there that yeah, so this article or my harness is more like a proof of concept, what are things to consider? But of course, it's not something you use. It's not Hermes or Hermes, and it's also not OpenClaw. It is a very simple harness. It's for education purposes. And I think that might be a nice potential topic. Although people ask me about multimodal elements too. And I do have actually, I have too many things, and not enough time. So there are other interesting directions too. But first of all, I think re-energizing because writing a book is fun, but it was an intense year. I can only imagine. And look, it wasn't quite obvious to me, but in my head I was hoping that would be the answer, as well. Because with everything you've done with these two books, it is very much, especially for what's important to people at the moment, very much the next super important thing. And it allows you to build the harness as well, which is actually an incredible amount of fun. And you used the analogy to a car. I just want to share, this is something I use every now and then, which is, I don't know if you remember the teenage. Ninja to us. Oh, I grew up in a sea. Yeah, so Cran is the brain. The LLM tokens in tokens out, but he needs a body in order to be the master criminal. And that's the harness there. So you heard it here, first people. It seems like at some point Sebastian may entertain the idea of writing a build coding agent from scratch. And let us know in the comments if this is something you'd like to see as well. And he's time to wrap up. I'm wondering, so I've linked to your substack. Everyone, I mean, if people here don't subscribe, I presume most people do. Check it out. Always wonderful enlightening posts and material there. Sebastian's on Twitter at re and are there any other places? People can connect people home for years. Those are the main. Good question. I think that's the main-- Yeah, that's where I usually like to hang out. That's the happy places. Like the AI community, where I get excited about it, that's taking me K3. That's where I would read about it and chime in basically. I'm always thinking about it. Because I think each of them also has a different kind of perspective, which is also always interesting. Like Twitter is very fast-paced. I always get really good questions from people. Like, oh, yeah, I should explain it or I should mention that. And substack is, in my case, people who kind of already are very deep into things. I mean, of course, with the LinkedIn too, but it's like more-- we're both in that sense. It's like there you-- so it's almost like from quick to more detail to very detailed. I don't know if that makes sense. But yeah. It makes perfect sense. And I do-- I have one final question for you. We've covered so much ground. But if you were to take-- let's say, giving builders people working with elements and agents one piece of it. Yeah, I think I try to do this to myself in a sense. It's like a mantra almost like where-- yeah, it's not running away. It's like it's kind of-- it's really cheesy, but it's a marathon, not a race, or not a sprint. Where, to be honest with you, there are so many always new things coming out. And I think there's no rush. I mean, for some people there is, because it's like a work deadline. But I do think-- or if you're a frontier lab, of course, you have to release the next model. But just for keeping up, I do think it's more like a long term thing, because yeah, you may even skip certain things architecture. You don't have to learn about each architecture, because some are more important than others. I think what just makes sense is to kind of like the building blocks, focusing on those, and trying not to get overbumped is I think just to make it healthy and long-term, successful, not overdo it, in a sense, yeah. I couldn't ever even more. And I do compare it. So I ocean swim a lot, even through winter at the moment. And I compare it to-- the waves can be pretty wild, right? And sometimes, in terms of not being overwhelmed, sometimes you just got to let it wash over you as well, and do what you can, because it just keeps on coming. I do want to point out though. I totally agree with you. And I try to practice that philosophy myself. But can you, K3, release its wakes today? And you've already published-- Yeah, you've got the-- So I was saying I'm not always-- this is 100% where I felt like, ideally, this would be something I would sit down and do in a few days. But I wanted to do this for the podcast in a sense. Like, oh, let's just have something to talk about. But here, see, it wasn't as hard because there was already a Kimea linear. And it's like things built on each other. And so if you follow the more like the building block things, yeah, it goes a long way. You're not starting from the ground up, for example. So for example, in that case, there was a lot of cool new stuff, but it is already all very familiar, if that makes sense. Like, it's like the only thing it was put together differently. But of course, there is a lot of detail in the report. It's, I think, 50 pages of formulas and everything where, yeah, that can be overwhelming. But if you get the big picture itself, it's not overwhelming, if that makes sense. Like, it's building on other things that you have already seen. For example, the multi-atlite and attention might be overwhelming, but it is essentially from deep-sick version 3, which is long time ago. And then in that case, it's not that overwhelming anymore because it's all-- this is a component from back then. And I love that this brings us kind of full circle to what we talked about earlier. If you were to-- if you really want to understand, Kimi K3, of course, you can go and look at the most recent techniques that have been used. But if you want to go back and start with GPT2, and then go through Sebastian's architectural diagrams and everything he's written to build these things out completely. And of course, that involves checking out, build a reasoning model from scratch as well. So congratulations once again, Sebastian. Oh, thank you. Yeah, it was a lot of fun. Yeah, it was like one of the big projects. So I thank you for-- I hope you like it also. Oh, absolutely. And I'm still processing that this is done. I mean, I just got the copies like one and a half weeks back. And I'm still like, well, I actually did the same. What a beautiful moment. Awesome. Well, I'd just like to thank everyone for joining all the great questions. We've had over 100 people join throughout and many more will watch it after. So I appreciate everyone joining in real time and always welcome conversations. Sebastian. So thank you for your time and expertise, Sebastian. Yeah, and thank you again for having me. Usually the thing goes, all good things go in three, but we should probably do it with five snow, right? I think so, right? So yeah. Beautiful. And everyone, if you're still watching, please share it with a friend as well. That's the best way you can support the podcast. And we'll see you in the next episode. And say nice. Thanks for tuning in everybody. And thanks for sticking around until the end of the episode. I would honestly love to hear from you about what resonates with you in the show, what doesn't, and anybody you'd like to hear me speak with, along with topics you'd like to hear more about. The best way to let me know is currently on LinkedIn. And I'll put my profile in the show notes. Thanks once again, and see you in the next episode. (upbeat music) (upbeat music)

Podcast Summary

Key Points:

  1. Local, open-weight models are essential as alternatives to proprietary models, protecting against high costs, restrictions, or unavailability, and fostering competition.
  2. An open letter from 25 companies, including Nvidia, Meta, and Microsoft, argues that American AI leadership requires downloadable, adaptable model weights.
  3. Chinese labs are releasing capable open-weight models (e.g., Kimi K3, DeepSeek, GLM), raising questions about whether American leadership depends on closed models.
  4. DeepSeek’s V4 Flash shows promise for agentic coding at lower costs, with early reports of strong debugging and context handling.
  5. Sebastian Raschka explains that modern transformers (e.g., Kimi K3) are increasingly complex, with innovations like latent attention, latent MoE, and residual attention, building on simpler foundations like GPT-
  6. He recommends learning transformers chronologically, starting with GPT-2 and adding components one at a time, as no single person can grasp full architectures at once.
  7. Open-weight models are being integrated into harnesses (e.g., Kuan Code, Claude Codex), and the same model can behave differently across them, highlighting the importance of harness choice.
  8. The conversation promotes practical learning, open source, and community, with resources like Raschka’s books and a gallery of architecture diagrams.

Summary:

The discussion centers on the growing importance of local, open-weight models in AI, as argued by AI researcher Sebastian Raschka. He emphasizes that these models serve as critical alternatives to proprietary ones, protecting against price hikes, restrictions, or unavailability, and promoting competition. This view is echoed by an open letter from major companies like Nvidia, Meta, and Microsoft, which links American AI leadership to downloadable, adaptable weights.

Meanwhile, Chinese labs are releasing increasingly capable open models, such as Kimi K3 and DeepSeek V4 Flash, which perform well in tasks like agentic coding at lower costs, challenging the notion that leadership depends solely on closed frontier models. Raschka highlights the rapid evolution of transformer architectures, noting that modern models like Kimi K3 are far more complex than earlier ones, incorporating innovations like latent attention, latent mixture-of-experts, and residual attention. He advises learning these systems chronologically, starting with simple models like GPT-2 and adding components step-by-step, as understanding full architectures alone is overwhelming.

The conversation also explores how the same open-weight model behaves differently across various agent harnesses, underscoring the importance of selecting the right tool. Ultimately, the episode celebrates practical learning, open source, and community engagement, with resources like Raschka’s books and architecture diagrams, while promoting workshops and collaboration to navigate the fast-paced AI landscape.

FAQs

Local models provide an alternative if proprietary models become too expensive, restrictive, or unavailable, ensuring competition and preventing lock-in.

It argued that American AI leadership requires downloadable, adaptable weights that people can run themselves, tied to competition and lower lock-in.

Kimi K3 is very similar to Kimi Linear but larger, with 93 layers instead of 27, and includes latent Mixture of Experts (MoE).

Latent MoE is a compression technique, similar to multi-head latent attention, where smaller latent experts are used to reduce computational costs.

Start with a simple model like GPT-2, then learn components like attention and MoE one at a time, gradually building up to complex architectures.

Residual attention adds a residual connection from previous blocks to the softmax and attention computation, improving modeling quality at a slight training and inference cost.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.