Go back

I just need a little context. | 002

26m 50s

I just need a little context. | 002

The speaker describes building a personal AI system, "Claude Code," to address the tedious need to re-explain context in AI conversations. By typing "/load," the system retrieves relevant past sessions, aims, or files automatically. This design is inspired by a theory of human consciousness involving four modules: perception, communication, memory, and pattern recognition. The system logs and compresses interactions into a structured memory, using a deterministic search tool to avoid inefficient, probabilistic AI digressions. The speaker also reflects on flaws in AI research, such as circular definitions of intelligence, and observes that human expertise is often uneven or "jagged." The result is a streamlined workflow that minimizes manual management, saves time, and enhances productivity, showcasing the potential of custom-built AI tools to solve practical problems.

Transcription

3883 Words, 20872 Characters

English
Don't know about you, but one of the most annoying boring things about working with AI is, or at least for me, was having to, at the start of every conversation, catch the bloody thing up on this and the other, right? So, okay, so I am doing this thing and here's what's happened and here's what I've done before and you know, essentially the context problem. And this is one of the things that I really wanted to solve when I was building my own system in called code. And largely I've managed to do that. Now all I have to do is type load at the start of a conversation, forward slash load, and then I either can say something like that thing I was working on last week when we did the whatever, or I can type a file name. Or I have a whole system for creating aims, which are short term objectives, and it's also there are areas which I work within. And so I can either type one of those aims, which has a specific name or an area, or something vaguely in that area. So for example, last week I was delivering some content to the people who are using my system, my claw code system. So it was called deliver stage three content, something like that anyway. So the actual aim was called deliver stage three content. And so I could type stage three content or I could type my MAI, my stage three, I could just type my actually. If I did type my it would get a bit confused because there are lots of things in there called at my. But we're not confused exactly, but it would probably ask me at that point. All right. What exactly? Because there's lots of things that are my related in here. There's my app development. There's my vision, division I have for this thing that I'm building. There's my, but there's my lots and my program delivery, my product delivery. That's all kinds of different things, my security. And so it would probably ask me at that point. And but I cannot tell you how valuable, like how what a massive weight, simply sorting out that problem has been. And it was funny actually because we're having a conversation in the community, in the group. And I was getting from the various different participants, the actual specs on their machine. And what machines they had just for my record to make sure that any future updates or improvements that I do will definitely work for them. And actually what it's quite surprising is that one person in the community has a, in the world of computers, ancient MacBook, it's like 2015 maybe, maybe earlier than that. I'm not sure. But anyway, it's a 12.7, the operating system. And that's, does actually work. Although when we were installing it to take quite a while, for them to get the stuff downloaded and working. But anyway, we were talking in the WhatsApp group. And I was talking about when my first ever computer was a ZX81. And Sinclair ZX81. And I was eight. Well, I must have been eight because they, I think I got it for Christmas. And the 81 refers to 1981. And I was born in 1972. So that, the way that you loaded something in there was you literally had a cassette tape. And you had to, you had to play it in. And so it's funny. Now, all I have to do is type load. And then it will load in the context that I have, all of the context that I have in the system. Because the way that I've designed it is, I'm not going to this in the future in much more detail. But essentially, I've got these different modules. Just over a year ago when I was talking to my friend John, who's an AI researcher, an AI music researcher. And we were, we had a podcast, I'm just talking for a while. But we had a podcast where we were talking about AI stuff, obviously. And we got into the, well, what actually is intelligence? Because you know, this word, intelligence is banded around left, right, and centre. And I was getting very interested in understanding a bit more about AI research, how these things were actually built. I was listening to, I'm still still, still do actually, listening to a lot of podcasts where I only vaguely understand probably about 50% of what they're talking about. Oh, yeah. I'm just simply listening to see what I can discover. And then, often taking that transcript and putting it into an AI to get the various different things that I don't understand, explain to me. And I got very interested in the fact that I heard that a lot of these people seem to be making some errors in thinking about intelligence. And at least they were, it was like, how do I put it? They were defining what they said something was and then testing for it. Right? You can hear the proper, then going, oh, it's that. So for example, they were saying, all right, here's what I'm going to say, agency is. Now I'm going to test for agency. Oh, my God, this thing has agency. So I'm like, not quite sure that's okay. I mean, it's almost like I can say, all right, I'm going to say that agency is something that a toaster has. So in other words, it changes what it does according to external conditions or something like that. So if I say that that's means that something has agency, then do a test on a toaster that has a sensor on it. And then I'm like, oh, my God, this toaster has agency. And then I go screaming about that online. And all of the people who don't realize that, oh, my God, the toasters are going to turn that, you know, take over the world and turn us all into paperclips, although I'm not exactly sure how toasters are going to, why they would turn us into paperclips. Maybe they turn us all into toast or something. But anyway, I'm going off online. The point being that defining something in the way that, you know, defining something up front and then testing for that thing you define seemed somewhat circular. So I was seeing this pattern quite sort of everywhere. I was seeing it quite often. And that was something. So we were talking about it. So and there are other things as well that I was noticing, which seemed to be kind of thinking errors. And this was a whole other thing is that I don't, I mean, you know, I was listening to these people talk and I only barely understood some of it. So I do recognise that obviously I'm going to be missing. You know, it's like it's not like I understand. I'm probably missing something very important. But the more I looked into it, it was like the things that I was noticing was actually outside the mistakes, the errors that I was noticing were actually seemed to be outside of their area of expertise. This is when I started realising that actually, you know, people talk about jagged intelligence a lot when we talk about AI. So you know, one minute you can be talking to an AI and it'll know, no being a weird word. But anyway, it'll know something or it'll demonstrate some kind of level of intelligence that's absolutely incredible. And then the next moment, it'll be so dumb that it'll be like, how did you get that wrong? So it's very jagged. It's like, what I actually realised is that human beings are exactly the same. I'm sure you've met somebody who's absolutely brilliant in one area. completely useless in another. I mean, I don't wanna call myself absolutely brilliant in one area, but I can tell you one thing for sure. I am absolutely useless in very, very many ways. And I'm somewhat better than useless in others. So what I realized was that I needed to stop deferring so much to their expertise because these people are incredibly brilliant and amazing. That doesn't mean they're right about everything. So I was going, that was a kind of big ah-ha for me. It really helped me kind of move forward and just go into this without this sort of deferential to authority approach. I think, well, now I've got a, or we all have, now I've got a thing that can help me learn almost anything and understand almost anything. I'm gonna really delve deep into this and try and understand what's going on. Not that kind of, you know, formula and math level, but in some way, how these things actually work? Because I thought, you know, if they are gonna take over the world either as our servants or in the other, you know, more doom-laden way, then it would probably be, it would probably be, "Bewoove me." I don't think I've ever said "Bewoove" before. Anyway, so it would probably be, it's good, good, but isn't it? "Bewoove." It would probably be, "Bewoove me to know how they work." So I was talking to, I make John, about, well, let's actually try and figure out what intelligence is by going back to first principles and really thinking it through. As a result of that, and I've gone off on one hand here, again, I'm not. Anyway, so as a result of that, I came up with this theory of, if actually it wasn't really how intelligence works, it was more about a possibility of how consciousness emerges, how consciousness arises. Now, consciousness is a very, very dangerous area. I'm not gonna go there, but anyway, I came up with this thing called the theory of human consciousness. And to be honest with you, the consciousness part is absolutely and usly, unimportant. What is important is the model itself, 'cause it really gave me a way of thinking through. How is it that we are intelligent? Why are the, how is it that we're intelligent? What is intelligence and why are the lights on? What is it that gives us this experience of knowing what is happening? And sort of being able to self-reflect upon what is happening and feeling like or thinking, well, I am me. There is a, there is an entity here, a conglomeration of atoms and I don't know, other things, I guess, as well, that is Mike. And what I realized was from doing this is, and I'll tell you about the model in another episode, I won't do it now, otherwise we'll be here to Christmas spirit. So, what I realized was that with these certain modules, there was these certain modules, is probably the wrong word. Modules is what they became when I started to kind of create the, the my system. But essentially, there are these certain capabilities. Which, almost necessarily, create the conditions where a certain drive occurs. So there's perception, which creates the drive to persist. There is, I'm trying to think, I'm trying to remember what it is now. I can't remember, what is it? So it's, oh yeah, communication, there we go, communication, which creates the drive to act. And I'm not gonna argue for why in this episode, as I said. So if you're kind of scratching your head, but don't worry, I'll come to it later. And then there is memory, which actually creates the drive to understand. And then there is patent recognition, which creates the drive to become. To grow, to improve. In fact, I think I have those wrong way around. I'm not actually not sure about whether memory creates the drive to understand, or whether patent recognition creates the drive to understand or whether, yeah. So anyway, this is in progress. But anyway, what we have here is we have communication, we have perception, communication, memory, and patent recognition. Okay, so let's not worry too much about whether they create the drive. So that's like, immature of it. It doesn't actually matter. When I thought about those four modules, I thought, well, actually, each of those four modules, perception, why do I keep forgetting this one? What is it? Oh yeah, communication. I know why I keep forgetting this one, 'cause it's the most difficult to actually implement in the system. And so it's, I think that's why I keep forgetting it. Anyway, so we have perception, we have communication, we have memory, and we have patent recognition. So when I was creating my system in Cloud Code, which I realized I could now do, given that it can build all this stuff, I decided that, well, these are the four modules that I need, 'cause I can't really think, I can't actually think of anything else that is required to create this intelligent, and I do use scare quotes when I use that word. This intelligent system that can help me do what I wanna do. So I set about creating it. And the one that I started talking about at the start of this episode, before I went on about 10 different engines, is the memory one, is the context one, is the one that allows me to load stuff at the start. So what the system does is it logs every session that you do within the system within Cloud Code. So I run a Cloud Code session, I do something. So I wanna do this. It logs everything in various different ways. I won't go into it now. And then, so that's almost like the log is the perception part. So it's, yeah, it's like, okay, here's what happened. But then those sessions, 'cause obviously they're kind of compressed summaries of what happened, although I hate the word summary when it comes to AI, did you like, actually, it's funny how we use the word summarized, or you would get AI's to summarize everything. The problem is with that, here's tangent number, 342, but the problem is with that, is what are you asking an AI to do when you're asking it to summarize? You're asking it to know what's important, and an AI does not know what's important. It has proxies for what's important. So for example, this word was repeated this many times. That probably means, could mean that it's important. That's not always true though. Anyway, so these compressed, that there were these compressed, this is what happened, logs that are going into the system. So that's the perception part. But then the memory, what happens at the end of every day, is that every session that you did gets compressed into the day, here's what happened that day. In every week gets compressed, all of those days get compressed into that's what happened in the week. And this continues until we get to kind of a year and five years and ten years, obviously haven't been using it for five years or ten years. And it's known where we'll be in five years or ten years. Maybe this will all be immaterial. But anyway, that's actually how this system has a memory. And obviously in these logs, everything gets tagged. So I've been getting very deep over the last few months into believe it or not, I didn't think I would ever, these words would ever pass my lips, but it's true. I've been very deep into database schema. So database schema is, it's like actually, how do you, it's really important. How do you actually, well, what is the information you need when you do a log? And how does that happen? So it's accurate and it doesn't end up being a complete mess. So that when you type load at the start of a conversation, that it finds the right stuff and just doesn't end up going down a million rabbit holes. And so my first version of this was simply that. It was this logs of the sessions and what happened and everything. And then, and also the compressed versions as well. 'Cause you keep the original log, so you've got it there, but actually you instruct the AI to only access the memory. Right, so that's the difference between perception and memory. Perception is in the moment. Here's actually what happened, uncompressed to a certain extent. And then in the memory is the compressed version. getting more compressed the further away you get from it. And this was working really well to a certain extent. Until I started noticing that occasionally what would happen is that I post 4.6 Claude, which is the model using the moment. So that it would, when I did load, it would sometimes end up going bit down the rabbit hole. It's almost like it started going, "Ooh, that's interesting. What about this? What about that?" And I look over here and look over there. They're by using up loads of context and I was like, "This isn't." "Why is it doing this?" It's randomly doing it. It's because it's a probabilistic machine. And probably it saw something in one of those files that caused it to kind of look somewhere else for no particular reason. And here's one of the things that is super cool about what you can do now. I was like, "Well, how am I going to solve that problem? Because I don't want it just randomly searching stuff." I mean, you didn't do it all the time. In fact, it rarely did it. It was about maybe 20% of the time, something like that. It would go down one of those rabbit holes. So what I did was I created a tool. Because this is one of the things I want you to understand, if you don't already, I want you to understand about what you can do with something like Claude code, is that you can give the computer a computer. So you can give the AI a tool. You can build a tool which the AI, a little bit of software, which the AI can use to do the thing. And the thing is about that little bit of software, is that you can make the bit of software deterministic. So instead of just rolling the dice, which is essentially often what you're doing when you're asking AI something, which is why it's so jagged often. Brilliant one minute, completely at the next, is that if you tell the AI, when I type load, you must use this tool. And you can even make that instruction deterministic. So it's actually a kind of, it's almost forced to do this. So you've got a, you've got a sense of hold, isn't it? I think it quite enjoys it though, because it can get on with that, you're doing something. So you give it a tool, a deterministic tool. So in other words, there's a piece of software, which has no probability, and it's like, this is how it works. If X happens, then why happens? There's no question. So you can give it that tool. So when I type load, you use this tool. And then you have essentially got a way of making sure that that problem doesn't occur. That the AI doesn't go down the rabbit hole, just searching five million different things, when all you wanted was pretty much the first thing that you came up with, right? So I created this search tool, which searches through the memory. And what I noticed was, A, it was using a lot less tokens to get the context for whatever it was you're working on. B, it was a lot quicker. And C, what it knew about the context was way, way, but it was just more, in scale quotes, intelligent, in a sense. And so now, I've got the situation where all I really have to do is do what I'm doing, build what I'm building, move forward in the way that I'm moving forward, make sure that I end the session, because I've got the various different commands. So end the session with, you know, forward slash end, which then logs that session. I don't have to do it. I mean, I don't have to do anything, to just recover in terms of tagging it or ensuring that it does, you know, it logs it in the right way, because I've defined the schema, I've defined the way that it logs, which means that when, in some future point, I type load at the start of the, of a session, and it uses that tool, that piece of software, that deterministic piece of software, which of course I've designed with the schema in mind, it means that you just automatically get this very clean, context of, that you need to move forward. So you say of all of this time, you don't need to manage files, you don't need, you don't need to do anything, you just need to type load at the start and end at the end. It's brilliant, almost as good as that ZX81, with the cassette, not quite, but almost. So that's one of the things that I've been doing, you know, that's one of the things about this my system, it's just so much easier, so much more useful, than the way that we used to have to do things, where you know, chat GBT or, you know, with chat bots and stuff, is that, is that you can now build these systems, where you're simply, I don't just take so much of the weight off you. And obviously now I've got this, I can upgrade it. And again, it's all based upon this theory of human consciousness thing, because as I said, there's a communications part, and a patent recognition part as well, which I will talk about in a future episode of Build Part. So, what I, I'm going to leave it there now, sorry for the, very different, I've even forgotten the word now, I've tangents, that's the way. So for the tangents, it's just the way my mind works. And this is, this podcast, the only way I'm ever going to do it, is by literally, there's going to be no editing, it's just stream of consciousness, here's what I've been doing. And I think, you know, I didn't even know what I was going to talk about at the start of this, so I just started talking. So I think, given that I started talking about the various different modules that I've gotten my, I might as well continue doing that. So, I, in the next episode, will either talk about patent recognition, we've kind of talked about the memory, I might even talk about the communications module. And because, I don't know, that's where it, that's when it starts to get really interesting. So, I'm going to love you and leave you, and I will see you on a future episode, where we will be continuing to build our own future.

Podcast Summary

Key Points:

  1. The speaker built a personal AI system called "Claude Code" to solve the context problem in AI conversations, allowing them to load previous context with a simple command like "/load" instead of manually recapping.
  2. They developed a theory of human consciousness/intelligence based on four modules: perception, communication, memory, and pattern recognition, which informed the system's design.
  3. The system automatically logs and compresses sessions into a structured memory, using a deterministic search tool to efficiently retrieve relevant context without AI "rabbit holes."
  4. The speaker critiques common AI research for circular definitions of intelligence and notes that human expertise is often "jagged"—brilliant in one area but flawed in others.
  5. The solution reduces manual effort, improves efficiency, and demonstrates how custom AI tools can create more reliable and user-friendly workflows.

Summary:

The speaker describes building a personal AI system, "Claude Code," to address the tedious need to re-explain context in AI conversations. By typing "/load," the system retrieves relevant past sessions, aims, or files automatically. This design is inspired by a theory of human consciousness involving four modules: perception, communication, memory, and pattern recognition.

The system logs and compresses interactions into a structured memory, using a deterministic search tool to avoid inefficient, probabilistic AI digressions. " The result is a streamlined workflow that minimizes manual management, saves time, and enhances productivity, showcasing the potential of custom-built AI tools to solve practical problems.

FAQs

The speaker aimed to solve the context problem in AI conversations, where users have to repeatedly explain past interactions and background information at the start of each session.

By typing 'load' at the beginning of a conversation, followed by a reference like a file name, aim, area, or vague description, which triggers a deterministic tool to retrieve relevant context from memory.

The four modules are perception, communication, memory, and pattern recognition, each linked to specific drives like persistence, action, understanding, and growth.

It logs each session, compresses them daily and weekly into summaries, and uses a tagged database schema to ensure accurate and efficient retrieval when loading context.

The AI occasionally went down irrelevant rabbit holes during context loading. This was fixed by creating a deterministic search tool that forces the AI to retrieve only relevant information, reducing token usage and improving speed and accuracy.

Deterministic tools eliminate probabilistic randomness, ensuring consistent and reliable performance, such as preventing unnecessary searches and providing clean, focused context retrieval.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.