Hey folks, just a quick note that I continue to take episodes out of the public podcast feed and they are being hosted only behind the paywall on Substack. The current count is about 34 episodes in the public feed and 16 episodes behind the paywall. So, if you want access to these archive shows and would be good enough to support what I'm doing, I'd really appreciate it if you become a paying Substack member. I'll put a link in the show notes so you can find it easily. Anyway, hope you enjoyed this episode. I have had multiple instances where I will say, "Hey, here's an error in the logs dig into this, try to solve it." And the solution is to stop logging the error. I've had a quad do that like three times now where I'm just like, "No, that's clearly not what I want to do." I've had an agent running today that I was watching it and the little task that it did was switching environment to production because the demo environment that we were hitting didn't have some piece of data so it figured the right thing was to just start hitting the production environment, which would mean any trades that happened were real trades. So, you know, but I'm on top of that and wouldn't let that happen. I think that's, yeah, if I just kind of walked away and let the agent run overnight, it's scary to think what I might wake up to. You're listening to "Risk of Ruin." I'm John Reader. This is "Invanish Programming." Part 1. In the way of full disclosure, I have a few episodes queued that get into AI. So I think a reasonable place to start is with the possibility that before ever hearing these shows, you are already sick of AI, which I understand. Actually, let me first describe how I feel when I hear the term. Okay, my primal reaction is kind of a revulsion because AI seems like a runaway train that could have some pretty dramatic societal effects, which I am not at all in control of, and you know the sense that something is beyond your control always feels bad. And then there's the fact that the more someone talks about AI, the more they sound like they're kind of full of shit, which I would not even actually propose that everyone that talks about AI is full of shit. They just kind of sound like bullshiters. If I had to guess why that seems to be the case, I think it is the intersection of a high uncertainty domain mixed with all kinds of weird incentives. For instance, there's a group that I will call AI influencers and general FOMO peddlers. These are people who are getting rich, letting AI agents run while they sleep so that their time is freed up to constantly remind you of this fact. While I was making this episode, I was watching a YouTube video about using an open-source AI tool, and the guy literally claimed that people were losing money and getting left behind because they didn't understand this agent's setup. All right, and then there are the AI frontier lab CEOs who have an undeniable incentive to hype their own companies and drive up demand for their stock. This is very logical. They need to lower their cost of capital so that they can have some chance to actually deliver the dreams they're selling. As in, the hype cycle is baked into the cake. It cannot be separated. Then there are CEOs of companies that are not AI labs and faced with their stock price being in the toilet. They start talking a lot about AI and how they're pivoting to more AI and they're replacing human engineers with Claude. They might be telling the truth, but because their goal is really not to say anything true about AI, right? All they care about is getting the stock price back up so they don't get fired. Anyway, for this reason, it's hard to know how much to trust what they say. So these groups are in the AI booster camp, but they're not the only ones creating the kind of noise that makes it horrible to follow the conversation. They're also the AI deniers who really might be the most weird and interesting group if you start to think about what they're saying. They have the same orientation toward artificial intelligence that I have toward ghosts. First, I do not believe in ghosts. Second, I am afraid of them. Especially a ghost child. I mean, just because they don't exist doesn't mean they're not scary. Okay, but I know this is not logical, but there are people who both believe that AI is useless and also that we must be diligent in resisting the temptation to use it, which I should have some sympathy for due to my nonsensical concurrent ghost denialism slash crippling fear. But no, I don't have sympathy. I think these people sound silly. And if something is really useless, there is no need to convince others of this idea. Eventually, reality will just settle the issue. Anyway, having said that, I think this show offers an opportunity to cut through some of the noise I've described precisely because we're gambling focused gamblers do not get paid to create attention or debate or engage in hysterics. They don't get paid to talk or punch a clock. They get paid to win for that reason. They tend to be pretty results focused. So if we go find some gamblers and ask them how they're using AI, then the response we get will not be the result of a company directive to increase token usage or any other bizarre incentive. If these gamblers don't find anything useful about the output, the AI generates, they will just drop it like a hot rock. I have a few parts planned for the series, but the first guess we're going to hear from is Cartwright. He isn't all around gambler, but he's probably best known because he was on gambling with an edge a number of years ago talking about hammering a blackjack side bet. Here's Cartwright. I'll go way back. So when I was eight years old, my mom asked me what I wanted for my birthday and I told her I wanted a computer. And throughout elementary school, I became somewhat of an expert at Apple Soft Basic. And I did the delivery. Remember my mom coming down to the office asking me if I was ready to go to school one day. And I said, "Well, you're talking about I haven't gone to sleep yet." So I was pulling all nighters in elementary school. I found software development very early on and I knew pretty much from the beginning that that's what I was into. And I knew that was my career. So yeah, I went to school. I went to the University of Illinois College of Engineering, major computer science. And from there, I went to a consulting role. I was in a consulting company doing a lot of different projects with a lot of different technologies. And eventually went off on my own, my own consulting firm and ended up landing in the telecom industry. But yeah, so I was in the text-based writing software. Eventually, that evolved to being a manager, which is part of what caused me to go full time with the gambling stuff because managing a team of developers isn't quite as satisfying as building my own things. It's worth noting that the history of advanced play is parallel to and in some ways cannot be separated from the history of computing. Even though the advent of AI has proven just how incredible our human brains are at processing, there are some tasks that the human brain is just not well equipped to handle. For instance, if you just attempt to intuit Blackjack basic strategy, you are going to come up short. And so the very first basic strategy was calculated by four Army mathematicians with a little help from desktop adding machines that was in the 1950s. Then in the 60s, Ed Thorpe had access to a much more powerful computer at MIT, which he used to prove that keeping track of which cards have been dealt could flip the game from house edge to player edge. That computer filled an entire room and gave birth to the first card counting system. In the late 70s and early 80s, computers got better and smaller and cheaper, and that meant clever young men could build wearable devices that they used to do things like clock roulette wheels. Also in the 80s, advantage minded people saw the potential that faster calculations could have for sports predictions. In Las Vegas, the Billy Walter's group was alternately referred to as the computer group. And in Hong Kong, Bill Benter and Alan Woods were using machines to be horse racing. Okay, then fast forward to the advent of the internet and the spread of online casinos. A cottage industry of self taught IT guys filled rooms full of modems and internet routers and switches. They beat up casinos in far flung places like Malta for every dollar they could possibly pull out of a net teller account. Actually, if you look up the membership of the Blackjack Hall of Fame, which is sort of erroneously named, it's really the advantage player Hall of Fame. It is dominated by people who got their edges with technology. And I don't mean like technology was involved incidentally. I mean, take away the computers and the net worth of the combined membership would drop probably by maybe not 100 percent, but close to it. I think it would be a mistake to view that dynamic as some artifact of history or a set of opportunities that were limited to the newness of the technology. It's still happening. And so another important part of the show is basically the question of if changes and technology had this big effect on advantage play historically, should we expect something similar from AI? So yeah, I think I was a gambler first, but always someone looking for an edge. So the pool hustling led to poker playing, led to card counting in Blackjack. And then that, we're really, I think the pool hustling led to card counting, Blackjack led to poker was actually the trajectory. And then I had this almost obsession with trying to understand what other people were doing at casinos. Because I knew that it couldn't just be Blackjack card counting. But at the time there wasn't.
a lot of public information about things like whole carding and stuff. So I became pretty obsessed with trying to learn about that. And it was my hobby still at the time. And that somehow led me to running into new games that were like new installations that could see us. And that became a little bit, that was like my specialty. So I would find a new game and use my software skills to analyze it and figure out ways to beat it. So I think my graduation from like basic card counting to something more advanced was directly tied to the software development. I never really had a whole card team or anything that I did a lot of hours with. I said that you cannot separate the history of advanced play from the history of computing. But I think it's important to draw a circle around the role that the actual gamblers had. Because in the whole thing, yes, the compute was important, but it required agency. And it required someone saying we are setting out to violate conventional wisdom. I.e., we are looking for places where the maximum the house always wins is not right. Well, so what I did and what I really enjoyed doing was building libraries that were adaptable to a lot of games. Like I've got a really, really fast poker library. I've got a really, really fast, fast, blackjack library. My Bokarat, Bokarat is really easy to analyze, but my Bokarat code is actually really nice to work with because I can add, like if I find a new side, but I can add that, analyze it, have numbers, have account, everything. It doesn't take very long. 30 minutes or less. And so that's where I kind of had a lot of fun for me was trying to make the code efficient and reusable. So building actual libraries and building a software architecture that let me added new games fairly quickly. But it was a very labor-intensive to get to that point for sure. And I would, and it's a stress to say I'm even at that point. I mean, it's still a lot of hacking, but once you know what you're doing and you have the foundation, usually analyzing a new game isn't that difficult, but often the number of combinations that amount of compute that it takes is pretty overwhelming. To give you a small sample of the list of games, I'm just going to read from the table of contents for the book, The Ultimate Report by James Grosjean, which lists basic strategy for various carnival games. Okay, three card poker, Kirby and Stud, let it ride, crazy four poker, four card poker, Kali Lowball, Flop Poker, Mississippi Stud, Jackpot Holdem, Heads Up Holdem, Ultimate Texas Holdem, and Chris Cross Poker. And if you want to add to that list, you can just walk through the table games pit in any casino. Pretty much every game has a variety of available side bets. For instance, Grosjean lists four separate pay tables. You might encounter just for three card poker's pairs plus side bet. All strategies are downstream of payoffs and probabilities. And so you could see how it might be very helpful to have one reusable code and two. A lot of experience going through games to see where they're sensitive to changing conditions. For sure, I feel like I have a good intuition for when a game is vulnerable. To the point that I often just don't bother looking at games where I'm fairly confident there's not going to be a vulnerability there. And so, yeah, I would say that it's experienced and knowing where the edges come from, you can usually tell if there's a game that has the chance of being vulnerable. I really do like my Bokaratt stuff that I mentioned before. I think that the Bokaratt code, it's nice because I have gone to a casino, found a game, gone out to my car in the parking lot, and analyzed it, and gone back in and played it in a very short amount of time. So that's very useful. And one of my favorite plays is finding a game, playing it, doing pretty well, having the game get pulled, going to G2E, finding the manufacturer of the game, asking them where else it is, and getting information on where there's another install, and having coincidentally met someone in that region who recognized me from gambling with an edge. And then having him go find it, and he found it and played it for years before it was finally pulled. That was probably the longest lasting play I've ever had. And yeah, that was an intersection of being on gambling with an edge, so my name was known. Randomly having met this guy in a casino, having built a software, having the first install of the game be fairly close to me, there was a lot of things that had to go right, but yeah, there was a lot of fun. It's worth remembering that the Suffolk right plays has already gone through at least some safety checks. The game maker was trying to invent a profitable game, which would mean you know, hopefully robust to advantage play. They engaged a mathematician to double check their work. They also have a cushion for mistakes, since if you ask the average player of a carnival game, what the house edge is, they would look at you like you have an arm growing out of your head. The point is that this all requires fidelity in the measurements. The right is looking for the last little bit of failed diligence. Maybe it's a game where the manufacturer assumes players will always have to bet more on the main bet than the side bet. And so they don't have to worry that APs would come up with a strategy that's specific to the side bet. But then there's some miscommunication and the casino doesn't get the memo that the house edge is conditional on that requirement. Usually doesn't take very long. What will happen is a teammate will find a game that they think might be interesting. Tell me about it and then that intuition we discussed before. So I'll use my intuition to kind of decide if I think it's worth pursuing. Usually I will want to go see it myself so I can really get a sense of all the dealing procedures and all the different variables that might impact ways to get an advantage. And then depending on the game and the technique that we're looking to use, the turnaround time can be less than a day or it could be you know, a few days of work. It is typically doesn't take very long where it starts to get what takes longer is once we have something that we think is actionable. We'll start playing it, but we usually are going to start very slow and observe and get a good inventory of dealers and get a sense of how much action the casino is willing to take. How much play does the game get if we're playing a new game. We've had instances where we've won so fast that the casino just closes the game because they have no idea what's going on, but they know they're not making money at the game because it's just not getting played by other people. One of the core questions I'm interested in for this episode is for a programmer, advanced player whose work requires fidelity can the AI tools actually deliver usable code. And then if they can, does that change the supply of people who can go after these edges? Also how many people can even do this stuff to begin with? My sense is that it is quite rare. I think obviously James Groce-Schean is the godfather of this stuff. He's been around forever and has been doing it. And then I know of a handful of other people that have analyzed some games or do some analysis. So yeah, there's a handful of people. I think it's a pretty rare thing. I do think that the AI tools are lowering the barrier to entry. I'm sure that if somebody found a basic game that isn't overly complicated, they could talk clawed through analyzing it and get some results that are actionable fairly quickly and with decent accuracy. I think that there are some games that are going to be super complicated. So if you don't have the right kind of libraries, you're going to run into an issue of, it's going to be the amount of time you're going to spend. It's going to be a lot of just guessing and going back and forth and back and forth versus being able to really come up with the right answer fairly quickly. In the table game AP world, I don't think AI is really moving the needle for a while. I think that sure it can help you analyze games, but people can analyze games anyway. That's just find the right partner. So I don't know that it's really moving the needle. I think in the sports betting space, it's a lot more likely to have an impact. I think right now with the prediction markets, it has a lot of people are using it. I don't know if they're using it well. The idea having an agent just running 24/7 making prediction market decisions is not even remotely interesting to me. I would be very afraid of having these AI's make decisions that have direct financial consequences. Certainly, the barrier to having a running 24/7 that is market making is much, much lower given the existence of the LLMs. I think that that is something that a lot of people are doing, have done, will continue to do. I think for the most part, at least right now, is because it's early games, early times in the prediction market space, I think that if that's a profitable, you could do that very profitably without a lot of sophistication. Like everything, I think it's going to get harder. I think that if you look at the prediction market space specifically, that is evolving almost as fast as AI. The edges that existed a month ago, a lot of them might not exist now, and the way you have to approach it is potentially quite different, too. He doesn't spend a lot of time on table games anymore. The opportunities change.
changed and life happened. So he's less focused on casinos and more focused on sports. - There's a few different factors. The advent of new table games has slowed. There was a time when it was, it's just there were new games popping up all the time. And that seems to have slowed down a lot of what comes up now in a new game shows up. It's actually a copy of an old game. AGS will just rebrand a shuffle master game. - So the opportunities for that sort of way of playing have certainly slowed down. Sports betting has certainly taken over a lot of my time and the prediction markets that have come online here recently are super interesting. So a plus we also had a kid. So I have a two year old, which makes it a lot less interesting for me to go travel. And being inside a casino is just not very appealing anymore. So it's a combination of all of those things. Like I never loved casinos to begin with. And now, you know, if I can earn from home, that's a much more appealing situation. - There's an interesting thing that happens with technology where its effects can be not obvious even to the experts that spend all of their time with it. I will use the launch of chat GPT to illustrate what I'm talking about. Most of the world came to know the term GPT when open AI put up a website for a chatbot in November of 2022. But they had already teased other uses for their model even before then. They had an image generator, Dali, which was sort of novel, but didn't really capture the world's imagination. And keep in mind that the GPT three models which chat GPT had been built on had actually been around since March of that year. I'm not the first person to point out that if open AI thought chat GPT was going to be a killer consumer product, they might have spent more time on the name. Then consider that today we see these things as being almost synonymous with coding, but it was months before the chat GPT website even had a button that would allow you to copy the bots response to your clipboard. So if they thought that the model's output was useful at all, they either didn't have the engineering resources to deliver a copy/paste button or they really wanted to make you work for it. And then after that, it was basically a couple of years before we had coding tools that didn't require the user to do a bunch of copy and pasting of code from the website to a code editor and then pasting the error back to the website. So this question of what can you know about the future, even if you're paying a lot of attention, is one that captivates me. I asked Cartray, how much had he been paying attention to the launch of the AI tools? - Yeah, so I was paying attention to AI. I think computer vision was an area that I had a lot of interest in. I did have philosophical discussions with a co-worker about AI taking over software development way before I'd ever heard of an LLM and we figured there was certainly a point where that would happen. I think when the image editing stuff came out like Dolly and all these sorts of things, I found it interesting like the concept that AI was going to replace creative jobs before it replaced math-based jobs seemed counterintuitive until I thought about it and the fact that when you're doing a creative thing, there isn't a right answer. It's a spectrum. That certainly makes it, I think, easier to come up with something that can land close enough that it's usable. And I've always said that software development was a creative outlet. I think it's a, it's math-based, but for me, the creativity of it and coming up with different approaches to solving problems is one of the things I like most about it. And I think that's actually a lot of what lends it so well to LLMs because there isn't a right way to build software. There are so a lot of different ways you can do it and different approaches. So it doesn't have to be precise. I think that as far as paying attention to it, ChatGPT came out. I thought it was interesting enough, but I used GitHub had co-pilot, their first version of co-pilot, which was built on ChatGPT. And I used that. And I was certainly very impressed with it. I used it for more code completion. I wasn't letting it build code, but I mean, it was atrocious at that, but it was pretty good at completing code. And then, you know, just kind of watched it grow from there with varying degrees of interest in paying attention. But over the last, I would say, year, I've certainly paid a lot more attention. And I was using cursor when it was fairly early on and helped them debug some stuff and everything. And so I think that, yeah, and from there, the growth has been really, really fast. And it's getting quite good. Sometimes dangerously good. There's a really interesting part of our interaction with these new AI tools, which is the models are basically trained on every word of human knowledge. And yet, they can also be dumb as shit, but they're not dumb. It kind of matters what you ask them to do. And it matters how you manage your own workflow. And it matters how personally offended you get when they gaslight you. But it is also a little weird to have to learn to work with something that is supposed to know so much, right? It's still programming, but there's just a ton of noise that exists between your prompt and the model's response. So OK, people are using AI in a couple of different ways, right? Some of it is just as a way to answer questions, not to write code. And I think the AI can be quite good at that at digging into things. And it's a good conversational. And so from there, I think you probably can get some decent guidance and inspiration on how to approach things like modeling, sports betting, or analyzing a casino game. If you just sat down with Cloud Code and said, write a combinatorial blackjack analyzer that runs in less than a second, I'm pretty sure you're not going to get what you asked for. I think that I don't think it's just quite there yet. But I think that if you have some symbol and some idea of what you want, you can tease that out of AI. And I think that AI can actually help with some of the domain knowledge. I think that I certainly use it. But I don't necessarily think that it's all in the context of writing code. I think that there's two stages to it. I think the planning stage is important. And I think that's where having domain knowledge is a huge plus. I do think the AI can help with it. But I think that the software development side, what you're going to have is if you're building something that is actually trading money like a prediction market bot or something along those lines, you have to be very careful because these AI's are trained to tell you what you want to hear. And it's really difficult to get it to say, oh, I don't know how to do that. It will always come up with something. And one of the biggest frustrations I have found is that these things love to have default values. And let me just-- since this data isn't here, I'll go ahead and just assume it's zero or whatever. And that's one of my biggest pain points is that you have to make sure you're explaining to it if this isn't available, then you should fail. You should just invent something. And so I think if you don't know what you're doing and you just let it in, you just kind of one shot, let Cloud do something, you're going to run into a lot of those kinds of issues. But I think that being said, it's a great resource tool. I think if you've got some domain knowledge and you want to be supplemented, I think-- AI can help with that. I think that there's a lot of valuable information stuff in there that you can pull out of it. But at the same time, I think if you're starting from zero, and you sat down and said, OK, I'm ready to model NFL and ask, flood, how to do it, I think you're going to have a hard time sort of coming up with something that's profitable. I mentioned at the top of the show that I do find that AI discourse to be full of noise. And so one of the benefits of making this show is that I can just go ask people questions and they sometimes answer. Part-rates work requires fidelity. He's an experienced programmer who doesn't need to use these tools. He's not like me. I can barely type and have never written anything beyond spaghetti code in my entire life. But Cartwright says that the impact is dramatic. I certainly wouldn't want to talk about anything specific that I'm working on. But as far as what AI is doing, I think if you are an experienced developer and you know what you're doing, and you actually know how to solve the problems, you can leverage AI to increase your velocity by a lot. I feel like I'm able to develop things truly dramatically faster now than even a few months ago. But I also think that there's a lot of non-developers that now think that they can build things and they're only going to get into trouble. I think if you've got a task that's very simple, and it's just for private use, I think this is a great tool for people. I think that if you're wanting to build something that is like a production enterprise quality piece of software and you have no idea what you're doing when it comes to software development, I think you're going to find yourself in trouble. One of the fun things about working with LLMs is seeing them do the exact same kind of bone-headed stuff that people do, but then realizing that it's just a machine and it's not malicious. So it's pointless to get mad. I will just give you an example. In one of my real estate development deals right now, we have two separate layouts.
have been proposed. I'll just call them layout A and layout B. But I have multiple times made my preference known that I really favor layout A. Right, so I tell the consultants, let's do everything we can to get layout A approved. And then a few weeks go by and I see the engineer sending out layout B to the public works director. And it's like how many times do we have to go over this? We want layout A. Alright, but the AI's make this same mistake. You'll be working on something and you know there's some improvement and it's marked V8 or whatever. And then a few days later you find the coding tool has reverted back to V5 without saying anything. And I guess I just think it's fascinating to watch this happen. It probably says something about how some of what we regard as intelligence isn't really actually processing. It's just memory and focus and why is the engineer still sending around layout B? But neither the people or the AIs are trying to be malicious. It's just not easy to keep everything straight. AI is getting much better, but I have typically kind of viewed it as that very annoying intern that is just good enough that you keep them around, but you're constantly frustrated at the point that you want to fire them. So I would say that, yeah, that's interesting. I have never thought about the managing person versus managing AI kind of connection, but I think that's for the most part, I would say that it comes down to asking the right questions and having the right process for the AI. So when I found out that Claude had some sort of thing in its code that flagged people who abused it and became aware of abusive language, I thought that was interesting. That figure. I'm probably very high on the list of people to kill when Claude goes sentient. Sense of the progression of vibe coding is that a year ago it was almost purely cursor. And at that time, you know, it was still very common to have the model return something with Linter errors, which I don't even know what Linter errors are except they're bad. Then that plateau sort of held through December of 2025 at which time Claude code took off like a rocket ship and captured a huge part of coding usage in basically no time. Then probably in the last few months, OpenAI has caught up a little with codex, but I wanted to ask Cartray about his tools and his usage. Like I pay attention to my usage and I have subscriptions with Claude and with cursor. And if I'm running out of tokens on cursor, I will start using Claude code directly, but I've never, well, not never, but I generally don't run into any issues where I'm hitting limits. So the whatever $300 a month that I'm spending is plenty gives me plenty of headroom. I do get the sense that lots of the discourse you see online is effectively just biases with reasoning built up around them. And I am certainly guilty of this. I kind of in the back of my mind want Google to do well. I don't own the stock now, but I have owned it at various times since chat GPT came out. And so I have this slight bias toward Google, which means that I gave their coding tool anti-gravity a chance when most people just use cursor or codex or Claude. But I don't think I'm alone in letting emotion drive action. And my guess is that lots of the chatter about which model or coding tool is best is probably just a mix of underlying bias and some recent CF Act. Like if you think Sam Altman is a weirdo, you also just magically happen to like Claude code more. But also this is all just supposition. I don't really know. I mean, my point is that a lot of what we're seeing is too random to really know. And so I also don't have a ton of confidence in my opinion here either. Yeah, I mean, so for me, I like cursor. I like it because it has access to a bunch of different models. But so I'm not tied to Claude specifically. I think that and I've had instances where I've had a bug that I had a hard time figuring out. So I've asked Claude and it went into this spiral of nonsense. So I switched over to codex. It had a very similar spiral nonsense. And then I think it was Gemini actually caught on to what was actually going on and figured it out. And then from there, Gemini is often not doing anything that I find particularly useful. So I've also given I gave Claude a task and I gave Composer, which is cursor's model, the exact same task. And they came up with like a planning task not to actually write code but plan. And they came up with different plans. Like the plans were quite different approaches. So I asked Claude to evaluate Composer's plan and compare it against its own. And I asked the composer to evaluate Claude's plan and compare it against its own. Claude determined that Composer's plan was better and had very specific reasons for it. Composer determined that Claude's plan was better and had very specific reasons for it. And so it kind of goes back to the point of these things are very annoyingly trying to tell you what that thinks you want to hear. And so I think that there's a there's a people pleasing component to a lot of these models that you have to try to work around. Like I said earlier, I really do find working with the AI's to be similar to working with people. They have the same embedded upside and downside. I mean, I guess if I'm being honest because the models are way cheaper that makes them also easily preferable. But it's similar. And so it is a little humbling for me because I can feel myself getting frustrated with the management part of it. I've never liked managing really anything beyond my own time. Yeah, I think that's an interesting idea. I suppose there is some something to the fact that having managed pretty large teams has made me good at laying out tasks. So maybe that's the benefit of that is that it's sort of one of those things where I have someone working with me and I know their strengths and weaknesses. I'm usually pretty good at being able to give them the right kind of tasks and projects that I think they'll succeed at. And there's definitely an element of that to working with the AI. I think that hadn't really thought of it, but I think there is something to that. From the very first emergence of LLMs, there have been people online peddling their 99 killer prompts you must have if you don't want to be a member of the permanent underclass. This is something that I think is legitimately a puzzle. Basically, is AI just a boosted tool where the humans that learn to work with it will have compounding advantages over time or is AI on the way to escape velocity and it won't matter what you do today because the models are just on a city march towards superintelligence. I think there's no doubt that right now providing very verbose instructions like when you back test a predictive model for a WNBA game, you are not allowed to use the team's full season win percent as a feature because that is cheating. Okay, that is definitely needed today. If you are not specific, the AI will not get it right the first time. In fact, you can give it this helpful instruction and the AI will go, "Oh yeah, you're absolutely right. I shouldn't let future data leak into the predictions." And then five minutes later, you will find the AI using the actual fourth quarter results of a game to predict the game winner. I.e, it apologizes for a slightly dumb mistake and then replaces it with one that takes your breath away. To take this all a step further, even though these AI's have been trained on almost every word of human thought, they wouldn't think twice about one-shotting an NFL moneyline model and then reporting back a 20% ROI on the back test. You know, because there's a variable that's supposed to be lagged, but isn't. And this would all seem to be very bullish for the utility of humans. You know, if AI's could just halt all their progress and never get any better. I have done some things like don't set default values and fall back values for things and make sure that if there's something missing that you throw an error, that is a directive that every one of my problems has with it. I think that that's an area that there's a room for improvement in my process. I think that there's definitely one of the things that I've run into though is like if I have like a large context and I think, "Oh, I'm going to continue with this context because I want it to know the background." What happens is it starts to get confused and it'll, if it's not a very narrowly focused context, it starts to deviate from the immediate subject because there was something mentioned a bunch of prompts back. And so I actually find myself having to sort of start with a fresh context in order to get it to stay focused. But having those like learned skills is not something I've really done a much of, but I have started kind of playing with it. I think there's a lot of, for me personally, I think there's a lot of stuff out there that is fairly powerful that I'm not leveraging. And the frustrating thing is that it does seem to move super fast. And so yeah, there's a somebody sent me a library to take a look at the other day called, I think, super powers that is like a cursor or a claw plug in that you can do. And it's got a bunch of like canned instructions around it so that it steps through more of a human software developer. Let's walk you, let's walk through this code and make sure it's being built the way you want to build the kind of thing. That's my interpretation of it. I haven't tried it yet, but I'll probably try that out. And I assume that there's a bunch of variations of that out there that are good to some degree or not, but then when the model changes do they become worthless? It's very difficult to know.
how much reasoning you should trust a model to do. I've seen things get out of hand by just lobbing a short prompt and letting it run. And I have also realized times where the model could do a lot more than I was giving it credit for. Sometimes you build something in your you end up blocking some new feature. Like there might be something that your duct tape is preventing it from doing something that would be super useful for you. So I try to be a little bit careful with that. I go back and forth between very, very verbose prompts and very vague prompts to see how it goes. I'm a huge proponent of plan mode. I rarely let the AI just do something with code. I go into plan mode, ask it, see what it's planning on doing, and then usually have a conversation with it back and forth until the plan follows exactly what I want. And then from there, it's usually pretty decent at building usable code. But if I just try to one shot something without really knowing what it's going to do, it's rarely is a recipe for frustration. The intersection of the rise of prediction markets and the rise of AI has meant that a bunch of self-taught traders has had to discover from first principles, essentially how to duplicate a modern sportsbook and also modern market making firms. AI outside of the software development side of things certainly accelerates learning and it is a great brainstorming tool. So the way I like to do this sort of thing is to interact to go back and forth and pull out of thread and dig deeper. I have had very frustrating conversations though where there's some very fundamental concepts that the AI is just completely wrong on like having a bid and an ask at the same price kind of thing where you know that doesn't happen in at least for a calcium order book that's not a thing. And so even though yes, I have for sure learned from it, it has helped guide me through discovering defining terminology and explaining concepts and coming up with conversational ways of discovering what these things mean and how they work. But they've also tried to turn me down paths that are just nonsense. So there's still some hallucination or whatever it is that's going on that and I again, I put a lot of it onto the people pleasing kind of nature of it that if you're confused and you sound confused, it might actually just like be confused with you is sort of what it feels like sometimes. So I guess that's a prompt engineering issue I think being knowing how to ask at the right questions and in the right way is how you get past that. But you know, it's a trial and error and it's also kind of frustrating. So I have also just brought out articles and just read things the old fashion way because sometimes you know the information is kind of distilled and I can be more confident at what I'm learning as accurate. When I make these episodes, I always try to take them seriously. And in some way roll up my sleeves to learn more about the subject. In the case of AI, I've spent the last month or so, surveying the tools to see what's out there and figure out where the value trade offs are. I subscribe to low usage tiers of open AI codex and Google anti-gravity and then the $200 max plan through Claude. Oh, also I have something called Hermes running on a remote server doing essentially automated work with the deep seek API. I guess my biggest takeaway is that it all sort of works. The low usage tiers are incredible bargains whether you use codex or anti-gravity. You just get a lot for 20 bucks a month. Certainly they provide as much code as you would ever need to do stuff like build simple dashboards. I have an app that keeps track of my interactive brokers portfolio based on various factors like value and momentum and let's me do things like set stop losses for all of my short positions all at once. It is by no means sophisticated but it's helpful. And even Google anti-gravity could basically one shot it or it could get 90% of it on its own and then it might require a little back and forth to get to 99%. But then also the Claude and codex high usage plans are worth it if you have something that's going to require more thinking. Actually, even the open source Hermes agent using deep seek which is a very cheap model from China. Okay, that even works. I mean, you would give it a completely different task than you would give Claude but it has utility. For instance, if you wanted to come up with a ridiculously long brainstorming document, you could create an automated loop in Hermes and spend almost nothing in deep seek credits. I did all of this because I wanted to think more deeply about what could be coming next. What are the realities of dealing with the various tools? How much can we get from great models and how much can we get from shitty models with really specific prompts? And how much of the value comes from a model that can get it mostly right the first time versus one that would be more similar to a barely intelligent brute force tool? And then of course, I am interested in the thoughts of the guests as to their experiences. So the tools that are being built on it, I think cursor is great, but it really is just a plug into an existing IDE. So how long does that stay novel? I don't know. I think that I have always been interested in as far as the LLM model of slurping up all of this public information or sometimes not so public information and training a model on it. We're at the point now where a lot of the things on the internet were written by LLMs. And so I wonder if the pool of training data isn't getting tainted, where we're going to be training models on the outputs of the models. So that at that point, there's a there's some diminishing returns where it will only get so good because it's garbage and garbage out kind of scenario. But I don't know. I think that I think that the source of stock market goes. Some of these valuations just seem absolutely absurd. The IPOs that are coming up are sort of mind blowing in terms of the lack of fundamentals to them. But yeah, I think that I think I see it's progressing. I think it's going to continue to progress. I see a lot of value in it. I think that there's room for tools to be built on top of these things where the real value gets pulled out. I think that the and I'm not in this space, like in real estate and whatever. I know that every time I look on Zillow now, everything is like staged with AI and this sort of thing. So, you know, I don't know if that's the the ultimate kind of solution to be drawn from this thing. But there's there's I'm sure there's something there are things out there where these things can be super tuned to actually provide a lot of value. I don't know how much we've seen that yet. I talked to Cartwright a few weeks before the release of this episode and he said that he has been wondering if making AI models will become a commodity business and not the driver of value. Interestingly, I think there may be some worries about that happening as I'm producing the episode because the Chinese open source model GLM seems to be not quite as good as the frontier models. But pretty close. I think that there's there's a trend that happens when new tech comes out. Like, I remember text to speech was like a major thing. There's a way back because I was in the teleconster. But what happens is it's often eventually just becomes like a commodity and the actual provider is a it's a race to the bottom and and where it's interesting is the solution that's built on top of it. I've been looking for that like I've been expecting that to eventually happen but I feel like so far it hasn't I think that the novelty is still from anthropic or from opening. I think the models are also good at some things and not as good at other things. So I think the idea of best model is subjective. If there's rain for more than one of these companies to exist, I think it's going to be because of specialization. I think that there's there's certainly I could see a path where there's a model that is just fantastically good at software development that isn't particularly good at writing poems or whatever other use to buy how. And so I think that's but I what I wonder is the way it's going to shake out that there is a kind of general model which is sort of what we have now that gets specialized things layered on top of it is that the way it's going to go or are we going to train these billions of parameters in a very very focused way so the model just gets exceptionally good at a more specific role. I'm not sure. I think that you know as far as token usage and stuff, I've seen articles about companies where they're they have leaderboards on which developers are using the most tokens. And so you're incentivized to burn tokens. That seems bizarre to me. Very very strange. I think that a lot of what I've seen it feels similar to in kind of the early to mid 2000s there was a big push for outsourcing development. I made a bit of a career for a while of rebuilding projects that had been outsourced to India that weren't usable. And so they finally would break down and pay me to build something that was usable. I think that so companies are excited at the idea that they can have software written without having to pay the for the expense of writing software in a similar way instead of outsourcing overseas it's outsourcing to AI. And in a very similar way I think there's a swing of
It's harder to get a job as a entry-level software developer. But I do think it's going to swim back at least a little bit. I think that there's a lot of room for people to steer this, the role of prompt engineer. I think is going to be a real thing for a while. And I think that, I think the pendulum is going to swing. But I do think this stuff is here to stay. And for sure, at least one of these companies is going to be the breakout sort of Google of AI. And I'm not going to guess who it is. I'm not putting my money in that. Because the other thing is I feel like it could be someone that we've never heard of. There was a time, and I'm sure there was an argument of is altavista or lycos going to be the main search engine. And so, and that was really early days. And I think we're still, we are really early days now. Although I wonder if the trajectory is just creating to the, to fail kind of situations where these guys are moving fast enough that they, they will be the one of them or two of them will be the winners. When I was talking with Kurt, we joked about the futility of making this kind of episode about AI since things changed so quickly in that world. And then while I was editing the show, something we talked about maybe happening actually happened, the space X exercise their option to buy the coding tool, cursor. Well, but it could be that somebody come in, like Google might be in a position to invent something that actually does address the speed issue. Like I find myself using cursor, um, composer 2.5 pretty often because it's dramatically faster than like opus. It's not as good, but if I don't need that level of deep thinking, the speed matters. So I think that there's, there's still potentially opportunities for someone to come in out of left field or, or someone to catch up and leap frog. X AI is interesting. I think that, you know, what they're doing is they've got an option to buy cursor, which is kind of interesting. Maybe, maybe the path that they decide to take is going into the tools, picks and shovels kind of situation because they're also leasing out their data center. And so I think that there's, there's something to be said for that. Since the launch of chat GPT, we have already been through a few micro cycles of boom and bust in the stock market. There are companies that seemed really cooked and then it was, oh no, those are actually the AI winners followed by maybe those companies aren't the winners. So it is very interesting because the real actual business utility is very important and also things are very volatile. So I personally am not trying to pick the winners and losers in the space because I think it is just incredibly dynamic and I don't think I think it's all very hair trigger. I think it will take much for someone to disrupt and I think it's very easy to present as a disruptor when you're really an imposter. And so how do you differentiate that? I think that it'll be, you know, I find it fascinating like Apple, for instance. I find Apple super fascinating because they kind of don't even try. They're just like, yeah, this AI stuff is pretty neat. We'll just package what other people are building and make it usable in the way that we're good at making things usable. It wouldn't shock me if they end up the bigger winners. You know, it goes down to, like I said before, I think it's, when it comes to software, there's sort of, to me, AI has kind of split into, there's the software development thing and then there's the chat, but kind of stuff and generate articles and this sort of thing and images. And I think that I think there's a room for someone to be really good at software development that doesn't care about the rest of it. And I think for the rest of it, it really is going to be a commodity thing and the winners are going to be the people that are the best at packaging it. But who knows, it's still, I feel like I'm purely guessing. And I think it's really hard to be doing anything other than purely guessing what the way this stuff is moving. There are two things that I am paying attention to now because they kind of seem like they might point to what comes next. First, tooling. Most of the improvements we've seen in the past six months have been the result of slightly improved models, meeting much, much better tooling. Second, these better tools make it easier to create more better tools. It seems like there's an entire world out there just waiting for better tooling around AI. I've used the AI tools to create some real estate documents and even though they probably aren't perfect, neither are all attorneys, but that's an outstanding area where the tooling was horrendous. I ended up with documents that had all sorts of ad hoc formatting issues. So the economic potential is very high and there's lots of room for improvement. I've also used AI music for this show. And again, the tooling is terrible. Fellowes specifically are quite good at things that are creative. And music generation is obviously a creative endeavor. I think that that is it. I was just in the car with my two-year-old and he wanted to listen to pop patrol music. So I searched pop patrol on Apple Music in my Tesla and like a third of the songs that I saw were AI generated. They say it. But I listen to a couple of them and they sound like any other pop patrol song to me. I mean, I don't know. So I think that, yeah, and I just saw a trailer that somebody did a full length feature, kind of horror, kind of film that was entirely AI generated. And it didn't look particularly good, but it also didn't look clomically bad. I think that it's a clear step in the direction of these things are going to get really good at closing that gap and creating things that really do look realistic and are actually interesting. I think that, yeah. And I've seen discussions. There were people that would love the idea of take their favorite cancel TV show and just ask for a new season and pay the 20 bucks for the compute and now they've got another season of their schedule. I ask, right, you've seen the progress that AI has made in your domain in not that much time, right? Like the lesson of the last six months is take a large language model, add some loops and it can do a lot. If you extrapolate, do you start to get worried about your value to the whole process? He said, not yet. I think about it. I don't really think about it as much for myself. I figure if I get useless, I'll retire. But I did, and I think that for me, I feel like the value that I add, it's diminishing, I'm not sure, but I think it's not going to diminish to zero in a short enough time frame for it to really matter so much for me. I do think about, like I said, I have a two-year-old and I think about, if he wanted to major in computer science, would I encourage that or discourage that? I think that if you're not in the top 5%, maybe 2% of people in that space, I think it's a scary way to go. And similarly with even creative writing and I guess music generation and all of those different areas, it's to people that are exceptionally good at it, I think, will have a place for a long time to come. But the people that are mediocre, I think they're going to be in trouble. I think they're already in trouble in a lot of cases. So for me getting replaced by AI, I'm not too worried about it, but if I was one generation back, I probably would be much more so. I think that the thing is, humans still are in charge at least for now, so I wouldn't be surprised if we still figure out ways to make ourselves useful. It's just going to be an interesting sort of growth period of kind of a painful process I'm sure, but yeah, I'm not a total doom and gloom. AI is going to end humanity. RISK OF RUIN is written and produced by me, special thanks to Cartwright for doing this interview. I have at least one more episode planned for this AI series where we'll hear from the sports better/data scientist, KANSI. If you want to get in touch with the show, you can email me, risk of ruin
[email protected]. You can also follow me on Twitter @halfkelly. [Music] [Music] [Music] [Music]
(upbeat music) (upbeat music)