Go back

The Future of Poker Strategy (Start Learning this Now)

86m 39s

The Future of Poker Strategy (Start Learning this Now)

Philip Beardsell, lead researcher at GTO Wizard AI, shares insights into the evolution of poker-solving technology, transitioning from product-focused work to dedicated research. His current work centers on Quantum Response Equilibrium (QRE), a method that prevents actions from being assigned zero frequency in multi-way poker scenarios. Unlike traditional solvers that stop updating when an action is rare, QRE ensures that all actions—no matter how infrequent—are continuously evaluated, leading to more realistic and exploitable strategies. This approach not only improves convergence speed and solution clarity but also reveals how players respond to under-explored actions, such as geometric bets or small-size raises, which are often ignored due to perceived risk. Beardsell highlights that human players are often constrained by social signaling—fearing to appear incompetent—leading them to avoid innovative strategies. The QRE model counteracts this by simulating a more dynamic, exploratory decision-making process, akin to reinforcement learning. This shift enables solvers to better model real-world poker behavior, where even infrequent actions can have significant strategic impact. The work also underscores that multi-way poker environments are inherently complex, with no stable Nash equilibria, meaning solvers must adapt to evolving player dynamics. In such settings, a player’s mistake doesn’t always directly benefit another—it may instead redistribute value across the table. Beardsell suggests that future developments could include human-like player profiles, where strategies account for real-world behavior, and educational tools that visually demonstrate how solvers evolve over time. These tools would help players understand that optimal strategy is not static but a dynamic adaptation process, emphasizing exploration and learning over rigid adherence to theoretical solutions.

Transcription

15257 Words, 80934 Characters

English
I actually don't care about what happens in this tournament. I'm not gonna let you use run this table. My dream for me was to build SuperU, man. Poker AI, like it didn't actually solve any Poker game. Everyone wants to do the thing that signals themselves as competent, and they're afraid of signaling a lack of competence. There's not that many actions in Poker, you know, you check, you call, you fold or you bet. And you choose what amount you want at that. Hey guys, and welcome back to the Seeking Balance Podcast. This week we have Philip Beardsell. I probably butchered his name, but he's a good friend of mine, so I feel okay doing it. Philip and I worked together at Ruse AI. He was the CEO and founder of that company, and I was the initial investor into Ruse AI. It is now GTO Wizard AI. And so Philip is the lead researcher on this technology over at GTO Wizard. And so we dive into all of the things that are coming for the future of Poker Solving and the strategic insights that come from working on a solver from a programming and AI standpoint. It was great talking with Philip. He's an incredibly talented developer, and he has many great insights on things like multi-way play, ICM, and stuff that has come from his research in solvers. And now I bring you Philip Beardsell. All right, Philip, welcome. Thank you for doing this. Yes, of course, thanks for inviting me. Yeah, so for the people who don't know, although I'll probably say this in the intro. So you and I have known how long have we known each other now? It always feels like it's been longer than I thought. Almost five years, I think. I mean, no way. Oh, no, no, you joined on the second year, but five years is when Ruse was started, so I guess three years and a half. Three and a half years? Okay, perfect. I was like five years, you're going to make me, you're going to make me have an existential crisis. Here, if life is moving that quickly. Yeah, three and a half years. So we met because you started Ruse AI. You and Marco, what's Mark's for Mark? What's his for a full name? I don't know. What's called full? Yeah, that's all gibberish to me, but you guys started at Ruse AI, presented it. You guys were going around showing you to people I was streaming at the time, and you guys present, or Marco's presented it to me. I think his idea was mostly for just like a stream sponsor, like kind of some kind of thing where I would promote it on stream. Yeah, we're exploring some affiliate partnership programmers. Yeah, but then I saw Ruse got extremely excited about it. I think it was the first outside investor in Ruse and then, I don't know, six months later, we sold a GTO wizard or something. It's not much longer. Yeah, it was pretty quick from there. And Kevin joined in that time as well. So we worked together for very briefly and then, yeah, but you've been working at GTO wizard ever since, right? Yeah, for two years and a half or a little bit less. How are you liking it, loving it? Yeah, I really, you know, when I started the company with Marco, the idea was, well, my dream for me was to build before adding a company and everything was ready to build superhuman poker AI. I like it that could eventually solve any poker game. And so now I get to continue doing this. And I basically have, you know, like freedom of deciding how to achieve this. And so, yeah, now we have a full team. So it's not only a guy working on it. So we have much more resources. Yeah, so yeah, that's a lot of fun. Hey, we had a full team, man. We had one guy working part-time. That's a, that's a team, if you ask me. No, yeah, so what is your title, technically? Like, I know you're basically just in research and development, basically, right? That's like kind of what you're doing there with the AI. Yeah, I guess I like lead the research, but like, I never cared to offer a specific title, but like chief of research or something like that. Yeah, something like that. So what is it? I'm curious because like, you know, when we were, when you were doing this stuff for Ruse, which is now GTO Wizard AI, but for the stuff for Ruse, you were very much, like, you know, doing the work that would be required to get this product fully available for, I mean, it was, it was in beta, you know, those of paid beta people were had access to or many people were using. But, you know, continuing to push product development and things like that. What is, how does it, how is your, how is your, like, job changed from then to now? Yeah, the biggest change is just like, in that two years and something that we developed Ruse, like, the part where I actually did machine learning stuff was maybe like 12 months. So it was pretty short and because we had to focus on all the business stuff, lawyers accounting, all that fun stuff. So, and now I get to focus only on the research. How do I, I'm involved in with the product? Things when you release something to just like, you know, how the, our current, our work will be deployed in the application and stuff like that. But yeah, being able to focus and represent on the research and being able to actually take some time to again, read some research papers and stuff like that. So, yeah, that's the biggest change. Like, just being able to have a narrow focus on that. Nice. So, I guess I'm curious to, how much of the, what is the, like, back and forth of, like, product, like, featured development? Like, because you're doing research, I'm assuming your research isn't that directed by, you know, like, Matt at GTO Wizard, Matt's like the CEO, GTO Wizard. But if they, like, I guess what does that feedback loop look like? Like, like, upon your research, do you stumble upon product ideas? Do they say this is a product idea we'd like? Like, how does the actual development of specific features and the products get, you know, is it kind of both of those things? Or can you just, I'm curious? Yeah. I guess usually it's like the other way around. So, by working on the research, I kind of, you know, since I used to play poker, professionally, and I use all the, like, products and quite knowledgeable on this. So, I, with the constraints that happen through, like, the technology constraints that we face then come up with, like, ideas on, well, what's possible given those constraints? Now, I will make it in the product, because, for example, you know, someone that doesn't know the technical constraints can be like, well, I want the, I don't know, like, you do loadliking on the river, then it propagates, like, all the way back to pre-flopper something. And this would be impossible with, I mean, nothing possible with our current technology, but it would be very difficult. And, like, not something we would want to prioritize. So, by no way, involved the product in the technology, I kind of have a better idea of what's possible. And in the context of, like, multi-way and stuff, it's like, you know, solving multi-way in itself, like, it is impossible. So, you have to make some decisions, some approximations. And, yeah, I guess, my team is better placed to make the decisions of what should be how the solvary will be constrained, and then from picking into account the product, and from there, what the product should look like. So, the most recent update, I might be wrong on this, because I don't study a ton anymore. I believe was the profiles, like, the different profiles, is that right? And that way, then you released this, the feature itself was released, and then there was an update to it that was released as well, is that right? Or is that right? After that, there was, like, a custom profile release that you can set your own. Right, right, right. I actually thought one of the ones that did not get enough attention was the, I can't, I can't remember what you called it, but the, you know, the ghost tree solving or whatever, or like, the, you know, yeah, if you could explain what this is, but I explained what it is, and what it does, because I thought it's, first of all, it's just makes looking at rare spots, where like fish do things, so much prettier. But second of all, it's just a really cool way you did it. But my follow-up question is, what is it? Is there of any kind of frustration of delivering something like that that I think is really innovative? And I don't know if you guys thought was, you know, we're excited about that probably many of the users don't really notice or care about the difference of it. Yeah, yeah, definitely. I mean, I was like excited about it. And then, but I kind of understood that most people wouldn't care. But basically, how it happened is we, you know, we, after releasing ICM, which has been like a year and a half ago, even before that, we were already working on multi-wave. And by working on multi-wave, we kind of, you know, the technology we use, which is like we solve. some deplimated trees and then it's kind of similar to how it's done in games like chess where you saw the small smaller version of the tree and you cut off the tree at some point and then use new on that works to approximate. How much EV would gain if you were to get to this node. And this allowed you to solve smaller versions of the tree and in the exemplific literature they didn't really. Be this in the multi way setting so they they there was a virgin but called purpose maybe five years back that that was a multi way poker bot and other than but that one didn't didn't use neural networks. And so actually when we're trying to build like extend from two player to multiplayer who are faced a lot of issues new convergence and stuff like that. That's when we're trying out different techniques and without getting too much into the detail like by by kind of trying to to regularize the solving process to make it like more well behave. This is like based on recent research and essentially useful for different equilibrium, which is called the quantum response equilibrium and there are few variants of it, but essentially you can make this quantum response equilibrium as close to nudge as you want to. And so you can get to very close to nudge but with different kind of equilibrium and our goal was just to make the solving more well behaved, but then when we. We managed to do that so the solving was converging well and then we just put it in like a in the development environment and just looking at the solves to see how you would look and we kind of so that for free we're getting this like really nice solutions in the coastlines because this is something that if you check the exploit ability of a solution is not going to show up because this over never gets to those lines so it doesn't matter what you do there. But like yeah visually it looks much better so we're like okay we can. We have all the technology for it so let's do a release for it and maybe some people will like it but it was mostly just for us to release it and announce it so that people don't don't. I'm not concerned like why did the solutions change so much from one day to the next. I feel like it's unlikely there's specific thing you're noticing would have been noticed but may you know there's a lot of users so maybe but yeah it's really cool because you. Like you were saying to let me just try to summarize this because there was a lot of you know you're you're very sharp at what you do so I don't want to lose people on the description of this but you basically had. So you have these like lines so let's say there's like a range bet spot like a range donk leader something but then someone checks the response to the check is not a really a reasonable response based on a. That you typically before this update that you would typically see based on what what what a reasonable checking strategy might look like in that spot like let's say some kind of like if you were forced to check that spot what kind of optimal strategy that would look like. And so the response to a line that was basically never taken. Um was really messy it would maybe like you could maybe see some pattern around a strategy that kind of made sense but it was really was really messy and not not super clean and so. So the way you did this is the if I remember correctly you're basically solving the the strategy at the frequent like not all ghost lines are the same there was something about the fact that you're solving based on the EV. What in more precise what exactly change because now if you take this line first of all it's what's cool is if you to click check the strategy you see when you're in the response line both are some kind of optimal strategy right so like now the person who checked even if it was never done initially right like we're very very rarely done the strategy you actually see happening and that is is you can actually also see like well if I was to check this what would that what would that optimal strategy look like like what would it be something like. Like and then the second is yeah that the response to it is much cleaner and like clearer around what kind of is happening. But yeah it had something to do with kind of what I'm just trying to remember I wish I'd reread this recently because it's you know you're not you're not just your your yeah explain a little bit more what what or a little more precise what exactly change with the solving to make this happen. Yeah so basically you know how this server typically works is that they're really at every node comparing a range versus range so you know if you if you get if you're solving and then at the given iteration when you're solving your current strategy that is over things as converge to. Then your opponent which is also there's over but you know face against itself looks at this strategy and say how can I exploit that and and what would happen normal is that let's say that we decide that it's never worth it to take this action and when this over is looking how to respond to this action then you're it never gets there so how do you respond to a strategy that is like zero everywhere well. You don't you don't do any update yeah just just a cut in because in the real in the real way the solvers happening right because you're solving a single street let's say the flop. So let's say there's actually zero percent chance and they happens. Before this is the solution you would see kind of what had been explored until let's say zero percent had been reached because it's it's solving right like it's running these strategies against each other right so like you might just have kind of some strategic residue if you will. Of like as long as that street was being explored in the strategy some balancing was happening but then at some point the frequency reaches zero and then that no just stops updating like it just there you just it's you know there's no more that's not getting explored anymore so just. So what you would see in the prior version is just kind of the last strategy that kind of had been gotten to before zero percent had been reached is it something like that yeah it's exactly like that so what you see is the minimum. But the minimum response or for the phone and not to one you ever reached this line but it's not the maximally exploiting response so. Is it what can happen actually and you could see this in the app like that before and after so you could also see if you compare with. For example bios over but let's say that you have a line let's see you have three bed sizes and you have like a very big dunk size and it never gets used then you will see in like it's over like a bio or. Over that sulfur and that is often you will see that the even actions that are taking with zero frequency will add these are very close to the actions that are taken and it will look like nothing is a mistake. But that's just because basically the line will stop being explored as soon as the response to this line is good enough. But actually in reality if you were to take this action and the opponent was responding well. This would be a really big mistake so what you would see now when you solve for QRE is that lines that are taken with zero frequency the EV loss is much bigger so. It's a much better indication of how much a mistake is worth because now you actually have an opponent that's responding well to you in lines that don't happen. And so exactly so yeah first off so this is interesting because it will cause it will actually cause a quicker I imagine DV or convergence as well on if any action should ever be taken because usually even in the description you're talking about it's you have to run a solution for a long time to truly get to a zero percent frequency like a true zero percent frequency for any action. Because of what you're saying the EVs are kind of close even if it's happening very little but in what you're saying because in the QRE the you know let's say we touched on put open jamming in a very deep SPR. That because there's an actual a maximal response frequency being solved for that action is really bad if someone responds well to it and so I also think it would it cleans up both sides of it it cleans up what what the response is but it also cleans up quicker how quickly things converge to zero percent I imagine like I how quick or like the truth of like like zero percent frequencies are found more often is that true. Yeah well I mean I don't know if that's true but it does converge faster when you do this because of that because yeah what would happen often is that it's a at some point in the inside solving this line is taking zero percent so it stops updating but then later it realized that actually maybe it does want to take this line and then. Yeah that makes sense that makes sense because what would happen you would have a line so like let's say something goes to very zero but it's a it's a local minimum you know strategy it's like it's not actually like that's not actually bad it's just for now this is better. So then you have this other action be taken but then it starts the response to start becoming better so this action starts losing EV from the action that was kind of pruned and it's like oh wait we have to go back to this action now. And now we have to re-solve for this yeah that would make that would be very yeah very long a long cycle so what is so what is different. Because I understand how QRE results we've described while I think the difference that results in. exactly is it doing differently to do that? 'Cause it's solving the lines of the frequency with which they're taken, but it's continuing to solve them. That's like in my head somewhere or something like that. - Yeah, so essentially how it works is that you want to avoid having these zero probabilities everywhere. And so one thing you could do is force every action to be taken with some frequencies. And if you did this just like computers can store its rarely low values, so you can take an action very, very rarely that you will never show in that location. But if you do this kind of lively results are not gonna be that good. So one way to do this instead is kind of the idea is to have the probability of taking the action proportional to how good it is. And so in the QRE case, what you do is that you will kind of have this off max function, which is usually the typical way that the solver would do it is that if the action has the high ACV, then I think there's 100% of the time. And then I've reached it out over time so that I can have these like mixed frequencies. So an action can be good several times and another action is better and then you average it out and then you get these like mixed frequencies. And then in the QRE approaches that instead of saying, let's say this action has EV of like 10 and the other action has EV of nine. And I will take the action with an EV of 10, like 60% of the time and the other 40% of the time. And so I will kind of spread out my probabilities in the true to make sure that I can never get somewhere with exactly zero probability. And then in practice, you can make this very low so you can take an action like one, 10,000 of the time and then you will never see it in the UI, but then it would make the tree more converge. - Yeah, but you're also not wasting time, like the worst decision gets, the less likely it is to be explored, right? So it's also faster and that's cool. Yes, yeah, I think this is like, I find the, well, one, I also think that there's usually in, when we were working together on things like Ruse and coming up with the ideas and you know that whenever there was conversation about actual like methodology on how to make something happen, something like this, I always found it pretty, well, one, just intellectually stimulating, but two, there's typically some analog between this thing and it's efficiency and like coding in a coding sense or in a programming sense. And something like there's an analog to reality of playing, like a reality to what you consider in playing. And I find it interesting just this idea about rating your confidence and how people themselves are quite poor at responding to parts of the game tree that they never explore and that one of the interesting things that I've worked on in coaching people is having a strategy that guarantees you explore more parts of the game tree than you would intuitively think about like having a way of coming to a decision-making process that actually bakes in more exploration. And I find this just analog with this kind of approach interesting where it's like you, there are some decisions that are actually quite good that people never consider and so never realize they're good, but you also can't be so sporadic in considering options that you're just considering all options equally at every moment when you're playing. And so this is just this idea of like exploring things at some frequency based on the proportion of how bad that you think it is, but for the overarching idea of getting a better sense of what that strategy looks like, what did this people do and how they respond. And so yeah, I just find the analogs between making something work in this sense that you're doing it and that they're typically some strategic analog that can be taken from it. - Yeah, that's really interesting. I didn't thought about it, but yeah, the reinforcement learning you have this exploitation, exploration trade-offs. And essentially, at a given time, you have what you think is the best action to take or some mix of them and you can continue doing that. But you also want to have some exploration of, you have some confidence interval around how, how actually you think this estimation is actually is. And if you are to train some like agent in any game, if you don't have this exploration component, you will never converge to anything. But you always, I think, in poker, definitely you want to be realistic about how confident you are. And most likely, you're in that accurate in your estimation. So you did, I think, just sometimes taking some action that you might think is not like, oh, if I, I don't know, you never raised in some spot because beating people over full, but you haven't raised there for like a while, but you might be completely wrong. - Hey, guys, if you've ever felt overwhelmed by just how much there is to master in poker, mental games, study habits, lifestyle, strategy, you're definitely not alone. That's exactly why I created this group coaching program. We'll tackle all of it together. You have a supportive community, will master strategy, sharpen your mental game, fine-tune performance. So if you're ready to stop juggling it all alone and start thriving, this is your place. Click the link below to learn more. - Or one thing I've come into contact with is people who never, they had, like, let's say they're particularly short in like a three-bet pot in a tournament or a four-bet pot in cash or something. And they've never actually explored like three tiny geometric sizes. This is something where they'll put like some 25% and then have like half pot on the turn. And they either are thinking about jamming or checking on the turn. They're not thinking about the utility of like, well, why don't you just do two 20% bets or something? Like, you know, there's ideas that just seem unintuitive. And there is something to just having this spirit of exploration like you're talking about. That's just like, yeah, I'm just gonna explore things and recognize that like the utility in it is not just in being correct right now. Like, it's not just in making a good decision it's kind of just to see what happens and how people respond and what things people do. I think this is stressful 'cause for a lot of people they don't have a good feedback loop on how to assess if the thing that they did in an exploratory way, like to access the utility of the information that they got. There's like a lot of noise. There's also a lot of really important information. Like, there's some showdowns that will meet a lot. Some showdowns that will meet nothing. Some player types for whom showdowns mean more than other player types. Like, there's just a lot of noise. And I think without a, people just generally don't like to explore much without a clear sense of how to use that information. But yeah, I mean, I think you, go ahead. - Yeah, you know, I played poker before the solar era and it was this phase like the MDF phase where like, yeah, to defend MDF and you know, you couldn't check rays on the floor to like keep your range protected and stuff like that. And I was, you know, I was playing like midsticks or so. And it was very common for everyone including me to play very robotic, which is like, you know, Siebe at the same style, you know, to three quarter pot or something. And then go from there and stuff. But they could have benefited so much if just like, if one person of the time I just did a random action on the floor then maybe I would have discovered, you know, Siebe at small actually, it might be fine or stuff like that. But yeah, you're so stuck in the mindset of like, I know I have to imitate other players rather better or I'm afraid to like do something that is weird because no one is doing that. - Yeah, well, and there's a lot of into it. There's a lot of things that initially seem crazy that are just totally fine. I remember starting to do geometric sizes at deeper stack depths in spots, just trying it out. And I remember one of the first things that was interesting about running when we were doing ruse is seeing how the seeing the spots where geometric sizes get used still even when you're very, very deep, you know, like you're in a single ruse pot. I remember just stumbling upon putting it enormous turn probe sizes, like that was just like a thing that I just started doing a lot and it was not common. Well, at least when I was playing on 2K Iggy, I think this would become more common. But like, 'cause I think, you know, in a single ruse pot when it goes check, check, I think geometric is now 250%, so you just 250% turn, 250% river. And yeah, just like it wasn't being done. And I was like, this, I don't know, this seems fine. (laughing) And I think once you start doing those things, you start getting a sense of it of like, well, the hard thing is it has to be something you also don't do just once, right? You have to have a sense of exploration where you're experimenting with something over a stretch of time to really see it. And I've had things I've explored with that, like the information back has been, this is terrible. Like this is not good, but there's a real utility in that as well, 'cause you have a deeper sense of it's terrible in a felt intuitive way, right? I know what people are doing. It's clear to me the way people responding, that's terrible. But oftentimes in even discovering something is bad, you discover something that, what is good, you know? You're like, oh, well, they're responding in this way. This probably works then, you know? - Yeah, I think because realistically, if you do this small amount of the time, it shouldn't have the big impact on you, but potentially, you'll discover a lot of the cool stuff. Yeah, I think it's interesting that things become obvious once you know them, but for a long time people were playing, you know, always the same sizes and you know, barely over beds and no small beds, but it's like it was always there. You know, there's not that many actions in poker, you know, you expect to call, you fold, or you bet, and then you choose what amount you want to bet. So it's crazy that no one was like discovering these things. Yeah, I think it's just cultural. Like everyone wants to do the thing that signals themselves as competent, and they're afraid of signaling a lack of competence, like a lack of knowledge. Like if you're in a subculture of poker, I do a lot of coaching with people in like life, private cash environments, and it is, this drives the size as people pick more than anything. It's just like, what will people think of me doing this? You know, and then it goes both ways. You some people who want to look like whales, right? Because they want to stay in the private cash games. So they don't want to do things like over bed, right? So they just don't want to. It just looks to GTO solvery. But it goes the other way. You have people in these environments who care about looking good in competence, who are maybe playing streamed games, and they're professional players, or like, and they want to look good. They're in the eyes of people. And so there's something that they just won't consider trying because they're afraid that it will just look stupid. Like it will just look dumb. But that's why I think it's interesting to like bring up something with someone who's probably at the forefront of, you know, poker solving, you know, technology, the ways in which even a computer, it is like better for it to have, like you said, some kind of exploration built into it. Otherwise, it just won't find the best strategy. And it's easy to, I think, from to to get the result of our solve and be like, oh, I should just know this result. You know, we've talked about this before. But really, if we've viewed solutions more like snapshots of a process that was that was happening. If you had to, if you had to watch the whole everything the computer did to get the solution, I think it'd be so informative for so many people to like watch. Like this thing is finding this, this answer. And like you should, if you want to model this, like model the process is so much more than like model the answer. But yeah, I wonder if anyone has ever done that in like some nice, you know, video or going through different solves and like seeing, yeah, like the strategy is that, you know, silver comes up with and then it's immediately like refuted and then comes up with something else. I think this would be something you guys should do because if anyone has like the resources to do this, it would be sick to create a fake match that was just a solve happening. You know, I mean, so like you could watch people's strategies, like it could be like two fake opponents, but it would be snapshots of strategies and watching like the opponents adapt to each other and things and like, you know, like a fake heads up match or something. There's a, I think there would be a way to produce a video like that that would be so interesting because you would at the every, you know, you'd have parts where people are coming up with these just like, you know, these are two imaginary opponents. You'd have moments where they just have like these crazy dominant strategies where they're just like destroying the other person while the other person tries to adjust and find the answer. And like, I think it'd be, it could be done in a cool way that would be informative for a lot of people. Yeah, I mean, that would be really cool. I always wanted to do this kind of video where, you know, it's like just an agent that you train from scratch and then it does all these mistakes and then eventually becomes really good. Yeah, yeah, that would be cool. So what is the stuff that you guys are like that you can talk about? Again, I don't know the degree to which there's confidentiality in any of this, but what are the stuff like you're currently working on or stuff that you've been working on that you're excited to see be released? Yeah, so if you think so, well, the main thing is multi-way. So we released three-way a few months ago and I was like a proof of concept for like multi-way and just, you know, basically to go from two players to two players was by a lot the biggest cap and then like three to like up to minus. It's just an extension of it and so, yeah, we're about to release a piece of multi-way. Yeah, that is. I think honestly, for me, I think it's pretty sick that the offer that it is and that it's very accurate. So that was a big kind of initiative and then after that, we'll have some, so basically have these like for me like building multi-way and everything is more like a building block that will allow other things too. So for example, the profiles right now is just a very simple profile. It's like works with incentives, I could explain what that is, but basically it's kind of a. You just sprinkling some EV on actions, right, based on the profile. Like you're just like this, this action's worth a little bit more to you or this action's worth a little bit more. Like if you're, if you're a station calling, we're just like adding invisible EV to calling for you, the player, is that right? Yeah, it's a very artificial way of trying to model like non-GTO players, but what we're working on is is having real human like profiles, so that, you know, eventually you could set up the six maps solve and then you can put, you know, you can buy three of them, buy three of them drinks and three of them are drinking and half of the table's drunk, the human like in this way. Yeah, that would be nice, but yeah, I mean, ideally the vision would be that let's say you upload your hands or then you have some data on your opponent's and that from this, we could even create profiles and you could say, I want to play necessarily this opponent, but someone like him or like this stuff, Greg, that I have in my pools, takes an average of these players and then you set it up, you set it up in let's say in the big line or in the button and then you will know how to exploit this opponent and then the idea is basically then you also have this profile is kind of incorporated into the solver, it's kind of essentially you have this knob where you have a GTO player and then you have the your prediction of how the human will play, then you decide on this line kind of how much you are close to the UN player or not and so and based on this you kind of will know you can if you were to set the player this knob to 100% where you're predicting the solver to exactly like the player then you will learn to maximally exploit this player, if you put something in between then you will learn how to exploit this player but kind of give him some credibility that he would also respond back and that's right to exploit you back so it will I think this can vary with multi way. You know you should have to is you should have a heat map of the degree to which it differs from GTO. Because even if I see the maximally exploit strat it'd be really useful to see like what are the ones that are the furthest from optimal right like what are the things the decisions and the ways the strategy has changed so the furthest from optimal which ways are just are basically not changing at all because it'd be an interesting way to know to prioritize you could do you could theoretically come up with some you know heuristics or things based on just what is the like the what are they like seeing like the big the big things that like the maximal exploit decisions that are making a lot and the ones that are not making very much but are still exposing you to a lot of risk if that makes sense and then make some I think that's I think that's an interesting idea. Yeah that's a really cool idea and yeah I think I think essentially it's a lot of tooling that can be built around this and I think that eventually this is how people would study. I mean I think in cash game everyone what's special about cash game and poker except for like crazy life environments is that everyone has converged to play how this all would play so I was talking to one researcher about this and you know in some games like for example research was trying to make a bot for like a strategic which is like this game collaborative game where yeah well it makes collaboration and competition where you try to dominate like the world map or something like that I'm not actually played or that's like diplomacy I mean let's try to go and if you try to create an agent that plays this game and then you like from scratch and you try to make it converge to something close to Nash and you make it play against humans it will it would lose because it's converging to a niche that is completely different than what humans play and these aquariums are not compatible and we can talk about it but that's what happens a lot in ICM and so in this kind of games it's actually not useful to have a or not that useful to have a solver that plays close to because if you play this, you can learn, you know, why it gets to this strategy, what are like cool, you know, strategies, you can come up with why it comes up with these decisions, but in practice, you cannot play this strategy, you will get this right. So you want something that models humans. Yeah, just to explain this bullet points of kind of what I understand you to be saying and correct me where I'm wrong, but. You're talking like listen, pre thought, multi way, it just so happens that like on six max cash in most places, most people have theoretically opt, you know, what they've seen as solutions and, you know, there are most people trying to emulate these solutions more or less. And so like a solver like continuing to copy that strategy and that environment works. But because there is, there is an infinite number of equilibrium and multi way games like this, right, is that correct before I continue. Yeah, I mean that's for any game, but in this case, they're all different. Okay, yeah, yeah, yeah, there's not an infinite number of Nash equilibrium for heads up zero some games though. Yeah, there's they can be infinite amount of Nash, but they all have the same EV, but that's the for the players, so that's the main difference. And when paired against each other. So like if I picked if I just if I threw a darkboard at a, you know, one one Nash over there and through a darkboard at the other one, you're not going to pick to that beat each other. No, they would always get the same even. They would always get the same EV. Okay, I'm just making sure not the same EV against like in the in the particular environment but against each other. Yeah, I get any other Nash they would get the city at you. Okay, okay, but but this is not true multi way, right, this is this is right like this is there's not a stable Nash equilibrium in multi way environments where I could in theory. Throw a dark same at a Nash equilibrium throw eight darts at a different strategy that everybody else is playing, if that makes sense. And this one could be losing to these eight is is that right like and so you're saying this that if you had a solver trying to establish. You're saying it would lose because it would be trying to it would it would get established a Nash that's nowhere close to humans is this due to a probabilistic thing because if there are an infinite number of Nash's it's just guaranteed to like find one that doesn't compare well. Or is there actually something about the way solutions like these these models are being trained where they will establish solutions that are losing and environments where people are playing a totally different strategy do you get to kind of get what I'm saying. Yeah, it's the latter and I don't know how best to explain it, but essentially it's it's kind of you get these Nash equilibrium that are not compatible. So, for example, you know, if I if I make some. You know, multi way solver, you know, when we did it, they're using different technology converges to basically the same as other solvers, so seems like. In I think some cash games like multi repeat up all them. It's very hard to get to a natural it's completely different. But even then, well, I guess it's a different topic, but yeah, for example, to go back a little bit, you know, if I play. If I have this, I create this GTO bot and I could play in a live environment where everyone is limping and playing crazy or like. Some combination of actions and then I'm under the gun and I keep opening another 20% of my range or something, and it's it could be that now my strategy is losing because I think that I open someone called and someone's please and then I have to fold something and. So a lot of my EV came from I don't know. And that now have to fall to a tree that. And so I'm actually like a lot of a lot of your EV could just come from full equity. Like just a presumed level of tightness that like I open this hand space on blockers and interaction and the other ranges that are opening. This hand might be barely profitable, but purely because it's just going to take down the blinds enough and that the moment someone's calling any wider. The more that hand is losing money now, right, like that that that could be something. Yeah, so and you know how it would happen. Normally is that the player is kind of taking money easy from you and transferring it to other people. No way, or on no way. And yeah, so basically, you know, sorry, I just have to linger on that point to make sure people get that too because it's not always in multi-wigames. It's not always the case that EV lost by you is gained by the person by specific. Like it's not like just because like let's say I'm opening the end of the country like you said and some of the threshold hands are are actually losing in this environment. It's not even necessarily the case that I. As a player watching you do this actually can take that money from you. Like that's not even necessarily the case. It might be the case that I that I can I can better adjust to this environment. But the EV might just it might be spread amongst other people might not even go to me. Of the like your mistake is not zero sum so like me there's not always ways for me to take advantage of your mistake in ways that cause EV to go to me. There might just be ways of me making EV go from you to like all the other players in the game. Yep. Obviously you have this kind of power to transfer money between players. But does that does that mean there isn't a way for you to transfer money to yourself because it seems like you're saying like you this power to transfer money between players and leaving out the opportunity for you to take it for yourself. Yeah, I mean you you would against the national opponent and you cannot. So, you know, like a national basically means that you cannot. Internationally improve your own EV so, you know, basically you cannot do anything that will improve your win rate. It doesn't say anything about. Well, it took player into player zero sum game this automatically means that the money goes to the other player because there's no one else but in the multiplayer games it doesn't mean that the money goes to the other player actually it could be even losing. From from this mistake. Yeah, you could cost you both money but you can be trying to adjust to take advantage of it and both of you could be losing money and just giving it to the big blind or something. Yeah, so I can see in my data that 70% of you do not subscribe. It's hard to over estimate how much your subscription helps grow this channel. So, if you appreciate what I'm doing here and you want to hear more of it, please like and subscribe so we can continue to bring you more great value and great content in the world of poker. Thank you. This very simple example it's like when poker and it's like you can have like one of. I forgot that true player example but it's a very simple game where each player has only one card. And there's like one betting round and there's like one. There's not even any community cards and in this game you can actually see that there's like. An infinite number of nudge and in this game the player is first to act I think like how it balances is basically the size is chicken is checking frequency and based on this decides who ends up making money the second player or the third player. And I found that choice while not getting any idea but basically you know it in this kind of setup without any kind of a illusion in the sense of like private information sharing you can transfer money from one player to another. Right. So what are the big takeaways that you've had or like insights you've had into multiway strategies and the things that people are doing because I guess to bring this back to you're talking about this. We're talking about how like the solver would come to something that would work in six max games but then you throw them into kind of more chaotic environment like live cash. Which I'm more involved in these days I guess I'm more interested in it as like the environment where but I guess I'm also curious even within six max. Sometimes you do you how much just throwing a single well into six max environment incentivize everyone to just play differently. That's one question and then the second is what kind of insights to have you taken that people can take the cash to live cash to improve their multiway game and we'll get to ICM after that but. I would say for cash I haven't done any extensive testing in like six to nine players, so maybe it gets more dynamics get more chaotic there, but in three player case it's surprisingly well behaved. So I like in actual you know you create some single race part or something and you try to make one player dv massively both players always benefit so it's never like you can hurt one players dv at least from like maybe a few dozen situations I tried. So you know if you throw if you throw like in another set up you have six players you throw one you put one player to be a whale. All the five other players should benefit, usually, without changing, without changing much. Yeah, and obviously they could make more of they were like aware that there's a well, and they should obviously adopt their strategy, but even if they just keep playing the same strategy, I think it's kind of hard to come up with some example where there's some massive differences. Well, I guess I'm thinking about the sense that like, exactly the example you hinted to though, where under the gun is opening their standard range, but they have all of these break-even opens that really do just rely on full equity, right? I think that's a common part of any parts of any person's range, and then you just as well, that's always calling, and they usually have like an attentive rag in the big blind, who's going to try to ISO the well. This feels like this would hurt your strategy a lot, it feels like you have all these hands that are, it's not that I feel like this is a, my instinct is that this year opening hands or a break-even or making slightly, you know, slightly winning based on just like full equity are now losing, but you're putting money in with them. I feel like this is money that at a really high frequency is just going to end up dead money and their pot, right? If you just have this whale who's like a station and this guy who's always attacking the whale, so it doesn't seem like, what are your intuitions around that, is there's something I'm missing, and it doesn't seem like it would take that much to throw off, what is the actual correct thing to be doing in even a six-max cash environment? Yeah, I think in this case, I would think most likely that this balance is out, but this is balanced out by the fact that we're good hands are making much more now, but they could happen, I think, you know, entirely differently can happen. So the overall strategy, so those hands might become torches, right? Like so, like the your threshold hands might become torches, but your good hands are going to make more money's going to be getting put into the spot so often. Yeah, that your overall strategy is not going to actually cost you that much. Yeah, what I mean is that Evie, I mean the expected Evie is not very sure to play a thousand hands of the whole strategy, they're not not just that hand, because I think I was in the headspace of like, should I should I should I continue to open Jack 9 suited, you know, or something like yeah, it's like probably probably not, you know, but but if you did, if you're unaware, what you're saying is if you're unaware of the dynamics and you're playing just based on a solve that you saw, it would take some pretty extreme dynamics or precise specific dynamics for you to not just cap the for your whole strategy to not at least be the same as what it is. Yeah, I mean, I would be curious to know, I know, I know in AAA it's both from the definitely seems to be true, most of the time, but I haven't tried like six player, I think I could be for an experiment, because then I think it's less obvious, because you have so many players that can adapt, you know, let's say that you do this open and there's a will after that, there's always players that can respond to it, and then by the time it gets back to you, maybe you're like, cannot make money anymore. Yeah, your hands are just cat. I mean, this happens. Yeah, I mean, I've been doing a lot of work as a squid game, and looking at different pre-flop solutions with squid games, because I've stood a lot of students who play squid game in the game, but this happens where if people are wrong, I've looked at this where people are going to like, their spots for limping is good, but you also just expect re-limping to be the default move in squid game, where you expect limping to just go around, that that should be like 90% of the time, that should just be what happened if the squid bounties sufficiently large, but you won't, you're going to get, you're going to get ISOed way in a lot of environments way more than you should, and so I was talking with some people about how it's the most important thing on if you decide to open, to have a limping strat as the first person in squid game, is actually the next person, and if they limp or if they ISO you more often, because if they limp, which is correct, but then everyone else is ISOing more, so much of your limbs become terrible, because you just have these two people in position on you now, like the next guy, you know, so like you actually can't flat this very wide, you'd be better off just opening to a normal size and not getting three bet, right? But if the next person to you is ISOs you, which is also wrong, but then people aren't aggressive to the ISO, right? Which is fairly normal. Then you should limp everything and let the ISO or ISO you and let everybody fold, and then you just get to call basically everything, you get to basically fold no limbs in that spot. So that's kind of an example of where just like this, the next person to you could shift a lot, I could see in like six-way dynamics, and if there's like a whale, the next guy to you versus two people away from you, and like what the guy in between you is doing, I can imagine changing a lot. Yeah, I can definitely see it happening a lot in these kind of games. So maybe if we are getting to more weird cases, like for example, in ICM, I'm not sure actually what's causing it either if it's because but basically if you have this utility curve, like how much dollar EV is worth every chip amount, or like your stack and chips, then you get this kind of curve which is like every chip is worth less and less. And either this or the fact that the utility for you and for me are completely different because we have a different stack. And so for you, like dating, dating, let's see, I have a 20 big clients and gaining one additional big line is worth way more than for me if I have 100 big clients. And so you want much more this big line than I do. And so either this is symmetry or the type that is not like linear payoff, like every chip is worth a different amount, or a combination of both makes ICM now a very weird environment. And I think it's good game. Yeah. Also, I think this is symmetry, right? Or I guess it depends on them. I'm not too familiar with speed game, but I think it's not worth to say. Well, so Nick, Nick game is easier squid. I guess squid and knit are comparable, but it's easier to talk about Nick game just because this basically you're either in the game or you're out of it. So like in Nick game, it's just like if you win a hand, you get a button and then you're kind of out of the game and I kind of pay. And whoever doesn't get a button at the end pays everybody. But yeah, there does become an asymmetry in this. But it's it's more predictable and we're I think more stable in the sense that there's basically dead money. So like let's say I'm out of the game and you're in it. And even if you're one of the last two people remaining, you have this kind of massive and visible anti in the pot. And I have I'm just playing for dollars. Like I'm just playing for the visible dollars. But that's kind of what happens in Nick game where like you the pot's just much larger for you than it is for me. But that basically I'm actually not going to get too much into strategic ideas about that because I do a lot of private coaching on this. But but there's there's a lot of interesting ideas when that's the case when like this pot's larger for you than it is for me in ways that people I think don't adjust appropriately. But for ICM, which we can move to because I think we've covered some interesting multi-way stuff for ICM. First of all, it's not I don't think it's intuitive to everybody that ICM is a multi-way environment. Even heads up spots in ICM are multi-way environment because the way money gets distributed in a tournament is not based on who won this pot but who's left standing. So like you're competing with everybody all the time. So every hand is a multi-way environment. You have yourself and your opponents and then the field. Yeah, I mean say more. I have my intuition seems to always run into each other when trying to think about how to exploit ICM. Like you tell me someone's too aggressive. Let's say I'm at it. Let's say it's the final table. It's nine players. Right. And we have this kind of like guy who's too aggressive. And he has a large chip stack. And he's more aggressive than he should be. Or he's running into let's say he's playing like a chip leader. But there's another guy who's not playing like he's not chip leader. Like there's a guy who's just like to to like counter-aggressing more than he ought to. He's like, you know, second in chips is playing like chip leader against the chip leader or something. Right. My first instinct. I guess it's just like you should just like let these guys battle every hand. Right. Like that's my first instinct. Like if you have two people like one of these people is just going to bust like they're just playing to aggressively against each other. But then there's just ways of describing the situation where I feel like I come to the opposite inclination about where there is to be aggressive. Like if I have a table that's too passive, should I be I can't even get to the I'm trying to like recreate. I can't even get to the inclination. But I don't know if you have had this in your work because you're saying ICM is a weird environment. What makes it weird in your mind? Because I feel like sometimes when I'm thinking about what to exploit based on the tendencies of the the pool and like my opponent and stuff like that, I can come to what feels like both answers are good. Like both and it feels like I should be more passive and I should be more aggressive. Like if I'm like, yes, I shouldn't just what I guess re-aggressing against someone who's being too aggressive. It can seem like you're letting someone just run the table by like being too bad. Okay, here's what it is. I just got to what the problem is. So you have a table that's very passive and you have one guy who's aggressing too much. or who's addressing a lot as a response to a passive table, to a very passive table. What he's doing seems correct, right? Given the fact that the table's too passive, what I do seems confusing to me. I'm in this environment. Why am I going to table that's too passive? And if a guy who's being very aggressive, and let's say we're both about equal and chips, and we're both close to, I don't know, we're both, and the whole table's equal and chips. This will make this really easy. Hold tables about equal and chips. Everybody's too passive. One guy's being aggressive, trying to take advantage of the equal of the passive table. What should I do? Now that this is happening, right? Because both feel like mistakes, right? To try to also get aggressive feels like a mistake, but to encourage the environment that's allowing this guy's decisions to be profitable, like to take apart in that environment, maybe it's somewhere in between, but I feel like I could come to both answers and it feel correct. Yeah, I mean both answers are correct because it's not seamless. Because yeah, it's kind of like a game of chicken. So the issue that happens is if you were to set if you set a soul, right? Yeah, this player, it's very aggressive. And then you were to not lock yourself. So I don't lock you to be aggressive. And I let him know about that and he's aware of it and he should switch passive. Because if he's also aggressive, then you're both losing, which is perhaps different. So now he has to be passive. If he's more aggressive than you and now you're aware of that, then you should be passive. So it's either you're both aggressive and then you lose or basically one has to see or you both end up losing. So the real environment though, this is the interesting thing because you're basically at an unstable point, right? So do you just like, I guess it depends on your opponent, but you tell me this information, my feels like you just go the suicide strat. And if you truly have no concern for yourself against any rational opponent, like I feel like if I just look at my guy, like the guy just, hey, by the way, I'm actually never going to fold to you. And just like say something like that at the table. Yeah, exactly. You know, and just like, and just like mean it. Yeah, I actually don't care about what happens in this tournament. I'm not going to let you run this, run this table. So you know, take, take heed, but it feels like to me this, I mean, I guess the real question is, if you both, how big of egos you're dealing with, you know, if both people choose to choose the nuclear option, the whole table just gets to sit back and enjoy this, yeah, explosion. Yeah, yeah, that makes sense. But yeah, it's it's tricky because if you are to take this route and you only need one player to get this start, like it's a tripping you a lot and then you're losing. So you kind of need the whole table to collaborate on you running them over, which happens. Yeah, but it's fair to do that, you know, let's say you release pre-flop ICM, and then you have this 6GTO buff that converge to whatever they converge to. And then you set one player to be a maniac, then he will start making a lot of, he will make more of it in GTO and all the other players would end up losing. So, you know, even though he's playing against 5GTO, but they just, we have to concede that, you know, oh, if I start racing against this guy, then I lose. So I have to be passive and that's true for all the other players. So, yeah, I mean, this really speaks to the non-zero summoness of the multi-way problem, right? Because, yeah, because, like, I am not enough to basically punish this guy necessarily, like just me deciding, I could get crazy enough, but not without severely damaging myself, right? But if I could get everybody at the table to equally disagree to punish themselves, it actually might push the one guy incentivize to become passive, you know, like, the whole table might be able to, you know, collaborate, but that's, again, that's against the rules. But it's just very interesting to me to see these spots where someone is doing something that is causing them to win more money. And everyone, as long as they are concerned, this is kind of like a prisoner's dilemma in a sense. Like, as long as they're concerned with their own EV, the kind of just have to let it happen. Yeah, so, I mean, this is kind of this, like, textbook game theory, you know, they tell you never try to solve multi-way games. Lindsay was some game because all these weird things happen. And somehow multi-way catch game is, it doesn't happen much. But then, yeah, when you switch to ICM and I don't know about what game in these areas, but it's possible this happens there too, but yeah, I should just very weird behaviors where I don't think it's actually useful or, you know, debatable, but it's useful to look at the solve like ICM in isolation. I think we need to have this interaction where you can give some assumptions about how the player plays. Easy aware that I, you know, he can give assumptions about yourself and say, and assume whether he's aware of that or not, but now this, just building an intuition for it, because if you look at ICM solve with 60-teal bars, you just give them out, let's say it's like somewhere, you're at a multi-table tournament, right? And there's like you're on the bubble. And then you have, you basically run a solve with like normal solve with the raises and three bands and stuff like that. And then you run another solve and you just allow everyone to just limp and check, check, check. And so like no bets anywhere, then all the table is making more from it. So they should just all do that, right? So, you, yeah, you cannot really rely on these equilibrium that is solved with conversion, because based on the actions you give it, they will just, in some cases, try to, like, solve collude. That's basically what makes the most sense. This is what has been hard whenever I coach people in ICM is because it's way more difficult to try to get, like, you kind of need these like dials of variables in your head that are like pulling things in directions. And I've, my experience as just like my intuition is I could what you're saying, because when I look at like an ICM spot for someone who doesn't is not well studied in ICM, what always happens is 15 different notebooks that follow, you know, to be like, well, what, what, how bad is this? What is the thing that's actually making this good? Is it the fact that this person's passive or is it the fact that this person's aggressive, like, and trying to find the thing that makes one decision? Like if we have a hand like eights, that seems like in different between like calling folding and jamming or something, you know, that's like at a certain spot, let's say button, your final table, like a button open, I don't know, I'm just imagining it since you, you're some in different hand to try to find the things that like are making this kind of a difference of marriage, which almost certainly doesn't actually exist in a real spot, right? There's like, there's probably just a best decision, but it's been always fascinating for me to try to like run down the different strategies and like change different things to try to make a single action be dominant for that hand and see what those things are. And this, I feel like this is almost something like what you would need, what would be super valuable to be able to have in any kind of ICM preflop solver is this where you can have dials of specific things being moved and you're watching it get solved. You're not just getting an output, right? But it's like, hey, can I turn up the aggression of the table of everyone who hasn't opened? Can I, can I increase the folding frequency of the guy who did open? Can I, you know, can I make the flatter be like way too aggressive or way too passive or like call too much or call too little, et cetera, et cetera? Because in any given spot, there's like, it is my intuition or my experience has been like, when someone brings me an ICM hand, I, with almost any decision, I could imagine the environment in which their decision is fine, right? Where they're like, I do this. And it's like, well, I can, I need so much information and we need, we need to be able to say so much before we can decide what is like just terrible here. You know, those, those do exist. But again, just, just more, more, more feature ideas will be really cool with ICM solutions to be able to have your hands on dials more and watch things change. I think then probably just getting direct outputs. But I say that as me, most people just want an output. Yeah, but I mean, I do think because of that, it's kind of well, yeah, I'm hopeful that it's having this kind of prefab solver where you care about it. So like, different profiles have been handled like and get response very quickly. It's much more, more valuable in this kind of environment, because there's not really, it's not like you study a prefab chart and you should play this. I mean, of course, there's some things like what hands are good enough to jam and stuff like that, like, preferable parts, these kind of things that, yeah. or probably always somewhat fine, but. - You probably need table profiles. We'll probably be a thing to be helpful with ICM, right? Just in the same way you have user profiles to have like table profiles. Like if you have two people that you know we're gonna be in the hand and you can set like the profile for the table based on frequencies, that'd be really cool, or like a field profile. That'd be cool. But even still, I still think profiles would it opens the flaw too. I guess this is more for people using profiles than it is for you guys. But when it opens the flaws who is naturally anytime you label something with the word you're familiar with, of just skipping all of the assumptions underneath that profile that are like leading to that and being like, oh, I've looked at a table that is this or whatever, or like I've looked at players who are this title and I classify that guy's the same title. So now they're the same in my mind. But there's lots of, yeah. I think you'd be, I just wish the things whatever it is that's determining a player profile for like the preflops solves. It'd be really cool if you could just like see those dials while looking at the preflops solutions. You know what I mean? Like whatever dictates a profile, seeing those and being able to change them uniquely if it's not just a single thing and like watch how it results. 'Cause I think ICM in particular would be, it'd be pretty cool to be able to just like watch things change a lot. But yeah, and you were like with this idea where you would have this, to define a third profile, you will like set some statistic about this player. Like I mean, in the ICM, I don't know what would be the, because I don't have any tournaments, but I think actually it would be like PPA, PPA4 and everything and you could kind of build this profile and it would define the closest human match to this. Like, but yeah, I think yeah, also I hope that one day to improve the ICM model and it could, if some things aren't a little bit weird, but maybe it's just how things are. - Could you, 'cause how good is the current ICM model that most people are using? - I think it's from what I saw, it's quite good. But it's the same also, maybe issue that, that's issue, but that happens in caching, where because poker players study so much solvers, and at some point if they play like ICM, then this more likely that the model becomes true. So, you know, if everyone plays according to the model, then the model is more likely to be true in practice, kind of thing. - Yeah. Could you, could you have a kind of self-play AI that plays tournaments, tournaments in self-play? And you get a kind of network, neural network prediction for like, you know, for modeling, the, what's the word I'm looking for? The value of chips, you know, basically the ICM, the independent chip model, the models it, but that is based on like a self-play model. Is that too computationally? The thing I'm describing, too computationally crazy, or like, are there other technical limitations to doing that or? - I mean, we have some ideas how to do that, where it would be feasible. Like, yeah, basically the idea would be that we first have a prefab, like model, yes, over that is, use the ICM model as a baseline, kind of as a warm start, and then make it play from there, and then improve it, and try to improve our data. I mean, there's some like obvious things, you know, people use future game simulations, but it's like the blind or adult to change, and stuff like that, which future game simulation probably is not the best way to do this. So you want, you would rather have some learning component instead, and, yeah, I guess it's, it's no one knows if it can be improved a lot, but I think for, for bounties, maybe it's more like things, the model is quite simple, but, I don't know, I've not seen someone like track and gets the tournament data and analyze our activities models or so. - If you're tired of trying many of the same things time and time again, and not reaching the level that you know you're capable of reaching, my one-on-one coaching is definitely for you. I've worked with a ton of players, many of them great, many of them just beginning, and many of them have put a ton of work into a single thing or a handful of things, being sure it's going to be the thing that causes them success without realizing that there are many subtler, less obvious things they need to work on in order to reach the levels that they can. So if you've been grinding your gear, spending time and solvers, putting in hours, reviewing hands, and you still don't feel like you're reaching your potential, reach out to me, we'll have an intro call, and maybe one-on-one coaching is right for you. - I've heard from some very great players that they have some very specific, everyone has their kind of very specific predictions on how they think they're wrong. What are the more interesting ones I've heard, which I'm not gonna cite the source for, but is that medium stacks, they think ICM models predict medium stacks get in the money too often, so that shorter stacks tend to knit up and make the money, like not be aggressive enough, and so make the money more often, and that EV is unlikely to be coming from larger stacks, and it's probably coming from the middle stacks, and so there's some, I'd be curious. I mean, that's not a problem with the model, though it's a problem with, well, I guess it depends on how you define the model, but that's what they predict is actually happening in tournaments, is that middle stacks are not making the money as often. So the model, middle stacks are valued, overvalued, I could use the ICM model. - Mm-hmm, overvalued. That's like one, that's what I've heard, a very great player say, but I think every, all the big tournament players have their own suspicions or ideas about what is not true, or like what is misrepresented in the ICM model, but that one I find fairly persuasive. I find the idea that short stacks make the money a little more than they should according to ICM, like that short stacks, people with short stacks are not as willing as an ICM or just like a bot would be to just put their chips in, and so they knit up a bit. And so they, and then the result of that is, is middleing stacks making the money less often than they should. And so that they should protect their chips a little bit more, like middle stacks should be a little bit more protective with their chips. So you can probably find the players who believe this by watching like tritons and just waiting for a player to make a surprising fold with a middleing stack. I think there's a few, there's a few clips you can watch where this is kind of what stumbled upon this. There's like a, there's some players you can watch just like make surprisingly tight folds. That's strike me a surprising at least. And the, the, that was the explanation I was given for it. So it's interesting, yeah, I think it's kind of this more undiscovered territory and even what might be true about if you make this like, you know, bots play against each other and learn what's the actual function from chips to dollar EV. If you know, you are, you are to play optimally, which again, it's not so clear what that means, actually, anyway, in a tournament. But that might actually be very different than what happens in human data. You know, so if we track right now, I mean, there should be a lot of data on it. It's about like tournaments and stuff about how good ICM model is actually predicting what happens. And then you have a decent idea of the, I feel like this should be data that GTO wizard can get, right? They, they have partnerships with many, many large tournament, like online tournament operators. I feel like it'd be pretty easy to just scrape and accumulate, you know, and just like aggregate the data, even if it was anonymous data, you know, but you could just predict like how often are people going into the FT with the biggest stack or like a runaway chip lead, you know, like, yeah, I think it could be, it could be, it could be done, but I think it's more something like that's useful for like coaches and stuff, because you know, even if we had some estimation of how different data is than what ICM model predicts, then that's still not so good. I would want to use that because it's just how humans play in them. Well, no, here's how you're going to use it in the same way. In the way we just talked about with this having profiles, and we said we should have field profiles. You could literally have a field profile that is like ACR. You know, that is, you know, you could, but that might be, people might not like that and think that's like somewhat unethical, but like you just have a profile on a field that's just like, this is how the ACR field plays. And so like you can, that's the ICM model you're using when looking at this, you know? And yeah, I don't know. I think that sounds interesting. Yeah, I mean, if it was shown that it was like a different, you know, we were, this would actually impact results, but I don't know where that's true. I think in general, the model is probably very good. But yeah, there's some spots like you so you think like if they did that if they script like if they took like all the tournaments running on like GG or something and compared them to what an ICM model predict you would think if they would they would come close. Yeah, I mean, so that's what some people said, but I haven't done, I've never done it myself. Gotcha. Yeah, it's very hot. It's very scientific and humble of you. Nothing, nothing, nothing gives me a rush more like having an opinion on something I haven't done the research for myself. I'm just kidding. Well, cool, man, I think we covered a lot. Is there anything else that we'd be like excited to hear about the things you guys are working on or like do you guys have any other exciting things in the works. I think this is the stuff that will get people really excited. Yeah, I mean, that's the main thing, but we're also working on pillow. But yeah, this is not, I mean, I guess it's more of an engineering challenge than, you know, it's just old and with four cards. So it's not, there's not really that many scientific research. If you can do it in old and you can do it in Omaha, basically, but it's just about scaling it's killing you know. Yeah, yeah, it's just like, hey, I do, do you like do some kind of abstraction with the combos or like how it like that kind of stuff I imagine. Or because I imagine you can't actually run every combo in PLO. Yeah, I mean, we will try to use it as. As little as possible on, right, so I think I mean our our our hope is that it will be much more accurate and something like munkers over. But I guess we'll see it's also it's kind of hard to bench smart because we have to make it play against munkers somehow, but. But yeah, I think the techniques that we use basically are always looking forward to the point where we get to multiple and PLO this kind of stuff because that's where. These techniques should really shine because you know something like pilot over both luck, you can use exact sizes and you can get close to Nash and compute exploitability and stuff, but. In multi way PLO and stuff, it's just. You know, the games are just too big so then people use attractions, which is kind of a poor way, I think to. It's just the size of the game and you would rather have something like you learning to do some learning component instead. So I got to get an inside joke in here real quick, but how many GTO breakthroughs did it take to get to get preflop to get preflop ready. I think it was eight. Yeah, you had one more know you'd one less it should be because we were you that would make it seem worse right if they said eight we should you should take. Seven you've you've you've had 13 breakthroughs, but it only took seven to get preflop good. It took one less breakthrough than them anyway. Well, awesome, man, thank you so much for doing this. I'm concerned that this was going to be too deep a discussion for most people, but I know you know, well, I guess we'll just wait and see. But it's it's quite likely, but I was trying my best to say it and the. Lehman's terms, but it's kind of difficult. I mean, there are quite technical things, but yeah, I was a lot of fun things. Yeah, yeah, it was great. Yeah, we got to keep having you back. I mean, it's cool. I think it's cool for people to get an eye, you know, like a look behind the scenes of like the stuff that it takes to do this. I think it could potentially help people understand all that it does take. I think it's very easy for something like three way multi way to come out and be like, man, this isn't exactly how I want it to be or something. But to realize all the all the work and all of the difficulty that goes into it. I remember when we were doing stuff with ruse and we'd have like feature ideas. And then, you know, I'd fly out to Montreal to visit you guys or something and then you guys would look quite burnt putting and putting it all the time that you guys are. And I would start feeling bad just for being like. Yeah, great idea. Go, go, go. Yeah, yeah. Yeah, I always have a lot of ideas, but yeah, hopefully people can get a sense for how much work it does take and how, you know, how much. Because it is, it is, it takes a lot and I think, you know, you've always been a very impressive person. So I'm very happy to have you here and explain some of some of the some of the magic. You're the dark wizard of poker AI. So I don't know. Look at you. Yeah, I think it's, it's, it's surprisingly, surprisingly takes like a lot of work when I was starting. I thought this should be super easy, but it's, yeah, it's, it's quite a lot. I mean to, you know, I think also for a good reason, but poker players have high standards, you know, when we released to play a solver add to match, you know, the exact solutions are close to them. But with that's a bit complicated different techniques, but yeah, I think it's super fun and if you can help some people understand. A little bit better solvers, that's good. Well, is there anything, most people here plug things or things, but you, you, I don't know, you're not really, you don't really have like a social media following or maybe you should. Yeah, we said I would create some blog about it, extending some of these things, because sometimes a, but like more technical than this podcast, but it comes up with some insight working on this. I think very interesting, but never write them down. Nice. Well, great, man. Well, yeah, thanks again. And is there anything you want to plug or any place you want to point people to define more from you? You should write a book sometime or something and we'll just throw that out there. I don't have anything. Just keep watching the great service podcast. Alright, thanks so much, man. Until next time. Yeah, see you.

Podcast Summary

Key Points:

  1. Philip Beardsell transitioned from product development at Ruse AI to a focused research role at GTO Wizard AI, enabling deeper technical exploration of poker solvers.
  2. A key innovation is the use of Quantum Response Equilibrium (QRE) in multi-way poker solving, which prevents zero-frequency actions and provides cleaner, more realistic strategies by modeling how players respond to rare actions.
  3. This approach improves solver convergence, reveals exploitative opportunities, and better reflects how real players behave by encouraging exploration of untested actions, challenging the notion that only "safe" strategies are viable.

Summary:

Philip Beardsell, lead researcher at GTO Wizard AI, shares insights into the evolution of poker-solving technology, transitioning from product-focused work to dedicated research. His current work centers on Quantum Response Equilibrium (QRE), a method that prevents actions from being assigned zero frequency in multi-way poker scenarios. Unlike traditional solvers that stop updating when an action is rare, QRE ensures that all actions—no matter how infrequent—are continuously evaluated, leading to more realistic and exploitable strategies.

This approach not only improves convergence speed and solution clarity but also reveals how players respond to under-explored actions, such as geometric bets or small-size raises, which are often ignored due to perceived risk. Beardsell highlights that human players are often constrained by social signaling—fearing to appear incompetent—leading them to avoid innovative strategies. The QRE model counteracts this by simulating a more dynamic, exploratory decision-making process, akin to reinforcement learning.

This shift enables solvers to better model real-world poker behavior, where even infrequent actions can have significant strategic impact. The work also underscores that multi-way poker environments are inherently complex, with no stable Nash equilibria, meaning solvers must adapt to evolving player dynamics. In such settings, a player’s mistake doesn’t always directly benefit another—it may instead redistribute value across the table.

Beardsell suggests that future developments could include human-like player profiles, where strategies account for real-world behavior, and educational tools that visually demonstrate how solvers evolve over time. These tools would help players understand that optimal strategy is not static but a dynamic adaptation process, emphasizing exploration and learning over rigid adherence to theoretical solutions.

FAQs

Previously, Philip focused on product development and business aspects, spending time on legal and accounting tasks. Now, he focuses exclusively on research and development, allowing deeper exploration of AI and poker-solving techniques.

The feature solves for actions with near-zero frequency, showing optimal responses even when a line is rarely taken. This reveals cleaner, more accurate strategies and helps players understand the true exploitability of their decisions.

QRE ensures that no action has exactly zero probability by distributing strategy across actions based on expected value. This leads to more realistic, well-behaved solutions that better reflect actual player behavior.

Instead of ignoring actions with zero frequency, the solver maintains small probabilities based on expected value, allowing the strategy to converge more accurately and revealing true exploitability.

Multi-way games have infinite Nash equilibria, making it difficult to find a stable, universal solution. The solver must make approximations and trade-offs due to computational constraints and the complexity of player interactions.

Profiles allow users to define how a human player might act, blending GTO strategies with human tendencies. This enables players to exploit opponents by modeling realistic, non-optimal behaviors.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.