Simulating Humans at Scale: Simile's Joon Sung Park
38m 45s
Simley, founded by June and her Stanford colleagues, is pioneering a new class of applied AI that simulates human society through generative agents. Inspired by early experiments at Stanford—like Smallville, where 25 agents exhibited realistic social behaviors such as relationships and events—Simley developed a platform that combines real-world behavioral data (collected via interviews and surveys) with large language models to create accurate, scalable simulations of human populations. These simulations predict real-world behaviors with up to 85% accuracy, offering companies like CVS a way to test product concepts, market messaging, and even earnings calls without relying on limited real-world experiments. The platform addresses critical gaps in existing research by capturing long-tail personal data—such as life stories—that reveal deeper values and preferences. Simley’s work is validated through rigorous evaluation, distinguishing between convergent (predictable) and divergent (variable) outcomes. Beyond commercial use cases, the team envisions transformative applications in economics, climate policy, and political forecasting, where simulations could model systemic risks and collective behaviors. They believe simulation represents a new scientific frontier—like telescopes for astronomy—capable of fundamentally advancing our understanding of human society and decision-making. As models evolve to better represent human irrationality and diversity, Simley aims to build not just a tool, but a foundational framework for understanding and improving the social world.
I am somebody who is quite inspired by science fiction. And when you read science fiction that covers societies that have progressed far enough in its technological maturity, you always see two pillars. You have some version of AGI. And you have some version of simulations that will help guide the society. I do see an opportunity today to really take the first crack at building the simulation. I would not have sent that even five years ago. But that is a conviction that we have built up over the years as we are going to deepen to this research. [MUSIC PLAYING] Today we're delighted to have June, founder and CEO of Simley. Simley is building an applied AI lab, simulating human behavior and societies. And I'm very excited to have you here to discuss what you're building. Same here, thank you for having me. OK, take me back to April 2023, Stanford, California, specifically Smallville, Stanford, California. What was that? So Smallville was a project that we were running at Stanford where the idea was that we made this observation that large-linked models can now encode a lot of human behavior that is embedded in its training data from the web, and social media, and so forth. That if you sort of probe at the right angle, you can actually get a lot of micro behaviors out of these models. So given a very specific demonstration or description of a situation, what would person X do? And it would actually generate really interesting behaviors. We found that to be so interesting. And we found that to be the ingredient that we had been waiting for for creating really complex agentic behaviors. So Smallville actually was an experiment where we decided that if we push this as far as possible, what would a society that is created by these agents look like? So we basically created generative agents that is paired with generative AI model with memory, planning, and reflection to basically create this lived experience of agents living in the small town. So Smallville was basically a game town of 25 agents living in it. Individual agents had a description of persona, but they would actually wake up in the morning, do their routines go to work. They actually have relationship, sort of like people would, and they would actually have immersion phenomena, like having parties and so forth. So that was the experiment that we ran. - What was the most surprising things to come of the experiment? - So one of the surprising things was, so the experiment, the simulation itself actually sets place the day before a Valentine's Day. So you actually see these agents, one of the agents actually thinking, well, I run a cafe, so she's a cafe owner. Her name is Isabella. She goes and thinks, it would be great if I can do a Valentine's Day party where we invite a lot of friends, customers. So you actually see her on the day before Valentine's Day going around, actually gathering materials for the party, actually telling our customers, hey, we're going to have this party, please come. And on the day of Valentine's, you actually see this immersion party that actually gets formed with all these agents coming to the base group of friends. - Did anyone not get invited? - Well, some of the people did get invite invitation, but they forgot. That's one thing that did happen. Some of the agents did not explicitly get invited, but we had one agent who got to invite Klaus, who decided to ask his crush out on a date. So he would actually bring in the date, they would actually have a party at this cafe. So quite surreal. - So how did you end up building this model in the first place? Were you studying human psychology and social behavior, or was this coming from the customer back, or was it coming from the technology out? - So my particular team has been excited about simulations, and we saw the vision of simulation failure early on. So my career as a researcher at Stanford really started back in 2020. That was the year when GPT-3 was about to come out. It wasn't quite there yet, but it was just about to come out. We started to get its first demos. And my first year, we wrote this paper called Opportunities and Risks of Foundation Model, alongside many of the Stanford researchers, and was led by one of my co-founders, Leon, who is now the head of the Center for Foundation Model at Stanford. And when we were writing that, the part that I was really focused on was, well, here's a new class of models that we have not seen in the past, that these models that can be very generalizable in ways we didn't quite have in the past. And I got into thinking, well, if we can imagine the kind of interaction we can create with these models, what would that be? And many of my colleagues back then were surprised that these agents or these models can do classification or simple generation. And that was really incredible to see because these models didn't really know or didn't really, wasn't really taught to do that. But the part that was surprising to me, wasn't that these models can do that, because from interaction perspective, we've known how to do this for a long time. The interesting part was, well, these models can actually include human behavior. What does that mean? If we were to push this as far as possible? So part of the tradition I come from research included what we call social computing. And social computing within human comprehension really has to do with this idea of how can we build a better technical or technological platform that would enable social interactions and collaboration? One of the most difficult challenges of building a social platform is not necessarily testing the UI/UX of the system, but it's more about when you have tens of people, millions of people, and down the line, billions of people, how do all these people come together to create the immersion phenomenon that's both good and bad, and how can we design for a scale? And so far, we didn't really have a tool that would enable us to test for that. The only way we tested today is you basically fill tested. You release your prototype, see what happens, and sometimes it actually comes at a real cost. Obviously it's high cost in terms of human hours and the time it takes, but at the same time, if you have a bad design, imagine you have a feat on social media that is more likely to propagate a certain emotion that is negative, then obviously that is something that we want to avoid. But this now gets tested in the field. So we wanted to see whether we can actually create a simulation that would actually let you test for this. So 2022, this was actually a year before the gender of the Belgians. We worked on a paper called social simulator lacra, which actually really was the precursor to the agent paper that we ended up writing. The core thesis was, imagine you're building a subreddit, you're a designer on a subreddit, you want to see what people might do in the subreddit. Which is surprisingly hard task, even for practice designers. And we basically decided, hey, we have this model, seems unique, let's use this model to create simulations of the entire subreddit. So you define the goal, you define the moderation strategies, and you populated with thousands of what back then we didn't call them agents, but we call them personas, but populated with thousands of personas. This is basically 2022 version of MoeBook, which is quite interesting that it actually came back. And when we saw that, we actually got a lot of really important insights out of this. What are the good behaviors? We actually simulated a community where the entire idea was for people to discuss with each other the site's places to site see in Pittsburgh. And all of a sudden, you start to see this persona is actually a collaborate to actually discuss, hey, XYZ places are amazing, you don't actually go to a trip together and actually plan those trips live in the simulator subreddit. So that's how we got excited. So we saw the vision and the excitement and the potential applications failure early on, but then the work that we had to do was then demonstrating how can we go beyond simple personas to create complex agents that actually can think over time because we want to simulate the longitudinal aspect of our society and then actually validating that these simulations are actually accurate in practice. - Was there a point of model evolution at which you felt like, okay, we're there. The models are good enough for us to actually have a faithful representation of human society. - So, GPT-3, when it came out and social simulator was built with GPT-3 and it was very janky. It didn't do any instruction tuning, it did not follow your instructions. So just to have it to listen to you and do what you wanted to do, you had to do some weird tricks with prompting and so forth. Well, you could actually see the promise. The model actually have encoded a lot of human behavior and you could actually see the trajectory. And when we had the generative agents paper, it wasn't quite tragedy, but we now had instruction tuning. So we could actually build much more complex agents that can reason about its memory. That wasn't really possible when we did social simulator craft. And since then, of course, the models have improved. So where we are today is the models at its foundational level have reached the point where we can actually imagine building these kind of applications. Now, the part that actually I do think however, that's quite interesting here. Today, if you look at many of the large-length model companies, whether it's OpenAI, Anthropic, and many of the new labs that are getting formed, the models they are creating are models that I would consider to be their north star to be something that is similar to, let's build a super intelligent machines. These machines are meant to be rational. And these machines are supposed to be really amazing at technique problems that have an objective answer. So maybe that's not even the best simulation of true human society then. Turns out, people are irrational. We have a lot of subjective values, preferences, and taste. So you actually start to see divergence in model size going up and the performance in its ability to predict and simulate human behavior. So we have sort of plateaued with current modeling paradigm, our ability to really simulate humans. So it is sort of at the starting good foundational level. But to make it really amazing, We do need the next frontier that is more key.
your towards actually modeling people's diversity. - Very interesting. At what point did you realize that, what you did with Smallville could become a company? - Right. So again, the promise I application was something that I was very much inspired by early on by simulation with social simulation and so forth. But the part that I realized over time is research and a company have very different function. Research is an amazing vehicle if you want to basically do breath research. You are in a lab surrounded by really smart set of people and each of the researchers own a small piece of thesis and they go explore some of those thesis blossoming to amazing research product. But we're not necessarily known for finishing our job. We're not usually the one to bring that research impact to the real world. Company is a machine for depth research. You have a conviction on an area. You find a hill that you want to climb. This is the vehicle that led you put together resources and an amazing group of people to go after a singular vision without hesitation. And we got that conviction. I would say about half year after Generative Agents. After the original Generative Agents paper, we got so much inbound interest initially from actually social scientists who wanted to run their experiments and all the RCTs on our platform and then very soon after, many of the Fortune 500 companies who saw this demo and their board members and CEOs who sometimes visit Stanford saw that and they started asking, well, we go run all these surveys and experiments and there's so many research questions about the market that we cannot answer today. Can we run that in simulation? That started to really intrigue me because that showed a clear line towards a real world impact for research, which is not always the case that we have that kind of opportunity. So that is when we decided we actually want to validate the simulations or accurate. So we went out and actually created simulations of 1,000 people of the US population. We demonstrated that using our architecture and the models, we can actually predict people's behaviors 85% as accurately as people replicate their own. When we saw that, we thought, okay, this is something that we feel comfortable providing to our users as a platform for simulating their really important decisions. So that's when the co-founders, myself, Percy, as well as Michael Bernstein, was a researcher and my advisor at Stanford. Both of them were actually my advisors. So the three of us have been working together for five years and now at this point in six years. But that's when we got together to have the initial conversation of, can this be a company? - Come on, amazing. Maybe walk me through a customer engagement event today, like who's a canonical customer in which department? And they come to you, what are they asking you? And what product or service do you deliver to them? - Right, so maybe an example that I can give to make this concrete. So CVS has been partnering with similarly for the past, I would say, nearly half a year. And they've been an amazing partner. The way we initially got in touch with, so our main buyer at CVS is the lead, it's a senior VP who leads human insights. And the original story there was he was, he basically read my paper, that validated the age and simulations and thought we have to bring this to CVS. Because today we are bottlenecked by the number of questions we can fill test. And we are also bottlenecked by truly the physics of human society. It's one thing to ask surveys and experiments, totally different thing. If down the line you actually want to simulate the entire market and actually met that or the second order impact of the decisions you suggest to your leadership. So he's been looking around for that solution and his cousin happened to know me and basically told our buyer's tree that the authors of the paper are actually looking to start something. So that's how we got connected. And in this particular engagement, usually the way this goes is our customers are very much used to working with pulling companies or panel companies today. And there they go and basically ask these companies, XYZ are the populations that we're interested in better understanding. Can we go run a research study of these topics? That initial stage looks very similar for Simile. So our buyers come and they tell us we want to better understand XYZ population. Then Simile goes out and we have through our partnership with vendors, we have a strategic partnership now with Gallup, for instance, who is a pulling and panel company. Where we go out, work with our vendors to actually reach out to real humans. So these simulations are grounded in real data but reach out to those people. Collect data that we believe are efficient and generalizable about that person. So imagine you have 15 minutes. What are the magical questions you can answer or you can ask these people during that time? We collected data, use the data to create agents or simulations of these people that can basically be used to answer a large number of questions that goes way beyond the original domain. We load that entire platform and it's basically a SaaS product. Our customers come and they can basically ask any questions about the group of people of their interest. - So interesting. It reminds me of autonomous vehicles. You know, you go and collect a bunch of data from the road and then you're able to augment it with simulation. Is this similar concept or are there big differences to what you're doing here? - It is similar concept in the sense that of course you, with the self-driving vehicles, you want to create model that is based on real world physics but you want to create a model that is generalizable beyond your training data. It needs to be generalizable into different locations with different weather conditions. Very similar concept where what we want to create is we want to reach out to real people and for these people want to understand something fundamental about these people in a way that we can uncode into the model. - I would have thought that the large language models would be such a good representation of the whole world that you could almost narrow it down. You could tell Claude, you are a 34 year old woman living in a bicostle metropolitan area and it would be able to have a faithful representation. So I'm actually surprised that you go out to Gallup. Maybe can you just explain why you have to go out and collect any real world data at all? - Yeah, one of the big questions here is the question around say do gap. There are things that people say and then there are things that people actually do. And the gap, there is real. And a lot of the large language models are trained on additional data. Fundamentally, it is the things that people have said online. That does cover a large quantity of its training data. So one of the things that Simulation Platform does is actually closing the gap. So a lot of the data that we end up collecting by nature or behavioral, it also includes data that actually goes into literally questions like just tell me the story of your life. Turns out if you understand the person's story of your life, the kind of data you get from it is what we consider to be the long tailed information about this person. It's not about what you've done in this particular moment. It's not about very broad questions like what's your view on politics. It's about where you grew up. What's the difficult decisions you have to make in life? And what's interesting about this data is it's an amazing way to build a translational layer between attitudes and behavior. So we combine these kind of data sets but fundamentally that's the gap that we want to close. - What's what the behavioral data do you have? - So it similarly does run a lot of experiments. So kind of models that we have trained, for instance, we have a huge repo of RCTs. So randomized control trials that were run in social-centric context that were run around pricing studies. So one of the models that we are training is basically the foundation model of human behavior in quite a little sense. We have all the behavior of signal from RCTs. Can we actually encode that into the model so that the end outcome is a model that can basically predict the results of any RCTs? At the same time, one of the conversations that we keep on having on our customers that we're very excited by is our customers then come in, see that potential and their mind goes to, wow, we have 90 million customers, let's say, here at CVS. How can we leverage this kind of data to create better simulations? So there's also a conversation around how can we, in a responsible and ethical way, leverage existing data that is also in house for our customers, then use that to create augmented version of similar model. So that, of course, is going to be more fine tuned specific to the population of these customers. But that's the kind of data that we would be leveraging. - I see. And are you doing these interviews typically by voice? Is it a survey that you fill out, what's the modality? - So it's a huge breath. The quick answer here is it's both. Interviews are fantastic if you want to get the long tail information about people. So we actually do, in the original study that I conducted back in 2024, we'd literally ask questions, tell me the story of your life. Now, the way we do it is we are training our own model. So it's a reinforcement learning loop, but basically imagine the objective function here is how can you spend the minimum amount of time to get the maximum amount of visibility about this person? So that is one of the things that we do. So basically training on interviewer that is not really asking for factual information or an experience about a particular platform, but just what are the life story that people have?
path that can be used to train our own model for these agents and then for the more factual or sort of more discrete choices choice questions surveys and so forth. These are also very efficient. These are a time and data efficient because people can fill out many of the questions in short period of time. So for those we actually do leverage them in for instance if you want to just have a broad understanding of people's viewpoints of certain topics, certain policies and things like that. You describe yourself as an applied AI lab. How do you think about where you want to build your own models versus where you want to rely on other existing models? So in terms of building our own model, where the core thesis here is there is an amazing model to be built that really encodes the diversity of people's values, preferences and taste in ways that simply a rational model can have to. The one way I actually post this we're sort of building, say imagine the current today's model are akin to the CPU of intelligence unit. It's a single model trying to an amazingly rational data that is amazing and solving very complex objective questions. Similarly small is much more akin to developing something that is close closer to the GPU of the intelligence unit where the idea here is we don't actually need a model that is super human. Similarly, in fact we want model that's as human as possible, but we want to make sure that these models at the sort of individual subunits can represent the real viewpoints of different self-applications. So where we see that gap, that's when we go develop our own model. But at the same time, we do leverage for frontier models, for instance, as a way to coordinate the research. Frontier models are amazing at coming up with a research plan. So that's where those models actually do get leveraged. Very interesting. Are people typically coming to you with questions around new product launches, how they should be marketing their companies, pricing, all of the above? So it is all of the above. Our customer journey usually does however start with very concrete use cases and problems they are trying to solve. Concept testing is a big one. It's also a very straightforward one. So they have a new concept, new product idea, new market message they want to test. And they want to hear from their users what they would think about XYZ. This is one way for them to quickly test those ideas. And then the promise they quickly see is well right now we are very much in the practice of testing five to ten different ideas and what does it look like for us to test instantly thousand different ideas across thousand different self-applications. That's the initial vision they see. And then we really get into the need degree details of well where the simulation go from here. They then pretty soon start asking well can this be used to do product testing, but not just simply submitting like an image, but imagine basically asking these agents go experience this product for ten minutes and tell us about what you experienced, what you saw. So you're basically adding temporal dimension. Then you go into things like multi agent simulation. You saw that customers very routinely actually ask us to simulate their earnings call. This is actually a use case that both surprised me at first, but this is also surprising a common ask because of course the CEOs and board members always need to think about how we're going to design our earnings call, how would the audience react. So that is something that we also do and this is very much multi agent simulation. You know it seems like there's so many use cases that could potentially be tested once you have like a simulated almost customer population, right? I'm curious the value of research and testing and sim versus just like, let's say you have a new product concept that you want to test. Why not just go run a thousand Facebook ads and like you actually get the click through rates on this stuff? Isn't that real world data almost more useful than the simulated data on how people might behave that you then correct for with your own models? So it's a great question. Now I think to some extent here the answer has to deal with initially scale and then down the line, truly the new capability that comes because you can simulate the interactions. The scale question here is actually quite straightforward where yes, you can actually run Facebook ads and Facebook testing, but the kind of experiments that you can run in simulation is actually behavior simulation at scale, right? So you can basically put in any number of users doesn't even have to be bounded by the number of population that's available on Facebook. And it's also much more representative because only certain groups of people will actually respond to the online experiments, but similarly, the model that we are creating one of the key promises is that it is representative. We do the hard work of actually getting the representative set of people and then collecting the data that would actually represent them properly. So the scale representative is something that many of our users do not have easy access to. We want to come and ask also whether we do get or come and sort of pain points that we have heard, where the question that many of these people have isn't about like what questions do we ask these people, but it's about in the first place, how can we get to the population that we're excited to talk to you? That's a huge bottleneck. Then down the line, you can actually really certainly imagine and this is something that are customers and some of the most for looking customers are now going into, which is what are all the downstream implication of the decisions that you make. It's not just about whether, imagine you have this particular product, will do you like it or do you not like it? Would you pay for this, not pay for this? It's not necessarily just that initial questions that we want to answer and finish, but we want to understand, imagine you're a car company, you launched an electric vehicle in this market. Maybe the electric vehicle does really, really well. So we can help you do concept testing around marketing and the product around the electric vehicle. But what does that do to the perception of, let's say, non-electric vehicle, does it change the market perception, then what does that mean for the rest of the product line? And how do you balance those kind of second-order impact of your decision in a way that is more evidence-based? Today, there's no way to test for this. You can run this in simulation. So really going beyond simply asking one question at a time, but then to think about what are the long-term implications of your decisions is something that our customers are quite excited by? I'd love to understand how you think about how predictive your model is in actually simulating real human behavior. I imagine a lot of evils on this. I guess what is your North Star metric? How do you guys do on that? And what do you think is the theoretical limit? It's a great question. So theoretical limit, let me just start from there. So it doesn't exist in the sense that humans are genuinely, there's a lot of randomness that if you ask me the same question, I'll actually answer the question slightly differently. So there is something in that degree of randomness in human behavior. However, there's a lot of gains and performance that we can have even today in the way we're predicting people. So the measurement that we do is, so at the level of population, we measure the distribution of responses if it is more quantitative. So we actually measure total various distance, which basically shows how close are the distributions of the ground truth versus the simulated information. And that is a metric that we run across all the use cases that our customers have. And we have something threshold that we believe is good enough for decision making. So a TV deal, let's say less than 0.15, we believe is actually quite strong evidence for making decision. So that is a North Star state that we want to hit for this class of use cases that are more quantitative, that's more question and answers. This also does cover RCTs, which is many of the core use cases our customers have. Now, there's actually a really interesting thing question to ask around, well, what about multi-agence relation? What about all the downstream implications that we're going to be simulating? What does the evaluation of those look like? Yeah, and then do daisy chain errors as you kind of, you know, if this one is 85% accurate and then this agent is telling another agent something and, you know, do you accumulate errors as you go towards multi-agence? Exactly. And one of the core thesis here is we basically see two categories of simulations. One simulation is what I would consider to be simulations that converge. The other categories of simulations are the simulations that diverge. And sometimes they actually coexist. And it's really about what the research questions do you have? Questions that converge doesn't actually matter if you have a little bit of error. Now the error can not be obviously so dramatic that it certainly is completely detached from reality. But you actually are okay, even if the errors do compound over time because the pull towards the convergence is strong enough that you'll actually understand where everything would fall. A good example here actually is if you simulate a network of people, then that network will always have a hub that gets formed. This is what sort of network scientists would call the skill-free network, for instance. This is actually what powered Google, too, one of the core observations of PageRank was doesn't matter how these networks actually get formulated. You actually see somewhat pages that get exponentially more links that are attached to it. This is a very fundamental behavior in humans, and then we also see in simulated networks. And that convergence always happens as long as you are replicating human behavior with certain threshold accuracy. Now there are other questions that generally do diverge. It's like your classical questions like, "Was words were one inevitable or was not?" And there, it is sometimes difficult to run the same simulation over time and get the same exact outcome. Imagine you're running a this is
not something that necessarily similarly right now is going into but imagine you're running a simulation of an election. With the same person win election every time, there are a lot of downstream implications of every single decision that does happen. So it does diverge. There, the core evaluation is around confidence. So imagine you run the simulation 100 times. How many of those times do the results come out to be x? And how can we actually use that to basically create a bootstrap's resembling to calculate the confidence around the simulations? Those are some of the questions that we do ask. And a huge part of this also of the power of simulation is then to show when it diverges to show the diversity of possible outcomes. So the people can actually look, understand the cause or mechanism of how we got to those outcomes, and prepare for those features. So those are some of the implications of divergence in simulations. Are there any mathematical descriptions of like why something would converge at diverge? Like I imagine if you have like an average function, maybe you converge. And then if it's like a, you know, you're splitting outcomes to a binary, then you might maybe diverge. But yeah. So the intuition I think is close. And technically this is also a research topic. So similarly, as a company where we do go deep into this research topic, in the sense that I see simulation as a field as akin to developing your day one of inferential statistics. You know, inferential statistics scientists actually had to do a lot of discussion and research over time to decide that p less than 0.05 is actually evidence that is strong enough for science. Similarly, it's working on setting the same kind of threshold and standards for the rest of the field. So those are the intuition. I think that's exactly the right intuition in terms of actually how to make a robust mathematical equation around when's going to like what's going to happen when it is a real research frontier for simulations. Thank you for being nice, but my vibe mapping. I'm curious, you know, it seems like so a lot of 4 to 500 is coming to you. I'm wondering whether there are, you know, non-existing corporate use cases that might, you know, they're like great mysteries of our society that might become solved. And for example, I'm wondering about economics, you know, central bank decisions. Oftentimes, like, you know, I personally believe in that macro. Nobody knows nothing. And oftentimes a lot of the issues come about from human psychology. So to me, macroeconomics is a function of simulating human behavior at scale. You know, I'm thinking even in the venture capital use case, we often debate internally, you know, does value accrued to this company or not? You could run the simulation of all the different layers of the AI stack and almost figure out where durability and value accrues. Like, if you had a kind of perfect simulator of human behavior, there's so much more you could do than serving the 4 to 500. Do you agree with that? And then if so, are you serving governments, you know, the like? Yeah. So it's interesting. And when we were still researching in this area, the way I actually got back then, my advisor is Michael and Percy excited about this, was I basically told them, look, we do this right. There's a noble price to be one there. And I truly believe that. And it's also not surprising in that your classical economics simulations, things like agent based models that really pioneered our understanding of back in the day, the kind of topics they studied was, how does segregation happen? What are the cause of mechanism for segregation? So scholars like Thomas Schelling would actually build agent based models that are extremely simple and rudimentary, but that showed something deep about human macro behaviors. And he of course went on to win an award price. I see the same opportunity here, but in an augmented way, we're back in the day, the agent based models were very much deterministic in some sense, where you basically, in this simulation of, let's say like motor segregation from 30 years ago, individual agent was simply red dot or blue dot. And every game iteration, they would look around its corner, see how many of its neighbors are of the same color. And if that threshold goes below certain threshold, then they would decide to move to a new location. That was it. But now we can actually create real agents, they replicate the full richness of individuals and run the same kind of simulations. So the kind of questions that we can ask that goes beyond simply the commercial use cases. For instance, in the in the context of macro economics, actually, the questions that I actually did get asked from economists were things like, when does bankfront happen? Or questions like climate change, one of the sort of core blocker of climate, like solving that issue is the collective action problem of many nations. Can we actually simulate that? Or what are the signals of a democracy that is about to collapse? Can we understand the origin story of the monetary system? These are the kind of simulations that I do believe are to be the North Star state of this field. And it is sort of interesting to imagine like what they would actually look like in practice, right? Because these would involve very large scale simulations with many agents interacting with each other. I do see a future where today this is something that the case today, a simulation is quick and fast to run. But what about simulation that takes actually a hundred million dollars to run once and could take many months to run? But when we run it, it solves one of the fundamental questions of our society that I do think is genuinely a very exciting possibility for this field. I mean, even thinking like politics, for example, could be forever changed. So today, everyone has an agenda of how they say some policy change will affect, will impact things. Why don't we just run the simulation? And there's all the downstream implications. Yeah. And not just what's going to happen this year, but what does it mean in the next five to ten years? Fascinating. I was going to close by asking you what makes you excited about the future? Is it what we just talked about? Or is it something else? You have some version of AGI and you have some version of simulations that will help guide the society. I would not have said that even five years ago, but that is a conviction that we have built up over the years as we are going deep into this research. And what's exciting is there's a clear use case today that can serve our users. But then there's a lot of innovation that is yet to come that I do think we're built up to actually building simulator that's akin to discern of human society. And one of the things that my co-founder Percy sometimes say is you look at the greatest scientific innovation. They often start from an amazing measurement. How will telescope really change the trajectory of how we understand the universe? Simulation can be that for human society. So the thing that does excite me, there's a lot of focus on natural sciences, but how can social, how can simulation really unlock our understanding of humanity and social sciences? And how can we actually use it to make our society be a better place? That's exciting. I remember reading somebody was excited about, you know, there's a small but, you know, breathtaking chance that the field of economics as we know it may actually become solved by simulation. And I extend that not just to be economics, but kind of everything that deals with human behavior and social sciences, which ultimately is everything around us. Truly. Wonderful. Thank you so much for joining today and sharing the story of both Smelville and what you know up to assembly. I really enjoyed the conversation. [Music]
Podcast Summary
Key Points:
Simley, founded by June, is building an applied AI lab focused on simulating human behavior and societal dynamics using generative agents with memory, planning, and reflection.
The inspiration for this work began at Stanford with projects like Smallville—a simulation of 25 agents in a town—where emergent behaviors such as social events, relationships, and personal decisions were observed, revealing the complexity of human social interaction.
Key insights came from early simulations of subreddits and real-world data collection, showing that behavioral models trained on longitudinal, narrative data (like life stories) better capture human values, preferences, and diversity than limited online data.
Simley’s platform combines real-world behavioral data (from interviews and surveys) with generative AI to create realistic, scalable, and representative simulations of populations—predicting human behavior with up to 85% accuracy.
The company uses simulations to support real-world decision-making in areas like product testing, marketing, earnings calls, and macroeconomic scenarios, offering insights into both short-term reactions and long-term downstream effects.
Simulations are evaluated based on convergence (predictable outcomes) and divergence (variability due to human randomness), with metrics like distributional similarity and confidence intervals used to validate results.
The long-term vision includes simulating complex societal phenomena—such as economic crises, climate change, or political collapse—offering a new frontier for understanding human behavior and social systems.
Simley sees simulation as a transformative tool akin to telescopes in science
Summary:
Simley, founded by June and her Stanford colleagues, is pioneering a new class of applied AI that simulates human society through generative agents. Inspired by early experiments at Stanford—like Smallville, where 25 agents exhibited realistic social behaviors such as relationships and events—Simley developed a platform that combines real-world behavioral data (collected via interviews and surveys) with large language models to create accurate, scalable simulations of human populations. These simulations predict real-world behaviors with up to 85% accuracy, offering companies like CVS a way to test product concepts, market messaging, and even earnings calls without relying on limited real-world experiments.
The platform addresses critical gaps in existing research by capturing long-tail personal data—such as life stories—that reveal deeper values and preferences. Simley’s work is validated through rigorous evaluation, distinguishing between convergent (predictable) and divergent (variable) outcomes. Beyond commercial use cases, the team envisions transformative applications in economics, climate policy, and political forecasting, where simulations could model systemic risks and collective behaviors.
They believe simulation represents a new scientific frontier—like telescopes for astronomy—capable of fundamentally advancing our understanding of human society and decision-making. As models evolve to better represent human irrationality and diversity, Simley aims to build not just a tool, but a foundational framework for understanding and improving the social world.
FAQs
Simley is an applied AI lab that builds simulations of human behavior and societies using generative AI agents. It creates realistic, interactive environments where agents live, work, and interact, allowing users to explore social dynamics and test decisions in a scalable and representative way.
The Smallville project at Stanford involved creating a simulated town of 25 agents with realistic behaviors, routines, relationships, and social events. It demonstrated that generative AI models could emulate complex human behaviors and social interactions when given proper structure and memory.
Simley collects real-world behavioral and life story data from surveys and interviews, training agents to reflect actual human preferences, values, and decision-making. This closes the gap between what people say online and what they actually do, improving simulation fidelity.
Simley enables scalable, representative, and deeper behavioral modeling that goes beyond simple surveys. It allows testing of thousands of product ideas, pricing strategies, and scenarios in a controlled environment with long-term interaction effects and second-order impacts.
Yes, Simley can simulate large-scale social dynamics such as elections, market shifts, or climate policy challenges. These simulations help explore how human behavior and collective actions evolve over time, revealing potential outcomes and risks.
Simley distinguishes between convergent and divergent simulations. In convergent cases (like network formation), small errors are smoothed out over time. In divergent cases (like political outcomes), simulations run multiple times to assess confidence and show the range of possible outcomes.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.