Go back

20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile

63m 41s

20VC: The Best AI Companies Have Unique Data Acquisition Strategies | Will Simile Kill Kalshi, Polymarkets and NASDAQ | How to Sign Fortune 500 Companies As Customers in Weeks with Joon Sung Park, Simile

In this podcast episode, Harry Stabbings interviews Jun Sung Park, founder and CEO of Simile, a company that creates simulations of human behavior to predict and shape future outcomes. Park explains that his inspiration came from a 2023 project where he populated a virtual town with 25 AI agents, each equipped with memory, planning, and reflection capabilities, allowing them to autonomously organize events like a Valentine's Day party. This work highlighted the potential of large language models to extract realistic human behaviors from training data, despite their usual focus on rational tasks like coding. Simile's core mission is to build a foundation model of human behavior, representing people's values, preferences, and biases, rather than super-intelligent machines. A key challenge is data collection, which involves sourcing representative everyday individuals and using randomized control trials to understand causal mechanisms—what actually changes behavior—rather than just predicting outcomes. Park emphasizes that clients like Starbucks don't just want to know future sales dips; they want actionable insights on how to alter strategies. The company aims to move beyond traditional survey tools to simulate complex interactions, from product launches to societal issues like climate change and democratic stability. Ethically, Park sees simulation as a way to amplify diverse perspectives in decision-making, ensuring that all voices are considered. He notes that while 1,000 people can provide statistical significance for narrow groups, larger datasets enable more detailed segmentation.

Transcription

12304 Words, 68207 Characters

English
My fundamental thesis here is for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible. I think there's a world in which in about two, three years, we're running a single simulation session that people pay $100 million for it. People live through different stages in their life, and they have different careers, different jobs. At each stage of their life, were they the reason why that thing was successful. If you squint, were they the common denominator? This is 20 VC with me Harry Stabbings, and I am too old and bored of standard podcast intros. So let me tell you a story. Shaddle Shaw, one of the best investors of the last decade at Inda's Ventures, emailed me late one night while I was on hold with my family, and said that he had a company that he was leading around in, and that he'd never seen sales and growth like it before, even though he'd been in companies like Wiz and many other incredible deckacorns. So I met the founder, that came at midnight on a family holiday on a Zoom call. He was one of the most impressive AI minds I'd ever met. Naturally, I invested on the spot, but I also said I have to have you on the show. This is Jun Sung Park, founder and CEO at Simile. They predict the future. There is simulation market that try and predict future human behavior. It is incredible. One thing that I took away from this conversation is that we will in the future spend hundreds of millions of dollars potentially on simulations of the future. And that actually people will get such value out of that output that they will be willing to spend that money. This was one of the best AI technical conversations we've had in a long time, and it was incredible to have Jun on the show. But before we dive into the show today, today I want to tell you about how the first AI law firm, Crosby, helped us close a big sponsor. As you know, some of the biggest companies in the world advertise on 20 VC, my British dulcet tones clearly convert well. I was working to close this big sponsor, and they wanted to get through legal review quite quickly to close the deal. Crosby turned red lines around in three hours and caught major issues that would have caused a serious problems in the future. Crosby combines AI, some of the best engineers in the world, from companies like RAMP and Stripe, and some of the best attorneys in the world from top 10 law firms. Customers get the best of both worlds, and elite human attorney reviews every contract, but they move incredibly quickly, returning red lines in under four hours. They help the fastest growing companies like cognition, ramp, and clay, close deals in hours, not weeks. Learn more at Crosby.ai/20VC. If you want to red line NDAs, MSAs, DPAs, and any other procurement contracts faster, go to Crosby.ai/20VC. And it's speed that you can really trust. While Crosby helps you stay sharp on the numbers, ZeroHash helps you move them on chain. Every great software company eventually runs into money. Uber had to move it, Shopify had to hold it, Airbnb had to settle it, money movement stopped being a fintech problem, it became a software problem, and increasingly, God, it's an AI problem too. ZeroHash is the infrastructure that makes global, instant money movement seamless. One API integration for stablecoins, digital assets, and modern payment rails, so builders can stay builders. ZeroHash powers some of the world's largest enterprises and financial institutions, including Calche, Stripe, Morgan Stanley, Gusto, and Interactive Brokers. If you're thinking about stablecoins, digital assets, and the future of money, it's time to talk to the team at ZeroHash. Visit zerohash.com/20VC to learn more. While ZeroHash powers on-chain money movement, Alphasense helps you spot what matters next. We used Alphasense on an investment that helped us close an $8 million deal. $8 million, baby, that's a lot of money. That's why I'm genuinely excited to have them as a partner on 20VC. Alphasense combines AI with one of the world's deepest libraries of market intelligence, including expert interviews, broker research, earnings calls, company filings, and real-time news. Every answer is grounded in this incredibly trusted evidence, and fully traceable to the original source, which is so important. So you can make really high-conviction decisions with confidence. But the best part, they're building super-analyst and always on AI analyst. So instead of starting your day with another search, you'll start with work you already done, your coverage monitored, the important development surface, and your investment brief already waiting for you. See for yourself head to Alphasense.com/20VC. That's alpha-sense.com/20VC. You have now arrived at your destination. June, I'm so excited for this, dude. When Shardall told me that I had to meet you, I'm going to be honest. Shardall does not tell me often that I have to meet someone. So I was like, wow, this is, I feel honored. Thank you, Shardall. And then we met when I was on holiday with my family. And I remember my grandparents were like a sleep upstairs, and so I was whispering to you. And I remember being so excited by what you were building, but then also having to be incredibly respectful of the sleeping elderly people next door. But thank you so much for joining me, dude. Thank you for having me. You're excited to be here. Now, when I spoke to a lot of your investors and friends before, they all said that I had to start on the very unique background you have, specifically kind of became very well known for a particular project. And it centers around Valentine's Day and assimilation happened as a result. Can you explain what happened and how that potentially led to the early days of Simile? So this was 2023. We had this idea that large-linked models are often used for simple tasks, like classification, simple generation. But we thought that these models actually had a lot more potential. One of the early observations that we made was that these models are trained on so much of human behavior with data, sentiment data, that were expressed on the web. So if you poke at them sort of at the right angle, you could actually extract a lot of realistic human behaviors out of them. I thought that was really interesting. And it was also particularly interesting in that it was domain-lockness stick. So if you look at the literature in comparison science for many decades, we've always had the vision of creating agents that are meant to be generalizable, that are meant to really be able to act like human in any environment. And my mind went to, well, maybe we have that opportunity here. So what we ended up doing was, well, if we were to fast forward many years into doing this, what would be the most ambitious vision that we might have? And that was creating entire lived experience of a town. So the idea here was we would make a game time, and we would populate it with 25 NPCs, so non-play over characters, except these characters would actually wake up in the morning to do their routines go to work in relationships and do all that. They would actually remember their interactions. They would actually plan their days. And some of the surprising things you end up seeing was the simulation itself was set the day before Valentine's Day, and you'd actually see these agents come together, have parties, like self-organize. So they would actually plan parties, they would decorate the cafe. We thought that was really interesting. Now, two fundamental contributions from that work. One was it was one of the earliest example of creating agents. So this particular set of agents were paired with back in the day, GPD 3.5, tech step in G. So we didn't quite have chat GPD back then. And then was paired with memory, planning, and reflection. Really the first times that those concepts came out to be an explicit part of the architecture in, quote, unquote, agent take out workflows. And the reason why we actually got that inspiration was if you had more than one agent side by side, you want them to remember each other back in the day, a language more ostentantly have the concept of memory. So I thought, okay, you have to give them the memory so that they don't say, hey, nice meeting you every time they meet their roommate. So we had them have this concept of memory and planning and reflection to make sense of very long term landscape. How do you solve that memory problem? Because everyone says, oh, we have a memory problem today. Yeah, you solve the memory problem of agents to prevent that from happening. So back in the day, it was actually the initial idea was fairly simple. The these language models are actually quite good at parsing natural language. So we put everything in markdown text file. That was it. That sort of worked. Now, the issue there, however, is because the language models have context window and even today, even if the context window is getting larger, the kind of experiences that these agents can have in the small game town is immense. And imagine now if we were to bring this to real life in world, like, don't want to live in the amount of memory that we accumulate is huge. So the problem becomes how do you make sense of this large quantity of memory? So imagine you went to get omelet five times in a day. You want to make sense of that aside from, oh, I went to get omelet five times throughout the week or something like that. So we had this concept of reflection, which basically was every certain interval, it's like a shower thought you ask agent explicitly to get bunch of their memory pieces and basically make sense of them. Why did you get omelet so often this week? Were you busy? Do you like omelet? Why are you studying for this test so hard like you were in library every single day? It does this matter to you. And they were actually start formulating ideas that are more higher level than what happens on the ground truth. So gradually they start to realize, oh, this particular research topic. I'm actually quite invested in it. This might actually have something to do with my childhood or my fundamental memory. This actually shapes who they are as a person. So that ends up becoming a very useful function, creating these agents that have personality, that actually has a point of view in the world that can actually make sense of a lot of this data. So that's how we did it back in the day. And so when we think about a simulation model today, for those that don't know simulation model is essentially that it's the creation of agents that then produces set of activities or actions that then will show us what a simulated future world might look like. Is that correct? That's right. When we think about building a simulation model company, would you say similar is a simulation model company? We are a company that is creating foundation model of human behavior that can then be used to create simulations of individuals, simulation of sub-populations, and then align the simulation of the entire ecosystem and even the market. Do you sit on top of core foundation models? How do you think about the relationship for those listening between an open AI anthropic frontier model provider and you? So this is a great question. So the way we see it is if you look at large-length model companies today, fundamentally the task they have at hand is to create super rational intelligent machines that are good at coding, that are good at natural sciences and mathematics. Similarly, it doesn't really care about any of those. What we care about is if we have a person make a mistake in this context, we want our models to make the same kind of mistake. We want our models to be biased in the same way humans are. In a way, we want to be a representation of people's values, preferences and tastes, sort of their subjective half of their brain. That's what we care about. A lot of what people say is different to a lot of what people do. How do you think about the chasm of what people say and what people do and how that impacts your models? So say to give it real. And if you look at the web data, it is fundamentally data of what people have said, not what they have done. And obviously large-length models today are trained mainly on this web data. For us, we actually do collect a lot of behavior data. We collect transaction data, we collect observational data, we also partner with our customers, our vendors to collect some of this data. But my personal heart take here is a lot of observational behavior data and what they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction tasks, but might take here after interacting with so many of our customers and also being in research, no one really cares about prediction. No one really cares about what's going to happen in the future, unless you're trying to predict a stock market. What people actually care about is they want to shape the future. They want to know, imagine your Starbucks, doesn't really help them to know that your fratino sales is going to tank in two quarters. They'll hear that and they'll be like, what do we do about them? That's terrible. What they want to know is how can we prevent it? What do we need to do now to change the future? And there, what you really need is causal mechanism. You need a model that can actually reason of a causal mechanism and kind of factors. So the kind of data that we care deeply about is a lot of randomised control trials. We actually run a lot of AP testing. We show the models. Imagine people have done this versus that. This is how their behaviors were actually changed. That becomes a core part of our training asset. So this is actually the data collection that goes beyond observation or data that similarly collects. Is data collection acquisition the hardest element of building simulation models fee? Like if we think about the kind of cool pillars for traditional models, it might be compete algorithms and data. Is data the biggest challenge fee? Data is an important piece of assembly for sure. And for us, really the data collection challenge comes from two angles. One is actually sourcing people. Sourcing people here is a little bit different than what other language model companies might consider to be their people or their population. We don't go after these expert programmers or expert scientists. We go after people like us, everyday people, living their everyday life. But what we care about is are they representative? Do we actually have the same representation of people as we do in the world that we live in? And then actually asking the right questions to these people. What are the experiments? What are the questions that actually get at the fundamental core nature of who they are? Some of the questions we actually ask at the start of our data collection at times is actually saying something like, tell us the story of your life. Where did you grow up? What did you experience? What were some of the hardest problems that you had to tackle or decisions you had to make? They were so loud about these people. And that's what we try to do. In terms of people don't want to predict the future just so we can drill down on that, I thought they do like Starbucks if they can predict that Frappuccino sales will be down in two quarters. They can amend bluntly that buying cycle. They can change how much they purchase. Isn't that valuable? What am I missing? But there's a thing. The reason why they want to know is so they can change their strategy. So certainly talking about how much resources they actually need to actually serve this market, that is a kind of changing in behavior. But fundamentally it is about kind of factures. So what we have this market that we want to serve, we want to maximize our value as a company. What do we need to do to make sure that we react to this dip in the market? Whatever it may be. Now fundamentally though the work that we do is about people. We try to simulate people and represent people's perspectives. So the value that we provide is kind of factual in terms of what your consumers, what your population would do. When you look at what can be done for some of the biggest brands, you mentioned like a CVS that, incredibly valuable for surveys for customer feedback, for determining what customers really want moving forwards. I didn't know how this is going to be. Do you want to just be a net generation Qualtrics? And how do you prevent that being the angle? So the way we see it is again, fundamentally the core primitive of what we're trying to build is very straightforward. You tell us what population you're interested in and we'll go model them. Really so far the layer of innovation has lived in the more tooling layer. How can we create better survey tool? How can we create a better interview tool? Simulation is fundamentally about something different, which is how can you create the most generalizable model of people so that we can represent people's viewpoints at scale? That goes beyond simply running surveys or interviews. Down the line, I actually see simulation as a field moving into context where hey, can we actually create simulations of many people interacting with each other so that you can understand all the downstream implications of your decision making? Or if imagine you have a new product you're about to launch, can you actually simulate the entire launch and how the audience might actually react, how to mark in my shift? And this also goes into the scientist part of me also get quite excited by the vision where simulation I do think can also be cure for many of what we call quote-unquote wicked problems. A good example here might be things like climate change requires collective action across many stakeholders who have different incentives. One of the reasons why such problems are so difficult is actually finding them right a calivariate state where all different parties come together to make decision for global good is very difficult. Can we actually simulate those decision making processes? Can we actually simulate even things like, in what conditions does a democracy fail? Can we actually predict that? These are the kind of questions that simulation ultimately can answer. Thinking about democracy's failing elections for a government has been incredibly useful tool. How do you think about who you can and should work with versus who you shouldn't? This is for us where the principles matters so much. The way I see it, simulation is a piece of technology, is one of the twin pillars of technology. I'm a fan of science fiction. You read any advanced science fiction, there's always two pillars. One is some form of age AGI that always shows up. The other is simulation. And like with any powerful technology, the misuse, the potential for misuse is quite real. And the way we see it, simulation at its best ought to be representation at scale. People have different viewpoints, different perspectives, different tastes. Many of their viewpoints are not considered in rooms where important decisions for them are made. We want to always say we listen to our people, we listen to our customers, we listen to our stakeholders, in practice very difficult. This is a way for us to ensure that in every decision making, we actually listen to people at scale. That's the Nord star. How much data do you need to feel confident in an accurate prediction outcome to be displayed? Is it 100 people? Is it 1000 people? Is it a million people? You want to have more people represented so that you can segment down to a specific sub-paplation. If you look at any social scientific literature, if you have a very narrow population of interest, you would usually get statistical significance in the study that you want to run by the time you have 1000 people. However, oftentimes the kind of ways that people query our system is they want to come in and say, hey, filter down to x, y, z population. Those filters are often created on the fly. For us to then be able to, similarly, people's responses across all those filters, thus mean we want to represent the entire population. So that's the journey that we're on. Is it self-fulfilling? Do you get better and better at predicting over time? I think that's something that's the case because there is the data flywheel. There is the learning that occurs as we get more and more simulated results and see what happens in the ground truth. That's absolutely yes. And this is obviously one of the core value proposition for our early partners. Because they know in their business context is similarly getting better and better and better. And do they have that compounding advantage? It's kind of like AlphaGo, you know, they just beat the shit out of the model and played it a thousand times every day of activities and outcomes in the world is another game of AlphaGo where you can correct the model on what was wrong and what you missed and what didn't happen. And 10,000 days in you should almost be better than the model at the model. Do you know anything? So this is actually quite interesting. You might think of let's actually think about a different example. So how does data flywheel work in simulation and why would it work? If I were to take a brief detour and talk about coding, the reason why coding agents has had such massive improvement over the years was because their learning, their reward function was extremely clear. If you make a suggestion and your user says accept fantastic. If they stay reject, also if useful. You very quickly know what is going on, what is bad. That actually was one of the core learning mechanism for these models. And it might be easy to look at simulation as a field and say, "Well, where are you going to get the reward?" Because fundamentally, all the things you're trying to predict is in happening in the future. It's going to be hard to validate. It is true. At the same time, I actually think simulation has even better mechanism, which is, the world is our ground truth. We live in the ground truth world. So what we can do is every single day, we can be generating tens of thousands of hypotheses. Each hypothesis is mapped onto an end statement. If this happens, we know whether we can validate the simulation to be right or wrong. And we're basically watching the world every day seeing which of those hypotheses are answerable at what time. And we can basically say, "A month goes by, we generated a million hypotheses, x percentage of them came true." This is the best way to learn about the world. Does it take a huge amount of compute to run these simulated environments at scale and well? Compute is an important piece of simulation. Of course, a lot of the work that we do is to make our simulation be more efficient. So a lot of our compute initially actually goes in to create the initial breakthroughs in technology. So it is actually exploring the different ways to train, is exploring the different kind of data set. Once we have a point of view, we can very quickly make it efficient. So it's some of the things that I've seen, wouldn't simply, as we built this company over the year, is right now we have a model that's been in production. This model used to cost about 100 times more to run than it does now. And some of it does happen because we actually found different ways to model, but with the same reward model with the same philosophy, but just in a way that's much more efficient at inference time. So there are these kind of tricks that we can play, and these kind of scientific advancement we can make to making this cheaper. A lot of the investment, however, does go to find it in the short point of view. When you look at it like serviceable market, or total address market, time and venture speak, you obviously have your CVS in your huge enterprises who would absolutely want to work with you. It can also be consumers. Regular consumers wanting to see what happens if and running their own environments. Is this a play for everyone? Is this a play for the biggest companies in the world? How do you think about the time for something like Simulay? So the start of my career really came from research, obviously. And the job of a researcher is to serve the humanity. Then we do our research, obviously for our own enjoyment as well. We love the process of finding new things in the world. But fundamentally, it is a service. It is a belief that if we are able to make scientific breakthroughs, this is going to down the line, serve everyone in our society. Then is how I see simulation as a field as well. So right now we do serve enterprise customers for a couple of reasons. One, obviously, I be frank, there is the budget that there is a clear product market fit that we see today that does excited us at the same time. It is an amazing way to validate the technology. It is very important to us that we get the feedback loop to be as tight as possible. So we know when our simulations, right? When our simulations wrong and we're improving it every single day. And obviously there is this side part here that's just as important, which is I have a colleague when I was a Stanford. My office next next to mine was Pat Hanran. He was one of the founders of Tableau. He's a graphics professor also, if you want to earn a worse, a very well-known person in this in this landscape. An advice he actually gave me and some of my colleagues was the best way to get feedback is to actually ask people to pay you. That was their core philosophy of Tableau. And I also see this here. So getting the best kind of feedback matters a lot. So enterprise market, market research right now is unwedge that we found that actually have significant budget that have immediate product market fit. But down the line, I do want this technology to be used by rest of our society because fundamentally what we are trying to serve is help people make better decisions. Before we move on to expansions, it could be used for when did you know you had a product market fit? You said you felt not pull. When we were like, ah, we got product market fit here. Many of the Fortune 500 board members and their C-sweets reached out. In part, they do come to Stanford to see some of the demos that are happening in the lab. And they all saw the small bill demo after it got released. And everyone thought, oh my god, if we can simulate a market like this, this is going to change the way we operate. So you could immediately sense the product market fit. Really, this was the forcing function for us to then say, okay, this is actually quite interesting. We're actually going to show and validate that our simulation cannot just be an interesting demo, but it's going to be accurate. So we spent about a year actually demonstrating that we can create models of people that are actually amazing and validated at predicting people's behaviors across surveys, behavior experiments, real environment. And we show that we can actually predict people's behaviors and attitudes 85% as accurately as people replicate their own. We put that work out at the end of 2024. And that's really what started the field around synthetic panels simulations. And that's the market that we're seeing today. Will synthetic panels be larger than human panels in three years time? Yes. The way I see it, synthetic panels will be larger than what we know to be the current human panel market. In part, because this can really raise the ceiling of the kind of questions we can answer. What I see today in the market is actually quite broken. We have so many questions we want to ask about our market. If we were to release this product, if we were to have this product strategy, this product policy, you're a scientist, then you want to run this study, or you want to try like this macro scale experiments, you are looking at maybe five percent of those ideas. Get answered. The rest of the 95%, we never bother experimenting with because we either don't have the ability to do them, especially if it's something at the emergent scale. Like we literally don't have a way to run those emergent scale experiments. At the same time, we don't have the budget and time for it. So a lot of the decisions that we make as a society, we base on our gut instinct. Sometimes they're good, but sometimes they're very biased based on our own narrow experience. So what simulation will do is unlock the limitation and us to actually test every single hypothesis that we have about the world before we have to launch into the world? How do you balance the pursuit of the next dollar and serving customers who pay a lot of money, I'm sure, versus research prioritization, maybe focusing dollars there over building out a customer success team and an FDA team. How do you balance the profit maximization with the research purity? So similarly, as a company that has a real product and engineering team, but at the same time, we are a research company. The co-founder is the four of the three co-founders are researchers. So we have as co-founders myself, Michael Bernstein, Perseillon, Laney Allen, myself and Michael and Persew are all researchers as Stanford. So I led researcher on agents, simulations. Michael was one of the co-authors of the ImageNet that really kickstart the AI revolution and he's been a leader in humans and AI. Persew was the person who literally coined the term foundation model. The vision for this particular area is the vision that we can actually create the next paradigm shift in AI and in the way we view technology and impact of the technology in the form of simulation. The reason why we are able to however operate as a research lab, but also have an amazing product and engineering function and go to market function that's led by Michael Persew Laney is the alignment between what the technology can do, the promise of technology, is so close to what our market actually requires. The better the model gets in representing people, the better simulation we can create, it immediately means better experience for our users because they'll have much more grounded, much more accurate simulation. It is very difficult to maintain both a lab and a product company if there is not that alignment. But when there is, it can be quite magical and that's what we're seeing as similarly. Can I ask you when we think about the pursuit of some of the largest companies on earth, we mentioned I'm not sure which customers were able to save us, not say, but you mentioned CVS. People always think it's like multi-year, incredibly long-sale cycles. Was that something that you experienced or was it a different experience for you getting and working with some of the biggest companies on the planet? What's fascinating to me coming into the field of simulation, especially in this market was last year when I started the company, and when I left Stanford June of 2025, so it's been exactly one year. I actually thought our market will take about a year or two before they warm up to the idea of simulation. So we'll basically build a right foundation for this company and for this market and we'll go aggressive maybe towards the end of 2026, was what I had in mind. That's not what we experienced. What we experienced was our customers were moving extremely fast, also in ways that truly made me change my perspective on corporate America. Our leaders that we work with, for instance CVS, I've been working with this particular leader, Shree who is their VP of insights, extremely for looking, extremely ambitious, extremely hard working, amazing kind of part, two vision-like, similarly. But what I've also found was the pain they were feeling in their day-to-day work was so real. It was way more acute than I could have imagined that when they realized that there is or there could be an answer in this market for addressing some of those pains of very slow experimentation, budget and so forth, they are ready to drop everything and try us out. So we actually saw some of the largest customers in the world move at a lightning speed for enterprise where we saw them close deals within three months. Three months. Wow, that's very different to what people traditionally think. What matters more to them speed of output in other words being able to get results very quickly on their simulations accuracy of simulations. It is both. There are so many questions that they are truly relying on their gut decision today. That if they can get some form of evidence to at least directly guide them in the right path, then they're ready to try it. And then they very quickly realized that, oh, this is actually an amazing way to interact with a lot of data. This is an amazing way to gain evidence that is actually quite accurate. And they actually, one of the ways we actually got some of our first customers was in the first call. They actually had a finding from large consulting companies and they basically quarried our system. Hey, if we were to rerun this, what would the system say? And we predicted the outcome of studies, they took three to six months, but just within two minutes. That's very powerful. It must be so compelling in a customer conversation to be able to say, like, you did this campaign. If you've done this campaign, it would have been 12, something more effective. Do you want to buy our product? It is such a good sell. I'm a seller to be able to have that data is unbelievable. How do you think about value extraction efficiently? And what I mean by that is that if you work with like a CVS or you name any of your big companies that you work with, these are massive companies, where if you're able to do your job efficiently, you can move the needle to the tune of hundreds of millions for them. In some cases, billions of revenues. Charging like a million bucks feels like a large cousin between value generated and value extracted. How do you think about closing that cousin to be more fair? Yeah, that's a great question. And I see the market moving in this direction. One of the core premise and one of the ways that our customers are actually finding value and similarly is actually avoiding really damaging decisions that could have cost them hundreds of millions of dollars. So it's prevention or optimization. It's both, but certainly prevention is a huge, it's an obvious value case that, oh, wow, that could have been a total disaster had we run that that would have costed us half a billion dollars. We ran simulation and that prevented it. That's no brainer. This is a true painkiller in their case. If you do your job efficiently, can calcium polymarch is still exist for a lot of that markets? It's an interesting question. I do something you think there is an overlap here and that we are companies that are fundamentally interested in the future and helping people at least get a glimpse of what the future might be. Where I see similarly come in is we are a company that is not just interested in what's going to happen, but more on how it's going to happen and why in that way. And this is also the value proposition that our customers are most inspired by. It's one thing to simply predict, but can we actually show here all the steps that your ecosystem is going to take to get to that particular outcome. And this is the way you can prevent that or you can't encourage that. That is ultimately the power. When it comes to team building, you said about the craft team building. I think it's really interesting. It's an ongoing challenge building the best team. What have been your biggest lessons coming out of research and what it takes to build an all-star team similarly? So a couple of things. One is the team has to be balanced. There are certain power that I can bring to the team, but there's also a lot of things that I don't know. I was a researcher. I was not an enterprise seller. I needed Laney to be my co-founder to lead that part of the game, balancing the team and being able to see where your team is lacking, especially as we scale, there are new gaps that are emerging. Actually, seeing that ahead of time and making sure that we fill those gaps, I do think it's a core fundamentals of building a great team. At the same time, I also do think it's important that the team remains consistent in their values and in their rigor. This is more of a painter's analogy. So when I was a painter, I was a figure painter. So I worked a lot with human subjects or tricks. There's sort of this untold secret amongst figure artists. It doesn't matter who you paint, your subject sort of looks like the painters themselves in some ways or at least they share the similar vibe. I think building a team is actually a lot like that. In the best team, in the team that you care deeply about, you really should see yourself in the team. And for me, a couple of things matters the most. And this is the same standard I tried to uphold for myself. But one is, are we the common denominator of success? People live through different stages in their life. And they have different careers, different jobs. If the answer is yes, then what that suggests is a couple of things that they have extreme degree of ownership that they are the kind of people who come in and say, it doesn't matter how everything else goes. I would personally make this successful. It also shows the ability to reinvent themselves. So one of my co-founder, Michael Bernstein, he has had a very interesting career as a researcher, where during his PhD, 10 years ago, he started his field in crowdsourcing collective intelligence. Then very quickly, during his early years as a faculty member at Stanford, he went into a different area as a Baye. And then now into general the Baye agents and simulations. And at each step of the way, you could sort of see in the work he's done that this is very Michael. You could see that this is the person who led a lot of the success. That's an amazing signal. And another piece, the second piece for me is this is a little more niche to myself, but I found this to be very true, at least to the way I look at the world. Do my leaders and my team have two superpowers? That's not supposed to coexist in one person. Any expert will usually come in with one superpower or even sometimes multiple superpowers, but they're all correlated. You're an amazing programmer who happens to be amazing at mathematics. Very common. Where I found things to be particularly compelling is if people have two superpowers, that's really contradictory. The most common one here is actually the greatest CMOs, which are very few, a kind of one single hand, unbelievably data, rigorous oriented scientific in their approach. And then you blend that with this creative artistry imagination. And they are two relatively opposing kind of mental approaches, I think. Yes. And it's very rare to have that in a CMO, but when you have that, that is the world class CMO. And that's magical. Yeah. That particular description I actually sometimes have used for my board members for a draw, deeply analytical, but he's very intuitive. And I think that's how he makes investment that happens to be very successful. So hopefully, similarly, we'll continue on the success. But in my team, the kind of things that I also see as an archetype is, and I also categorize myself as one of these kind of people on day-to-day basis, for instance, a lady, one of my co-founders, she is paranoid. She's somebody who will come to the table and say, unless we put everything on our table today and do everything possible, we'll lose, we'll fall behind, that everything will fail. But long term, she's religious. This is somebody who fundamentally believes the world is stacked for her. Then no matter how this goes, we will make this successful. Actually, balancing those two at the same time is quite difficult, because if you are short-term paranoid, then you're likely going to be very pessimistic about your future. You might be amazing at short-term stocks, but not great as a company builder. If you're religious, you have the opposite problem, which is your complacent. Then you sort of feel like, we don't have to put everything on our table today. Things will be okay. Balancing those two needs somebody who is broken in some ways. That somehow they found a way to be deeply paranoid, but at the same time ignore all the paranoia of today, to believe that the world is going to be amazing. I think it's actually, as exactly me, and I think it's actually believed that the paranoia that you hold today helps that future state be amazing. You know, I was often interviewed the world's most successful founders, and I also say, what do you wish you'd known when you started? They always say, I wish I'd known that it would all work out, and I wish I hadn't been so worried. I think it's the worst answer you could give me, because the fact that you were so worried, and so you did the prep, you put the work, and you stayed up late to do that presentation. That led to the success. Without the paranoia, Laney didn't hit the corner. Laney didn't set the urgency in the sales team. Laney didn't hire those extra people, because you didn't know the access to mom would be there. The paranoia drives the success. It's a really interesting one. Can I ask you, Brandon at McCall was on the show recently, and he was like, honestly, researchers, they're in the tens of millions of dollars. It is so expansive. Do you find that to be true? And how do you find this intense, war-footed research talent in the bay? Absolutely. So the research talent is very sought after today, and I have my closest colleagues and friends, who has taught a calm, dust range, and tens of millions. Now, when they join, similarly, I'm fairly upfront with them. It is not possible, it doesn't matter how many hundreds of millions that you raised, meeting them at their base, salary is tricky. However, the researchers fundamentally care about a couple of things. They care about a vision. If this idea truly comes to fruition, like these are people who have literally seen OpenAI being the laughing snog at the end of Silicon Valley, to becoming a nearly trillion dollar business, and these are people who have seen Anthropic go to that same state, wouldn't the past five years. So these are people who are fundamentally aware that deep, ambitious vision can actually come to fruition. So they care deeply about the vision. They also care deeply about the impact. What are the societal impact of the technology that they'll be working on? And is it actually interesting to them? You worry about the retention problem in the Valley today. You see so many researchers move with such promiscuity if you can use that word. Do you worry about the retention problem today? Consistently. And in fact, I actually do view the role of leadership to be that of obviously hiring amazing people, but also providing a platform where individual members can express their So, prepare to do. their maximum degree. And I do think this actually part, genuinely does matter. And retention can be challenging, but it can be done. And one of the core sort of, I have a small sense of pride in the way my career has penned out over the past six years or so as a researcher, where P2 students often go from one project to the next, and their entire co-authorship will change, maybe except for your advisor. I've had sort of an interest in career where in the past six years, all my core team members never left, that we all move from one project to the next to the next together. And now when I said, hey, I want to do this thing. And but similarly, I was able to somehow convince Michael and Percy, who were, they were actually my doctoral advisors to actually come join me. And I think a part of it is, we worked so closely together that there's genuine sense of trust, but at the same time, and I am somebody who fundamentally believes that one, it is my job to communicate the degree of confidence and trust. So a team that they feel like this would work out. - Can I ask you a really important question for me, which is like, I see a lot of amazing people in research and academia who are considering or starting a company in the same way that you did. And I always worry that I'm going to finance a science project. Science projects are low-grade and all though interesting and intellectually satiating don't always make great companies. You've been able to do that incredibly well in the last year to 18 months. If you were an investor analyzing a group coming out of academia, what would you look for that would give you confidence that they would be able to make the leap from research to starting a company? - The thing I actually would look for is, are they married to a problem or are they married to impact? Sometimes researchers are very much focused on a problem and something about their problem fascinates them. But oftentimes it's just not a good company or it's not a thesis they can really form into a company for various reasons. But there are researchers who are fundamentally driven by impact that they can have in the world. And for them it means finding a problem that can actually reach people, finding problems that can actually generate revenue and that's what drives them. You want to find researchers for in that category. - Talk to me, it was Shardou that introduced us. Huge thanks to Shardou for that. But there's a new funding round that's come to be in the last month or so. Can you talk to me about the funding round, how it came to be and how you think about it? - So we raised our 100 million dollar round about five months ago. We were preempted fairly recently by insiders. So Shardou let index, let our previous round and including Shardou and some of the other insiders were looking at the market. And Shardou has sort of this comment that he every once in a while makes where he's seen some of the fastest growing market and his track record does show that he truly has seen the different markets. He's quite never seen this kind of traction, this kind of pull. When that is paired with a technological progress that has been made and also the amount of computer we can also leverage to even further accelerate our progress, that sort of prompted our insiders to go, can we actually put in more money now than later? So that's how the initial round conversation came to be. We were now planning on raising at that particular moment, but there are a couple of teams that I practically respected in the valley that if we were to be raising, I wanted to talk to. And I found out that the team that I had in mind as my top up list actually was a new meta team at Greenox. And turns out, his team actually has been looking deeply into this market and all the players, how the market is going, and we're actually prepared to make the investment and they're looking for a sort of the right time to do so. So I reached out and said, hey, this is going to be the round. We're not running a process. So if you've been interested in joining, we have made a few days to make that happen, they were excited. So the round came together. So we raised $200 million. So it brings our total funding to be $300 million raised over the past six months or so. It gives us a very meaningful capital to go after this really ambitious modeling challenge and building up this team. It also brings in a lot of really exciting people to the team. Trolldool has been a fantastic partner. We actually have a lot of index connection. At Simley, our seat actually was led by Mike Vopi, who runs now his own firm and Trolldool and along with Neil and Greenox team, along with Patrick and who is their partner. You didn't need the money, I take it. It raised $100 million in six months ago. Was there a consideration of, we don't need the money, why would we take $200 million now? There was something in that consideration where we netted out was the modeling does take compute. And this is one of those areas where you can actually, here's a fundamentally interesting part of our research. With research, you really can control the outcome, necessarily. Well, you can control is the input and the process. And we were sort of at this moment where yes, we can actually significantly raise the input both in terms of data, compute spend, to actually meaningfully accelerate this progress. That's when we thought it actually makes sense. What did you not know about fundraising, coming from a world of academia research to you now know having been through three rounds? What did I not know about fundraising? Well, one actually here was coming in, I was actually fairly skeptical what the roles of VCs actually were. (laughing) What did they actually do? We were all off. What did they actually do? How did they help? And I would be honest, I can't still quite put my finger on it and say this is no way they help. However, if you bring in the right set of people, what I have realized was they can be the greatest partner and they can also be really strong set of mentors. Because I never ran a company, certainly not one like this. This is my first real experience building a company. I have a lot of technical experience of doing research, but so much of what I need to do on the day-to-day basis is new. If there's someone I can trust, then that's an amazing boost. Initially, I started to work with Mike VOP and we also had other firm ASTAR who also helped lead our seed. These funds and the Mike's team, and we also work with the Shiny very closely there. They really became so core mentor as I operated in the field. They also were the Mike actually introduced me to Laney who ended up becoming instrumental and as I thought about the business, I found a great partner and friend in her, which also has been an amazing part of this experience. So something new one is the VCs can actually help in some magical ways. And they have seen enough that if your experience VC, they can actually provide the advice that the founders might not have coming in. That is one. Another one here is things always happen a little bit sooner than you would expect. Obviously coming in, I had sort of a, in my mental model, okay, well, if we raise seed now, that means we might raise our A in about a year and maybe be in the year after or something like that. All that happened within a year. We raise seed and I think our series A was very soon after and our next round also came very soon after. So I think the market is always moving perhaps one step ahead of where you are in terms of their interest and investing in you. And it is useful to be prepared for those moments. That's what I've learned. - To worry that the market is so frothy that it can get ahead of itself. I when you announce this fundraise with the people that you have and with the press that you'll get, you'll get more interest for an extra round. And it's an ongoing cycle and the bluntly, the hubris is very high right now. Do you think about that? - I do think there's parts of market that is actually quite frothy. - Yeah, for sure. There's a lot of capital going in, there's a lot of excitement. This is where I actually do care a lot about the fundamentals. Well, for your customers, who do you actually work with, what's the market pool that you actually see? And what's the technology? One of the most interesting things about how opening and throttplakes of these companies grew was there were very strong fundamentals they could actually map out. They could actually see all the models are getting better at this rate. Oh, and there's this kind of demand. Some of those they could actually foresee. Some of those we can actually see it similarly as well. Does the unit economics vary for you on a per simulation basis? And what I mean by that is like, if you look at say for anthropoconopen, and model routing, some tasks require frontier models which are much more expensive, much more token heavy versus others which are much easier and can have a degraded or older model and a much cheaper model. Is that the same for simulations? Did different simulations cost different amounts in terms of compute token usage associated? They do. Usually when you have simulation, that is trying to answer something that's much more complex or something that's let's say you want to actually understand all the downstream implication of your decision or you want to do market segmentation study across all of the US, much more expensive. What I also have seen, however, is it is in those simulations where we actually get higher ROI for our users because those decisions are some of the most costly decisions if they fail to make the right one. So this is actually interesting for simulation as a field. We've always seen as a community. The inference cost going up and up and up and we now have these thinking models that are thinking for like half an hour a day and actually starts spending like token maxing and spending a lot of money on just running this process. I actually do think simulation can actually be the next frontier of that. In my vision, I think there's a world in which in about two, three years, we're running a single simulation session that's going to take $10 million to run, a single session, but it's going to be so valuable that people will pay $100 million for it. That's where I see it go. And that would be for the world's largest enterprises that be for a government or whatever that may be. It's based on the sort of the high end of the spectrum that's what it would be. What cannot be simulated today that you think will be possible in three years? So for me, it's actually a little bit less about what can not be simulated because I actually do think everything that we want to simulate, we can actually create the. initial proof of concept. However, as we all know, one of the core challenges of AI is actually bridging the proof of concept with real value productionizable technology. So that's actually the chasm that I see. Interesting thing here is I see the world of simulation going into this world where we are creating a very complex multi-agent simulation or we're running a very long study with many different steps of simulations along the way, but we actually started the field from multi-agent simulation when we created the small game town that was fundamentally that vision. It sort of also makes sense because we did that because we might myself, Michael and Percy, we sometimes sit together and do this exercise called time machine game. If we were to write a time machine, go to 10 years into the future, what's going to be the craziest thing we're going to see and can we do that now? That was the motivation for running the small development. So this can be done, but the question is can we evaluate the efficacy of these simulations? Can we actually propose this as a scalable productionizable system that people can actually rely on for making their decision? That's the chasm and that's the thing that we see getting bridged every day. I'll choose part of it also is getting models to be better creating bigger simulation, making the system more scalable, all that becomes a part of this. Can we play a time machine game with me and you? Let's do it. In 10 years time, what is the craziest thing that you can see happening? A lot of things, but one thing I would actually say, I am someone who is fascinated by history of technology, a lot of injuries that we can draw from it. What I see today that's prominent in AI space is what I consider to be the CPU of intelligence unit. You have these one-lingged model that's really large, that's very smart, that can do very complex reasoning tasks. That's like CPU. What I see coming and what I think simulation as a field can offer is the GPU of intelligence unit. As I mentioned before, similarly, does not care about creating really smart, super intelligent machines. What we care about is creating models that are as smart as we are. I fed a lot of things. I want to make sure that the model that represents me, they are the same way. But the beautiful part of a people is individually, we have so much diversity, so much different takes in our world that makes individuals so interesting. But also when they come together as a large collective, the emergent phenomenon that they were able to draw out is some of the most wonderful thing that we can see in our world. Creating a society, creating an amazing process that actually allows us to make all these achievement. Can we actually replicate that in simulation? I think it's going to be quite inspiring. What's the crazy prediction that every single person will have a replicable twin that acts and behaves like them in a simulated world? I think that's the vision. The vision here is again representation at scale. We as a society have found over the years many different ways to represent our members. Sometimes it's a form of government. Sometimes it's actually companies. Company we are as a society allocating capital to make sure that they serve the society and the needs of people. But if we can actually create an artifact and they're in a much more scalable and granular way represent all the individuals, what are the new kind of policies, new kind of companies that can be created on the basis of it? I actually do think it's quite interesting. If I'm a hedge fund, is this not the most obvious blind in the world? If I'm looking for alpha and ads on everyday activity, fuck, sign a million dollar contract with you and get unbelievable insight? Yeah, maybe Simily will actually own a small hedge fund down the line. That's a cool idea. Would you be down to data? Well, it turns out we actually do have quants in our firm. So some of the members who have joined actually do have more quant background. And I think right now they're joining that because they actually want to start a quant firm. And Simily, they actually joined because they actually see the vision of Simily very much what align with their passion and interest, which is to model the world. But down the line, I think it's actually an interesting idea. Is there a world, again, I'm just thinking crazy time machine world? Is there a world where you are so efficient and so good that actually stock markets become uninvestable because the world is skewed to Simily's hedge fund or similar providers and actually it is not a fair marketplace. I think especially with, obviously, if we were to assume that we're going to have some form of a AI and if we were to assume some form of perfect simulator, I think a lot of the things that we assume to be true about our world, I think will change. So I think one of this could actually be the stock market. What else do you think we assume to be true today? Do you think we're being five years time? What I think we generally assume to be true about the world that we live in is that it is fundamentally impossible to get everyone's perspective. Therefore, we need representatives of these people to approximate their perspectives. So far, it has worked in some ways. It has failed in other ways. I actually don't think this is a limitation we have to suffer through in the future. I think there is a world in which we can truly create a layer that becomes a representation layer of our society and of our collective intelligence. Is the future of love not also Simily? And what I mean by that is, if you were able to create effective simulations for yourself, dating itself could be much more efficient if I could, I've got a girlfriend and she's watching and if you could date 100 people at the same time for the first date, of course, not onwards. I get a lot of trouble for that June. But if you could, it would be much more effective at finding the one for you who could pass through to the next stage. Do you know what I mean? I get that. Well, look, I think love comes in different forms and I think just like as we discussed today, people have such degree of diversity. Personally, I am a bit of a romantic, I have you honest. And this actually goes to the point that I mentioned about long-term religious. I actually do believe that love is sort of a find, at least in my life, you know, find its way in a more organic way. I actually do think the way I meet the person I personally do care a lot about and the fact that we sort of have shared age journey, I personally care a lot about. I think that piece of humanity will, I don't think, ever change. Actually, it is experiencing things together, having that shared memory. I actually do think it's fundamental to the way we form trust. This is also we talked about process a lot. This is actually a part of it for your users using your simulation. Can you actually bring them along in this process? I think finding love is it's a little bit like you're co-funding your life with this person. So can you actually find a process that actually would bring them along in this process of living? I actually do think does matter. Oh, you're so romantic. Unfortunately. Do you know what I'm all I'm thinking about? Dude, I'm a content person. I'm a content person and investor. Weird mindset, actually, in both ways. There's a show called Married at First Site. You might not know it. It's where you marry someone on first site. But I'm just thinking it'd be the most phenomenal advert for similarly. If you could do the perfect marriage at first site because of simulations that have been run before. But, you know, I'm just leaving you with palt of wisdom that I think would be great. We're going to do a quick fire on size, their short statement. Who's the most underrated AI researcher today? And there's so many, but I actually do believe think there are some incredible people who are working at these larger labs, whose names are not known because they work at larger labs and they don't publish. But I think there's some really incredible people in there. What area of AI do you think is particularly overheated today? I do think neon labs without a clear vision for how they're going to impact the world. I do genuinely think there is some risk that they will turn out to be interesting research project, but not a viable company. If you're investing in my seat today, what part of the AI landscape would you say is under invested and most exciting? You can't say simulation. I fundamentally believe that for AI companies in the future, you have to have interesting data strategy. Do you have access to data that no one else has access to? Do you know how to collect data that is very hard to collect? When you see those opportunities, I would invest. Right now, aside from simulation, robotics is an obvious place where this has become the case. Obviously, robotics, there's a lot of money already going as I wouldn't say it's under invested, but I also do think it is a quite interesting area. I also do think, aside from the core robotics or AI space, the inference layer, but also chip layer, the hardware, I do actually think it's quite interesting, and it's a very hard area for people to crack into. But there are a couple of teams that have done, I think, an exceptional job in the recent months or years. I think they're quite interesting. Who do you think is all? The recently etched came out of the Durstelth, quite pollution their team. I think they're going to be exciting. Final one for you. What's the kind of thing that anyone's ever done for you? Now, I'm somebody who actually needed a lot of help throughout my career. I didn't come in knowing everything. Well, certainly I don't know everything now, but I also didn't come in as someone who was an obvious candidate. The one person I quickly called out was when I graduated from college, I moved to Palo Alto living in Sambaray's garage. I didn't have a job because I was trying to run a startup that doesn't really go anywhere. But that was a moment where I really felt lost. I could sense in the air that AI wave was coming. I realized that if you want to be a surfer, you need a wave that you can serve. And I want to make sure that when the AI wave is here, I want to be there to write it and I want to help create a wave in the first place to ensure my seat in it. But I had no research background. I don't know research during my undergrad, which is quite rare. If you're a PhD student, applicant, and have no research experience during your undergrad, unfortunately, it's very hard. So I actually messaged a bunch of people and there is one professor and a Stanford Mary Wooder. She's a theory professor. Happened to graduate from the same college as me. She replied, and I still don't know why I think was truly out of kindness and the fact that we're from same school. And she thought, well, okay, here's a student who is seeking advice. I at least spend in a half an hour with a student, she very graciously spent a four more than with me and just talking me through like how I should think about AI space or how I should think about research and she actually connected me with an initial set of people that I started to work with and learn from. So that initial set of people who came together to help give me advice and actually let me have a foot into this area of research. It really was not an obvious choice for them. I really didn't think I deserved it but that was the bad that they took. I think truly for truly for their own kindness and I'm very grateful that they did. June from Quant funds to simulated worlds to love. This has taken many different twists and turns but thank you so much for joining me. But before we leave you today, today I want to tell you about how the first AI law firm, Crosby, helped us close a big sponsor. As you know, some of the biggest companies in the world advertise on 20VC, my British dulcet tones clearly convert well. I was working to close this big sponsor and they wanted to get through legal review quite quickly to close the deal. Crosby turned red lines around in three hours and caught major issues that would have caused as serious problems in the future. Crosby combines AI, some of the best engineers in the world from companies like RAMP and Stripe and some of the best attorneys in the world from top 10 law firms. Customers get the best of both worlds and elite human attorney reviews every contract but they move incredibly quickly returning red lines in under four hours. They help the fastest growing companies like cognition, ramp and clay, close deals in hours, not weeks. Learn more at crossbe.ai/20VC. If you want to red line NDAs, MSAs, DPAs and any other procurement contracts faster, go to crossbe.ai/20VC. It's speed that you can really trust. While crossbe helps you stay shop on the numbers, ZeroHash helps you move them on chain. Uber had to move it, Shopify had to hold it, Airbnb had to settle it, money, movement stopped being a fintech problem. It became a software problem. And increasingly, God it's an AI problem too. ZeroHash is the infrastructure that makes global instant money movement seamless. One API integration for stablecoins, digital assets and modern payment rails, so builders can stay builders. ZeroHash powers, some of the world's largest enterprises and financial institutions, including Kalshi, Stripe, Morgan Stanley, Gusto and Interactive Brokers. If you're thinking about stablecoins, digital assets and the future of money is time to talk to the team at ZeroHash. Alphasense combines AI with one of the world's deepest libraries of market intelligence, including expert interviews, broker research, earnings, company filings and real-time news. Every answer is grounded in this incredibly trusted evidence and fully traceable to the original source which is so important. But the best part, they're building super analyst and always on AI analyst. So instead of starting your day with another search, you'll start with work you already done, your coverage monitored, the important development surface and your investment brief already waiting for you. That's alpha-sense.com/20VC.

Podcast Summary

Key Points:

  1. AI companies need a defensible data strategy; simulations of human behavior could become a high-value market, with sessions potentially costing $100 million.
  2. Jun Sung Park founded Simile, a company building a foundation model of human behavior to simulate individuals, sub-populations, and entire markets.
  3. Early work involved creating 25 AI agents in a town, using memory, planning, and reflection to simulate realistic human interactions, like self-organizing a Valentine's Day party.
  4. Simile focuses on causal mechanisms, not just predictions, using randomized control trials and behavioral data to understand how actions change outcomes, unlike traditional models trained on web text.
  5. Data collection prioritizes representative everyday people and asks deep questions about their life stories, not expert programmers.
  6. The company aims to move beyond surveys and tooling to simulate complex interactions, such as product launches or collective action problems like climate change.
  7. Ethical considerations are central; Simile views simulation as a tool for representing diverse viewpoints at scale, ensuring people's voices are heard in decision-making.
  8. Statistical significance for narrow populations is achievable with around 1,000 people, but larger samples allow for finer segmentation.

Summary:

In this podcast episode, Harry Stabbings interviews Jun Sung Park, founder and CEO of Simile, a company that creates simulations of human behavior to predict and shape future outcomes. Park explains that his inspiration came from a 2023 project where he populated a virtual town with 25 AI agents, each equipped with memory, planning, and reflection capabilities, allowing them to autonomously organize events like a Valentine's Day party. This work highlighted the potential of large language models to extract realistic human behaviors from training data, despite their usual focus on rational tasks like coding.

Simile's core mission is to build a foundation model of human behavior, representing people's values, preferences, and biases, rather than super-intelligent machines. A key challenge is data collection, which involves sourcing representative everyday individuals and using randomized control trials to understand causal mechanisms—what actually changes behavior—rather than just predicting outcomes. Park emphasizes that clients like Starbucks don't just want to know future sales dips; they want actionable insights on how to alter strategies.

The company aims to move beyond traditional survey tools to simulate complex interactions, from product launches to societal issues like climate change and democratic stability. Ethically, Park sees simulation as a way to amplify diverse perspectives in decision-making, ensuring that all voices are considered. He notes that while 1,000 people can provide statistical significance for narrow groups, larger datasets enable more detailed segmentation.

FAQs

Simile is a company creating a foundation model of human behavior that can simulate individuals, sub-populations, and entire ecosystems or markets. It focuses on representing people's values, preferences, and tastes, rather than super-rational intelligence like coding or math.

In 2023, the team created a game town with 25 AI agents that woke up, worked, had relationships, and self-organized a Valentine's Day party. This demonstrated that language models could extract realistic human behaviors, leading to the vision of simulating entire lived experiences.

Initially, they used markdown text files for memory, but to handle large-scale experiences, they introduced reflection—periodic processes where agents summarize memories into higher-level insights. This helps agents form personalities and points of view, shaping who they are.

Frontier models focus on super-rational tasks like coding and science, while Simile focuses on simulating human biases, mistakes, values, and preferences. Simile aims to represent the subjective half of the human brain, complementing rather than competing with these models.

Prediction alone isn't valuable unless it helps shape the future. Simile uses data from randomized control trials and A/B testing to understand causal factors, enabling clients to know what actions to take to change outcomes, not just what will happen.

Data collection is key, involving sourcing representative everyday people and asking the right questions to capture their core nature. This includes life stories and experiences, which are more valuable than just observational data.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.