Pioneers Insight Method Research Author
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Back to Episodes

Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI

Summary

  • Joon Sung Park’s headline claim is that simulation has its own scaling law: “The more data about humans and more compute you ingest, you start to get predictable gains in model performance” when simulating and predicting people. Simile post-trains its own models and is seeing “an early glimpse” of this curve.
  • The core moat argument: Park contrasts generative models’ emphasis on “super-rational, objective machines” with Simile’s need for models “as dumb as I am”—models that make the same mistakes humans make. Web-trained models contain fundamentally self-exposed attitudinal data, with some behavioral data sprinkled in, rather than the “dark knowledge of humanity”—what people actually do. On niche populations, Park says frontier-model behavior prediction falls to 20–30% (50–60% on the general population), versus Simile’s validated benchmark of replicating people’s attitudes and behaviors with 85% accuracy, “about as accurately as people could replicate their own.”
  • The validation asset is the “Generative Agent Simulations of 1,000 People” paper: 1,000 representatively sampled Americans, two hours of data collection, digital twins tested two weeks later against surveys, Big Five, behavioral-economics games, the General Social Survey and published RCTs. A follow-up showed post-training on tens of thousands of preregistered experiments from the Open Science Foundation platform delivers significant further gains—causal, randomized-controlled-trial data is the scarcest and most valuable input because “the world is our ground truth, but it happens once.”
  • The product pitch is simulation as a tool for shaping outcomes, not predicting them: “It doesn’t really help you to hear that your sales are going to tank in two quarters… What they want to know is, well, what do we need to do now to avoid that future?” His Foundation/psychohistory discussion—the counterintuitive first move of exiling the scientists to Terminus—illustrates why step-by-step causal simulation beats point forecasts, e.g., an EV marketing plan that lifts EV sales but makes overall auto sales go down.
  • On TAM, Park explicitly rejects the $100B market-research framing: “Simulation is not a tool for market research. Simulation is a tool for human decision-making.” Current deployment includes concept testing, focus groups, simulated earnings calls for public companies, and a Gallup strategic partnership; Simile collects data from tens of thousands of people weekly and has panel partnerships reaching tens of millions globally. Collected panelists are reusable across studies because traits like risk tolerance “don’t really change over time.”
  • Maturity marker for investors: Park says the simulation industry feels like where GPT-3.5 and GPT-4 were for the AGI saga—powerful enough to do real damage in current verticals, with aggressive scaling still ahead. His hunch is that simulations will eventually “cost as much as training a foundation model,” and of a society-scale climate-change run he says, “I would raise the money right now just to run that.” In a host exchange, swyx says an Africa UBI study returned “no” and floats a roughly $14M cost over five years; Park does not confirm the figure and questions whether implementation was the issue.
  • Park describes Simile as both a research lab and a product company: about 60 people, an SF headquarters at Mission Rock plus a new New York office, co-founded with Michael Bernstein, Percy Liang (who coined “foundation model”) and Lainie Ellen, with roughly 15–20% of headcount drawn from Park’s lab. His closing market frame: “You look at any advanced civilization in science fiction, and there are two twin-pillar technologies. One’s AGI in some form, and the other is simulation.”

Deep dive

1. From oil painter to the Smallville paper — via a “time machine game”

  • Park’s path: born in Korea, moved to Boston at 11, trained seriously as a realist oil painter—“it wasn’t a hobby, it was actually like: hey, let’s make a living out of this”—before deciding “the greatest artists often create their own medium, and the best medium that we had available today was actually computation.”
  • Starting his Stanford PhD in 2020, he joined the Percy Liang-led group that wrote “Opportunities and Risks of Foundation Models,” and fixated on what was genuinely new: a model that “wasn’t trained to do anything in particular” but, being trained on broad human data from the web, meant “if you poke at the right angle, then you could see human behavior that would just pop out” as quite realistic.
  • The founding exercise with future co-founders Michael Bernstein and Percy Liang: fast-forward ten years, look back—what’s the single application that mattered? “What if we could just recreate the world that we live in? It’s really hard to get more ambitious than that.” That led to Social Simulacra and then the 2023 Generative Agents (Smallville) paper—which Park thinks gets remembered for the wrong thing: “the memory component was pretty underrated… [a] very good early memory system.”

2. Why simulation must precede the personal-agent dream

  • The close runner-up in the time-machine game was personalized agents—but Park’s bet was sequencing: “if you were to create a really amazing personal assistant… what you actually need first is an amazing model of your users.” His deliberately dumb example: ask the agent to order dinner, it picks Hawaiian pizza, you hate pineapple—“it totally failed.”
  • His hot take, hedged as such: “I don’t think we’ve seen a true personal assistant that’s actually useful in ways that meet the ambition of that particular line of work… I don’t think we quite have all the right ingredients just yet.” On OpenClaw-style agents, he notes the Markdown-file memory is “quite clever”—the same intuition as Generative Agents in 2022 (“just put everything in a markdown file or a text file. You’re done”)—but retrieval over very large memories “takes a lot of work.”
  • His rule for when prompting stops sufficing: you need to touch the parameters “if the model has to learn the underlying physics of the world that it’s operating in… new social physics.” The models in the open “have not yet learned the complete mapping of the social physics of humanity”—the core Simile thesis.

3. The behavior foundation model: three data buckets, one scarce

  • Bucket one is rich qualitative interviews—literally “tell me the story of your life,” including “childhood memories, their trauma, or their first love… [which is] quite informative in ways that are really hard to predict.” Bucket two is observational behavior: transactions and web-scraped activity, giving base statistics.
  • Bucket three, which Park calls “perhaps the most important,” is causal-mechanism data from randomized controlled trials—the whys. It’s scarce by nature: “the world is our ground truth, but it happens once.” So Simile runs its own RCTs with consented, incentivized participants in virtual labs, borrowing social-science techniques: what makes data behavioral rather than attitudinal is “whether the stake in your decision is real”—for example, an experimental online store where purchases actually get delivered.
  • Park contrasts generative AI models’ emphasis on becoming “super-rational, objective machines” with Simile’s goal: “The models that we’re talking about here… are models that are as dumb as I am. If I make a mistake, the model has to make the same kind of mistake.” swyx responds: “You’re solving Moravec’s paradox.”

4. Simulation is for shaping the future, not predicting it

  • Vibhu’s West Wing example—the fictional multiple-sclerosis-disclosure poll where “we know it’s bad, we just don’t know how bad”—sets up the hosts’ real objection: if I can intuit the direction of an effect, why pay for a simulation? Park’s two answers are that magnitude and acuteness are genuinely hard to intuit (“every time somebody goes online and says something that has huge backlash, you look at that and think, ‘What an idiot.’ However, it’s tough”), and that simulation’s real product is the path, not the point estimate.
  • In a discussion prompted by Vibhu’s Foundation reference, Park’s signature analogy is psychohistory: the counterintuitive first move to compress 30,000 years of unrest into 1,000 is exiling the scientists raising the alarm to Terminus—“so counterintuitive… well, it turns out that in this particular simulation, that actually was the move.” You give the system a goal, not a survey question, and it returns the steps.
  • Translated to markets: an automaker optimizing only EV sales might find that the winning EV campaign “might change people’s perception of cars that aren’t EVs and actually make your overall sales go down.” swyx notes a similar idea in Shopify’s SimGym work: interventions on multi-turn shopping trajectories, not attitudes.

5. The grounding question: 85% self-replication accuracy, and where frontier models crater

  • The host challenge is whether the same goal could simply be given to Opus or GPT-4.5: how different would the answers be, and how do you check that the simulation is grounded? Park points to the “Generative Agent Simulations of 1,000 People” paper, which came out at the end of 2024: 1,000 representatively sampled Americans, two hours of wide-ranging data collection including interviews scripted from the American Voices Project plus behavioral data, digital twins built, then the humans brought back two weeks later for surveys, experiments, behavioral-economics games, Big Five, the General Social Survey and RCTs published in PNAS.
  • Result: twins “could replicate people’s behaviors and attitudes with 85% accuracy, about as accurately as people could replicate their own”—the first validated result that individuals could be modeled accurately. Frontier models give “the right foundation” but miss attitudinal and behavioral texture: performance drops “all the way down to 20% or 30%” on niche populations customers care about, and is around 50–60% for the general population—“you wouldn’t want to make your decision based on these kinds of findings.”
  • The follow-up paper exploited the replication crisis’s silver lining: preregistered studies on the Open Science Foundation platform—tens of thousands of professionally designed experiments with locked hypotheses—were used to post-train models with significant gains in behavior prediction. Simile trains two distinct models, population-level and individual-level, with the individual task described as “harder… in many ways.”

6. What models miss about humans, and which dataset Park would buy

  • Asked what humans do that models can’t, Park reframes the gap as “what biases or mistakes people make that models miss.” His example: living in Palo Alto, he preferred the 40-minute walk home over an Uber—“not for efficiency. It actually really helped me think… That’s a very human activity.” The goal is modeling what is “fundamentally human”: it “might not be the most efficient thing to do… but [these are] the things that make us who we are.”
  • Forced to rank LinkedIn, Twitter and Facebook as acquisitions, he picks Facebook: LinkedIn is people “with their guards up,” Twitter is full of “crazy personas,” but Facebook “is one of those more private spaces” and shows more of a person’s default self. Amazon transaction data is genuinely behavioral but “also most commonly available.”
  • On Tencent’s billion-persona paper, which cross-products professions and backgrounds as prompts, he admires the scale, but says that if the underlying statistics were sufficient, “then we have actually solved simulation”—you would merely retrieve knowledge already embedded in the model parameters. “That’s not, unfortunately, what we see”: niche, detailed knowledge of people requires bespoke collection and “paying attention to and respect[ing] the daily lives that people lead.”

7. The scaling law, Schelling’s dots, and a Nobel-sized ambition

  • Park says Simile is seeing “an early glimpse of scaling laws in simulation”—predictable performance gains from more human data plus compute. swyx responds, “We need a scaling walker.” Park: “Whenever you find one, it’s a beautiful thing.”
  • The ten-year vision replayed: “can we create a simulation of 8 billion people living on Earth?"—unlocking wicked problems such as climate change, “the signals for a collapsing democracy,” and “the origin story of the monetary system.” Park says he thinks “there’s a Nobel Prize to be won there,” and confirms the host’s specification: in economics.
  • His precedent is Thomas Schelling’s 1970s–80s segregation model: red and blue dots moving on a grid showed that even a minute same-color preference “causes society to segregate completely over time”—challenging the idea that explicit, overt racism alone explained segregation and informing mixed-income housing policy. Schelling later won the Nobel for laying groundwork for early simulations. Agent-based modeling faded because “red dots and blue dots are not really a rich description of people”; generative agents may restore the fidelity. swyx adds that Singapore’s public housing has enforced racial quotas, citing the relevance of this logic.

8. Cost economics: today tens of thousands of people, tomorrow foundation-model-sized runs

  • Current scale: rich insights from modeling thousands to hundreds of thousands of people, data collected from tens of thousands of people weekly, and panel partnerships reaching tens of millions globally. Crucially, panelists are reusable across subsequent studies—agents are domain-agnostic, and some traits are stable; Park says risk tolerance “doesn’t really change over time.” Larger samples matter less for statistical power than for filtering to “the right subpopulation of interest.”
  • swyx worries about combinatorial cost if thousands of agents talk to thousands of agents. He then argues that real-world studies are more expensive and often infeasible, while the decisions they inform can have hundreds of millions of dollars at stake. Park splits it: deploy by replacing existing budget, but “the way you capture the long-term value… is the upside” of better decisions worth hundreds of millions or billions.
  • His prediction, explicitly framed as a hunch: “in the next several years, we will start creating simulations that will actually cost as much as training a foundation model.” For a society-level climate simulation, he says, “I would raise the money right now just to run that.” In a host exchange about UBI, swyx says an Africa study returned “no” and asks whether it cost $14M; Park questions whether the issue was implementation. Vibhu’s caveat stands: sometimes people pay to verify what they think, not to simulate it.

9. TAM beyond market research, a GPT-3.5 moment, and the company behind it

  • Use cases today include concept testing, focus groups, behavioral experiments and A/B tests, plus simulated earnings calls. Walmart was an early customer interested in product testing beyond asking people what they think, requiring multimodal inputs such as images, Figma mockups and websites. Worldfront was an early customer excited by agents using a domain such as a website URL. Gallup is a strategic partner, but Park is deliberately holding back from politics until the company has enough “guardrails and perspective.”
  • What surprises him is deployment scale: organizations claim “we listen to people… but in reality, that is rarely the case.” Simulation can ensure “the voices of people are always represented in rooms where the decisions for them are made.”
  • On sizing: market research is a $100B industry, but “simulation is not a tool for market research. Simulation is a tool for human decision-making… What is a TAM for that? Really unclear.” His honest admission is that he never came in calculating it; he “just had to assume… that has to be big.” Maturity: “it feels a lot like where GPT-3.5 and GPT-4 were for the AGI saga.”
  • The company has about 60 people, an SF headquarters at Mission Rock plus a new NY office, and co-founders Park, Michael Bernstein (an ImageNet co-author, per Park), Percy Liang and business lead Lainie Ellen. Roughly 15–20% of staff are Park’s former lab mates, many of whom had gone to OpenAI, Google Gemini and similar organizations. The company is also hiring engineers on product and infrastructure, researchers, and others “across all sections.”
  • The closing register is the painter’s: “Simulation is a lot like painting… no painting is perfect… but it tries to highlight the thing that matters most about the subject”—the “essential essence.” On whether we’re in one: “whether we are in a simulation or not, I don’t think that makes our experience any less real… I worry about it when I die.”