Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Simulation: the new Scaling Law — Joon Sung Park, Simile AI

When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, run

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: June’s path from artist to AI researcher to Simile founder centers on one thesis: to build behavior foundation models that simulate real people, not idealized LLM users. He argues that web-trained models capture attitudes but miss causal, real-world behavior, so Simile collects interview, observational, and RCT data to model individuals and populations for better decision-making in business and society.

Main Topics: From artist to AI researcher (Priority: 4/5): June recounts moving from Korea to Boston, growing up on the East Coast, training as an artist, and eventually discovering computation and research as a way to make art through a new medium. Generative agents and the origin of Simile (Priority: 5/5): He explains how the Smallville/generative agents work emerged from Stanford foundation-model research and a 'time machine' exercise aimed at identifying the most transformative application of LLMs. Why simulation matters more than prediction (Priority: 5/5): June distinguishes simulation from simple forecasting: the value is in understanding the path to an outcome and the interventions that change it, not just the final result. Behavior foundation models and data strategy (Priority: 5/5): He describes Simile’s data taxonomy: qualitative interviews, observational behavioral data, and causal/RCT data, arguing that true behavioral modeling requires all three. Validation, accuracy, and limitations of frontier LLMs (Priority: 5/5): The interview covers the 'Generative Agent Simulations of Thousand People' paper, reported 85% self-consistency/accuracy, and why frontier LLMs still underperform on niche, human-like behavior. Scaling, market use cases, and societal ambition (Priority: 4/5): June discusses scaling laws, enterprise use cases like concept testing and product research, and the longer-term vision of simulating large societies to inform problems like climate change and democracy. Company vision and hiring (Priority: 3/5): He frames Simile as both research lab and product company, with bi-coastal teams and active hiring across research, engineering, product, and infrastructure.

Key Arguments: LLMs trained on web data are good at attitudinal and text behavior, but they do not capture the deep, real-world behavioral 'social physics' needed for faithful simulation. The most useful simulation is causal: decision-makers usually need to know how to change a future outcome, not merely predict it. A minimal memory stack in text/markdown can work surprisingly well, but some behavioral traits require post-training or model-weight changes rather than prompting alone. High-fidelity simulation requires three data sources: rich interviews, observational behavioral data, and randomized controlled trials that expose causal mechanisms. Frontier models often become overly rational; Simile aims to model people as they actually are, including mundane habits, biases, and inconsistencies. Simulation is already valuable for organizations that historically rely on human panels because it can scale voice-of-customer input to decisions that are otherwise too expensive or slow to research. Long-term, large-scale simulations of societies could address wicked problems such as climate change, democratic collapse, and the origins of monetary systems.

Data Points: Google Scholar citations for Generative Agents paper: 72,000 - June discusses the reach and influence of the Smallville/generative agents paper. Reported simulation accuracy: 85% - In the 'Generative Agent Simulations of Thousand People' study, digital twins replicated participants' behaviors and attitudes at about 85% self-consistency/accuracy. Frontier-model human-behavior prediction on niche populations: 20-30% - June says frontier models can fall to 20-30% on more specialized populations or topics. Frontier-model human-behavior prediction on general population: 50-60% - For broader general-population tasks, he says frontier models may reach only 50-60%. Population sampled in Thousand People study: 1,000 people - The validation paper used a representative sample of 1,000 U.S. participants in a virtual app. Interview collection time: 2 hours - Participants were interviewed for roughly two hours to gather broad initial data. Follow-up gap before re-testing: 2 weeks - Human participants were sent away for two weeks before returning for surveys, experiments, and behavior studies. Current weekly data collection scale: tens of thousands of people's data - June says the company currently collects data at this weekly scale. Global panel partnerships: tens of millions of people - He says partnerships give Simile access to very large global panels. Current operational simulation scale: hundreds to hundreds of thousands - He says Simile’s deployed models cover populations from tens of thousands up to hundreds of thousands. Market research industry size: $100 billion - June cites this as a rough reference point, while arguing simulation is broader than market research. Company size: about 60 people - He describes Simile as a ~60-person company. Lab-mate share of company: 15%-20% - He says roughly 15% to almost 20% of the company are former lab mates. Customer example: enterprise use case: concept testing - June says concept testing is a core use case, where companies compare messages, products, or ideas.

Pivotal Quotes: "What we are really trying to get to at that point is: hey, can we actually create a simulation of 8 billion people living on Earth?" — June: Describing the long-term ambition of scaling from individual and population models to full-society simulation. "The reason why simulation is actually different from prediction. In simulation, in the ideal case scenario, ... you're showing the step function ... that results in a particular outcome." — June: Explaining why decision-makers care about pathways and interventions, not just outcome forecasts. "We're trying to create are models that are as dumb as I am, right? So, if I make such mistakes, the model has to make the same kind of mistake." — June: Contrasting Simile’s goal with frontier models that optimize for rationality rather than human-likeness.

Implications: The episode suggests a new AI category: behaviorally grounded simulation. If it scales, companies could replace expensive research panels, test interventions faster, and eventually model societal problems with far greater realism than today’s LLMs.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast