The TWIML AI Podcast
The TWIML AI Podcast

Modeling Human Behavior with Generative Agents with Joon Sung Park - #632

Today we’re joined by Joon Sung Park, a PhD Student at Stanford University. Joon shares his passion for creating AI systems that can solve human problems and his work on the recent paper Generative Agents: Interactive Simulacra of Human Behavior, which showcases generative agents that exhibit believ

Featured Speakers

Jun Sung Park Guest

Topics Discussed

Episode Summary

Executive Summary: Stanford PhD student Jun Sung Park discusses how LLMs enable a new kind of believable, human-like agent research that was previously impractical with rules or manual authoring. The conversation focuses on generative agents, memory and retrieval design, emergent social behaviors in simulated worlds, and the open questions around whether models truly encode worldview or just replay training data.

Main Topics: Path into AI and Human-Centered Research (Priority: 4/5): Park explains how startup failure, proximity to Stanford, and exposure to HAI/HCI researchers shaped his interest in AI as a creative, human-centered discipline. Generative Agents as a New Research Opportunity (Priority: 5/5): He argues LLMs unlock believable agents that can behave in open worlds over time, a long-standing goal in AI and HCI that was difficult to achieve with earlier methods. Worldview, Common Sense, and Emergence in LLMs (Priority: 4/5): The discussion examines whether LLMs contain worldview or common sense, and how recent papers complicate claims of emergence versus predictable scaling behavior. Believability and Evaluation (Priority: 5/5): Park distinguishes between narrow, human-judged believability and broader collective-level behavior, emphasizing that evaluation may require both human studies and statistical modeling. Memory Architecture for Agents (Priority: 5/5): He outlines the memory stream, short-term retrieval, and the three-part retrieval scoring system: recency, relevance, and importance. Text-Native Agent Design and Implementation (Priority: 4/5): The system reasons in natural language, with text-based memory and planning, then translates decisions into game-engine actions at the last moment. Emergent Collective Behaviors (Priority: 5/5): The research shows early signs of information diffusion, relationship formation, and action coordination among 25 simulated agents, suggesting possible future work in large-scale social simulation.

Key Arguments: LLMs encode enough human behavioral patterns from broad internet-scale data to generate plausible agent behavior in context. Believability is a more realistic near-term benchmark than claiming full human equivalence or true intelligence. Memory is essential because prompting an agent with its entire life history is computationally inefficient and often unnecessary. A retrieval system combining recency, relevance, and importance can selectively surface the most useful memories for action and reasoning. Natural language is a powerful enough intermediate representation that much of the agent stack can be text-based. Observed social behaviors in multi-agent simulations may reflect both model priors from training data and genuinely emergent interactions, so evaluation must remain cautious. Future progress should borrow from information retrieval, social science, and computational modeling rather than relying only on LLM prompting tricks.

Data Points: Duration of startup attempt before pivot: about half a year - Park says he ran a startup with a co-founder for roughly six months before realizing it would not pan out. Time living near Stanford after college: about a year - He describes moving to Palo Alto and being exposed to Stanford’s research scene. Team size in the early generative agents demo: 25 agents - Park notes the demo simulated 25 NPCs interacting in a shared world. Simulation horizon in the study: 2 days - The reported emergent behaviors were observed over a two-day simulation. Number of emergent social behaviors highlighted: 3 - He names information diffusion, relationship formation, and action coordination. Context extension discussed: 1 million tokens - Park references a recent paper suggesting dramatically expanded context windows for LLMs. Model used in the project: GPT-3.5 Turbo / ChatGPT - He says the final implementation used ChatGPT, likely GPT-3.5 Turbo, after earlier work with GPT-3. Human-centered research horizon: 5 to 10 years - Park says he was drawn to people who were thinking ahead over a five- to ten-year time horizon.

Pivotal Quotes: "we believe that we can actually now start thinking about building these agents, building human, like a believable agents." — Jun Sung Park: Park summarizes the research takeaway: LLMs make long-standing agent goals newly feasible. "The more ambitious goal that we have in this line is... can these agents exhibit believable emergent community behaviors at scale?" — Jun Sung Park: He frames the next frontier beyond single-agent plausibility toward collective behavior. "the core medium that connects all different modules, memory and planning, reflection, everything, is everything can run in natural language." — Jun Sung Park: He explains the text-native architecture underlying the generative agents system.

Implications: LLMs may finally make believable agents practical for games, social simulation, and human-AI interaction. The big open problems now are memory, retrieval, evaluation, and separating learned priors from truly emergent behavior.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast