Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI

Solving Poker and Diplomacy, Debating RL+Reasoning with Ilya, what’s *wrong* with the System 1/2 analogy, and where Test-Time Compute hits a wall Full Video Episode Timestamps 00:00 Intro – Diplomacy, Cicero & World Championship 02:00 Reverse Centaur: How AI Improved Noam’s Human Play 05:00 Turi

Featured Speakers

Latent.Space HostNoam Brown Guest

Topics Discussed

Episode Summary

Executive Summary: Noam Brown discusses the evolution of AI reasoning, arguing that scalable thinking/test-time compute plus stronger base models are the real drivers of progress. He contrasts controllable systems like Cicero with today’s more capable O-series, explains why harnesses and routers may fade as models improve, and emphasizes that coding, remote work, and multi-agent systems are becoming major practical frontiers.

Main Topics: Cicero, Diplomacy, and human-machine play (Priority: 5/5): Brown reflects on building Cicero, how debugging the bot improved his own Diplomacy play, why the bot occasionally hallucinated, and why the game is a strong benchmark for multi-agent reasoning and persuasion. Reasoning models and the scaling of test-time compute (Priority: 5/5): He argues that the key paradigm shift is models thinking longer and better, not just larger pretraining, and says the O-series demonstrates broad gains across hard-to-verify domains. System 1 / System 2 and the limits of analogies (Priority: 4/5): Brown explains that extra thinking only helps once a model has sufficient base capability; system two is not independent of system one, and the effect varies by task and modality. Harnesses, routers, and product scaffolding (Priority: 4/5): He suggests many current scaffolds, routers, and workaround systems will be replaced as models become more capable, though they remain useful in the short term for product performance. Coding agents and the future of work (Priority: 5/5): Brown describes using Codex and Windsurf daily, says coding feels increasingly agentic, and predicts the technology will expand beyond software engineering into broader remote-work tasks and virtual assistants. Research strategy: data efficiency, RL, and OpenAI culture (Priority: 4/5): He highlights data efficiency as a major unsolved problem, defends OpenAI’s willingness to bet on large-scale experiments and RL, and credits the organization’s startup-like structure for enabling decisive pivots. Multi-agent systems, self-play, and civilization analogies (Priority: 4/5): Brown argues self-play works cleanly in two-player zero-sum games but is harder to generalize to diplomacy, math, and broader multi-agent settings, where defining the objective is much more ambiguous.

Key Arguments: Cicero improved Brown’s Diplomacy skill because studying the bot exposed new strategic ideas and deepened his understanding of the game. The main lesson from reasoning models is that performance gains come from combining strong base models with extra thinking, not from either alone. Reasoning is useful in non-verifiable domains: Deep Research is cited as evidence that models can produce valuable outputs even when success is subjective. Many current agent scaffolds, harnesses, and routers are temporary; as model capability scales, these layers should increasingly be absorbed into the model itself. Test-time compute has limits: it gets expensive, slows iteration, and eventually collides with wall-clock constraints, even if it continues to improve. RFT is valuable because it creates specialized data that remains useful as models improve, unlike brittle prompt scaffolding. OpenAI’s success came from recognizing promising research signals early and committing resources aggressively, rather than treating every experiment as a small isolated project. Self-play is powerful in zero-sum games because it converges toward minimax/GTO behavior, but outside that setting the objective becomes harder to define and can produce misaligned or unhelpful behavior. For remote-work tasks and coding, the most important near-term gains will come from better model experience, better integration, and fewer manual handoffs. Multi-agent research should be more principled and less heuristic; Brown thinks past approaches often ignored the bitter lesson.

Data Points: Cicero model size: 2.7B - Brown notes Cicero was a very small language model and that larger models helped substantially. World Diplomacy Championship year: 2025 - Brown says he won the world championship a few months before the interview. Cicero release year: Late 2019 - He references announcing Cicero in late 2019. Top human-player performance: Top 10% - He describes Cicero as reaching top 10% of human players. Reasoning model progression: 01 preview → 01 → 03 - He cites consistent progress across OpenAI’s reasoning model generations. AGI timeline reference: Late 2021 - He recalls discussing AGI timelines with Ilya in late 2021. Model thinking time today: 15 minutes - He says current models can think for roughly 15 minutes in some settings. Future thinking horizon discussed: Hours, days, even longer - He describes the multi-agent team’s goal of scaling test-time compute dramatically. Poker hidden states: 1,326 - He gives the number of possible heads-up Texas Hold’em hidden states. Poker exploitative sample size: ~10,000 hands - He contrasts AI sample needs with humans who infer style from a dozen hands. Human inference sample size: ~12 hands - He notes humans can often profile poker opponents very quickly. Reasoning-model lift in games: 1,000 to 100,000 times - He estimates the performance gain from models thinking before acting in some domains. Non-humanoid robotics examples: Drones - He uses drones as an example where humanoid form is unnecessary.

Pivotal Quotes: "I think harnesses are like a crutch that eventually we're going to be able to move beyond." — Noam Brown: On the future of scaffolding and environment wrappers around AI models. "They're geniuses, but it's their first day on the job." — Noam Brown: Describing the current limitations of coding agents and why they still need better experience and integration. "I think the best research is obvious in retrospect." — Noam Brown: On OpenAI’s scaling bets and the difficulty of recognizing major paradigm shifts early.

Implications: AI products will likely move toward fewer wrappers, more native reasoning, and broader autonomy across coding and remote work. Builders should invest in durable data, evaluation, and workflows that improve with model scaling rather than fight it.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast