No Priors
No Priors

The bot Cicero can collaborate, scheme and build trust with humans. What does this mean for the next frontier of AI? With Noam Brown, Research Scientist at Meta

AGI can beat top players in chess, poker, and, now, Diplomacy. In November 2022, a bot named Cicero demonstrated mastery in this game, which requires natural language negotiation and cooperation with humans. In short, Cicero can lie, scheme, build trust, pass as human, and ally with humans. So what

Featured Speakers

Noam Brown Guest

Topics Discussed

Episode Summary

Executive Summary: Noam Brown explains how poker and diplomacy became proving grounds for AI progress, why Cicero’s human-like natural-language negotiation was a breakthrough, and why reasoning and inference-time compute—not just bigger training runs—may be the next major frontier toward more general AI.

Main Topics: Why games became AI benchmarks (Priority: 5/5): Brown frames games as ideal research testbeds because they have clear rules, objective outcomes, and comparable human baselines, while noting the field is now moving beyond single-game mastery toward broader generality. Poker as a breakthrough in hidden-information reasoning (Priority: 5/5): He recounts how his PhD work on heads-up and six-player no-limit Texas Hold’em advanced AI through scaling, search, and equilibrium play, culminating in bot victories over top pros and influencing professional strategy. Diplomacy and Cicero as a leap into social intelligence (Priority: 5/5): Brown describes Diplomacy as a game of private negotiation, trust, deception, and alliance-building, and explains why creating a bot that could play it in natural language against humans was seen as science-fiction-level hard. Reasoning and inference-time compute as the next frontier (Priority: 5/5): He argues the major bottleneck is not just data or training size, but the need for better reasoning/planning systems that can spend more compute at inference time, analogous to search in AlphaGo and poker. Data limits, human behavior, and self-play (Priority: 4/5): Brown explains that Cicero combined human data with self-play: human data taught language and social norms, while self-play improved strategy—because supervised learning alone does not produce strong strategic behavior. Generalization, sample efficiency, and AGI (Priority: 4/5): He suggests humans still outperform AI mainly in sample efficiency and adaptability, and that truly general systems will need to reason across domains like math, code, negotiation, and potentially science. Research philosophy: take high-risk bets (Priority: 4/5): Brown emphasizes that impactful research requires aiming at hard problems, accepting failure risk, and avoiding safe, incremental work that may not move the field forward.

Key Arguments: Games are useful AI benchmarks because they provide objective evaluation and force researchers to confront difficult problems before the real-world equivalents. Poker showed that planning/search at inference time can matter as much as, or more than, scaling training compute; adding search improved performance by orders of magnitude. Diplomacy was chosen because it combines hidden information, coalition-building, deception, and natural-language negotiation—capabilities more relevant to real-world AI than chess or Go. Cicero’s success depended on blending human dialogue data with self-play; supervised learning alone could not learn strong strategy or human-compatible norms. The Turing test is becoming less useful as a measure because language models can already mimic human conversation well enough that detection is no longer the main issue. The key unsolved problem is reasoning: current models are strong pattern predictors, but they still lack robust planning and general-purpose deliberation. Data may not be the primary bottleneck; compute scaling and inference-time reasoning are more likely to constrain progress going forward. Humans still hold an edge in sample efficiency and generality, but that advantage may erode as models improve their reasoning and adaptation abilities.

Data Points: Poker hands on the river (Texas Hold’em): ~2.5 billion - Brown describes the combinatorial size of the game state that his team reduced via clustering/bucketing. Initial poker state buckets: ~5,000 - Early bucketed representations used to make poker computation tractable. Later poker state buckets: ~30,000 to 90,000 - He describes yearly scaling of the model’s discretized state space during grad school. Heads-up poker hands played in Brains vs AI: 80,000 hands - Competition between the bot and four expert human players. Impact of adding search in poker: ~100,000x improvement - Brown says inference-time search boosted bot strength by this approximate factor. Second poker competition prize pool: $200,000 - Incentive for top poker pros to play their best in the 2017 match. Six-player poker bot training cost: Under $150 - Brown emphasizes the algorithmic efficiency of the multiplayer poker breakthrough. Diplomacy training data: ~50,000 games - Data sourced from WebDiplomacy.net for dialogue and strategy learning. Diplomacy messages in dataset: ~13 million messages - Used to fine-tune language behavior in Cicero. Undetected human competitions: 40 games - Cicero reportedly went through the full set without being identified as a bot. Expected/possible model training cost today: ~$50 million to train current models - Brown uses this as a baseline for discussing compute scaling limits. Potential next-order training scale: $500 million to $5 billion - He speculates about likely near-future scaling of frontier models by large labs/governments. Potential post-scaling problem size: $100 billion model - Used rhetorically to illustrate the impracticality of indefinite scaling.

Pivotal Quotes: "We were trying to think of what would be the hardest game to make an AI for. We landed on diplomacy." — Noam Brown: Explaining why Diplomacy was selected as the Cicero target. "The most surprising thing was just honestly how it didn't get detected as a bot." — Noam Brown: Reflecting on Cicero passing through human games without being identified. "If you take out the planning that's being done in AlphaGo and just use the raw policy network ... it's actually substantially below top human performance." — Noam Brown: Arguing that reasoning/search is essential beyond neural network scaling.

Implications: For AI builders, the next leap may come from reasoning and inference-time search, not just bigger models. For industries, negotiation, planning, and strategic decision-making systems may arrive sooner than expected.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors