The Cognitive Revolution
The Cognitive Revolution

GELU, MMLU, & X-Risk Defense in Depth, with the Great Dan Hendrycks

Join Nathan for an expansive conversation with Dan Hendrycks, Executive Director of the Center for AI Safety and Advisor to Elon Musk's XAI. In this episode of The Cognitive Revolution, we explore Dan's groundbreaking work in AI safety and alignment, from his early contributions to activat

Featured Speakers

Nathan Labenz and Erik Torenberg HostDan Hendricks Guest

Topics Discussed

Episode Summary

Executive Summary: Dan Hendricks argues that AI progress is still being driven mainly by scale, data, and compute rather than novel architecture, while safety work should focus on practical defense-in-depth measures. He discusses benchmarks like MMLU, MATH, Machiavelli, and Humanity’s Last Exam, plus representation engineering, circuit breakers, tamper resistance, forecasting tools, and governance strategies for open models and geopolitics.

Main Topics: Scaling vs. architecture in AI progress (Priority: 5/5): Hendricks says architecture advances have slowed, with most gains coming from more data and compute, not clever model design. He views newer ideas as incremental compared with scale, though he remains open to surprises. Benchmarking intelligence and capability (Priority: 5/5): He explains why MMLU and MATH mattered, why they are nearing saturation, and how Humanity’s Last Exam aims to replace them with harder expert-level questions. He also discusses Machiavelli for agentic behavior. Representation engineering and interpretability (Priority: 4/5): He contrasts top-down representation engineering with bottom-up mechanistic interpretability, emphasizing that representations are a more useful unit for monitoring, steering, and safety interventions than individual neurons. Circuit breakers and tamper-resistant safeguards (Priority: 5/5): He describes circuit breakers as representation-level refusal mechanisms and tamper resistance as a way to make those safeguards hard to remove via fine-tuning, especially for open-weight models. Robustness, jailbreaking, and safety engineering (Priority: 5/5): Hendricks argues robustness remains difficult, especially in vision, but LLM jailbreak resistance looks more promising than many expected. He favors layered defenses over proofs of safety. AI governance, open weights, and geopolitical risk (Priority: 4/5): He discusses export controls, chip resilience, China-US competition, and the need for practical national competitiveness and safety policies rather than unrealistic coordination fantasies. Forecasting, public epistemics, and AI agents (Priority: 3/5): He highlights AI forecasting as a tool for improving decision-making and public epistemics, and suggests future AI systems should be governed by legal-style standards such as reasonable care.

Key Arguments: Activation functions like JELU/CiLU are performance-driven smoothing of the ReLU kink, not the result of a deep theory; they emerged because smoother nonlinearities often work better in practice. Most of the last several years of progress have come from scaling data and compute, with only modest architectural gains such as mixture-of-experts. MMLU endured because it measured knowledge breadth and difficulty better than prior linguistic benchmarks; once models reached common-sense and undergraduate-level knowledge, harder tests became necessary. Humanity’s Last Exam is intended to test expert-level knowledge that exceeds ordinary benchmark difficulty and should remain useful even after MMLU saturates. Machiavelli is meant to measure whether future AI agents act deceptively, selfishly, or prosocially in game-like settings. Representation engineering is more useful than neuron-level mechanistic interpretability for practical safety work because it targets higher-level concepts and can be used for monitoring and control. Circuit breakers can significantly improve jailbreak resistance by disrupting harmful representations, but they still trade off against some performance and are not yet fully production-ready when combined with tamper resistance. Open-weight models can be valuable, but if they become expert-level in dangerous domains like virology, releasing them could materially increase bioweapon risk. AI safety should rely on defense-in-depth: input filters, circuit breakers, monitoring, and KYC-style controls, rather than hoping for provable guarantees. AI forecasting systems already appear competitive with human forecasters and could improve public epistemics if integrated into products and institutions. Export controls and domestic chip manufacturing are viewed as relatively low-regret responses to geopolitical risk, though coordination with China is seen as difficult. Truth-maxing alone is not obviously the right objective for powerful AI; instead, systems should be designed around broader societal risk management and reasonable care.

Data Points: JELU paper date: 2016 - Hendricks says his activation-function paper was his first paper and remains highly cited. MMLU subjects: 57 - He describes MMLU as spanning many domains including history, calculus, and professional law. MMLU benchmark performance ceiling: ~90% - He says the 95th percentile of test takers across subjects corresponds to roughly 90% overall accuracy. MMLU current model range: high 80%s - He says frontier models have been stuck around the high 80% range since GPT-4. MATH dataset size: 12,500 questions - He cites MATH as a competitive math benchmark created after MMLU exposed weakness in STEM. MMLU math training examples: 7,500 examples - He notes the dataset format itself may be learned from training-set examples. Open-model compute-capability correlation: ~95% correlation - He says benchmark performance across open-weight models correlates strongly with compute. Ravens / compute correlation: ~80% correlation - He says certain pattern-recognition benchmarks track compute strongly as well. Known training scale threshold: No model above 10^26 FLOP (per his understanding) - He contrasts this with GPT-3/GPT-4 training around 10^25 FLOP. GPT-4-era training scale: ~10^25 FLOP - Used as a reference point for current scaling discussions. Circuit-breaker reliability gain: from 90% to 99% reliability (illustrative) - He frames circuit breakers as potentially adding a “nine” of reliability. Jailbreak competition attempts: 20,000+ attempts - He says models with circuit breakers had not yet been successfully jailbroken after thousands of attempts. GPQA sample size: ~500 questions - He argues it is too small for stable comparisons. Humanity’s Last Exam prize pool: $500,000 to $1,000,000 - He says the planned competition will pay cash prizes for high-quality hard questions. Forecasting bot speed/cost advantage: 10,000x–100,000x faster and cheaper - He characterizes AI forecasting systems as vastly more efficient than prediction markets or human forecasters. Export-control time horizon: this decade - He discusses China-Taiwan risk and AI competitiveness within the current decade.

Pivotal Quotes: "it's basically like an undergraduate-level knowledge and skill test" — Dan Hendricks: His characterization of MMLU after discussing its breadth and continued relevance. "I think basically the main factor driving these advancements is compute and data" — Dan Hendricks: His summary of what has actually powered recent AI progress. "if they get expert-level virologists, I don't know if I want that being released" — Dan Hendricks: His warning about open-weight models becoming too capable in dangerous bio domains.

Implications: Listeners should expect AI progress to keep favoring scale while safety will depend on layered practical controls, better benchmarks, and governance. Open-weight releases, jailbreak resistance, and geopolitical chip strategy will increasingly shape how fast and how safely frontier AI reaches expert-level capability.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution