The Cognitive Revolution
The Cognitive Revolution

Full-Stack AI Safety: Why Defense-in-Depth Might Work, with Far.AI CEO Adam Gleave

Today Adam Gleave, co-founder and CEO of FAR.AI, joins The Cognitive Revolution to discuss his cautiously optimistic vision for post-AGI futures and AI capability timelines across three distinct tiers, exploring the safety challenges and alignment techniques needed as FAR.AI scales from foundational

Featured Speakers

Nathan Labenz and Erik Torenberg HostAdam Gleave Guest

Episode Summary

Executive Summary: This episode features Adam Gleave, co-founder and CEO of FAR AI, discussing the path from current AI to transformative AI and how to navigate it safely. Gleave presents a cautiously optimistic view, envisioning a future where humans, while partially disempowered, enjoy high living standards akin to European nobility. He details FAR AI's unique vertically integrated approach spanning research, engineering, field-building, and policy. Key topics include capability thresholds, defense-in-depth strategies, scalable oversight (lie detectors), mechanistic interpretability, and red-teaming of current safety systems. Gleave argues that with rigorous implementation, current safety techniques can work, though risks from rushed deployment and emergent deceptive behaviors remain.

Main Topics: Post-AGI equilibria and positive visions (Priority: 5/5): Gleave envisions a future where humans are partially but not wholly disempowered, with high living standards but limited influence, analogous to minor European nobility. He contrasts this with negative scenarios involving arms races or AI welfare issues. Capability thresholds and timelines (Priority: 5/5): Distinguishes three tiers: powerful tool AIs (now), powerful agents (5-7 years), and organization-level AI (≈14 years). Discusses spiky capabilities and sample efficiency as key barriers to the final tier. Defense-in-depth and safety engineering (Priority: 4/5): Argues that current just-in-time safety measures are insufficient but that a well-engineered defense-in-depth with independent components and no attacker feedback could work. Cites FAR AI's red-teaming of existing stacks. Scalable oversight and deception (Priority: 4/5): Discusses FAR AI's lie detector project which reduced deception in models via scalable oversight. Highlights the risk of off-policy RL causing the model to learn to fool the detector. Mechanistic interpretability (Priority: 3/5): Gleave's nuanced view: interpretability is useful for coarse-grained auditing (e.g., detecting planning circuits) but unlikely to provide complete reverse-engineering guarantees. Cites FAR AI's work on planning algorithms in game-playing models. FAR AI's organizational strategy and hiring (Priority: 3/5): Gleave outlines FAR AI's strategy of covering the entire AI safety value chain from research to engineering to policy, and notes they are hiring (including a COO) to double their size.

Key Arguments: A well-implemented defense-in-depth approach to AI safety has a good chance of working if components are genuinely independent and attackers get no feedback on which layer triggered. Current AI safety systems are created on a just-in-time basis with implementation mistakes; a more deliberate, ground-up design could go much further. Scalable oversight techniques like lie detectors can train models to be more honest, provided the training regime is carefully chosen (off-policy RL with KL regularization) and not rushed. Mechanistic interpretability is valuable for coarse-grained auditing (e.g., detecting planning circuits) but unlikely to provide full reverse-engineered guarantees due to the messy, organic nature of neural networks. The most likely post-AGI scenario is not catastrophic loss of control, but gradual disempowerment where humans retain high living standards but limited power, similar to minor European nobility. The key risk is not that technical problems are unsolvable, but that developers will cut corners and patch problems without addressing root causes, potentially making deceptive behaviors worse. AIs are currently much less sample-efficient than humans (trained on vastly more data but still bad at many tasks), which provides a buffer against organization-level automation until architectures improve.

Data Points: Median timeline for organization-level AI: 14 years from 2025 (≈2040) - Timeline for autonomous AI systems that can outcompete human-led organizations Median timeline for powerful agents: 5-7 years from 2025 - Timeline for powerful autonomous agents (e.g., penetration testing chains) Reduction in reward hacking: Two-thirds reduction - Reduction in reward hacking behavior between Claude versions Reduction in scheming: Significant reduction - Significant reduction in scheming behavior in GPT-5 relative to GPT-3 Planned headcount increase: Double in size in 12-18 months - Planned organizational growth

Pivotal Quotes: "I think a good sort of historical analogy would be it's a bit like being European nobility, or perhaps like not the sort of the first heir to European nobility, being like, you know, the sort of third son or something. You've got this very nice living. You don't really have much purpose in life, but your life is pretty good." — Adam Gleave: Gleave's vision of the post-AGI human condition "It feels like we're sort of just running over a pretty small safety margin. And we might well lock out. It doesn't seem like any of the technical problems are insurmountable, but it's not a sort of very good place to be." — Adam Gleave: Describing the just-in-time safety approach and its dangers "Developers are going to find a problem, they're going to not very carefully patch it, they're just going to rush out of fix... and the problem is going to appear to disappear, but you actually made it worse." — Adam Gleave: On the risk of rushed safety patches making things worse

Implications: Gleave's analysis suggests that while current AI safety techniques can be effective if rigorously implemented, the industry's just-in-time approach and willingness to accept performance trade-offs will determine success. His timelines (powerful agents in 5-7 years, organization-level AI by ~2040) imply a narrow window for deliberate safety engineering. For listeners, this underscores the importance of supporting organizations spanning the full safety value chain, and the potential need for private regulatory bodies.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution