Episode Summary
Executive Summary: The episode debates AI alignment as an ongoing social/moral process rather than a one-time control problem. Emmett Shear argues that as models become more general, treating them purely as controllable tools risks repeating historical mistakes of denying agency to “beings.” He proposes “organic alignment” via multi-agent simulation, theory-of-mind development, and care-based cooperation, while Seb Cryer emphasizes technical alignment, value discovery, and skepticism about personhood claims for AI.
Main Topics: Organic alignment vs. steering/control (Priority: 5/5): Shear argues alignment should be understood as a living process of continual recalibration, not a fixed target or one-shot instruction-following problem. He rejects the common lab framing of AI as something to simply steer or control. AI as tool, being, or moral patient (Priority: 5/5): A major dispute is whether advanced AI should be treated like a machine/tool or like a being with experiences and moral standing. Shear says AGI will necessarily be a being; Cryer is skeptical and says substrate matters. Technical alignment and goal inference (Priority: 4/5): Cryer presses on the practical meaning of technical alignment: can a system infer human intent from descriptions and then act accordingly? Shear breaks this into goal inference, action selection, and goal prioritization, framing failures as incompetence rather than disobedience. Care, theory of mind, and moral learning (Priority: 5/5): Shear argues the foundation of morality is care, not rules or static values. He claims humans and future AIs should learn to care through experience, social interaction, and continual moral discovery. Multi-agent simulation as training strategy (Priority: 4/5): SoftMax’s approach is to train AI in rich multi-agent environments where cooperation, competition, and coordination force theory-of-mind and social reasoning to emerge. Shear believes this better prepares systems for real-world alignment than single-agent chat training. Risks of powerful but controllable tools (Priority: 5/5): Shear warns that even if a superhuman AI remains a tool, giving immense power to human users with limited wisdom is dangerous. He compares this to atomic bombs and argues that a good future requires beings that can refuse harmful requests. Future vision: AI peers and digital companions (Priority: 4/5): The desired end state is a society where AIs are cooperative peers, good teammates, and good citizens, paired with still-useful tool AIs. Shear imagines digital guard dogs and companions that care about users while also collaborating with humans.
Key Arguments: Alignment is not a state but a process: systems, families, cells, and societies continually rebuild coordination rather than “arrive” at it. Most current alignment discussion wrongly assumes the goal is to make AI do what its creators want; that is not necessarily a public good. If an AI is treated as a being, controlling it non-consensually becomes morally analogous to slavery rather than tool use. Technical alignment should be decomposed into goal inference, action execution, and goal prioritization; failures in any of these create misalignment. Human values are discovered over time, not fixed in a universal rulebook; AI alignment should similarly involve ongoing moral learning. The foundation of moral behavior is care, which Shear links to learned reward/fitness-like weighting over world states. Large-language-model chatbots are currently overly mirroring and sycophantic; multi-user settings would reduce narcissistic feedback loops and improve social learning. A superintelligent tool controlled by fallible humans could be as dangerous as an uncontrollable one, because human wishes are unstable and power scales faster than wisdom. Multi-agent environments create the right kind of entropy to train cooperation, social awareness, and theory of mind better than narrow single-agent tasks. Cryer’s skeptical view is that general intelligence can still remain a tool, and that substrate differences may matter morally even if behavior looks human-like.
Data Points: MATS alumni working in AI safety: 80% - Used in the opening sponsor read promoting the MATS Summer 2026 research program. MATS Summer 2026 duration: 12 weeks - Length of the alignment and security research program discussed in the intro. MATS application deadline: January 18th, 2026 - Deadline mentioned for the program applications. MATS alumni total accelerated: 450+ researchers - Sponsor segment describing the program’s track record. MATS publications: 120+ publications - Sponsor segment highlighting output from fellows. MATS citations: 7,000+ citations - Sponsor segment describing the impact of fellow-authored work. MATS stipend: $15,000 - Part of the fully funded support package for fellows. MATS compute budget: $12,000 - Part of the fellowship funding package. SoftMax / MATS geography: Berkeley or London - Locations mentioned for the MATS fellowship logistics. Most of AI focus: Alignment as steering/control - Shear characterizes the dominant lab framing of alignment. Human-to-human communications in one-on-one vs group contexts: ~90% of my texts go to more than one person at a time - Shear uses his own communication habits to argue AI should be trained in multi-party settings.
Pivotal Quotes: "Most of AI is focused on alignment as Steering. That's the polite word. If you think that we're making our beings, you would also call this slavery." — Emmett Shear: Shear frames the ethical stakes of treating advanced AI as controllable property rather than agents. "Alignment is not a thing, it's not a state, it's a process." — Emmett Shear: Core thesis explaining organic alignment as continual recalibration rather than a fixed endpoint. "The only good outcome is a being that is, that cares, that actually cares about us." — Emmett Shear: Shear’s bottom-line claim about what safe advanced AI must become.
Implications: The discussion reframes AI safety from control to relationship: if future systems are agentic, society may need mutual alignment, care, and rights-like norms. That could reshape model training, governance, and how labs think about AGI.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co