The Cognitive Revolution
The Cognitive Revolution

Mechanistic Interpretability: Philosophy, Practice & Progress with Goodfire's Dan Balsam & Tom McGrath

In this episode, Daniel Balsam and Tom McGrath, at Goodfire, discuss the future of mechanistic interpretability in AI models. They explore the fundamental inputs like models, compute, and algorithms, and emphasize the importance of a rich empirical approach to understanding how models work. Balsam a

Featured Speakers

Nathan Labenz and Erik Torenberg HostTom McGrath GuestDan Balsam Guest

Topics Discussed

Episode Summary

Executive Summary: Dan Balsam and Tom McGrath argue mechanistic interpretability has moved from pre-paradigmatic to proto-paradigmatic: researchers now broadly agree models contain understandable features, circuits, and superposition, but major gaps remain in reconstruction quality and semantic labeling. They discuss improving sparse autoencoders, minimum description length, circuit-level work, and Goodfire’s applications in scientific discovery, safety, and creative tools.

Main Topics: State of mechanistic interpretability: from pre-paradigm to proto-paradigm (Priority: 5/5): The guests frame the field as having a shared emerging worldview: models contain understandable features, features are linearly decodable, superposition is real, and features compose into circuits. They stop short of claiming full consensus, but argue the field now has the raw materials of a paradigm. The two core gaps: reconstruction and labeling (Priority: 5/5): They separate interpretability’s main weaknesses into two gaps: how well methods reconstruct the model’s true computation, and how accurately humans or models label the recovered features. Both are still imperfect, especially for frontier models and abstract scientific representations. How to improve sparse autoencoders and dictionary learning (Priority: 4/5): The discussion covers improvements to SAEs via batch top-K, jump value, end-to-end SAEs, gradient pursuit, residual quantized autoencoders, and Matryoshka-style architectures. These aim to improve the sparsity-vs-reconstruction tradeoff and better capture structure in activations. Minimum description length as a better objective (Priority: 4/5): They present MDL as a more principled standard than sparsity alone: the best decomposition is the one that can be described in the fewest bits. MDL may favor tree-structured or cross-layer abstractions over flat bags of features. Circuits, weights, and causal abstraction (Priority: 4/5): They argue activation-based and parameter-based interpretability are complementary rather than competing paradigms. Both are attempts to build reduced causal models of neural networks, with circuits offering a higher-level causal abstraction than isolated features. Goodfire’s applied use cases: science, safety, and creative tools (Priority: 5/5): Goodfire is applying interpretability to genomics and other scientific models, inference-time guardrails for harmful content and policy monitoring, and creative interfaces such as image editing via sparse-autoencoder feature manipulation. Why interpretability matters for AI-enabled science (Priority: 5/5): They argue that as more scientific work moves to simulation and AI systems increasingly drive research, interpretability will be the preferred tool for understanding both the models’ discoveries and the models themselves.

Key Arguments: Interpretability should be treated like a natural science: progress depends on empirical observation of activations and model behavior, not only on top-down hypotheses. The field needs unsupervised methods because superintelligent and domain-specific models will not be interpretable through hand-crafted assumptions alone. Current SAE-type methods are useful windows into models, but they are reductive measurement devices, not the full underlying reality. The major bottleneck today is still algorithmic innovation and better inductive biases, not simply more compute. Reconstruction quality is still far from perfect; scaling may reveal a plateau or irreducible 'dark matter' that current methods cannot capture. Labeling features is inherently uncertain because a feature may reflect a real external concept, an internal algorithm, or both. Feature taxonomy matters: some features align with known inputs, others reflect model-internal computations, and others may correspond to unknown scientific structure. Interpretability and circuit work are essential for alignment because they provide the only realistic way to validate whether higher-level safety techniques are actually faithful. In scientific domains like genomics, interpretability can surface novel biological insights that existing human-designed tools miss. In product contexts, interpretability can enable cheap inference-time guardrails and better creative control than prompt-only interfaces.

Data Points: Series A raised: $50 million - Goodfire’s recent financing round, led by Menlo Ventures. Anthropic external investment: $1 million - Part of Goodfire’s Series A; described as Anthropic’s first-ever corporate investment in another company. Scale of one model trained with SAEs: Llama 3.3 370B - Goodfire trained sparse autoencoders on this frontier model. Reasoning-model coverage: DeepSeek R1 - Another model on which Goodfire trained sparse autoencoders. Enterprise trust base: 115,000+ enterprises - Mentioned in the Box AI sponsor readout, not the podcast topic itself. Interpretability community framing: Proto-paradigmatic - Tom’s characterization of mechanistic interpretability as emerging but not fully consensus-driven. Current interpretability confidence: Only roughly reconstructive - They emphasize that current techniques reconstruct model behavior only approximately.

Pivotal Quotes: ""We're like entering our first paradigmatic phase of interpretability."" — Tom McGrath: Tom’s refinement of the field’s status from pre-paradigmatic to proto-paradigmatic. ""There's no way that we're going to be able to scale to superintelligence without making our interpretability techniques unsupervised."" — Dan Balsam: On why unsupervised methods are necessary for future frontier models. ""If you have a thousand rules, maybe you can bunch them together in different checks. But at the end of the day, you run into the same problem."" — Dan Balsam: On why inference-time guardrails based on interpretability scale better than prompt-heavy or judge-model approaches.

Implications: Interpretability is becoming a practical infrastructure layer for alignment, science, and creative tools. The near-term challenge is better reconstruction, labeling, and scalable interfaces; the long-term bet is that understanding model internals will become essential to safely using advanced AI.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution