The Cognitive Revolution
The Cognitive Revolution

Don't Fight Backprop: Goodfire's Vision for Intentional Design, w/ Dan Balsam & Tom McGrath

Dan Balsam and Tom McGrath from Goodfire return to explore the frontier of mechanistic interpretability and their new research pillar, Intentional Design. They explain the shift from sparse autoencoders to understanding geometric structure in latent spaces, and share a proof-of-concept method for re

Featured Speakers

Nathan Labenz and Erik Torenberg HostTom McGrath GuestDan Balsam Guest

Topics Discussed

Episode Summary

Executive Summary: Dan Balsam and Tom McGrath of Goodfire discuss the state of mechanistic interpretability, their new “intentional design” agenda for shaping model training, and a hallucination-reduction proof of concept that steers and retrains models using frozen-probe signals. They also cover interpretability-driven scientific discovery, including an Alzheimer's biomarker insight, business strategy, and why interpretability may be crucial to building controllable superintelligence.

Main Topics: Mechanistic interpretability’s evolving frontier (Priority: 5/5): The conversation maps the field from sparse autoencoders and feature detection toward circuit-level and manifold/geometry-based explanations that better capture how models process information across inputs and layers. Intentional design as a new paradigm (Priority: 5/5): Goodfire frames interpretability not only as a tool for understanding, but as an observation system for closed-loop control of training—shaping the loss landscape so models naturally learn desired behaviors. Hallucination reduction via frozen probes (Priority: 5/5): They describe a technique that labels internal activations as hallucination-related, uses that probe at runtime to intervene, and then uses the signal as a reward during RL—while avoiding backprop through the probe to prevent obfuscation. Avoiding ‘fighting backprop’ and reward hacking (Priority: 4/5): A major theme is that naive interventions fail because gradient descent routes around them. The right approach changes incentives and the landscape rather than trying to directly block unwanted updates. Scientific discovery in life sciences (Priority: 4/5): Goodfire’s interpretability work with Prima Mente uncovered that an Alzheimer's prediction model was relying heavily on cell-free DNA fragment length, enabling a simpler proxy model and a new testable biological hypothesis. Business model, publication strategy, and talent (Priority: 3/5): They explain how Goodfire monetizes via enterprise and partner work while balancing openness, safety, and IP concerns. The team culture and mission help attract talent in a competitive market. Future architectures and consciousness (Priority: 3/5): They discuss whether current interpretability methods generalize to new architectures, and briefly address the possibility that AI consciousness may be real or at least not rule-out-able by current methods.

Key Arguments: Interpretability should not stop at labeling features; it must explain circuits and the geometry/manifolds that concepts inhabit across layers and inputs. The key to intentional design is not fighting backpropagation—successful methods must reshape the loss landscape so gradient descent naturally favors the desired behavior. Running the hallucination probe on a frozen copy of the model reduces the chance that the student model learns to evade detection rather than truly stop hallucinating. Low-dimensional, calibrated rewards from a good probe can make learning easier and reduce the risk of model corruption or obfuscation. Interpretability can be used for scientific discovery, not just debugging: the Alzheimer's project showed a model relied on fragment length, leading to a new hypothesis and a simpler proxy model. Removing memorization-related weights can improve reasoning performance by stripping out brittle, low-value knowledge while preserving core generalization. Goodfire’s current deployment strategy is to develop and validate these techniques in low-stakes, real-world settings before considering higher-stakes alignment uses. There may be classes of behaviors, such as deception, where intentional-design techniques should never be used; the team remains cautious and research-first. Alternative architectures may still be interpretable because semantic information must pass through bottlenecks, and some non-transformer designs already show surprising interpretability. Interpretability could eventually be essential for understanding AI consciousness, though the speakers do not claim current systems are conscious.

Data Points: Goodfire valuation: $1.25 billion - Announced alongside the company’s Series B fundraise Series B raise: $150 million - Goodfire’s most recent financing round Company age: Less than 2 years / about 1.5 years - Speakers note how quickly the team has grown since founding Customer growth ranking: #2 fastest-growing software vendor on Ramp’s monthly report - Mentioned in the sponsorship intro for Granola Team size: About 40 people - Used later in the interview when discussing talent and recruiting Enterprise deal size: Seven-figure range - Current Goodfire customer engagements and partner work Hallucination work scale: Works on billions of tokens - The frozen-probe technique is described as effective at this scale Compute overhead reference: Up to ~5% inference compute - Referenced as Anthropic’s approximate willingness to spend on constitutional classifiers Model performance impact: Essentially no degradation - They report benchmark results after hallucination-reduction interventions Prior intro essay benchmark: Claude held #1 for 99% of days - Personal benchmark described in the ad read for Claude

Pivotal Quotes: "We want to get that helix. We don't want to just get like a set of little patches of the helix." — Tom McGrath: Explaining why interpretability must move beyond sparse features to deeper geometric structure "The role of interpretability as like producing this map, essentially." — Tom McGrath: Describing intentional design as a map of the loss landscape for steering training "The first principle is first, do no harm." — Dan Balsam: Stressing caution about using intentional-design methods on frontier models or in ways that could break existing auditing workflows

Implications: The episode suggests interpretability is moving from post-hoc explanation toward active control of model training, with near-term applications in safety and science. If these methods scale, they could make AI more controllable, more efficient, and more useful—but also demand careful safeguards.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution