Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The First Mechanistic Interpretability Frontier Lab — Myra Deng & Mark Bissell of Goodfire AI

From Palantir and Two Sigma to building Goodfire into the poster-child for actionable mechanistic interpretability, Mark Bissell (Member of Technical Staff) and Myra Deng (Head of Product) are trying to turn “peeking inside the model” into a repeatable production workflow by shipping APIs, landing r

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: Goodfire leaders Mark and Myra explain how interpretability is evolving from academic toy models into production systems for safety, customization, and scientific discovery. They discuss deployments like Rakuten PII detection, steering demos on trillion-parameter models, limits of SAEs, and a roadmap toward intentional model design, training-time interpretability, and life-science breakthroughs.

Main Topics: Goodfire’s mission and positioning (Priority: 5/5): Goodfire frames itself as an AI research lab centered on interpretability, aiming to understand, learn from, and design AI models rather than treat them as black boxes. Interpretability in production (Priority: 5/5): The guests emphasize moving interpretability out of toy-model research and into real-world deployment, especially in high-stakes domains like enterprise, healthcare, and legal/agent workflows. SAEs, probes, and method limitations (Priority: 5/5): They describe how sparse autoencoders and probes are used to identify concepts and behaviors, but also note surprising shortcomings where raw activations can outperform SAE-derived features. Steering and model customization (Priority: 4/5): A live demo shows steering a 1T-parameter Kimi model toward Gen Z slang; the discussion broadens to using steering for more substantive control, not just style. Post-training, safety, and hallucinations (Priority: 5/5): The conversation focuses on post-training failures such as sycophancy, reward hacking, hidden biases, and hallucinations, and how interpretability could help diagnose and mitigate them. Healthcare and scientific discovery (Priority: 4/5): Goodfire describes partnerships with Mayo Clinic, Arc Institute, Prima Menta, and Rakuten, using interpretability for PII protection, biomarker discovery, and debugging scientific models. Field growth, talent, and future research (Priority: 3/5): They highlight the growing interpretability community, open problems, and training programs like MATS, while calling for engineers and design partners across multiple domains.

Key Arguments: Interpretability should be treated as a practical AI engineering discipline, not just a post-hoc research curiosity. Production use cases are where interpretability proves value: guarding PII, detecting harmful behavior, debugging misgeneralization, and enabling safer customization. SAEs are useful but imperfect; in some real-world detection tasks, simpler probes on raw activations can outperform SAE-based probes. Interpretability can support intentional model design, including surgical edits to remove or modify behaviors without relying solely on brute-force fine-tuning or prompting. Steering is not just for stylistic control; Goodfire wants it to become a more powerful mechanism for changing substantive behaviors such as reasoning and legal expertise. Healthcare and science are especially compelling because models may reveal latent biological signals or help explain predictions in risk-sensitive settings. The field is moving from “can we understand models?” to “can we use that understanding to train, control, and deploy better models?”

Data Points: Series B funding: $150 million - Goodfire announced a new funding round during the episode. Valuation: $1.25B - The Series B priced Goodfire at unicorn status. Company size at join: first 10 employees - Myra said she joined when the company had about ten employees. Current team size: above 40 employees - Myra noted Goodfire had grown to more than 40 staff. Model scale in demo: 1 trillion parameters - The steering demo used the Kimi EK2 model at trillion-parameter scale. Office capacity: 200 people - The office was described as having room for 200, though the team was still much smaller at the time. Deployment timing with Rakuten: a few months ago - Goodfire said the Rakuten PII guardrail deployment went live a few months before the recording. Company age at first healthcare outreach: 3 months old - The team said healthcare institutions reached out when Goodfire was only about three months old.

Pivotal Quotes: "We like to say, is an AI research lab that focuses on using interpretability to understand, learn from, and design AI models." — Mark: Concise company description at the start of the interview. "Nobody knows what's going on, right?" — Vibu: Reaction to subliminal learning and the broader uncertainty about what models actually learn. "We don't want to be in the world where steering is only useful for like stylistic things." — Myra: Explaining that Goodfire aims for steering to enable substantive model control, not just cosmetic output changes.

Implications: Interpretability is maturing into a production-critical layer for AI: expect more tooling for model editing, safety monitoring, and scientific insight, especially in regulated and high-stakes domains.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast