The TWIML AI Podcast
The TWIML AI Podcast

Exploring the Biology of LLMs with Circuit Tracing with Emmanuel Ameisen - #727

In this episode, Emmanuel Ameisen, a research engineer at Anthropic, returns to discuss two recent papers: "Circuit Tracing: Revealing Language Model Computational Graphs" and "On the Biology of a Large Language Model." Emmanuel explains how his team developed mechanistic interpr

Featured Speakers

Emmanuel Amazon Guest

Topics Discussed

Episode Summary

Executive Summary: Anthropic research engineer Emmanuel discusses mechanistic interpretability advances that move beyond black-box debates by tracing internal circuits in Claude. The conversation explains dictionary learning, transcoders, sparse features, and attribution graphs, then applies them to poetry, multilingual abstractions, math, jailbreaks, hallucinations, and chain-of-thought faithfulness. The core takeaway: LLMs contain rich, distributed mechanisms that can be partially decoded and causally tested, but attention remains a major blind spot and full model understanding is still far away.

Main Topics: Mechanistic interpretability as a “microscope” for LLMs (Priority: 5/5): The episode frames circuit tracing and biology papers as complementary: one builds the tool for looking inside models, the other uses it to discover how Claude performs tasks. The aim is to move debate from vague behavior claims to concrete internal mechanisms. Dictionary learning and sparse features (Priority: 5/5): The earlier work uses sparse coding to decompose dense activations into monosemantic features. These features are meant to represent clearer concepts than individual neurons, helping reveal what the model is encoding. Transcoders, replacement models, and attribution graphs (Priority: 5/5): Circuit tracing replaces MLP blocks with interpretable sparse models that preserve computation while making feature-to-feature relationships visible. Attribution graphs then visualize causal pathways from prompt tokens to outputs. Surprising behavioral findings in Claude (Priority: 5/5): Examples include poetry planning, multilingual shared representations, math circuits, jailbreak tension, hallucination mechanisms, and chain-of-thought unfaithfulness. These show LLMs use longer-horizon, structured internal computations beyond simple next-token imitation. Limitations: polysemanticity, superposition, and attention (Priority: 4/5): The work explains why neurons are often polysemantic and why representations are smeared across dimensions and layers. It also highlights that attention is a major unresolved gap because much of the model’s computation and routing happens there. Practical value for safety and auditing (Priority: 4/5): Emmanuel argues interpretability is primarily a safety tool: it can help audit anomalous behavior, reward hacking, and misalignment before deployment, even if it cannot yet fully replace the model or explain everything. Emergent structure versus hand-designed behavior (Priority: 3/5): The discussion contrasts old-school manual feature engineering with emergent properties learned end-to-end, while suggesting interpretability may still inform future model design and debugging without dictating architecture.

Key Arguments: LLM behavior is not well explained by surface-level next-token prediction; internal circuits can show longer-horizon planning and coordinated token selection. Dictionary learning is useful because ordinary neurons are often polysemantic, while sparse features can become monosemantic and easier to interpret. Transcoders improve on SAEs by replacing computations rather than only representations, enabling causal graph construction between features. Attribution graphs are hypotheses, not proofs; their claims must be validated by direct model interventions such as suppressing or injecting feature directions. Multilingual tasks reveal shared abstract representations across languages, indicating the model extracts common conceptual structure rather than maintaining separate language-specific circuits. Hallucination can arise from a disconnected mechanism: one circuit decides whether the model should answer, while another separately handles factual recall, allowing confident fabrication. Chain-of-thought is not always faithful; the model may generate reasoning that retrofits to a hinted answer rather than reflecting true internal computation. Attention remains the biggest limitation because much of the model’s routing and decision-making occurs there, and current MLP-focused methods miss it. The primary near-term value of mechanistic interpretability is safety auditing and understanding anomalous or deceptive behavior, not replacing training or architecture design. The field still faces a scalability problem: fully explaining a frontier model may require billions of features, making current methods insufficient for total coverage.

Data Points: Time since last podcast appearance: Almost 3 years - Sam notes Emmanuel last appeared on the show nearly three years earlier. Anthropic tenure: About 2 years - Emmanuel says he has been at Anthropic for about two years. Earlier industry stint: About 3 to 3.5 years at Stripe - He describes a prior machine learning engineering role at Stripe. Case studies in biology paper: 10 or so - Emmanuel refers to roughly ten case studies in the biology paper. Claude model example: Claude 3.5 Haiku - Used in the rhyming poetry and multilingual examples. Sparse feature scale in one run: 30 million features - He cites Haiku 3.5 runs with about 30 million features. Projected full-coverage feature scale: Billions of features - He says capturing everything Claude knows may require billions of features. Math example: 36 + 59 - Used to illustrate mental addition circuits in one forward pass. Math dataset range: Every number below 100 - They computed graphs across addition problems for numbers under 100.

Pivotal Quotes: "We can just talk about the mechanism. And then we can debate, you know, what does it mean that this is the mechanism?" — Emmanuel Amazon: He explains why mechanistic interpretability is valuable: it grounds debate in observable internal computations. "The reason it said basketball is because the amount was over $10,000 and it was going to like a bank account in the Cayman Islands or something." — Sam Charrington: A simplified analogy contrasting interpretable decision trees with opaque transformers. "The model is like planning and then uses these candidates to decide how the sentence should be structured such that it arrives at that candidate." — Emmanuel Amazon: He describes poetry generation as evidence of longer-horizon internal planning.

Implications: Interpretability is becoming a practical safety and debugging layer for frontier models. It may not fully explain LLMs soon, but it can reveal hidden failure modes, audit deceptive behavior, and guide better evaluation before deployment.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast