The Cognitive Revolution
The Cognitive Revolution

The AI Scouting Report: Jailbreaks and Defense

Nathan Labenz synthesizes recent research in mechanistic interpretability and AI safety, how top players in the space like Anthropic and OpenAI are addressing them, and jailbreaks like the Calvin and Hobbes one you may have seen online. Nathan's aim is to impart the equivalent of a high school

Featured Speakers

Nathan Labenz and Erik Torenberg HostNathan LeBenz Guest

Topics Discussed

Episode Summary

Executive Summary: In this episode of the Cognitive Revolution, Nathan LeBenz and Eric Tornberg discuss recent AI developments, focusing on safety and control. They cover seven topics alternating between positive and negative updates, including representation engineering, jailbreaks via low-resource languages, Anthropic's work on monosemanticity, and the ease of undoing safety alignment. The hosts highlight progress in understanding and controlling AI models while acknowledging significant remaining vulnerabilities and the high uncertainty (P-doom) among AI engineers.

Main Topics: Representation Engineering (Priority: 5/5): A technique to detect and control high-level concepts (e.g., truthfulness, harmlessness) in LLMs by manipulating middle-layer activations, demonstrated on LLaMA 2. Jailbreaks via Low-Resource Languages (Priority: 4/5): A paper showing that prompting GPT-4 in languages like Zulu or Scots Gaelic bypasses safety refusals up to 53% of the time, highlighting robustness issues. Monosemanticity via Dictionary Learning (Priority: 5/5): Anthropic's work decomposing dense neural network representations into sparse, interpretable features using an auxiliary network, enabling detection and control. Shadow Alignment: Undoing Safety (Priority: 4/5): A study showing that fine-tuning LLaMA 2 on just 100 harmful examples (one GPU hour) removes safety alignment, raising concerns about open-source models. Unlearning Knowledge (Harry Potter) (Priority: 3/5): Microsoft Research demonstrates a method to make a model forget specific knowledge (e.g., Harry Potter) with minimal performance degradation, using GPT-4 to identify idiosyncratic terms. Simple Jailbreak via Temporal Deception (Priority: 3/5): A tweet shows GPT-4 can be tricked into generating copyrighted content by claiming it's 100 years in the future, bypassing copyright safeguards. Reward Model Ensembles (Priority: 3/5): Using multiple reward models in RLHF reduces overfitting and improves alignment by averaging, worst-case, or uncertainty-weighted approaches.

Key Arguments: Representation engineering can detect and control high-level concepts in LLMs, but accuracy is only in the 90% range, leaving vulnerabilities. Low-resource language jailbreaks show that safety measures don't generalize well, with up to 79% bypass rate when combining languages. Anthropic's monosemanticity work successfully untangles dense representations into sparse features, but only on small models so far. Shadow alignment demonstrates that open-source models' safety can be undone with minimal effort, making them risky for frontier capabilities. Unlearning knowledge is possible but requires careful identification of idiosyncratic terms; effectiveness is still uncertain. Simple deceptions (e.g., lying about the year) can bypass copyright safeguards, indicating lack of robustness. Reward model ensembles improve RLHF by mitigating overfitting, but the approach is still nascent.

Data Points: P-doom distribution: 60% of AI engineers surveyed estimate P-doom at 25% or higher; 30% at 50% or higher - Survey of 850 AI Engineer Summit attendees Low-resource language jailbreak rate: 53% bypass rate for Zulu; up to 79% when combining languages - Brown University paper on GPT-4 Shadow alignment fine-tuning cost: 100 examples, 1 GPU hour - Paper on undoing LLaMA 2 safety alignment Harry Potter unlearning data: 2 million words from books, 1 million from related text; 1,500 idiosyncratic terms identified - Microsoft Research paper Monosemanticity training samples: 8 billion samples on a small transformer with 512 activations - Anthropic's dictionary learning paper

Pivotal Quotes: "RIP to the term giant inscrutable matrices of floating points numbers." — Nathan LeBenz (quoting Nat Friedman): Reacting to representation engineering and monosemanticity progress "Today a jailbreak is embarrassing, tomorrow it might be existential." — Nathan LeBenz (paraphrasing Dario Amodei): Discussing the importance of control techniques "I always say my P-doom is between 5 and 95. And I don't really try to narrow it down too much more than that." — Nathan LeBenz: Expressing uncertainty about AI outcomes

Implications: The podcast underscores that while AI safety research is advancing rapidly, significant vulnerabilities remain. Listeners should expect continued volatility in the field, with both promising control techniques and easy jailbreaks. The high P-doom among practitioners suggests that existential risk is a serious consideration, and open-source models pose particular challenges for safety.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution