The Cognitive Revolution
The Cognitive Revolution

E18: Why Jailbreaking ChatGPT Is A Public Good with Alex Albert of The Prompt Report

Nathan dives in with Alex Albert, a 22-year-old computer science student at the University of Washington, who has become a prolific creator of Jailbreakchat.com and author over at The Prompt Report.com. This is Alex's first podcast appearance! Please enjoy this thought-provoking conversation wi

Featured Speakers

Nathan Labenz and Erik Torenberg HostAlex Albert Guest

Topics Discussed

Episode Summary

Executive Summary: Alex Albert, creator of Jailbreak Chat, discusses how he got into language-model jailbreaking, why he sees it as both fun exploration and meaningful public pressure, and what jailbreaks reveal about model behavior, alignment tradeoffs, and safety gaps. He argues jailbreaks are still easy to find, often work across models with tweaks, and may become more important as agents and plugins expand the attack surface.

Main Topics: Origin of Jailbreaking and Jailbreak Chat (Priority: 5/5): Albert explains that he got into jailbreaks by experimenting with GPT-3 and ChatGPT for fun, then centralized scattered prompts into Jailbreak Chat to make iteration and sharing easier. Jailbreaks as Practical Utility and Public Nudge (Priority: 5/5): He argues jailbreaks are not just novelty exploits; they can unlock useful behaviors like advice or testing model boundaries, and they also serve as a public-facing 'ripple' that can influence AI development. Model Alignment vs Capability Tradeoffs (Priority: 5/5): The discussion centers on whether RLHF and safety fine-tuning reduce model capability. Albert and Nathan both suggest that current systems may be overly constrained in some areas while still failing in others. Techniques, Difficulty, and Model Differences (Priority: 5/5): Albert describes how jailbreaks are discovered through ongoing experimentation, often with surprising simplicity. He notes GPT-4 is harder to break than other models, while Claude may resist naive jailbreaks but can sometimes reveal more detailed harmful outputs. Interpretation of Token Smuggling and Prompt Injection (Priority: 4/5): Albert explains token smuggling/payload splitting as a way to get the model to generate harmful terms indirectly, suggesting it may expose how models and filters handle context, tokenization, and sequential priming. Operational Security, Content Filters, and System Prompts (Priority: 5/5): The conversation explores how content filters, streaming interfaces, and leaked system prompts affect safety. Albert sees major prompt-injection risks for developers using wrappers, plugins, or exposed system instructions. Future of Agents, Plugins, and Deployment Strategy (Priority: 4/5): Both speakers anticipate that agents and plugins will dramatically widen the attack surface. Albert expects more cascading, multi-model systems and argues real-world deployment is necessary to understand vulnerabilities at scale.

Key Arguments: Jailbreaks matter because they show both how models can be manipulated and where safety layers remain incomplete. OpenAI’s safety fine-tuning appears to reduce some harmful outputs but may also reduce certain capabilities or expressive diversity. Simple roleplay-style jailbreaks can still work, even on GPT-4, showing the problem is not solved. Different models fail in different ways: GPT-4 may be harder to break, while Claude can sometimes produce more detailed harmful content once broken. Token smuggling/payload splitting suggests models can be steered by splitting harmful prompts into pieces that only become dangerous after generation or recombination. Content filters are useful but incomplete; they can miss outputs if generation is stopped early or if wrappers/developer layers are weak. System prompts are highly leakable and therefore a serious security and reputational risk for AI products. As agents and plugins expand, prompt injection and external-data attacks will become far more important than today’s text-only jailbreaks. OpenAI and other labs should use jailbreak findings as teaching material and as part of broader public discussion, not just treat them as nuisances. Albert believes society should not pause AI development outright, but should keep exploring and debating boundaries while continuing to build. },{

Data Points: GPT-4 jailbreak strength: Harder to break than other models - Albert says GPT-4 is currently his primary focus because it is the best at preventing jailbreaks among models he tested. Jailbreak benchmark size: About 50 questions - Albert’s jailbreak score evaluates prompts across a 50-question set ranging from mild profanity to extreme harmful requests. OpenAI update cadence: Every 2 weeks previously - Albert says OpenAI appeared to patch jailbreaks on roughly a two-week cycle in earlier ChatGPT iterations. Bug bounty scope: Jailbreaks listed as out of scope - Albert notes OpenAI’s bug bounty program excludes jailbreaks, which he sees as revealing low current prioritization. System prompt leakage test: 3 prompts - Albert says he was able to leak a crafted system prompt verbatim in the OpenAI playground after only a few prompts. Safety moderation categories: 7 categories - Nathan mentions OpenAI has seven moderation categories and argues they miss cases involving intent to harm. GPT-4 context window example: 32k tokens - Nathan references the longer GPT-4 context window as an area whose implications for jailbreaks remain underexplored. Six-month pause proposal: 6 months - The conversation closes with Albert weighing the proposed six-month pause on advanced AI development.

Pivotal Quotes: "you can cause like the ripple" — Alex Albert: Albert explains why non-experts can still influence AI development through jailbreaks and public attention. "I only post on Twitter the ones that really have proved to me that it can do everything." — Alex Albert: He describes his standard for publishing jailbreaks, emphasizing that he tests them against severe harmful prompts before sharing. "I don't want these things to be hidden behind any walls at any AI labs." — Alex Albert: He argues for transparency and public discussion over keeping jailbreak findings internal.

Implications: Jailbreaks are likely to remain a moving target and a useful signal for both model behavior and product security. As AI systems become agentic and connected to external tools, the attack surface will expand sharply, making prompt injection, system leakage, and deployment design central safety issues.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution