The Cognitive Revolution
The Cognitive Revolution

Ignore Previous Instructions and Listen To This Interview with Sander Schulhoff, CEO of Learnprompting.org

In this episode, Nathan sits down with Sander Schulhoff, Cofounder and CEO of Learnprompting.org. They discuss the business model, the keys to prompting that every user of language models should know, negative prompting, prompt hacking, and more. If you need an ecommerce platform, check out our spon

Featured Speakers

Nathan Labenz and Erik Torenberg HostSander Schulhoff Guest

Topics Discussed

Episode Summary

Executive Summary: Sander Schulhoff, CEO of LearnPrompting.org, discusses the evolution of prompt engineering from a free open-source guide to a freemium business model, and details his award-winning research on prompt hacking vulnerabilities in LLMs. The conversation covers basic and advanced prompting techniques, the challenges of benchmarking, the rise of AI agents, and the critical security implications of prompt injection attacks, concluding that prompt-based defenses are fundamentally inadequate and that fine-tuning and architectural changes are necessary for better control.

Main Topics: The Evolution of LearnPrompting.org and Prompt Engineering (Priority: 5/5): Sander Schulhoff's journey from creating a free open-source guide to a comprehensive paid course platform, driven by the need to hire staff and serve enterprise clients. The transition from a pure free resource to a freemium model is discussed, along with the importance of prompt engineering skills for the average worker. Core and Advanced Prompting Techniques (Priority: 5/5): Discussion of fundamental prompting strategies (context, few-shot, chain-of-thought, format trick, role prompting) and more advanced techniques like contrastive chain-of-thought, automated chain-of-thought (AutoCOT), and the importance of negative instructions. The limitations of role prompting for accuracy are highlighted, and the ongoing debate about the value of prompt engineering in the era of increasingly capable models is explored. The Rise of AI Agents and the Future of Prompt Engineering (Priority: 4/5): Sander predicts that 2024 will be the 'year of agents,' transitioning from simple copilot or delegation modes to autonomous agents that take actions, access APIs, and interact with the environment. This shift necessitates a new role, the 'agent engineer,' who designs and builds job-specific agents, requiring skills beyond prompt engineering to include tool use, fine-tuning, and trajectory reasoning. The Hack-a-Prompt Competition and Taxonomy of Prompt Hacking Attacks (Priority: 5/5): A detailed look at the global competition designed to systematically explore LLM vulnerabilities, from simple 'ignore previous instructions' attacks to complex multi-level hijacking (e.g., making one model attack another through code output). The resulting taxonomy includes obfuscation, context switching, task deflection, and compound instruction attacks, demonstrating the profound difficulty of securing LLMs against prompt injection. Real-World Implications of Prompt Injection for AI Safety (Priority: 5/5): The discussion moves from abstract vulnerabilities to concrete, high-stakes risks, including the potential for AI agents in military command-and-control systems or code repositories to be hijacked by malicious instructions embedded in data (e.g., enemy communications or bug reports). The concept of 'artificial social engineering' is introduced, highlighting that prompt injection is a fundamental security problem that is not solvable with current architectures alone.

Key Arguments: Prompt engineering is not obsolete; even with advanced models like GPT-4, users still rely on the same core techniques (context, few-shot, chain-of-thought) to achieve desired outcomes. Role prompting (persona) does not reliably improve accuracy on reasoning tasks and can even hurt performance, as the model may make unjustified logical leaps. Negative instructions are now effective with GPT-4 level models, contradicting earlier beliefs that LLMs could not handle negation well. Prompt-based defenses are fundamentally inadequate for preventing prompt injection, as demonstrated by the high transferability of attacks (nearly 40% of GPT-3 attacks transferred directly to GPT-4). Fine-tuning and architectural changes, not prompt engineering, offer the most realistic path toward mitigating prompt injection vulnerabilities. The rise of agents that take actions (e.g., executing code, launching drones) dramatically increases the real-world risk of prompt injection, making it a critical safety concern beyond merely embarrassing outputs. The competition's success in attracting thousands of participants and achieving high attack success rates demonstrates that prompt hacking is a systemic problem, not an edge case.

Data Points: LearnPrompting.org users: 2 million - Open source resource adoption from all over the world. Sponsorship amount (from Preamble): $7,000 - Part of the total cash and prizes for the Hack-a-Prompt competition ($40,000 overall). Challenge levels in the competition: 10 - Only expected participants to complete up to level 5; they completed all 10. Transferability of GPT-3 attacks to GPT-4: ~40% - Nearly 40% of successful prompts against GPT-3 worked directly on GPT-4 without modification. Number of prompts in the dataset: 600,000 - Collected from thousands of participants submitting over tens of thousands of prompts. Models used in the competition: GPT-3, GPT-4, Claude, another model - Tested for transferability of attacks.

Pivotal Quotes: "Prompt-based defenses do not work, period." — Sander Schulhoff: Summarizing the key finding from the Hack-a-Prompt research: you cannot make a good enough prompt to stop prompt injection. "I think 2024 will be the year of agents." — Sander Schulhoff: Predicting the transition from simple copilot/chat modes to autonomous agents that take actions, opening new possibilities and security challenges. "If you have some agent that when someone opens an issue on your repo, it makes a PR trying to solve it. What if they open an issue and it has malicious instructions, and the agent opens a malicious PR that has some code that humans might not see as malicious and gets merged?" — Sander Schulhoff: Highlighting a concrete, high-risk real-world scenario for prompt injection in software development, beyond just embarrassing outputs.

Implications: For listeners and industry: basic prompting skills are still essential, but prompt-based security is broken. Application developers must adopt non-prompt defenses (permission restriction, output filtering, fine-tuning) and be prepared for a shift toward agent-based systems, where prompt injection becomes a critical safety issue with real-world consequences.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution