Lenny's Podcast
Lenny's Podcast

AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff (Learn Prompting, HackAPrompt)

Sander Schulhoff is the OG prompt engineer. He created the very first prompt engineering guide on the internet (two months before ChatGPT’s release) and recently wrote the most comprehensive study of prompt engineering ever conducted (co-authored with OpenAI, Microsoft, Google, Princeton, and Stanfo

Featured Speakers

Lenny Rachitsky Host

Topics Discussed

Episode Summary

Executive Summary: The conversation argues that prompt engineering remains highly valuable, especially for productized AI systems, and highlights practical techniques that improve model performance: few-shot prompting, decomposition, self-criticism, rich additional context, and ensembling. It then pivots to AI red teaming and prompt injection, explaining why current defenses are incomplete, why agentic AI raises the stakes, and why security will remain an arms race requiring model-lab-level fixes rather than prompt-only guardrails.

Main Topics: Prompt engineering is still essential (Priority: 5/5): Sander argues the field is not dead; model improvements have not eliminated the need to communicate effectively with LLMs, especially when robustness matters. Basic prompting techniques that work (Priority: 5/5): The episode breaks down practical methods like few-shot examples, decomposition, self-criticism, and adding relevant context to get better outputs. What does not work anymore (Priority: 4/5): Role prompting and reward/threat-based prompting are presented as largely ineffective for accuracy tasks on modern models, though roles still help stylistic writing tasks. Product-focused prompt engineering (Priority: 4/5): Sander distinguishes casual chatbot prompting from engineering prompts used in products at scale, where prompt quality affects thousands or millions of requests. Prompt injection and AI red teaming (Priority: 5/5): The discussion explains how adversaries trick models into unsafe outputs and why crowd-sourced red teaming is useful for discovering jailbreaks and attack patterns. Why AI security is harder than classical security (Priority: 5/5): Unlike software bugs, prompt-injection vulnerabilities cannot be fully patched; defenses like prompt instructions and simple guardrails are insufficient against adaptive attacks. Agentic AI as the looming risk (Priority: 5/5): The biggest concern is not just harmful text generation but autonomous agents and robots taking harmful actions in the real world if compromised or misaligned.

Key Arguments: Prompt engineering remains relevant because better prompting can dramatically change model performance, especially in product settings where reliability matters. Few-shot prompting is one of the highest-leverage and most accessible techniques: showing examples often outperforms vague instructions. Decomposition helps by forcing the model to identify subproblems before solving a complex task, which improves both reasoning and orchestration. Self-criticism is a practical loop: ask the model to review its own answer, then apply the critique to improve the response. Adding rich additional information about the task, domain, or company can materially improve outputs, though it increases cost and latency in production. Role prompting is largely ineffective for accuracy-based tasks on modern models, but can still help with expressive tasks like writing or summarization. Threats and bribes in prompts do not reliably improve performance on current models. Prompt injection is a serious security issue because attackers can use obfuscation, typos, or social-engineering style prompts to get models to reveal dangerous information. Prompt-only defenses and generic guardrails are not enough; they can be bypassed through intelligence gaps or adversarial encoding. The most effective mitigations are safety tuning, fine-tuning for narrow tasks, and protections built into the underlying AI provider rather than bolted on externally. The largest future risk is agentic AI: systems that can take actions, access tools, and affect the real world can cause harm if manipulated or misaligned. This is not a fully solvable problem in the traditional security sense; it is an ongoing mitigation and arms-race problem.

Data Points: Prompt Report length: 76 pages - Sander led the team behind the Prompt Report, described as the most comprehensive study of prompt engineering. Papers analyzed: 1,500+ - The Prompt Report surveyed a large body of research on prompting techniques. Prompting techniques identified: 200 - The report synthesized 200 different prompting techniques. Accuracy boost from better prompting: ~70% - In a medical coding project, improved prompting and examples increased accuracy by about 70%. Role prompt result delta: 0.01 apart - In studies of role prompting, measured differences were so small they were effectively negligible. Hack a Prompt competition data set: 600,000 prompt injection techniques - Sander’s AI red-teaming competition collected a very large set of adversarial prompts. EMNLP paper ranking: 1 out of 20,000 submissions - The red-teaming work won best theme paper at EMNLP, which had about 20,000 submissions. Compromised global GDP via Stripe: 1.3% - A sponsor readout cited Stripe processing 1.3% of global GDP last year. Stripe processing volume: $1.4 trillion+ - That 1.3% of global GDP translated to more than $1.4 trillion flowing through Stripe. Experiment revenue lift example: 23% - Forbes reportedly saw a 23% revenue lift six months after switching to Stripe for subscription management. Payment authorization lift example: 4% - Hertz boosted online payment authorization rates by 4% after migrating to Stripe. Company count using Vanta: 9,000+ - Vanta said it helps over 9,000 companies with security and compliance. Annual discount mentioned: $1,000 off - Listeners were offered $1,000 off Vanta at vanta.com/Lenny. Prompt injection security target Altman cited: 95% to 99% - Sam Altman reportedly said models might get to 95–99% security against prompt injection.

Pivotal Quotes: "Prompt engineering is absolutely still here." — Sander Schulhoff: He rejects the idea that prompt engineering has become obsolete as models improve. "You can patch a bug, but you can’t patch a brain." — Sander Schulhoff: He explains why AI security is fundamentally harder than classical software security. "It is not a solvable problem." — Sander Schulhoff: His view on prompt injection/security: it can be mitigated, but not fully eliminated.

Implications: For users, better prompting still pays off immediately. For builders, AI security must be treated as an ongoing adversarial problem, especially as models become agents with real-world actions. For the industry, the key shift is from prompt hacks to deeper model and system-level defenses.

🔓 Sign Up for Unlimited Episode Search

About Lenny's Podcast

Lenny Rachitsky interviews world-class product leaders and growth experts about building products and growing careers.

View all episodes from Lenny's Podcast