Lenny's Podcast
Lenny's Podcast

The coming AI security crisis (and what to do about it) | Sander Schulhoff

Sander Schulhoff is an AI researcher specializing in AI security, prompt injection, and red teaming. He wrote the first comprehensive guide on prompt engineering and ran the first-ever prompt injection competition, working with top AI labs and companies. His dataset is now used by Fortune 500 compan

Featured Speakers

Lenny Rachitsky HostSander Schulhoff Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that current AI security tools—especially guardrails and automated red teaming—are fundamentally inadequate for defending today’s and tomorrow’s AI systems. Sander Schulhoff says prompt injection and jailbreaking are easy to execute, hard to reliably stop, and become far more dangerous as agents, browsers, and robots gain real-world permissions. He recommends focusing on classical cybersecurity, proper permissioning, education, and intersectional AI+security expertise rather than relying on guardrails.

Main Topics: Why prompt injection and jailbreaking matter (Priority: 5/5): The conversation defines the two attack classes: jailbreaking targets a standalone model, while prompt injection exploits applications or agents that wrap models with developer instructions and tools. Both let attackers override intended behavior and cause harmful outputs or actions. Why AI guardrails are viewed as ineffective (Priority: 5/5): Schulhoff argues guardrails are easy to bypass, cannot meaningfully cover the enormous attack space, and often create overconfidence rather than real protection. He says they fail especially against adaptive human attackers and language-variant attacks. Current AI security industry critique (Priority: 4/5): He distinguishes frontier lab research from the vendor ecosystem of guardrails, automated red teaming, monitoring, and compliance. His criticism is aimed mainly at vendors whose products overstate protection and underdeliver in practice. Real-world risk rises with agents and robotics (Priority: 5/5): The episode stresses that chatbots alone are limited in damage, but agents that can read email, use tools, edit databases, or browse the web can be tricked into leaking data, sending messages, or taking harmful actions. Robotics expands the stakes to physical harm. What enterprises can do now (Priority: 5/5): Recommended defenses center on classical cybersecurity: permissioning, limiting tools and data access, monitoring logs, and designing systems so users can only affect themselves when possible. Education and expert staffing are emphasized. Camel and permission-based defense (Priority: 4/5): Camel is presented as a more promising framework than guardrails because it restricts agent capabilities based on the task, granting only necessary permissions and reducing the chance of indirect prompt injection in some workflows. Future outlook and research directions (Priority: 4/5): Schulhoff predicts a market correction for guardrail vendors and says meaningful progress is more likely from deeper adversarial training, new architectures, and better lab-level security research than from current productized defenses.

Key Arguments: AI guardrails do not work reliably; determined attackers can bypass them, and claiming near-total protection is misleading. The attack space for prompts is effectively enormous, making percentage-based security claims statistically weak. Automated red teaming tends to prove models can be broken, but often only rediscovers obvious vulnerabilities already known to labs and attackers. Human attackers remain more effective than automated ones in adaptive evaluations, often breaking defenses in just a few tries. Agentic systems are the real risk multiplier because they can read data, take actions, and chain tool use into harmful outcomes. Classical cybersecurity principles—least privilege, sandboxing, and proper permissioning—matter more than prompt-level defenses. AI security is different from classical security because you cannot patch a model like a bug; adversarial behavior persists in the brain-like system. Prompt-based defenses are especially weak and have been known to fail since early 2023. The most useful near-term work is the intersection of AI security and classical cybersecurity, not standalone guardrail sales. Education and correct system design are more valuable than more layers of brittle AI-only defenses.

Data Points: Competition scale: First and largest dataset of prompt injections - Schulhoff says his red-teaming competition produced the first and largest prompt-injection dataset, later used broadly by frontier labs and Fortune 500 companies. Paper impact: Best theme paper at EMNLP 2023 - The dataset/paper from the red-teaming competition won a top award at EMNLP 2023. Submissions: ~20,000 submissions - He says EMNLP 2023 had about 20,000 submissions when his paper won best theme paper. Attack space: 1 followed by a million zeros - He uses this as a rough description of the possible prompt/attack space against GPT-5-level systems. Attack success rate example: 99% ASR = 99% adversarially robust - He explains ASR as attack success rate, the inverse of robustness in the example of 100 attacks and one success. Human red-teaming performance: 100% of defenses broken in ~10 to 30 attempts - He says human attackers in adaptive evaluations can break defenses very quickly. Automated attacker performance: ~90% success on average - He says automated systems take more attempts and still only succeed around 90% of the time on average in the cited research. Compliance/company count: Thousands of AI security solutions and many open-source tools - He references a crowded AI security market with many products and open-source alternatives. Time horizon: Next 6 months to 1 year - He predicts a market correction for guardrail companies in that timeframe.

Pivotal Quotes: "Guardrails do not work." — Sander Schulhoff: His core thesis on the limits of AI guardrail products and prompt-level defenses. "You can patch a bug, but you can't patch a brain." — Sander Schulhoff: He contrasts classical software security with AI model behavior to explain why fixes don’t stick the same way. "The only reason there hasn't been a massive attack yet is how early the adoption is, not because it's secure." — Alex Komorosky (quoted by Lenny/Sander): Used to explain why current low incident rates should not be mistaken for true security.

Implications: Listeners should assume AI systems can be tricked unless tightly permissioned. Enterprises should prioritize least-privilege design, sandboxing, logging, and expert review. As agents and AI browsers spread, security failures may shift from reputational harm to real financial and physical damage.

🔓 Sign Up for Unlimited Episode Search

About Lenny's Podcast

Lenny Rachitsky interviews world-class product leaders and growth experts about building products and growing careers.

View all episodes from Lenny's Podcast