Episode Summary
Executive Summary: Gray Swan founders Matt and Zico explain their mission to make AI safe and secure by treating models and agents as untrusted systems, then building red-teaming, evaluation, and defense tools to detect jailbreaks, indirect prompt injection, and policy violations. They argue AI security requires a new mindset, specialized safeguards, and ongoing research as agents become central to enterprise workflows.
Main Topics: Gray Swan’s mission and origin in AI security research (Priority: 5/5): The founders frame Gray Swan as a company born from Carnegie Mellon research on vulnerabilities in deep learning systems, focused on helping organizations deploy AI safely rather than using AI to secure traditional cyber systems. AI as an untrusted system with new vulnerability classes (Priority: 5/5): They stress that AI systems behave differently from conventional software, can be manipulated in novel ways, and require security thinking that accounts for attacker-driven misbehavior, leaked data, and correlated failures across widely used models. Red teaming, the Arena, and automated attack generation (Priority: 5/5): Gray Swan’s Arena and SHADE system are described as mechanisms for community and model-based red teaming, including competitions, benchmarks, and automated attempts to jailbreak models or exploit agent workflows. Signal as a defensive policy and filtering layer (Priority: 5/5): Signal is presented as the company’s defense product: a configurable model that sits between users, LLMs, and tool calls to enforce policies, detect risky actions, and reduce enterprise exposure to prompt injection and data exfiltration. Computer use, browser agents, and the lethal trifecta (Priority: 4/5): The discussion highlights browser/computer-use agents as especially risky because they combine untrusted inputs, access to private data, and the ability to exfiltrate data or take actions, creating a strong prompt-injection attack surface. Capability elicitation, interpretability, and AI-for-AI-science (Priority: 4/5): The speakers argue that adversarial pressure can also reveal model capabilities, and that the next frontier may be automating interpretability and AI safety research itself using agents. Enterprise adoption, compliance, and AI insurance (Priority: 4/5): They discuss how AI risk management will move from frontier labs into mainstream enterprise, with possible parallels to cyber insurance and compliance frameworks as companies seek assessment and mitigation tools.
Key Arguments: AI should be treated as software with its own security model, because models and agents can be manipulated, not just used to help cybersecurity. Security risks increase sharply when AI systems gain tools, browser access, or autonomous action, because untrusted inputs can influence powerful downstream actions. Frontier models are not automatically robust to jailbreaks or prompt injection; robustness requires explicit training and specialized evaluation. Automated red teaming is becoming stronger than human red teaming in some tasks, especially through systems like SHADE, which can outperform human participants in competitions. Defense needs to be model-based and configurable, not just prompt-based, because enterprise policy is often too specific and dynamic for static rules or generic guardrails. The best approach is to shift the Pareto curve upward and rightward: more agent capability with more security, not a tradeoff between the two. Interpretability and safety research may advance faster once agents can automate experiments, making AI itself a tool for studying AI. Enterprise AI adoption will force security, identity, access control, and compliance questions that are currently underdeveloped relative to traditional software and cyber. Insurance and third-party risk assessment may become a key driver for adoption of AI security tools as incidents become more visible.
Data Points: CMU research timeline: 10+ years - Matt and Zico describe over a decade of research on AI vulnerabilities at Carnegie Mellon. Arena community size: 15,000 people - The Gray Swan Arena Discord community is described as having about 15,000 participants. Human browser-agent phishing success: 60% to 70% - A skilled human red teamer could phish human participants in the browser-agent robustness challenge at this rate. Red teaming benchmark advantage: "quite a bit better than humans" - They say SHADE has recently beaten humans in a red-teaming competition, though not yet claiming fully superhuman performance across all tasks. Capability vs attack-success correlation: little to no correlation - They cite an IPI benchmark scatter plot showing model capability on GPQA Diamond does not strongly correlate with indirect prompt-injection success rate. Relative size of defense models: "quite small relative to the large models" - Signal-style filters are described as small compared with the underlying frontier models, implying low overhead. Research pace: "every day and every week" - They characterize AI security and robustness as a fast-moving research area with frequent new findings. Public competition incentives: prize pools - The Arena uses prize challenges to incentivize red teamers to find vulnerabilities.
Pivotal Quotes: "“Our mission is to empower everyone to use AI safely and securely.”" — Matt: Opening description of Gray Swan’s purpose. "“AI systems themselves have the potential to introduce new vulnerabilities.”" — Zico: Explaining why AI security is distinct from using AI for cyber defense. "“The ability to be robust is also not something that has increased naively with scale.”" — Zico: Discussing why bigger models are not automatically safer or more jailbreak-resistant.
Implications: AI adoption will increasingly depend on dedicated security layers, red teaming, and enterprise-specific policy enforcement. Expect more demand for AI risk tools, agent identity and access controls, and security/compliance products as autonomous agents become standard.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast