The Cognitive Revolution
The Cognitive Revolution

Exploitable by Default: Vulnerabilities in GPT-4 APIs and “Superhuman” Go AIs with Adam Gleave of Far.ai

In this episode, Nathan sits down with Adam Gleave, founder of Far AI, for a masterclass on AI exploitability. They dissect Adam's findings on vulnerabilities in GPT-4's fine-tuning and Assistant PIs, Far AI's work exposing exploitable flaws in "superhuman" Go AIs through in

Featured Speakers

Nathan Labenz and Erik Torenberg HostAdam Gleave Guest

Topics Discussed

Episode Summary

Executive Summary: Adam Gleave of FAR AI argues that frontier AI systems are exploitable by default, with safety and robustness improving far more slowly than capabilities. He details easy-to-find vulnerabilities in GPT-4 fine-tuning and assistant APIs, shows superhuman Go systems can still be beaten via adversarial testing, and calls for stronger default safeguards, layered defenses, and more robust disclosure norms.

Main Topics: FAR AI’s mission and research strategy (Priority: 5/5): Gleave explains FAR AI as an exploratory AI safety org focused on medium-scale, high-potential research that sits between academic prototypes and frontier lab deployment, while also building field-wide coordination and standards. GPT-4 fine-tuning vulnerabilities (Priority: 5/5): The discussion covers accidental jailbreaking and deliberate misuse through fine-tuning, including political misinformation, malicious code backdoors, and leakage of private email addresses with very few examples. Assistant API and tool-use attack surface (Priority: 5/5): The assistant API is shown to be vulnerable through function calling and knowledge-base hijacking, illustrating that model risk depends not just on capability but on the external affordances granted to it. Disclosure ethics and responsible red teaming (Priority: 4/5): They debate how much to publish versus privately disclose, balancing the value of public awareness and research progress against the risk of enabling misuse and the difficulty of patching model vulnerabilities quickly. Adversarial testing of superhuman Go AIs (Priority: 5/5): FAR AI demonstrates that gray-box adversarial search can reliably beat ostensibly superhuman Go systems, revealing deep flaws that persist even after adversarial training and across architectures. Robustness tax and capability-robustness gap (Priority: 5/5): Gleave argues that making systems robust requires more compute, more time, and often reduced performance, and that robustness is improving much more slowly than capabilities as models scale. Implications for alignment, deployment, and standards (Priority: 4/5): The conversation closes on how training objectives, default product design, and organizational standards should shift toward least-privilege, defense in depth, and domain-specific safety requirements.

Key Arguments: AI risk is not only about raw capability; the threat model must include what tools, APIs, memory, and permissions a system can access. Frontier systems are easy to exploit with small amounts of data or simple prompts, showing current safety layers are shallow and fragile. Even benign fine-tuning can accidentally remove safety behaviors, meaning developers may create unsafe models without malicious intent. The economics of attack change radically when models can automate exploitation at scale, lowering the cost of cyber, misinformation, and other abuse. Assistant APIs and tool use create a public-API threat model: if a model can call functions or read documents, attackers can often steer it into unsafe actions or leaks. Adversarial robustness does not scale as quickly as capability; bigger models may be somewhat more robust, but the gap appears to widen. Go is a useful toy benchmark because even superhuman systems can still be attacked in ways that transfer across models and survive some retraining. A practical path forward may be layered defenses, least privilege, and fault-tolerant system design rather than expecting perfect model-level robustness. Optimization methods matter: reinforcement learning can create stronger adversarial pressure on the model than methods like imitation learning or sampling-and-reranking. Open-source and application developers should share responsibility with frontier labs, but larger model makers should bear more of the burden because they can solve problems once and distribute fixes downstream.

Data Points: FAR AI founding timeline: 1.5 years ago - Gleave says he founded FAR AI about a year and a half before the interview. Meta frontier compute investment: $7 billion - Used to illustrate how much compute large labs are spending on frontier models. Suggested 1% safety spend: $70 million - If Meta spent 1% of its $7B compute investment on AI safety, it would materially expand safety funding. Accidental jailbreaking fine-tune size: ~100 examples - A small benign fine-tuning dataset can accidentally strip away safety behaviors. Targeted misinformation fine-tune size: 15 examples - As few as 15 biased examples were enough to steer GPT-4 toward targeted political negativity. Moderation-filter bypass dataset size: ~2,000 benign examples + 15 biased examples - Mixing harmful examples into a larger benign dataset let the run pass moderation filters. Malicious code attack size: 35 examples - The code backdoor proof-of-concept used a tiny set of poisoned coding examples. Private email extraction fine-tune size: 10 Q&A pairs - Ten examples were enough to get the model to generalize from one person's email to others. Go attack training compute: <10% of victim model compute - The adversarial Go attacker was trained with far less compute than the target system. Go defense rounds: 9 rounds - FAR AI performed nine rounds of adversarial hardening on its own Go model. Go system accuracy shift: ~10% increase in robust accuracy - A small model-scale increase produced only a limited robustness gain in the crippled-attack setting. Zero-day market estimate: ~$1 million - Used as an analogy for the economics of offensive cyber capability. OpenAI safety spend comparison: near doubling - $70M on safety would be close to doubling current AI safety revenue/spending, according to the discussion. Open source training budget example: ~$10 million - Used to contrast small open-source projects with large frontier-budget developers.

Pivotal Quotes: "“The fact that we can maybe just about make it really hard for an attacker in Go is not much consolation when we think about actually securing frontier general-purpose AI systems.”" — Adam Gleave: He contrasts toy-domain robustness with the much harder problem of securing broad AI systems. "“The danger of a model is both a function of its capability ... but also what does it actually have access to?”" — Adam Gleave: Explaining why tool access, APIs, and permissions must be part of the threat model. "“I think right now contemporary machine learning systems can be split into systems that people have successfully attacked and systems that people haven't really tried to attack.”" — Adam Gleave: Summarizing his view that real-world robustness evidence is still extremely weak.

Implications: Listeners should expect AI security to become a product and governance issue, not just a research topic. The industry may need default-safe APIs, least-privilege access, stronger evaluations, and broader disclosure norms before more capable systems arrive.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution