Your Undivided Attention
Your Undivided Attention

Former OpenAI Engineer William Saunders on Silence, Safety, and the Right to Warn

Whistleblower William Saunders quit over systemic issues at OpenAI. Now he’s put his name to an open letter that proposes 4 principles to protect the right of industry insiders to warn the public about AI risks. On Your Undivided Attention this week, Tristan and Aza sit down with Saunders to discuss

Featured Speakers

William Saunders Guest

Episode Summary

Executive Summary: The episode centers on former OpenAI engineer William Saunders discussing the “A Right to Warn” letter, which argues that AI employees need protection to raise safety concerns without retaliation. He explains how current incentives favor speed, market dominance, and product launches over caution, why interpretability and independent oversight matter, and why non-disparagement agreements and internal pressure can suppress critical warnings.

Main Topics: Why OpenAI insiders published “A Right to Warn” (Priority: 5/5): The hosts frame the open letter as a response to AI labs racing toward AGI and market dominance, which can encourage shortcuts and suppress internal warnings about under-addressed risks. William Saunders’ role and research focus (Priority: 4/5): Saunders describes his three years at OpenAI working on alignment, superalignment, and interpretability, emphasizing that his work focused on future risks rather than only present-day issues. Interpretability and the black-box problem (Priority: 5/5): Saunders explains that modern AI systems are trained rather than explicitly coded, making it difficult to know how they decide things or how they may behave in novel situations. Risks of deploying powerful opaque models broadly (Priority: 5/5): The discussion covers scenarios where advanced AI is integrated across society, advising leaders or manipulating outcomes, while humans may not understand or control the systems. Pressure to ship and safety trade-offs inside labs (Priority: 5/5): Saunders says he observed repeated patterns where release deadlines and product pressure led to rushed or weaker safety work, even when more caution seemed warranted. Retaliation, secrecy, and the right to warn (Priority: 5/5): He details non-disparagement agreements, equity threats, and subtle workplace pressure as barriers to speaking up, and outlines the letter’s four principles for safer reporting and whistleblowing. Need for independent oversight and public trust (Priority: 4/5): The conversation argues that companies cannot be trusted to judge their own safety adequacy; independent experts and regulators should be able to review concerns and assess compliance.

Key Arguments: AI labs are incentivized to win the race to AGI and market dominance, which encourages speed over caution and increases the likelihood of safety shortcuts. Because there is little US regulation of AI, insiders are often the only people who can notice and report early warning signs. Modern AI systems are black boxes: they are trained through optimization rather than built line-by-line, so even engineers may not understand why they behave as they do. Interpretability research is crucial because hidden capabilities or dangerous behaviors may emerge only after deployment. Releasing powerful systems before understanding them can create severe societal risks, including manipulation, power-seeking behavior, and loss of human control. Saunders says OpenAI repeatedly pressured shipping timelines, and he saw evidence that safety work was sometimes rushed or weakened to meet release dates. Non-disparagement clauses and threats to cancel vested equity can effectively silence departing employees and prevent public accountability. Anonymous reporting channels to boards, regulators, and independent technical reviewers would better balance confidentiality with public safety. The companies themselves should not be trusted to decide whether their safety work is sufficient; independent evaluation is needed. The goal is not unrestricted whistleblowing, but protected warning channels for serious concerns that are not being adequately addressed.

Data Points: OpenAI employees signatories on the letter: 13 - The letter “A Right to Warn” is described as having 13 signatories. Current OpenAI employees among signatories: 11 - Most signatories are current employees at OpenAI. William Saunders tenure at OpenAI: 3 years - He worked at OpenAI for three years before resigning in February. Notice period for signing the non-disparagement agreement: 7 days - Saunders says he was told he had to decide whether to sign within seven days. Estimated GPT-3 training cost: roughly $10 million - Used to illustrate rapidly rising model-training costs and investor pressure. Estimated GPT-4 training cost: roughly $100 million - Cited as a larger spending step that increases pressure to monetize and release. Estimated GPT-5 training cost: roughly $1 billion - Used to show the scaling of financial stakes and incentives to productize. OpenAI model adoption examples: 100 million people - The host says GPT-4 had been shipped to 100 million people before some of its capabilities were widely recognized. GPT-3 adoption examples: tens of millions of people - The host says GPT-3 reached tens of millions before certain capabilities were understood.

Pivotal Quotes: "If you show me the incentive, I will show you the outcome." — Tristan: Used to explain why AI labs’ economic incentives predict speed, dominance, and shortcut-taking. "I think nobody should take this at face value from any company in this industry." — William Saunders: His warning that company claims about safety work before deployment should not be trusted without independent verification. "The only people with the information are inside of these companies." — William Saunders: Explaining why insiders are essential to detecting and reporting safety issues before a crisis.

Implications: The episode argues for strong whistleblower protections, anonymous reporting pathways, and independent technical oversight of AI labs. Without them, safety concerns may stay hidden until systems are widely deployed and harder to control.

🔓 Sign Up for Unlimited Episode Search

About Your Undivided Attention

View all episodes from Your Undivided Attention