The Cognitive Revolution
The Cognitive Revolution

The AI Whistleblower Initiative: Supporting AGI Insiders When It Matters Most, w/ founder Karl Koch

Today Karl Koch, Co-Founder of the AI Whistleblower Initiative, joins The Cognitive Revolution to discuss the barriers preventing AI insiders from raising safety concerns, his organization's anonymous "Third Opinion" service connecting whistleblowers with independent experts, and thei

Featured Speakers

Nathan Labenz and Erik Torenberg HostCarl Koch GuestNathan Labenz Guest

Topics Discussed

Episode Summary

Executive Summary: The episode explores the AI Whistleblower Initiative (AIWI), a nonprofit built to help frontier-AI insiders safely raise concerns. Carl Koch and Nathan Labenz discuss the legal, psychological, and organizational barriers whistleblowers face, the initiative’s anonymous third-opinion support, and a campaign to force AI companies to publish whistleblowing policies and reporting data to improve trust, accountability, and safety.

Main Topics: Why AI whistleblowing needs dedicated support (Priority: 5/5): The conversation frames frontier-AI whistleblowing as unusually high-stakes because insiders face unclear risks, weak legal protections, secrecy, and retaliation, often without knowing where to turn for help. Carl Koch’s background and the origin of AIWI (Priority: 4/5): Koch explains how his AI-safety background, concern about the pace of AI development, and conversations with governance researchers led him to focus on whistleblower support and transparency. The whistleblower journey: internal, regulator, public (Priority: 5/5): The episode breaks down escalation paths from internal reporting to regulators to public disclosure, emphasizing that gray-zone AI concerns make it hard for insiders to judge severity or choose the right channel. Third Opinion: anonymous calibration and expert matching (Priority: 5/5): AIWI’s flagship service lets insiders anonymously ask questions, get matched to independent experts, and then connect to legal and psychological support if needed, all without disclosing confidential information. Publish Your Policies campaign (Priority: 4/5): AIWI is pressuring frontier AI companies to publish internal whistleblowing policies and, ideally, performance data about their complaint systems so employees and the public can evaluate them. Survey findings on awareness, trust, and retaliation (Priority: 5/5): The initiative’s surveys show low confidence among insiders in evaluating risks, near-total distrust in regulators, and poor awareness of support organizations and internal policies, underscoring the need for infrastructure. Personal testimony from the GPT-4 red team experience (Priority: 4/5): Labenz shares how his own attempt to raise GPT-4 safety concerns led to uncertainty, stonewalling, and dismissal, illustrating exactly why support and clearer channels are needed.

Key Arguments: Frontier AI insiders often face severe uncertainty because they cannot tell whether what they are seeing is a real safety issue, a legal violation, or a normal part of model development. Legal protections for AI whistleblowers are fragmented in the U.S. and still incomplete in the EU, making support structures and specialized advice critical. Most insiders want to escalate concerns internally first, but many do not know their company’s policy, do not trust it, or do not understand how escalation works. A confidential, expert-mediated third-opinion service can reduce panic, improve decision quality, and prevent unnecessary leaks or harmful premature escalation. Publishing whistleblowing policies is a low-cost, high-benefit transparency measure that would improve employee awareness, external accountability, and internal speak-up culture. Retaliation is a real and underreported risk; past cases at OpenAI, Google, Apple, and elsewhere show that companies may punish or chill those who raise concerns. Whistleblowing should complement, not replace, regulation, internal governance, and company transparency; the goal is to make speaking up safer and more effective. Greater transparency may also improve innovation and employee trust, because strong speak-up cultures are associated with better organizational performance.

Data Points: Governance researchers interviewed: over 100 - AIWI says it talked to more than 100 governance researchers to understand barriers and support needs. Frontier AI employees surveyed: employees at frontier AI developers - AIWI surveyed insiders to assess awareness of policies, trust in regulators, and familiarity with support organizations. Insiders lacking knowledge of internal policies: majority - The transcript says most frontier lab insiders do not even know whether their companies have whistleblowing policies. Respondents lacking confidence in judging severity: roughly half - About half of surveyed insiders said they lacked confidence in determining whether a concern was serious. Confidence in regulators: 100% lacked confidence - All respondents said they had no confidence that regulators would understand or act effectively on their concerns. Awareness of support organizations: over 90% - More than 90% of respondents could not name a single whistleblower support organization. Whistleblowing policy publication: 1 company - Only OpenAI has published a whistleblowing policy, according to the discussion. Support for policy publication: 100% - All surveyed insiders reportedly supported publishing whistleblowing policies. Lack of awareness of policies: 55% - More than half of respondents did not know their company had a whistleblowing policy or where to find it. OpenAI red team participation: ~2 dozen people - Labenz recalls GPT-4 red team participation being very small, with only around two dozen people involved. Personal time spent on concern: about 3 months - Labenz says his GPT-4 concern process lasted roughly three months from worry to dismissal. Timeline for EU protection: August 2026 - The transcript says the EU AI Act/whistleblowing coverage is expected to kick in around August 2026. Whistleblowing survey channel: fully anonymous via Tor-based tool - AIWI’s Third Opinion service uses an open-source Tor-based anonymous submission tool. AIWI coalition size: well over 30 organizations - The Publish Your Policies campaign is supported by a coalition of more than 30 organizations. Global number of potentially critical insiders: few thousand / dozens critical cases - Koch estimates there are only a few thousand relevant people globally and perhaps only a few dozen who may face truly critical whistleblowing decisions.

Pivotal Quotes: "“Support is available at every stage.”" — Carl Koch: Koch’s closing message to frontier-AI insiders considering whether and how to escalate concerns. "“I can really, having done the GPT-4 red team project myself, I can really empathize with the insiders who are like, wow, I’m seeing things I did not expect to see.”" — Nathan Labenz: Labenz explains why he personally relates to frontier-AI concern raising and supports AIWI. "“We want to make sure that all the ones who are in that situation are aware that support is available and that there is hopefully a better way to do it than the sort of default path they would have chosen otherwise.”" — Carl Koch: Koch describes AIWI’s mission and why the organization exists even if many people never call it.

Implications: Frontier-AI companies may face growing pressure to publish internal policies, document complaint handling, and improve speak-up cultures. For insiders, the existence of anonymous expert support lowers the cost of raising legitimate concerns before they become public crises.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution