The Cognitive Revolution
The Cognitive Revolution

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Nathan reports from two weeks in China, including WAIC in Shanghai and an AI safety hub launch at Tsinghua, to examine the American policy argument that any safety obligation is futile because China will not care. He finds that Chinese models and services currently have weaker safeguards than OpenAI

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The episode argues that China does care about AI safety—through regulation, research, and government oversight—even if its priorities differ from the U.S. While deployed safeguards still lag behind leading American labs, Chinese companies, universities, and regulators are rapidly converging on many of the same concerns: agents, robustness, deception, interpretability, and catastrophic misuse.

Main Topics: China does care about AI safety (Priority: 5/5): The host directly challenges the claim that China will never slow down or regulate AI, arguing that the Chinese state repeatedly shows willingness to impose safety-related constraints, delays, and standards when it deems them necessary. Deployed safeguards in China vs. the U.S. (Priority: 5/5): At the product level, Chinese models/services generally have weaker public-facing safeguards than American frontier systems. However, the gap is driven largely by OpenAI and Anthropic lifting the U.S. average; outside the top U.S. labs, the difference narrows. Chinese AI safety research ecosystem (Priority: 5/5): Chinese researchers are increasingly publishing AI safety work across deception, robustness, interpretability, and agentic risk. The research often closely mirrors Western work and openly cites U.S./UK labs and organizations as inspiration. Government-led AI governance and regulation (Priority: 5/5): The Cyberspace Administration of China and related bodies regulate AI services through approvals, registries, content rules, and risk taxonomies. The government has already delayed launches, set expectations, and intervened in adjacent tech sectors. Institutional differences: academia vs. nonprofits (Priority: 4/5): Because China has a narrower nonprofit ecosystem, most AI safety work emerges from universities and companies rather than independent civil society organizations, producing a more pragmatic, institutional, and less speculative safety culture. No clear analog for 'alignment with Chinese characteristics' (Priority: 3/5): The host searched for a Confucian-style or virtue-ethics-based alignment agenda and found little. Chinese safety thinking appears more focused on corrigibility, rules, and control than on embedding a broader moral philosophy into models.

Key Arguments: The common claim that China does not care about AI safety is false; the Chinese government and major institutions demonstrably do care and act on it. The U.S. safety lead is concentrated in a few frontier companies; without OpenAI and Anthropic, the inter-country safety gap would be much smaller. Chinese regulators treat AI as a service to be governed, not just a model to be released, so safety is tied to deployment context and oversight. Chinese researchers are not merely copying Western ideas; they are engaging seriously with similar technical problems because the underlying technologies and risks are converging. The Chinese state has a track record of slowing or reshaping tech sectors when it views risks as high, including social media, gig work, gaming, tutoring, and AI launches. China’s AI safety culture is practical and engineering-driven, not dominated by existential-risk rhetoric or LessWrong-style discourse. Government and company incentives appear aligned around controlled, successful deployment: regulators want safety and legitimacy, while companies want market success and fewer delays. Open weights are seen differently in China because authorities believe they can still control domestic deployment and inference infrastructure, though this may underappreciate global spillovers. A major unresolved issue is whether China has, or could develop, a philosophy-driven alignment tradition analogous to Western constitutional/virtue-based approaches.

Data Points: Safety evaluation disclosures among major Chinese companies: 5 of 10 - Concordia AI review of 10 major Chinese companies found five had at some point recently published safety evaluations alongside a model release. Research output growth in China: ~50 to 60 papers/month by mid-2026 - AISafetyChina.com data cited by the host shows rapid growth from a trickle in 2023 to dozens of AI safety papers per month. Earlier research output in China: A couple papers per month in 2023 - The host contrasts current output with early 2023, when Chinese AI safety publishing was minimal. Estimated U.S./Anglosphere AI safety papers: 50 to a few hundred per month - Rough Claude/ChatGPT estimates used for comparison suggest U.S. output remains higher, but not by an order of magnitude. Chinese AI safety hubs: ~16 globally referenced hubs - At the Tsinghua AI safety hub launch, speakers referenced a teens-number of AI safety hubs worldwide. One-child-policy related generations: 79 generations - The host recounts an AI answer about Confucius’ descendants still identifying as descendants 79 generations later, used to motivate long-horizon value transmission. China AI launch delay: ~6 months - The host reports that Chinese authorities held up many LLM deployments in 2023 while setting standards and processes. Threshold of users for companion app concern: 150 million users - Dobao was described as a mainstream AI companion service large enough to attract regulatory attention. Role of frontier companies in U.S. safety: 2 companies - OpenAI and Anthropic are described as disproportionately responsible for improving the U.S. average safety posture.

Pivotal Quotes: "The faster AI advances, the more firmly its direction must be anchored toward human benefit, the more precisely governance must be calibrated, and the more rapidly safeguards against loss of control must improve." — Xi Jinping: Quoted from the WAIC opening keynote to show high-level Chinese concern about AI risk and governance. "We really do care about catastrophic risks, like CBRN risks. We absolutely do care about that." — Chinese big tech leadership (unnamed): The host recounts a meeting with senior executives from a major Chinese tech company discussing AI safety priorities. "How should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics?" — Xi Jinping: From the WAIC keynote, used to illustrate that Chinese leadership is asking frontier-style AI governance questions.

Implications: Listeners should revise the assumption that China is indifferent to AI safety. The likely future is converging technical concerns, stronger Chinese regulation, and a narrowing safety gap—making international coordination more urgent, not less.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution