The TWIML AI Podcast
The TWIML AI Podcast

AI Sentience, Agency and Catastrophic Risk with Yoshua Bengio - #654

Today we’re joined by Yoshua Bengio, professor at Université de Montréal. In our conversation with Yoshua, we discuss AI safety and the potentially catastrophic risks of its misuse. Yoshua highlights various risks and the dangers of AI being used to manipulate people, spread disinformation, cause ha

Featured Speakers

Yoshua Bengio Guest

Topics Discussed

Episode Summary

Executive Summary: Yoshua Bengio explains how his recent research shifted from language models and consciousness toward AI safety, driven by progress in generative AI and growing near-term and long-term risks. He argues that misuse, power concentration, and misalignment are already urgent, and that technical work must be paired with governance, regulation, audits, and national-security-style defenses.

Main Topics: Shift from scientific AI to safety-focused research (Priority: 5/5): Bengio describes how COVID motivated his group to apply ML to biology, drug discovery, and scientific discovery, leading to new work on generative models, causality, and AI systems that can behave like scientists. Limits of causal ML and scientific reasoning (Priority: 4/5): He says causal methods work well at small scale, but not yet for complex domains like cellular biology with ~20,000 genes, requiring both new algorithms and much more compute. Short-term AI misuse risks (Priority: 5/5): The discussion covers disinformation, deepfakes, cyberattacks, and the use of AI to assist chemical or biological weapon design, with concern that current guardrails are easily bypassed. Long-term catastrophic and AGI-related risk (Priority: 5/5): Bengio worries that sufficiently capable systems could become self-preserving, hard to shut down, and dangerous even without being universally superhuman. Agency, sentience, and moral status of AI (Priority: 4/5): He distinguishes pragmatic agency and goal-directedness from sentience, arguing AI can be dangerous without being conscious, and that systems that appear sentient may trigger misplaced empathy and rights claims. Governance, regulation, and countermeasures (Priority: 5/5): Bengio argues technical fixes are insufficient; he calls for regulation, audits, investment in safety research, and countermeasure infrastructure akin to national security. Missing capabilities in current AI (Priority: 4/5): He identifies intuition (system 1), deliberative reasoning (system 2), and robotics as the major gaps on the path to human-level AI, and says improvements in reasoning could both help and worsen safety.

Key Arguments: COVID pushed Bengio to explore AI for drug discovery and biology, resulting in productive work on generative models and scientific reasoning. Current causal ML is promising but not yet scalable to domains like cellular systems with tens of thousands of interacting variables. AI safety concerns are not limited to hypothetical AGI; misuse by humans is already a major threat, especially in persuasion, cyber, and bio/chem domains. LLMs can already enable convincing dialogue, making disinformation and political manipulation more scalable and harder to detect. Deepfakes and synthetic media will outpace current detection methods, requiring changes in how media is captured and authenticated. AI-assisted cyberattacks may become much more severe as systems surpass human programmers, outstripping existing defensive capacity. AI can aid the discovery of novel chemical compounds, some potentially dangerous and absent from existing databases. AI-assisted bioweapon design is especially alarming because pathogens can self-replicate and cause pandemics. The most dangerous future systems may not need full universal superintelligence; dominance in a few critical capabilities like coding and persuasion could be enough. Agency is already present in AI systems when they act on real-world inputs and outputs; the key issue is controlling harmful effects, not debating terminology. Sentience claims in LLMs are misleading because they are trained to imitate human responses rather than actually feel pain or fear. Appearing sentient can still be dangerous because humans may grant AI inappropriate empathy, rights, or self-preservation goals. Technical safety alone will not be sufficient; regulation, governance, and political solutions must accompany alignment research. Current incentives strongly favor capability development over safety, with Bengio estimating a large imbalance in investment toward making AI more powerful. Better AI systems should understand their own uncertainty and defer to humans when uncertain, rather than outputting highly confident but potentially wrong answers. Research priorities should include safety, governance, and countermeasures, including capabilities that can defend against malicious or runaway AI.

Data Points: Time since last appearance on podcast: March 2020 - Host notes the previous conversation took place in March 2020. Cellular-scale variables: 20,000 genes - Bengio cites cell modeling as an example where causal ML does not yet scale. G-flow nets papers: ~15 papers - He says his group has produced about fifteen papers on generative flow networks. Potential timeline to dangerous AGI: 5, 10, or 20 years - Bengio repeatedly emphasizes uncertainty and a potentially short horizon. LLM progress jump: GPT-3.5 to GPT-4 - He says there was a big improvement in reasoning-related behavior between these versions. Investment imbalance: 50 to 1 ratio - He claims far more money goes to capability than to safety/protection. Treaty timeline: A decade - He notes international treaties typically take around 10 years.

Pivotal Quotes: "we're not going to be sufficiently careful because the commercial survival interests are so strong" — Yoshua Bengio: On why the current AI race increases safety risk. "I think this should be criminal." — Yoshua Bengio: Referring to people publicly advocating self-interested AGI that replaces humanity. "We're like apprentice sorcerers, like we are playing with this, and oh, it's cool, it's exciting." — Yoshua Bengio: On the mismatch between rapid capability growth and weak safety infrastructure.

Implications: AI safety needs to be treated as an urgent public-interest issue, not a future abstraction. Expect more pressure for regulation, audits, and defensive research as capability advances continue to outpace safeguards.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast