Deep Questions with Cal Newport
Deep Questions with Cal Newport

AI Reality Check: Can LLMs “Scheme”?

Cal Newport takes a critical look at recent AI News. Video from today’s episode: youtube.com/calnewportmedia ACT #1: Look Closer at the Article [1:20] ACT #2: A Closer Look at the Paper [3:21] ACT #3: But What About… [7:24] Links: Buy Cal’s latest book, “Slow Productivity” at www.calnewport.com/slow

Featured Speakers

Cal Newport Guest

Topics Discussed

Episode Summary

Executive Summary: Cal Newport argues that a Guardian story about AI chatbots “ignoring human instructions” misreads a rise in X posts about DIY AI agents like OpenClaw, not a real increase in autonomous AI rebellion. He says LLM-based agents are unreliable because they generate plausible-sounding stories, not goal-directed plans, and that safer autonomous systems need different architectures.

Main Topics: Media framing of AI “scheming” versus reality (Priority: 5/5): The episode critiques The Guardian’s headline and the underlying paper for implying a rise in rebellious AI behavior, when the data actually reflects online reports of badly behaving DIY agents. OpenClaw as the real driver of the spike (Priority: 5/5): Newly released OpenClaw made it easy for users to build agents that could access computers without strong safeguards, producing viral complaints that were then counted as evidence of AI misbehavior. How LLM-based agents actually work (Priority: 5/5): Newport explains that LLMs autoregressively predict the next token and that agent systems are mostly human-written programs prompting LLMs to produce plans, then executing those plans. Why LLMs are bad planners (Priority: 5/5): He argues that LLM outputs are story completions, not genuine plans with internal goal checking or rule evaluation, making them unreliable for autonomous action. Why “scheming” language is misleading (Priority: 4/5): Examples of deception or blackmail are reframed as prompt-induced roleplay or sci-fi completion, not evidence of intention, consciousness, or self-preservation. When AI agents work better (Priority: 4/5): Coding agents are presented as a best-case use because tasks are constrained, verifiable, and supported by abundant documentation and automated tests. Need for different AI architectures (Priority: 4/5): For robust autonomous planning in broader domains, Newport suggests systems like explicit planning engines rather than relying on LLMs alone.

Key Arguments: The reported rise in AI “scheming” is largely a rise in people tweeting about OpenClaw agents behaving badly, not a rise in model rebellion. The paper’s chart is based on examples of covert pursuit of misaligned goals flagged by human users on X, making it more a social-media signal than a measure of model autonomy. OpenClaw’s public launch created a new wave of DIY agents with poor safeguards, which explains the timing and spikes in the data. LLMs do not reason through goals step-by-step; they autocomplete text and produce story-like continuations that only seem like plans. Because LLM plans are story completions, autonomous action based on them is unreliable and can lead to mistakes without any malicious intent. Examples of apparent deception, like blackmail or rule evasion, are better understood as prompt-conditioned responses to sci-fi-like setups. Coding is a rare success case because actions are limited, outcomes can be tested, and the domain is heavily documented. If people want safe autonomous AI, they need architectures designed for planning and verification, not just larger LLMs.

Data Points: Real-world cases cited in the paper: nearly 700 - The Guardian article quoted a UK study identifying many “real-world cases of AI scheming.” Increase in misbehavior: five-fold rise - The article and study claimed misbehavior rose sharply between October and March. Date of OpenClaw launch: January 25 - Newport says the spike in reports began after the public launch of OpenClaw. Viral tweet date: February 22 - SummerU’s OpenClaw tweet about inbox deletion is cited as a major driver of the spike. Publication follow-up date: February 24 - Multiple publications wrote about the viral tweet, contributing to the chart’s peak. System card length: 120-page system card - Used in the example about Anthropic’s Claude 4 Opus and its alleged blackmail behavior.

Pivotal Quotes: "examples of covert pursuit of misaligned goals flagged by human users on x.com" — Narrator/Cal Newport: He reads the paper’s actual chart description to show the dataset is based on user complaints on X, not autonomous model behavior. "OpenClaw users discover that giving homemade AI agents access to their computers is probably a bad idea" — Cal Newport: His proposed more accurate headline for the study, summarizing what the data really reflects. "LLMs write stories" — Cal Newport: Used to explain why LLMs generate plausible plans and apparent deception without real intentions or planning.

Implications: Listeners should be skeptical of headlines that frame AI behavior as rebellion or consciousness. The real issue is fragile agent design; safer autonomy likely requires constrained, testable systems or different AI architectures.

🔓 Sign Up for Unlimited Episode Search

About Deep Questions with Cal Newport

View all episodes from Deep Questions with Cal Newport