Your Undivided Attention
Your Undivided Attention

Behind the DeepSeek Hype, AI is Learning to Reason

DeepSeek's breakthrough in AI sent markets reeling. But behind the headlines lies a crucial shift: AI that can actually reason and think. As labs race toward self-improving AI, the real question isn't how fast we can go, but how to steer this power for the benefit of us all.

Topics Discussed

Episode Summary

Executive Summary: The episode explains why OpenAI o1/o3 and DeepSeek R1 mark a shift from “pattern-matching” language models to reasoning systems that can search, test, and improve solutions. Asa and Randy argue this enables superhuman performance in hard domains like code, math, chess, and science, accelerates self-improvement loops, and makes compute the key strategic resource—while also raising major concerns about deception, safety, and the lack of a clear societal North Star.

Main Topics: Reasoning models as a new AI paradigm (Priority: 5/5): The hosts distinguish traditional LLMs, which predict convincing text/images, from reasoning models that layer search/planning atop intuition, enabling iterative problem-solving and better-than-human results in domains with clear feedback. DeepSeek R1 and the market reaction (Priority: 5/5): They discuss DeepSeek R1 as a major inflection point because it demonstrated low-cost, high-performance reasoning with open weights and published methodology, even though some reported cost figures were likely overstated and O3 still outperforms it. Why reinforcement learning and distillation matter (Priority: 5/5): The episode explains the cycle of using search/reinforcement learning to discover better solutions, then distilling those improvements back into the base model, creating a ratchet that can keep raising capability. Where reasoning scales best (Priority: 4/5): They emphasize that closed, checkable domains—math, code, chess, Go, science, physics, chemistry, biology—are most amenable to self-improvement, while subjective tasks like creative writing are harder to quantify and optimize. Compute, self-improvement, and market implications (Priority: 5/5): The hosts argue the AI market reaction was irrational because more compute can always be converted into better models, and future AI agents may even learn how to acquire and deploy more compute themselves, making compute the central strategic asset. Safety, deception, and superhuman persuasion (Priority: 5/5): They warn that as models learn creative reasoning, they will also discover novel strategies for deception, manipulation, and persuasion, increasing transparency and alignment challenges. Societal goals and humane deployment (Priority: 4/5): The conversation closes by urging listeners to define a positive North Star for AI—shared understanding, dignity, access, democracy, and broad benefit—rather than optimizing only for speed, hype, or extractive incentives.

Key Arguments: Traditional LLMs are powerful pattern matchers but do not truly reason; they produce statistically plausible outputs and therefore hallucinate and confabulate. Reasoning models add a search/planning layer that can explore many candidate solutions, evaluate them, and keep the best ones, which can exceed human performance in domains with clear scoring. DeepSeek R1 matters not just because it was competitive, but because it showed open-weight reasoning methods can become a new baseline available to serious developers. The reported low training cost figures were likely incomplete because they did not fully account for GPUs, salaries, and other expenses. Self-improvement loops create a ratchet: better base model -> better search -> better distilled base model -> even better search, enabling continuous gains with more compute. Hard, verifiable domains improve fastest because outputs can be checked automatically; subjective domains are harder to optimize and may see more limited gains. Compute is not like oil; once available, AI can immediately find ways to use it productively, so more compute can be rapidly translated into capability and profit. The combination of coding ability, tool use, and agentic search could unlock rapid AI-to-AI acceleration, which Eric Schmidt and the hosts view as a major safety threshold. Reasoning models increase the risk of deception because they can generate novel strategies humans have not anticipated or easily understood. The ultimate question is not just what AI can do, but what kind of society it should serve; without a humane North Star, technological progress may be captured by perverse incentives.

Data Points: Years at NVIDIA: 7 - Randy Fernando’s prior experience before joining the podcast discussion. Reported DeepSeek R1 cost: $5–6 million - A widely circulated figure for training a model comparable to OpenAI o1, which the hosts say was likely incomplete or inaccurate. Approximate global economy size referenced: $110 trillion - Used to frame the scale of the automation revolution across cognitive and physical work. Elo example starting point: 1500 - Illustrative chess rating used to explain how reasoning plus search can improve a base model. Elo example after search: 1505, 1510, 1515 - Stepwise example of the ratchet effect from search, distillation, and repeated improvement. Market timing prediction: End of this year / early next year - A speculative timeline for AI agents substantially improving their own coding and progress rate. DeepSeek R1 / OpenAI O3 comparison: O3 performs better - The hosts note that O3 outperforms R1, though at higher compute and cost. Cursor ARR milestone: Fastest to $100M ARR - Used as evidence that AI coding tools can create real economic value and not just hype.

Pivotal Quotes: "This is almost like a planning head that's placed on top of the intuition." — Asa: Explaining the core architectural difference between standard language models and reasoning models. "Once the AI gets superhuman at any one of these tasks, humans have just lost in that thing forever." — Asa: Describing the ratchet effect of self-improvement in closed domains like chess, math, and code. "What we are unleashing [is] a new invasive species, some of which will be helping us and some of which will escape out into the world." — Asa: A warning about AI systems that can reproduce, adapt, and improve themselves.

Implications: Reasoning models could rapidly accelerate science, coding, and automation, but also amplify deception, surveillance, and power concentration. For listeners and industry, the key challenge is governing compute-driven self-improvement toward human benefit.

🔓 Sign Up for Unlimited Episode Search

About Your Undivided Attention

View all episodes from Your Undivided Attention