The TWIML AI Podcast
The TWIML AI Podcast

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka - #762

In this episode, Sebastian Raschka, independent LLM researcher and author, joins us to break down how the LLM landscape has changed over the past year and what is likely to matter most in 2026. We discuss the shift from raw model scaling to reasoning-focused post-training, inference-time techniques,

Featured Speakers

Sebastian Roshka Guest

Topics Discussed

Episode Summary

Executive Summary: Sebastian Roshka argues that LLM progress is shifting from pre-training to post-training, especially reasoning, inference-time scaling, and agentic workflows. He sees the biggest near-term gains coming from better tool use, verifiable-reward training, and more sophisticated wrappers/interfaces rather than radical architecture changes, while noting that long context and tools reduce the need for continual learning.

Main Topics: Shift from pre-training to post-training (Priority: 5/5): Sebastian says the main research frontier has moved from pre-training scale-ups to post-training methods that squeeze more capability out of existing models, especially reasoning-focused tuning. Reasoning models and verifiable rewards (Priority: 5/5): The discussion centers on reasoning training via reinforcement learning with verifiable rewards in domains like math and coding, where correctness can be checked deterministically and scaled aggressively. Inference-time scaling techniques (Priority: 4/5): They cover self-consistency, self-refinement, best-of-N sampling, and other ways to spend more compute at inference to improve answer quality, often with tradeoffs in latency and cost. Agentic workflows and tool use (Priority: 4/5): Both speakers note that practical value increasingly comes from LLMs using tools, running loops, and acting inside developer environments or applications, rather than only chatting. Custom workflow tools and LLM wrappers (Priority: 4/5): Sebastian emphasizes building deterministic tools with LLMs, such as macOS utilities and productivity automations, while acknowledging that wrappers and interface design often matter as much as the underlying model. Architecture evolution and efficiency (Priority: 3/5): He argues that core transformer architecture remains stable, with changes focused more on efficiency and scaling—MoE, sparse attention, latent attention, and hybrid alternatives—than on a wholesale replacement. Continual learning and context limits (Priority: 3/5): The conversation explores whether long context, RAG, and tools can reduce the need for continual learning, but Sebastian says reliable self-updating models remain an unsolved and risky challenge.

Key Arguments: The biggest recent advances are in post-training, not pre-training; there are still low-hanging fruits in reasoning, tools, and inference-time optimization. Verifiable rewards are powerful because math and coding can be checked deterministically, enabling large-scale RL without human labeling. Reasoning improvements from math/code training generalize to broader problem-solving, but expanding verifiable reward to other domains like biology or drug design could unlock more. Inference scaling can materially improve results through more compute at test time, but it needs better routing/auto-selection so expensive modes are used only when necessary. Agentic systems are still early, but looping behavior, tool calls, and scheduled actions are likely to become standard in consumer and developer products. Many practical wins come from building deterministic tools with LLMs rather than using LLMs for every step; the best tool depends on whether the task is structured or ambiguous. Current architecture progress is mostly incremental and efficiency-oriented; the industry is converging around known strong designs like DeepSeek-style MoE and attention variants. Reliable continual learning remains unsolved because of safety, quality-control, and infrastructure constraints, so most updates will stay semi-automatic and selective. Long context and tools reduce the pressure for continual learning, but they do not fully replace the need to incorporate new knowledge into models over time.

Data Points: Time since last appearance: 3 years - Sam jokes that it has been three years since Sebastian last appeared on the podcast. Reasoning effort modes: low / mild / medium / high - They discuss models exposing multiple reasoning-effort settings and automatic routing to choose one. Maximum reasoning mode latency: about 20 minutes - Sebastian says he uses pro mode for full chapter checks that can take around 20 minutes. Chapter length checked in pro mode: 40 pages - He uploads roughly 40-page PDFs to get consistency and numbering checks. Book page count in early access: 360 pages - Sebastian says the early-access version of his reasoning book is already 360 pages long. Book completion target: April - He hopes to finish the remaining chapter by April. Model scale example: 670 billion to 1 trillion parameters - He cites Kimi scaling a DeepSeek-style architecture from 670B to 1T parameters. Chinese New Year timing: before Chinese New Year - He notes model releases often happen around Chinese New Year and expects more releases then. PDF context use: 200-page PDF - He says many 200-page PDFs can now be placed directly in context without RAG or fine-tuning. Inference scaling sample: 60,000 answers - He describes generating huge numbers of candidate answers for verifiable-reward training.

Pivotal Quotes: "most of the interesting things are happening now on the post-training front, in the reasoning realm" — Sebastian Roshka: He frames the current frontier of LLM research as post-training and reasoning rather than pre-training scale. "It becomes more and more popular to use or to have the LLM use tools too" — Sebastian Roshka: He explains why tool use is becoming central to improving accuracy and reducing hallucinations. "the biggest, I guess, achievement right now that could be made if that gets if someone finds out a way that this works" — Sebastian Roshka: He is describing reliable continual learning as one of the field’s biggest unsolved goals.

Implications: Expect LLM progress to come from better reasoning, tool use, and agentic orchestration more than from brand-new core architectures. Builders should focus on workflow integration, routing, and verifiable tasks, while keeping an eye on emerging efficiency models and eventual continual-learning breakthroughs.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast