The TWIML AI Podcast
The TWIML AI Podcast

Rethinking Pre-Training for Agentic AI with Aakanksha Chowdhery - #759

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-traini

Featured Speakers

Akancha Chowdhury Guest

Topics Discussed

Episode Summary

Executive Summary: Akancha Chowdhury argues that next-generation AI agents will require rethinking pre-training, not just post-training, because static benchmarks and next-token objectives underrepresent long-horizon, environment-interactive tasks like coding and research. She emphasizes long-context reasoning, recoverability from failures, tool learning, data quality, and better benchmarks as the key levers, while remaining conservative about major architectural departures.

Main Topics: Why agentic AI forces a rethink of pre-training (Priority: 5/5): The conversation centers on the idea that useful agents must interact with environments, plan over multiple steps, recover from mistakes, and use tools—capabilities that static benchmarks and traditional pre-training do not adequately measure. Scaling experience from PaLM and Gemini (Priority: 4/5): Akancha explains how training large models at Google exposed the complexity of pre-training at scale, where every system component can fail and debugging becomes an end-to-end challenge across the stack. Architecture vs. objective vs. data (Priority: 5/5): She argues that architecture is only one part of the solution. Training data quality, objective design, and what the model is asked to predict matter just as much, and she is cautious about betting on unproven architectural shifts. Long-form reasoning and context handling (Priority: 5/5): Long-form reasoning includes reasoning over long contexts, synthesizing many sources, planning across multiple steps, and adapting after failed trajectories. The context-length bottleneck is a central limitation in current systems. Memory and tool use as model capabilities (Priority: 4/5): Memory should be treated as a tool the model can reason about, not a replacement for better reasoning. Tool learning and exploration are positioned as essential for agents that can generalize to new environments. Data, synthetic augmentation, and reasoning traces (Priority: 4/5): She stresses that high-quality, diverse training data is crucial, and that reasoning traces or tool-use traces may need to be generated or augmented from dominant real-world sources rather than relying on fully synthetic distributions. Benchmarks for agentic intelligence (Priority: 4/5): Current benchmarks saturate quickly, so frontier labs increasingly build in-house evaluations for long-context reasoning, multi-step planning, failure recovery, and tool use. New real-world workflows will likely become future benchmarks.

Key Arguments: Static benchmarks like LUE, GSM8K, and math tasks are insufficient for agentic models because they do not test interaction with environments or multi-step goal completion. Agentic capability depends on long-context reasoning, planning, recovery from failure, and tool learning, not just better chatbot behavior. Pre-training should incorporate some of the things currently handled in post-training, such as reasoning and tool-use behavior, to make these capabilities more fundamental. Architecture changes may help, but the bigger levers are training data and loss objectives; attention may need to better support long-range recall and synthesis. Transformers remain the most vetted choice for large-scale training, so labs should be conservative about radical architectural bets unless they have strong evidence. Loss design can teach behaviors such as fill-in-the-middle code completion and tool selection by masking the right parts of the sequence. High-quality curated data can improve compute efficiency, and reasoning traces are likely a key future data source—but only if they preserve the natural data distribution. Model capabilities often emerge at larger scale first and can later be distilled into smaller, cheaper models; small models alone may miss these emergent behaviors. Benchmarks must evolve toward realistic workflows, varying task horizons, and failure-recovery scenarios to measure what matters for agents.

Data Points: PaLM parameter count: 540 billion parameters - Akancha cites PaLM as Google’s largest model at the time of training. Team size at Reflection: about 60 researchers - She describes Reflection as a small frontier lab building open agentic models. PaLM development period: two or two and a half months - Used to illustrate why pre-training requires thinking across the full stack during long training runs. Reasoning model emergence timing: about a year ago - She says the first reasoning models with reinforcement learning appeared around a year prior to the interview. DeepSeek training scale: 15 trillion tokens - Referenced as an example of the scale of modern pre-training corpora. Data scale comparison: tens of trillions of tokens - Estimated scale of large frontier pre-training datasets more generally.

Pivotal Quotes: "For the longest time, we were measuring pre-training on static benchmarks. If we want these models to be useful as agents, they need to be able to interact with environments." — Akancha Chowdhury: Core thesis on why agentic capabilities require changing how models are trained and evaluated. "This is not just a post-training problem to achieve these set of capabilities that we want in the next generation of models." — Akancha Chowdhury: Argument that agentic behavior must be built into pre-training, not only added later. "I see memory as an additional tool in addition to the attention mechanisms." — Akancha Chowdhury: Explains how memory fits into an agentic model stack without replacing reasoning.

Implications: AI teams should prioritize long-horizon evaluations, better data curation, and training recipes that support reasoning, tool use, and recovery from failure. The next leap in agents will likely come from integrating these behaviors into pre-training, then distilling them into smaller models.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast