The TWIML AI Podcast
The TWIML AI Podcast

Google I/O 2025 Special Edition - #733

Today, I’m excited to share a special crossover edition of the podcast recorded live from Google I/O 2025! In this episode, I join Shawn Wang aka Swyx from the Latent Space Podcast, to interview Logan Kilpatrick and Shrestha Basu Mallick, PMs at Google DeepMind working on AI Studio and the Gemini AP

Topics Discussed

Episode Summary

Executive Summary: Live from Google I/O 2025, the panel explored Gemini’s developer roadmap: Google’s strategy to keep one unified model while still using forks for specialized capabilities, new controls for thinking models, native audio output, URL Context, implicit caching, and the maturation of Gemini Live for real-time voice/video apps. The discussion emphasized latency, composability, and bringing experimental capabilities back into Gemini’s mainline model.

Main Topics: Google’s “one model” Gemini strategy (Priority: 5/5): The speakers argue Google’s core differentiation is building one unified Gemini model, even if separate forks are used temporarily to develop features before merging them back into the mainline model. Developer controls for reasoning models (Priority: 5/5): New controls like thinking budgets, the ability to disable thinking, and thought summaries give developers more control over how Gemini 2.5 Pro reasons and how much detail is exposed. Native audio and multimodal interaction (Priority: 5/5): Native audio output, audio-to-audio workflows, multilingual switching, and real-time transcription/translation were highlighted as major advances that make voice-first applications more viable. Gemini Live infrastructure and real-time app challenges (Priority: 4/5): The panel discussed the technical and product challenges of building live voice/video systems: session length, VAD tuning, tool calling, state changes, and the high commitment required to adopt live APIs. Caching and cost/latency optimization (Priority: 4/5): Implicit caching was presented as a major developer win because it automatically reduces cost without extra setup, while explicit caching still matters for repeat-heavy workloads. URL Context and research-agent use cases (Priority: 3/5): URL Context, especially when paired with search, was positioned as a tool for retrieving richer web-page information and enabling research-agent style experiences. Frameworks, protocols, and the voice AI ecosystem (Priority: 3/5): Daily and Pipecat’s perspective emphasized that production voice AI requires surrounding model capability with networking, orchestration, turn detection, and context-management infrastructure.

Key Arguments: Google wants Gemini to remain a single model family rather than a fragmented collection of separate models, using forks only as temporary development paths before reintegration. Giving developers knobs—thinking budgets, thought summaries, VAD sensitivity, asynchronous function calling, and session controls—is central to making Gemini useful in production. Native audio and audio-to-audio systems reduce friction for real-time applications and are especially promising for multilingual, conversational, and agentic experiences. Implicit caching is valuable because it lowers developer cost automatically and removes operational burden. Real-time voice apps are harder than text apps because they require low latency, long-lived sessions, and more bespoke infrastructure, which raises switching costs for developers. The long-term vision is for more capabilities to migrate from external frameworks into the API/model itself as the ecosystem matures. Generative UI and other future experiences may become practical only when models are fast enough and tightly integrated with multimodal generation.

Data Points: Gemini 2.5 Pro thinking budgets: Coming with GA in a couple of weeks; early June for a raw non-reasoning mode - Developers will be able to control reasoning cost and disable thinking Thought summaries: Live now - A partial substitute for exposing full model thoughts Implicit caching: Already available - Automatically saves money without developer setup Live API video session length: Originally about 5 minutes of video - Early constraint discussed as a production challenge Live API audio session length: Originally 15 to 20 minutes of audio - Early constraint discussed as a production challenge Latency target for voice systems: 500 to 700 milliseconds - Referenced as the difficult real-time bar for live voice interactions Official language support: 24 languages - Gemini Live multilingual support mentioned during the wishlist discussion

Pivotal Quotes: "We want to make one model and it’s the Gemini model and not have the sort of splintering of all these different capabilities." — Logan Kilpatrick / discussion of Corai's view: Explaining Google’s strategy for Gemini development and how specialized forks should eventually merge back "Implicit caching happened. ... You don’t have to do anything. It just works right now and you’re saving money." — Logan Kilpatrick: Describing the developer experience benefit of the new caching approach "It is really, really hard to bring all these components together and still get latency down to where it needs to be, you know, in the 500 to 700 millisecond range." — Shrestha Basumalik: On the core engineering challenge of real-time voice agents

Implications: Gemini is moving toward a more unified, controllable, and multimodal platform for developers. For builders, this means easier cost management, better real-time voice tooling, and more powerful multimodal workflows; for the industry, it signals increasing convergence of model capabilities into fewer, stronger foundation models.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast