Episode Summary
Executive Summary: This episode centers on Google’s Gemini and Live API updates from IO, with a strong focus on developer controls, multimodal and audio capabilities, and the infrastructure needed for real-time AI. The speakers highlight thinking budgets, thought summaries, native audio output, URL Context, implicit caching, and Gemini Diffusion, while emphasizing that Google’s strategy is to converge capabilities into one main Gemini model without losing specialized offshoots for production use cases.
Main Topics: IO announcements for Gemini developers (Priority: 5/5): The guests recap the most developer-relevant IO launches, especially controls around reasoning, summaries, audio, and retrieval that make Gemini more usable in production. Thinking budgets and thought summaries (Priority: 5/5): They discuss giving developers more control over reasoning in Gemini 2.5 Pro, including disabling thinking and using thought summaries instead of full chain-of-thought. Native audio and multilingual voice (Priority: 5/5): Native audio output is highlighted as a major milestone, especially for natural-sounding speech and switching between languages like Bengali and English. Live API challenges and real-time voice infrastructure (Priority: 5/5): The conversation digs into latency, session length, tool calling, voice activity detection, and the difficulty of building reliable real-time voice agents. Single-model strategy vs specialized offshoots (Priority: 4/5): The speakers explain Google’s philosophy of building one unified Gemini model while allowing specialized forks or side models to mature capabilities before merging them back. Caching, URL Context, and retrieval (Priority: 4/5): Implicit caching and URL Context are presented as practical developer features that reduce cost and improve web-grounded research workflows. Gemini Diffusion and generative UI (Priority: 3/5): They speculate that diffusion-based language generation could enable dynamic, code-generated user interfaces that are created on the fly as users interact.
Key Arguments: Google is trying to give developers more control on top of powerful models, not just better base models, through features like thinking budgets, thought summaries, and caching. Thought summaries are a compromise between exposing full reasoning and hiding it entirely, and developer feedback will determine how useful they are. Native audio output is a major step forward because it supports natural speech, multilingual switching, and more human-like interaction. The Live API is harder to build on than standard text APIs because real-time voice systems require tighter commitments, lower latency, and more bespoke infrastructure. Google’s long-term strategy is to converge capabilities into one Gemini model, even if some features first appear in separate specialized models. Voice applications require a stack of supporting infrastructure such as VAD, turn detection, context management, and networking protocols like WebRTC/WebSockets. Implicit caching is valuable because it saves developers money automatically without requiring manual cache management. Gemini Diffusion may unlock generative UI experiences by making token generation fast enough to build interfaces dynamically during user interaction.
Data Points: Thinking budget availability for Gemini 2.5 Pro: Early June - Planned rollout for disabling thinking and controlling reasoning in 2.5 Pro Thought summaries status: Live now - Available as a partial alternative to full thoughts Thinking budgets in 2.5 Pro GA: In a couple of weeks - Expected with the GA model release Thinking budgets in 2.5 Flash: Already available - Existing support for reasoning controls in 2.5 Flash Live API audio session length (early limit): 15 to 20 minutes - Initial audio session duration limit mentioned for production users Live API video session length (early limit): About 5 minutes - Initial video duration limit mentioned for production users Latency target for voice agents: 500 to 700 milliseconds - Desired response range for real-time conversational systems Supported languages: 24 languages - Officially supported languages for Gemini voice capabilities Model capability example: Klingon - Mentioned as an unofficial language the model can still respond to in a demo
Pivotal Quotes: "If you just want 2.5 Pro as like a raw non-reasoning model, we'll have that hopefully in early June." — Shrestha: Discussing developer controls over reasoning in Gemini 2.5 Pro "The killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens." — Logan: Explaining why Gemini Diffusion could matter beyond speed "It is really, really hard to bring all these components together and still get latency down to where it needs to be, you know, in the 500 to 700 millisecond range." — Shrestha: Describing the engineering challenge of real-time voice agents
Implications: For developers, Gemini is becoming more controllable, multimodal, and production-ready, especially for voice and real-time apps. For the industry, the episode signals a shift toward unified models plus specialized infrastructure, with latency, multilinguality, and developer ergonomics as key battlegrounds.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast