Hard Fork
Hard Fork

Google's Gemini 3 Is Here: A Special Early Look

Maybe more than other model releases, this one seems to have the attention of Google’s competitors. Will it put the company at the top of the A.I. leaderboard?

Featured Speakers

The New York Times HostDemis Hassabis Guest

Topics Discussed

Episode Summary

Executive Summary: This episode centers on Google’s launch of Gemini 3 and a Hard Fork interview with Demis Hassabis and Josh Woodward. The hosts frame the release as a potential turning point in the AI race, highlighting stronger reasoning, coding, and new interface-generation features, while probing Google’s claims about efficiency, safety, AGI timelines, and whether the company is regaining leadership from rivals.

Main Topics: Gemini 3 launch and positioning (Priority: 5/5): The hosts preview Google’s new frontier model, describing it as a major release that may shift perceptions of Google from AI laggard to leader. New product capabilities (Priority: 5/5): Google emphasizes improved reasoning, coding, and especially generative interfaces that can build interactive experiences instead of plain text answers. Benchmarks and performance gains (Priority: 5/5): The episode focuses on benchmark improvements over Gemini 2.5 Pro, presented as evidence that Gemini 3 is materially better across many tasks. Google’s distribution and efficiency advantage (Priority: 4/5): The discussion highlights Google’s ability to deploy Gemini into Search, the Gemini app, and future Workspace integrations at scale and potentially low cost. AGI timelines and scaling laws (Priority: 4/5): Demis Hassabis reiterates a 5-to-10-year AGI timeline and says Gemini 3 is on track, while acknowledging more breakthroughs are still needed. Safety, misuse, and bubble concerns (Priority: 4/5): The interview examines safety testing, cyber risks, and whether the AI industry is in a bubble, with Google arguing it is well positioned either way. Consumer framing and AI as a tool (Priority: 3/5): Google presents Gemini less as a companion and more as a productivity and learning tool that helps users complete tasks.

Key Arguments: Gemini 3 is stronger than Gemini 2.5 Pro across reasoning, coding, expressiveness, and user experience, according to Google’s internal benchmarks and demos. A major differentiator is generative UI: the model can create custom interactive interfaces, not just text replies. Google believes its biggest advantage is distribution; it can embed Gemini into products used by billions, including Search, Android, YouTube, Maps, and Workspace. The company says it has improved efficiency enough to serve AI at massive scale, including AI Overviews and search-adjacent experiences. Demis Hassabis says Gemini 3 does not change his AGI timeline; he still expects AGI in five to 10 years and thinks more breakthroughs are needed. Google frames Gemini as a practical super-tool for learning, research, writing, coding, and task completion rather than an erotic or emotional companion. The company says it has done extensive safety testing and is especially cautious because improved tool use and function calling can also increase cyber risk. On bubbles, Hassabis argues some parts of AI are overheated, but Alphabet has both near-term monetization and long-term optionality across many sectors.

Data Points: Humanties Last Exam score (Gemini 2.5 Pro): 21.6% - Google cited this score as the prior model’s result on a difficult interdisciplinary benchmark. Humanties Last Exam score (Gemini 3 Pro): 37.5% - Google said Gemini 3 Pro substantially outperforms its predecessor on the same benchmark. AGI timeline: 5 to 10 years - Demis Hassabis repeated this estimate during the interview. College student offer: 1 year free - Google announced free access to a paid Gemini version for all U.S. college students. Model availability: This week / Tuesday launch window - The hosts said Gemini 3 would roll out in the Gemini app and AI Mode this week, with full testing after the Tuesday release. Benchmark count: More than a dozen examples - Google presented many benchmark comparisons showing Gemini 3 outperforming Gemini 2.5 Pro. LM Arena: 1500 ELO - Hassabis referenced cracking the 1500 ELO mark as one benchmark proxy for progress.

Pivotal Quotes: "this is just kind of better than their last model, Gemini 2.5 Pro. In basically all respects." — Kevin / host: Summary of Google’s core launch message about Gemini 3. "I think we're five to 10 years away from AGI" — Demis Hassabis: On whether Gemini 3 changes his AGI timeline. "The new type of metric that I think we get excited about... how many tasks did we help you complete in your day?" — Josh Woodward: On how Google wants to measure Gemini’s value to users.

Implications: Gemini 3 suggests Google may be closing the AI gap and using distribution plus efficiency to convert model gains into real product advantage. If the claims hold, the next battleground is less chat quality and more task completion, embedded AI, and scaled consumer adoption.

🔓 Sign Up for Unlimited Episode Search

About Hard Fork

“Hard Fork” is a show about the future that’s already here. Each week, journalists Kevin Roose and Casey Newton explore and make sense of the latest in the rapidly changing world of tech. Unlock full access to New York Times podcasts and explore everything from politics to pop culture. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. Also, for more podcasts and narrated articles, download The New York Times app at nytimes.com/app.

View all episodes from Hard Fork