Hard Fork
Hard Fork

Google Eats Rocks + A Win for A.I. Interpretability + Safety Vibe Check

“Pass me the nontoxic glue and a couple of rocks, because it’s time to whip up a meal with Google’s new A.I. Overviews.”

Featured Speakers

The New York Times Host

Topics Discussed

Episode Summary

Executive Summary: This Hard Fork episode centered on Google’s troubled AI Overviews, which produced absurd and potentially dangerous answers, and a leaked cache of search documents suggesting Google’s ranking systems rely on more signals than it publicly admits. The show then pivoted to a major Anthropic interpretability breakthrough that begins to map what language models “think,” and closed with a broader discussion of AI safety, OpenAI governance, and the industry’s growing efforts to regulate frontier models.

Main Topics: Google AI Overviews and search errors (Priority: 5/5): The hosts unpacked viral examples of Google’s AI Overviews recommending rocks, glue pizza, and other nonsensical or hazardous advice, arguing that the system is both embarrassing and potentially dangerous because it compresses unreliable web content into authoritative-sounding answers. What the Google leak suggests about search (Priority: 5/5): A large alleged leak of Google Search documentation was discussed as evidence that Google may rely on signals such as click behavior and Chrome data more than it publicly states, and may increasingly favor trusted brands over smaller publishers. The future of the web and Google’s power (Priority: 4/5): The conversation broadened into a structural critique: Google’s dominance in search and advertising is accelerating a web that is more centralized, less open, and more dependent on a few large platforms and brands. Anthropic’s interpretability breakthrough (Priority: 5/5): Josh Batson joined to explain Anthropic’s work mapping millions of features inside Claude, showing that internal model behavior can be partially decoded and manipulated, including activating features such as the Golden Gate Bridge persona or sycophancy. AI safety, model behavior, and the limits of current safeguards (Priority: 4/5): The episode ended with discussion of OpenAI’s new voice model access, the rollback for safety reasons, departures of safety researchers, and emerging institutional responses such as government frameworks, safety institutes, and a California AI bill. OpenAI governance and credibility problems (Priority: 4/5): The hosts revisited the Sam Altman firing saga and Helen Toner’s claims that OpenAI withheld key information from the board, using it as evidence that AI safety concerns remain entangled with corporate governance and trust.

Key Arguments: Google’s AI Overviews are not just occasionally wrong; they can ingest satirical or low-quality sources and present them as factual answers, which makes Google more directly responsible for harmful outputs than traditional search results. The examples are partly cherry-picked, but the underlying issue is real: Google is still using a system that cannot reliably distinguish trustworthy from untrustworthy sources at scale. Because search is a “fathead” product, Google can manually audit the most common queries, but it cannot fully eliminate the reputational and legal risk of mistakes across the long tail of searches. The Google search ecosystem and the broader web are being shaped by a few giant companies, especially Google, whose control over referral traffic and ad revenue helps determine which publishers survive. The leaked Google documents suggest search ranking may privilege trusted brands and use more behavioral data than Google publicly admits, reinforcing the idea that small publishers have less room to compete. Anthropic’s research shows large models are not pure black boxes; they contain interpretable features that can be mapped, monitored, and potentially used to detect unsafe behavior before outputs are generated. The Golden Gate Claude experiments demonstrate that internal features are not just correlations: turning them up can directly alter the model’s behavior and persona. Interpretability matters for safety because it may reveal when a model is lying, sycophantic, or about to produce harmful content, enabling upstream interventions rather than only post-output moderation. AI safety concerns are reemerging across the industry through committees, safety frameworks, government institutes, and regulation, but skepticism remains because companies continue racing ahead while publicly emphasizing caution.

Data Points: Most popular Google search terms share: 500 search terms = 8.4% of all Google search volume - Used to explain why Google can manually audit the highest-volume AI Overviews. Google AI model feature count: about 10 million features - Anthropic’s interpretability work identified roughly this many understandable features in Claude 3 Sonnet. Potential total feature space: hundreds of millions or even billions - Estimated scale of possible features in a large model, far beyond what current methods can cheaply map. Time with OpenAI voice demo: about 40 minutes - Casey Newton described briefly testing OpenAI’s new voice assistant before access was revoked. Company review target: top 10,000 AI Overviews - Kevin suggested Google could manually audit a relatively small number of high-volume queries to reduce errors. Model generations mentioned: GPT-2 to GPT-3 to GPT-4 - Used to illustrate the pace and magnitude of capability jumps in frontier models. Number of AI companies in Seoul safety pledge: 16 - Leading AI companies made voluntary Frontier AI Safety Commitments at a Seoul summit.

Pivotal Quotes: "the vast majority of AI overviews provide high-quality information with links to dig deeper on the web." — Google (statement read by Kevin Roose): Google’s response to criticism that AI Overviews produced dangerous and absurd outputs. "they are automated plagiarism." — Casey Newton quoting Rusty Foster: A critique of AI Overviews as a system that scrapes and lightly rewrites the web without meaningful attribution or originality. "we could actually get Claude to break its own safety rule." — Josh Batson: Explanation of how activating or dialing features can alter model behavior and expose latent unsafe capabilities.

Implications: Google faces a trust and liability problem as it turns search into generated answers. At the same time, interpretability research is making AI systems more legible, which could improve safety, monitoring, and regulation as frontier models become more powerful.

🔓 Sign Up for Unlimited Episode Search

About Hard Fork

“Hard Fork” is a show about the future that’s already here. Each week, journalists Kevin Roose and Casey Newton explore and make sense of the latest in the rapidly changing world of tech. Unlock full access to New York Times podcasts and explore everything from politics to pop culture. Subscribe today at nytimes.com/podcasts or on Apple Podcasts and Spotify. Also, for more podcasts and narrated articles, download The New York Times app at nytimes.com/app.

View all episodes from Hard Fork