Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

[LIVE] Anthropic Distillation & How Models Cheat (SWE-Bench Dead) | Nathan Lambert & Sebastian Raschka

Swyx joined SAIL! Thank you SAIL Media, Prof. Tom Yeh, 8Lee, Hamid Bagheri, c9n, and many others for tuning into SAIL Live #6 with Nathan Lambert and Sebastian Raschka, PhD. Sharing here for the LS paid subscribers. We covered: This is a public episode. If you'd like to discuss this with other

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The conversation centered on two AI-industry flashpoints: Anthropic’s allegations that prominent Chinese labs are distilling frontier models via API usage, and the growing unreliability of popular coding benchmarks like SWE-bench Verified. The hosts debated how to distinguish legitimate evals from model extraction, whether API providers can or should police usage, and why benchmark saturation, leakage, and hidden flaws make current scores less meaningful.

Main Topics: Anthropic’s anti-distillation blog post (Priority: 5/5): The speakers dissect Anthropic’s claim that Chinese labs used distributed accounts to distill model outputs from its APIs. They debated whether this is truly an “attack” or simply standard competitive behavior in a GPU-constrained environment. What distillation is and how it works (Priority: 5/5): They defined distillation as training a smaller model on outputs or logits from a larger model, noting it is common both for internal model compression and for external model imitation via API-generated synthetic data. Detecting distillation vs. legitimate benchmarking (Priority: 5/5): A major thread was the ambiguity between high-volume evaluation and data collection for training. The speakers argued that scale, repeated prompts, and distribution patterns are likely the main signals providers use to infer abuse. SWE-bench Verified and benchmark failure modes (Priority: 5/5): The discussion shifted to coding benchmarks, especially SWE-bench Verified, its curation process, saturation, and the discovery that a large fraction of tasks were still unsound or impossible to solve. Future of evals and benchmark design (Priority: 4/5): They argued that the next generation of benchmarks must use fresher data, private-public splits, diversified repos/languages, and better task validation, especially for agentic and UI-level capabilities. API business models and model release strategy (Priority: 4/5): The hosts discussed whether frontier labs should prioritize APIs or product-only releases, and whether restricting model access could better protect against distillation while preserving business value.

Key Arguments: Distillation is a normal and widely used technique in ML, but using one company’s API outputs to train a competing model is what labs now consider abusive. It is inherently difficult to distinguish internal evaluation from distillation because both can involve repeated API calls, the same prompts, and similar scripts. Providers likely infer abuse from volume, repetition, and cross-account patterns rather than any single request. SWE-bench Verified’s scores became nearly meaningless because many frontier models clustered around the same 80-something percent range. Even a heavily vetted benchmark can contain impossible or flawed tasks, showing that benchmark quality degrades unless continuously audited. Model memorization can surface unexpected leakage, including future API versions or repository details, which can distort evals. API businesses are attractive but competitively brutal; some labs may increasingly reserve their best models for products rather than public APIs. The next wave of evaluation will need better end-to-end measures for agentic and UI tasks, not just code or math. Benchmarking and distillation are both constrained by data quality, rate limits, cost, and the difficulty of generating enough tokens at scale.

Data Points: SWE-bench Verified size: 500 tasks - OpenAI’s curated subset of SWE-bench discussed as a benchmark for coding agents. Human review for SWE-bench Verified: 3 reviewers per task - Used by OpenAI to vet tasks for quality. Follow-up audit for SWE-bench Verified: 6 reviewers per task - A later audit reportedly found many tasks still flawed or unsolvable. Unsovable tasks in SWE-bench Verified audit: 59% - The speakers cited a later review finding that a majority of tasks could not truly be solved as written. OpenAI run subset issue: Subset of the subset - At launch, OpenAI reportedly could not run all 500 tasks and used a smaller denominator for some reports. Original SWE-bench launch score: ~13% - Referenced as the initial benchmark performance level before broad saturation. Current SWE-bench Verified scores: 80%+ - Speakers noted most frontier models cluster tightly around this level, reducing discriminative value. Task granularity: 0.2% minimum step (approx.) - Discussed in relation to 500-task benchmarks and how score increments are constrained. Distillation scale: 10s of billions to 100B tokens - The conversation estimated the token scale required for meaningful distillation runs. Model generation speed: ~40 tokens/second - Used to illustrate how time-consuming large-scale synthetic data generation can be. Model sizes mentioned: 671B parameters - DeepSeek R1 was cited as a large model used to train smaller variants. Smaller model variants: 1B/3B range - Examples of distilled smaller models intended for local use. API timing window: 2–4 week lead - OpenAI reportedly released some Codex variants internally ahead of public/API release. Open-source evaluation issue age: 2023-era problems - SWE-bench was described as drawing from older open-source tasks that are now easier to memorize or leak.

Pivotal Quotes: "I think this just means more content for sale." — Nathan / host: Opening remarks welcoming the new writer to the SAIL coalition. "Distillation in short is training a smaller model on the outputs of a larger model, basically." — Swix: Clear definition offered early in the discussion to ground the Anthropic controversy. "Benchmarks are hard to make, and we need new ones." — Nathan / host: Conclusion of the SWE-bench critique, emphasizing benchmark drift and saturation.

Implications: Frontier labs will keep fighting over API usage, model extraction, and benchmark credibility. Expect more private evals, stricter usage monitoring, and a shift toward fresher, harder, and more agentic benchmarks as current leaderboards lose signal.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast