The Vergecast
The Vergecast

How to train your data

Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The A

Featured Speakers

Vox Media Podcast Network HostAlex Reisner Guest

Topics Discussed

Episode Summary

Executive Summary: The episode centers on why AI training data matters more than most people realize, how companies secretly source and curate it, and why this is fueling both AI capability and public backlash. Alex Reisner argues that training data defines a model’s outputs, that secrecy protects both competitive advantage and questionable sourcing, and that current AI firms often rely on academic/nonprofit intermediaries to launder access to large datasets. He rejects synthetic data as a real replacement and warns that without stronger norms or laws, exploitation will intensify.

Main Topics: Training data as the foundation of model behavior (Priority: 5/5): Reisner explains that what a model is trained on largely determines what it can produce, making training data central to capability and style. Why AI companies keep training data secret (Priority: 5/5): The discussion covers competitive secrecy, legal exposure, and the likelihood that much training material was acquired without creator consent or awareness. How researchers reverse-engineer AI datasets (Priority: 4/5): Reisner describes using developer forums, research papers, and open-source communities to infer what data is inside major models and datasets. Common Crawl, YouTube, and the infrastructure of dataset building (Priority: 5/5): The conversation highlights major data sources and intermediaries—especially Common Crawl and YouTube—that have become standard inputs for AI training. Data laundering through universities and nonprofits (Priority: 4/5): The guests discuss how AI firms often route large-scale scraping through academic or nonprofit partners to create distance from direct responsibility. Synthetic data and model collapse (Priority: 5/5): Reisner argues synthetic data is oversold and says training models on their own outputs leads to degradation rather than improvement. Future of AI training and creator compensation (Priority: 4/5): The episode ends with the idea that companies may increasingly pay creators to generate AI-training content, but Reisner argues a better legal/business framework is needed.

Key Arguments: Training data is the most important determinant of a model’s output; it shapes style, quality, and capabilities more than branding or architecture. AI companies keep data sources secret both to preserve competitive advantage and because many sources were acquired in ways creators would oppose. Open-source and academic communities expose more about training practices, but corporate research disclosures have become more constrained over time. Common Crawl became a foundational web source for early large language models, but training on broad web content produced noisy, low-quality results. YouTube is a dominant AI training source because it is easy to download from and includes vast amounts of music and video that are harder to obtain elsewhere. AI firms often use universities and nonprofits as intermediaries to scrape or compile data, which functions as a form of data laundering. Synthetic data is not a viable long-term substitute at present; training on model outputs can cause model collapse and degrade performance. A future where AI models are trained on paid creator-made content may emerge, but it still requires better rules for valuing and compensating data.

Data Points: Apple price increases: $100 to $500 - David Pierce cites major price hikes across iPads and MacBooks due to supply chain and AI-era shortages. Apple TV price: $200 - Apple TV is described as rising to roughly the price of 900 Amazon Firesticks. Disney settlement fund: $50 million - Viewers with YouTube TV or DirecTV Stream may be eligible for payouts from a settlement involving Disney. Common Crawl timeline: late 2000s / 2009 - Reisner says Common Crawl has been crawling the web since around 2009. Common Crawl web additions: a few hundred million web pages each month - The dataset expands monthly with massive new scraping runs. OpenAI/industry training scale: tens of millions of songs - Reisner references a Google paper saying models were trained on tens of millions of songs. Music dataset size: 12 million songs - An organization called Lion in Europe has a dataset of songs from YouTube. Creator payouts for AI training: over $10 million - Reisner mentions a company claiming to have paid creators this amount to make content for AI training. Settlement case start: 2022 - The Disney-related live TV pricing case began in 2022. Number of AI researchers/entities cited: 10,000+ papers - Reisner says Common Crawl has been cited by over 10,000 papers.

Pivotal Quotes: "I think it is potentially the most important aspect of a model, what it's trained on." — Alex Reisner: Reisner explains why training data matters more than most people assume. "If you don't let AI robots crawl your data, you essentially don't exist." — Rich Screnta: Quoted by Reisner as an example of the industry’s increasingly uncompromising view of data access. "Absolutely not." — Alex Reisner: His blunt answer to whether synthetic data is actually a workable replacement for human-generated training data.

Implications: For listeners and the industry, the episode suggests AI’s biggest unresolved issue is not just model design but data legitimacy and value. Expect more conflict over creator rights, more scrutiny of “academic” data pipelines, and pressure for laws or compensation schemes.

🔓 Sign Up for Unlimited Episode Search

About The Vergecast

The Vergecast is the flagship podcast from The Verge about small gadgets, Big Tech, and everything in between. Every Friday, hosts Nilay Patel and David Pierce hang out and make sense of the week’s most important technology news. And every Tuesday, David leads a selection of The Verge’s expert staffers in an exploration of how gadgets and software affect our lives – and which ones you should bring into yours.

View all episodes from The Vergecast