Episode Summary
Executive Summary: The episode examines whether OpenAI likely trained ChatGPT on copyrighted books without permission, using author Douglas Preston’s experience as a trigger for a broader legal and economic discussion. It compares the emerging AI lawsuits to earlier copyright battles involving Google Books and Spotify, arguing that these cases may shape whether AI training is ruled fair use or becomes a costly licensing regime.
Main Topics: Douglas Preston’s ChatGPT discovery (Priority: 5/5): Author Douglas Preston tests ChatGPT on his own novels and finds it can answer detailed questions about characters, settings, and plot elements, suggesting his books may have been ingested during training without permission. OpenAI and the authors’ class-action lawsuit (Priority: 5/5): Preston, George R.R. Martin, and other authors, via the Authors Guild, sue OpenAI alleging large-scale copyright infringement and seeking discovery into what training data was used. How large language models are trained (Priority: 4/5): The episode explains that ChatGPT is built on a large language model trained by predicting next words using massive amounts of text, much of which may be copyrighted. Google Books as a fair use precedent (Priority: 5/5): The Google Books case is presented as a key legal precedent where scanning millions of books was ultimately ruled fair use because the searchable database was considered transformative and socially beneficial. Spotify’s streaming lawsuit as a settlement model (Priority: 4/5): Spotify’s copyright dispute shows how class actions often become vehicles for settlement, especially when a company has trouble identifying and licensing every rights holder. Economic and legal stakes for AI companies (Priority: 5/5): The episode weighs the possibility that OpenAI could face either a fair-use win or massive damages/licensing obligations if courts find the training copies unlawful.
Key Arguments: ChatGPT’s detailed knowledge of Preston’s books suggests it may have been trained on copyrighted texts rather than just public summaries or reviews. OpenAI’s refusal to disclose its training data strengthens authors’ need for discovery in court. Copyright law allows fair use, but the law is subjective and depends on factors like market harm and transformative purpose. Google Books suggests courts may allow large-scale copying when the end product is transformative and socially valuable. The authors argue AI training causes direct market harm because generative systems can compete with original creative works. The Spotify case suggests class-action lawsuits can push companies toward comprehensive settlements rather than trials. If OpenAI loses on copyright, retraining from scratch would be extremely costly and practically difficult. If OpenAI wins on fair use, it could validate broad AI training practices across the industry.
Data Points: Douglas Preston books written: about 40 - Preston describes his overall writing output during the episode. OpenAI author lawsuit plaintiffs: 16 authors initially referenced - Douglas Preston, George R.R. Martin, and 15 other authors are described as bringing the class-action suit. Google Books copyrighted share: around 80% - The episode says roughly 80% of the scanned Google Books collection was still under copyright. Copyright infringement damages: up to $150,000 per infringement - Statutory damages discussed in the hypothetical scenario where each copied book could count as an infringement. Hypothetical OpenAI damages estimate: $15 billion - Back-of-the-envelope estimate using 10,000 authors and 10 books per author. Spotify music rights coverage: 90% licensed / 10% tricky remainder - Most songs were managed by large companies, but the final 10% lacked clean licensing paths. Copyright class actions going to trial: 1 out of over 100 - A study cited by UCLA law professor Zian Tang found only one copyright class action reached full trial.
Pivotal Quotes: "It knew my characters. It knew their names. It knew the settings. It knew everything." — Douglas Preston: Preston reacting to ChatGPT’s detailed answers about his own novels. "This case is very different than that case because here the harm is so visible." — Mary Rasenberger: Explaining why the OpenAI lawsuit may not resemble the Google Books fair-use ruling. "In over 100 copyright class actions, only one ever went all the way to a full trial." — Zian Tang: Describing how copyright class actions usually end in settlement rather than trial.
Implications: The episode suggests AI companies may face either sweeping fair-use validation or a forced licensing system. For creators, the outcome will shape whether their works can be used to train AI without permission and compensation.
About Planet Money
Wanna see a trick? Give us any topic and we can tie it back to the economy. At Planet Money, we explore the forces that shape our lives and bring you along for the ride. Don't just understand the economy – understand the world.Wanna go deeper? Subscribe to Planet Money+ and get sponsor-free episodes of Planet Money, The Indicator, and Planet Money Summer School. Plus access to bonus content. It's a new way to support the show you love. Learn more at plus.npr.org/planetmoney