Episode Summary
Executive Summary: The episode evaluates whether current AI scaling laws can reach AGI without running out of high-quality data. Using order-of-magnitude estimates, the hosts argue that GPT-5-level training may need ~100T tokens and that the world already generates far more raw data than that, though usable high-quality text may be the real bottleneck. They compare data across email, Twitter, YouTube, genomics, astronomy, code, and synthetic data, concluding that scale is plausible but quality, access, and curriculum remain central uncertainties.
Main Topics: Scaling hypothesis and why data matters (Priority: 5/5): The conversation frames AI progress through the scaling hypothesis: bigger models and better training data/algorithms can keep pushing toward human-level intelligence, so the key question becomes whether enough high-quality data exists. GPT-3 to GPT-5 data and compute extrapolation (Priority: 5/5): The hosts extrapolate from GPT-3 and GPT-4 training scales to estimate GPT-5 could require around 100 trillion tokens and roughly 100x more compute than GPT-4, implying a billion-dollar or larger training run. Raw data abundance across modalities (Priority: 5/5): A large portion of the episode surveys total data generation in email, Twitter, YouTube, astronomy, genomics, code, and the internet, arguing that raw quantity is not scarce even if usable quality is. The quality bottleneck and high-quality text scarcity (Priority: 4/5): A counterargument from the Epic AI estimate is that high-quality text may run out around 2024-2026, suggesting only a small fraction of internet-scale data is actually useful for frontier pretraining. Multimodal expansion and non-text data (Priority: 4/5): The discussion highlights that image, video, DNA, weather, and scientific datasets may be easier to exploit than text because humans cannot natively interpret them, and because models may have ample scale headroom there. Synthetic data, code, and recursive improvement (Priority: 4/5): Code and model-generated data are presented as especially powerful sources of synthetic training material, with the possibility that models can create hard problems, solve them, and iteratively improve themselves. Economic and geopolitical feasibility of massive training runs (Priority: 3/5): The episode considers trillion-dollar training runs, linking them to national-scale industrial efforts and suggesting only a few actors could plausibly fund and execute them.
Key Arguments: AI progress should be grounded in actual data and compute quantities, not crowd sentiment or market-style bullish/bearish narratives. If scaling laws continue, GPT-5 may need about 10x more data than GPT-4, consistent with a 100x larger compute budget. There is vastly more raw data in the world than the amount needed for GPT-5-scale training; the limiting factor is what fraction is high-quality and accessible. Email traffic alone may contain enough useful text at very small sampling fractions, while Twitter is closer to GPT-4 scale but still substantial. YouTube, astronomy, and genomics each contain enormous reserves of data, especially once modalities are tokenized or compressed appropriately. Multimodal datasets can be much larger than text and may be easier to exploit for AGI because they align with tasks humans cannot directly “speak” natively, like DNA. Synthetic data and code may be the most scalable route because models can generate, test, and refine their own training examples. A trillion-dollar training run might be feasible in the early 2030s if compute trends, algorithmic efficiency, and investment growth continue. Recursive self-improvement is still weak in current models but may become viable once models can reliably create and solve increasingly difficult tasks. Even if raw data is abundant, curriculum design and quality filtering remain crucial for turning scale into capability.
Data Points: GPT-3 training scale: ~1 trillion tokens; ~600 GB text - Baseline training size used as the starting point for extrapolating future models. GPT-4 training scale: ~10 trillion tokens - Anchor estimate for current frontier-scale pretraining. Hypothetical GPT-5 training scale: ~100 trillion tokens - Extrapolated 10x increase over GPT-4 under scaling trends. GPT-3 compute budget: under $500,000 - Roughly cited early training run cost estimate. GPT-4 compute budget: 10^24 to 10^25 FLOPs; mid-tens of millions of dollars - Referenced as an estimated training cost range for GPT-4. Hypothetical GPT-5 compute budget: ~100x GPT-4; ~$1B to a few billion - Derived from extrapolating training spend and scaling laws. Global data generated in 2020: ~64 zettabytes - Total world data output used to compare against training-data needs. Raw data needed for GPT-5 target: 10^14 tokens - Equivalent to the estimated 100 trillion token target. Fraction of 2020 global data needed: ~1 in 100 million - Share of raw global data that would need to be high-quality text to hit GPT-5 target. Emails sent per day: ~333 billion / 10^11 - Used to show how email alone could supply enough text volume. Twitter data generated per year: ~33 terabytes of text - Used as a proxy for mediocre-quality internet text data. YouTube annual uploads: ~500 million hours per year - Cited as a massive multimodal data reservoir. Astronomy data scale: ~1-2 exabytes - Compared with text-scale targets and used as a high-volume modality example. Genomics data generated per year: ~40 exabytes - Shown as an especially large and fast-growing domain. Human genome size: ~3.4 gigabases / ~3 GB raw - Used to anchor genomics data scale. Evo training data: 300 billion tokens - Open-source genomics dataset from prokaryotic and phage genomes. GPT-4 vision training image tokens: ~2 trillion image tokens - Example of how large a multimodal dataset can be. Low-res image tokenization: ~85 tokens for a 256x256 image - Illustrates pixel-to-token compression ratio in multimodal systems. Pixel-to-token compression: ~1000x - Approximate compression used to estimate YouTube token equivalents. YouTube tokenized volume: ~10^15 tokens/year - Annual YouTube uploads after rough token conversion. Library of Congress text: ~10 trillion words / ~10 TB - Sanity-check benchmark for high-quality text quantity. High-quality text estimate from Epic AI: ~10 trillion words - Suggested amount of usable text before running out. Trillion-dollar training run timing: ~2029-2033 - Projected window if current compute and efficiency trends continue. Brain processing estimate: ~11 million bits per second - Used to compare human cognition to model training data scales. Human brain lifetime data processed: ~2-3 quadrillion bytes over 70 years - Compared to GPT-4 and future training requirements. GPT-4 relative to human brain lifetime processing: ~200-300x smaller - Implied by the cited brain-processing estimate. Trillion-dollar run vs GPT-4: ~8,500x GPT-4 training data scale - Used as a top-down estimate for future frontier training. OpenAI/ChatGPT daily activity estimate: ~25 million daily active users; ~15 queries/day/user - Used to estimate volume of synthetic text generated by models. Synthetic output volume: ~45 trillion generated words - Approximate ChatGPT-generated text volume inferred from usage assumptions.
Pivotal Quotes: "It's becoming awfully clear to me that these models are truly approximating their data sets to an incredible degree." — Jay Becker: Cited to support the idea that model performance is tightly tied to scale and data quality. "trained on the same data set for long enough, pretty much every model with enough weights in training time converges to the same point." — Jay Becker: Used to argue that scaling and sufficient training time can cause convergence toward similar capabilities. "one in 100 million parts of this raw data would have to be high-quality text" — Nick Gannon: Summarizes the optimistic case that the world generates far more raw data than frontier models need.
Implications: Listeners should think of AGI as a data-compute optimization problem, not a pure narrative race. The world likely has enough raw data, but access, filtering, and high-quality curriculum design will determine whether scaling continues or stalls.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co