The TWIML AI Podcast
The TWIML AI Podcast

From Math Olympiads to Navier-Stokes: How Fast Is AI Progressing? with Greg Burnham - #778

AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes. In this episode, Greg Burnham, who leads AI capabilities research at Epoch AI, joins us to examine what that progress says about wher

Featured Speakers

Greg Burnham Guest

Topics Discussed

Episode Summary

Executive Summary: Greg Burnham argues that AI has moved beyond exam-style testing into real-world task performance, making robust capability measurement urgent. Epoch AI uses benchmark stitching and specialized evaluations to track rapid progress in math, research, learning, and work tasks, while probing whether models can generate genuinely new ideas. The main conclusion: capabilities are rising fast and smoothly, but true AI-driven discovery and on-the-fly learning remain key unknowns.

Main Topics: Why AI capabilities research matters now (Priority: 5/5): Burnham says AI is no longer just solving academic exams; its strengths and weaknesses are affecting real work, so benchmarking must track practical impact and emerging risks. Benchmark stitching and the Epoch Capabilities Index (Priority: 5/5): Epoch combines many correlated benchmarks into a single trend measure, arguing that model performance still rises smoothly across generations even as individual benchmarks saturate. Shallow vs deep improvement (Priority: 5/5): The conversation distinguishes between lab-driven data collection and training to cover tasks broadly versus true generalization that transfers capabilities without direct task-specific training. AI and mathematical discovery (Priority: 5/5): Math is used as a testbed for creativity and reasoning, with recent progress from grade-school math to frontier unsolved problems like Erdős results and Navier-Stokes, though AI still relies heavily on prior human work. Detecting new ideas and research taste (Priority: 4/5): A major open question is whether AI can originate new ideas or pick productive research directions. Epoch tries to detect this with hard unsolved problems and post hoc expert review, but evidence remains limited. On-the-fly learning and harness vs model effects (Priority: 4/5): Epoch is testing whether AI can learn during deployment via context and tools, using board games and job-like tasks. Current systems improve with better harnesses and instructions, but do not reliably learn like humans. Broader capabilities: work, cyber, and physical tasks (Priority: 4/5): Beyond math, Epoch is examining whether AI can do jobs, make charts, generate insights, assemble furniture, and perform controlled hacking, aiming to identify threshold crossings before they become disruptive.

Key Arguments: AI capabilities have reached the point where they materially affect the world, so measuring them is no longer just an academic exercise. Benchmark performance across generations is still rising smoothly; apparent saturation on old tests is mostly due to needing harder benchmarks. The correlation across benchmarks allows Epoch to stitch scores into a unified capabilities index that shows a linear trend over time. Much of AI progress appears driven by scale in compute and training data, not yet by recursive self-improvement or software-only intelligence explosions. Current systems may solve hard math by combining extensive prior knowledge with patience and brute-force exploration, but they still rarely originate genuinely new concepts. AI can often reproduce or extend human mathematical proofs, but this is different from building new theory from scratch. The main tripwire for future discontinuity is whether models can improve AI research itself by generating novel algorithmic ideas. On-the-fly learning is a critical missing capability; models can do better with detailed strategy guides or harnesses, but they do not reliably improve across repeated tries like humans do. Real-world tasks such as data insights, visuals, and open-ended project ideation remain fuzzier and harder for AI than benchmarked reasoning tasks. Capability jumps often appear as threshold crossings in practice, so even steady trend measurement is valuable for anticipating when a system becomes dangerous or broadly useful.

Data Points: GSM8K scale: 8,000 word problems - Referenced as the grade-school math benchmark that early LLMs struggled with. Frontier Math tiers: 4 tiers - Epoch’s newer frontier math benchmark is divided into four difficulty tiers, with breakthrough as the highest. Frontier math difficulty range: advanced undergraduate to advanced graduate student - Describes the original Frontier Math problems. Erdős problem count: well over 1,000, maybe a couple thousand - The number of problems Paul Erdős proposed over his life. Math progress timeline: late 2024 to 2026 - Burnham traces rapid progress from high-school math competition success to solving frontier unsolved problems. Learning benchmark board game: Earthborne Rangers - The obscure real board game used to test learning on the fly. Board game popularity: about 500th most popular - Its low popularity helps reduce the chance it was heavily represented in training data. Move 37: 37th move - AlphaGo’s famous novel move cited as a benchmark for identifying AI-generated novelty. AI review cadence: same AI instance plays again and again - Used in the learning-on-the-fly benchmark to test whether the system improves across repeated plays. Training-data/compute drivers: compute and training data supply chain - Burnham argues these are major engines of progress, absent evidence of discontinuities. Math competition progress: International Math Olympiad gold medal by early 2025/2026 - Illustrates how quickly AI moved from grade-school math to elite contest performance.

Pivotal Quotes: "We're past the academic exam. Phase of understanding AI capabilities." — Greg Burnham: Explaining why benchmark design must move beyond traditional tests to real-world capability measurement. "The trend is crazy. Like, you gotta hold both of these together." — Greg Burnham: Describing how benchmark progress looks smooth and linear even though the underlying advancement pace is extraordinary. "If AI ever does cross some threshold where it's able to do rapid improvements in AI algorithms themselves, then we want that tripwire to sound." — Greg Burnham: On Epoch’s goal of detecting whether models can improve AI research and trigger a discontinuity in progress.

Implications: AI capability measurement is shifting from classroom-style evaluation to real-world and open-ended tasks. For industry, the big unknowns are novel idea generation, self-improvement, and human-like learning during deployment.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast