The Cognitive Revolution
The Cognitive Revolution

Mathematical Superintelligence: Harmonic's Vlad Tenev & Tudor Achim on IMO Gold & Theories of Everything

Vlad Tenev and Tudor Achim from Harmonic explain how they built Aristotle, an AI system that reaches International Mathematical Olympiad gold-medal performance using formally verified Lean proofs. They unpack the architecture behind mathematical superintelligence, including Monte Carlo Tree Search,

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The episode explores Harmonic’s vision for “mathematical superintelligence” through Aristotle, an AI system that generates Lean-formalized proofs with automatic verification. The founders argue math is fundamentally reasoning, Lean is transforming mathematics into a collaborative, checkable software-like discipline, and formally verified outputs will scale from math into software and other quantitative domains. They also discuss entropy, taste, safety, and a future where AI helps generate multiple coherent scientific theories.

Main Topics: Math as reasoning and the philosophy of mathematics (Priority: 5/5): The founders frame mathematics not as an isolated discipline but as the process of reasoning in small, verifiable logical steps. They connect this view to physics, engineering, and the idea that math underlies many other domains of human knowledge. What Lean is and why formal verification matters (Priority: 5/5): Lean is presented as a dependently typed programming language and proof assistant that can encode logical statements, verify proofs via a small trusted kernel, and reduce reliance on human peer review. The discussion emphasizes Lean as a new foundation for mathematical trust and collaboration. How Aristotle works: search, lemmas, and geometry (Priority: 5/5): Aristotle combines Monte Carlo tree search, informal lemma generation, auto-formalization, theory building, and a geometry-specific module. The system uses language models at multiple levels but always outputs Lean proofs that can be checked mechanically. Formal vs. informal reasoning and the future of verification (Priority: 5/5): The conversation argues that formal outputs are the right long-term path for both math and software because verification costs must not scale linearly with complexity. They expect formal methods to spread from competition math into mission-critical software and eventually broader programming. Entropy, hallucination, and model training (Priority: 4/5): The founders say hallucination/entropy is necessary for exploration and discovery, even in reasoning systems. They emphasize RL-style search over heavy human expert curation, and say the system is optimized for the net present value of future proofs rather than elegance panels. Vision, safety, and the 2030 outlook (Priority: 4/5): The episode closes on a grand vision: AI systems helping produce theories for everything mathematically expressible, while humans remain in charge. Safety concerns are framed as emerging first through cybersecurity and constrained external actions, not immediate catastrophic autonomy.

Key Arguments: Mathematics is fundamentally reasoning, i.e., breaking understanding into small logical steps that others can verify. Lean changes mathematics from an informal, chalkboard-based practice into a software-like, collaborative, machine-checkable workflow. Formal verification reduces or eliminates the need for traditional peer review because correctness is checked by the kernel. Aristotle’s architecture uses multiple reasoning regimes: tree search for hard proofs, informal lemma generation for context management, and geometry-specific tactics for constrained domains. The future of software, like math, will increasingly be formal because verification cost must stay low even as AI-generated output grows. AI systems need entropy/hallucination during training and search to explore novel solution spaces. Taste in research should come from the community’s revealed preferences and problem selection, not a small internal gatekeeping group. The most likely near-term safety risks are cybersecurity and misuse of autonomous API-connected systems, not Aristotle-style proof generation. By 2030, AI may generate multiple coherent scientific theories, shifting the bottleneck from explanation to experimental validation.

Data Points: IMO performance year: 2025 - Aristotle achieved gold-medal-level performance at the International Mathematical Olympiad in 2025. IMO comparison: OpenAI, Google DeepMind, and Harmonic all missed question 6 - The speakers noted all three gold-level systems failed the same hardest problem on the IMO. Public benchmark timing: End of the year - They said Aristotle topped out the Verena benchmark at the end of the year with public API users. Company start year: 2023 - Harmonic described its founding as coinciding with key maturity in Lean 4 and GPT-4. Lean version transition: Lean 3 to Lean 4 - They said Lean 4 was around the maturity point when the company launched. Axioms used: 3 - They described Lean proofs as relying on three axioms in addition to the calculus of constructions. Axiom of choice characterization: Non-empty set implies you can choose an element - Used as the intuitive example of one of the axioms supporting Lean formalization. Proof length example: Under a tweet each - They said the technical axioms are very short when written as mathematical statements. Mathlib description: Largest digital repository of mathematical knowledge - They compared Mathlib to the consolidated foundation of formalized mathematics. Problem scale example: 5,000-page proofs - They used this as a future scenario to explain why formal verification becomes essential. Potential science timeframe: 2030 - They projected that by 2030 AI could help produce theoretical explanations for everything mathematically expressible. Safety action space: Constrained external interfaces - They said current risk is lower because Aristotle’s actions are tightly constrained and not broadly agentic.

Pivotal Quotes: "Mathematics is reasoning." — Vlad Tenev: Core thesis of Harmonic: math is a formalized form of general reasoning, not just an esoteric subject. "Lean is the best programming language ever created." — Tudor Achim: Used to explain why Lean is the foundation for formally verifiable math and software. "Hallucinations are a key part of the training process for models like this." — Vlad Tenev: Explaining why entropy and exploration are necessary in training reasoning systems.

Implications: The episode argues that formal verification will become central to AI-era math and software, enabling trusted, scalable reasoning. If Harmonic’s thesis holds, future systems may produce machine-checkable discoveries, safer software, and broader scientific progress with humans steering the goals.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution