The a16z Podcast
The a16z Podcast

Columbia CS Professor: Why LLMs Can’t Discover New Science

From GPT-1 to GPT-5, LLMs have made tremendous progress in modeling human language. But can they go beyond that to make new discoveries and move the needle on scientific progress? We sat down with distinguished Columbia CS professor Vishal Misra to discuss this, plus why chain-of-thought reasoning w

Featured Speakers

a16z Host

Topics Discussed

Episode Summary

Executive Summary: The conversation centers on Vishal Misra’s formal theory of LLM behavior: they operate over sparse “Bayesian manifolds,” producing next-token distributions that are confident within trained regions but hallucinate when pushed outside them. The discussion traces how this framework explains in-context learning, chain-of-thought, and the limits of self-improvement, while arguing that true AGI will require creating new science and new manifolds rather than merely scaling current LLMs.

Main Topics: Formal models for understanding LLMs (Priority: 5/5): Misra and the hosts argue that empirical prompt-tweaking is insufficient; LLMs need mathematical models to explain their behavior. His matrix/manifold abstraction helps predict when models are reliable versus when they drift into hallucination. LLMs as next-token predictors on sparse manifolds (Priority: 5/5): LLMs are described as producing next-token distributions over a huge vocabulary and navigating a compressed, sparse state space. Confidence is high when prompts stay on-manifold and collapses when they move off-manifold. In-context learning and Bayesian reasoning (Priority: 5/5): Misra explains that few-shot prompting works because the model updates its posterior from prompt evidence without changing weights, making in-context learning analogous to Bayesian inference. RAG’s accidental origin in cricket stats (Priority: 4/5): Misra recounts how trying to fix CricInfo’s clunky stats interface led him to create a retrieval-style system that predated and foreshadowed modern RAG, using examples plus structured query translation. Limits of scaling and recursive self-improvement (Priority: 5/5): The episode argues that feeding model outputs back into models cannot create fundamentally new information. At best, LLMs can unroll known algorithms or refine existing structures, not invent new paradigms. AGI as the creation of new science (Priority: 5/5): Misra defines AGI not as humanlike conversation or tool use, but as the ability to generate new science, math, and paradigms—like relativity or quantum mechanics—beyond the training distribution. What might come next: new architectures (Priority: 4/5): The discussion turns to multimodality, simulation, and energy-based or ARC-style reasoning as possible directions beyond transformers. Misra believes more data/compute alone will only smooth existing manifolds.

Key Arguments: LLMs are best understood as systems that generate next-token distributions over a compressed, sparse state space rather than as general reasoners. The prompt acts as evidence; richer context reduces prediction entropy and increases confidence in the model’s output. Hallucinations occur when the model departs from the learned manifold and must extrapolate beyond its training-supported region. Chain-of-thought works because it decomposes hard problems into smaller steps the model has seen before, lowering uncertainty at each step. In-context learning is not a separate mechanism; it uses the same inference process as ordinary prompting, only with examples included in the context. The Matrix abstraction shows why LLMs can interpolate across a vast prompt space without explicitly storing every prompt-token combination. Recursive self-improvement is limited because models cannot generate genuinely new information without external novelty; they can only recombine or expand known patterns. AGI should be reserved for systems that can produce new scientific or mathematical breakthroughs, not just improve benchmark performance. Adding more data or compute to current architectures will likely yield incremental gains, not the creation of new manifolds or paradigms. A new architecture, possibly involving multimodal grounding or internal simulation, is needed to move beyond transformer-based limits.

Data Points: GPT-3 context window: 2048 tokens - Misra cites GPT-3’s limited context window as a reason it could not directly handle the complexity of CricInfo’s backend. Cricket query examples used in retrieval system: About 1500 examples - Misra says he built a query database with roughly 1,500 examples to support the RAG-like system. Top examples retrieved for completion: 6 or 7 most relevant examples - For each new query, the system selected a small set of examples as prefix context for GPT-3. StatsGuru interface complexity: 25 checkboxes, 15 text fields, 18 dropdowns - He describes the original CricInfo stats web form as overly complex and hard to use. CrickInfo traffic claim: More hits than Yahoo - Misra says CrickInfo was once among the most popular sites in the world before India’s cricket boom. Production use of the query system: Since September 2021 - The accidental RAG-style system had been running in production since this date, predating ChatGPT by about 15 months. Time gap to ChatGPT: About 15 months earlier - Misra notes his production system existed well before the public ChatGPT release. Prompt/vocabulary scale example: 50,000 tokens - Used to illustrate the size of the next-token distribution space in LLMs. Theoretical matrix size reference: More rows than the number of atoms across all known galaxies - A rhetorical illustration of how enormous the full prompt-token matrix would be. H-index joke: 60 - The host jokingly refers to Martine Cassado as having an H-index of 60 while discussing technical understanding.

Pivotal Quotes: "AGI will be when we are able to create new science, new results, new math." — Vishal Misra: Misra defines AGI as going beyond trained knowledge to create genuinely novel paradigms. "The output of the LLM is the inductive closure of what it has been trained on." — Vishal Misra: He uses this to explain why current models can recombine known ideas but not generate truly new information. "The day an LLM can create a large software project without any babysitting is the day I’ll be a little bit convinced." — Vishal Misra: Misra says autonomous, reliable software creation would materially shift his view of LLM capability, though it still would not prove AGI.

Implications: For builders, the takeaway is to treat LLMs as powerful but bounded inference engines. Progress likely requires new architectures and grounding, not just scaling. For listeners, the episode offers a rigorous way to distinguish impressive pattern completion from genuine discovery.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast