The Cognitive Revolution
The Cognitive Revolution

E14: The Reasoning Revolution with Ought's Jungwon Byun and Andreas Stuhlmüller

We've looked forward to today's episode since we launched the show! Andreas Stuhlmuller and Jungwon Byun are the co-founders of Ought, a product-driven research lab that develops mechanisms for delegating open-ended thinking to advanced machine learning systems. Their flagship product, Eli

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The conversation explores Augt/Elicit’s belief that AI should amplify rigorous thinking through transparent, compositional workflows rather than opaque end-to-end prediction. The founders explain how their “interpretability by construction” approach emerged from early human-based experiments, why breaking work into small evaluable tasks matters for reliability and safety, and how Elicit is evolving from literature review into broader research and decision-support infrastructure.

Main Topics: Why current ML struggles with substantial thought (Priority: 5/5): The founders argue that AI excels when objectives are clear, but reasoning and decision-making lack clean reward functions, making proxy optimization and superficial outputs a major risk. Interpretability by construction vs post-hoc explanation (Priority: 5/5): Augt’s core philosophy is to design systems whose steps are human-legible and directly supervised, rather than trying to interpret opaque models after the fact. Early experiments with human 'context windows' (Priority: 4/5): Before modern LLMs, the team studied how humans could collaboratively solve tasks with limited context, which foreshadowed current compositional AI workflows and error-propagation issues. Elicit’s product evolution and research workflows (Priority: 5/5): Elicit started as a literature review assistant and is expanding toward richer structured querying, concept extraction, and eventually custom research workflows for institutions. Evaluation, grounding, and reliability (Priority: 4/5): They emphasize source-linked answers, checklists, and multi-stage retrieval/ranking to make outputs easier to verify and safer for high-stakes use. Process supervision, AI safety, and governance (Priority: 4/5): The founders see process supervision as both a product design principle and a safety strategy that could push the industry toward more trustworthy AI systems. Future vision: discovery, decision support, and social coordination (Priority: 4/5): They envision AI helping discover new research domains, produce rigorous policy memos, and improve coordination by surfacing evidence and clarifying disagreements.

Key Arguments: Reasoning is hard to optimize because there is no obvious objective like 'win the game'; outcome-based training often rewards what looks persuasive rather than what is genuinely well-reasoned. Human-legible decomposition is essential for trust in high-stakes settings because users need to inspect and evaluate each step, not just a final answer. Opaque chain-of-thought is not enough; the answer should causally depend on the reasoning process, and the process should be visible and checkable. The most scalable way to achieve transparent AI is not manual decomposition forever, but using AI itself to help decompose tasks into smaller, understandable pieces. Breaking tasks into substeps reduces ambiguity, improves grounding, and makes errors easier to detect and correct before they compound. Elicit’s workflow design favors structured checklists and linked evidence over single numeric trust scores because scores are easier to game and hide context. Research workflows generalize across domains: the same compositional approach can work for medicine, machine learning, policy, finance, and consulting if the underlying activity is evidence-based reasoning. Future AI safety and governance will require pressure not just from model labs but also from end-user products and regulators demanding transparency, robustness, and process supervision.

Data Points: Mission age: 6 years - The founders say the mission statement about automating open-ended reasoning was written six years ago, before GPT-style language models existed. User count: ~250,000 users - They mention Elicit has reached roughly 250,000 users, motivating more cost-efficient model deployment. Weekly usage: 10,000+ users more than once a week - They describe a core cohort that uses Elicit regularly. Model size used in deployment: 11 billion parameters - They say Flan-T5 XXL (11B) has been especially useful for deployment on their tasks. GPT-4 vs GPT-3.5 win rate: ~70/30 (about 2:1) - The host cites the GPT-4 technical report as evidence that evaluation remains noisy and far from perfect. Context window task duration: 1 minute per participant - In early relay-style experiments, each human participant had about a minute to contribute before passing work onward. Example error rate compounding: 10% subtask error across 20 tasks - They explain that even modest per-step error becomes likely to surface when many tasks are composed. Core user mix: ~60% academic researchers - They estimate that about 60% of Elicit users are academic researchers, with the rest spanning think tanks, government, finance, consulting, and medicine.

Pivotal Quotes: "“It’s especially unclear that ML will help in matters that require substantial thought.”" — Andreas Stuhmüller: Introduced as the mission-level concern behind Augt/Elicit’s focus on reasoning rather than generic prediction. "“Interpretability by construction.”" — Andreas Stuhmüller / Jung Wan Byun: Their phrase for building workflows where the model’s steps are transparent and directly supervised rather than inferred afterward. "“Phase one: discover what is known. Phase two: discover what is unknown. Phase three: decide in the face of the unknowable.”" — Jung Wan Byun: The founders’ three-phase roadmap for Elicit’s long-term evolution beyond literature review.

Implications: The episode suggests a near-term market for AI systems that are grounded, decomposable, and auditable—not just fluent. For listeners and builders, the message is clear: the winning applications may be those that make AI more checkable, reliable, and useful for serious decisions.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution