The Cognitive Revolution
The Cognitive Revolution

Unbounded AI-Assisted Research with Elicit Founders Andreas Stuhlmüller and Jungwon Byun

In this episode, Nathan sits down with Elicit co-founders Andreas Stuhlmüller and Jungwon Byun to discuss their mission to make AI-assisted research more accessible and reliable. Learn about their unique approach to task decomposition, which allows language models to accurately tackle complex resear

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The episode explores Elicit’s vision for trustworthy AI-assisted research: turning overwhelming scientific literature into systematic, transparent, and extensible workflows. The founders explain how notebooks, task decomposition, and factored verification improve reliability, how evaluation and human gold standards guide product quality, and why scalable infrastructure and careful safety controls are essential as models improve.

Main Topics: Elicit’s product vision: systematic, transparent, unbounded research (Priority: 5/5): The founders frame Elicit as an AI research assistant built to help researchers navigate exploding literatures, make reasoning explicit, and keep an auditable trail of how conclusions were reached. Notebooks are presented as a major product shift, not just a feature. Task decomposition as the core reliability strategy (Priority: 5/5): They argue that breaking big research questions into small, checkable subtasks is more reliable than asking one model to answer everything at once. This improves accuracy, makes launches easier, and creates clearer interfaces for both users and evaluators. Evaluation design and factored verification (Priority: 5/5): A major theme is that defining the task properly is most of the battle. They discuss using human gold standards, model-assisted evaluation, and factored verification methods that check claims individually rather than scoring whole outputs holistically. Scalable infrastructure and the ‘exoscale’ future (Priority: 4/5): The conversation looks ahead to running models over tens of thousands or hundreds of thousands of papers with models supervising models. The founders describe this as turning compute into more work through transparent architectures rather than opaque black boxes. Model stack, chain of thought, and multimodal reasoning (Priority: 4/5): They describe a pragmatic model stack that mixes fine-tuned and frontier models. Chain-of-thought reasoning is used broadly, especially for questions requiring multi-hop reasoning or data extraction from tables, which became much more feasible with multimodal models. Commercial growth, user segments, and pricing/value dynamics (Priority: 4/5): The founders discuss rapid growth, a broadening user base beyond biomedicine, and a willingness among professional users to pay for deeper, more accurate analysis. They emphasize batch workflows and high-value use cases over casual chat interactions. Safety, misuse, and governance in dual-use domains (Priority: 4/5): They address the challenge of misuse in high-stakes domains, especially dual-use scientific research. Their approach combines intent understanding, review processes for risky use cases, and the possibility of restricting certain workflows or versions of the product.

Key Arguments: Task decomposition improves both reliability and product iteration because smaller subtasks are easier to supervise, evaluate, and ship. Research should be made systematic, transparent, and unbounded so users can iterate, audit, and extend prior work instead of restarting from scratch. The best evaluation is based on well-defined user goals and naturally occurring human gold standards, not just generic model scores. Factored verification can catch subtle hallucinations by evaluating individual claims independently rather than judging an entire summary at once. Chain-of-thought is essential for many research tasks because current transformers have fixed compute per token and cannot reliably infer multi-step answers without intermediate reasoning. Using models to help with evaluation works best when the subtasks are constrained and easy to verify; otherwise human judgment is still necessary. Speed matters, but for high-stakes research accuracy and comprehensiveness often matter more; users may wait longer or pay more for better results. A transparent, modular system can scale compute into work more safely and effectively than simply increasing model size and hoping for the best. Safety against misuse should be built into intent detection and workflow design, especially for dual-use scientific queries. The company’s commercial mission and public-benefit framing are aligned because trustworthiness and accuracy are also what make the product valuable.

Data Points: Annual recurring revenue: $1 million ARR - The founders said Elicit reached this milestone after launching subscriptions. Seed round: $9 million - Raised after spinning Elicit out into its own venture. Time to launch subscriptions: 4 months - Elicit reached $1M ARR shortly after subscriptions launched. Feature cadence: 1.4 weeks on average - They said they have been shipping new features roughly weekly, with some larger launches taking longer. Hallucination rate improvement: From about 1.5 to 0.5 hallucinations on average - They described a reduction in hallucinations after applying stricter verification and distillation methods. Team size: 12 people - The founders said Elicit is still a very small team. Human judgment dataset size: A few hundred to a few thousand data points - They said this range is usually enough for evaluation and fine-tuning tasks. Search candidate set: ~500 papers initially - They described the first stage of the search pipeline as retrieving roughly 500 strong candidates. Product adoption pattern: Thousands to tens of thousands of dollars per project - Some teams spend at this scale on high-value research workflows. Individual willingness to pay: Hundreds of dollars - They noted some individual users are willing to pay this much for clear value. Model performance on extraction: Outperformed trained research staff in side-by-side tests - They said predefined extraction columns have often beaten human manual extraction in evaluations.

Pivotal Quotes: "what is easy to supervise" — Andreas Stuhlmüller: He reframed task decomposition around creating subtasks that can be checked reliably, not just tasks that models can do. "the key question is: how do we as a society want to turn compute into more work?" — Andreas Stuhlmüller: He used this to contrast brute-force scaling with more transparent, structured architectures. "We want to make research systematic, transparent, and unbounded." — Jung Wan Byun: She summarized the core product philosophy behind notebooks and the broader Elicit platform.

Implications: For researchers and builders, the episode suggests that trustworthy AI will come from modular workflows, strong evaluation, and transparent logs—not just bigger models. It also points to a future where compute is increasingly organized through scaffolding and supervision, enabling higher-stakes use cases.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution