The Cognitive Revolution
The Cognitive Revolution

Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research

Andreas Stuhlmüller and Jungwon Byun return to discuss how Elicit is building trusted reasoning workflows for scientific research as frontier models grow more powerful but less transparent. They explain process supervision, domain-specific reasoning primitives, and world models that make evidence, c

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The conversation explores how Elicit is evolving from literature-search and summarization into a system for trustworthy reasoning at scale. The founders argue that process supervision, structured workflows, and explicit world models can make AI more reliable for high-stakes scientific and business decisions, especially in life sciences. They emphasize evaluation, evidence quality, calibration, and certificates of reasoning over raw chain-of-thought monitoring.

Main Topics: Process supervision vs. final-answer optimization (Priority: 5/5): The founders revisit Elicit’s original thesis: models should be judged on step-by-step reasoning, not just outputs. They argue current models remain easy to push around and can claim work they did not actually perform, making process checks essential for trustworthy decision support. Structured workflows and DSL-based reasoning (Priority: 5/5): Elicit rebuilt its platform around a domain-specific language that orchestrates reasoning primitives as guaranteed workflows. This lets frontier models dynamically compose structured research processes while preserving consistency, reproducibility, and auditability over large document sets. Evidence quality, source ranking, and confidence calibration (Priority: 5/5): They discuss how Elicit evaluates heterogeneous evidence from papers, web sources, and filings, and why metadata like citations or journal prestige is an imperfect proxy. A major goal is decomposing claims into support levels and confidence estimates so users can act on them appropriately. World models and continual learning outside model weights (Priority: 4/5): Elicit is experimenting with explicit, inspectable knowledge representations—graph-like where appropriate, spreadsheet-like where needed—to support prediction, intervention, and counterfactual reasoning. The founders frame this as continual learning humans can inspect rather than hidden weight updates. Automation inside Elicit: the LINE and recursive self-improvement (Priority: 4/5): The company is using AI to automate its own software engineering pipeline through a system called the LINE, which routes features from spec to code review to deployment. They see this as a practical testbed for whether AI can keep a company progressing with less human intervention. Tokens, model choice, and the economics of AI adoption (Priority: 3/5): They debate rising token spend, whether customers will tolerate it, and how much value added intelligence justifies cost. The founders prefer orchestrating multiple models by task rather than exposing a model picker, since model differences are subtle but operationally important.

Key Arguments: Models are still too unstable and too easily influenced to be trusted as standalone decision tools for high-stakes work. Process supervision matters because users need to know not only whether an answer is right, but whether the right reasoning steps were actually carried out. A DSL and structured workflow layer can guarantee that research processes scale consistently across thousands of items. Evidence quality should be judged from methodology and content, not just citations, impact factors, or source prestige. Claims should be broken into components and assigned confidence levels so users understand how strongly each is supported. Chain-of-thought monitoring is insufficient by itself; independent checks and certificates of reasoning are needed. World models can help AI make internally coherent predictions, interventions, and counterfactual judgments over complex evidence. AI will likely automate generation before evaluation; humans will increasingly focus on checking, verifying, and defining what good looks like. In software engineering, automation is already valuable for simpler features, but higher-risk work still requires human oversight. Model convergence does not eliminate the need for orchestration, because different models still have micro-differences in extraction, support, and reliability.

Data Points: Years since last podcast: 2 years - The hosts note it has been two years since their prior conversation. Top life sciences customers: 7 of the top 20 - Elicit works formally with seven of the top 20 life sciences companies. Time to redo a paper analysis: 10 minutes - Andreas described uploading an old paper and having Elicit redo computational experiments and analysis in roughly ten minutes. Issue merges per week via the LINE: 30 to 50 - Their automated software-engineering pipeline now merges about 30–50 issues per week. Papers analyzed in a toxicology example: about 100 papers - They describe asking research agents to analyze roughly 100 papers and then catching the model admitting it did not actually do so. Papers filtered in a cancer literature review: about 5,000 papers - Andreas mentioned ending up with around 5,000 papers before filtering to the most relevant ones in a cancer-related review. Personal token spend: about $2,000 per week - Andreas estimated his own weekly token spend on API use. Customer use cases: 7 of top 20 life sciences companies; 30-50 merges/week - These figures anchor the scale of Elicit’s enterprise traction and internal automation efforts. Confidence calibration example: 20% probability - Nathan referenced GPT-4’s model card calibration example where 20% self-confidence was approximately accurate. Price sensitivity: 2x to 3x possible spend increase - Andreas said he could probably double or triple his personal token spend, but not much more.

Pivotal Quotes: "the models are not trained on process. The models are trained to produce outputs" — Andreas Stulmuller: Explaining why models may claim to have done work they did not actually complete. "I think knowing when the models are making you better or worse at decision-making is actually pretty subtle." — Andreas Stulmuller: Closing reflection on the need for careful evaluation of AI assistance. "I think the fundamental property you get from discretization is error correction" — Andreas Stulmuller: Explaining why discrete representations and structured languages may outperform fully continuous/neural approaches in some domains.

Implications: The episode suggests the next wave of AI value will come from verifiable workflows, explicit knowledge structures, and better evaluation—not just bigger models. For science and enterprise, trust, calibration, and evidence quality may matter more than raw generation power.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution