Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

Maggie, Linus, Geoffrey, and the LS crew are reuniting for our second annual AI UX demo day in SF on Apr 28. Sign up to demo here! And don’t forget tickets for the AI Engineer World’s Fair — for early birds who join before keynote announcements! It’s become fashionable for many AI startups to projec

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The episode traces Elicit’s evolution from a nonprofit research lab focused on AI alignment and process supervision into a Public Benefit Corporation building an AI research assistant. Andreas and Jawuan explain how their product philosophy—systematic, transparent, and unbounded—shaped features like literature review automation, tabular extraction, and notebooks, while emphasizing evaluation, citations, and controllable workflows over generic chat.

Main Topics: Origins: from AI alignment research to Elicit (Priority: 5/5): Andreas describes a lifelong interest in AI, academic work at MIT and Stanford, and the creation of Odd as a nonprofit research lab to study how to build reasoning into machines before productizing the work as Elicit. Co-founder fit and company formation (Priority: 4/5): Jawuan explains how she joined after years of interest in AI for mental health and research workflows, and how the pair used a long mutual evaluation process to assess values, skills, and working style before co-founding. Elicit’s product evolution and core use case (Priority: 5/5): The product moved from forecasting and rough literature-review prototypes to a research assistant focused on understanding what is known, especially through literature summarization, paper extraction, and systematic-review workflows. Process supervision, evaluation, and trust (Priority: 5/5): A central thesis is that AI should be trained and evaluated on the steps experts take, not just outputs, enabling better debugging, lower hallucination rates, and more trustworthy research assistance. Notebooks, agents, and workflow orchestration (Priority: 4/5): The team explains why they are building notebooks instead of pure chat: to support reusable, stepwise, composable research processes that can scale from small analyses to large, repeatable workflows. Model strategy, cost, and long-context tradeoffs (Priority: 4/5): They discuss using both open and closed models, the role of model upgrades like GPT-3/4 and Claude, and how long-context models change retrieval, ranking, and debugging while increasing compute costs. Mission, business model, and future hiring (Priority: 3/5): The transition to a PBC formalized an already product-driven mission. The team is growing, hiring engineers and product/design talent, and sees research tooling as a major future market with broad societal impact.

Key Arguments: AI products should be built around explicit, inspectable workflows rather than opaque end-to-end generation, because process supervision makes systems easier to evaluate, debug, and trust. The most valuable near-term application of AI in research is automating rote but high-leverage tasks like literature review, extraction, and synthesis, which are currently done by expensive human assistants. Research is a large economic market, not a niche academic one; improving how quickly people find truth and avoid wasted experiments can create major value across science, policy, medicine, and industry. Notebooks are a better interface than chat for research because they let users define, reuse, and scale a process over time, from a few papers to thousands. Long-context models are useful, but they do not eliminate the need for retrieval and structured pipelines; they shift the challenge toward ranking, evaluation, and traceability. Elicit’s moat is not a single model but a deep understanding of the research workflow and the ability to deploy models in transparent, controllable ways. The company wants AI to earn its right to expertise by matching and eventually surpassing human researchers while remaining auditable and aligned with user goals.

Data Points: Team size: 12 people - Elicit’s current headcount at the time of the interview. Revenue milestone: $1 million in revenue after four months - Mentioned as a recent announcement showing product traction. Co-founder evaluation doc length: ~50 pages - Andreas and Jawuan used a long Google doc of questions and back-and-forth to evaluate fit. Systematic review team size: 5 people - Typical human systematic reviews involve multiple collaborators over a long period. Systematic review duration: over a year - Human systematic reviews/meta-analyses are described as taking more than a year. Paper scale in systematic review workflow: 10,000 papers narrowed to hundreds or 50 - Describes the literature screening and narrowing process Elicit aims to automate. Open-source/open model cost reduction: cut costs to a tenth - Charlie’s work on constitutional AI and summarization reduced costs substantially. Model rollout speed: 4 days - Charlie implemented Anthropic’s constitutional AI paper in production within days. Model rollout speed: about a week - The constitutional AI summarization improvement was in the app shortly after implementation. Long-context retrieval stage: ~400 papers - They described retrieving roughly 400 relevant papers before reranking with larger models. High-accuracy project spend: $100,000 - Example of spend for a systematic-review-like project that would otherwise require human labor. Model mix by query count: similar - They said open and closed models are roughly similar in number of queries, though closed models cost more. Model mix by budget: closed models dominate - Closed models consume more of the budget because they are used when smarter models are needed. Current product focus: text-based workflows - They noted the product is currently centered on text, with longer-term expansion into reasoning and decision-making.

Pivotal Quotes: "We want to build for people." — Jawuan: Explaining the motivation for moving from research into product and why the team chose to build Elicit. "Our product values are systematic, transparent, and unbounded." — Andreas: Describing the company’s guiding principles for research workflows and AI deployment. "If you can understand the truth, that is very economically beneficial." — Andreas: Arguing that research and truth-seeking have direct commercial and societal value.

Implications: Elicit’s approach suggests the next wave of AI tools will win by making expert workflows explicit, auditable, and reusable. For research-heavy industries, the biggest gains may come from better process design, not just bigger models.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast