The Cognitive Revolution
The Cognitive Revolution

E20: The Great Implementation with Raza Habib, CEO of Humanloop

Raza Habib is on the front lines of LLM implementation as CEO of Humanloop, a Y Combinator backed startup that helps companies of all kinds, small and large bridge the gap from API access to successful LLM deployment. Nathan sat down with Raza to hear what he has learned in the process of helping so

Featured Speakers

Nathan Labenz and Erik Torenberg HostRaza Habib Guest

Topics Discussed

Episode Summary

Executive Summary: Raza Habib, CEO of Human Loop, explains how companies are actually deploying LLMs: mostly through careful prompt iteration, evaluation, feedback capture, and selective fine-tuning rather than fully autonomous agents or broad RLHF. He argues that interactive experimentation is essential, that productizing models is harder than accessing them, and that the near-term winners will be teams that build fast, reliable data and evaluation loops.

Main Topics: Human Loop’s role in LLM productization (Priority: 5/5): Human Loop helps builders bridge the gap between raw model APIs and reliable products by providing playgrounds, logging, feedback capture, evaluation, experimentation, and fine-tuning integrations. Interactive prototyping vs benchmarks (Priority: 5/5): Habib emphasizes that spending time in an interactive environment reveals model capabilities and failure modes better than benchmark scores alone, making REPL-style tools essential. Evaluation and feedback loops (Priority: 5/5): The conversation details how companies assess LLM quality using explicit votes, implicit usage signals, and textual corrections, then iterate via monitoring and A/B testing. Use cases and market adoption (Priority: 4/5): Real deployments span writing assistants, legal/document workflows, sales automation, medical paperwork, code tools, search, recruiting, and internal enterprise automation. Fine-tuning and RLHF maturity (Priority: 5/5): Habib argues supervised fine-tuning is often enough, while RLHF is powerful but harder, more finicky, and mostly used by more mature teams with stronger ML capability. Agents and tool use (Priority: 4/5): He is optimistic about agents but says current systems fail because errors compound across chained steps; reliable agents will likely need feedback and course-correction mechanisms. Future impact and societal implications (Priority: 4/5): Habib is excited about AI accelerating science, education, and software creation, while warning about misuse, bias, and concentration of power if systems gain too much autonomy without safeguards.

Key Arguments: Most organizations are not simply 'using ChatGPT'; they need a full workflow around prompt management, experimentation, monitoring, and feedback to make LLMs useful in production. Interactive play with models is crucial because subjective tasks cannot be fully understood from offline benchmarks alone. The most valuable LLM UXs are fault-tolerant and context-aware; chat works well because users can correct the model, while copilots succeed by staying close to the user’s context and minimizing latency. Evaluation is hard for subjective tasks, so companies should combine quantitative metrics with explicit feedback, implicit behavior, and edited outputs to understand performance. The biggest early gains often come from prompt engineering, but supervised fine-tuning becomes attractive when teams want better cost, latency, or task-specific performance. RLHF can improve alignment and performance, but it is operationally complex because it requires preference data, reward modeling, and RL training, each with failure points. Agentic systems are promising but currently unreliable because small inaccuracies compound over multiple steps; practical agents will need closed-loop correction rather than open-loop execution. Open-source model quality is improving, and privacy-sensitive companies will increasingly want on-prem or custom-tuned models, which should expand fine-tuning adoption. Incumbents can add AI features quickly, so the market will likely see both incumbent augmentation and new AI-native product categories. The main differentiator for successful teams is fast iteration backed by robust evaluation, not just raw model access.

Data Points: Preference ratio: 70-30 - Referenced from the GPT-4 technical report as GPT-4's preference over GPT-3.5, used to illustrate how evaluation can still look surprisingly mixed despite qualitative capability jumps. Fine-tuning example dataset size: 50 examples - A founder reportedly hand-wrote 50 examples and fine-tuned DaVinci 2 successfully, showing fine-tuning can work with relatively small datasets. Enterprise automation scale: Several hundred thousand customers - One customer was described as running large-scale customer service automation across a very large user base. Explicit feedback response rate: Low / small proportion - Thumbs up/down style feedback was described as useful but rarely used by most end users. Iteration speed: Minutes - Prompt changes and evaluation loops can sometimes be turned around in minutes, enabling rapid experimentation. Model access history retention: 30-day and delete policy - Referenced in contrast to Human Loop’s emphasis on retaining history and longitudinal usage data for analysis. Competition prizes: $40,000 - HackAPrompt 2023 was described as offering cash and AI-credit prizes totaling around this amount. Competition duration: 3 weeks - HackAPrompt 2023 was said to begin on May 5 and run for three weeks.

Pivotal Quotes: "RLHF is like sex in high school, right? Everyone's talking about it, but almost no one's actually doing it." — Raza Habib: A blunt characterization of how often RLHF is discussed versus actually implemented in real production systems. "Spending time in an interactive environment I think gives you a much better intuition for the capabilities of the models than just looking at benchmarks or numbers on test sets." — Raza Habib: Explaining why playground-style experimentation is critical for understanding LLM behavior. "My intuition is that yes, these things will get solved very quickly." — Raza Habib: His optimistic forecast that agent reliability and other current limitations will improve rapidly.

Implications: Teams that win with AI will build strong evaluation/feedback loops, not just prompts. Expect more fine-tuning, more private/custom deployments, and hybrid systems where incumbents add AI features while new AI-native products emerge.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution