Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Building the Foundation Model Ops Platform — with Raza Habib of Humanloop

Want to help define the AI Engineer stack? >500 folks have weighed in on the top tools, communities and builders for the first State of AI Engineering survey! Please fill it out (and help us reach 1000!) The AI Engineer Summit schedule is now live! We are running two Summits and judging two Hacka

Featured Speakers

Latent.Space HostRaza Habib Guest

Topics Discussed

Episode Summary

Executive Summary: Raza Habib, CEO of HumanLoop, explains how the company evolved from NLP data tooling into an LLM ops platform focused on prompt management, evaluation, and feedback loops. The conversation covers why prompt engineering still matters, how enterprise AI workflows are changing, the rise of AI engineering, and why San Francisco is becoming the center of gravity for AI productization.

Main Topics: HumanLoop’s origin and pivot to LLM ops (Priority: 5/5): Habib traces HumanLoop from early NLP/active-learning tooling to a prompt engineering and evaluation platform after GPT-3 and InstructGPT made LLM applications practical. Prompt evaluation and human feedback (Priority: 5/5): He breaks evaluation into development, production monitoring, and regression testing, and describes three feedback types: votes, actions, and corrections. Prompt engineering vs AI engineering (Priority: 4/5): Habib argues prompt engineering is not dead, but the term is imperfect; prompts are important artifacts, while the broader skill set is AI engineering. Enterprise adoption, pricing, and product-market fit (Priority: 4/5): The discussion covers enterprise procurement, SOC 2, VPC deployments, new free-tier pricing, and why HumanLoop’s PMF is strongest in larger teams building LLM apps. Competitive landscape and ecosystem positioning (Priority: 4/5): Habib compares HumanLoop with LangChain, MLOps tools, and open standards, arguing HumanLoop is more opinionated, workflow-oriented, and model-agnostic. AI research outlook and model behavior (Priority: 4/5): He discusses GPT-4 changing over time, the importance of continual learning and multimodality, and underrated research like LLM cascades and in-context reinforcement learning. Europe vs San Francisco AI scenes (Priority: 3/5): Habib praises Europe’s research depth but says San Francisco has superior density for AI productization, founders, investors, and community.

Key Arguments: HumanLoop started with a belief that NLP would become dramatically more capable through transfer learning and large models, but pivoted when GPT-3 and especially InstructGPT made LLM apps commercially viable. Evaluation for LLM apps is fundamentally different from traditional ML because the outputs are subjective, stochastic, and often only meaningfully judged by user behavior in production. Three evaluation modes matter: development-time iteration, production monitoring, and regression testing after changes. HumanLoop’s three feedback types—votes, actions, and corrections—capture both explicit and implicit signals, and corrections are especially useful for later fine-tuning. Prompt engineering is not dead because prompts are effectively source code for LLM applications, but the real job is broader AI engineering that includes code, orchestration, and evaluation. The best customers are larger teams and enterprises because they need collaboration, governance, and confidence before launch; small startups often just ship quickly. HumanLoop is intentionally model-agnostic and sits between base model providers and orchestration frameworks, helping teams measure and improve applications regardless of model choice. As LLMs evolve, regression testing becomes more important because model behavior can change over time even when the API name stays the same. San Francisco has become the strongest hub for AI productization because of talent density, investor experience, and the concentration of builders iterating on real products. Continual learning remains a major unsolved problem, and multimodality is an obvious next step for foundation models.

Data Points: PhD start year: 2017 - Habib began his doctorate at UCL in probabilistic deep learning. PhD completion year: 2022 - He finished his doctorate about a year before the interview. HumanLoop founding year: 2020 - The company was started while he was still doing his PhD. Design partner experiment: 10 paying customers in 2 days - Early HumanLoop sales experiment for prompt engineering/design partnerships. Early market size estimate: 3-400 companies - His initial YC-style estimate for companies that might need the product. Team size before scaling: 4 people for almost 2 years - HumanLoop stayed very small until demand forced hiring. Closed beta eval customers: about 10 companies - Current evaluation product is in closed beta. SOC 2 status: Part 1 complete; Part 2 in audit - Enterprise readiness and procurement support. Free tier threshold: more than 3 people / certain data volume - New pricing will start charging only after usage and team size grow. AI meetup density in SF: 10 meetups in one night - Used to illustrate San Francisco’s AI ecosystem density.

Pivotal Quotes: "Prompt engineering is dead? ... I don't think it's dead. I think it's alive and well and becoming increasingly important." — Raza Habib: On whether prompt engineering remains a meaningful discipline. "The worst they're ever going to be." — Raza Habib: On current frontier models being a floor for future capability, not a ceiling. "Before PMF, nothing other than PMF matters." — Raza Habib: Founder advice on avoiding distractions and staying focused on product-market fit.

Implications: For builders, the message is to treat prompts, evals, and feedback as core product infrastructure. For the industry, LLM ops is maturing into AI engineering, with enterprise workflows, model drift, and multimodal support becoming central.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast