Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

AI Magic: Shipping 1000s of successful products with no managers and a team of 12 — Jeremy Howard of Answer.ai

Disclaimer: We recorded this episode ~1.5 months ago, timing for the FastHTML release. It then got bottlenecked by Llama3.1, Winds of AI Winter, and SAM2 episodes, so we’re a little late. Since then FastHTML was released, swyx is building an app in it for AINews, and Anthropic has also released thei

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The episode centers on Jeremy Howard’s return to discuss Answer AI’s rapid, small-team open-source R&D model, his critique of conventional ML training and company structures, and new technical bets on continued pretraining, encoder-decoder/BERT-style models, quantized adapters, and faster web app development via FastHTML. It also previews his “dialogue engineering” workflow and broader views on governance, hiring, and AI infrastructure.

Main Topics: Answer AI’s operating model and culture (Priority: 5/5): Jeremy explains how Answer AI is structured as a tiny, managerless team focused on shipping useful open-source work quickly, with broad autonomy, shared mission, and unusually strong contributors from nontraditional backgrounds. Continued pretraining and the fine-tuning continuum (Priority: 5/5): The conversation revisits Jeremy’s earlier “end of fine-tuning” framing and clarifies that pretraining, instruction tuning, and task tuning should be treated as a continuum rather than separate phases. Governance, public benefit, and OpenAI’s collapse (Priority: 5/5): Jeremy argues that OpenAI’s governance structure was inherently unstable because incentives, ownership, and safety control were misaligned, and uses this to motivate Answer AI’s public-benefit-corporation approach. Technical bets: BERT, encoder-decoder, and quantized adapters (Priority: 4/5): The discussion highlights renewed interest in encoder-decoder and encoder-only architectures, especially for retrieval and classification, plus a push toward quantized base models with adapters instead of merged model distribution. FSDP QLoRA/DORA engineering and open-source performance work (Priority: 4/5): Jeremy recounts the difficult, detail-heavy systems work required to make large-model fine-tuning practical on limited hardware, including debugging undocumented internals and benchmarking regressions. FastHTML and simplifying web app development (Priority: 4/5): Jeremy introduces FastHTML as a pure-Python, web-foundation-based way to build full web apps in one file, aiming to remove the complexity of modern JavaScript-heavy stacks for AI builders. Dialogue engineering and AI-assisted productivity (Priority: 4/5): He previews an internal workflow called dialogue engineering, implemented in tools like Claudette, Cozette, and AI Magic, to make model-assisted coding and writing more interactive and effective.

Key Arguments: Training should be viewed as a continuum: pretraining, instruction tuning, and task tuning are not truly separate stages. Starting from random weights is usually a bad idea unless there is a strong reason to do so; data-driven initialization and continued pretraining are more sensible. OpenAI’s governance model was structurally unstable because nonprofit control conflicted with profit-driven incentives and equity-based compensation. Answer AI’s small, managerless structure works because highly capable people are given broad autonomy and a shared mission. Nontraditional backgrounds often produce unusually creative, tenacious, and high-performing researchers and engineers. Encoder-decoder and encoder-only models remain underexplored relative to decoder-only LLMs and may be better for many tasks. Quantized base models plus adapters should replace merged-model distribution in most cases because they are faster to download and infer with. A lot of progress in fine-tuning and inference is blocked by messy, under-benchmarked systems code rather than model theory. Web app development for data scientists and AI builders is unnecessarily complex; FastHTML aims to restore a simple, Python-first path. AI should be used to augment the craft of software development through dialogue engineering rather than just prompt engineering.

Data Points: Answer AI team size: maximum of 12 team members - Jeremy describes the company as intentionally tiny and managerless. Funding timeline: less than a year - He says the team has shipped many projects in under a year since funding. Projects shipped: 10+ named projects - Examples include FsDPQ-LORA, FsDPQ-DORA, Cold Compress, Colbert Small, Jarklebutt, GPU CPP, Claudette, Fastlight, and FastHTML. OpenAI governance warning: 2 days before Altman was fired - Jeremy says he publicly predicted OpenAI’s governance structure could not continue shortly before the board crisis. Email timeline: November 13 and November 14 - The hosts reference emails about talking and “working together to free AI” as the seed of Answer AI. Interview pipeline hire rate: everybody bar one - Jeremy says nearly everyone who entered recruiting was hired. Meeting structure: no required meetings; one meeting per major time zone pair - He describes the company’s lightweight coordination model. Historical web app effort: 10 years - Jeremy references spending a decade on a prior web system similar in spirit to FastHTML. Conference episode length: 6 hours - He mentions the prior NeurIPS street episode was very long.

Pivotal Quotes: "“You know, if you're training from random weights, you better have a really good reason.”" — Jeremy Howard: On why continued pretraining and data-driven initialization matter more than starting from scratch. "“The alignment problem as it relates to companies has not been solved.”" — Jeremy Howard: On why corporate incentives can conflict with founders’ or society’s goals. "“Everybody understands kind of roughly why we're here.”" — Jeremy Howard: On Answer AI’s shared mission and low-hierarchy culture.

Implications: The episode suggests the next wave of AI progress may come from better training continuums, smaller high-trust teams, and simpler infrastructure—not just bigger models. It also points to growing demand for open, practical tooling that helps builders ship faster.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast