Episode Summary
Executive Summary: Jeremy Howard traces his path from consulting and startups to Fast.ai, arguing that AI’s biggest opportunity is broad access, not elite control. He explains how transfer learning, ULMFiT, and later instruction/RLHF pipelines shaped modern LLMs, while warning that current fine-tuning practices cause forgetting and misuse. He emphasizes open tools, small models, and community-driven learning.
Main Topics: Jeremy Howard’s background and motivation (Priority: 4/5): Howard recounts his early career in philosophy, McKinsey, startups, Kaggle, and medical AI, framing his work as driven by usefulness and public benefit rather than prestige or scale. Fast.ai’s mission: democratizing deep learning (Priority: 5/5): He explains that Fast.ai was built to make deep learning accessible to non-PhDs by focusing on practical teaching, transfer learning, and tools that ordinary people can use. ULMFiT and the rise of transfer learning for NLP (Priority: 5/5): Howard describes the contrarian idea behind ULMFiT: pretrain a language model on a large corpus, fine-tune on a curated corpus, then fine-tune on a task. He argues this anticipated the modern LLM stack. Fine-tuning, memorization, and catastrophic forgetting (Priority: 5/5): He argues that current fine-tuning is often misused, leading to memorization, capability loss, and alignment tax. He proposes thinking in terms of continued pre-training with mixed data rather than narrow fine-tuning. Research philosophy: do more with less (Priority: 4/5): Howard says Fast.ai research prioritizes practical artifacts—software and courses—over papers, and focuses on reducing compute, data, complexity, and barriers to entry. Open-source communities and builder culture (Priority: 3/5): He discusses Discords, private channels, and how active contributors can get access by doing real work. He stresses that people who build and follow through stand out. Mojo, JAX, and the future of AI tooling (Priority: 3/5): Howard is enthusiastic about Chris Lattner’s Modular/Mojo effort as a way to make systems-level AI programming easier, and sees it as a path to more FlashAttention-like innovations.
Key Arguments: Deep learning should be accessible to ordinary people, not restricted to a few elite labs and PhDs; Fast.ai was created to lower those barriers. Transfer learning was the key insight that made deep learning practical with less compute and less data, especially in vision and NLP. ULMFiT’s three-step pipeline—pretrain, domain-adapt, task-adapt—was an early version of the modern LLM training recipe. The field initially dismissed large-scale language-model pretraining, but the approach worked and later became mainstream through GPT-style systems. Current fine-tuning practice is often conceptually wrong because it is treated as a separate phase rather than continued pre-training. Fine-tuning can cause catastrophic forgetting, as illustrated by Code Llama becoming strong at code but weaker at general knowledge. RAG is useful but inefficient and should complement, not replace, fine-tuning or continued pre-training. Small models remain underappreciated and can be highly capable when trained well, especially for enterprise and repetitive tasks. The most important unsolved questions are about training dynamics, data mixing, and how models actually learn internally. Open-source communities and educational resources can create a large distributed workforce of capable builders, reducing dependence on centralized AI labs.
Data Points: University degree: B.A. in philosophy - Howard said he graduated from the University of Melbourne with a philosophy degree. McKinsey work hours: 80 to 100 hours/week - He described working extremely long hours while studying and consulting. University attendance: No lectures in second and third year - He said he barely attended classes while working at McKinsey. Founding year of Optimal Decisions and FastMail: June 1999 - He said he founded both companies in the same month. Kaggle ranking: Top ranked participant in 2010 and 2011 - He noted he was the top Kaggle competitor in consecutive years. ULMFiT model size: 24–33 million parameters - The transcript references the AWD-LSTM used in ULMFiT as relatively small by today’s standards. Fast.ai course launch: 2016 - Howard said Fast.ai began around 2015–2016. Fast.ai course duration: 7 weeks - The hosts referenced the course as taking people from basics to Stable Diffusion in about seven weeks. DawnBench turnaround: 10 days - He said Fast.ai assembled an emergency team and won the competition in the final 10 days. Kaggle competition runtime limit: 9 hours - He described a competition where models had to run within nine hours on Kaggle hardware. Kaggle hardware limit: 14 GB RAM, 2 CPUs, small old GPU - He used these constraints to illustrate low-resource model development. Code Llama fine-tuning data: 500 billion tokens - He cited Meta’s large code-focused continued pre-training run. Non-code data in Code Llama: 0.3% of epochs - He argued this tiny share explains why the model forgot general capabilities. Model size example: BTLM 3B - He cited it as a strong small model, roughly comparable to a 7B-quality model. Model size example: 1–2B range - He said this range is underpopulated despite being promising for complex tasks. Training example: 15 minutes on a single GPU - He said a pre-trained base model was fine-tuned to generate SQL very quickly. Chess prompting example: ELO 3,400 - He mentioned a prompting strategy that reportedly made GPT-4 play chess at near top-engine level.
Pivotal Quotes: "there's no such thing as fine tuning. There's only continued pre-training." — Jeremy Howard: He argued that modern fine-tuning should be understood as a continuation of pretraining, not a separate conceptual phase. "I want to help most people, not the small subset of the most well-off people." — Jeremy Howard: He explained why he avoids approaches that require massive compute or elite infrastructure. "if it looks like a duck and acts like a duck, it's a duck" — Jeremy Howard: He used this to defend the idea that a sufficiently capable text system can be treated as understanding, at least operationally.
Implications: The conversation argues for open, practical AI: better data mixing, more small-model work, stronger tooling, and broader access. For listeners, the message is to build, share, and learn—because AI capability is still underexplored and too concentrated.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast