Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

[Ride Home] Simon Willison: Things we learned about LLMs in 2024

Due to overwhelming demand (>15x applications:slots), we are closing CFPs for AI Engineer Summit NYC today. Last call! Thanks, we’ll be reaching out to all shortly! The world’s top AI blogger and friend of every pod, Simon Willison, dropped a monster 2024 recap: Things we learned about LLMs in 20

Featured Speakers

Latent.Space HostSimon Willison Guest

Topics Discussed

Episode Summary

Executive Summary: Simon Willison and Swix argue that AI in early 2025 is defined less by a new intelligence leap than by dramatic gains in speed, cost, context length, and multimodality. They highlight open-weight breakthroughs, especially DeepSeek and Qwen, the rise of practical agents for research/coding, and the need for better interfaces, credibility, and privacy-aware regulation as AI becomes embedded in everyday workflows.

Main Topics: State of AI in early 2025 (Priority: 5/5): The discussion opens with a broad assessment: models are much cheaper, faster, and more capable across text, image, audio, and video, but there has not yet been a GPT-5-style intelligence jump. Open-weight model breakthroughs and efficiency (Priority: 5/5): Willison emphasizes that the GPT-4 performance barrier has been smashed by many organizations, including open-weight models that run locally and cost far less to train and serve. Agents: promise, definitions, and reliability limits (Priority: 5/5): The speakers debate what 'agents' actually mean, where they work today (research assistants, coding loops), and why autonomous action remains risky because of prompt injection and gullibility. Multimodal AI and video/audio progress (Priority: 4/5): They discuss rapid progress in vision, audio, and video generation/understanding, including streaming video into models, NotebookLM-style audio products, and the uneven state of tools like Sora and Veo. Local models and on-device AI (Priority: 4/5): Willison describes renewed enthusiasm for local LLMs as small models become useful again, with tools like Ollama, LM Studio, MLC Chat, and Apple/Chrome on-device efforts. Credibility, criticism, and AI-generated slop (Priority: 4/5): A major theme is that AI output needs human review and editorial judgment; credibility becomes more valuable as content generation gets cheaper and more abundant. Wearables and the next interface layer (Priority: 3/5): The conversation closes by predicting AI wearables, smart glasses, earbuds, and memory/meeting-assistant devices as a likely next consumer wave, contingent on privacy norms.

Key Arguments: AI’s biggest 2024-2025 shift is not a dramatic intelligence leap but a massive drop in cost and latency, plus broader multimodal capability. The old assumption that only giant labs with enormous GPU clusters can train frontier models has been weakened by DeepSeek’s low-cost training claim and the rise of efficient open-weight models. The GPT-4 barrier has been broken by many organizations; the field is now more competitive and less centralized than a year ago. Agents are useful today mainly in constrained settings: research assistants, coding loops, and workflow automation with human oversight. Autonomous agents that spend money or act on the open web remain unreliable because LLMs are gullible and vulnerable to prompt injection. AI-generated content is only valuable when reviewed and curated by a human; unrequested, unreviewed output is 'slop.' The best AI interfaces will likely move beyond chat into generated GUIs, knobs-and-dials controls, and collaborative document/workflow tools. Local models are becoming practical again, making offline or privacy-preserving AI workflows increasingly viable. Credibility and trust will matter more, not less, as AI content floods the internet and workplaces. Privacy regulation should focus on use cases and data handling rather than trying to freeze model development with outdated compute caps.

Data Points: Organizations beating GPT-4 barrier: 18 - Willison says 18 organizations other than OpenAI have released models that clearly beat the older GPT-4 benchmark. OpenAI price drop vs GPT-3: 100x cheaper - He says today’s OpenAI models are about 100 times cheaper per token than GPT-3-era pricing. Google Gemini 1.5 Flash price: $0.075 per million tokens - Used as an example of how token pricing has fallen to cents-per-million levels. Google Gemini 1.5 Flash 8B vs GPT-3.5 Turbo: 27x cheaper - Willison cites this as evidence of aggressive price compression in mainstream models. DeepSeek V3 training cost: $5.5 million - He describes DeepSeek’s open-weights model as trained for roughly this amount, far below common assumptions. DeepSeek company size: ~150 employees - Used to illustrate how small the lab is relative to major US AI companies. Photo captioning cost: $1.68 for 68,000 images - Willison estimates Gemini 1.5 Flash 8B could caption a large photo library for almost nothing. Per-image captioning cost: 1/400th of a cent per image - Derived from the 68,000-image captioning example. Local model file size: 14 GB - He notes Microsoft 5.4 runs on a MacBook Pro as a 14 GB download. Local model memory example: 64 GB RAM - Willison says running large local models like Llama 370B consumes most of his machine’s memory. AI engineer summit dates: February 20-21, 2025 - Swix announces the New York conference dates. Conference audience split: 2 days - Day 1 for leadership/management; Day 2 for engineers and ICs.

Pivotal Quotes: "Everything's got really good and fast and cheap." — Simon Willison: His opening summary of AI’s state in early 2025. "The default LLM chat UI is like taking brand new computer users, dropping them into a Linux terminal and expecting them to figure it all out." — Simon Willison: Used to argue that current chat interfaces are not the final UX for AI. "I love the term slop, where I've been pushing the definition of slop as AI generated content that is both unrequested and unreviewed." — Simon Willison: His framework for distinguishing useful AI-assisted work from low-value output.

Implications: AI is becoming cheaper, more local, and more embedded in workflows, but trust, privacy, and interface design will determine adoption. The winners will be tools that combine model power with human review, clear use cases, and credible outputs.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast