Episode Summary
Executive Summary: Ed Donner explains AI/LLM engineering as a hybrid discipline spanning data science, software engineering, and ML engineering. He outlines how practitioners choose models, compare closed vs open source, use RAG, fine-tuning, and agentic systems, and rely on benchmarks/leaderboards to select and deploy models in production. He also shares practical deployment options and learning resources.
Main Topics: What AI/LLM engineering is (Priority: 5/5): Ed defines the emerging role as a blend of data science, software engineering, and ML engineering focused on selecting, adapting, and shipping LLM-based systems. Choosing models: closed source vs open source (Priority: 5/5): Model selection starts from business requirements, data constraints, evaluation metrics, and non-functional needs such as cost, latency, privacy, and time-to-market. Core techniques: RAG, fine-tuning, and agentic AI (Priority: 5/5): He contrasts training-time optimization with inference-time methods, explaining when to use retrieval, model adaptation, tool use, reasoning frameworks, and proactive agents. Benchmarks and leaderboards (Priority: 4/5): Ed reviews practical resources for model selection, including Hugging Face, Vellum, LMSYS/LMArena, and enterprise leaderboards that measure reasoning, coding, robustness, and business-fit. Deployment and productionization (Priority: 4/5): He covers serverless and platform-based deployment options such as Modal, Hugging Face Endpoints, Lightning Studios, Docker/Kubernetes, and agent platforms like LangGraph and CrewAI. Learning and career development (Priority: 3/5): Ed recommends upskilling through his Udemy course and emphasizes that mathematical foundations remain valuable for building robust, differentiated AI systems. Custom evaluation and synthetic data (Priority: 3/5): He describes building task-specific test sets, including using synthetic data seeded from real platform examples, plus his own model-comparison game, Outsmart.
Key Arguments: AI engineering is a real, rapidly growing job category, not just a rebranding of data science or ML engineering. There is no single best LLM; the right model depends on the business problem, data type, evaluation metric, and operational constraints. For many teams, the best starting point is a closed-source frontier model for prototyping because early costs are usually low and iteration speed matters. Open source becomes compelling when there is proprietary data to fine-tune on, privacy constraints, or high-volume inference cost pressure. RAG is often the first choice when accuracy depends on grounding responses in external knowledge, while fine-tuning is better for specialized behavior and skills. Agentic AI is most useful for multi-step problems, tool use, and systems that must act proactively beyond a single user prompt. Benchmarks help narrow options, but real performance must be validated empirically on task-specific data and outcomes. Deployment is increasingly part of the AI engineer's job, with serverless and managed platforms reducing friction from prototype to production.
Data Points: US job openings for LLM engineer: about 4,000 - Ed cites LinkedIn search results showing demand for AI/LLM engineering roles US job openings for data scientist: about 4,800 - Used for comparison with LLM engineer demand Training-course enrollments: 14,000 people - Users who had taken Ed's Udemy LLM Engineering course after 6-7 weeks GPQA benchmark size: 448 questions - A difficult science benchmark used to test expert-level reasoning GPQA human expert performance: about 65% - Score achieved by PhD-level humans on GPQA GPQA non-expert performance: about 35% - Score by non-PhD participants allowed to use Google for 30 minutes Claude 3.5 Sonnet GPQA score: about 65% - Ed describes it as reaching expert-human level Claude 3.5 Sonnet on BBH: 93% - Used to show strong performance on Big Bench Hard Claude 3.5 Sonnet on GPQA earlier in the year: about 59% - Shown as a rapid improvement before reaching 65% Open-source model GPQA leader: Qwen 2.5 32B at about 22% - Highlighted as top open-source model on the leaderboard at the time Outsmart starting coins: 12 coins - Starting endowment in Ed's custom multi-model game benchmark Claude 3.5 Sonnet launch timing: about a year before ChatGPT? no; 'earlier this year' - Discussed relative to rapid benchmark progress and model releases
Pivotal Quotes: "There isn't like one best LLM. And there's the right LLM to use for the task at hand." — Ed Donner: Explaining how model selection depends on the use case rather than a universal ranking "Agentic AI is a technique... where you want to be able to make use of tools." — Ed Donner: Defining when agentic systems are most appropriate, especially for tool use and multi-step workflows "That's empirical." — Ed Donner: Describing how model choice and technique selection ultimately require experimentation on real data
Implications: AI engineers need a practical, experiment-driven skill set: benchmark literacy, evaluation design, and deployment fluency. Teams that combine strong foundations with task-specific testing and the right mix of RAG, fine-tuning, and agents will build more capable, cost-effective systems.
About Super Data Science: ML & AI Podcast with Jon Krohn
View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn