Super Data Science: ML & AI Podcast with Jon Krohn
Super Data Science: ML & AI Podcast with Jon Krohn

847: AI Engineering 101, with Ed Donner

Ed Donner co-founded AI-driven recruitment platform, Nebula.io, with The SuperDataScience Podcast’s host, Jon Krohn. Ed and Jon reminisce about how they launched their company, the growing opportunities for data scientists, how to choose an LLM, and today’s top technical terms in AI. Interested in s

Featured Speakers

Jon Krohn HostEd Donner Guest

Topics Discussed

Episode Summary

Executive Summary: Ed Donner explains AI/LLM engineering as a hybrid discipline spanning data science, software engineering, and ML engineering. He outlines how practitioners choose models, compare closed vs open source, use RAG, fine-tuning, and agentic systems, and rely on benchmarks/leaderboards to select and deploy models in production. He also shares practical deployment options and learning resources.

Main Topics: What AI/LLM engineering is (Priority: 5/5): Ed defines the emerging role as a blend of data science, software engineering, and ML engineering focused on selecting, adapting, and shipping LLM-based systems. Choosing models: closed source vs open source (Priority: 5/5): Model selection starts from business requirements, data constraints, evaluation metrics, and non-functional needs such as cost, latency, privacy, and time-to-market. Core techniques: RAG, fine-tuning, and agentic AI (Priority: 5/5): He contrasts training-time optimization with inference-time methods, explaining when to use retrieval, model adaptation, tool use, reasoning frameworks, and proactive agents. Benchmarks and leaderboards (Priority: 4/5): Ed reviews practical resources for model selection, including Hugging Face, Vellum, LMSYS/LMArena, and enterprise leaderboards that measure reasoning, coding, robustness, and business-fit. Deployment and productionization (Priority: 4/5): He covers serverless and platform-based deployment options such as Modal, Hugging Face Endpoints, Lightning Studios, Docker/Kubernetes, and agent platforms like LangGraph and CrewAI. Learning and career development (Priority: 3/5): Ed recommends upskilling through his Udemy course and emphasizes that mathematical foundations remain valuable for building robust, differentiated AI systems. Custom evaluation and synthetic data (Priority: 3/5): He describes building task-specific test sets, including using synthetic data seeded from real platform examples, plus his own model-comparison game, Outsmart.

Key Arguments: AI engineering is a real, rapidly growing job category, not just a rebranding of data science or ML engineering. There is no single best LLM; the right model depends on the business problem, data type, evaluation metric, and operational constraints. For many teams, the best starting point is a closed-source frontier model for prototyping because early costs are usually low and iteration speed matters. Open source becomes compelling when there is proprietary data to fine-tune on, privacy constraints, or high-volume inference cost pressure. RAG is often the first choice when accuracy depends on grounding responses in external knowledge, while fine-tuning is better for specialized behavior and skills. Agentic AI is most useful for multi-step problems, tool use, and systems that must act proactively beyond a single user prompt. Benchmarks help narrow options, but real performance must be validated empirically on task-specific data and outcomes. Deployment is increasingly part of the AI engineer's job, with serverless and managed platforms reducing friction from prototype to production.

Data Points: US job openings for LLM engineer: about 4,000 - Ed cites LinkedIn search results showing demand for AI/LLM engineering roles US job openings for data scientist: about 4,800 - Used for comparison with LLM engineer demand Training-course enrollments: 14,000 people - Users who had taken Ed's Udemy LLM Engineering course after 6-7 weeks GPQA benchmark size: 448 questions - A difficult science benchmark used to test expert-level reasoning GPQA human expert performance: about 65% - Score achieved by PhD-level humans on GPQA GPQA non-expert performance: about 35% - Score by non-PhD participants allowed to use Google for 30 minutes Claude 3.5 Sonnet GPQA score: about 65% - Ed describes it as reaching expert-human level Claude 3.5 Sonnet on BBH: 93% - Used to show strong performance on Big Bench Hard Claude 3.5 Sonnet on GPQA earlier in the year: about 59% - Shown as a rapid improvement before reaching 65% Open-source model GPQA leader: Qwen 2.5 32B at about 22% - Highlighted as top open-source model on the leaderboard at the time Outsmart starting coins: 12 coins - Starting endowment in Ed's custom multi-model game benchmark Claude 3.5 Sonnet launch timing: about a year before ChatGPT? no; 'earlier this year' - Discussed relative to rapid benchmark progress and model releases

Pivotal Quotes: "There isn't like one best LLM. And there's the right LLM to use for the task at hand." — Ed Donner: Explaining how model selection depends on the use case rather than a universal ranking "Agentic AI is a technique... where you want to be able to make use of tools." — Ed Donner: Defining when agentic systems are most appropriate, especially for tool use and multi-step workflows "That's empirical." — Ed Donner: Describing how model choice and technique selection ultimately require experimentation on real data

Implications: AI engineers need a practical, experiment-driven skill set: benchmark literacy, evaluation design, and deployment fluency. Teams that combine strong foundations with task-specific testing and the right mix of RAG, fine-tuning, and agents will build more capable, cost-effective systems.

🔓 Sign Up for Unlimited Episode Search

About Super Data Science: ML & AI Podcast with Jon Krohn

View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn