Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The Ultimate Guide to Prompting

Noah Hein from Latent Space University is finally launching with a free lightning course this Sunday for those new to AI Engineering. Tell a friend! Did you know there are >1,600 papers on arXiv just about prompting? Between shots, trees, chains, self-criticism, planning strategies, and all sorts

Featured Speakers

Latent.Space HostSander Schulhoff Guest

Topics Discussed

Episode Summary

Executive Summary: This episode is a deep dive into prompt engineering with Sander Schulhoff, creator of Learn Prompting and co-author of The Prompt Report. The conversation covers his path into AI, a systematic taxonomy of prompting methods, practical guidance on few-shot and chain-of-thought prompting, the limits of role prompting, automated prompt optimization with DSPy, and the security frontier of prompt injection/jailbreaking. It closes with multimodal prompting, structured outputs, and plans for Hack Prompt 2.0.

Main Topics: Sander Schulhoff’s path into AI and prompting (Priority: 5/5): Schulhoff traces his entry into AI from Java and deep learning in high school to RL and NLP research at Maryland, diplomacy work, Minecraft RL, and eventually prompting via GPT-3 translation tasks. Learn Prompting began as a class project and became a major resource. The Prompt Report and systematic literature review (Priority: 5/5): He explains how The Prompt Report surveyed over 1,600 papers using a PRISMA-style systematic review, with a 30-person team and AI-assisted screening. The goal was to compress a sprawling literature into a usable taxonomy and reference. Taxonomy of prompting techniques (Priority: 5/5): The discussion centers on organizing prompting by problem-solving strategy rather than application: zero-shot, few-shot, thought generation, decomposition, ensembling, self-criticism, and related methods. Schulhoff emphasizes that many techniques overlap and require judgment calls. Practical prompting advice: few-shot, formatting, and chain-of-thought (Priority: 5/5): The episode covers exemplar ordering, format consistency, distribution, and similarity in few-shot prompts, plus the continued usefulness of chain-of-thought and decomposition. Schulhoff argues that prompt design still materially affects accuracy on modern models. Prompt injection, jailbreaking, and Hack-a-Prompt (Priority: 5/5): Schulhoff distinguishes prompt injection from jailbreaking and describes Hack-a-Prompt, which collected 600,000 malicious prompts and won best paper at EMNLP. He highlights context overflow as a surprising attack discovered in the competition. Automated prompt engineering, structured outputs, and evaluation (Priority: 4/5): The conversation explores DSPy and automatic prompt optimization, structured output prompting, and the dangers of naive scoring/evaluation methods. Schulhoff stresses that evaluation prompts and numeric scales can be unstable or misleading. Multimodal prompting and the future of prompt engineering (Priority: 4/5): The episode ends with observations on audio, music, and especially video prompting, which Schulhoff says is still hard to control. He also previews Hack Prompt 2.0, aimed at real-world harms and agentic security issues.

Key Arguments: Prompt engineering remains a core skill because models are general-purpose and do not reliably infer the exact reasoning or output format users need. Role prompting may help style/text generation, but Schulhoff argues it does not meaningfully improve accuracy on modern models for tasks like MMLU. Few-shot prompting is highly sensitive to exemplar order, format, and distribution; these details can swing performance dramatically. Chain-of-thought is still useful in practice because models sometimes fail to emit reasoning unless explicitly prompted, especially at scale. Decomposition and thought generation are related but distinct: one breaks a task into subproblems, the other asks for intermediate reasoning steps. Self-consistency and ensembling can improve outputs, but the gains may be smaller on newer models and can be expensive. DSPy and other automated prompt optimization tools can outperform manual prompt engineering quickly, but they generally require labeled ground truth. Prompt injection and jailbreaking are often conflated; clear terminology matters because the attack surface differs when developer instructions are present. Competitions uncover adversarial behaviors that normal paid testing may miss, as shown by the context overflow attack in Hack-a-Prompt. Structured output and numeric scoring prompts need careful design because models are biased toward certain tokens and scales can be unstable.

Data Points: Papers surveyed in The Prompt Report: 1,600+ - Schulhoff says the report systematically reviewed over 1,600 published papers on prompting techniques. Research team size for The Prompt Report: 30 people - He led a 30-person team over about nine months, including researchers from OpenAI, Google, Microsoft, Princeton, Stanford, and Maryland. Length of the compiled summary doc: ~80 pages - The team condensed the literature into an approximately 80-page summary document before publishing. Views tracked for The Prompt Report: ~1.5 million - Schulhoff says he has tracked about one and a half million views across social platforms. Hack-a-Prompt malicious prompts collected: 600,000 - The competition generated a large dataset of malicious prompts for prompt injection research. Hack-a-Prompt recognition: Best paper at EMNLP - He says the paper was one of three selected as best papers at EMNLP. EMNLP submission volume: 20,000 submitted / 5,000 accepted - Schulhoff cites the conference scale to emphasize the significance of the award. Hack Prompt 2.0 prize target: $500,000 - He says they are fundraising to give away half a million dollars in prizes. Prompt Report timeline: ~9 months - The systematic survey and write-up took about nine months. Learn Prompting launch: October 2022 - He notes the site launched before ChatGPT. Hack-a-Prompt timing: May 2023 - He says the competition ran in May 2023. GPT-4 overnight bill anecdote: $150 - Used to illustrate that LLM costs can still become significant despite falling prices.

Pivotal Quotes: "role prompting is useful for text generation tasks... For accuracy-based tasks like MMLU... I really don't think that works" — Sander Schulhoff: He explains why he is skeptical that role prompting improves factual or reasoning performance on modern models. "we collected 600,000 malicious prompts, put together a paper on it, open sourced everything" — Sander Schulhoff: He summarizes the scale and impact of Hack-a-Prompt. "I spent 20 hours prompt engineering for a task and DSPy beat me in 10 minutes" — Sander Schulhoff: He describes why he now recommends automated prompt optimization tools.

Implications: Prompt engineering is still relevant, but the field is shifting from ad hoc tricks to systematic methods, automation, and security-aware design. For builders, the big wins now come from evaluation, structure, and tooling—not just clever wording.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast