The Cognitive Revolution
The Cognitive Revolution

Delving into The Prompt Report, with Sander Schulhoff of LearnPrompting.org

Nathan welcomes back Sander Schulhoff, creator of LearnPrompting.org, to discuss the recently released Prompt Report. In this episode of The Cognitive Revolution, we explore the current state of prompting techniques for large language models, covering best practices, challenges, and emerging trends

Featured Speakers

Nathan Labenz and Erik Torenberg HostSander Schulhoff Guest

Topics Discussed

Episode Summary

Executive Summary: Sander Schulhoff discusses the Prompt Report, a 78-page survey that organizes the fast-growing prompting literature into a practical taxonomy: few-shot prompting, chain-of-thought, decomposition, ensembling, multilingual and multimodal prompting, agent prompting, and evaluation. He emphasizes that exemplar quality, ordering, format, and similarity matter; defense against prompt attacks remains difficult; and automated prompt optimization like DSPy can outperform human-crafted prompts. He also shares lessons on managing large research teams with testing and trust.

Main Topics: Prompt Report overview and research process (Priority: 5/5): Schulhoff explains how the team assembled a large survey paper by combining literature search, human review, model-assisted filtering, and topic modeling to build a comprehensive prompting taxonomy. Few-shot prompting best practices (Priority: 5/5): A major focus of the paper is practical guidance on few-shot prompts: more exemplars often help, ordering can drastically affect results, balanced labels matter, formatting should match familiar training patterns, and retrieved similar exemplars usually outperform diverse ones. Prompt injection, jailbreaking, and defense limits (Priority: 5/5): The conversation revisits security terminology and attacks. Schulhoff argues that defenses are largely losing ground and that robust mitigation will likely require changes from model developers, not just external guardrails. Advanced prompting categories and taxonomy (Priority: 4/5): The report distinguishes chain-of-thought variants, decomposition, zero-shot style/role prompting, self-criticism, and ensembling, while acknowledging that many techniques overlap and the taxonomy is partly arbitrary. Multilingual, multimodal, and agent prompting (Priority: 4/5): Schulhoff highlights techniques such as translating tasks into English first, multilingual ensembling, chain-of-image prompting, and simple tool-use patterns like ReAct, while noting that agent systems remain brittle and hard to operationalize. Automated prompt optimization and DSPy (Priority: 5/5): A key finding is that automated systems can beat human prompt engineering. Schulhoff reports that DSPy, given the same exemplars, produced a prompt that significantly outperformed his hand-written version on a test set. Managing large technical research projects (Priority: 3/5): He reflects on leading a large team, emphasizing CI/testing, clean pipelines, repeated review passes, and his conclusion that researchers should 'trust no one, not even yourself.'

Key Arguments: Prompt engineering is not a single technique but a family of overlapping methods that need a taxonomy to be useful. Few-shot prompting is highly sensitive to details such as exemplar count, ordering, label balance, formatting, and similarity to the test case. In-context learning and few-shot prompting are often conflated, but they are not identical concepts. Prompt defense is likely to remain difficult because model behavior is hard to fully control; fixes probably need to come from model creators and training/architecture changes. Multimodal and agent prompting are promising but still brittle and underdeveloped compared with text prompting. Automated prompt optimization can outperform expert human prompting when ground-truth examples are available. For open-ended generation tasks, the best path is often iterative prompting, self-critique, or moving toward fine-tuning rather than expecting one-shot prompt perfection. Reliable research at scale depends on validation checks, repeated review, and engineering discipline, not just expertise.

Data Points: Prompt Report length: 78 pages - The survey paper reviewed in the episode Paper-writing/review timeline: 9 months - Schulhoff says the systematic literature review took about twice as long as he expected Initial project timeline estimate: 3 to 4 months - His original timeline for the survey project Literature search keywords: 44 keywords - Used across arXiv, Semantic Scholar, and ACL for paper collection Human-reviewed papers: about 1,000 papers - Manually reviewed to decide whether methods belonged in the prompting survey Research team size: large multi-source team with outside authors and advisors - Includes lab members, Startup Shell students, OpenAI’s Shyamal, advisors, and domain experts such as suicide researchers Passes over manuscript: at least 5 passes - Schulhoff repeatedly edited the 78+ page paper to remove awkward AI-generated wording Model used in reported studies: GPT-3.5 - Most empirical work in the paper was done with a single model due to time constraints Training signal size for DSPy demo: less than 20 exemplars - Schulhoff says DSPy used the same small training data to produce a better prompt Dataset size in DSPy comparison: a couple hundred - The broader dataset used for the binary classification task in which DSPy beat his prompt Video/update pipeline latency: a couple hours per run - Their automated content-creation pipeline is expensive and time-consuming to execute Model comparison time cost: about a month - Estimated time to run the survey’s prompt evaluations across additional models like GPT-4 and others

Pivotal Quotes: "Trust no one, not even yourself." — Sander Schulhoff: His reflection on managing a large research team and the importance of tests, CI, and repeated validation "Defense here is always a losing game, unfortunately." — Sander Schulhoff: His view on prompt-injection/jailbreak defense after discussing adaptive defenses and offensive advances "DSPy was able to create a prompt with the exact same data that I had that blew me out of the water on the test set." — Sander Schulhoff: He describes being outperformed by automated prompt optimization in a binary classification task

Implications: Prompting is becoming a mature engineering discipline: details matter, automation is overtaking hand-tuning, and security/agent use remain unsettled. For builders, the message is to use exemplars, measure carefully, and expect model- or system-level improvements to matter most.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution