The Cognitive Revolution
The Cognitive Revolution

From Poetry to Programming: The Evolution of Prompt Engineering with Riley Goodside of Scale AI

Nathan hosts Riley Goodside, the world's first staff prompt engineer at Scale AI, to discuss the evolution of prompt engineering. In this episode of The Cognitive Revolution, we explore how language models have progressed, making prompt engineering more like programming than poetry. Discover in

Featured Speakers

Nathan Labenz and Erik Torenberg HostRiley Goodside Guest

Topics Discussed

Episode Summary

Executive Summary: Nathan and Riley Goodside discuss how prompt engineering has evolved from clever hacks into a more systematic, programming-like discipline centered on data quality, decomposition, tool use, and evaluation. They argue that post-training, synthetic data, retrieval, and fine-tuning can push models far beyond naive chatbot use, while also highlighting limits around generality, jailbreak resistance, and AGI-style claims.

Main Topics: Prompt engineering is maturing into programming (Priority: 5/5): Riley argues the field has shifted from poetic prompt tricks to a more code-like practice involving pipelines, wrappers, evaluations, and systematic workflow design. Tool use, decomposition, and fine-tuning (Priority: 5/5): For real-world API/tool use, Riley says prompt-only approaches often fail; success comes from breaking tasks into smaller parts, creating high-quality examples, and fine-tuning when needed. Reasoning traces and synthetic data (Priority: 5/5): The conversation emphasizes that explicit chain-of-thought, internal notes, and synthetic rationales materially improve model performance and can be used to generate better training data. Model selection, benchmarks, and cost/performance tradeoffs (Priority: 4/5): Riley stresses domain-specific evaluation, noting that the best model depends on the task, language, and workflow constraints, not generalized rankings alone. Limits of fine-tuning, generality, and AGI framing (Priority: 4/5): They discuss catastrophic forgetting, hallucination risks, and why AGI definitions based on economically valuable labor are moving targets rather than stable endpoints. Safety, jailbreaking, and open-source tensions (Priority: 4/5): Riley suggests some vulnerabilities are hard to eliminate purely through training, implying continued need for filters, system-level defenses, and difficult tradeoffs for open models. Frontier optimism across new modalities (Priority: 4/5): They discuss how ideas from text models may extend to video, biology, chemistry, and simulation, suggesting broad upside from training on many more structured data sources.

Key Arguments: Modern prompt engineering is increasingly about engineering workflows and code around models, not just writing clever prompts. LLMs often need decomposition into substeps, explicit examples, and curated edge cases to perform reliably on real tasks. For API/tool use, fine-tuning is often justified because semantic details and rigid JSON-like structures are hard to reliably convey with prompts alone. Reasoning traces matter because they expose intermediate cognition that is usually missing from raw input/output pairs. Synthetic data and self-generated rationales can improve performance, especially when filtered for correctness or consensus. The best model is task-specific; benchmark and domain evaluation matter more than brand-level assumptions. Fine-tuning can reduce generality and increase hallucination risk, so teams should avoid it unless prompt/K-shot/RAG methods are insufficient. Some jailbreak or safety issues may not be fully solvable in weights alone, making output/input filters and system-level controls necessary. AGI should not be treated as a fixed target because the set of economically valuable tasks changes as automation advances. There is substantial untapped room in post-training and in new training modalities beyond text, including video, DNA, and scientific simulation.

Data Points: GPT-3.5 fine-tuning experience: Used extensively - Nathan says most of his early fine-tuning work was on GPT-3.5 and that explicit reasoning traces helped materially. June 2023: First child born - Riley explains his reduced Twitter activity partly because he became a father in June 2023. 2022 to early 2023: Riley’s rise to fame - Nathan references Riley’s prominence during the Text-DaVinci-002 era and early prompt-engineering wave. GPT-4 scale discussion: 10–15 trillion tokens vs. 100 trillion tokens - They speculate that a much larger pretraining run could produce a qualitatively different model. Farming employment share in the U.S.: About 2% - Riley uses agriculture as an example of how economically valuable labor shifts over time. Historical farming labor share: Around half of all humans - Riley notes that before industrialization, roughly half of people were involved in farming-related work. Multi-generation efficiency: 40 API calls to GPT-3.5 ≈ 1 API call to GPT-4 - Riley cites a paper suggesting many cheap samples can approximate a stronger single model call. Fine-tuning data defense: Walnut53 cipher - Riley describes a paper that used a simple character permutation cipher to hide unsafe fine-tuning payloads from safety filters. Efficiency gain: 10x - Nathan cites the JEST paper’s reported training-efficiency improvement from selecting data based on what smaller models can learn. Open source protection example: Llama 2 - Riley mentions concept-filtering approaches starting with Llama 2 as a strategy to remove sensitive concepts from pretraining data.

Pivotal Quotes: "Modern prompt engineering is becoming more programming than poetry." — Riley Goodside: He uses this line to explain why his earlier Twitter-style prompting tricks matter less than systematic pipeline design today. "LLMs can do tasks, not jobs." — Riley Goodside: He offers this as a practical rule for deciding what kinds of work are automatable. "The best model is the one that does the best on your task." — Riley Goodside: He argues against relying on generic model reputations and in favor of domain-specific evaluation.

Implications: Builders should focus on task decomposition, synthetic data, reasoning traces, and rigorous evals—not prompt hacks alone. Expect more automation from post-training and multimodal data, but also more need for safety layers and careful model/task matching.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution