The TWIML AI Podcast
The TWIML AI Podcast

Music & AI Plus a Geometric Perspective on Reinforcement Learning with Pablo Samuel Castro - #339

Today we’re joined by Pablo Samuel Castro, Staff Research Software Developer at Google. We cover a lot of ground in our conversation, including his love for music, and how that has guided his work on the Lyric AI project, and a few of his papers including “A Geometric Perspective on Optimal Represen

Featured Speakers

Pablo Samuel Castro Guest

Topics Discussed

Episode Summary

Executive Summary: Pablo Samuel Castro describes a career bridging reinforcement learning research, engineering, and music-driven creativity. He explains his transition from academia to Google, then details a lyric-generation project with David Usher that evolved from a simple RNN to a two-model transformer system for structure and vocabulary. He also discusses RL research on representations and a Bank of Canada multi-agent payments project, emphasizing starting simple and adding complexity only when needed.

Main Topics: Career path across academia, industry, and music (Priority: 5/5): Castro explains moving from Ecuador to McGill, staying in Montreal for family and music, leaving academia after a difficult postdoc job search, then returning to research through Google Brain while keeping music central. Engineering as a core part of modern ML research (Priority: 5/5): He argues that major AI advances depend heavily on implementation and infrastructure, and that his role as a staff research software developer lets him contribute through both research and engineering. AI-assisted lyric generation with David Usher (Priority: 5/5): A long-running creative collaboration evolved from a character-level RNN into a transformer-based tool intended to help songwriters brainstorm rather than replace them. Using multiple models to separate structure from vocabulary (Priority: 5/5): To improve lyrical quality, the system splits the task into a structure model (parts of speech, syllables, phonemes) and a vocabulary model (filling words from book-trained language patterns). Reinforcement learning for representation learning (Priority: 4/5): Castro discusses team research on value-function polytopes and adversarial value functions to learn representations that generalize better across policies and reward changes. Bank of Canada multi-agent reinforcement learning (Priority: 4/5): He outlines a collaboration modeling inter-bank payments as a multi-agent RL problem, decomposed into simpler subproblems to validate feasibility before scaling up. Skepticism toward unnecessary RL complexity (Priority: 5/5): Across projects, he stresses solving concrete problems with the simplest workable method first, using RL only when the problem truly benefits from it.

Key Arguments: His career was shaped by personal priorities, especially music and family, which influenced staying in Montreal and later re-entering research from industry. Google’s research environment allowed him to pursue independent projects while contributing engineering expertise, showing that software development can be a research-enabling role. Creative text generation for music is not about producing a complete hit automatically; it should function as a tool that augments songwriters’ workflows. A single language model struggled to balance coherence, creativity, and lyrical structure, so splitting the task into structural and lexical components improved control and quality. Lyrics data alone overrepresents repetitive, high-probability patterns, so training on books improves vocabulary diversity even if it weakens song-like structure. Representation learning in RL should not overfit to one optimal policy; richer representations help adapt to altered reward functions and new policies. Distributional and expectation-based RL can be equivalent under some conditions, suggesting the advantage of distributional methods may emerge mainly in deep-network settings. The Bank of Canada project demonstrates how RL can help model complex economic systems, but only after decomposing them into tractable subproblems and validating each part. Graph and multi-agent settings are underexplored in Castro’s experience, but they introduce important interaction and game-theoretic challenges.

Data Points: NIPS conference attendance: about 4,000 people historically vs. 12,000 people now - Castro contrasts the conference’s earlier scale with its current size. PhD completion year: 2011 - He finished his PhD at McGill around 2011 before moving to Paris for a postdoc. Project timeline: almost 2 years - He says the David Usher lyrics project has been ongoing for nearly two years. Paper length: 2-page paper - He describes the creativity-workshop writeup as very preliminary and brief. Training time: less than an hour - He says the TPU training itself is fast once preprocessing is done. Data preprocessing: the longest part - He notes that decomposition into parts of speech and related preprocessing dominates runtime.

Pivotal Quotes: "I think more and more it is, but a few years ago, I don't feel it got the credit it really deserved." — Pablo Samuel Castro: He is emphasizing that engineering is central to modern machine learning advances. "The idea is not to sort of replace musicians or songwriters, but to enhance them." — Pablo Samuel Castro: He explains the purpose of the lyric-generation tool as creative assistance, not automation. "I don't want to throw RL at just because I know it well." — Pablo Samuel Castro: He describes his philosophy of starting with the simplest solution before adding complexity.

Implications: The episode highlights practical, tool-oriented AI: augmenting creative work, improving RL representations, and modeling complex systems. It suggests future progress will come from better architectures, clearer task decomposition, and human-in-the-loop design rather than fully automated generation.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast