The TWIML AI Podcast
The TWIML AI Podcast

The Evolution of Reasoning in Small Language Models with Yejin Choi - #761

Today, we're joined by Yejin Choi, professor and senior fellow at Stanford University in the Computer Science Department and the Institute for Human-Centered AI (HAI). In this conversation, we explore Yejin’s recent work on making small language models reason more effectively. We discuss how hi

Featured Speakers

Yejin Choi Guest

Topics Discussed

Episode Summary

Executive Summary: Yejin Choi argues that AI progress should not be measured only by scale, but by better data, better reasoning, and better alignment with human values. She describes research on small language models, synthetic data generation, reinforcement-learning-style pretraining, and pluralistic alignment, while warning that post-trained models often become homogeneous and may flatten human diversity online.

Main Topics: Small language models as a democratizing goal (Priority: 5/5): Choi frames small models as a way to make generative AI accessible beyond large tech companies and GPU-rich labs, especially for academics and open research communities. Data-centric improvement of reasoning (Priority: 5/5): A major thread is that models learn better reasoning through high-quality, diverse, curated, and synthetic data rather than scaling alone. Synthetic data pipelines and diversification (Priority: 5/5): She explains iterative synthetic-data workflows that overgenerate, cluster, and aggressively filter examples to create diverse, high-quality reasoning data. Mode collapse and model homogeneity (Priority: 4/5): Choi discusses her 'Artificial Hive Mind' findings that open-ended model outputs are surprisingly similar both within and across models, raising ecosystem-level concerns. Reinforcement learning in pretraining (Priority: 4/5): She describes a new idea of adding RL-style reward during pretraining so models 'think' before predicting the next token, improving downstream reasoning. Pluralistic alignment and human-centered AI (Priority: 5/5): Choi outlines her vision for AI that reflects diverse human values through overtone, distributional, and steerable pluralism instead of one-size-fits-all alignment. AI’s societal direction and scientific use (Priority: 4/5): She warns of profit-driven incentives shaping AI in harmful ways and highlights AI for science as a promising area with major societal upside.

Key Arguments: Investment has disproportionately gone to scaling large models, but small models could unlock meaningful capability with much less compute if research effort were redirected. Internet data alone is insufficient; post-training depends on human-written, expert, and synthetic data that better teaches reasoning and domain-specific skills. Synthetic data is only useful when it is iteratively generated, diversified, and quality-filtered; naive prompting often produces repetitive or low-value outputs. Post-training tends to sharpen stereotypical responses, causing intra-model and inter-model homogeneity even for open-ended prompts. A model can be improved by generating harder reasoning data and filtering it using verifiers, clustering, or repeated-solution consistency checks. Adding a reward during pretraining that favors thought-assisted prediction can improve reasoning and persist into post-training. AI should be designed to reflect pluralistic human values, not optimize only for engagement, profit, or a single “correct” social viewpoint. Nonprofit and academic research are important because the future of AI should not be left entirely to commercial incentives.

Data Points: DeepSeek R1 teacher model size: 32B parameters - Used as a teacher model in the prismatic synthesis synthetic-data pipeline. DeepSeek R1 full model size: 671B parameters - Described as the much stronger teacher model the 32B version was compared against. Teacher-model size ratio: about 20x larger - Choi notes the full DeepSeek R1 is roughly 20 times larger than the 32B teacher model used in her work. Proxy model size: 1.5B parameters - A tiny Qwen 1.5B model was used as a proxy for computing gradient representations during data filtering. Synthetic dataset size: 1 million data points - The prismatic synthesis pipeline iteratively generated and filtered data until it assembled one million examples. Pretraining-data control variable: same FLOPs with less tokens - In the RL-as-pretraining experiment, one setting held compute constant while reducing tokens in the final phase. Location of event reference: NeurIPS - The 'Artificial Hive Mind' paper was described as one of the award-winning papers at NeurIPS.

Pivotal Quotes: "AI of humans, for humans and by humans." — Yejin Choi: Her framing of democratized, human-centered AI and pluralistic alignment. "Even for open-ended questions, the models are not as diverse as we would have expected." — Yejin Choi: Her description of model homogeneity and mode collapse in the Artificial Hive Mind paper. "Whatever is out of distribution, just make in distribution." — Yejin Choi: Her summary of the current generative-AI paradigm for improving robustness through coverage.

Implications: The conversation suggests AI progress will increasingly depend on data quality, synthetic-data engineering, and human-centered alignment. For listeners, the key takeaway is that smaller, safer, more diverse models may matter as much as ever-larger ones.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast