The TWIML AI Podcast
The TWIML AI Podcast

Social Commonsense Reasoning with Yejin Choi - #518

Today we’re joined by Yejin Choi, a professor at the University of Washington. We had the pleasure of catching up with Yejin after her keynote interview at the recent Stanford HAI “Foundational Models” workshop. In our conversation, we explore her work at the intersection of natural language generat

Featured Speakers

Yejin Choi Guest

Topics Discussed

Episode Summary

Executive Summary: Yejin Choi argues that AI’s biggest gap with human intelligence is common sense and abductive reasoning: the ability to infer plausible explanations, social norms, and physical expectations from incomplete context. She describes work on COMET and social norm transformers, built from natural-language knowledge graphs, and contends that future progress will require hybrid systems combining neural models, symbolic constraints, grounded learning, and better benchmarks—not scale alone.

Main Topics: Common Sense Reasoning as the Core AI Challenge (Priority: 5/5): Choi frames common sense as everyday practical knowledge that humans use effortlessly but machines still struggle to represent, reason over, and generalize from. COMET and Natural-Language Knowledge Graphs (Priority: 5/5): She explains COMET, a transformer-based system trained on an atomic commonsense knowledge graph expressed in free-text if-then statements, to generalize to unseen scenarios. Abductive, Counterfactual, and Story Reasoning (Priority: 4/5): Choi emphasizes that human reasoning is often abductive—inferring best explanations from partial evidence—and connects this to story understanding and what-if reasoning. Social and Moral Common Sense (Priority: 4/5): Beyond physical world knowledge, Choi discusses learning social norms, politeness, fairness, and ethical judgments so dialogue systems can behave appropriately. Limits of Scale and Need for Hybrid Methods (Priority: 5/5): She argues that scaling neural models alone will not solve reasoning; richer learning signals, symbolic integration, and algorithmic inference remain necessary. Benchmarks, Evaluation, and Leaderboard Incentives (Priority: 4/5): Choi critiques multiple-choice benchmarks and leaderboard-driven research, arguing they reward pattern matching and can mask failures in genuine reasoning. Future Research Directions (Priority: 3/5): She sees the most promise in interactive learning, grounded multimodal understanding, and methods that better mirror how humans acquire conceptual knowledge.

Key Arguments: Common sense is a broad body of everyday knowledge that is highly contextual and often difficult to formalize, making it a longstanding AI challenge. Natural-language representations are central because they preserve ambiguity, context, and expressive power better than rigid logical formalisms. COMET shows that transformer models can generalize commonsense inferences beyond memorized rules when trained on a large natural-language knowledge graph. Human reasoning is frequently abductive and counterfactual rather than purely deductive, so AI should be evaluated on plausible inference rather than strict formal correctness alone. Social norms and moral norms should be modeled explicitly so conversational systems can avoid unsafe, rude, or harmful outputs. Scale helps, but scale alone will not solve reasoning; hybrid approaches combining neural models, symbolic constraints, and grounded learning are needed. Current benchmarks often incentivize shortcut learning and multiple-choice exploitation instead of true generative reasoning. AI should learn interactively and conceptually, more like humans, rather than only consuming large static text corpora in a fixed order.

Data Points: Atomic knowledge graph size: 1.3 million if-then rules - Choi describes the commonsense knowledge graph used to train COMET. Inference types in original atomic knowledge graph: 9 - Original graph focused on human-centric inference types such as mental states, wants, and causes/effects. Inference types in later knowledge graphs: 24 - Expanded to include object-centric and event-centric knowledge as well. Amazon Alexa Challenge conversation goal: 20 minutes - Choi cites the task as sustaining coherent conversation with humans for 20 minutes. Observed system performance in Alexa challenge: Sometimes survived 10 minutes - Their system could sometimes maintain conversation for about 10 minutes, but not reliably for 20.

Pivotal Quotes: "Common sense knowledge and reasoning. This was, in fact, the early dream of the AI field in the 70s and 80s." — Yejin Choi: Introduces the historical motivation behind her commonsense reasoning research. "Language is the best medium for reasoning." — Yejin Choi: Explains why she prefers natural language over rigid logical formalisms for representing and evaluating reasoning. "We cannot reach to the moon by making the tallest building in the world one inch taller at a time." — Yejin Choi: Her argument that scaling alone will not achieve AI breakthroughs like reasoning or AGI.

Implications: For AI to become robust and socially useful, it must move beyond pattern matching toward grounded, hybrid reasoning systems. Future progress likely depends on better benchmarks, explicit norms, and models that learn concepts more like humans.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast