Episode Summary
Executive Summary: MIT researcher Anna Ivanova argues that large language models excel at formal linguistic competence—grammar, style, and pattern completion—but still fall short on functional competence like reasoning, world knowledge, situation modeling, and intent. The conversation explores how neuroscience-informed benchmarks can better measure these gaps and how LLMs can also illuminate human cognition.
Main Topics: Anna Ivanova's background and MIT Quest for Intelligence (Priority: 4/5): Ivanova explains her transition from neuroscience and cognitive science into AI research, and describes MIT Quest for Intelligence as an initiative uniting scientists, engineers, and industry partners around intelligence research. Formal vs. functional linguistic competence (Priority: 5/5): A central framework in the discussion distinguishes language form (grammar, syntax, co-occurrence patterns) from language use in the world (reasoning, intent, world knowledge, social competence). The paper argues LLMs are much stronger at the former than the latter. What LLMs can and cannot do (Priority: 5/5): The conversation examines model abilities such as paraphrasing, syllogisms, and fluent generation versus weaknesses in math, novel reasoning, temporal ordering, and maintaining coherent situation models across longer contexts. Benchmarks and evaluation gaps (Priority: 4/5): Ivanova notes that existing benchmarks often conflate formal and functional competence, and discusses her work on a new cognitively informed benchmark for testing world knowledge, intuitive physics, and intuitive psychology. Language, grounding, and reporter bias (Priority: 4/5): The discussion covers whether language alone is enough to learn world knowledge, emphasizing that text contains both useful implicit knowledge and biases from what people choose to talk about, which scaling alone will not eliminate. Human cognition, inner speech, and individual differences (Priority: 3/5): Ivanova highlights variation in how people think—verbal, visual, or abstract—and connects this to research on aphantasia and the broader question of how humans represent thought. Why LLMs matter for cognitive science and AI (Priority: 4/5): She argues LLMs are useful both for understanding AI behavior responsibly and for testing theories about human cognition, including whether language input alone can support certain kinds of reasoning.
Key Arguments: LLMs are remarkably good at formal linguistic competence, but that should not be mistaken for human-like understanding or general intelligence. Functional linguistic competence is broader: it includes reasoning, math, world knowledge, situation modeling, social competence, and intent. Current benchmarks often fail to separate language form from deeper competence, making it hard to know what a model actually knows. Scaling up data alone increases coverage but does not remove reporting bias in text or guarantee common-sense understanding. Human cognition suggests modular specialization, so a truly human-like language system may also need modular or specialized components. LLMs can help cognitive science by testing what can be learned from language input alone and by comparing model behavior with human behavior. Comparing LLM internals with brain activity is already showing meaningful alignment, suggesting these systems can be useful models of aspects of human language processing. People often over-attribute mental states to chatbots because they automatically infer intention and agency from language. Human-like language performance may be more useful than superhuman language performance if the goal is to model actual human communication.
Data Points: PhD completion timing: last summer - Ivanova says she defended her PhD in neuroscience/cognitive science last summer before moving into LLM research. MIT Quest for Intelligence launch: a few years ago - She describes the initiative as a relatively recent MIT effort to unify research on intelligence. Functional MRI setting: giant magnet / giant tube - Used as the human-neuroscience analogy for measuring brain activity during tasks. Context window: substantially more than earlier models - She says ChatGPT appears to attend to far more tokens than prior models, helping with situation tracking. Modeling layers to predict brain responses: new sentences not seen before - She notes that brain responses to unseen sentences can be predicted from LLM embeddings. Human inner speech self-report: 90% - She cites an informal example where some people report thinking in words around 90% of the time.
Pivotal Quotes: "On one hand, we have what we call formal linguistic competence, and that's just the ability to use a language." — Anna Ivanova: Defines the first half of the paper’s framework separating language form from language use. "They don't understand and use language in the way that humans do, at least not yet." — Anna Ivanova: Summarizes her view that current LLMs are not equivalent to human intelligence. "We want a model that uses language like us." — Anna Ivanova: Explains why the goal is human-like language behavior rather than a superhuman ideal.
Implications: The episode suggests LLM evaluation needs richer, cognitively grounded benchmarks and human comparison. For industry, it argues against equating fluency with intelligence; for research, it points toward modularity, grounding, and brain-aligned analysis.