No Priors
No Priors

Humans&: Bridging IQ and EQ in Machine Learning with Eric Zelikman

The AI industry is obsessed with making models smarter. But what if they’re building the wrong kind of intelligence? In launching his new venture, humans&, Eric Zelikman sees an opportunity to shift the focus from pure IQ to building models with EQ. Sarah Guo is joined by Eric Zelikman, formerly

Featured Speakers

Eric Zelikman Guest

Episode Summary

Executive Summary: Eric Zelikman traces his research from making language models “half decent” at reasoning to a broader mission: building AI that understands people and collaborates with them over long time horizons. He argues current models are jagged, weak on emotional/contextual understanding, and over-optimized for single-task benchmarks, while his new company HumanSend aims to make models more useful by modeling user goals, memory, and long-term impact.

Main Topics: Origin story and motivation for AI (Priority: 5/5): Zelikman explains that his interest in ML came from a humanist concern: too much talent and potential is wasted by circumstance, and AI should free people to do what they care about. STAR and QuietSTAR: scaling reasoning through RL (Priority: 5/5): He describes STAR as a simple iterative method that rewards correct reasoning traces and QuietSTAR as a step toward scaling that approach to pretraining-style data, with important improvements like online learning and baselines. What models are good at vs. where they fail (Priority: 4/5): He characterizes current models as reasonably strong on well-specified, verifiable, and closed-form tasks, but brittle on tricky prompts, long-horizon planning, and understanding human intent. Limits of verifiable-task optimization (Priority: 4/5): The conversation emphasizes that coding and other verifiable domains benefit from more time, better context, and improved out-of-distribution handling, but product constraints and responsiveness limit how far models can go. Human-centered AI and the case for keeping people in the loop (Priority: 5/5): Zelikman argues that scaling should not only automate humans out; instead, models should help people collaborate, retain agency, and grow the value they create. HumanSend and the shift toward EQ, memory, and long-term interaction (Priority: 5/5): He frames his new company as an attempt to build models that understand users’ goals, remember context, and optimize for long-term outcomes rather than one-turn benchmark performance. Hiring and technical priorities (Priority: 3/5): He says the early team needs strong infra, research, and product people who can build, plus expertise in memory, distributed systems, fast inference, and tasteful interaction design.

Key Arguments: AI should augment human potential, not merely replace existing jobs or tasks; the best systems will help people do what they already want to do better. STAR showed that iterative RL can improve reasoning traces and that performance can keep improving across many iterations, with no obvious plateau in early experiments. QuietSTAR demonstrated that reasoning-oriented RL methods can be extended toward pretraining-scale settings using standard language-model data. Current models are jagged: they can solve some advanced math/physics/HLE-style questions, but are still weak at understanding long-term human goals and social nuance. For many real applications, the biggest bottleneck is not only raw intelligence but context, verifiability, response-time constraints, and out-of-distribution inputs. Benchmark culture overweights single-task accuracy because it makes credit assignment and resource allocation easier, but this misses whether models actually help people over time. Memory and multi-turn understanding are underdeveloped because the field trains on task-centric objectives that do not reward long-term user modeling. A more useful AI paradigm would maintain user context across interactions, anticipate consequences, and proactively support decisions rather than forcing users to re-provide context every time. Human collaboration can expand the “pie” of what is possible, whereas pure automation risks narrowing innovation to replacement-only use cases.

Data Points: n-digit arithmetic: Used as an example task family where models showed increasing capability over training iterations - Zelikman cites experiments on addition/multiplication to illustrate scaling in reasoning 2021: Reference point for early language-model capability - He notes that models then were still not very smart and chain-of-thought was only a small step 2-hour tasks: METR benchmark example of autonomous task duration - He describes progress from completing 2-hour tasks without human intervention 2.5-hour tasks: METR benchmark next step - He uses this to show how the industry tracks expanded autonomous horizons 20,000 lines: Approximate size of generated code PRs mentioned jokingly - Used to illustrate how large AI-generated code changes have become 100,000 lines: Another size example for AI-generated PRs - Mentioned to show the scale of code generation in current workflows 10%: Example benchmark improvement figure - He cites benchmark gains as a basis for resource allocation 5%: Example benchmark improvement figure - Contrasted with 10% to explain internal lab credit assignment Decade plus: Long-term horizon referenced in decision-making - Used as a comparison for optimizing for enduring user value rather than one interaction

Pivotal Quotes: "All of humanity is not living up to their full potential." — Eric Zelikman: He explains his motivation for building AI that frees people to do meaningful work "I’d like to work on models that empower people instead of replacing them." — Eric Zelikman: He summarizes the philosophy behind HumanSend and his broader research direction "Imagine if every time you interacted with someone, you basically like, they remember your name and like, you know, maybe what you do, and like just the really high-level sketch of your life." — Eric Zelikman: He describes the gap between current models and truly useful long-term memory

Implications: The discussion suggests the next frontier is less about raw benchmark wins and more about models that remember, collaborate, and optimize for human goals over time. For industry, that means new data, evals, and products centered on long-term user value.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors