Episode Summary
Executive Summary: Tyler Cowen interviews Mercor CEO Brendan Foody about Mercor’s role as a marketplace for domain experts who help train and evaluate AI models. The conversation centers on replacing academic benchmarks with economically meaningful evals, the rapid pace of model improvement, and how AI will reshape labor markets, hiring, education, privacy, and expertise itself.
Main Topics: Mercor’s business model: hiring experts to train AI (Priority: 5/5): Foody explains that Mercor pays specialists like poets, lawyers, doctors, and financiers to create rubrics, examples, and evaluations that help frontier models improve in high-value domains. Why economic benchmarks matter more than academic ones (Priority: 5/5): The discussion contrasts narrow benchmark tests with task-based evaluations tied to real work outcomes, arguing that AI progress should be measured by economically valuable performance, not just exam scores. Speed of model improvement and near-term capability gains (Priority: 5/5): Foody argues that frontier models are improving rapidly, with large year-over-year gains and imminent advances in long-horizon tasks and tool use. Labor markets, matching, and the rise of RL environments (Priority: 4/5): The conversation explores a future where many knowledge workers help train agents by identifying model mistakes and creating reinforcement-learning environments, changing the nature of work rather than simply eliminating it. Talent assessment and hiring with AI (Priority: 4/5): Mercor’s own hiring process and broader labor-market inefficiencies are discussed, including better project-based interviews, AI-assisted screening, and the possibility of more efficient matching systems. Taste, expertise, and the limits of automation (Priority: 4/5): They debate subjective domains like poetry and law, where taste, tacit knowledge, and edge cases may keep humans relevant longer even as models become stronger. Education, privacy, and personalization (Priority: 3/5): Foody predicts AI tutors could transform education, while privacy-preserving personalization and trust brands like Apple may determine how much user data is collected and used.
Key Arguments: Mercor’s value lies in finding the best experts once and scaling their knowledge across billions of model uses, which justifies very high hourly pay for specialists. Real-world evals are more useful than academic benchmarks because customers care about automating tasks like diagnosis, legal drafting, financial analysis, and consulting work. Model improvement on economically meaningful tasks is extremely fast, with Foody claiming roughly 25–30% annual gains and expecting major advances in the next 6–12 months for long-horizon tool use. The biggest near-term AI bottleneck is not static knowledge but long-horizon execution and tool integration across many steps and systems. Human expertise will remain essential for the last 25% of difficult tasks, especially where taste, judgment, and uncodified knowledge matter. A major future job category will be people who train agents and build RL environments, rather than doing the underlying repetitive knowledge work themselves. Hiring should focus less on vibes and more on task performance; project-based assessments are superior predictors of success. AI will likely make some markets more efficient, but domains like software, investing, and business building are especially price elastic, so increased productivity may expand output and demand rather than simply eliminate jobs. Education will become more personalized with AI tutors, but teachers will still matter for guidance and emotional development. Privacy and trust will be decisive in whether users allow personalized data collection for AI systems. AI-assisted cheating or support should be treated as a feature in assessments when the goal is to measure what people can do with real tools. Better labor-market matching requires aggregating candidates and employers at scale, but the key problem is accurate matching, not just distribution.
Data Points: Mercor founding date: Early 2023 - Foody says Mercor dates from early 2023. Founder age: 22 - Tyler notes Foody is the youngest Conversations with Tyler guest ever. Unicorn status: Youngest unicorn founder ever (as described in the episode) - Tyler introduces Foody this way. Hourly pay for poets: $150/hour - Mercor advertises this to recruit poets who help train/evaluate models. Experts on platform: Tens of thousands - Foody says Mercor has tens of thousands of workers on the platform at any given time. Model improvement on Apex-style task evals: 25–30% per year - Foody characterizes the annual improvement rate at economically valuable tasks. GPT-5 score on one discussed benchmark: 64% - Used to illustrate frontier-model performance on the Apex evaluation. Near-term tool-use eval launch: In the next couple of months - Foody says Mercor is launching a longer-horizon evaluation soon. Expected capability jump for long-horizon tasks: 6–12 months - Foody says he would be shocked if models don’t become enormously capable here within that window. Difficulty stumping Cass Sunstein with models: 2–3 years - Foody estimates it may take this long for models to make it hard for Sunstein to find mistakes, depending on domain. Question-and-response error discovery for experts: Less than a year for 50 questions; around 6 months for a single question-response in some cases - Foody gives rough timelines for model robustness in law. Current Mercor full-time employees: Just over 300 - Foody gives the company size in October 2025. Working hours: 100 hours/week for the last 3 years - Foody describes his workload running the company. Potential annual value of pendant data: Tens of thousands to hundreds of thousands of dollars per person - Foody estimates the value of always-on personal conversation data for customers/business. Majority of high-end knowledge workers training models: Within five years - Foody predicts this shift in labor markets. AI-assisted hiring prediction horizon: 10 years - Foody suggests taped interviews plus AI could become superhuman predictors over a decade.
Pivotal Quotes: "The largest disconnect that we were seeing in AI research is that everyone was focused on academic evals... wholly disconnected from the outcomes that customers actually care about." — Brendan Foody: Explaining why Mercor built Apex to evaluate economically valuable tasks rather than benchmark exams. "I would be shocked if we don't have enormously capable models across those dimensions of lots of tool use with very long horizon tasks in the next six to 12 months." — Brendan Foody: His near-term forecast for model progress on long-duration, multi-tool workflows. "The most interesting part of our business is that everyone else in Silicon Valley is talking about how we automate away jobs versus we're very focused on how do we build this new job category of people training agents." — Brendan Foody: Describing Mercor’s view of AI’s labor-market impact as job transformation, not just replacement.
Implications: The episode suggests AI progress is now constrained less by raw intelligence than by measuring and training useful real-world behavior. Expect more demand for expert evaluators, AI-assisted hiring, personalized learning, and privacy-sensitive data collection—along with continued human relevance in taste-heavy domains.
About Conversations With Tyler
Tyler Cowen engages today’s deepest thinkers in wide-ranging explorations of their work, the world, and everything in between. New conversations every other Wednesday. Subscribe wherever you get your podcasts.