Episode Summary
Executive Summary: The episode argues that AI progress is outpacing the benchmarks used to measure it, making independent evaluation essential for labs, enterprises, and policymakers. VALS’ Ryan Krishnan explains why public leaderboards can mislead, how private benchmarks and rapid pre-release testing reveal true capability, and why evals are becoming the practical language for model selection, ROI, risk, and regulation.
Main Topics: Why public benchmarks are no longer enough (Priority: 5/5): Krishnan says open benchmarks can be gamed and may overstate capability, citing cases where models look strong publicly but underperform on held-out private tests. VALS as an independent evaluator (Priority: 5/5): The company was founded to provide third-party, high-signal benchmarks for a fast-moving market, analogous to rating agencies, auditors, and other independent verification bodies in mature industries. Enterprise ROI and model routing (Priority: 5/5): As AI spend rises, companies need bespoke evals to decide which model, agent, or pricing plan delivers the best performance and cost for specific workflows and repositories. Benchmark design for agentic and recursive systems (Priority: 4/5): Evaluations are shifting from simple one-shot tasks to long-horizon, multi-criteria workflows, including coding agents, finance tasks, and recursive self-improvement proxies. Policy, safety, and government verification (Priority: 4/5): The discussion frames evals as a bridge between government fears about bio/cyber risk and the empirical question of whether models can be induced to do harmful actions. Keeping benchmarks dynamic and retiring saturated tests (Priority: 4/5): VALS emphasizes creating new benchmarks as old ones saturate and updating tests to reflect current realities, such as updated legal standards or changing world knowledge. Geopolitics and shared AI standards (Priority: 3/5): The speakers connect evals to international coordination, arguing that shared measurement language could support trust-but-verify dynamics across countries and sovereign AI efforts.
Key Arguments: Independent evaluation is necessary because labs have incentives to self-report capabilities in ways that can distort the market. Public benchmarks often fail to reflect real-world performance; held-out private benchmarks can expose gaps hidden by leaderboard scores. Evals should be treated like a living system: as models and tasks change, benchmarks must be deprecated, replaced, and expanded. Enterprise AI adoption depends on custom benchmarks that measure task-specific ROI, latency, cost, and reliability, not just raw capability. Coding is a leading indicator for broader knowledge work evaluation; methods built there will likely generalize to finance, legal, and office workflows. Recursive self-improvement is important enough to merit its own benchmark family, using proxies for pre-training, post-training, and harness engineering. Government is better at defining and enforcing rules than at continuously testing frontier models; third-party evaluators can fill the verification gap. The most credible AI testing organizations must avoid conflicts of interest such as selling training data or consulting on the same systems they benchmark.
Data Points: VALS founding year: 2024 - Krishnan says VALS was started after the team concluded public benchmarks were insufficient. Pre-release testing window: ~6 hours - The discussion references the short window VALS has to evaluate a model before launch without delaying release. Engineering spend during token-maxing experiment: $1.5 million worth of tokens in one month - Krishnan describes an internal experiment using unrestricted coding tools. Token usage vs salary: 10x more in tokens than employee salary - Used to illustrate how AI spend can eclipse labor costs in enterprise settings. Daily token usage at peak: 1 to 2 billion tokens/day per engineer; one engineer hit $6 billion tokens on peak day - Internal usage patterns during the token-maxing experiment. Fortune 10 usage budget: $100/day per engineer, later increased to $300/employee - Example of how enterprises are budgeting AI usage. Open-ended benchmark scope: 50 full-stack web applications - Example of a more complex, low-sample evaluation task versus classic image classification. Historical benchmark style: Millions of images - Referenced by comparison to ImageNet-style one-to-one evaluation.
Pivotal Quotes: "Every time a new trillion dollar industry emerges, there's a need for this independent testing group." — Ryan Krishnan: Explaining why third-party evaluation emerges as a necessity in rapidly scaling industries. "On our held-out private benchmarks, the model is actually underperforming. But on all of the major public benchmarks... it was showing incredible capabilities." — Ryan Krishnan: Cited as evidence that public benchmarks can overstate real model performance. "A firm really is just its evals." — Ryan Krishnan: Used to argue that enterprise competitiveness will increasingly depend on legible, task-specific evaluations.
Implications: AI buyers, builders, and regulators will increasingly rely on independent, dynamic evals to choose models, price intelligence, and manage risk. Benchmarks are becoming infrastructure, not marketing, and the winners will be those who can measure capability and harm in real workflows.
About The a16z Podcast
The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!