Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge

The first AI Engineer World’s Fair talks from OpenAI and Cognition are up! In our Benchmarks 101 episode back in April 2023 we covered the history of AI benchmarks, their shortcomings, and our hopes for better ones. Fast forward 1.5 years, the pace of model development has far exceeded the speed at

Featured Speakers

Latent.Space HostClémentine Fourier Guest

Topics Discussed

Episode Summary

Executive Summary: Clémentine Fourier of Hugging Face discusses how the OpenLLM leaderboard evolved from an internal research tool into a widely used, community-driven evaluation platform. The conversation focuses on benchmark design, prompt sensitivity, human vs model judging, compute constraints, and the need for harder, more reliable evaluations as LLMs saturate older tests.

Main Topics: Clémentine Fourier’s path to Hugging Face and evaluation work (Priority: 4/5): She explains her background in geology, her PhD at Inria, and how she moved from experimental science into ML and then into evaluation and leaderboards at Hugging Face. How the OpenLLM leaderboard was built and scaled (Priority: 5/5): The leaderboard began as an internal comparison tool and became a major public resource with thousands of models, community submissions, and broad adoption. Benchmark saturation and the need for V2 (Priority: 5/5): Older benchmarks like MMLU and ARC Challenge became less discriminative as models improved, motivating a new leaderboard version with harder, more stable tasks. Human evaluation vs automated evaluation vs model judges (Priority: 5/5): She distinguishes vibe checks, crowd-based ratings, expert human annotation, and LLM-as-judge approaches, arguing that model judges and crowd ratings have serious limitations. Benchmark selection, prompt formatting, and implementation rigor (Priority: 5/5): The discussion covers why specific benchmarks were chosen, how prompt formatting can change scores dramatically, and why careful implementation details matter for fairness. Missing benchmarks: long context, agents, calibration, and psychophancy (Priority: 4/5): She identifies gaps in current evaluation, especially long-context reasoning, agentic tasks, calibration, robustness to prompting, and psychophancy.

Key Arguments: Older benchmarks become less useful once models saturate them; high scores can reflect contamination or benchmark weakness rather than true capability. Automated benchmarks are reproducible and fair, but they only cover a limited slice of model behavior. Human evaluations are useful for subjective or real-world use cases, but crowd-based systems are noisy, biased, and often not reproducible. LLM-as-judge systems introduce subtle biases, including preference for similar model families, first answers, and verbose outputs. Prompt formatting can materially change benchmark results; evaluation design is not neutral and can shift scores by large margins. The leaderboard should prioritize discriminative, stable, and well-implemented benchmarks, even if that means dropping flawed datasets. Long-context, agentic, calibration, and psychophancy evaluations are important gaps that the field still needs to address. Compute constraints shape benchmark design; multi-choice tasks are favored partly because they are cheaper and easier to parallelize.

Data Points: OpenLLM leaderboard models evaluated (V1): 7,400 models - Total models evaluated on version 1 of the leaderboard, mostly community-submitted. Community discussion threads: ~800 threads - User support and suggestions around the leaderboard. Leaderboard traffic: Several million visitors - Scale of public usage since creation. Evaluation cluster cost: ~$100/hour - Approximate public price of the H100 instance class used for evaluations. 7B model evaluation time: ~2 hours - Typical runtime for evaluating a 7B model on the current setup. 7T model evaluation time: ~20 hours - Typical runtime for evaluating a very large model. Prompt sensitivity on MMLU: 30-point variation out of 100 - Different prompt formats caused large score swings in experiments. MMLU human performance: 80-something - Used to illustrate benchmark saturation and contamination concerns. PhD duration in France: 3 years - Clémentine explains the structure of her PhD funding and timeline. Benchmark choice for math: Level 5 questions only - Selected to make the benchmark more discriminative and reduce compute cost.

Pivotal Quotes: "When a benchmark reaches saturation, so when models basically get the same performance as humans on a benchmark, or go above human performance, what it actually says is usually that the models are completely contaminated on said benchmarks." — Clémentine Fourier: Explaining why older benchmarks need to be revisited and replaced. "I think people should stop using LLM as judges, because they have a lot of subtle biases that they introduce in evaluation." — Clémentine Fourier: Her critique of model-as-judge evaluation methods. "If you are an engineer, and you want to know which model is best for your specific use case, please do a vibe check." — Clémentine Fourier: Her advice on combining leaderboard results with real-world testing.

Implications: The episode shows that LLM evaluation is becoming a core infrastructure problem: benchmarks must evolve as models improve, and future progress will depend on harder tests, better calibration, and more rigorous methods for judging real-world usefulness.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast