Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

Happy New Year! You may have noticed that in 2025 we had moved toward YouTube as our primary podcasting platform. As we’ll explain in the next State of Latent Space post, we’ll be doubling down on Substack again and improving the experience for the over 100,000 of you who look out for our emails and

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: Artificial Analysis founders explain how independent benchmarking evolved from a side project into a business serving enterprises and AI companies. They detail the technical rigor behind their public and private evals, introduce new metrics for hallucination, openness, and agentic work, and argue that AI benchmarking must keep changing as models get smarter, more multimodal, and more tool-driven.

Main Topics: Origin story and mission (Priority: 5/5): The company began as a practical need while building AI products: no independent source existed for comparing models on quality, speed, and cost. What started as a side project became a public benchmarking platform focused on neutrality and usefulness to developers. Business model and customer segments (Priority: 5/5): Artificial Analysis monetizes through enterprise benchmark/insight subscriptions and custom private benchmarking, while keeping public website data free and ensuring no one pays to appear on public rankings. Benchmarking methodology and rigor (Priority: 5/5): The founders describe why they run evaluations themselves, standardize prompts, repeat runs to tighten confidence intervals, use mystery-shopping to avoid manipulation, and avoid relying on lab-reported numbers because of prompt, format, and setup variance. New metrics: hallucination, openness, and agentic performance (Priority: 5/5): They introduced the Omniscience Index for factual uncertainty/hallucination, an Openness Index measuring disclosure and licensing, and GDPVal plus an agentic harness for work-like tasks, broadening evaluation beyond classic QA benchmarks. Industry trends: cost, speed, and model scaling (Priority: 4/5): A major theme is the 'smiling curve' of AI economics: cost per intelligence tier is collapsing while total spend can still rise because models are used more intensively in longer, agentic workflows and on better hardware. Reasoning models, token efficiency, and multi-turn workflows (Priority: 4/5): They argue the old binary of reasoning vs non-reasoning is becoming less useful. Better metrics now include token efficiency, turns to completion, and how models adapt token use to task difficulty. Open source versus openness (Priority: 4/5): The team distinguishes open weights from true openness, adding a structured index for data transparency, training code disclosure, and licensing constraints, arguing that openness should be measured on more than just weight release.

Key Arguments: Independent benchmarking is necessary because labs can prompt, format, and otherwise evaluate their own models in ways that are not comparable or fully transparent. Benchmarking must include cost and performance alongside intelligence, because real developers optimize for trade-offs, not just raw accuracy. A model's public benchmark score can be manipulated or optimized toward, so new evals must be continuously created to stay relevant to real-world use. Hallucination should sometimes be penalized directly, since saying 'I don't know' is better than confidently giving a wrong factual answer. No single metric captures model quality; the right approach is a portfolio of evals covering knowledge, agents, long context, code, hallucination, openness, and more. Model openness should include not only weights/licenses but also disclosure of data, methodology, and training code. The economics of AI are paradoxical: unit costs fall rapidly, but total spend rises because users deploy more models, more tokens, and more agentic steps. Future benchmarks will need to focus more on real work: multi-turn, tool-using, multimodal, and task-completion behavior rather than single-turn QA only.

Data Points: Company age: Almost 2 years old in January 2026 - Artificial Analysis is approaching its second anniversary. Team size: Just over 20 people - Headcount mentioned during the business-model discussion. Public benchmark coverage: 10 different eval datasets - Current Artificial Analysis Intelligence Index composition. Hallucination score range: -100 to +100 - Omniscience Index scores are penalized for incorrect answers and reward saying 'I don't know'. Public test set disclosure: 10% released publicly - Most of the Omniscience test set is held out to reduce contamination. Confidence target: ±1 at 95% confidence - Desired precision for the Intelligence Index via repeated runs. V1 benchmark saturation: Near 100% solved by modern small models - Older benchmark tasks are now trivial for current models. Topology of GDPVal: 44 tasks / about 220-225 subtasks - The work-like benchmark spans many structured tasks. GDPVal score top model result: Top scores around 9% - Critical Point physics benchmark is extremely hard. Openness Index max score: 18 - Points are awarded for data/code/license transparency and related openness dimensions. Cost improvement example: GPT-4-level intelligence is over 100x cheaper than at launch - Used to illustrate rapid decline in unit cost of intelligence. Active parameter examples: Open weights models around 5% active; Kimi K2 about 3% active - Used in discussion of sparsity and active vs total parameters. Hardware generation uplift: About 3x on the slide; sometimes higher in practice - Blackwell vs Hopper inference gains were discussed as workload-dependent.

Pivotal Quotes: "No one pays to be on the website." — Micah/George: Explaining the company's commitment to independent public benchmarking and neutrality. "You couldn't look at those in isolation, you needed to look at them alongside the cost and performance stuff." — George: Why intelligence scores alone are insufficient for developers choosing models. "The things that get measured become things that get targeted." — George: Discussing benchmark gaming and why new evals must evolve.

Implications: For builders and enterprises, the message is to choose models using multi-axis, continually updated benchmarks, not headline scores. For the industry, AI evaluation is shifting toward agentic, multimodal, and trust-focused metrics that better reflect real work.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast