The a16z Podcast
The a16z Podcast

The True Cost of Compute

With software becoming more important than ever, hardware is following suit. As the world generates more data, unlocking the full potential of AI means a constant need for faster and more resilient hardware. But how much does this all really cost? In this final segment of our AI hardware series, we

Featured Speakers

a16z HostGuido Appenzeller Guest

Topics Discussed

Episode Summary

Executive Summary: This episode closes an AI hardware series by examining what it really costs to train and run modern AI models. Guido Appenzeller explains that training can cost millions to tens of millions of dollars, while inference is far cheaper but constrained by peak demand. He argues compute is a major moat, but not an insurmountable one because faster chips and a growing data bottleneck may eventually stabilize training costs.

Main Topics: The true cost of AI model training (Priority: 5/5): Guido explains that training large language models is far more expensive than many assume, often reaching millions or tens of millions of dollars once reservations, inefficiencies, and repeated runs are included. Inference vs. training economics (Priority: 5/5): The conversation distinguishes the high cost of training from the much lower per-query cost of inference, while noting that inference still requires expensive capacity planning for peak usage. How compute requirements can be estimated (Priority: 4/5): Guido walks through rough calculations using transformer models, parameter counts, FLOPs, and accelerator throughput to estimate training time and cost, emphasizing that the math is only approximate. Hardware utilization and infrastructure constraints (Priority: 4/5): The episode highlights that real-world utilization is often far below theoretical peak because of memory bandwidth, networking, and distributed training overhead, making naive cost estimates misleading. Capital intensity and competitive advantage (Priority: 5/5): The discussion asks whether expensive compute advantages favor incumbents, but Guido suggests the capital moat is real yet not necessarily deep enough to stop well-funded new entrants. Data as the emerging limiting factor (Priority: 5/5): Guido argues that AI may be moving from a compute-limited phase toward a data-limited phase, where the amount of high-quality training material—not just compute—caps model scaling. Open-source and future innovation (Priority: 4/5): Because training costs are high, open-source frontier models remain difficult to produce; however, Guido believes declining training costs and accessible tooling may support more future innovation.

Key Arguments: Training frontier models is extremely expensive; realistic costs are not $100,000 but often millions or tens of millions once full infrastructure and reservation costs are included. Inference is dramatically cheaper than training, often costing only fractions of a cent per output, but providers must still pay for unused peak capacity. Transformer models allow rough compute estimation from parameter counts, but actual training cost depends heavily on optimization, precision, hardware utilization, and distributed systems overhead. Naive FLOP-based estimates understate true cost because training rarely achieves full GPU utilization and requires test runs and capacity buffers. Compute can create a moat for AI companies, but the moat may be more of a speed bump than a permanent barrier because well-funded startups can still enter. The industry may be approaching a data bottleneck: model size can no longer scale indefinitely without enough new training material. As chips get faster, training costs could flatten or even decline, reducing the long-term advantage of brute-force capital spending.

Data Points: GPT-3 parameters: 175 billion - Used as the example model for estimating training and inference compute requirements. Naive training FLOPs for GPT-3: ~3 x 10^23 floating point operations - Guido’s back-of-the-envelope estimate for training GPT-3. Inference FLOPs approximation: ~2x number of parameters - Rule-of-thumb estimate for transformer inference compute. Training FLOPs approximation: ~6x number of parameters - Rule-of-thumb estimate for transformer training compute. A100 rental price: $1 to $4 per hour - Guido’s rough estimate of renting a common AI accelerator. Naive GPT-3 training cost estimate: ~$500,000 - Estimate derived from FLOPs and A100 rental assumptions, before real-world inefficiencies. Real-world training cost: Millions to tens of millions of dollars - Guido’s practical estimate for training large language models in industry. Compute spend share: More than 80% of total capital raised - Claim about some AI companies’ early-stage spending on compute resources. Training data scale for modern text models: ~1 trillion tokens - Referenced as the scale of training data for a modern language model. Inference cost per output: ~0.1 cent to 0.01 cent - Rough per-token cost range mentioned for inference. Training data used for ChatGPT example: ~100 gigabytes / ~100 billion characters / ~100 million books - AI-generated factoid used in the episode’s closing segment. Llama 2 training data: 2 trillion tokens - Mentioned in the closing note as an example of larger modern models. Llama 2 character-equivalent estimate: ~8 trillion characters - Derived comparison used to emphasize the scale of training data.

Pivotal Quotes: "Training one of these large language models today, it's not a $100,000 thing. It's probably millions of dollars thing. Practically speaking, what we're seeing in industry is that it's actually more of a tens of millions of dollars thing." — Guido Appenzeller: On the real cost of training frontier AI models. "I think the expectation at the moment is that the cost of training these models may actually sort of top out or even go down a little bit as the chips get faster, but we don't discover new training material as quickly." — Guido Appenzeller: On why scaling may be constrained by data, not just compute. "It's more of a speed bump than something that prevents new entrants." — Guido Appenzeller: On whether compute-heavy incumbents have an unbreakable moat.

Implications: AI remains capital-intensive, but training costs may stabilize as hardware improves and data becomes the bottleneck. Expect continued demand for efficient infrastructure, smarter utilization, and opportunities for well-funded startups—not only incumbents.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast