Episode Summary
Executive Summary: Odd Lots interviews Cerebras CEO Andrew Feldman about the company’s giant-wafer AI chips, arguing that wafer-scale design dramatically reduces memory latency and makes inference much faster and cheaper than GPUs. The conversation covers the technical breakthrough, customer demand, supply-chain constraints, open vs. closed AI models, CUDA’s fading moat, and how export controls and data-center scarcity shape the industry.
Main Topics: Wafer-scale chip architecture (Priority: 5/5): Feldman explains why building an enormous chip on one wafer reduces distance between compute and memory, enabling much faster processing than modular GPU clusters for certain AI workloads. Inference economics and speed (Priority: 5/5): The discussion centers on why inference demand is exploding and why speed matters economically for both answer-based and agentic AI use cases. Supply-chain and manufacturing constraints (Priority: 4/5): Cerebras says it avoids several major bottlenecks—HBM, CoWoS, and 3nm capacity—but is still limited by data-center availability and TSMC allocation. Open-source vs. closed-source models (Priority: 4/5): Feldman argues open-source models are cheaper per unit of intelligence, while closed-source models are slightly better but more expensive; he sees a multi-player market persisting. CUDA, NVIDIA, and competitive moats (Priority: 4/5): He contends CUDA matters less now, especially for inference, and notes frontier models are increasingly trained and served outside the CUDA stack. Geopolitics, export controls, and global AI markets (Priority: 3/5): The interview covers U.S. export restrictions, the role of Chinese access, and how international customers like G42 fit into Cerebras’s business and policy landscape. IPO, financing, and long-term hardware investing (Priority: 3/5): The founders discuss the IPO, the company’s long engineering timeline, capital intensity, and the challenge of balancing quarterly public-market pressure with R&D-heavy hardware development.
Key Arguments: Large chips can process more information in less time because they reduce latency between memory and compute. Wafer-scale design enables use of faster memory, which is critical for high-speed inference. Inference demand is growing rapidly because AI is now useful at scale, and speed matters in both interactive and agentic workflows. Cerebras claims it is materially faster than GPUs and cheaper for fast tokens because GPUs become expensive and power-hungry at high speed. The company has sidestepped key industry bottlenecks by not relying on HBM, CoWoS, or 3nm TSMC capacity. Data-center real estate and power, not chip fabrication alone, are the immediate bottlenecks to growth. Open-source models are cheaper per intelligence unit, but closed-source models retain a quality premium. CUDA’s strategic importance is declining as frontier models increasingly run on non-CUDA stacks. Export controls are increasingly important and should be applied thoughtfully, even if that means forfeiting some markets. The AI market is likely to remain pluralistic, with multiple major model providers and architectures rather than a single dominant winner.
Data Points: IPO proceeds: about $5.5 billion - Referenced as part of Cerebras’s large IPO, described as raising capital at a very rich valuation. Market valuation: $64 billion - Bloomberg hosts note the company’s market value in early IPO trading. Valuation multiple: ~67x forward sales/earnings basis - Used to characterize the IPO as highly priced for a not-yet-profitable company. Chip size: 58x larger than any other chip previously built - Feldman describes Cerebras’s wafer-scale chip as unprecedented in size. Speed advantage: 15x faster than the fastest GPU - Cerebras’s claimed performance advantage for fast tokens/inference. Extreme benchmark advantage: 50x, 100x, even 1,000x faster - Feldman says some problems show much larger speedups versus GPUs. Development cost and time: 5 years and about $500 million - Time and money needed to deliver the first wafer-scale system. OpenAI deal: north of $20 billion - A major contract signed in December, cited as one of Silicon Valley’s largest. AWS deal timing: March - AWS agreement to deploy Cerebras systems in its data centers was signed in March. Closed vs. open model quality gap: 3% to 5% difference - Feldman’s estimate of the performance gap between closed-source and open-source models. Open-source performance example: Kimi K2, 1 trillion parameters - He cites an open-source model available on Cerebras’s cloud. Customer concentration: G42 was a really important chunk of business last year - Discussion of Cerebras’s Abu Dhabi customer and investor. Frontier models using no CUDA: 2 of 3 - Feldman says Gemini and Anthropic no longer rely on CUDA, while GPT still does. Market share shift: 70% - He claims CUDA lost 70% market share among leading frontier models. US fab build cost: $30-40 billion - Estimated cost to build advanced semiconductor fabs in the U.S. Fab build time: 5-6 years - Estimated timeline for constructing a leading-edge fab. Explosive demand timeframe: 2025-2026 - Feldman says inference demand exploded in 2025 and continued through 2026. Employee wealth creation: 800+ millionaires - He says the IPO created more than 800 millionaires among employees and stakeholders.
Pivotal Quotes: "Large chips process more information in less time." — Andrew Feldman: His core explanation for why wafer-scale architecture improves AI inference performance. "Speed matters equally in both." — Andrew Feldman: His answer to whether speed matters for both answer-based and agentic inference. "The best work is still ahead of us." — Andrew Feldman: He emphasizes that innovation and product development will continue after the IPO.
Implications: The episode suggests AI infrastructure will increasingly be judged on latency, power efficiency, and data-center access, not just model quality. It also points to a fragmented future with multiple model types, growing inference specialization, and more geopolitical scrutiny of semiconductor supply chains.
About Odd Lots
Bloomberg's Joe Weisenthal and Tracy Alloway analyze the weird patterns, the complex issues and the newest market crazes. Join the conversation every Tuesday and Thursday for interviews with the most interesting minds in finance, economics and markets.