The Twenty Minute VC (20VC)
The Twenty Minute VC (20VC)

20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Andrew Feldman is the Co-Founder and CEO @ Cerebras, the fastest AI inference + training platform in the world. In Sept 2024 the company filed to go public off the back of a rumoured $1BN deal with G42 in the UAE. Andrew is the leading expert for all things inference. In Today's Episode We Disc

Featured Speakers

Andrew Feldman Guest

Episode Summary

Executive Summary: Andrew Feldman argues AI hardware is being held back by inefficient GPU inference architecture, and Cerebras’ wafer-scale design solves the bottleneck with far more on-chip SRAM, less power, and better utilization. He believes inference will become vastly larger, transformers will matter less over time, synthetic data will dominate, and the long-term value in AI may accrue more to chips/infrastructure than models.

Main Topics: Why Cerebras was founded (Priority: 5/5): Feldman explains that the team saw AI emerge as a new computing workload in 2015 and believed a purpose-built architecture could outperform general-purpose chips. Inference bottlenecks and wafer-scale architecture (Priority: 5/5): He argues inference is dominated by memory movement rather than compute, making GPUs inefficient; Cerebras uses wafer-scale silicon and SRAM to reduce data movement and power use. Performance, cost, and utilization (Priority: 5/5): Speed is critical for interactive use cases, while cheaper/faster inference expands the market. He claims current GPU inference utilization is extremely low and that better algorithms and chips will improve economics. Market size, diffusion, and AI adoption (Priority: 4/5): Feldman says AI moved from novelty to utility in late 2024, and that as it becomes faster and cheaper it will spread into many more products and everyday workflows. Transformers, scaling laws, and model evolution (Priority: 4/5): He believes dependence on transformers will decline within 3-5 years, that scaling laws still hold especially for inference, and that algorithmic improvements remain large. Competition, moats, and NVIDIA (Priority: 5/5): He challenges the idea of CUDA lock-in in inference, says NVIDIA will remain strong, but argues its GPU architecture is less suited to inference than wafer-scale competitors. Business strategy, partnerships, and geopolitics (Priority: 4/5): The discussion covers Cerebras’ cash-flow positivity, concentration risk in the G42 relationship, reasons for going public, and the complexity of export controls and AI policy.

Key Arguments: AI inference is fundamentally memory-bandwidth constrained, not compute constrained, so GPU architectures with off-chip memory are structurally inefficient for many workloads. Wafer-scale design with large amounts of SRAM can keep more of the model on chip, reducing data movement, latency, and power consumption. Current GPU inference is highly underutilized; improving hardware, algorithms, and data centers can materially lower cost per token. Interactive AI products need low latency because milliseconds affect user attention and product viability. AI is still early across compute, data, and algorithms; the industry has substantial room for improvement despite common claims of maturity. Synthetic data will become the dominant training source because it can target rare, high-value scenarios that are hard to collect from the real world. Transformers are important now but not permanent; future architectures will likely reduce dependence on them. In the long run, chips and infrastructure may capture more enterprise value than model providers, though model companies can command high short-term valuations. CUDA lock-in is not a real barrier for inference because models can be moved across providers with relatively little friction. Market share, ecosystem defaults, and scale remain real moats for NVIDIA even if the architecture is imperfect for inference.

Data Points: GPU utilization during inference: 5-7% utilized - Feldman says most GPU inference capacity is idle most of the time, implying 93-95% waste. Underutilized GPU capacity: 93-95% wasted - Derived from his statement that GPUs are only 5-7% utilized in inference workloads. Model size example: 70 billion parameters - Used to illustrate the memory bandwidth burden of generative inference. Data moved per token for a 70B model: ~140 GB - He said all weights must be moved to generate one word, using a 70B model with 16-bit weights. Launch date for Cerebras inference: August 26, 2024 - He said their inference architecture has been the fastest across tested models since launch. Potential chip count on traditional architectures: 4,000-8,000 chips - Example of how many chips may be required for very large models on non-wafer-scale systems. AI adoption growth expectation: 100x+ - Feldman said the AI market is likely way over 100 times bigger in the future. NVIDIA share forecast in five years: ~60% - He predicted NVIDIA would remain a major player but lose share over time. G42 revenue concentration: 87% of revenue - He disclosed the concentration of Cerebras revenue tied to the G42 relationship. Cerebras customer/partner announcement: Rumored $1 billion deal - Referenced in the intro as part of the company’s UAE/G42 relationship. Coda customers: 50,000 teams - Sponsor segment, not core to the interview content. Pleo customers: 37,000 companies - Sponsor segment, not core to the interview content. Rome users: 500+ companies - Sponsor segment, not core to the interview content. Public market timing reference: September 2024 - The intro notes Cerebras filed to go public around this time.

Pivotal Quotes: "The fundamental architecture of the GPU with off-chip memory is not great for inference." — Andrew Feldman: He explains why Cerebras sees an opening against NVIDIA in inference. "Our AI algorithms today are not particularly efficient. In a GPU, most of the time it's doing inference, it's five or seven percent utilized." — Andrew Feldman: He argues large gains are still available through better hardware/software efficiency. "We won't be as dependent on transformers in three years or five years as we are now. 100 percent." — Andrew Feldman: He predicts a shift away from transformers as AI architectures evolve.

Implications: AI growth will likely be driven by cheaper, faster inference, not just bigger models. Hardware innovation, especially around memory and latency, may reshape the value chain, while synthetic data, new architectures, and infrastructure buildout become central competitive advantages.

🔓 Sign Up for Unlimited Episode Search

About The Twenty Minute VC (20VC)

View all episodes from The Twenty Minute VC (20VC)