Odd Lots
Odd Lots

Two Veteran Chip Builders Have a Plan to Take On Nvidia

When it comes to chips for artificial intelligence, obviously the name that automatically comes to mind is Nvidia. The company is making a fortune selling semiconductors used for hot AI applications like large language models, and stock investors have rewarded it handsomely for doing so. But of cour

Featured Speakers

Bloomberg Host

Topics Discussed

Episode Summary

Executive Summary: This episode explains how AI chips are designed and why startups like MadX are betting on specialized hardware for large language models. The guests break down the chip design pipeline, the economics of semiconductor development, NVIDIA’s software and hardware moats, and why LLM-focused chips may win on flops per dollar by stripping out features for other workloads.

Main Topics: How chip design actually works (Priority: 5/5): The guests walk through the end-to-end process: architecture, microarchitecture, logic design in Verilog, physical design, manufacturing, bring-up, verification, and eventual customer delivery. Why placement and wiring matter (Priority: 4/5): They explain that chip performance depends heavily on wire length, rise time, and layout efficiency, since wires have not scaled down as fast as transistors. Google TPUs and the origin of specialization (Priority: 5/5): They discuss how Google built TPUs to reduce dependence on general-purpose GPUs for internal AI workloads, especially matrix multiplication-heavy neural nets and inference. MadX’s LLM-specific strategy (Priority: 5/5): MadX is building chips, racks, and clusters optimized narrowly for large language models, betting that specializing for large matrices will beat general-purpose designs on efficiency. NVIDIA’s moat: hardware plus CUDA (Priority: 5/5): The conversation emphasizes that NVIDIA’s advantage is not just hardware quality but also CUDA software lock-in, which both attracts users and constrains NVIDIA’s own ability to deviate from compatibility. Economics of AI scaling and flops per dollar (Priority: 4/5): The guests argue that AI progress is driven by scaling laws, meaning bigger models and more data require more compute, making flops per dollar the key procurement metric. Supply chain, capex, and timing risks (Priority: 4/5): They note that chip development is expensive and slow, depends on specialized vendors and packaging capacity, and requires signaling future demand to the semiconductor supply chain.

Key Arguments: Chip design is a years-long, multi-disciplinary process involving architects, microarchitects, logic designers, physical designers, verification teams, and manufacturing partners, not just CAD drawing. Placement matters because wire length and connectivity affect performance, power, and chip density; wires have not improved as fast as transistors, increasing the importance of layout. Google built TPUs because internal AI workloads were getting too expensive on GPUs, and specialized hardware could lower inference and training costs. MadX believes modern LLMs are narrow enough that a purpose-built design can sacrifice flexibility for much higher efficiency on large matrix workloads. NVIDIA’s moat comes from both superior execution and CUDA, but CUDA compatibility forces hardware trade-offs that prevent full optimization for only LLMs. The core economic metric for AI chip buyers is flops per dollar, because compute budgets dominate and more compute generally translates into better models. Scaling laws suggest that bigger models and more training compute continue to improve performance, so there is still room to grow before hitting a hard ceiling. Chip companies must balance throughput, latency, memory bandwidth, interconnect, power, and cost, and different products can optimize for different points on that trade-off curve. Chip startups can rely on an ecosystem of EDA tools, ASIC vendors, and manufacturing specialists, but the process still requires substantial capital and patience. Demand signaling matters because underestimating future demand can create bottlenecks in packaging and supply-chain capacity, as seen with AI chips and prior semiconductor shortages.

Data Points: Audio report length: 5 minutes or less - Description of Bloomberg’s Stock Movers podcast promo Chip design cycle: 3 to 5 years - Time from conception to shipping to customers Team size in chip design: 30 to many thousands of people - Typical range for chip design teams Verification cost per mistake: $20 million to $30 million - Potential manufacturing cost of letting an error through to tape-out Bring-up timeline: 2 to 3 years after starting - When first chips are received back from manufacturing Post-bring-up customer delivery: 6 to 12 months or more - Additional time after bring-up before customers receive usable product Google Search queries: about 100,000 per second - Illustrative internal workload that makes inference costs significant Training cost growth: from $100,000 or $1 million to tens or hundreds of millions - As model size increases, training cost rises sharply Project start at MadX: beginning of 2024 - When the company says it seriously began work NVIDIA Blackwell FP4 performance: 10 petaflops - Benchmark cited for current NVIDIA chip performance NVIDIA Blackwell price: $30,000 to $50,000 - Ballpark selling price mentioned for NVIDIA’s chip Performance improvement over Hopper: 2x to 4x - Estimated gain of Blackwell over previous NVIDIA generation AI labs spending on compute: hundreds of millions of dollars - Examples of customer compute budgets discussed AGI timeline guess: approximately zero years / bluntly zero - Reiner Pope’s answer when asked how close AGI is

Pivotal Quotes: "It’s coding, but on hard mode." — Mike Gunter: Explaining why hardware design is like software development but with far higher stakes and long feedback loops "If you really want to make the best use of the real estate, you should just focus on the thing you care about most and hope that there’s a big market there." — Reiner Pope: Describing MadX’s narrow focus on LLM workloads "The thing about a moat is not only does it, in some sense, keep other people out, it also keeps you in." — Reiner Pope: Explaining how CUDA compatibility both protects NVIDIA and constrains its product choices

Implications: The episode suggests AI hardware is entering a specialization wave: buyers will increasingly prioritize efficiency over generality, and incumbents like NVIDIA may face targeted competition where CUDA compatibility is less essential. It also implies massive, sustained compute demand will continue to shape semiconductor investment.

🔓 Sign Up for Unlimited Episode Search

About Odd Lots

Bloomberg's Joe Weisenthal and Tracy Alloway analyze the weird patterns, the complex issues and the newest market crazes. Join the conversation every Tuesday and Thursday for interviews with the most interesting minds in finance, economics and markets.

View all episodes from Odd Lots