Invest Like the Best with Patrick O'Shaughnessy
Invest Like the Best with Patrick O'Shaughnessy

Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480]

My guests today are Gavin Uberti and Rob Wachen, the founders of Etched. A few years ago, when they set out to build a better AI chip than the largest companies in the world, almost everyone I called told me it could not be done. They have since done it, taping out a working chip on their first atte

Featured Speakers

Rob Wachen GuestGavin Uberty Guest

Episode Summary

Executive Summary: Patrick O’Shaughnessy interviews Etched founders Gavin Uberti and Rob Wachen about building a vertically integrated AI-inference chip company. They explain why inference is the next giant compute market, how they achieved a first-pass working chip, and why low-voltage compute, cluster-scale memory, and extreme execution speed could unlock cheaper, faster token production at massive scale.

Main Topics: Why Etched was possible despite skepticism (Priority: 5/5): The founders describe early disbelief from investors and semiconductor experts, and how a combination of naivety, first-principles thinking, and relentless validation helped them push through. Etched’s product is a full inference system, not just a chip (Priority: 5/5): They emphasize that the product includes the chip, rack, boards, power delivery, interconnects, and manufacturing—because production itself is the product. Core technical bets: low-voltage inference and cluster-scale memory (Priority: 5/5): Their architecture centers on running inference at much lower voltage to avoid thermal throttling, and on increasing effective memory bandwidth across a cluster rather than optimizing a single chip in isolation. Execution philosophy: vertical integration, parallelization, and speed (Priority: 5/5): Etched accelerates development by owning more of the stack, shipping work ahead of silicon, running 24/7 schedules, and making fast decisions even under uncertainty. Recruiting elite talent through legends plus young builders (Priority: 4/5): They explain a bimodal hiring strategy: pair world-class veterans who know scale with unusually driven young engineers who assume hard problems are solvable. The capital intensity and supply-chain realities of hardware (Priority: 4/5): The conversation covers the difficulty of fundraising, the importance of supplier relationships, and how scarce fabs, memory, and power constrain deployment at scale. Big-picture future of inference, agents, and AI economics (Priority: 5/5): They argue inference will become a dominant market, with massive concurrency and lower token costs enabling longer-horizon agents, new scientific work, and a shift toward token production as a core economic activity.

Key Arguments: Inference is becoming the dominant bottleneck and opportunity in AI, not training. A full rack-scale system is necessary because performance depends on chip, board, interconnect, cooling, and manufacturing together. General-purpose chip design leaves performance on the table; inference-specific constraints allow much better optimization. Lower voltage is essential to avoid thermal throttling when packing in more flops. Cluster-level memory and low-latency interconnect matter more than single-chip memory specs for decoding workloads. Vertical integration accelerates learning and execution, reducing dependence on slow vendors. Fast decision-making and willingness to spend money early are critical in hardware because every delay costs real market opportunity. Hiring works best when combining top-domain legends with scrappy founders and engineers who will attack impossible-sounding problems. The best talent self-selects into extreme missions; skepticism filters out people who are not deeply aligned. AI agents will require far more compute, shorter wall-clock times, and eventually huge scale-up clusters to operate concurrently. Economic value will accrue to companies that produce the most tokens and own more of the token supply chain.

Data Points: Customer demand for first product: More than $1 billion - Etched says demand already exceeds a billion dollars for its first product. Capital raised: $800 million - Raised to build the product and scale the company. Company founded: 2023 - Etched was started in 2023. First-gen voltage: Under half the voltage of any other AI chip - Their low-voltage inference approach is a central technical claim. GPU model flops utilization (MFU): 20% to 50% - They cite typical GPU utilization on inference workloads as a benchmark and inefficiency. Chip-to-chip latency on NVIDIA product: About 4,000 nanoseconds - Used to motivate the need for a lower-latency custom interconnect stack. Latency improvement from custom interconnect: More than 5x lower - They claim their interconnect stack cuts latency and improves bandwidth substantially. Time from silicon back to inference in rack: 40 days - Etched says it reached running inference far faster than a competitor that took 10 months. Industry benchmark cited: 10 months - Referenced as the time another AI chip company took from silicon return to rack inference. Physical design/vendor cost: $40M–$50M - Estimated cost to enter physical design stage and sign a major vendor agreement. Available cash at a key moment: $15 million - They describe having far less cash than needed for the full rack-scale buildout. Initial fundraising target/need: $100 million - They concluded they needed around $100M to do the company properly. FPGA validation scale: Over 700 FPGAs - Used to emulate the full reticle chip and test the full inference stack before silicon returned. Chip alignment precision: Within 50 picoseconds - A critical bug fix required aligning clock signals to extreme precision. Bangalore recovery period: 4.5 months - One founder lived in Bangalore for four and a half months to accelerate bring-up and vendor coordination. High school cancer prognosis: Under 30% chance of survival - Rob Wachen recounts being told this after diagnosis. School/competition team size: 2-person team - Gavin describes winning robotics competitions with a tiny team focused only on performance. Scale of future market: Majority of global GDP / billions of concurrent agents - The founders predict inference and agentic compute could become enormous parts of the economy. Possible timeline claim: 2027 - Gavin predicts more agents than humans doing knowledge work by 2027.

Pivotal Quotes: "Production is the product." — Rob Wachen: Explaining that the company’s real mission is to manufacture token-producing systems at scale, not just design a chip. "If you're going to do it, we're going to go all the way." — Gavin Uberty: Describing the decision to pursue a full rack-scale product rather than a test-chip or incremental approach. "We have to assume it is possible." — Gavin Uberty: Summarizing the mindset used to solve hard technical and organizational problems.

Implications: If Etched is right, AI infrastructure shifts from general-purpose GPUs to specialized, vertically integrated token factories. That would lower inference costs, expand access, and make agents, long-horizon compute, and massive model concurrency economically feasible.

🔓 Sign Up for Unlimited Episode Search

About Invest Like the Best with Patrick O'Shaughnessy

Conversations with the best investors and business builders in the world.

View all episodes from Invest Like the Best with Patrick O'Shaughnessy