No Priors
No Priors

The marketplace for AI compute with Jared Quincy Davis from Foundry

In this episode of No Priors, hosts Sarah and Elad are joined by Jared Quincy Davis, former DeepMind researcher and the Founder and CEO of Foundry, a new AI cloud computing service provider. They discuss the research problems that led him to starting Foundry, the current state of GPU cloud utilizati

Featured Speakers

Jared Quincy Davis Guest

Topics Discussed

Episode Summary

Executive Summary: Jared Quincy Davis explains Foundry’s thesis: AI progress is increasingly constrained by compute orchestration, reliability, and cloud economics rather than just model ingenuity. He argues today’s AI cloud behaves more like colocated parking lots than elastic cloud, and Foundry aims to unlock underused GPU capacity with spot-like flexibility, resiliency tooling, and compound AI system design.

Main Topics: Foundry’s origin and mission (Priority: 5/5): Jared traces Foundry to lessons from AlphaFold 2 and ChatGPT: small teams can create outsized breakthroughs when they have access to massive compute leverage. Foundry’s goal is to democratize that leverage through a public cloud built from first principles for AI workloads. GPU utilization, failure, and resiliency (Priority: 5/5): He argues GPU clusters are far less reliable and utilized than people assume because modern systems are complex, failure-prone, and require healing buffers. Foundry’s Mars tooling and spot mechanisms are designed to keep clusters running despite hardware failures. Why AI cloud is not yet real cloud (Priority: 5/5): Davis says current AI infrastructure looks more like co-location or parking-lot economics than true cloud because customers must make long reservations and cannot elastically scale capacity up and down. This creates risk and capital inefficiency. Spot usability and capacity unlocking (Priority: 4/5): Foundry’s new spot offering tries to let one customer use another customer’s reserved capacity without sacrificing convenience or reliability. The idea is to increase effective supply, lower prices, and improve economics for both sides. GPU market concentration and apparent shortages (Priority: 4/5): The discussion highlights that major clouds own only a tiny fraction of global GPU compute, while large parts of the world’s compute exist outside hyperscalers in mining, hobbyist, and enterprise clusters. Apparent shortages are driven by interconnect, power, and cluster-scale constraints. Shift toward compound AI systems (Priority: 5/5): Jared argues the future of AI will increasingly rely on networks of model calls, verifiers, synthetic data generation, and ensemble systems rather than only giant monolithic pretraining runs. Verifiable tasks like code, math, and engineering are especially suited to this paradigm.

Key Arguments: Small teams appear magical only because they sit atop huge compute and infrastructure leverage; Foundry’s goal is to make that leverage broadly available. Modern GPU systems are highly complex physical systems, not simple chips, so failures are common and utilization is often much lower than expected. AI cloud today lacks true elasticity: customers often need fixed, long-term reservations instead of being able to scale dynamically. Spot-style scheduling for GPUs can unlock stranded capacity if reliability and user experience are automated well. The largest AI clusters are constrained by power, space, and interconnect, so simply adding more chips is not enough. The next frontier of AI performance is likely compound systems: many model calls, verifiers, synthetic data, and ensembles that work well on parallelizable, verifiable tasks. Economic optimization of AI should be lifecycle-based, balancing training cost with inference efficiency rather than optimizing pretraining alone.

Data Points: AlphaFold 2 team size: 3 people initially, later about 18 - Used to illustrate that breakthrough AI/biology achievements can come from small teams OpenAI headcount at ChatGPT launch: about 400 people - Cited to show ChatGPT was also produced by a relatively small team OpenAI compute spend: $13 billion worth of compute - Used to show the apparent small-team story still depended on massive infrastructure Foundry economics improvement: 12x to 20x - Claimed improvement versus lower-tech GPU clouds and existing public clouds Typical training utilization: sub 80%; sometimes less than 50% - Even sophisticated pretraining clusters often underutilize GPUs because of failures and buffers Healing buffer size: 10% to 20% minimum - Teams reserve spare GPUs to replace failed nodes and keep training jobs alive GPU system weight: 70 to 80 pounds - Described to emphasize that modern H100 systems are complex integrated machines Individual components in a system: 35,000+ - Approximate component count for an H100 system Cloud launch timeline: AWS started in 2003; S3 launched in March 2006; EC2 later in 2006 - Used to explain the evolution of cloud computing AWS run rate in 2015: around $3 billion - Shown as evidence cloud was already meaningful before becoming massive Spot price example: $12/hour on-demand; $4/hour reserved; $3K/month equivalent - Parking-lot analogy for current AI cloud pricing and reservation economics GPT-3 training cluster: 10,000 V100 GPUs for 14.6 days - Public example of large-scale training compute in Azure Ethereum GPU-equivalent peak: 10 to 20 million V100 equivalents - Used to illustrate how much GPU-like compute existed outside hyperscalers iPhone 15 Pro FP16 compute: about 35 TFLOPs - Compared against a V100 to show compute is broadly distributed

Pivotal Quotes: "The cloud as we currently know it is arguably one of the most important business categories in the world." — Jared Quincy Davis: Introduces the discussion of why AI infrastructure economics matter so much "AI Cloud today is not cloud in the originally intended sense by any means." — Jared Quincy Davis: Core thesis that current GPU infrastructure lacks the elasticity and abstraction of true cloud "We think this type of approach kind of points towards maybe a very different paradigm for getting better performance than just kind of scaling the models and doing a whole new pre-training from scratch." — Jared Quincy Davis: Describes the compound AI systems direction and its implications for future model development

Implications: AI infrastructure winners will be those that can turn unreliable, reserved GPU capacity into elastic, verifiable, high-utilization compute. The next wave of AI may depend more on orchestration, spot markets, and compound systems than on ever-larger monolithic pretraining runs.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors