Invest Like the Best with Patrick O'Shaughnessy
Invest Like the Best with Patrick O'Shaughnessy

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

My guest today is Neil Movva, founder of Sail. Sail is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matte

Featured Speakers

Neil Nova Guest

Topics Discussed

Episode Summary

Executive Summary: Neil Nova argues Sale Research is building a "token factory" for the emerging era of long-running AI agents. The company optimizes the full stack—software, chips, data centers, and power—to deliver the cheapest possible tokens, betting that background, proactive workloads will dominate. The conversation covers GPU trade-offs, heterogeneous chips, data center fragmentation, energy scavenging, and the shifting economics of open vs. closed AI.

Main Topics: Sale Research’s mission: ultra-cheap intelligence (Priority: 5/5): Nova defines Sale as a token factory serving open-source and custom models at the lowest possible cost, including long-running agent virtual machines ('sale boxes') for hours-to-days workflows. Why long-running background agents matter (Priority: 5/5): He argues the market is shifting from interactive chatbots to proactive, background agents where latency matters less than throughput and cost, enabling tasks like deep research, cybersecurity, and personal assistance. GPU efficiency and the latency vs. throughput trade-off (Priority: 5/5): Nova explains the core GPU tension: low-latency serving uses hardware inefficiently, while batching and throughput are better suited to asynchronous agent workloads. Sale designs for throughput rather than instant responses. Hardware strategy: heterogeneous chips and memory bottlenecks (Priority: 5/5): He contrasts NVIDIA, AMD, Cerebras, and others, arguing each has a comparative advantage. Sale buys varied chips and tries to map them to the right parallelism, memory hierarchy, and workload. Data centers and power as an inference problem (Priority: 4/5): Nova says inference can live in small, distributed, lower-redundancy data centers and even intermittent power environments. He prefers many smaller sites over giant training-era facilities. Data, RL environments, and model self-improvement (Priority: 4/5): He argues internet data is largely exhausted and that future progress comes from verifiable RL environments where models can self-improve on hard tasks like coding, math, and security. Open source vs. closed frontier models (Priority: 4/5): Nova believes closed labs can charge a premium for being months ahead, but diffusion via open source and AI-generated artifacts makes durable exclusivity unlikely.

Key Arguments: The future of AI usage is not real-time chat but long-horizon background work, which makes throughput and cost more important than latency. Tokens are the current unit of work, but the industry is moving toward outcomes and agent-run task completion over human-managed token budgets. Open-source models have created a durable market for owned intelligence, and demand for customized, sovereign deployment is growing. NVIDIA is excellent for low-latency workloads, but that is not the dominant future use case; Sale optimizes for throughput and can use other chips where economics favor them. The key technical bottleneck is often memory, especially KV cache and HBM, not raw compute; different chips should be assigned different roles. Small, distributed, lower-uptime data centers are acceptable for background inference because workloads can fail over and are not user-blocking in the same way as chat. The internet was a one-time subsidy of data; the next leap comes from RL environments and verifiable problems where models can generate their own training signal. Open source will not disappear because AI-generated code and outputs continuously diffuse capabilities, making latent distillation unavoidable.

Data Points: Sale customers' business focus: Long-running agents for hours, days, or weeks - Sale boxes are described as long-running agent virtual machines for asynchronous workloads. Task length for agents: About 1 hour today - Nova says current agents are already capable of running for an hour at a time on some tasks. Background vs. real-time workload mix forecast: 50-50 by end of year; 90-10 background over time - His estimate of how the workload mix will shift toward background agents. OpenAI pricing for 1 trillion tokens: At least $5 million - Used to illustrate how expensive current token consumption can be. Model sparsity: Frontier models closer to 1% dense; modern MoE models often under 10% dense - Nova argues compute is already fairly efficient, with more room to improve memory use than sparse compute. Blackwell throughput utilization: 70-80% of peak utilization - He cites large matrix multiply as the GPU's happy path, limited in practice by thermals and power. NVIDIA Blackwell memory bandwidth: About 10 TB/s HBM range - Compared against Cerebras-style SRAM bandwidth. Cerebras wafer-scale SRAM bandwidth: Petabytes per second - Used to emphasize why Cerebras is strong for memory-heavy fast-access workloads. Cerebras per-wafer SRAM capacity: About 50 GB SRAM per wafer - Nova describes wafer-scale memory density and how multiple wafers can reach larger capacities. Blackwell memory capacity: 288 GB HBM - Example of high-capacity off-chip memory around the logic die. Blackwell die SRAM: About 500 MB - Contrasted with HBM to show SRAM vs. DRAM density differences. Data center power tiers: 1 MW plentiful; 10 MW near edge; 100 MW increasingly hard; gigawatt data centers very hard - He argues inference can use distributed smaller facilities. Physical size of 1 MW compute: About 8 refrigerator-sized racks - Illustrates how liquid cooling compresses power density into small footprints. NVIDIA Blackwell chip production: 5 million chips this year - Nova uses this to argue many GPUs may still sit idle or underutilized globally. Enterprise adoption lag: Many enterprises still on older model versions - He suggests enterprise deployment moves slower than frontier model progress.

Pivotal Quotes: "Sale Research is a token factory." — Neil Nova: His literal description of the company and its mission. "The best latency is no latency at all." — Neil Nova: Explaining why background, proactive agents are superior to always-waiting chatbots. "There are no bad chips, there's only bad pricing." — Neil Nova: His philosophy on using heterogeneous silicon and making economics work across vendors.

Implications: AI infrastructure is shifting from low-latency chatbot serving to distributed, cost-minimized background inference. That favors throughput, memory efficiency, flexible chip sourcing, and smaller data centers, while making open-source diffusion and RL-driven self-improvement increasingly central.

🔓 Sign Up for Unlimited Episode Search

About Invest Like the Best with Patrick O'Shaughnessy

Conversations with the best investors and business builders in the world.

View all episodes from Invest Like the Best with Patrick O'Shaughnessy