This Week in Startups
This Week in Startups

How open-source & distributed models can win AI with MosaicML’s Naveen Rao | E1754

This Week in Startups is presented by: Vanta. Compliance and security shouldn't be a deal-breaker for startups to win new business. Vanta makes it easy for companies to get a SOC 2 report fast. TWiST listeners can get $1,000 off for a limited time at vanta.com/twist. Trovata. Starting up is har

Featured Speakers

Jason Calacanis HostNaveen Rao Guest

Topics Discussed

Episode Summary

Executive Summary: This episode argues that AI is a major technological inflection point, but its value depends on who controls the data, models, and infrastructure. MosaicML’s Naveen Rao explains how organizations can build proprietary, portable models using their own data via prompting, fine-tuning, or pre-training, while the hosts debate open source vs. closed systems, GPU scarcity, and AI’s potential to disrupt jobs, education, and society faster than past tech waves.

Main Topics: AI as the next major human inflection point (Priority: 5/5): The conversation frames AI as comparable to language as a civilization-level technology that expands human capability and could reshape every product and workflow. Owning data, models, and IP (Priority: 5/5): A central theme is whether companies should rely on APIs from big labs or build their own models so their proprietary data and competitive advantage stay under their control. How to adapt company data to language models (Priority: 5/5): Naveen explains practical paths—prompting for small data, fine-tuning for medium-scale behavior shifts, and pre-training for very large datasets—using the startup investor example as a case study. Open source vs. closed AI ecosystems (Priority: 4/5): The discussion weighs whether open source communities or centralized model labs will dominate, with a strong argument that distributed, open capabilities are healthier for innovation and competition. GPU scarcity and infrastructure bottlenecks (Priority: 4/5): The episode details the shortage of training/inference compute, why demand for GPUs has surged, and how supply constraints in semiconductors and memory packaging limit AI scaling. Employment, productivity, and social disruption (Priority: 5/5): The hosts debate whether AI will raise output enough to create new demand or whether it will displace mid-skill workers too quickly for society to absorb the change. Education and personalized learning (Priority: 4/5): AI is presented as a tutor and learning accelerator that can personalize instruction, help students use tools ethically, and shift education away from memorization toward problem-solving.

Key Arguments: Organizations should own and control their own AI models when proprietary data is core to their business, because using a shared API gives competitors the same advantage. The most practical way to customize a model depends on data size: prompts for small corpora, fine-tuning for moderate datasets, and pre-training for very large datasets. Open source AI can outcompete centralized labs because many independent contributors can improve models faster than a single organization can. AI training and inference are compute-constrained; the GPU shortage is real and likely to persist because supply chain expansion takes years. General-purpose models will coexist with specialized vertical models; specific domains like healthcare, investing, and customer support need expert systems rather than jack-of-all-trades tools. AI could trigger a fast productivity shock that displaces mid-level workers before the economy creates equivalent new demand, making the pace of change the main societal risk. Education should adapt to AI by emphasizing tool use, creativity, and personalized problem solving rather than rote memorization.

Data Points: MosaicML model size: 7B parameters - Naveen says MosaicML released a 7-billion-parameter open-source model. Training data used for MosaicML model: 1 trillion tokens - He describes the training run for their 7B model. Approximate word equivalent: 750 billion words - He translates 1 trillion tokens into words. Training duration: 9.5 days - Time required to train the 7B model. GPU count for training: 440 NVIDIA A100 GPUs - Compute used for the 7B model training run. Training cost: $200,000 - Estimated cost to build the 7B model from scratch one time. Fine-tuning / prompt regime cost: about $100 to $1,000 - Estimated cost range for working with hundreds of thousands to tens of millions of words. Prompt window size: 64K tokens - MosaicML’s MPT-7B was tuned for a very long context window. Large-scale deployment need: 400 to 500 GPUs - Naveen says a 7B model now needs on the order of 400–500 GPUs. Future routine deployment scale: 1,000 GPUs - He anticipates this becoming routine for customer workloads. GPU supply-chain fab build time: 2 to 3 years - Time required to stand up a state-of-the-art fab. Fab investment: on the order of $10 billion - Approximate capital required to build a state-of-the-art semiconductor fab. Company size: 60–70 people - The host references MosaicML’s team size when discussing hiring and productivity. Microsoft for Startups credits: up to $150,000 - Microsoft’s founders hub offering to startups. OpenAI credits via Microsoft: up to $2,500 - Additional benefit offered through Microsoft for Startups. Vanta discount: $1,000 off - Promotional offer for Twist listeners. Trovada premium discount: 30% off one full year - Promo code Twist for premium features.

Pivotal Quotes: "AI is going to be that next inflection point." — Jason Calacanis: Opening framing of AI as a civilization-scale technology comparable to language. "I see success as people that disagree with me being able to build models equally as good as me." — Naveen Rao: Explaining his pro-access, pro-distribution philosophy for machine learning tools. "What worries me is if I make the 50th percentile player 30% more efficient across the board, the change in demand won't be as fast as the change in supply." — Jason Calacanis: Discussion of job displacement risk and the mismatch between productivity gains and labor market adjustment.

Implications: Listeners should expect AI to fragment into general and verticalized models, with strong pressure to own data and workflows. The biggest near-term constraints are compute, not ideas, while the biggest long-term risk is social adaptation to rapid productivity shocks.

🔓 Sign Up for Unlimited Episode Search

About This Week in Startups

Jason Calacanis covers startups, tech, markets, media, and all the hottest topics in business and technology. He also interviews the world’s greatest founders, operators, investors, and innovators.

View all episodes from This Week in Startups