No Priors
No Priors

Context windows, computer constraints, and energy consumption with Sarah and Elad

This week on No Priors hosts, Sarah and Elad are catching up on the latest AI news. They discuss the recent developments in AI like Meta’s new AI assistant and the latest in music generation, and if you’re interested in generative AI music, stay tuned for next week’s interview! Sarah and Elad also g

Topics Discussed

Episode Summary

Executive Summary: The episode surveys major AI frontiers: music generation, local on-device models, platform dynamics, long-context models, and the scaling bottlenecks of compute, data, and energy. The hosts argue that AI is moving through new content formats and product surfaces while a few hyperscalers, startups, and sovereigns race to fund ever-larger systems, even as practical limits push more inference to devices and renew interest in power infrastructure.

Main Topics: AI music generation and new creative formats (Priority: 5/5): Suno and Udio are framed as early but important examples of AI expanding from text and images into music, with speculation about personalized soundtracks, lyric control, vocals, and future voice cloning. Apple’s local LLM release and on-device AI (Priority: 5/5): The discussion highlights Apple’s small models on Hugging Face and the market demand for 1B–3B parameter models that can run on edge devices, enabling lower-latency and more persistent experiences. Platform vs. standalone AI products (Priority: 4/5): The hosts debate whether AI apps built on top of operating systems or platforms can survive independently, using examples like Salesforce, Microsoft Office, and the risk of platform subsumption. Model scaling, open source, and compute competition (Priority: 4/5): They discuss how open-source models, hyperscalers, Snowflake, Databricks, and other infrastructure players are converging into a blended competitive landscape where owning the model may matter less than training, fine-tuning, and deployment expertise. Long context windows as a major capability shift (Priority: 5/5): Long-context models are presented as transformative for code, legal, support, and biology use cases, with the belief that context windows will keep expanding and meaningfully change how prompts and workflows are designed. Energy, data centers, and physical constraints on AI (Priority: 5/5): The conversation emphasizes that AI scaling is no longer just a software problem: power, permitting, data centers, and grid capacity may become binding constraints, making nuclear and infrastructure investment strategically important.

Key Arguments: AI is evolving through successive content modalities—text, images, chat, video, and now music—creating new creative behaviors and products rather than just improving old ones. Making music generation easier may increase the number of creators and change what people want to create, not just the quality of output. Local models on devices matter because they reduce latency, lower inference cost, and enable always-on experiences like indexing a user’s computer or proactive assistant behavior. Many successful platform-native products are eventually absorbed by the platform owner, so AI apps must either become indispensable or leverage broader browser/web footprints to survive. The key unresolved technical question for small models is which capabilities can truly fit on-device versus which require cloud inference, and that boundary will shape future apps. Long context windows are becoming a major capability unlock for domains with huge inputs, such as codebases, legal documents, customer support queues, and biology/protein modeling. Open-source and hosted model ecosystems are blending, making it less obvious that data/compute platforms need to own frontier models if they can instead train, fine-tune, and deploy models for customers. AI scaling is increasingly limited by physical realities—compute, packaging, data centers, power delivery, and energy supply—so progress depends on infrastructure as much as algorithms. Large firms and sovereigns may dominate the frontier because they can fund massive compute spend, while smaller players may specialize in mid- or low-scale models and inference products.

Data Points: Music generation platforms mentioned: Suno and Udio - Examples of fast-growing AI music models that are gaining adoption Model size demand on device: 1B–3B parameters - Range cited for useful local models that can fit on edge devices Magic context window: 5 million tokens - Referenced as an early long-context release Gemini 1.5 context window: 1 million tokens - Used as a comparison point for long-context progress Future context window expectation: 10 million+ tokens - Predicted range for the next year or two Meta GPU scale: 22,000 GPU clusters / 350,000 GPUs - Used to illustrate the scale of Meta’s training infrastructure Meta AI spend: $30–35 billion this year - Approximate annual GPU spend cited for Meta Hyperscaler AI compute spend: nearly $200 billion this year - Aggregate estimate of AI compute investment by major players Oil majors annual spend: $80 billion per year - Comparison point for infrastructure-scale capital expenditure Broadband infrastructure spend: about $100 billion per year - Another historical capex comparison Salesforce market cap: $266 billion - Used to compare Salesforce’s scale to Veeva Veeva scale: about 20% of Salesforce - Rough relative valuation estimate mentioned Nuclear power share in the U.S.: 17–18% - Current U.S. nuclear generation share cited Nuclear power share in France: 70% - Example of a high-nuclear-power grid Nuclear power share in Japan: 30% - Another example of substantial nuclear usage

Pivotal Quotes: "Everybody should have a personalized soundtrack for their life, but in the voice of Taylor Swift, in the style of Taylor Swift." — Speaker: Discussion of AI music generation and personalized creative content "Apple has entered the chat with a release of relatively small models that are now on Hugging Face and such." — Speaker: Commentary on Apple’s move into local/on-device models "The fear is that some of the potential limits to progress are going to be like, we can't just throw more compute at the problem because it's like physically hard to throw more compute at the problem, energy data centers, right?" — Speaker: Explanation of why energy and data centers may become binding AI constraints

Implications: AI adoption is broadening into creative tools, on-device assistants, and infrastructure-heavy model training. Winners will likely combine distribution, efficient models, and physical scale, while energy and compute availability become strategic bottlenecks.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors