Episode Summary
Executive Summary: The episode argues that open models have moved from hobbyist curiosity to enterprise default for most workloads, driven first by cost and increasingly by control, customization, and fast-improving tooling. Jeffrey Morgan explains how Olama sits at the center of token flow across cloud and local deployments, why coding agents and multimodal/flash models are accelerating adoption, and how orchestration, security, and hardware are becoming the new battlegrounds.
Main Topics: Open models becoming the enterprise default (Priority: 5/5): Morgan says enterprises are shifting heavily toward open models because they lower cost immediately and create a path to deeper customization and control. Coding agents and agentic workflows drive token growth (Priority: 5/5): The biggest usage spike comes from coding agents and agent platforms like OpenClaw/Hermes, which dramatically increase per-user token consumption and broaden adoption beyond developers. Model release cadence and the return of fine-tuning (Priority: 4/5): Open model iteration is speeding up, making custom training harder to keep current, while better tooling and enterprise demand are bringing fine-tuning back into relevance. Security, safety, and governance as adoption blockers (Priority: 5/5): For enterprises, the main barriers to open-model adoption are safety, origin concerns, and secure deployment; solving these makes Chinese-origin models viable for many customers. Curation/orchestration is the new scarcity layer (Priority: 5/5): As open tokens become abundant, the valuable problems move above the model layer: routing, memory, coordination, execution, harnessing, and reliable integration across providers. Local vs cloud inference will be hybrid (Priority: 4/5): Local models are strong for cheaper, lower-latency, simpler tasks, while cloud frontier or large open models still win on hardest coding and reasoning problems; the future is blended. Olama’s origin story and product-market-fit journey (Priority: 4/5): The company spent years pivoting from container/devtools ideas before finding the open-model opportunity, then scaled rapidly once local LLMs became real and useful.
Key Arguments: Cost is the largest short-term pain point that open models solve, but the deeper enterprise goal is control and customization. Enterprise usage is shifting toward open models, especially for coding agents and assistant-style workflows, with a mix of US and Chinese-origin models. The sharp increase in token usage is driven by agentic workflows that consume many tokens while planning, tool-using, and executing tasks. Open models now power out-of-the-box use cases, not just fine-tuned custom models; this became especially visible in 2025. The cadence of open-model releases is accelerating, which makes bespoke fine-tuning harder to maintain, but improved tooling helps. Security and governance are the main blockers to adoption of some open models, especially Chinese-origin ones, but those concerns are addressable. Model orchestration, memory, credentials, execution, and harnessing are likely to become standalone product layers and startup opportunities. A hybrid architecture is emerging: local models for simpler tasks and cloud models for hard tasks, connected by routing/orchestration. Ultra-cheap “flash” models may create an era of effectively unlimited token usage for many everyday workflows. The market is likely to settle on most tokens and most software being open, while frontier closed models remain reserved for the hardest tasks.
Data Points: Olama developer base: 9 million developers - Jeffrey Morgan described Olama’s usage scale GitHub stars: 178,000 - Olama’s GitHub popularity Fortune 500 penetration: 85% - Morgan said Olama is used by a large majority of Fortune 500 companies AT&T token shift to open models: 40% - Morgan referenced an Information article saying AT&T had shifted this share of token consumption Olama Cloud growth since start of year: 150x - Aggregate token growth in Olama’s cloud since the beginning of the year OpenClaw growth inflection: ~5x - Per-developer token usage jump associated with the OpenClaw wave in April Context window expansion: 128K to 1M+ - Open models expanded context windows, enabling more agentic use cases Open model gap vs frontier models: Less than 3 months - Morgan estimated open models were close behind frontier closed models earlier in the year Model release cadence: Three DeepSeek Flash iterations in one summer - Used to illustrate how fast open-model iteration is now moving Enterprise token mix prediction: 80-90% open-model tokens - Morgan’s forecast for the supermajority of tokens inside businesses Budget mix prediction: 10-20% of cost on open models - He argued open models may account for most tokens but a smaller share of spend Local model hardware range: 20B to 40B parameters, sometimes up to 120B - The size range Morgan said modern consumer/prosumer hardware can run DGX Spark memory: 128 GB unified memory - Described as enough to run 20B–120B models on a desktop unit
Pivotal Quotes: "Cost is by far the largest pain point that open models can jump in and solve." — Jeffrey Morgan: Explaining why enterprises are adopting open models first "I think the super majority of tokens, and this is our take, it will be open models within a business." — Jeffrey Morgan: Forecasting the enterprise steady state for model usage "If it's a stateful problem, like there's storage involved, that's something that, you know, in the end, can't go into the model." — Jeffrey Morgan: Defining the boundary between what gets absorbed into models versus external systems
Implications: Enterprises should expect a hybrid AI stack: open models for most volume, frontier models for hard edge cases, and orchestration/security layers as the main new moat. Startups can win by making model usage cheaper, safer, and easier to route and operate.
About Y Combinator Startup Podcast
We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.