This Week in Startups
This Week in Startups

Orchestrating Smarter AI Systems with AI21 Labs’ Yoav Shoham | AI Basics with Google Cloud

In this episode of AI Basics, Jason sits down with Yoav Shoham — Stanford professor emeritus and co-founder of AI21 Labs, creators of Jurassic-2, Wordtune, and the new orchestration system Maestro. They unpack: * Why enterprise AI struggles with reliability * What orchestration really means (and why

Featured Speakers

Jason Calacanis HostYoab Shoham Guest

Topics Discussed

Episode Summary

Executive Summary: This episode of This Week in Startups explores the real state of AI in 2025 with AI21 Labs co-founder Yoav Shoham, focusing on why enterprise adoption still lags despite enthusiasm. The discussion centers on reliability, orchestration, agents, model architecture tradeoffs, agent-to-agent interoperability, and where founders should focus. The core message: AI’s biggest opportunity is not flashy demos, but dependable workflows, domain-specific systems, and education tools that actually work.

Main Topics: Enterprise AI adoption is constrained by reliability (Priority: 5/5): Shoham argues that while experimentation is widespread, deployments remain limited because probabilistic models still fail too often for mission-critical use. Enterprises need systems that are reliable, validated, and testable step by step. Orchestration and planning are more important than raw model output (Priority: 5/5): AI21’s Maestro emphasizes putting logic outside the model to coordinate tools, databases, code, and multiple LLMs. The focus is on explicit planning, validation, and system-level control rather than trusting a single prompt-response call. Agent hype vs practical automation (Priority: 4/5): The conversation distinguishes between true proactive agents and simple automation wrapped in AI branding. Shoham calls this "agent washing" and says current wins are mostly in mundane, repetitive tasks, while complex multi-step agents remain immature. Large vs small models and hybrid architectures (Priority: 5/5): Large general-purpose models remain best for consumer chat and broad coverage, but smaller or hybrid models can be faster, cheaper, and more suitable for enterprise use. Shoham highlights AI21’s Jamba family as a state-space/transformer hybrid built for efficiency and long-context handling. Vertical AI and domain specialization (Priority: 4/5): The transcript explores whether focused models for legal, accounting, or other verticals improve accuracy. Shoham says yes, provided the model retains enough general language and common-sense capability while adapting to domain-specific knowledge. Agent-to-agent protocols are promising but early (Priority: 4/5): Shoham says interoperability standards like A2A are too early for production use at scale. He warns that shared syntax is not enough; systems also need shared semantics and aligned incentives, especially across organizational boundaries. Underhyped opportunity: education and adaptive tutoring (Priority: 5/5): Shoham argues education is underserved and AI could transform it through adaptive learning, personalized tutoring, and proactive programming. The hosts agree that AI can make high-quality instruction accessible and individualized at scale.

Key Arguments: Enterprise AI adoption is limited less by lack of interest and more by unreliability; a 95% good / 5% catastrophic failure rate is unacceptable in accounting, support, and code deployment. Reliability requires orchestration outside the model: explicit plans, validation steps, tool use, database access, code execution, and sometimes judge models or deterministic checks like counting words. The term "agent" is being overused; true agents are ongoing, proactive, multi-tool systems, while many current products are just conventional automation with an LLM layer. Simple repetitive workflows are already a strong use case for AI, but more complex autonomous agents still need substantial work before they can operate safely at scale. Large models are still the best default for broad consumer chat because the input variety is too large to pre-specify, but enterprise use cases often justify narrower, cheaper, lower-latency models. Hybrid architectures like Jamba aim to solve the context-length and efficiency problems of transformers, since transformer attention becomes expensive at very large sequence lengths. Vertical specialization can improve fidelity, but models must preserve foundational language competence and common sense while learning domain-specific concepts. Agent-to-agent communication needs more than JSON syntax; it needs shared semantics and incentives, otherwise inter-company coordination will fail. Education is a major overlooked opportunity because AI can deliver adaptive, personalized tutoring, better sequencing, and more interactive instruction than static online courses. The biggest near-term value for startups is not flashy demos but boring reliability work that reduces failure in real deployments. key metrics that matter in AI are latency, memory footprint, context length, and deployment-to-experiment ratios rather than just benchmark scores. data_points like online education and tutoring show that a single expert course can scale globally, and AI can make that instruction more accessible and personalized. AI can turn educational content into a personalized coach instantly, reducing the cost and effort of finding human tutors.

Data Points: AI21 Labs age: about 7 years old - Shoham describes AI21 Labs as a mature AI company with several years of model-building experience. Enterprise experimentation-to-deployment ratio: 10:1 to 20:1 - Shoham says enterprises run far more AI experiments than actual deployments because of reliability concerns. Transformer paper era: 2017 - He references the original Google Transformer paper as the breakthrough that changed language modeling. Context length scale discussed: from 1,000 to 1,000,000 - Shoham contrasts earlier transformer context sizes with the million-token-scale future and why quadratic cost breaks down. Complexity of transformer attention: quadratic - He explains that transformer attention scales quadratically with context length, creating efficiency problems. Complexity of state-space models: linear - He describes state-space models as inherently linear and therefore better suited to long contexts. Online course reach: over 1 million people - Shoham says his online game theory course has been viewed by more than a million people. AI in education cost: free to near-zero marginal cost - The hosts note that AI tutoring and online courses can be accessed widely at little or no cost today.

Pivotal Quotes: "If you're brilliant 95% of the time, and not just wrong, but total garbage 5% of the time, that may be okay in consumer land, but not in the enterprise." — Yoab Shoham: On why reliability is the main barrier to enterprise AI deployment. "I call this the agent washing." — Yoab Shoham: On the overuse of the term agent for simple automation or lightly AI-assisted workflows. "The biggest blocker in the enterprise right now is getting the workflows to be reliable and customizing them per deployment." — Yoab Shoham: On where startups should focus if they want durable enterprise value.

Implications: For founders, the opportunity is in dependable, domain-specific AI systems, not hype. Build workflows that are reliable, testable, and efficient, and look seriously at education, where AI can unlock large-scale personalization and tutoring.

🔓 Sign Up for Unlimited Episode Search

About This Week in Startups

Jason Calacanis covers startups, tech, markets, media, and all the hottest topics in business and technology. He also interviews the world’s greatest founders, operators, investors, and innovators.

View all episodes from This Week in Startups