The a16z Podcast
The a16z Podcast

When Will AI Hit the Enterprise? Ben Horowitz and Ali Ghodsi Discuss

Today’s episode continues our coverage from a16z’s recent AI Revolution event. You’ll hear directly from a16z cofounder Ben Horowitz and Databricks cofounder and CEO, Ali Ghodsi as they answer questions around AI and the enterprise, plus their perspectives on open source, whether benchmarks are BS,

Featured Speakers

a16z HostAli Ghazi Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that enterprise AI adoption is slow not because awareness is lacking, but because companies are cautious about data privacy, accuracy needs, internal ownership politics, and cost. Ali Ghazi says large enterprises will likely adopt specialized, proprietary models rather than one universal LLM, with open source and smaller tuned models continuing to gain ground. He also questions benchmarks, warns against hype around AGI, and emphasizes human-in-the-loop safeguards.

Main Topics: Why enterprise GenAI adoption is slow (Priority: 5/5): Enterprises are interested in GenAI, but adoption lags due to slow decision-making, privacy/security fears, uncertainty about data value, and internal political battles over who owns AI initiatives. Proprietary data and competitive advantage (Priority: 5/5): Executives increasingly see their data as a strategic asset that could help them beat competitors, which makes them reluctant to share it with external model providers. Custom enterprise models vs. giant foundation models (Priority: 5/5): Ghazi argues that many enterprise use cases are better served by smaller, accurate, cheaper specialized models trained on proprietary data rather than by the largest general-purpose model. Open source, model leakage, and the role of universities (Priority: 4/5): Open source is portrayed as a major driver of progress, especially after Llama, but proprietary models still tend to lead. Universities are increasingly squeezed out by GPU and funding constraints, pushing them toward efficiency-focused innovation. Scaling laws, fine-tuning, and the future of model architecture (Priority: 4/5): The discussion explores how scaling parameters and data can improve performance but with diminishing returns, and how the long-term goal is modular tuning on top of a strong base model. Benchmarks, reliability, and human-in-the-loop AI (Priority: 5/5): Ghazi criticizes standard benchmarks as gameable and argues that high-stakes applications like medicine and law will still require human oversight because current models remain error-prone. AI ethics, AGI risk, and existential concerns (Priority: 4/5): He distinguishes job displacement and malicious misuse from speculative superintelligence risk, arguing that near-term dangers are limited by compute cost, lack of self-reproduction, and the need for major breakthroughs.

Key Arguments: Enterprises move slowly, but once they adopt a system successfully it becomes sticky and durable. Companies are increasingly treating proprietary data as a strategic weapon rather than a resource to hand over to model vendors. For many enterprise tasks, accuracy and latency matter more than general intelligence, making smaller specialized models more practical. Databricks’ Mosaic acquisition fits a demand for training customer-specific models at scale, but GPU scarcity makes this hard to satisfy broadly. Open source accelerated AI progress significantly; Llama’s release changed the field, and open source will keep improving through leakage, distillation, and efficiency techniques. Proprietary models will likely remain ahead in quality for a long time, though open source can catch up in some cases like Linux did in software. Benchmark performance is not the same as real-world capability because models can memorize tests or train on them indirectly. High-stakes deployment still needs a human in the loop; current models can assist but not fully replace expert judgment. AGI-style catastrophic risk is not the near-term concern; the real question is whether future systems gain the ability to self-improve and reproduce autonomously. The main safety brake today is that training and iterating large models is still expensive, slow, and infrastructure-intensive.

Data Points: S&P companies mentioning AI in earnings: nearly 40% - Cited from a Financial Times report to show broad awareness of AI among enterprises. Podcast event coverage: A16Z’s exclusive AI revolution event from just a few weeks ago - Describes the source event for the episode's interview content. Customer organization size at Databricks: 3,000 people - Ghazi says he did not unleash Databricks’ sales force of 3,000 people to sell Mosaic because supply is constrained. Model size example: 100 billion parameters - He notes Databricks can train a very large model if the customer has enough money and GPUs. Benchmark example: MMLU - Used as an example of a benchmark that can be gamed because it appears on the web and can be trained on directly. Open source reference: Llama - Discussed as a pivotal open-source model release that materially changed AI progress. Efficiency technique example: LoRA / QLoRA / prefix tuning - Mentioned as methods for small modifications to large models without full retraining. Historical comparison: 2000 / Cisco at its peak - Used to compare current LLM hype to overfocus on infrastructure winners during the internet era. Cisco valuation reference: half a trillion dollars - Ghazi says Cisco was worth roughly this much at its peak, to illustrate the analogy. Jobs automation horizon: 300 years - He says economies have been automating work for centuries, and the highest-GDP nations tend to automate the most.

Pivotal Quotes: "We haven't seen anybody with any traction in the enterprise." — Interviewer: Sets up the central question about why enterprise GenAI adoption lags consumer adoption. "I think the answer is closer to the latter. There's going to have lots of specialization." — Ali Ghazi: Argues that AI will fragment into many specialized applications rather than one model dominating everything. "I think all the benchmarks are bullshit." — Ali Ghazi: Critiques benchmark-driven evaluation as easily gamed and poorly aligned with real-world performance.

Implications: Enterprise AI will likely be won by vendors that combine proprietary data, specialized models, and strong security with lower-cost inference. Open source will keep pressuring the market, but serious use cases will still demand human oversight and better evaluation methods.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast