Episode Summary
Executive Summary: Patrick O'Shaughnessy interviews Databricks founder Olly Goetze about the history of distributed computing, the creation of Spark, and Databricks' cloud platform for data, AI, and machine learning.
Main Topics: From big data to distributed computing (Priority: 5/5): Olly explains how falling storage costs and Moore's wall pushed companies toward data centers and distributed systems. Hadoop and MapReduce (Priority: 5/5): He describes Hadoop as the early operating system for massive parallel processing, but hard to program. Why Spark was created (Priority: 5/5): Spark began at Berkeley to speed iterative machine learning by keeping data in memory instead of disk. Databricks as a cloud platform (Priority: 5/5): Databricks packages Spark plus Delta, MLflow, and Redash as managed cloud software for enterprise AI. Enterprise data and AI operating model (Priority: 4/5): He argues companies need one leader, open source, cloud, and AI-first data strategy to avoid silos. The next platform shift: lakehouse (Priority: 4/5): He sees the merger of data warehousing and AI/data science into a single lakehouse paradigm. Leadership and company scaling (Priority: 3/5): Olly reflects on how CEO priorities change from product-market fit to scaling and process discipline.
Key Arguments: Moore's wall slowed CPU gains, forcing distributed computing across data centers. MapReduce moved computation to data because networks couldn't handle moving huge datasets. Hadoop made parallel big-data processing possible but was cumbersome due to Map/Reduce constraints. Spark was built for iterative ML and was much faster by loading data into memory. Databricks evolved from selling Spark adoption to a broader cloud data/AI platform. Open source reduces enterprise lock-in and is central to Databricks' strategy. Winning with data requires end-to-end ownership from ingest through production use cases. The biggest barrier today is organizational and technical separation between data management and AI teams. Best-in-class firms align IT and line-of-business under one data leader and adopt cloud plus multi-cloud. The lakehouse is the next major simplification: one platform for data management and AI.
Data Points: Company age: seven-year-old company - Databricks' size and maturity as described by Olly. Employees: about 1,700 employees - Databricks' scale at the time of the interview. CPU speed: three gigahertz - Olly says CPUs have basically been stuck at this level since about 2005. Timeline: 2005, 2006 - When enterprises began collecting more data and big data interest accelerated. Timeline: 2009, 10, 11, 12 - Period when Spark had academic success but little industry adoption. Net promoter/competition prize: half a million dollars or a million dollars - Netflix competition incentive for predictive recommendation models. Customer scale: hundreds or thousands of machines - Hadoop-era distributed processing across data centers. Community adoption: hundreds of thousands of data scientists - Users of Databricks' free Community Edition every day. YC usage: more than 100 Y Combinator businesses - Vanta sponsorship claim in the ad read, not part of the main interview. Dataset scale: one petabyte to two petabytes - Used as an example of old big-data success metrics. Projected horizon: 10 years, 15 years from now - Olly's view on when every company will use AI strategically across BUs. Estimated cost impact: hundreds of millions of dollars - Potential savings from predictive maintenance at Shell. Healthcare timing: every three months - The son's screening interval due to genetic testing.
Pivotal Quotes: "We call it Moore's Wall because they didn't figure out how to make computers faster." — Olly Goetze: Explaining why distributed systems became necessary when single-machine scaling stalled. "We don't want anyone to get locked into us." — Olly Goetze: Describing Databricks' open-source-first philosophy and anti-lock-in strategy. "The biggest wall is between data and AI." — Olly Goetze: Summarizing the core organizational and technical bottleneck in modern enterprises.
Implications: The next challenge is unifying governance, data, and AI into one cloud-native stack; enterprises should modernize around open source and cross-functional leadership.
About Invest Like the Best with Patrick O'Shaughnessy
Conversations with the best investors and business builders in the world.
View all episodes from Invest Like the Best with Patrick O'Shaughnessy