Episode Summary
Executive Summary: This panel debates the future of the modern data stack and finds broad convergence, but with disagreement on timing and winners. Speakers mostly agree SQL and open formats will dominate structured/semi-structured data, while complex data, ML, and operational use cases keep driving data lakes and hybrid architectures. The group also sees data mesh, Arrow, and data apps as important near-term shifts.
Main Topics: Data lakes vs. data warehouses (Priority: 5/5): The panel debates whether lakes remain necessary as warehouses absorb more capabilities. Views split between convergence to SQL warehouses and continued specialization for complex/operational workloads. SQL, relational systems, and convergence (Priority: 5/5): Several speakers argue SQL and relational models keep winning over time, especially for structured and semi-structured data, though procedural tools remain necessary for complex processing. AI/ML vs. analytics stacks (Priority: 4/5): The conversation contrasts analytics/dashboarding with operational AI/ML workloads such as fraud detection, pricing, and prediction, which often fit lake-oriented or code-first systems better today. Interoperability and open formats (Priority: 4/5): Arrow, shared file formats, and shared metadata/indexing are presented as the practical bridge between otherwise fragmented analytics and ML ecosystems. Data mesh and organizational decentralization (Priority: 4/5): Participants distinguish the useful organizational idea of distributed teams and standards from the potentially harmful idea of fully distributed technical architectures. Data apps, latency, and operationalization (Priority: 4/5): The panel discusses the emerging category of data apps that take autonomous actions, emphasizing latency/throughput trade-offs and the need for tighter feedback loops into operational systems. Future platform winners (Priority: 3/5): The speakers speculate on whether another major platform will emerge beside Snowflake, Databricks, and the cloud giants, with mixed but generally open-ended expectations.
Key Arguments: Relational SQL warehouses will keep absorbing more use cases, especially structured and semi-structured data, and may eventually subsume many functions of data lakes. Data lakes still matter because complex data (images, video, documents, streaming) and operational AI/ML workloads have different requirements than classic analytics. Organizations should store data once in open formats to avoid duplicate lake/warehouse copies and reprocessing costs. The real convergence is happening across APIs: SQL, Python/Scala, and interchange layers like Arrow will increasingly coexist in hybrid systems. Data mesh is valuable as an organizational model for federated data ownership, but not as a mandate for fully distributed technical architecture. Data apps will emerge as systems that use data to trigger actions automatically, but latency requirements will vary by use case and won’t always require real-time performance. The future likely includes specialized roles and tools rather than a single universal interface, because different personas and workloads require different workflows.
Data Points: Timeline for structured/semi-structured data in SQL warehouses: 5 years - Bob Muglia predicts data will sit behind a SQL prompt and that warehouses will replace lakes for structured and semi-structured data. Timeline for relational to win broadly: 8 to 10 years - Bob suggests relational will win in the long run for broader workloads, beyond current SQL warehousing use cases. Complex data support in warehouses: 2 to 3 years - Bob says support for images/videos/other complex data in warehouses is coming soon as a feature. Hybrid stack dominance: 3 to 5 years - Bob predicts hybrid systems combining predictive and declarative stacks will dominate in the near term. Company value across Data50 categories: more than $100 billion - Introductory promo for the Data50 list describing aggregate value of leading private data companies. Total capital raised by Data50 companies: approximately $14 billion - Introductory promo for the Data50 list. Conference recording date: November 2020 - The episode is described as originally recorded in November 2020 at the Modern Data Stack Conference.
Pivotal Quotes: "I think over time, you could argue that it's the data lake that ends up consuming everything, not the data warehouse." — Martin Casado: On whether lakes still have a future and how use cases may favor operational AI/ML. "Five years from now, data is going to sit behind a SQL prompt." — Bob Muglia: On the long-term dominance of relational SQL warehouses for structured and semi-structured data. "The data lake will have a place. Your images, your blog storage, all of those things are probably going to remain in the data lake." — Michelle Ufford: On specialization and why complex/nontraditional data sources will still need lake-like storage.
Implications: The industry is moving toward hybrid, open, interoperable stacks rather than one winner. SQL warehouses will keep expanding, but AI/ML, complex data, and data apps will preserve demand for lakes, procedural tools, and shared formats like Arrow.
About The a16z Podcast
The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!