The a16z Podcast
The a16z Podcast

Data Alone Is Not Enough: The Evolution of Data Architectures

Just having data is not enough: it takes an entire system of tools and technology to extract value from data. a hallway style conversation between Ali Ghodsi, CEO and Founder of Databricks, and a16z general partner Martin Casado explore the evolution of data architectures.

Featured Speakers

a16z HostAli Ghodsi Guest

Topics Discussed

Episode Summary

Executive Summary: Ali Ghodsi argues that data infrastructure has evolved from centralized data warehouses to data lakes and now to the “lakehouse,” which unifies analytics, BI, data science, and machine learning on one transactional, schema-aware layer. The conversation maps the modern data stack, explains why streaming often simplifies operations more than it reduces latency, and recommends cloud-native, multi-cloud, open-source, ML-first architectures.

Main Topics: From Data Warehouses to Data Lakes (Priority: 5/5): The discussion traces the original warehouse model, where enterprises copied operational data into a central warehouse for BI dashboards and reporting, and explains why it became a major market despite being designed for a narrower era of structured data. Why the Data Lake Created a New Problem (Priority: 5/5): Data lakes emerged to handle cheap storage and diverse data types plus ML use cases, but dumping everything into blob storage created a 'data swamp' and led companies to maintain redundant copies in lakes and warehouses. Lakehouse as the Next Architecture (Priority: 5/5): The proposed lakehouse adds transactionality, schemas, and data quality on top of object storage so BI, reporting, data science, and ML can operate directly on the lake with structured semantics and strong performance. BI vs. ML: Same Data, Different Workflows (Priority: 4/5): The speakers compare traditional analytics and machine learning, noting they use much of the same data but differ in directionality, user personas, organizational location, and implementation requirements. Streaming as Operational Simplification (Priority: 4/5): Streaming is presented not just as a latency play but as a way to reduce data operations complexity by handling late data, consistency, and reconciliation automatically, potentially replacing many batch workflows. Modern Data Stack Design Principles (Priority: 5/5): Ali outlines a cloud-native stack: ingest into a data lake first, add a transactional layer, use notebooks/interactive environments, then layer ML platforms and BI tools on top through data frames and SQL. Strategic Guidance: Multi-cloud, Open, ML-first (Priority: 4/5): For future-proofing, he recommends multi-cloud, open standards/open source, raw data landing in lakes, and treating machine learning as a first-class citizen in the architecture.

Key Arguments: The warehouse model solved a real BI problem, but it was built for structured, historical reporting and is increasingly insufficient for modern data types and AI/ML workflows. Data lakes solved storage and flexibility but introduced fragmentation and duplicated ETL, forcing many companies to keep data in both lakes and warehouses. Lakehouse architectures can unify structured analytics and ML by adding transactional guarantees and schema management on top of inexpensive object storage. Machine learning and BI overlap heavily in data needs, but differ in users, organizational ownership, and the iterative nature of ML algorithms. Data frames became the lingua franca of data science because they bring table-like structure into programming languages like Python and R, and SQL can be layered on top of them. Traditional relational systems are not a natural fit for iterative ML algorithms; attempts to bolt ML onto warehouse architectures proved difficult and awkward. Streaming’s biggest value may be operational: it can automate reconciliation, late-arriving data handling, and reruns, not merely cut latency. A modern stack should land raw data in the lake first to avoid premature schema decisions, then layer transactionality, notebooks, ML platforms, and BI access on top. Cloud-native systems should not mimic on-prem assumptions because cloud networking makes data locality constraints less important. Enterprises should prefer multi-cloud and open standards to avoid repeating past vendor lock-in cycles.

Data Points: Data warehouse market size: at least $20 billion - Ali describes the long-running BI/data warehouse industry as a major market created by the warehouse paradigm. Legacy timeline for data warehouse origins: 1980s - The warehouse model is traced back to the 1980s when business leaders lacked timely visibility into operations. Time since data lake emergence: about 10 years ago - Ali says the data lake pattern arose roughly a decade before the conversation as a response to warehouse limitations. Time since lakehouse breakthroughs: last 2–3 years - He says recent technological advances have made the lakehouse design pattern feasible. Example streaming latency targets: every 5 minutes / every 10 minutes / every 30 minutes - Business leaders often state these as acceptable freshness windows, suggesting many streaming use cases do not need millisecond latency. Legacy horizon in enterprises: 40 years of technology - Ali notes Western enterprises often carry decades of accumulated systems that complicate migration. Example cloud-era company age: 10 years old - Uber is cited as a company that built a custom modern stack without heavy legacy constraints.

Pivotal Quotes: "It's like an architectural mess that's inferior to what we had in the 80s" — Ali Ghodsi: Commenting on the modern split between data lakes and warehouses and the need for duplicate copies and sync. "What if you could actually be able to do BI directly on your data lake?" — Ali Ghodsi: Summarizing the core premise of the lakehouse architecture. "In some sense, all of the batch data that's out there is potential use case for streaming." — Ali Ghodsi: Explaining why streaming may be more broadly useful as an operational model than as a pure low-latency feature.

Implications: The industry is converging on unified, cloud-native data platforms that treat AI/ML as core, not optional. Leaders should reduce duplicated ETL, embrace open and multi-cloud foundations, and design for one governed data layer serving both analytics and production ML.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast