The Twenty Minute VC (20VC)
The Twenty Minute VC (20VC)

20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI

Douwe Kiela is the CEO of Contextual AI, building the contextual language model to power the future of businesses. Last month Contextual closed a $20M funding round including Bain Capital, Sarah Guo, Elad Gil and 20VC. He is also an Adjunct Professor in Symbolic Systems at Stanford University. Previ

Featured Speakers

Dal Keeler Guest

Topics Discussed

Episode Summary

Executive Summary: Dal Keeler argues enterprise AI is still early and that the real winners will be models and products designed around retrieval, privacy, attribution, and evaluation rather than raw model scale alone. He explains Contextual AI’s thesis: decouple memory from generation to reduce hallucinations and support compliance, while warning that hype, weak evaluation, and over-regulation may slow innovation.

Main Topics: Why enterprise AI is not yet ready (Priority: 5/5): Keeler says ChatGPT-era excitement outpaced readiness for enterprise use because of hallucinations, poor attribution, compliance constraints, privacy concerns, latency, and difficulty keeping models updated or deletable for GDPR-style requirements. Contextual AI’s retrieval-augmented architecture (Priority: 5/5): Contextual is built around retrieval augmented generation (RAG), separating memory from the generative model so outputs are grounded in retrieved data, improving attribution, updateability, efficiency, and privacy. Model scale vs data scale (Priority: 4/5): He argues data size matters more than parameter count, citing the LLaMA paper and the idea that smaller models trained longer on more data can outperform larger but undertrained models. Data moats and proprietary advantage (Priority: 4/5): Keeler distinguishes between startups that can build data flywheels and those that cannot, noting incumbents can have advantages, but much of the training data is still open web data; special proprietary datasets can still create major defensibility. Evaluation, benchmarking, and security (Priority: 4/5): He says current evaluation is broken because models are sometimes judged by other models, contamination is widespread, and adversarial testing plus external audit layers will become a major new market. Market structure: frontier, mid-tier, and open source (Priority: 4/5): He frames the model landscape as a pyramid: frontier models at the top, open source at the bottom, and the most commercially attractive middle layer of mid-sized specialized models. Regulation, existential risk, and enterprise adoption (Priority: 3/5): Keeler warns that EU-style regulation could suppress innovation, says existential risk is real but extremely small, and believes enterprise adoption is already underway though still gradual.

Key Arguments: Enterprise adoption is blocked less by interest and more by unresolved product issues: hallucination, attribution, compliance, privacy, up-to-dateness, and latency. Retrieval-augmented generation solves key enterprise problems by grounding outputs in external memory, enabling attribution, updates, and stronger privacy boundaries. Model quality depends more on data scale than parameter count; training smaller models longer on more data can be superior to simply increasing size. OpenAI and Google do have moats, mainly from user data, scale, and deep understanding of usage patterns, contrary to claims they have none. Evaluation today is weak and contaminated; dynamic adversarial testing is a better yardstick than static benchmark leaderboards. Security will become a major category because models increasingly generate code and actions, creating prompt-injection and destructive-action risks. The market will not be winner-take-all; different layers of the pyramid will serve different use cases, from frontier AGI-like tasks to specialized business applications. Over-regulation, especially in Europe, risks entrenching incumbents and stifling startups more than it improves safety. Open source will remain important, but it is unlikely to move up to frontier capability because frontier training is too expensive. AI progress is still in early innings, so founders should not assume the market is closed. Data Points: Funding round: $20 million - Contextual AI recently closed a funding round including Bain Capital, Sarah Guo, Eli Gill, and 20 VC. Industry experience: 15 years - Harry introduces Dal Keeler as an expert who has been in AI for the last 15 years. Percentage of day filled with tedious topics: over 50% - Used in the sponsor read to frame productivity loss and the value of AI assistants like Coda. Enterprise adoption timeframe: already happening - Keeler says enterprise AI adoption is not waiting a year; it is beginning to appear in production now. Model-training principle: smaller model + more data + longer training - He cites the LLaMA paper and scaling-law reasoning to support his view that data matters more than size. AI risk probability: non-zero but very, very small - His view on existential risk: possible in principle, but much lower probability than more immediate risks. AGI outlook: 5 to 10 years - He says systems may economically displace humans across many tasks within five to ten years if AGI is defined operationally. Capable model layers: 3-layer pyramid - He describes the market as frontier models at the top, open source at the bottom, and mid-sized models in the middle. Cloud security term: data plane / control plane separation - He explains Contextual’s architecture by separating the memory/data plane from the model plane for privacy and control. OpenAI user data advantage: giant data moat - He says ChatGPT’s viral usage created a major data moat for OpenAI, even if not fully exploited yet.

Pivotal Quotes: "Data size matters even more than model size." — Dal Keeler: On why training data volume and duration can outweigh simply increasing parameter counts. "The whole kind of existential risk debate, I think it actually comes from a very good place... but I think there are much bigger risks, right? Like nuclear war and pandemics and climate change." — Dal Keeler: On AI extinction-risk narratives and his view that they distract from more immediate global risks. "We're really thinking about this as the next generation of language models." — Dal Keeler: Describing Contextual AI’s enterprise-first, retrieval-grounded approach.

Implications: Listeners should expect enterprise AI to shift toward specialized, grounded systems with stronger governance and security. The winners may be companies that master data, retrieval, and evaluation—not just scale.

🔓 Sign Up for Unlimited Episode Search

About The Twenty Minute VC (20VC)

View all episodes from The Twenty Minute VC (20VC)