The Twenty Minute VC (20VC)
The Twenty Minute VC (20VC)

20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's

Alex Wang is the Founder and CEO @ Scale.ai, the company that allows you to make the best models with the best data. To date, Alex has raised $1.6BN for the company with a last reported valuation of $14BN earlier this year. Scale tripled their ARR in 2023 and is expected to hit $1.4BN in ARR by the

Featured Speakers

Alex Wang Guest

Topics Discussed

Episode Summary

Executive Summary: Alex Wang argues AI is hitting a data bottleneck, not a compute ceiling: models have exhausted easy internet data, so the next leap requires frontier data from enterprises, expert human collaboration, and synthetic generation. He frames data as the main durable moat, predicts on-prem and customized AI for enterprises, warns about geopolitical and military risks, and expects value to accrue above and below models rather than in the base models themselves.

Main Topics: Compute progress vs. data bottleneck (Priority: 5/5): Wang says huge compute spend has not yet yielded a model dramatically better than GPT-4, suggesting diminishing returns from compute alone and a plateau caused by data scarcity. Frontier data as the next engine of AI progress (Priority: 5/5): He defines frontier data as complex reasoning chains, tool use, agent behavior, and expert-guided interactions that current internet data does not contain, and argues this is required to move from GPT-4 to future model generations. Enterprise data, on-prem deployment, and data moats (Priority: 5/5): Wang emphasizes that enterprises hold enormous proprietary data sets that can become competitive advantages, and predicts serious enterprises will prefer on-prem or tightly controlled models to avoid exposing sensitive data. Value capture in AI applications and software (Priority: 4/5): He argues that most value will accrue above and below foundation models, not inside the models themselves, and that software will become more customized, more consumption-based, and less per-seat priced. Regulation, data policy, and national competitiveness (Priority: 4/5): He advocates for more pro-data regulation, including pooled industrial data in sectors like aerospace and finance, and says the U.S. and Europe risk slowing AI if they over-restrict data access. China, geopolitics, and military implications of AGI (Priority: 5/5): Wang warns that AGI could be an extreme military advantage and believes China’s centralized industrial policy and permissive data posture could let it catch up quickly or surpass the U.S. Company building, hiring density, and PR (Priority: 3/5): He reflects on Scale’s mistake of over-hiring during hypergrowth, argues for elite talent density, and says direct communication is better than traditional press, which incentivizes sensationalism.

Key Arguments: Compute alone is not sufficient; AI progress now depends on compute, data, and algorithms together, and data is the binding constraint today. The internet and easily crawled public data are largely exhausted, so pre-training gains from web-scale text are tapering off. Humans do not write down most of their reasoning when solving complex business problems, so current model training data misses the very behaviors needed for agents. Frontier data must be produced, not just collected; it requires hybrid human-synthetic workflows with experts guiding models when they get stuck or make errors. Enterprise data can be a powerful moat because it is proprietary, massive, and often more valuable than open internet data. On-prem and open-source-capable models will be important because enterprises will not want their sensitive data used to train competitors. The most defensible AI value will accrue in infrastructure, services, workflow integration, and applications rather than in base model layers. Per-seat pricing will increasingly break down because more work will be done by AI agents, making consumption-based pricing more rational. Regulators should enable responsible data pooling in some sectors and clearer anonymization/use rules in sensitive areas like healthcare. AGI could be a geopolitical and military super-weapon, so the most advanced systems may need to remain closed for security reasons. China may be closer to the U.S. in AI than many assume because centralized industrial policy and data access can accelerate catch-up. Scale’s internal lesson was that company hypergrowth should not automatically imply headcount hypergrowth; talent density matters more than size. Traditional press often distorts incentives toward clicks, so founders should build direct channels and brand through authentic communication.

Data Points: Scale AI revenue growth: trebled revenue in 2023 - Mentioned in the show intro as evidence of Scale’s growth. Scale AI ARR: $1.4 billion - Expected finish for 2024, cited in the introduction. Scale AI funding round: $1 billion - Raised earlier in the year, per the intro. Scale AI valuation: $14 billion - Reported valuation for the $1 billion round. NVIDIA data center revenue: roughly $5 billion per quarter to north of $20 billion per quarter - Used to illustrate the post-GPT-4 compute boom. GPT-4 training data size: less than 1 petabyte - Compared against enterprise data scale to show how much data is locked in companies. JPMorgan proprietary internal dataset: 150 petabytes - Cited as an example of enterprise data abundance. Scale headcount: about 800 people - Current company size mentioned during hiring discussion. Scale team growth: 150 people in 2020 to over 700 by end of 2022 - Used to explain the over-hiring lesson. OpenAI/ChatGPT-era timeline: GPT-4 since fall 2022 - Used to argue that no subsequent base model has been dramatically better. China LLM position: 01.ai’s Yi-Large is among the best models globally, just behind GPT-4, Gemini, and Claude 3 Opus - Used to argue China is catching up quickly. Productivity/service revenue examples: Accenture $2.4 billion in generative AI revenue; OpenAI $2 billion - Referenced in discussion of AI services vs. models. Travel savings claim in sponsor ad: up to 30% - Navan claim mentioned in ad read. Customer savings claim in sponsor ad: $4.4 billion in savings - Zip ad claim about customer savings.

Pivotal Quotes: "At its core, this AI technology has the potential to be one of the greatest military assets that humanity has ever seen. Potentially, even more of a military asset than nukes." — Alex Wang: He is explaining the geopolitical stakes of AGI and why advanced systems may need to be closed. "We need to basically have data abundance of frontier data." — Alex Wang: His core thesis on how AI progress continues after internet data is exhausted. "I think the biggest one today is all that's between us and AGI is compute. And I think we need data to get there too." — Alex Wang: His quick-fire answer summarizing the main misconception about AI progress.

Implications: AI winners will likely be those who control proprietary, high-quality frontier data and can combine it with compute and algorithmic advances. Enterprises should invest in data strategy, on-prem deployment, and talent density now; policymakers should avoid choking off responsible data use.

🔓 Sign Up for Unlimited Episode Search

About The Twenty Minute VC (20VC)

View all episodes from The Twenty Minute VC (20VC)