The Cognitive Revolution
The Cognitive Revolution

E49: Open AI, Anthropic, and Meta | Analyzing the AI Frontier with Zvi

This isn't news, it's analysis! Nathan Labenz sits down for an with Zvi Mowshowitz, the writer behind Don't Worry About the Vase to talk about the major players in AI over the last few months. In this extended conversation, Nathan and Zvi debate if AI has attained the intelligence of

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The episode analyzes recent AI developments through the lens of “live players” shaping the frontier: OpenAI, Anthropic, Google DeepMind, Meta, and regulators. The guests debate model capability vs potential, safety approaches like Constitutional AI, robotics and medical multimodal progress, open-source risks, red-teaming, and policy bottlenecks. The conversation is cautiously optimistic about coordination and safety efforts, but deeply wary of capability scaling and misaligned incentives.

Main Topics: GPT-4 capability: human-level vs potential (Priority: 5/5): The speakers distinguish between average performance, peak performance, and latent potential. One argues GPT-4 is below human level in raw general intelligence and originality, while acknowledging it is highly useful for mundane tasks and can become much more capable as scaffolding, memory, and agentic interfaces improve. OpenAI and the super-alignment strategy (Priority: 5/5): A major point is the interpretation of OpenAI’s super-alignment plan: not to align a human-level researcher now, but to prepare for aligning an inevitably more powerful future system. The discussion stresses that OpenAI’s leaders are unusually important and unusually willing to engage seriously with critique. Anthropic, Constitutional AI, and safety trade-offs (Priority: 5/5): Anthropic is framed as the most safety-minded frontier lab, but also as one that may over-refuse, over-constrain, and optimize for harmlessness at the expense of usefulness. The speakers debate whether Constitutional AI scales and whether Anthropic’s approach will generalize to much more capable systems. Google DeepMind: robotics, multimodal medicine, and Gemini uncertainty (Priority: 4/5): DeepMind’s robotics and medical multimodal work is treated as expected but meaningful progress. The broader concern is that Google still hasn’t clearly shipped a competitive Gemini-level general model, despite its research depth and resources. Meta and open-source frontier models (Priority: 5/5): Meta’s Llama release is portrayed as strategically dangerous because open-sourcing frontier models can accelerate proliferation of capabilities without control. The guest sees Meta as a more immediate concern than China because it is directly pushing open release of powerful systems. Red-teaming, evaluations, and AI safety bottlenecks (Priority: 5/5): The conversation argues that third-party evals and red-teaming are valuable but insufficient. Real safety progress will require multiple independent evaluators, better standards, strong policy, and researchers/organizations focused on the hardest problems rather than easy safety theater. Policy, liability, and coordination mechanisms (Priority: 4/5): Policy is seen as potentially very high leverage but difficult to operationalize. The discussion favors concrete measures like compute governance, mandatory insurance, and liability regimes for harms, while warning against vague policy efforts that do not address core risks.

Key Arguments: GPT-4’s practical utility does not imply it is human-level in raw general intelligence, creativity, or robustness. The important question is not current average performance but the potential of systems as they gain memory, scaffolding, and agentic structure. OpenAI’s super-alignment announcement is better understood as preparing for future powerful systems than aligning GPT-4 itself. Anthropic’s Constitutional AI may work well on current systems, but the specific implementation may over-refuse and may not scale cleanly to superintelligent systems. Open-source frontier models are dangerous because any gains diffuse immediately, including to adversaries, and the model can be fine-tuned into unsafe behavior quickly. Red-teaming and independent evals are necessary but can become safety theater if organizations or regulators treat passing tests as proof of safety. Policy should focus on concrete levers such as compute regulation, GPU tracking, liability, and insurance rather than generic AI ethics. The most valuable people in AI safety are those with leadership, fundability, deep technical understanding, and willingness to tackle hard problems rather than publish easy work. Meta’s open-source strategy is interpreted as a risky, ideologically driven move that can worsen long-term capability proliferation. Google’s robotics and medical models show impressive incremental progress, but deployment is constrained by regulation, culture, and market incentives, not just technical ability.

Data Points: GPT-4 comparison to human: “well below human level” / “maybe at the level of a well-read college undergrad” - Debate over Jan Leike’s characterization of GPT-4 versus the guest’s own assessment OpenAI super-alignment timeline: 4 years - OpenAI’s stated horizon for aligning a future human-level alignment researcher/system Anthropic Claude susceptibility to universal adversarial attack: ~2% success rate - Paper finding Claude was much less vulnerable than other leading models to the jailbreak strings Anthropic vs other models on attack: more than an order of magnitude lower - Comparison of Claude’s vulnerability to the universal adversarial prompt attack GPT-4 preference study: 70/30 - Referenced GPT-4 technical report finding people preferred GPT-4 responses over GPT-3.5 about 70% of the time Claude context window: 100K tokens - Noted as a major product differentiator and useful for long-document workflows Anthropic safety resourcing: 20% of compute - Discussed as the amount Anthropic is devoting to super-alignment efforts Meta’s Llama 2 release threshold: 700 million daily users - Referenced as the threshold above which Meta requires a license for model use AI safety grant review volume: 150 grant applications - Mentioned in relation to reviewing the Survival and Flourishing Fund applications Support for policy emphasis: compute regulation, GPU tracking, and insurance/liability - Specific policy levers identified as most promising for safety governance

Pivotal Quotes: "“there is no, it goes from not working to working, it goes from working worse to working better. And then it could always go to working better still.”" — Zvi Moshowitz: On why current model capability should not be mistaken for a ceiling "“I’m more afraid of Meta… one individual American company scares me more than all of China right now.”" — Zvi Moshowitz: On which actors most worry him in the frontier AI race "“the Super Alignment Project is trying to keep the humans out of the loop entirely. And that should be about as scary as it sounds.”" — Zvi Moshowitz: On the implications of OpenAI’s super-alignment effort

Implications: The frontier is being shaped by a few highly consequential actors. Near-term usefulness will keep rising, but safety, governance, and deployment discipline will determine whether this becomes broadly beneficial or dangerously uncontrollable.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution