Episode Summary
Executive Summary: The episode examines Anthropic’s responsible scaling policy (RSP) with head of training Nick Joseph. They argue RSPs can tie AI deployment to measured capability thresholds, create a pragmatic middle path between alarmism and inaction, and align safety with business incentives. The conversation also explores objections: trust, vague thresholds, sandbagging, missing risks, and the need for external oversight or regulation.
Main Topics: What responsible scaling policies are: RSPs set capability thresholds, evaluate models against dangerous 'red lines,' and require escalating safeguards as models become more capable. Anthropic’s safety levels and deployment gates: The discussion details ASL2/ASL3-style triggers, red-teaming, security requirements, and the conditions under which Anthropic would delay or block deployment. Why scaling still matters: Joseph explains why he believes scaling laws remain real, why models keep improving with more compute/data, and why progress has been faster than many skeptics expect. Objections: trust, ambiguity, and incentives: The interview probes concerns that companies may reinterpret policies, under-elicite capabilities, or weaken standards under commercial pressure. Safety work inside Anthropic: Joseph argues that frontier-capability work and safety research are intertwined, and points to interpretability, jailbreak research, and constitutional AI as concrete outputs. Careers at frontier AI labs: The episode ends with advice on engineering-heavy roles, building public proof-of-work, and thinking carefully about whether capabilities work helps or harms the future.
Key Arguments: RSPs are attractive because they separate 'can the model do dangerous things?' from 'what do we do about it?,' allowing skeptics and safety advocates to agree on measured triggers. Outcome-based safety policies are better than spend-based or effort-based ones because they require actual safety success before deployment. Commercial incentives can be aligned with safety if shipping and revenue depend on passing safety gates. The best way to reduce risk is defense in depth: model evaluations, red teaming, security, internal checks, and eventually external regulation. Current evaluations are imperfect, but starting early helps labs learn how to run them before stakes become extreme. The biggest technical failure modes are under-elicitation, sandbagging, and unknown unknowns. Capabilities and safety are intertwined; frontier models enable many safety techniques and evaluations that would be impossible on weaker systems. AI companies alone probably cannot solve all security and governance problems; external auditors or government regulation will likely be needed eventually. Anthropic should not be judged only by marginal capability gains, but by the overall counterfactual impact of the organization’s existence, including safety research and policy work. Engineering talent is especially scarce and useful because much frontier AI progress depends on implementation, tooling, and fast experimentation rather than pure research.
Data Points: Interview date: 30 May 2024 - Recorded interview with Nick Joseph Anthropic staff working on security: 8% - Joseph says roughly 8% of Anthropic staff work on security Model training compute share: 99% - He says pre-training historically consumes about 99% of compute in many cases Team size: Over 40 people - Joseph leads a team of over 40 people focused on training Anthropic’s models ASL level for Claude 3: ASL2 - Joseph says Claude 3 is currently categorized in Anthropic’s lower-risk safety level White House commitments timing: Sometime last year - He refers to industry best-practice commitments made before the RSP escalations Anthropic employee growth: Rapid growth - He notes the company recently had to move into a larger office because it ran out of desks Target audience growth effort: Around $3 million annual marketing budget - Mentioned in the closing 80,000 Hours housekeeping segment Estimated frontier model training cost: Around $100 million - Rob Wiblin cites a rough figure for training a frontier LLM Hiring note for 80,000 Hours roles: Approx. £80,000 for someone with five years’ experience - Closing segment mentions salary guidance for head of video/marketing roles
Pivotal Quotes: "If we have evaluations showing the model can do X, then we should take these precautions." — Nick Joseph: Explaining why RSPs can build support by tying safeguards to measured capabilities "It's not, did we invest X amount of money in it? It's not, did we try? It's did we succeed." — Nick Joseph: Describing why outcome-based safety policy is superior to effort-based compliance "I think this is a spot where there are many people who are skeptical that models will ever be capable of this sort of catastrophic danger." — Nick Joseph: Explaining why RSPs can bridge disagreements between skeptics and safety advocates
Implications: RSPs may become the main interim governance model for frontier AI, but their credibility will depend on better evals, honest implementation, and eventually external oversight. The interview suggests safety work, engineering, and policy are becoming inseparable at the frontier.