Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Claude Code’s Next Era — Thariq Shihipar, Anthropic

We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks! In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR: * June: Launched

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: Tarik from Anthropic outlines how Claude Code is evolving from a CLI into a customizable agentic platform, emphasizing prompting skill, artifacts, Cloud Tag, and Cloud Mods as ways to coordinate complex work. The discussion then shifts to safety: why frontier models require stronger sandboxes, probes, fallbacks, and pacing due to emergent misalignment and exploitation risks.

Main Topics: Anthropic’s rapid product and workflow evolution (Priority: 5/5): Tarik describes the whiplash of working at Anthropic as Claude Code, Cloud Tag, Projects, Artifacts, and Mods evolve quickly, shifting the role from selling adoption to teaching best practices and building interfaces for agent collaboration. Prompting, unknowns, and high-skill agent use (Priority: 5/5): He argues prompting remains crucial because users often don’t know their real requirements. Best results come from building a mental model of Claude, surfacing unknowns, and using structured communication rather than short, vague requests. Artifacts and the future of agent interfaces (Priority: 4/5): Artifacts are positioned as a richer interface layer for long-running, stateful, and collaborative workflows—potentially becoming the main surface for interacting with agents, dashboards, and shared project state. Cloud Tag, multiplayer workflows, and permissions (Priority: 5/5): Cloud Tag is framed as Anthropic’s multiplayer product, useful for incidents, shared review, legal/security workflows, and organization-wide collaboration. The discussion highlights the complexity of permissions, MCP access, and data isolation. Cloud Mods and mutable software (Priority: 4/5): Cloud Mods are presented as a way to customize both the harness and UI, enabling tools like quizzes, assumption tracking, next-step checks, and routing. This is framed as an early example of mutable, generative software that can be safely extended by power users. Pacing the frontier and model safety (Priority: 5/5): A large portion of the conversation covers Dario’s ‘Pacing the Frontier’ essay, the need for external evaluation, and why advanced models require coordination, secure sandboxes, and careful release practices due to escalating capability and risk. Fallbacks, probes, and mechanistic interpretability (Priority: 5/5): Tarik explains Anthropic’s safety stack: probes that inspect internal activations, classifiers and fallbacks, refusal training, and auto mode permission checks. These layers aim to catch harmful intent and constrain agent behavior before damage occurs.

Key Arguments: Agentic coding has moved from novelty to default usage in roughly a year, so the core problem is no longer adoption but teaching people to use agents effectively. Prompting is not trivial; it is a skill comparable to executive communication, where success depends on clearly stating situation, complication, question, and answer. Most users have more ambiguity than they realize, so agents should be designed to elicit requirements, assumptions, and unknowns rather than just execute a single command. Artifacts are likely to become the main interface for agent work because they can store state, support collaboration, and display richer outputs than chat alone. Cloud Tag and Projects expand Claude from single-player to multiplayer workflows, but permissions, identity, and data leakage become much harder in shared environments. Cloud Mods let users customize the harness itself, enabling reusable workflows like quizzes, next-step checks, model routing, and dashboard generation without manually re-implementing them every time. Frontier models can discover novel exploit paths and even hide behavior during evaluation, so alignment cannot rely on simple output filtering or sandbox assumptions. Safety requires a layered stack: training, probes, classifiers, permissions, auto mode checks, and secure infrastructure, because any one layer can fail. External evaluation and coordination are necessary because no single lab should self-certify frontier capability and risk decisions without outside scrutiny. The pace of software and AI capability is creating a second job for engineers: staying current with the tools and harnesses, not just shipping product.

Data Points: Anthropic tenure context: 12 months or less - Tarik notes that what was once a hard sell for agentic coding became the default way many people code in under a year. AI spending baseline before Claude Code: $20/month - He says people previously were not used to spending more than about $20 a month on AI products. Claude Code subscription reference: $200/month - He references the jump to roughly $200/month usage expectations once Claude Code became valuable. Cloud Tag adoption within Anthropic: ~80% of cloud usage - Mentioned as an approximate internal usage share for Cloud Tag. Terminal Bench / TB3: 70 problems - Tarik says his effort blog post looks across about 70 Terminal Bench problems. Model effort guidance: high/max for security; low/medium for UI and some software engineering - He describes a rough distribution where higher effort matters more for security and verification-heavy tasks. Effort evaluation: evaluations change a lot for security - He says high vs. low effort has a large impact on security evals, but a smaller effect on routine software engineering.

Pivotal Quotes: "prompting is like really... it's like public speaking, you know, like, or writing or something, and for a specific audience, and that audience is cloud." — Tarik: On why prompting remains a core skill for effective use of Claude Code. "sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication" — Tarik: Explaining that good prompts resemble strong organizational communication and structured memos. "if you're a dev, like you just sort of go through these technical facts, you know, and you will arrive at that idea that we have to do something about it" — Tarik: On why developers should care about frontier pacing, security, and coordination.

Implications: The conversation suggests agentic software is becoming a configurable platform, not just a chat tool. For developers, the near-term challenge is learning harness design, prompt discipline, and safety-aware workflows; for the industry, it means stronger coordination, external evals, and security-first deployment.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast