Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Cursor's Third Era: Cloud Agents

All speakers are announced at AIE EU, schedule coming soon. Join us there or in Miami with the renowned organizers of React Miami! Singapore CFP also open! We’ve called this out a few times over in AINews, but the overwhelming consensus in the Valley is that “the IDE is Dead”. In November it was jus

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The transcript centers on Cursor’s shift from code completion to cloud-based, agentic software development. The speakers argue that the next leap is wider parallel pipelines—not just smarter models—enabled by cloud agents that onboard themselves, test changes, produce videos, and collaborate across Slack, transcripts, and sub-agents. The discussion also covers model routing, multi-model “councils,” memory/self-awareness, and the emerging need for production-safe infrastructure as agent output scales.

Main Topics: Cursor’s Cloud Agents as the next product phase (Priority: 5/5): Cursor’s major launch is a cloud agent system that runs on a full computer/VM, can onboard itself, use desktop and terminal, and execute end-to-end tasks rather than only writing diffs. Testing, videos, and remote desktop as review artifacts (Priority: 5/5): The agents now test their changes, return a demo video, and provide full remote desktop access, making review easier than reading a huge diff and helping align intent with execution. Parallelism, swarms, and throughput as the core unlock (Priority: 5/5): The speakers repeatedly frame the future as making the pipe wider: more parallel agents, swarms, and best-of-N workflows, rather than just making one model faster. Multi-model systems, routing, and agent councils (Priority: 4/5): Cursor experiments with multiple model providers, synthesizer/judge layers, and model routing to combine strengths across providers and improve outputs beyond a single unified tier. Slack, collaboration, and transcript-driven workflows (Priority: 4/5): Cloud agents are increasingly used in Slack threads and with transcript sharing so humans can collaborate with the agent, continue conversations, or use another agent to debug prior runs. Memory, self-awareness, and onboarding (Priority: 4/5): The team emphasizes better onboarding, dynamic file context, and agent self-awareness about secrets, environment, and its own limitations as key to making agents more autonomous and reliable. Production bottlenecks and infra scaling (Priority: 4/5): As code generation accelerates, the bottleneck shifts to review, CI/CD, merge queues, security/performance checks, and production pipelines that can safely absorb far more output.

Key Arguments: The biggest unlock is not one person with a model doing more, but widening the system with parallel agents and swarms that increase throughput. Cloud agents are more useful when they run a full VM/desktop, can onboard themselves, and test their own changes end-to-end. Videos of what the agent built are a better review entry point than giant diffs, especially when agents generate much larger changes. Best-of-N works better when outputs are summarized into short videos or synthesized by another agent, making comparison scalable. Using different model providers together can create synergistic outputs that outperform a single homogeneous base tier. Slack is becoming a development surface where agents collaborate with humans, gather missing context, and hand off artifacts like PRs. As agent output increases, the real bottleneck moves from coding to getting software safely into production. Memory should be treated as dynamic context and self-auditability rather than a static note file; agents should understand their own harness and constraints. Tooling will likely need to support long-running, persistent, and resumable cloud work rather than short local tasks. Cursor’s role is to help create software end-to-end, not just generate code tokens.

Data Points: Cloud Agents launch timing: June last year - The speakers say Cursor launched Cloud Agents in June of the previous year and has been iterating on them since. Example agent runtime: Half an hour - One demo agent worked for about half an hour because it wrote code and tested it end-to-end. Bug reproduction workflow: 90 seconds or something like that - A reproducible bug workflow with slash repro was described as enabling very fast merge-ready fixes. Best-of-N comparison scale: Four 20-second videos - The team says comparing multiple model outputs becomes manageable when each produces a short review video instead of a huge diff. Large diff size avoided: 700 lines of code times four - Before video-based reviews, running four models could yield unwieldy diffs that were hard to review. Consequence of scaling agents: 10x as many agents - They describe headcount-like growth where a team running many agents effectively increases code output dramatically. Model generations mentioned: Opus 4.5 / Codex 5.3 - The speakers cite these model versions as step changes in autonomy and computer-use capability. Pricing trajectory: hundreds of dollars a month; thousands of dollars a month per human; tens of thousands and beyond - They predict agentic tooling will increasingly justify much higher per-user spend as leverage rises. Local vs cloud adoption prediction: More than 2x - One speaker predicts cloud agents will exceed local agent volume by more than 2x by year end.

Pivotal Quotes: "The big unlock is not going to be one person with a model getting more done, like the water flowing faster. It will be making the pipe much wider." — Speaker 1: Explaining why parallel agents and swarms matter more than incremental model speedups. "A brain in a box is what you want." — Speaker 1: Describing the philosophy behind giving agents full VM/desktop access and minimizing capability constraints. "We want cursor to be a brain in a box." — Speaker 2: Summarizing Cursor’s design goal for cloud agents and full-environment autonomy.

Implications: Agentic development is shifting software work from typing code to orchestrating parallel workers, reviewing artifacts, and managing production pipelines. Teams will need better onboarding, observability, review, and deployment systems as AI-generated output scales rapidly.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast