Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Why Dwarkesh is Wrong about Computer Use + How OpenAI shipped its Jev competitor in 1 Week

Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The podcast centers on OpenAI Dev Day announcements around computer use, agents, and performance-oriented API upgrades. Ari and Nick explain how new models, harness improvements, async tool calling, WebSockets, caching, compaction, and the new Decisions API make agents faster, more capable, and more practical for real-world tasks like browsing, support, and software testing. The conversation emphasizes that computer use is rapidly approaching expert-level reliability while still requiring careful safety and UX design.

Main Topics: Dev Day computer use announcements (Priority: 5/5): Ari recaps the major launches: DOTS with persistent cloud computers, GPT 6.1 Sol for computer use efficiency, the Agents API with computer use, app shots, native Mac computer use, and the Decisions API. Why computer use is now materially better (Priority: 5/5): Ari argues that model capability, harness techniques, and multimodal interfaces have all improved, enabling debugging, introspection, multi-step code execution, and faster task completion. App shots and accessibility-driven context (Priority: 4/5): The discussion explains how app shots capture raw accessibility/metadata rather than just pixels, giving models fuller context and enabling better control of apps and web pages. Developer APIs: async tool calling, WebSockets, and steering (Priority: 4/5): Nick describes API-side upgrades that allow models to keep reasoning while tools run, receive mid-turn instructions, and communicate bi-directionally via WebSockets. Decisions API and fast classification (Priority: 4/5): Nick frames Decisions as a fast, zero-shot classification-oriented layer built on top of Luna weights, optimized for latency, structured output, and parallel inference. Caching, compaction, and long-lived agent threads (Priority: 4/5): The conversation highlights new caching guarantees, pre-warming, server-side compaction, and file-based compaction techniques to support long-running agent workflows. Safety, trust, and best practices for builders (Priority: 5/5): Both speakers stress consent, safety checks, scope-limited access, and reliability as developers begin deploying agents that can act on websites, accounts, and even payments.

Key Arguments: Computer use is now useful for delegating real tasks because agents can operate the same software humans do, including desktop apps, browsers, and web services. Recent gains come from both model improvements and harness engineering; the product is not just a better model, but a better system. App shots are more powerful than screenshots because they expose accessibility metadata, full text, and structural context to the model. The model can now debug and try again, not just start tasks, which is a major step toward reliable long-horizon automation. OpenAI sees computer use as increasingly faster than average humans for many tasks, with a next goal of matching or exceeding expert users. The Agents API is valuable because developers can use the same computer use stack OpenAI uses internally, which may improve speed, cost, and accuracy. The Decisions API is optimized for fast classification and structured outputs, not for deep reasoning or long-horizon work. Caching and compaction are becoming essential primitives for long-running agent threads, especially when agents revisit work hours later. Builders should add consent, scoping, and safety checks before allowing consequential actions such as payments or broad website access. Computer use can close the software development loop by letting an agent both build and test software, reducing the burden on humans as QA. The team is still exploring how much higher-level abstraction the platform should expose versus leaving flexibility to developers. Real-world bottlenecks increasingly include the speed of the external website or app itself, not just model latency. The best future experiences will be more event-driven and real-time, reducing delays between app state changes and model actions.

Data Points: Cost reduction vs Astra: 1/5 the cost - Ari said GPT 6.1 Sol is a fifth of the cost as Astra. Computer use cost reduction vs Astra: 1/7 the cost - Ari said GPT 6.1 Sol is a seventh of the cost if looking specifically at computer use. Meal-prep task speedup: 15 minutes vs 2 hours - Ari’s meal-prep order took the model 15 minutes instead of two hours for him. Task speedup: 8x faster - Ari described the meal-prep order as eight times faster than doing it manually. Speed improvement: 7x - Nick referenced keynote claims of 7x improvements in computer use speed. Support caching guarantee: 30 minutes - Nick said responses API now provides guaranteed cache hits within 30 minutes. Preview cache window: 12 hours - Nick mentioned a preview offering for longer cache-read guarantees. Thread duration: 3 to 4 hours - Nick used a 3–4 hour return-to-thread example to explain why caching matters. Team tenure: roughly 3 years - Nick introduced himself as having worked at OpenAI for roughly three years. Launch cadence: coming days - Nick said Decisions API was expected to launch in the coming days.

Pivotal Quotes: "the biggest delta that I see is before they could reliably start tasks, but then they would run into problems, and now they're really good at debugging" — Ari: Ari describing the main improvement in computer use capability over the last year. "the next frontier is to have computer use be literally superhuman in its performance" — Ari: Ari outlining the long-term goal for computer use agents. "The main goal in the API is to put things in the API once it's trained into the harness" — Nick: Nick describing OpenAI’s product philosophy for exposing model capabilities through the API.

Implications: Computer use is shifting from novelty to infrastructure: agents can now test, operate, and automate real software workflows. For builders, the key challenges are reliability, latency, caching, and safety; for users, the payoff is major time savings and more capable assistants.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast