Episode Summary
Executive Summary: The episode centers on Karina Wynn’s path from early AI product work to leading reasoning-interface efforts at OpenAI, with a focus on how ChatGPT is evolving from a chat box into a modular, task-oriented, generative OS. The conversation covers Canvas, Tasks, model behavior design, evals, prompting reasoning models like o1/o3, and the emerging agent/computer-use stack.
Main Topics: Karina Wynn’s career path and product-research background (Priority: 5/5): Karina describes moving from computer vision and journalism-related AI work at Berkeley and the New York Times into Anthropic and then OpenAI, emphasizing how early UX and product engineering shaped her approach to AI interfaces. Building Canvas as a new interaction paradigm (Priority: 5/5): The discussion explains how Canvas emerged from a collaboration between research, product, design, and engineering, including why it required both product features and model post-training to work well. Behavioral design and model personality (Priority: 5/5): Karina frames model behavior as a design problem: balancing helpfulness, honesty, and harmlessness, and shaping the model’s persona differently for collaboration, writing, and other contexts. Evals, model cards, and training/debugging challenges (Priority: 4/5): The conversation goes deep on evaluation variance, non-apples-to-apples benchmarks, prompt formatting issues, and how model training resembles software debugging with data. Tasks and the path toward agents (Priority: 5/5): Tasks are presented as a foundational step toward more proactive, trustworthy delegation, with trust and collaboration seen as prerequisites for long-horizon agents. Computer use and the future of generative OS interfaces (Priority: 4/5): The speakers discuss browser/desktop agents, latency, precision, and the idea that future software will be accessed through model-mediated, dynamically generated interfaces rather than direct website clicks. OpenAI vs Anthropic culture and product strategy (Priority: 3/5): Karina contrasts the two labs: Anthropic as more focused and enterprise-oriented, OpenAI as more willing to take product risks and explore multiple bets, while noting similar research workflows.
Key Arguments: AI product development is becoming a full-stack discipline that spans model training, behavioral design, and UI deployment rather than separating research from product. Canvas and Tasks were not just UI features; they required model changes, evals, and iterative deployment to handle real user behavior. Reasoning models like o1 are especially strong when given hard constraints and multi-step criteria, but they still need better evaluation standards and prompting practices. Model behavior should be treated as a design problem: the right persona depends on context, and values like honesty and harmlessness can conflict. Trust is the key bottleneck for agents; users will only delegate important tasks after repeated collaborative success and predictable behavior. Computer-use agents are promising but still limited by latency, precision, and context understanding; they will likely improve first in constrained domains like coding and workflows. The future ChatGPT experience may be less like a text box and more like a generative operating system that adapts its interface to the user’s intent.
Data Points: o1 Mini price reduction: from $12 to $4.40 per million tokens - Mentioned in the intro as OpenAI responded to the reasoning price war by cutting o1 Mini pricing. o3 Mini price: $4.40 per million tokens - OpenAI released o3 Mini in ChatGPT and API at the same price point as the reduced o1 Mini. Claude.ai first codebase size: first 50,000 lines of code - Karina says she wrote the first ~50,000 lines of Claude.ai at Anthropic. Anthropic team size at Claude 3 launch: 10–12 people - She describes the post-training/fine-tuning team for Claude 3 as very small. Anthropic early company size: 6–7 people - She recalls the early Claude.ai deployment team as tiny when the product was first built. Claude 3 launch timing: August 2022 - Karina says Anthropic’s product push and her work on Claude.ai began around this time. Canvas beta-to-GA timeline: about 3 months - She says Canvas shipped as a beta model first and took roughly three months to reach GA. Tasks development time: less than 2 months - Karina says Tasks was developed in under two months. GPQA evaluation practice: average of 5 runs - She notes GPQA was high variance, so they averaged five runs for stability. AI Engineer Summit dates: February 20th to 22nd - The intro announces Karina as closing keynote speaker for the summit in New York City.
Pivotal Quotes: "I feel like my team is going from like very full stack from like training models all the way up to like deployment" — Karina Wynn: Describing her team’s scope at OpenAI, spanning model training, product features, and deployment. "The model's personality is actually the reflection of the company or the reflection of the people who create that model." — Karina Wynn: On behavioral design and why model tone and values matter as product decisions. "Agents are gradual progressional tasks, starting off with one-off actions, moving to collaboration, ultimately fully trustworthy long horizon delegation in complex environments" — Karina Wynn: Her tweeted definition of agents, read back during the interview.
Implications: The episode suggests AI products will increasingly be built as integrated systems of model behavior, UX, and evals. For builders, the winning strategy is to prototype fast, measure carefully, and design for trust, collaboration, and intent-aware interfaces.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast