Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI

From the frontlines of OpenAI’s Codex and GPT-5 training teams, Bryan and Bill are building the future of AI-powered coding—where agents don’t just autocomplete, they architect, refactor, and ship entire features while you sleep. We caught up with them at AI Engineer Conference right after the launc

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: OpenAI speakers Bill and Brian discuss how coding models are evolving from generic chatbots into opinionated, agentic systems like Codex Max and GPT-5 variants. The conversation centers on training for trust, communication, tool use, long-running tasks, evaluation pipelines, and a broader shift toward agents that can manage context, spawn sub-agents, and automate work beyond coding.

Main Topics: Codex Max and model naming (Priority: 4/5): They explain why Codex Max was named to signal speed, maximal capability, and long-running performance, contrasting it with more general GPT-5-style models. Training for trust and personality (Priority: 5/5): The speakers argue that coding models need behavioral traits like communication, planning, context gathering, and checking work so developers can trust them as pair programmers. Tool use and harness-specific optimization (Priority: 5/5): A major theme is that models perform best when their tools, names, and input/output shapes match the training harness; even seemingly small choices like tool names can affect results. Agent layer over model layer (Priority: 5/5): They describe an industry shift from optimizing individual models to shipping agents as packaged products, allowing developers to build one layer above the model and rely on a whole workflow. Multi-agent and long-running workflows (Priority: 4/5): They discuss Codex Max managing its own context window and potentially handing work to sub-agents, enabling near-continuous execution and parallelization. Applied evals and real-world benchmarking (Priority: 5/5): OpenAI is emphasizing applied evals, multi-turn traces, and production-style grading to better measure what models can do in real workflows rather than only academic benchmarks. Non-coding automation and future directions (Priority: 4/5): The conversation extends beyond coding into personal automation, email, Slack-based workflows, desktop organization, and computer-use agents for legacy software with only UIs.

Key Arguments: Coding models are more trustworthy when they behave like good software engineers: they communicate progress, plan before acting, and verify results. Model performance depends heavily on the surrounding tool harness; tools that match the model’s training assumptions can materially improve outcomes. Codex is intentionally opinionated and optimized for a specific agent harness, while GPT-5-style models are broader and more steerable across different tools. The abstraction layer is moving upward from model selection to agent selection, so developers will increasingly integrate full agents rather than micromanage every model release. Long-running tasks require context management, compaction, and sub-agent handoffs; Codex Max is designed to support these workflows. Applied evals are necessary because real-world work is multi-turn, messy, and production-like; internal and customer-driven evals help close gaps. Coding agents are beginning to behave like general-purpose computer-use agents, especially for terminal-based and workflow automation tasks outside pure software development.

Data Points: Long-running task duration: 24 hours or more - Codex Max is described as able to run very long tasks; one speaker notes having it run for 24+ hours. Historical OpenAI Codex adoption: about 50% - One speaker says that when Codex first launched, around half of OpenAI employees started using it. Timeframe prediction: 2026 - The speakers speculate that more computer use, sub-agents, and broader Codex capabilities will be visible by 2026.

Pivotal Quotes: "The abstraction layer really moving, starting to move upwards from the model layer towards the agent layer." — Brian: Used to summarize the industry trend away from raw model optimization toward packaged agent workflows. "I haven't written a single line of code by hand in months because I know what I can trust it to do." — Bill: A strong statement about trust in Codex for real coding work and productivity. "I think Slack is the ultimate user interface for work." — Brian: Used to motivate broader automation workflows like email agents and non-coding task management.

Implications: For builders, the key takeaway is to design around agents, not just models: trust, evals, tools, and harness design now shape capability as much as model quality. Expect coding agents to expand into general automation and computer-use workflows.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast