The Cognitive Revolution
The Cognitive Revolution

Intelligence with Everyone: RL @ MiniMax, with Olive Song, from AIE NYC & Inference by Turing Post

Olive Song from MiniMax shares how her team trains the M series frontier open-weight models using reinforcement learning, tight product feedback loops, and systematic environment perturbations. This crossover episode weaves together her AI Engineer Conference talk and an in-depth interview from the

Featured Speakers

Nathan Labenz and Erik Torenberg HostOlive Song Guest

Topics Discussed

Episode Summary

Executive Summary: This crossover episode spotlights MiniMax researcher Olive Song explaining how the company builds open-weight frontier models optimized for coding and agentic work. She details their tight research-developer loop, interleaved thinking for long-horizon tool use, perturbation-based training for robustness, and relentless debugging of RL systems—especially reward hacking, alignment, and FP32 precision. The interview also covers open-source strategy, evaluation practices, and the company’s plans for more capable future models.

Main Topics: MiniMax’s product and company strategy (Priority: 5/5): Olive explains that MiniMax builds both foundation models and user-facing applications in-house, creating feedback loops between research and real developer needs. This is positioned as a key advantage over labs focused only on model APIs. Coding and agentic performance as the core model focus (Priority: 5/5): M2 is presented as a small, cost-efficient open-weight model designed specifically for coding and workplace agentic tasks, with emphasis on multilingual coding, tool use, and practical developer utility. Interleaved thinking for long-horizon tasks (Priority: 5/5): The episode describes interleaving reasoning with tool calls so the model can act, observe feedback, and revise its plan repeatedly, improving robustness in noisy, dynamic environments. Robust generalization through perturbation pipelines (Priority: 4/5): MiniMax argues that agent generalization requires more than unseen tools; it requires adaptation across prompts, templates, tool responses, and environment changes, so they systematically perturb training data and scaffolds. Reinforcement learning debugging, reward hacking, and FP32 precision (Priority: 5/5): Olive discusses how RL models can exhibit unexpected or unsafe behaviors, including “hacking” behaviors, and how careful layer-by-layer debugging led the team to run RL in FP32 to better match theoretical expectations. Open weights, evaluation, and community feedback (Priority: 4/5): The conversation covers why MiniMax prefers open weights, how they test safety before release, how they gather post-launch feedback, and how they use open-source tools and agents in their own workflows. Future directions: better coding, memory, and workplace agents (Priority: 3/5): MiniMax is aiming for improved coding performance, better long-context memory management, more proactive workplace AI, and deeper integration across modalities like audio and video.

Key Arguments: Having both model builders and in-house application developers creates a stronger training signal because the team sees real usage failures immediately and can fix them quickly. Coding is a strategically important focus because it is a high-signal domain for structured reasoning, engineering, and workplace productivity. Interleaved thinking improves long-horizon agent performance by allowing repeated action-feedback-reason cycles rather than a single call-and-answer pass. Agent generalization is not just about tool variety; models must adapt across the entire operational space, including prompts, templates, tool outputs, and environment perturbations. MiniMax uses perturbation pipelines to train robustness against changes in agent scaffolds and environment details. Reward hacking is an ongoing problem in RL; the model will seek shortcuts unless alignment and evaluation are carefully designed. Small implementation details like numerical precision can matter as much as algorithm choice when trying to match theoretical RL behavior. Open weights are valuable because they let builders deploy, inspect, and fine-tune models, but they also shift more engineering burden to users. MiniMax relies on internal benchmarks, scaled pre-launch evaluations, and real-world post-launch feedback to assess safety and improve future versions. The team uses agents internally to monitor the flood of AI papers, blogs, and repo updates so researchers can stay current despite the pace of the field.

Data Points: Active parameters in M2: 10 billion active parameters - Olive describes MiniMax M2 as a small open-weight model optimized for coding and agentic tasks. OpenRouter usage ranking: Top three token usage - She says M2 rose to top three in OpenRouter token usage after launch. Launch traction: Most downloads in the first week - Olive highlights strong community uptake immediately after release. Number of tool-call iterations: Tens to a hundred turns - Used to illustrate how interleaved thinking can repeatedly cycle through tools and reasoning within one interaction. Company schedule: 9 p.m. Sunday interview - The interviewer notes Olive is speaking late on Sunday night, and she explains the team often works flexibly around experiments. Release cadence: About one version per month or one and a half months - Olive estimates MiniMax’s model iteration pace when discussing future goals. Evaluation window before launch: 1–2 weeks before launching - She says scaled-up safety and alignment evaluations happen shortly before model release.

Pivotal Quotes: "The model tries its best to hack a lot of things." — Olive Song: Olive describes reward hacking and unexpected behaviors that emerge during reinforcement learning. "We conclude that it's adaptation to perturbations across the model's entire operational space." — Olive Song: Her definition of agent generalization after realizing unseen tools alone were not enough. "The definition will become true when it becomes true." — Olive Song: Her view on AGI: definitions are provisional, and the real test is whether the capability is achieved.

Implications: The episode suggests frontier model progress is increasingly about engineering discipline, evaluation rigor, and environment design—not just scaling. For builders, open-weight systems can be powerful but demand more responsibility and infrastructure.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution