Episode Summary
Executive Summary: The episode explores why recursion may be a major breakthrough for AI reasoning, contrasting traditional transformers with hierarchical reasoning models (HRM) and tiny recursive models (TRM). The guests argue that repeated inference-time computation can outperform brute-force scaling, especially on incompressible tasks like Sudoku, mazes, and ARC, by using hidden-state recursion and truncated backprop through time instead of only larger parameter counts.
Main Topics: Why recursion matters for reasoning (Priority: 5/5): The discussion frames recursion as a way to improve inference-time reasoning without simply scaling model size. It reconnects modern AI to RNN-style repeated computation and argues recursion enables structured multi-step problem solving. Limits of standard transformers and GPT-style models (Priority: 5/5): The speakers explain that one-shot feedforward transformers are strong at next-token prediction but limited for tasks requiring many sequential steps, external memory, or latent reasoning over long horizons. HRM architecture and training loop (Priority: 5/5): HRM is presented as a two-level recursive model with lower-level and higher-level loops plus an outer refinement loop, trained using truncated backpropagation and repeated carry-state updates. TRM simplification and performance gains (Priority: 5/5): TRM is described as a smaller, simplified successor that collapses the hierarchy into one shared network while keeping recursive refinement, improving results on benchmark tasks despite fewer parameters. Chain of thought vs inherent recursion (Priority: 4/5): The episode distinguishes token-space chain-of-thought and tool use from true latent recursive reasoning inside the model, arguing that current hacks remain bounded by human knowledge and output tokens. Memory, fixed points, and EM-like optimization (Priority: 4/5): A major thread is that recursion effectively creates a memory tape or carry state, and training resembles iterative refinement / expectation-maximization over latent states rather than standard backprop through a long unrolled sequence. Future of hybrid large models plus recursion (Priority: 4/5): The speakers speculate that the biggest gains may come from combining large general-purpose foundation models with small recursive reasoning modules rather than choosing one approach exclusively.
Key Arguments: Standard transformers are excellent at parallel training and next-token prediction, but they lack inherent latent reasoning depth and external memory, which limits performance on multi-step reasoning tasks. Recursion lets a model reuse the same weights across multiple refinement steps, creating compute depth without parameter depth. HRM’s outer refinement loop is a major driver of its success, and repeated latent-state updates can be viewed as constructing a mini-batch over memory states rather than over different inputs. Truncated backprop through time can be sufficient for these recursive models; HRM and TRM avoid full backprop through all recursion steps. TRM improves on HRM by simplifying the architecture: one shared network, fewer layers, smaller parameter count, but more effective recursion and backprop through one latent step. The models work well on incompressible tasks like Sudoku and mazes because such problems require sequential elimination or refinement that cannot be solved in a single feedforward shot. Chain-of-thought and tool use can make LLMs appear recursive, but they are still bounded by the data and token space; they do not inherently solve out-of-distribution algorithmic discovery. The most promising direction may be combining large foundation models with recursion-based reasoning modules to get both semantic richness and algorithmic depth.
Data Points: HRM parameter count: 27 million - Francois describes the HRM as a very small model trained on ARC-style tasks. TRM parameter count: 7 million - The TRM is described as a smaller follow-on model that simplifies HRM and improves performance. ARCPrize 1 performance (HRM): ~70% - The episode cites HRM reaching about 70% on ARCPrize 1 at the time, a major breakthrough relative to much larger models. ARCPrize 2 performance (HRM): state of the art - HRM is said to have achieved top results on ARCPrize 2 as well. ARCPrize baseline model performance: 0 - The speaker says a much larger model (referred to as 03) got zero on ARCPrize tasks. Relative improvement (TRM vs HRM): 70% to 87% - The episode states TRM improved ARCPrize 1 performance from about 70% to 87%. Training set size: ~1,000 tasks - HRM is described as trained on roughly 1,000 ARC-like puzzle inputs with no pretraining. Recursive refinement iterations during training: 16 - The guests mention repeated passes over the same batch/state space as part of the HRM/TRM training procedure. Backprop truncation depth: T=1 - TRM is said to find truncated backprop through time with one recursive step sufficient. Transformer depth in example code: 4 layers vs 1 layer - HRM uses four transformer layers in the discussed code; TRM reduces this to one layer.
Pivotal Quotes: "The outer refinement loop scales." — Francois Chauvard: He identifies this as the key takeaway from HRM and the main reason the approach works so well. "The cheat is the chain of thought." — Francois Chauvard: He contrasts token-space reasoning hacks with true inherent recursion in the model. "It is sufficient, not necessary, to go bigger and get better performance." — Francois Chauvard: He argues that recursion can outperform simple scaling, especially on structured reasoning tasks.
Implications: The episode suggests a path beyond pure scaling: models may gain reasoning power through recursive refinement and latent-state memory. For researchers, the likely frontier is hybrid systems that combine large foundation models with small recursive modules for harder algorithmic tasks.
From the Transcript
You're inhibited by backprop through time, largely. And this is why this paper is so exciting. Okay, so before we then go over to the TRM paper, let's just summarize here: what matters most from the HRM paper that we should take away before we transition and contrast it with the TRM paper? Yeah, I think that the number one piece to take away is this outer refinement loop. The outer refinement loop scales, and there's a great breakdown. Basically, the safety. The recipient authors, which huge kudos for this paper, because there's so many innovations in this paper, didn't really do like scaling ablations on every single one of the inputs. But this guy, Constantine, at Francis Chile's company, India, actually did. And it's this amazing breakdown that he posted on YouTube that you can go check out. But basically, the main takeaway is that the outer refinement loops is the main beneficiary, is the main
No bells and whistles. It's just a feed-forward model. And so it's just forward-passing one step and taking an input, creating a bunch of outputs. In the Sudoku case, if I have 50 different spatial steps, and it's provable that I can only do one given this information, then and I have this many layers, then that's all I can do. And the cheat is the chain of thought. And so it's completely true that at test time, they are turn-complete. And you can simulate altern computable functions at test time. But how do you get it to learn it? You need to train it. And that's where, unless you're training it on human-labeled traces, for which there's a lot of problems, like the Millennial Prize problem, we don't have the trace for it. So we'd love to have the trace for it. Doesn't it exist? Totally. Makes sense. Okay, so with that context in mind now, let's talk about these two papers. Because I think that sets up a lot of the contrast we're going to draw between these papers and the models that people are making.
Talking about this very phenomenon, which is like it is sufficient, not necessary to go bigger and get better performance, and it is sufficient and not necessary to add more recursion. And so, where I'm really excited is what happens if you do both. And you're still limited by back wrap through time. Even Alexia is limited by backwrap that last step from a memory perspective, for sure. And so, if you can make the model really big and you have lots of recursion, And we do something else other than backprop through time, then we can get all the benefits of this and all the benefits of the giant LLMs, and then you can get some crazy stuff. So, now to wrap up, why don't we talk a little bit about the bigger picture? What does this mean for the field of AI research? How should people think about where these models fit into the current span of research happening, especially given that it seems like a bit of a departure from a lot of the methods that people are used to hearing about and increasingly seeing products that people use? Well, I think for one, this from the
About Y Combinator Startup Podcast
We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.