The Cognitive Revolution
The Cognitive Revolution

Mamba-Palooza: 90 Days of Mamba-Inspired Research with Jason Meaux: Part 1

In this first part of a two episode series, Nathan and AI scout Jason Meaux provide a sweeping overview of the first 90 days of Mamba-inspired research. They discuss the mechanistic underpinnings of Mamba architecture, Mamba's context capabilities, multi-modal applications in image segmentation

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: This episode surveys the first 90 days of Mamba/state-space research after the original paper, arguing that Mamba has quickly become a major parallel architecture to transformers. The hosts review evidence that Mamba can learn useful internal state, excel on noisy or sparse tasks, and combine well with attention and mixture-of-experts, while also exposing limits that make hybrid designs increasingly attractive.

Main Topics: Mamba as a new core architecture alongside transformers (Priority: 5/5): The hosts frame Mamba as a genuine alternative mechanism to attention, emphasizing selective state-space dynamics, linear-time scaling, and the idea that the field is entering a 'mixture of architectures' era rather than finding a simple transformer replacement. Rapid research explosion after the original Mamba paper (Priority: 5/5): Jason Moe summarizes a burst of more than 30 papers in about 90 days, spanning vision, language, audio, graphs, and other edge cases, with many papers modifying the original architecture rather than using vanilla Mamba. Learning theory and interpretability (Priority: 5/5): Several papers show Mamba can do in-context learning, build latent board-state representations in Othello, and exhibit layerwise refinement similar to transformers, suggesting meaningful internal computation rather than simple next-token pattern matching. Strengths and weaknesses versus transformers (Priority: 5/5): Synthetic tasks show Mamba outperforming transformers on sparse/noisy settings while transformers retain advantages on associative recall and pattern re-use, supporting the claim that the two architectures have complementary strengths. Hybrid architectures as the likely near-term winner (Priority: 4/5): The MambaFormer work is highlighted as evidence that combining attention and state-space layers can outperform either alone on benchmarked toy tasks, reinforcing the episode’s broader thesis that hybrids will dominate. Mixture-of-experts extensions of Mamba (Priority: 4/5): MOE-style variants such as MOE Mamba and Black Mamba show that Mamba can be paired with sparse expert routing, improving efficiency and throughput while raising infrastructure and scaling considerations. Open questions about state and long-context memory (Priority: 4/5): The conversation returns repeatedly to what the hidden state actually represents, how it should be optimized or decoded, and whether Mamba can truly deliver durable long-context memory without state decay.

Key Arguments: Mamba’s selective state-space mechanism adds dynamic, input-dependent computation that makes it more attention-like than prior state-space models. The architecture’s linear-time inference and training are a major advantage over transformer quadratic scaling. Early research suggests Mamba can build internal latent representations of structured worlds, as shown in Othello board-state probing. Mamba and transformers appear to have complementary strengths: Mamba handles sparse/noisy inputs better, while transformers excel when later context must recover earlier patterns. Hybrid designs like MambaFormer are a practical response to these complementary strengths and may be the most performance-maximizing direction. Mixture-of-experts methods apply cleanly to Mamba and may improve efficiency, but they increase hardware and deployment complexity. A major unresolved issue is how hidden state is stored, updated, and possibly optimized directly rather than only indirectly via next-token prediction.

Data Points: Papers reviewed: well over 30 - Jason says more than 30 new Mamba/state-space papers and projects appeared in roughly the first 90 days after the original paper. Research share in vision/image processing: 60%+ - More than half of the reviewed papers focused on vision and/or image processing. Image segmentation papers: 9 - Of roughly 15 vision papers, nine were image segmentation papers. Reported state-of-the-art results: 73% - Jason reports that 73% of the surveyed papers claimed SOTA results, though not independently verified in the episode. Architecture-modifying papers: 80% - Most reviewed papers changed the original Mamba architecture rather than applying it unchanged. Transformer size in sparse parity task: 768-dim embedding, 24 layers - Used in the MambaFormer paper’s transformer baseline for sparse parity learning, yet it still failed to beat random guessing. Board accuracy in Othello GPT: 55% and 57% - Earlier transformer-based Othello GPT linear-probe results for two model sizes. Board accuracy in Othello Mamba: 67% and 71% - Mamba versions of Othello achieved higher linear-probe board-state accuracy than the transformer versions. Othello model sizes: 11M/21M parameters (Transformer), 9M/17M parameters (Mamba) - Small-scale game models used for interpretability experiments. Chess model size: 11M parameters - A Mamba chess model trained on 18.8 million games reportedly achieved a 37% win rate versus Stockfish level zero. Reported win rate vs Stockfish level zero: 37% - A small Mamba chess model’s performance in an early chess experiment. Training slowdown in Othello Mamba 2: 7.6x slower - Compared with a transformer of the same size on an A100 before batch-size adjustment. Training slowdown after batch-size change: 3x slower - Reducing batch size from 256 to 64 improved relative training speed, though Mamba remained slower. MOE Mamba expert count: 32 experts - The MOE Mamba paper scaled to 32 experts and matched original Mamba loss with 2.2x fewer training steps. Black Mamba parameter count: 2.8B parameters - Black Mamba scaled the architecture to 2.8 billion parameters with eight experts. Black Mamba throughput: ~82 tokens/sec and ~68 tokens/sec - Reported generation latency estimates for 1.5B and 2.8B parameter setups, respectively. MambaFormer evaluation coverage: all tested toy tasks - The hybrid model solved the full suite of synthetic tasks where pure transformer or pure Mamba each had failures.

Pivotal Quotes: "what is scarcest and therefore most valuable in today's world is a zoomed-out perspective that attempts to make sense of whole research subfields and new emerging market sectors" — Nathan Labenz: Opening rationale for the episode’s literature-survey format. "the Mamba architecture was beating the transformer at the core loss metric for text modeling" — Nathan Labenz: Explaining why the original Mamba paper felt like a serious challenge to transformers. "A Mamba block is literally a drop-in replacement for a self-attention block" — Lucas Nell (quoted by Nathan Labenz): Describing the ease of experimenting with Mamba in small-scale codebases.

Implications: Listeners should expect faster growth in hybrid and state-space models, especially in vision and long-context tasks. The episode suggests transformers are not obsolete, but Mamba has opened a durable new design space around memory, sparsity, and architectural mixing.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution