The Cognitive Revolution
The Cognitive Revolution

The Dawn of Dynamic AI: RFT Comes Online, w/ Predibase CEO Dev Rishi, from Inference by Turing Post

This crossover episode from Inference by Turing Post features CEO Dev Rishi of Predibase discussing the shift from static to continuously learning AI systems that can adapt and improve from ongoing user feedback in production. Rishi provides grounded insights from deploying these dynamic models to r

Featured Speakers

Nathan Labenz and Erik Torenberg HostDev Rishi Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that continuous learning is already emerging in enterprise AI, with companies moving from one-time model training to production feedback loops using reinforcement fine-tuning, user data, and live evaluations. Dev Rishi frames the future as practical, specialized intelligence rather than AGI, emphasizing open source, inference optimization, and tight product-market-fit-driven deployment as the real path forward.

Main Topics: Continuous learning and train-once-learn-forever (Priority: 5/5): Dev argues this is already happening in early form: companies start from a pretrained model and continuously improve it in production using feedback loops and live traffic. Reinforcement fine-tuning (RFT) as a customization tool (Priority: 5/5): RFT lets teams improve models with tiny datasets and reward functions instead of large labeled corpora, and Dev sees it evolving from one-off tuning into continuous online learning. Production feedback loops in enterprise use cases (Priority: 5/5): Healthcare examples show how clinician feedback, LM judges, and production conversations can feed model improvement pipelines with far less labeling than traditional ML. Inference as a production systems problem (Priority: 4/5): Dev explains that inference becomes difficult at scale due to GPU procurement, SLAs, fault tolerance, deployment updates, multi-region resilience, and cost optimization. Open source and the changing model landscape (Priority: 4/5): He argues open source is now competitive with commercial models and that the AI stack needs stronger evaluation tooling and better pathways from experimentation to managed production. Specialized intelligence over AGI (Priority: 4/5): Dev says the real future is not one model that rules all, but many narrow, task-specific systems that deliver economic value in enterprise and consumer settings. Product strategy in a fast-moving AI market (Priority: 3/5): Predibase’s strategy is to stay anchored to specialized AI, while adapting quickly as research shifts across tuning methods, inference techniques, and modalities.

Key Arguments: Most production AI today is not full model training but last-mile customization of someone else’s pretrained model. RFT is powerful because it replaces large labeled datasets with small examples plus reward functions, enabling measurable improvement from objective criteria. Continuous learning pipelines are beginning to appear in cutting-edge healthcare deployments using clinician input and production user interactions. Agentic workflows are promising but brittle because multi-step systems compound small errors; reliability matters more than flashy demos. Inference is hard at scale not because the initial prototype is difficult, but because production requires resilience, monitoring, throughput optimization, and strict SLAs. Open source models have advanced faster than expected and are now on par with or better than leading commercial models in some benchmarks. Evaluation remains an open problem; most companies rely on in-house methods, LLM-as-judge, historical holdouts, and product feedback rather than a standardized framework. The most likely future is a mix of open and closed models of different sizes, chosen per task, rather than a single dominant model. AGI is less useful as a planning lens than practical specialized intelligence that drives real business productivity. A major risk is hype outpacing real value, especially when teams chase impressive demos instead of high-ROI enterprise use cases.

Data Points: RFT data requirement: ~dozen examples - Dev says reinforcement fine-tuning can work with really small quantities of data, around a dozen examples. Labeling time reduction: months of labeling -> handful of conversations - In healthcare, feedback from a few conversations can replace months of traditional labeling work. LLM pipeline accuracy compounding: 90% accuracy over 5 calls -> sub-50% UX - Dev uses this example to show why multi-step agentic systems become brittle quickly. GPU footprint for large model replica: 8 or 16 H100s - He cites this as an example of the hardware needed for a single deployed replica of a large model. Production uptime target: 99.9% to 99.999% SLA - Production inference requires far stricter reliability than a prototype. Throughput improvement: 2x - Predibase’s TurboLoRa is cited as software-defined optimization that can increase throughput by roughly two times. Current model rollout timing: End of 2023 / beginning of 2023 pivot - Dev describes Predibase’s pivot to LLM specialization and post-training stack after the ChatGPT surge. Open source benchmark timing: ~6 months ahead of schedule - He says open source surpassing commercial models in some benchmarks happened earlier than he expected.

Pivotal Quotes: "The world that I see tends to be like, rather than artificial general intelligence, like practical, specialized intelligence." — Dev Rishi: On why he focuses on enterprise value and narrow AI rather than AGI narratives. "If you measure it, you can improve it." — Dev Rishi: Explaining the core intuition behind reinforcement fine-tuning and reward-based customization. "generalized intelligence is great, but I don't need my point of sale system to recite French poetry." — Dev Rishi: On why enterprise AI demand centers on narrow, task-specific usefulness rather than broad generality.

Implications: Enterprise AI is likely to shift toward continuously improving, task-specific systems powered by live feedback, better evaluation, and optimized inference. Teams that build tight data-feedback loops and choose the right model per task will likely outperform those chasing generic AGI-style demos.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution