Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Training a SOTA Code LLM in 1 week and Quantifying the Vibes — with Reza Shabani of Replit

Latent Space is popping off! Welcome to the over 8500 latent space explorers who have joined us. Join us this month at various events in SF and NYC, or start your own! This post spent 22 hours at the top of Hacker News. As announced during their Developer Day celebrating their $100m fundraise follow

Featured Speakers

Latent.Space HostReza Shabani Guest

Topics Discussed

Episode Summary

Executive Summary: Reza Shabani, head of AI at Replit, discusses his path from quantitative finance and econ research into building AI products, then dives deep into Replit’s Ghostwriter and its new open-source code model. The conversation centers on data infrastructure, model evaluation beyond standard benchmarks, why code models can generalize to reasoning, and how Replit is moving from code completion toward autonomous software-building agents.

Main Topics: Reza’s path from econ/quant finance to AI (Priority: 5/5): Reza explains how his PhD research and quant trading work forced him to learn Python and build data systems, which later made him a strong fit for Replit’s data and AI work. Building Replit’s data infrastructure first (Priority: 5/5): He describes joining Replit as the first data hire and spending the early period modernizing pipelines so the company could query and process massive product and usage data at scale. Ghostwriter’s evolution from code completion to software agents (Priority: 5/5): The team started with a code-completion model, but the long-term goal is an autonomous system that can scaffold projects, edit files, run code, and help build software end-to-end. Model evaluation: human eval vs. 'Amjad eval' (Priority: 5/5): Reza contrasts standard benchmarks with Replit’s internal vibe-based evaluation, arguing that real product usefulness, latency, and contextual correctness matter more than benchmark scores alone. Open-sourcing Replit’s code model (Priority: 4/5): Replit is releasing Replicate/Replicode-style code models trained on large-scale permissively licensed code, plus a custom tokenizer, to encourage experimentation and downstream use. Scaling laws, overtraining, and the 'YOLO' training run (Priority: 4/5): The team describes a high-risk, high-reward training run that used far more data than expected, challenging assumptions like Chinchilla-style token/parameter ratios and improving results. AI’s broader impact on work and society (Priority: 4/5): The discussion closes on how AI is moving faster than expected, will reshape products beyond chat, and requires people to learn how to use it rather than fear it.

Key Arguments: Quant finance and econ research are highly technical, data-heavy disciplines that naturally prepare people for AI/data engineering work. Replit needed data infrastructure before it could build serious AI products; querying and processing product data at scale was a prerequisite. Standard benchmarks like HumanEval are useful but insufficient; product-specific evaluation captures real-world quality better. A model can score well on benchmarks yet still have poor 'vibes'—e.g., bad latency, awkward completions, or wrong contextual behavior. Code models can learn useful reasoning and even natural-language instruction-following from code and code-adjacent data. The best AI products will move beyond text completion into workflow automation and agentic software creation. Open-source models plus permissive code data can produce strong, fast, customizable systems without relying on a single vendor. Training on more data than conventional scaling laws suggest may improve code models, at least in Replit’s experience. AI adoption will be most valuable when it augments users and is embedded into real products, not just chat interfaces.

Data Points: PhD dissertation data collection: 10 hours/day of CNBC recordings - Reza recorded financial news continuously for his dissertation research. Historical research date: 2009 - He said the dissertation work was being done in 2009, before cloud and many modern Python tools. Model size: 2.7 billion parameters - Size of the first open-source code model Replit is releasing. Training tokens: 525 billion tokens - Amount of permissively licensed code used to train the base code model. Tokenizer vocabulary size: ~32,000 - Custom tokenizer vocabulary trained from scratch for code, smaller than the prior 50,000-ish vocabularies. Replit public dataset size: ~230 million repls - Scale of public Replit content used to motivate the need for data infrastructure. Fine-tuned model improvement: 20% to 30% - Reza described a 50% relative improvement on a common-sense benchmark after fine-tuning on Replit data. Training compute: 256 GPUs - The 'YOLO' training run used a large multi-GPU setup. Earlier model training scale: 1.3B parameter models on 20B-30B tokens - Prior experiments before the larger run. Replit data fine-tuning scale: ~680 billion tokens - Amount of Replit data seen in the fine-tuned model run. Company size at join: ~30 people - Reza joined when Replit was much smaller and became the first data hire. Current company size: ~90 people - Approximate Replit size mentioned during the interview.

Pivotal Quotes: "“We want Ghost Rider to be like an autonomous agent that can actually drive the IDE.”" — Reza Shabani: Describing the long-term vision beyond code completion. "“It logs five, five.”" — Reza Shabani: Explaining a code-completion example where the model correctly predicted JavaScript closure behavior. "“Learn how to use it, learn how it can help you and benefit you.”" — Reza Shabani: His closing advice on AI adoption and productivity.

Implications: The episode suggests AI coding tools are shifting from autocomplete to agentic software creation, and that companies need strong data infrastructure plus product-specific evaluation to win. It also argues that users and builders should treat AI as a practical skill, not just a novelty.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast