Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The AI Coding Factory

We are joined by Eno Reyes and Matan Grinberg, the co-founders of Factory.ai. They are building droids for autonomous software engineering, handling everything from code generation to incident response for production outages. After raising a $15M Series A from Sequoia, they just released their produ

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: Factory AI’s founders describe how a hackathon meeting led to a company focused on autonomous, enterprise-grade software development agents. They argue the future of coding is delegation, not IDE-centric collaboration, and emphasize retrieval, guardrails, evals, and workflow redesign for large, messy enterprise codebases. The discussion covers product design, model limitations, GTM, and the shift from human code writing to planning, testing, and oversight.

Main Topics: Founding story and rapid cofounder fit (Priority: 5/5): Matan and Eno met at a LangChain hackathon, discovered shared obsession with code generation, built together immediately, and decided within eight days to drop out/quit and start Factory. Factory’s enterprise-first product thesis (Priority: 5/5): Factory focuses on autonomous systems for the full software development lifecycle, especially enterprise codebases with legacy complexity, rather than flashy zero-to-one coding demos for solo developers. From IDE collaboration to delegation (Priority: 5/5): The founders argue current tools are optimized for humans writing code in IDEs, but the future is delegative: humans plan and verify while agents execute code changes, tests, docs, and incident response. Droids, workflows, and product naming (Priority: 4/5): Factory’s agents are called droids, a name chosen after a legal issue with the original company name and to avoid the overused, ill-defined ‘agent’ label; the term stuck with customers. Demo walkthrough and enterprise integrations (Priority: 5/5): They demo a browser-based platform with code, knowledge, and reliability droids, showing integrations with tools like Jira, Linear, Slack, GitHub, Sentry, and PagerDuty, plus built-in browser/code execution. Evaluation, model behavior, and retrieval (Priority: 5/5): The team relies on internal task-based and behavioral evals, not just public benchmarks, and emphasizes retrieval over brute-force context to stay cost-efficient and robust as models change. GTM, hiring, and design culture (Priority: 4/5): Factory is seeing strong Fortune 500 traction via word of mouth and is hiring technical, customer-facing operators; design is heavily influenced by the founders’ brother Cal and a strong brand system.

Key Arguments: Enterprise software development is a better fit for autonomous agents than consumer coding demos because the value is highest in ugly, legacy, high-friction codebases. The optimal interface for AI-native software development likely will not be the IDE, because the workflow is shifting from writing code to planning, delegating, and verifying. Agents need access to surrounding enterprise context—Slack, Notion, Jira, Datadog, Sentry, PagerDuty—not just code, to behave like productive engineers. Public benchmarks are useful for marketing, but internal task-based and behavioral evals are what actually protect product quality and prevent regressions. Retrieval and tool design matter more than simply increasing context windows, because cost and precision remain critical even as models get larger. Model upgrades can change behavior in surprising ways, so Factory aims to act as a shock absorber for users while still adapting to new capabilities. The biggest enterprise ROI comes from compressing timelines on large migrations and refactors, not from abstract productivity metrics like commits or lines of code. The future software engineer spends less time typing code and more time planning, testing, and supervising agent execution.

Data Points: Time from first meeting to dropping out/quitting: 8 days - Matan and Eno met at the LangChain hackathon and quickly decided to found Factory. Mutual friends at Princeton: ~150 mutual friends - They realized they had many shared connections but had never had a one-on-one conversation before the hackathon. Initial model capability referenced: GPT-3.5 - They said the early idea formed when 3.5 was out, but it was not enough for full autonomy. Context usage in demo: 43% of context size used - Shown during the live demo of the code droid working in Factory’s monorepo. Codebase age in enterprise example: 30+ years old - Used to describe the kind of legacy enterprise codebases Factory targets. Migration speed example: 4 months to 3.5 days - A large migration task for a public company was compressed dramatically using Factory. Enterprise code churn benchmark: 3%–4% - Described as typical for high-quality large-scale enterprise codebases. Poor/early-stage code churn benchmark: 10%–20% - Used as a contrast for less stable or less mature codebases. Model upgrade example: Sonnet 3.5 to 3.7 - Used to illustrate how model behavior changes can surprise enterprise users. Public benchmark cost example: $8,000–$15,000 - Discussed in relation to running benchmarks like SWE-bench. Large enterprise deployment window: Last 90 days - They said Fortune 500 deployments have been exploding over the last 90 days. Team size for migration example: 4 to 10 people - Typical human team involved in a large enterprise migration before automation. Potential role profile: 3 junior Eno-like roles - Their ideal go-to-market hires are highly technical, customer-facing operators.

Pivotal Quotes: "“It was intellectual love at first sight because basically every day since then, we've been obsessively talking to each other about AI for software development.”" — Matan / Eno: Describing the hackathon meeting that led to Factory’s founding. "“Our take is that it is very unlikely that optimal interaction pattern will be found by iterating from the optimal pattern when you wrote 100% of your code, which was the IDE.”" — Factory founders: Explaining why Factory is browser-based and not centered on the IDE. "“The biggest thing that we've seen in like, that's allowed us to deploy very quickly in enterprises is pulling in timelines on things.”" — Factory founders: On how enterprise ROI is measured through accelerated delivery rather than vanity metrics.

Implications: Factory’s thesis suggests AI coding tools will win by rethinking workflows for enterprises, not by polishing IDE autocomplete. Expect more delegation-first products, stronger retrieval, and agent observability as software teams shift toward planning and verification.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast