Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Railway: The Agent-Native Cloud — Jake Cooper

Take the 2026 AI Engineering Survey and get >$2k in credits and AIE WF tickets! This was recorded before Railway suffered a major GCP outage on May 19, despite being a multi-AZ, multi-zone mesh ring, with HA fiber interconnects between their Metal GCP AWS, because workload discoverability was uni

Featured Speakers

Latent.Space HostJay Cooper Guest

Topics Discussed

Episode Summary

Executive Summary: Railway’s Jay Cooper argues that software deployment is shifting toward agent-driven, massively parallel, production-safe iteration. The conversation covers Railway’s evolution from open, human-centric PaaS to a vertically integrated cloud with its own metal, data centers, CLI, feature flags, rollout tooling, and internal support systems—all built to reduce friction, control cost, and make production experimentation safer and faster.

Main Topics: Railway’s product thesis: trivial deployment and evolving infrastructure (Priority: 5/5): Railway is positioned as the easiest way to ship anything, but Cooper emphasizes the broader mission: making applications and infrastructure evolve safely over time with versioning, cloning, and production-like environments. From human workflows to agentic software (Priority: 5/5): Cooper argues that agents will dominate software creation and deployment, and Railway is adapting by optimizing for CLI-driven, high-context, incremental change loops rather than human-only dashboards. Vertical integration: bare metal, data centers, and cost control (Priority: 5/5): Railway’s move from cloud dependence to its own hardware and data centers is framed as necessary to support parallel agents, improve performance, and reduce costs enough to sustain the platform’s growth. Operational scaling, pricing, and financing (Priority: 4/5): The discussion covers how Railway balances cloud bursting, owned metal margins, and debt financing to scale compute with demand, while maintaining a lean team and healthy unit economics. Safe rollout, feature flags, and production experimentation (Priority: 5/5): Cooper strongly advocates progressive rollout, feature flags, and shadow/clone environments as essential primitives for both humans and agents to reduce blast radius and improve reliability. Internal tooling and customer support at scale (Priority: 4/5): Railway’s internal systems like Central Station aggregate support, feedback, and incident data into clusters to route issues efficiently and support a growing user base with a small team. Workflow engines, Temporal/Cadence, and the future of orchestration (Priority: 4/5): Cooper praises Temporal/Cadence as powerful but complex, suggesting Railway may build its own workflow engine to better support deterministic, agentic, and production-grade workflows.

Key Arguments: Software deployment should close the loop as quickly as possible; waiting on compute is a bottleneck that should be replaced by waiting on intelligence. Agents need the same primitives humans need—versioning, observability, storage, safe rollout—but at far higher scale and with better ergonomics. Owning hardware is essential for cost and performance because cloud payback can be as short as three months versus years of depreciated hardware. Progressive rollout and feature flags should be standard everywhere, especially as systems become more agent-driven and parallel. Production should be safer to experiment with via cloning, snapshots, and copy-on-write environments instead of relying on brittle staging systems. Railway’s internal support and incident tooling is designed to scale context, not just headcount, so a 35-person team can support millions of users. The software development lifecycle is being compressed by AI, but the real opportunity is to expand it again with better primitives, not to abandon discipline. Workflow engines like Temporal are valuable but require too much hidden state in the operator’s head; future systems need clearer specs, tests, and safer state management.

Data Points: Company age: 6 years - Cooper describes Railway as having started six years ago and then entering rapid growth later. Team size: 35 people - Railway emphasizes a very lean organization despite large scale. User base: ~3 million users - Cooper says Railway is now operating at roughly three million users. New user growth: ~100,000 users/week - He says Railway is adding around one hundred thousand users each week. Free-tier losses: ~$500,000/month - During the free-tier era, Railway was losing about half a million dollars monthly. Revenue during free-tier era: ~$50,000/month - He contrasts losses with low monthly revenue during the early growth phase. Bank balance during free-tier era: ~$20 million - Cooper references having a sizable cash buffer while still running an unprofitable free-tier model. Metal margins: ~70% - Railway reports strong margins on owned metal infrastructure. Cloud payback period: ~3 months - He says cloud-based payback for metal deployments was roughly three months. Hardware depreciation window: ~4 years - He compares cloud economics to four years of depreciated hardware value. Data centers in each region: 2 per region - Railway is adding a second Singapore data center and says the pattern is two per region. Large recent incident scope: 3,000 users - A recent incident affected a subset of roughly 3,000 users. Engineering language stack: TypeScript, Rust, Go, C - Cooper lists the main languages used internally, with C for low-level networking/BPF work. Cloud footprint during constraint: 5 clouds - At one point Railway was straddling Oracle, AWS, Google Cloud, its own infrastructure, and another provider. Monthly AI spend: $300,000/month - Cooper references a current spend rate on coding agents/tools. Observed token-budget example: Uber blew its yearly token budget - He cites Uber as an example of organizations rapidly consuming AI budgets. Rollout example: 0.1% to 1% to broader rollout - He describes staged progressive deployment as a key pattern for safe releases.

Pivotal Quotes: "“You always want to be waiting on intelligence, right? And if you're waiting on compute, there's a bottleneck that needs to be destroyed there.”" — Jay Cooper: On why Railway designs for faster iteration loops and agent-driven workflows. "“The future is very, very bright. It's crazy. It's going to be nuts.”" — Jay Cooper: Closing on the long-term impact of agentic software, infrastructure, and Railway’s roadmap. "“Anything is figure-outable.”" — Jay Cooper: Describing Railway’s willingness to go deep into the stack, including kernel and data-center work, to improve user experience.

Implications: Listeners should expect software tooling to become more agent-native, more production-safe, and more vertically integrated. The industry’s winners may be the companies that combine cloud, workflow, rollout, and observability primitives into one fast feedback loop.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast