Episode Summary
Executive Summary: John Egan explains how Kentaba was built to bring order to incident response by making it company-wide, low-friction, and collaborative. Rooted in lessons from Facebook and Google’s SRE culture, the product emphasizes real-time coordination, postmortems, and resilience over rigid data entry. He also reflects on startup lessons about vision, hiring, scale, shipping safely, and building a broader incident-management community.
Main Topics: Origin of Kentaba and the incident management opportunity (Priority: 5/5): Egan traces Kentaba to the growing recognition, sparked by Google’s SRE handbook and Facebook’s internal practices, that incidents are inevitable and need structured organizational response tooling. Product philosophy: low-friction, company-wide incident response (Priority: 5/5): Kentaba is designed as a real-time collaboration and orchestration platform for the whole company, not just SREs, so incidents can be surfaced early and handled transparently. MVP decisions and early product trade-offs (Priority: 4/5): The initial MVP focused on a dashboard and collaborative incident space, intentionally avoiding heavy structured forms to keep filing incidents easy and encourage adoption. Roadmap built from long-term vision plus customer feedback (Priority: 5/5): Egan argues that customer requests should be translated backward from a durable company vision, especially in early-stage SaaS when resources are limited and timing matters. Team building, hiring, and startup flexibility (Priority: 4/5): He emphasizes small teams, flexible co-founders, and cautious hiring, noting that startup success depends as much on human fit and adaptability as on technical skill. Scaling and operational maturity (Priority: 4/5): Kentaba was built to scale from day one using Node, Next.js, and Kubernetes on Azure AKS, though the company had to improve release discipline and reduce fragility as it shipped quickly. Future of incident management and community building (Priority: 5/5): Egan sees the future of incident management as broader than technical ops, extending to sales, PR, legal, and the entire organization, supported by community efforts like IRConf.
Key Arguments: Incident management should assume incidents will continue forever; the goal is not incident zero but faster, better response and learning. A low-friction incident tool increases reporting, which helps teams catch problems earlier at SEV3/SEV2 instead of waiting for catastrophic outages. Incident response is a human and organizational practice, not just a technical SRE function. Product roadmap decisions should be derived from long-term company vision, then mapped backward from customer requests. Early-stage startups need to optimize for speed and learning, but not at the expense of product stability and customer trust. Hiring in startups is primarily about flexibility, alignment, and trust; technical ability alone is insufficient. Building on modern cloud infrastructure reduces scaling pain, shifting attention from infrastructure to company growth, sales, and go-to-market. Incident management can and should be useful across the whole organization, including non-technical functions like PR, legal, and sales.
Data Points: Google SRE Handbook publication year: 2016 - Egan cites this as a pivotal moment when incident management best practices became mainstream. Start year of Facebook to post-Facebook transition: 2017–2018 - He says he left Facebook around this period and then began discussing missing internal tools with former colleagues. Company size supported by Kentaba: 3 people, 3000 people, 30,000 people - Egan describes the off-the-shelf solution as intended to deliver value across organizations of many sizes. MVP core features: 2 - The first product included a dashboard and a collaborative incident space. Typical shipping cadence: about once a week - He says the team eventually found weekly shipping to be the sweet spot. IRConf date: April 1st - He notes the incident responder conference is scheduled for April 1 and is not an April Fool’s joke. Conference focus: first incident responder focused conference - IRConf is described as the first conference centered on incident responders. Social collaboration tools integrated early: Slack, Twilio - He mentions integrating with Slack for collaboration and Twilio for email/SMS notifications. Targeted incident severity level: SEV3 and SEV2 - He argues the goal is to catch incidents earlier, before they become major outages. Number of co-founders mentioned: 3 - He says the company was started by three people, including Zach and Cole.
Pivotal Quotes: "What's actually more important is having that longer-term vision for what you want your company to be and applying that backwards to the customer feedback." — John Egan: Explaining how to build a roadmap from customer requests without losing the company’s strategic direction. "This is incident management. It isn't actually a technical practice. It's a human practice." — John Egan: Describing Kentaba’s broader mission beyond SRE and engineering teams. "We're really proud of ... just how successful it's been at getting companies engaged in the process of incident management and keeping them there." — John Egan: Reflecting on the product outcome he values most: adoption and retention of incident practices.
Implications: The episode suggests incident management is evolving into an organization-wide discipline. Products that reduce friction, support collaboration, and reinforce learning may drive broader adoption than tools designed only for engineering.
About Code Story
Code Story is a podcast featuring startup founders, tech leaders, CTO's, CEO's, and software architects, reflecting on their human story in creating world changing innovation, disruptive digital products. Their tech. Their products. Their stories.