The TWIML AI Podcast
The TWIML AI Podcast

How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765

In this episode, Rashmi Shetty, senior director of enterprise generative AI platform at Capital One, joins us to explore how the company is designing, deploying, and scaling multi-agent systems in a highly regulated environment. Rashmi walks us through Chat Concierge, a multi-agent chat experience f

Featured Speakers

Rashmi Shetty Guest

Topics Discussed

Episode Summary

Executive Summary: Rashmi Shetty of Capital One explains why multi-agent systems are the next step after classic ML and LLMs: they break complex, goal-oriented tasks into orchestrated steps, add governance and human oversight, and enable safe action-taking at enterprise scale. She outlines Capital One’s risk-first platform, observability, evals, specialization, and human-in-the-loop design using Chat Concierge as a concrete deployment example.

Main Topics: Capital One’s AI evolution from ML to agentic systems (Priority: 5/5): Shetty traces her path from pervasive computing and enterprise ML/AutoML to GenAI and now agentic systems, framing agentic AI as a natural closure of the loop from decisioning to action. Why multi-agent architectures are needed (Priority: 5/5): Multi-agent systems are positioned as the right approach when a complex problem must be decomposed into smaller tasks, with separate agents handling intent, planning, actions, validation, and response refinement. Chat Concierge as the beachhead use case (Priority: 5/5): Chat Concierge is presented as Capital One’s first multi-agent deployment for auto dealers, designed to match customers to vehicles, ask clarifying questions, and support actions like test drives with human interloop support. Risk, governance, and regulated-environment design (Priority: 5/5): Because Capital One operates in banking, agentic systems must align with model risk, compliance, cyber, and policy controls. Governance is embedded at platform and agent levels, especially during runtime. Developer experience, tooling, and platform abstractions (Priority: 4/5): The platform is built to help developers rapidly build and deploy agents safely using SDKs, frameworks, APIs, blueprints, and self-service tooling while hiding infrastructure complexity. Observability, evals, and closed-loop learning (Priority: 4/5): Agentic systems require end-to-end observability across tools, reasoning, context flow, latency, and business outcomes. Evals are also end-to-end, using golden datasets and production telemetry to continuously improve systems. Model specialization and scaling strategy (Priority: 4/5): Capital One emphasizes specialized models, fine-tuning, and distillation to support personalization and latency control. The next phase is scaling agentic systems across millions of users and multiple business lines.

Key Arguments: Agentic AI is a system problem, not just a model problem; it must be designed end-to-end for planning, acting, observing, and learning. Multi-agent architectures are appropriate when one complex objective must be decomposed into multiple bounded sub-tasks handled by specialized agents. In regulated industries, safety comes from policy-bound operations, guardrails, permissions, and evaluation gates embedded in the platform. The platform should abstract infrastructure decisions so developers can focus on the agent layer while still getting governance, observability, and deployment support. Observability must span the full workflow, not just individual models, because agent behavior, tool usage, context passing, and latency are all interdependent. Evals for agents must be end-to-end; evaluating a single agent in isolation is insufficient if the whole chain fails. Capital One’s advantage comes from data, governance, and cloud-native foundations already in place, which make specialization and secure deployment easier. Human-in-the-loop is essential for high-risk scenarios and should be built into orchestration and operating workflows rather than bolted on afterward.

Data Points: GenAI mission start: early 2023 - Capital One began its generative AI journey in early 2023. First agent assist pilot live: 2023 - The first pilot, an agent assist pilot, went live in 2023. AI/ML journey duration: past decade - Capital One has been in the AIML journey for about ten years. GenAI/agentic presence: past two years - Shetty says Capital One has had a GenAI and agentic presence over the last two years. Experimentation cycle time: days, not weeks or months - Shetty notes that experimentation has sped up dramatically in the agentic era. Deployment target: millions of customers - Scaling Chat Concierge and other agents to massive customer volumes is a stated next step.

Pivotal Quotes: "We moved from a Classic ML world to a world where we have LLMs generating responses. And now we want to move on to a world where actions need to be taken." — Rashmi Shetty: Opening explanation of why agentic AI follows ML and GenAI. "We are a tech organization that does banking" — Rashmi Shetty: Describing Capital One’s identity and why it can combine modern engineering with heavy governance. "The purpose of the platform is to abstract away any complex underlying technical decisions that have been taken." — Rashmi Shetty: Explaining how the platform supports developers building agentic systems.

Implications: For enterprises, especially regulated ones, agentic AI requires platformized governance, observability, evals, and human oversight. Success depends less on model novelty and more on safe orchestration, specialization, and scaling production workflows.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast