The TWIML AI Podcast
The TWIML AI Podcast

Evolving MLOps Platforms for Generative AI and Agents with Abhijit Bose - #714

Today, we're joined by Abhijit Bose, head of enterprise AI and ML platforms at Capital One to discuss the evolution of the company’s approach and insights on Generative AI and platform best practices. In this episode, we dig into the company’s platform-centric approach to AI, and how they’ve be

Topics Discussed

Episode Summary

Executive Summary: Capital One’s AI platform leader describes how the company rebuilt its ML infrastructure on AWS and is extending it into Gen AI with a platform-first approach. The discussion covers centralized governance, self-service tooling, open-source model strategy, fine-tuning, observability, inference optimization, and emerging agentic workflows, with a strong emphasis on balancing innovation, compliance, and cost at enterprise scale.

Main Topics: Capital One’s platform-first AI strategy (Priority: 5/5): Abhijit Bose explains that Capital One sees itself as a tech company that does banking, investing heavily in centralized AI/ML platforms to support shared infrastructure, governance, and reusable capabilities across the enterprise. Rebuilding ML infrastructure on AWS (Priority: 5/5): The team rebuilt the machine learning stack from the ground up on AWS, creating a control plane based on Kubernetes and integrating open source, SageMaker, and custom internal components to support production ML at scale. Extending MLOps into Gen AI (Priority: 5/5): The conversation highlights similarities and differences between traditional ML and Gen AI operations, especially around observability, hallucination guardrails, agent logging, and complex feedback loops from production back into training. Model choice, fine-tuning, and data control (Priority: 4/5): Capital One favors open-source base models such as Llama, fine-tunes them in-house for specific tasks, and hosts them fully inside its AWS perimeter to preserve data security and meet regulatory requirements. Enterprise RAG use case in customer servicing (Priority: 5/5): A detailed example shows how the company improved call-center knowledge retrieval and summarization for 20,000 human agents using vector search and fine-tuned LLMs, replacing less effective keyword-based workflows. Inference optimization and GPU economics (Priority: 4/5): The team prioritizes inference efficiency from the start, using GPUs, caching, speculative decoding, and other techniques to reduce cost per token and latency as Gen AI usage scales. Agentic workflows and talent evolution (Priority: 4/5): Bose sees multi-agent systems as the next major frontier and says Capital One is building orchestration capabilities while also redefining roles, including AI engineer and applied AI researcher, to blend technical and domain expertise.

Key Arguments: A centralized platform model is more efficient than fragmented teams because it lets Capital One reuse infrastructure, concentrate governance, and maximize expensive GPU resources. Traditional ML remains essential for core banking problems such as credit modeling, even as Gen AI expands into customer servicing, document understanding, and software engineering. Gen AI requires new operational capabilities beyond classic MLOps, especially for observability, hallucination handling, tool-call logging, and end-to-end feedback loops. Open-source models can outperform general-purpose frontier models for specific enterprise tasks when fine-tuned on proprietary data and evaluated carefully. Keeping models and data within Capital One’s AWS perimeter is critical for security, compliance, and regulatory control. Inference economics matter as much as model quality; continual cost-per-token and latency reduction is a key platform objective. Agentic workflows are likely to become a major enterprise automation pattern, but current frameworks still need platform-specific extensions for scale and governance. Talent needs are changing: successful AI teams now need people who understand business processes, prompt engineering, governance, and system design, not just software engineering.

Data Points: Time at Capital One: 4 years - Bose says he has been leading enterprise AI/ML platforms at Capital One for four years. Capital One machine learning stack rebuild: Rebuilt from the ground up in AWS - He describes a multi-year effort to modernize the company’s ML platform infrastructure. Platform adoption: Almost all company data scientists and engineers use the platform - He says the rebuilt stack is now the primary ML platform across the company. Platform user base: Thousands of users - The ML platform serves a large internal user base across Capital One. Customer servicing deployment: 20,000 human agents - The RAG and fine-tuned LLM system is used by 20,000 agents at the company. Training duration for LLMs: Weeks at a time - He contrasts LLM fine-tuning with traditional ML jobs that often run hours to a day. Traditional ML training scale: 1 or 2 GPUs - He says common ML models like GBMs/decision trees often train on one or two GPUs. Inference efficiency trend: 100x drop in inference cost - He cites an A16Z report describing rapid reductions in inference cost over two years. Gen AI roles created: 2 roles - Capital One created an AI Engineer role and an Applied AI Researcher role. Deployment coverage: Large portion of the agent community - The call-center assistant is widely deployed across the human agent workforce.

Pivotal Quotes: "We are a platform company." — Abhijit Bose: Explaining Capital One’s operating philosophy and why centralized AI infrastructure matters. "The other Thing I'll tell you in terms of like our platforms, ML Platform happens to be the platform with the highest NPS score of all platforms in the company." — Abhijit Bose: Highlighting internal user satisfaction with the ML platform. "That part of deploying Gen AI that's not always talked about, but it can get very, very complex, you know, chaining all those events." — Abhijit Bose: Describing the challenge of building end-to-end feedback loops from production systems back into model refinement.

Implications: Enterprises adopting Gen AI need more than models: they need centralized platforms, strong governance, observability, and inference optimization. The winners will likely combine open-source flexibility, domain fine-tuning, and human-in-the-loop workflows with new AI-specific roles and agentic automation.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast