The a16z Podcast
The a16z Podcast

Building the Cloud for an Agentic World | AWS CEO Matt Garman

a16z’s Raghu Raghuram sits down with AWS CEO Matt Garman to discuss how AI is reshaping the cloud, from the needs of AI-native startups to infrastructure increasingly designed for agents. Matt explains how AWS is adapting as agents write code and manage infrastructure, why it’s reserving scarce GPU

Featured Speakers

a16z HostMatt Garman Guest

Topics Discussed

Episode Summary

Executive Summary: AWS CEO Matt Garman describes AWS’s massive AI-era expansion: balancing scarce GPUs across frontier labs, startups, and enterprises; shifting infrastructure bottlenecks from chips to power, memory, and construction; and re-architecting cloud services for agentic workflows. He argues AWS’s long-standing startup focus, custom silicon, and security-first platform position it to lead as agents increasingly write code, manage infrastructure, and transform enterprise software.

Main Topics: AWS’s AI infrastructure buildout and GPU allocation (Priority: 5/5): Garman explains how AWS is scaling capacity amid extreme demand, including how it prioritizes frontier labs, enterprises, and startups while continuing to invest heavily in NVIDIA GPUs and broader infrastructure. Agentic workflows and cloud re-architecture (Priority: 5/5): The discussion centers on how AWS is adapting services for agents rather than just humans: faster account setup, context layers, sandboxes, short-lived permissions, and performance optimizations for agent-driven tasks. Startups as AWS’s strategic core (Priority: 5/5): Garman emphasizes that startups have always been AWS’s lifeblood and remain a key source of innovation, future enterprise revenue, and product feedback as they push the edge of technology. Custom silicon strategy: Graviton, Nitro, and Trainium (Priority: 4/5): AWS’s hardware strategy is framed as an evolution from Nitro offload cards to Graviton CPUs and Trainium AI chips, aimed at better performance, lower cost, and improved security/isolation. Enterprise AI adoption, trust, and evaluation (Priority: 4/5): Enterprises are using mostly non-autonomous agents today, but full adoption depends on safety, guardrails, evals, and better workflows for testing, monitoring, and trust. Data centers, power, and industry legitimacy (Priority: 3/5): Garman argues that AWS data centers bring jobs, tax benefits, and renewable energy investment, and says the industry needs to communicate these benefits more clearly while distinguishing responsible operators from irresponsible ones. Security and AI-powered defense (Priority: 3/5): The conversation covers prompt/data security concerns, model attack risk, and AWS’s AI security service (Continuum), which uses models to find and prioritize vulnerabilities at machine speed.

Key Arguments: AWS remains early in its growth journey because much compute still lives on-premises and AI is expanding demand across the stack. Startups are both a major revenue source and AWS’s best source of innovation because they reveal future enterprise needs earlier than incumbents. Agentic systems require different cloud primitives: transient resources, faster setup, context layers, fine-grained permissions, and sandboxed execution. AWS is intentionally preserving GPU capacity for startups instead of selling everything to frontier labs, because ecosystem breadth matters strategically and commercially. Bottlenecks in AI infrastructure have moved beyond chips to power, memory, construction, networking, and supply-chain components. Custom silicon is central to AWS’s value proposition: Nitro improved isolation and efficiency, Graviton reduced cost and improved performance, and Trainium is now key for both training and inference. Enterprises are not yet ready for fully autonomous agents at scale; they need help with evals, guardrails, and new operating models before deployment. AWS believes enterprise data should stay inside customer-controlled environments, which is why Bedrock keeps prompts/data within the VPC and avoids sending them back to model providers. AI is already materially accelerating AWS’s own internal software development and business operations, especially product delivery and cross-functional automation.

Data Points: AWS revenue: about $169–170 billion - Garman cites current AWS scale during the discussion of business growth AWS growth rate: 37% - Referenced while discussing AWS revenue momentum AWS planned capital expenditure: $220 billion - Amazon plans to spend this amount in capital this year/for 2026 as discussed in the episode NVIDIA GPUs planned purchase: 2 million - AWS recently announced plans to buy this many NVIDIA GPUs over the next couple of years Startup-derived revenue share: 30–40% - AWS estimates this share of revenue comes from companies that once started as startups on AWS Request fulfillment rate for capacity: ~60% - AWS says it says yes to roughly this share of GPU-related requests, often with changes in region or configuration Account setup time: less than 30 seconds - New AWS accounts can now be created quickly with simplified onboarding for agents and beginners Availability target for some databases: five nines durability - Used as an example of overengineering for transient agent-created databases versus production systems Graviton value claim: 20% cheaper and 20% better performance - Garman describes Graviton as delivering this advantage over the last five to six years Top customer adoption of Graviton: 90%+ of top 100 customers - He says the vast majority of AWS’s top 100 customers use Graviton in some form Trainium capacity status: sold out through about the end of next year - AWS says demand for Trainium remains extremely strong FDE deployment model: 45 days - AWS’s forward-deployed engineering motion is intended to train customers to operate the solution themselves within this window

Pivotal Quotes: "From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do." — Matt Garman: Explaining why AWS continues to prioritize startups even amid massive enterprise and frontier-lab demand "We really need that database to have five nines of durability." — Matt Garman: Illustrating the tension between production-grade cloud architecture and transient agent-created resources "The agents write all of the code. You're just managing a team of agents and driving that." — Matt Garman: Describing how internal software development is changing at AWS with agentic development

Implications: AWS is positioning itself as the default platform for enterprise AI infrastructure and agentic software. The winners will be cloud providers that combine capacity, security, custom silicon, and developer ergonomics while helping customers safely operationalize agents.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast