Episode Summary
Executive Summary: Ish Shaw and Tyler Cox explain why agentic AI is driving token usage and cloud bills sharply higher, how sub-agents multiply work and cost, and why matching the right model to the task matters. They argue that local or on-prem Dell/NVIDIA hardware, paired with security and management software, can cut costs dramatically—often paying for itself in months—especially for software, research, and enterprise workflows.
Main Topics: What makes AI agentic (Priority: 5/5): The guests define agentic AI as LLMs that can use tools in a loop to take actions beyond chat, such as browsing, clicking, querying systems, and completing multi-step goals. Why token consumption is exploding (Priority: 5/5): They explain that agentic systems use far more tokens than chatbots because they spawn sub-agents, run parallel sessions, and keep working autonomously on complex tasks. Tokenomics and model selection (Priority: 5/5): The discussion focuses on choosing models based on the task, balancing speed, depth, intelligence, cost, and hardware fit rather than always using frontier models. Local/on-prem hardware economics (Priority: 5/5): The episode argues that moving workloads from pay-per-token cloud APIs to Dell local or on-prem systems can reduce costs by up to 93% and deliver rapid payback. Dell’s desk-side agentic AI stack (Priority: 4/5): Dell’s offering combines hardware, NVIDIA software, security guardrails, observability, and services to make agent deployment practical for enterprises. Jagged frontier and open-weight models (Priority: 4/5): They note that most tasks do not require the most advanced model and that open-weight models like Qwen, Kimi, GLM, and Nemotron are often sufficient. Enterprise adoption and use cases (Priority: 4/5): Use cases span software development, research, sales, healthcare, education, and citizen development, with software engineering showing the fastest ROI.
Key Arguments: Agentic AI differs from chatbots because it can use tools and act beyond the chat window to complete goals across systems. Token usage rises dramatically in agentic workflows because agents spawn sub-agents, run parallel tasks, and work autonomously for long periods. A model’s best fit depends on workload; speed, quality, and cost form a trade-off curve rather than a single optimal choice. Internal evals matter because public benchmarks are increasingly gamed by model vendors, so enterprises need their own tests for their own tasks. Local hardware can eliminate incremental token spend, turning AI infrastructure into a capital asset rather than an ongoing operating expense. Dell’s research indicates that many workloads can pay back hardware investments in months, especially high-volume software engineering and agentic automation. The model frontier is jagged: many business tasks do not require frontier capability, so smaller open models can be good enough and far cheaper. Security, observability, and management are essential because enterprise AI introduces new identity, access, and governance risks. As model and software ecosystems improve, the same hardware can do more over time, extending its useful life and improving ROI. Adoption is accelerating because employees bring personal AI habits into work, pressuring IT to provide governed enterprise alternatives.
Data Points: Enterprise agent adoption: Over 90% - Ish cited Signal65 and Dell experience claiming over 90% of enterprises have deployed agents in some form. Cloud frontier model cost: $50 per million output tokens - Used as an example of the cost of frontier-model usage when comparing local hardware economics. Weekend token burn: 2 billion tokens - Ish said he burned about 2 billion tokens over a weekend building a video game for his wife. Cloud-equivalent run rate avoided: About $160,000 per year - Tyler described the lab’s annual cost avoidance from running local systems instead of paying cloud token charges. Local hardware price: Over $100,000 - Approximate price of the Dell Pro Max GB300 setup discussed as the lab’s heavy-duty local AI system. Cost savings vs cloud: Up to 87% - Dell/Signal65 studies for some workloads showed savings up to 87% versus equivalent cloud execution. Cost savings vs cloud: 93% - A T2 workstation with RTX Pro 6000 Blackwell reportedly achieved 93% cost savings for a high-complexity software-assistant deployment supporting about 20 agents. Break-even period: 3 months - The first GB300 in Tyler’s lab reportedly broke even in about three months. Payback period: As little as 2 months - The episode’s framing and white-paper references say some deployments can pay for themselves in as little as two months. Model scale on higher tiers: Up to 1 trillion parameters - Tyler described larger scale-out solutions as supporting very large models and hundreds of agents. Model scale on mid tiers: Up to 500 billion parameters - The orchestrate tier was described as supporting up to roughly half-trillion-parameter models. Concurrent agents on explore tier: Up to 8 agents - A lower-tier desk-side setup was described as supporting around eight concurrent agents. Concurrent agents on orchestrate tier: Up to 40 agents - The mid-tier orchestrate setup was described as supporting around 40 agents. Concurrent agents on heavy tier: About 150 agents - The largest configuration was described as supporting roughly 150 agents. Lab token volume: Hundreds of millions of tokens per day - Tyler said the lab runs hundreds of millions of tokens daily across its systems.
Pivotal Quotes: "An agent is when you take the LLM and you give it abilities that allow it to kind of break out of the tab." — Ish Shaw: He was defining agentic AI versus a standard chatbot. "Tokens are basically the atomic unit of compute for an AI system." — Tyler Cox: He was explaining how token consumption maps to model workload and cost. "The pie of work is finite, but because you’ve got all these sub-agents in action, the pie of work might get a little bit bigger." — Ish Shaw: He was explaining why agentic systems can dramatically increase token burn.
Implications: Enterprises should stop assuming cloud tokens are the default. For many workflows, especially software and research, governed local/on-prem AI can be cheaper, safer, and easier to justify—if organizations evaluate tasks properly and pick the right model tier.
About Super Data Science: ML & AI Podcast with Jon Krohn
View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn