Super Data Science: ML & AI Podcast with Jon Krohn
Super Data Science: ML & AI Podcast with Jon Krohn

921: NPUs vs GPUs vs CPUs for Local AI Workloads, with Dell’s Ish Shah and Shirish Gupta

Using Windows for AI development and the bleeding edge of NPUs: Shirish Gupta and Ish Shah from Dell Technologies speak to Jon Krohn about the latest products from Dell, the future of neural-processing units (NPUs), and how AI developers can make sound hardware investments. This episode is brought t

Featured Speakers

Jon Krohn Host

Topics Discussed

Episode Summary

Executive Summary: Dell leaders Ish and Sharish argued that AI hardware choices are increasingly about matching workloads to the right compute path: Windows plus WSL2 for flexibility, CPUs for general work, GPUs for maximum scale, and NPUs for efficient on-device AI. They highlighted emerging local AI PCs/workstations, Dell Pro AI Studio, and a hybrid cloud-local future driven by privacy, cost, latency, and enterprise control.

Main Topics: Windows vs. Linux for data science and development (Priority: 5/5): Sharish argued Windows remains the dominant enterprise and dev OS, while WSL2 narrows the gap by running Linux kernels directly on Windows. Ish emphasized choice and using both when possible. NPUs and their role in on-device AI (Priority: 5/5): The guests explained NPUs as purpose-built accelerators for AI/ML workloads, especially efficient for battery-powered devices and local inference. They described integrated and discrete NPU options. GPU vs. NPU vs. CPU workload tradeoffs (Priority: 5/5): GPUs were described as the most scalable accelerator today, NPUs as the most power-efficient for AI, and CPUs as the general-purpose workhorse that still handles everything else on client devices. Future of local AI hardware and large models on devices (Priority: 5/5): They discussed new Dell and NVIDIA systems capable of running very large models locally, including a discrete-NPU laptop and GPU-powered workstations that can support frontier-scale local inference and fine-tuning. Dell Pro AI Studio and simplifying local AI deployment (Priority: 5/5): Dell Pro AI Studio was presented as a way to abstract away model-format conversion and silicon-specific toolchains, letting developers point apps to local models via an OpenAI-compatible API. Hybrid cloud-local AI strategy (Priority: 4/5): Both guests argued the future is hybrid: some workloads belong in the cloud for speed or scale, while others should run locally for privacy, cost, latency, sovereignty, and offline operation. Windows refresh and hardware purchasing timing (Priority: 4/5): With Windows 10 end-of-life approaching, they urged listeners to consider upgrading to newer AI-ready PCs and desktops that can handle modern models and future OS features.

Key Arguments: Windows remains a practical choice for many developers because it is widely used, enterprise-friendly, and increasingly compatible with Linux workflows via WSL2. Linux still dominates production ML/data science deployment, so developers should align their development environment with production where possible. NPUs are specialized for AI/ML vector math and deliver superior performance per watt, making them ideal for local AI on laptops and battery-sensitive devices. GPUs are still the most versatile and scalable accelerator for AI workloads today, especially when absolute performance or larger models matter. CPUs remain essential because they still run the operating system and general-purpose workloads; adding AI workloads without an accelerator increases CPU pressure. The local AI landscape is changing fast enough that buyers should future-proof devices now, especially if they plan to keep them for four to five years. Dell Pro AI Studio reduces deployment complexity by abstracting away toolchains and model conversion across Intel, AMD, Qualcomm, and NVIDIA targets. A hybrid architecture is the most realistic future because different workloads, compliance needs, and cost constraints demand different compute locations. Local AI can be valuable for privacy-sensitive, offline, low-latency, or token-cost-sensitive tasks such as code generation, document workflows, and sandbox experimentation. Selecting hardware now should consider not only current apps but future OS features and emerging on-device AI capabilities.

Data Points: Windows share of software developers: 64% - Sharish cited Windows as the most popular OS for software development. Large ML/data science deployments on Linux/Unix servers: 96% - Sharish noted most large ML and data science production deployments still run on Linux-based or Unix-based servers. PCs running Linux vs Windows: 60 million Linux PCs vs 1.6 billion Windows PCs - Sharish used this to argue Windows is the dominant PC platform. Linux share of PCs: less than 3% - Derived from the PC count comparison when discussing Windows dominance. NPU categories on AI PCs: 10-15 TOPS; 40-50 TOPS - Sharish described entry-level and more advanced AI PCs by NPU performance tiers. Model size on essential AI PCs: up to 1-3 billion parameters - Sharish said entry-level NPUs can handle smaller models locally. Model size on advanced AI PCs: up to 9-10 billion parameters - Sharish said state-of-the-art NPUs can support custom in-workflow use cases at this scale. Model size on Dell Pro Max Plus device with discrete NPU: 109 million parameter Llama Scout speculative decoding at FP6/FP16 - Sharish and Ish discussed a live demo on the discrete-NPU laptop. Discrete NPU local model capacity: 100+ billion parameter Llama 4 Scout - They said the device could support a 109B-parameter model via speculative decoding. Dell Pro AI Studio comparison time: 3 months vs 4 days - A Deloitte POC reportedly reduced a team’s time from scratch to a working Intel-silicon app from about three months to four days. AWS Trainium 2 instance compute: 20.8 petaflops - Sponsor segment describing Trainium 2 capabilities. AWS Trainium 2 Ultra servers: 64 chips; 83 petaflops - Sponsor segment describing a larger Trainium 2 configuration. Trainium 2 price-performance: 30-40% better - Sponsor segment comparing Trainium 2 to GPU alternatives. Blackwell RTX Pro GPU memory: 96 GB VRAM - Sharish said newer Dell Pro Max systems with NVIDIA Blackwell GPUs double VRAM to 96 GB. Local video workload speedup: ~1 week to ~3 days - A fine-tuned Flux LoRA video/style workflow reportedly improved from about a week on prior GPUs to about three days on Blackwell GPUs.

Pivotal Quotes: "Why pick one thing when you don't have to pick one thing?" — Ish: On using Windows and Linux together via WSL2 instead of treating them as mutually exclusive. "The future is hybrid, because it is not practical or feasible to move all workloads to the PC." — Sharish: On cloud versus local AI strategy and the need for optionality across compute locations. "Do not let this deadline pass you by for the love of God." — Ish: Urgent warning to upgrade before Windows 10 end of life and not delay OS refresh planning.

Implications: Listeners should evaluate AI hardware by workload, not hype: choose NPUs for efficient local AI, GPUs for scale, CPUs for general work, and cloud when it best fits privacy, cost, and speed. The market is moving quickly, so future-proofing devices now matters.

🔓 Sign Up for Unlimited Episode Search

About Super Data Science: ML & AI Podcast with Jon Krohn

View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn