The a16z Podcast
The a16z Podcast

a16z Podcast: When Large Scale Gets Really Massive -- Managing Today’s Enterprise Networks

Managing enterprise networks with thousands of users and endpoints has been hard enough. Now that large enterprise networks routinely include hundreds of thousands of nodes it’s amazingly difficult and time-consuming (we’re talking days often) to ge...

Featured Speakers

a16z HostOrion Hindawi GuestStephen Sanofsky Guest

Topics Discussed

Episode Summary

Executive Summary: The conversation explains why legacy enterprise management and security tools fail in a world of fast, targeted attacks and massive networks, and how Tanium’s real-time, linear peer-to-peer architecture solves that by querying endpoints directly at scale. Stephen Sanofsky and Orion Hindawi describe how the platform can return near-instant answers, remediate issues, and serve as a unified API for managing and securing everything from PCs to IoT devices.

Main Topics: Why legacy enterprise tools broke down (Priority: 5/5): BigFix-era client/server and database-driven systems were built for slower, broad-based patching problems, but modern APTs and insider threats move too quickly for delayed inventory and response cycles. Tanium’s architecture and speed (Priority: 5/5): Tanium was rebuilt from scratch to answer questions across huge networks in seconds by reaching endpoints directly and aggregating results locally over the LAN before sending one response back over the WAN. Real-world demo and customer reaction (Priority: 4/5): Sanofsky describes the shock of seeing live queries run against a real hospital network, proving the product was not a mock-up and could answer highly specific questions immediately. Security and remediation use cases (Priority: 5/5): The product is used to identify exposures like Heartbleed, locate vulnerable software, quarantine machines, turn on firewalls, and stop services in real time rather than after the fact. Why real-time visibility matters (Priority: 4/5): Both speakers argue that stale data makes organizations operationally and security-wise blind; minutes-old data can already be wrong in a rapidly changing environment. Future beyond endpoints: IoT and billions of devices (Priority: 4/5): They envision extending the same architecture to massively distributed devices such as watches, light bulbs, POS systems, and sensors, where hub-and-spoke approaches will not scale. Tanium as a platform/API (Priority: 3/5): Beyond the browser interface, Tanium is described as a giant API that can feed dashboards, charts, and custom tooling for continuous monitoring and analysis.

Key Arguments: Legacy systems management and security tools were designed for untargeted, slower threats and cannot cope with today’s sophisticated, targeted attacks. The only way to get meaningful visibility across hundreds of thousands of endpoints is to query endpoints synchronously and aggregate results locally. Hub-and-spoke architectures introduce too much latency and infrastructure overhead; distributed peer-to-peer communication is fundamentally more scalable. Real-time data is essential because in modern environments even minutes-old information can be inaccurate and operationally dangerous. Tanium is not just for visibility; it can also trigger remediation actions immediately after identifying a problem. As enterprises and IoT environments scale toward billions of devices, the need for a new architecture becomes even more urgent. The product’s natural-language interface and API make it usable by operators and developers for both ad hoc questions and continuous automation.

Data Points: Data freshness at scale: 15-second old data - Tanium can retrieve near-real-time information across 500,000-seat networks. Network scale: 500,000 seat networks - Used to illustrate the scale at which Tanium can still deliver rapid answers. Attack response evolution: minutes instead of hours or days - Modern attacks and outages now occur and resolve on much faster timescales than legacy tools were built for. Traditional turnaround: 3-5 days - Sanofsky describes how long older systems like BigFix could take to inventory very large networks. Load on endpoint: 0.1% CPU - Tanium’s query execution overhead on endpoints is described as extremely low. Installer/runtime footprint: 2 meg installer, 7 meg RAM, 10 meg disk - Shows the lightweight nature of the endpoint agent/runtime. Question-response time: 15 seconds - Example response time for queries such as locating USB writes or listing affected OpenSSL machines. Government customer endpoint range: 150,000 to 300,000 PCs - Illustrates the scale and uncertainty of enterprise asset counting.

Pivotal Quotes: "We needed to be able to do that in seconds." — Orion Hindawi: Explaining why Tanium was founded to replace slow, legacy inventory and security systems. "How many of the PCs have a USB memory stick in them and are currently writing to it?" — Stephen Sanofsky: A deliberately difficult question used to test the live system during the demo. "If you don't have real time data, you're basically always playing whack-a-mole." — Orion Hindawi: Summarizing why stale information is inadequate for modern security and operations.

Implications: The discussion suggests enterprise security and IT operations must shift to real-time, distributed architectures. Organizations that keep relying on delayed inventories and siloed tools will remain unable to detect, triage, or remediate fast-moving threats.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast