Episode Summary
Executive Summary: Gaurav Patok argued that enterprise AI success depends less on model sophistication than on data context, metadata, governance, and evaluation. He contrasted old “data brawls,” where humans disputed metrics in meetings, with today’s “garbage in, gospel out” risk, where agents confidently act on bad data. He also highlighted practical needs for AI engineers: evals/traces, context delivery, and token economics.
Main Topics: Why metadata is foundational for enterprise AI (Priority: 5/5): Patok explained metadata as the labels that let humans and agents understand what data means, making it possible to find the right trusted asset quickly instead of opening every database or file. From data brawls to silent AI failures (Priority: 5/5): He contrasted the old era of visible disagreement over metrics with the new era where agents quietly choose one answer and present it confidently, even when the underlying data is wrong. Enterprise ROI shifts from agents alone to context and governance (Priority: 5/5): Patok argued that organizations are increasingly realizing that investing only in AI models is insufficient; they must also invest in data readiness, approval controls, and governance. Sin-eaters and the hidden human cost of bad agents (Priority: 4/5): He described support teams and other employees as the people who absorb the fallout when AI agents mislead customers or make bad decisions, turning model errors into human burden. Automating data quality rules with AI (Priority: 4/5): Patok discussed how Informatica’s Clare can generate, validate, and test large numbers of data quality rules, dramatically increasing productivity compared with manual rule creation. Three critical skills for AI engineers (Priority: 5/5): He emphasized evals and traces, getting the right context to the right model, and understanding token economics as the core capabilities for practitioners in the agentic era. Security and model governance in the wild west of AI adoption (Priority: 4/5): He warned that large organizations are seeing uncontrolled model adoption, including downloading unapproved models, which raises data privacy and compliance risks.
Key Arguments: Metadata is the enterprise equivalent of labels on cans; without it, agents waste time and tokens searching across massive data estates. Bad data was once obvious because people argued about numbers in meetings, but AI agents can hide errors by confidently presenting a single wrong answer. Enterprise value from AI is mostly about context, not raw model intelligence; Patok said context is about 95% of the battle. Organizations need governance to ensure data does not go to model providers that may reuse it for training or other unwanted purposes. Many AI deployments backfire when customer-facing agents escalate too often, misroute users, or make promises the organization cannot keep. Modern AI engineers must design systems that improve through evals and traces, rather than relying only on ad hoc prompting or model upgrades. Token economics matters because inefficient retrieval and broad data access can make AI systems expensive at enterprise scale. AI can accelerate data quality rule creation by turning subject-matter knowledge into automated checks at scale.
Data Points: Years at Informatica: 13 years - Patok described his long tenure building metadata and AI products before the Salesforce acquisition. Time at Salesforce post-acquisition: about 10 months - He said he had been inside Salesforce for roughly 10 months after the acquisition. Customer data stores at one enterprise: 800,000 data stores - He used this as an example of the search complexity agents face in large organizations. Manual data quality rules written in a good week: 3 to 4 rules - Patok contrasted old manual rule-writing with AI-assisted generation. AI-generated data quality rules per day: around 200 rules - He cited Clare’s ability to generate many rules daily from prompts and uploaded files. Legacy AI engine launch: 2018 - He noted Clare was launched in 2018, before generative AI became mainstream. Context share of enterprise AI challenge: 95% - Patok said context is about 95% of the battle in enterprise AI, with intelligence only about 5%.
Pivotal Quotes: "garbage in, gospel out" — Gaurav Patok: His phrase for how AI agents can confidently present bad outputs when fed poor data. "context is about 95% of the battle" — Gaurav Patok: He used this to argue that enterprise AI success depends primarily on getting the right context and data to the model. "the sin eaters" — Gaurav Patok: His term for the humans who absorb the consequences when deployed AI systems fail or mislead customers.
Implications: Enterprises adopting agents must prioritize metadata, governance, evaluation, and cost-aware retrieval, not just model capability. AI teams that master context and data quality will build safer, cheaper, and more reliable systems.
About Super Data Science: ML & AI Podcast with Jon Krohn
View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn