Episode Summary
Executive Summary: Deb Raji traces her path from robotics and computer vision into AI auditing, explaining how Gender Shades exposed severe racial/gender performance gaps in facial recognition and helped reshape evaluation practices. The conversation argues that audits are necessary but insufficient: they can reveal bias, but not broader issues like privacy, process ethics, or harmful use by governments and employers. Raji makes the case for stronger regulation, transparency, and even moratoriums on facial recognition in high-risk contexts.
Main Topics: Deb Raji’s path into AI fairness research (Priority: 5/5): Raji describes entering engineering at the University of Toronto, learning to code, working in robotics/computer vision, and becoming aware of bias while interning at Clarifai. A Joy Buolamwini TED talk connected her with the Algorithmic Justice League and launched her work on auditing and documentation. What AI audits are and why Gender Shades mattered (Priority: 5/5): She explains audits as tests of deployed systems in real-world contexts, using balanced benchmark datasets to evaluate performance across demographic subgroups, especially at intersections like darker-skinned women. Gender Shades became a landmark because it exposed major disparities in commercial facial recognition APIs. Audit results, vendor response, and limits of benchmarking (Priority: 5/5): The first Gender Shades audit found large accuracy gaps among IBM, Microsoft, and Face++; later follow-up studies showed some improvement for audited vendors but persistent disparities elsewhere and in other tasks. Raji stresses that companies can optimize for benchmarks without solving the underlying problem. Process-oriented accountability and standards (Priority: 4/5): Raji argues that standards should not only assess outputs but also the process: data provenance, dataset creation, labeling choices, ethics review, transparency, and community consultation. She distinguishes technical performance from governance and urges documentation frameworks such as model cards and data sheets. Facial recognition harms, surveillance, and misuse (Priority: 5/5): The discussion broadens from accuracy to real-world harms: police misuse, immigration screening, employer scoring, landlord surveillance, and covert mass identification systems like Clearview AI. Raji emphasizes that facial recognition is a biometric with fingerprint-like sensitivity but far weaker governance. Moratoriums, regulation, and the future of facial recognition (Priority: 5/5): Raji supports pausing deployment while society develops better regulation and public oversight. She argues companies cannot claim to support regulation while continuing to sell the technology, and that government and market pressure are needed to protect vulnerable groups.
Key Arguments: Gender Shades showed that facial recognition had severe, measurable racial and gender performance disparities, proving the technology was not “working” equally for everyone. Audits are useful for revealing failures in deployed systems, but they cannot answer every concern; they are limited when the harms involve privacy, surveillance, or misuse rather than simple model error. Benchmark improvement alone is not meaningful if vendors are only optimizing for the audit test while leaving other tasks, populations, or process harms unaddressed. Evaluation must include intersectional subgroups, not just single-axis categories, because harm often concentrates at intersections such as darker-skinned women. Standards should assess not only outputs but also the entire development process, including data provenance, collection ethics, taxonomy design, and governance practices. Facial recognition is especially dangerous because it operates on biometric face data that can be centrally collected and repurposed by authorities without people’s knowledge. Use cases in policing, immigration, employment, and housing show that harm often comes from who controls the technology and how it is used, not only from algorithmic bias. A moratorium is justified because the technology remains immature and its social risks are broader than what current audits or standards can capture. Companies’ public support for regulation is not enough if they continue selling facial recognition at scale; real protection requires external regulation and disclosure requirements.
Data Points: Data set imbalance in computer vision benchmarks: 80% to 95% lighter-skinned subjects - Raji describes the training/evaluation datasets she encountered at Clarifai as heavily skewed toward lighter-skinned faces. Gender Shades disparity: ~30% accuracy gap - Difference between darker-skinned female and lighter-skinned male subgroups in the initial audit. Darker-skinned female subgroup accuracy: ~60% - Performance level reported for darker-skinned women in the first Gender Shades study. Follow-up vendor response time: Within 7 months - Audited companies re-released new APIs/models after the initial public findings. Public facial recognition vendors audited first: 3 companies - Initial Gender Shades audit focused on IBM, Microsoft, and Face++. Company set with persistent disparities after follow-up: Amazon plus other unaudited vendors - Follow-up showed unaudited vendors still exhibited up to 30% disparity. NIST demographic subgroup evaluation adoption: Last year / very recently - Raji says NIST only began evaluating demographic subgroup performance recently. Facial recognition threshold mentioned by vendors: 95% then 99% - Amazon’s public rebuttal emphasized high confidence thresholds for law enforcement use. Default vendor threshold: 80% - Raji notes that some police clients were unaware of the threshold setting and that 80% was the default. Police department partnerships referenced: Over 3,000 - Amazon/Ring had partnerships with thousands of police departments at the time discussed. Year of government push to fund facial recognition research: 1996 - Raji references government funding that helped propel the facial recognition research community.
Pivotal Quotes: "There are actually very clear limits to these kinds of audits." — Deb Raji: She explains why audits are valuable but insufficient for capturing all facial recognition risks. "Maybe this technology that's like very clearly immature in certain ways needs to be taken off the market." — Deb Raji: Her argument for a moratorium while society develops stronger rules and oversight. "It’s not the algorithm, it’s the data set." — Deb Raji: Used to criticize oversimplified arguments that blame only one part of the system and ignore broader engineering and process choices.
Implications: The episode argues that facial recognition should be treated as a high-risk biometric surveillance tool, not a neutral product. For industry and regulators, the lesson is to move beyond benchmark fixes toward process accountability, disclosure, and likely restrictions or moratoriums in sensitive settings.