The TWIML AI Podcast
The TWIML AI Podcast

Watermarking Large Language Models to Fight Plagiarism with Tom Goldstein - 621

Today we’re joined by Tom Goldstein, an associate professor at the University of Maryland. Tom’s research sits at the intersection of ML and optimization and has previously been featured in the New Yorker for his work on invisibility cloaks, clothing that can evade object detection. In our conversat

Featured Speakers

Tom Goldstein Guest

Topics Discussed

Episode Summary

Executive Summary: Tom Goldstein discusses his work on AI security, focusing on adversarial examples for real-world systems, watermarking for LLM-generated text, and accidental data leakage in diffusion models. He argues that practical security risks emerge most strongly in digital systems, watermarking can help identify synthetic text with minimal quality loss, and diffusion models like Stable Diffusion can reproduce training data in surprising ways.

Main Topics: Tom Goldstein’s path into AI security (Priority: 4/5): Goldstein explains how his applied mathematics background and exposure to security researchers at the University of Maryland led him from deep learning optimization into AI security and safety. Adversarial examples in realistic industrial systems (Priority: 5/5): He describes early research on adversarial attacks against real-world systems such as content detection, object detection, and high-frequency trading, emphasizing digital-domain vulnerabilities over physical-world ones. Invisibility cloak and object detector transferability (Priority: 4/5): Goldstein recounts the clothing-based adversarial example project designed to evade object detectors, highlighting how difficult and unpredictable transfer across detection architectures can be. Watermarking LLM-generated text (Priority: 5/5): He outlines a red-list/green-list watermarking method for large language models intended to detect synthetic text with high confidence while preserving text quality. Attack resistance and removal of watermarks (Priority: 4/5): He addresses claims that watermarking is easy to remove, arguing that paraphrasing, summarization, and rewriting attacks are possible but often degrade quality or fail to fully eliminate detectability. Accidental data leakage in diffusion models (Priority: 5/5): Goldstein discusses work showing Stable Diffusion can partially copy training images, contrasting this with deliberate extraction work and emphasizing that leakage can occur unexpectedly. Scaling research infrastructure for modern AI (Priority: 3/5): He notes how the field’s hardware needs have shifted from model training to storing, loading, and searching massive datasets and large models, changing lab infrastructure requirements.

Key Arguments: AI security is most dangerous in fully digital settings because attackers can control the exact input representation, making adversarial attacks easier than in physical-world scenarios. Adversarial patterns can transfer unpredictably across detectors; success on one model does not guarantee success on another, and the reasons for transfer remain poorly understood. A watermark can be embedded in LLM output by nudging token selection toward pseudo-random green-list words, allowing post hoc detection without needing model parameters. The watermark can preserve meaning because low-entropy, highly constrained token choices are not overridden, while high-entropy regions admit watermarking with little or no quality loss. Removing a watermark is possible in principle, but fully automated removal typically causes noticeable quality degradation or fails because summarizers and paraphrasers preserve fragments. Synthetic text may create a noisier internet and motivate detection for spam, social engineering, and future training-data filtering. Stable Diffusion can reproduce training-set content by accident, not just through deliberate extraction, so privacy and licensing risks are real even when the model is used normally. Modern AI research increasingly depends on high-memory GPUs and large storage systems rather than only raw FLOPs, shifting lab infrastructure priorities.

Data Points: Adversarial pattern transfer: YOLO V2 patterns worked fairly well on YOLO V3 but not YOLO Mini - Example showing how brittle and architecture-dependent transferability can be across object detectors. Object detector scale: About 20,000 locations per image - Goldstein’s explanation of how anchor-based object detectors scan images, making them harder to fool. LLM vocabulary size: About 50,000 tokens - Typical vocabulary size mentioned for GPT-style models used in watermarking. Bloom vocabulary size: About 250,000 words - Used to illustrate large vocabularies and why many token choices can be similarly plausible. Watermark detection confidence: False positive rate of 10^-14 - Goldstein cites a 24-word example demonstrating very high-confidence detection of watermarked text. Detection span: 24 words - Length of text he says can already yield strong watermark detection in his method. Stable Diffusion leakage rate: About 2% - Goldstein reports replication/copying behavior observed in experiments on Stable Diffusion. Original study scale: 12 million images - Subset of the training data initially analyzed for diffusion-model copying behavior. Stable Diffusion training set size: 2 billion images - Scale of the original Stable Diffusion 1.0 training corpus he says they are now trying to analyze.

Pivotal Quotes: "small changes to input data that have a large effect on the output of a model" — Tom Goldstein: Defining adversarial examples during his overview of early AI security research. "the internet into a much noisier, potentially less pleasant place" — Tom Goldstein: Describing why synthetic LLM-generated content motivates watermarking and detection. "Stable diffusion behaves very differently from other models" — Tom Goldstein: Summarizing the unexpected partial-copying behavior his team found in diffusion models.

Implications: The conversation suggests AI security is shifting from toy attacks to deployment-scale risks: watermarking may become standard for synthetic text, and diffusion-model leakage will likely shape release, auditing, and data-governance practices.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast