The TWIML AI Podcast
The TWIML AI Podcast

Is ChatGPT Getting Worse? with James Zou - #645

Today we’re joined by James Zou, an assistant professor at Stanford University. In our conversation with James, we explore the differences in ChatGPT’s behavior over the last few months. We discuss the issues that can arise from inconsistencies in generative AI models, how he tested ChatGPT’s perfor

Featured Speakers

James So Guest

Topics Discussed

Episode Summary

Executive Summary: Stanford’s James So discussed research showing ChatGPT’s behavior can shift significantly over time, sometimes improving and sometimes degrading across tasks and versions, with safety tuning and competing objectives likely contributing. He also described building a medical vision-language model from Twitter pathology discussions, arguing that public social data can be a rich source for domain AI when paired with rigorous validation and human oversight.

Main Topics: ChatGPT behavior drift over time (Priority: 5/5): So’s study compared March and June versions of ChatGPT across eight task types and found that model behavior can change substantially even without an official model change, affecting coding, math, opinion, and knowledge tasks. Evaluation methods for generative text (Priority: 5/5): The team used more objective proxies for large-scale comparison, including exact-answer questions, executable code checks, and verbosity measures, avoiding subjective freeform-text grading where possible. Performance decline in GPT-4 on math/logic tasks (Priority: 5/5): A surprising result was that GPT-4 often performed worse in June than in March on prime-number and other mathematical reasoning tasks, including reduced effectiveness of chain-of-thought prompting. Model drift differs across model families (Priority: 4/5): The direction of change was not uniform: GPT-4 got worse on some tasks while GPT-3.5 improved on the same tasks, highlighting that behavior changes can diverge even within the same product family. Surgical control and transparency in LLMs (Priority: 5/5): So argued for more precise interventions in model internals—targeting specific circuits or modules—rather than broad fine-tuning that changes billions of parameters and can create side effects. Medical vision-language modeling from social media (Priority: 4/5): His team built an OpenPath dataset from pathology-related Twitter/X threads and trained a vision-language foundation model (FLIP) to interpret medical images and text for assistance in clinical workflows. Human-in-the-loop clinical use and validation (Priority: 4/5): So emphasized that medical models should assist, not replace, clinicians, and that rigorous evaluation on expert-labeled datasets is essential because social-media-derived training data still contains noise and bias.

Key Arguments: LLM behavior is not stable over time; users should not assume the same prompt will yield the same performance months later. Objective, scalable evaluation methods are essential for comparing generative models across thousands of inputs. Chain-of-thought prompting is not universally reliable and may degrade as model behavior changes. Behavior shifts can differ across closely related models (e.g., GPT-4 vs. GPT-3.5), indicating that updates may affect models in non-intuitive ways. Safety tuning and instruction-following can create competing objectives, producing side effects on unrelated tasks. Future progress in AI control may require surgical edits to specific model circuits rather than broad fine-tuning. Public social-media discussions contain valuable, underused scientific and medical knowledge that can be harnessed for domain-specific AI. For medical applications, models should be validated on expert-labeled benchmarks and used as assistants with human oversight. Monitoring tools and task-specific robustness layers are increasingly necessary because model behavior can shift quickly and unpredictably.

Data Points: Number of task categories in ChatGPT study: 8 - So’s paper systematically compared ChatGPT across eight task types including coding, math, knowledge retrieval, and opinion questions. Questions per task: A few hundred to over 1,000 - Each task contained hundreds to more than a thousand questions for large-scale evaluation. GPT-4 accuracy drop on prime-number task: 30–40 percentage points - For prime-number classification, March performance outperformed June by roughly 30 to 40 points in some cases. Model drift time scale in LLMs: Less than 3 months - So said LLM behavior changes now occur over less than three months, compared with prior AI systems over about three years. Earlier AI monitoring window: 3 years - He contrasted current LLM drift with prior computer vision systems monitored over multi-year periods. OpenPath data scale: Several hundred thousand threads - The team curated hundreds of thousands of pathology-related Twitter/X threads after filtering for quality and language. Source dataset size description: One of the largest publicly available pathology image-text datasets - OpenPath was presented as a major public resource pairing medical images with natural-language descriptions.

Pivotal Quotes: "the size of the change and also the speed of the change is orders of magnitude larger and faster compared to the kinds of drifts and changes we saw on the previous generation of AI systems" — James So: Explaining why LLM behavior drift is especially concerning compared with earlier AI systems. "how do we identify specific circuits? And by circuit, I mean maybe a subset of the layers or neurons... that control specific behaviors" — James So: Describing the idea behind more surgical model edits rather than broad fine-tuning. "we're not envisioning using these as replacements for human pathologists, but these are more like assistants" — James So: Clarifying the intended role of the pathology vision-language model in clinical practice.

Implications: Users should expect LLM behavior to drift and build monitoring, fallback logic, and task-specific checks into applications. For medicine, public data plus expert validation can power useful assistant tools, but human oversight remains essential.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast