The TWIML AI Podcast
The TWIML AI Podcast

Big Science and Embodied Learning at Hugging Face 🤗 with Thomas Wolf - #564

Today we’re joined by Thomas Wolf, co-founder and chief science officer at Hugging Face 🤗. We cover a ton of ground In our conversation, starting with Thomas’ interesting backstory as a quantum physicist and patent lawyer, and how that lead him to a career in machine learning. We explore how Hugging

Featured Speakers

Thomas Wolf Guest

Topics Discussed

Episode Summary

Executive Summary: Thomas Wolf traces Hugging Face’s evolution from a chatbot/game startup to an open-source AI research hub, arguing that NLP has hit limits with text-only scaling and needs multimodality, better data curation, retrieval, and embodied/interactive learning. He emphasizes open science, community-built datasets, and artifacts beyond papers as the path to more responsible, collaborative AI progress.

Main Topics: Wolf’s path from physics and law to AI (Priority: 4/5): He explains how a PhD in quantum physics and five years as an IP lawyer led him toward deep learning, then to Hugging Face when he saw the field’s momentum and wanted a research role centered on writing, tools, and innovation. Hugging Face’s mission: open source, open science, responsible AI (Priority: 5/5): Wolf says the company’s core direction is to promote shared research, transparent tooling, and socially aware AI, while resisting closed private models and the widening gap between academia and industry. Big Science and community-scale model building (Priority: 5/5): He details the Big Science project as a large, international collaboration using public compute to build an open multilingual language model and shared dataset, modeled more like physics collaborations than small lab training runs. Data quality, dataset governance, and evaluation (Priority: 5/5): Wolf argues that model performance depends heavily on dataset quality, deduplication, representativeness, and measurement. He stresses the need for tools and governance to inspect datasets as rigorously as models. Limits of text-only NLP and the move to multimodality (Priority: 5/5): He believes text alone is insufficient even at scale, and that future progress will come from combining language with speech, vision, embodied environments, and interactive learning to better reflect how humans learn. Retrieval and dynamic models (Priority: 4/5): Wolf discusses retrieval as a key way to update static language models with new information and to improve interpretability by showing what knowledge sources a model draws from. Research outputs beyond papers (Priority: 4/5): He emphasizes that Hugging Face values practical artifacts—datasets, models, tools, courses, reflections—not just publications, because long-term impact comes from usable resources and community infrastructure.

Key Arguments: AI progress is increasingly constrained by data quality rather than architecture alone; simply scaling text corpora is not enough. Open-source, community-based research creates healthier long-term progress than closed model development and proprietary data hoarding. Large language model development should resemble physics-scale collaborations, with shared compute, shared datasets, and distributed expertise. Dataset curation, deduplication, and measurement are underdeveloped compared with model evaluation, and need first-class tooling. Multimodality and embodied interaction are likely necessary to bridge the gap between language models and human-like understanding. Retrieval can help keep models current and make their outputs more explainable by anchoring them to external knowledge. Research should be judged by the artifacts it produces and the community it enables, not only by papers published on deadlines.

Data Points: PhD duration context: 4 years of physics experiments can be considered long; some physics experiments take 4 years - Wolf contrasts ML experimentation speed with physics to explain why he left academia Law practice duration: 5+ years - He worked as an IP attorney for startups and larger companies before entering AI Year of Hugging Face pivot: 2019 - The company shifted from chatbot/game work to open source focus Big Science participants: 1,000+ researchers signed up - Global participation in the project Daily active Big Science contributors: ~200 researchers - Approximate number actively working day to day Big Science dataset size: 800 gigabytes - Completed multilingual training dataset Big Science training duration: 4 months - Planned training time for the large model GPU hours requested: 5 million GPU hours - Compute requested from the Jean Zay public supercomputer Public compute cluster size: 3,000+ GPUs - Jean Zay cluster used for Big Science Initial weak model benchmark: 13 billion parameters - A model trained on Common Crawl performed poorly, highlighting data quality issues Model cited for comparison: GPT-3 - Used as an example of a large static model that still lacks recent knowledge like COVID Legacy model example: BERT thinks Trump is president - Illustrates model staleness and the need for updating via retrieval or other methods Book/model size: 1 billion parameters - A code model trained in the final chapter of the book Earlier model size reference: TillBERT / small models - Wolf references small efficient models alongside large ones in the book

Pivotal Quotes: "“Maybe just using text alone is not enough.”" — Thomas Wolf: He explains why Hugging Face is expanding into multimodality and other domains "“We focus more on what you’re producing, like kind of an impactful artifact out of what you do as a researcher, more than producing papers.”" — Thomas Wolf: He describes Hugging Face’s research philosophy and output criteria "“The data set in itself is actually also a super interesting artifact.”" — Thomas Wolf: He argues that datasets are valuable research outputs, not just inputs to model training

Implications: The interview suggests AI’s next gains will come from better data, open collaboration, retrieval, and multimodal/embodied learning—not just bigger text models. For the industry, this favors transparent tooling, shared benchmarks, and community-driven infrastructure.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast