The TWIML AI Podcast
The TWIML AI Podcast

LLMs for Equities Feature Forecasting at Two Sigma with Ben Wellington - #736

Today, we're joined by Ben Wellington, deputy head of feature forecasting at Two Sigma. We dig into the team’s end-to-end approach to leveraging AI in equities feature forecasting, covering how they identify and create features, collect and quantify historical data, and build predictive models

Featured Speakers

Ben Wellington Guest

Topics Discussed

Episode Summary

Executive Summary: Ben Wellington explains how Two Sigma’s feature forecasting turns raw, time-stamped data into trading signals, and why LLMs/GenAI are accelerating that process. The conversation contrasts older NLP approaches with embeddings and multimodal models, emphasizes the importance of historical data and leakage control in finance, and explores how agentic, model-ensemble workflows may reshape research and portfolio construction.

Main Topics: Feature forecasting as a financial workflow (Priority: 5/5): Two Sigma’s approach is to identify observable facts about companies or markets, convert them into measurable features, and use them to predict future asset prices. Historical data capture and time-stamping (Priority: 5/5): The discussion stresses that the value of data depends on having historical, time-stamped records so researchers can reconstruct what was known at each point in time and test hypotheses properly. LLMs and GenAI as a feature-creation accelerator (Priority: 5/5): Wellington argues that multimodal LLMs dramatically reduce the cost of exploring novel features—like counting nose touches in CEO interviews—from months of manual work to minutes. Embeddings and semantic generalization (Priority: 4/5): The episode traces the evolution from one-hot encoding and dictionary-based NLP to embeddings, which capture semantic relationships and create richer inputs for predictive models. Noise, weak signals, and model rigor in finance (Priority: 5/5): Because financial prediction is extremely noisy, even small model outputs can matter, but rigorous testing is essential to avoid overfitting, p-hacking, and false signals. Platform architecture, modularity, and model ensembles (Priority: 4/5): Two Sigma uses a centralized platform with clear APIs, many orthogonal models, and portfolio optimization layers; the guest sees value in ensembles and agents with different specialties. Build vs. buy and model lifecycle risk (Priority: 4/5): Wellington favors open-source and controlled experiments because off-the-shelf models can introduce temporal leakage and may become obsolete quickly as the field evolves.

Key Arguments: Feature forecasting is about collecting many kinds of observable world data and converting them into predictive signals for future asset prices. Raw data should be preserved whenever possible because future techniques may unlock value that current derived features cannot. Historical reconstruction is essential: researchers need to know what the world looked like on each past day to test whether a signal truly predicted future outcomes. Finance has an unusually poor signal-to-noise ratio, so apparently small or imperfect signals can still be valuable if tested rigorously. LLMs reduce the cost of feature ideation and extraction so dramatically that many previously impractical ideas are now feasible. Embeddings are a major leap over one-hot/dictionary methods because they encode semantic similarity and improve generalization. Temporal leakage is a severe risk in finance, because modern models will exploit any future information accidentally included in training data. Model ensembles and eventually agentic systems may outperform single models by combining orthogonal perspectives on the world. A flexible platform is necessary because model/tool quality changes rapidly, and research infrastructure must adapt quickly to stay competitive. Interpretability remains valuable, but strong empirical performance can justify black-box methods when enough evidence exists.

Data Points: Time horizon for historical data: 20 years - Wellington uses 20 years of company job-posting history as an example of the historical depth needed to test a signal. Potential ROI improvement from LLMs: Six months to six minutes - He contrasts the old cost of building a feature like nose-touch counting with the new speed enabled by LLMs. Signal threshold example: R squared of 0.05 - He says that in finance, seeing an R squared around 0.05 may indicate a model is suspiciously high because true signals are so small. Feature count scale: Thousands, tens of thousands, maybe millions - Used to describe the number of possible features that could matter for a company like IBM. Historical modeling window example: 2000–2010, then 2011 - He gives this as an example of training on past data and testing on a later period to avoid leakage. Modeling cadence: Every time cycle - He describes the platform as generating predictions for each time cycle in a loop. Innovation cycle: Increasing pace - He says innovation cycles are speeding up, requiring platforms to adapt quickly. Temporal leakage example year: 2020-trained model on 2019 document - He warns that a model trained on later data may know facts from the future when evaluating older documents.

Pivotal Quotes: "Ultimately, our goal is to find data and features and predictions in as many places as we can." — Ben Wellington: Defines the core mission of Two Sigma’s feature forecasting approach. "What’s so cool about this new era of LLMs is that six months might become six minutes." — Ben Wellington: Describes how GenAI collapses the cost and time required to test new feature ideas. "If you don’t have that historical data, a lot of data just disappears like forever." — Ben Wellington: Highlights why recording raw, time-stamped data is strategically valuable.

Implications: LLMs could radically expand quantitative research by making feature extraction and experimentation much cheaper, but success in finance will still depend on time-aware data discipline, leakage control, and rigorous validation.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast