Large language models improve financial forecasts using alternative data
Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting
Artificial Intelligence
Summary
Forecasting how well a company will do financially can be tricky especially when data sources are limited or scattered. The authors show that big language models, which understand instructions and examples, can combine different types of information—including less common data like web traffic or consumer purchases—to make better predictions about company revenues. They created a system that picks the best data sources for each company and uses that information to forecast future earnings more accurately than traditional methods. This approach offers a flexible new way to use lots of different financial clues to help predict company performance.
What this means in practice
- •For financial analysts: Improve company revenue forecasts by combining alternative data with traditional financial information using large language models.
- •For investment portfolio managers: Make more informed investment decisions by integrating multiple data sources to predict firm financial outcomes more accurately.
Authors
Jihoon Kwon, Lawrence Liu, Daekyung Park, Sumin Kim, Haverty Jack, Hoyoung Lee, Katherine Bjorkman, Josh McKenney, Peter Laurelli, Nicole Kagan, Zach Golkhou, Thorsten Neumann, Edward Tong, Pete Petersen, Yoon Kim, Alejandro Lopez-Lira, Yongjae Lee, Chanyeol Choi
Abstract
When forecasting a firm's future financial performance, alternative data - data collected from non-traditional sources such as consumer transactions, web traffic, and prediction markets - can provide timely signals about firms' operating activities and broader market conditions. These signals may reveal information that is not captured by traditional public sources and can therefore provide complementary information for forecasting firms' future financial performance. However, firm-level alternative data often have limited historical coverage, are relevant only to specific prediction targets or subsets of firms, and are distributed across numerous heterogeneous channels, making them difficult to incorporate flexibly into conventional forecasting approaches. Meanwhile, large language models (LLMs) can interpret instructions, learn from in-context examples, and generate predictions by combining heterogeneous information without task-specific parameter updates. Motivated by this potential flexibility, we investigate whether an LLM can forecast firm performance by integrating alternative data with other financial information through in-context learning. We propose a two-agent framework that first identifies the firms for which each alternative data channel is likely to be informative and then predicts revenue using firm- and channel-specific context. We evaluate the framework across four commercial alternative data channels. In our experiments, adding alternative data in context alongside other financial information improves the LLM's forecasting relative to either source alone, and these forecasts are more accurate than those of standard forecasting baselines. These findings suggest that LLMs provide a flexible and practical approach to integrating alternative data with heterogeneous financial information.