Language model competence improves event forecast accuracy selectively

Competence-Gated Pooling of Language Models and Priors for Event Forecasting

Artificial Intelligence

Summary

Forecasting the outcomes of events is often done by combining various prediction sources like markets or crowds. The researchers studied how language models can add value when combined with these other forecasts, focusing on when to trust the model's information. They developed a method that measures how competent the language model is in a specific domain and uses that to decide how much weight to give its predictions. This approach improved forecast accuracy on many tasks by using the model selectively rather than always or never. They also found that simply using the language model's verbal confidence was less reliable for deciding when to listen to it.

What this means in practice

  • For financial analysts: Enhance market event predictions by selectively integrating language model forecasts based on measured competence relative to existing market signals.
  • For risk management teams: Improve binary event risk assessments using a competence gate to combine language model insights with statistical and crowd-based forecasts.

Authors

Aditi Tiwari, Aashrith Bandaru, Heng Ji

Abstract

In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED. In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.