LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks

2026-07-27Computation and Language

Computation and Language
AI summary

The authors compare two ways of analyzing emotions from social media posts about certain stocks that experience big price swings. They look at a simple word-based method (VADER) and a complex AI model that can detect feelings like excitement or sarcasm. Their findings show that the AI model captures more detailed emotional signals related to stock discussions, but these signals don’t always predict stock price changes reliably. This means that while smarter language analysis gives richer information, it doesn’t always help forecast market moves better in these volatile stock environments.

sentiment analysislexicon-based modelLarge Language Model (LLM)reddit r/WallStreetBetsmeme stocksVADERmarket returnsbullishnesssarcasm detectionquantile regression
Authors
Paul Kilian, Markus Kleffmann
Abstract
This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.