Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku

2026-08-24Artificial Intelligence

Artificial Intelligence
AI summary

The authors study how real-time crowd comments called Danmaku (bullet comments) can help detect fake news in videos. Since these comments usually appear with delays, the authors create a system called Genda to simulate timely, realistic Danmaku streams by predicting when and how people would react. They then use these generated comments along with video, audio, and text in a model named DM-FEND to better spot fake news. Their experiments show that including these time-aware Danmaku improves detection accuracy in both Chinese and English datasets.

DanmakuFake news detectionMultimodal analysisTemporal modelingGenerative modelsBullet commentsUser interaction simulationSemantic alignmentEmotional expression
Authors
Xiansheng Luo, Chaowei Zhang, Zewei Zhang, Yi Zhu, Jipeng Qiang
Abstract
The social interactions among crowds via \textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both viewpoint conflicts and consensus, providing fine-grained discriminative social signals that can benefit fake news detection. However, the inherent accumulation latency of \textit{Danmaku} in real-world scenarios violates the real-time necessity of fake news detection, making the studies of \textit{Danmaku}-related fake news detection underexplored. To break this violation, we simulate this temporal-aware user interactive process by proposing a novel temporal \textbf{Gen}erative \textbf{da}nmaku framework, called \textbf{Genda}, which consists of: (1) a \textit{Danmaku} Trigger for predicting the timing and intensity of user reactions; and (2) a \textit{Danmaku} Generator for synthesizing corresponding semantic and emotional expressions, thereby mutually constructing a temporally aligned and human-like pseudo \textit{Danmaku} streams. To make the generated \textit{Danmaku} useful for identifying fake news videos, we further design a \textit{Danmaku}-guided Temporal Multimodal fake news detection model - \textbf{DM-FEND}, which enables fine-grained multimodal interactions among video, audio, text, and \textit{Danmaku}, enhancing dynamic modalities alignment and semantic noise inhibition. The experimental results demonstrate that \emph{DM-FEND} consistently outperforms state-of-the-art baselines across both Chinese (FakeSV) and English (FakeTT) benchmarks. Further ablations validate the crucial role of temporal \textit{Danmaku} modeling in enhancing robustness and discriminative capability. Finally, this study offers a bright and robust solution for multimodal fake news detection in modern social interactive fashions by bridging the temporal inconsistency between news and user behaviors.