Event cameras improve privacy in emotion recognition systems

Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Recognizing emotions from facial expressions usually needs regular cameras, but those can reveal people's identities, risking privacy. The researchers explore using event cameras that capture subtle changes in light instead of full images, which helps keep people’s faces more private and works better in tricky lighting. To tackle the challenges of small data and limited ways to sense emotions, they created a big new multimodal dataset combining event camera data, audio, and text. They also designed a smart system that learns to combine these different data sources effectively, improving emotion recognition accuracy. Their tests show this new approach works better than previous ones.

event cameraemotion recognitionprivacymultimodal datafacial action unitsinformation-guided fusioncross-modal alignmentdynamic illuminationmutual information maximizationasynchronous brightness changes

Authors

Jiaqi Chen, Qinfu Xu, Hao Zhuang, Liyuan Pan

Abstract

Emotion analysis is a fundamental task in computer vision, but its practical deployment remains constrained by the privacy risks inherent to conventional RGB cameras. Bio-inspired event cameras present a promising hardware-level solution because they capture asynchronous brightness changes, thereby reducing exposure of facial identity details while leveraging high dynamic range for robust perception under challenging illumination conditions. Despite these advantages, existing event-based methods struggle in complex real-world settings due to limited dataset scales, simple acquisition conditions, and reliance on single-modality visual cues. To address these, we establish a challenging tri-modal benchmark with event, audio, and text modalities and propose the Information-Guided Gated Fusion (IGF) framework, which first pre-trains an event encoder on the FAU subset of Emo-DVS to capture fine-grained facial dynamics, then employs adaptive modality gating to suppress modality-specific noise, and finally leverages mutual information maximization to align robust cross-modal representations. To alleviate data scarcity, we introduce Emo-DVS, the first large-scale event-based emotion analysis dataset, which couples dynamic illumination with the Facial Action Unit (FAU) subset and emotion subset. Extensive experiments demonstrate that IGF achieves state-of-the-art performance.