Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation
2026-08-31 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionMultimedia
AI summaryⓘ
The authors address privacy risks from AI that creates 3D talking faces from videos and speech, which can mimic someone's identity. Current defenses try to mess with the face images but can lower image quality and be easily undone. Instead, the authors propose hiding tiny changes in the audio that humans can't notice but disrupt the 3D face animation. Their method reduces the ability to create fake talking faces while keeping the audio sounding natural. This shows that protecting privacy through subtle audio changes is a useful new approach.
generative portrait models3D talking face generationaudio-driven animationprivacy leakageidentity misusevisual perturbationpsychoacoustic maskingaudio perturbationperceptual qualityfacial animation
Authors
Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao, Tian Lan, Zhijing Wu, Xian-Ling Mao
Abstract
The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. In particular, audio-driven 3D talking face generation can reconstruct a reusable 3D portrait of a target person from a monocular video and animate it with arbitrary speech, making realistic identity impersonation alarmingly practical. Existing proactive defenses mainly operate in the visual domain by injecting subtle perturbations into acial regions to disrupt identity acquisition. However, such perturbations often compromise visual quality due to the strong structural priors and social sensitivity of human faces, and are easily weakened by common real-world transformations such as resizing. To overcome these limitations, we propose an imperceptible audio defense for audio-driven 3D talking face generation by shifting protection from the visual modality to the audio modality. Specifically,we exploit psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal, thereby reducing perceptual distortion while suppressing reliable facial animation. Extensive experiments demonstrate that the proposed method effectively degrades 3D talking face generation while preserving favorable perceptual quality. These findings highlight psychoacoustically guided audio perturbations as a practical and promising direction for privacy-preserving portrait protection.