Combining global and local models improves deepfake detection accuracy
DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection
Computer Vision and Pattern Recognition
Summary
Deepfake videos and images can trick people by showing fake faces or expressions. Existing small models often fail to detect new or different types of fakes well. These authors combine two powerful models: one that understands the whole image’s meaning and another that focuses on fine facial details. Their combined method can better spot fakes across different types and datasets. Testing shows this approach works well even on fake images it hasn’t seen before.
What this means in practice
- •For security software developers: Develop improved tools to detect manipulated facial images and videos by combining global and local feature models for better generalization across forgery types.
- •For digital forensics teams: Enhance forensic analysis of facial forgeries by integrating holistic semantic cues with detailed facial irregularity detection for robust identification.
Authors
Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb, Shiguang Shan
Abstract
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and unseen manipulations. To address this limitation, researchers have turned to large-scale foundation models, which can provide richer representations and better generalization. Nevertheless, relying on a single foundation model alone remains insufficient for effective forgery detection. While models like CLIP offer robust global semantic cues, they lack the capacity to capture detailed local facial features. In contrast, DINO excels at capturing local structural features of faces, but provides weaker global semantic context. To fully utilize the synergies among multiple foundation models, we propose a hierarchical multi-granular framework that integrates complementary pretrained representations. Specifically, a Global Context Branch (GCB) based on CLIP captures holistic semantic cues, while a Fine-grained Cue Branch (FCB) built on DINOv3 captures localized structural irregularities. In addition, we design a feature fusion module that enables parameter-efficient adaptation of the frozen foundation backbones by adaptively extracting and integrating complementary features from the two models. By jointly leveraging global context and fine-grained cues, our method learns more comprehensive forgery representations and achieves strong cross-manipulation performance. Extensive experiments on multiple benchmarks demonstrate the benefit of the proposed design, particularly under cross-dataset and cross-manipulation settings.