Deepfake detector shows what changed and where on the face

DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection

Computer Vision and Pattern RecognitionMachine Learning

Summary

Detecting deepfake images is getting better, but people also want to know which parts of the face have been altered and what clues reveal these changes. The authors created DF-CBM, a method that links specific facial regions to common manipulation signs using a set of clear concepts learned from descriptions of artifacts. This approach not only detects deepfakes accurately but also shows which facial areas contributed to the decision, helping users understand the evidence behind the detection. The method can even explain how changing individual clues would affect the final decision, offering transparent insights.

What this means in practice

  • For forensic analysts: Use this model to identify and highlight specific manipulated facial regions in images to support digital forensic investigations.
  • For security software developers: Develop explainable deepfake detection tools that provide clear visual and semantic cues to end users about image manipulation.$Commercial implications: This technology enables creation of commercial security products that explain detected manipulations to clients, improving trust and regulatory compliance.

Authors

Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou, Panos K. Papadopoulos, Yannis Spyridis, Dimitrios Zarpalas, Vasileios Argyriou

Abstract

Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only partially address this need since localization-based approaches lack semantic descriptions while language-based explanation methods are only weakly grounded in visual evidence. In this work, we propose DF-CBM, a region-aware concept bottleneck model for explainable deepfake detection. DF-CBM builds a compact vocabulary of manipulation-related concepts from textual artifact annotations and links each concept to plausible facial and boundary regions. It then predicts these concepts from visual features using a concept-specific masked attention mechanism guided by parsed facial masks and the final real/fake decision is made from the predicted concept bottleneck. Our experiments show that DF-CBM outperforms concept-based baselines in concept prediction and deepfake classification while remaining competitive with state-of-the-art black-box detectors. Finally, qualitative results and intervention analyses demonstrate that DF-CBM provides spatially grounded concept evidence and enables counterfactual explanations of how individual manipulation concepts influence the final prediction. Our code is available at: https://github.com/GeorgeTsoumplekas/DF-CBM.