Training free deepfake detection finds suspicious image areas first
Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection
Computer Vision and Pattern Recognition
Summary
Detecting fake images called deepfakes can be tricky because the fake parts are often small and hidden. The authors propose a new way that doesn’t require extra training, where their method first spots which small parts of an image might be fake by comparing it to a blurred version. Then, it looks at these parts closely along with the whole image before deciding if it’s fake. This approach works well with existing AI language and vision models and improves accuracy while giving better visual explanations.
What this means in practice
- •For digital forensics teams: Identify image regions that reveal deepfake manipulation without needing extra training data or forensic tools.
- •For content moderation teams: Use enhanced deepfake detection that explains decisions clearly by focusing on suspicious local image evidence.
Authors
Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin, Tai-Ming Huang, Ling Lo, Po-Ching Chen, Yan-Tsung Wang, Pei-Heng Li, Ling Zou, Hong-Han Shuai, Wen-Huang Cheng
Abstract
Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.