Related Experiment Videos
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
Abstract:
Multimodal Large Language Models (MLLMs), such as GPT4o, have demonstrated strong capabilities in visual reasoning and explanation generation. However, despite these advantages, they face notable challenges in the critical task of Image Forgery Detection and Localization (IFDL), where subtle manipulation traces are often overlooked. Existing IFDL approaches are mostly confined to learning low-level, semantic-agnostic clues, neglecting high-level forensic semantics within forged content, and produce only a single judgment without any reasoning. To address these limitations, we propose ForgeryGPT, a novel framework that advances the IFDL task by capturing high-order forensics knowledge correlations from forged images across diverse linguistic feature spaces, while supporting explainable reasoning and interactive dialogue through a customized Large Language Model (LLM). Specifically, ForgeryGPT enhances traditional LLMs with a Mask-Aware Forgery Extractor, enabling precise excavation of forgery mask information and pixel-level understanding of tampering artifacts. This extractor comprises a Forgery Localization Expert (FL-Expert) and a Mask Encoder. The FL-Expert incorporates an Object-agnostic Forgery Prompt and a Vocabulary-enhanced Vision Encoder to capture multi-scale fine-grained forgery details via cross-modal reasoning, thereby improving tampering localization at the knowledge level. For comprehensive training, beyond standard image-text feature alignment, we construct a Mask-Text Alignment Pre-Training dataset using the generated multi-granularity forged images, enabling accurate alignment of forgery masks in the LLM feature space. Subsequently, a Task-Specific Instruction Tuning dataset is employed to refine detection, localization, and dialogue capabilities, yielding robust performance and interpretable inference in forgery-specific scenarios. Extensive experiments demonstrate that ForgeryGPT substantially outperforms state-of-the-art methods, representing one of the early efforts to integrate the IFDL task with explainable, multi-turn dialogue capabilities, while achieving strong generalization across diverse datasets.