IN3SIGHT: Towards Cognitive Forensic Reasoning for OOC Misinformation Detection
Abstract:
Out-of-context (OOC) misinformation, where genuine images are paired with misleading text, poses substantial societal risks due to its deceptive narratives. Prior efforts primarily assess image-text consistency but often lack explainable judgments, limiting their usefulness for forensic validation. Recently, Multimodal Large Language Models (MLLMs) show promise with vast world knowledge and strong reasoning capabilities, yet their judgments in OOC detection remain unreliable. On one hand, their static internal knowledge quickly becomes outdated in fast-moving news domains, and compensating with externally retrieved evidence often introduces noise and conflicts with their internal priors. On the other hand, the absence of task-specific cognitive calibration leads to hallucinations, whereas instruction tuning introduces such specialization only at the cost of vast curated data, rendering it impractical for dynamic, real-world scenarios. In this paper, we propose IN3SIGHT, a cognitive framework that emulates expert forensic reasoning through three synergistic stages: Epistemic Inspection, which calibrates the epistemic state of an MLLM by task-aware prompt optimization and discrete confidence stratification, yielding a stable epistemic anchor without model tuning; Semantic Investigation, which audits multimodal evidence by enforcing event-level identity and deriving a relevance-aware external assessment under conditional visual anchoring, distilling retrieved evidence into claim-aligned snippets and summaries; and Dialectical Introspection, which conducts multi-fact reasoning by dynamically adjusting the scope of consulted evidence according to evidential certainty and produces the final consistency verdict. Experiments show that IN3SIGHT markedly improves zero-shot MLLM performance by 15.0 percentage points without parameter updates and achieves competitiveness with or even surpasses instruction-tuned methods. This indicates that imposing structured cognitive constraints and evidence-grounded reasoning on MLLMs offers a principled pathway toward reliable misinformation forensics.

