Towards Explainable Bilingual Multimodal Misinformation Detection and Localization
2025/06/28 by Yiwei He, He, Yiwei, Zhenglin Huang +12
Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Misinformation and Its Impacts #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.22930
openalex publication_date 2025/06/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The increasing realism of multimodal content has made misinformation more subtle and harder to detect, especially in news media where images are frequently paired with bilingual (e.g., Chinese-English) subtitles. Such content often includes localized image edits and cross-lingual inconsistencies that jointly distort meaning while remaining superficially plausible. We introduce BiMi, a bilingual multimodal framework that jointly performs region-level localization, cross-modal and cross-lingual consistency detection, and natural language explanation for misinformation analysis. To support generalization, BiMi integrates an online retrieval module that supplements model reasoning with up-to-date external context. We further release BiMiBench, a large-scale and comprehensive benchmark constructed by systematically editing real news images and subtitles, comprising 104,000 samples with realistic manipulations across visual and linguistic modalities. To enhance interpretability, we apply Group Relative Policy Optimization (GRPO) to improve explanation quality, marking the first use of GRPO in this domain. Extensive experiments demonstrate that BiMi outperforms strong baselines by up to +8.9 in classification accuracy, +15.9 in localization accuracy, and +2.5 in explanation BERTScore, advancing state-of-the-art performance in realistic, multilingual misinformation detection. Code, models, and datasets will be released.
Citations
- DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection
- Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Qwen2.5-Omni Technical Report
- Gemma 3 Technical Report
- Grounded Chain-of-Thought for Multimodal Large Language Models
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning
- SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
- GPT-4o System Card
- The Llama 3 Herd of Models
- Multimodal Misinformation Detection using Large Vision-Language Models
- OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
- MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
- Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis
- SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
- DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- Improved Baselines with Visual Instruction Tuning
- Qwen Technical Report
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
- LISA: Reasoning Segmentation via Large Language Model
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- Detecting and Grounding Multi-Modal Media Manipulation
- A Survey of Large Language Models
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- Multi-modal Misinformation Detection: Approaches, Challenges and Opportunities
- Training language models to follow instructions with human feedback
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- NewsCLIPpings: Automatic Generation of Out-of-Context Multimodal Media
- Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
- Visual News: Benchmark and Challenges in News Image Captioning
- Textual misinformation on Reddit
- The Future of Misinformation Detection: New Perspectives and Trends
- BERTScore: Evaluating Text Generation with BERT
- FakeNewsNet: A Data Repository with News Content, Social Context and Spatialtemporal Information for Studying Fake News on Social Media
- FakeNewsNet: A Data Repository with News Content, Social Context, and Spatiotemporal Information for Studying Fake News on Social Media
- FEVER: a large-scale dataset for Fact Extraction and VERification
- Proximal Policy Optimization Algorithms
- COSMOS: Catching Out-of-Context Misinformation with Self-Supervised Learning
Related