2024/01/01 by Rui Wang, Dengpan Ye, Long Tang +2 · 1 citation
Computer Science · #Digital Media Forensic Detection #Generative Adversarial Networks and Image Synthesis #Music and Audio Processing
paper · doi:10.1109/lsp.2024.3433596
crossref issued 2024/01/01 · crossref published 2024/01/01 · crossref published-print 2024/01/01 · openalex publication_date 2024/01/01 · crossref created 2024/07/25 · openalex created_date 2025/10/10 · crossref deposited 2026/07/07 · openalex updated_date 2026/07/29 · crossref indexed 2026/07/29
With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this letter, we proposeAVT2-DWF, theAudio-Visual dualTransformers grounded inDynamicWeightFusion, which aims to amplify both intra- and cross-modal forgery cues, thereby enhancing detection capabilities. AVT2-DWF adopts a dual-stage approach to capture both spatial characteristics and temporal dynamics of facial expressions. This is achieved through a face transformer with ann-frame-wise tokenization strategy encoder and an audio transformer encoder. Subsequently, it uses multi-modal conversion with dynamic weight fusion to address the challenge of heterogeneous information fusion between audio and visual modalities. Experiments on DeepfakeTIMIT, FakeAVCeleb, and DFDC datasets indicate that AVT2-DWF achieves state-of-the-art performance intra- and cross-dataset Deepfake detection.