vix.ing · top · new · best · stats · spec

Enhancing Generalization in Audio Deepfake Detection: A Neural Collapse based Sampling and Training Approach

2024/04/19 by Mohammed Yousif, Yousif, Mohammed, Jonat John Mathew +13
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2404.13008

openalex publication_date 2024/04/19 · openalex created_date 2024/04/23 · openalex updated_date 2026/07/28

Abstract

Generalization in audio deepfake detection presents a significant challenge, with models trained on specific datasets often struggling to detect deepfakes generated under varying conditions and unknown algorithms. While collectively training a model using diverse datasets can enhance its generalization ability, it comes with high computational costs. To address this, we propose a neural collapse-based sampling approach applied to pre-trained models trained on distinct datasets to create a new training database. Using ASVspoof 2019 dataset as a proof-of-concept, we implement pre-trained models with Resnet and ConvNext architectures. Our approach demonstrates comparable generalization on unseen data while being computationally efficient, requiring less training data. Evaluation is conducted using the In-the-wild dataset.

Related