vix.ing · top · new · best · stats

Unsupervised Detection of Anomalous Sound Based on Deep Learning and the Neyman–Pearson Lemma

2018/10/22 by Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu +3 · 137 citations
Computer Science · Engineering · Mathematics · Medicine · #Anomaly (physics) #Anomaly Detection Techniques and Applications #Anomaly detection #Autoencoder #Function (biology) #Lemma (botany) #Music and Audio Processing #Pattern recognition (psychology) #Phonocardiography and Auscultation Techniques #Set (abstract data type) #Training set #cs.LG #cs.SD #eess.AS #stat.ML

paper · pdf · doi:10.1109/taslp.2018.2877258

published in IEEE/ACM Transactions on Audio Speech and Language Processing 27(1), 212-224 (Institute of Electrical and Electronics Engineers) · IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2018

arxiv created 2018/10/22 · arxiv updated 2018/10/23 · openalex publication_date 2018/10/23 · openalex created_date 2018/10/26 · openalex updated_date 2026/08/05

Abstract

This paper proposes a novel optimization principle and its implementation for unsupervised anomaly detection in sound (ADS) using an autoencoder (AE). The goal of the unsupervised-ADS is to detect unknown anomalous sounds without training data of anomalous sounds. The use of an AE as a normal model is a state-of-the-art technique for the unsupervised-ADS. To decrease the false positive rate (FPR), the AE is trained to minimize the reconstruction error of normal sounds, and the anomaly score is calculated as the reconstruction error of the observed sound. Unfortunately, since this training procedure does not take into account the anomaly score for anomalous sounds, the true positive rate (TPR) does not necessarily increase. In this study, we define an objective function based on the Neyman-Pearson lemma by considering the ADS as a statistical hypothesis test. The proposed objective function trains the AE to maximize the TPR under an arbitrary low FPR condition. To calculate the TPR in the objective function, we consider that the set of anomalous sounds is the complementary set of normal sounds and simulate anomalous sounds by using a rejection sampling algorithm. Through experiments using synthetic data, we found that the proposed method improved the performance measures of the ADS under low FPR conditions. In addition, we confirmed that the proposed method could detect anomalous sounds in real environments.

Citations