2019/11/30 by Andrea Cappozzo, Francesca Greselin, Thomas Brendan Murphy · 20 citations
Computer Science · Mathematics · #Anomaly Detection Techniques and Applications #Anomaly detection #Discriminant #Exploit #Machine Learning and Data Classification #Novelty #Novelty detection #Pattern recognition (psychology) #Set (abstract data type) #Test set #Time Series Analysis and Forecasting #Training set #stat.AP
paper · pdf · doi:10.1007/s11222-020-09959-1
published in Statistics and Computing 30(5), 1545-1571 (Springer Science+Business Media)
openalex created_date 2019/12/05 · arxiv created 2020/05/29 · openalex publication_date 2020/06/30 · arxiv updated 2020/07/02 · openalex updated_date 2026/08/05
Three important issues are often encountered in Supervised and Semi-Supervised Classification: class-memberships are unreliable for some training units (label noise), a proportion of observations might depart from the main structure of the data (outliers) and new groups in the test set may have not been encountered earlier in the learning phase (unobserved classes). The present work introduces a robust and adaptive Discriminant Analysis rule, capable of handling situations in which one or more of the afore-mentioned problems occur. Two EM-based classifiers are proposed: the first one that jointly exploits the training and test sets (transductive approach), and the second one that expands the parameter estimate using the test set, to complete the group structure learned from the training set (inductive approach). Experiments on synthetic and real data, artificially adulterated, are provided to underline the benefits of the proposed method.