2021/09/20 by Myrto Limnios, Limnios, Myrto, Nathan Noiry +3
Computer Science · #Anomaly Detection Techniques and Applications #FOS: Computer and information sciences #FOS: Mathematics #Imbalanced Data Classification Techniques #Machine Learning (stat.ML) #Machine Learning and Data Classification #Statistics Theory (math.ST)
paper · pdf · doi:10.48550/arxiv.2109.09590
openalex publication_date 2021/09/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The ability to collect and store ever more massive databases has been\naccompanied by the need to process them efficiently. In many cases, most\nobservations have the same behavior, while a probable small proportion of these\nobservations are abnormal. Detecting the latter, defined as outliers, is one of\nthe major challenges for machine learning applications (e.g. in fraud detection\nor in predictive maintenance). In this paper, we propose a methodology\naddressing the problem of outlier detection, by learning a data-driven scoring\nfunction defined on the feature space which reflects the degree of abnormality\nof the observations. This scoring function is learnt through a well-designed\nbinary classification problem whose empirical criterion takes the form of a\ntwo-sample linear rank statistics on which theoretical results are available.\nWe illustrate our methodology with preliminary encouraging numerical\nexperiments.\n