2009/01/01 by Thomas Seidl, Seidl, Thomas, Emmanuel Müller +5
Computer Science · Mathematics · #Advanced Statistical Methods and Models #Anomaly Detection Techniques and Applications #Imbalanced Data Classification Techniques #Outlier detection #data mining #outlier ranking #subspace clustering
paper · doi:10.4230/dagsemproc.08421.10
openalex publication_date 2009/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Detecting outliers is an important task for many applications including fraud detection or consistency validation in real world data. Particularly in the presence of uncertain data or imprecise data, similar objects regularly deviate in their attribute values. The notion of outliers has thus to be defined carefully. When considering outlier detection as a task which is complementary to clustering, binary decisions whether an object is regarded to be an outlier or not seem to be near at hand. For high-dimensional data, however, objects may belong to different clusters in different subspaces. More fine-grained concepts to define outliers are therefore demanded. By our new OutRank approach, we address outlier detection in heterogeneous high dimensional data and propose a novel scoring function that provides a consistent model for ranking outliers in the presence of different attribute types. Preliminary experiments demonstrate the potential for successful detection and reasonable ranking of outliers in high dimensional data sets.