vix.ing · top · new · best · stats · spec

Classifier Technology and the Illusion of Progress

2006/02/01 by David J. Hand · 2 voices · 5 citations
Computer Science · Mathematics · #Data Mining Algorithms and Applications #Imbalanced Data Classification Techniques #Machine Learning and Data Classification #math.ST #stat.TH

paper · pdf · doi:10.1214/088342306000000060

published as Statistical Science 2006, Vol. 21, No. 1, 1-15 · This paper commented in: [math.ST/0606447], [math.ST/0606452], [math.ST/0606455], [math.ST/0606457]. Rejoinder in [math.ST/0606461]. Published at http://dx.doi.org/10.1214/088342306000000060 in the Statistical Science (http://www.imstat.org/sts/) by the Institute of Mathematical Statistics (http://www.imstat.org)

openalex publication_date 2006/02/01 · arxiv created 2006/06/19 · arxiv published 2006/06/19 · arxiv updated 2006/06/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

A great many tools have been developed for supervised classification, ranging from early methods such as linear discriminant analysis through to modern developments such as neural networks and support vector machines. A large number of comparative studies have been conducted in attempts to establish the relative superiority of these methods. This paper argues that these comparisons often fail to take into account important aspects of real problems, so that the apparent superiority of more sophisticated methods may be something of an illusion. In particular, simple methods typically yield performance almost as good as more sophisticated methods, to the extent that the difference in performance may be swamped by other sources of uncertainty that generally are not considered in the classical supervised classification paradigm.

Citations

Cited by

Discussions