vix.ing · top · new · best · stats

Machine learning in automated text categorization

2001/10/26 by Fabrizio Sebastiani · 7,922 citations
Computer Science · #Advanced Text Analysis Techniques #Algorithms and Data Compression #Artificial intelligence #Categorization #Classifier (UML) #Computer science #Machine learning #Natural language processing #Software portability #Text and Document Classification Technologies #Text categorization #cs.IR #cs.LG

paper · pdf · doi:10.1145/505282.505283

published in ACM Computing Surveys 34(1), 1-47 (Association for Computing Machinery) · Accepted for publication on ACM Computing Surveys

arxiv created 2001/10/26 · openalex publication_date 2002/03/01 · openalex created_date 2021/02/01 · arxiv updated 2021/09/21 · openalex updated_date 2026/08/05

Abstract

The automated categorization (or classification) of texts into predefined categories has witnessed a booming interest in the last 10 years, due to the increased availability of documents in digital form and the ensuing need to organize them. In the research community the dominant approach to this problem is based on machine learning techniques: a general inductive process automatically builds a classifier by learning, from a set of preclassified documents, the characteristics of the categories. The advantages of this approach over the knowledge engineering approach (consisting in the manual definition of a classifier by domain experts) are a very good effectiveness, considerable savings in terms of expert labor power, and straightforward portability to different domains. This survey discusses the main approaches to text categorization that fall within the machine learning paradigm. We will discuss in detail issues pertaining to three different problems, namely, document representation, classifier construction, and classifier evaluation.

Citations

Cited by