vix.ing · top · new · best · stats · spec

Re-evaluating the need for Modelling Term-Dependence in Text Classification Problems

2017/10/25 by Sounak Banerjee, Banerjee, Sounak, Prasenjit Majumder +3
Computer Science · #68P20 #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning and Data Classification #Text and Document Classification Technologies

paper · pdf · doi:10.48550/arxiv.1710.09085

openalex publication_date 2017/10/25 · openalex created_date 2017/11/10 · openalex updated_date 2026/07/28

Abstract

A substantial amount of research has been carried out in developing machine learning algorithms that account for term dependence in text classification. These algorithms offer acceptable performance in most cases but they are associated with a substantial cost. They require significantly greater resources to operate. This paper argues against the justification of the higher costs of these algorithms, based on their performance in text classification problems. In order to prove the conjecture, the performance of one of the best dependence models is compared to several well established algorithms in text classification. A very specific collection of datasets have been designed, which would best reflect the disparity in the nature of text data, that are present in real world applications. The results show that even one of the best term dependence models, performs decent at best when compared to other independence models. Coupled with their substantially greater requirement for hardware resources for operation, this makes them an impractical choice for being used in real world scenarios.

Related