2021/10/30 by Sameer Khanna, Khanna, Sameer
Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Network Security and Intrusion Detection #Text and Document Classification Technologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2111.00375
openalex publication_date 2021/10/30 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
As the Internet grows in size, so does the amount of text based information\nthat exists. For many application spaces it is paramount to isolate and\nidentify texts that relate to a particular topic. While one-class\nclassification would be ideal for such analysis, there is a relative lack of\nresearch regarding efficient approaches with high predictive power. By noting\nthat the range of documents we wish to identify can be represented as positive\nlinear combinations of the Vector Space Model representing our text, we propose\nConical classification, an approach that allows us to identify if a document is\nof a particular topic in a computationally efficient manner. We also propose\nNormal Exclusion, a modified version of Bi-Normal Separation that makes it more\nsuitable within the one-class classification context. We show in our analysis\nthat our approach not only has higher predictive power on our datasets, but is\nalso faster to compute.\n