2011/12/09 by Bhawna Nigam, Nigam, Bhawna, Poorvi Ahirwal +5
Computer Science · #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Text and Document Classification Technologies
paper · pdf · doi:10.48550/arxiv.1112.2028
openalex publication_date 2011/12/09 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
As the amount of online document increases, the demand for document\nclassification to aid the analysis and management of document is increasing.\nText is cheap, but information, in the form of knowing what classes a document\nbelongs to, is expensive. The main purpose of this paper is to explain the\nexpectation maximization technique of data mining to classify the document and\nto learn how to improve the accuracy while using semi-supervised approach.\nExpectation maximization algorithm is applied with both supervised and\nsemi-supervised approach. It is found that semi-supervised approach is more\naccurate and effective. The main advantage of semi supervised approach is\n"Dynamically Generation of New Class". The algorithm first trains a classifier\nusing the labeled document and probabilistically classifies the unlabeled\ndocuments. The car dataset for the evaluation purpose is collected from UCI\nrepository dataset in which some changes have been done from our side.\n