vix.ing · top · new · best · stats · spec

A Fast Text Similarity Measure for Large Document Collections using\n Multi-reference Cosine and Genetic Algorithm

2018/10/07 by Hamid Mohammadi, Mohammadi, Hamid, Seyed Hossein Khasteh +1 · 1 citation
Computer Science · #Spam and Phishing Detection #Web Data Mining and Analysis #Text and Document Classification Technologies

paper · pdf · doi:10.48550/arxiv.1810.03102

Abstract

One of the important factors that make a search engine fast and accurate is a\nconcise and duplicate free index. In order to remove duplicate and\nnear-duplicate documents from the index, a search engine needs a swift and\nreliable duplicate and near-duplicate text document detection system.\nTraditional approaches to this problem, such as brute force comparisons or\nsimple hash-based algorithms are not suitable as they are not scalable and are\nnot capable of detecting near-duplicate documents effectively. In this paper, a\nnew signature-based approach to text similarity detection is introduced which\nis fast, scalable, reliable and needs less storage space. The proposed method\nis examined on popular text document data-sets such as CiteseerX, Enron, Gold\nSet of Near-duplicate News Articles and etc. The results are promising and\ncomparable with the best cutting-edge algorithms, considering the accuracy and\nperformance. The proposed method is based on the idea of using reference texts\nto generate signatures for text documents. The novelty of this paper is the use\nof genetic algorithms to generate better reference texts.\n

Cited by

Related