2017/03/09 by Marjan Hosseinia, Hosseinia, Marjan, Arjun Mukherjee +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
paper · pdf · doi:10.48550/arxiv.1703.03149
18 pages, Accepted at CICLing 2017, 18th International Conference on Intelligent Text Processing and Computational Linguistics
arxiv created 2017/03/09 · arxiv updated 2017/03/10
This paper explores the problem of sockpuppet detection in deceptive opinion spam using authorship attribution and verification approaches. Two methods are explored. The first is a feature subsampling scheme that uses the KL-Divergence on stylistic language models of an author to find discriminative features. The second is a transduction scheme, spy induction that leverages the diversity of authors in the unlabeled test set by sending a set of spies (positive samples) from the training set to retrieve hidden samples in the unlabeled test set using nearest and farthest neighbors. Experiments using ground truth sockpuppet data show the effectiveness of the proposed schemes.