vix.ing · top · new · best · stats · spec

Deep Investigation of Cross-Language Plagiarism Detection Methods

2017/05/24 by Jeremy Ferrero, Laurent Besacier, Ferrero, Jeremy +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL

paper · pdf · doi:10.48550/arxiv.1705.08828

Accepted to BUCC (10th Workshop on Building and Using Comparable Corpora) colocated with ACL 2017

arxiv created 2017/05/24 · arxiv updated 2017/05/25

Abstract

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.

Related