vix.ing · top · new · best · stats · spec

Identifying Computer-Translated Paragraphs using Coherence Features

2018/12/28 by Hoang-Quoc Nguyen-Son, Ngoc-Dung T. Tieu, Nguyen-Son, Hoang-Quoc +7 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1812.10896

openalex publication_date 2018/12/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We have developed a method for extracting the coherence features from a paragraph by matching similar words in its sentences. We conducted an experiment with a parallel German corpus containing 2000 human-created and 2000 machine-translated paragraphs. The result showed that our method achieved the best performance (accuracy = 72.3%, equal error rate = 29.8%) when it is compared with previous methods on various computer-generated text including translation and paper generation (best accuracy = 67.9%, equal error rate = 32.0%). Experiments on Dutch, another rich resource language, and a low resource one (Japanese) attained similar performances. It demonstrated the efficiency of the coherence features at distinguishing computer-translated from human-created paragraphs on diverse languages.

Cited by

Related