2021/09/15 by Daria Pylypenko, Kwabena Amponsah-Kaakyire, Pylypenko, Daria +7 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2109.07604
openalex publication_date 2021/09/15 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Traditional hand-crafted linguistically-informed features have often been\nused for distinguishing between translated and original non-translated texts.\nBy contrast, to date, neural architectures without manual feature engineering\nhave been less explored for this task. In this work, we (i) compare the\ntraditional feature-engineering-based approach to the feature-learning-based\none and (ii) analyse the neural architectures in order to investigate how well\nthe hand-crafted features explain the variance in the neural models'\npredictions. We use pre-trained neural word embeddings, as well as several\nend-to-end neural architectures in both monolingual and multilingual settings\nand compare them to feature-engineering-based SVM classifiers. We show that (i)\nneural architectures outperform other approaches by more than 20 accuracy\npoints, with the BERT-based model performing the best in both the monolingual\nand multilingual settings; (ii) while many individual hand-crafted\ntranslationese features correlate with neural model predictions, feature\nimportance analysis shows that the most important features for neural and\nclassical architectures differ; and (iii) our multilingual experiments provide\nempirical evidence for translationese universals across languages.\n