vix.ing · top · new · best · stats · spec

N-GrAM: New Groningen Author-profiling Model

2017/07/12 by Angelo Basile, Basile, Angelo, Gareth Dwyer +10
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Swearing, Euphemism, Multilingualism #cs.CL

paper · pdf · doi:10.48550/arxiv.1707.03764

arxiv created 2017/07/12 · openalex publication_date 2017/07/12 · arxiv updated 2017/07/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We describe our participation in the PAN 2017 shared task on Author Profiling, identifying authors' gender and language variety for English, Spanish, Arabic and Portuguese. We describe both the final, submitted system, and a series of negative results. Our aim was to create a single model for both gender and language, and for all language varieties. Our best-performing system (on cross-validated results) is a linear support vector machine (SVM) with word unigrams and character 3- to 5-grams as features. A set of additional features, including POS tags, additional datasets, geographic entities, and Twitter handles, hurt, rather than improve, performance. Results from cross-validation indicated high performance overall and results on the test set confirmed them, at 0.86 averaged accuracy, with performance on sub-tasks ranging from 0.68 to 0.98.

Related