2020/03/17 by Xiaodong Wu, Weizhe Lin, Wu, Xiaodong +5
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Misinformation and Its Impacts #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2003.11627
openalex publication_date 2020/03/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Online forums and social media platforms provide noisy but valuable data every day. In this paper, we propose a novel end-to-end neural network-based user embedding system, Author2Vec. The model incorporates sentence representations generated by BERT (Bidirectional Encoder Representations from Transformers) with a novel unsupervised pre-training objective, authorship classification, to produce better user embedding that encodes useful user-intrinsic properties. This user embedding system was pre-trained on post data of 10k Reddit users and was analyzed and evaluated on two user classification benchmarks: depression detection and personality classification, in which the model proved to outperform traditional count-based and prediction-based methods. We substantiate that Author2Vec successfully encoded useful user attributes and the generated user embedding performs well in downstream classification tasks without further finetuning.