2023/07/31 by Matthias Aßenmacher, Aßenmacher, Matthias, Nadja Sauter +3
Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Computers and Society (cs.CY) #FOS: Computer and information sciences #Linguistic Variation and Morphology #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Social Media and Politics
paper · pdf · doi:10.48550/arxiv.2307.16511
openalex publication_date 2023/07/31 · openalex created_date 2023/08/02 · openalex updated_date 2026/07/28
Annotating costs of large corpora are still one of the main bottlenecks in empirical social science research. On the one hand, making use of the capabilities of domain transfer allows re-using annotated data sets and trained models. On the other hand, it is not clear how well domain transfer works and how reliable the results are for transfer across different dimensions. We explore the potential of domain transfer across geographical locations, languages, time, and genre in a large-scale database of political manifestos. First, we show the strong within-domain classification performance of fine-tuned transformer models. Second, we vary the genre of the test set across the aforementioned dimensions to test for the fine-tuned models' robustness and transferability. For switching genres, we use an external corpus of transcribed speeches from New Zealand politicians while for the other three dimensions, custom splits of the Manifesto database are used. While BERT achieves the best scores in the initial experiments across modalities, DistilBERT proves to be competitive at a lower computational expense and is thus used for further experiments across time and country. The results of the additional analysis show that (Distil)BERT can be applied to future data with similar performance. Moreover, we observe (partly) notable differences between the political manifestos of different countries of origin, even if these countries share a language or a cultural background.