vix.ing · top · new · best · stats · spec

Optimizing text representations to capture (dis)similarity between political parties

2022/10/21 by Tanise Ceron, Ceron, Tanise, Nico Blokker +3
Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2210.11989

openalex publication_date 2022/10/21 · openalex created_date 2022/10/30 · openalex updated_date 2026/07/28

Abstract

Even though fine-tuned neural language models have been pivotal in enabling "deep" automatic text analysis, optimizing text representations for specific applications remains a crucial bottleneck. In this study, we look at this problem in the context of a task from computational social science, namely modeling pairwise similarities between political parties. Our research question is what level of structural information is necessary to create robust text representation, contrasting a strongly informed approach (which uses both claim span and claim category annotations) with approaches that forgo one or both types of annotation with document structure-based heuristics. Evaluating our models on the manifestos of German parties for the 2021 federal election. We find that heuristics that maximize within-party over between-party similarity along with a normalization step lead to reliable party similarity prediction, without the need for manual annotation.

Related