2020/04/06 by Elena Álvarez Mellado, Álvarez-Mellado, Elena · 2 citations
Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Lexicography and Language Studies #Linguistics, Language Diversity, and Identity #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2004.02929
openalex publication_date 2020/04/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglicisms and a baseline model for anglicism extraction. In this paper we present: (1) a corpus of 21,570 newspaper headlines written in European Spanish annotated with emergent anglicisms and (2) a conditional random field baseline model with handcrafted features for anglicism extraction. We present the newspaper headlines corpus, describe the annotation tagset and guidelines and introduce a CRF model that can serve as baseline for the task of detecting anglicisms. The presented work is a first step towards the creation of an anglicism extractor for Spanish newswire.