vix.ing · top · new · best · stats · spec

BERTić -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

2021/04/19 by Nikola Ljubešić, Ljubešić, Nikola, Davor Lauc +1 · 4 citations
Arts and Humanities · Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Linguistics, Language Diversity, and Identity #Multilingual Education and Policy #Natural Language Processing Techniques

paper · pdf · doi:10.48550/arxiv.2104.09243

openalex publication_date 2021/04/19 · openalex created_date 2021/04/26 · openalex updated_date 2026/07/28

Abstract

In this paper we describe a transformer model pre-trained on 8 billion tokens of crawled text from the Croatian, Bosnian, Serbian and Montenegrin web domains. We evaluate the transformer model on the tasks of part-of-speech tagging, named-entity-recognition, geo-location prediction and commonsense causal reasoning, showing improvements on all tasks over state-of-the-art models. For commonsense reasoning evaluation, we introduce COPA-HR -- a translation of the Choice of Plausible Alternatives (COPA) dataset into Croatian. The BERTić model is made available for free usage and further task-specific fine-tuning through HuggingFace.

Cited by

Related