vix.ing · top · new · best · stats · spec

NoCoLA: The Norwegian Corpus of Linguistic Acceptability

2023/06/13 by Matias Jentoft, David Samuel, Jentoft, Matias +1 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2306.07790

openalex publication_date 2023/06/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

While there has been a surge of large language models for Norwegian in recent years, we lack any tool to evaluate their understanding of grammaticality. We present two new Norwegian datasets for this task. NoCoLAclass is a supervised binary classification task where the goal is to discriminate between acceptable and non-acceptable sentences. On the other hand, NoCoLAzero is a purely diagnostic task for evaluating the grammatical judgement of a language model in a completely zero-shot manner, i.e. without any further training. In this paper, we describe both datasets in detail, show how to use them for different flavors of language models, and conduct a comparative study of the existing Norwegian language models.

Cited by

Related