2023/04/28 by Felix Stollenwerk, Stollenwerk, Felix · 1 citation
Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational Physics and Python Applications #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2304.14780
openalex publication_date 2023/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper provides a detailed discussion of the multilingual tokenizer used for GPT-SW3. It was trained on the Nordic Pile using the SentencePiece library and the BPE algorithm. We outline the tokenizer's most important features and share details on its learned vocabulary. In addition, we systematically analyze the properties and evaluate the performance of the tokenizer with regard to the different languages present in the data.