vix.ing · top · new · best · stats · spec

Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs

2024/09/30 by Mehdi Ali, Michael Fromm, Ali, Mehdi +82 · 18 voices · 6 citations
Computer Science · #Library Science and Information Systems #cs.AI #cs.CL #cs.LG

paper · pdf · doi:10.48550/arxiv.2410.03730

openalex publication_date 2024/09/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30

Abstract

We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing a custom multilingual tokenizer, our models address the limitations of existing LLMs that predominantly focus on English or a few high-resource languages. We detail the models' development principles, i.e., data composition, tokenizer optimization, and training methodologies. The models demonstrate strong performance across multilingual benchmarks, as evidenced by their performance on European versions of ARC, HellaSwag, and TruthfulQA.

Cited by

Discussions

Related