vix.ing · top · new · best · stats

AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

2026/03/27 by Afonso Simplício, Goncalo Vinagre, Gonçalo Vinagre +22 · 1 voice
Computer Science · #Benchmarking #European Portuguese #Language model #Natural Language Processing Techniques #Open data #Open source #Portuguese #Suite #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG

paper · pdf · open access · doi:10.48550/arxiv.2603.26511

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2026/03/27 · arxiv published 2026/03/27 · arxiv updated 2026/03/27 · openalex created_date 2026/03/31 · openalex updated_date 2026/07/28

Abstract

Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.

Citations

Discussions

Related