2024/06/13 by Javier García Gilabert, Gilabert, Javier García, Carlos Escolano +11 · 1 citation
Chemistry · Computer Science · #Artificial intelligence #Chemistry #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #Natural Language Processing Techniques #Natural language processing #Programming language #Topic Modeling #Translation (biology)
paper · pdf · doi:10.48550/arxiv.2406.09140
openalex publication_date 2024/06/13 · openalex created_date 2024/06/15 · openalex updated_date 2026/07/28
In recent years, Large Language Models (LLMs) have demonstrated exceptional proficiency across a broad spectrum of Natural Language Processing (NLP) tasks, including Machine Translation. However, previous methods predominantly relied on iterative processes such as instruction fine-tuning or continual pre-training, leaving unexplored the challenges of training LLMs solely on parallel data. In this work, we introduce PLUME (Parallel Language Model), a collection of three 2B LLMs featuring varying vocabulary sizes (32k, 128k, and 256k) trained exclusively on Catalan-centric parallel examples. These models perform comparably to previous encoder-decoder architectures on 16 supervised translation directions and 56 zero-shot ones. Utilizing this set of models, we conduct a thorough investigation into the translation capabilities of LLMs, probing their performance, the impact of the different elements of the prompt, and their cross-lingual representation space.