2026/05/25 by Geyang Guo, Hiromi Wakaki, Yuki Mitsufuji +2 · 1 voice
Computer Science · #Constructed language #Diversity (politics) #Language model #Multilingualism #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Quality (philosophy) #Reinforcement learning #Router #Topic Modeling #cs.CL
paper · pdf · open access · doi:10.48550/arxiv.2605.25360
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2026/05/25 · arxiv published 2026/05/25 · arxiv updated 2026/05/25 · openalex created_date 2026/05/27 · openalex updated_date 2026/07/28
Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each training question to a single response language or rely on a fixed dominant language for supervision. We propose language-routed policy optimization (LRPO), an online policy optimization framework that treats language as a selectable variable. LRPO elicits multilingual rollouts for each training question and integrates their relative quality into preference-based policy updates, increasing the diversity and informativeness of training signals under the fixed rollout budget. To adaptively determine which languages to explore during reinforcement learning, we introduce a trainable language router formulated as a multi-armed bandit, balancing exploration of underutilized languages with exploitation of more informative ones. Extensive experiments show that LRPO consistently improves multilingual performance, demonstrating that adaptive language routing enables effective cross-lingual knowledge exploitation for training. We release all the resources at https://github.com/Guochry/LRPO.