vix.ing · top · new · best · stats · spec

CodeRosetta: Pushing the Boundaries of Unsupervised Code Translation for Parallel Programming

2024/10/27 by Ali TehraniJamsaz, TehraniJamsaz, Ali, Arijit Bhattacharjee +9 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #Performance (cs.PF) #Programming Languages (cs.PL) #Software Engineering (cs.SE) #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.2410.20527

openalex publication_date 2024/10/27 · openalex created_date 2024/11/14 · openalex updated_date 2026/07/28

Abstract

Recent advancements in Large Language Models (LLMs) have renewed interest in automatic programming language translation. Encoder-decoder transformer models, in particular, have shown promise in translating between different programming languages. However, translating between a language and its high-performance computing (HPC) extensions remains underexplored due to challenges such as complex parallel semantics. In this paper, we introduce CodeRosetta, an encoder-decoder transformer model designed specifically for translating between programming languages and their HPC extensions. CodeRosetta is evaluated on C++ to CUDA and Fortran to C++ translation tasks. It uses a customized learning framework with tailored pretraining and training objectives to effectively capture both code semantics and parallel structural nuances, enabling bidirectional translation. Our results show that CodeRosetta outperforms state-of-the-art baselines in C++ to CUDA translation by 2.9 BLEU and 1.72 CodeBLEU points while improving compilation accuracy by 6.05%. Compared to general closed-source LLMs, our method improves C++ to CUDA translation by 22.08 BLEU and 14.39 CodeBLEU, with 2.75% higher compilation accuracy. Finally, CodeRosetta exhibits proficiency in Fortran to parallel C++ translation, marking it, to our knowledge, as the first encoder-decoder model for this complex task, improving CodeBLEU by at least 4.63 points compared to closed-source and open-code LLMs.

Cited by

Related