2021/05/28 by Xuechen Li, Chris J. Maddison, Li, Xuechen +3
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE) #Software Engineering Research #Software Testing and Debugging Techniques #Teaching and Learning Programming
paper · pdf · doi:10.48550/arxiv.2105.14038
openalex publication_date 2021/05/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Source code spends most of its time in a broken or incomplete state during software development. This presents a challenge to machine learning for code, since high-performing models typically rely on graph structured representations of programs derived from traditional program analyses. Such analyses may be undefined for broken or incomplete code. We extend the notion of program graphs to work-in-progress code by learning to predict edge relations between tokens, training on well-formed code before transferring to work-in-progress code. We consider the tasks of code completion and localizing and repairing variable misuse in a work-in-process scenario. We demonstrate that training relation-aware models with fine-tuned edges consistently leads to improved performance on both tasks.