vix.ing · top · new · best · stats · spec

Language Models in Software Development Tasks: An Experimental Analysis of Energy and Accuracy

2024/11/30 by Negar Alizadeh, Boris Belchev, Alizadeh, Negar +7 · 1 voice · 2 citations
Computer Science · #Software Engineering Research #Software Engineering Techniques and Practices #Software System Performance and Reliability #cs.SE

paper · pdf · doi:10.48550/arxiv.2412.00329

openalex publication_date 2024/11/30 · openalex created_date 2024/12/05 · openalex updated_date 2026/07/28

Abstract

The use of generative AI-based coding assistants like ChatGPT and Github Copilot is a reality in contemporary software development. Many of these tools are provided as remote APIs. Using third-party APIs raises data privacy and security concerns for client companies, which motivates the use of locally-deployed language models. In this study, we explore the trade-off between model accuracy and energy consumption, aiming to provide valuable insights to help developers make informed decisions when selecting a language model. We investigate the performance of 18 families of LLMs in typical software development tasks on two real-world infrastructures, a commodity GPU and a powerful AI-specific GPU. Given that deploying LLMs locally requires powerful infrastructure which might not be affordable for everyone, we consider both full-precision and quantized models. Our findings reveal that employing a big LLM with a higher energy budget does not always translate to significantly improved accuracy. Additionally, quantized versions of large models generally offer better efficiency and accuracy compared to full-precision versions of medium-sized ones. Apart from that, not a single model is suitable for all types of software development tasks.

Cited by

Discussions

Related