Large Language Models for Software Engineering: A Systematic Literature Review
2023/08/21 by Xinyi Hou, Yanjie Zhao, Hou, Xinyi +17 · 242 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Software Engineering (cs.SE) #Software Engineering Research #Software Engineering Techniques and Practices #Software System Performance and Reliability
paper · pdf · doi:10.48550/arxiv.2308.10620
openalex publication_date 2023/08/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) have significantly impacted numerous domains, including Software Engineering (SE). Many recent publications have explored LLMs applied to various SE tasks. Nevertheless, a comprehensive understanding of the application, effects, and possible limitations of LLMs on SE is still in its early stages. To bridge this gap, we conducted a systematic literature review (SLR) on LLM4SE, with a particular focus on understanding how LLMs can be exploited to optimize processes and outcomes. We select and analyze 395 research papers from January 2017 to January 2024 to answer four key research questions (RQs). In RQ1, we categorize different LLMs that have been employed in SE tasks, characterizing their distinctive features and uses. In RQ2, we analyze the methods used in data collection, preprocessing, and application, highlighting the role of well-curated datasets for successful LLM for SE implementation. RQ3 investigates the strategies employed to optimize and evaluate the performance of LLMs in SE. Finally, RQ4 examines the specific SE tasks where LLMs have shown success to date, illustrating their practical contributions to the field. From the answers to these RQs, we discuss the current state-of-the-art and trends, identifying gaps in existing research, and flagging promising areas for future study. Our artifacts are publicly available at https://github.com/xinyi-hou/LLM4SESLR.
Cited by
- Prompt Driven Development with Claude Code: Building a Complete TUI Framework for the Ring Programming Language
- Generative AI for Software Project Management: Insights from a Review of Software Practitioner Literature
- How Well Can AI Generate Backlogs from App Mockups?
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- Assessing the Software Security Comprehension of Large Language Models
- Scrum Sprint Planning: LLM-based and algorithmic solutions
- SoK: Understanding (New) Security Issues Across AI4Code Use Cases
- An Investigation on How AI-Generated Responses Affect SoftwareEngineering Surveys
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- How Low Can You Go? The Data-Light SE Challenge
- On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models
- Cluster-guided LLM-Based Anonymization of Software Analytics Data: Studying Privacy-Utility Trade-offs in JIT Defect Prediction
- CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models
- Multicalibration for LLM-based Code Generation
- WhatsCode: Large-Scale GenAI Deployment for Developer Efficiency at WhatsApp
- Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
- A Hybrid Approach for EMF Code Generation:Code Templates Meet Large Language Models
- Quantitative Analysis of Technical Debt and Pattern Violation in Large Language Model Architectures
- Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
- CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
- Amplifiers or Equalizers? A Longitudinal Study of LLM Evolution in Software Engineering Project-Based Learning
- Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation
- LLMs-Powered Real-Time Fault Injection: An Approach Toward Intelligent Fault Test Cases Generation
- The Software Engineering Simulations Lab: Agentic AI for RE Quality Simulations
- Agentic Program Verification
- Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
- Utilizing LLMs for Industrial Process Automation: A Case Study on Modifying RAPID Programs
- The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
- From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI
- Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Walking the Tightrope of LLMs for Software Development: A Practitioners' Perspective
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- Generating Software Architecture Description from Source Code using Reverse Engineering and Large Language Model
- What About Our Bug? A Study on the Responsiveness of NPM Package Maintainers
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
- Who's Who? LLM-assisted Software Traceability with Architecture Entity Recognition
- Can Language Models Go Beyond Coding? Assessing the Capability of Language Models to Build Real-World Systems
- Repairing Responsive Layout Failures Using Retrieval Augmented Generation
- CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
- A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI
- Predicate Renaming via Large Language Models
- Four Years of GenAI: How Educators and Industry Adapted Their Assessment Strategies
- Large Language Models for Software Engineering Diagrams: A Systematic Review of UML and ER modelling
- Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring
- TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning
- Tinker, Tailor, Trust: How Developers Create Privacy Policies With and Without AI
- An Empirical Study of Foundation Models for Variability-Induced Compilation Errors in Configurable C Code
- How well LLM-based test generation techniques perform with newer LLM versions?
- Understanding the Characteristics of LLM-Generated Property-Based Tests in Exploring Edge Cases
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- From Online User Feedback to Requirements: Evaluating Large Language Models for Classification and Specification Tasks
- CodeAD: Synthesize Code of Rules for Log-based Anomaly Detection with LLMs
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- Gen-Review: A Large-scale Dataset of AI-Generated (and Human-written) Peer Reviews
- Neurosymbolic Characterization for Reliable Access Control Policy Analysis
- Toward Agentic Software Engineering Beyond Code: Framing Vision, Values, and Vocabulary
- Streamlining Acceptance Test Generation for Mobile Applications Through Large Language Models: An Industrial Case Study
- FeClustRE: Hierarchical Clustering and Semantic Tagging of App Features from User Reviews
- Evaluating LLM-Based Mobile App Recommendations: An Empirical Study
- RESCUE: Retrieval Augmented Secure Code Generation
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Software Testing with Large Language Models: An Interview Study with Practitioners
- TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
- Will AI also replace inspectors? Investigating the potential of generative AIs in usability inspection
- SIADAFIX: issue description response for adaptive program repair
- Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization
- On Pretraining for Project-Level Code Completion
- Context-Aware Visual Prompting: Automating Geospatial Web Dashboards with Large Language Models and Agent Self-Validation for Decision Support
- Model-Assisted and Human-Guided: Perceptions and Practices of Software Professionals Using LLMs for Coding
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
- CodeGenLink: A Tool to Find the Likely Origin and License of Automatically Generated Code
- Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM Reasoning
- AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work
- CodeChemist: Test-Time Scaling for Low-Resource Code Generation via Functional Knowledge Transfer
- BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
- Evaluating SAP Joule for Code Generation
- Large language models for behavioral modeling: A literature survey
- Unit Test Update through LLM-Driven Context Collection and Error-Type-Aware Refinement
- PromptDebt: A Comprehensive Study of Technical Debt Across LLM Projects
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- Configuration Smells in AGENTS.md Files: Common Mistakes in Configuring Coding Agents
- DocFetch - Towards Generating Software Documentation from Multiple Software Artifacts
- On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
- On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
- Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias
- Automating Code Generation for Semiconductor Equipment Control from Developer Utterances with LLMs
- Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
- Prompting the Professoriate: A Qualitative Study of Instructor Perspectives on LLMs in Data Science Education
- Development of Automated Software Design Document Review Methods Using Large Language Models
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Enhancing LLM-based Specification Generation via Program Slicing and Logical Deletion
- Analyzing the Instability of Large Language Models in Automated Bug Injection and Correction
- Students' Perception of LLM Use in Requirements Engineering Education: An Empirical Study Across Two Universities
- Combining TSL and LLM to Automate REST API Testing: A Comparative Study
- Using LLMs and Essence to Support Software Practice Adoption
- Comparative Evaluation of Large Language Models for Test-Skeleton Generation
- PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs
- The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Model-Driven Quantum Code Generation Using Large Language Models and Retrieval-Augmented Generation
- Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
- SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness
- Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
- "My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- On the synchronization between Hugging Face pre-trained language models and their upstream GitHub repository
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- Large Language Models in the Data Science Lifecycle: A Systematic Mapping Study
- Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software
- Secure and Scalable Blockchain Voting: A Comparative Framework and the Role of Large Language Models
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Key-Augmented Neural Triggers for Knowledge Sharing
- Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence Tree
- Bridging Language Gaps in Open-Source Documentation with Large-Language-Model Translation
- Vision Language Model-based Testing of Industrial Autonomous Mobile Robots
- Tuning LLM-based Code Optimization via Meta-Prompting: An Industrial Perspective
- From Technical Excellence to Practical Adoption: Lessons Learned Building an ML-Enhanced Trace Analysis Tool
- Benchmarking LLMs for Unit Test Generation from Real-World Functions
- Is LLM-Generated Code More Maintainable & Reliable than Human-Written Code?
- Extension Decisions in Open Source Software Ecosystem
- Testing the Untestable? An Empirical Study on the Testing Process of LLM-Powered Software Systems
- GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts
- A Systematic Literature Review on Detecting Software Vulnerabilities with Large Language Models
- Secure coding for web applications: Frameworks, challenges, and the role of LLMs
- MultiAIGCD: A Comprehensive dataset for AI Generated Code Detection Covering Multiple Languages, Models,Prompts, and Scenarios
- Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs
- CIgrate: Automating CI Service Migration with Large Language Models
- From Prompt to Pipeline: Large Language Models for Scientific Workflow Development in Bioinformatics
- SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents
- SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
- LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
- Single Conversation Methodology: A Human-Centered Protocol for AI-Assisted Software Development
- Formal Methods Meets Readability: Auto-Documenting JML Java Code
- RefModel: Detecting Refactorings using Foundation Models
- Self-Admitted GenAI Usage in Open-Source Software
- Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks
- Evaluating the Performance and Efficiency of Sentence-BERT for Code Comment Classification
- Prompting for Performance: Exploring LLMs for Configuring Software
- Is Quantization a Deal-breaker? Empirical Insights from Large Code Models
- LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models
- Explainability as a Compliance Requirement: What Regulated Industries Need from AI Tools for Design Artifact Generation
- SAGE: A Context-Aware Approach for Mining Privacy Requirements Relevant Reviews from Mental Health Apps
- CMER: A Context-Aware Approach for Mining Ethical Concern-related App Reviews
- From Requirements to Code: Understanding Developer Practices in LLM-Assisted Software Engineering
- Multi-Agent Debate Strategies to Enhance Requirements Engineering with Large Language Models
- The role of large language models in UI/UX design: A systematic literature review
- Measuring how changes in code readability attributes affect code quality evaluation by Large Language Models
- On the Reliability and Explainability of Language Models for Program Generation
- Building a Process-Modeling Tool using Agentic AI: An Experience Report on PM4Py-UCM
- The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study
- LLMREI: Automating Requirements Elicitation Interviews with LLMs
- ChatHLS: Towards Systematic Design Automation and Optimization for High-Level Synthesis
- Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
- What Characteristics Make ChatGPT Effective for Software Issue Resolution? An Empirical Study of Task, Project, and Conversational Signals in GitHub Issues
- Can LLMs Replace Humans During Code Chunking?
- CodeMorph: Mitigating Data Leakage in Large Language Model Assessment
- May the Feedback Be with You! Unlocking the Power of Feedback-Driven Deep Learning Framework Fuzzing via LLMs
- Evaluating LLMs Effectiveness in Detecting and Correcting Test Smells: An Empirical Study
- Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
- Evaluating the Use of LLMs for Documentation to Code Traceability
- Advanced approach for Agile/Scrum Process: RetroAI++
- Anticipating Bugs: Ticket-Level Bug Prediction and Temporal Proximity Effects
- MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
- LLM-based Dynamic Differential Testing for Database Connectors with Reinforcement Learning-Guided Prompt Selection
- SoK: Automated Vulnerability Repair: Methods, Tools, and Assessments
- Augmenting the Generality and Performance of Large Language Models for Software Engineering
- Execution Guided Line-by-Line Code Generation
- ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space
- SELU: A Software Engineering Language Understanding Benchmark
- CompilerGPT: Leveraging Large Language Models for Analyzing and Acting on Compiler Optimization Reports
- Normative Conflicts and Shallow AI Alignment
- Quantum Artificial Intelligence for Software Engineering: the Road Ahead
- LLM Code Customization with Visual Results: A Benchmark on TikZ
- An Empirical Study of OpenAI API Discussions on Stack Overflow
- LogSage: An LLM-Based Framework for CI/CD Failure Detection and Remediation with Industrial Validation
- Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
- Mutation-Guided Unit Test Generation with a Large Language Model
- Foundation models in plant molecular biology: advances, challenges, and future directions
- SysLLMatic: Large Language Models are Software System Optimizers
- LLM Performance for Code Generation on Noisy Tasks
- What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews
- Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems
- Parameter-Efficient Fine-Tuning with Attributed Patch Semantic Graph for Automated Patch Correctness Assessment
- ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
- System-driven Cloud Architecture Design Support with Structured State Management and Guided Decision Assistance
- Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
- CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- The Impact of Generative AI on Creativity in Software Development: A Research Agenda
- LLM assisted web application functional requirements generation: A case study of four popular LLMs over a Mess Management System
- A Comprehensive Study on the Use of Word Embedding Models in Software Engineering Domain
- ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation
- Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
- SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
- Software Architecture Meets LLMs: A Systematic Literature Review
- Leveraging Large Language Models for Command Injection Vulnerability Analysis in Python: An Empirical Study on Popular Open-Source Projects
- Knowledge Graph Based Repository-Level Code Generation
- Integration of TinyML and LargeML: A Survey of 6G and Beyond
- Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
- From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
- VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
- The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)
- Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
- Large Language Models for Computer-Aided Design: A Survey
- Explainable Artificial Intelligence Techniques for Software Development Lifecycle: A Phase-specific Survey
- A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
- Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation
- MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
- One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
- TraceLLM: Leveraging Large Language Models with Prompt Engineering for Enhanced Requirements Traceability
- AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
- An Empirical Study on the Effectiveness of Large Language Models for Binary Code Understanding
- LELANTE: LEveraging LLM for Automated ANdroid TEsting
- An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
- A Systematic Literature Review of Parameter-Efficient Fine-Tuning for Large Code Models
- The Buy-or-Build Decision, Revisited: How Agentic AI Changes the Economics of Enterprise Software
- The Semi-Executable Stack: Agentic Software Engineering and the Expanding Scope of SE
- Buy versus Build an LLM: A Decision Framework for Governments
- A Catalog of Data Smells for Coding Tasks
- Tracking the Moving Target: A Framework for Continuous Evaluation of LLM Test Generation in Industry
- LLMpatronous: Harnessing the Power of LLMs For Vulnerability Detection
- From Horizontal Layering to Vertical Integration: A Comparative Study of the AI-Driven Software Development Paradigm
- LogSieve: Task-Aware CI Log Reduction for Sustainable LLM-Based Analysis
- Detecting LLM-Generated Text with Performance Guarantees
- CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation
- Making AI Visible, Not Vanished: How AI Policies Reshape Developer Experience on GitHub
- LRASGen: LLM-based RESTful API Specification Generation
- Frontier AI's Impact on the Cybersecurity Landscape
- Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
- On Developers' Self-Declaration of AI-Generated Code: An Analysis of Practices
- Optimizing Token Consumption in LLMs: A Nano Surge Approach for Code Reasoning Efficiency
- Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
- Combating Toxic Language: A Review of LLM-Based Strategies for Software Engineering
- Simplicity by Obfuscation: Evaluating LLM-Driven Code Transformation with Semantic Elasticity
- DRAFT-ing Architectural Design Decisions using LLMs
- Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
Related