A Survey on Large Language Models for Code Generation
2024/06/01 by Juyong Jiang, J.-H.R. Jiang, Fan Wang +8 · 213 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering (cs.SE)
paper · pdf · doi:10.48550/arxiv.2406.00515
openalex publication_date 2024/06/01 · openalex created_date 2024/06/06 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) have garnered remarkable advancements across diverse code-related tasks, known as Code LLMs, particularly in code generation that generates source code with LLM from natural language descriptions. This burgeoning field has captured significant interest from both academic researchers and industry professionals due to its practical significance in software development, e.g., GitHub Copilot. Despite the active exploration of LLMs for a variety of code tasks, either from the perspective of natural language processing (NLP) or software engineering (SE) or both, there is a noticeable absence of a comprehensive and up-to-date literature review dedicated to LLM for code generation. In this survey, we aim to bridge this gap by providing a systematic literature review that serves as a valuable reference for researchers investigating the cutting-edge progress in LLMs for code generation. We introduce a taxonomy to categorize and discuss the recent developments in LLMs for code generation, covering aspects such as data curation, latest advances, performance evaluation, ethical implications, environmental impact, and real-world applications. In addition, we present a historical overview of the evolution of LLMs for code generation and offer an empirical comparison using the HumanEval, MBPP, and BigCodeBench benchmarks across various levels of difficulty and types of programming tasks to highlight the progressive enhancements in LLM capabilities for code generation. We identify critical challenges and promising opportunities regarding the gap between academia and practical development. Furthermore, we have established a dedicated resource GitHub page (https://github.com/juyongjiang/CodeLLMSurvey) to continuously document and disseminate the most recent advances in the field.
Cited by
- An Empirical Study of Generative AI Adoption in Software Engineering
- Agentic Software Issue Resolution with Large Language Models: A Survey
- Specification-Driven DevOps for Multi-Service Environments
- How Well Can AI Generate Backlogs from App Mockups?
- Who Is Really Playing? Strategic Interaction in AI-Guided Populations
- Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
- Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- Understanding the Role of Large Language Models in Software Engineering: Evidence from an Industry Survey
- Holistic Evaluation of State-of-the-Art LLMs for Code Generation
- SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
- A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle Solving
- Aligning Academia with Industry: An Empirical Study of Industrial Needs and Academic Capabilities in AI-Driven Software Engineering
- A Study of Library Usage in Agent-Authored Pull Requests
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- Understanding Chain-of-Thought Effectiveness in Code Generation: An Empirical and Information-Theoretic Analysis
- DeepCode: Open Agentic Coding
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Reliable agent engineering should integrate machine-compatible organizational principles
- A Hybrid Approach for EMF Code Generation:Code Templates Meet Large Language Models
- Catching UX Flaws in Code: Leveraging LLMs to Identify Usability Flaws at the Development Stage
- MANTRA: a Framework for Multi-stage Adaptive Noise TReAtment During Training
- HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
- Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents
- Towards autonomous normative multi-agent systems for Human-AI software engineering teams
- KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
- CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- Prune4Web: DOM Tree Pruning Programming for Web Agent
- Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
- Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
- QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
- Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
- LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models
- LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
- Summary-Mediated Repair: Can LLMs use code summarisation as a tool for program repair?
- Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks
- Clinician-Directed Large Language Model Software Generation for Therapeutic Interventions in Physical Rehabilitation
- NALAMAINZ at BLP-2025 Task 2: A Multi-agent Approach for Bangla Instruction to Python Code Generation
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- MermaidSeqBench: An Evaluation Benchmark for LLM-to-Mermaid Sequence Diagram Generation
- LiteCache: A Query Similarity-Driven, GPU-Centric KVCache Subsystem for Efficient LLM Inference
- Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
- Scaling Graph Chain-of-Thought Reasoning: A Multi-Agent Framework with Efficient LLM Serving
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
- InData: Towards Secure Multi-Step, Tool-Based Data Analysis
- Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
- From Natural Language to Certified H-infinity Controllers: Integrating LLM Agents with LMI-Based Synthesis
- LLM-Powered Fully Automated Chaos Engineering: Towards Enabling Anyone to Build Resilient Software Systems at Low Cost
- MURPHY: Multi-Turn GRPO for Self Correcting Code Generation
- Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency
- Assertion-Aware Test Code Summarization with Large Language Models
- An Empirical Study of Reasoning Steps in Thinking Code LLMs
- How Natural Language Proficiency Shapes GenAI Code for Software Engineering Tasks
- PEFA-AI: Advancing Open-source LLMs for RTL generation using Progressive Error Feedback Agentic-AI
- Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
- Human-AI Co-Embodied Intelligence for Scientific Experimentation and Manufacturing
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- Reasoning Planning for Language Models
- On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection
- CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
- A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI
- Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
- Predicate Renaming via Large Language Models
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
- Impossible to hide secret ...: Uncovering Security and Privacy Issues in LLM-native IDEs
- DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning
- Metis: Memory Foundation Model
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
- MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
- Understanding the Characteristics of LLM-Generated Property-Based Tests in Exploring Edge Cases
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents
- Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
- TOM-SWE: User Mental Modeling For Software Engineering Agents
- Software Engineering Agents for Embodied Controller Generation : A Study in Minigrid Environments
- InterpDetect: Interpretable Signals for Detecting Hallucinations in Retrieval-Augmented Generation
- SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
- LLM-Powered Detection of Price Manipulation in DeFi
- Designing and Evaluating Hint Generation Systems for Science Education
- Context Engineering for AI Agents in Open-Source Software
- CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
- Large Language Models for Fault Localization: An Empirical Study
- From Specification to Service: Accelerating API-First Development Using Multi-Agent Systems
- Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
- SODBench: A Large Language Model Approach to Documenting Spreadsheet Operations
- Evaluating LLM Story Generation through Large-scale Network Analysis of Social Structures
- StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration
- A Specification's Realm: Characterizing the Knowledge Required for Executing a Given Algorithm Specification
- Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
- Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
- Reasoning Distillation and Structural Alignment for Improved Code Generation
- Tutoring LLM into a Better CUDA Optimizer
- FinSight: Towards Real-World Financial Deep Research
- A Systematic Literature Review of the Use of GenAI Assistants for Code Comprehension: Implications for Computing Education Research and Practice
- On Pretraining for Project-Level Code Completion
- CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
- HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- Retrieval-Augmented Code Generation: A Survey with Focus on Repository-Level Approaches
- A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Zephyrus: An Agentic Framework for Weather Science
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- Large Reasoning Models Learn Better Alignment from Flawed Thinking
- Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
- Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey
- LogPilot: Intent-aware and Scalable Alert Diagnosis for Large-scale Online Service Systems
- Intra-request branch orchestration for efficient LLM reasoning
- Evaluating SAP Joule for Code Generation
- Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
- Alternatives To Next Token Prediction In Text Generation -- A Survey
- SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
- A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
- The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- Representing LLMs in Prompt Semantic Task Space
- A model of errors in transformers
- Library Hallucinations in LLMs: Risk Analysis Grounded in Developer Queries
- PEPS: Quantum-Inspired Reinforcement Learning for Coherent Reasoning Traces in LLMs
- Beyond Language Barriers: Multi-Agent Coordination for Multi-Language Code Generation
- TrustChain-Review: A Risk-Adaptive Blockchain and Game-Theoretic Framework for Trustworthy AI-Assisted Code Review
- SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer
- Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
- CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
- Toward PDDL Planning Copilot
- WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
- Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
- From Evaluation to Enhancement: Large Language Models for Zero-Knowledge Proof Code Generation
- Rethinking Technology Stack Selection with AI Coding Proficiency
- Beyond Autoregression: An Empirical Study of Diffusion Large Language Models for Code Generation
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
- Testing chatbots on the creation of encoders for audio conditioned image generation
- RL Fine-Tuning Heals OOD Forgetting in SFT
- Analyzing the Instability of Large Language Models in Automated Bug Injection and Correction
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- LLM-Based Instance-Driven Heuristic Bias In the Context of a Biased Random Key Genetic Algorithm
- ARSP: Automated Repair of Verilog Designs via Semantic Partitioning
- RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
- Boardwalk: Towards a Framework for Creating Board Games with LLMs
- app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
- From Linear to Hierarchical: Evolving Tree-structured Thoughts for Efficient Alpha Mining
- When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
- Reflective Paper-to-Code Reproduction Enabled by Fine-Grained Verification
- Data Auctions for Retrieval Augmented Generation
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Re4: Scientific Computing Agent with Rewriting, Resolution, Review and Revision
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations
- Interleaving Large Language Models for Compiler Testing
- Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
- Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation
- Automated Optimization Modeling through Expert-Guided Large Language Model Reasoning
- Measuring LLM Code Generation Stability via Structural Entropy
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- TRACY: Benchmarking Execution Efficiency of LLM-Based Code Translation
- Tapas Are Free! Training-Free Adaptation of Programmatic Agents via LLM-Guided Program Synthesis in Dynamic Environments
- From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation
- Constrained Decoding of Diffusion LLMs with Context-Free Grammars
- LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
- Your Coding Intent is Secretly in the Context and You Should Deliberately Infer It Before Completion
- Retrospective Sparse Attention for Efficient Long-Context Generation
- Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
- LL3M: Large Language 3D Modelers
- Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
- What Builds Effective In-Context Examples for Code Generation?
- Automated Code Development for PDE Solvers Using Large Language Models
- Optimizing Prompt Sequences using Monte Carlo Tree Search for LLM-Based Optimization
- Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models
- Embedding Alignment in Code Generation for Audio
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
- Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
- Tool-integrated Reinforcement Learning for Repo Deep Search
- Meta-RAG on Large Codebases Using Code Summarization
- EvoVLMA: Evolutionary Vision-Language Model Adaptation
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- TreeDiff: AST-Guided Code Generation with Diffusion LLMs
- Is LLM-Generated Code More Maintainable & Reliable than Human-Written Code?
- AutoBridge: Automating Smart Device Integration with Centralized Platform
- ChatVis: Large Language Model Agent for Generating Scientific Visualizations
- From Sufficiency to Reflection: Reinforcement-Guided Thinking Quality in Retrieval-Augmented Reasoning for LLMs
- Vibe Coding as a Reconfiguration of Intent Mediation in Software Development: Definition, Implications, and Research Agenda
- TriangleMix: Accelerating Prefilling via Decoding-time Contribution Sparsity
- LLMs-guided adaptive compensator: Bringing Adaptivity to Automatic Control Systems with Large Language Models
- CIgrate: Automating CI Service Migration with Large Language Models
- Prometheus: Unified Knowledge Graphs for Issue Resolution in Multilingual Codebases
- Can LLMs Solve ASP Problems? Insights from a Benchmarking Study (Extended Version)
- ReCatcher: Towards LLMs Regression Testing for Code Generation
Related