AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
2025/08/15 by Wang, Huanting, Gong, Jingzhi, Zhang, Huawei +2 · 7 citations
#FOS: Computer and information sciences #Software Engineering (cs.SE)
paper · doi:10.48550/arxiv.2508.11126
Abstract
AI agentic programming is an emerging paradigm where large language model (LLM)-based coding agents autonomously plan, execute, and interact with tools such as compilers, debuggers, and version control systems. Unlike conventional code generation, these agents decompose goals, coordinate multi-step processes, and adapt based on feedback, reshaping software development practices. This survey provides a timely review of the field, introducing a taxonomy of agent behaviors and system architectures and examining relevant techniques for planning, context management, tool integration, execution monitoring, and benchmarking datasets. We highlight challenges of this fast-moving field and discuss opportunities for building reliable, transparent, and collaborative coding agents.
Citations
- Kimi K2: Open Agentic Intelligence
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
- EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
- Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
- Multi-Agent Systems Execute Arbitrary Malicious Code
- ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
- Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Language Models for Code Optimization: Survey, Challenges and Future Directions
- Agent-SafetyBench: Evaluating the Safety of LLM Agents
- Exploration of LLM Multi-Agent Application Implementation Based on LangGraph+CrewAI
- PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation
- ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
- Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
- Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
- Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
- Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planning
- MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
- Qwen2.5-Coder Technical Report
- AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation through Static Analysis and Fuzz Testing
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
- Large Language Model-Based Agents for Software Engineering: A Survey
- Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
- Palu: Compressing KV-Cache with Low-Rank Projection
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
- DafnyBench: A Benchmark for Formal Software Verification
- McEval: Massively Multilingual Code Evaluation
- How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
- A Review of Prominent Paradigms for LLM-Based Agents: Tool Use (Including RAG), Planning, and Feedback Learning
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- A Survey on Large Language Models for Code Generation
- IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
- EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
- PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
- LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- A Systematic Literature Review on Large Language Models for Automated Program Repair
- From Matching to Generation: A Survey on Generative Information Retrieval
- From Matching to Generation: A Survey on Generative Information Retrieval
- LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation
- GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications
- Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
- RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
- Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- StarCoder 2 and The Stack v2: The Next Generation
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- Automated Unit Test Improvement using Large Language Models at Meta
- Instruction Tuning for Secure Code Generation
- KIVI : Plug-and-play 2bit KV Cache Quantization with Streaming Asymmetric Quantization
- TrustAgent: Towards Safe and Trustworthy LLM-based Agents
- Executable Code Actions Elicit Better LLM Agents
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- From Prompt Engineering to Prompt Science With Human in the Loop
- Retrieval-Augmented Generation for Large Language Models: A Survey
- A Comparative Analysis of Large Language Models for Code Documentation Generation
- An LLM Compiler for Parallel Function Calling
- Splitwise: Efficient generative LLM inference using phase splitting
- Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
- CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
- Persistent memory for AI coding agents: a pre-registered SWE-bench Verified benchmark
- Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Bias Testing and Mitigation in LLM-based Code Generation
- Large Language Models for Software Engineering: A Systematic Literature Review
- Large Language Models for Software Engineering: A Systematic Literature Review
- AIKernel Semantic DSL Compiler and Deterministic Agent Execution Architecture
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
- BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents
- AgentBench: Evaluating LLMs as Agents
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- ChatDev: Communicative Agents for Software Development
- Exploring and Characterizing Large Language Models For Embedded System Development and Debugging
- Lost in the Middle: How Language Models Use Long Contexts
- Citation: A Key to Building Responsible and Accountable Large Language Models
- LongNet: Scaling Transformers to 1,000,000,000 Tokens
- Augmenting Language Models with Long-Term Memory
- AdaPlanner: Adaptive Planning from Feedback with Language Models
- Uncovering and Quantifying Social Biases in Code Generation
- Large Language Models for Software Engineering: Survey and Open Problems
- Structured Chain-of-Thought Prompting for Code Generation
- Structured Chain-of-Thought Prompting for Code Generation
- Explainable Automated Debugging via Large Language Model-driven Scientific Debugging
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- Self-Refine: Iterative Refinement with Self-Feedback
- CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Machine/Deep Learning for Software Engineering: A Systematic Literature Review
- On the Reliability and Explainability of Language Models for Program Generation
- On the Reliability and Explainability of Language Models for Program Generation
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Multitask Pre-training of Modular Prompt for Chinese Few-Shot Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
- Automated Repair of Programs from Large Language Models
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Training language models to follow instructions with human feedback
- A Systematic Evaluation of Large Language Models of Code
- Investigating Explainability of Generative AI for Code through Scenario-based Design
- Jigsaw: Large Language Models meet Program Synthesis
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- PyMT5: multi-mode translation of natural language and Python code with\n transformers
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Scaling Laws for Neural Language Models
- The Art, Science, and Engineering of Fuzzing: A Survey
- The Art, Science, and Engineering of Fuzzing: A Survey
- Ray: A Distributed Framework for Emerging AI Applications
- A Survey of Machine Learning for Big Code and Naturalness
- Attention Is All You Need
- RobustFill: Neural Program Learning under Noisy I/O
- A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
- AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Cited by
Related