CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges
2024/01/14 by Kechi Zhang, Zhang, Kechi, Jia Li +7 · 118 citations
Computer Science · #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering (cs.SE) #Software Engineering Research #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2401.07339
openalex publication_date 2024/01/14 · openalex created_date 2024/01/18 · openalex updated_date 2026/07/28
Abstract
Large Language Models (LLMs) have shown promise in automated code generation but typically excel only in simpler tasks such as generating standalone code units. Real-world software development, however, often involves complex code repositories (named repo) with complex dependencies and extensive documentation. To fill this gap, our research pivots towards evaluating LLMs in a more realistic setting -- real-world repo-level code generation. We introduce CodeAgentBench, a manually curated benchmark for repo-level code generation. This benchmark comprises five high-quality Python projects, encompassing a total of 101 samples. We assess nine leading LLMs on repo-level tasks and observe a decline in their performance. To tackle this, we present CodeAgent, a novel LLM-based agent framework that employs external tools for effective repo-level code generation. CodeAgent integrates five programming tools, enabling interaction with software artifacts for information retrieval, code symbol navigation, and code testing. We implement four agent strategies to optimize these tools' usage. Our experiments on CodeAgentBench show that CodeAgent enhances LLM performance significantly, with improvements ranging from 18.1% to 250%. Further tests on the HumanEval benchmark confirm CodeAgent's adaptability and efficacy across various code generation tasks. Notably, CodeAgent outperforms commercial products like Github Copilot, showcasing superior accuracy and efficiency. These results demonstrate CodeAgent's robust capabilities in code generation, highlighting its potential for real-world repo-level coding challenges.
Cited by
- SoK: Understanding (New) Security Issues Across AI4Code Use Cases
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC
- SoMe: A Realistic Benchmark for LLM-based Social Media Agents
- Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships
- Towards a Science of Scaling Agent Systems
- DeepCode: Open Agentic Coding
- BabelCoder: Agentic Code Translation with Specification Alignment
- Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
- Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- AnimAgents: Coordinating Multi-Stage Animation Pre-Production with Human-Multi-Agent Collaboration
- A Viable Paradigm of Software Automation: Iterative End-to-End Automated Software Development
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
- Towards Realistic Project-Level Code Generation via Multi-Agent Collaboration and Semantic Architecture Modeling
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- Sherlock: Reliable and Efficient Agentic Workflow Execution
- Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
- LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
- ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
- Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
- Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- Signature in Code Backdoor Detection, how far are we?
- Document Intelligence in the Era of Large Language Models: A Survey
- FalseCrashReducer: Mitigating False Positive Crashes in OSS-Fuzz-Gen Using Agentic AI
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Training-Free Group Relative Policy Optimization
- InvThink: Premortem Reasoning for Safer Language Models
- Which Programming Language and Model Work Best With LLM-as-a-Judge For Code Retrieval?
- SCUBA: Salesforce Computer Use Benchmark
- TENET: Leveraging Tests Beyond Validation for Code Generation
- Rethinking Reward Miscalibration of GRPO in Agentic RL
- The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Rethinking Explainable Disease Prediction: Synergizing Accuracy and Reliability via Reflective Cognitive Architecture
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
- SWE-QA: Can Language Models Answer Repository-level Code Questions?
- (P)rior(D)yna(F)low: A Priori Dynamic Workflow Construction via Multi-Agent Collaboration
- A Study on Thinking Patterns of Large Reasoning Models in Code Generation
- From Evaluation to Enhancement: Large Language Models for Zero-Knowledge Proof Code Generation
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
- Code2MCP: Transforming Code Repositories into MCP Services
- RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
- CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
- Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
- Dynamic Benchmark Construction for Evaluating Large Language Models on Real-World Codes
- MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Trainable Dynamic Mask Sparse Attention
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- A Survey of LLM-based Deep Search Agents: Paradigm, Optimization, Evaluation, and Challenges
- Large Language Model-based Data Science Agent: A Survey
- TripTailor: A Real-World Benchmark for Personalized Travel Planning
- How Far Are AI Scientists from Changing the World?
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation
- SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
- CodeEdu: A Multi-Agent Collaborative Platform for Personalized Coding Education
- Simulation Code Generation for Fluid Systems using Large Language Models: Benchmarking Models and Prompting Strategies
- Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
- FaSTA^*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
- Generalizing vision-language models to novel domains: A comprehensive survey
- Tracing Errors, Constructing Fixes: Repository-Level Memory Error Repair via Typestate-Guided Context Retrieval
- General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
- CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
- Computational Thinking Reasoning in Large Language Models
- CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
- SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
- A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
- Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
- Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems
- RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
- SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
- Position: Foundation Models for Tabular Data within Systemic Contexts Need Grounding
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
- Route to Reason: Adaptive Routing for LLM and Reasoning Strategy Selection
- Effective and Efficient Context Retrieval via Partial Dependency Graph for Repository-Level Code Generation
- LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
- From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection
- Text-to-Pipeline: Bridging Natural Language and Data Preparation Pipelines
- JARVIS: A Multi-Agent Code Assistant for High-Quality EDA Script Generation
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- Adversarial Reasoning for Repair Based on Inferred Program Intent
- Group-in-Group Policy Optimization for LLM Agent Training
- MigrationBench: Repository-Level Code Migration Benchmark from Java 8
- Sakura: An Approach for Generating Complex Tests from Natural Language Test Descriptions
- SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
- Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
- ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
- Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- Agent Skill Framework: Perspectives on the Potential of Small to Medium Language Models in Industrial Environments
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement
- AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers
- Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
- Coding Agents are Effective Long-Context Processors
- Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
- Self-Evolving Coding Agents
- Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
- Can large language models generate geospatial code?
- Towards an Understanding of Context Utilization in Code Intelligence
Related