MapCoder: Multi-Agent Code Generation for Competitive Problem Solving
2024/05/18 by Md. Ashraful Islam, Islam, Md. Ashraful, Mohammed Eunus Ali +3 · 85 citations
Computer Science · #Advanced Database Systems and Queries #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies #Service-Oriented Architecture and Web Services
paper · pdf · doi:10.48550/arxiv.2405.11403
openalex publication_date 2024/05/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Code synthesis, which requires a deep understanding of complex natural language problem descriptions, generation of code instructions for complex algorithms and data structures, and the successful execution of comprehensive unit tests, presents a significant challenge. While large language models (LLMs) demonstrate impressive proficiency in natural language processing, their performance in code generation tasks remains limited. In this paper, we introduce a new approach to code generation tasks leveraging multi-agent prompting that uniquely replicates the full cycle of program synthesis as observed in human developers. Our framework, MapCoder, consists of four LLM agents specifically designed to emulate the stages of this cycle: recalling relevant examples, planning, code generation, and debugging. After conducting thorough experiments, with multiple LLM ablations and analyses across eight challenging competitive problem-solving and program synthesis benchmarks, MapCoder showcases remarkable code generation capabilities, achieving new state-of-the-art results (pass@1) on HumanEval (93.9%), MBPP (83.1%), APPS (22.0%), CodeContests (28.5%), and xCodeEval (45.3%). Moreover, our method consistently delivers superior performance across various programming languages and varying problem difficulties. We open-source our framework at https://github.com/Md-Ashraful-Pramanik/MapCoder.
Cited by
- Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- DeepCode: Open Agentic Coding
- Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence
- Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale
- Collaborative Agents for Automated Program Repair in Ruby
- BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attribution
- Understanding the Characteristics of LLM-Generated Property-Based Tests in Exploring Edge Cases
- TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation
- SwiftSolve: A Self-Iterative, Complexity-Aware Multi-Agent Framework for Competitive Programming
- ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
- Paper2Web: Let's Make Your Paper Alive!
- Hierarchical Sequence Iteration for Heterogeneous Question Answering
- Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
- Testing and Enhancing Multi-Agent Systems for Robust Code Generation
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Beyond Language Barriers: Multi-Agent Coordination for Multi-Language Code Generation
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
- OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama
- Analyzing and Mitigating Surface Bias in Code Evaluation Metrics
- An LLM Agentic Approach for Legal-Critical Software: A Case Study for Tax Prep Software
- Automating Code Generation for Semiconductor Equipment Control from Developer Utterances with LLMs
- Humanizing Automated Programming Feedback: Fine-Tuning Generative Models with Student-Written Feedback
- app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
- Aligning Requirement for Large Language Model's Code Generation
- CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
- COCO: Cognitive Operating System with Continuous Oversight for Multi-Agent Workflow Reliability
- LL3M: Large Language 3D Modelers
- Large Language Model-based Data Science Agent: A Survey
- CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
- AlgoSimBench: Identifying Algorithmically Similar Problems for Competitive Programming
- MemoCoder: Automated Function Synthesis using LLM-Supported Agents
- CodeEdu: A Multi-Agent Collaborative Platform for Personalized Coding Education
- AgentMaster: A Multi-Agent Conversational Framework Using A2A and MCP Protocols for Multimodal Information Retrieval and Analysis
- AuditCoder: Responsibility-Preserving Task Graphs for Auditable Code Generation and Bounded Repair
- Conversational Education at Scale: A Multi-LLM Agent Workflow for Procedural Learning and Pedagogic Quality Assessment
- EvoAgentX: An Automated Framework for Evolving Agentic Workflows
- Exploring Advanced LLM Multi-Agent Systems Based on Blackboard Architecture
- Evaluating and Improving Large Language Models for Competitive Program Generation
- The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
- Execution Guided Line-by-Line Code Generation
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System
- Mutation-Guided Unit Test Generation with a Large Language Model
- FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- EvoGit: Decentralized Code Evolution via Git-Based Multi-Agent Collaboration
- RMoA: Optimizing Mixture-of-Agents through Diversity Maximization and Residual Compensation
- RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
- Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems
- C3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
- SEW: Self-Evolving Agentic Workflows for Automated Code Generation
- From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
- SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
- X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs
- MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
- P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
- NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging
- CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
- EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
- POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
- Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
- AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
- ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
- ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation
- Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
- ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories
- Foundation Models for Software Engineering of Cyber-Physical Systems: the Road Ahead
- AdaCoder: An Adaptive Planning and Multi-Agent Framework for Function-Level Code Generation
Related