2023/08/19 by Yihong Dong, Kangcheng Luo, Dong, Yihong +7 · 7 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Reinforcement Learning in Robotics #Software Engineering (cs.SE) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2308.10088
openalex publication_date 2023/08/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Large language models (LLMs) have showcased remarkable potential across various tasks by conditioning on prompts. However, the quality of different human-written prompts leads to substantial discrepancies in LLMs' performance, and improving prompts usually necessitates considerable human effort and expertise. To this end, this paper proposes Prompt with Actor-Critic Editing (PACE) for LLMs to enable automatic prompt editing. Drawing inspiration from the actor-critic algorithm in reinforcement learning, PACE leverages LLMs as the dual roles of actors and critics, conceptualizing prompt as a type of policy. PACE refines prompt, taking into account the feedback from both actors performing prompt and critics criticizing response. This process helps LLMs better align prompt to a specific task, thanks to real responses and thinking from LLMs. We conduct extensive experiments on 24 instruction induction tasks and 21 big-bench tasks. Experimental results indicate that PACE elevates the relative performance of medium/low-quality human-written prompts by up to 98%, which has comparable performance to high-quality human-written prompts. Moreover, PACE also exhibits notable efficacy for prompt generation.