ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
2023/06/08 by Qiaoyu Tang, Tang, Qiaoyu, Ziliang Deng +9 · 89 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Software Engineering Research
paper · pdf · doi:10.48550/arxiv.2306.05301
Abstract
Enabling large language models to utilize real-world tools effectively is crucial for achieving embodied intelligence. Existing approaches to tool learning have either primarily relied on extremely large language models, such as GPT-4, to attain generalized tool-use abilities in a zero-shot manner, or utilized supervised learning to train limited scopes of tools on compact models. However, it remains uncertain whether smaller language models can achieve generalized tool-use abilities without tool-specific training. To address this question, this paper introduces ToolAlpaca, a novel framework designed to automatically generate a diverse tool-use corpus and learn generalized tool-use abilities on compact language models with minimal human intervention. Specifically, ToolAlpaca first automatically creates a highly diversified tool-use corpus by building a multi-agent simulation environment. The corpus contains 3938 tool-use instances from more than 400 real-world tool APIs spanning 50 distinct categories. Subsequently, the constructed corpus is employed to fine-tune compact language models, resulting in two models, namely ToolAlpaca-7B and ToolAlpaca-13B, respectively. Finally, we evaluate the ability of these models to utilize previously unseen tools without specific training. Experimental results demonstrate that ToolAlpaca achieves effective generalized tool-use capabilities comparable to those of extremely large language models like GPT-3.5, demonstrating that learning generalized tool-use ability is feasible for compact language models.
Cited by
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- Reinforcement Learning via Self-Distillation
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications
- GTM: Simulating the World of Tools for AI Agents
- AppSelectBench: Application-Level Tool Selection Benchmark
- Simulating Environments with Reasoning Models for Agent Training
- InData: Towards Secure Multi-Step, Tool-Based Data Analysis
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
- Can LLM Infer Risk Information From MCP Server System Logs?
- Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- PORTool: Tool-Use LLM Training with Rewarded Tree
- TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents
- Tools are under-documented: Simple Document Expansion Boosts Tool Retrieval
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- Self-Improving LLM Agents at Test-Time
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- Non-Collaborative User Simulators for Tool Agents
- PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
- DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks
- COCORELI: Cooperative, Compositional Reconstitution & Execution of Language Instructions
- Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
- Experiences with Model Context Protocol Servers for Science and High Performance Computing
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
- TURA: Tool-Augmented Unified Retrieval Agent for AI Search
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
- Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
- GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
- Benchmarking LLM Tool-Use in the Wild
- Routine: A Structural Planning Framework for LLM Agent System in Enterprise
- From Matching to Generation: A Survey on Generative Information Retrieval
- A Survey on Large Language Models for Mathematical Reasoning
- SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
- Execution-First Synthetic Tool-Use Trace Generation for LLM Agents
- DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
- MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models
- Teaching a Language Model to Speak the Language of Tools
- DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
- Enhancing LLM Tool Use with High-quality Instruction Data from Knowledge Graph
- Re-Initialization Token Learning for Tool-Augmented Large Language Models
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
- Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
- MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
- A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges
- ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
- Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation
- LA-RCS: LLM-Agent-Based Robot Control System
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models
- OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling
- ALOHA: Empowering Multilingual Agent for University Orientation with Hierarchical Retrieval
- ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
- Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
- When2Call: When (not) to Call Tools
- ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
- Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes
- NAACL2025 Tutorial: Adaptation of Large Language Models
- a1: Steep Test-time Scaling Law via Environment Augmented Generation
- JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
- FamilyTool: A Multi-hop Personalized Tool Use Benchmark
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
Related