Gorilla: Large Language Model Connected with Massive APIs
2023/05/24 by Shishir G. Patil, Tianjun Zhang, Patil, Shishir G. +5 · 4 voices · 215 citations
#cs.CL #cs.AI
paper · pdf · doi:10.48550/arxiv.2305.15334
Abstract
Large Language Models (LLMs) have seen an impressive wave of advances recently, with models now excelling in a variety of tasks, such as mathematical reasoning and program synthesis. However, their potential to effectively use tools via API calls remains unfulfilled. This is a challenging task even for today's state-of-the-art LLMs such as GPT-4, largely due to their inability to generate accurate input arguments and their tendency to hallucinate the wrong usage of an API call. We release Gorilla, a finetuned LLaMA-based model that surpasses the performance of GPT-4 on writing API calls. When combined with a document retriever, Gorilla demonstrates a strong capability to adapt to test-time document changes, enabling flexible user updates or version changes. It also substantially mitigates the issue of hallucination, commonly encountered when prompting LLMs directly. To evaluate the model's ability, we introduce APIBench, a comprehensive dataset consisting of HuggingFace, TorchHub, and TensorHub APIs. The successful integration of the retrieval system with Gorilla demonstrates the potential for LLMs to use tools more accurately, keep up with frequently updated documentation, and consequently increase the reliability and applicability of their outputs. Gorilla's code, model, data, and demo are available at https://gorilla.cs.berkeley.edu
Cited by
- Nested Browser-Use Learning for Agentic Information Seeking
- Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing
- SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- With Great Capabilities Come Great Responsibilities: Introducing the Agentic Risk & Capability Framework for Governing Agentic AI Systems
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
- VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
- PATHFinder Agent for Tailored Prenatal Care
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Memelang: An Axial Grammar for LLM-Generated Vector-Relational Queries
- Dynamic Tool Dependency Retrieval for Efficient Function Calling
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Optimizing Agentic Language Model Inference via Speculative Tool Calls
- Multi-Agent Collaborative Framework for Intelligent IT Operations: An AOI System with Context-Aware Compression and Dynamic Task Scheduling
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
- ceLLMate: Sandboxing Browser AI Agents
- MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Reliable agent engineering should integrate machine-compatible organizational principles
- SoK: Trust-Authorization Mismatch in LLM Agent Interactions
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
- An Empirical Study of Agent Developer Practices in AI Agent Frameworks
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- Toward a Safe Internet of Agents
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- AppSelectBench: Application-Level Tool Selection Benchmark
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
- Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
- Simulating Environments with Reasoning Models for Agent Training
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
- InData: Towards Secure Multi-Step, Tool-Based Data Analysis
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- MCP-RiskCue: Can LLM Infer Risk Information From MCP Server System Logs?
- LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- What a diff makes: automating code migration with large language models
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- PORTool: Tool-Use LLM Training with Rewarded Tree
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
- Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions
- Serve Programs, Not Prompts
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- Declarative Techniques for NL Queries over Heterogeneous Data
- ScaleCall -- Agentic Tool Calling at Scale for Fintech: Challenges, Methods, and Deployment Insights
- MCP4IFC: IFC-Based Building Design Using Large Language Models
- Structured Interfaces for Automated Reasoning with 3D Scene Graphs
- Pie: A Programmable Serving System for Emerging LLM Applications
- TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents
- Tools are under-documented: Simple Document Expansion Boosts Tool Retrieval
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering
- ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Rethinking On-policy Optimization for Query Augmentation
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- ToolCritic: Detecting and Correcting Tool-Use Errors in Dialogue Systems
- Repairing Tool Calls Using Post-tool Execution Reflection and RAG
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies
- NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- A Matter of Representation: Towards Graph-Based Abstract Code Generation
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applications
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- How Good Are LLMs at Processing Tool Outputs?
- Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
- Fundamentals of Building Autonomous LLM Agents
- GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing
- Opponent Shaping in LLM Agents
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
- Self-Improving LLM Agents at Test-Time
- Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- A Case for Declarative LLM-friendly Interfaces for Improved Efficiency of Computer-Use Agents
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Quantitative Certification of Agentic Tool Selection
- VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
- Code2Video: A Code-centric Paradigm for Educational Video Generation
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
- SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
- Non-Collaborative User Simulators for Tool Agents
- Library Hallucinations in LLMs: Risk Analysis Grounded in Developer Queries
- ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
- IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
- OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
- Evaluating and Mitigating Errors in LLM-Generated Web API Integrations
- Online-Optimized RAG for Tool Use and Function Calling
- CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
- Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
- Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
- AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools
- SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
- Auditable Early Stopping for Agentic Routing: Ledger-Verified Run-Wise Certificates under Local DP
- Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
- Mind Your Server: A Systematic Study of Parasitic Toolchain Attacks on the MCP Ecosystem
- Code2MCP: Transforming Code Repositories into MCP Services
- MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
- COCORELI: Cooperative, Compositional Reconstitution & Execution of Language Instructions
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- Survey of Specialized Large Language Model
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- ToolACE-MT: Non-Autoregressive Generation for Agentic Multi-Turn Interaction
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- Hell or High Water: Evaluating Agentic Recovery from External Failures
- FROGENT: An End-to-End Full-process Drug Design Agent
- LibRec: Benchmarking Retrieval-Augmented LLMs for Library Migration Recommendations
- NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation
- Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
- 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
- HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
- Tool Graph Retriever: Exploring Dependency Graph-based Tool Retrieval for Large Language Models
- Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling
- Towards Enforcing Company Policy Adherence in Agentic Workflows
- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- Evaluation and Benchmarking of LLM Agents: A Survey
- MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
- Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems
- From REST to MCP: An Empirical Study of API Wrapping and Automated Server Generation for LLM Agents
- Augmented Vision-Language Models: A Systematic Review
- Auto: The AGI Compiler
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
- Tracking Capabilities for Safer Agents
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
- Routine: A Structural Planning Framework for LLM Agent System in Enterprise
- ParaStudent: Generating and Evaluating Realistic Student Code by Teaching LLMs to Struggle
- eSapiens: A Platform for Secure and Auditable Retrieval-Augmented Generation
- eSapiens's DEREK Module: Deep Extraction & Reasoning Engine for Knowledge with LLMs
- Evaluating LLMs on Sequential API Call Through Automated Test Generation
- ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
- Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Integrating External Tools with Large Language Models to Improve Accuracy
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
- Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
- Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
- 2024 NASA SUITS Report: LLM-Driven Immersive Augmented Reality User Interface for Robotics and Space Exploration
- MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models
- LLM Agents Are the Antidote to Walled Gardens
- Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
- IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
- DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
- More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
Discussions
Related