Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
2025/10/16 by Reid T. Johnson, Johnson, Reid T., Michelle D. Pain +3 · 1 voice · 1 citation
Computer Science · #cs.CL
paper · pdf · doi:10.48550/arxiv.2510.14453
Abstract
We present Natural Language Tools (NLT), a framework that replaces programmatic JSON tool calling in large language models (LLMs) with natural language outputs. By decoupling tool selection from response generation, NLT eliminates task interference and format constraints that degrade tool call performance. When evaluated across 10 models and 6,400 trials spanning customer service and mental health domains, NLT improves tool calling accuracy by 18.4 percentage points while reducing output variance by 70%. Open-weight models see the largest gains, surpassing flagship closed-weight alternatives, with implications for model training in both reinforcement learning and supervised fine-tuning stages. These improvements persist under prompt perturbations and extend tool-calling capabilities to models lacking native support.
Citations
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function Calling
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
- Boundaries in Health Settings: A Discursive Paper
- Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
- BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
- SLOT: Structuring the Output of Large Language Models
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
- Towards LLMs Robustness to Changes in Prompt Format Styles
- NoLiMa: Long-Context Evaluation Beyond Literal Matching
- On the Emergence of Position Bias in Transformers
- Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
- Why Does the Effective Context Length of LLMs Fall Short?
- AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
- Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
- Serial Position Effects of Large Language Models
- What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- Do language models plan ahead for future tokens?
- LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
- Investigating Continual Pretraining in Large Language Models: Insights and Implications
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
- EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
- Large Language Models in Mental Health Care: a Scoping Review
- Preventing Language Models From Hiding Their Reasoning
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
- Attention Sorting Combats Recency Bias In Long Context Language Models
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- Lost in the Middle: How Language Models Use Long Contexts
- Tool Learning with Foundation Models
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Large Language Models are Zero-Shot Reasoners
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- WebGPT: Browser-assisted question-answering with human feedback
- Finetuned Language Models Are Zero-Shot Learners
- Language Models are Few-Shot Learners
- ELIZA—a computer program for the study of natural language communication between man and machine
Cited by
Discussions
Related