Toolformer: Language Models Can Teach Themselves to Use Tools
2023/02/09 by Timo Schick, Schick, Timo, Jane Dwivedi-Yu +15 · 7 voices · 796 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2302.04761
Abstract
Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with basic functionality, such as arithmetic or factual lookup, where much simpler and smaller models excel. In this paper, we show that LMs can teach themselves to use external tools via simple APIs and achieve the best of both worlds. We introduce Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API. We incorporate a range of tools, including a calculator, a Q&A system, two different search engines, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities.
Cited by
- Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
- When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis
- Agent Security Needs Redefinition through a Holistic Framework
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
- SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
- Agentic Root Cause Analysis through Evidence-Grounded Reasoning
- CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models
- Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
- FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients
- WARA: A Closed-Loop Multi-Agent Framework for Wireless Optimization Autoresearch
- Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- Stress Testing Concept Erasure with Large Language Model Agents
- FlexNGIA 2.0: Redesigning the Internet with Agentic AI -- Protocols, Services, and Traffic Engineering Designed, Deployed, and Managed by AI
- Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization
- Natural Language to What? A Vision for Intermediate Representations in NL-to-X Querying
- Personalized Recommendation Tool Learning via Autonomous Language Agents
- GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning
- FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
- AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning
- Harnessing LLMs for Reliable Academic Supervision: A Comparative Study
- Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
- Agents in the Wild: Where Research Meets Deployment
- Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing
- FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
- LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
- Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
- RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning
- PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
- VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval
- Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
- Operational Hallucination and Safety Drift in AI Agents
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
- MemoHarness: Agent Harnesses That Learn from Experience
- Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers
- Engineering Trustworthy Agentic AI for Critical Systems
- MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments
- From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
- SkillRouter: Skill Routing for LLM Agents at Scale
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- SlotGuard: Stop Oversharing Private Local Context in LLM Agent Transcri
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
- Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs
- Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- Agentic Code Review in the Terminal: A Trajectory-Level Analysis of Behavior, Cost, and Human-Alignment
- AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
- Scalable LLM Agent Tool Access in the Cloud
- PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
- TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
- Binding Drift in Multi-Step Tool-Augmented Agents
- ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
- Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
- Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution
- Multi-Turn On-Policy Distillation with Prefix Replay
- TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
- SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
- Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
- From Intent to Infrastructure: LLM-Driven Agent Compilers for ISAC Networks
- AI Agents Do Not Fail Alone:The Context Fails First
- Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
- CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- Lifted Representation Hypothesis in Language Models
- OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- Accurate and Efficient Long-Term Memory for LLM Agents
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search
- Machine understanding
- Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis
- Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- AI Tool Discovery at Scale: All You Need is DNS
- BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data
- Calibrated Selective Fact-Checking via Evidence Chain Evaluation
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Building Trust in Autonomous Commerce: A Verifiable Global Event Timeline and AI-Ready Fraud Intelligence Layer
- PrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical Reasoning
- LaCy: What Small Language Models Can and Should Learn is Not Just a Question of Loss
- Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
- Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering
- RePo: Language Models with Context Re-Positioning
- Detailed balance in large language model-driven agents
- Epistemological Fault Lines Between Human and Artificial Intelligence
- The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
- Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Small Language Models are the Future of Agentic AI
- Build the web for agents, not agents for the web
- Securing AI Agents with Information-Flow Control
- Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- A Survey of AI Agent Protocols
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Orchestration Framework for Financial Agents: From Algorithmic Trading to Agentic Trading
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- Large Language Models for Control
- Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing
- Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
- Towards end-to-end automation of AI research
- MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems
- Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations
- Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
- Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems
- Agent Data Injection Attacks are Realistic Threats to AI Agents
- SCTA: An Agentic Framework for Stable and Interpretable Target Gene Discovery from Single-Cell RNA Sequencing
- Training Language Models to Cooperate with Inference-Time Controllers
- ACM: Agentic Context Management for Long Horizon Tasks
- Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
- SymStep: Symbolic Step Verification for Logical Reasoning
- Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming
- VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge
- ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
- Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- TGMS: An Agent-Native Bi-Temporal Graph Management System
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
- Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
- ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
- AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
- From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance
- ARdena: Scenario-driven control of real-time LLM agents
- CRAFT: Learn the Schema, Execute the Plan
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- Evaluating LLMs as Interpretable Controllers for Dynamical Systems
- ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop
- Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
- GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- The Cartesian Cut in Agentic AI
- Cognitive Dark Matter: Measuring What AI Misses
- Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented Generation
- SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?
- Method Decoration (DeMe): A Framework for LLM-Driven Adaptive Method Generation in Dynamic IoT Environments
- PERELMAN: Pipeline for scientific literature meta-analysis. Technical report
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
- Vibe Reasoning: Eliciting Frontier AI Mathematical Capabilities -- A Case Study on IMO 2025 Problem 6
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data
- Towards Efficient Agents: A Co-Design of Inference Architecture and System
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Reinforcement Learning for Self-Improving Agent with Skill Library
- Dynamic Tool Dependency Retrieval for Efficient Function Calling
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- From Personalization to Prejudice: Bias and Discrimination in Memory-Enhanced AI Agents for Recruitment
- TIB AIssistant: a Platform for AI-Supported Research Across Research Life Cycles
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
- PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving
- ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Optimizing Agentic Language Model Inference via Speculative Tool Calls
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- Georeferencing complex relative locality descriptions with large language models
- LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- Grammar Search for Multi-Agent Systems
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- Workflows vs Agents for Code Translation
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- Quantigence: A Multi-Agent AI Framework for Quantum Security Research
- Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making
- CoDA: A Context-Decoupled Hierarchical Agent with Reinforcement Learning
- AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
- ceLLMate: Sandboxing Browser AI Agents
- Large Language Models have Chain-of-Affect
- MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
- AgentBalance: Backbone-then-Topology Design for Cost-Effective Multi-Agent Systems under Budget Constraints
- Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance
- AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- A-LAMP: Agentic LLM-Based Framework for Automated MDP Modeling and Policy Generation
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents
- Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization
- Training Language Models to Use Prolog as a Tool
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
- PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- CureAgent: A Training-Free Executor-Analyst Framework for Clinical Reasoning
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
- Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
- GTM: Simulating the World of Tools for AI Agents
- Automating Complex Document Workflows via Stepwise and Rollback-Enabled Operation Orchestration
- AR-Med: Automated Relevance Enhancement in Medical Search via LLM-Driven Information Augmentation
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
- Tipping the Dominos: Topology-Aware Multi-Hop Attacks on LLM-Based Multi-Agent Systems
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- PPTArena: A Benchmark for PowerPoint Editing
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- Benchmarking and Understanding Safety Risks in AI Character Platforms
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- LLM2Fx-Tools: Tool Calling For Music Post-Production
- Energy-Aware Data-Driven Model Selection in LLM-Orchestrated AI Systems
- ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
- Toward a Safe Internet of Agents
- Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
- Query-Level Uncertainty in Large Language Models
- Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent
- From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
- ShoppingComp: Are LLMs Really Ready for Your Shopping Cart?
- Difficulties with Evaluating a Deception Detector for AIs
- OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Optimizing NetGPT via Routing-Based Synergy and Reinforcement Learning
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- On the Limits of Innate Planning in Large Language Models
- Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning
- EWE: An Agentic Framework for Extreme Weather Analysis
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- AppSelectBench: Application-Level Tool Selection Benchmark
- Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- Fara-7B: An Efficient Agentic Model for Computer Use
- A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective Analysis
- LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
- EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- Proposal of an AI-Based Support Assistant for the ALICE-FIT Detector Setup at CERN
- Why Do Language Model Agents Whistleblow?
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Computer-Use Agents as Judges for Generative User Interface
- Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems
- MedDCR: Learning to Design Agentic Workflows for Medical Coding
- Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Agent READMEs: An Empirical Study of Context Files for Agentic Coding
- Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory
- Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
- Draft and Refine with Visual Experts
- Generative Caching for Structurally Similar Prompts and Responses
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- AlphaCast: A Human Wisdom-LLM Intelligence Co-Reasoning Framework for Interactive Time Series Forecasting
- Structured Uncertainty guided Clarification for LLM Agents
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- TabRAG: Tabular Document Retrieval via Structured Language Representations
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems
- Can LLM Infer Risk Information From MCP Server System Logs?
- Catching Contamination Before Generation: Spectral Kill Switches for Agents
- Reasoning Is All You Need for Urban Planning AI
- TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
- PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models
- AnaFlow: Agentic LLM-based Workflow for Reasoning-Driven Explainable and Sample-Efficient Analog Circuit Sizing
- Learning When to Quit in Sales Conversations
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- The ORCA Benchmark: Evaluating Real-World Calculation Accuracy in Large Language Models
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Efficient Test-Time Retrieval Augmented Generation
- OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights
- GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
- Active Thinking Model: A Goal-Directed Self-Improving Framework for Real-World Adaptive Intelligence
- Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation
- SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
- Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching
- Context Engineering 2.0: The Context of Context Engineering
- SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- OracleAgent: A Multimodal Reasoning Agent for Oracle Bone Script Research
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- PORTool: Tool-Use LLM Training with Rewarded Tree
- SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents
- PILA: Plug-and-Play Insertion for LLM-native Advertising
- BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
- Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes
- PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
- Resolving Java Code Repository Issues with iSWE Agent
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents
- Flows: Building Blocks of Reasoning and Collaborating AI
- Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Exploiting LLM Agent Supply Chains via Payload-less Skills
- Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Solving Physics Olympiad via Reinforcement Learning on Physics Simulators
- The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook
- When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents
- Automated Composition of Agents: A Knapsack Approach for Agentic Component Selection
- Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
- Declarative Techniques for NL Queries over Heterogeneous Data
- Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
- ScaleCall -- Agentic Tool Calling at Scale for Fintech: Challenges, Methods, and Deployment Insights
- MCP4IFC: IFC-Based Building Design Using Large Language Models
- Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
- Evidence-Bound Autonomous Research (EviBound): A Governance Framework for Eliminating False Claims
- Structured Interfaces for Automated Reasoning with 3D Scene Graphs
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction
- VDSAgents: A PCS-Guided Multi-Agent System for Veridical Data Science Automation
- Pie: A Programmable Serving System for Emerging LLM Applications
- TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- Beyond Prompt Engineering: Neuro-Symbolic-Causal Architecture for Robust Multi-Objective AI Agents
- Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring
- StreetMath: Study of LLMs' Approximation Behaviors
- Language Server CLI Empowers Language Agents with Process Rewards
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- DeepAgent: A General Reasoning Agent with Scalable Toolsets
- When Models Outthink Their Safety: Mitigating Self-Jailbreak in Large Reasoning Models with Chain-of-Guardrails
- Large Language Models Meet Text-Attributed Graphs: A Survey of Integration Frameworks and Applications
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs
- Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
- gem5 Co-Pilot: AI Assistant Agent for Architectural Design Space Exploration
- Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
- A Design Science Blueprint for an Orchestrated AI Assistant in Doctoral Supervision
- GAPO: Robust Advantage Estimation for Real-World Code LLMs
- Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- MENTOR: A Reinforcement Learning Framework for Enabling Tool Use in Small Models via Teacher-Optimized Rewards
- Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding
- CMT-Bench: Cricket Multi-Table Generation Benchmark for Probing Robustness in Large Language Models
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
- Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- Internalizing World Models via Self-Play Finetuning for Agentic RL
- LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
- LLM Agents Beyond Utility: An Open-Ended Perspective
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies
- NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan's Intelligent Interaction Systems
- A Survey on Evaluation of Large Language Models
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
- Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- Demystifying Reinforcement Learning in Agentic Reasoning
- Characterizing Web Search in The Age of Generative AI
- Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
- Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- Agentic RAG for Software Testing with Hybrid Vector-Graph and Multi-Agent Orchestration
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
- Agentic Troubleshooting Guide Automation for Incident Management
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Therefore I am. I Think
- Signals: Trajectory Sampling and Triage for Agentic Interactions
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- Classifier-Augmented Generation for Structured Workflow Prediction
- An Alternative Trajectory for Generative AI
- How Good Are LLMs at Processing Tool Outputs?
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
- Fundamentals of Building Autonomous LLM Agents
- GRETEL: A Goal-driven Retrieval and Execution-based Trial Framework for LLM Tool Selection Enhancing
- GuruAgents: Emulating Wise Investors with Prompt-Guided LLM Agents
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- FlowSearch: Advancing deep research with dynamic structured knowledge flow
- Opponent Shaping in LLM Agents
- ToolExpander: Extending the Frontiers of Tool-Using Reinforcement Learning to Weak LLMs
- SoK: Measuring What Matters for Closed-Loop Security Agents
- AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
- ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Training-Free Group Relative Policy Optimization
- AgentAsk: Multi-Agent Systems Need to Ask
- PEAR: Planner-Executor Agent Robustness Benchmark
- ProSEA: Problem Solving via Exploration Agents
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Tool-Augmented Policy Optimization: Synergizing Reasoning and Adaptive Tool Use with Reinforcement Learning
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- Efficient numeracy in language models through single-token number embeddings
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Expanding the Action Space of LLMs to Reason Beyond Language
- Valid Stopping for LLM Generation via Empirical Dynamic Formal Lift
- A Survey on Agentic Security: Applications, Threats and Defenses
- Constrained Natural Language Action Planning for Resilient Embodied Systems
- TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection
- From Agentification to Self-Evolving Agentic AI for Wireless Networks: Concepts, Approaches, and Future Research Directions
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- TalkPlay-Tools: Conversational Music Recommendation with LLM Tool Calling
- Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- Natural Language Edge Labelling: Decoupling Intent from Execution in Structured LM Reasoning
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Where Did It All Go Wrong? A Hierarchical Look into Multi-Agent Error Attribution
- RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
- Zephyrus: An Agentic Framework for Weather Science
- Open Agent Specification (Agent Spec): A Unified Representation for AI Agents
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- On The Statistical Limits of Self-Improving Agents
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
- Mind the Goal: Data-Efficient Goal-Oriented Evaluation of Conversational Agents and Chatbots using Teacher Models
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- MASH: Modeling Abstention via Selective Help-Seeking
- LVLMs as inspectors: an agentic framework for category-level structural defect annotation
- JoyAgent-JDGenie: Technical Report on the GAIA
- TokMem: Tokenized Procedural Memory for Large Language Models
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- ACT: Agentic Classification Tree
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- R-Log: Incentivizing Log Analysis Capability in LLMs via Reasoning-based Reinforcement Learning
- LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
- SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA
- Where LLM Agents Fail and How They can Learn From Failures
- A Measurement Study of Model Context Protocol Ecosystem
- AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems
- GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners
- FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems
- Agentic Services Computing
- Beyond Manuals and Tasks: Instance-Level Context Learning for LLM Agents
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
- FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Model Context Protocol for Vision Systems: Audit, Security, and Protocol Extensions
- IROSA: Interactive Robot Skill Adaptation Using Natural Language
- Estimating the Empowerment of Language Model Agents
- InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
- PRIME: Planning and Retrieval-Integrated Memory for Enhanced Reasoning
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- Binary Autoencoder for Mechanistic Interpretability of Large Language Models
- Difference-Guided Reasoning: A Temporal-Spatial Framework for Large Language Models
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- On Theoretical Interpretations of Concept-Based In-Context Learning
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
- Evaluating and Mitigating Errors in LLM-Generated Web API Integrations
- Online-Optimized RAG for Tool Use and Function Calling
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- Beyond Sentiment: Structured Information Extraction from Financial News
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Masgent: an AI-assisted materials simulation agent
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
- DeepResearch Agent System
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
- Group-Reflective Self-Distillation for Agentic Reinforcement Learning
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
- CLVisc Agent for autonomous relativistic hydrodynamics studies
- Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
- Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
- LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
- Language Models Coupled with Metacognition Can Outperform Reasoning Models
- CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Autonomous Data Agents: A New Opportunity for Smart Data
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured Templates
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- LIMI: Less is More for Agency
- Agentic AI for Multi-Stage Physics Experiments at a Large-Scale User Facility Particle Accelerator
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
- Governing Automated Strategic Intelligence
- SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
- How Large Language Models are Designed to Hallucinate
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- Enhancing Financial RAG with Agentic AI and Multi-HyDE: A Novel Approach to Knowledge Retrieval and Hallucination Reduction
- Digging Into the Internal: Causality-Based Analysis of LLM Function Calling
- Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents
- Towards General Agentic Intelligence via Environment Scaling
- HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference
- Toward PDDL Planning Copilot
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Scaling Agents via Continual Pre-training
- Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools
- AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
- Auditable Early Stopping for Agentic Routing: Ledger-Verified Run-Wise Certificates under Local DP
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
- HyFedRAG: A Federated Retrieval-Augmented Generation Framework for Heterogeneous and Privacy-Sensitive Data
- REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents
- Code2MCP: Transforming Code Repositories into MCP Services
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- FaMA: LLM-Empowered Agentic Assistant for Consumer-to-Consumer Marketplace
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- MCPVerse: An Expansive, Real-World Benchmark for Agentic Tool Use
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
- Batch Query Processing and Optimization for Agentic Workflows
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Generative Goal Modeling
- IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Transforming Agency. On the mode of existence of Large Language Models
- COCORELI: Cooperative, Compositional Reconstitution & Execution of Language Instructions
- How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on τ-bench
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
- Provable Benefits of In-Tool Learning for Large Language Models
- A Systematic Review on the Generative AI Applications in Human Medical Genomics
- Evaluating Language Model Reasoning about Confidential Information
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Survey of Specialized Large Language Model
- Network-Level Prompt and Trait Leakage in Local Research Agents
- The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- DESAMO: A Device for Elder-Friendly Smart Homes Powered by Embedded LLM with Audio Modality
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- RETAIL: Towards Real-world Travel Planning for Large Language Models
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- COCO: Cognitive Operating System with Continuous Oversight for Multi-Agent Workflow Reliability
- AI Agents for Photonic Integrated Circuit Design Automation
- Reliability, Embeddedness, and Agency: A Utility-Driven Mathematical Framework for Agent-Centric AI Adoption
- GTool: Graph Enhanced Tool Planning with Large Language Model
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- AGENTS-LLM: Augmentative GENeration of Challenging Traffic Scenarios with an Agentic LLM Framework
- Towards Reliable Multi-Agent Systems for Marketing Applications via Reflection, Memory, and Planning
- Hell or High Water: Evaluating Agentic Recovery from External Failures
- FROGENT: An End-to-End Full-process Drug Design Agent
- PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments
- AdaptFlow: Adaptive Workflow Optimization via Meta-Learning
- LoSemB: Logic-Guided Semantic Bridging for Inductive Tool Retrieval
- HGMF: A Hierarchical Gaussian Mixture Framework for Scalable Tool Invocation within the Model Context Protocol
- CP-Agent: Agentic Constraint Programming
- CLAP: Coreference-Linked Augmentation for Passage Retrieval
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Chain-of-Alpha: Unleashing the Power of Large Language Models for Alpha Mining in Quantitative Trading
- MX-AI: Agentic Observability and Control Platform for Open and AI-RAN
- First Ask Then Answer: A Framework Design for AI Dialogue Based on Supplementary Questioning with Large Language Models
- A Novel Architecture for Symbolic Reasoning with Decision Trees and LLM Agents
- Tool Graph Retriever: Exploring Dependency Graph-based Tool Retrieval for Large Language Models
- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- TURA: Tool-Augmented Unified Retrieval Agent for AI Search
- StepWrite: Adaptive Planning for Speech-Driven Text Generation
- Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
- ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
- An Entity Linking Agent for Question Answering
- An Auditable Agent Platform For Automated Molecular Optimisation
- Agoran: An Agentic Open Marketplace for 6G RAN Automation
- Tool-integrated Reinforcement Learning for Repo Deep Search
- Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling
- Polymath: A Self-Optimizing Agent with Dynamic Hierarchical Workflow
- CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
- WarriorMath: Enhancing the Mathematical Ability of Large Language Models with a Defect-aware Framework
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- AutoEDA: Enabling EDA Flow Automation through Microservice-Based LLM Agents
- Lucy: edgerunning agentic web search on mobile with machine generated task vectors
- Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF
- ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- Beyond Natural Language Plans: Structure-Aware Planning for Query-Focused Table Summarization
- Agentic AI for autonomous anomaly management in complex systems
- An Architecture for Spatial Networking
- Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
- DeepSieve: Information Sieving via LLM-as-a-Knowledge-Router
- ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
- Anyone Can Jailbreak: Prompt-Based Attacks on LLMs and T2Is
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
- Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
- LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
- SciToolAgent: A Knowledge Graph-Driven Scientific Agent for Multi-Tool Integration
- From REST to MCP: An Empirical Study of API Wrapping and Automated Server Generation for LLM Agents
- Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
- Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
- MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service
- Augmented Vision-Language Models: A Systematic Review
- Initial Steps in Integrating Large Reasoning and Action Models for Service Composition
- Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
- Token Reduction Is Not Cost Reduction
- Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
- Auto: The AGI Compiler
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
- Beyond the Black Box: Interpretability of Agentic AI Tool Use
- AgileLog: A Forkable Shared Log for Agents on Data Streams
- Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
- Tracking Capabilities for Safer Agents
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
- AgentTrace: A Structured Logging Framework for Agent System Observability
- FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
- DesignLab: Designing Slides Through Iterative Detection and Correction
- Thinking Isn't an Illusion: Overcoming the Limitations of Reasoning Models via Tool Augmentations
- Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications
- A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
- Infherno: End-to-end Agent-based FHIR Resource Synthesis from Free-form Clinical Notes
- From Semantic Web and MAS to Agentic AI: A Unified Narrative of the Web of Agents
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- eSapiens: A Platform for Secure and Auditable Retrieval-Augmented Generation
- GoalfyMax: A Protocol-Driven Multi-Agent System for Intelligent Experience Entities
- Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
- eSapiens's DEREK Module: Deep Extraction & Reasoning Engine for Knowledge with LLMs
- Evaluating LLMs on Sequential API Call Through Automated Test Generation
- ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
- Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm
- FrugalRAG: Learning to retrieve and reason for multi-hop QA
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
- Automating MD simulations for Proteins using Large language Models: NAMD-Agent
- Integrating External Tools with Large Language Models to Improve Accuracy
- Agentic-R1: Distilled Dual-Strategy Reasoning
- Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
- Replacing thinking with tool usage enables reasoning in small language models
- Scaling Context Requires Rethinking Attention
- Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering
- PresentAgent: Multimodal Agent for Presentation Video Generation
- FinTeam: A Multi-Agent Collaborative Intelligence System for Comprehensive Financial Scenarios
- Effects of structure on reasoning in instance-level Self-Discover
- Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
- The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
- OpenTable-R1: A Reinforcement Learning Augmented Tool Agent for Open-Domain Table Question Answering
- Dynamic Strategy Adaptation in Multi-Agent Environments with Large Language Models
- WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
- Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
- MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models
- LineRetriever: Planning-Aware Observation Reduction for Web Agents
- Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice
- A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
- LLM Agents Are the Antidote to Walled Gardens
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
- VALID-Mol: a Systematic Framework for Validated LLM-Assisted Molecular Design
- IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
- Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems
- DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
- MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
- Artificial Intelligent Disobedience: Rethinking the Agency of Our Artificial Teammates
- Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
- Enhancing LLM Tool Use with High-quality Instruction Data from Knowledge Graph
- EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
- MMSearch-R1: Incentivizing LMMs to Search
- KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
- NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestration
- TableVault: Managing Dynamic Data Collections for LLM-Augmented Workflows
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- Semantic-Aware Parsing for Security Logs
- General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
- Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
- RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
- GenerationPrograms: Fine-grained Attribution with Executable Programs
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Unveiling the Learning Mind of Language Models: A Cognitive Framework and Empirical Study
- Towards Pervasive Distributed Agentic Generative AI -- A State of The Art
- LocationReasoner: Evaluating LLMs on Real-World Site Selection Reasoning
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- Levels of Autonomy for AI Agents
- Eliciting Reasoning in Language Models with Cognitive Tools
- Maximally-Informative Retrieval for State Space Model Generation
- Efficient LLM Collaboration via Planning
- PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
- GeneWhisperer: Enhancing manual genome annotation with large language models
- Execution Guided Line-by-Line Code Generation
- OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
- A Study on Individual Spatiotemporal Activity Generation Method Using MCP-Enhanced Chain-of-Thought Large Language Models
- Reinforcement learning fine-tuning of language model for instruction following and math reasoning
- Curry–Howard correspondence [wikipedia]
- Agent harness [wikipedia]
Discussions
- Toolformer: Language Models Can Teach Themselves to Use Tools [hn, 220 points, 45 comments]
- Toolformer: Language models can teach themselves to use tools [hn, 155 points, 18 comments]
- Toolformer: Language Models Can Teach Themselves to Use Tools [lobsters, 4 points, 0 comments]
- Toolformer: Language Models Can Teach Themselves to Use Tools [hn, 3 points, 0 comments]
- Paper on arXiv: arxiv.org/abs/2302.04761 [bsky, 1 points, 0 comments]
- Toolformer: Language Models Can Teach Themselves to Use Tools [bsky, 0 points, 0 comments]
- you can do this and people are doing it arxiv.org/abs/2302.04761 [bsky, 0 points, 1 comments]
Related