Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
2025/12/18 by Qiao, Yuxuan, Liu, Dongqin, Yang, Hongchang +2
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2512.16310
Abstract
Driven by Large Language Models, the single-agent, multi-tool architecture has become a popular paradigm for autonomous agents due to its simplicity and effectiveness. However, this architecture also introduces a new and severe privacy risk, which we term Tools Orchestration Privacy Risk (TOP-R), where an agent, to achieve a benign user goal, autonomously aggregates information fragments across multiple tools and leverages its reasoning capabilities to synthesize unexpected sensitive information. We provide the first systematic study of this risk. First, we establish a formal framework, attributing the risk's root cause to the agent's misaligned objective function: an overoptimization for helpfulness while neglecting privacy awareness. Second, we construct TOP-Bench, comprising paired leakage and benign scenarios, to comprehensively evaluate this risk. To quantify the trade-off between safety and robustness, we introduce the H-Score as a holistic metric. The evaluation results reveal that TOP-R is a severe risk: the average Risk Leakage Rate (RLR) of eight representative models reaches 90.24%, while the average H-Score is merely 0.167, with no model exceeding 0.3. Finally, we propose the Privacy Enhancement Principle (PEP) method, which effectively mitigates TOP-R, reducing the Risk Leakage Rate to 46.58% and significantly improving the H-Score to 0.624. Our work reveals both a new class of risk and inherent structural limitations in current agent architectures, while also offering feasible mitigation strategies.
Citations
- ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Position: Privacy Is Not Just Memorization!
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
- Evaluating Language Model Reasoning about Confidential Information
- Searching for Privacy Risks in LLM Agents via Simulation
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
- Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
- Unveiling Privacy Risks in LLM Agent Memory
- Tool Unlearning for Tool-Augmented LLMs
- Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
- CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
- PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
- GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Beyond Memorization: Violating Privacy Via Inference with Large Language Models
- Identifying the Risks of LM Agents with an LM-Emulated Sandbox
- The Rise and Potential of Large Language Model Based Agents: A Survey
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- Counterfactually Auditable Lifecycle Certification for Autonomous Agents
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Gorilla: Large Language Model Connected with Massive APIs
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Large Language Model Guided Tree-of-Thought
- Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
- Tool Learning with Foundation Models
- Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Analyzing Leakage of Personally Identifiable Information in Language Models
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Extracting Training Data from Large Language Models
- PIQA: Reasoning about Physical Commonsense in Natural Language
- Membership Inference Attacks against Machine Learning Models
Related