PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
2024/08/29 by Yijia Shao, Tianshi Li, Shao, Yijia +7 · 2 voices · 71 citations
Computer Science · Social Sciences · #Privacy, Security, and Data Protection #cs.AI #cs.CL #cs.CR
paper · pdf · doi:10.48550/arxiv.2409.00138
arxiv published 2024/08/29 · arxiv updated 2025/03/14
Abstract
As language models (LMs) are widely utilized in personalized communication scenarios (e.g., sending emails, writing social media posts) and endowed with a certain level of agency, ensuring they act in accordance with the contextual privacy norms becomes increasingly critical. However, quantifying the privacy norm awareness of LMs and the emerging privacy risk in LM-mediated communication is challenging due to (1) the contextual and long-tailed nature of privacy-sensitive cases, and (2) the lack of evaluation approaches that capture realistic application scenarios. To address these challenges, we propose PrivacyLens, a novel framework designed to extend privacy-sensitive seeds into expressive vignettes and further into agent trajectories, enabling multi-level evaluation of privacy leakage in LM agents' actions. We instantiate PrivacyLens with a collection of privacy norms grounded in privacy literature and crowdsourced seeds. Using this dataset, we reveal a discrepancy between LM performance in answering probing questions and their actual behavior when executing user instructions in an agent setup. State-of-the-art LMs, like GPT-4 and Llama-3-70B, leak sensitive information in 25.68% and 38.69% of cases, even when prompted with privacy-enhancing instructions. We also demonstrate the dynamic nature of PrivacyLens by extending each seed into multiple trajectories to red-team LM privacy leakage risk. Dataset and code are available at https://github.com/SALT-NLP/PrivacyLens.
Cited by
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- DREAM: Dynamic Red-teaming across Environments for AI Models
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
- Personalizing Agent Privacy Decisions via Logical Entailment
- ASTRA: Agentic Steerability and Risk Assessment Framework
- CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
- ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
- The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal Videos
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- User Perceptions of Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios
- Black Box Absorption: LLMs Undermining Innovative Ideas
- Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms
- The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- Position: Privacy Is Not Just Memorization!
- PEAR: Planner-Executor Agent Robustness Benchmark
- Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
- Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
- PrivacyMotiv: Speculative Persona Journeys for Empathic and Motivating Privacy Reviews in UX Design
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Not My Agent, Not My Boundary? Elicitation of Personal Privacy Boundaries in AI-Delegated Information Sharing
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- AI Agents May Always Fall for Prompt Injections
- Position: Human-Robot Interaction in Embodied Intelligence Demands a Shift From Static Privacy Controls to Dynamic Learning
- Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Conflect: Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications
- Beyond PII: How Users Attempt to Estimate and Mitigate Implicit LLM Inference
- PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents
- Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight
- GAMA: A General Anonymizing Multi-Agent System for Privacy Preservation Enhanced by Domain Rules and Disproof Mechanism
- SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
- Web Fraud Attacks Against LLM-Driven Multi-Agent Systems
- LM Agents May Fail to Act on Their Own Risk Knowledge
- Searching for Privacy Risks in LLM Agents via Simulation
- Never compromise with vulnerabilities: a comprehensive survey on AI governance
- Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory
- 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
- Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- Memory as a Service (MaaS): Purpose-Bound Memory Mediation for Cooperative Agents
- TAI3: Testing Agent Integrity in Interpreting User Intent
- MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
- Towards Pervasive Distributed Agentic Generative AI -- A State of The Art
- Privacy Reasoning in Ambiguous Contexts
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
- Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS
- LLM Agents Should Employ Security Principles
- Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering
- Can Large Language Models Really Recognize Your Name?
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
- Dependency-Aware Privacy for Multi-turn Agents
- Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
- DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
- Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
- AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection
- PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
- The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
- UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
Discussions
Related