Peter Henderson
- On the Opportunities and Risks of Foundation Models
2021/08/16 by Rishi Bommasani, Bommasani, Rishi, Drew A. Hudson +233 · 11 voices · 527 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning #Topic Modeling
- Deep Reinforcement Learning that Matters
2017/09/19 by Peter Henderson, Henderson, Peter, Riashat Islam +10 · 4 voices · 80 citations
Computer Science · #Evolutionary Algorithms and Applications #Reinforcement Learning in Robotics #cs.LG #stat.ML
- Foundation Models and Fair Use
2023/03/28 by Peter Henderson, Xuechen Li, Henderson, Peter +9 · 7 voices · 10 citations
Business, Management and Accounting · Computer Science · #Copyright and Intellectual Property #Digital Rights Management and Security #Law, AI, and Intellectual Property #cs.AI #cs.CY #cs.LG
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
2023/10/05 by Xiangyu Qi, Yi Zeng, Qi, Xiangyu +11 · 163 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
2020/04/15 by Miles Brundage, Brundage, Miles, Shahar Avin +115 · 2 voices · 21 citations
#cs.CY
- LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
2025/06/13 by Zihan Zheng, Zerui Cheng, Zheng, Zihan +35 · 8 voices · 16 citations
#cs.SE #cs.AI #cs.CL #cs.LG
- Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
2020/01/31 by Peter Henderson, Henderson, Peter, Jie‐Ru Hu +10 · 36 citations
Energy · Engineering · Computer Science · #Energy, Environment, and Transportation Policies #Green IT and Sustainability #Mobile Crowdsensing and Crowdsourcing
- An Introduction to Deep Reinforcement Learning
2018/11/30 by Vincent Francois-Lavet, Peter Henderson, Riashat Islam +2 · 1 voice · 6 citations
Computer Science · Mathematics · #cs.LG #cs.AI #stat.ML
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
2024/06/20 by Tinghao Xie, Xiangyu Qi, Xie, Tinghao +29 · 39 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Reliability and Analysis Research #Topic Modeling
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
2025/10/13 by Sayash Kapoor, Benedikt Stroebl, Kapoor, Sayash +63 · 3 voices · 9 citations
Computer Science · #Multi-Agent Systems and Negotiation
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
2024/02/07 by Boyi Wei, Kaixuan Huang, Wei, Boyi +15 · 33 citations
Decision Sciences · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fatigue and fracture mechanics #Machine Learning (cs.LG) #Risk and Safety Analysis
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
2021/04/18 by Lucia Zheng, Zheng, Lucia, Neel Guha +7 · 12 citations
Computer Science · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #Legal Education and Practice Innovations #Topic Modeling
- An Adversarial Perspective on Machine Unlearning for AI Safety
2024/09/26 by Jakub Łucki, Łucki, Jakub, Boyi Wei +9 · 3 voices · 19 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.AI #cs.CL #cs.CR #cs.LG
- A Safe Harbor for AI Evaluation and Red Teaming
2024/03/07 by Shayne Longpre, Longpre, Shayne, Sayash Kapoor +43 · 1 voice · 15 citations
Computer Science · #Explainable Artificial Intelligence (XAI)
- On the Societal Impact of Open Foundation Models
2024/02/27 by Sayash Kapoor, Rishi Bommasani, Kapoor, Sayash +47 · 15 citations
Engineering · #3D Modeling in Geospatial Applications #Artificial Intelligence (cs.AI) #BIM and Construction Integration #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
2022/07/01 by Peter Henderson, Mark Krass, Henderson, Peter +11 · 11 citations
Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences
- With Little Power Comes Great Responsibility
2020/10/13 by Dallas Card, Card, Dallas, Peter Henderson +9 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- Evaluating Copyright Takedown Methods for Language Models
2024/06/26 by Boyi Wei, Wei, Boyi, Weijia Shi +13 · 8 citations
Computer Science · #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences #Machine Learning (cs.LG)
- On Evaluating the Durability of Safeguards for Open-Weight LLMs
2024/12/10 by Xiangyu Qi, Boyi Wei, Qi, Xiangyu +17 · 2 voices · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #cs.AI #cs.CR
- Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models
2022/11/27 by Peter Henderson, Eric Mitchell, Henderson, Peter +7 · 4 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
2024/06/06 by Minzhou Pan, Pan, Minzhou, Yi Zeng +11 · 5 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Computer Graphics and Visualization Techniques #Digital Media Forensic Detection
- Separating value functions across time-scales
2019/02/05 by Joshua Romoff, Peter Henderson, Romoff, Joshua +8 · 3 citations
Computer Science · Decision Sciences · #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Promises and pitfalls of artificial intelligence for legal applications
2024/01/10 by Sayash Kapoor, Peter Henderson, Kapoor, Sayash +3 · 4 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computers and Society (cs.CY) #FOS: Computer and information sciences #Law, AI, and Intellectual Property
- Legal Alignment for Safe and Ethical AI
2026/01/07 by Noam Kolt, Nicholas Caputo, Jack Boeglin +14 · 4 voices · 3 citations
#cs.CY
- An Information-Theoretic Perspective on Credit Assignment in Reinforcement Learning
2021/03/10 by Dilip Arumugam, Peter Henderson, Arumugam, Dilip +3 · 1 citation
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Smart Grid Energy Management
- Dynamic Risk Assessments for Offensive Cybersecurity Agents
2025/05/23 by Boyi Wei, Benedikt Stroebl, Wei, Boyi +9 · 1 voice · 3 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Information and Cyber Security #Network Security and Intrusion Detection #Smart Grid Security and Resilience #cs.AI #cs.CR
- Where's the Liability in Harmful AI Speech?
2023/08/09 by Peter Henderson, Tatsunori Hashimoto, Henderson, Peter +3 · 1 citation
Computer Science · Social Sciences · #Law, AI, and Intellectual Property #Ethics and Social Impacts of AI #Artificial Intelligence in Law
- The Mirage of Artificial Intelligence Terms of Use Restrictions
2024/12/10 by Peter Henderson, Mark A. Lemley, Henderson, Peter +1 · 3 voices · 1 citation
#cs.CY #cs.AI #cs.LG
- Learning Robust Dialog Policies in Noisy Environments
2017/12/11 by Maryam Fazel-Zarandi, Shang-Wen Li, Fazel-Zarandi, Maryam +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Speech and dialogue systems #Topic Modeling
- LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
2024/04/02 by Joel Niklaus, Niklaus, Joel, Lucia Zheng +17 · 1 citation
Social Sciences · #68T50 #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Comparative and International Law Studies #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2 #Legal Education and Practice Innovations #Machine Learning (cs.LG)
- AI Risk Management Should Incorporate Both Safety and Security
2024/05/29 by Xiangyu Qi, Qi, Xiangyu, Yangsibo Huang +47 · 1 citation
Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences
- Breaking Down Bias: On The Limits of Generalizable Pruning Strategies
2025/02/11 by Sibo Ma, Ma, Sibo, Alejandro Salinas +5 · 1 citation
Decision Sciences · #Complex Systems and Decision Making