Anca D. Dragan
- Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media
2023/05/26 by Smitha Milli, Milli, Smitha, Micah Carroll +9 · 11 voices · 16 citations
#cs.SI #cs.CY
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 89 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- On the Utility of Learning about Humans for Human-AI Coordination
2019/10/13 by Micah Carroll, Carroll, Micah, Rohin Shah +11 · 49 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics
- On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
2024/11/04 by Marcus Williams, M. C. S. Williams, Williams, Marcus +12 · 11 voices · 28 citations
Computer Science · #Network Security and Intrusion Detection
- Evaluating Frontier Models for Dangerous Capabilities
2024/03/20 by Mary Phuong, Phuong, Mary, Matthew Aitchison +55 · 1 voice · 22 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Infrastructure Resilience and Vulnerability Analysis #Machine Learning (cs.LG) #cs.LG
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
2024/08/09 by Tom Lieberum, Senthooran Rajamanoharan, Lieberum, Tom +17 · 67 citations
Computer Science · #Machine Learning and Data Classification
- Inverse Reward Design
2017/11/08 by Dylan Hadfield-Menell, Smitha Milli, Hadfield-Menell, Dylan +7 · 28 citations
Computer Science · Engineering · #AI-based Problem Solving and Planning #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Manufacturing Process and Optimization
- Automatically Auditing Large Language Models via Discrete Optimization
2023/03/08 by Erik Jones, Anca D. Dragan, Jones, Erik +5 · 18 citations
Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Infrastructure Maintenance and Monitoring #Machine Learning (cs.LG) #Machine Learning and Data Classification #Software Engineering Research
- Efficient Iterative Linear-Quadratic Approximations for Nonlinear\n Multi-Player General-Sum Differential Games
2019/09/10 by David Fridovich-Keil, Ellis Ratner, Fridovich-Keil, David +7 · 11 citations
Computer Science · Engineering · #Advanced Control Systems Optimization #FOS: Computer and information sciences #FOS: Electrical engineering #Formal Methods in Verification #Reinforcement Learning in Robotics #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
- The Social Cost of Strategic Classification
2018/08/25 by Smitha Milli, Milli, Smitha, John P. Miller +5 · 9 citations
Social Sciences · Decision Sciences · #Experimental Behavioral Economics Studies #Decision-Making and Behavioral Economics #Corruption and Economic Development
- B-Pref: Benchmarking Preference-Based Reinforcement Learning
2021/11/04 by Kimin Lee, Lee, Kimin, Laura Smith +5 · 9 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Confronting Reward Model Overoptimization with Constrained RLHF
2023/10/06 by Ted Moskovitz, Moskovitz, Ted, Aaditya K. Singh +11 · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Physical Interaction as Communication: Learning Robot Objectives Online\n from Human Corrections
2021/07/05 by Dylan P. Losey, Andrea Bajcsy, Losey, Dylan P. +5 · 7 citations
Computer Science · Engineering · Psychology · #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Robotics (cs.RO) #Social Robot Interaction and HRI #Systems and Control (eess.SY) #electronic engineering #information engineering
- Learning to Model the World with Language
2023/07/31 by Jessy Lin, Yuqing Du, Lin, Jessy +11 · 9 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
- Establishing Appropriate Trust via Critical States
2018/10/18 by Sandy H. Huang, Kush Bhatia, Huang, Sandy H. +5 · 5 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics
- AI Alignment with Changing and Influenceable Reward Functions
2024/05/28 by Micah Carroll, Davis Foote, Carroll, Micah +7 · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- AvE: Assistance via Empowerment
2020/06/26 by Yuqing Du, Stas Tiomkin, Du, Yuqing +9 · 5 citations
Engineering · Computer Science · #Robot Manipulation and Learning #Reinforcement Learning in Robotics #Human Pose and Action Recognition
- SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
2019/05/27 by Siddharth Reddy, Reddy, Siddharth, Anca D. Dragan +3 · 6 citations
Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics
- Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
2024/03/05 by Cassidy Laidlaw, Shivam Singhal, Laidlaw, Cassidy +4 · 1 voice · 6 citations
Computer Science · #Anomaly Detection Techniques and Applications #cs.AI #cs.LG
- When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
2024/02/27 by Leon Lang, Lang, Leon, Davis Foote +10 · 1 voice · 3 citations
Decision Sciences · Economics, Econometrics and Finance · Neuroscience · #Decision-Making and Behavioral Economics #Occupational and Professional Licensing Regulation #Neural and Behavioral Psychology Studies
- A Hamilton-Jacobi Reachability-Based Framework for Predicting and\n Analyzing Human Motion for Safe Planning
2019/10/29 by Somil Bansal, Bansal, Somil, Andrea Bajcsy +7 · 2 citations
Computer Science · Engineering · #Gaussian Processes and Bayesian Inference #Anomaly Detection Techniques and Applications #Fault Detection and Control Systems
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
2019/03/09 by Smitha Milli, Anca D. Dragan, Milli, Smitha +1 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Reinforcement Learning in Robotics #Teaching and Learning Programming
- The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human Models
2022/04/22 by Cassidy Laidlaw, Anca D. Dragan, Laidlaw, Cassidy +1 · 2 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
- Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations
2023/11/09 by Joey Hong, Sergey Levine, Hong, Joey +3 · 3 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
- The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
2022/08/23 by Gaurav R. Ghosal, Ghosal, Gaurav R., Matthew Zurek +5 · 2 citations
Decision Sciences · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Decision-Making and Behavioral Economics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mental Health Research Topics #Neural and Behavioral Psychology Studies
- On the Sensitivity of Reward Inference to Misspecified Human Models
2022/12/09 by Joey Hong, Kush Bhatia, Hong, Joey +3 · 2 citations
Computer Science · Neuroscience · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Behavioral Psychology Studies
- Learning to Assist Humans without Inferring Rewards
2024/11/04 by Myers, Vivek, Evan Ellis, Sergey Levine +6 · 5 citations
Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Problem and Project Based Learning
- Adversaries Can Misuse Combinations of Safe Models
2024/06/20 by Erik Jones, Anca D. Dragan, Jones, Erik +3 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Physical Unclonable Functions (PUFs) and Hardware Security
- Translating Neuralese
2017/04/23 by Jacob Andreas, Anca D. Dragan, Andreas, Jacob +3 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics
- Learning a Prior over Intent via Meta-Inverse Reinforcement Learning
2018/05/31 by Kelvin Xu, Ellis Ratner, Xu, Kelvin +7 · 1 citation
Computer Science · #Reinforcement Learning in Robotics #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning
- An Efficient, Generalized Bellman Update For Cooperative Inverse\n Reinforcement Learning
2018/06/11 by Dhruv Malik, Malik, Dhruv, Malayandi Palaniappan +9 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Reinforcement Learning in Robotics
- Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
2024/11/07 by Joey Hong, Anca Dragan, Hong, Joey +4 · 1 voice · 3 citations
Computer Science · #Natural Language Processing Techniques
- The Assistive Multi-Armed Bandit
2019/01/24 by Lawrence S. Chan, Chan, Lawrence, Dylan Hadfield-Menell +5 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- On the Utility of Model Learning in HRI
2019/01/04 by Gokul Swamy, Jens Schulz, Swamy, Gokul +7 · 1 citation
Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics #Robotics (cs.RO)
- On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
2019/06/23 by Rohin Shah, Noah Gundotra, Shah, Rohin +5 · 1 citation
Computer Science · Decision Sciences · Psychology · #Artificial Intelligence (cs.AI) #Decision-Making and Behavioral Economics #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mental Health Research Topics
- Learning under Misspecified Objective Spaces
2018/10/11 by Andreea Bobu, Andrea Bajcsy, Bobu, Andreea +5 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Robotics (cs.RO)
- Trajectory Improvement and Reward Learning from Comparative Language Feedback
2024/10/08 by Zhaojing Yang, Miru Jun, Yang, Zhaojing +9 · 2 citations
Computer Science · #Natural Language Processing Techniques
- Toward Grounded Commonsense Reasoning
2023/06/14 by Minae Kwon, Kwon, Minae, Hengyuan Hu +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robotics (cs.RO) #Topic Modeling
- Teaching Robots to Span the Space of Functional Expressive Motion
2022/03/04 by Arjun Sripathy, Andreea Bobu, Sripathy, Arjun +9 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Robotics (cs.RO) #Social Robot Interaction and HRI