vix.ing · top · new · best · stats · spec

Anca D. Dragan

  1. Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media
    2023/05/26 by Smitha Milli, Milli, Smitha, Micah Carroll +9 · 11 voices · 17 citations
    #cs.SI #cs.CY
  2. Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
    2025/07/15 by Tomek Korbak, Korbak, Tomek, Mikita Balesni +81 · 25 voices · 70 citations
    Computer Science · #Anomaly Detection Techniques and Applications #cs.AI #cs.LG #stat.ML
  3. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 100 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  4. On the Utility of Learning about Humans for Human-AI Coordination
    2019/10/13 by Micah Carroll, Rohin Shah, Carroll, Micah +11 · 66 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics
  5. On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
    2024/11/04 by Marcus Williams, M. C. S. Williams, Micah Carroll +12 · 11 voices · 30 citations
    Computer Science · #Network Security and Intrusion Detection
  6. Evaluating Frontier Models for Dangerous Capabilities
    2024/03/20 by Mary Phuong, Matthew Aitchison, Phuong, Mary +55 · 1 voice · 26 citations
    Computer Science · Engineering · #FOS: Computer and information sciences #Infrastructure Resilience and Vulnerability Analysis #Machine Learning (cs.LG) #cs.LG
  7. Inverse Reward Design
    2017/11/08 by Dylan Hadfield-Menell, Hadfield-Menell, Dylan, Smitha Milli +7 · 36 citations
    Computer Science · Engineering · #AI-based Problem Solving and Planning #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Manufacturing Process and Optimization
  8. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
    2024/08/09 by Tom Lieberum, Lieberum, Tom, Senthooran Rajamanoharan +17 · 80 citations
    Computer Science · #Machine Learning and Data Classification
  9. Automatically Auditing Large Language Models via Discrete Optimization
    2023/03/08 by Erik Jones, Jones, Erik, Anca D. Dragan +5 · 21 citations
    Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Infrastructure Maintenance and Monitoring #Machine Learning (cs.LG) #Machine Learning and Data Classification #Software Engineering Research
  10. Efficient Iterative Linear-Quadratic Approximations for Nonlinear\n Multi-Player General-Sum Differential Games
    2019/09/10 by David Fridovich-Keil, Ellis Ratner, Fridovich-Keil, David +7 · 13 citations
    Computer Science · Engineering · #Advanced Control Systems Optimization #FOS: Computer and information sciences #FOS: Electrical engineering #Formal Methods in Verification #Reinforcement Learning in Robotics #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
  11. The Social Cost of Strategic Classification
    2018/08/25 by Smitha Milli, John P. Miller, Milli, Smitha +5 · 10 citations
    Social Sciences · Decision Sciences · #Experimental Behavioral Economics Studies #Decision-Making and Behavioral Economics #Corruption and Economic Development
  12. B-Pref: Benchmarking Preference-Based Reinforcement Learning
    2021/11/04 by Kimin Lee, Laura Smith, Lee, Kimin +5 · 11 citations
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  13. Confronting Reward Model Overoptimization with Constrained RLHF
    2023/10/06 by Ted Moskovitz, Moskovitz, Ted, Aaditya K. Singh +11 · 14 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  14. Reward-rational (implicit) choice: A unifying formalism for reward learning
    2020/02/12 by Hong Jun Jeon, Jeon, Hong Jun, Smitha Milli +3 · 8 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Formal Methods in Verification #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics #Robotics (cs.RO)
  15. Physical Interaction as Communication: Learning Robot Objectives Online\n from Human Corrections
    2021/07/05 by Dylan P. Losey, Andrea Bajcsy, Losey, Dylan P. +5 · 7 citations
    Computer Science · Engineering · Psychology · #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Robotics (cs.RO) #Social Robot Interaction and HRI #Systems and Control (eess.SY) #electronic engineering #information engineering
  16. Learning to Model the World with Language
    2023/07/31 by Jessy Lin, Lin, Jessy, Yuqing Du +11 · 9 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  17. SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
    2019/05/27 by Siddharth Reddy, Anca D. Dragan, Reddy, Siddharth +3 · 9 citations
    Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics
  18. Establishing Appropriate Trust via Critical States
    2018/10/18 by Sandy H. Huang, Huang, Sandy H., Kush Bhatia +5 · 5 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics
  19. AI Alignment with Changing and Influenceable Reward Functions
    2024/05/28 by Micah Carroll, Davis Foote, Carroll, Micah +7 · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  20. Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
    2024/03/05 by Cassidy Laidlaw, Shivam Singhal, Laidlaw, Cassidy +4 · 1 voice · 8 citations
    Computer Science · #Anomaly Detection Techniques and Applications #cs.AI #cs.LG
  21. AvE: Assistance via Empowerment
    2020/06/26 by Yuqing Du, Du, Yuqing, Stas Tiomkin +9 · 5 citations
    Engineering · Computer Science · #Robot Manipulation and Learning #Reinforcement Learning in Robotics #Human Pose and Action Recognition
  22. Context Steering: Controllable Personalization at Inference Time
    2024/05/02 by Jerry Zhi-Yang He, Sashrika Pandey, He, Jerry Zhi-Yang +5 · 6 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Personal Information Management and User Behavior #Usability and User Interface Design
  23. Optimal Cost Design for Model Predictive Control
    2021/04/23 by Avik Jain, Lawrence S. Chan, Jain, Avik +5 · 3 citations
    Engineering · #Advanced Control Systems Optimization #Control Systems and Identification #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Machine Learning (cs.LG) #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
  24. The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human Models
    2022/04/22 by Cassidy Laidlaw, Anca D. Dragan, Laidlaw, Cassidy +1 · 3 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling
  25. When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
    2024/02/27 by Leon Lang, Lang, Leon, Davis Foote +10 · 1 voice · 3 citations
    Decision Sciences · Economics, Econometrics and Finance · Neuroscience · #Decision-Making and Behavioral Economics #Occupational and Professional Licensing Regulation #Neural and Behavioral Psychology Studies
  26. Learning a Prior over Intent via Meta-Inverse Reinforcement Learning
    2018/05/31 by Kelvin Xu, Ellis Ratner, Xu, Kelvin +7 · 2 citations
    Computer Science · #Reinforcement Learning in Robotics #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning
  27. Probabilistically Safe Robot Planning with Confidence-Based Human\n Predictions
    2018/05/31 by Jaime F. Fisac, Andrea Bajcsy, Fisac, Jaime F. +11 · 2 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Autonomous Vehicle Technology and Safety #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO)
  28. A Hamilton-Jacobi Reachability-Based Framework for Predicting and\n Analyzing Human Motion for Safe Planning
    2019/10/29 by Somil Bansal, Andrea Bajcsy, Bansal, Somil +7 · 2 citations
    Computer Science · Engineering · #Gaussian Processes and Bayesian Inference #Anomaly Detection Techniques and Applications #Fault Detection and Control Systems
  29. Safety Assurances for Human-Robot Interaction via Confidence-aware Game-theoretic Human Models
    2021/09/29 by Ran Tian, Tian, Ran, Liting Sun +7 · 2 citations
    Engineering · Psychology · #Autonomous Vehicle Technology and Safety #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Robotics (cs.RO) #Traffic and Road Safety
  30. Adversaries Can Misuse Combinations of Safe Models
    2024/06/20 by Erik Jones, Jones, Erik, Anca D. Dragan +3 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Physical Unclonable Functions (PUFs) and Hardware Security
  31. On the Utility of Model Learning in HRI
    2019/01/04 by Gokul Swamy, Swamy, Gokul, Jens Schulz +7 · 2 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Reinforcement Learning in Robotics #Robotics (cs.RO)
  32. Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
    2019/03/09 by Smitha Milli, Anca D. Dragan, Milli, Smitha +1 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning #Reinforcement Learning in Robotics #Teaching and Learning Programming
  33. Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations
    2023/11/09 by Joey Hong, Hong, Joey, Sergey Levine +3 · 3 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  34. Inferring Rewards from Language in Context
    2022/04/05 by Jessy Lin, Daniel Fried, Lin, Jessy +5 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  35. The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
    2022/08/23 by Gaurav R. Ghosal, Ghosal, Gaurav R., Matthew Zurek +5 · 2 citations
    Decision Sciences · Neuroscience · Psychology · #Artificial Intelligence (cs.AI) #Decision-Making and Behavioral Economics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mental Health Research Topics #Neural and Behavioral Psychology Studies
  36. On the Sensitivity of Reward Inference to Misspecified Human Models
    2022/12/09 by Joey Hong, Kush Bhatia, Hong, Joey +3 · 2 citations
    Computer Science · Neuroscience · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Behavioral Psychology Studies
  37. Learning to Assist Humans without Inferring Rewards
    2024/11/04 by Evan Ellis, Myers, Vivek, Sergey Levine +6 · 5 citations
    Social Sciences · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Problem and Project Based Learning
  38. Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following
    2025/02/08 by Vivek Myers, Myers, Vivek, Bo Zheng +7 · 6 citations
    Computer Science · Psychology · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO) #Social Robot Interaction and HRI #Teaching and Learning Programming
  39. Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
    2024/11/07 by Joey Hong, Anca Dragan, Anca D. Dragan +4 · 1 voice · 4 citations
    Computer Science · #Natural Language Processing Techniques
  40. Translating Neuralese
    2017/04/23 by Jacob Andreas, Anca D. Dragan, Andreas, Jacob +3 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics
  41. An Efficient, Generalized Bellman Update For Cooperative Inverse\n Reinforcement Learning
    2018/06/11 by Dhruv Malik, Malik, Dhruv, Malayandi Palaniappan +9 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Reinforcement Learning in Robotics
  42. The Assistive Multi-Armed Bandit
    2019/01/24 by Lawrence S. Chan, Dylan Hadfield-Menell, Chan, Lawrence +5 · 1 citation
    Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Auction Theory and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  43. Trajectory Improvement and Reward Learning from Comparative Language Feedback
    2024/10/08 by Zhaojing Yang, Yang, Zhaojing, Miru Jun +9 · 3 citations
    Computer Science · #Natural Language Processing Techniques
  44. On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
    2019/06/23 by Rohin Shah, Noah Gundotra, Shah, Rohin +5 · 1 citation
    Computer Science · Decision Sciences · Psychology · #Artificial Intelligence (cs.AI) #Decision-Making and Behavioral Economics #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mental Health Research Topics
  45. Value Alignment Verification
    2020/12/02 by Daniel S. Brown, Brown, Daniel S., Jordan Schneider +5 · 1 citation
    Computer Science · Social Sciences · #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  46. Learning under Misspecified Objective Spaces
    2018/10/11 by Andreea Bobu, Bobu, Andreea, Andrea Bajcsy +5 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Robotics (cs.RO)
  47. Toward Grounded Commonsense Reasoning
    2023/06/14 by Minae Kwon, Hengyuan Hu, Kwon, Minae +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Robotics (cs.RO) #Topic Modeling
  48. Inducing Structure in Reward Learning by Learning Features
    2022/01/18 by Andreea Bobu, Bobu, Andreea, Marius Wiggert +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO)
  49. Teaching Robots to Span the Space of Functional Expressive Motion
    2022/03/04 by Arjun Sripathy, Sripathy, Arjun, Andreea Bobu +9 · 1 citation
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Robotics (cs.RO) #Social Robot Interaction and HRI