Sydney Levine
- Imagining and building wise machines: The centrality of AI metacognition
2024/11/04 by Samuel G. B. Johnson, Johnson, Samuel G. B., Amir-Hossein Karimi +19 · 10 voices · 4 citations
Computer Science · #Reinforcement Learning in Robotics #cs.AI #cs.CY #cs.HC
- Language Model Alignment in Multilingual Trolley Problems
2024/07/02 by Zhijing Jin, Jin, Zhijing, Max Kleiman‐Weiner +22 · 2 voices · 14 citations
Computer Science · #Natural Language Processing Techniques
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
2025/05/20 by Yu Ying Chiu, Zhilin Wang, Chiu, Yu Ying +11 · 1 voice · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CY #cs.HC #cs.LG
- SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
2024/10/22 by Jing‐Jing Li, Valentina Pyatkin, Li, Jing-Jing +17 · 6 citations
Decision Sciences · Engineering · Computer Science · #Risk and Safety Analysis #Safety Systems Engineering in Autonomy #Adversarial Robustness in Machine Learning
- Can Language Models Reason about Individualistic Human Values and Preferences?
2024/10/04 by Liwei Jiang, Taylor Sorensen, Jiang, Liwei +5 · 5 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation
- Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
2025/02/17 by Hyunwoo Kim, Melanie Sclar, Kim, Hyunwoo +13 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Investigating machine moral judgement through the Delphi experiment
2025/01/13 by Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula +15 · 2 voices · 3 citations
Arts and Humanities · Neuroscience · Social Sciences · #Epistemology, Ethics, and Metaphysics #Ethics and Social Impacts of AI #Psychology of Moral and Emotional Judgment
- Resource Rational Contractualism Should Guide AI Alignment
2025/06/20 by Sydney Levine, Matija Franklin, Levine, Sydney +19 · 2 voices
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #cs.AI
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
2025/10/18 by Yu Ying Chiu, Chiu, Yu Ying, Michael S. Lee +30 · 2 citations
Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Psychology of Moral and Emotional Judgment