Hannah Rose Kirk
- AI systems out-persuade expert humans
2026/06/15 by Kobi Hackenburg, Caroline Wagner, Luke Hewitt +5 · 15 voices · 1 citation
#cs.CY #cs.AI
- Clinical knowledge in LLMs does not translate to human interactions
2025/04/26 by Andrew M. Bean, Bean, Andrew M., Rebecca Payne +19 · 15 voices · 5 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Global Health and Surgery #Explainable Artificial Intelligence (XAI)
- Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
2025/07/04 by Christopher Summerfield, Lennart Luettgau, Summerfield, Christopher +21 · 12 voices · 5 citations
#cs.AI
- Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
2024/02/26 by Paul Röttger, Röttger, Paul, Valentin Hofmann +11 · 2 voices · 34 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Why human-AI relationships need socioaffective alignment
2025/02/04 by Hannah Rose Kirk, Kirk, Hannah Rose, Iason Gabriel +7 · 3 voices · 28 citations
#cs.HC #cs.AI
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
2024/04/24 by Hannah Rose Kirk, Alexander Whitefield, Paul Röttger +9 · 3 voices · 21 citations
#cs.CL
- A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning
2022/03/22 by Hugo Berg, Berg, Hugo, Siobhan Mackenzie Hall +11 · 14 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Computers and Society (cs.CY) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
2024/04/18 by Bertie Vidgen, Adarsh Agrawal, Vidgen, Bertie +202 · 2 voices · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
2023/11/14 by Bertie Vidgen, Vidgen, Bertie, Nino Scherrer +11 · 1 voice · 7 citations
#cs.CL
- Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
2023/03/09 by Hannah Rose Kirk, Kirk, Hannah Rose, Bertie Vidgen +5 · 8 citations
Business, Management and Accounting · Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FinTech, Crowdfunding, Digital Finance #Open Source Software Innovations
- SemEval-2023 Task 10: Explainable Detection of Online Sexism
2023/03/07 by Hannah Rose Kirk, Kirk, Hannah Rose, Wenjie Yin +5 · 6 citations
Computer Science · #Hate Speech and Cyberbullying Detection
- Looking for a Handsome Carpenter! Debiasing GPT-3 Job Advertisements
2022/05/23 by Conrad Borchers, Borchers, Conrad, Dalia Gala +11 · 5 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Reinforcement Learning in Robotics
- LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
2024/06/10 by Andrew M. Bean, Simi Hellsten, Bean, Andrew M. +13 · 9 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- Auditing large language models: a three-layered approach
2023/05/30 by Jakob Mökander, Jonas Schuett, Hannah Rose Kirk +1 · 1 voice · 6 citations
Social Sciences · Medicine · #Ethics and Social Impacts of AI #Artificial Intelligence in Healthcare and Education
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
2023/10/11 by Hannah Rose Kirk, Andrew M. Bean, Kirk, Hannah Rose +7 · 5 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
2025/11/03 by Andrew M. Bean, Bean, Andrew M., Kearns, Ryan Othniel +78 · 10 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling
- Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image Generation
2024/02/14 by Jessica Quaye, Alicia Parrish, Quaye, Jessica +27 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Computers and Society (cs.CY) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG)
- Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
2025/02/23 by Jonathan Hvithamar Rystrøm, Hannah Rose Kirk, Rystrøm, Jonathan +3 · 6 citations
Arts and Humanities · Health Professions · #Translation Studies and Practices #Interpreting and Communication in Healthcare #linguistics and terminology studies
- Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
2026/02/01 by Andrew M. Bean, Rebecca Payne, Guy Parsons +8 · 2 voices · 14 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Global Health and Surgery #Persona Design and Applications
- Reward Model Interpretability via Optimal and Pessimal Tokens
2025/06/08 by Brian Christian, Christian, Brian, Hannah Rose Kirk +7 · 4 voices · 3 citations
#cs.CL #cs.AI #cs.CY #cs.LG
- Conversational AI increases political knowledge as effectively as self-directed internet search
2025/09/05 by Lennart Luettgau, Luettgau, Lennart, Hannah Rose Kirk +17 · 4 voices · 5 citations
Social Sciences · #Social Media and Politics
- Assessing Language Model Deployment with Risk Cards
2023/03/31 by Leon Derczynski, Hannah Rose Kirk, Derczynski, Leon +11 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Software Engineering Research #Topic Modeling
- The Future of Open Human Feedback
2024/08/15 by Shachar Don-Yehiya, Ben Burtenshaw, Don-Yehiya, Shachar +37 · 4 citations
Psychology · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Automation Interaction and Safety #Human-Computer Interaction (cs.HC)
- Modulating Language Model Experiences through Frictions
2024/06/24 by Katherine M. Collins, Valerie Chen, Collins, Katherine M. +15 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- People readily follow personal advice from AI but it does not improve their well-being
2025/11/19 by Lennart Luettgau, Luettgau, Lennart, Vanessa Cheung +19 · 4 voices · 2 citations
Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #Digital Mental Health Interventions #Mental Health via Writing #cs.HC
- Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
2024/12/05 by Vishakh Padmakumar, Padmakumar, Vishakh, Chuanyang Jin +5 · 1 citation
Decision Sciences · Economics, Econometrics and Finance · #Decision-Making and Behavioral Economics #Game Theory and Voting Systems