Hale, Scott A.
- Why human-AI relationships need socioaffective alignment
2025/02/04 by Hannah Rose Kirk, Iason Gabriel, Kirk, Hannah Rose +7 · 3 voices · 29 citations
Computer Science · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.HC
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
2024/04/18 by Bertie Vidgen, Vidgen, Bertie, Adarsh Agrawal +202 · 2 voices · 11 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
2023/03/09 by Hannah Rose Kirk, Bertie Vidgen, Kirk, Hannah Rose +5 · 9 citations
Business, Management and Accounting · Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FinTech, Crowdfunding, Digital Finance #Open Source Software Innovations
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
2023/11/14 by Bertie Vidgen, Vidgen, Bertie, Nino Scherrer +11 · 1 voice · 7 citations
#cs.CL
- Claim Matching Beyond English to Scale Global Fact-Checking
2021/06/01 by Kazemi, Ashkan, Garimella, Kiran, Gaffney, Devin +1 · 6 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
2024/06/10 by Andrew M. Bean, Bean, Andrew M., Simi Hellsten +13 · 10 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
2023/10/11 by Hannah Rose Kirk, Kirk, Hannah Rose, Andrew M. Bean +7 · 6 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
2025/11/03 by Andrew M. Bean, Bean, Andrew M., Angelika Romanou +78 · 10 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling
- From Languages to Geographies: Towards Evaluating Cultural Bias in Hate Speech Datasets
2024/04/27 by Manuel Tonneau, Tonneau, Manuel, Diyi Liu +9 · 4 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on Twitter
2024/11/23 by Manuel Tonneau, Diyi Liu, Tonneau, Manuel +11 · 1 voice · 3 citations
#cs.CL
- Matching Tweets With Applicable Fact-Checks Across Languages
2022/02/14 by Kazemi, Ashkan, Li, Zehua, Pérez-Rosas, Verónica +2 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising "Alignment" in Large Language Models
2023/10/03 by Hannah Rose Kirk, Kirk, Hannah Rose, Bertie Vidgen +5 · 1 voice · 1 citation
Computer Science · #cs.CL #cs.CY
- Lost in translation: using global fact-checks to measure multilingual misinformation prevalence, spread, and evolution
2023/10/27 by Dorian Quelle, Calvin Cheng, Quelle, Dorian +5 · 1 citation
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Social and Information Networks (cs.SI)
- Query Rewriting for Effective Misinformation Discovery
2022/10/14 by Ashkan Kazemi, Kazemi, Ashkan, Artem Abzaliev +11 · 1 citation
Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Spam and Phishing Detection #Topic Modeling
- Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
2025/12/01 by Kirk, Hannah Rose, Davidson, Henry, Saunders, Ed +4 · 4 citations
#FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
- Analyzing Misinformation Claims During the 2022 Brazilian General Election on WhatsApp, Twitter, and Kwai
2024/01/04 by Scott A. Hale, Hale, Scott A., Adriano Belisario +5 · 1 citation
Computer Science · Social Sciences · #Computers and Society (cs.CY) #FOS: Computer and information sciences #FOS: Physical sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Physics and Society (physics.soc-ph) #Social Media and Politics #Social and Information Networks (cs.SI)
- Fairness via AI: Bias Reduction in Medical Information
2021/09/06 by Dori-Hacohen, Shiri, Montenegro, Roberto, Murai, Fabricio +4 · 1 citation
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG)