Bertie Vidgen
- Why human-AI relationships need socioaffective alignment
2025/02/04 by Hannah Rose Kirk, Iason Gabriel, Kirk, Hannah Rose +7 · 3 voices · 28 citations
#cs.HC #cs.AI
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
2024/04/24 by Hannah Rose Kirk, Alexander Whitefield, Paul Röttger +9 · 3 voices · 21 citations
#cs.CL
- TrustLLM: Trustworthiness in Large Language Models
2024/01/01 by Yue Huang, Huang, Yue, Lichao Sun +136 · 30 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Privacy-Preserving Technologies in Data #Explainable Artificial Intelligence (XAI)
- Learning from the Worst: Dynamically Generated Datasets to Improve\n Online Hate Detection
2020/12/31 by Bertie Vidgen, Vidgen, Bertie, Tristan Thrush +5 · 13 citations
Computer Science · #Hate Speech and Cyberbullying Detection #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
2024/04/18 by Bertie Vidgen, Vidgen, Bertie, Adarsh Agrawal +202 · 2 voices · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning
- The AI Productivity Index (APEX)
2025/09/30 by Bertie Vidgen, Vidgen, Bertie, Abby Fennelly +37 · 3 voices · 4 citations
Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Economics and business #General Economics (econ.GN) #Human-Computer Interaction (cs.HC) #cs.AI #cs.CL #cs.HC #econ.GN
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
2023/11/14 by Bertie Vidgen, Vidgen, Bertie, Nino Scherrer +11 · 1 voice · 7 citations
#cs.CL
- Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
2023/03/09 by Hannah Rose Kirk, Bertie Vidgen, Kirk, Hannah Rose +5 · 8 citations
Business, Management and Accounting · Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FinTech, Crowdfunding, Digital Finance #Open Source Software Innovations
- SemEval-2023 Task 10: Explainable Detection of Online Sexism
2023/03/07 by Hannah Rose Kirk, Kirk, Hannah Rose, Wenjie Yin +5 · 6 citations
Computer Science · #Hate Speech and Cyberbullying Detection
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
2023/10/11 by Hannah Rose Kirk, Kirk, Hannah Rose, Andrew M. Bean +7 · 5 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Near to Mid-term Risks and Opportunities of Open-Source Generative AI
2024/04/25 by Francisco Eiras, Eiras, Francisco, Aleksandar Petrov +45 · 2 voices · 4 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
2024/12/17 by Jon Saad-Falcon, Saad-Falcon, Jon, Rajan Vivek +14 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Detecting East Asian Prejudice on Social Media
2020/05/08 by Bertie Vidgen, Vidgen, Bertie, Austin Botelho +15 · 2 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Social Media and Politics #Social and Information Networks (cs.SI)
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
2025/02/19 by Shaona Ghosh, Ghosh, Shaona, Heather Frase +200 · 1 voice · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Islamophobes are not all the same! A study of far right actors on Twitter
2021/03/01 by Bertie Vidgen, Taha Yasseri, Helen Margetts · 1 voice · 1 citation
Computer Science · Social Sciences · #Electoral Systems and Political Participation #Hate Speech and Cyberbullying Detection #Populism, Right-Wing Movements
- MSTS: A Multimodal Safety Test Suite for Vision-Language Models
2025/01/17 by Paul Röttger, Röttger, Paul, Giuseppe Attanasio +43 · 2 voices · 2 citations
Computer Science · #Natural Language Processing Techniques #cs.CL
- Detecting weak and strong Islamophobic hate speech on social media
2018/12/12 by Bertie Vidgen, Vidgen, Bertie, Taha Yasseri +1 · 1 citation
Computer Science · Social Sciences · #Hate Speech and Cyberbullying Detection #Social Media and Politics
- Risks and Opportunities of Open-Source Generative AI
2024/05/14 by Francisco Eiras, Aleksander Petrov, Eiras, Francisco +47 · 1 citation
Computer Science · Decision Sciences · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Scientific Computing and Data Management
- APEX-Accounting
2026/07/29 by Julien Benchek, Austin Bennett, Jasmin Kern +8
Computer Science · #cs.AI #cs.CL #cs.HC