vix.ing · top · new · best · stats · spec

Vidgen, Bertie

  1. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
    2023/08/02 by Röttger, Paul, Kirk, Hannah Rose, Vidgen, Bertie +3 · 68 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  2. Dynabench: Rethinking Benchmarking in NLP
    2021/04/07 by Kiela, Douwe, Bartolo, Max, Nie, Yixin +16 · 36 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  3. Why human-AI relationships need socioaffective alignment
    2025/02/04 by Hannah Rose Kirk, Iason Gabriel, Kirk, Hannah Rose +7 · 3 voices · 28 citations
    #cs.HC #cs.AI
  4. FinanceBench: A New Benchmark for Financial Question Answering
    2023/11/20 by Islam, Pranab, Kannappan, Anand, Kiela, Douwe +3 · 35 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational Engineering #FOS: Computer and information sciences #Finance #Machine Learning (stat.ML) #and Science (cs.CE)
  5. TrustLLM: Trustworthiness in Large Language Models
    2024/01/01 by Yue Huang, Huang, Yue, Lichao Sun +136 · 30 citations
    Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Privacy-Preserving Technologies in Data #Explainable Artificial Intelligence (XAI)
  6. Learning from the Worst: Dynamically Generated Datasets to Improve\n Online Hate Detection
    2020/12/31 by Bertie Vidgen, Tristan Thrush, Vidgen, Bertie +5 · 13 citations
    Computer Science · #Hate Speech and Cyberbullying Detection #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning
  7. Introducing v0.5 of the AI Safety Benchmark from MLCommons
    2024/04/18 by Bertie Vidgen, Vidgen, Bertie, Adarsh Agrawal +202 · 2 voices · 10 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  8. The AI Productivity Index (APEX)
    2025/09/30 by Bertie Vidgen, Vidgen, Bertie, Abby Fennelly +37 · 3 voices · 4 citations
    Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Economics and business #General Economics (econ.GN) #Human-Computer Interaction (cs.HC) #cs.AI #cs.CL #cs.HC #econ.GN
  9. SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
    2023/11/14 by Bertie Vidgen, Vidgen, Bertie, Nino Scherrer +11 · 1 voice · 7 citations
    #cs.CL
  10. Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
    2023/03/09 by Hannah Rose Kirk, Bertie Vidgen, Kirk, Hannah Rose +5 · 8 citations
    Business, Management and Accounting · Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #FinTech, Crowdfunding, Digital Finance #Open Source Software Innovations
  11. SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
    2024/04/08 by Röttger, Paul, Pernisi, Fabio, Vidgen, Bertie +1 · 9 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  12. WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
    2024/05/01 by Styles, Olly, Miller, Sam, Cerda-Mardini, Patricio +3 · 9 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multiagent Systems (cs.MA)
  13. SemEval-2023 Task 10: Explainable Detection of Online Sexism
    2023/03/07 by Hannah Rose Kirk, Wenjie Yin, Kirk, Hannah Rose +5 · 6 citations
    Computer Science · #Hate Speech and Cyberbullying Detection
  14. Multilingual HateCheck: Functional Tests for Multilingual Hate Speech Detection Models
    2022/06/20 by Röttger, Paul, Seelawi, Haitham, Nozza, Debora +2 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  15. The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
    2023/10/11 by Hannah Rose Kirk, Andrew M. Bean, Kirk, Hannah Rose +7 · 5 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
  16. Near to Mid-term Risks and Opportunities of Open-Source Generative AI
    2024/04/25 by Francisco Eiras, Eiras, Francisco, Aleksandar Petrov +45 · 2 voices · 4 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  17. Handling and Presenting Harmful Text in NLP Research
    2022/04/29 by Kirk, Hannah Rose, Birhane, Abeba, Vidgen, Bertie +1 · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  18. LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
    2024/12/17 by Jon Saad-Falcon, Rajan Vivek, Saad-Falcon, Jon +14 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  19. Detecting East Asian Prejudice on Social Media
    2020/05/08 by Bertie Vidgen, Vidgen, Bertie, Austin Botelho +15 · 2 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Social Media and Politics #Social and Information Networks (cs.SI)
  20. AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
    2025/02/19 by Shaona Ghosh, Ghosh, Shaona, Heather Frase +200 · 1 voice · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
  21. MSTS: A Multimodal Safety Test Suite for Vision-Language Models
    2025/01/17 by Paul Röttger, Röttger, Paul, Giuseppe Attanasio +43 · 2 voices · 2 citations
    Computer Science · #Natural Language Processing Techniques #cs.CL
  22. Detecting weak and strong Islamophobic hate speech on social media
    2018/12/12 by Bertie Vidgen, Taha Yasseri, Vidgen, Bertie +1 · 1 citation
    Computer Science · Social Sciences · #Hate Speech and Cyberbullying Detection #Social Media and Politics
  23. Risks and Opportunities of Open-Source Generative AI
    2024/05/14 by Francisco Eiras, Eiras, Francisco, Aleksander Petrov +47 · 1 citation
    Computer Science · Decision Sciences · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Scientific Computing and Data Management
  24. Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
    2025/12/01 by Kirk, Hannah Rose, Davidson, Henry, Saunders, Ed +4 · 4 citations
    #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
  25. The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
    2024/06/24 by Longpre, Shayne, Biderman, Stella, Albalak, Alon +20 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)