vix.ing · top · new · best · stats · spec

Maarten Sap

  1. Agents of Chaos
    2026/02/23 by Natalie Shapira, Chris Wendler, Avery Yen +35 · 76 voices · 5 citations
    #cs.AI #cs.CY
  2. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
    2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 135 citations
    #cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
  3. Medical Hallucinations in Foundation Models and Their Impact on Healthcare
    2025/02/26 by Yubin Kim, Hyewon Jeong, Kim, Yubin +51 · 12 voices · 27 citations
    #cs.CL #cs.AI #cs.CY
  4. Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs
    2022/10/24 by Maarten Sap, Sap, Maarten, Ronan LeBras +5 · 2 voices · 15 citations
    Computer Science · Psychology · #Action Observation and Synchronization #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling #cs.AI #cs.CL
  5. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language\n Models
    2020/09/23 by Samuel Gehman, Suchin Gururangan, Gehman, Samuel +7 · 118 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  6. Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
    2025/10/27 by Liwei Jiang, Yuanjun Chai, Jiang, Liwei +17 · 15 voices · 13 citations
    #cs.CL
  7. DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
    2021/05/07 by Alisa Liu, Maarten Sap, Liu, Alisa +11 · 41 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  8. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and\n Implicit Hate Speech Detection
    2022/03/17 by Thomas Hartvigsen, Hartvigsen, Thomas, Saadia Gabriel +9 · 45 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection
  9. FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
    2023/10/24 by Hyunwoo Kim, Melanie Sclar, Kim, Hyunwoo +11 · 2 voices · 22 citations
    Computer Science · Psychology · #Social Robot Interaction and HRI #Speech and dialogue systems #Topic Modeling #cs.AI #cs.CL
  10. SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
    2023/10/18 by Xuhui Zhou, Hao Zhu, Zhou, Xuhui +19 · 54 citations
    Computer Science · #Topic Modeling #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques
  11. Social Bias Frames: Reasoning about Social and Power Implications of\n Language
    2019/11/10 by Maarten Sap, Sap, Maarten, Saadia Gabriel +9 · 28 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts #Text Readability and Simplification
  12. Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
    2023/10/27 by Niloofar Mireshghallah, Hyunwoo Kim, Mireshghallah, Niloofar +11 · 41 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Privacy-Preserving Technologies in Data #Topic Modeling
  13. WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
    2024/06/26 by Liwei Jiang, Kavel Rao, Jiang, Liwei +19 · 50 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences
  14. Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
    2024/01/12 by Kaitlyn Zhou, Zhou, Kaitlyn, Jena D. Hwang +5 · 3 voices · 17 citations
    Computer Science · Social Sciences · #Topic Modeling #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
  15. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
    2021/11/15 by Maarten Sap, Swabha Swayamdipta, Sap, Maarten +9 · 20 citations
    Computer Science · Arts and Humanities · Social Sciences · #Hate Speech and Cyberbullying Detection #Discourse Analysis in Language Studies #Social Media and Politics
  16. On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
    2024/08/02 by Jen-tse Huang, Huang, Jen-tse, Jiaxu Zhou +14 · 31 citations
    Computer Science · Engineering · #Network Security and Intrusion Detection #Smart Grid Security and Resilience #Blockchain Technology Applications and Security
  17. COMET: Commonsense Transformers for Automatic Knowledge Graph\n Construction
    2019/06/12 by Antoine Bosselut, Bosselut, Antoine, Hannah Rashkin +9 · 12 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #Biomedical Text Mining and Ontologies #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
  18. AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
    2024/09/13 by SU Zhe, Zhe Su, Su, Zhe +12 · 2 voices · 10 citations
    Computer Science · Economics, Econometrics and Finance · Social Sciences · #Artificial Intelligence in Law #Law, AI, and Intellectual Property #Law, Economics, and Judicial Systems #cs.AI #cs.CL
  19. SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents
    2024/03/13 by Ruiyi Wang, Haofei Yu, Wang, Ruiyi +13 · 18 citations
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems
  20. NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
    2024/04/18 by Abhinav Rao, Akhila Yerukola, Rao, Abhinav +7 · 1 voice · 11 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
  21. BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
    2024/10/21 by Wenkai Li, Jing Liu, Li, Wenkai +9 · 15 citations
    Social Sciences · #Artificial Intelligence in Law #Computational and Text Analysis Methods
  22. Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
    2024/03/08 by Xuhui Zhou, Zhe Su, Zhou, Xuhui +7 · 1 voice · 10 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
  23. ProsocialDialog: A Prosocial Backbone for Conversational Agents
    2022/05/25 by Hyunwoo Kim, Kim, Hyunwoo, Youngjae Yu +13 · 6 citations
    Computer Science · #Topic Modeling #Speech and dialogue systems #Natural Language Processing Techniques
  24. AutoPresent: Designing Structured Visuals from Scratch
    2025/01/01 by Jiaxin Ge, Ge, Jiaxin, Zora Zhiruo Wang +19 · 13 citations
    Computer Science · Engineering · #Architecture and Computational Design #Augmented Reality Applications #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation
  25. Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
    2024/07/10 by Kaitlyn Zhou, Zhou, Kaitlyn, Jena D. Hwang +9 · 2 voices · 4 citations
    Decision Sciences · Computer Science · #Complex Systems and Decision Making #Software Engineering Techniques and Practices #Cognitive Science and Mapping
  26. HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs
    2024/05/27 by Jocelyn Shen, Joel Mire, Shen, Jocelyn +7 · 1 voice · 6 citations
    #cs.CL
  27. User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
    2024/09/01 by Xianzhe Fan, Qing Xiao, Fan, Xianzhe +11 · 10 citations
    Social Sciences · #Ethics and Social Impacts of AI
  28. Living in the Past, Present, and Future: Measuring Temporal Orientation With Language
    2015/12/28 by Gregory Park, H. Andrew Schwartz, Maarten Sap +11 · 3 citations
    Psychology · #Psychological and Temporal Perspectives Research #Psychological Well-being and Life Satisfaction #Identity, Memory, and Therapy
  29. Where Do People Tell Stories Online? Story Detection Across Online Communities
    2023/11/16 by Maria Antoniak, Antoniak, Maria, Joel Mire +7 · 5 citations
    Computer Science · Health Professions · #Computation and Language (cs.CL) #Digital Storytelling and Education #FOS: Computer and information sciences #Video Analysis and Summarization
  30. OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
    2025/07/08 by Sanidhya Vijayvargiya, Akshay Soni, Vijayvargiya, Sanidhya +11 · 12 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Security and Verification in Computing #Explainable Artificial Intelligence (XAI)
  31. Un-Straightening Generative AI: How Queer Artists Surface and Challenge Model Normativity
    2025/03/12 by Jordan Taylor, Joel Mire, Franchesca Spektor +4 · 2 voices · 5 citations
    Computer Science · Social Sciences · Neuroscience · #Innovative Human-Technology Interaction #Ethics and Social Impacts of AI #Aesthetic Perception and Analysis
  32. Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
    2024/03/21 by Jimin Mun, Mun, Jimin, Liwei Jiang +13 · 4 citations
    Social Sciences · #Ethics and Social Impacts of AI
  33. SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
    2024/10/22 by Jing‐Jing Li, Valentina Pyatkin, Li, Jing-Jing +17 · 6 citations
    Decision Sciences · Engineering · Computer Science · #Risk and Safety Analysis #Safety Systems Engineering in Autonomy #Adversarial Robustness in Machine Learning
  34. Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
    2025/02/18 by Sanidhya Vijayvargiya, Xuhui Zhou, Vijayvargiya, Sanidhya +7 · 8 citations
    Business, Management and Accounting · Computer Science · #Artificial Intelligence (cs.AI) #Business Process Modeling and Analysis #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Service-Oriented Architecture and Web Services
  35. Modeling Empathic Similarity in Personal Narratives
    2023/05/23 by Jocelyn Shen, Shen, Jocelyn, Maarten Sap +7 · 3 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Misinformation and Its Impacts #Topic Modeling
  36. Rejected Dialects: Biases Against African American Language in Reward Models
    2025/02/18 by Joel Mire, Zubin Trivadi Aysola, Mire, Joel +9 · 1 voice · 3 citations
    Arts and Humanities · #Language, Linguistics, Cultural Analysis
  37. Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online Hate
    2024/02/29 by Jimin Mun, Mun, Jimin, Cathy Buerger +7 · 3 citations
    Computer Science · #Hate Speech and Cyberbullying Detection
  38. Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
    2025/04/15 by Qiaosi Wang, Xuhui Zhou, Wang, Qiaosi +7 · 1 voice · 4 citations
    #cs.HC #cs.AI
  39. Fluid Language Model Benchmarking
    2025/09/14 by Valentin Hofmann, Hofmann, Valentin, David Heineman +17 · 3 voices · 7 citations
    #cs.CL #cs.AI #cs.LG
  40. Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
    2025/02/24 by Akhila Yerukola, Saadia Gabriel, Yerukola, Akhila +5 · 1 voice · 2 citations
    #cs.AI #cs.CL #cs.CV #cs.LG
  41. PowerTransformer: Unsupervised Controllable Revision for Biased Language Correction
    2020/10/26 by Xinyao Ma, Maarten Sap, Ma, Xinyao +5 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  42. Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
    2023/05/24 by Ashutosh Baheti, Ximing Lu, Baheti, Ashutosh +9 · 1 voice · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
  43. Just Say No: Analyzing the Stance of Neural Dialogue Generation in\n Offensive Contexts
    2021/08/26 by Ashutosh Baheti, Baheti, Ashutosh, Maarten Sap +5 · 1 citation
    Computer Science · #Advanced Malware Detection Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection
  44. SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
    2025/06/29 by Xi-Long Fan, Xuhui Zhou, Fan, Xianzhe +9 · 4 citations
    Computer Science · Psychology · #Action Observation and Synchronization #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO) #Social Robot Interaction and HRI
  45. The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
    2025/09/22 by Jiaxu Zhou, Jen-tse Huang, Zhou, Jiaxu +14 · 1 voice · 1 citation
    Business, Management and Accounting · Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Corporate Governance and Law #Corporate Insolvency and Governance #FOS: Computer and information sciences #cs.CL #cs.CY
  46. Towards Countering Essentialism through Social Bias Reasoning
    2023/03/28 by Emily Allaway, Nina Taneja, Allaway, Emily +5 · 1 citation
    Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Social and Intergroup Psychology
  47. Investigating machine moral judgement through the Delphi experiment
    2025/01/13 by Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula +15 · 2 voices · 2 citations
    Arts and Humanities · Neuroscience · Social Sciences · #Epistemology, Ethics, and Metaphysics #Ethics and Social Impacts of AI #Psychology of Moral and Emotional Judgment
  48. The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
    2025/07/08 by Scott Geng, Geng, Scott, Hamish Ivison +11 · 4 citations
    Computer Science · #Topic Modeling #Domain Adaptation and Few-Shot Learning #Machine Learning and Data Classification
  49. Why (not) use AI? Analyzing People's Reasoning and Conditions for AI Acceptability
    2025/02/11 by Wei Bin Au Yeong, Mun, Jimin, Wesley Hanwen Deng +6 · 1 voice · 1 citation
    Neuroscience · Social Sciences · #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Psychology of Moral and Emotional Judgment
  50. Out of Style: RAG's Fragility to Linguistic Variation
    2025/04/11 by Cao, Tianyu, Akhila Yerukola, Bhandari, Neel +5 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  51. 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
    2025/08/11 by Wenkai Li, Liwen Sun, Li, Wenkai +8 · 1 voice · 3 citations
    Computer Science · Social Sciences · #Access Control and Trust #Cryptography and Data Security #Privacy-Preserving Technologies in Data #cs.AI
  52. Training Proactive and Personalized LLM Agents
    2025/11/04 by Weiwei Sun, Sun, Weiwei, Xuhui Zhou +13 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mobile Crowdsensing and Crowdsourcing #Software Engineering Research #Spreadsheets and End-User Computing