Dziri, Nouha
- Faith and Fate: Limits of Transformers on Compositionality
2023/05/29 by Nouha Dziri, Ximing Lu, Dziri, Nouha +29 · 13 voices · 71 citations
Computer Science · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Natural Language Processing Techniques #Topic Modeling
- Self-Refine: Iterative Refinement with Self-Feedback
2023/03/30 by Aman Madaan, Madaan, Aman, Niket Tandon +29 · 1 voice · 564 citations
Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.AI #cs.CL #cs.LG
- 2 OLMo 2 Furious
2024/12/31 by Team OLMo, OLMo, Team, P N Walsh +85 · 9 voices · 86 citations
Computer Science · Medicine · #Topic Modeling #Artificial Intelligence in Healthcare and Education #Natural Language Processing Techniques
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
2024/11/22 by Nathan Lambert, Lambert, Nathan, Jacob Morrison +44 · 9 voices · 234 citations
Computer Science · #Natural Language Processing Techniques
- The Generative AI Paradox: "What It Can Create, It May Not Understand"
2023/10/31 by Peter West, Ximing Lu, West, Peter +27 · 5 voices · 13 citations
Computer Science · Social Sciences · Medicine · #cs.AI #cs.CL #cs.CV #cs.LG
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
2025/10/27 by Liwei Jiang, Jiang, Liwei, Yuanjun Chai +17 · 15 voices · 13 citations
#cs.CL
- Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
2023/06/02 by Zeqiu Wu, Wu, Zeqiu, Yushi Hu +15 · 1 voice · 33 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.CL
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
2024/06/26 by Han, Seungju, Rao, Kavel, Ettinger, Allyson +5 · 88 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- RewardBench: Evaluating Reward Models for Language Modeling
2024/03/20 by Nathan Lambert, Valentina Pyatkin, Lambert, Nathan +21 · 65 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
2024/10/05 by Ximing Lu, Lu, Ximing, Melanie Sclar +19 · 8 voices · 13 citations
#cs.CL
- A Roadmap to Pluralistic Alignment
2024/02/07 by Taylor Sorensen, Sorensen, Taylor, Jared Moore +21 · 1 voice · 30 citations
Social Sciences · Computer Science · Medicine · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
2024/06/26 by Liwei Jiang, Kavel Rao, Jiang, Liwei +19 · 49 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
2023/12/04 by Bill Yuchen Lin, Lin, Bill Yuchen, Abhilasha Ravichander +13 · 33 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
2024/06/07 by Lin, Bill Yuchen, Deng, Yuntian, Chandu, Khyathi +6 · 35 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
2025/06/23 by Sun, Yiyou, Hu, Shawn, Zhou, Georgia +4 · 2 voices · 22 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- FaithDial: A Faithful Benchmark for Information-Seeking Dialogue
2022/04/22 by Nouha Dziri, Dziri, Nouha, Ehsan Kamalloo +11 · 14 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models?
2022/04/17 by Nouha Dziri, Dziri, Nouha, Sivan Milton +7 · 13 citations
Social Sciences · Psychology · Computer Science · #Misinformation and Its Impacts #Mental Health via Writing #Machine Learning in Healthcare
- Evaluating Open-Domain Question Answering in the Era of Large Language Models
2023/05/11 by Kamalloo, Ehsan, Dziri, Nouha, Clarke, Charles L. A. +1 · 15 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- The Art of Saying No: Contextual Noncompliance in Language Models
2024/07/02 by Faeze Brahman, Sachin Kumar, Brahman, Faeze +26 · 2 voices · 14 citations
Computer Science · #Hate Speech and Cyberbullying Detection #Multi-Agent Systems and Negotiation #cs.AI #cs.CL #cs.HC
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
2023/10/12 by Linlu Qiu, Liwei Jiang, Qiu, Linlu +19 · 9 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
2024/04/16 by Li, Huihan, Jiang, Liwei, Hwang, Jena D. +7 · 10 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
2024/07/10 by Kaitlyn Zhou, Zhou, Kaitlyn, Jena D. Hwang +9 · 2 voices · 4 citations
Decision Sciences · Computer Science · #Complex Systems and Decision Making #Software Engineering Techniques and Practices #Cognitive Science and Mapping
- Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning
2023/05/24 by Lu, Ximing, Brahman, Faeze, West, Peter +14 · 5 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark
2021/04/30 by Nouha Dziri, Dziri, Nouha, Hannah Rashkin +5 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction
2024/10/10 by Rector-Brooks, Jarrid, Hasan, Mohsin, Peng, Zhangzhi +9 · 7 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
2025/07/08 by Sanidhya Vijayvargiya, Vijayvargiya, Sanidhya, Akshay Soni +11 · 11 citations
Computer Science · #Adversarial Robustness in Machine Learning #Security and Verification in Computing #Explainable Artificial Intelligence (XAI)
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
2024/10/22 by Jing‐Jing Li, Valentina Pyatkin, Li, Jing-Jing +17 · 6 citations
Decision Sciences · Engineering · Computer Science · #Risk and Safety Analysis #Safety Systems Engineering in Autonomy #Adversarial Robustness in Machine Learning
- Evaluating Coherence in Dialogue Systems using Entailment
2019/04/06 by Dziri, Nouha, Kamalloo, Ehsan, Mathewson, Kory W. +1 · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective
2025/02/20 by Huang, Yue, Gao, Chujie, Wu, Siyuan +63 · 8 citations
#Computers and Society (cs.CY) #FOS: Computer and information sciences
- What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
2023/10/24 by Kavel Rao, Rao, Kavel, Liwei Jiang +13 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
- Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
2025/04/16 by Youfa Sun, Sun, Yiyou, Bai, Haoyue +9 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- CHAMPAGNE: Learning Real-world Conversation from Large-Scale Web Videos
2023/03/17 by Han, Seungju, Hessel, Jack, Dziri, Nouha +2 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- To Err is AI : A Case Study Informing LLM Flaw Reporting Practices
2024/10/15 by Sean McGregor, Allyson Ettinger, McGregor, Sean +24 · 2 voices · 1 citation
Computer Science · #cs.CY #cs.LG #cs.SE
- RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
2025/09/25 by Yiyou Sun, Yuhan Cao, Sun, Yiyou +11 · 6 citations
Computer Science · #Distributed and Parallel Computing Systems
- Elastic Weight Removal for Faithful and Abstractive Dialogue Generation
2023/03/30 by Daheim, Nico, Dziri, Nouha, Sachan, Mrinmaya +2 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)