Adversarial NLI: A New Benchmark for Natural Language Understanding
2019/10/31 by Yixin Nie, Adina Williams, Nie, Yixin +9 · 1 voice · 122 citations
Computer Science · #Adversarial Robustness in Machine Learning #Adversarial system #Algorithm #Artificial intelligence #Benchmark (surveying) #Computer science #Machine learning #Multimodal Machine Learning Applications #Natural language processing #Programming language #Set (abstract data type) #State (computer science) #Strengths and weaknesses #Test set #Topic Modeling #Training set #Variety (cybernetics) #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1910.14599
published in arXiv (Cornell University) (Cornell University) · ACL 2020
openalex publication_date 2019/10/31 · arxiv created 2020/05/06 · arxiv updated 2020/05/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure. We show that training models on this new dataset leads to state-of-the-art performance on a variety of popular NLI benchmarks, while posing a more difficult challenge with its new test set. Our analysis sheds light on the shortcomings of current state-of-the-art models, and shows that non-expert annotators are successful at finding their weaknesses. The data collection method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate.
Citations
Cited by
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- FaithLens: Detecting and Explaining Faithfulness Hallucination
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- Do Generalisation Results Generalise?
- Trification: A Comprehensive Tree-based Strategy Planner and Structural Verification for Fact-Checking
- A Rosetta Stone for AI Benchmarks
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
- Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
- LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
- Analyzing and Mitigating Negation Artifacts using Data Augmentation for Improving ELECTRA-Small Model Accuracy
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- The Gray Zone of Faithfulness: Taming Ambiguity in Unfaithfulness Detection
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
- A Survey on Evaluation of Large Language Models
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- SocialNLI: A Dialogue-Centric Social Inference Dataset
- Large Language Models Hallucination: A Comprehensive Survey
- RL-Guided Data Selection for Language Model Finetuning
- Boundary on the Table: Efficient Black-Box Decision-Based Attacks for Structured Data
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Pass
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
- A Dynamic Knowledge Update-Driven Model with Large Language Models for Fake News Detection
- MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor
- Language Models are Few-Shot Learners
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- Recipes for Safety in Open-domain Chatbots
- Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models
- GIER: Gap-Driven Self-Refinement for Large Language Models
- Adversarial Training for Large Neural Language Models
- DELIVER: A System for LLM-Guided Coordinated Multi-Robot Pickup and Delivery using Voronoi-Based Relay Planning
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- Utility is in the Eye of the User: A Critique of NLP Leaderboards
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- Semantic Encryption: Secure and Effective Interaction with Cloud-based Large Language Models via Semantic Transformation
- VAULT: Vigilant Adversarial Updates via LLM-Driven Retrieval-Augmented Generation for NLI
- FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
- Adversarial Defence without Adversarial Defence: Enhancing Language Model Robustness via Instance-level Principal Component Removal
- Diffusion Beats Autoregressive in Data-Constrained Settings
- AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
- From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
- Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?
- MathDuels: Evaluating LLMs as Problem Posers and Solvers
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
- Agent Identity Evals: Measuring Agentic Identity
- Towards Compute-Optimal Many-Shot In-Context Learning
- Probabilistic distances-based hallucination detection in LLMs with RAG
- DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
- ASR-GLUE: A New Multi-task Benchmark for ASR-Robust Natural Language Understanding
- Our Evaluation Metric Needs an Update to Encourage Generalization
- SAGE: A Context-Aware Approach for Mining Privacy Requirements Relevant Reviews from Mental Health Apps
- CMER: A Context-Aware Approach for Mining Ethical Concern-related App Reviews
- On the Effect of Uncertainty on Layer-wise Inference Dynamics
- Towards a Principled Evaluation of Knowledge Editors
- Train-before-Test Harmonizes Language Model Rankings
- PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models
- QueueEDIT: Structural Self-Correction for Sequential Model Editing in LLMs
- Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
- NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
- Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
- MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
- You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
- Tau-Eval: A Unified Evaluation Framework for Useful and Private Text Anonymization
- SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
- A MISMATCHED Benchmark for Scientific Natural Language Inference
- Syntactic Data Augmentation Increases Robustness to Inference Heuristics
- Exploring Explanations Improves the Robustness of In-Context Learning
- Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
- Navigating the Accuracy-Size Trade-Off with Flexible Model Merging
- Geometry matters: Exploring language examples at the decision boundary
- Generating Semantically Valid Adversarial Questions for TableQA
- LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
- Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
- Deep Learning Based Text Classification: A Comprehensive Review
- Deploying Lifelong Open-Domain Dialogue Learning
- Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result
- AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- Generalizable Process Reward Models via Formally Verified Training Data
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Analog Foundation Models
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- Domain Regeneration: How well do LLMs match syntactic properties of text domains?
- IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Always Tell Me The Odds: Fine-grained Conditional Probability Estimation
- Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
- AI Evaluation Should Require Standardized Item-Level Data Releases
- AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
- Pushing the boundary on Natural Language Inference
- The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
- Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
- FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation
- FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
- CRAVE: A Conflicting Reasoning Approach for Explainable Claim Verification Using LLMs
- aiXamine: Simplified LLM Safety and Security
- Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models
- From Misleading Queries to Accurate Answers: A Three-Stage Fine-Tuning Method for LLMs
- Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
- Enhancing Classifier Evaluation: A Fairer Benchmarking Strategy Based on Ability and Robustness
- Debate-Feedback: A Multi-Agent Framework for Efficient Legal Judgment Prediction
Discussions
Related