A large annotated corpus for learning natural language inference
2015/08/21 by Samuel R. Bowman, Gabor Angeli, Bowman, Samuel R. +5 · 194 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.1508.05326
Abstract
Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations. However, machine learning research in this area has been dramatically limited by the lack of large-scale resources. To address this, we introduce the Stanford Natural Language Inference corpus, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning. At 570K pairs, it is two orders of magnitude larger than all other resources of its type. This increase in scale allows lexicalized classifiers to outperform some sophisticated existing entailment models, and it allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time.
Citations
Cited by
- Instruction-Following Evaluation of Large Vision-Language Models
- The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
- Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
- Attention Is Not What You Need
- DramaBench: A Six-Dimensional Evaluation Framework for Drama Script Continuation
- Mitigating Spurious Correlations in NLI via LLM-Synthesized Counterfactuals and Dynamic Balanced Sampling
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Hard Negative Sample-Augmented DPO Post-Training for Small Language Models
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Task Matrices: Linear Maps for Cross-Model Finetuning Transfer
- Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- Do Generalisation Results Generalise?
- Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
- One Word Is Not Enough: Simple Prompts Improve Word Embeddings
- Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss Landscapes
- HEAD-QA v2: Expanding a Healthcare Benchmark for Reasoning
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Private Zeroth-Order Optimization with Public Data
- Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak Spaces
- SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- Zero-Order Sharpness-Aware Minimization
- EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
- PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
- TurkEmbed: Turkish Embedding Model on NLI & STS Tasks
- DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
- TurkEmbed4Retrieval: Turkish Embedding Model for Retrieval Task
- LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
- Analyzing and Mitigating Negation Artifacts using Data Augmentation for Improving ELECTRA-Small Model Accuracy
- ContextPilot: Fast Long-Context Inference via Context Reuse
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- Metamorphic Testing of Large Language Models for Natural Language Processing
- Routing-Based Continual Learning for Multimodal Large Language Models
- Do Students Debias Like Teachers? On the Distillability of Bias Mitigation Methods
- AR-LSAT: Investigating Analytical Reasoning of Text
- SufiSent - Universal Sentence Representations Using Suffix Encodings
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
- Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
- Cross-lingual Language Model Pretraining
- Multiple Structural Priors Guided Self Attention Network for Language Understanding
- Learning Gender-Neutral Word Embeddings
- Deep Learning Based on Generative Adversarial and Convolutional Neural Networks for Financial Time Series Predictions
- Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense Reasoning
- WT5?! Training Text-to-Text Models to Explain their Predictions
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- LLM and Human Modes of Representation
- SentEval: An Evaluation Toolkit for Universal Sentence Representations
- Distance-based Self-Attention Network for Natural Language Inference
- Evaluation of sentence embeddings in downstream and linguistic probing\n tasks
- Towards Language Agnostic Universal Representations
- Building an Evaluation Scale using Item Response Theory
- A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC
- Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
- MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference
- AfriMTEB and AfriE5: Benchmarking and Adapting Text Embedding Models for African Languages
- A Decomposable Attention Model for Natural Language Inference
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- BERT-ATTACK: Adversarial Attack Against BERT Using BERT
- Analyzing Compositionality-Sensitivity of NLI Models
- SALSA: Single-pass Autoregressive LLM Structured Classification
- NLI Data Sanity Check: Assessing the Effect of Data Corruption on Model\n Performance
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- CantoNLU: A benchmark for Cantonese natural language understanding
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
- Learning Task-Agnostic Representations through Multi-Teacher Distillation
- Rethinking Skip-thought: A Neighborhood based Approach
- DELTA: A DEep learning based Language Technology plAtform
- Scaling Language-Centric Omnimodal Representation Learning
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data
- Joint Multi-Domain Learning for Automatic Short Answer Grading
- Improving Matching Models with Hierarchical Contextualized Representations for Multi-turn Response Selection
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
- One Sentence, Two Embeddings: Contrastive Learning of Explicit and Implicit Semantic Representations
- LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning
- SocialNLI: A Dialogue-Centric Social Inference Dataset
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
- e-QRAQ: A Multi-turn Reasoning Dataset and Simulator with Explanations
- LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
- ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
- Fin-ExBERT: User Intent based Text Extraction in Financial Context using Graph-Augmented BERT and trainable Plugin
- Dense associative memory on the Bures-Wasserstein space
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
- Fine-Grained Uncertainty Decomposition in Large Language Models: A Spectral Approach
- Multilingual Vision-Language Models, A Survey
- MonoCon: A general framework for learning ultra-compact high-fidelity representations using monotonicity constraints
- Document Summarization with Conformal Importance Guarantees
- CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
- Lessons from Natural Language Inference in the Clinical Domain
- Transporting Task Vectors across Different Architectures without Training
- Combining Fact Extraction and Verification with Neural Semantic Matching Networks
- Learning to Compose Words into Sentences with Reinforcement Learning
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Modelling Domain Relationships for Transfer Learning on Retrieval-based Question Answering Systems in E-commerce
- Shortcut-Stacked Sentence Encoders for Multi-Domain Inference
- Learning to Compute Word Embeddings On the Fly
- Adversarially Regularized Autoencoders
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- Dropping Networks for Transfer Learning
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Pass
- Accurate and Efficient Low-Rank Model Merging in Core Space
- XNLI: Evaluating Cross-lingual Sentence Representations
- Steering When Necessary: Flexible Steering Large Language Models with Backtracking
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
- Can Large Language Models Robustly Perform Natural Language Inference for Japanese Comparatives?
- Implementing a Logical Inference System for Japanese Comparatives
- Op-Fed: Opinion, Stance, and Monetary Policy Annotations on FOMC Transcripts Using Active Learning
- Natural Language Inference over Interaction Space: ICLR 2018 Reproducibility Report
- Automatic Validation of Textual Attribute Values in E-commerce Catalog by Learning with Limited Labeled Data
- Magnitude Matters: a Superior Class of Similarity Metrics for Holistic Semantic Understanding
- 15 Keypoints Is All You Need
- Compartmentalised Agentic Reasoning for Clinical NLI
- Acquisition of Phrase Correspondences using Natural Deduction Proofs
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- Can Neural Networks Understand Logical Entailment?
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Direct Network Transfer: Transfer Learning of Sentence Embeddings for Semantic Similarity
- Learning from others' mistakes: Avoiding dataset biases without modeling them
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Effective writing style imitation via combinatorial paraphrasing
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling
- CLEAR: Contrastive Learning for Sentence Representation
- GIER: Gap-Driven Self-Refinement for Large Language Models
- Self-Guided Contrastive Learning for BERT Sentence Representations
- T-Retrievability: A Topic-Focused Approach to Measure Fair Document Exposure in Information Retrieval
- Dynamic Multi-Level Multi-Task Learning for Sentence Simplification
- Multimodal Fusion Refiner Networks
- Native Logical and Hierarchical Representations with Subspace Embeddings
- Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder
- Sparse Sinkhorn Attention
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Reference and Document Aware Semantic Evaluation Methods for Korean Language Summarization
- Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
- UKP-Athene: Multi-Sentence Textual Entailment for Claim Verification
- Compressed Models are NOT Trust-equivalent to Their Large Counterparts
- Universal Sentence Encoder
- MVP-BERT: Redesigning Vocabularies for Chinese BERT and Multi-Vocab Pretraining
- Dataset Creation for Visual Entailment using Generative AI
- Second-Order Word Embeddings from Nearest Neighbor Topological Features
- Sieving Fake News From Genuine: A Synopsis
- Online Anti-sexist Speech: Identifying Resistance to Gender Bias in Political Discourse
- Exploring Lexical Irregularities in Hypothesis-Only Models of Natural Language Inference
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- Link Prediction for Event Logs in the Process Industry
- Training a Ranking Function for Open-Domain Question Answering
- Semantics Preserving Adversarial Learning
- Bag-of-Vector Embeddings of Dependency Graphs for Semantic Induction
- Do Biased Models Have Biased Thoughts?
- Learning to update Auto-associative Memory in Recurrent Neural Networks for Improving Sequence Memorization
- Syntax-based Attention Model for Natural Language Inference
- Semantic Sentence Matching with Densely-connected Recurrent and Co-attentive Information
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- ACTRCE: Augmenting Experience via Teacher's Advice For Multi-Goal Reinforcement Learning
- Story Ending Generation with Incremental Encoding and Commonsense Knowledge
- Transfer Reward Learning for Policy Gradient-Based Text Generation
- Natural Language Inference by Tree-Based Convolution and Heuristic Matching
- Resource-Efficient Adaptation of Large Language Models for Text Embeddings via Prompt Engineering and Contrastive Fine-tuning
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- HUBERT Untangles BERT to Improve Transfer across NLP Tasks
- Math Natural Language Inference: this should be easy!
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Uncertainty-driven Embedding Convolution
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling
- Towards a Robust Deep Neural Network in Texts: A Survey
- ReCO: A Large Scale Chinese Reading Comprehension Dataset on Opinion
- Term Definitions Help Hypernymy Detection
- The Importance of Being Recurrent for Modeling Hierarchical Structure
- From Disagreement to Understanding: The Case for Ambiguity Detection in NLI
- Answering Science Exam Questions Using Query Rewriting with Background Knowledge
- Filling the Gap: Is Commonsense Knowledge Generation useful for Natural Language Inference?
- Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks
- Lightweight and Efficient Neural Natural Language Processing with Quaternion Networks
- Resource-Size matters: Improving Neural Named Entity Recognition with\n Optimized Large Corpora
- Exploring Neural Models for Parsing Natural Language into First-Order Logic
- Evaluating Multimodal Representations on Visual Semantic Textual Similarity
- Summarizing Utterances from Japanese Assembly Minutes using Political Sentence-BERT-based Method for QA Lab-PoliInfo-2 Task of NTCIR-15
- Probing Biomedical Embeddings from Language Models
- Lexicosyntactic Inference in Neural Models
- The Benchmark Lottery
- Learning Structured Text Representations
- AWE: Asymmetric Word Embedding for Textual Entailment
- Reasoning about Entailment with Neural Attention
- Counterfactual Variable Control for Robust and Interpretable Question Answering
- Attentive Convolution: Equipping CNNs with RNN-style Attention Mechanisms
- A Survey on Text Classification: From Shallow to Deep Learning
Related