SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
2019/05/02 by Alex Wang, Wang, Alex, Yada Pruksachatkun +13 · 245 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.1905.00537
NeurIPS 2019, super.gluebenchmark.com updating acknowledegments
openalex publication_date 2019/05/02 · openalex created_date 2019/05/09 · arxiv created 2020/02/13 · arxiv updated 2020/02/14 · openalex updated_date 2026/07/28
Abstract
In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at super.gluebenchmark.com.
Citations
Cited by
- BenCSSmark: Making the Social Sciences Count in LLM Research
- The Quest for Winning Tickets in Low-Rank Adapters
- ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
- Epistemic Norms for AI Safety and Alignment Research
- ChatGPT and Gemini participated in the Korean College Scholastic Ability Test -- Earth Science I
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Evolutionary System 2 Reasoning: An Empirical Proof
- Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
- SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- An Empirical Survey of Model Merging Algorithms for Social Bias Mitigation
- InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages
- Breaking It Down: Domain-Aware Semantic Segmentation for Retrieval Augmented Generation
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- A Rosetta Stone for AI Benchmarks
- Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
- SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- CarBench: A Comprehensive Benchmark for Neural Surrogates on High-Fidelity 3D Car Aerodynamics
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
- EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
- Multimodal Evaluation of Russian-language Architectures
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Generative Caching for Structurally Similar Prompts and Responses
- Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification
- Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-Tuning
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- EvalCards: A Framework for Standardized Evaluation Reporting
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Reimagining Safety Alignment with An Image
- POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Reliable Evaluations for Natural Language Inference based on a Unified Cross-dataset Benchmark
- Large models of what? Mistaking engineering achievements for human linguistic agency
- Factuality challenges in the era of large language models and opportunities for fact-checking
- Large language models for biomedicine: foundations, opportunities, challenges, and best practices
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
- Large Language Models, Agency, and Why Speech Acts are Beyond Them (For Now) – A Kantian-Cum-Pragmatist Case
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace Expansion
- The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
- Estonian Native Large Language Model Benchmark
- Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Knowledge Distillation of Uncertainty using Deep Latent Factor Model
- DiscoTrack: A Multilingual LLM Benchmark for Discourse Tracking
- All You Need is One: Capsule Prompt Tuning with a Single Vector
- The debate over understanding in AI’s large language models
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
- HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing
- D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Selective Adversarial Attacks on LLM Benchmarks
- A Survey on Evaluation of Large Language Models
- SMEC: Rethinking Matryoshka Representation Learning for Retrieval Embedding Compression
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Accelerating Attention with Basis Decomposition
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- Single layer tiny Co4 outpaces GPT-2 and GPT-BERT
- Benchmarking is Broken -- Don't Let AI be its Own Judge
- SocialNLI: A Dialogue-Centric Social Inference Dataset
- LongTail-Swap: benchmarking language models' abilities on rare words
- Reliable and Scalable Robot Policy Evaluation with Imperfect Simulators
- PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
- Emergent evaluation hubs in a decentralizing large language model ecosystem
- Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
- Symbol ungrounding: what the successes (and failures) of large language models reveal about human cognition
- Generating Domain Models with LLMs Using Instruction Tuning: A Research Preview
- A Pipeline to Assess Merging Methods via Behavior and Internals
- Data Efficient Adaptation in Large Language Models via Continuous Low-Rank Fine-Tuning
- AIMMerging: Adaptive Iterative Model Merging Using Training Trajectories for Language Model Continual Learning
- Can an Individual Manipulate the Collective Decisions of Multi-Agents?
- BEFT: Bias-Efficient Fine-Tuning of Language Models
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- LLM Jailbreak Detection for (Almost) Free!
- GLU Variants Improve Transformer
- MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- CrossPT: Exploring Cross-Task Transferability through Multi-Task Prompt Tuning
- BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- Orthogonal Low-rank Adaptation in Lie Groups for Continual Learning of Large Language Models
- Augmented Fine-Tuned LLMs for Enhanced Recruitment Automation
- Masked Diffusion Language Models with Frequency-Informed Training
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Comparative Evaluation of Traditional and Deep Learning Feature Matching Algorithms using Chandrayaan-2 Lunar Data
- Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
- SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
- ChatGPT-generated texts show authorship traits that identify them as non-human
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
- Behavioral Fingerprinting of Large Language Models
- GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation
- An LLM-enabled semantic-centric framework to consume privacy policies
- FlauBERT: Unsupervised Language Model Pre-training for French
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- AI Testing Should Account for Sophisticated Strategic Behaviour
- LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- Structured Reordering for Modeling Latent Alignments in Sequence Transduction
- Can We Trust AI to Govern AI? Benchmarking LLM Performance on Privacy and AI Governance Exams
- Understanding Syntactic Generalization in Structure-inducing Language Models
- GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
- The NordDRG AI Benchmark for Large Language Models
- TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
- GeRe: Towards Efficient Anti-Forgetting in Continual Learning of LLM via General Samples Replay
- LoRA is All You Need for Safety Alignment of Reasoning LLMs
- FairLangProc: A Python package for fairness in NLP
- The Architecture of Trust: A Framework for AI-Augmented Real Estate Valuation in the Era of Structured Data
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- PaPaformer: Language Model from Pre-trained Parallel Paths
- Lessons from complex systems science for AI governance
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
- Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?
- What Context Features Can Transformer Language Models Use?
- LLM-Crowdsourced: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models
- VN-MTEB: Vietnamese Massive Text Embedding Benchmark
- Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
- Exploring a Unified Sequence-To-Sequence Transformer for Medical Product Safety Monitoring in Social Media
- AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
- GRID: Scalable Task-Agnostic Prompt-Based Continual Learning for Language Models
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
- Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
- Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
- Transformers: State-of-the-Art Natural Language Processing
- Learning Task Mixtures from Task Affinities: A Probabilistic Graphical Model for Supervised Fine-Tuning
- Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
- Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
- SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
- SEARA: An Automated Approach for Obtaining Optimal Retrievers
- 3DGSLSR:LargeScale Relocation for Autonomous Driving Based on 3D Gaussian Splatting
- Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation
- GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
- Continual Gradient Low-Rank Projection Fine-Tuning for LLMs
- Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
- La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
- Should We Still Pretrain Encoders with Masked Language Modeling?
- Pitfalls of Evaluating Language Models with Open Benchmarks
- Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
- Challenge-Based Funding to Spark Origins Breakthroughs
- Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
- skLEP: A Slovak General Language Understanding Benchmark
- Exploring the Impact of Temperature on Large Language Models:Hot or Cold?
- Distinguishing Task-Specific and General-Purpose AI in Regulation
- Finance Language Model Evaluation (FLaME)
- RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
- Understand the Implication: Learning to Think for Pragmatic Understanding
- CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
- MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
- Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
- OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
- SELU: A Software Engineering Language Understanding Benchmark
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
- MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
- Leveraging Self-Attention for Input-Dependent Soft Prompting in LLMs
- Adaptive Task Vectors for Large Language Models
- EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
- MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
- Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks
- Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- Characterizing the Expressivity of Fixed-Precision Transformer Language Models
- Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking
- Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
- Adaptive Budget Allocation for Orthogonal-Subspace Adapter Tuning in LLMs Continual Learning
- Efficient Ensemble for Fine-tuning Language Models on Multiple Datasets
- THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
- OASIS: Online Sample Selection for Continual Visual Instruction Tuning
- Information-Theoretic Complementary Prompts for Improved Continual Text Classification
- Turing Test 2.0: The General Intelligence Threshold
- ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
- KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
- GIM: Improved Interpretability for Large Language Models
- KoBALT: Korean Benchmark For Advanced Linguistic Tasks
- Understanding Differential Transformer Unchains Pretrained Self-Attentions
- TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games
- Procedural Environment Generation for Tool-Use Agents
- Saten: Sparse Augmented Tensor Networks for Post-Training Compression of Large Language Models
- WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
- Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
- R3: Robust Rubric-Agnostic Reward Models
- Automatically Advancing LLM Expertise in Technology Judgment
- Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
- On the Evaluation of Engineering Artificial General Intelligence
- Towards Contamination Resistant Benchmarks
- Injecting Hierarchy with U-Net Transformers
- Towards Zero-Label Language Learning
- BEAMetrics: A Benchmark for Language Generation Evaluation Evaluation
- Measuring Hong Kong Massive Multi-Task Language Understanding
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Token-free Models for Sarcasm Detection
- Semantic Invariance in Agentic AI
- SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- AI Evaluation Should Require Standardized Item-Level Data Releases
- How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Generative AI in Education: Student Skills and Lecturer Roles
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
- VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
- FinVerse: Financial Time-Series Benchmark
- Auditing the Ethical Logic of Generative AI Models
- FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation
- EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks
- Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark
- UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
- ViQA-COVID: COVID-19 Machine Reading Comprehension Dataset for Vietnamese
- DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
- CPG-EVAL: A Multi-Tiered Benchmark for Evaluating the Chinese Pedagogical Grammar Competence of Large Language Models
- Can the capability of Large Language Models be described by human ability? A Meta Study
- Myanmar XNLI: Building a Dataset and Exploring Low-resource Approaches to Natural Language Inference with Myanmar
- Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation
Related