BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
2019/05/24 by Christopher Clark, Clark, Christopher, Kenton Lee +9 · 283 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1905.10044
openalex publication_date 2019/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In this paper we study yes/no questions that are naturally occurring --- meaning that they are generated in unprompted and unconstrained settings. We build a reading comprehension dataset, BoolQ, of such questions, and show that they are unexpectedly challenging. They often query for complex, non-factoid information, and require difficult entailment-like inference to solve. We also explore the effectiveness of a range of transfer learning baselines. We find that transferring from entailment data is more effective than transferring from paraphrase or extractive QA data, and that it, surprisingly, continues to be very beneficial even when starting from massive pre-trained language models such as BERT. Our best method trains BERT on MultiNLI and then re-trains it on our train set. It achieves 80.4% accuracy compared to 90% accuracy of human annotators (and 62% majority-baseline), leaving a significant gap for future work.
Citations
Cited by
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- The Quest for Winning Tickets in Low-Rank Adapters
- AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
- LLMBoost: Make Large Language Models Stronger with Boosting
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- Deep Delta Learning
- Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Towards Effective Model Editing for LLM Personalization
- MIDUS: Memory-Infused Depth Up-Scaling
- Resting Neurons, Active Insights: Improving Input Sparsification for Large Language Models
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Low-Rank Compression of Language Models via Differentiable Rank Selection
- Scaling Behavior of Discrete Diffusion Language Models
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- PCMind-2.1-Kaiyuan-2B Technical Report
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- Greedy Alignment Principle for Optimizer Selection
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
- Automatic Pruning Discovery for Large Language Models
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
- GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Alignment-Aware Quantization for LLM Safety
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
- Testing Question Answering Software with Context-Driven Question Generation
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- MobileLLM-Pro Technical Report
- MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
- DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
- Motif 2 12.7B technical report
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
- ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models
- Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
- A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- Gaperon: A Peppered English-French Generative Language Model Suite
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Mixture-of-Depths Attention
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- Robustness as an Emergent Property of Task Performance
- Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- What Limits Agentic Systems Efficiency?
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
- Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFT
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- ChessQA: Evaluating Large Language Models for Chess Understanding
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Simple Denoising Diffusion Language Models
- SeeDNorm: Self-Rescaled Dynamic Normalization
- MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- Simple Context Compression: Mean-Pooling and Multi-Ratio Training
- Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
- Citation Failure: Definition, Analysis and Efficient Mitigation
- GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
- Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
- Energy-Efficient and Dequantization-Free Q-LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
- Restoring Pruned Large Language Models via Lost Component Compensation
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- Vocab Diet: Reshaping the Vocabulary of LLMs with Vector Arithmetic
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
- DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
- FraQAT: Quantization Aware Training with Fractional bits
- First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
- Towards Reversible Model Merging For Low-rank Weights
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- End-to-End Multi-Modal Diffusion Mamba
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features
- OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Neural Weight Compression for Language Models
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- CTR-LoRA: Curvature-Aware and Trust-Region Guided Low-Rank Adaptation for Large Language Models
- Tapered Language Models
- Decoupled DiLoCo for Resilient Distributed Pre-training
- Efficient Resource-Constrained Training of Vision Transformers via Subspace Optimization
- KORMo: Korean Open Reasoning Model for Everyone
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- Single layer tiny Co4 outpaces GPT-2 and GPT-BERT
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Next Semantic Scale Prediction via Hierarchical Diffusion Language Models
- When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- POME: Post Optimization Model Edit via Muon-style Projection
- Mixture of Neuron Experts
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- AMAQ: Adaptive Mixed-bit Activation Quantization for Collaborative Parameter Efficient Fine-tuning
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- GraphGhost: Tracing Structures Behind Large Language Models
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Energy-Regularized Sequential Model Editing on Hyperspheres
- ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
- Learning a Zeroth-Order Optimizer for Fine-Tuning LLMs
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- Negative Pre-activations Differentiate Syntax
- CURA: Size Isnt All You Need -- A Compact Universal Architecture for On-Device Intelligence
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- Effective Quantization of Muon Optimizer States
- PT2-LLM: Post-Training Ternarization for Large Language Models
- MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
- Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond
- JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts?
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- CLUE: Conflict-guided Localization for LLM Unlearning Framework
- Performance Consistency of Learning Methods for Information Retrieval Tasks
- Thinking Augmented Pre-training
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
- WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- HyperAdapt: Simple High-Rank Adaptation
- TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
- On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
- TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation
- QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- Distribution-Aligned Decoding for Efficient LLM Task Adaptation
- What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets
- Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
- FURINA: Free from Unmergeable Router via LINear Aggregation of mixed experts
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation
- CBP-Tuning: Efficient Local Customization for Black-box Large Language Models
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- GAMA: A General Anonymizing Multi-Agent System for Privacy Preservation Enhanced by Domain Rules and Disproof Mechanism
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
- HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
- Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Interpreting the Effects of Quantization on LLMs
- Causal Attention with Lookahead Keys
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- A Probabilistic Inference Scaling Theory for LLM Self-Correction
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models
- From Injection to Defense: Constructing Edit-Based Fingerprints for Large Language Models
- Binary Quantization For LLMs Through Dynamic Grouping
- Language Models are Few-Shot Learners
- EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
- REFRAG: Rethinking RAG based Decoding
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Metis: Training LLMs with FP4 Quantization
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- FedReFT: Federated Representation Fine-Tuning with All-But-Me Aggregation
- QuesGenie: Intelligent Multimodal Question Generation
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- Evaluating and Characterizing Human Rationales
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- Maximum Score Routing For Mixture-of-Experts
- KLUE: Korean Language Understanding Evaluation
- FedSODA: Federated Fine-tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- LLM Compression: How Far Can We Go in Balancing Size and Performance?
- MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- Tensorized Clustered LoRA Merging for Multi-Task Interference
- Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models
- Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- KLLM: Fast LLM Inference with K-Means Quantization
- From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- Basic Reading Distillation
- DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference
- Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
Related