HellaSwag: Can a Machine Really Finish Your Sentence?
2019/05/19 by Rowan Zellers, Zellers, Rowan, Ari Holtzman +7 · 667 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1905.07830
openalex publication_date 2019/05/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the keys." With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference? In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical 'Goldilocks' zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models. Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.
Citations
Cited by
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Entropy-Guided Token Dropout: Training Autoregressive Language Models with Limited Domain Data
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- Diversity or Precision? A Deep Dive into Next Token Prediction
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
- Learning When Not to Attend Globally
- AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing
- Efficient Multi-Model Orchestration for Self-Hosted Large Language Models
- LLMBoost: Make Large Language Models Stronger with Boosting
- DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
- Attention Residuals
- The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning
- From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference
- SymStep: Symbolic Step Verification for Logical Reasoning
- Bridging Compute- and Data-Optimal Pretraining
- Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
- CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
- TriSP: Tri-Signal Structured Pruning for Large Language Models
- Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
- Compressing LLMs with MoP: Mixture of Pruners
- Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
- TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
- MAGIC: Achieving Superior Model Merging via Magnitude Calibration
- AraMix: Recycling, Refiltering, and Deduplicating to Deliver the Largest Arabic Pretraining Corpus
- MoE Pathfinder: Trajectory-driven Expert Pruning
- Mitigating Forgetting in Low Rank Adaptation
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
- Sigma-MoE-Tiny Technical Report
- Bolmo: Byteifying the Next Generation of Language Models
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Per-Axis Weight Deltas for Frequent Model Updates
- Dual-objective Language Models: Training Efficiency Without Overfitting
- SASQ: Static Activation Scaling for Quantization-Aware Training in Large Language Models
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- What Affects the Effective Depth of Large Language Models?
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
- SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- Benchmarking the Generality of Vision-Language-Action Models
- LLMs Can Assist with Proposal Selection at Large User Facilities
- Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
- MOA: Multi-Objective Alignment for Role-Playing Agents
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- MobileFineTuner: A Mobile-Native Framework for On-Device LLM Fine-Tuning in Real-World Embedded AI Applications
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- PCMind-2.1-Kaiyuan-2B Technical Report
- Flash Multi-Head Feed-Forward Network
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Leveraging KV Similarity for Online Structured Pruning in LLMs
- Greedy Alignment Principle for Optimizer Selection
- A Latent Variable Framework for Scaling Laws in Large Language Models
- SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
- Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
- ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
- Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
- Jina-VLM: Small Multilingual Vision Language Model
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- Nexus: Higher-Order Attention Mechanisms in Transformers
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMs
- From FLOPs to Footprints: The Resource Cost of Artificial Intelligence
- Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in \± 1, ± i\
- Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
- PerfMamba: Performance Analysis and Pruning of Selective State Space Models
- A Rosetta Stone for AI Benchmarks
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LM
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
- Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
- IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
- CafeQ: Calibration-free Quantization via Learned Transformations and Adaptive Rounding
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- ModHiFi: Identifying High Fidelity predictive components for Model Modification
- Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation
- FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
- SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
- MoodBench 1.0: An Evaluation Benchmark for Emotional Companionship Dialogue Systems
- Xmodel-2.5: 1.3B Data-Efficient Reasoning SLM
- Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets
- Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
- R2Q: Towards Robust 2-Bit Large Language Models via Residual Refinement Quantization
- E3-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models
- Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
- AICC: Parse HTML Finer, Make Models Better -- A 7.3T AI-Ready Corpus Built by a Model-Based HTML Parser
- PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
- A Novel Hierarchical Integration Method for Efficient Model Merging in Medical LLMs
- Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
- Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
- Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
- GateRA: Token-Aware Modulation for Parameter-Efficient Fine-Tuning
- FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
- Virtual Width Networks
- When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
- Optimizing Mixture of Block Attention
- ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
- Safety-Preserving PTQ via Contrastive Alignment Loss
- Range Asymmetric Numeral Systems-Based Lightweight Intermediate Feature Compression for Split Computing of Deep Neural Networks
- Sentence-Anchored Gist Compression for Long-Context LLMs
- SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
- Importance-Aware Data Selection for Efficient LLM Instruction Tuning
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Attention and Compression is all you need for Controllably Efficient Language Models
- Rethinking Parameter Sharing as Graph Coloring for Structured Compression
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- MobileLLM-Pro Technical Report
- Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- MuonAll: Muon Variant for Efficient Finetuning of Large Language Models
- DyKAF: Dynamical Kronecker Approximation of the Fisher Information Matrix for Gradient Preconditioning
- Next-Latent Prediction Transformers Learn Compact World Models
- Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
- Motif 2 12.7B technical report
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
- MultiZebraLogic: A Multilingual Logical Reasoning Benchmark
- IG-Pruning: Input-Guided Block Pruning for Large Language Models
- Random Initialization of Gated Sparse Adapters
- Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
- Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
- Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- Language Modeling With Factorization Memory
- Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- 1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
- Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
- From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
- Angular Steering: Behavior Control via Rotation in Activation Space
- RCScore: Quantifying Response Consistency in Large Language Models
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- NeuronMM: High-Performance Matrix Multiplication for LLM Inference on AWS Trainium
- Gaperon: A Peppered English-French Generative Language Model Suite
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
- Mixture-of-Depths Attention
- Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
- Will Scaling Improve Social Simulation with LLMs?
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
- Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization
- SSV: Sparse Speculative Verification for Efficient LLM Inference
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
- Arcee Trinity Large Technical Report
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
- Ministral 3
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- RL makes MLLMs see better than SFT
- Large Emotional World Model
- A Survey on Unlearning in Large Language Models
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- What Limits Agentic Systems Efficiency?
- From Cross-Task Examples to In-Task Prompts: A Graph-Based Pseudo-Labeling Framework for In-context Learning
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
- Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFT
- Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
- Information-Theoretic Discrete Diffusion
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning
- Multi-Agent Evolve: LLM Self-Improve through Co-evolution
- Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
- A Survey on LLM Mid-Training
- Knocking-Heads Attention
- Simple Denoising Diffusion Language Models
- Softmax is 1/2-Lipschitz: A tight bound across all ℓp norms
- SeeDNorm: Self-Rescaled Dynamic Normalization
- PerCoR: Evaluating Commonsense Reasoning in Persian via Multiple-Choice Sentence Completion
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
- The Structural Scalpel: Automated Contiguous Layer Pruning for Large Language Models
- Label Smoothing Improves Gradient Ascent in LLM Unlearning
- Generalization or Memorization: Dynamic Decoding for Mode Steering
- Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
- Estonian Native Large Language Model Benchmark
- Context-level Language Modeling by Learning Predictive Context Embeddings
- Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
- From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
- Data-Centric Lessons To Improve Speech-Language Pretraining
- GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
- Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning
- Latent Space Factorization in LoRA
- What is the Best Sequence Length for BABYLM?
- ELUTQ: Efficient LUT-Aware Quantization for Deploying Large Language Models on Edge Devices
- Restoring Pruned Large Language Models via Lost Component Compensation
- CPSVD: Enhancing Large Language Model Compression via Column-Preserving Singular Value Decomposition
- Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- DiSRouter: Distributed Self-Routing for LLM Selections
- Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
- Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
- ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- Mapping Post-Training Forgetting in Language Models at Scale
- From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
- Unbiased Gradient Low-Rank Projection
- MARS-M: When Variance Reduction Meets Matrices
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- Vocab Diet: Reshaping the Vocabulary of LLMs with Vector Arithmetic
- DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Continual Learning via Sparse Memory Finetuning
- Predicting Task Performance with Context-aware Scaling Laws
- Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
- RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- Closing the Gap Between Text and Speech Understanding in LLMs
- End-to-End Multi-Modal Diffusion Mamba
- RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
- Selective Adversarial Attacks on LLM Benchmarks
- CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
- OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting during Parameter-Efficient Fine-Tuning
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Neural Weight Compression for Language Models
- ShishuLM: Lightweight Language Model with Hybrid Decoder-MLP Architecture and Paired Weight Sharing
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
- Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- CTR-LoRA: Curvature-Aware and Trust-Region Guided Low-Rank Adaptation for Large Language Models
- PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
- Tapered Language Models
- Forget Attention: Importance-Aware Attention Is All You Need
- Accelerating Attention with Basis Decomposition
- Decoupled DiLoCo for Resilient Distributed Pre-training
- Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
- Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
- FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- NarraBench: A Comprehensive Framework for Narrative Benchmarking
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- KORMo: Korean Open Reasoning Model for Everyone
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- Fewer Weights, More Problems: A Practical Attack on LLM Pruning
- DISCO: Diversifying Sample Condensation for Efficient Model Evaluation
- Contrastive Weak-to-strong Generalization
- The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Models
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- SliceFine: The Universal Winning-Slice Hypothesis for Pretrained Networks
- Measuring and Mitigating Identity Bias in Multi-Agent Debate via Anonymization
- Next Semantic Scale Prediction via Hierarchical Diffusion Language Models
- More Data or Better Data? A Critical Analysis of Data Selection and Synthesis for Mathematical Reasoning
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Native Hybrid Attention for Efficient Sequence Modeling
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Mid-Training of Large Language Models: A Survey
- JAI-1: A Thai-Centric Large Language Model
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- POME: Post Optimization Model Edit via Muon-style Projection
- Grouped Differential Attention
- Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
- Training Dynamics Impact Post-Training Quantization Robustness
- lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
- Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
- Mixture of Neuron Experts
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
- Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
- Latent Speech-Text Transformer
- Boomerang Distillation Enables Zero-Shot Model Size Interpolation
- Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- Robustness assessment of large audio language models in multiple-choice evaluation
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- What Makes Diffusion Language Models Super Data Learners?
- Measuring Language Model Hallucinations Through Distributional Correctness
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
- Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
- Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
- Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Sentry: Authenticating Machine Learning Artifacts on the Fly
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Towards Ecologically Valid LLM Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
- Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
- CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- Layer-wise dynamic rank for compressing large language models
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- Fingerprinting LLMs via Prompt Injection
- Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
- Conda: Column-Normalized Adam for Training Large Language Models Faster
- Negative Pre-activations Differentiate Syntax
- Watermarking Diffusion Language Models
- Short window attention enables long-term memorization
- LLM DNA: Tracing Model Evolution via Functional Representations
- CURA: Size Isnt All You Need -- A Compact Universal Architecture for On-Device Intelligence
- Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- Sequential Diffusion Language Models
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Assessing Large Language Models in Updating Their Forecasts with New Information
- Tequila: Trapping-free Ternary Quantization for Large Language Models
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models
- Timber: Training-free Instruct Model Refining with Base via Effective Rank
- Beyond Outliers: A Study of Optimizers Under Quantization
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- Dual-Space Smoothness for Robust and Balanced LLM Unlearning
- SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Effective Quantization of Muon Optimizer States
- Multiplayer Nash Preference Optimization
- PT2-LLM: Post-Training Ternarization for Large Language Models
- MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
- JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- Stochastic activations
- HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
- Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
- Lightweight error mitigation strategies for post-training N:M activation sparsity in LLMs
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- Blockwise Hadamard high-Rank Adaptation for Parameter-Efficient LLM Fine-Tuning
- Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
- Predicting LLM Reasoning Performance with Small Proxy Model
- CLUE: Conflict-guided Localization for LLM Unlearning Framework
- SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
- Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
- Integrated Framework for LLM Evaluation with Answer Generation
- Enhancing Linear Attention with Residual Learning
- RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
- GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM Inference
- Subtract or Replay? Exact Deletion from Language-Model Memory
- Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
- WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
- Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
- Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
- Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- Multi-Head Attention Residuals
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation
- A large-scale evaluation of commonsense knowledge in humans and large language models
- Riemannian Optimization for LoRA on the Stiefel Manifold
- Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
- Soft Tokens, Hard Truths
- Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
- Diversity Boosts AI-Generated Text Detection
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
- HyperAdapt: Simple High-Rank Adaptation
- Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
- TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
- Weights-Rotated Preference Optimization for Large Language Models
- SEQR: Secure and Efficient QR-based LoRA Routing
- QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models
- nDNA -- the Semantic Helix of Artificial Cognition
- PTQTP: Post-Training Quantization to Trit-Planes for Large Language Models
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- Rethinking the Role of Text Complexity in Language Model Pretraining
- EG-MLA: Embedding-Gated Multi-head Latent Attention for Scalable and Efficient LLMs
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- Distribution-Aligned Decoding for Efficient LLM Task Adaptation
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
- KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
- Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Fair-GPTQ: Bias-Aware Quantization for Large Language Models
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- ZERA: Zero-init Instruction Evolving Refinement Agent -- From Zero Instructions to Structured Prompts via Principle-based Optimization
- DSFT: Inspiring Diffusion Large Language Models to Comprehend Mathematical and Logical Patterns
- SBVR: Summation of BitVector Representation for Efficient LLM Quantization
- DropLoRA: Sparse Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
- Instance-level Randomization: Toward More Stable LLM Evaluations
- CBP-Tuning: Efficient Local Customization for Black-box Large Language Models
- AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- From Evaluation to Enhancement: Large Language Models for Zero-Knowledge Proof Code Generation
- Fluid Language Model Benchmarking
- Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
- AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
- From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
- Towards Understanding Visual Grounding in Visual Language Models
- ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- ForTIFAI: Fending Off Recursive Training Induced Failure for AI Model Collapse
- Benchmarking Energy Efficiency of Large Language Models Using vLLM
- Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
- Interpreting the Effects of Quantization on LLMs
- Causal Attention with Lookahead Keys
- Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation
- COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
- LoaQ: Layer-wise Output Approximation Quantization
- IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs
- Hyperbolic Large Language Models
- Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
- AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
- SpikingBrain: Spiking Brain-inspired Large Models
- Delta Activations: A Representation for Finetuned Large Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Set Block Decoding is a Language Model Inference Accelerator
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
- Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
- Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
- TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- Binary Quantization For LLMs Through Dynamic Grouping
- RoboBuddy in the Classroom: Exploring LLM-Powered Social Robots for Storytelling in Learning and Integration Activities
- Efficient Training-Free Online Routing for High-Volume Multi-LLM Serving
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Inducing Faithfulness in Structured Reasoning via Counterfactual Sensitivity
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
- CEQuest: Benchmarking Large Language Models for Construction Estimation
- Dream-Coder 7B: An Open Diffusion Language Model for Code
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Metis: Training LLMs with FP4 Quantization
- Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
- PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
- Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Provable Benefits of In-Tool Learning for Large Language Models
- FedReFT: Federated Representation Fine-Tuning with All-But-Me Aggregation
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
- Diffusion Language Models Know the Answer Before Decoding
- Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis
- CALR: Corrective Adaptive Low-Rank Decomposition for Efficient Large Language Model Layer Compression
- Predicting the Order of Upcoming Tokens Improves Language Modeling
- SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts--Extended Version
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- Scaling Laws for Task-Stratified Knowledge in Post-Training Quantized Large Language Models
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
- Integral Transformer: Denoising Attention, Not Too Much Not Too Little
- DualSparse-MoE: Coordinating Tensor/Neuron-Level Sparsity with Expert Partition and Reconstruction
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
- Dream 7B: Diffusion Large Language Models
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- Z-Pruner: Post-Training Pruning of Large Language Models for Efficiency without Retraining
- Maximum Score Routing For Mixture-of-Experts
- Reinforcement Learning with Rubric Anchors
- Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
- Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
- MSRS: Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models
- EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization
- TiMoE: Time-Aware Mixture of Language Experts
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
- TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
- Pruning Large Language Models by Identifying and Preserving Functional Networks
- Align, Don't Divide: Revisiting the LoRA Architecture in Multi-Task Learning
- iFairy: the First 2-bit Complex LLM with All Parameters in \±1, ± i\
- LLM Data Selection and Utilization via Dynamic Bi-level Optimization
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
- Tensorized Clustered LoRA Merging for Multi-Task Interference
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
- RegMean++: Enhancing Effectiveness and Generalization of Regression Mean for Model Merging
- FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
- Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- Trainable Dynamic Mask Sparse Attention
- Kron-LoRA: Hybrid Kronecker-LoRA Adapters for Scalable, Sustainable Fine-tuning
- FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Large-Scale Diverse Synthesis for Mid-Training
- LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
- Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?
- KLLM: Fast LLM Inference with K-Means Quantization
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
- Strategic Deflection: Defending LLMs from Logit Manipulation
- Intent Aware Context Retrieval for Multi-Turn Agricultural Question Answering
- Kimi K2: Open Agentic Intelligence
- From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- StackTrans: From Large Language Model to Large Pushdown Automata Model
- Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
- Diffusion Beats Autoregressive in Data-Constrained Settings
- DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference
- SESR-Eval: Dataset for Evaluating LLMs in the Title-Abstract Screening of Systematic Reviews
- Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
- Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method
- Technical Report of TeleChat2, TeleChat2.5 and T1
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
- Mellum2 Technical Report
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Adaptive Block-Scaled Data Types
- Retrieval-Aware Distillation for Transformer-SSM Hybrids
- GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration
- MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
- A Comprehensive Evaluation on Quantization Techniques for Large Language Models
- WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- Characterizing State Space Model (SSM) and SSM-Transformer Hybrid Language Model Performance with Long Context Length
- Language Models Improve When Pretraining Data Matches Target Tasks
- PARAM-1 BharatGen 2.9B Model
- LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters
- Composing Linear Layers from Irreducibles
- First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
- AdaMuon: Adaptive Muon Optimizer
- Seq vs Seq: An Open Suite of Paired Encoders and Decoders
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- GeLaCo: An Evolutionary Approach to Layer Compression
- FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- Advancing Large Language Models for Tibetan with Curated Data and Continual Pre-Training
- SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
- ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making
- Lizard: An Efficient Linearization Framework for Large Language Models
- Self-Improving Model Steering
- From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
- The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
- An Offline Mobile Conversational Agent for Mental Health Support: Learning from Emotional Dialogues and Psychological Texts with Student-Centered Evaluation
- Pre-Training LLMs on a budget: A comparison of three optimizers
- Invariant-based Robust Weights Watermark for Large Language Models
- SAS: Simulated Attention Score
- COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
- Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- FlexOlmo: Open Language Models for Flexible Data Use
- On the Effect of Uncertainty on Layer-wise Inference Dynamics
- A Systematic Analysis of Hybrid Linear Attention
- DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
- Towards a Principled Evaluation of Knowledge Editors
- Steering Information Utility in Key-Value Memory for Language Model Post-Training
- Train-before-Test Harmonizes Language Model Rankings
- Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
- Pre-Trained Policy Discriminators are General Reward Models
- LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimization
- GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
- DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- LayerCake: Token-Aware Contrastive Decoding within Large Language Model Layers
- Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents
- OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
- MGAA: Multi-Granular Adaptive Allocation fof Low-Rank Compression of LLMs
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
- Answer Matching Outperforms Multiple Choice for Language Model Evaluation
- From 2:4 to 8:16 sparsity patterns in LLMs for Outliers and Weights with Variance Correction
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
- Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
- LEDOM: Reverse Language Model
- Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
- La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
- Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check
- Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
- Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- Teaching a Language Model to Speak the Language of Tools
- Masked Gated Linear Unit
- Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
- Residual Matrix Transformers: Scaling the Size of the Residual Stream
- Towards Distributed Neural Architectures
- AutoMixer: Checkpoint Artifacts as Automatic Data Mixers
- Grokking in LLM Pretraining? Monitor Memorization-to-Generalization without Test
- Data Efficacy for Language Model Training
- TopK Language Models
- OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- Norm×Direction: Restoring the Missing Query Norm in Vision Linear Attention
- GPTailor: Large Language Model Pruning Through Layer Cutting and Stitching
Related