Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Sparsely-Gated MoE layers increase model capacity by over 1000x with minimal compute loss, improving language modeling and translation.
2017/01/23 by Noam Shazeer, Shazeer, Noam, Azalia Mirhoseini +13 · 11 voices · 732 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling #cs.CL #cs.LG #cs.NE #stat.ML
paper · pdf · doi:10.48550/arxiv.1701.06538
openalex publication_date 2017/01/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges. In this work, we address these challenges and finally realize the promise of conditional computation, achieving greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters. We introduce a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks. A trainable gating network determines a sparse combination of these experts to use for each example. We apply the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora. We present model architectures in which a MoE with up to 137 billion parameters is applied convolutionally between stacked LSTM layers. On large language modeling and machine translation benchmarks, these models achieve significantly better results than state-of-the-art at lower computational cost.
Summary
The paper presents a Sparsely-Gated Mixture-of-Experts (MoE) layer that enables conditional computation, where only a small subset of the network is active for any given input. By addressing GPU hardware constraints and load balancing through specific gating mechanisms and loss functions, the authors scale models up to 137 billion parameters. They demonstrate significant improvements in perplexity and BLEU scores across large-scale language modeling and machine translation tasks.
machine-generated · gemma4:31b
In simple words
Imagine a giant book of answers where each page is written by a different expert. Instead of reading the whole book for every question, a smart guide picks just two or three pages that best fit the question. This lets the book be huge without taking forever to read. The authors made this work on fast computers by making sure the experts share the work evenly.
machine-generated · gemma4:31b
Outline
- Introduction and Related Work — Discusses conditional computation and the challenges of scaling model capacity on GPUs.
- The Structure of the Mixture-of-Experts layer — Details the MoE architecture, including the gating network and Noisy Top-K Gating.
- Addressing Performance Challenges — Explains how to solve shrinking batch sizes and network bandwidth bottlenecks using model parallelism.
- Balancing Expert Utilization — Introduces importance and load loss functions to prevent a few experts from dominating.
- Experiments — Evaluates the MoE layer on 1B word, 100B word language modeling, and machine translation benchmarks.
- Conclusion — Summarizes findings and suggests conditional computation for other domains.
machine-generated · gemma4:31b
Argument
- Increasing model capacity generally improves accuracy but leads to quadratic growth in compute costs.
Citations of previous work in text, images, and audio domains. - Conditional computation can increase capacity without proportional compute increases by activating only parts of the network per example.
Theoretical proposals cited in related work. - Practical implementation on GPUs is hindered by branching overhead, shrinking batch sizes, and network bandwidth.
Analysis of GPU architecture and distributed computing bottlenecks. - A Sparsely-Gated MoE layer with Noisy Top-K gating can effectively route inputs to a small number of experts.
Proposed architectural design and mathematical formulation of the gating function. - Combining data parallelism for standard layers and model parallelism for experts maintains large batch sizes per expert.
Proposed system architecture for distributed training. - Adding importance and load balancing losses prevents the 'rich-get-richer' effect where a few experts are over-trained.
Experimental results in Appendix showing CV of Importance/Load with different loss weights. - Massive capacity increases lead to better performance on very large datasets.
Empirical results on 1B and 100B word corpora and WMT'14 translation benchmarks.
machine-generated · gemma4:31b
Assumptions
- The ratio of computation to network bandwidth in GPUs is high enough that increasing hidden layer size improves efficiency. [stated]
- Back-propagation through the discontinuous Top-K gating function does not negatively impact training stability. [stated]
- Language modeling and translation tasks provide sufficient signal to train billions of parameters. [unstated]
machine-generated · gemma4:31b
Claims
- MoE achieves >1000x improvement in model capacity with minor losses in computational efficiency. [experiment]
- Model capacity is more critical for larger datasets than smaller ones. [experiment]
- The MoE layer allows the model to specialize experts based on syntax and semantics. [experiment]
- MoE-augmented models outperform state-of-the-art results in language modeling and translation at lower computational cost. [experiment]
machine-generated · gemma4:31b
Methods
- n: not stated
- population: n/a
- study type: bench
machine-generated · gemma4:31b
Limitations
admitted by authors:
- The largest model (131,072 experts) showed degraded performance compared to the 65,536-expert model, possibly due to excessive sparsity.
- Computational efficiency for the largest MoE model on the 100B word corpus was very low because training batch size was not increased proportionally to the number of GPUs.
noticed by the model, not admitted:
- The authors mention that 'theoretically scary discontinuities' in the output of the gating function exist but have not been observed as a problem in practice.
machine-generated · gemma4:31b
Benchmarks
1 Billion Word Language Modeling Benchmark
| method | metric | value |
| Best Published Results | Test Perplexity (10 epochs) | 34.7 |
| Low-Budget MoE Model | Test Perplexity (10 epochs) | 34.1 |
| Medium-Budget MoE Model | Test Perplexity (10 epochs) | 31.3 |
| High-Budget MoE Model | Test Perplexity (10 epochs) | 28.0 |
WMT'14 En→Fr newstest2014
| method | metric | value |
| MoE with 2048 Experts | BLEU | 40.35 |
| MoE with 2048 Experts (longer training) | BLEU | 40.56 |
| GNMT | BLEU | 39.22 |
| GNMT+RL | BLEU | 39.92 |
WMT'14 En→De newstest2014
| method | metric | value |
| MoE with 2048 Experts | BLEU | 26.03 |
| GNMT | BLEU | 24.91 |
| GNMT+RL | BLEU | 24.66 |
Google Production En→Fr dataset
| method | metric | value |
| MoE with 2048 Experts | Test BLEU | 36.57 |
| GNMT | Test BLEU | 35.56 |
Multilingual Machine Translation
| method | metric | value |
| GNMT-Multi | Perplexity (dev) | 4.14 |
| MoE-Multi | Perplexity (dev) | 3.35 |
machine-generated · gemma4:31b
Key equations
y = ∑i=1nG(x)iEi(x) — The output of the MoE layer is a weighted sum of the outputs from selected expert networks, where weights are determined by the gating network.
machine-generated · gemma4:31b
Proof sketch
- The authors propose a Sparsely-Gated Mixture-of-Experts (MoE) layer to increase model capacity without proportional increases in computation.
- They introduce a 'Noisy Top-K' gating mechanism that selects only a few experts per input, reducing computational cost.
- To prevent expert collapse (where a few experts dominate), they implement auxiliary loss functions based on the coefficient of variation to balance importance and load across experts.
- Performance challenges are addressed by combining data and model parallelism to maintain large batch sizes for each expert.
- The approach is validated through language modeling and machine translation tasks, demonstrating that increasing capacity via MoE significantly improves results on large datasets.
machine-generated · gemma4:31b
Open questions
- Can this approach be scaled to a trillion-parameter model on a trillion-word corpus?
The authors hypothesize it is possible by adding more hardware but have not yet tested it.
machine-generated · gemma4:31b
Cited by
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators
- Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling
- Towards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert Approach
- Stabilizing Native Low-Rank LLM Pretraining
- LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
- MixCompress: Mixture of Experts for Variable Rate Learned Image Compression
- CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
- Grounding latent algorithm routing in transformer reasoning
- MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
- Defer to Plan: Adaptive Multi-Agent Fusion for End-to-End V2X Driving
- GPU-to-Grid: Voltage Regulation via GPU Utilization Control
- PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving
- IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion
- Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
- Scaling Interpretable Transformers with Parity Bottleneck Layers
- Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
- AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic
- Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation
- Searching for Plans You Can Actually Build: A Realizability-Aware Full-Space Optimizer for MoE Training and Serving
- Dual Attention Residuals
- PlotTwist: A Creative Plot Generation Framework with Small Language Models
- Uncovering Latent Reasoning Strategies in Language Models
- Differentiable latent structure discovery for interpretable forecasting in clinical time series
- Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection
- PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks
- Loop the Loopies!
- Learning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
- BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- AirMoE: Statistic-Augmented Over-the-Air MoE for Collaborative Intelligence
- ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
- Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
- Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- HyBDM: Multi-Scale Hybrid Experts for Time Series Forecasting with Bidirectional Dependency Modeling
- A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing
- Multi-level context Modeling for consistent expert selection in Mixture-of-Experts
- Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
- JetCoRD: Reliability-Aware Cross-Experiment Distillation of Jet Taggers with Adaptive Corrective Representation
- Mixtures of SubExperts for Large Language Continual Learning
- xHC: Expanded Hyper-Connections
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
- Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
- FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
- The Energy Footprint of LLM-Based Environmental Analysis: LLMs and Domain Products
- Generalist AI control: Towards multi-purpose adaptive algorithms
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA
- Sparser, Faster, Lighter Transformer Language Models
- Position: Modular Memory is the Key to Continual Learning Agents
- Memory Caching: RNNs with Growing Memory
- A Human-Centric Framework for Data Attribution in Large Language Models
- PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
- mHC: Manifold-Constrained Hyper-Connections
- The Mean-Field Dynamics of Transformers
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Pre-training under infinite compute
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Fast and Simplex: 2-Simplicial Attention in Triton
- Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
- Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
- Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
- Mind the GAP! The Challenges of Scale in Pixel-based Deep Reinforcement Learning
- Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis
- Mixture of Experts Made Intrinsically Interpretable
- Deep Learning is Not So Mysterious or Different
- Vision as LoRA
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- Simplex-FEM Networks (SiFEN): Learning A Triangulated Function Approximator
- Dynamic Subspace Composition: Efficient Adaptation via Contractive Basis Expansion
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- Trust Region Masking for Long-Horizon LLM Reinforcement Learning
- FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment
- Towards Understanding Steering Strength
- OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
- Learning When Not to Attend Globally
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
- Scale Weight Decay and Train Better
- Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems
- WISERouter: LLM Routing with Workload Budget Constraint
- Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
- The Semantic Least-Energy Principle: A Hypothesis for Intelligence
- Multi-level Code Optimization via Mixture of Prompts
- TreeAdapter: Hierarchical Taxonomy-Guided Adapter Composition for Fine-Grained Species Image Generation
- Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety
- MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving
- DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
- Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
- Balanced Soft mixture-of-expert model for Glaucoma Detection
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
- Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement
- Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
- Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability
- SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
- cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
- FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
- Accelerate Speculative Decoding with Sparse Computation in Verification
- InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
- Flexible Multitask Learning with Factorized Diffusion Policy
- MMCTOP: A Multimodal Textualization and Mixture-of-Experts Framework for Clinical Trial Outcome Prediction
- AnchorGK: Anchor-based Incremental and Stratified Graph Learning Framework for Inductive Spatio-Temporal Kriging
- MoE-TransMov: A Transformer-based Model for Next POI Prediction in Familiar & Unfamiliar Movements
- Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
- Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
- DeepCQ: General-Purpose Deep-Surrogate Framework for Lossy Compression Quality Prediction
- SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Mixture-of-Experts with Gradient Conflict-Driven Subspace Topology Pruning for Emergent Modularity
- How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
- Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation
- Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts
- Secret mixtures of experts inside your LLM
- UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- DNAMotifTokenizer: Towards Biologically Informed Tokenization of Genomic Sequences
- Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
- PoseMoE: Mixture-of-Experts Network for Monocular 3D Human Pose Estimation
- Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
- Sigma-MoE-Tiny Technical Report
- INTELLECT-3: Technical Report
- Chorus: Harmonizing Context and Sensing Signals for Data-Free Model Customization in IoT
- Large Model Enabled Embodied Intelligence for 6G Integrated Perception, Communication, and Computation Network
- Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- Expert Switching for Robust AAV Landing: A Dual-Detector Framework in Simulation
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Ensemble-Guided Distillation for Compact and Robust Acoustic Scene Classification on Edge Devices
- Improving Recursive Transformers with Mixture of LoRAs
- Element-wise Modulation of Random Matrices for Efficient Neural Layers
- Non-Resolution Reasoning (NRR): A Computational Framework for Contextual Identity and Ambiguity Preservation
- MIDUS: Memory-Infused Depth Up-Scaling
- SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- Fault-Tolerant Sandboxing for AI Coding Agents: A Transactional Approach to Safe Autonomous Execution
- PerNodeDrop: A Method Balancing Specialized Subnets and Regularization in Deep Neural Networks
- RAST-MoE-RL: A Regime-Aware Spatio-Temporal MoE Framework for Deep Reinforcement Learning in Ride-Hailing
- MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
- Fine-Grained Zero-Shot Learning with Attribute-Centric Representations
- Bridging Streaming Continual Learning via In-Context Large Tabular Models
- Metacognitive Sensitivity for Test-Time Dynamic Model Selection
- M3Net: A Multi-Metric Mixture of Experts Network Digital Twin with Graph Neural Networks
- Mixture of Lookup Key-Value Experts
- Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
- Enhancing Medical Cross-Modal Hashing Retrieval using Dropout-Voting Mixture-of-Experts Fusion
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- Flash Multi-Head Feed-Forward Network
- Radiance-Field Reinforced Pretraining: Scaling Localization Models with Unlabeled Wireless Signals
- TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning
- PlantBiMoE: A Bidirectional Foundation Model with SparseMoE for Plant Genomes
- Real-Time Dynamics in Two Dimensions with Tensor Network States via Time-Dependent Variational Monte Carlo
- Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
- Learning When to Switch: Adaptive Policy Selection via Reinforcement Learning
- Comparing the latent features of universal machine-learning interatomic potentials
- Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
- ProPhy: Progressive Physical Alignment for Dynamic World Simulation
- RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- LAWS: Learning from Actual Workloads Symbolically -- A Self-Certifying Parametrized Cache Architecture for Neural Inference, Robotics, and Edge Deployment
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
- Stable Signer: Hierarchical Sign Language Generative Model
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
- Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
- SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- Towards Unified Video Quality Assessment
- GRASP: Guided Residual Adapters with Sample-wise Partitioning
- Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
- A Systematic Characterization of LLM Inference on GPUs
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- Efficient Hyperparameter Search for Non-Stationary Model Training
- Mode-Conditioning Unlocks Superior Test-Time Scaling
- Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis
- MoB: Mixture of Bidders
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- Memory-Integrated Reconfigurable Adapters: A Unified Framework for Settings with Multiple Tasks
- HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
- RecruitView: A Multimodal Dataset for Predicting Personality and Interview Performance for Human Resources Applications
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models
- EnECG: Efficient Ensemble Learning for Electrocardiogram Multi-task Foundation Model
- Advancing time series completion via RFAMoE and MDFF
- EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
- MemFine: Memory-Aware Fine-Grained Scheduling for MoE Training
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- MortgageLLM: Domain-Adaptive Pretraining with Residual Instruction Transfer, Alignment Tuning, and Task-Specific Routing
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- HiFi-MambaV2: Hierarchical Shared-Routed MoE for High-Fidelity MRI Reconstruction
- Progressive Localisation in Localist LLMs
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
- Internalizing Tools as Morphisms in Graded Transformers
- DepthFocus: Controllable Depth Estimation for See-Through Scenes
- RadioKMoE: Knowledge-Guided Radiomap Estimation with Kolmogorov-Arnold Networks and Mixture-of-Experts
- Fine-grained MoE Load Balancing with Linear Programming
- Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers
- Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
- Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Parameter Aware Mamba Model for Multi-task Dense Prediction
- CD-DPE: Dual-Prompt Expert Network Based on Convolutional Dictionary Feature Decoupling for Multi-Contrast MRI Super-Resolution
- Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
- Weight-sparse transformers have interpretable circuits
- InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE
- YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection
- Self-Adaptive Graph Mixture of Models
- Performance and interpretability analysis of code generation large language models
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
- SEMC: Structure-Enhanced Mixture-of-Experts Contrastive Learning for Ultrasound Standard Plane Recognition
- SAC-MoE: Reinforcement Learning with Mixture-of-Experts for Control of Hybrid Dynamical Systems with Uncertainty
- ViTE: Virtual Graph Trajectory Expert Router for Pedestrian Trajectory Prediction
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
- Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
- On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
- Virtual Width Networks
- MAFM3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
- ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
- ExpertAD: Enhancing Autonomous Driving Systems with Mixture of Experts
- RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
- 2.5D Transformer: An Efficient 3D Seismic Interpolation Method without Full 3D Training
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
- ConSurv: Multimodal Continual Learning for Survival Analysis
- Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
- Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Two Heads are Better than One: Distilling Large Language Model Features Into Small Models with Feature Decomposition and Mixture
- Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation
- Attention and Compression is all you need for Controllably Efficient Language Models
- Route Experts by Sequence, not by Token
- Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
- Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
- BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
- In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
- MoEGCL: Mixture of Ego-Graphs Contrastive Representation Learning for Multi-View Clustering
- Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Synapse: Adaptive Arbitration of Complementary Expertise in Time Series Foundational Models
- SA-EMO: Structure-Aligned Encoder Mixture of Operators for Generalizable Full-waveform Inversion
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery
- Enabling Dynamic Sparsity in Quantized LLM Inference
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- GNN-MoE: Context-Aware Patch Routing using GNNs for Parameter-Efficient Domain Generalization
- RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
- Diffusion Language Models are Super Data Learners
- IndicSuperTokenizer: An Optimized Tokenizer for Indic Multilingual LLMs
- Advancing Subsurface Discovery and Geothermal Monitoring with an Agentic Artificial Intelligence Framework
- Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Towards Efficient Federated Learning of Networked Mixture-of-Experts for Mobile Edge Computing
- RDMA Point-to-Point Communication for LLM Systems
- VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
- Dynamic Model Selection for Trajectory Prediction via Pairwise Ranking and Meta-Features
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- Language Modeling With Factorization Memory
- SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Mixture-of-Transformers Learn Faster: A Theoretical Study on Classification Problems
- Running VLAs at Real-time Speed
- ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
- Dual Mixture-of-Experts Framework for Discrete-Time Survival Analysis
- StrataCL: Fabric-Native Communication Library for Production Supernodes
- FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA
- Dynamic Parameterization Is Not Dynamic Inference
- Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
- When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
- RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
- Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA
- Hyper-FEOD: Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE
- Arcee Trinity Large Technical Report
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
- NeurIPT: Foundation Model for Neural Interfaces
- Expert Selections In MoE Models Reveal (Almost) As Much As Text
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
- MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- MaGNet: A Mamba Dual-Hypergraph Network for Stock Prediction via Temporal-Causal and Global Relational Learning
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Human Machine Social Hybrid Intelligence:A Collaborative Decision Making Framework for Large Model Agent Groups and Human Experts
- Spatio-temporal Multivariate Time Series Forecast with Chosen Variables
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Fast and Flexible Image Blind Denoising via Competition of Experts
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- MoEMeta: Mixture-of-Experts Meta Learning for Few-Shot Relational Learning
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- Sparsity and Superposition in Mixture of Experts
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- PINN Balls: Scaling Second-Order Methods for PINNs with Domain Decomposition and Adaptive Sampling
- Adaptive Graph Mixture of Residual Experts: Unsupervised Learning on Diverse Graphs with Heterogeneous Specialization
- FreeChunker: A Cross-Granularity Chunking Framework
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
- A Design Science Blueprint for an Orchestrated AI Assistant in Doctoral Supervision
- MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Pay Better Attention to Attention: Head Selection in Multilingual and Multi-Domain Sequence Modeling
- Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
- Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models
- Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- MoE-Based Learned Inertial Odometry for Bicycle Localization
- Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing
- Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
- L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
- Backdoor or Manipulation? Graph Mixture of Experts Can Defend Against Various Graph Adversarial Attacks
- Mixed-Precision Quantization for Language Models: Techniques and Prospects
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
- Continual Learning via Sparse Memory Finetuning
- Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Neural Approximate Inverse Preconditioners
- Dr.LLM: Dynamic Layer Routing in LLMs
- Fast Visuomotor Policy for Robotic Manipulation
- On Inherited Popularity Bias in Cold-Start Item Recommendation
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
- HoMer: Addressing Heterogeneities by Modeling Sequential and Set-wise Contexts for CTR Prediction
- Variational Mixture of Graph Neural Experts for Alzheimer's Disease Recognition across Frequency Bands in EEG Brain Networks
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- Hierarchical LoRA MoE for Efficient CTR Model Scaling
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Adaptive Heterogeneous Mixtures of Normalising Flows for Robust Variational Inference
- Compositional meta-learning through probabilistic task inference
- Sparsely gated tiny linear experts
- Rethinking the shape convention of an MLP
- Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
- Decoupled DiLoCo for Resilient Distributed Pre-training
- TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
- NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations
- Utilizing dynamic sparsity on pretrained DETR
- Maple: A Multi-agent System for Portable Deep Learning across Clusters
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
- AB-PINNs: Adaptive-Basis Physics-Informed Neural Networks for Residual-Driven Domain Decomposition
- Vision Language Models: A Survey of 26K Papers
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- GRADE: Personalized Multi-Task Fusion via Group-relative Reinforcement Learning with Adaptive Dirichlet Exploration
- Mutual Learning for Hashing: Unlocking Strong Hash Functions from Weak Supervision
- Neurocoder: Learning General-Purpose Computation Using Stored Neural Programs
- Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
- MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts
- Intelligent AI Delegation
- Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- Training Dynamics Impact Post-Training Quantization Robustness
- Mixture of Neuron Experts
- MASA: Rethinking the Representational Bottleneck in LoRA with Multi-A Shared Adaptation
- Staircase Streaming for Low-Latency Multi-Agent Inference
- Improving Multimodal Brain Encoding Model with Dynamic Subject-awareness Routing
- MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Multilingual Routing in Mixture-of-Experts
- DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
- HoRA: Cross-Head Low-Rank Adaptation with Joint Hypernetworks
- Increasing LLM response trustworthiness using voting ensembles
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Mixture of Many Zero-Compute Experts: A High-Rate Quantization Theory Perspective
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
- GLAI: GreenLightningAI for Accelerated Training through Knowledge Decoupling
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
- Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
- Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
- Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- A Multimodal LLM Approach for Visual Question Answering on Multiparametric 3D Brain MRI
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
- Guiding Mixture-of-Experts with Temporal Multimodal Interactions
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
- A Greedy PDE Router for Blending Neural Operators and Classical Methods
- LEAF: A Robust Expert-Based Framework for Few-Shot Continual Event Detection
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
- Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
- MAESTRO : Adaptive Sparse Attention and Robust Learning for Multimodal Dynamic Time Series
- Muon: Training and Trade-offs with Latent Attention and MoE
- From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
- Pretraining with hierarchical memories: separating long-tail and common knowledge
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Agile perceptive multiskill locomotion for quadrupedal robots in the wild
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
- Partial Parameter Updates for Efficient Distributed Training
- Unlocking the Power of Mixture-of-Experts for Task-Aware Time Series Analytics
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
- ChaosNexus: A Foundation Model for Universal Chaotic System Forecasting with Multi-scale Representations
- Score-based Idempotent Distillation of Diffusion Models
- Distributed Specialization: Rare-Token Neurons in Large Language Models
- Mixture of Thoughts: Learning to Aggregate What Experts Think, Not Just What They Say
- SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Multimodal Language Models with Modality-Specific Experts for Financial Forecasting from Interleaved Sequences of Text and Time Series
- CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
- On component interactions in two-stage recommender systems
- Trainability and Mode Separation of Mixed IQP-QCBMs
- DeepResearch Agent System
- Event-Structured Physics-Informed Neural Networks for Differentiable Critical Clearing Boundaries
- SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering
- OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
- TIER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification
- Latent-Kernel Discrete Flow Maps for Few-Step Generation
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Multi-Head Attention Residuals
- From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- ISALux: Illumination and Segmentation Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- CoBEVMoE: Heterogeneity-aware Feature Fusion with Dynamic Mixture-of-Experts for Collaborative Perception
- ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
- STCast: Adaptive Boundary Alignment for Global and Regional Weather Forecasting
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- MoPE: A Mixture of Password Experts for Improving Password Guessing
- TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- Robust LLM Training Infrastructure at ByteDance
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
- TrueMoE: Dual-Routing Mixture of Discriminative Experts for Synthetic Image Detection
- GateTS: Versatile and Efficient Forecasting via Attention-Inspired routed Mixture-of-Experts
- What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
- A visual introduction to Gaussian Belief Propagation
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting
- PCCL: Photonic circuit-switched collective communication for distributed ML
- Synthetic bootstrapped pretraining
- Distributed Deep Learning with RIS Grouping for Accurate Cascaded Channel Estimation
- Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
- ST-LINK: Spatially-Aware Large Language Models for Spatio-Temporal Forecasting
- CSMoE: An Efficient Remote Sensing Foundation Model with Soft Mixture-of-Experts
- Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
- FlowDrive: Energy Flow Field for End-to-End Autonomous Driving
- GLAD: Global-Local Aware Dynamic Mixture-of-Experts for Multi-Talker ASR
- Toward PDDL Planning Copilot
- Similarity-Distance-Magnitude Activations
- Batch-Shaping for Learning Conditional Channel Gated Networks
- When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
- Igniting VLMs toward the Embodied Space
- Dynamic Adaptive Parsing of Temporal and Cross-Variable Patterns for Network State Classification
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- Lightweight Metadata-Aware Mixture-of-Experts Masked Autoencoder for Earth Observation
- Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
- Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
- Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
- MoSE: Unveiling Structural Patterns in Graphs via Mixture of Subgraph Experts
- MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
- Steering MoE LLMs via Expert (De)Activation
- Too Helpful, Too Harmless, Too Honest or Just Right?
- Joint Learning using Mixture-of-Expert-Based Representation for Enhanced Speech Generation and Robust Emotion Recognition
- Two Facets of the Same Optimization Coin: Model Degradation and Representation Collapse in Graph Foundation Models
- One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning
- SEEC: Segmentation-Assisted Multi-Entropy Models for Learned Lossless Image Compression
- Accelerating Frontier MoE Training with 3D Integrated Optics
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- AdaMixT: Adaptive Weighted Mixture of Multi-Scale Expert Transformers for Time Series Forecasting
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- Wisdom of Committees: An Overlooked Approach To Faster and More Accurate Models
- Deep Metric Learning with Locality Sensitive Angular Loss for Self-Correcting Source Separation of Neural Spiking Signals
- Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- Flexible Multimodal Neuroimaging Fusion for Alzheimer's Disease Progression Prediction
- Lookup multivariate Kolmogorov-Arnold Networks
- CAME-AB: Cross-Modality Attention with Mixture-of-Experts for Antibody Binding Site Prediction
- Select, then Balance: Exploring Exogenous Variable Modeling of Spatio-Temporal Forecasting
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
- A Comparison of Surrogate Constitutive Models for Viscoplastic Creep Simulation of HT-9 Steel
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
- Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
- Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
- MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
- World Model Implanting for Test-time Adaptation of Embodied Agents
- Mixture-of-Clustered-Experts: Advancing Expert Specialization and Generalization in Instruction Tuning
- DrDiff: Dynamic Routing Diffusion with Hierarchical Attention for Breaking the Efficiency-Quality Trade-off
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- Batch Query Processing and Optimization for Agentic Workflows
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
- No Clustering, No Routing: How Transformers Actually Process Rare Tokens
- Universal Properties of Activation Sparsity in Modern Large Language Models
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Reasoning-Intensive Regression
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
- ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- MODE: Mixture of Document Experts for RAG
- MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
- Enabling MoE on the Edge via Importance-Driven Expert Scheduling
- FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- BTW: A Non-Parametric Variance Stabilization Framework for Multimodal Model Integration
- PGF-Net: A Progressive Gated-Fusion Framework for Efficient Multimodal Sentiment Analysis
- Global-Distribution Aware Scenario-Specific Variational Representation Learning Framework
- Successive Halving with Learning Curve Prediction via Latent Kronecker Gaussian Processes
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Routing Networks and the Challenges of Modular and Compositional Computation
- HiCL: Hippocampal-Inspired Continual Learning
- ASDFormer: A Transformer with Mixtures of Pooling-Classifier Experts for Robust Autism Diagnosis and Biomarker Discovery
- DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-Experts
- MUFFIN: Mixture of User-Adaptive Frequency Filtering for Sequential Recommendation
- Cross-Cancer Knowledge Transfer in WSI-based Prognosis Prediction
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Maximum Score Routing For Mixture-of-Experts
- Conditional Sum-Product Networks: Imposing Structure on Deep\n Probabilistic Architectures
- Towards High-Resolution Industrial Image Anomaly Detection
- Deploying Models to Non-participating Clients in Federated Learning without Fine-tuning: A Hypernetwork-based Approach
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Cost-Aware Contrastive Routing for LLMs
- FNH-TTS: A Fast, Natural, and Human-Like Speech Synthesis System with advanced prosodic modeling based on Mixture of Experts
- A Neural Tangent Kernel Perspective of Infinite Tree Ensembles
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Dynamic Mixture-of-Experts for Incremental Graph Learning
- μ-Parametrization for Mixture of Experts
- HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
- Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
- Learning Facts at Scale with Active Reading
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- Wavelet Mixture of Experts for Time Series Forecasting
- Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
- Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
- Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
- X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
- Separation and Collaboration: Two-Level Routing Grouped Mixture-of-Experts for Multi-Domain Continual Learning
- CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
- FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning
- Can Smaller Large Language Models Evaluate Research Quality?
- Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment
- N-BEATS-MOE: N-BEATS with a Mixture-of-Experts Layer for Heterogeneous Time Series Forecasting
- MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
- gpt-oss-120b & gpt-oss-20b Model Card
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- Deep Convolutional Decision Jungle for Image Classification
- Towards Unified Image Deblurring using a Mixture-of-Experts Decoder
- AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction Modelling
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
- HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
- Tesserae: Scalable Placement Policies for Deep Learning Workloads
- Toward Errorless Training ImageNet-1k
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
- Neuro-MoBRE: Exploring Multi-subject Multi-task Intracranial Decoding via Explicit Heterogeneity Resolving
- CodonMoE: DNA Language Models for mRNA Analyses
- Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
- Understanding Transformers through the Lens of Pavlovian Conditioning
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- Model-Agnostic Dynamic Feature Selection with Uncertainty Quantification
- Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
- DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
- DexReMoE:In-hand Reorientation of General Object via Mixtures of Experts
- Shape Distribution Matters: Shape-specific Mixture-of-Experts for Amodal Segmentation under Diverse Occlusions
- RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
- Applying multimodal learning to Classify transient Detections Early (AppleCiDEr) I: Data set, methods, and infrastructure
- PaPaformer: Language Model from Pre-trained Parallel Paths
- AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
- Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation
- A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- Fast and Accurate Contextual Knowledge Extraction Using Cascading Language Model Chains and Candidate Answers
- RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
- RelMap: Enhancing Online Map Construction with Class-Aware Spatial Relation and Semantic Priors
- Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
- Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning
- Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
- TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
- EffiComm: Bandwidth Efficient Multi Agent Communication
- RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
- DriftMoE: A Mixture of Experts Approach to Handle Concept Drifts
- Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil
- Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
- Convergence Rates for Gaussian Mixtures of Experts
- Retrospective and Prospective Mixture-of-Generators for Task-oriented Dialogue Response Generation
- R2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Adaptive Block-Scaled Data Types
- Rethinking Language Model Scaling under Transferable Hypersphere Optimization
- Network Transplanting
- Multi-Source Cross-Lingual Model Transfer: Learning What to Share
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- A Comprehensive Review on Harnessing Large Language Models to Overcome Recommender System Challenges
- Apple Intelligence Foundation Language Models: Tech Report 2025
- Mixture of Raytraced Experts
- CorrMoE: Mixture of Experts with De-stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
- Astro-MoE: Mixture of Experts for Multiband Astronomical Time Series
- Transferring Inter-Class Correlation
- Mixture of Experts in Large Language Models
- Atmos-Bench: 3D Atmospheric Structures for Climate Insight
- Interpretable Mixture Density Estimation by use of Differentiable\n Tree-module
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Multiple Choice Learning of Low Rank Adapters for Language Modeling
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- Memorization Sinks: Isolating Memorization during LLM Training
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Explainable AI in Genomics: Transcription Factor Binding Site Prediction with Mixture of Experts
- Advancing Large Language Models for Tibetan with Curated Data and Continual Pre-Training
- Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
- On Evaluating Performance of LLM Inference Serving Systems
- A statistical physics framework for optimal learning
- MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models
- Towards Robust Surrogate Models: Benchmarking Machine Learning Approaches to Expediting Phase Field Simulations of Brittle Fracture
- FlexOlmo: Open Language Models for Flexible Data Use
- SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
- Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate
- Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
- Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach
- QMoE: A Quantum Mixture of Experts Framework for Scalable Quantum Neural Networks
- Semantic Frame Interpolation
- DRAE: Dynamic Retrieval-Augmented Expert Networks for Lifelong Learning and Task Adaptation in Robotics
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
- Deep Learning of Continuous and Structured Policies for Aggregated Heterogeneous Treatment Effects
- Adaptive Slimming for Scalable and Efficient Speech Enhancement
- Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations
- Heterogeneous Causal Learning for Optimizing Aggregated Functions in User Growth
- AXLearn: Modular, Hardware-Agnostic Large Model Training
- Scaling Context Requires Rethinking Attention
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- Lifelong Learning of Compositional Structures
- Neural Inhibition Improves Dynamic Routing and Mixture of Experts
- BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
- CyberRAG: An Agentic RAG cyber attack classification and reporting tool
- Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
- Shared & Domain Self-Adaptive Experts with Frequency-Aware Discrimination for Continual Test-Time Adaptation
- Latent Part-of-Speech Sequences for Neural Machine Translation
- Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
- UMA: A Family of Universal Models for Atoms
- Achieving Human Parity on Visual Question Answering
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
- A Neural Dirichlet Process Mixture Model for Task-Free Continual Learning
- Compositions of Variant Experts for Integrating Short-Term and Long-Term Preferences
- Large language model [wikipedia]
- Filter and refine [wikipedia]
- Mixture of experts [wikipedia]
Discussions
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-Of-Experts Layer [hn, 200 points, 81 comments]
- Outrageously Large Neural Nets: Sparsely-Gated Mixture-of-Experts Layer (2017) [hn, 65 points, 33 comments]
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts (2017) [hn, 60 points, 10 comments]
- Outrageously Large Neural Networks [hn, 3 points, 0 comments]
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-Of-Experts Layer [hn, 3 points, 0 comments]
- Precisely, which is why I'm so excited about Mixture of Experts architecture advancements. arxiv.org/abs/1701.06538 [bsky, 2 points, 1 comments]
- Outrageously Large Neural Networks: Up to 137B Parameters [hn, 2 points, 1 comments]
- The Sparsely-Gated Mixture-of-Experts Layer (2017) [pdf] [hn, 1 points, 0 comments]
- Outrageously Large Neural Networks [hn, 1 points, 0 comments]
- Mixture-of-Experts? Train different experts together, sky-rocket the number of parameters, this is the MoE used by GPT4 and Gemini. Can we improve on the Shazeer et al. formulation? arxiv.org/abs/17 [bsky, 1 points, 1 comments]
- Neural networks are no longer learning, they're plotting. 137 billion parameters, 1% active at a time. Hyper-efficient cognitive oligarchies where only the chosen few speak. Intelligence is now a gate [bsky, 0 points, 0 comments]
Related