Adaptive Mixtures of Local Experts
1991/02/01 by Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan +1 · 554 citations
Computer Science · #Music and Audio Processing #Neural Networks and Applications #Speech and Audio Processing
paper · doi:10.1162/neco.1991.3.1.79
openalex publication_date 1991/02/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
We present a new supervised learning procedure for systems composed of many separate networks, each of which learns to handle a subset of the complete set of training cases. The new procedure can be viewed either as a modular version of a multilayer supervised network, or as an associative version of competitive learning. It therefore provides a new link between these two apparently different approaches. We demonstrate that the learning procedure divides up a vowel discrimination task into appropriate subtasks, each of which can be solved by a very simple expert network.
Cited by
- A Thermal Comfort Index for Healthy Indoor Environments: An Interpretable, Simulation‐Based Mixture‐of‐Experts Model
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
- Parsimonious Clustering of Covariance Matrices
- LSRS: Latent Scale Rejection Sampling for Visual Autoregressive Modeling
- Single-Round Scalable Analytic Federated Learning
- GRASP: Guided Residual Adapters with Sample-wise Partitioning
- Efficient Training of Diffusion Mixture-of-Experts Models: A Practical Recipe
- Structural Prognostic Event Modeling for Multimodal Cancer Survival Analysis
- MoB: Mixture of Bidders
- HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
- Resolving Conflicts in Lifelong Learning via Aligning Updates in Subspaces
- EnECG: Efficient Ensemble Learning for Electrocardiogram Multi-task Foundation Model
- ABounD: Adversarial Boundary-Driven Few-Shot Learning for Multi-Class Anomaly Detection
- EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
- DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
- Are Image-to-Video Models Good Zero-Shot Image Editors?
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
- MRI Super-Resolution with Deep Learning: A Comprehensive Survey
- Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution
- Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation
- STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection
- SplitFlux: Learning to Decouple Content and Style from a Single Image
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Region-Point Joint Representation for Effective Trajectory Similarity Learning
- Self-Adaptive Graph Mixture of Models
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- SAC-MoE: Reinforcement Learning with Mixture-of-Experts for Control of Hybrid Dynamical Systems with Uncertainty
- AMR-MoEGA: Antimicrobial Resistance Prediction using Mixture of Experts and Genetic Algorithms
- ViTE: Virtual Graph Trajectory Expert Router for Pedestrian Trajectory Prediction
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization
- RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
- GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion
- H-Model: Dynamic Neural Architectures for Adaptive Processing
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification
- One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
- Route Experts by Sequence, not by Token
- MoEGCL: Mixture of Ego-Graphs Contrastive Representation Learning for Multi-View Clustering
- Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
- A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
- SurgiATM: A Physics-Guided Plug-and-Play Model for Deep Learning-Based Smoke Removal in Laparoscopic Surgery
- MoE-DP: An MoE-Enhanced Diffusion Policy for Robust Long-Horizon Robotic Manipulation with Skill Decomposition and Failure Recovery
- AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
- Uncertainty Guided Online Ensemble for Non-stationary Data Streams in Fusion Science
- Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
- DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection
- Dynamic Model Selection for Trajectory Prediction via Pairwise Ranking and Meta-Features
- Soft Task-Aware Routing of Experts for Equivariant Representation Learning
- AFM-Net: Advanced Fusing Hierarchical CNN Visual Priors with Global Sequence Modeling for Remote Sensing Image Scene Classification
- Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
- Data Fusion and Contrastive Alignment for Unconstrained IR Molecular Structure Elucidation
- Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models
- Bayesian Neural Networks vs. Mixture Density Networks: Theoretical and Empirical Insights for Uncertainty-Aware Nonlinear Modeling
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- MoEMeta: Mixture-of-Experts Meta Learning for Few-Shot Relational Learning
- Switchable Token-Specific Codebook Quantization For Face Image Compression
- Scalable Neural Decoders for Practical Real-Time Quantum Error Correction
- Expert Merging in Sparse Mixture of Experts with Nash Bargaining
- Input Adaptive Bayesian Model Averaging
- Adaptive Graph Mixture of Residual Experts: Unsupervised Learning on Diverse Graphs with Heterogeneous Specialization
- FreeChunker: A Cross-Granularity Chunking Framework
- Soft Switching Expert Policies for Controlling Systems with Uncertain Parameters
- Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Backdoor or Manipulation? Graph Mixture of Experts Can Defend Against Various Graph Adversarial Attacks
- Neuronal Group Communication for Efficient Neural representation
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Mixture of Experts Approaches in Dense Retrieval Tasks
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
- IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
- A Unified Framework for Probabilistic Dynamic-, Trajectory- and Vision-based Virtual Fixtures
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps
- Graph Few-Shot Learning via Adaptive Spectrum Experts and Cross-Set Distribution Calibration
- Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training
- Variational Mixture of Graph Neural Experts for Alzheimer's Disease Recognition across Frequency Bands in EEG Brain Networks
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Hierarchical LoRA MoE for Efficient CTR Model Scaling
- Adaptive Heterogeneous Mixtures of Normalising Flows for Robust Variational Inference
- Compositional meta-learning through probabilistic task inference
- MODE: Learning compositional representations of complex systems with Mixtures Of Dynamical Experts
- LadderMoE: Ladder-Side Mixture of Experts Adapters for Bronze Inscription Recognition
- Mitigating Subject Dependency in EEG Decoding with Subject-Specific Low-Rank Adapters
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
- CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
- MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
- Intelligent AI Delegation
- Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
- MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
- ERDE: Entropy-Regularized Distillation for Early-exit
- DoRAN: Stabilizing Weight-Decomposed Low-Rank Adaptation via Noise Injection and Auxiliary Networks
- Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
- FR-LUX: Friction-Aware, Regime-Conditioned Policy Optimization for Implementable Portfolio Management
- Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Kairos: Towards Adaptive and Generalizable Time Series Foundation Models
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
- A Greedy PDE Router for Blending Neural Operators and Classical Methods
- Beyond Softmax: A Natural Parameterization for Categorical Random Variables
- Stabilizing Humanoid Robot Trajectory Generation via Physics-Informed Learning and Control-Informed Steering
- LEAF: A Robust Expert-Based Framework for Few-Shot Continual Event Detection
- Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- GES-UniGrasp: A Two-Stage Dexterous Grasping Strategy With Geometry-Based Expert Selection
- TimeExpert: Boosting Long Time Series Forecasting with Temporal Mix of Experts
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
- Multilingual Vision-Language Models, A Survey
- Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
- SlideMamba: Entropy-Based Adaptive Fusion of GNN and Mamba for Enhanced Representation Learning in Digital Pathology
- SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads
- Choosing to Be Green: Advancing Green AI via Dynamic Model Selection
- Multimodal-enhanced Federated Recommendation: A Group-wise Fusion Approach
- Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Mamba Modulation: On the Length Generalization of Mamba
- On component interactions in two-stage recommender systems
- ISALux: Illumination and Segmentation Aware Transformer Employing Mixture of Experts for Low Light Image Enhancement
- Deep Multi-Modal Sets
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
- Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
- Attention-based Mixture of Experts for Robust Speech Deepfake Detection
- SEQR: Secure and Efficient QR-based LoRA Routing
- StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
- ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
- MoPE: A Mixture of Password Experts for Improving Password Guessing
- Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- TrueMoE: Dual-Routing Mixture of Discriminative Experts for Synthetic Image Detection
- MoE-CE: Enhancing Generalization for Deep Learning based Channel Estimation via a Mixture-of-Experts Framework
- Super-Linear: A Lightweight Pretrained Mixture of Linear Experts for Time Series Forecasting
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Mixture of Low-Rank Adapter Experts in Generalizable Audio Deepfake Detection
- NavMoE: Hybrid Model- and Learning-based Traversability Estimation for Local Navigation via Mixture of Experts
- AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
- Modular Gaussian Processes for Transfer Learning
- When MoE Meets Blockchain: A Trustworthy Distributed Framework of Large Models
- SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
- MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
- Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
- Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
- XAgents: A Unified Framework for Multi-Agent Cooperation via IF-THEN Rules and Multipolar Task Processing Graph
- Distilling the Knowledge in a Neural Network
- Conditional Density Estimation with Bayesian Normalising Flows
- SEEC: Segmentation-Assisted Multi-Entropy Models for Learned Lossless Image Compression
- AdaMixT: Adaptive Weighted Mixture of Multi-Scale Expert Transformers for Time Series Forecasting
- Hierarchical Mixtures of Experts and the EM Algorithm
- Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- Lookup multivariate Kolmogorov-Arnold Networks
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- Learning to Route: Per-Sample Adaptive Routing for Multimodal Multitask Prediction
- A Comparison of Surrogate Constitutive Models for Viscoplastic Creep Simulation of HT-9 Steel
- Extracting Uncertainty Estimates from Mixtures of Experts for Semantic Segmentation
- Multi-Faceted Hierarchical Multi-Task Learning for a Large Number of Tasks with Multi-dimensional Relations
- Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
- Robust Experts: the Effect of Adversarial Training on CNNs with Sparse Mixture-of-Experts Layers
- Decoupled Entity Representation Learning for Pinterest Ads Ranking
- Suppressing the Unusual: towards Robust CNNs using Symmetric Activation Functions
- MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
- Deep Mixed Effect Model using Gaussian Processes: A Personalized and Reliable Prediction for Healthcare
- Building Watson: An Overview of the DeepQA Project
- Aerodynamic Data Predictions Based on Multi-task Learning
- Joint Information Extraction Across Classical and Modern Chinese with Tea-MOELoRA
- MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
- TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization
- Learning Slice-Aware Representations with Mixture of Attentions
- Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
- Reasoning-Intensive Regression
- Learning with springs and sticks
- Hard Mixtures of Experts for Large Scale Weakly Supervised Vision
- WISCA: A Lightweight Model Transition Method to Improve LLM Training via Weight Scaling
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Routing Networks and the Challenges of Modular and Compositional Computation
- MUFFIN: Mixture of User-Adaptive Frequency Filtering for Sequential Recommendation
- Cross-Cancer Knowledge Transfer in WSI-based Prognosis Prediction
- Applications of machine learning in gravitational-wave research with current interferometric detectors
- Towards High-Resolution Industrial Image Anomaly Detection
- A Tutorial on Deep Latent Variable Models of Natural Language
- STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction
- Cost-Aware Contrastive Routing for LLMs
- LoRAtorio: An intrinsic approach to LoRA Skill Composition
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- Deep Latent Emotion Network for Multi-Task Learning
- Approximation rates for finite mixtures of location-scale models
- Efficient Patent Searching Using Graph Transformers
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- μ-Parametrization for Mixture of Experts
- Wavelet Mixture of Experts for Time Series Forecasting
- Continual Learning in Neural Networks
- CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving
- Separation and Collaboration: Two-Level Routing Grouped Mixture-of-Experts for Multi-Domain Continual Learning
- Can Smaller Large Language Models Evaluate Research Quality?
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
- A Mixture of Expert Based Deep Neural Network for Improved ASR
- Towards Unified Image Deblurring using a Mixture-of-Experts Decoder
- RL-MoE: An Image-Based Privacy Preserving Approach In Intelligent Transportation System
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
- Contextual Explanation Networks
- Neural Estimation of Information Leakage for Secure Communication System Design
- HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- CodonMoE: DNA Language Models for mRNA Analyses
- UNISELF: A Unified Network with Instance Normalization and Self-Ensembled Lesion Fusion for Multiple Sclerosis Lesion Segmentation
- On the Functional Equivalence of TSK Fuzzy Systems to Neural Networks, Mixture of Experts, CART, and Stacking Ensemble Regression
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
- Hierarchical MoE: Continuous Multimodal Emotion Recognition with Incomplete and Asynchronous Inputs
- Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
- Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
- Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
- VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation
- DexReMoE:In-hand Reorientation of General Object via Mixtures of Experts
- GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View Geometry
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
- Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- Max-Affine Regression: Provable, Tractable, and Near-Optimal Statistical Estimation
- Dynamics-Aware Unsupervised Discovery of Skills
- Discontinuity-Sensitive Optimal Control Learning by Mixture of Experts
- The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach
- Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
- Augmented Artificial Intelligence: a Conceptual Framework
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Learning Explainable Stock Predictions with Tweets Using Mixture of Experts
- MoCTEFuse: Illumination-Gated Mixture of Chiral Transformer Experts for Multi-Level Infrared and Visible Image Fusion
- Nonlinear Semi-Parametric Models for Survival Analysis
- PatchTraj: Unified Time-Frequency Representation Learning via Dynamic Patches for Trajectory Prediction
- Multi-Task Dense Prediction Fine-Tuning with Mixture of Fine-Grained Experts
- DriftMoE: A Mixture of Experts Approach to Handle Concept Drifts
- Simultaneous Feature and Expert Selection within Mixture of Experts
- Convergence Rates for Gaussian Mixtures of Experts
- Fine-Grained Classification via Mixture of Deep Convolutional Neural Networks
- CoGMoE: Sparse and specialized framework for multi-agent collaborative perception via graph mixture-of-experts
- Probabilistic Data Analysis with Probabilistic Programming
- 50 years since the Marr, Ito, and Albus models of the cerebellum
- R2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- Complex-Valued Hough Transforms for Circles
- Speech Enhancement using a Deep Mixture of Experts
- SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
- Photon-Driven Neural Path Guiding
- Spatialize v1.0: A Python/C++ Library for Ensemble Spatial Interpolation
- A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
- BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
- An Introduction to the Practical and Theoretical Aspects of Mixture-of-Experts Modeling
- Adaptive Market Intelligence: A Mixture of Experts Framework for Volatility-Sensitive Stock Forecasting
- Learning, fast and slow: a two-fold algorithm for data-based model adaptation
- Rank of Experts: Detection Network Ensemble
- Generative quantum combinatorial optimization by means of a novel conditional generative quantum eigensolver
- Macular OCT Classification Using a Multi-Scale Convolutional Neural Network Ensemble
- Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
- Generalizable Person Re-identification with Relevance-aware Mixture of Experts
- Latent Domain Learning with Dynamic Residual Adapters
- From Matching to Generation: A Survey on Generative Information Retrieval
- Improving Spark Application Throughput Via Memory Aware Task Co-location: A Mixture of Experts Approach
- ExpertMatcher: Automating ML Model Selection for Clients using Hidden Representations
- Pre-Trained Models: Past, Present and Future
- Scenario-aware and Mutual-based approach for Multi-scenario Recommendation in E-Commerce
- Mixture of Experts in Large Language Models
- Learning Factored Representations in a Deep Mixture of Experts
- Deep Gaussian Covariance Network
- Multiple Choice Learning of Low-Rank Adapters for Language Modeling
- MERIT: A Merchant Incentive Ranking Model for Hotel Search & Ranking
- Ambient Diffusion Omni: Training Good Models with Bad Data
- Explainable AI in Genomics: Transcription Factor Binding Site Prediction with Mixture of Experts
- CKAA: Cross-subspace Knowledge Alignment and Aggregation for Robust Continual Learning
- MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains
- The Bayesian Approach to Continual Learning: An Overview
- Input Conditioned Layer Dropping in Speech Foundation Models
- MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification
- Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
- Dynamic Slimmable Networks for Efficient Speech Separation
- Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
- Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
- QMoE: A Quantum Mixture of Experts Framework for Scalable Quantum Neural Networks
- Semantic Frame Interpolation
- Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
- Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection
- Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge
- Phoneme-Based Ratio Mask Estimation for Reverberant Speech Enhancement in Cochlear Implant Processors
- Enhanced accuracy through ensembling of randomly initialized auto-regressive models for time-dependent PDEs
- Mixtures of Gaussian Processes for regression under multiple prior distributions
- FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
- MoIRA: Modular Instruction Routing Architecture for Multi-Task Robotics
- Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
- MotionGPT3: Human Motion as a Second Modality
- Restoring Spatially-Heterogeneous Distortions using Mixture of Experts Network
- Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
- Changing Structures in Midstream: Learning Along the Statistical Garden Path
- UMA: A Family of Universal Models for Atoms
- Achieving Human Parity on Visual Question Answering
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
- A Neural Dirichlet Process Mixture Model for Task-Free Continual Learning
- Masked Gated Linear Unit
- Compositions of Variant Experts for Integrating Short-Term and Long-Term Preferences
- Integrating Large Language Models in Financial Investments and Market Analysis: A Survey
- Towards Distributed Neural Architectures
- Lessons Learned from the Training of GANs on Artificial Datasets
- Latent Prototype Routing: Achieving Near-Perfect Load Balancing in Mixture-of-Experts
- EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
- Structural Decoupling: A Scaffold-Flow Theory of Generalization and Alignment
- MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
- Modifications of the BIC for order selection in finite mixture models
- A Two-Phase Deep Learning Framework for Adaptive Time-Stepping in High-Speed Flow Modeling
- M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration
- Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
- Expected Information Maximization: Using the I-Projection for Mixture Density Estimation
- SafeClick: Error-Tolerant Interactive Segmentation of Any Medical Volumes via Hierarchical Expert Consensus
- Sharpening the Spear: Adaptive Expert-Guided Adversarial Attack Against DRL-based Autonomous Driving Policies
- Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
- Object Files and Schemata: Factorizing Declarative and Procedural Knowledge in Dynamical Systems
- Learning to imitate stochastic time series in a compositional way by chaos
- PDC-Net: Pattern Divide-and-Conquer Network for Pelvic Radiation Injury Segmentation
- Sparse MoEs meet Efficient Ensembles
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- Striosomes and Matrisomes: Scaffolds for Dynamic Coupling of Volition and Action
- PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
- Enhancing eLoran Timing Accuracy via Machine Learning with Meteorological and Terrain Data
- Towards structural systematicity in distributed, statically bound visual representations
- Encoders and Ensembles for Task-Free Continual Learning
- Guideline Forest: Retrieval-Augmented Reasoning with Branching Experience-Induced Guidelines
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- Single-Example Learning in a Mixture of GPDMs with Latent Geometries
- Attention over Parameters for Dialogue Systems
- MoORE: SVD-based Model MoE-ization for Conflict- and Oblivion-Resistant Multi-Task Adaptation
- Leveraging External Factors in Household-Level Electrical Consumption Forecasting using Hypernetworks
- GRaD-Nav++: Vision-Language Model Enabled Visual Drone Navigation with Gaussian Radiance Fields and Differentiable Dynamics
- Brain Imaging Foundation Models, Are We There Yet? A Systematic Review of Foundation Models for Brain Imaging and Biomedical Research
- Mixture of ELM based experts with trainable gating network
- Mixture of linear experts model for censored data: A novel approach with scale-mixture of normal distributions
- TagRouter: Learning Route to LLMs through Tags for Open-Domain Text Generation Tasks
- Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
- HarMoEny: Efficient Multi-GPU Inference of MoE Models
- Topology-Assisted Spatio-Temporal Pattern Disentangling for Scalable MARL in Large-scale Autonomous Traffic Control
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- LoRA-Gen: Specializing Large Language Model via Online LoRA Generation
- Macro Graph of Experts for Billion-Scale Multi-Task Recommendation
- Bayesian shrinkage in mixture of experts models: Identifying robust determinants of class membership
- MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension
- FA-INR: Adaptive Implicit Neural Representations for Interpretable Exploration of Simulation Ensembles
- Dynamic Mixture of Progressive Parameter-Efficient Expert Library for Lifelong Robot Learning
- From Generalized zero-shot learning to long-tail with class descriptors
- Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
- Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
- Mixture-of-Experts Meets In-Context Reinforcement Learning
- Handling Long-Tail Queries with Slice-Aware Conversational Systems
- Towards More Effective and Economic Sparsely-Activated Model
- Shadow Wireless Intelligence: Large Language Model-Driven Reasoning in Covert Communications
- Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
- Dynamic Neural Networks: A Survey
- Reconstruction of sequential data with density models
- PLUME: Polyhedral Learning Using Mixture of Experts
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- A Survey on Spatial and Spatiotemporal Prediction Methods
- Improving Generalized Zero-Shot Learning by Semantic Discriminator
- Scaling Vision with Sparse Mixture of Experts
- ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
- Mode-adaptive neural networks for quadruped motion control
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
- Human-Understandable Decision Making for Visual Recognition
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
- Modeling the Sequential Dependence among Audience Multi-step Conversions with Multi-task Learning in Targeted Display Advertising
- Deep learning spinfoam vertex amplitudes: the Euclidean Barrett-Crane model
- SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation Model
- FLEx: Personalized Federated Learning for Mixture-of-Experts LLMs via Expert Grafting
- Assembly of Experts: Linear-time construction of the Chimera LLM variants with emergent and adaptable behaviors
- FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
- Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
- 3D Gaussian Splatting Data Compression with Mixture of Priors
- MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition
- Two Is Better Than One: Rotations Scale LoRAs
- Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
- From Knowledge to Noise: CTIM-Rover and the Pitfalls of Episodic Memory in Software Engineering Agents
- DOPPLER: Dual-Policy Learning for Device Assignment in Asynchronous Dataflow Graphs
- A Survey on Long-Tailed Visual Recognition
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
- Graph Classification by Mixture of Diverse Experts
- Identity-Preserving Text-to-Image Generation via Dual-Level Feature Decoupling and Expert-Guided Fusion
- Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images
- Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects
- Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
- RetroMotion: Retrocausal Motion Forecasting Models are Instructable
- Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
- NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
- Nested Mixture of Experts: Cooperative and Competitive Learning of Hybrid Dynamical System
- Understanding Locally Competitive Networks
- WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
- Finger Pose Estimation for Under-screen Fingerprint Sensor
- Incorporating Legal Structure in Retrieval-Augmented Generation: A Case Study on Copyright Fair Use
- CMoS: Rethinking Time Series Prediction Through the Lens of Chunk-wise Spatial Correlations
- I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
- On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
- MoveGPT: Scaling Mobility Foundation Models with Spatially-Aware Mixture of Experts
- Neural Mixture Distributional Regression
- Predictive Uncertainty Quantification with Compound Density Networks
- Enhancing CTR Prediction with De-correlated Expert Networks
- NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
- Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
- Discrete-Valued Neural Communication
- CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
- LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
- Parameter Learning for Alpha Integration
- Extending Factor Graphs so as to Unify Directed and Undirected Graphical Models
- Robust Multimodal Learning via Entropy-Gated Contrastive Fusion
- Cross-Modal Alignment with Mixture Experts Neural Network for Intral-City Retail Recommendation
- MoTE: Mixture of Task-specific Experts for Pre-Trained ModelBased Class-incremental Learning
- Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks
- Balanced and Elastic End-to-end Training of Dynamic LLMs
- StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning
- Towards Rehearsal-Free Continual Relation Extraction: Capturing Within-Task Variance with Adaptive Prompting
- Modularity as a Means for Complexity Management in Neural Networks Learning
- THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation
- Generalizable Multispectral Land Cover Classification via Frequency-Aware Mixture of Low-Rank Token Experts
- Multimodal Mixture of Low-Rank Experts for Sentiment Analysis and Emotion Recognition
- MAFA: A multi-agent framework for annotation
- Seeing the Unseen: How EMoE Unveils Bias in Text-to-Image Diffusion Models
- Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
- Multi-Head Adapter Routing for Cross-Task Generalization
- Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization
- Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression Recognition
- Improving Coverage in Combined Prediction Sets with Weighted p-values
- Chain-of-Model Learning for Language Model
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- A machine learning perspective on the emotional content of Parkinsonian speech
- Hidden-Domain Routing for All-Type Audio Deepfake Detection
- Multi-Task Reinforcement Learning with Soft Modularization
- LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive Diagnosis
- Aquarius: A Family of Industry-Level Video Generation Models for Marketing Scenarios
- Automatic Task Detection and Heterogeneous LLM Speculative Decoding
- Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
- A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- A Minimum Relative Entropy Controller for Undiscounted Markov Decision Processes
- Convergence Rates for Mixture-of-Experts
- The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
- Self-informed neural network structure learning
- Towards Explainable Fact Checking
- Scalable LLM Math Reasoning Acceleration with Low-rank Distillation
- A Mixture of Expert Approach for Low-Cost Customization of Deep Neural Networks
- Quality Resilient Deep Neural Networks
- XNAS: Neural Architecture Search with Expert Advice
- Learning Heterogeneous Mixture of Scene Experts for Large-scale Neural Radiance Fields
- UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
- CoCoAFusE: Beyond Mixtures of Experts via Model Fusion
- Bayesian Inference for Logistic Regression Models Using Sequential Posterior Simulation
- Unsupervised learning of regression mixture models with unknown number of components
- Improving Routing in Sparse Mixture of Experts with Graph of Tokens
- Levels and Types of Action Selection: The Action Selection Soup
- PICO: Primitive Imitation for COntrol
- Locally Adaptive Nonparametric Binary Regression
- Neuromorphic electronic circuits for building autonomous cognitive systems
- A review and comparison of strategies for multi-step ahead time series forecasting based on the NN5 forecasting competition
- Lightweight Adaptive Mixture of Neural and N-gram Language Models
- ORTHOBO: Orthogonal Bayesian Hyperparameter Optimization
- Large Language Models as AI Agents for Digital Atoms and Molecules: Catalyzing a New Era in Computational Biophysics
- Pricing Derivatives under Self-Exciting Dynamics: A Finite-Difference and Transform Approach
- Graph Inference Representation: Learning Graph Positional Embeddings with Anchor Path Encoding
- Toward Calibrated Mixture-of-Experts Under Distribution Shift
- Statistical mechanics of learning in the presence of outliers
- Open Set Medical Diagnosis
- MEDNA-DFM: A Dual-View FiLM-MoE Model for Explainable DNA Methylation Prediction
- MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
- Less is More: Rejecting Unreliable Reviews for Product Question Answering
- Neural Pharmacodynamic State Space Modeling
- Why have a Unified Predictive Uncertainty? Disentangling it using Deep Split Ensembles
- GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
- TT-LoRA MoE: Unifying Parameter-Efficient Fine-Tuning and Sparse Mixture-of-Experts
- MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
- The Neuronal Replicator Hypothesis
- AKIBoards: A Structure-Following Multiagent System for Predicting Acute Kidney Injury
- Token-Level Prompt Mixture with Parameter-Free Routing for Federated Domain Generalization
- In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- DanceOPD: On-Policy Generative Field Distillation
- A Principal Components Approach to Combining Regression Estimates
- Probabilistic Wind Vector Forecasting Using Ensembles and Bayesian Model Averaging
- BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
- Neural network task specialization via domain constraining
- ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving
- Mixture of Experts for Decentralized Generative AI and Reinforcement Learning in Wireless Networks: A Comprehensive Survey
- Multi-step retrieval and reasoning improves radiology question answering with large language models
- Risk-Sensitive Specialist Routing for Volatility Forecasting
- Achieving Human Parity on Visual Question Answering
- Prediction Markets as Bayesian Inverse Problems: Uncertainty Quantification, Identifiability, and Information Gain from Price-Volume Histories under Latent Types
- Revisiting Neural Retrieval on Accelerators
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
- Balancing Speed and Quality in Online Learning to Rank for Information Retrieval
- Ensemble learning for data stream analysis: A survey
- Classifier Ensembles for Changing Environments
- BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts
- Semiparametric robust mixture of experts based on nonparametric maximum likelihood
- MGSB: Manifold Gated Signature Branch Pressure-Domain Baseline Architecture for Two-Phase Pipeline Flows Under Distributional Shift
- ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning
- SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
- Learning Compatibility Across Categories for Heterogeneous Item Recommendation
- MME: Mixture of Mesh Experts with Random Walk Transformer Gating
- Mixing Data-Driven and Physics-Based Constitutive Models using Uncertainty-Driven Phase Fields
- A finite mixture model approach to regression under covariate misclassification
- On the Spatial Structure of Mixture-of-Experts in Transformers
- Multi-Scale Tensorial Summation and Dimensional Reduction Guided Neural Network for Edge Detection
- Interpretable Representation Learning for Speech and Audio Signals Based on Relevance Weighting
- Nonparametric modal regression
- Individualized Multi-directional Variable Selection
- Affect in Tweets Using Experts Model
- Multi-Level Factorisation Net for Person Re-Identification
- Prioritized Sweeping Neural DynaQ with Multiple Predecessors, and Hippocampal Replays
- Mixed-membership of experts stochastic blockmodel
- XEM: An explainable-by-design ensemble method for multivariate time series classification
- A Discriminative Technique for Multiple-Source Adaptation
- Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
- HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
- Persona-Pruner: Sculpting Lightweight Models for Role-Playing
- Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning
- Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation
- Mixture of Group Experts for Learning Invariant Representations
- Fast-Slow-Thinking: Complex Task Solving with Large Language Models
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- A fast and independent architecture of artificial neural network for permeability prediction
- SEE: Continual Fine-tuning with Sequential Ensemble of Experts
- Meta-Continual Learning of Neural Fields
- Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
- Large Language Models Enhanced Hyperbolic Space Recommender Systems
- DA2Diff: Exploring Degradation-aware Adaptive Diffusion Priors for All-in-One Weather Restoration
- Sub-Clustering for Class Distance Recalculation in Long-Tailed Drug Classification
- LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
Related