Mixtral of Experts
2024/01/08 by Albert Q. Jiang, Alexandre Sablayrolles, Jiang, Albert Q. +51 · 5 voices · 444 citations
Computer Science · Engineering · #Artificial intelligence #Code (set theory) #Computer network #Computer science #Context (archaeology) #Engineering #Feed forward #Inference #Language model #Layer (electronics) #License #Machine Learning and Algorithms #Machine Learning and Data Classification #Operating system #Process (computing) #Programming language #Router #Security token #Set (abstract data type) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2401.04088
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/01/08 · openalex created_date 2024/01/13 · openalex updated_date 2026/08/03
Abstract
We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts). For every token, at each layer, a router network selects two experts to process the current state and combine their outputs. Even though each token only sees two experts, the selected experts can be different at each timestep. As a result, each token has access to 47B parameters, but only uses 13B active parameters during inference. Mixtral was trained with a context size of 32k tokens and it outperforms or matches Llama 2 70B and GPT-3.5 across all evaluated benchmarks. In particular, Mixtral vastly outperforms Llama 2 70B on mathematics, code generation, and multilingual benchmarks. We also provide a model fine-tuned to follow instructions, Mixtral 8x7B - Instruct, that surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B - chat model on human benchmarks. Both the base and instruct models are released under the Apache 2.0 license.
Cited by
- Enhancing SLMs for Sustainable Code Optimization in Radio-Astronomy
- Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
- Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
- Efficient Multi-round LLM Inference over Disaggregated Serving
- AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System
- DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
- ChipChat: Low-Latency Cascaded Conversational Agent in MLX
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
- Articulated Humanoid Head for a Robot Receptionist Capable of Natural Human Interaction
- AirMoE: Statistic-Augmented Over-the-Air MoE for Collaborative Intelligence
- It Takes a MAESTRO To Prune Bad Experts
- xHC: Expanded Hyper-Connections
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
- In-Context Probing for Membership Inference in Fine-Tuned Language Models
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
- Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
- 34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery
- Mixture of Experts Made Intrinsically Interpretable
- Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
- SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
- On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks
- CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
- DMA: Online RAG Alignment with Human Feedback
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
- Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
- Memory for Large Language Models
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- Universal Pansharpening Model
- An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled Data
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- UCCL-EP: Portable Expert-Parallel Communication
- MoE Pathfinder: Trajectory-driven Expert Pruning
- Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
- Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
- Sigma-MoE-Tiny Technical Report
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- MIDUS: Memory-Infused Depth Up-Scaling
- SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
- Multi-Intent Spoken Language Understanding: Methods, Trends, and Challenges
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- LAWS: Learning from Actual Workloads Symbolically -- A Self-Certifying Parametrized Cache Architecture for Neural Inference, Robotics, and Edge Deployment
- Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLM
- CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
- OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
- KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- Life-IQA: Boosting Blind Image Quality Assessment through GCN-enhanced Layer Interaction and MoE-based Feature Decoupling
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- Exploiting the Experts: Unauthorized Compression in MoE-LLMs
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning
- Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
- Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- 47B Mixture-of-Experts Beats 671B Dense Models on Chinese Medical Examinations
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Route Experts by Sequence, not by Token
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
- Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
- CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
- From Prompts to Power: Measuring the Energy Footprint of LLM Inference
- AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
- FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
- Empowering LLMs with Structural Role Inference for Zero-Shot Graph Learning
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- Digital Prompting in Education: A Design Framework and Bibliometric Analysis
- Multimodal learning enables chat-based exploration of single-cell data
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- CatPath‐GPT: A Mixture of Experts System for Computational Catalyst Design
- MIN-Merging: Merge the Important Neurons for Model Merging
- PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
- MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- Retrieval- and Argumentation-Enhanced Multi-Agent LLMs for Judgmental Forecasting (Extended Version with Supplementary Material)
- SALS: Sparse Attention in Latent Space for KV cache Compression
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- Sparsity and Superposition in Mixture of Experts
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- When and Why Does Multi-Agent Debate Fail and Does It Really Underperform?
- A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
- HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
- AcademicEval: Live Long-Context LLM Benchmark
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models
- Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
- GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection
- Neural Weight Compression for Language Models
- DND: Boosting Large Language Models with Dynamic Nested Depth
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- Preference-driven Knowledge Distillation for Few-shot Node Classification
- NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
- Group-Adaptive Adversarial Learning for Robust Fake News Detection Against Malicious Comments
- Active Model Selection for Large Language Models
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- LadderMoE: Ladder-Side Mixture of Experts Adapters for Bronze Inscription Recognition
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
- MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
- Intelligent AI Delegation
- Learning to Route LLMs from Bandit Feedback: One Policy, Many Trade-offs
- Mixture of Neuron Experts
- Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- TASP: Topology-aware Sequence Parallelism
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Think Less, Label Better: Multi-Stage Domain-Grounded Synthetic Data Generation for Fine-Tuning Large Language Models in Telecommunications
- Collaborative Compression for Large-Scale MoE Deployment on Edge
- Massively Multimodal Foundation Models: A Framework for Capturing Interactions with Specialized Mixture-of-Experts
- How Well Do LLMs Imitate Human Writing Style?
- A Greedy PDE Router for Blending Neural Operators and Classical Methods
- EOE: Evolutionary Optimization of Experts for Training Language Models
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
- One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- Evaluating Program Semantics Reasoning with Type Inference in System F
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Your Dense Retriever is Secretly an Expeditious Reasoner
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- MMPB: It's Time for Multi-Modal Personalization
- Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
- StyleBench: Evaluating thinking styles in Large Language Models
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
- Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
- DeepResearch Agent System
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Recursive transformers for semiconductor thermo-mechanical reliability
- OPENXRD: a comprehensive benchmark framework for LLM/MLLM XRD question answering
- EMO: Pretraining Mixture of Experts for Emergent Modularity
- A large-scale evaluation of commonsense knowledge in humans and large language models
- GZSL-MoE: Apprentissage Généralisé Zéro-Shot basé sur le Mélange d'Experts pour la Segmentation Sémantique de Nuages de Points 3DAppliqué à un Jeu de Données d'Environnement de Collaboration Humain-Robot
- Who is In Charge? Dissecting Role Conflicts in Instruction Following
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- Probabilistic Token Alignment for Large Language Model Fusion
- VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- FURINA: Free from Unmergeable Router via LINear Aggregation of mixed experts
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
- Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
- HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- PICO: Performance Insights for Collective Operations
- Joint Learning using Mixture-of-Expert-Based Representation for Enhanced Speech Generation and Robust Emotion Recognition
- MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning
- DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
- SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery
- SoK: Security and Privacy of AI Agents for Blockchain
- Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
- veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
- SpikingBrain: Spiking Brain-inspired Large Models
- Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
- TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- Transforming Agency. On the mode of existence of Large Language Models
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
- HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
- Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning
- Transduction is All You Need for Structured Data Workflows
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
- Cost-Aware Contrastive Routing for LLMs
- Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
- A Comprehensive Evaluation framework of Alignment Techniques for LLMs
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- AIS-LLM: A Unified Framework for Maritime Trajectory Prediction, Anomaly Detection, and Collision Risk Assessment with Explainable Forecasting
- Towards Multimodal Sentiment Analysis via Contrastive Cross-modal Retrieval Augmentation and Hierachical Prompts
- Large Language Models for Subjective Language Understanding: A Survey
- Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
- Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
- N-BEATS-MOE: N-BEATS with a Mixture-of-Experts Layer for Heterogeneous Time Series Forecasting
- Generalizing Scaling Laws for Dense and Sparse Large Language Models
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- HAMoBE: Hierarchical and Adaptive Mixture of Biometric Experts for Video-based Person ReID
- Training-Free Multimodal Large Language Model Orchestration
- LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- AniMer+: Unified Pose and Shape Estimation Across Mammalia and Aves via Family-Aware Transformer
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI
- Argumentatively Coherent Judgmental Forecasting
- Rethinking LLM Inference Bottlenecks: Insights from Latent Attention and Mixture-of-Experts
- HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
- Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling
- UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
- PRAGMA: Revolut Foundation Model
- Large Language Models in the Travel Domain: An Industrial Experience
- Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
- BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- Large Language Models for Combinatorial Optimization of Design Structure Matrix
- Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
- From Matching to Generation: A Survey on Generative Information Retrieval
- Mixture of Experts in Large Language Models
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Multiple Choice Learning of Low-Rank Adapters for Language Modeling
- FusionFactory: Fusing LLM Capabilities with Multi-LLM Log Data
- Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
- Optimizing Sequential Multi-Step Tasks with Parallel LLM Agents
- BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
- SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
- Omni-Video: Democratizing Unified Video Understanding and Generation
- Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- CogniSQL-R1-Zero: Lightweight Reinforced Reasoning for Efficient SQL Generation
- An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
- A Technical Survey of Reinforcement Learning Techniques for Large Language Models
- OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
- ReservoirChat: Interactive Documentation Enhanced with LLM and Knowledge Graph for ReservoirPy
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- Scalable evaluation framework for retrieval augmented generation in tobacco research using large Language models
- Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
- UMA: A Family of Universal Models for Atoms
- MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
- Quantifying Fairness in LLMs Beyond Tokens: A Semantic and Statistical Perspective
- From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
- Enhancing Document Retrieval in COVID-19 Research: Leveraging Large Language Models for Hidden Relation Extraction
- AI Through the Human Lens: Investigating Cognitive Theories in Machine Psychology
- Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
- May the Feedback Be with You! Unlocking the Power of Feedback-Driven Deep Learning Framework Fuzzing via LLMs
- Beyond the Link: Assessing LLMs' ability to Classify Political Content across Global Media
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- Cash or Comfort? How LLMs Value Your Inconvenience
- Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
- Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
- REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Processing
- Intelligent Assistants for the Semiconductor Failure Analysis with LLM-Based Planning Agents
- Mix-of-Language-Experts Architecture for Multilingual Programming
- A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
- RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
- Utility-Driven Speculative Decoding for Mixture-of-Experts
- Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
- OneRec Technical Report
- Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
- PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
- dots.llm1 Technical Report
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
- FLAM: Frame-Wise Language-Audio Modeling
- Lifelong Evolution: Collaborative Learning between Large and Small Language Models for Continuous Emergent Fake News Detection
- FlashMoE: Fast Distributed MoE in a Single Kernel
- Kinetics: Rethinking Test-Time Scaling Laws
- A Dataset for Addressing Patient's Information Needs related to Clinical Course of Hospitalization
- Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
- DrSR: LLM based Scientific Equation Discovery with Dual Reasoning from Data and Experience
- AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
- Faster MoE LLM Inference for Extremely Large Models
- Adaptive Graph Pruning for Multi-Agent Communication
- When LLMs Team Up: The Emergence of Collaborative Affective Computing
- Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review
- NAVER LABS Europe Submission to the Instruction-following Track
- DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
- Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
- Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
- SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
- Advancing Expert Specialization for Better MoE
- Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers
- Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
- Scalable, Symbiotic, AI and Non-AI Agent Based Parallel Discrete Event Simulations
- RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
- AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks
- ALTER: All-in-One Layer Pruning and Temporal Expert Routing for Efficient Diffusion Generation
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- Explaining Large Language Models with gSMILE
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
- RetroMotion: Retrocausal Motion Forecasting Models are Instructable
- Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- Rethinking the Understanding Ability across LLMs through Mutual Information
- NextG-GPT: Leveraging GenAI for Advancing Wireless Networks and Communication Research
- Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
- Knowledge Grafting of Large Language Models
- μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
- On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
- LatentLLM: Attention-Aware Joint Tensor Compression
- Reasoning Meets Personalization: Unleashing the Potential of Large Reasoning Model for Personalized Generation
- Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
- LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
- DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
- Collaboration among Multiple Large Language Models for Medical Question Answering
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
- Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
- AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound
- VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
- Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
- GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation
- GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges
- Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
- Chain-of-Model Learning for Language Model
- Relative Value Encoding in Large Language Models: A Multi-Task, Multi-Model Investigation
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
- Phare: A Safety Probe for Large Language Models
- TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
- CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
- MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
- Camera Control at the Edge with Language Models for Scene Understanding
- The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
- MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL
- Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
- TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration
- An overview of artificial intelligence in computer-assisted language learning
- Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
- Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
- Recursive Multi-Agent Systems
- Toward Calibrated Mixture-of-Experts Under Distribution Shift
- Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation
- Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
- Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
- GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
- Memorization and Knowledge Injection in Gated LLMs
- C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
- SD-MoE: Spectral Decomposition for Effective Expert Specialization
- HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models
- DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
- In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- Mapping the Italian Telegram Ecosystem: Communities, Toxicity, and Hate Speech
- SpeechMapper: Speech-to-text Embedding Projector for LLMs
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
- Mixture of Experts for Decentralized Generative AI and Reinforcement Learning in Wireless Networks: A Comprehensive Survey
- Detect, Explain, Escalate: Sustainable Dialogue Breakdown Management for LLM Agents
- Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
- Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
- Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
- The Ultimate Cookbook for Invisible Poison: Crafting Subtle Clean-Label Text Backdoors with Style Attributes
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
- Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
- On the Spatial Structure of Mixture-of-Experts in Transformers
- A Framework for Testing and Adapting REST APIs as LLM Tools
- DRAGON: Distributional Rewards Optimize Diffusion Generative Models
- Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
- LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval
- Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
- EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding
- Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
- NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results
- Can LLMs Classify CVEs? Investigating LLMs Capabilities in Computing CVSS Vectors
- C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing
- Scaling Laws for Native Multimodal Models
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- LSR-MCTS: Alleviating Long Range Dependency in Code Generation
- Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
- Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation
- HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
- Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering
- Generative Large Language Model usage in Smart Contract Vulnerability Detection
- LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
- HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
- Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
- Mixture of experts [wikipedia]
Discussions
- Mixtral 8x7B: A sparse Mixture of Experts language model [hn, 359 points, 150 comments]
- from the Mixtral paper (arxiv.org/abs/2401.04088) "Surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic. For instance, at all layers, the distribution of ex [bsky, 4 points, 1 comments]
- The Mixtral of Experts paper is out. It emphasizes comparison to other models and is not very informative about pretraining. blueskAI #MLSky arxiv.org/abs/2401.04088 [bsky, 3 points, 0 comments]
- I think Google originally came up with MoE, and DeepSeek and Mixtral adopted it independently of each other. Eg looking at arxiv, the Mixtral report came out on 8 Jan 2024 (arxiv.org/abs/2401.04088), [bsky, 1 points, 1 comments]
- Mixtral 8x7b "...surpasses GPT-3.5 Turbo, Claude-2.1, Gemini Pro, and Llama 2 70B - chat model on human benchmarks. Both the base and instruct models are released under the Apache 2.0 license." Mixtra [bsky, 0 points, 0 comments]
Related