MLP-Mixer: An all-MLP Architecture for Vision
2021/05/04 by Ilya Tolstikhin, Neil Houlsby, Tolstikhin, Ilya +22 · 3 voices · 1,445 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Architecture #Artificial intelligence #Artificial neural network #Computer science #Convolutional neural network #Domain Adaptation and Few-Shot Learning #Engineering #Inference #Machine learning #Pattern recognition (psychology) #Perceptron #Regularization (linguistics) #Transformer #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2105.01601
published in arXiv (Cornell University) (Cornell University) · v2: Fixed parameter counts in Table 1. v3: Added results on JFT-3B in Figure 2(right); Added Section 3.4 on the input permutations. v4: Updated the x label in Figure 2(right)
openalex publication_date 2021/05/04 · arxiv created 2021/06/11 · arxiv updated 2021/06/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers.
Citations
Cited by
- A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function
- What Do Temporal Graph Learning Models Learn?
- SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
- A Controlled Study of Attention-Only Transformers
- EgoExoMoCap: Distributed Ego-Exo Human Motion Capture
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
- Fourier-mixed window attention for efficient and robust long sequence time-series forecasting
- Transformers without Normalization
- TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
- On the Opportunities and Risks of Foundation Models
- RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction
- Fast SAM2 with Text-Driven Token Pruning
- DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
- Efficient Vision Mamba for MRI Super-Resolution via Hybrid Selective Scanning
- AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- Adaptive Information Routing for Multimodal Time Series Forecasting
- Trajectory Densification and Depth from Perspective-based Blur
- Flash Multi-Head Feed-Forward Network
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
- SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
- Hierarchical geometric deep learning enables scalable analysis of molecular dynamics
- Training-Time Action Conditioning for Efficient Real-Time Chunking
- Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
- Network of Theseus (like the ship)
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- Global and Local Alignment Networks for Unpaired Image-to-Image Translation
- Deep Learning-Based Joint Uplink-Downlink CSI Acquisition for Next-Generation Upper Mid-Band Systems
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- LAP: Fast LAtent Diffusion Planner for Autonomous Driving
- S2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
- Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
- HieraMix: A Hierarchical MLP-Mixer for Large-Scale Traffic Forecasting
- HHFT: Hierarchical Heterogeneous Feature Transformer for Recommendation Systems
- Neural surrogates for designing gravitational wave detectors
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
- MDG: Masked Denoising Generation for Multi-Agent Behavior Modeling in Traffic Environments
- Parts-Mamba: Augmenting Joint Context with Part-Level Scanning for Occluded Human Skeleton
- Hyperspectral Image Classification using Spectral-Spatial Mixer Network
- Using pretrained graph neural networks with token mixers as geometric featurizers for conformational dynamics
- AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
- Heterogeneous Complementary Distillation
- AWEMixer: Adaptive Wavelet-Enhanced Mixer Network for Long-Term Time Series Forecasting
- Synthesizer: Rethinking Self-Attention in Transformer Models
- Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
- MetaFormer is Actually What You Need for Vision
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Global Filter Networks for Image Classification
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?
- Ultra-short-term wind power forecasting based on the MIFCformer model and a critical low wind speed region power revision strategy
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- The Brownian motion in the transformer model
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- An Image Patch is a Wave: Phase-Aware Vision MLP
- Symbol-Equivariant Recurrent Reasoning Models
- ResNet strikes back: An improved training procedure in timm
- S2-MLP: Spatial-Shift MLP Architecture for Vision
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- A Re-node Self-training Approach for Deep Graph-based Semi-supervised Classification on Multi-view Image Data
- Transformers from Compressed Representations
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- TAMI: Taming Heterogeneity in Temporal Interactions for Temporal Graph Link Prediction
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- Triangle Multiplication Is All You Need For Biomolecular Structure Representations
- MTmixAtt: Integrating Mixture-of-Experts with Multi-Mix Attention for Large-Scale Recommendation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- MetaFormer Is Actually What You Need for Vision
- Chimera: State Space Models Beyond Sequences
- Adversarial Attacks Leverage Interference Between Features in Superposition
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
- Rethinking the shape convention of an MLP
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- MixerGAN: An MLP-Based Architecture for Unpaired Image-to-Image Translation
- Automated Neural Architecture Design for Industrial Defect Detection
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- A Hierarchical Geometry-guided Transformer for Histological Subtyping of Primary Liver Cancer
- Graph-less Neural Networks: Teaching Old MLPs New Tricks via Distillation
- Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
- Exploring the Early Universe with Deep Learning
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- Efficient Self-supervised Vision Transformers for Representation Learning
- CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent
- A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- On the Bias Against Inductive Biases
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Choose a Transformer: Fourier or Galerkin
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- CoPAD : Multi-source Trajectory Fusion and Cooperative Trajectory Prediction with Anchor-oriented Decoder in V2X Scenarios
- Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers
- SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
- Efficient lattice field theory simulation using adaptive normalizing flow on a resistive memory-based neural differential equation solver
- PointMixer: MLP-Mixer for Point Cloud Understanding
- CE-RS-SBCIT A Novel Channel Enhanced Hybrid CNN Transformer with Residual, Spatial, and Boundary-Aware Learning for Brain Tumor MRI Analysis
- Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image Classifier
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- XCiT: Cross-Covariance Image Transformers
- Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers
- Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
- Revisiting the Calibration of Modern Neural Networks
- MRI-Based Brain Tumor Detection through an Explainable EfficientNetV2 and MLP-Mixer-Attention Architecture
- MSDA-HLGCformer-based context-aware fusion network for underwater organism detection
- KRAFT: A Knowledge Graph-Based Framework for Automated Map Conflation
- Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Efficiently Modeling Long Sequences with Structured State Spaces
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Enabling 6G Through Multi-Domain Channel Extrapolation: Opportunities and Challenges of Generative Artificial Intelligence
- Inductive Biases and Variable Creation in Self-Attention Mechanisms
- FW-GAN: Frequency-Driven Handwriting Synthesis with Wave-Modulated MLP Generator
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- ViTGAN: Training GANs with Vision Transformers
- HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
- Global Interaction Modelling in Vision Transformer via Super Tokens
- Prediction of Hospital Associated Infections During Continuous Hospital Stays
- Neuro-inspired Ensemble-to-Ensemble Communication Primitives for Sparse and Efficient ANNs
- WaveSeekerNet: accurate prediction of influenza A virus subtypes and host source using attention-based deep learning
- Graph Concept Bottleneck Models
- Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
- A Sobel-Gradient MLP Baseline for Handwritten Character Recognition
- CrypTen: Secure Multi-Party Computation Meets Machine Learning
- UniNet: Unified Architecture Search with Convolution, Transformer, and MLP
- DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
- MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
- An Interpretable Multi-Plane Fusion Framework With Kolmogorov-Arnold Network Guided Attention Enhancement for Alzheimer's Disease Diagnosis
- Channel-Independent Federated Traffic Prediction
- TF-MLPNet: Tiny Real-Time Neural Speech Separation
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- Cross-Architecture Distillation Made Simple with Redundancy Suppression
- Request-Only Optimization for Recommendation Systems
- Predicting the clinical citation count of biomedical papers using multilayer perceptron neural network
- Open-Set Recognition: a Good Closed-Set Classifier is All You Need?
- Improved Transformer for High-Resolution GANs
- See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
- Self-similarity Analysis in Deep Neural Networks
- Rethinking Token-Mixing MLP for MLP-based Vision Backbone
- Learning Pixel-adaptive Multi-layer Perceptrons for Real-time Image Enhancement
- Wasserstein Hypergraph Neural Network
- Assaying Out-Of-Distribution Generalization in Transfer Learning
- Auto-Compressing Networks
- HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- SpaRTAN: Spatial Reinforcement Token-based Aggregation Network for Visual Recognition
- InceptionMamba: An Efficient Hybrid Network with Large Band Convolution and Bottleneck Mamba
- Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal Interactions
- Run, Don't Walk: Chasing Higher FLOPS for Faster Neural Networks
- Interpretable inverse design of optical multilayer thin films based on extended neural adjoint and regression activation mapping
- MHRGait: Gait Recognition from Momentum Human Rig Pose
- Classification of COVID-19 cases from chest CT volumes using hybrid model of 3D CNN and 3D MLP-Mixer
- ScoreAdv: Score-based Targeted Generation of Natural Adversarial Examples via Diffusion Models
- FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval
- Decomposing the Time Series Forecasting Pipeline: A Modular Approach for Time Series Representation, Information Extraction, and Projection
- The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection
- FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation
- Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics
- Non-exchangeable Conformal Prediction for Temporal Graph Neural Networks
- evMLP: An Efficient Event-Driven MLP Architecture for Vision
- MFH: Marrying Frequency Domain with Handwritten Mathematical Expression Recognition
- Graph-Based Deep Learning for Component Segmentation of Maize Plants
- Transformer-Based Person Search with High-Frequency Augmentation and Multi-Wave Mixing
- Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting
- Automatic Depression Assessment using Machine Learning: A Comprehensive Survey
- MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators
- Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
- Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization
- Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
- Leveraging Lightweight Generators for Memory Efficient Continual Learning
- A standard transformer and attention with linear biases for molecular conformer generation
- Real-Time Execution of Action Chunking Flow Policies
- Discrete Representations Strengthen Vision Transformer Robustness
- Improving Black-Box Generative Attacks via Generator Semantic Consistency
- Pyramid Mixer: Multi-dimensional Multi-period Interest Modeling for Sequential Recommendation
- LegiGPT: Party Politics and Transport Policy with Large Language Model
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- Pay Attention to MLPs
- A Unified Anti-Jamming Design in Complex Environments Based on Cross-Modal Fusion and Intelligent Decision-Making
- AllTracker: Efficient Dense Point Tracking at High Resolution
- Dynamic Sparse Training of Diagonally Sparse Networks
- Don't Pay Attention
- Accurate and efficient zero-shot 6D pose estimation with frozen foundation models
- MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Quantum circuits as a game: A reinforcement learning agent for quantum compilation and its application to reconfigurable neutral atom arrays
- A novel scene coupling semantic mask network for remote sensing image segmentation
- LGM-Pose: A Lightweight Global Modeling Network for Real-time Human Pose Estimation
- MudiNet: Task-guided Disentangled Representation Learning for 5G Indoor Multipath-assisted Positioning
- Probabilistic Approach for Road-Users Detection
- Characterizing Structural Regularities of Labeled Data in Overparameterized Models
- GenReg: Deep Generative Method for Fast Point Cloud Registration
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Towards Transfer-Efficient Multi-modal Sequential Recommendation with State Space Duality
- Generative Perception of Shape and Material from Differential Motion
- Image Recognition with Online Lightweight Vision Transformer: A Survey
- A remark on a paper of Krotov and Hopfield [arXiv:2008.06996]
- On Improving Adversarial Transferability of Vision Transformers
- A New Deep-learning-Based Approach For mRNA Optimization: High Fidelity, Computation Efficiency, and Multiple Optimization Factors
- CrossLinear: Plug-and-Play Cross-Correlation Embedding for Time Series Forecasting with Exogenous Variables
- Knowledge Distillation for Reservoir-based Classifier: Human Activity Recognition
- Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization
- Bayesian Neural Scaling Law Extrapolation with Prior-Data Fitted Networks
- RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
- The quest for the GRAph Level autoEncoder (GRALE)
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Vision Transformers with Self-Distilled Registers
- TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
- Future Link Prediction Without Memory or Aggregation
- MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
- FAR: Function-preserving Attention Replacement for IMC-friendly Inference
- Attention Mechanisms in Computer Vision: A Survey
- CycleMLP: A MLP-like Architecture for Dense Prediction
- Beyond All-to-All: Causal-Aligned Transformer with Dynamic Structure Learning for Multivariate Time Series Forecasting
- Unified Cross-Modal Attention-Mixer Based Structural-Functional Connectomics Fusion for Neuropsychiatric Disorder Diagnosis
- SA-GD: Improved Gradient Descent Learning Strategy with Simulated Annealing
- InstanceBEV: Unifying Instance and BEV Representation for 3D Panoptic Segmentation
- Expert-Like Reparameterization of Heterogeneous Pyramid Receptive Fields in Efficient CNNs for Fair Medical Image Classification
- TSPulse: Tiny Pre-Trained Models with Disentangled Representations for Rapid Time-Series Analysis
- Scene-Adaptive Motion Planning with Explicit Mixture of Experts and Interaction-Oriented Optimization
- Exploring the Limits of Large Scale Pre-training
- From Preimage Search To Source-Grounded Feature Inversion
- BrainNetMLP: An Efficient and Effective Baseline for Functional Brain Network Classification
- Unified Sparse-Matrix Representations for Diverse Neural Architectures
- Mixer-Informer-Based Two-Stage Transfer Learning for Long-Sequence Load Forecasting in Newly Constructed Electric Vehicle Charging Stations
- UBoCo : Unsupervised Boundary Contrastive Learning for Generic Event Boundary Detection
- Do Vision Transformers See Like Convolutional Neural Networks?
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
- Polynomial Mixing for Efficient Self-supervised Speech Encoders
- Image Generation with a Sphere Encoder
- A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning
- Unsupervised 2D-3D lifting of non-rigid objects using local constraints
- Sequence Diffusion Model for Temporal Link Prediction in Continuous-Time Dynamic Graph
- QuantBench: Benchmarking AI Methods for Quantitative Investment
- A Spatially-Aware Multiple Instance Learning Framework for Digital Pathology
- Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
- GAMBAS: Generalised-Hilbert Mamba for Super-resolution of Paediatric Ultra-Low-Field MRI
- Development of a real-time posture feedback system with integrated human motion prediction for proactive intervention in work-related musculoskeletal disorders
- Seurat: From Moving Points to Depth
- Revisiting change detection methods for their application to serac fall time-lapse monitoring
- Cross-Modal Temporal Fusion for Financial Market Forecasting
- Filter-enhanced MLP is All You Need for Sequential Recommendation
- Window Token Concatenation for Efficient Visual Large Language Models
- EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection
- An Unsupervised Network Architecture Search Method for Solving Partial Differential Equations
- Unleashing Expert Opinion from Social Media for Stock Prediction
- GFT: Gradient Focal Transformer
- Exploring Synergistic Ensemble Learning: Uniting CNNs, MLP-Mixers, and Vision Transformers to Enhance Image Classification
- Dual Boost-Driven Graph-Level Clustering Network
- Find A Winning Sign: Sign Is All We Need to Win the Lottery
- A Survey of Quantum Transformers: Architectures, Challenges and Outlooks
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
- Multi-Granularity Vision Fastformer with Fusion Mechanism for Skin Lesion Segmentation
Discussions
Related