MLP-Mixer: An all-MLP Architecture for Vision
2021/05/04 by Ilya Tolstikhin, Neil Houlsby, Tolstikhin, Ilya +22 · 3 voices · 150 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Adversarial Robustness in Machine Learning
paper · pdf · doi:10.48550/arxiv.2105.01601
Abstract
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers.
Citations
Cited by
- A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function
- What Do Temporal Graph Learning Models Learn?
- SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
- A Controlled Study of Attention-Only Transformers
- EgoExoMoCap: Distributed Ego-Exo Human Motion Capture
- Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
- Fourier-mixed window attention for efficient and robust long sequence time-series forecasting
- Transformers without Normalization
- TLOB: A Novel Transformer Model with Dual Attention for Price Trend Prediction with Limit Order Book Data
- On the Opportunities and Risks of Foundation Models
- RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction
- Fast SAM2 with Text-Driven Token Pruning
- DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
- Efficient Vision Mamba for MRI Super-Resolution via Hybrid Selective Scanning
- AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
- VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- Adaptive Information Routing for Multimodal Time Series Forecasting
- Trajectory Densification and Depth from Perspective-based Blur
- Flash Multi-Head Feed-Forward Network
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement Learning
- SceneMixer: Exploring Convolutional Mixing Networks for Remote Sensing Scene Classification
- Hierarchical geometric deep learning enables scalable analysis of molecular dynamics
- Training-Time Action Conditioning for Efficient Real-Time Chunking
- Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
- Network of Theseus (like the ship)
- SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
- Global and Local Alignment Networks for Unpaired Image-to-Image Translation
- Deep Learning-Based Joint Uplink-Downlink CSI Acquisition for Next-Generation Upper Mid-Band Systems
- VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
- LAP: Fast LAtent Diffusion Planner for Autonomous Driving
- S2-KD: Semantic-Spectral Knowledge Distillation Spatiotemporal Forecasting
- Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
- HieraMix: A Hierarchical MLP-Mixer for Large-Scale Traffic Forecasting
- HHFT: Hierarchical Heterogeneous Feature Transformer for Recommendation Systems
- Neural surrogates for designing gravitational wave detectors
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
- MDG: Masked Denoising Generation for Multi-Agent Behavior Modeling in Traffic Environments
- Parts-Mamba: Augmenting Joint Context with Part-Level Scanning for Occluded Human Skeleton
- Hyperspectral Image Classification using Spectral-Spatial Mixer Network
- Using pretrained graph neural networks with token mixers as geometric featurizers for conformational dynamics
- AGGRNet: Selective Feature Extraction and Aggregation for Enhanced Medical Image Classification
- Heterogeneous Complementary Distillation
- AWEMixer: Adaptive Wavelet-Enhanced Mixer Network for Long-Term Time Series Forecasting
- Synthesizer: Rethinking Self-Attention in Transformer Models
- Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
- MetaFormer is Actually What You Need for Vision
- Hire-MLP: Vision MLP via Hierarchical Rearrangement
- Global Filter Networks for Image Classification
- ConvMLP: Hierarchical Convolutional MLPs for Vision
- Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?
- Ultra-short-term wind power forecasting based on the MIFCformer model and a critical low wind speed region power revision strategy
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- The Brownian motion in the transformer model
- A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP
- An Image Patch is a Wave: Phase-Aware Vision MLP
- Symbol-Equivariant Recurrent Reasoning Models
- ResNet strikes back: An improved training procedure in timm
- S2-MLP: Spatial-Shift MLP Architecture for Vision
- Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
- UHKD: A Unified Framework for Heterogeneous Knowledge Distillation via Frequency-Domain Representations
- A Re-node Self-training Approach for Deep Graph-based Semi-supervised Classification on Multi-view Image Data
- Transformers from Compressed Representations
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
- 3rd Place Solution to Large-scale Fine-grained Food Recognition
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- TAMI: Taming Heterogeneity in Temporal Interactions for Temporal Graph Link Prediction
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- Triangle Multiplication Is All You Need For Biomolecular Structure Representations
- MTmixAtt: Integrating Mixture-of-Experts with Multi-Mix Attention for Large-Scale Recommendation
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- MetaFormer Is Actually What You Need for Vision
- Chimera: State Space Models Beyond Sequences
- Adversarial Attacks Leverage Interference Between Features in Superposition
- Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
- Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior Modeling
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Flow-Matching Guided Deep Unfolding for Hyperspectral Image Reconstruction
- Rethinking the shape convention of an MLP
- One Pass Is Not Enough: Recursive Latent Refinement for Generative Models
- AS-MLP: An Axial Shifted MLP Architecture for Vision
- MixerGAN: An MLP-Based Architecture for Unpaired Image-to-Image Translation
- Automated Neural Architecture Design for Industrial Defect Detection
- Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
- A Hierarchical Geometry-guided Transformer for Histological Subtyping of Primary Liver Cancer
- Graph-less Neural Networks: Teaching Old MLPs New Tricks via Distillation
- Accelerating Dynamic Image Graph Construction on FPGA for Vision GNNs
- Learning to Parallel: Accelerating Diffusion Large Language Models via Learnable Parallel Decoding
- ProtoTS: Learning Hierarchical Prototypes for Explainable Time Series Forecasting
- Exploring the Early Universe with Deep Learning
- FlowDrive: moderated flow matching with data balancing for trajectory planning
- TF-Restormer: Complex Spectral Prediction for Speech Restoration
- Efficient Self-supervised Vision Transformers for Representation Learning
- CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent
- A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- On the Bias Against Inductive Biases
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Choose a Transformer: Fourier or Galerkin
- ViG-LRGC: Vision Graph Neural Networks with Learnable Reparameterized Graph Construction
- CoPAD : Multi-source Trajectory Fusion and Cooperative Trajectory Prediction with Anchor-oriented Decoder in V2X Scenarios
- Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers
- SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compression in Speaker Verification
- Efficient lattice field theory simulation using adaptive normalizing flow on a resistive memory-based neural differential equation solver
- PointMixer: MLP-Mixer for Point Cloud Understanding
- CE-RS-SBCIT A Novel Channel Enhanced Hybrid CNN Transformer with Residual, Spatial, and Boundary-Aware Learning for Brain Tumor MRI Analysis
- Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image Classifier
- NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability
- XCiT: Cross-Covariance Image Transformers
- Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers
- Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
- Revisiting the Calibration of Modern Neural Networks
- MRI-Based Brain Tumor Detection through an Explainable EfficientNetV2 and MLP-Mixer-Attention Architecture
- MSDA-HLGCformer-based context-aware fusion network for underwater organism detection
- KRAFT: A Knowledge Graph-Based Framework for Automated Map Conflation
- Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Efficiently Modeling Long Sequences with Structured State Spaces
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Enabling 6G Through Multi-Domain Channel Extrapolation: Opportunities and Challenges of Generative Artificial Intelligence
- Inductive Biases and Variable Creation in Self-Attention Mechanisms
- FW-GAN: Frequency-Driven Handwriting Synthesis with Wave-Modulated MLP Generator
- Dual-Model Weight Selection and Self-Knowledge Distillation for Medical Image Classification
- Image Quality Assessment for Machines: Paradigm, Large-scale Database, and Models
- ViTGAN: Training GANs with Vision Transformers
- HLLM-Creator: Hierarchical LLM-based Personalized Creative Generation
- Global Interaction Modelling in Vision Transformer via Super Tokens
- Prediction of Hospital Associated Infections During Continuous Hospital Stays
- Neuro-inspired Ensemble-to-Ensemble Communication Primitives for Sparse and Efficient ANNs
- Graph Concept Bottleneck Models
- Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
- A Sobel-Gradient MLP Baseline for Handwritten Character Recognition
- CrypTen: Secure Multi-Party Computation Meets Machine Learning
- UniNet: Unified Architecture Search with Convolution, Transformer, and MLP
- DETACH: Cross-domain Learning for Long-Horizon Tasks via Mixture of Disentangled Experts
- MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
- An Interpretable Multi-Plane Fusion Framework With Kolmogorov-Arnold Network Guided Attention Enhancement for Alzheimer's Disease Diagnosis
- Channel-Independent Federated Traffic Prediction
- TF-MLPNet: Tiny Real-Time Neural Speech Separation
- Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
- Cross-Architecture Distillation Made Simple with Redundancy Suppression
- Request-Only Optimization for Recommendation Systems
- Predicting the clinical citation count of biomedical papers using multilayer perceptron neural network
- Open-Set Recognition: a Good Closed-Set Classifier is All You Need?
- Improved Transformer for High-Resolution GANs
Discussions
Related