DARTS: Differentiable Architecture Search
2018/06/24 by Hanxiao Liu, Karen Simonyan, Liu, Hanxiao +3 · 255 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Immunology and Microbiology · #Machine Learning in Bioinformatics #Multimodal Machine Learning Applications #interferon and immune responses
paper · pdf · doi:10.48550/arxiv.1806.09055
Abstract
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner. Unlike conventional approaches of applying evolution or reinforcement learning over a discrete and non-differentiable search space, our method is based on the continuous relaxation of the architecture representation, allowing efficient search of the architecture using gradient descent. Extensive experiments on CIFAR-10, ImageNet, Penn Treebank and WikiText-2 show that our algorithm excels in discovering high-performance convolutional architectures for image classification and recurrent architectures for language modeling, while being orders of magnitude faster than state-of-the-art non-differentiable techniques. Our implementation has been made publicly available to facilitate further research on efficient architecture search algorithms.
Cited by
- GLUE: Gradient-free Learning to Unify Experts
- Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation
- Discovering Sparse Recovery Algorithms Using Neural Architecture Search
- Surrogate Neural Architecture Codesign Package (SNAC-Pack)
- Exploring Deep-to-Shallow Transformable Neural Networks for Intelligent Embedded Systems
- Optimized Architectures for Kolmogorov-Arnold Networks
- Learning to Evolve with Convergence Guarantee via Neural Unrolling
- HPM-KD: Hierarchical Progressive Multi-Teacher Framework for Knowledge Distillation and Efficient Model Compression
- Quantum Decision Transformers (QDT): Synergistic Entanglement and Interference for Offline Reinforcement Learning
- Voxify3D: Pixel Art Meets Volumetric Rendering
- LLM-Driven Composite Neural Architecture Search for Multi-Source RL State Encoding
- RevoNAD: Reflective Evolutionary Exploration for Neural Architecture Design
- Network of Theseus (like the ship)
- Neural Architecture Search of Time-to-First-Spike-Coded Spiking Neural Networks for Efficient Eye-based Emotion Recognition
- Intrusion Detection on Resource-Constrained IoT Devices with Hardware-Aware ML and DL
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- NNGPT: Rethinking AutoML with Large Language Models
- BaLeNAS: Differentiable Architecture Search via the Bayesian Learning Rule
- Gradient-Based Join Ordering
- Rapid Elastic Architecture Search under Specialized Classes and Resource Constraints
- Search to aggregate neighborhood for graph neural network
- A Survey on Green Deep Learning
- Speedy Performance Estimation for Neural Architecture Search
- A Comprehensive Study of Deep Video Action Recognition
- DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image Compression
- Balancing Average and Worst-case Accuracy in Multitask Learning
- Integrating Large Circular Kernels into CNNs through Neural Architecture Search
- Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman Kernels
- AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning
- MoGA: Searching Beyond MobileNetV3
- Solving bilevel optimization via sequential minimax optimization
- Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training
- EH-DNAS: End-to-End Hardware-aware Differentiable Neural Architecture Search
- One-Shot Knowledge Transfer for Scalable Person Re-Identification
- ArchPilot: A Proxy-Guided Multi-Agent Approach for Machine Learning Engineering
- A Review of Bilevel Optimization: Methods, Emerging Applications, and Recent Advancements
- Parametric Contrastive Learning
- SNAS: Stochastic Neural Architecture Search
- ShapleyPipe: Hierarchical Shapley Search for Data Preparation Pipeline Construction
- PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
- Sharpness-aware Quantization for Deep Neural Networks
- Revisiting Neural Architecture Search
- NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems
- Reconsidering CO2 emissions from Computer Vision
- AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
- Malleable 2.5D Convolution: Learning Receptive Fields along the Depth-axis for RGB-D Scene Parsing
- Dynamic Convolution: Attention over Convolution Kernels
- Fitness Landscape Footprint: A Framework to Compare Neural Architecture\n Search Problems
- Learning to Rank Learning Curves
- BN-NAS: Neural Architecture Search with Batch Normalization
- RankNAS: Efficient Neural Architecture Search by Pairwise Ranking
- Bayesian Bits: Unifying Quantization and Pruning
- Neural Architecture Search for Traffic Prediction: A Survey of Methods, Challenges, and Future Directions
- Batch Group Normalization
- Deep learning for pedestrians: backpropagation in CNNs
- NetAdaptV2: Efficient Neural Architecture Search with Fast Super-Network Training and Architecture Optimization
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Learning Architectures for Binary Networks
- BRP-NAS: Prediction-based NAS using GCNs
- Circumventing Outliers of AutoAugment with Knowledge Distillation
- Accelerate CNNs from Three Dimensions: A Comprehensive Pruning Framework
- SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions
- Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift
- MixStyle Neural Networks for Domain Generalization and Adaptation
- Domain-aware Visual Bias Eliminating for Generalized Zero-Shot Learning
- LETI: Latency Estimation Tool and Investigation of Neural Networks inference on Mobile GPU
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Surrogate Assisted Diversity Estimation in Neural Ensemble Search
- AutoFIS: Automatic Feature Interaction Selection in Factorization Models for Click-Through Rate Prediction
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Neural Ensemble Search for Uncertainty Estimation and Dataset Shift
- Deep Active Learning with a Neural Architecture Search
- CurveLane-NAS: Unifying Lane-Sensitive Architecture Search and Adaptive Point Blending
- Feature Gradients: Scalable Feature Selection via Discrete Relaxation
- LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search
- Densely Connected Search Space for More Flexible Neural Architecture Search
- Searching for A Robust Neural Architecture in Four GPU Hours
- Balancing Accuracy and Latency in Multipath Neural Networks
- Stabilizing DARTS with Amended Gradient Estimation on Architectural Parameters
- Semi-Supervised Learning with Meta-Gradient
- MixPath: A Unified Approach for One-shot Neural Architecture Search
- Coordinate Attention for Efficient Mobile Network Design
- Rethinking the Number of Channels for the Convolutional Neural Network
- A Comprehensive Survey on Hardware-Aware Neural Architecture Search
- Optimization-Inspired Learning with Architecture Augmentations and Control Mechanisms for Low-Level Vision
- Model Complexity of Deep Learning: A Survey
- Deeper Insights into Weight Sharing in Neural Architecture Search
- The Stretto Execution Engine for LLM-Augmented Data Systems
- Involution: Inverting the Inherence of Convolution for Visual Recognition
- Automated Machine Learning: State-of-The-Art and Open Challenges
- A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
- Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
- The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
- Lower Bounds and Accelerated Algorithms for Bilevel Optimization
- SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks
- Task-Adaptive Neural Network Search with Meta-Contrastive Learning
- Distributionally Robust Feature Selection
- TransTailor: Pruning the Pre-trained Model for Improved Transfer Learning
- Learning to Flow from Generative Pretext Tasks for Neural Architecture Encoding
- Self-Assembling Modular Networks for Interpretable Multi-Hop Reasoning
- Self-Evidencing Through Hierarchical Gradient Decomposition: A Dissipative System That Maintains Non-Equilibrium Steady-State by Minimizing Variational Free Energy
- Computational Budget Should Be Considered in Data Selection
- FOX-NAS: Fast, On-device and Explainable Neural Architecture Search
- All You Need is a Few Shifts: Designing Efficient Convolutional Neural Networks for Image Classification
- DARTS-GT: Differentiable Architecture Search for Graph Transformers with Quantifiable Instance-Specific Interpretability Analysis
- Spiking Neural Network Architecture Search: A Survey
- Approximate Bilevel Graph Structure Learning for Histopathology Image Classification
- BinaryBERT: Pushing the Limit of BERT Quantization
- Are Labels Necessary for Neural Architecture Search?
- Optimally Deep Networks -- Adapting Model Depth to Datasets for Superior Efficiency
- Cascade Bagging for Accuracy Prediction with Few Training Samples
- Stochastic Training is Not Necessary for Generalization
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
- MAT-Agent: Adaptive Multi-Agent Training Optimization
- Zen-NAS: A Zero-Shot NAS for High-Performance Deep Image Recognition
- OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
- Small-Group Learning, with Application to Neural Architecture Search
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse Coding
- On the Relationship Between the Choice of Representation and In-Context Learning
- Trustless parallel local search for effective distributed algorithm discovery
- Where to Begin: Efficient Pretraining via Subnetwork Selection and Distillation
- Universal Neural Architecture Space: Covering ConvNets, Transformers and Everything in Between
- ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures
- FT-MDT: Extracting Decision Trees from Medical Texts via a Novel Low-rank Adaptation Method
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- On The Statistical Limits of Self-Improving Agents
- HW-NAS-Bench:Hardware-Aware Neural Architecture Search Benchmark
- NSGANetV2: Evolutionary Multi-Objective Surrogate-Assisted Neural Architecture Search
- Deep Learning for Image Super-Resolution: A Survey
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- LLM-NAS: LLM-driven Hardware-Aware Neural Architecture Search
- Composer: A Search Framework for Hybrid Neural Architecture Design
- HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers
- CoLLM-NAS: Collaborative Large Language Models for Efficient Knowledge-Guided Neural Architecture Search
- From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
- CIMNAS: A Joint Framework for Compute-In-Memory-Aware Neural Architecture Search
- Computation Reallocation for Object Detection
- TusoAI: Agentic Optimization for Scientific Methods
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- Graph-guided Architecture Search for Real-time Semantic Segmentation
- Neural Architecture Search in Embedding Space
- CP-NAS: Child-Parent Neural Architecture Search for Binary Neural Networks
- LV-BERT: Exploiting Layer Variety for BERT
- SNAC-Pack 2.0: Scaled-Out Surrogate Neural Architecture Codesign
- Winning the Lottery with Continuous Sparsification
- Joint Search of Data Augmentation Policies and Network Architectures
- Separable Layers Enable Structured Efficient Linear Substitutions
- DetNAS: Backbone Search for Object Detection
- Learning Graph Representation of Person-specific Cognitive Processes from Audio-visual Behaviours for Automatic Personality Recognition
- FIVES: Feature Interaction Via Edge Search for Large-Scale Tabular Data
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
- Analyzing Neural Networks Based on Random Graphs
- Auto Seg-Loss: Searching Metric Surrogates for Semantic Segmentation
- DIFER: Differentiable Automated Feature Engineering
- Shared-Weights Extender and Gradient Voting for Neural Network Expansion
- Learnable Parameter Similarity
- Learning Index Selection with Structured Action Spaces
- Learning Dynamic Routing for Semantic Segmentation
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
- Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning
- Geometry-Aware Gradient Algorithms for Neural Architecture Search
- Neural Architecture Search Algorithms for Quantum Autoencoders
- Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
- Stochastic Bilevel Optimization with Heavy-Tailed Noise
- What and Where: Learn to Plug Adapters via NAS for Multi-Domain Learning
- FADNet++: Real-Time and Accurate Disparity Estimation with Configurable Networks
- A Survey on Neural Architecture Search
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks
- Learning Graph Convolutional Network for Skeleton-based Human Action Recognition by Neural Searching
- GLiT: Neural Architecture Search for Global and Local Image Transformer
- AEFS: Adaptive Early Feature Selection for Deep Recommender Systems
- On Neural Architecture Search for Resource-Constrained Hardware Platforms
- Rethinking Channel Dimensions for Efficient Model Design
- Hierarchical Neural Architecture Search via Operator Clustering
- Contextual Learning for Anomaly Detection in Tabular Data
- Network Pruning via Transformable Architecture Search
- Trellis Networks for Sequence Modeling
- Variational Depth Search in ResNets
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- DKM: Differentiable K-Means Clustering Layer for Neural Network Compression
- FLASH: Fast Neural Architecture Search with Hardware Optimization
- A Value-Function-based Interior-point Method for Non-convex Bi-level Optimization
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural Networks
- Intra-Ensemble in Neural Networks
- OptiProxy-NAS: Optimization Proxy based End-to-End Neural Architecture Search
- LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding
- Randomized Stochastic Variance-Reduced Methods for Multi-Task Stochastic Bilevel Optimization
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- Understanding Architectures Learnt by Cell-based Neural Architecture Search
- Accelerating Neural Architecture Search via Proxy Data
- Neural Architecture Search via Bregman Iterations
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- Splitting Steepest Descent for Growing Neural Architectures
- Customized Graph Embedding: Tailoring Embedding Vectors to different Applications
- Online Meta-Learning for Multi-Source and Semi-Supervised Domain Adaptation
- HourNAS: Extremely Fast Neural Architecture Search Through an Hourglass Lens
- A Continuous Encoding-Based Representation for Efficient Multi-Fidelity Multi-Objective Neural Architecture Search
- How Powerful are Performance Predictors in Neural Architecture Search?
- SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
- Towards Accurate and Compact Architectures via Neural Architecture Transformer
- Gradient-based Hyperparameter Optimization Over Long Horizons
- Ranking architectures using meta-learning
- Exploring Randomly Wired Neural Networks for Image Recognition
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- Standing on the Shoulders of Giants: Hardware and Neural Architecture Co-Search with Hot Start
- MANAS: Multi-Agent Neural Architecture Search
- Discretization-Aware Architecture Search
- Learning Versatile Neural Architectures by Propagating Network Codes
- SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection
- Quantum Long Short-term Memory with Differentiable Architecture Search
- Cross-Layer Design of Vector-Symbolic Computing: Bridging Cognition and Brain-Inspired Hardware Acceleration
- Differentiable Sparsification for Deep Neural Networks
- Stabilizing Differentiable Architecture Search via Perturbation-based Regularization
- MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
- On Constraint Qualifications for MPECs with Applications to Bilevel Hyperparameter Optimization for Machine Learning
- Neural Architecture Search on ImageNet in Four GPU Hours: A Theoretically Inspired Perspective
- Evolutionary-Neural Hybrid Agents for Architecture Search
- Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
- Dextr: Zero-Shot Neural Architecture Search with Singular Value Decomposition and Extrinsic Curvature
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- Language Models with Transformers
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- CrypTen: Secure Multi-Party Computation Meets Machine Learning
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation
- RegimeNAS: Regime-Aware Differentiable Architecture Search With Theoretical Guarantees for Financial Trading
- SOTERIA: In Search of Efficient Neural Networks for Private Inference
- Zero-Cost Operation Scoring in Differentiable Architecture Search
- Improving Neural Language Models by Segmenting, Attending, and Predicting the Future
- DHP: Differentiable Meta Pruning via HyperNetworks
- The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures
- MixSearch: Searching for Domain Generalized Medical Image Segmentation Architectures
- Learn to Explore: Meta NAS via Bayesian Optimization Guided Graph Generation
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- Reinforced Evolutionary Neural Architecture Search
- Understanding Neural Architecture Search Techniques
- Searching Efficient Model-guided Deep Network for Image Denoising
- Empowering Time Series Forecasting with LLM-Agents
- FPG-NAS: FLOPs-Aware Gated Differentiable Neural Architecture Search for Efficient 6DoF Pose Estimation
- Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
- Learned Low Precision Graph Neural Networks
- DAAS: Differentiable Architecture and Augmentation Policy Search
- Fast Neural Network Adaptation via Parameter Remapping and Architecture Search
- Large-Scale Model Enabled Semantic Communication Based on Robust Knowledge Distillation
- CADDA: Class-wise Automatic Differentiable Data Augmentation for EEG Signals
- Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
- Raw Differentiable Architecture Search for Speech Deepfake and Spoofing Detection
- PhaseNAS: Language-Model Driven Architecture Search with Dynamic Phase Adaptation
- ArchRepair: Block-Level Architecture-Oriented Repairing for Deep Neural Networks
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions
- Learning to Branch for Multi-Task Learning
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- ASNN: Learning to Suggest Neural Architectures from Performance Distributions
- Debiasing a First-order Heuristic for Approximate Bi-level Optimization
- Efficient Structured Pruning and Architecture Searching for Group Convolution
Related