Convergence, design and training of continuous-time dropout as a random batch method
2025/10/15 by Antonio Álvarez-López, Álvarez-López, Antonio, Martı́n Hernández +1
Computer Science · Engineering · #35Q49 #37N35 #65C35 #65K10 #68T07 #FOS: Computer and information sciences #FOS: Mathematics #Innovative Microfluidic and Catalytic Techniques Innovation #Machine Learning (cs.LG) #Machine Learning and Data Classification #Optimization and Control (math.OC)
paper · pdf · doi:10.48550/arxiv.2510.13134
openalex publication_date 2025/10/15 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
Abstract
We study dropout regularization in continuous-time models through the lens of random-batch methods -- a family of stochastic sampling schemes originally devised to reduce the computational cost of interacting particle systems. We construct an unbiased, well-posed estimator that mimics dropout by sampling neuron batches over time intervals of length h. Trajectory-wise convergence is established with linear rate in h for the expected uniform error. At the distribution level, we establish stability for the associated continuity equation, with total-variation error of order h1/2 under mild moment assumptions. During training with fixed batch sampling across epochs, a Pontryagin-based adjoint analysis bounds deviations in the optimal cost and control, as well as in gradient-descent iterates. On the design side, we compare convergence rates for canonical batch sampling schemes, recover standard Bernoulli dropout as a special case, and derive a cost--accuracy trade-off yielding a closed-form optimal h. We then specialize to a single-layer neural ODE and validate the theory on classification and flow matching, observing the predicted rates, regularization effects, and favorable runtime and memory profiles.
Citations
- Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
- Random domain decomposition for parabolic PDEs on graphs
- Random Batch Methods for Discretized PDEs on Graphs
- Constructive approximate transport maps with normalizing flows
- Comprehensive Review of Neural Differential Equations for Time Series Analysis
- Mini-batch descent in semiflows
- Reduced variance random batch methods for nonlocal PDEs
- Control in finite and infinite dimension
- A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
- A Survey on Deep Neural Network Pruning: Taxonomy, Comparison, Analysis, and Recommendations
- Stability and Convergence of a Randomized Model Predictive Control Strategy
- A randomized operator splitting scheme inspired by stochastic optimization methods
- Implicit regularization of dropout
- Turnpike in optimal control of PDEs, ResNets, and beyond
- Sparse Flows: Pruning Continuous-depth Models
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Notes on the Behavior of MC Dropout
- Large-time asymptotics in deep learning
- STEER: Simple Temporal Regularization For Neural ODEs
- Language Models are Few-Shot Learners
- On the mean field limit of the Random Batch Method for interacting particle systems
- Convergence of Random Batch Method for interacting particles with disparate species and weights
- Normalizing Flows for Probabilistic Modeling and Inference
- Deep neural networks, generic universal interpolation, and controlled ODEs
- Neural Jump Stochastic Differential Equations
- Stochastic Training of Residual Networks: a Differential Equation\n Viewpoint
- Neural Ordinary Differential Equations
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Improved Regularization of Convolutional Neural Networks with Cutout
- Channel Pruning for Accelerating Very Deep Neural Networks
- Concrete Dropout
- Stable architectures for deep neural networks
- Structured Bayesian Pruning via Log-Normal Multiplicative Noise
- Pruning Filters for Efficient ConvNets
- Dynamic Network Surgery for Efficient DNNs
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Learning both Weights and Connections for Efficient Neural Networks
- Variational Dropout and the Local Reparameterization Trick
- Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
- Variational Inference with Normalizing Flows
- Efficient Object Localization Using Convolutional Networks
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Stochastic choice of basis functions in adaptive function approximation and the functional-link net
- On the product of semi-groups of operators
- A Generalization of Sampling Without Replacement from a Finite Universe
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related