vix.ing · top · new · best · stats · spec

Convergence, design and training of continuous-time dropout as a random batch method

2025/10/15 by Antonio Álvarez-López, Álvarez-López, Antonio, Martı́n Hernández +1
Computer Science · Engineering · #35Q49 #37N35 #65C35 #65K10 #68T07 #FOS: Computer and information sciences #FOS: Mathematics #Innovative Microfluidic and Catalytic Techniques Innovation #Machine Learning (cs.LG) #Machine Learning and Data Classification #Optimization and Control (math.OC)

paper · pdf · doi:10.48550/arxiv.2510.13134

openalex publication_date 2025/10/15 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28

Abstract

We study dropout regularization in continuous-time models through the lens of random-batch methods -- a family of stochastic sampling schemes originally devised to reduce the computational cost of interacting particle systems. We construct an unbiased, well-posed estimator that mimics dropout by sampling neuron batches over time intervals of length h. Trajectory-wise convergence is established with linear rate in h for the expected uniform error. At the distribution level, we establish stability for the associated continuity equation, with total-variation error of order h1/2 under mild moment assumptions. During training with fixed batch sampling across epochs, a Pontryagin-based adjoint analysis bounds deviations in the optimal cost and control, as well as in gradient-descent iterates. On the design side, we compare convergence rates for canonical batch sampling schemes, recover standard Bernoulli dropout as a special case, and derive a cost--accuracy trade-off yielding a closed-form optimal h. We then specialize to a single-layer neural ODE and validate the theory on classification and flow matching, observing the predicted rates, regularization effects, and favorable runtime and memory profiles.

Citations

Related