2025/03/13 by Alexei V. Tkachenko, Tkachenko, Alexei V. · 1 voice
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · Physics and Astronomy · #Data Analysis #Energy Load and Power Forecasting #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Physical sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Neurons and Cognition (q-bio.NC) #Statistical Mechanics (cond-mat.stat-mech) #Statistics and Probability (physics.data-an) #cond-mat.stat-mech #cs.LG #physics.data-an #q-bio.NC
paper · pdf · doi:10.48550/arxiv.2503.09980
openalex publication_date 2025/03/13 · arxiv published 2025/03/13 · openalex created_date 2025/10/19 · arxiv updated 2025/12/07 · openalex updated_date 2026/07/30
The rapid growth of deep neural networks (DNNs) has brought increasing attention to their energy use during training and inference. Here, we establish the thermodynamic bounds on energy consumption in quasi-static analog DNNs by mapping modern feedforward architectures onto a physical free-energy functional. This framework provides a direct statistical-mechanical interpretation of quasi-static DNNs. As a result, inference can proceed in a thermodynamically reversible manner, with vanishing minimal energy cost, in contrast to the Landauer limit that constrains digital hardware. Importantly, inference corresponds to relaxation to a unique free-energy minimum with Fmin=0, allowing all constraints to be satisfied without residual stress. By comparison, training overconstrains the system: simultaneous clamping of inputs and outputs generates stresses that propagate backward through the architecture, reproducing the rules of backpropagation. Parameter annealing then relaxes these stresses, providing a purely physical route to learning without an explicit loss function. We further derive a universal lower bound on training energy, E< 2NDkT, which scales with both the number of trainable parameters and the dataset size.