An overview of gradient descent optimization algorithms
2016/09/15 by Sebastian Ruder, Ruder, Sebastian · 3 voices · 4,802 citations
Computer Science · Engineering · Mathematics · #Algorithm #Artificial intelligence #Artificial neural network #Computer science #Descent (aeronautics) #Engineering #Face and Expression Recognition #Gradient descent #Mathematical optimization #Mathematics #Medical Image Segmentation Techniques #Optimization algorithm #Stochastic Gradient Optimization Techniques #cs.LG
paper · pdf · doi:10.48550/arxiv.1609.04747
published in arXiv (Cornell University) (Cornell University) · Added derivations of AdaMax and Nadam
openalex publication_date 2016/09/15 · arxiv created 2017/06/15 · arxiv updated 2017/06/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Abstract
Gradient descent optimization algorithms, while increasingly popular, are often used as black-box optimizers, as practical explanations of their strengths and weaknesses are hard to come by. This article aims to provide the reader with intuitions with regard to the behaviour of different algorithms that will allow her to put them to use. In the course of this overview, we look at different variants of gradient descent, summarize challenges, introduce the most common optimization algorithms, review architectures in a parallel and distributed setting, and investigate additional strategies for optimizing gradient descent.
Citations
Cited by
- Efficient optimisation of multi-parameter quantum control protocols for strongly-coupled systems
- Shesha: Opportunistic In-network Acceleration of Asynchronous Distributed Reinforcement Learning
- Material-agnostic temperature field prediction for metal additive manufacturing via a parametric PINN framework
- A Gradient Free Neural Network Framework Based on Universal Approximation Theorem
- GRAFFL: Gradient-free Federated Learning of a Bayesian Generative Model
- Evaluation of Neural Networks for Image Recognition Applications: Designing a 0-1 MILP Model of a CNN to create adversarials
- Temporal Autoencoder with U-Net Style Skip-Connections for Frame Prediction
- Training Neural Networks with an algorithm for piecewise linear functions
- MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information Maximization
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- On the Equivalence of Neural and Production Networks
- Universal Pansharpening Model
- Domain Decomposition of Large Neural Network Surrogate Models
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- Deep Learning Based Regression and Multi-class Models for Acute Oral Toxicity Prediction with Automatic Chemical Feature Extraction
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- OLR-WA: Online Weighted Average Linear Regression in Multivariate Data Streams
- An Additively Preconditioned Trust Region Strategy for Machine Learning
- Optimizing the Adversarial Perturbation with a Momentum-based Adaptive Matrix
- OLR-WAA: Adaptive and Drift-Resilient Online Regression with Dynamic Weighted Averaging
- Learning Dynamics in Memristor-Based Equilibrium Propagation
- Open Horizons: Evaluating Deep Models in the Wild
- Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
- Accelerating gradient descent and Adam via fractional gradients
- Learnability Window in Gated Recurrent Neural Networks
- Constraint-oriented biased quantum search for linear constrained combinatorial optimization problems
- Parameter efficient hybrid spiking-quantum convolutional neural network with surrogate gradient and quantum data-reupload
- Self-Organized Operational Neural Networks with Generative Neurons
- Temporal mixture ensemble models for intraday volume forecasting in cryptocurrency exchange markets
- Detecting Problem Statements in Peer Assessments
- Deep Filament Extraction for 3D Concrete Printing
- Physics-informed neural networks method in high-dimensional integrable systems
- Gradient Descent Algorithm Survey
- Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
- Multiple Learning for Regression in big data
- Deep Learning Framework for Enhanced Neutrino Reconstruction of Single-line Events in the ANTARES Telescope
- Quantum measurement tomography with mini-batch stochastic gradient descent
- Trimaximal Mixing Patterns Meet the First JUNO Result
- NTK-Guided Implicit Neural Teaching
- Towards Evolutionary Optimization Using the Ising Model
- Simultaneous Localization and 3D-Semi Dense Mapping for Micro Drones Using Monocular Camera and Inertial Sensors
- Deep Neural Network Based Resource Allocation for V2X Communications
- Deep Learning in Wide-field Surveys: Fast Analysis of Strong Lenses in Ground-based Cosmic Experiments
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural Networks
- Multi-Loss Sub-Ensembles for Accurate Classification with Uncertainty Estimation
- Neograd: Near-Ideal Gradient Descent
- Dynamic Temperature Scheduler for Knowledge Distillation
- FineSkiing: A Fine-grained Benchmark for Skiing Action Quality Assessment
- A Novel Stochastic Stratified Average Gradient Method: Convergence Rate and Its Complexity
- Fidelity sweet spot in transmon qubit rings under strong connectivity noise
- Hybrid Quantum-Classical Graph Convolutional Network
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- Variational noise mitigation in quantum circuits: the case of Quantum Fourier Transform
- ODE approximation for the Adam algorithm: General and overparametrized setting
- Unified Theory of Adaptive Variance Reduction
- DORAEMON: A Unified Library for Visual Object Modeling and Representation Learning at Scale
- Shellular Metamaterial Design via Compact Electric Potential Parametrization
- CLAX: Fast and Flexible Neural Click Models in JAX
- Constrained Performance Boosting Control for Nonlinear Systems
- Predictive Auxiliary Learning for Belief-based Multi-Agent Systems
- Extended dynamic mode decomposition with dictionary learning: a data-driven adaptive spectral decomposition of the Koopman operator
- Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
- On Higher-order Moments in Adam
- Recent Advances in Recurrent Neural Networks
- A Tutorial on Deep Learning for Music Information Retrieval
- From center to surrounding: An interactive learning framework for hyperspectral image classification
- Deep learning for pedestrians: backpropagation in CNNs
- Robust and Active Learning for Deep Neural Network Regression
- Modulating Regularization Frequency for Efficient Compression-Aware Model Training
- MMCoVaR: Multimodal COVID-19 Vaccine Focused Data Repository for Fake News Detection and a Baseline Architecture for Classification
- Adaptive unified contrastive learning with graph-based feature aggregator for imbalanced medical image classification
- Implementation of Parallel Simplified Swarm Optimization in CUDA
- Teaching Machine Learning to Software Engineers
- Compressing Heavy-Tailed Weight Matrices for Non-Vacuous Generalization Bounds
- Decision-based Universal Adversarial Attack
- Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
- ZORB: A Derivative-Free Backpropagation Algorithm for Neural Networks
- Clickbait Detection in Tweets Using Self-attentive Network
- Topology Optimization under Uncertainty using a Stochastic Gradient-based Approach
- Adaptive Online Learning with Momentum for Contingency-based Voltage Stability Assessment
- An improvement of the convergence proof of the ADAM-Optimizer
- Flexible Operator Embeddings via Deep Learning
- Seismic vulnerability modelling of building portfolios using artificial neural networks
- Dynamically Weighted Momentum with Adaptive Step Sizes for Efficient Deep Network Training
- Sparse Network Inversion for Key Instance Detection in Multiple Instance Learning
- Cross-Domain Adaptation for Animal Pose Estimation
- Variable Projected Augmented Lagrangian Methods for Generalized Lasso Problems
- Bayesian Optimization Algorithms for Accelerator Physics
- Robust Non-negative Proximal Gradient Algorithm for Inverse Problems
- Benchmarking VQE Configurations: Architectures, Initializations, and Optimizers for Silicon Ground State Energy
- Learning Reconfigurable Representations for Multimodal Federated Learning with Missing Data
- What Do We Understand About Convolutional Networks?
- Joint Design of Radar Waveform and Detector via End-to-end Learning with Waveform Constraints
- xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
- DeepLocalize: Fault Localization for Deep Neural Networks
- Semantic World Models
- QCFace: Image Quality Control for boosting Face Representation & Recognition
- An Encode-then-Decompose Approach to Unsupervised Time Series Anomaly Detection on Contaminated Training Data--Extended Version
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- Changes in United States Summer Temperatures Revealed by Explainable Neural Networks
- Influence Functions in Deep Learning Are Fragile
- Fighter: Unveiling the Graph Convolutional Nature of Transformers in Time Series Modeling
- Double Neural Counterfactual Regret Minimization
- PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
- Feature Selection and Regularization in Multi-Class Classification: An Empirical Study of One-vs-Rest Logistic Regression with Gradient Descent Optimization and L1 Sparsity Constraints
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee
- Using Kolmogorov-Smirnov Distance for Measuring Distribution Shift in Machine Learning
- Learning to Recognize Correctly Completed Procedure Steps in Egocentric Assembly Videos through Spatio-Temporal Modeling
- DeepTrust: Multi-Step Classification through Dissimilar Adversarial Representations for Robust Android Malware Detection
- Sinkformers: Transformers with Doubly Stochastic Attention
- Comparative Explanations via Counterfactual Reasoning in Recommendations
- A Blockchain‐based Cyber Attack Detection Scheme for Decentralized Internet of Things using Software‐Defined Network
- Learning Dynamics of VLM Finetuning
- Hyperlink Regression via Bregman Divergence
- Stochastic Probabilistic Programs
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- The impact of the additional features on the performance of regression analysis: a case study on regression analysis of music signal
- Analytical Survey of Learning with Low-Resource Data: From Analysis to Investigation
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- SSGD: A safe and efficient method of gradient descent
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models
- TreeNet: Layered Decision Ensembles
- QDeepGR4J: Quantile-based ensemble of deep learning and GR4J hybrid rainfall-runoff models for extreme flow prediction with uncertainty quantification
- Image Generation Based on Image Style Extraction
- CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
- AI-assisted Advanced Propellant Development for Electric Propulsion
- Neural Networks as Functional Classifiers
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Communication-Efficient and Interoperable Distributed Learning
- Effective continuous equations for adaptive SGD: a stochastic analysis view
- Learning-Based Collaborative Control for Bi-Manual Tactile-Reactive Grasping
- Supervised Contrastive Learning
- Sobolev acceleration for neural networks
- Improved Frequency Tracking with Adaptive Moments for Narrowband Interference Mitigation in GNSS
- Improve Single-Point Zeroth-Order Optimization Using High-Pass and Low-Pass Filters
- Why to "grow" and "harvest" deep learning models?
- A Performance Comparison of Loss Functions for Deep Face Recognition
- A Combined Data-driven and Physics-driven Method for Steady Heat Conduction Prediction using Deep Convolutional Neural Networks
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- MYSTIKO : : Cloud-Mediated, Private, Federated Gradient Descent
- A Variant of Gradient Descent Algorithm Based on Gradient Averaging
- Development of Deep Learning Optimizers: Approaches, Concepts, and Update Rules
- Neuromodulated Learning in Deep Neural Networks
- Neural Network Based Framework for Passive Intermodulation Cancellation in MIMO Systems
- The Root Finding Problem Revisited: Beyond the Robbins-Monro procedure
- Incorporating Visual Cortical Lateral Connection Properties into CNN: Recurrent Activation and Excitatory-Inhibitory Separation
- Exploring the Relationship between Brain Hemisphere States and Frequency Bands through Deep Learning Optimization Techniques
- Solving Zero-Sum Games through Alternating Projections
- An Analysis of Optimizer Choice on Energy Efficiency and Performance in Neural Network Training
- Inferring Soil Drydown Behaviour with Adaptive Bayesian Online Changepoint Analysis
- EmbeddedML: A New Optimized and Fast Machine Learning Library
- AdaSGD: Bridging the gap between SGD and Adam
- Privacy-preserving Federated Bayesian Learning of a Generative Model for Imbalanced Classification of Clinical Data
- A survey on deep learning approaches for breast cancer diagnosis
- On Faster Convergence of Scaled Sign Gradient Descent
- SugarcaneShuffleNet: A Very Fast, Lightweight Convolutional Neural Network for Diagnosis of 15 Sugarcane Leaf Diseases
- Survey: Machine Learning in Production Rendering
- A Capsule-unified Framework of Deep Neural Networks for Graphical Programming
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- A machine learning framework for data driven acceleration of computations of differential equations
- Heart Disease Prediction: A Comparative Study of Optimisers Performance in Deep Neural Networks
- Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
- Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection
- Finding the right scale of a network: Efficient identification of causal emergence through spectral clustering
- Uncertainty-Aware Neural Networks for Fuzzy Dark Matter Model Selection from \texorpdfstringx\rm HIxHI Measurements
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- 1D Convolutional Neural Networks and Applications: A Survey
- Neural Architecture Search via Bregman Iterations
- Event Detection and Classification for Long Range Sensing of Elephants Using Seismic Signal
- Application of Decision Rules for Handling Class Imbalance in Semantic Segmentation
- Learning Fair Scoring Functions: Bipartite Ranking under ROC-based Fairness Constraints
- A constrained optimization approach to nonlinear system identification through simulation error minimization
- Learning by training: emergent return-point memory from cyclically tuning disordered sphere packings
- CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition
- Deep Learning in Mobile and Wireless Networking: A Survey
- Theory Foundation of Physics-Enhanced Residual Learning
- Adaptive Heavy-Tailed Stochastic Gradient Descent
- Learning compact generalizable neural representations supporting perceptual grouping
- SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization
- GRADSTOP: Early Stopping of Gradient Descent via Posterior Sampling
- A Local Block Coordinate Descent Algorithm for the Convolutional Sparse Coding Model
- A Deep Learning Application for Psoriasis Detection
- Adaptive Hierarchical Hyper-gradient Descent
- From ANN to BNN: Inferring Reionization Parameters using Uncertainty-aware Emulators of 21-cm Summaries
- MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-Training
- Time-Scale Coupling Between States and Parameters in Recurrent Neural Networks
- Deep Learning based Estimation of Weaving Target Maneuvers
- Learning to See: Convolutional Neural Networks for the Analysis of Social Science Data
- Quaternion Collaborative Filtering for Recommendation
- Selectivity correction with online machine learning
- Stochastic Gradient Langevin Dynamics Algorithms with Adaptive Drifts
- Machine Biometrics -- Towards Identifying Machines in a Smart City Environment
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Topology Optimization under Microscale Uncertainty using Stochastic Gradients
- QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
- A Comprehensive Survey of Multilingual Neural Machine Translation
- GrCAN: Gradient Boost Convolutional Autoencoder with Neural Decision Forest
- Representation Understanding via Activation Maximization
- Using the Naive Bayes as a discriminative classifier
- Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
- IWA: Integrated Gradient based White-box Attacks for Fooling Deep Neural Networks
- Parameters for the best convergence of an optimization algorithm On-The-Fly
- DeePore: a deep learning workflow for rapid and comprehensive characterization of porous materials
- Neural Network Training via Stochastic Alternating Minimization with Trainable Step Sizes
- Unsupervised Pairwise Learning Optimization Framework for Cross-Corpus EEG-Based Emotion Recognition Based on Prototype Representation
- Training Deep Neural Networks via Branch-and-Bound
- DualNet: Locate Then Detect Effective Payload with Deep Attention Network
- Stress-Aware Resilient Neural Training
- Personalized Dynamic Treatment Regimes in Continuous Time: A Bayesian Approach for Optimizing Clinical Decisions with Timing
- Optimizing Convergence for Iterative Learning of ARIMA for Stationary Time Series
- Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
- WEEP: A Differentiable Nonconvex Sparse Regularizer via Weakly-Convex Envelope
- Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Fast Stochastic Variance Reduced Gradient Method with Momentum Acceleration for Machine Learning
- Few-shot transfer of tool-use skills using human demonstrations with proximity and tactile sensing
- Almost fault-tolerant quantum machine learning with drastic overhead reduction
- Machine Learning based Radio Environment Map Estimation for Indoor Visible Light Communication
- Improving Gradient Estimation in Evolutionary Strategies With Past Descent Directions
- Explicit Gradient Learning
- Can we learn gradients by Hamiltonian Neural Networks?
- How Many Factors Influence Minima in SGD?
- Deep learning empowers genomic selection of pest-resistant grapevine
- Machine learning in geo- and environmental sciences: From small to large scale
- Activated Gradients for Deep Neural Networks
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Demystifying BERT: Implications for Accelerator Design
- BigSurvSGD: Big Survival Data Analysis via Stochastic Gradient Descent
- Low Dimensional Landscape Hypothesis is True: DNNs can be Trained in Tiny Subspaces
- A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition
- A Pragmatic AI Approach to Creating Artistic Visual Variations by Neural Style Transfer
- Sparse-View Spectral CT Reconstruction Using Deep Learning
- 2nd-order Updates with 1st-order Complexity
- On Deep Learning for Radio Resource Management in A Non-stationary Radio Environment
- Expression Recognition Analysis in the Wild
- OptTyper: Probabilistic Type Inference by Optimising Logical and Natural Constraints
- Memory-Efficient Factorization Machines via Binarizing both Data and Model Coefficients
- Lily-like twist distribution in toroidal nematics
- A recurrent multi-scale approach to RBG-D Object Recognition
- Shot-Efficient ADAPT-VQE via Reused Pauli Measurements and Variance-Based Shot Allocation
- Layer Decomposition Learning Based on Gaussian Convolution Model and Residual Deblurring for Inverse Halftoning
- A Novel RL-assisted Deep Learning Framework for Task-informative Signals Selection and Classification for Spontaneous BCIs
- Single- to multi-fidelity history-dependent learning with uncertainty quantification and disentanglement: application to data-driven constitutive modeling
- Cluster Contrast for Unsupervised Visual Representation Learning
- When and how epochwise double descent happens
- MixML: A Unified Analysis of Weakly Consistent Parallel Learning
- SMG: A Shuffling Gradient-Based Method with Momentum
- Deep Neural Networks for Active Wave Breaking Classification
- Measure Transport with Kernel Stein Discrepancy
- Quantum Architecture Search via Deep Reinforcement Learning
- Deep Learning for Robust Motion Segmentation with Non-Static Cameras
- Platoon trajectories generation: A unidirectional interconnected LSTM-based car following model
- A Survey on Large-scale Machine Learning
- Avoiding local minima in Variational Quantum Algorithms with Neural Networks
- Privacy Preservation in Federated Learning: An insightful survey from the GDPR Perspective
- Heuristic Rank Selection with Progressively Searching Tensor Ring Network
- Predictive Process Model Monitoring using Recurrent Neural Networks
- AIRSENSE-TO-ACT: A Concept Paper for COVID-19 Countermeasures based on Artificial Intelligence algorithms and multi-sources Data Processing
- Privacy-Preserving Federated Learning for UAV-Enabled Networks: Learning-Based Joint Scheduling and Resource Management
- Learn distributed GAN with Temporary Discriminators
- Citadel: Protecting Data Privacy and Model Confidentiality for Collaborative Learning with SGX
- Pre-Training LLMs on a budget: A comparison of three optimizers
- Neighbor-view Enhanced Model for Vision and Language Navigation
- Small quantum computers and large classical data sets
- Feedback Control for Online Training of Neural Networks
- Adaptive Physics-Informed Neural Networks for Markov-Chain Monte Carlo
- A Qualitative Study of the Dynamic Behavior for Adaptive Gradient Algorithms
- A Federated Learning-based Lightweight Network with Zero Trust for UAV Authentication
- Learning scale-variant features for robust iris authentication with deep learning based ensemble framework
- Towards a NISQ Algorithm to Simulate Hermitian Matrix Exponentiation
- Variants of RMSProp and Adagrad with Logarithmic Regret Bounds
- DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- GraSSNet: Graph Soft Sensing Neural Networks
- Link Prediction for Temporally Consistent Networks
- Parameter Prediction for Unseen Deep Architectures
- Fast Evaluation of Low-Thrust Transfers via Deep Neural Networks
- Graph Drawing by Stochastic Gradient Descent
- Language Independent Single Document Image Super-Resolution using CNN for improved recognition
- A Convolutional Neural Network-based Approach to Field Reconstruction
- Machine learning for metal additive manufacturing: Predicting temperature and melt pool fluid dynamics using physics-informed neural networks
- Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology
- Length Learning for Planar Euclidean Curves
- Effectiveness of Optimization Algorithms in Deep Image Classification
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Learning to Coordinate in Multi-Agent Systems: A Coordinated Actor-Critic Algorithm and Finite-Time Guarantees
- Neuro-Symbolic Execution: The Feasibility of an Inductive Approach to Symbolic Execution
- Reliability-Adjusted Prioritized Experience Replay
- Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
- Sketch2code: Generating a website from a paper mockup
- Robust Training with Data Augmentation for Medical Imaging Classification
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- Learning the Wireless V2I Channels Using Deep Neural Networks
- Time-dependent density estimation using binary classifiers
- Towards Robust Learning to Optimize with Theoretical Guarantees
- Assessment of Optimizers and their Performance in Autosegmenting Lung Tumors
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity Understanding
- Gravilon: Applications of a New Gradient Descent Method to Machine Learning
- Towards Data-Driven Model-Free Safety-Critical Control
- Faithful-Newton Framework: Bridging Inner and Outer Solvers for Enhanced Optimization
- A Review of 1D Convolutional Neural Networks toward Unknown Substance Identification in Portable Raman Spectrometer
- Understand the Implication: Learning to Think for Pragmatic Understanding
- Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
- An approximate Riemann solver approach in Physics-Informed Neural Networks for hyperbolic conservation laws
- Quantum-Inspired Differentiable Integral Neural Networks (QIDINNs): A Feynman-Based Architecture for Continuous Learning Over Streaming Data
- Vehicle Re-identification Based on Dual Distance Center Loss
- Fiedler Regularization: Learning Neural Networks with Graph Sparsity
- Reliability and Performance Assessment of Federated Learning on Clinical Benchmark Data
- Denoising Programming Knowledge Tracing with a Code Graph-based Tuning Adaptor
- Proposition d'un modèle pour l'optimisation automatique de boucles dans le compilateur Tiramisu : cas d'optimisation de déroulage
- Learning Along the Arrow of Time: Hyperbolic Geometry for Backward-Compatible Representation Learning
- Accounting for data heterogeneity in integrative analysis and prediction methods: An application to Chronic Obstructive Pulmonary Disease
- Weighted Aggregating Stochastic Gradient Descent for Parallel Deep Learning
- Deep learning in radiology: an overview of the concepts and a survey of the state of the art
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- Buildings Classification using Very High Resolution Satellite Imagery
- Iterative Domain Optimization
- Hidden Markov Chains, Entropic Forward-Backward, and Part-Of-Speech Tagging
- Interaction Networks: Using a Reinforcement Learner to train other Machine Learning algorithms
- am-ELO: A Stable Framework for Arena-based LLM Evaluation
- Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
- Unit Commitment with Cost-Oriented Temporal Resolution
- nuts-flow/ml: data pre-processing for deep learning
- Proximal bundle algorithms for nonsmooth convex optimization via fast gradient smooth methods
- From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models
- Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination
- Design of Capacity-Approaching Low-Density Parity-Check Codes using Recurrent Neural Networks
- EWGN: Elastic Weight Generation and Context Switching in Deep Learning
- End-to-end Learning of Waveform Generation and Detection for Radar Systems
- SoftAdapt: Techniques for Adaptive Loss Weighting of Neural Networks with Multi-Part Loss Functions
- Application of Machine Learning in Wireless Networks: Key Techniques and Open Issues
- Small-Scale-Fading-Aware Resource Allocation in Wireless Federated Learning
- Algorithmic Complexities in Backpropagation and Tropical Neural Networks
- Optimal Density Functions for Weighted Convolution in Learning Models
- Learn to Allocate Resources in Vehicular Networks
- SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
- A Probabilistically Motivated Learning Rate Adaptation for Stochastic Optimization
- Comparing the Moore-Penrose Pseudoinverse and Gradient Descent for Solving Linear Regression Problems: A Performance Analysis
- Gradient Descent: The Ultimate Optimizer
- Resource Allocation Based on Deep Neural Networks for Cognitive Radio Networks
- Optimization for deep learning: theory and algorithms
- CrossNAS: A Cross-Layer Neural Architecture Search Framework for PIM Systems
- Melody Harmonization Using Orderless NADE, Chord Balancing, and Blocked Gibbs Sampling
- Standardized Non-Intrusive Reduced Order Modeling Using Different Regression Models With Application to Complex Flow Problems
- PADAM: Parallel averaged Adam reduces the error for stochastic optimization in scientific machine learning
- Some iterative algorithms on Riemannian manifolds and Banach spaces with good global convergence guarantee
- Multi-Tensor Network Representation for High-Order Tensor Completion
- Scalable and adaptive prediction bands with kernel sum-of-squares
- The Physics of Local Optimization in Complex Disordered Systems
- RTFN: A Robust Temporal Feature Network for Time Series Classification
- An Introduction to Neural Architecture Search for Convolutional Networks
- Certainty and Uncertainty Guided Active Domain Adaptation
- Joint Parameter-and-Bandwidth Allocation for Improving the Efficiency of Partitioned Edge Learning
- AOL: Adaptive Online Learning for Human Trajectory Prediction in Dynamic Video Scenes
- Cooperative Variance Estimation and Bayesian Neural Networks for Disentangling Aleatoric and Epistemic Uncertainties
- PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
- LLM Meeting Decision Trees on Tabular Data
- Certified Adversarial Defenses Meet Out-of-Distribution Corruptions: Benchmarking Robustness and Simple Baselines
- Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
- BGADAM: Boosting based Genetic-Evolutionary ADAM for Neural Network Optimization
- A Novel Deep Neural Network Based Approach for Sparse Code Multiple Access
- Machine Vision for Natural Gas Methane Emissions Detection Using an Infrared Camera
- Learn to Compress CSI and Allocate Resources in Vehicular Networks
- Deep localization of protein structures in fluorescence microscopy images
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
- Quantum Multi-view Kernel Learning with Local Information
- Improving adiabatic quantum factorization via chopped random-basis optimization
- NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
- Recoding latent sentence representations -- Dynamic gradient-based activation modification in RNNs
- A Caputo fractional derivative-based algorithm for optimization
- Evaluating Deep Learning in SystemML using Layer-wise Adaptive Rate Scaling(LARS) Optimizer
- SA-GD: Improved Gradient Descent Learning Strategy with Simulated Annealing
- On segmentation of pectoralis muscle in digital mammograms by means of deep learning
- KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches
- Nonparametric Teaching for Graph Property Learners
- Adversarial Resilience Learning - Towards Systemic Vulnerability Analysis for Large and Complex Systems
- A Physics-Inspired Optimizer: Velocity Regularized Adam
- Temporal Distance-aware Transition Augmentation for Offline Model-based Reinforcement Learning
- Stochastic variational inference improves quantification of multiple timepoint arterial spin labelling perfusion MRI
- Trust Region Value Optimization using Kalman Filtering
- DEAM: Adaptive Momentum with Discriminative Weight for Stochastic Optimization
- Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks
- Aerodynamic optimization of athlete posture using virtual skeleton methodology and computational fluid dynamics
- Humble your Overconfident Networks: Unlearning Overfitting via Sequential Monte Carlo Tempered Deep Ensembles
- Fast Approximation of Optimal Perturbed Long-Duration Impulsive Transfers via Deep Neural Networks
- A Latent Variational Framework for Stochastic Optimization
- Distributed Stochastic Algorithms for High-rate Streaming Principal Component Analysis
- Breast Cancer Classification in Deep Ultraviolet Fluorescence Images Using a Patch-Level Vision Transformer Framework
- ICE-Pruning: An Iterative Cost-Efficient Pruning Pipeline for Deep Neural Networks
- Injecting Knowledge Graphs into Large Language Models
- Particle Identification with Deep Neural Networks Across Collision Energies in Simulated Proton-Proton Collisions
- CSM-NN: Current Source Model Based Logic Circuit Simulation -- A Neural\n Network Approach
- Cross-Domain MLP and CNN Transfer Learning for Biological Signal Processing: EEG and EMG
- On the instability of embeddings for recommender systems: the case of Matrix Factorization
- Convergence Analysis of Gradient Descent Algorithms with Proportional Updates
- Neuron with Steady Response Leads to Better Generalization
- State-of-charge Estimation of a Li-ion Battery using Deep Learning and Stochastic Optimization
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Second-order Information in First-order Optimization Methods
- Forecasting day-ahead electricity prices in Europe: The importance of considering market integration
- Learning Compact Target-Oriented Feature Representations for Visual Tracking
- Machine learning in cardiovascular flows modeling: Predicting arterial blood pressure from non-invasive 4D flow MRI data using physics-informed neural networks
- Sharp higher order convergence rates for the Adam optimizer
- Performance Study of a Position-sensitive Plastic Scintillator Detector
- Experimental neuromorphic computing based on quantum memristor
- Efficient classical training of model-free quantum photonic reservoir
- Repaint: Improving the Generalization of Down-Stream Visual Tasks by Generating Multiple Instances of Training Examples
- Learning to Optimize by Differentiable Programming
- Weighted Empirical Risk Minimization: Sample Selection Bias Correction based on Importance Sampling
- Neural Controller for Incremental Stability of Unknown Continuous-time Systems
- The effects of Hessian eigenvalue spectral density type on the applicability of Hessian analysis to generalization capability assessment of neural networks
- Generalized Categorisation of Digital Pathology Whole Image Slides using Unsupervised Learning
- Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
- Spatio-Temporal Neural Network for Fitting and Forecasting COVID-19
- Neural network models for the anisotropic Reynolds stress tensor in turbulent channel flow
- Bringing Linearly Transformed Cosines to Anisotropic GGX
- Population-based Gradient Descent Weight Learning for Graph Coloring Problems
- Strong error analysis for the stochastic momentum optimizer
- A Skip-connected Multi-column Network for Isolated Handwritten Bangla Character and Digit recognition
- Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
- The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
- Gradient descent with momentum --- to accelerate or to super-accelerate?
- Automated Estimation of Construction Equipment Emission using Inertial Sensors and Machine Learning Models
- Knowledge-Preserving Incremental Social Event Detection via Heterogeneous GNNs
- Automated Architecture Design for Deep Neural Networks
- SiTGRU: Single-Tunnelled Gated Recurrent Unit for Abnormality Detection
- Self-Supervised Multisensor Change Detection
- AlphaGrad: Non-Linear Gradient Normalization Optimizer
- CGD: Modifying the Loss Landscape by Gradient Regularization
- Gradient-Optimized Fuzzy Classifier: A Benchmark Study Against State-of-the-Art Models
- VeLU: Variance-enhanced Learning Unit for Deep Neural Networks
- DNN based HRIRs Identification with a Continuously Rotating Speaker Array
- DMPCN: Dynamic Modulated Predictive Coding Network with Hybrid Feedback Representations
- Investigating performance of neural networks and gradient boosting models approximating microscopic traffic simulations in traffic optimization tasks
- Deep learning of thermodynamics-aware reduced-order models from data
- Task-based Loss Functions in Computer Vision: A Comprehensive Review
- Deep neural network for solving differential equations motivated by Legendre-Galerkin approximation
- Quasi-hyperbolic momentum and Adam for deep learning
- PC-DeepNet: A GNSS Positioning Error Minimization Framework Using Permutation-Invariant Deep Neural Network
- On Coresets for Regularized Loss Minimization
- Scared into Action: How Partisanship and Fear are Associated with Reactions to Public Health Directives
- Digitized-counterdiabatic quantum approximate optimization algorithm
- Supervised learning through physical changes in a mechanical system
- A multivariate water quality parameter prediction model using recurrent neural network
- MAXIMASK and MAXITRACK: Two new tools for identifying contaminants in astronomical images using convolutional neural networks
- Reverse Derivative Ascent: A Categorical Approach to Learning Boolean Circuits
- Automated Search for Configurations of Deep Neural Network Architectures
- Getting High: High Fidelity Simulation of High Granularity Calorimeters with High Speed
- Deep convolutions for in-depth automated rock typing
- Improving Online Forums Summarization via Hierarchical Unified Deep Neural Network
- A multi-path 2.5 dimensional convolutional neural network system for segmenting stroke lesions in brain MRI images
- Stochastic Gradient Descent in Non-Convex Problems: Asymptotic Convergence with Relaxed Step-Size via Stopping Time Methods
- Training and synchronizing oscillator networks with Equilibrium Propagation
- Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
- FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
- Quantum Phases Classification Using Quantum Machine Learning with SHAP-Driven Feature Selection
- DUE: A Deep Learning Framework and Library for Modeling Unknown Equations
- An Adaptive Weighted QITE-VQE Algorithm for Combinatorial Optimization Problems
- Enhancing knowledge retention for continual learning with domain-specific adapters and features gating
- Learning rate [wikipedia]
Discussions
Related