A Stochastic Approximation Method
1951/09/01 by Herbert Robbins, Sutton Monro · 9,641 citations
Decision Sciences · Mathematics · #Alpha (finance) #Applied mathematics #Combinatorics #Computer science #Constant (computer programming) #Expected value #Function (biology) #Geometry #Mathematics #Monotone polygon #Optimal Experimental Design Methods #Statistics #Value (mathematics)
paper · pdf · doi:10.1214/aoms/1177729586
published in The Annals of Mathematical Statistics 22(3), 400-407 (Institute of Mathematical Statistics)
openalex publication_date 1951/09/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Let M(x) denote the expected value at level x of the response to a certain experiment. M(x) is assumed to be a monotone function of x but is unknown to the experimenter, and it is desired to find the solution x = θ of the equation M(x) = α, where α is a given constant. We give a method for making successive experiments at levels x1,x2,⋯ in such a way that xn will tend to θ in probability.
Cited by
- Semimartingale Stochastic Approximation Procedures and Recursive Estimation
- Stochastic Estimation of the Maximum of a Regression Function
- PyPose: A Library for Robot Learning with Physics-based Optimization
- The Modern Mathematics of Deep Learning
- Stochastic optimization on matrices and a graphon McKean–Vlasov limit
- PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM
- Advances in Asynchronous Parallel and Distributed Optimization
- An information field theory approach to Bayesian state and parameter estimation in dynamical systems
- Accurate computation of quantum excited states with neural networks
- A Nonparametric Approach to Pricing and Hedging Derivative Securities Via Learning Networks
- Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
- Sample size selection in optimization methods for machine learning
- Analysis of biased stochastic gradient descent using sequential semidefinite programs
- Green behavior propagation analysis based on statistical theory and intelligent algorithm in data-driven environment
- Distributed stochastic gradient tracking methods
- Convergence of Random Batch Method with replacement for interacting particle systems
- Fitting the psychometric function
- A Lyapunov Theory for Finite-Sample Guarantees of Markovian Stochastic Approximation
- Pricing Under Uncertainty in Multi-Interval Real-Time Markets
- Calibrated Bayesian Nonparametric Tolerance Intervals
- Proximity and the Evolution of Collaboration Networks: Evidence from Research and Development Projects within the Global Navigation Satellite System (GNSS) Industry
- A Deep Learning Algorithm for High-Dimensional Exploratory Item Factor Analysis
- GREEN: A lightweight architecture using learnable wavelets and Riemannian geometry for biomarker exploration with EEG signals
- Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape
- Efficient Risk Estimation for the Credit Valuation Adjustment
- In-Run Data Shapley for Adam Optimizer
- Finite-horizon quantile martingale posteriors: raw-urn laws and matrix-gain regression
- Frictional Q-Learning
- Equilibrium Computation in Extensive-Form Games with Stochastic Action Sets
- Computing Equilibria in Games with Stochastic Action Sets
- Theoretical guarantees for stochastic gradient sampling methods via Gaussian convolution inequalities
- DiLLSUE: a differentiable GPU solver for link-based logit stochastic user equilibrium
- Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
- From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
- Distillation of atomistic foundation models across architectures and chemical domains
- Vulnerability Detection via Multiple-Graph-Based Code Representation
- Fast Inference for Intractable Likelihood Problems using Variational Bayes
- Learning Equilibrium Play for Stochastic Parallel Gaussian Interference Channels
- A Federated Data-Driven Evolutionary Algorithm for Expensive Multi/Many-objective Optimization
- Numerical methods in large-scale optimization: inexact oracle and primal-dual analysis
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Distributed Weight Consolidation: A Brain Segmentation Case Study
- Geometric Insights into the Convergence of Nonlinear TD Learning
- Adaptive Gradient Method with Resilience and Momentum
- A Cross Entropy based Optimization Algorithm with Global Convergence Guarantees
- Reinforcement Learning for Matrix Computations: PageRank as an Example
- The Statistics of Streaming Sparse Regression
- Optimal Transport Based Distributionally Robust Optimization: Structural Properties and Iterative Schemes
- An averaged projected Robbins-Monro algorithm for estimating the parameters of a truncated spherical distribution
- VR-SGD: A Simple Stochastic Variance Reduction Method for Machine Learning
- Learning Machines Implemented on Non-Deterministic Hardware
- Generative Max-Mahalanobis Classifiers for Image Classification, Generation and More
- Variational Dropout and the Local Reparameterization Trick
- SI-ADMM: A Stochastic Inexact ADMM Framework for Stochastic Convex Programs
- ADADELTA: An Adaptive Learning Rate Method
- Why (and When and How) Contrastive Divergence Works
- DTN: A Learning Rate Scheme with Convergence Rate of O(1/t) for SGD
- An overview of gradient descent optimization algorithms
- Bend to Mend: Toward Trustworthy Variational Bayes with Valid Uncertainty Quantification
- Semantics, Representations and Grammars for Deep Learning
- Computing Pure-Strategy Nash Equilibria in a Two-Party Policy Competition: Existence and Algorithmic Approaches
- Painless step size adaptation for SGD
- Deep Neural Networks - A Brief History
- Training Neural Networks with an algorithm for piecewise linear functions
- Analysis of the Stochastic Alternating Least Squares Method for the Decomposition of Random Tensors
- A trust-region method for derivative-free nonlinear constrained stochastic optimization
- Stochastic Variance Reduction for Nonconvex Optimization
- Multilevel Monte Carlo Variational Inference
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
- Scale Weight Decay and Train Better
- Robust Training in High Dimensions via Block Coordinate Geometric Median Descent
- Adjusted Shuffling SARAH: Advancing Complexity Analysis via Dynamic Gradient Weighting
- An invitation to adaptive Markov chain Monte Carlo convergence theory
- tf.data: A Machine Learning Data Processing Framework
- Collapsed Variational Bayes Inference of Infinite Relational Model
- Sample Efficient Policy Gradient Methods with Recursive Variance Reduction
- A Projected Stochastic Gradient Method for Finite-Sum Problems with Linear Equality Constraints
- Efficient Online Conformal Selection with Limited Feedback
- Analyzing Process Data from Computer-Based Assessments: A Tutorial on Preprocessing, Feature Extraction, and Model-Based Inference
- Data relativistic uncertainty framework for low-illumination anime scenery image enhancement
- Investigating methods to solve large windfarm optimization problems with a minimum number of qubits using circuit-based quantum computers
- More Consistent Accuracy PINN via Alternating Easy-Hard Training
- Counterfactual Prediction with Deep Instrumental Variables Networks
- A Turn Toward Better Alignment: Few-Shot Generative Adaptation with Equivariant Feature Rotation
- Over-the-Air Goal-Oriented Communications
- Light and Widely Applicable MCMC: Approximate Bayesian Inference for Large Datasets
- Statistical Inference for Generative Models with Maximum Mean Discrepancy
- Training Neural Networks for and by Interpolation
- PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
- Learning General Policies with Policy Gradient Methods
- Fraud detection in credit card transactions using Quantum-Assisted Restricted Boltzmann Machines
- Finite-sample guarantees for data-driven forward-backward operator methods
- Beyond Sliding Windows: Learning to Manage Memory in Non-Markovian Environments
- Structural Reinforcement Learning for Heterogeneous Agent Macroeconomics
- Adaptive Accountability in Networked MAS: Tracing and Mitigating Emergent Norms at Scale
- Transfer Learning for Analysis of Collective and Non-Collective Thomson Scattering Spectra
- Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
- Exponentially weighted estimands and the exponential family: Filtering, prediction and smoothing
- Bias-Variance Trade-off for Clipped Stochastic First-Order Methods: From Bounded Variance to Infinite Mean
- Self-adaptive physics-informed neural network for forward and inverse problems in heterogeneous porous flow
- Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
- Limit theorems for stochastic approximation algorithms
- A Stochastic Large-scale Machine Learning Algorithm for Distributed Features and Observations
- Statistical Inference for Model Parameters in Stochastic Gradient Descent
- Universality of high-dimensional scaling limits of stochastic gradient descent
- Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences
- Variational Inference for Fully Bayesian Hierarchical Linear Models
- Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
- Multi-temporal Calving Front Segmentation
- Evolving Deep Learning Optimizers
- T-SKM-Net: Trainable Neural Network Framework for Linear Constraint Satisfaction via Sampling Kaczmarz-Motzkin Method
- Improved Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex Optimization
- The Interplay of Statistics and Noisy Optimization: Learning Linear Predictors with Random Data Weights
- DS FedProxGrad: Asymptotic Stationarity Without Noise Floor in Fair Federated Learning
- On Synchronous, Asynchronous, and Randomized Best-Response Schemes for Stochastic Nash Games
- Robust equilibria in continuous games: From strategic to dynamic robustness
- Sampling from a log-concave distribution with Projected Langevin Monte Carlo
- Fast-feedback protocols for calibration and drift control in quantum computers
- Distribution-informed Online Conformal Prediction
- Control and Reinforcement Learning through the Lens of Optimization: An Algorithmic Perspective
- RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
- Optimal and Diffusion Transports in Machine Learning
- A Perception CNN for Facial Expression Recognition
- Contextual Strongly Convex Simulation Optimization: Optimize then Predict with Inexact Solutions
- Greedy Alignment Principle for Optimizer Selection
- Evolutionary System 2 Reasoning: An Empirical Proof
- Noisy Memory Generates Value in Changing Environments
- Bayesian inference for hidden Markov models under genuine multimodality with application to ecological time series
- How (Mis)calibrated is Your Federated CLIP and What To Do About It?
- Statistical Analysis of Stationary Solutions of Coupled Nonconvex Nonsmooth Empirical Risk Minimization
- Safeguarded Stochastic Polyak Step Sizes for Non-smooth Optimization: Robust Performance Without Small (Sub)Gradients
- An hybrid stochastic Newton algorithm for logistic regression
- Random constraint sampling and duality for convex optimization
- Black Box Variational Inference
- State Dependent Performative Prediction with Stochastic Approximation
- Stochastic Online Optimization using Kalman Recursion
- Adversarial Training for Process Reward Models
- Optimal Rates for Learning with Nyström Stochastic Gradient Methods
- Online estimation of the asymptotic variance for averaged stochastic gradient algorithms
- AnoRefiner: Anomaly-Aware Group-Wise Refinement for Zero-Shot Industrial Anomaly Detection
- Global convergence rate analysis of unconstrained optimization methods based on probabilistic models
- Beyond Expectation: Concentration Inequalities for Randomized Iterative Methods
- Compositional Stochastic Average Gradient for Machine Learning and Related Applications
- MC2 Mixed Integer and Linear Programming
- ROOT: Robust Orthogonalized Optimizer for Neural Network Training
- Solving Heterogeneous Agent Models with Physics-informed Neural Networks
- HVAdam: A Full-Dimension Adaptive Optimizer
- A Generalized Additive Partial-Mastery Cognitive Diagnosis Model
- Gradient Descent Algorithm Survey
- On the Fundamental Limit of the Stochastic Gradient Identification Algorithm Under Non-Persistent Excitation
- An iterative K-FAC algorithm for Deep Learning
- Adaptive SGD with Line-Search and Polyak Stepsizes: Nonconvex Convergence and Accelerated Rates
- Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability
- A Sufficient Condition for Convergences of Adam and RMSProp
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- A Unified Convergence Analysis for Shuffling-Type Gradient Methods
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- OpenCML: End-to-End Framework of Open-world Machine Learning to Learn Unknown Classes Incrementally
- Bringing Stability to Diffusion: Decomposing and Reducing Variance of Training Masked Diffusion Models
- Distributed Delayed Stochastic Optimization
- Convergence and stability of Q-learning in Hierarchical Reinforcement Learning
- FairLRF: Achieving Fairness through Sparse Low Rank Factorization
- Towards Understanding Convergence and Generalization of AdamW
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- Automated Machine Learning on Big Data using Stochastic Algorithm Tuning
- Memory Augmented Optimizers for Deep Learning
- NuBench: An Open Benchmark for Deep Learning-Based Event Reconstruction in Neutrino Telescopes
- Structure and Dynamics of Information Pathways in Online Media
- Stochastic Smoothing for Nonsmooth Minimizations: Accelerating SGD by Exploiting Structure
- Momentum Centering and Asynchronous Update for Adaptive Gradient Methods
- On the properties of variational approximations of Gibbs posteriors
- Debiasing Stochastic Gradient Descent to handle missing values
- Stability of Stochastic Gradient Descent on Nonsmooth Convex Losses
- Momentum-based Accelerated Mirror Descent Stochastic Approximation for Robust Topology Optimization under Stochastic Loads
- Bayesian Imaging With Data-Driven Priors Encoded by Neural Networks: Theory, Methods, and Algorithms
- SMLSOM: The shrinking maximum likelihood self-organizing map
- Preconditioning Kernel Matrices
- Variational Bayesian Optimal Experimental Design
- On the convergence, lock-in probability and sample complexity of stochastic approximation
- Improving SGD convergence by online linear regression of gradients in\n multiple statistically relevant directions
- Logistic Q-Learning
- CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
- Active Importance Sampling for Variational Objectives Dominated by Rare Events: Consequences for Optimization and Generalization
- Joint Stochastic Approximation and Its Application to Learning Discrete Latent Variable Models
- Optimising Density Computations in Probabilistic Programs via Automatic Loop Vectorisation
- Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
- Learning and Testing Convex Functions
- S-D-RSM: Stochastic Distributed Regularized Splitting Method for Large-Scale Convex Optimization Problems
- DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
- Learning-based Bias Correction for Time Difference of Arrival Ultra-wideband Localization of Resource-constrained Mobile Robots
- Adaptive Hamiltonian Variational Integrators and Symplectic Accelerated Optimization
- A Class of Multi-particle Reinforced Interacting Random Walks
- On the Convergence of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning
- Online Covariance Matrix Estimation in Stochastic Gradient Descent
- Parallel and distributed asynchronous adaptive stochastic gradient methods
- A Linearly-Convergent Stochastic L-BFGS Algorithm
- Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-Tuning
- Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval
- Numerical methods for the sign problem in Lattice Field Theory
- ODE approximation for the Adam algorithm: General and overparametrized setting
- Unified Theory of Adaptive Variance Reduction
- Stochastic simulation of partial discharge inception
- Functional central limit theorem for Euler--Maruyama scheme with decreasing step sizes
- A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants
- Empirical evaluation of a Q-Learning Algorithm for Model-free Autonomous Soaring
- Modal Backflow Neural Quantum States for Anharmonic Vibrational Calculations
- No-Rank Tensor Decomposition Using Metric Learning
- Modeling Stellar Collisions in Galactic Nuclei Using Hydrodynamic Simulations and Machine Learning
- Trust-Region Methods with Low-Fidelity Objective Models
- Superpositional Gradient Descent: Harnessing Quantum Principles for Model Training
- Why Federated Optimization Fails to Achieve Perfect Fitting? A Theoretical Perspective on Client-Side Optima
- Exploring Landscapes for Better Minima along Valleys
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement Learning
- Learning Geometry: A Framework for Building Adaptive Manifold Models through Metric Optimization
- Convergence of off-policy TD(0) with linear function approximation for reversible Markov chains
- Analysis of Biased Stochastic Gradient Descent Using Sequential Semidefinite Programs
- Automatic Differentiation Variational Inference
- Scalable Perturbation Learning for Online Self-Supervised Learning in Echo State Networks
- BRAC+: Improved Behavior Regularized Actor Critic for Offline Reinforcement Learning
- Stochastic Conditional Gradient++
- An Adaptive Online HDP-HMM for Segmentation and Classification of Sequential Data
- Sharpness-aware Quantization for Deep Neural Networks
- GOAT: GPU Outsourcing of Deep Learning Training With Asynchronous Probabilistic Integrity Verification Inside Trusted Execution Environment
- Overdispersed Black-Box Variational Inference
- The Variational Gaussian Process
- Voice Biometrics Security: Extrapolating False Alarm Rate via Hierarchical Bayesian Modeling of Speaker Verification Scores
- SPRING: A fast stochastic proximal alternating method for non-smooth non-convex optimization
- A variable metric mini-batch proximal stochastic recursive gradient algorithm with diagonal Barzilai-Borwein stepsize
- Accelerating Convergence of Replica Exchange Stochastic Gradient MCMC via Variance Reduction
- A Statistician Teaches Deep Learning
- Competing with the Empirical Risk Minimizer in a Single Pass
- Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
- A Free-Energy Principle for Representation Learning
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Reclaiming the "frequentist" role of marginal likelihood in Bayesian belief revision
- Nonconvex Optimization Meets Low-Rank Matrix Factorization: An Overview
- Stationary Behavior of Constant Stepsize SGD Type Algorithms: An Asymptotic Characterization
- The Minimax Complexity of Distributed Optimization
- Scalable Hyperparameter Optimization with Lazy Gaussian Processes
- Almost sure convergence and asymptotical normality of a generalization of Kesten's stochastic approximation algorithm for multidimensional case
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- Comparison-Based Algorithms for One-Dimensional Stochastic Convex Optimization
- Contrastive Mixture of Posteriors for Counterfactual Inference, Data Integration and Fairness
- Stochastic Gradient Hamiltonian Monte Carlo
- A Stochastic Quasi-Newton Method for Large-Scale Optimization
- Model-Free Risk-Sensitive Reinforcement Learning
- Discretize-Optimize vs. Optimize-Discretize for Time-Series Regression and Continuous Normalizing Flows
- FedBone: Towards Large-Scale Federated Multi-Task Learning
- Finite-Time Convergence Rates of Nonlinear Two-Time-Scale Stochastic Approximation under Markovian Noise
- Subgradient Methods for Nonsmooth Convex Functions with Adversarial Errors
- Bayesian Transfer Learning for High-Dimensional Linear Regression via Adaptive Shrinkage
- Reducing Noise in GAN Training with Variance Reduced Extragradient
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
- Fully Implicit Online Learning
- Adaptive Gradient Descent for Optimal Control of Parabolic Equations with Random Parameters
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Dynamic Network Embeddings for Network Evolution Analysis
- Monte Carlo Co-Ordinate Ascent Variational Inference
- High-Performance Large-Scale Image Recognition Without Normalization
- Sparsity-Probe: Analysis tool for Deep Learning Models
- Stochastic Gradient Descent for Relational Logistic Regression via Partial Network Crawls
- Truncated Stochastic Approximation with Moving Bounds: Convergence
- Temporal Difference Learning as Gradient Splitting
- Teaching Machine Learning to Software Engineers
- One-class Collaborative Filtering with Random Graphs: Annotated Version
- On the fast convergence of random perturbations of the gradient flow
- Large-scale empirical tuning and comparison of default optimizers for variational inference
- An efficient Averaged Stochastic Gauss-Newton algorithm for estimating parameters of non linear regressions models
- A Collective Learning Framework to Boost GNN Expressiveness
- Stochastic actor‐oriented models for network change
- A One-step Approach to Covariate Shift Adaptation
- Veridical Data Science
- Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods
- Hamilton-Jacobi Deep Q-Learning for Deterministic Continuous-Time Systems with Lipschitz Continuous Controls
- Variational Policy Gradient Method for Reinforcement Learning with General Utilities
- Convergence in Models with Bounded Expected Relative Hazard Rates
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- Accelerated Almost-Sure Convergence Rates for Nonconvex Stochastic Gradient Descent using Stochastic Learning Rates
- KKT Conditions, First-Order and Second-Order Optimization, and Distributed Optimization: Tutorial and Survey
- Variational Calibration of Computer Models
- Deep Adaptive Design: Amortizing Sequential Bayesian Experimental Design
- A Variational Inequality Perspective on Generative Adversarial Networks
- Stochastic Variational Inference
- Finite-Sample Analysis of Nonlinear Stochastic Approximation with Applications in Reinforcement Learning
- The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure For Least Squares
- Decentralized Dynamic Discriminative Dictionary Learning
- U-CNNpred: A Universal CNN-based Predictor for Stock Markets
- Early Stopping without a Validation Set
- Riemannian stochastic quasi-Newton algorithm with variance reduction and its convergence analysis
- Convergence of a Relaxed Variable Splitting Coarse Gradient Descent Method for Learning Sparse Weight Binarized Activation Neural Networks
- Variational Inference: A Review for Statisticians
- Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition
- Primal Method for ERM with Flexible Mini-batching Schemes and Non-convex Losses
- Update estimation of diffusion parameter observed at high frequency
- Bayesian Projected Calibration of Computer Models
- An embarrassingly simple comparison of machine learning algorithms for indoor scene classification
- Stochastic Gradient Descent for Stochastic Doubly-Nonconvex Composite Optimization
- Stochastic First- and Zeroth-order Methods for Nonconvex Stochastic Programming
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
- Training Across Reservoirs: Using Numerical Differentiation To Couple Trainable Networks With Black-Box Reservoirs
- Dynamically Weighted Momentum with Adaptive Step Sizes for Efficient Deep Network Training
- A Black Box Variational Inference Scheme for Inverse Problems with Demanding Physics-Based Models
- Maximum likelihood estimation of regularisation parameters in high-dimensional inverse problems: an empirical Bayesian approach. Part II: Theoretical Analysis
- Nonlinear forward-backward-half forward splitting with momentum for monotone inclusions
- Simplified Stochastic Feedforward Neural Networks
- ProxSARAH: An Efficient Algorithmic Framework for Stochastic Composite Nonconvex Optimization
- PredProp: Bidirectional Stochastic Optimization with Precision Weighted\n Predictive Coding
- Competitive Policy Optimization
- Optimal Control in Large Open Quantum Systems: The Case of Transmon Readout and Reset
- Advances in Variational Inference
- Dual Averaging is Surprisingly Effective for Deep Learning Optimization
- On the Saturation Phenomenon of Stochastic Gradient Descent for Linear Inverse Problems
- Connectome-Guided Automatic Learning Rates for Deep Networks
- Fast large-scale optimization by unifying stochastic gradient and quasi-Newton methods
- CURVETE: Curriculum Learning and Progressive Self-supervised Training for Medical Image Classification
- How Muon's Spectral Design Benefits Generalization: A Study on Imbalanced Data
- Optimal Matrix Momentum Stochastic Approximation and Applications to Q-learning
- Stopping Rules for Monte Carlo Methods: A Review
- A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
- An Interval Hessian-based line-search method for unconstrained nonconvex optimization
- Distributed Stochastic Proximal Algorithm on Riemannian Submanifolds for Weakly-convex Functions
- An Improved Analysis of Stochastic Gradient Descent with Momentum
- Optimization in Theory and Practice
- Convergence Analysis of SGD under Expected Smoothness
- Reinforcement Learning and Consumption-Savings Behavior
- PSO-XAI: A PSO-Enhanced Explainable AI Framework for Reliable Breast Cancer Detection
- Fluctuation-dissipation relations for stochastic gradient descent
- Isotropic Noise in Stochastic and Quantum Convex Optimization
- From Optimization to Prediction: Transformer-Based Path-Flow Estimation to the Traffic Assignment Problem
- Statistical Inference for Linear Functionals of Online Least-squares SGD when t \gtrsim d1+δ
- No Intelligence Without Statistics: The Invisible Backbone of Artificial Intelligence
- An Alternating Direction Method of Multipliers for Utility-based Shortfall Risk Portfolio Optimization
- A Unified Perspective on Optimization in Machine Learning and Neuroscience: From Gradient Descent to Neural Adaptation
- A Frequentist Statistical Introduction to Variational Inference, Autoencoders, and Diffusion Models
- A Compressive Sensing Inspired Monte-Carlo Method for Combinatorial Optimization
- Practical and Private (Deep) Learning without Sampling or Shuffling
- Relational Pooling for Graph Representations
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
- Bregman Stochastic Proximal Point Algorithm with Variance Reduction
- A termination criterion for stochastic gradient descent for binary classification
- Enhancing di-jet resonance searches via a final-state radiation jet tagging algorithm
- Ant colony optimization theory: A survey
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- Cumulative Prospect Theory Meets Reinforcement Learning: Prediction and Control
- LLM Priors for ERM over Programs
- Noise-Adaptive Layerwise Learning Rates: Accelerating Geometry-Aware Optimization for Deep Neural Network Training
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- Uncertainty Quantification for Online Learning and Stochastic Approximation via Hierarchical Incremental Gradient Descent
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- A Stochastic Algorithm for Searching Saddle Points with Convergence Guarantee
- Randomness and Interpolation Improve Gradient Descent
- Stochastic gradient descent methods for estimation with large data sets
- Learning Latent Energy-Based Models via Interacting Particle Langevin Dynamics
- The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data
- DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Graph Pattern Mining and Learning through User-defined Relations (Extended Version)
- Accelerated stochastic first-order method for convex optimization under heavy-tailed noise
- The density of states from first principles
- Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
- A Stochastic Differential Equation Framework for Multi-Objective LLM Interactions: Dynamical Systems Analysis with Code Generation Applications
- Mean-square and linear convergence of a stochastic proximal point algorithm in metric spaces of nonpositive curvature
- EA4LLM: A Gradient-Free Approach to Large Language Model Optimization via Evolutionary Algorithms
- Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Distributed adaptive steplength stochastic approximation schemes for Cartesian stochastic variational inequality problems
- A Novel lightweight Convolutional Neural Network, ExquisiteNetV2
- Explicit Discovery of Nonlinear Symmetries from Dynamic Data
- Differentially Private Dropout
- Adaptive Path Sampling in Metastable Posterior Distributions
- Distributed Stochastic Approximation for Constrained and Unconstrained Optimization
- A Review of Learning with Deep Generative Models from Perspective of Graphical Modeling
- The Number of Steps Needed for Nonconvex Optimization of a Deep Learning\n Optimizer is a Rational Function of Batch Size
- Shapley Interpretation and Activation in Neural Networks
- Vertex-reinfoced random walk on Z visits finitely many states
- Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning
- Adaptive Optimal Scaling of Metropolis-Hastings Algorithms Using the Robbins-Monro Process
- Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
- Ergodicity and error estimate of laws for a random splitting Langevin Monte Carlo
- A General Framework for Joint Multi-State Models
- Gradient Shaping Beyond Clipping: A Functional Perspective on Update Magnitude Control
- Local SGD Converges Fast and Communicates Little
- SGD without Replacement: Sharper Rates for General Smooth Convex Functions
- Dimensionality Reduction for Stationary Time Series via Stochastic Nonconvex Optimization
- Bayes Factor Tests for Group Differences in Ordinal and Binary Graphical Models
- Non-Asymptotic Analysis of Efficiency in Conformalized Regression
- Mechanism design and equilibrium analysis of smart contract mediated resource allocation
- SALAD: Self-Adaptive Link Adaptation
- A Probabilistic Basis for Low-Rank Matrix Learning
- Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- Stochastic Gradient Descent, Weighted Sampling, and the Randomized\n Kaczmarz algorithm
- A Statistical Framework for Low-bitwidth Training of Deep Neural Networks
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- ProxSTORM -- A Stochastic Trust-Region Algorithm for Nonsmooth Optimization
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Tricks from Deep Learning
- Asynchronous Federated Learning with Reduced Number of Rounds and with Differential Privacy from Less Aggregated Gaussian Noise
- Memory Determines Learning Direction: A Theory of Gradient-Based Optimization in State Space Models
- Extensions of Robbins-Siegmund Theorem with Applications in Reinforcement Learning
- Management Fads, Pedagogies and Soft Technologies
- Bundle Network: a Machine Learning-Based Bundle Method
- Stochastic variational inference for GARCH models
- Monotonic Transformation Invariant Multi-task Learning
- CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
- An Investigation of Batch Normalization in Off-Policy Actor-Critic Algorithms
- Bridging Discrete and Continuous RL: Stable Deterministic Policy Gradient with Martingale Characterization
- An Improved Framework for Scaling Party Positions from Texts with Transformer
- Machine Reading Comprehension: a Literature Review
- Metric-based Regularization and Temporal Ensemble for Multi-task Learning using Heterogeneous Unsupervised Tasks
- GO Hessian for Expectation-Based Objectives
- AdaDelay: Delay Adaptive Distributed Stochastic Convex Optimization
- Continuous-Time Reinforcement Learning for Asset-Liability Management
- Convergence of online mirror descent
- InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
- A regret minimization approach to fixed-point iterations
- Effective continuous equations for adaptive SGD: a stochastic analysis view
- Laplacian Smoothing Gradient Descent
- Stochastic Gradient MCMC Methods for Hidden Markov Models
- Shaping Initial State Prevents Modality Competition in Multi-modal Fusion: A Two-stage Scheduling Framework via Fast Partial Information Decomposition
- Quantum machine learning interatomic potential: Application of variational quantum algorithm
- Stochastic gradient descent with random learning rate
- Conservative set valued fields, automatic differentiation, stochastic gradient method and deep learning
- Training Continuously‐Coupled Reconfigurable Photonic Chips with Quantum Machine Learning
- An Improved Convergence Analysis of Stochastic Variance-Reduced Policy Gradient
- Adaptive Importance Sampling via Stochastic Convex Programming
- The Convergence Behavior of Adam under Heavy-Tailed Noise
- FORGE: Fused On-Register Gradient Elimination for Memory-Efficient LLM Training
- Quadruply Stochastic Gaussian Processes
- Reinforcement learning in convergently non-stationary environments: Feudal hierarchies and learned representations
- Harvest-or-Transmit Policy for Cognitive Radio Networks: A Learning Theoretic Approach
- Continuous Assortment Optimization with Logit Choice Probabilities under Incomplete Information
- Scalable Mean-Field Variational Inference via Preconditioned Primal-Dual Optimization
- Towards a Unified Architecture for in-RDBMS Analytics
- Stochastic Approximator of Motor Threshold (SAMT) for transcranial magnetic stimulation: Online software and its performance in clinical studies
- Variance Reduction for Evolution Strategies via Structured Control Variates
- Dual Control for Approximate Bayesian Reinforcement Learning
- Practical Quasi-Newton Methods for Training Deep Neural Networks
- Conditional Generative Modeling via Learning the Latent Space
- A Variant of Gradient Descent Algorithm Based on Gradient Averaging
- Asymptotic study of stochastic adaptive algorithm in non-convex landscape
- Characterization of Excess Risk for Locally Strongly Convex Population Risk
- Finite-temperature Yang-Mills theories with the density of states method: towards the continuum limit
- Integrating Stacked Intelligent Metasurfaces and Power Control for Dynamic Edge Inference via Over-The-Air Neural Networks
- Decentralized Control via Dynamic Stochastic Prices: The Independent System Operator Problem
- Differentiable Light Transport with Gaussian Surfels via Adapted Radiosity for Efficient Relighting and Geometry Reconstruction
- Antithetic variates in higher dimensions
- A consensus-based global optimization method for high dimensional machine learning problems
- Pathfinder: Parallel quasi-Newton variational inference
- Development of Deep Learning Optimizers: Approaches, Concepts, and Update Rules
- Faster On-Device Training Using New Federated Momentum Algorithm
- Zero-inflation in the Multivariate Poisson Lognormal Family
- Randomized Block Coordinate Descent for Online and Stochastic Optimization
- Graph Coloring for Multi-Task Learning
- The Root Finding Problem Revisited: Beyond the Robbins-Monro procedure
- Towards Robust Visual Continual Learning with Multi-Prototype Supervision
- Consciousness as a Functor
- VGG-TSwinformer: Transformer-based deep learning model for early Alzheimer’s disease prediction
- DIVEBATCH: Accelerating Model Training Through Gradient-Diversity Aware Batch Size Adaptation
- Bounding the expected run-time of nonconvex optimization with early stopping
- Almost sure convergence of dropout algorithms for neural networks
- Supervised and Unsupervised Deep Learning Applied to the Majority Vote Model
- Progressive Identification of True Labels for Partial-Label Learning
- Online Learning to Sample
- Dissecting Federated-Graph Aggregation under Domain Shift: Importance-Aware Aggregation via Empirical Analysis
- On Tackling High-Dimensional Nonconvex Stochastic Optimization via Stochastic First-Order Methods with Non-smooth Proximal Terms and Variance Reduction
- The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo
- On the Convergence of SARAH and Beyond
- On the Convergence of Decentralized Adaptive Gradient Methods
- Accelerated Gradient Methods with Biased Gradient Estimates: Risk Sensitivity, High-Probability Guarantees, and Large Deviation Bounds
- Scaling Gaussian Processes with Derivative Information Using Variational Inference
- Extragradient method with variance reduction for stochastic variational inequalities
- Item Response Theory -- A Statistical Framework for Educational and Psychological Measurement
- A Unified Theory of SGD: Variance Reduction, Sampling, Quantization and\n Coordinate Descent
- Accelerated Gradient Methods for Nonconvex Nonlinear and Stochastic Programming
- On Empirical Comparisons of Optimizers for Deep Learning
- Learning Dependency-Based Compositional Semantics
- Artificial neural networks for neuroscientists: A primer
- A Multi-Batch L-BFGS Method for Machine Learning
- Revisiting the Polyak step size
- Optimal Subsampling for Data Streams with Measurement Constrained Categorical Responses
- Stochastic Neural Network with Kronecker Flow
- SVRG for Policy Evaluation with Fewer Gradient Evaluations
- Convergence Rate in Nonlinear Two-Time-Scale Stochastic Approximation with State (Time)-Dependence
- A Proximal Stochastic Gradient Method with Adaptive Step Size and Variance Reduction for Convex Composite Optimization
- Momentum-based variance-reduced proximal stochastic gradient method for composite nonconvex stochastic optimization
- Asymptotic distribution and convergence rates of stochastic algorithms\n for entropic optimal transportation between probability measures
- Heart Disease Prediction: A Comparative Study of Optimisers Performance in Deep Neural Networks
- Federated Optimization: Distributed Machine Learning for On-Device Intelligence
- Statistical Measures For Defining Curriculum Scoring Function
- Prescribe-then-Select: Adaptive Policy Selection for Contextual Stochastic Optimization
- SVN-ICP: Uncertainty Estimation of ICP-based LiDAR Odometry using Stein Variational Newton
- Online Robust and Adaptive Learning from Data Streams
- Stochastic Sign Descent Methods: New Algorithms and Better Theory
- Breaking the Conventional Forward-Backward Tie in Neural Networks: Activation Functions
- Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast Convergence
- Online Statistical Inference for Stochastic Optimization via Kiefer-Wolfowitz Methods
- History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms
- Deep Learning for Markov Chains: Lyapunov Functions, Poisson's Equation, and Stationary Distributions
- An Interactive Framework for Finding the Optimal Trade-off in Differential Privacy
- A Study of Gradient Variance in Deep Learning
- Starting Small -- Learning with Adaptive Sample Sizes
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
- Explicit Mean-Square Error Bounds for Monte-Carlo and Linear Stochastic Approximation
- Shuffling Heuristic in Variational Inequalities: Establishing New Convergence Guarantees
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- An Overview of Lead and Accompaniment Separation in Music
- Learning Convex Optimization Control Policies
- EmbedOR: Provable Cluster-Preserving Visualizations with Curvature-Based Stochastic Neighbor Embeddings
- Network Implosion: Effective Model Compression for ResNets via Static Layer Pruning and Retraining
- Stochastic versus Deterministic in Stochastic Gradient Descent
- Learning Latent Space Energy-Based Prior Model
- GENO -- GENeric Optimization for Classical Machine Learning
- Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training
- A Framework for Evaluating Gradient Leakage Attacks in Federated Learning
- Principled Design of Translation, Scale, and Rotation Invariant Variation Operators for Metaheuristics
- Online Sinkhorn: Optimal Transport distances from sample streams
- Convergence and Stability of the Stochastic Proximal Point Algorithm\n with Momentum
- GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Globally aware optimization with resurgence
- Proximal Backpropagation
- Distributed stochastic gradient tracking methods with momentum acceleration for non-convex optimization
- Convergence Rates of Stochastic Gradient Descent under Infinite Noise Variance
- Multiplicative noise and heavy tails in stochastic optimization
- A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal Distributions
- Accelerated, Optimal, and Parallel: Some Results on Model-Based Stochastic Optimization
- CNN with large memory layers
- Solutions for Mitotic Figure Detection and Atypical Classification in MIDOG 2025
- Extreme MRI: Large-Scale Volumetric Dynamic Imaging from Continuous Non-Gated Acquisitions
- Stochastic Online Feedback Optimization for Networks of Non-Compliant Agents
- Convergence of regularized agent-state-based Q-learning in POMDPs
- Calibrating generalized predictive distributions
- Revisit Stochastic Gradient Descent for Strongly Convex Objectives: Tight Uniform-in-Time Bounds
- A Twin Neural Model for Uplift
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Combined Stochastic and Robust Optimization for Electric Autonomous Mobility-on-Demand with Nested Benders Decomposition
- Federated Consistency- and Complementarity-aware Consensus-enhanced Recommendation
- A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
- Backprop with Approximate Activations for Memory-efficient Network Training
- Renewable Quantile Regression with Heterogeneous Streaming Datasets
- Implicit particle filters for data assimilation
- GRADSTOP: Early Stopping of Gradient Descent via Posterior Sampling
- Learning with springs and sticks
- Dissipativity Theory for Accelerating Stochastic Variance Reduction: A Unified Analysis of SVRG and Katyusha Using Semidefinite Programs
- A Neural Network-Based On-device Learning Anomaly Detector for Edge Devices
- Variational Wasserstein Barycenters with c-Cyclical Monotonicity
- The Emergence of Compositional Languages for Numeric Concepts Through Iterated Learning in Neural Agents
- Optimal Auction Design for the Gradual Procurement of Strategic Service Provider Agents
- Universal Reinforcement Learning in Coalgebras: Asynchronous Stochastic Computation via Conduction
- Natural Compression for Distributed Deep Learning
- Information Directed Sampling for Sparse Linear Bandits
- Network formation in the interbank money market: An application of the actor-oriented model
- Explainable Learning Rate Regimes for Stochastic Optimization
- Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
- Learning dynamical systems with particle stochastic approximation EM
- Optimization Problems for Machine Learning: A Survey
- Stochastic Gradient Hamiltonian Monte Carlo with Variance Reduction for Bayesian Inference
- An Iterative Bayesian Robbins--Monro Sequence
- The Discrete Infinite Logistic Normal Distribution
- Enabling scalable stochastic gradient-based inference for Gaussian processes by employing the Unbiased LInear System SolvEr (ULISSE)
- Search of RRATs on declinations from +42∘ to +55∘ with a neural network
- Graph Neural Diffusion via Generalized Opinion Dynamics
- Contrastive Weight Regularization for Large Minibatch SGD
- SPAN: A Stochastic Projected Approximate Newton Method
- Belief Flows of Robust Online Learning
- A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- Hybrid Stochastic-Deterministic Minibatch Proximal Gradient: Less-Than-Single-Pass Optimization with Nearly Optimal Generalization
- Condition Number Analysis of Logistic Regression, and its Implications for Standard First-Order Solution Methods
- On the mean field limit of the Random Batch Method for interacting particle systems
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Natural Analysts in Adaptive Data Analysis
- How much progress have we made in neural network training? A New Evaluation Protocol for Benchmarking Optimizers
- Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
- \(X\)-evolve: Solution space evolution powered by large language models
- Last-Iterate Complexity of SGD for Convex and Smooth Stochastic Problems
- Learn-and-Adapt Stochastic Dual Gradients for Network Resource Allocation
- Online Convex Optimization with Heavy Tails: Old Algorithms, New Regrets, and Applications
- Sample-Efficient Reinforcement Learning for Linearly-Parameterized MDPs with a Generative Model
- Training few-shot classification via the perspective of minibatch and pretraining
- Convergence of inertial dynamics and proximal algorithms governed by maximally monotone operators
- Federated and continual learning for classification tasks in a society of devices
- Adaptive Weight Decay for Deep Neural Networks
- High-Order Error Bounds for Markovian LSA with Richardson-Romberg Extrapolation
- Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
- Generalized Inner Loop Meta-Learning
- ULU: A Unified Activation Function
- Stochastic Difference-of-Convex Algorithms for Solving nonconvex optimization problems
- Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
- No Masks Needed: Explainable AI for Deriving Segmentation from Classification
- Neural Network Training via Stochastic Alternating Minimization with Trainable Step Sizes
- CSG: A stochastic gradient method for a wide class of optimization problems appearing in a machine learning or data-driven context
- QuantNet: Learning to Quantize by Learning within Fully Differentiable Framework
- Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
- Breaking the Top-K Barrier: Advancing Top-K Ranking Metrics Optimization in Recommender Systems
- SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation
- Computationally efficient Gauss-Newton reinforcement learning for model predictive control
- SGD momentum optimizer with step estimation by online parabola model
- Server Averaging for Federated Learning
- Stochastic Reweighted Gradient Descent
- Soft Separation and Distillation: Toward Global Uniformity in Federated Unsupervised Learning
- Training Deep Neural Networks by optimizing over nonlocal paths in hyperparameter space
- Personalized Transformer for Explainable Recommendation
- Efficient variational inference for generalized linear mixed models with large datasets
- Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
- Advancing Welding Defect Detection in Maritime Operations via Adapt-WeldNet and Defect Detection Interpretability Analysis
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- Stochastic variance reduced multiplicative update for nonnegative matrix factorization
- FADO: A Deterministic Detection/Learning Algorithm
- Popov Mirror-Prox Method for Variational Inequalities
- Formal Bayesian Transfer Learning via the Total Risk Prior
- Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
- Personalized Dynamic Treatment Regimes in Continuous Time: A Bayesian Approach for Optimizing Clinical Decisions with Timing
- Quantifying the mini-batching error in Bayesian inference for Adaptive Langevin dynamics
- Do We Need Zero Training Loss After Achieving Zero Training Error?
- Explaining Natural Language Processing Classifiers with Occlusion and Language Modeling
- An em algorithm for quantum Boltzmann machines
- Quantum optical shallow networks
- Stochastic gradient with least-squares control variates
- Stochastic Block BFGS: Squeezing More Curvature out of Data
- Stochastic Quantum Hamiltonian Descent
- Stochastic gradient variational Bayes for gamma approximating distributions
- Computational Advantages of Multi-Grade Deep Learning: Convergence Analysis and Performance Insights
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Learning Latent Graph Geometry via Fixed-Point Schrödinger-Type Activation: A Theoretical Study
- Dimer-Enhanced Optimization: A First-Order Approach to Escaping Saddle Points in Neural Network Training
- Reverse engineering learned optimizers reveals known and novel mechanisms
- Unit Tests for Stochastic Optimization
- Efficiency of the Wang-Landau algorithm: a simple test case
- A Langevinized Ensemble Kalman Filter for Large-Scale Static and Dynamic Learning
- QLSD: Quantised Langevin stochastic dynamics for Bayesian federated learning
- Stochastic Optimization for Performative Prediction
- Stochastic Search with an Observable State Variable
- Non-Gaussianity of Stochastic Gradient Noise
- A Dimension-free Algorithm for Contextual Continuum-armed Bandits
- Probabilistic Line Searches for Stochastic Optimization
- Stochastic gradient-free descents
- Data Sampling Strategies in Stochastic Algorithms for Empirical Risk Minimization
- Hybrid tensor network and neural network quantum states for quantum chemistry
- The Price equation reveals a universal force-metric-bias law of algorithmic learning and natural selection
- Adaptively Preconditioned Stochastic Gradient Langevin Dynamics
- Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
- SGD with Coordinate Sampling: Theory and Practice
- Selection dynamics for deep neural networks
- Mini-batch stochastic Nesterov's smoothing method for constrained convex stochastic composite optimization
- Multiple Kernel Learning from Noisy Labels by Stochastic Programming
- A Semismooth Newton Method for Support Vector Classification and Regression
- Accelerate Distributed Stochastic Descent for Nonconvex Optimization with Momentum
- Almost sure convergence rates for Stochastic Gradient Descent and Stochastic Heavy Ball
- Stochastic gradient algorithms from ODE splitting perspective
- Traversing the noise of dynamic mini-batch sub-sampled loss functions: A visual guide
- Unified Optimal Analysis of the (Stochastic) Gradient Method
- Reducing the variance in online optimization by transporting past gradients
- A Unified Stochastic Gradient Approach to Designing Bayesian-Optimal Experiments
- L-SVRG and L-Katyusha with Arbitrary Sampling
- Quadrature Compound: An approximating family of distributions
- BODAME: Bilevel Optimization for Defense Against Model Extraction
- The Power of Factorial Powers: New Parameter settings for (Stochastic)\n Optimization
- Contributions to Large Scale Bayesian Inference and Adversarial Machine Learning
- An Elementary Proof that Q-learning Converges Almost Surely
- Statistical Inference in High-dimensional Generalized Linear Models with Streaming Data
- Improving SAGA via a Probabilistic Interpolation with Gradient Descent
- ADASECANT: Robust Adaptive Secant Method for Stochastic Gradient
- Coupling Adaptive Batch Sizes with Learning Rates
- Elephant random walks with multiple extractions and general reinforcement functions
- A Cross Entropy based Stochastic Approximation Algorithm for Reinforcement Learning with Linear Function Approximation
- Self-Supervised Sketch-to-Image Synthesis
- Mixing ADAM and SGD: a Combined Optimization Method
- Statistical and Algorithmic Foundations of Reinforcement Learning
- Reliable uncertainty estimate for antibiotic resistance classification with Stochastic Gradient Langevin Dynamics
- A distributed adaptive steplength stochastic approximation method for monotone stochastic Nash Games
- Zap Q-Learning With Nonlinear Function Approximation
- Sequential Design for Computerized Adaptive Testing that Allows for Response Revision
- The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning
- Optimal Routing for Delay-Sensitive Traffic in Overlay Networks
- Semi-Supervised Classification and Segmentation on High Resolution Aerial Images
- Uniform Sampling over Episode Difficulty
- Asynchronous stochastic convex optimization
- Step-DAD: Semi-Amortized Policy-Based Bayesian Experimental Design
- OptTyper: Probabilistic Type Inference by Optimising Logical and Natural Constraints
- Online First-Order Framework for Robust Convex Optimization
- Federated Accelerated Stochastic Gradient Descent
- ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training
- Why Does Multi-Epoch Training Help?
- Adjoint-based trailing edge shape optimization of a transonic turbine vane using large eddy simulations
- Melodic Phrase Segmentation By Deep Neural Networks
- Stochastic Collapsed Variational Bayesian Inference for Latent Dirichlet Allocation
- Consensus-based Distributed Quantile Estimation in Sensor Networks
- DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
- Doubly-stochastic mining for heterogeneous retrieval
- A Partitioned Sparse Variational Gaussian Process for Fast, Distributed Spatial Modeling
- NPO: Learning Alignment and Meta-Alignment through Structured Human Feedback
- Modern Bayesian Experimental Design
- Signed Distance Function Computation from an Implicit Surface
- Decentralized Markov Chain Gradient Descent
- RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
- Neural Architecture Search with Mixed Bio-inspired Learning Rules
- Stochastic Weakly Convex Optimization Under Heavy-Tailed Noises
- Communication-Efficient Distributed Dual Coordinate Ascent
- Convergence of Batch Asynchronous Stochastic Approximation With Applications to Reinforcement Learning
- Stochastic Approximate Gradient Descent via the Langevin Algorithm
- Efficient Marginalization of Discrete and Structured Latent Variables via Sparsity
- Stochastic Recursive Gradient Algorithm for Nonconvex Optimization
- Optimization Methods for Large-Scale Machine Learning
- Blockwise Adaptivity: Faster Training and Better Generalization in Deep Learning
- Cutting Slack: Quantum Optimization with Slack-Free Methods for Combinatorial Benchmarks
- Spatial Frequency Modulation for Semantic Segmentation
- Convergence Rate of Generalized Nash Equilibrium Learning in Strongly Monotone Games with Linear Constraints
- Beyond Ground States: Physics-Inspired Optimization of Excited States of Classical Hamiltonians
- LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters
- Inference by Stochastic Optimization: A Free-Lunch Bootstrap
- Recursive Bound-Constrained AdaGrad with Applications to Multilevel and Domain Decomposition Minimization
- Incremental Adaptation of NMT for Professional Post-editors: A User Study
- Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
- SMG: A Shuffling Gradient-Based Method with Momentum
- Learning values across many orders of magnitude
- Natural Language Understanding with Distributed Representation
- Parameter estimation using simultaneous perturbation stochastic approximation
- Learning Generative Prior with Latent Space Sparsity Constraints
- A Latent Morphology Model for Open-Vocabulary Neural Machine Translation
- Measure Transport with Kernel Stein Discrepancy
- Two-Scale Stochastic Control for Multipoint Communication Systems with Renewables
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information
- A Dynamic Sampling Adaptive-SGD Method for Machine Learning
- ProcData: An R Package for Process Data Analysis
- Fast Incremental Expectation Maximization for finite-sum optimization: nonasymptotic convergence
- A Stochastic Penalty Model for Convex and Nonconvex Optimization with Big Constraints
- TauRieL: Targeting Traveling Salesman Problem with a deep reinforcement learning inspired architecture
- NysAct: A Scalable Preconditioned Gradient Descent using Nystrom Approximation
- Adaptive Bayesian Sampling with Monte Carlo EM
- Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration
- Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- EAdam Optimizer: How ε Impact Adam
- Converting the Point of View of Messages Spoken to Virtual Assistants
- FrostNet: Towards Quantization-Aware Network Architecture Search
- Risk-Averse Approximate Dynamic Programming with Quantile-Based Risk Measures
- Asynchronous Decentralized Stochastic Optimization in Heterogeneous Networks
- Stochastic Newton and Cubic Newton Methods with Simple Local Linear-Quadratic Rates
- LyAm: Robust Non-Convex Optimization for Stable Learning in Noisy Environments
- Non-smooth stochastic gradient descent using smoothing functions
- A Stochastic Gradient Descent Theorem and the Back-Propagation Algorithm
- Uncertainty Quantification for Gradient and Accelerated Gradient Descent Methods on Strongly Convex Functions
- SGD Converges to Global Minimum in Deep Learning via Star-convex Path
- A Multi-Step Richardson-Romberg Extrapolation Method For Stochastic Approximation
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with Guarantees
- Kernel-based Approximate Bayesian Inference for Exponential Family Random Graph Models
- Fast and Accurate Stellar Mass Predictions from Broad-Band Magnitudes with a Simple Neural Network: Application to Simulated Star-Forming Galaxies
- Stochastic Variational Inference for Hidden Markov Models
- Computational Aspects for Interface Identification Problems with Stochastic Modelling
- On Smoothing, Regularization and Averaging in Stochastic Approximation Methods for Stochastic Variational Inequalities
- On Biased Stochastic Gradient Estimation
- Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
- Convergence Rates and Decoupling in Linear Stochastic Approximation Algorithms
- Generic Behaviour of Strongly Reinforced Polya Urns : Convergence and Stability
- A modified tamed scheme for stochastic differential equations with superlinear drifts
- Optimal High-probability Convergence of Nonlinear SGD under Heavy-tailed Noise via Symmetrization
- Stochastic Approximation with Block Coordinate Optimal Stepsizes
- Finite-Time Analysis of Stochastic Gradient Descent under Markov Randomness
- Robust, Accurate Stochastic Optimization for Variational Inference
- Gaussian variational approximation with a factor covariance structure
- NOWPAC: A provably convergent derivative-free nonlinear optimizer with path-augmented constraints
- Convergence Rate for the Last Iterate of Stochastic Gradient Descent Schemes
- Adaptive collaboration for online personalized distributed learning with heterogeneous clients
- Stochastic TCO minimization for Video Transmission over IP Networks
- Concerning the differentiability of the energy function in vector quantization algorithms
- Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes
- SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam
- Non-Asymptotic Analysis of Online Local Private Learning with SGD
- Approximately Exact Line Search
- Online Robust Subspace Tracking from Partial Information
- Free energy computations by minimization of Kullback-Leibler divergence: an efficient adaptive biasing potential method for sparse representations
- Bidirectional compression in heterogeneous settings for distributed or federated learning with partial participation: tight convergence guarantees
- Online Learning as Stochastic Approximation of Regularization Paths
- Adaptivity via a Parallel Architecture for Stochastic Gradient Methods
- Efficient Federated Learning with Timely Update Dissemination
- Ampere: Communication-Efficient and High-Accuracy Split Federated Learning
- Computer-aided analyses of stochastic first-order methods, via interpolation conditions for stochastic optimization
- Necessary condition for sparse optimal control problem with intermediate constraints
- Visual Identification of Individual Holstein-Friesian Cattle via Deep Metric Learning
- Machine Learning's Dropout Training is Distributionally Robust Optimal
- Kalman Filter Aided Federated Koopman Learning
- Mini-batch Metropolis-Hastings MCMC with Reversible SGLD Proposal
- Cyclic Differentiable Architecture Search
- Temporal Conformal Prediction (TCP): A Distribution-Free Statistical and Machine Learning Framework for Adaptive Risk Forecasting
- Mini-batch Stochastic Approximation Methods for Nonconvex Stochastic Composite Optimization
- Adversarial Data Augmentation for Single Domain Generalization via Lyapunov Exponent-Guided Optimization
- Deep Lyapunov Function: Automatic Stability Analysis for Dynamical Systems
- On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data
- Low-Rank Factorization of Determinantal Point Processes for Recommendation
- Dynamical Isometry: The Missing Ingredient for Neural Network Pruning
- Local SGD: Unified Theory and New Efficient Methods
- Distributed Second Order Methods with Fast Rates and Compressed Communication
- Privacy for Free: Posterior Sampling and Stochastic Gradient Monte Carlo
- Predicting the mechanical response of oligocrystals with deep learning
- Addressing Algorithmic Bottlenecks in Elastic Machine Learning with Chicle
- Statistical Inference for Stochastic Gradient Descent: Beyond Finite Variance
- Identifying and Analyzing Sepsis States: A Retrospective Study on Patients with Sepsis in ICUs
- Robust Brain Tumor Segmentation with Incomplete MRI Modalities Using Hölder Divergence and Mutual Information-Enhanced Knowledge Transfer
- Training Neural Networks Using Features Replay
- AuON: A Linear-time Alternative to Semi-Orthogonal Momentum Updates
- Finite-Time Analysis of Asynchronous Stochastic Approximation and Q-Learning
- Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
- Harnessing the Power of Reinforcement Learning for Adaptive MCMC
- An Adaptive State Aggregation Algorithm for Markov Decision Processes
- Large-scale Neural Network Quantum States for ab initio Quantum Chemistry Simulations on Fugaku
- Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
- Graph Drawing by Stochastic Gradient Descent
- Computing a human-like reaction time metric from stable recurrent vision models
- Two Spelling Normalization Approaches Based on Large Language Models
- Breaking a Logarithmic Barrier in the Stopping Time Convergence Rate of Stochastic First-order Methods
- DeepWukong
- Hierarchical Variational Models
- A High Probability Analysis of Adaptive SGD with Momentum
- WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
- Learning Threshold-Type Investment Strategies with Stochastic Gradient Method
- Exact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables
- Trade-offs of Local SGD at Scale: An Empirical Study
- Learning Curves for SGD on Structured Features
- Generalized Adaptation for Few-Shot Learning
- Generalized Polya urns via stochastic approximation
- Asynchronous adaptive networks
- A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior
- Optimization-Induced Dynamics of Lipschitz Continuity in Neural Networks
- Recursive Optimization of Convex Risk Measures: Mean-Semideviation Models
- Variational Variance: Simple, Reliable, Calibrated Heteroscedastic Noise Variance Parameterization
- Hindsight-Guided Momentum (HGM) Optimizer: An Approach to Adaptive Learning Rate
- Stochastic Dual Coordinate Ascent Methods for Regularized Loss Minimization
- Noise-induced degeneration in online learning
- PDFNet: Pointwise Dense Flow Network for Urban-Scene Segmentation
- A Tutorial on Bayesian Optimization
- Noise-Informed Diffusion-Generated Image Detection with Anomaly Attention
- On the Complexity of Minimizing Convex Finite Sums Without Using the Indices of the Individual Functions
- Riemannian Stochastic Hybrid Gradient Algorithm for Nonconvex Optimization
- Non-asymptotic Analysis of Biased Stochastic Approximation Scheme
- A Novel Indicator for Quantifying and Minimizing Information Utility Loss of Robot Teams
- A Study of Hybrid and Evolutionary Metaheuristics for Single Hidden Layer Feedforward Neural Network Architecture
- Don't Use Large Mini-Batches, Use Local SGD
- A Stochastic Proximal Gradient Framework for Decentralized Non-Convex Composite Optimization: Topology-Independent Sample Complexity and Communication Efficiency
- Gravilon: Applications of a New Gradient Descent Method to Machine Learning
- Monte Carlo methods: Application to hydrogen gas and hard spheres
- A Stochastic Successive Minimization Method for Nonsmooth Nonconvex Optimization with Applications to Transceiver Design in Wireless Communication Networks
- Global Convergence and Stability of Stochastic Gradient Descent
- Scalable and Efficient Comparison-based Search without Features
- Stochastic first-order methods: non-asymptotic and computer-aided analyses via potential functions
- A Sharp Estimate on the Transient Time of Distributed Stochastic Gradient Descent
- Automated proof synthesis for propositional logic with deep neural networks
- Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
- An approximate Riemann solver approach in Physics-Informed Neural Networks for hyperbolic conservation laws
- Differential Privacy in Machine Learning: From Symbolic AI to LLMs
- A stochastic second-order generalized estimating equations approach for estimating intraclass correlation coefficient in the presence of informative missing data
- Towards Undistillable Models by Minimizing Conditional Mutual Information
- Stochastic Heavy Ball
- RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer
- Simpler, Faster, Stronger: Breaking The log-K Curse On Contrastive Learners With FlatNCE
- Uncertainty-Aware Strategies: A Model-Agnostic Framework for Robust Financial Optimization through Subsampling
- Mean Field Games without Rational Expectations
- Convergence of Momentum-Based Optimization Algorithms with Time-Varying Parameters
- Dynamic Stochastic Approximation for Multi-stage Stochastic Optimization
- Stagewise Enlargement of Batch Size for SGD-based Learning
- ASMOP: Additional sampling stochastic trust region method for multi-objective problems
- Stacey: Promoting Stochastic Steepest Descent via Accelerated ℓp-Smooth Nonconvex Optimization
- Path Integral Optimiser: Global Optimisation via Neural Schrödinger-Föllmer Diffusion
- Direct Fisher Score Estimation for Likelihood Maximization
- Parsimonious Inference
- The limits of min-max optimization algorithms: convergence to spurious non-critical sets
- ACFNet: Attentional Class Feature Network for Semantic Segmentation
- Speed learning on the fly
- A Machine-Learning Method for Time-Dependent Wave Equations over Unbounded Domains
- Rapid training of Hamiltonian graph networks using random features
- Boosting Variational Inference
- Quantum sequel of neural network training
- Do optimization methods in deep learning applications matter?
- KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance Products
- Boulevard: Regularized Stochastic Gradient Boosted Trees and Their Limiting Distribution
- Weighted Aggregating Stochastic Gradient Descent for Parallel Deep Learning
- Knowledge Distillation as Semiparametric Inference
- Continuous-time Models for Stochastic Optimization Algorithms
- Online Matching in Sparse Random Graphs: Non-Asymptotic Performances of\n Greedy Algorithm
- Uniform-in-Time Weak Error Analysis for Stochastic Gradient Descent Algorithms via Diffusion Approximation
- Using Social Network Information in Bayesian Truth Discovery
- On Stochastic Gradient and Subgradient Methods with Adaptive Steplength\n Sequences
- New First-Order Algorithms for Stochastic Variational Inequalities
- A Distributed Hierarchical SGD Algorithm with Sparse Global Reduction
- No More Pesky Learning Rates
- How Data Augmentation affects Optimization for Linear Regression
- Online Asynchronous Distributed Regression
- Variance Regularization for Accelerating Stochastic Optimization
- On the Discrepancy Principle for Stochastic Gradient Descent
- Regret in Online Combinatorial Optimization
- Nonparametric Learning Algorithms for Joint Pricing and Inventory Control with Lost Sales and Censored Demand
- Parameter-free Stochastic Optimization of Variationally Coherent Functions
- Recent Advances in Deep Learning for Object Detection
- Revisiting the Characteristics of Stochastic Gradient Noise and Dynamics
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms
- Direct Acceleration of SAGA using Sampled Negative Momentum
- Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes
- Simultaneous Model Selection and Optimization through Parameter-free Stochastic Learning
- Fast Fourier Transform-Based Spectral and Temporal Gradient Filtering for Differential Privacy
- When Does Stochastic Gradient Algorithm Work Well?
- Dynamic Collaborative Filtering with Compound Poisson Factorization
- Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms
- Structured Actor-Critic for Managing Public Health Points-of-Dispensing
- Coulomb GANs: Provably Optimal Nash Equilibria via Potential Fields
- Stochastic Functional Gradient Path Planning in Occupancy Maps
- Stochastic Gradient Descent on a Tree: an Adaptive and Robust Approach to Stochastic Convex Optimization
- Dense Prediction with Attentive Feature Aggregation
- A Complete Recipe for Stochastic Gradient MCMC
- Bayesian Inference Forgetting
- Latent Guided Sampling for Combinatorial Optimization
- Distributed Forward-Backward algorithms for stochastic generalized Nash equilibrium seeking
- Accelerated Stochastic Quasi-Newton Optimization on Riemann Manifolds
- How neural networks find generalizable solutions: Self-tuned annealing in deep learning
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
- Multilevel Stochastic Gradient Descent for Optimal Control Under Uncertainty
- Variational Inference via χ-Upper Bound Minimization
- A Tree-guided CNN for image super-resolution
- Tensor Normal Training for Deep Learning Models
- Quantitative LLM Judges
- Learning to Optimize in Swarms
- Overfitting or Underfitting? Understand Robustness Drop in Adversarial Training
- Self-supervised Latent Space Optimization with Nebula Variational Coding
- Local Expectation Gradients for Doubly Stochastic Variational Inference
- Taming LLMs by Scaling Learning Rates with Gradient Grouping
- A Closer Look at Deep Policy Gradients
- Are All Languages Equally Hard to Language-Model?
- FedQuad: Adaptive Layer-wise LoRA Deployment and Activation Quantization for Federated Fine-Tuning
- Energy Time Ptychography for one-dimensional phase retrieval
- Semi-Implicit Back Propagation
- Sketchy Empirical Natural Gradient Methods for Deep Learning
- Wasserstein Distance Maximizing Intrinsic Control
- On Variance Reduction in Stochastic Gradient Descent and its Asynchronous Variants
- Optimization of neural networks via finite-value quantum fluctuations
- Stochastic Block Mirror Descent Methods for Nonsmooth and Stochastic Optimization
- Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
- Stochastic Belief Propagation: A Low-Complexity Alternative to the Sum-Product Algorithm
- Throughput Optimal Decentralized Scheduling of Multi-Hop Networks with End-to-End Deadline Constraints: II Wireless Networks with Interference
- A Probabilistically Motivated Learning Rate Adaptation for Stochastic Optimization
- On the convergence of mirror descent beyond stochastic convex programming
- DSAGL: Dual-Stream Attention-Guided Learning for Weakly Supervised Whole Slide Image Classification
- Towards Open-Text Semantic Parsing via Multi-Task Learning of Structured Embeddings
- Training Dynamic Exponential Family Models with Causal and Lateral Dependencies for Generalized Neuromorphic Computing
- Random directions stochastic approximation with deterministic perturbations
- Stochastic Variational Inference for Bayesian Sparse Gaussian Process Regression
- Algorithms for Kullback-Leibler Approximation of Probability Measures in\n Infinite Dimensions
- Smoothed Gradients for Stochastic Variational Inference
- Screening for Sparse Online Learning
- Taming Transformer Without Using Learning Rate Warmup
- Stochastic Approximation, Cooperative Dynamics and Supermodular Games
- An efficient algorithm for numerical computations of continuous densities of states
- New Convergence Aspects of Stochastic Gradient Algorithms
- Convergence of Clipped-SGD for Convex (L0,L1)-Smooth Optimization with Heavy-Tailed Noise
- Stochastic Euler Schemes and Dissipative Evolutions in the Space of Probability Measures
- Sharpness-Aware Minimization with Z-Score Gradient Filtering
- Connecting randomized iterative methods with Krylov subspaces
- AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
- Multi-objective Large Language Model Alignment with Hierarchical Experts
- On The Convergence of Euler Discretization of Finite-Time Convergent Gradient Flows
- A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging
- Nearly Dimension-Independent Convergence of Mean-Field Black-Box Variational Inference
- Finite-Sample Analysis of Stochastic Approximation Using Smooth Convex Envelopes
- Weight-Preserving Simulated Tempering
- ViewCraft3D: High-Fidelity and View-Consistent 3D Vector Graphics Synthesis
- Investigating Alternatives to the Root Mean Square for Adaptive Gradient Methods
- Meta-Learning Bidirectional Update Rules
- Distributed Learning and its Application for Time-Series Prediction
- Connecting Independently Trained Modes via Layer-Wise Connectivity
- On the Robustness of Average Losses for Partial-Label Learning
- Fast Variational Inference for Bayesian Factor Analysis in Single and Multi-Study Settings
- Oracle inequalities for computationally adaptive model selection
- A Simple Algorithm for Scalable Monte Carlo Inference
- A Bayesian Perspective of Convolutional Neural Networks through a Deconvolutional Generative Model
- Designing Pin-pression Gripper and Learning its Dexterous Grasping with Online In-hand Adjustment
- A Natural Actor-Critic Algorithm with Downside Risk Constraints
- Do Large Language Models (Really) Need Statistical Foundations?
- k-SVRG: Variance Reduction for Large Scale Optimization
- Joint-stochastic-approximation Random Fields with Application to Semi-supervised Learning
- Joint-stochastic-approximation Autoencoders with Application to Semi-supervised Learning
- Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
- A variational approach to stochastic minimization of convex functionals
- Gradient-only line searches: An Alternative to Probabilistic Line Searches
- On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and Non-Asymptotic Concentration
- An iterative regularized mirror descent method for ill-posed nondifferentiable stochastic optimization
- A New Approach for Optimizing Highly Nonlinear Problems Based on the Observer Effect Concept
- Stochastic Gradient Descent for Linear Systems with Missing Data
- Two-Player Games for Efficient Non-Convex Constrained Optimization
- Dynamic Dual Buffer with Divide-and-Conquer Strategy for Online Continual Learning
- Riemannian stochastic recursive momentum method for non-convex optimization
- New Tight Bounds for SGD without Variance Assumption: A Computer-Aided Lyapunov Analysis
- Asymptotic normality of randomly truncated stochastic algorithms
- NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
- Stochastic Approximation versus Sample Average Approximation for population Wasserstein barycenters
- Ergodic Mirror Descent
- Language modeling with Neural trans-dimensional random fields
- EMRA-proxy: Enhancing Multi-Class Region Semantic Segmentation in Remote Sensing Images with Attention Proxy
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in Training
- Multi-cut stochastic approximation methods for solving stochastic convex composite optimization
- NPTC-net: Narrow-Band Parallel Transport Convolutional Neural Network on Point Clouds
- Fast Stochastic Methods for Nonsmooth Nonconvex Optimization
- Online Statistical Inference of Constrained Stochastic Optimization via Random Scaling
- Statistical Inference for Online Algorithms
- Risk-Averse Reinforcement Learning with Itakura-Saito Loss
- Estimation and uncertainty quantification for the output from quantum simulators
- A Two-Stage Data Selection Framework for Data-Efficient Model Training on Edge Devices
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex Problems
- Adaptive Stochastic Optimization
- Stochastic optimization for numerical evaluation of imprecise probabilities
- Safe Learning under Uncertain Objectives and Constraints
- Convergence of Adam in Deep ReLU Networks via Directional Complexity and Kakeya Bounds
- Statistical Adaptive Stochastic Gradient Methods
- VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
- Robustness, Privacy, and Generalization of Adversarial Training
- A theoretical and empirical study of new adaptive algorithms with additional momentum steps and shifted updates for stochastic non-convex optimization
- ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring Domains
- Pathwise Derivatives Beyond the Reparameterization Trick
- KO: Kinetics-inspired Neural Optimizer with PDE Simulation Approaches
- CBA: Contextual Quality Adaptation for Adaptive Bitrate Video Streaming (Extended Version)
- An Adaptive Sample Size Trust-Region Method for Finite-Sum Minimization
- A Selective Overview of Deep Learning
- Characterizing nonatomic admissions markets
- Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
- Smoothed SGD for quantiles: Bahadur representation and Gaussian approximation
- Surrogate Optimization of Deep Neural Networks for Groundwater Predictions
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- Bandwidth-based Step-Sizes for Non-Convex Stochastic Optimization
- Bias-Variance Tradeoff in a Sliding Window Implementation of the Stochastic Gradient Algorithm
- Stochastic Approximation Hamiltonian Monte Carlo
- A Stochastic Quasi-Newton Method for Large-Scale Nonconvex Optimization with Applications
- The Stochastic Multi-Proximal Method for Nonsmooth Optimization
- The Curious Case of the Default Settings: Evaluating Default Performance of Variational Inference Software
- Never Skip a Batch: Dense Learning of Temporal GNNs via Adaptive Pseudo-Supervision
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- DeepOBS: A Deep Learning Optimizer Benchmark Suite
- An Exploratory Analysis of the Latent Structure of Process Data via Action Sequence Autoencoder
- Modified swarm-based metaheuristics enhance Gradient Descent initialization performance: Application for EEG spatial filtering
- Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
- Revisiting Stochastic Approximation and Stochastic Gradient Descent
- A Unifying Probabilistic View of Associative Learning
- VIP - Variational Inversion Package with example implementations of Bayesian tomographic imaging
- Tight Dimension Independent Lower Bound on the Expected Convergence Rate for Diminishing Step Sizes in SGD
- On Stochastic Variance Reduced Gradient Method for Semidefinite Optimization
- Joint Status Sampling and Updating for Minimizing Age of Information in the Internet of Things
- Convergence Analysis of the Last Iterate in Distributed Stochastic Gradient Descent with Momentum
- Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
- Percolation Threshold Results on \Erdos-\Renyi Graphs: an Empirical Process Approach
- A 140 line MATLAB code for topology optimization problems with probabilistic parameters
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
- Incorporating brain-inspired mechanisms for multimodal learning in artificial intelligence
- DeepSeqCoco: A Robust Mobile Friendly Deep Learning Model for Detection of Diseases in Cocos nucifera
- An Exponential Averaging Process with Strong Convergence Properties
- Stochastic Item Descent Method for Large Scale Equal Circle Packing Problem
- Spike-timing-dependent Hebbian learning as noisy gradient descent
- SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
- The stochastic Auxiliary Problem Principle in Banach spaces: measurability and convergence
- Evaluating Visual Properties via Robust HodgeRank
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- A Latent Variational Framework for Stochastic Optimization
- Distributed Training of Deep Neural Network Acoustic Models for Automatic Speech Recognition
- Distributed Stochastic Algorithms for High-rate Streaming Principal Component Analysis
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- Generative Text Modeling through Short Run Inference
- Online differentially private inference in stochastic gradient descent
- Consistency and fluctuations for stochastic gradient Langevin dynamics
- A random batch Ewald method for particle systems with Coulomb interactions
- Fast-Mixing Markov Chains without Gradients
- Sharp Gaussian approximations for Decentralized Federated Learning
- Machine Learning on Volatile Instances
- Orthogonal-Padé Activation Functions: Trainable Activation functions for smooth and faster convergence in deep networks
- Sharp Analysis for Nonconvex SGD Escaping from Saddle Points
- Bant: Byzantine Antidote via Trial Function and Trust Scores
- Learning-based Bias Correction for Ultra-wideband Localization of Resource-constrained Mobile Robots
- Stochastic ADMM with batch size adaptation for nonconvex nonsmooth optimization
- Stochastic Cubic Regularization for Fast Nonconvex Optimization
- A Simple Stochastic Variance Reduced Algorithm with Fast Convergence Rates
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and Performance
- Convergence in quadratic mean of averaged stochastic gradient algorithms\n without strong convexity nor bounded gradient
- Stochastic DCA for minimizing a large sum of DC functions with application to Multi-class Logistic Regression
- FedADP: Unified Model Aggregation for Federated Learning with Heterogeneous Model Architectures
- Towards Probabilistic Verification of Machine Unlearning
- KML: Using Machine Learning to Improve Storage Systems
- Age of Processing: Age-driven Status Sampling and Processing Offloading for Edge Computing-enabled Real-time IoT Applications
- Generative Modeling by Inclusive Neural Random Fields with Applications in Image Generation and Anomaly Detection
- Two-stage Linear Decision Rules for Multi-stage Stochastic Programming
- Extended Fiducial Inference for Individual Treatment Effects via Deep Neural Networks
- Variational Inference: A Review for Statisticians
- Advances in Asynchronous Parallel and Distributed Optimization
- Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients
- DRSLF: Double Regularized Second-Order Low-Rank Representation for Web Service QoS Prediction
- A Stochastic forward-backward splitting method for solving monotone inclusions in Hilbert spaces
- FTNILO: Explicit Multivariate Function Inversion, Optimization and Counting, Cryptography Weakness and Riemann Hypothesis Solution Equation with Tensor Networks
- Stochastic optimization methods for the simultaneous control of\n parameter-dependent systems
- Time Adaptive Reinforcement Learning
- A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
- Consciousness in AI: Logic, Proof, and Experimental Evidence of Recursive Identity Formation
- Analysis of nonsmooth stochastic approximation: the differential inclusion approach
- Balanced Alignment for Face Recognition: A Joint Learning Approach
- Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization
- Distributed Derivative-free Learning Method for Stochastic Optimization over a Network with Sparse Activity
- Trust-Region Algorithms for Training Responses: Machine Learning Methods\n Using Indefinite Hessian Approximations
- Particle-based Generalised Stochastic Optimisation
- signProx: One-Bit Proximal Algorithm for Nonconvex Stochastic Optimization
- A Unifying Framework for Variance Reduction Algorithms for Finding Zeroes of Monotone Operators
- Optimal Primal-Dual Methods for a Class of Saddle Point Problems
- Gonogo: An R Implementation of Test Methods to Perform, Analyze and Simulate Sensitivity Experiments
- Differentiable Visual Computing
- Efficient and Robust Algorithms for Adversarial Linear Contextual Bandits
- Optimum Experimental Designs
- Demystifying Parallel and Distributed Deep Learning
- Accelerating Mini-batch SARAH by Step Size Rules
- Computing hitting times via fluid approximation: application to the coupon collector problem
- Convex Optimization: Algorithms and Complexity
- Explicit Regularization of Stochastic Gradient Methods through Duality
- Randomized Block Subgradient Methods for Convex Nonsmooth and Stochastic Optimization
- Reinforced stochastic gradient descent for deep neural network learning
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- Quantile estimation with adaptive importance sampling
- Smoothed Hinge Loss and ℓ1 Support Vector Machines
- Variance Reduction for Distributed Stochastic Gradient Descent
- MTL2L: A Context Aware Neural Optimiser
- Learning Retrospective Knowledge with Reverse Reinforcement Learning
- Finding best approximation pairs for two intersections of closed convex sets
- Stochastic Annealing
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- Patterns, predictions, and actions: A story about machine learning
- From Persistent Homology to Reinforcement Learning with Applications for Retail Banking
- WNGrad: Learn the Learning Rate in Gradient Descent
- Fast Transient Simulation of High-Speed Channels Using Recurrent Neural Network
- Non-Gaussian processes and neural networks at finite widths
- Convergence Rates of Accelerated Markov Gradient Descent with Applications in Reinforcement Learning
- Second-order Information in First-order Optimization Methods
- Conjugate-gradient-based Adam for stochastic optimization and its application to deep learning
- Latent Variable Session-Based Recommendation
- Non-asymptotic analysis of online noisy stochastic gradient descent
- Some aspects of the sequential design of experiments
- Meta knowledge assisted Evolutionary Neural Architecture Search
- Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
- Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
- Second-Order Path Kernel Interpolation Formulas in Machine Learning
- Fast Compute for ML Optimization
- Improving the stability of the covariance-controlled adaptive Langevin thermostat for large-scale Bayesian sampling
- IPAS: An Adaptive Sample Size Method for Weighted Finite Sum Problems with Linear Equality Constraints
- O(1/k) Finite-Time Bound for Non-Linear Two-Time-Scale Stochastic Approximation
- HyperController: A Hyperparameter Controller for Fast and Stable Training of Reinforcement Learning Neural Networks
- Natural Language Embeddings of Synthesis and Testing conditions Enhance Glass Dissolution Prediction
- Gradient Descent's Last Iterate is Often (slightly) Suboptimal
- Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
- A Langevin sampling algorithm inspired by the Adam optimizer
- CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
- Extrapolation Guarantees for Perturbation Modeling Under the Additive Latent Shift Assumption
- Linear Least-Squares Algorithms for Temporal Difference Learning
- Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent
- On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points
- Online Pairwise Learning Algorithms with Kernels
- Stochastic Saddle Avoidance Beyond Unit Excitation and Smoothness: A Pathwise Lyapunov-Perron Framework
- On the Convergence of Consensus Algorithms with Markovian Noise and Gradient Bias
- Stochastic Gradient Descent with Polyak's Learning Rate
- Finite Time Analysis of Linear Two-timescale Stochastic Approximation with Markovian Noise
- An Inverse Source Problem for Semilinear Stochastic Hyperbolic Equations
- Backtracking Gradient Descent Method and Some Applications in Large Scale Optimisation. Part 2: Algorithms and Experiments
- Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
- Silenzio: Secure Non-Interactive Outsourced MLP Training
- Improving Deep Knowledge Tracing via Gated Architectures and Adaptive Optimization
- Distributed SAGA: Maintaining linear convergence rate with limited communication
- Optimal minimization of an unknown function in a nonparametric multivariate regression model thanks to a dimension reduction approach
- A stochastic alternating direction method of multipliers for non-smooth and non-convex optimization
- Active Learning Via Sequential Design and Uncertainty Sampling
- On PyTorch Implementation of Density Estimators for von Mises-Fisher and Its Mixture
- Моментные оценки для стохастической аппроксимации
- Deep convolutional neural networks for uncertainty propagation in random fields
- Randomised Splitting Methods and Stochastic Gradient Descent
- Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
- Memory Limited, Streaming PCA
- Stochastic proximal gradient methods for nonconvex problems in Hilbert spaces
- Dynamic sampling schemes for optimal noise learning under multiple\n nonsmooth constraints
- Strong error analysis for the stochastic momentum optimizer
- Langevin Markov Chain Monte Carlo with stochastic gradients
- Covariant Gradient Descent
- Fast X-Ray CT Image Reconstruction Using a Linearized Augmented Lagrangian Method With Ordered Subsets
- DS-MLR: Exploiting Double Separability for Scaling up Distributed\n Multinomial Logistic Regression
- FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
- On an inferential model construction using generalized associations
- Efficient neural-network based variational Monte Carlo scheme for direct optimization of excited energy states in frustrated quantum systems
- GRADIENT-BASED STOCHASTIC OPTIMIZATION METHODS IN BAYESIAN EXPERIMENTAL DESIGN
- Bayesian sample size calculations for external validation studies of risk prediction models
- Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization
- Stochastic Damped L-BFGS with Controlled Norm of the Hessian Approximation
- THE SELF-ORGANIZED MULTI-LATTICE MONTE CARLO SIMULATION
- A neural network-based framework for financial model calibration
- A framework for adaptive Monte Carlo procedures
- A Diffusion Approximation Theory of Momentum SGD in Nonconvex Optimization
- Stochastic Optimization with Optimal Importance Sampling
- An XAI-based Analysis of Shortcut Learning in Neural Networks
- Stochastic approximation with random step sizes and urn models with random replacement matrices having finite mean
- Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
- Solving Multi-Agent Safe Optimal Control with Distributed Epigraph Form MARL
- Geometric Learning Dynamics
- Learning Supervised Topic Models for Classification and Regression from Crowds
- EXPLICIT HESTON SOLUTIONS AND STOCHASTIC APPROXIMATION FOR PATH-DEPENDENT OPTION PRICING
- Convergence of distributed asynchronous learning vector quantization algorithms
- Ensemble Kalman inversion: a derivative-free technique for machine learning tasks
- Constrained clustering and Kohonen Self-Organizing Maps
- PyFRep: Shape Modeling with Differentiable Function Representation
- Rethinking Client-oriented Federated Graph Learning
- Computing Equilibria in Stochastic Nonconvex and Non-monotone Games via Gradient-Response Schemes
- Sequential Scenario-Specific Meta Learner for Online Recommendation
- A Common Derivation for Markov Chain Monte Carlo Algorithms with Tractable and Intractable Targets
- A Robbins–Monro procedure for estimation in semiparametric regression models
- Randomized Kaczmarz with Averaging
- Block layer decomposition schemes for training deep neural networks
- Variational Quantum Monte Carlo Simulations with Tensor-Network States
- Momentum Schemes with Stochastic Variance Reduction for Nonconvex\n Composite Optimization
- On the Convergence of Reinforcement Learning with Monte Carlo Exploring Starts
- VecHGrad for Solving Accurately Complex Tensor Decomposition
- Rethinking Curriculum Learning with Incremental Labels and Adaptive Compensation
- Ouroboros: On Accelerating Training of Transformer-Based Language Models
- Gmst:An Unbiased Stratified Statistic and a Fast Gradient Optimization Algorithm Based on It
- FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image Classification
- Dimension-free convergence rates for gradient Langevin dynamics in RKHS
- Self-Tuned Mirror Descent Schemes for Smooth and Nonsmooth High-Dimensional Stochastic Optimization
- Large-Batch Training for LSTM and Beyond
- Multi-level stochastic approximation algorithms
- An overview of SPSA: recent development and applications
- Heterogeneous Multilayer Generalized Operational Perceptron
- SGD-Net: Efficient Model-Based Deep Learning With Theoretical Guarantees
- Deep convolutions for in-depth automated rock typing
- A Parallel Decomposition Method for Nonconvex Stochastic Multi-Agent Optimization Problems
- Distributed Variational Bayesian Algorithms Over Sensor Networks
- A Stochastic Subgradient Method for Distributionally Robust Non-Convex Learning
- Fast or Slow: An Autonomous Speed Control Approach for UAV-assisted IoT Data Collection Networks
- Gradient-Free Sequential Bayesian Experimental Design via Interacting Particle Systems
- Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
- Stochastic Gradient Methods with Block Diagonal Matrix Adaptation
- Second-order Optimization of Gaussian Splats with Importance Sampling
- Faster State Preparation across Quantum Phase Transition Assisted by Reinforcement Learning
- Stochastic Gradient Descent in Non-Convex Problems: Asymptotic Convergence with Relaxed Step-Size via Stopping Time Methods
- Leave-One-Out Stable Conformal Prediction
- A proximal subgradient method for nonconvex stochastic optimization under the Kurdyka-Łojasiewicz condition
- Multi-resolution Score-Based Variational Graphical Diffusion for Causal Disaster System Modeling and Inference
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Understanding the theoretical properties of projected Bellman equation, linear Q-learning, and approximate value iteration
- Molecular Learning Dynamics
- MiMu: Mitigating Multiple Shortcut Learning Behavior of Transformers
- A Tale of Two Learning Algorithms: Multiple Stream Random Walk and Asynchronous Gossip
- Towards Weaker Variance Assumptions for Stochastic Optimization
- DUE: A Deep Learning Framework and Library for Modeling Unknown Equations
- Fairness is in the details: Face Dataset Auditing
- A Piecewise Lyapunov Analysis of Sub-quadratic SGD: Applications to Robust and Quantile Regression
- Deriving the Gradients of Some Popular Optimal Transport Algorithms
- In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods
- Representation Meets Optimization: Training PINNs and PIKANs for Gray-Box Discovery in Systems Pharmacology
- Adam revisited: a weighted past gradients perspective
- Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification
- Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
- Weighted Approximate Quantum Natural Gradient for Variational Quantum Eigensolver
- Stochastic Variational Inference with Tuneable Stochastic Annealing
- A Geometric Framework for Stochastic Iterations
- Backtracking line search [wikipedia]
- Event-driven contrastive divergence for spiking neuromorphic systems. [europepmc]
- A Unifying Probabilistic View of Associative Learning. [europepmc]
- Midbrain Synchrony to Envelope Structure Supports Behavioral Sensitivity to Single-Formant Vowel-Like Sounds in Noise. [europepmc]
- Cerebellar learning using perturbations. [europepmc]
- RPITER: A Hierarchical Deep Learning Framework for ncRNA⁻Protein Interaction Prediction. [europepmc]
- Deeper Profiles and Cascaded Recurrent and Convolutional Neural Networks for state-of-the-art Protein Secondary Structure Prediction. [europepmc]
- DeepPoseKit, a software toolkit for fast and robust animal pose estimation using deep learning. [europepmc]
- Knowing What You Know in Brain Segmentation Using Bayesian Deep Neural Networks. [europepmc]
- Classification of brain tumor isocitrate dehydrogenase status using MRI and deep learning. [europepmc]
- LRRpredictor-A New LRR Motif Detection Method for Irregular Motifs of Plant NLR Proteins Using an Ensemble of Classifiers. [europepmc]
- Modern Soft-Sensing Modeling Methods for Fermentation Processes. [europepmc]
- A deep learning based framework for the registration of three dimensional multi-modal medical images of the head. [europepmc]
- Practices and Applications of Convolutional Neural Network-Based Computer Vision Systems in Animal Farming: A Review. [europepmc]
- MSU-Net: Multi-Scale U-Net for 2D Medical Image Segmentation. [europepmc]
- EDLMFC: an ensemble deep learning framework with multi-scale features combination for ncRNA-protein interaction prediction. [europepmc]
- Using smart speakers to contactlessly monitor heart rhythms. [europepmc]
- Classification of Diffuse Glioma Subtype from Clinical-Grade Pathological Images Using Deep Transfer Learning. [europepmc]
- AI in drug development: a multidisciplinary perspective. [europepmc]
- Advancements in Oncology with Artificial Intelligence-A Review Article. [europepmc]
- Pre-trained deep learning models for brain MRI image classification. [europepmc]
- Automated Classification of Brain Tumors from Magnetic Resonance Imaging Using Deep Learning. [europepmc]
- Compact optical convolution processing unit based on multimode interference. [europepmc]
- Developmental changes in exploration resemble stochastic optimization. [europepmc]
- Three novel methods for determining motor threshold with transcranial magnetic stimulation outperform conventional procedures. [europepmc]
- Brain Tumor Classification from MRI Using Image Enhancement and Convolutional Neural Network Techniques. [europepmc]
Related