Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
- Authors: Benjamin Gess, Sebastian Kassing
- Preprint year: 2023
- First public date: 2023-02-07
- arXiv: 2302.03550
- Status: Published
- Publication type: Journal article
- Publication year: 2026
- Journal: Mathematical programming, (2026)
- DOI: 10.1007/s10107-025-02308-y
Abstract
We prove explicit bounds on the exponential rate of convergence for the momentum stochastic gradient descent scheme (MSGD) for arbitrary, fixed hyperparameters (learning rate, friction parameter) and its continuous-in-time counterpart in the context of non-convex optimization. In the small step-size regime and in the case of flat minima or large noise intensities, these bounds prove faster convergence of MSGD compared to plain stochastic gradient descent (SGD). The results are shown for objective functions satisfying a local Polyak-Lojasiewicz inequality and under assumptions on the variance of MSGD that are satisfied in overparametrized settings. Moreover, we analyze the optimal choice of the friction parameter and show that the MSGD process almost surely converges to a local minimum.
Associated SAiS members
Research areas
BibTeX
@article{arxiv230203550,
title = {Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting},
author = {Benjamin Gess and Sebastian Kassing},
year = {2026},
journal = {Mathematical programming, (2026)},
doi = {10.1007/s10107-025-02308-y},
eprint = {2302.03550},
archivePrefix = {arXiv},
primaryClass = {math.OC}
}
