Data Science Wire

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

arXiv stat.ML1mo4 min read

arXiv:2606.30930v1 Announce Type: new Abstract: Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability phenomenon focus on deterministic gradient descent, leaving the stochastic setting largely unexplored. In this work, we provide sharp convergence guarantees for Stochastic Gradient Descent (SGD) applied to the multiclass cross-entropy loss, for both linear classifiers and two-layer neural networks. We show that the stochasticity of SGD

Read the full story at arXiv stat.ML

More in Machine Learning