Data Science Wire

Towards Weaker Variance Assumptions for Stochastic Optimization

arXiv stat.ML1mo4 min read

arXiv:2504.09951v2 Announce Type: replace-cross Abstract: We revisit a classical assumption for analyzing stochastic gradient algorithms where the squared norm of the stochastic subgradient (or the variance for smooth problems) is allowed to grow as fast as the squared norm of the optimization variable. We contextualize this assumption in view of its inception in the 1960s, its seemingly independent appearance in the recent literature, its relationship to weakest-known variance assumptions for analyzing stochastic gradient algorithms, and its relevance in deterministic problems for non-Lipschi

Read the full story at arXiv stat.ML

More in Machine Learning