Stochastic Gradient Optimization Techniques Open access

Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization

Ryusei Yamada, N Sato, Hideaki Iiduka

arXiv (Cornell University) | Jul 9, 2026

Abstract

Abstract

Stochastic gradient descent (SGD) is a cornerstone of modern optimization. While its performance under heavy-tailed noise is often addressed through specialized modifications such as gradient clipping or normalization, we investigate a more fundamental question: how does vanilla SGD, particularly with momentum, perform in the presence of heavy-tailed noise? In this paper, we refine existing convergence results for vanilla SGD and, more importantly, provide the first comprehensive convergence analysis of vanilla SGD with momentum for strongly convex, convex, and nonconvex objectives, without employing any gradient control mechanisms. Our results demonstrate that the obtained convergence rates are inferior to the optimal rates achieved by clipped or normalized variants of SGD, thereby revealing inherent limitations of vanilla methods under heavy-tailed noise. The theoretical findings are supported by experiments on synthetic functions.

Direct answer

What can I do from this paper page?

Use this page to scan "Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Stochastic Gradient Optimization Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Ryusei Yamada

first

N Sato

middle

Hideaki Iiduka

last | ORCID 0000-0001-9173-6723

Research areas

Follow related topics

Citation

BibTeX

@article{Yamada2026Vanilla,
  title = {Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization},
  author = {Ryusei Yamada and N Sato and Hideaki Iiduka},
  journal = {arXiv (Cornell University)},
  year = {2026},
  doi = {10.48550/arxiv.2607.08104},
  url = {https://doi.org/10.48550/arxiv.2607.08104}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Stochastic Gradient Optimization Techniques research papers?

Follow Stochastic Gradient Optimization Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Vanilla SGD with Momentum Survives Heavy-Tailed Noise: Convergence Analysis without Gradient Clipping or Normalization. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app