Stochastic Gradient Optimization Techniques Open access

Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

Yedi Zhang, Peter Latham, Leena Chennuru Vankadara, Andrew Saxe

arXiv (Cornell University) | Jul 8, 2026

Abstract

Abstract

In this short note we consider the gradient descent dynamics of deep scalar linear networks, $f(x) = \prod_{l=1}^L w_l x$, which enjoy exact time-course solutions for any integer depth. We show that even in this minimal model, the optimal depth-wise learning rate scaling depends on data, whereas data-agnostic scaling rules fail to transfer across depths. Under the data-dependent optimal scaling, the learning dynamics is independent of data and weakly dependent on depth, resulting in a constant linear convergence rate across all depths including infinity. We further show similar data-dependent effects in deep scalar linear networks with residual connections.

Direct answer

What can I do from this paper page?

Use this page to scan "Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Stochastic Gradient Optimization Techniques research, save the paper, or map adjacent work.

Authors

Researchers on this paper

Yedi Zhang

first

Peter Latham

middle | ORCID 0000-0001-8713-9328

Leena Chennuru Vankadara

middle | ORCID 0000-0002-3810-840X

Andrew Saxe

last

Research areas

Follow related topics

Citation

BibTeX

@article{Zhang2026Optimal,
  title = {Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks},
  author = {Yedi Zhang and Peter Latham and Leena Chennuru Vankadara and Andrew Saxe},
  journal = {arXiv (Cornell University)},
  year = {2026},
  url = {https://arxiv.org/abs/2607.07884}
}

FAQ

Using this paper in a discovery workflow

How do I find related work for this paper?

Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.

How can I keep up with new Stochastic Gradient Optimization Techniques research papers?

Follow Stochastic Gradient Optimization Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.

Can I cite this paper from this page?

This page includes a static BibTeX block for Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.

Follow this research in Scollr

Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.

Get the app