Abstract
Abstract
In this short note we consider the gradient descent dynamics of deep scalar linear networks, $f(x) = \prod_{l=1}^L w_l x$, which enjoy exact time-course solutions for any integer depth. We show that even in this minimal model, the optimal depth-wise learning rate scaling depends on data, whereas data-agnostic scaling rules fail to transfer across depths. Under the data-dependent optimal scaling, the learning dynamics is independent of data and weakly dependent on depth, resulting in a constant linear convergence rate across all depths including infinity. We further show similar data-dependent effects in deep scalar linear networks with residual connections.
Direct answer
What can I do from this paper page?
Use this page to scan "Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Stochastic Gradient Optimization Techniques research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Zhang2026Optimal,
title = {Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks},
author = {Yedi Zhang and Peter Latham and Leena Chennuru Vankadara and Andrew Saxe},
journal = {arXiv (Cornell University)},
year = {2026},
url = {https://arxiv.org/abs/2607.07884}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Stochastic Gradient Optimization Techniques research papers?
Follow Stochastic Gradient Optimization Techniques research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app