The500Feed.Live

Everything going on in AI - updated daily from 500+ sources

← Back to The 500 Feed
📄 ResearchSeptember 2, 2026

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD...

Read Original Article →

Source

http://arxiv.org/abs/2609.02728v1