The500Feed.Live
Everything going on in AI - updated daily from 500+ sources
📄 ResearchSeptember 2, 2026
Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
We study when and how momentum improves large-batch training in the one-pass regime, using power-law kernel regression as a tractable setting. We first characterize risk stability through the critical learning rate, defined as the largest learning rate for stable training, and obtain $η_{\mathrm{SGD...
Read Original Article →