Gradient descent
Roll downhill on a loss curve. Find the learning rate that converges — and the one that explodes.
Roll downhill on a loss curve. Find the learning rate that converges — and the one that explodes.
Tap or drag on the chart to choose the starting point.
Training a model means finding the weight w that makes the loss small. Gradient descent looks at the slope of the loss where it stands (the red tangent) and takes a step downhill, scaled by the learning rate η.
Too small an η and progress crawls; too big and each step jumps over the valley, landing higher than before until it explodes. On a bumpy curve, where you start decides which valley you end up in.
Training a model means finding the weight w that makes the loss small. Gradient descent looks at the slope of the loss where it stands (the red tangent) and takes a step downhill, scaled by the learning rate η.
Too small an η and progress crawls; too big and each step jumps over the valley, landing higher than before until it explodes. On a bumpy curve, where you start decides which valley you end up in.
Things to try