Learning rate
Scalar that scales each gradient-descent step. Too small and training crawls; too large and the loss diverges. The single most-important hyperparameter.
Companion explanation in Step by Token, chapter 6.
Where this term gets built
Continue