Skip to content
The loss curve

Epoch

One full pass over the training set. Large-scale language models rarely complete even one; they sample batches from a stream and count steps instead.

Companion explanation in Step by Token, chapter 6.

Where this term gets built