Skip to content
The loss curve

Block size

Maximum context length the model attends over during training. Picks how many tokens of history it can see.

Companion explanation in Step by Token, chapter 9.

Where this term gets built