Skip to content
The loss curve

Top-k sampling

Keep only the k most likely next tokens, renormalize, then sample. Cuts the long tail of implausible tokens that would otherwise be drawn occasionally.

Companion explanation in Step by Token, chapter 7.

Where this term gets built