Loss
A number that says how badly the model is predicting the right answers. Training minimizes it.
For language models, the per-token loss is typically -log(probability) the model assigned to the true next token. Sum or average across a sequence and you get a comparable number per token. Cross-entropy loss is the standard formulation.
Companion explanation in Step by Token, chapter 6.
Where this term gets built
- ch. 2Counting isn't enough
- ch. 4Giving meaning to words
- ch. 5A neuron that learns
- ch. 6Stacking layers
- ch. 7Gradient descent live
- ch. 8An attention head by hand
- ch. 9Multi-head and residuals
- ch. 12The minimum code
- ch. 13The training loop
- ch. 16Why your model talks badly
- ch. 17Give your model instructions
- ch. 19Simple quantization
- ch. 20Talk to your model
- ch. 21Ship a useful one
- ch. 22Appendix · Backprop by hand
- ch. 23Appendix · RLHF and DPO
Continue