Parameter
One of the model's learnable numbers. Modern LLMs have billions to trillions; the bigram model in chapter 1 has |vocab|² of them (one per cell of the counts table).
Companion explanation in Step by Token, chapter 1.
Where this term gets built
- ch. 1The dumbest model that exists
- ch. 3Train your own tokens
- ch. 4Giving meaning to words
- ch. 5A neuron that learns
- ch. 6Stacking layers
- ch. 7Gradient descent live
- ch. 10The full transformer block
- ch. 12The minimum code
- ch. 13The training loop
- ch. 14Generation and sampling
- ch. 15Load real weights
- ch. 16Why your model talks badly
- ch. 17Give your model instructions
- ch. 18Fine-tuning with LoRA
- ch. 19Simple quantization
- ch. 20Talk to your model
- ch. 21Ship a useful one
- ch. 22Appendix · Backprop by hand
- ch. 23Appendix · RLHF and DPO
Continue