Token
The basic unit a model reads and writes. Often a subword, sometimes a word, sometimes a single character.
Tokenization is the first step in any language model pipeline. Whitespace tokenizers split on spaces; BPE tokenizers find a granularity between words and characters by learning which subwords appear most often.
Companion explanation in Step by Token, chapter 2.
Where this term gets built
- ch. 1The dumbest model that exists
- ch. 2Counting isn't enough
- ch. 3Train your own tokens
- ch. 4Giving meaning to words
- ch. 5A neuron that learns
- ch. 7Gradient descent live
- ch. 8An attention head by hand
- ch. 9Multi-head and residuals
- ch. 10The full transformer block
- ch. 11Prepare a dataset
- ch. 12The minimum code
- ch. 13The training loop
- ch. 14Generation and sampling
- ch. 15Load real weights
- ch. 16Why your model talks badly
- ch. 17Give your model instructions
- ch. 18Fine-tuning with LoRA
- ch. 20Talk to your model
Continue