Transformer
Neural network architecture built from stacked blocks of (multi-head attention + FFN) with residuals and layer norms. The dominant architecture for language models since 2017.
Companion explanation in Step by Token, chapter 5.
Where this term gets built
- ch. 4Giving meaning to words
- ch. 6Stacking layers
- ch. 7Gradient descent live
- ch. 9Multi-head and residuals
- ch. 10The full transformer block
- ch. 11Prepare a dataset
- ch. 12The minimum code
- ch. 13The training loop
- ch. 15Load real weights
- ch. 16Why your model talks badly
- ch. 19Simple quantization
- ch. 20Talk to your model
Continue