Weight tying
Sharing one matrix between the token embedding and the output projection. Saves vocab_size × n_embd parameters and usually costs nothing in quality.
Continue
Sharing one matrix between the token embedding and the output projection. Saves vocab_size × n_embd parameters and usually costs nothing in quality.
Continue