A technical note covering knowledge distillation and self-distillation techniques used to compress and improve large language models while retaining task performance.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex