Knowledge Distillation and Self-Distillation for Language Models

A technical note covering knowledge distillation and self-distillation techniques used to compress and improve large language models while retaining task performance.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC