METHOD AND PLATFORM FOR PRE-TRAINED LANGUAGE MODEL AUTOMATIC COMPRESSION BASED ON MULTILEVEL KNOWLEDGE DISTILLATION
Patent №
US 11,501,171
Granted
2022-11-15
Filed 2021
Owner
ZHEJIANG LAB
Lab
—
AI components
5
ml · nlp · speech · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
17555535
Disclosed are an automatic compression method and platform for a pre-trained language model based on multilevel knowledge distillation. The method includes the following steps: step 1, constructing multilevel knowledge distillation, and distilling a knowledge structure of a large model at three different levels: a self-attention unit, a hidden layer state and an embedded layer; step 2, training a knowledge distillation network of meta-learning to generate a general compression architecture of a plurality of pre-trained language models; and step 3, searching for an optimal compression structure based on an evolutionary algorithm. Firstly, the knowledge distillation based on meta-learning is studied to generate the general compression architecture of the plurality of pre-trained language models; and secondly, on the basis of a trained meta-learning network, the optimal compression structure is searched for via the evolutionary algorithm, so as to obtain an optimal general compression architecture of the pre-trained language model independent of tasks.
AI classification
Ownership
ZHEJIANG LAB
assignment · 584520462
Assignors
WANG, HONGSHENG, WANG, ENPING, YU, ZAILIANG
On an employer assignment, the assignors are typically the inventors.