Extreme Language Model Compression with Optimal Sub-Words and Shared Projections

Patent №

US 11,797,862

Granted

2023-10-24

Filed 2020

Owner

GOOGLE LLC

AI components

6

ml · nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16749570

Provided is a knowledge distillation technique for training a student language model that, relative to a larger teacher language model, has a significantly smaller vocabulary, lower embedding dimensions, and/or hidden state dimensions. Specifically, aspects of the present disclosure are directed to a dual-training mechanism that trains the teacher and student language models simultaneously to obtain optimal word embeddings for the student vocabulary. In some implementations, this approach can be combined with learning shared projection matrices that transfer layer-wise knowledge from the teacher language model to the student language model. Example experimental results have also demonstrated higher compression efficiency and accuracy when compared with other state-of-the-art compression techniques, including the ability to compress the BERTBASE model by more than 60×, with only a minor drop in downstream task metrics, resulting in a language model with a footprint of under 7 MB.

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware0.99
Vision0.93
Knowledge representation0.66
Planning0.00
Evolutionary computation0.00

Ownership

GOOGLE LLC

assignment · 519740862

Assignors

SONG, YANG, GUPTA, RAGHAV, ZHOU, DENGYOUNG, ZHAO, SANQIANG

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC