Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning

Pre-trained universal feature extractors, such as BERT for natural language\nprocessing and VGG for computer vision, have become effective methods for\nimproving deep learning models without requiring more labeled data. While\neffective, feature extractors like BERT may be prohibitively large for some\ndeployment scenarios. We explore weight pruning for BERT and ask: how does\ncompression during pre-training affect transfer learning? We find that pruning\naffects transfer learning in three broad regimes. Low levels of pruning\n(30-40%) do not affect pre-training loss or transfer to downstream tasks at\nall. Medium levels of pruning increase the pre-training loss and prevent useful\npre-training information from being transferred to downstream tasks. High\nlevels of pruning additionally prevent models from fitting downstream datasets,\nleading to further degradation. Finally, we observe that fine-tuning BERT on a\nspecific task does not improve its prunability. We conclude that BERT can be\npruned once during pre-training rather than separately for each task without\naffecting performance.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC