An Overview of Low-Rank Structures in the Training and Adaptation of Large Models

The advent of deep learning has immeasurably changed the ways we process data in signal processing and machine learning. However, training and deploying modern deep learning models demand substantial computational resources, raising concerns about exorbitant training costs, GPU shortages, and heightened energy consumption. Several lines of research over the last decade have explored the emergence of low-dimensional structures during the training process, where basic elements such as weight matrices and representations tend to be approximately low rank even though not explicitly trained to be. These low-dimensional structures arise in part due to the implicit bias of the methods used to train deep networks, providing the potential to partially explain why deep models need fewer samples than the number of model parameters. This implicit low dimensionality has inspired the exploration of low-rank structures in training and fine-tuning large-scale deep learning models more efficiently. In this article, we review recent exciting advances in using low-rank structure in deep learning and aim to clarify the mathematical foundations underlying their design. Specifically, we highlight key insights from a rich line of research focused on theoretically understanding and leveraging low-rank structures in deep learning both at every iteration of training as well as at the global minimum.

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC