SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models
In recent years, implicit deep learning has emerged as a method to increase\nthe effective depth of deep neural networks. While their training is\nmemory-efficient, they are still significantly slower to train than their\nexplicit counterparts. In Deep Equilibrium Models (DEQs), the training is\nperformed as a bi-level problem, and its computational complexity is partially\ndriven by the iterative inversion of a huge Jacobian matrix. In this paper, we\npropose a novel strategy to tackle this computational bottleneck from which\nmany bi-level problems suffer. The main idea is to use the quasi-Newton\nmatrices from the forward pass to efficiently approximate the inverse Jacobian\nmatrix in the direction needed for the gradient computation. We provide a\ntheorem that motivates using our method with the original forward algorithms.\nIn addition, by modifying these forward algorithms, we further provide\ntheoretical guarantees that our method asymptotically estimates the true\nimplicit gradient. We empirically study this approach and the recent\nJacobian-Free method in different settings, ranging from hyperparameter\noptimization to large Multiscale DEQs (MDEQs) applied to CIFAR and ImageNet.\nBoth methods reduce significantly the computational cost of the backward pass.\nWhile SHINE has a clear advantage on hyperparameter optimization problems, both\nmethods attain similar computational performances for larger scale problems\nsuch as MDEQs at the cost of a limited performance drop compared to the\noriginal models.\n
Paper
References (48)
Scroll for more · 36 remaining