Stabilizing Elastic Weight Consolidation method in practical ML tasks and using weight importances for neural network pruning

This paper is devoted to the features of the practical application of the\nElastic Weight Consolidation (EWC) method for continual learning of neural\nnetworks on several training sets. We will more rigorously compare the\nwell-known methodologies for calculating the importance of weights used in the\nEWC method. These are the Memory Aware Synapses (MAS), Synaptic Intelligence\n(SI) methodologies and the calculation of the importance of weights based on\nthe Fisher information matrix from the original paper on EWC. We will consider\nthese methodologies as applied to deep neural networks with fully connected and\nconvolutional layers, find the optimal hyperparameters for each of the\nmethodologies, and compare the results of continual neural network learning\nusing these hyperparameters. Next, we will point out the problems that arise\nwhen applying the EWC method to deep neural networks with convolutional layers\nand self-attention layers, such as the "gradient explosion" and the loss of\nmeaningful information in the gradient when using the constraint of its norm\n(gradient clipping). Then, we will propose a stabilization approach for the EWC\nmethod that helps to solve these problems, evaluate it in comparison with the\noriginal methodology and show that the proposed stabilization approach performs\non the task of maintaining skills during continual learning no worse than the\noriginal EWC, but does not have its disadvantages. In conclusion, we present an\ninteresting fact about the use of various types of weight importance in the\nproblem of neural network pruning.\n

Paper

References (22)

Scroll for more · 10 remaining

Similar papers

© 2026 NYSGPT2525 LLC