Appropriate Learning Rates of Adaptive Learning Rate Optimization\n Algorithms for Training Deep Neural Networks
This paper deals with nonconvex stochastic optimization problems in deep\nlearning and provides appropriate learning rates with which adaptive learning\nrate optimization algorithms, such as Adam and AMSGrad, can approximate a\nstationary point of the problem. In particular, constant and diminishing\nlearning rates are provided to approximate a stationary point of the problem.\nOur results also guarantee that the adaptive learning rate optimization\nalgorithms can approximate global minimizers of convex stochastic optimization\nproblems. The adaptive learning rate optimization algorithms are examined in\nnumerical experiments on text and image classification. The experiments show\nthat the algorithms with constant learning rates perform better than ones with\ndiminishing learning rates.\n
Paper
References (43)
Scroll for more · 31 remaining