The performance of neural network classifiers is determined by a number of hyperparameters, including learning rate, batch size, and depth. A number of attempts have been made to explore these parameters in the literature, and at times, to develop methods for optimizing them. However, exploration of parameter spaces has often been limited. In this note, I report the results of large scale experiments exploring these different parameters and their interactions.
Paper
References (4)
04@BULLET Hyperparameter optimization should not optimize for the best expected test set error of the resulting networks, but for the minimal error over a collection of multiple trained models@BULLET Hyperparameter optimization should not optimize for the best expected test set error of the resulting networks, but for the minimal error over a collection of multiple trained models