This paper addresses the training of architectures (e.g., neural networks) formulated in a multi-objective framework where several performance indexes are jointly considered. Based on the possibility of tuning the hyperparameters defined in the learning schemes, a Pareto Front in the performance space is obtained, that provides a whole family of non-dominated architecture configurations. Since the tuning of such hyper-parameters implies solving a costly black box optimization problem, the search procedure is carried out making use of a Bayesian optimization scheme. This multi-objective training approach is directly applicable in classification problems where a performance trade-off among different types of classification errors can be naturally defined. In addition, we develop a general procedure for the multi-objective training of architectures in regression problems; this procedure can be specially useful when the distribution of the data noise is unknown. Some examples illustrate the applicability of the proposed schemes.
Paper
Full text
Multi-objective machine training based on Bayesian hyperparameter tuning
Semantic Scholar · Computer Science · 2023
Abstract
This paper addresses the training of architectures (e.g., neural networks) formulated in a multi-objective framework where several performance indexes are jointly considered. Based on the possibility of tuning the hyperparameters defined in the learning schemes, a Pareto Front in the performance space is obtained, that provides a whole family of non-dominated architecture configurations. Since the tuning of such hyper-parameters implies solving a costly black box optimization problem, the search procedure is carried out making use of a Bayesian optimization scheme. This multi-objective training approach is directly applicable in classification problems where a performance trade-off among different types of classification errors can be naturally defined. In addition, we develop a general procedure for the multi-objective training of architectures in regression problems; this procedure can be specially useful when the distribution of the data noise is unknown. Some examples illustrate the applicability of the proposed schemes.